Hybrid clustering framework for large-scale scientific literature structuring
DOI: https://doi.org/10.3846/ntcs.2026.26284Abstract
This study addresses the challenge of organizing and interpreting large collections of scientific literature, focusing on a rapidly growing and heterogeneous AI research domain derived from object-detection-with-small-data literature. First, based on the set of keywords using Clarivate Analytics queries, over 18,500 research articles were retrieved, illustrating the problem of manual review infeasibility, leading to machine learning-based approaches. To begin literature filtration, a pipeline combining text preprocessing, TF-IDF feature extraction, and K-means++ clustering was implemented. Clusters were labelled using YAKE keyword extraction and thematically summarized with large language models (ChatGPT-4o and ChatGPT-o1), enabling the interpretation of research domains such as medical imaging, anomaly detection, etc. To compare clustering strategies, we employed neural autoencoder-based embeddings and deep embedded clustering architectures. The autoencoder latent space yielded both convergences and divergences relative to TF-IDF clusters, reflecting higher-level semantic structuring. Direct end-to-end DEC optimization produced unstable cluster assignments, thus we applied a three-phase IDEC training protocol using autoencoder pretraining, centroid initialization with K-means, and joint fine-tuning – this produced substantially more balanced and internally coherent clusters. Results demonstrate that integrating keyword-based and neural methods, supplemented by conversational AI summarization, offers a scalable and interpretable framework for structuring large scientific datasets.
This work highlights both methodological strengths and thematic overlaps across AI research, emphasizing the role of hybrid clustering approaches in managing the growing volume of scholarly publications.
Keywords:
automated literature sorting, TF-IDF feature extraction, K-means++ clustering, large language models, autoencoder embeddings, deep embedded clustering, scientific text miningHow to Cite
Share
License
Copyright (c) 2026 The Author(s). Published by Vilnius Gediminas Technical University.

This work is licensed under a Creative Commons Attribution 4.0 International License.
References
Arthur, D., & Vassilvitskii, S. (2007, January 7–9). K-means++: The advantages of careful seeding. In Proceedings of the Annual ACM-SIAM Symposium on Discrete Algorithms (pp. 1027–1035). Society for Industrial and Applied Mathematics.
Asmussen, C. B., & Møller, C. (2019). Smart literature review: a practical topic modelling approach to exploratory literature review. Journal of Big Data, 6(1), Article 93. https://doi.org/10.1186/s40537-019-0255-7
Blei, D. M., Ng, A. Y., & Jordan, M. I. (2003). Latent Dirichlet allocation. Journal of Machine Learning Research, 3(4–5). https://doi.org/10.7551/mitpress/1120.003.0082
Bradley, P. S., & Fayyad, U. M. (1998). Refining initial points for K-means clustering. In Proceedings of the Fifteenth International Conference on Machine Learning.
Campos, R., Mangaravite, V., Pasquali, A., Jorge, A., Nunes, C., & Jatowt, A. (2020). YAKE! Keyword extraction from single documents using multiple local features. Information Sciences, 509, 257–289. https://doi.org/10.1016/j.ins.2019.09.013
Chen, C. (2006). CiteSpace II: Detecting and visualizing emerging trends and transient patterns in scientific literature. Journal of the American Society for Information Science and Technology, 57(3), 359–377. https://doi.org/10.1002/asi.20317
de Winter, J. (2024). Can ChatGPT be used to predict citation counts, readership, and social media interaction? An exploration among 2222 scientific abstracts. Scientometrics, 129(4), 2469–2487. https://doi.org/10.1007/s11192-024-04939-y
Delgado-Chaves, F. M., Jennings, M. J., Atalaia, A., Wolff, J., Horvath, R., Mamdouh, Z. M., Baumbach, J., & Baumbach, L. (2025). Transforming literature screening: The emerging role of large language models in systematic reviews. Proceedings of the National Academy of Sciences of the United States of America, 122(2), Article e2411962122. https://doi.org/10.1073/pnas.2411962122
Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In J. Burstein, C. Doran, & T. Solorio (Eds.), NAACL HLT 2019 – 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies – Proceedings of the Conference (vol. 1, pp. 4171–4186). Association for Computational Linguistics.
Grootendorst, M. (2022). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv. http://arxiv.org/abs/2203.05794
Guo, X., Gao, L., Liu, X., & Yin, J. (2017). Improved deep embedded clustering with local structure preservation. IJCAI International Joint Conference on Artificial Intelligence (pp. 1753–1759). https://doi.org/10.24963/ijcai.2017/243
Hinton, G. E., & Salakhutdinov, R. R. (2006). Reducing the dimensionality of data with neural networks. Science, 313(5786), 504–507. https://doi.org/10.1126/science.1127647
Holst, D., Moenck, K., Koch, J., Schmedemann, O., & Schüppstuhl, T. (2025). Transparent reporting of AI in systematic literature reviews: Development of the PRISMA-trAIce checklist. JMIR AI, 4, Article e80247. https://doi.org/10.2196/80247
Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., & Liu, T. (2025). A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, 43(2), Article 2. https://doi.org/10.1145/3703155
Jiao, L., Zhang, F., Liu, F., Yang, S., Li, L., Feng, Z., & Qu, R. (2019). A survey of deep learning-based object detection. IEEE Access, 7, 128837–128868. https://doi.org/10.1109/ACCESS.2019.2939201
Jones, K. S. (1972). A statistical interpretation of term specificity and its application in retrieval. Journal of Documentation, 28(1), 11–21. https://doi.org/10.1108/eb026526
Kalibatiene, D., & Miliauskaitė, J. (2026). From manual to automated systematic review: Key attributes influencing the duration of systematic reviews in software engineering. Computer Standards & Interfaces, 96, Article 104073. https://doi.org/10.1016/J.CSI.2025.104073
Köhler, M., Eisenbach, M., & Gross, H.-M. (2024). Few-shot object detection: A comprehensive survey. IEEE Transactions on Neural Networks and Learning Systems, 35(9), 11958–11978. https://doi.org/10.1109/TNNLS.2023.3265051
Khraisha, Q., Put, S., Kappenberg, J., Warraitch, A., & Hadfield, K. (2024). Can large language models replace humans in systematic reviews? Evaluating GPT-4’s efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages. Research Synthesis Methods, 15(4), 616–626. https://doi.org/10.1002/jrsm.1715
Kitchenham, B. (2007). Guidelines for performing systematic literature reviews in software engineering (Technical Report, Ver. 2.3). EBSE.
Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). ImageNet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems, 25, 1097–1105.
Lam, M. S., Teoh, J., Landay, J., Heer, J., Bernstein, M. S. (2024). Concept induction: analyzing unstructured text with high-level concepts using LLooM. In Proceedings of CHI 2024. Association for Computing Machinery. https://doi.org/10.1145/3613904.3642830
Liu, L., Ouyang, W., Wang, X., Fieguth, P., Chen, J., Liu, X., & Pietikäinen, M. (2020). Deep learning for generic object detection: A survey. International Journal of Computer Vision, 128, 261–318. https://doi.org/10.1007/s11263-019-01247-4
Lloyd, S. P. (1982). Least squares quantization in PCM. IEEE Transactions on Information Theory, 28(2), 129–137. https://doi.org/10.1109/TIT.1982.1056489
MacQueen, J. (1967). Some methods for classification and analysis of multivariate observations. In Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability (Vol. 1, pp. 281–297). University of California Press.
Marshall, I. J., & Wallace, B. C. (2019). Toward systematic review automation: A practical guide to using machine learning tools in research synthesis. Systematic Reviews, 8, Article 163. https://doi.org/10.1186/s13643-019-1074-9
Min, E., Guo, X., Liu, Q., Zhang, G., Cui, J., & Long, J. (2018). A survey of clustering with deep learning: From the perspective of network architecture. IEEE Access, 6, 39501–39514. https://doi.org/10.1109/ACCESS.2018.2855437
O’Mara-Eves, A., Thomas, J., McNaught, J., Miwa, M., & Ananiadou, S. (2015). Using text mining for study identification in systematic reviews: A systematic review of current approaches. Systematic Reviews, 4, Article 5. https://doi.org/10.1186/2046-4053-4-5
OpenAI, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., Avila, R., Babuschkin, I., Balaji, S., Balcom, V., Baltescu, P., Bao, H., Bavarian, M., Belgum, J., … Zoph, B. (2024). GPT-4 Technical Report. arXiv. http://arxiv.org/abs/2303.08774
Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., Shamseer, L., Tetzlaff, J. M., Akl, E. A., Brennan, S. E., Chou, R., Glanville, J., Grimshaw, J. M., Hróbjartsson, A., Lalu, M. M., Li, T., Loder, E. W., Mayo-Wilson, E., McDonald, S., … Moher, D. (2021). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ, 372, Article 71. https://doi.org/10.1136/bmj.n71
Petticrew, M., & Roberts, H. (2008). Systematic reviews in the social sciences: A practical guide. Wiley. https://doi.org/10.1002/9780470754887
Pranckutė, R. (2021). Web of Science (WoS) and Scopus: The titans of bibliographic information in today’s academic world. Publications, 9(1), Article 12. https://doi.org/10.3390/publications9010012
Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using siamese BERT-networks. In EMNLP-IJCNLP 2019 – 2019 Conference on Empirical Methods in Natural Language Processing and 9th International Joint Conference on Natural Language Processing. Proceedings of the Conference (pp. 3982–3992). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1410
Ren, Y., Pu, J., Yang, Z., Xu, J., Li, G., Pu, X., Yu, P. S., & He, L. (2025). Deep clustering: A comprehensive survey. IEEE Transactions on Neural Networks and Learning Systems, 36(4), 5858–5878. https://doi.org/10.1109/TNNLS.2024.3403155
Rousseeuw, P. J. (1987). Silhouettes: A graphical aid to the interpretation and validation of cluster analysis. Journal of Computational and Applied Mathematics, 20, 53–65. https://doi.org/10.1016/0377-0427(87)90125-7
Salton, G., & McGill, M. J. (1983). Introduction to modern information retrieval. McGraw-Hill.
Scherbakov, D., Hubig, N., Jansari, V., Bakumenko, A., & Lenert, L. A. (2025). The emergence of large language models as tools in literature reviews: A large language model-assisted systematic review. Journal of the American Medical Informatics Association, 32(6), 1071–1086. https://doi.org/10.1093/jamia/ocaf063
Singh, A., D’Arcy, M., Cohan, A., Downey, D., & Feldman, S. (2023). SciRepEval: A Multi-format benchmark for scientific document representations. In EMNLP 2023 – 2023 Conference on Empirical Methods in Natural Language Processing. (pp. 5548–5566). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.emnlp-main.338
van de Schoot, R., de Bruin, J., Schram, R., Zahedi, P., de Boer, J., Weijdema, F., Kramer, B., Huijts, M., Hoogerwerf, M., Ferdinands, G., Harkema, A., Willemsen, J., Ma, Y., Fang, Q., Hindriks, S., Tummers, L., & Oberski, D. L. (2021). An open source machine learning framework for efficient and transparent systematic reviews. Nature Machine Intelligence, 3, 125–133. https://doi.org/10.1038/s42256-020-00287-7
van Dinter, R., Tekinerdogan, B., & Catal, C. (2021). Automation of systematic literature reviews: A systematic literature review. Information and Software Technology, 136, Article 106589. https://doi.org/10.1016/j.infsof.2021.106589
van Eck, N. J., & Waltman, L. (2010). Software survey: VOSviewer, a computer program for bibliometric mapping. Scientometrics, 84, 523–538. https://doi.org/10.1007/s11192-009-0146-3
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017, December). Attention is all you need. In Advances in Neural Information Processing Systems.
Vincent, P., Larochelle, H., Lajoie, I., Bengio, Y., & Manzagol, P. A. (2010). Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion. Journal of Machine Learning Research, 11, 3371–3408.
Wang, W., Wei, F., Dong, L., Bao, H., Yang, N., & Zhou, M. (2020, December). MINILM: Deep self-attention distillation for task-agnostic compression of pre-trained transformers. In Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS ‘20) (pp. 5776–5788). Curran Associates Inc.
Wei, X., Zhang, Z., Huang, H., & Zhou, Y. (2024). An overview on deep clustering. Neurocomputing, 590, Article 127761. https://doi.org/10.1016/j.neucom.2024.127761
Xie, J., Girshick, R., & Farhadi, A. (2016). Unsupervised deep embedding for clustering analysis. In 33rd International Conference on Machine Learning, ICML 2016. (pp. 478–487). JMLR.org.
Xu, Q., Gu, H., & Ji, S. W. (2023). Text clustering based on pre-trained models and autoencoders. Frontiers in Computational Neuroscience, 17, Article 1334436. https://doi.org/10.3389/fncom.2023.1334436
Zhao, Z. Q., Zheng, P., Xu, S. T., & Wu, X. (2019). Object detection with deep learning: A review. IEEE Transactions on Neural Networks and Learning Systems, 30(11), 3212–3232. https://doi.org/10.1109/TNNLS.2018.2876865
Zou, Z., Chen, K., Shi, Z., Guo, Y., & Ye, J. (2023). Object detection in 20 years: A survey. Proceedings of the IEEE, 111(3), 257–276. https://doi.org/10.1109/JPROC.2023.3238524
View article in other formats
Published
Issue
Section
Copyright
Copyright (c) 2026 The Author(s). Published by Vilnius Gediminas Technical University.
License

This work is licensed under a Creative Commons Attribution 4.0 International License.