Hybrid clustering framework for large-scale scientific literature structuring

DOI: https://doi.org/10.3846/ntcs.2026.26284

Abstract

This study addresses the challenge of organizing and interpreting large collections of scientific literature, focusing on a rapidly growing and heterogeneous AI research domain derived from object-detection-with-small-data literature. First, based on the set of keywords using Clarivate Analytics queries, over 18,500 research articles were retrieved, illustrating the problem of manual review infeasibility, leading to machine learning-based approaches. To begin literature filtration, a pipeline combining text preprocessing, TF-IDF feature extraction, and K-means++ clustering was implemented. Clusters were labelled using YAKE keyword extraction and thematically summarized with large language models (ChatGPT-4o and ChatGPT-o1), enabling the interpretation of research domains such as medical imaging, anomaly detection, etc. To compare clustering strategies, we employed neural autoencoder-based embeddings and deep embedded clustering architectures. The autoencoder latent space yielded both convergences and divergences relative to TF-IDF clusters, reflecting higher-level semantic structuring. Direct end-to-end DEC optimization produced unstable cluster assignments, thus we applied a three-phase IDEC training protocol using autoencoder pretraining, centroid initialization with K-means, and joint fine-tuning – this produced substantially more balanced and internally coherent clusters. Results demonstrate that integrating keyword-based and neural methods, supplemented by conversational AI summarization, offers a scalable and interpretable framework for structuring large scientific datasets.

This work highlights both methodological strengths and thematic overlaps across AI research, emphasizing the role of hybrid clustering approaches in managing the growing volume of scholarly publications.

Keywords:

automated literature sorting, TF-IDF feature extraction, K-means++ clustering, large language models, autoencoder embeddings, deep embedded clustering, scientific text mining
Published in Issue
September 29, 2026
Abstract Views
16

How to Cite

Teplov, D., Radzivil, A., & Bugajev, A. (2026). Hybrid clustering framework for large-scale scientific literature structuring. New Trends in Computer Sciences, 4(1), 161–195. https://doi.org/10.3846/ntcs.2026.26284

Share

References

Arthur, D., & Vassilvitskii, S. (2007, January 7–9). K-means++: The advantages of careful seeding. In Proceedings of the Annual ACM-SIAM Symposium on Discrete Algorithms (pp. 1027–1035). Society for Industrial and Applied Mathematics.

Asmussen, C. B., & Møller, C. (2019). Smart literature review: a practical topic modelling approach to exploratory literature review. Journal of Big Data, 6(1), Article 93. https://doi.org/10.1186/s40537-019-0255-7

Blei, D. M., Ng, A. Y., & Jordan, M. I. (2003). Latent Dirichlet allocation. Journal of Machine Learning Research, 3(4–5). https://doi.org/10.7551/mitpress/1120.003.0082

Bradley, P. S., & Fayyad, U. M. (1998). Refining initial points for K-means clustering. In Proceedings of the Fifteenth International Conference on Machine Learning.

Campos, R., Mangaravite, V., Pasquali, A., Jorge, A., Nunes, C., & Jatowt, A. (2020). YAKE! Keyword extraction from single documents using multiple local features. Information Sciences, 509, 257–289. https://doi.org/10.1016/j.ins.2019.09.013

Chen, C. (2006). CiteSpace II: Detecting and visualizing emerging trends and transient patterns in scientific literature. Journal of the American Society for Information Science and Technology, 57(3), 359–377. https://doi.org/10.1002/asi.20317

de Winter, J. (2024). Can ChatGPT be used to predict citation counts, readership, and social media interaction? An exploration among 2222 scientific abstracts. Scientometrics, 129(4), 2469–2487. https://doi.org/10.1007/s11192-024-04939-y

Delgado-Chaves, F. M., Jennings, M. J., Atalaia, A., Wolff, J., Horvath, R., Mamdouh, Z. M., Baumbach, J., & Baumbach, L. (2025). Transforming literature screening: The emerging role of large language models in systematic reviews. Proceedings of the National Academy of Sciences of the United States of America, 122(2), Article e2411962122. https://doi.org/10.1073/pnas.2411962122

Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In J. Burstein, C. Doran, & T. Solorio (Eds.), NAACL HLT 2019 – 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies – Proceedings of the Conference (vol. 1, pp. 4171–4186). Association for Computational Linguistics.

Grootendorst, M. (2022). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv. http://arxiv.org/abs/2203.05794

Guo, X., Gao, L., Liu, X., & Yin, J. (2017). Improved deep embedded clustering with local structure preservation. IJCAI International Joint Conference on Artificial Intelligence (pp. 1753–1759). https://doi.org/10.24963/ijcai.2017/243

Hinton, G. E., & Salakhutdinov, R. R. (2006). Reducing the dimensionality of data with neural networks. Science, 313(5786), 504–507. https://doi.org/10.1126/science.1127647

Holst, D., Moenck, K., Koch, J., Schmedemann, O., & Schüppstuhl, T. (2025). Transparent reporting of AI in systematic literature reviews: Development of the PRISMA-trAIce checklist. JMIR AI, 4, Article e80247. https://doi.org/10.2196/80247

Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., & Liu, T. (2025). A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, 43(2), Article 2. https://doi.org/10.1145/3703155

Jiao, L., Zhang, F., Liu, F., Yang, S., Li, L., Feng, Z., & Qu, R. (2019). A survey of deep learning-based object detection. IEEE Access, 7, 128837–128868. https://doi.org/10.1109/ACCESS.2019.2939201

Jones, K. S. (1972). A statistical interpretation of term specificity and its application in retrieval. Journal of Documentation, 28(1), 11–21. https://doi.org/10.1108/eb026526

Kalibatiene, D., & Miliauskaitė, J. (2026). From manual to automated systematic review: Key attributes influencing the duration of systematic reviews in software engineering. Computer Standards & Interfaces, 96, Article 104073. https://doi.org/10.1016/J.CSI.2025.104073

Köhler, M., Eisenbach, M., & Gross, H.-M. (2024). Few-shot object detection: A comprehensive survey. IEEE Transactions on Neural Networks and Learning Systems, 35(9), 11958–11978. https://doi.org/10.1109/TNNLS.2023.3265051

Khraisha, Q., Put, S., Kappenberg, J., Warraitch, A., & Hadfield, K. (2024). Can large language models replace humans in systematic reviews? Evaluating GPT-4’s efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages. Research Synthesis Methods, 15(4), 616–626. https://doi.org/10.1002/jrsm.1715

Kitchenham, B. (2007). Guidelines for performing systematic literature reviews in software engineering (Technical Report, Ver. 2.3). EBSE.

Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). ImageNet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems, 25, 1097–1105.

Lam, M. S., Teoh, J., Landay, J., Heer, J., Bernstein, M. S. (2024). Concept induction: analyzing unstructured text with high-level concepts using LLooM. In Proceedings of CHI 2024. Association for Computing Machinery. https://doi.org/10.1145/3613904.3642830

Liu, L., Ouyang, W., Wang, X., Fieguth, P., Chen, J., Liu, X., & Pietikäinen, M. (2020). Deep learning for generic object detection: A survey. International Journal of Computer Vision, 128, 261–318. https://doi.org/10.1007/s11263-019-01247-4

Lloyd, S. P. (1982). Least squares quantization in PCM. IEEE Transactions on Information Theory, 28(2), 129–137. https://doi.org/10.1109/TIT.1982.1056489

MacQueen, J. (1967). Some methods for classification and analysis of multivariate observations. In Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability (Vol. 1, pp. 281–297). University of California Press.

Marshall, I. J., & Wallace, B. C. (2019). Toward systematic review automation: A practical guide to using machine learning tools in research synthesis. Systematic Reviews, 8, Article 163. https://doi.org/10.1186/s13643-019-1074-9

Min, E., Guo, X., Liu, Q., Zhang, G., Cui, J., & Long, J. (2018). A survey of clustering with deep learning: From the perspective of network architecture. IEEE Access, 6, 39501–39514. https://doi.org/10.1109/ACCESS.2018.2855437

O’Mara-Eves, A., Thomas, J., McNaught, J., Miwa, M., & Ananiadou, S. (2015). Using text mining for study identification in systematic reviews: A systematic review of current approaches. Systematic Reviews, 4, Article 5. https://doi.org/10.1186/2046-4053-4-5

OpenAI, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., Avila, R., Babuschkin, I., Balaji, S., Balcom, V., Baltescu, P., Bao, H., Bavarian, M., Belgum, J., … Zoph, B. (2024). GPT-4 Technical Report. arXiv. http://arxiv.org/abs/2303.08774

Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., Shamseer, L., Tetzlaff, J. M., Akl, E. A., Brennan, S. E., Chou, R., Glanville, J., Grimshaw, J. M., Hróbjartsson, A., Lalu, M. M., Li, T., Loder, E. W., Mayo-Wilson, E., McDonald, S., … Moher, D. (2021). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ, 372, Article 71. https://doi.org/10.1136/bmj.n71

Petticrew, M., & Roberts, H. (2008). Systematic reviews in the social sciences: A practical guide. Wiley. https://doi.org/10.1002/9780470754887

Pranckutė, R. (2021). Web of Science (WoS) and Scopus: The titans of bibliographic information in today’s academic world. Publications, 9(1), Article 12. https://doi.org/10.3390/publications9010012

Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using siamese BERT-networks. In EMNLP-IJCNLP 2019 – 2019 Conference on Empirical Methods in Natural Language Processing and 9th International Joint Conference on Natural Language Processing. Proceedings of the Conference (pp. 3982–3992). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1410

Ren, Y., Pu, J., Yang, Z., Xu, J., Li, G., Pu, X., Yu, P. S., & He, L. (2025). Deep clustering: A comprehensive survey. IEEE Transactions on Neural Networks and Learning Systems, 36(4), 5858–5878. https://doi.org/10.1109/TNNLS.2024.3403155

Rousseeuw, P. J. (1987). Silhouettes: A graphical aid to the interpretation and validation of cluster analysis. Journal of Computational and Applied Mathematics, 20, 53–65. https://doi.org/10.1016/0377-0427(87)90125-7

Salton, G., & McGill, M. J. (1983). Introduction to modern information retrieval. McGraw-Hill.

Scherbakov, D., Hubig, N., Jansari, V., Bakumenko, A., & Lenert, L. A. (2025). The emergence of large language models as tools in literature reviews: A large language model-assisted systematic review. Journal of the American Medical Informatics Association, 32(6), 1071–1086. https://doi.org/10.1093/jamia/ocaf063

Singh, A., D’Arcy, M., Cohan, A., Downey, D., & Feldman, S. (2023). SciRepEval: A Multi-format benchmark for scientific document representations. In EMNLP 2023 – 2023 Conference on Empirical Methods in Natural Language Processing. (pp. 5548–5566). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.emnlp-main.338

van de Schoot, R., de Bruin, J., Schram, R., Zahedi, P., de Boer, J., Weijdema, F., Kramer, B., Huijts, M., Hoogerwerf, M., Ferdinands, G., Harkema, A., Willemsen, J., Ma, Y., Fang, Q., Hindriks, S., Tummers, L., & Oberski, D. L. (2021). An open source machine learning framework for efficient and transparent systematic reviews. Nature Machine Intelligence, 3, 125–133. https://doi.org/10.1038/s42256-020-00287-7

van Dinter, R., Tekinerdogan, B., & Catal, C. (2021). Automation of systematic literature reviews: A systematic literature review. Information and Software Technology, 136, Article 106589. https://doi.org/10.1016/j.infsof.2021.106589

van Eck, N. J., & Waltman, L. (2010). Software survey: VOSviewer, a computer program for bibliometric mapping. Scientometrics, 84, 523–538. https://doi.org/10.1007/s11192-009-0146-3

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017, December). Attention is all you need. In Advances in Neural Information Processing Systems.

Vincent, P., Larochelle, H., Lajoie, I., Bengio, Y., & Manzagol, P. A. (2010). Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion. Journal of Machine Learning Research, 11, 3371–3408.

Wang, W., Wei, F., Dong, L., Bao, H., Yang, N., & Zhou, M. (2020, December). MINILM: Deep self-attention distillation for task-agnostic compression of pre-trained transformers. In Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS ‘20) (pp. 5776–5788). Curran Associates Inc.

Wei, X., Zhang, Z., Huang, H., & Zhou, Y. (2024). An overview on deep clustering. Neurocomputing, 590, Article 127761. https://doi.org/10.1016/j.neucom.2024.127761

Xie, J., Girshick, R., & Farhadi, A. (2016). Unsupervised deep embedding for clustering analysis. In 33rd International Conference on Machine Learning, ICML 2016. (pp. 478–487). JMLR.org.

Xu, Q., Gu, H., & Ji, S. W. (2023). Text clustering based on pre-trained models and autoencoders. Frontiers in Computational Neuroscience, 17, Article 1334436. https://doi.org/10.3389/fncom.2023.1334436

Zhao, Z. Q., Zheng, P., Xu, S. T., & Wu, X. (2019). Object detection with deep learning: A review. IEEE Transactions on Neural Networks and Learning Systems, 30(11), 3212–3232. https://doi.org/10.1109/TNNLS.2018.2876865

Zou, Z., Chen, K., Shi, Z., Guo, Y., & Ye, J. (2023). Object detection in 20 years: A survey. Proceedings of the IEEE, 111(3), 257–276. https://doi.org/10.1109/JPROC.2023.3238524

View article in other formats

CrossMark check

CrossMark logo

Published

2026-09-29

Issue

Section

Articles

How to Cite

Teplov, D., Radzivil, A., & Bugajev, A. (2026). Hybrid clustering framework for large-scale scientific literature structuring. New Trends in Computer Sciences, 4(1), 161–195. https://doi.org/10.3846/ntcs.2026.26284

Share