Abstracts
Abstract
Topic modelling has become a prominent tool for the study of scientific fields, as they allow for a large-scale interpretation of research trends. Nevertheless, the output of these models is structured as a list of keywords, which requires a manual interpretation for the labelling. This paper proposes to assess the reliability of three LLMs, namely flan, GPT-4o, and GPT-4 mini for topic labelling. Drawing on previous research leveraging BERTopic, we generate topics from a dataset of all the scientific articles (n=34,797) authored by all biology professors in Switzerland between 2008 and 2020, as recorded in the Web of Science database. We assess the output of the three models both quantitatively and qualitatively and measure the effect of the temperature parameter in GPT models and find that, first, both GPT models are capable of correctly and precisely labelling topics from the models' output keywords at the default temperature. Second, 3-word labels are preferable to grasp the complexity of research topics.
Keywords:
- Topic modelling,
- labelling,
- automatization,
- bibliometrics,
- science of science
Résumé
La modélisation thématique est devenue un outil majeur pour l’étude des champs scientifiques, car elle permet une interprétation à grande échelle des tendances de recherche. Néanmoins, la sortie de ces modèles se présente sous la forme d’une liste de mots-clés, ce qui nécessite une interprétation manuelle pour leur étiquetage. Cet article propose d’évaluer la fiabilité de trois grands modèles de langage (LLM), à savoir Flan, GPT‑4o et GPT‑4 mini, pour l’étiquetage de thèmes. En s’appuyant sur des recherches antérieures utilisant BERTopic, nous générons des thèmes à partir d’un ensemble de données comprenant tous les articles scientifiques (n = 34 797) publiés par l’ensemble des professeurs de biologie en Suisse entre 2008 et 2020, tels que recensés dans la base de données Web of Science. Nous évaluons les résultats des trois modèles à la fois de manière quantitative et qualitative, et analysons l’effet du paramètre de température dans les modèles GPT. Nous constatons, premièrement, que les deux modèles GPT sont capables d’étiqueter correctement et précisément les thèmes à partir des mots-clés générés par les modèles à la température par défaut. Deuxièmement, les étiquettes composées de trois mots sont préférables pour saisir la complexité des thèmes de recherche.
Mots-clés :
- modélisation thématique,
- étiquetage,
- automatisation,
- bibliométrie,
- science de la science
Appendices
Bibliography
- Benz, P., Pradier, C., Kozlowski, D., Shokida, N. S., & Larivière, V. (2025). Mapping the unseen in practice: comparing latent Dirichlet allocation and BERTopic for navigating topic spaces. Scientometrics, 130(7), 3839-3870. https://doi.org/10.1007/s11192-025-05339-6
- Blei, D. M., Ng, A. Y., & Jordan, M. I. (2003). Latent dirichlet allocation. The Journal of Machine Learning Research, 3, 993–1022.
- Blei, D. M., & Lafferty, J. D. (2009). Topic Models. In A. N. Srivastava & M. Sahami (Eds.), Text mining: Classification, clustering, and applications (pp. 71–94). Taylor & Francis.
- Bogachek, O. (2025). Topic Labeling Using Large Language Models (SSRN Scholarly Paper No. 5369321). Social Science Research Network. https://doi.org/10.2139/ssrn.5369321
- Chen, C., Ibekwe-SanJuan, F., & Hou, J. (2010). The structure and dynamics of cocitation clusters: A multiple-perspective cocitation analysis. Journal of the American Society for Information Science and Technology, 61(7), 1386–1409. https://doi.org/10.1002/asi.21309
- Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., Webson, A., Gu, S. S., Dai, Z., Suzgun, M., Chen, X., Chowdhery, A., Narang, S., Mishra, G., Yu, A., … Wei, J. (2022). Scaling Instruction-Finetuned Language Models. arXiv. https://doi.org/10.48550/arXiv.2210.11416
- Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint. https://doi.org/10.48550/arXiv.1810.04805
- Grootendorst, M. (2022). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv preprint. https://doi.org/10.48550/arXiv.2203.05794
- Hjørland, B. (1992). The concept of ‘subject’ in information science. Journal of Documentation, 48(2), 172–200. https://doi.org/10.1108/eb026895
- Hjørland, B. (2017). Subject (of documents). Knowledge Organization, 44(1), 55–64. https://doi.org/10.5771/0943-7444-2017-1-55.
- Kang, D., & Evans, J. (2020). Against method: Exploding the boundary between qualitative and quantitative studies of science. Quantitative Science Studies, 1(3), 930–944. https://doi.org/10.1162/qss_a_00056
- Khandelwal, T. (2025). Using LLM-Based Approaches to Enhance and Automate Topic Labeling (No. arXiv:2502.18469). arXiv. https://doi.org/10.48550/arXiv.2502.18469
- Koopman, R., & Wang, S. (2017). Mutual information based labelling and comparing clusters. Scientometrics, 111(2), 1157–1167. https://doi.org/10.1007/s11192-017-2305-2
- Lancaster, F. W. (2003). Indexing and abstracting in theory and practice (3rd ed.). Facet Publishing.
- Lepori, B., Andersen, J. P. & Donnay, K. (2026). Opinion paper: generative AI and the future of scientometrics. Scientometrics. s11192-026-05667-1
- Marchetti, A., & Puranam, P. (2020). Interpreting Topic Models Using Prototypical Text: From ‘Telling’ To “Showing” (SSRN Scholarly Paper 3717437). https://doi.org/10.2139/ssrn.3717437
- Murray, D., Ni, C., Gu, W., & Hubbard, T. (2025). Using language models to label clusters of scientific documents. Scientometrics. https://doi.org/10.1007/s11192-025-05445-5
- OpenAI, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., Avila, R., Babuschkin, I., Balaji, S., Balcom, V., Baltescu, P., Bao, H., Bavarian, M., Belgum, J., … Zoph, B. (2024). GPT-4 Technical Report (arXiv:2303.08774). arXiv. https://doi.org/10.48550/arXiv.2303.08774
- Rijcken, E., Scheepers, F., Zervanou, K., Spruit, M., Mosteiro, P., & Kaymak, U. (2023). Towards Interpreting Topic Models with ChatGPT. The 20th World Congress of the International Fuzzy Systems Association, Daegu, Republic of Korea.
- Rüdiger, M., Antons, D., Joshi, A. M., & Salge, T.-O. (2022). Topic modeling revisited: New evidence on algorithm performance and quality metrics. PLoS One, 17(4), e0266325. https://doi.org/10.1371/journal.pone.0266325
- Sjögårde, P., Ahlgren, P., & Waltman, L. (2021). Algorithmic labeling in hierarchical classifications of publications: Evaluation of bibliographic fields and term weighting approaches. Journal of the Association for Information Science and Technology, 72(7), 853–869. https://doi.org/10.1002/asi.24452
- Suominen, A., & Toivanen, H. (2016). Map of science with topic modeling: Comparison of unsupervised learning and human‐assigned subject classification. Journal of the Association for Information Science and Technology, 67(10), 2464–2476. https://doi.org/10.1002/asi.23596
- Van Eck, N. J., & Waltman, L. (2024). An open approach for classifying research publications. Leiden Madtrics, 24.
- Velden, T., Yan, S., & Lagoze, C. (2017). Mapping the cognitive structure of astrophysics by infomap clustering of the citation network and topic affinity analysis. Scientometrics, 111(2), 1033–1051. https://doi.org/10.1007/s11192-017-2299-9

