Standard

An Experimental Study of Automating Explanatory Dictionary Compilation with Language Models. / Garipov, Timur; Morozov, Dmitry; Gubarkova, Yana et al.

Internet and Modern Society. IMS 2025.. Springer, 2026. p. 117-131 9 (Communications in Computer and Information Science; Vol. 2671 CCIS).

Research output: Chapter in Book/Report/Conference proceedingChapterResearchpeer-review

Harvard

Garipov, T, Morozov, D, Gubarkova, Y, Kozerenko, A & Glazkova, A 2026, An Experimental Study of Automating Explanatory Dictionary Compilation with Language Models. in Internet and Modern Society. IMS 2025.., 9, Communications in Computer and Information Science, vol. 2671 CCIS, Springer, pp. 117-131, Международная объединённая научная конференция «Интернет и современное общество» (Internet and Modern Society – IMS-2025), Санкт-Петербург, Russian Federation, 23.06.2025. https://doi.org/10.1007/978-3-032-04958-2_9

APA

Garipov, T., Morozov, D., Gubarkova, Y., Kozerenko, A., & Glazkova, A. (2026). An Experimental Study of Automating Explanatory Dictionary Compilation with Language Models. In Internet and Modern Society. IMS 2025. (pp. 117-131). [9] (Communications in Computer and Information Science; Vol. 2671 CCIS). Springer. https://doi.org/10.1007/978-3-032-04958-2_9

Vancouver

Garipov T, Morozov D, Gubarkova Y, Kozerenko A, Glazkova A. An Experimental Study of Automating Explanatory Dictionary Compilation with Language Models. In Internet and Modern Society. IMS 2025.. Springer. 2026. p. 117-131. 9. (Communications in Computer and Information Science). doi: 10.1007/978-3-032-04958-2_9

Author

Garipov, Timur ; Morozov, Dmitry ; Gubarkova, Yana et al. / An Experimental Study of Automating Explanatory Dictionary Compilation with Language Models. Internet and Modern Society. IMS 2025.. Springer, 2026. pp. 117-131 (Communications in Computer and Information Science).

BibTeX

@inbook{e0e45e7a303741f6a62e4631d910b302,
title = "An Experimental Study of Automating Explanatory Dictionary Compilation with Language Models",
abstract = "The creation of explanatory dictionaries has long been a cornerstone of classical linguistics. For the Russian language alone, dozens of such dictionaries have been compiled. However, the process of dictionary compilation is highly labor-intensive, requiring significant time and expertise. Moreover, the emergence of new words and the evolution of existing ones necessitate continuous updates to keep dictionaries relevant. Existing dictionaries are not without flaws; for instance, it is not uncommon for definitions to include terms that are more complex than the word being defined. Meanwhile, generative language models have reached a level of sophistication that allows them to tackle a wide range of applied tasks with near-expert proficiency. In this study, we explored whether the task of compiling a modern explanatory dictionary can be addressed using machine learning. We focused on two specific subtasks: 1) generating definitions for words not yet included in dictionaries, and 2) producing generalized definitions based on multiple existing dictionaries. We conducted a series of experiments that involved both fine-tuning and prompt-based approaches with language models. The quality of the generated definitions was evaluated using both automated metrics and human assessment. Our results demonstrate that while traditional sequence-to-sequence models like T5 and BART struggle with producing clear and accurate definitions, large language models (LLMs) yield significantly better results. At the same time, generating definitions from scratch works noticeably worse than generalizing existing ones. Our findings highlight the potential of LLM-based methods for automating dictionary compilation and suggest promising directions for further research in AI-assisted lexicography.",
keywords = "Explanatory Dictionary, Large Language Models, Natural Language Processing",
author = "Timur Garipov and Dmitry Morozov and Yana Gubarkova and Anastasia Kozerenko and Anna Glazkova",
note = "Garipov, T., Morozov, D., Gubarkova, Y., Kozerenko, A., Glazkova, A. (2026). An Experimental Study of Automating Explanatory Dictionary Compilation with Language Models. In: Bakaev, M., et al. Internet and Modern Society. IMS 2025. Communications in Computer and Information Science, vol 2671. Springer, Cham. https://doi.org/10.1007/978-3-032-04958-2_9; Международная объединённая научная конференция «Интернет и современное общество» (Internet and Modern Society – IMS-2025), IMS-2025 ; Conference date: 23-06-2025 Through 25-06-2025",
year = "2026",
doi = "10.1007/978-3-032-04958-2_9",
language = "English",
isbn = "978-3-032-04957-5",
series = "Communications in Computer and Information Science",
publisher = "Springer",
pages = "117--131",
booktitle = "Internet and Modern Society. IMS 2025.",
address = "United States",
url = "https://ims.itmo.ru/File/docs/IMS-2025_program.pdf",

}

RIS

TY - CHAP

T1 - An Experimental Study of Automating Explanatory Dictionary Compilation with Language Models

AU - Garipov, Timur

AU - Morozov, Dmitry

AU - Gubarkova, Yana

AU - Kozerenko, Anastasia

AU - Glazkova, Anna

N1 - Conference code: XXVIII

PY - 2026

Y1 - 2026

N2 - The creation of explanatory dictionaries has long been a cornerstone of classical linguistics. For the Russian language alone, dozens of such dictionaries have been compiled. However, the process of dictionary compilation is highly labor-intensive, requiring significant time and expertise. Moreover, the emergence of new words and the evolution of existing ones necessitate continuous updates to keep dictionaries relevant. Existing dictionaries are not without flaws; for instance, it is not uncommon for definitions to include terms that are more complex than the word being defined. Meanwhile, generative language models have reached a level of sophistication that allows them to tackle a wide range of applied tasks with near-expert proficiency. In this study, we explored whether the task of compiling a modern explanatory dictionary can be addressed using machine learning. We focused on two specific subtasks: 1) generating definitions for words not yet included in dictionaries, and 2) producing generalized definitions based on multiple existing dictionaries. We conducted a series of experiments that involved both fine-tuning and prompt-based approaches with language models. The quality of the generated definitions was evaluated using both automated metrics and human assessment. Our results demonstrate that while traditional sequence-to-sequence models like T5 and BART struggle with producing clear and accurate definitions, large language models (LLMs) yield significantly better results. At the same time, generating definitions from scratch works noticeably worse than generalizing existing ones. Our findings highlight the potential of LLM-based methods for automating dictionary compilation and suggest promising directions for further research in AI-assisted lexicography.

AB - The creation of explanatory dictionaries has long been a cornerstone of classical linguistics. For the Russian language alone, dozens of such dictionaries have been compiled. However, the process of dictionary compilation is highly labor-intensive, requiring significant time and expertise. Moreover, the emergence of new words and the evolution of existing ones necessitate continuous updates to keep dictionaries relevant. Existing dictionaries are not without flaws; for instance, it is not uncommon for definitions to include terms that are more complex than the word being defined. Meanwhile, generative language models have reached a level of sophistication that allows them to tackle a wide range of applied tasks with near-expert proficiency. In this study, we explored whether the task of compiling a modern explanatory dictionary can be addressed using machine learning. We focused on two specific subtasks: 1) generating definitions for words not yet included in dictionaries, and 2) producing generalized definitions based on multiple existing dictionaries. We conducted a series of experiments that involved both fine-tuning and prompt-based approaches with language models. The quality of the generated definitions was evaluated using both automated metrics and human assessment. Our results demonstrate that while traditional sequence-to-sequence models like T5 and BART struggle with producing clear and accurate definitions, large language models (LLMs) yield significantly better results. At the same time, generating definitions from scratch works noticeably worse than generalizing existing ones. Our findings highlight the potential of LLM-based methods for automating dictionary compilation and suggest promising directions for further research in AI-assisted lexicography.

KW - Explanatory Dictionary

KW - Large Language Models

KW - Natural Language Processing

U2 - 10.1007/978-3-032-04958-2_9

DO - 10.1007/978-3-032-04958-2_9

M3 - Chapter

SN - 978-3-032-04957-5

T3 - Communications in Computer and Information Science

SP - 117

EP - 131

BT - Internet and Modern Society. IMS 2025.

PB - Springer

T2 - Международная объединённая научная конференция «Интернет и современное общество» (Internet and Modern Society – IMS-2025)

Y2 - 23 June 2025 through 25 June 2025

ER -

ID: 80997234