Результаты исследований: Публикации в книгах, отчётах, сборниках, трудах конференций › глава/раздел › научная › Рецензирование
An Experimental Study of Automating Explanatory Dictionary Compilation with Language Models. / Garipov, Timur; Morozov, Dmitry; Gubarkova, Yana и др.
Internet and Modern Society. IMS 2025.. Springer, 2026. стр. 117-131 9 (Communications in Computer and Information Science; Том 2671 CCIS).Результаты исследований: Публикации в книгах, отчётах, сборниках, трудах конференций › глава/раздел › научная › Рецензирование
}
TY - CHAP
T1 - An Experimental Study of Automating Explanatory Dictionary Compilation with Language Models
AU - Garipov, Timur
AU - Morozov, Dmitry
AU - Gubarkova, Yana
AU - Kozerenko, Anastasia
AU - Glazkova, Anna
N1 - Conference code: XXVIII
PY - 2026
Y1 - 2026
N2 - The creation of explanatory dictionaries has long been a cornerstone of classical linguistics. For the Russian language alone, dozens of such dictionaries have been compiled. However, the process of dictionary compilation is highly labor-intensive, requiring significant time and expertise. Moreover, the emergence of new words and the evolution of existing ones necessitate continuous updates to keep dictionaries relevant. Existing dictionaries are not without flaws; for instance, it is not uncommon for definitions to include terms that are more complex than the word being defined. Meanwhile, generative language models have reached a level of sophistication that allows them to tackle a wide range of applied tasks with near-expert proficiency. In this study, we explored whether the task of compiling a modern explanatory dictionary can be addressed using machine learning. We focused on two specific subtasks: 1) generating definitions for words not yet included in dictionaries, and 2) producing generalized definitions based on multiple existing dictionaries. We conducted a series of experiments that involved both fine-tuning and prompt-based approaches with language models. The quality of the generated definitions was evaluated using both automated metrics and human assessment. Our results demonstrate that while traditional sequence-to-sequence models like T5 and BART struggle with producing clear and accurate definitions, large language models (LLMs) yield significantly better results. At the same time, generating definitions from scratch works noticeably worse than generalizing existing ones. Our findings highlight the potential of LLM-based methods for automating dictionary compilation and suggest promising directions for further research in AI-assisted lexicography.
AB - The creation of explanatory dictionaries has long been a cornerstone of classical linguistics. For the Russian language alone, dozens of such dictionaries have been compiled. However, the process of dictionary compilation is highly labor-intensive, requiring significant time and expertise. Moreover, the emergence of new words and the evolution of existing ones necessitate continuous updates to keep dictionaries relevant. Existing dictionaries are not without flaws; for instance, it is not uncommon for definitions to include terms that are more complex than the word being defined. Meanwhile, generative language models have reached a level of sophistication that allows them to tackle a wide range of applied tasks with near-expert proficiency. In this study, we explored whether the task of compiling a modern explanatory dictionary can be addressed using machine learning. We focused on two specific subtasks: 1) generating definitions for words not yet included in dictionaries, and 2) producing generalized definitions based on multiple existing dictionaries. We conducted a series of experiments that involved both fine-tuning and prompt-based approaches with language models. The quality of the generated definitions was evaluated using both automated metrics and human assessment. Our results demonstrate that while traditional sequence-to-sequence models like T5 and BART struggle with producing clear and accurate definitions, large language models (LLMs) yield significantly better results. At the same time, generating definitions from scratch works noticeably worse than generalizing existing ones. Our findings highlight the potential of LLM-based methods for automating dictionary compilation and suggest promising directions for further research in AI-assisted lexicography.
KW - Explanatory Dictionary
KW - Large Language Models
KW - Natural Language Processing
U2 - 10.1007/978-3-032-04958-2_9
DO - 10.1007/978-3-032-04958-2_9
M3 - Chapter
SN - 978-3-032-04957-5
T3 - Communications in Computer and Information Science
SP - 117
EP - 131
BT - Internet and Modern Society. IMS 2025.
PB - Springer
T2 - Международная объединённая научная конференция «Интернет и современное общество» (Internet and Modern Society – IMS-2025)
Y2 - 23 June 2025 through 25 June 2025
ER -
ID: 80997234