Результаты исследований: Научные публикации в периодических изданиях › статья по материалам конференции › Рецензирование
Belarusian Word Forms Morpheme Segmentation: Data and Evaluation. / Morozov, Dmitry A.; Shcherbakova, Olga A.; Astapenka, Lizaveta.
в: Компьютерная лингвистика и интеллектуальные технологии, № 24, 2026, стр. 429-437.Результаты исследований: Научные публикации в периодических изданиях › статья по материалам конференции › Рецензирование
}
TY - JOUR
T1 - Belarusian Word Forms Morpheme Segmentation: Data and Evaluation
AU - Morozov, Dmitry A.
AU - Shcherbakova, Olga A.
AU - Astapenka, Lizaveta
N1 - Conference code: 31
PY - 2026
Y1 - 2026
N2 - Morpheme segmentation is a crucial step for morphological analysis and linguistically informed subword tokenization, particularly for highly inflectional languages. However, research on low-resource languages like Belarusian has been hindered by the lack of word-form datasets, with existing resources limited strictly to dictionary lemmata. To address this gap, we introduce Slounik-Wordform, the first large-scale dataset of Belarusian word-form morpheme segmentations. It comprises 332,497 expertly validated entries generated semi-automatically from a lemmabased dictionary. Using this novel resource, we benchmark several segmentation architectures, including CNNs, LSTMs, and various monolingual and multilingual BERT-like models. To robustly assess the models’ generalization capability to out-of-vocabulary morphemes, we evaluate them under three distinct data-splitting scenarios: random, lemma-based, and root-based. In contrast to prior studies on lemmatized data, our results demonstrate that fine-tuning large multilingual BERT-like models significantly outperforms traditional neural networks. Specifically, the XLM-RoBERTa-large model achieves state-of-the-art performance, reaching a word-level accuracy of 99.1% on the random split, 92.5% on the lemma-based split, and 77.7% on the root-based split.
AB - Morpheme segmentation is a crucial step for morphological analysis and linguistically informed subword tokenization, particularly for highly inflectional languages. However, research on low-resource languages like Belarusian has been hindered by the lack of word-form datasets, with existing resources limited strictly to dictionary lemmata. To address this gap, we introduce Slounik-Wordform, the first large-scale dataset of Belarusian word-form morpheme segmentations. It comprises 332,497 expertly validated entries generated semi-automatically from a lemmabased dictionary. Using this novel resource, we benchmark several segmentation architectures, including CNNs, LSTMs, and various monolingual and multilingual BERT-like models. To robustly assess the models’ generalization capability to out-of-vocabulary morphemes, we evaluate them under three distinct data-splitting scenarios: random, lemma-based, and root-based. In contrast to prior studies on lemmatized data, our results demonstrate that fine-tuning large multilingual BERT-like models significantly outperforms traditional neural networks. Specifically, the XLM-RoBERTa-large model achieves state-of-the-art performance, reaching a word-level accuracy of 99.1% on the random split, 92.5% on the lemma-based split, and 77.7% on the root-based split.
KW - automated morpheme segmentation
KW - Belarusian language
KW - low-resource languages
KW - morphological analysis
KW - surface morpheme segmentation
KW - tokenization
KW - автоматическая морфемная сегментация
KW - белорусский язык
KW - малоресурсные языки
KW - морфологический анализ
KW - поверхностная морфемная сегментация
KW - токенизация
UR - https://www.mendeley.com/catalogue/23d83310-eff8-3474-bc7d-467b07f7b280/
U2 - 10.29003/2075-7182-2026-24-429-437
DO - 10.29003/2075-7182-2026-24-429-437
M3 - Conference article
SP - 429
EP - 437
JO - Компьютерная лингвистика и интеллектуальные технологии
JF - Компьютерная лингвистика и интеллектуальные технологии
SN - 2221-7932
IS - 24
T2 - International Conference “Dialogue 2026”: Computational Linguistics and Intellectual Technologies
Y2 - 24 June 2026 through 26 June 2026
ER -
ID: 82775810