Standard

Belarusian Word Forms Morpheme Segmentation: Data and Evaluation. / Morozov, Dmitry A.; Shcherbakova, Olga A.; Astapenka, Lizaveta.

In: Компьютерная лингвистика и интеллектуальные технологии, No. 24, 2026, p. 429-437.

Research output: Contribution to journalConference articlepeer-review

Harvard

Morozov, DA, Shcherbakova, OA & Astapenka, L 2026, 'Belarusian Word Forms Morpheme Segmentation: Data and Evaluation', Компьютерная лингвистика и интеллектуальные технологии, no. 24, pp. 429-437. https://doi.org/10.29003/2075-7182-2026-24-429-437

APA

Morozov, D. A., Shcherbakova, O. A., & Astapenka, L. (2026). Belarusian Word Forms Morpheme Segmentation: Data and Evaluation. Компьютерная лингвистика и интеллектуальные технологии, (24), 429-437. https://doi.org/10.29003/2075-7182-2026-24-429-437

Vancouver

Morozov DA, Shcherbakova OA, Astapenka L. Belarusian Word Forms Morpheme Segmentation: Data and Evaluation. Компьютерная лингвистика и интеллектуальные технологии. 2026;(24):429-437. doi: 10.29003/2075-7182-2026-24-429-437

Author

Morozov, Dmitry A. ; Shcherbakova, Olga A. ; Astapenka, Lizaveta. / Belarusian Word Forms Morpheme Segmentation: Data and Evaluation. In: Компьютерная лингвистика и интеллектуальные технологии. 2026 ; No. 24. pp. 429-437.

BibTeX

@article{429cae1984b543be801742ec144bd517,
title = "Belarusian Word Forms Morpheme Segmentation: Data and Evaluation",
abstract = "Morpheme segmentation is a crucial step for morphological analysis and linguistically informed subword tokenization, particularly for highly inflectional languages. However, research on low-resource languages like Belarusian has been hindered by the lack of word-form datasets, with existing resources limited strictly to dictionary lemmata. To address this gap, we introduce Slounik-Wordform, the first large-scale dataset of Belarusian word-form morpheme segmentations. It comprises 332,497 expertly validated entries generated semi-automatically from a lemmabased dictionary. Using this novel resource, we benchmark several segmentation architectures, including CNNs, LSTMs, and various monolingual and multilingual BERT-like models. To robustly assess the models{\textquoteright} generalization capability to out-of-vocabulary morphemes, we evaluate them under three distinct data-splitting scenarios: random, lemma-based, and root-based. In contrast to prior studies on lemmatized data, our results demonstrate that fine-tuning large multilingual BERT-like models significantly outperforms traditional neural networks. Specifically, the XLM-RoBERTa-large model achieves state-of-the-art performance, reaching a word-level accuracy of 99.1% on the random split, 92.5% on the lemma-based split, and 77.7% on the root-based split.",
keywords = "automated morpheme segmentation, Belarusian language, low-resource languages, morphological analysis, surface morpheme segmentation, tokenization, автоматическая морфемная сегментация, белорусский язык, малоресурсные языки, морфологический анализ, поверхностная морфемная сегментация, токенизация",
author = "Morozov, {Dmitry A.} and Shcherbakova, {Olga A.} and Lizaveta Astapenka",
year = "2026",
doi = "10.29003/2075-7182-2026-24-429-437",
language = "English",
pages = "429--437",
journal = "Компьютерная лингвистика и интеллектуальные технологии",
issn = "2221-7932",
publisher = "Komp'juternaja Lingvistika i Intellektual'nye Tehnologii",
number = "24",
note = "International Conference “Dialogue 2026”: Computational Linguistics and Intellectual Technologies, Dialogue 2026 ; Conference date: 24-06-2026 Through 26-06-2026",
url = "https://dialogue-conf.org/ru/dialogue-2026-information/",

}

RIS

TY - JOUR

T1 - Belarusian Word Forms Morpheme Segmentation: Data and Evaluation

AU - Morozov, Dmitry A.

AU - Shcherbakova, Olga A.

AU - Astapenka, Lizaveta

N1 - Conference code: 31

PY - 2026

Y1 - 2026

N2 - Morpheme segmentation is a crucial step for morphological analysis and linguistically informed subword tokenization, particularly for highly inflectional languages. However, research on low-resource languages like Belarusian has been hindered by the lack of word-form datasets, with existing resources limited strictly to dictionary lemmata. To address this gap, we introduce Slounik-Wordform, the first large-scale dataset of Belarusian word-form morpheme segmentations. It comprises 332,497 expertly validated entries generated semi-automatically from a lemmabased dictionary. Using this novel resource, we benchmark several segmentation architectures, including CNNs, LSTMs, and various monolingual and multilingual BERT-like models. To robustly assess the models’ generalization capability to out-of-vocabulary morphemes, we evaluate them under three distinct data-splitting scenarios: random, lemma-based, and root-based. In contrast to prior studies on lemmatized data, our results demonstrate that fine-tuning large multilingual BERT-like models significantly outperforms traditional neural networks. Specifically, the XLM-RoBERTa-large model achieves state-of-the-art performance, reaching a word-level accuracy of 99.1% on the random split, 92.5% on the lemma-based split, and 77.7% on the root-based split.

AB - Morpheme segmentation is a crucial step for morphological analysis and linguistically informed subword tokenization, particularly for highly inflectional languages. However, research on low-resource languages like Belarusian has been hindered by the lack of word-form datasets, with existing resources limited strictly to dictionary lemmata. To address this gap, we introduce Slounik-Wordform, the first large-scale dataset of Belarusian word-form morpheme segmentations. It comprises 332,497 expertly validated entries generated semi-automatically from a lemmabased dictionary. Using this novel resource, we benchmark several segmentation architectures, including CNNs, LSTMs, and various monolingual and multilingual BERT-like models. To robustly assess the models’ generalization capability to out-of-vocabulary morphemes, we evaluate them under three distinct data-splitting scenarios: random, lemma-based, and root-based. In contrast to prior studies on lemmatized data, our results demonstrate that fine-tuning large multilingual BERT-like models significantly outperforms traditional neural networks. Specifically, the XLM-RoBERTa-large model achieves state-of-the-art performance, reaching a word-level accuracy of 99.1% on the random split, 92.5% on the lemma-based split, and 77.7% on the root-based split.

KW - automated morpheme segmentation

KW - Belarusian language

KW - low-resource languages

KW - morphological analysis

KW - surface morpheme segmentation

KW - tokenization

KW - автоматическая морфемная сегментация

KW - белорусский язык

KW - малоресурсные языки

KW - морфологический анализ

KW - поверхностная морфемная сегментация

KW - токенизация

UR - https://www.mendeley.com/catalogue/23d83310-eff8-3474-bc7d-467b07f7b280/

U2 - 10.29003/2075-7182-2026-24-429-437

DO - 10.29003/2075-7182-2026-24-429-437

M3 - Conference article

SP - 429

EP - 437

JO - Компьютерная лингвистика и интеллектуальные технологии

JF - Компьютерная лингвистика и интеллектуальные технологии

SN - 2221-7932

IS - 24

T2 - International Conference “Dialogue 2026”: Computational Linguistics and Intellectual Technologies

Y2 - 24 June 2026 through 26 June 2026

ER -

ID: 82775810