Результаты исследований: Публикации в книгах, отчётах, сборниках, трудах конференций › статья в сборнике материалов конференции › научная › Рецензирование
TL; DR: Text Normalization for Social Media Corpus. / Feoktistov, Grigorii; Morozov, Dmitry.
Internet and Modern Society. Springer, 2026. стр. 168-176 13 (Communications in Computer and Information Science; Том 2671 CCIS).Результаты исследований: Публикации в книгах, отчётах, сборниках, трудах конференций › статья в сборнике материалов конференции › научная › Рецензирование
}
TY - GEN
T1 - TL; DR: Text Normalization for Social Media Corpus
AU - Feoktistov, Grigorii
AU - Morozov, Dmitry
N1 - Conference code: XXVIII
PY - 2026
Y1 - 2026
N2 - Text annotation is the process of enriching text with metalinguistic information. The most common types of annotation include morphological annotation, which involves assigning grammatical tags to words; syntactic annotation, which identifies syntactic relationships within sentences; and lemmatization, which determines the base form (lemma) of each word. Text annotation is a fundamental task in computational linguistics and plays a crucial role in both linguistic research and practical applications of natural language processing (NLP). Given the scale of modern text corpora (exceeding 1 billion words), manual annotation is virtually unfeasible, making automated annotation tools the only viable solution. However, most of these tools have been developed, trained, and tested on texts with standard orthography. As a result, their performance can significantly degrade when applied to non-standard texts, such as those from social media. This issue can be mitigated through automated text normalization prior to annotation. This task is related to automated spelling correction but is not equivalent to it, and remains significantly less studied. For instance, text normalization requires not only correcting mistyping and spelling errors but also restoring abbreviations commonly used in online communication. To explore approaches to text normalization, we compiled a corpus of social media sentences and manually paired them with their normalized versions. Using this corpus, we compared various text normalization methods, leveraging pre-trained language models. We examined both fine-tuning and prompt-based approaches. Our study has identified the most effective strategies in this domain, providing valuable insights for future research and applications.
AB - Text annotation is the process of enriching text with metalinguistic information. The most common types of annotation include morphological annotation, which involves assigning grammatical tags to words; syntactic annotation, which identifies syntactic relationships within sentences; and lemmatization, which determines the base form (lemma) of each word. Text annotation is a fundamental task in computational linguistics and plays a crucial role in both linguistic research and practical applications of natural language processing (NLP). Given the scale of modern text corpora (exceeding 1 billion words), manual annotation is virtually unfeasible, making automated annotation tools the only viable solution. However, most of these tools have been developed, trained, and tested on texts with standard orthography. As a result, their performance can significantly degrade when applied to non-standard texts, such as those from social media. This issue can be mitigated through automated text normalization prior to annotation. This task is related to automated spelling correction but is not equivalent to it, and remains significantly less studied. For instance, text normalization requires not only correcting mistyping and spelling errors but also restoring abbreviations commonly used in online communication. To explore approaches to text normalization, we compiled a corpus of social media sentences and manually paired them with their normalized versions. Using this corpus, we compared various text normalization methods, leveraging pre-trained language models. We examined both fine-tuning and prompt-based approaches. Our study has identified the most effective strategies in this domain, providing valuable insights for future research and applications.
KW - Corpus Linguistics
KW - Russian Language
KW - Social Networks
KW - Text Normalization
UR - https://www.mendeley.com/catalogue/ab62238a-917c-32ac-963c-d175362f45e2/
U2 - 10.1007/978-3-032-04958-2_13
DO - 10.1007/978-3-032-04958-2_13
M3 - Conference contribution
SN - 9783032049575
T3 - Communications in Computer and Information Science
SP - 168
EP - 176
BT - Internet and Modern Society
PB - Springer
T2 - Международная объединённая научная конференция «Интернет и современное общество» (Internet and Modern Society – IMS-2025)
Y2 - 23 June 2025 through 25 June 2025
ER -
ID: 81944763