Standard

TL; DR: Text Normalization for Social Media Corpus. / Feoktistov, Grigorii; Morozov, Dmitry.

Internet and Modern Society. Springer, 2026. p. 168-176 13 (Communications in Computer and Information Science; Vol. 2671 CCIS).

Research output: Chapter in Book/Report/Conference proceedingConference contributionResearchpeer-review

Harvard

Feoktistov, G & Morozov, D 2026, TL; DR: Text Normalization for Social Media Corpus. in Internet and Modern Society., 13, Communications in Computer and Information Science, vol. 2671 CCIS, Springer, pp. 168-176, Международная объединённая научная конференция «Интернет и современное общество» (Internet and Modern Society – IMS-2025), Санкт-Петербург, Russian Federation, 23.06.2025. https://doi.org/10.1007/978-3-032-04958-2_13

APA

Feoktistov, G., & Morozov, D. (2026). TL; DR: Text Normalization for Social Media Corpus. In Internet and Modern Society (pp. 168-176). [13] (Communications in Computer and Information Science; Vol. 2671 CCIS). Springer. https://doi.org/10.1007/978-3-032-04958-2_13

Vancouver

Feoktistov G, Morozov D. TL; DR: Text Normalization for Social Media Corpus. In Internet and Modern Society. Springer. 2026. p. 168-176. 13. (Communications in Computer and Information Science). doi: 10.1007/978-3-032-04958-2_13

Author

Feoktistov, Grigorii ; Morozov, Dmitry. / TL; DR: Text Normalization for Social Media Corpus. Internet and Modern Society. Springer, 2026. pp. 168-176 (Communications in Computer and Information Science).

BibTeX

@inproceedings{2790a9e9d77a43e99bbf8b8fcce3ef43,
title = "TL; DR: Text Normalization for Social Media Corpus",
abstract = "Text annotation is the process of enriching text with metalinguistic information. The most common types of annotation include morphological annotation, which involves assigning grammatical tags to words; syntactic annotation, which identifies syntactic relationships within sentences; and lemmatization, which determines the base form (lemma) of each word. Text annotation is a fundamental task in computational linguistics and plays a crucial role in both linguistic research and practical applications of natural language processing (NLP). Given the scale of modern text corpora (exceeding 1 billion words), manual annotation is virtually unfeasible, making automated annotation tools the only viable solution. However, most of these tools have been developed, trained, and tested on texts with standard orthography. As a result, their performance can significantly degrade when applied to non-standard texts, such as those from social media. This issue can be mitigated through automated text normalization prior to annotation. This task is related to automated spelling correction but is not equivalent to it, and remains significantly less studied. For instance, text normalization requires not only correcting mistyping and spelling errors but also restoring abbreviations commonly used in online communication. To explore approaches to text normalization, we compiled a corpus of social media sentences and manually paired them with their normalized versions. Using this corpus, we compared various text normalization methods, leveraging pre-trained language models. We examined both fine-tuning and prompt-based approaches. Our study has identified the most effective strategies in this domain, providing valuable insights for future research and applications.",
keywords = "Corpus Linguistics, Russian Language, Social Networks, Text Normalization",
author = "Grigorii Feoktistov and Dmitry Morozov",
note = "Feoktistov, G., Morozov, D. (2026). TL; DR: Text Normalization for Social Media Corpus. In: Bakaev, M., et al. Internet and Modern Society. IMS 2025. Communications in Computer and Information Science, vol 2671. Springer, Cham. https://doi.org/10.1007/978-3-032-04958-2_13; Международная объединённая научная конференция «Интернет и современное общество» (Internet and Modern Society – IMS-2025), IMS-2025 ; Conference date: 23-06-2025 Through 25-06-2025",
year = "2026",
doi = "10.1007/978-3-032-04958-2_13",
language = "English",
isbn = "9783032049575",
series = "Communications in Computer and Information Science",
publisher = "Springer",
pages = "168--176",
booktitle = "Internet and Modern Society",
address = "United States",
url = "https://ims.itmo.ru/File/docs/IMS-2025_program.pdf",

}

RIS

TY - GEN

T1 - TL; DR: Text Normalization for Social Media Corpus

AU - Feoktistov, Grigorii

AU - Morozov, Dmitry

N1 - Conference code: XXVIII

PY - 2026

Y1 - 2026

N2 - Text annotation is the process of enriching text with metalinguistic information. The most common types of annotation include morphological annotation, which involves assigning grammatical tags to words; syntactic annotation, which identifies syntactic relationships within sentences; and lemmatization, which determines the base form (lemma) of each word. Text annotation is a fundamental task in computational linguistics and plays a crucial role in both linguistic research and practical applications of natural language processing (NLP). Given the scale of modern text corpora (exceeding 1 billion words), manual annotation is virtually unfeasible, making automated annotation tools the only viable solution. However, most of these tools have been developed, trained, and tested on texts with standard orthography. As a result, their performance can significantly degrade when applied to non-standard texts, such as those from social media. This issue can be mitigated through automated text normalization prior to annotation. This task is related to automated spelling correction but is not equivalent to it, and remains significantly less studied. For instance, text normalization requires not only correcting mistyping and spelling errors but also restoring abbreviations commonly used in online communication. To explore approaches to text normalization, we compiled a corpus of social media sentences and manually paired them with their normalized versions. Using this corpus, we compared various text normalization methods, leveraging pre-trained language models. We examined both fine-tuning and prompt-based approaches. Our study has identified the most effective strategies in this domain, providing valuable insights for future research and applications.

AB - Text annotation is the process of enriching text with metalinguistic information. The most common types of annotation include morphological annotation, which involves assigning grammatical tags to words; syntactic annotation, which identifies syntactic relationships within sentences; and lemmatization, which determines the base form (lemma) of each word. Text annotation is a fundamental task in computational linguistics and plays a crucial role in both linguistic research and practical applications of natural language processing (NLP). Given the scale of modern text corpora (exceeding 1 billion words), manual annotation is virtually unfeasible, making automated annotation tools the only viable solution. However, most of these tools have been developed, trained, and tested on texts with standard orthography. As a result, their performance can significantly degrade when applied to non-standard texts, such as those from social media. This issue can be mitigated through automated text normalization prior to annotation. This task is related to automated spelling correction but is not equivalent to it, and remains significantly less studied. For instance, text normalization requires not only correcting mistyping and spelling errors but also restoring abbreviations commonly used in online communication. To explore approaches to text normalization, we compiled a corpus of social media sentences and manually paired them with their normalized versions. Using this corpus, we compared various text normalization methods, leveraging pre-trained language models. We examined both fine-tuning and prompt-based approaches. Our study has identified the most effective strategies in this domain, providing valuable insights for future research and applications.

KW - Corpus Linguistics

KW - Russian Language

KW - Social Networks

KW - Text Normalization

UR - https://www.mendeley.com/catalogue/ab62238a-917c-32ac-963c-d175362f45e2/

U2 - 10.1007/978-3-032-04958-2_13

DO - 10.1007/978-3-032-04958-2_13

M3 - Conference contribution

SN - 9783032049575

T3 - Communications in Computer and Information Science

SP - 168

EP - 176

BT - Internet and Modern Society

PB - Springer

T2 - Международная объединённая научная конференция «Интернет и современное общество» (Internet and Modern Society – IMS-2025)

Y2 - 23 June 2025 through 25 June 2025

ER -

ID: 81944763