Standard

CEFR Level Prediction for Short Russian L2 Texts: Evaluating Classifiers and Instruction-Based LLMs. / Glazkova, Anna; Laposhina, Antonina; Morozov, Dmitry.

Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026). European Language Resources Association, 2026. p. 1081-1091.

Research output: Chapter in Book/Report/Conference proceedingConference contributionResearchpeer-review

Harvard

Glazkova, A, Laposhina, A & Morozov, D 2026, CEFR Level Prediction for Short Russian L2 Texts: Evaluating Classifiers and Instruction-Based LLMs. in Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026). European Language Resources Association, pp. 1081-1091, The Fifteenth Language Resources and Evaluation Conference, Mallorca, Spain, 13.05.2026. https://doi.org/10.63317/27p9pbh4oods

APA

Glazkova, A., Laposhina, A., & Morozov, D. (2026). CEFR Level Prediction for Short Russian L2 Texts: Evaluating Classifiers and Instruction-Based LLMs. In Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026) (pp. 1081-1091). European Language Resources Association. https://doi.org/10.63317/27p9pbh4oods

Vancouver

Glazkova A, Laposhina A, Morozov D. CEFR Level Prediction for Short Russian L2 Texts: Evaluating Classifiers and Instruction-Based LLMs. In Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026). European Language Resources Association. 2026. p. 1081-1091 doi: 10.63317/27p9pbh4oods

Author

Glazkova, Anna ; Laposhina, Antonina ; Morozov, Dmitry. / CEFR Level Prediction for Short Russian L2 Texts: Evaluating Classifiers and Instruction-Based LLMs. Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026). European Language Resources Association, 2026. pp. 1081-1091

BibTeX

@inproceedings{1dfd23d0a57e45699cecdd40f976a286,
title = "CEFR Level Prediction for Short Russian L2 Texts: Evaluating Classifiers and Instruction-Based LLMs",
abstract = "This study explores the automated prediction of text complexity levels for short Russian texts on the Common European Framework of Reference for Languages (CEFR) scale. The dataset consists of 7,322 nonfictional fragments (15–30 words) extracted from textbooks for learners of Russian as a second language and filtered according to linguistic feature distributions typical of each CEFR level, with additional validation conducted by 4 human experts. Each text fragment was annotated with 127 linguistic features, including lexical, morphological, syntactic, and length-based characteristics. We evaluate several approaches to text complexity assessment: traditional machine learning classifiers, fine-tuned transformer models, and instruction-based large language models (LLMs). Among all models, RuBERT achieved the best strict F1-score (47.8%) and the lowest mean absolute error (0.56), while instruction-based LLMs such as YandexGPT captured overall complexity trends but underperformed in exact classification. Feature ablation experiments demonstrated that lexical features are the most informative for CEFR prediction. Our findings confirm that fine-tuned language models currently offer the most reliable results for short-text CEFR assessment in Russian, whereas instruction-based LLMs show potential for qualitative analysis of text difficulty patterns.",
keywords = "text complexity, CEFR levels, large language models, Russian as a second language",
author = "Anna Glazkova and Antonina Laposhina and Dmitry Morozov",
year = "2026",
doi = "10.63317/27p9pbh4oods",
language = "English",
isbn = "978-2-493814-49-4",
pages = "1081--1091",
booktitle = "Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026)",
publisher = "European Language Resources Association",
address = "Luxembourg",
note = "The Fifteenth Language Resources and Evaluation Conference, LREC 2026 ; Conference date: 13-05-2026 Through 15-05-2026",

}

RIS

TY - GEN

T1 - CEFR Level Prediction for Short Russian L2 Texts: Evaluating Classifiers and Instruction-Based LLMs

AU - Glazkova, Anna

AU - Laposhina, Antonina

AU - Morozov, Dmitry

N1 - Conference code: 15

PY - 2026

Y1 - 2026

N2 - This study explores the automated prediction of text complexity levels for short Russian texts on the Common European Framework of Reference for Languages (CEFR) scale. The dataset consists of 7,322 nonfictional fragments (15–30 words) extracted from textbooks for learners of Russian as a second language and filtered according to linguistic feature distributions typical of each CEFR level, with additional validation conducted by 4 human experts. Each text fragment was annotated with 127 linguistic features, including lexical, morphological, syntactic, and length-based characteristics. We evaluate several approaches to text complexity assessment: traditional machine learning classifiers, fine-tuned transformer models, and instruction-based large language models (LLMs). Among all models, RuBERT achieved the best strict F1-score (47.8%) and the lowest mean absolute error (0.56), while instruction-based LLMs such as YandexGPT captured overall complexity trends but underperformed in exact classification. Feature ablation experiments demonstrated that lexical features are the most informative for CEFR prediction. Our findings confirm that fine-tuned language models currently offer the most reliable results for short-text CEFR assessment in Russian, whereas instruction-based LLMs show potential for qualitative analysis of text difficulty patterns.

AB - This study explores the automated prediction of text complexity levels for short Russian texts on the Common European Framework of Reference for Languages (CEFR) scale. The dataset consists of 7,322 nonfictional fragments (15–30 words) extracted from textbooks for learners of Russian as a second language and filtered according to linguistic feature distributions typical of each CEFR level, with additional validation conducted by 4 human experts. Each text fragment was annotated with 127 linguistic features, including lexical, morphological, syntactic, and length-based characteristics. We evaluate several approaches to text complexity assessment: traditional machine learning classifiers, fine-tuned transformer models, and instruction-based large language models (LLMs). Among all models, RuBERT achieved the best strict F1-score (47.8%) and the lowest mean absolute error (0.56), while instruction-based LLMs such as YandexGPT captured overall complexity trends but underperformed in exact classification. Feature ablation experiments demonstrated that lexical features are the most informative for CEFR prediction. Our findings confirm that fine-tuned language models currently offer the most reliable results for short-text CEFR assessment in Russian, whereas instruction-based LLMs show potential for qualitative analysis of text difficulty patterns.

KW - text complexity

KW - CEFR levels

KW - large language models

KW - Russian as a second language

UR - https://www.mendeley.com/catalogue/b9e28e1d-f720-339c-bad6-343839de5418/

U2 - 10.63317/27p9pbh4oods

DO - 10.63317/27p9pbh4oods

M3 - Conference contribution

SN - 978-2-493814-49-4

SP - 1081

EP - 1091

BT - Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026)

PB - European Language Resources Association

T2 - The Fifteenth Language Resources and Evaluation Conference

Y2 - 13 May 2026 through 15 May 2026

ER -

ID: 81988204