Standard

An Improved Karlin Model Fit Test : Application to English and Uzbek Texts and Challenges. / Fayzullaev, Shahzod.

в: Glottometrics, Том 60, 2026, стр. 1-18.

Результаты исследований: Научные публикации в периодических изданиях › статья › Рецензирование

Harvard

APA

Vancouver

Fayzullaev S. An Improved Karlin Model Fit Test: Application to English and Uzbek Texts and Challenges. Glottometrics. 2026;60:1-18. doi: 10.53482/2026_60_430

Author

Fayzullaev, Shahzod. / An Improved Karlin Model Fit Test : Application to English and Uzbek Texts and Challenges. в: Glottometrics. 2026 ; Том 60. стр. 1-18.

BibTeX

@article{17231848a51e4a58ae4d54b4311b1868,
title = "An Improved Karlin Model Fit Test: Application to English and Uzbek Texts and Challenges",
abstract = "Zipf{\textquoteright}s law and similar frequency laws have been studied in many languages, but their behavior in Uzbek has not been investigated. In this paper, we examine the fit of the generalized Karlin model of Zipf{\textquoteright}s law to Uzbek texts. For 386 texts consisting of three different genres (prose, poetry, and newspapers), we compute the Tn and Hn statistics and their analogs in the first half of the text and propose a new goodness-of-fit statistic Qn based on their joint asymptotic behaviour. Our results show that model fit varies systematically with text length and genre: newspapers and poetry, which are typically shorter and more thematically compact, fit the model significantly better than long narrative prose. These results clarify how the Karlin model works in Uzbek texts and provide an empirical baseline for future comparative studies of understudied languages, including agglutinative ones. We also compare the new statistic with two previously proposed tests, based on type and hapax, on ten English reference texts. The results show that Qn produces virtually identical p-values to the second test and leads to the same accept/reject decisions at standard significance levels, while the first test is systematically more conservative.",
keywords = "Karlin model, Uzbek texts, goodness-of-fit, hapax legomena, quantitative linguistics, vocabulary growth, Модель Карлина, рост словарного запаса, гапакс легомена, критерий согласия, количественная лингвистика, узбекские тексты",
author = "Shahzod Fayzullaev",
year = "2026",
doi = "10.53482/2026_60_430",
language = "English",
volume = "60",
pages = "1--18",
journal = "Glottometrics",
issn = "2625-8226",
publisher = "International Quantitative Linguistics Association",

}

RIS

TY - JOUR

T1 - An Improved Karlin Model Fit Test

T2 - Application to English and Uzbek Texts and Challenges

AU - Fayzullaev, Shahzod

PY - 2026

Y1 - 2026

N2 - Zipf’s law and similar frequency laws have been studied in many languages, but their behavior in Uzbek has not been investigated. In this paper, we examine the fit of the generalized Karlin model of Zipf’s law to Uzbek texts. For 386 texts consisting of three different genres (prose, poetry, and newspapers), we compute the Tn and Hn statistics and their analogs in the first half of the text and propose a new goodness-of-fit statistic Qn based on their joint asymptotic behaviour. Our results show that model fit varies systematically with text length and genre: newspapers and poetry, which are typically shorter and more thematically compact, fit the model significantly better than long narrative prose. These results clarify how the Karlin model works in Uzbek texts and provide an empirical baseline for future comparative studies of understudied languages, including agglutinative ones. We also compare the new statistic with two previously proposed tests, based on type and hapax, on ten English reference texts. The results show that Qn produces virtually identical p-values to the second test and leads to the same accept/reject decisions at standard significance levels, while the first test is systematically more conservative.

AB - Zipf’s law and similar frequency laws have been studied in many languages, but their behavior in Uzbek has not been investigated. In this paper, we examine the fit of the generalized Karlin model of Zipf’s law to Uzbek texts. For 386 texts consisting of three different genres (prose, poetry, and newspapers), we compute the Tn and Hn statistics and their analogs in the first half of the text and propose a new goodness-of-fit statistic Qn based on their joint asymptotic behaviour. Our results show that model fit varies systematically with text length and genre: newspapers and poetry, which are typically shorter and more thematically compact, fit the model significantly better than long narrative prose. These results clarify how the Karlin model works in Uzbek texts and provide an empirical baseline for future comparative studies of understudied languages, including agglutinative ones. We also compare the new statistic with two previously proposed tests, based on type and hapax, on ten English reference texts. The results show that Qn produces virtually identical p-values to the second test and leads to the same accept/reject decisions at standard significance levels, while the first test is systematically more conservative.

KW - Karlin model

KW - Uzbek texts

KW - goodness-of-fit

KW - hapax legomena

KW - quantitative linguistics

KW - vocabulary growth

KW - Модель Карлина

KW - рост словарного запаса

KW - гапакс легомена

KW - критерий согласия

KW - количественная лингвистика

KW - узбекские тексты

UR - https://www.mendeley.com/catalogue/1ffbc2f7-b4c6-3381-a339-fff6bdcb8996/

UR - https://www.scopus.com/pages/publications/105038621669

U2 - 10.53482/2026_60_430

DO - 10.53482/2026_60_430

M3 - Article

VL - 60

SP - 1

EP - 18

JO - Glottometrics

JF - Glottometrics

SN - 2625-8226

ER -

ID: 83360536