Результаты исследований: Научные публикации в периодических изданиях › статья › Рецензирование
An Improved Karlin Model Fit Test : Application to English and Uzbek Texts and Challenges. / Fayzullaev, Shahzod.
в: Glottometrics, Том 60, 2026, стр. 1-18.Результаты исследований: Научные публикации в периодических изданиях › статья › Рецензирование
}
TY - JOUR
T1 - An Improved Karlin Model Fit Test
T2 - Application to English and Uzbek Texts and Challenges
AU - Fayzullaev, Shahzod
PY - 2026
Y1 - 2026
N2 - Zipf’s law and similar frequency laws have been studied in many languages, but their behavior in Uzbek has not been investigated. In this paper, we examine the fit of the generalized Karlin model of Zipf’s law to Uzbek texts. For 386 texts consisting of three different genres (prose, poetry, and newspapers), we compute the Tn and Hn statistics and their analogs in the first half of the text and propose a new goodness-of-fit statistic Qn based on their joint asymptotic behaviour. Our results show that model fit varies systematically with text length and genre: newspapers and poetry, which are typically shorter and more thematically compact, fit the model significantly better than long narrative prose. These results clarify how the Karlin model works in Uzbek texts and provide an empirical baseline for future comparative studies of understudied languages, including agglutinative ones. We also compare the new statistic with two previously proposed tests, based on type and hapax, on ten English reference texts. The results show that Qn produces virtually identical p-values to the second test and leads to the same accept/reject decisions at standard significance levels, while the first test is systematically more conservative.
AB - Zipf’s law and similar frequency laws have been studied in many languages, but their behavior in Uzbek has not been investigated. In this paper, we examine the fit of the generalized Karlin model of Zipf’s law to Uzbek texts. For 386 texts consisting of three different genres (prose, poetry, and newspapers), we compute the Tn and Hn statistics and their analogs in the first half of the text and propose a new goodness-of-fit statistic Qn based on their joint asymptotic behaviour. Our results show that model fit varies systematically with text length and genre: newspapers and poetry, which are typically shorter and more thematically compact, fit the model significantly better than long narrative prose. These results clarify how the Karlin model works in Uzbek texts and provide an empirical baseline for future comparative studies of understudied languages, including agglutinative ones. We also compare the new statistic with two previously proposed tests, based on type and hapax, on ten English reference texts. The results show that Qn produces virtually identical p-values to the second test and leads to the same accept/reject decisions at standard significance levels, while the first test is systematically more conservative.
KW - Karlin model
KW - Uzbek texts
KW - goodness-of-fit
KW - hapax legomena
KW - quantitative linguistics
KW - vocabulary growth
KW - Модель Карлина
KW - рост словарного запаса
KW - гапакс легомена
KW - критерий согласия
KW - количественная лингвистика
KW - узбекские тексты
UR - https://www.mendeley.com/catalogue/1ffbc2f7-b4c6-3381-a339-fff6bdcb8996/
UR - https://www.scopus.com/pages/publications/105038621669
U2 - 10.53482/2026_60_430
DO - 10.53482/2026_60_430
M3 - Article
VL - 60
SP - 1
EP - 18
JO - Glottometrics
JF - Glottometrics
SN - 2625-8226
ER -
ID: 83360536