Top

Linguistic Foundations Of Electronic Thesaurus Construction: Structure, Synset-Formation Methodology, And Corpus-Based Evidence From The Uzbek Language

Authors

Muhammadjon Najmiddinov

Files

PDF Certificate

Abstract

This article presents a comprehensive linguistic and technological analysis of the principles underlying the construction of electronic thesauri. The purpose of the study is threefold: (1) to characterise the structural and functional properties that distinguish an electronic thesaurus from a traditional dictionary; (2) to develop and empirically validate a step-by-step methodology for selecting thesaurus units and forming synsets, integrating corpus statistics, distributional semantics and transformer-based (BERT-type) models with expert verification; and (3) to describe how polysemy, synonymy, homonymy and ontological categorisation are represented within the resulting thesaurus model for four major word classes — object nouns, attributive adjectives, verbs of action/state, and adverbs. Descriptive, componential, distributive, ideographic, ontological and quantitative methods of analysis were applied to material drawn from the Uzbek Educational Corpus and the ARANEUM_UZBEKICUM corpus. As a result, 77 synsets were constructed across five lexical categories, three structural models of thesaurus organisation (hierarchical tree, network graph, and matrix) were formalised, and a six-stage synset-formation methodology was proposed and tested. The findings provide a replicable linguistic foundation for building electronic thesauri for Uzbek and typologically related agglutinative languages, with direct applications in information retrieval, machine translation and natural-language-processing systems.

PDF

References

1. Karaulov Yu.N. Общая и русская идеография. – Moscow: Nauka, 1976. – 356 p.

2. Miller G.A. WordNet: A Lexical Database for English // Communications of the ACM. – 1995. – Vol. 38, No. 11. – P. 39–41.

3. Fellbaum C. (ed.) WordNet: An Electronic Lexical Database. – Cambridge, MA: MIT Press, 1998. – 423 p.

4. Devlin J., Chang M.-W., Lee K., Toutanova K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding // Proceedings of NAACL-HLT. – 2019. – P. 4171–4186.

5. Vaswani A. et al. Attention Is All You Need // Advances in Neural Information Processing Systems. – 2017. – Vol. 30.

6. Mikolov T., Chen K., Corrado G., Dean J. Efficient Estimation of Word Representations in Vector Space // arXiv:1301.3781. – 2013.

7. Berners-Lee T., Hendler J., Lassila O. The Semantic Web // Scientific American. – 2001. – Vol. 284, No. 5. – P. 34–43.

8. Gruber T.R. A Translation Approach to Portable Ontology Specifications // Knowledge Acquisition. – 1993. – Vol. 5, No. 2. – P. 199–220.

9. Апресян Ю.Д. Лексическая семантика: синонимические средства языка. – Moscow: Nauka, 1974. – 367 p.

10. Kasares J. Diccionario ideológico de la lengua española. – Barcelona: Gustavo Gili, 1942.

11. Najmiddinov M.G'. Tezauruslar yaratishning lingvistik asoslari: Filologiya fanlari doktori (DSc) dissertatsiyasi avtoreferati. – Andijon, 2026. – 69 p.

12. Nešpore-Bērzkalne G. et al. ARANEUM corpora project documentation. – 2018.

13. “O'zbek tilining ta'limiy korpusi” [Uzbek Educational Corpus]. Electronic resource. Available at: https://uzbekcorpus.uz

14. Asadov T. Leksik-semantik guruhlarning tasnifi. – Toshkent: Fan, 2015. – 214 p.

Downloads

Download data is not yet available.

Details