Authors :
Raksana Aliashrafova; Namiq Abdurahmanov; Gunel Aghammadova; Vafa Atayeva
Volume/Issue :
Volume 11 - 2026, Issue 7 - July
Google Scholar :
https://tinyurl.com/mtrhwp9p
Scribd :
https://tinyurl.com/2mdhrh96
DOI :
https://doi.org/10.38124/ijisrt/26jul101
Note : A published paper may take 4-5 working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and ResearchGate.
Abstract :
Artificial intelligence (AI) systems, such as large-scale speech recognition systems, are vital for acquiring
linguistic competence and achieving this is essential for their use. Thus, for large-scale speech recognition systems,
acquisition of linguistic competence is essential for usage of the system. Even for the widely spoken language like English,
Mandarin and Spanish, these corpora are vast and have hundreds of billions of tokens, which can be used to power up
virtually all natural-language processing (NLP) tasks. In the online world, however, the Azerbaijani language is far less
well resourced: publicly available language corpora are limited to a small number of high-resource languages. In the
article, the main problems related to corpus scarcity, such as poor quality of machine translation, incorrect morphological
processing, recognition dialects of the Azerbaijan language and failure of chatbots to understand the Azerbaijan language
are discussed.
Keywords :
Azerbaijani Language, Natural Language Processing, Language Corpus, Low-Resource Languages, Machine Translation, Aİ Language Models, Turkic Languages.
References :
- Hajili, N. (2024). Open foundation models for the Azerbaijani language. arXiv preprint arXiv:2407.02337.
- Jafarov, Y. M., & Gasimova, R. T. (2020). Some problems in the design of domain names with the Azerbaijani alphabet on the internet. Journal of Information Society Problems, 11(1), 45–52.
- Kishiyev, H. (2023). Introducing AzCorpus: A game-changing resource for NLP applications in the Azerbaijani language. Medium.
- Liu, Y., Ott, M., Goyal, N., et al. (2019). RoBERTa: A robustly optimised BERT pretraining approach. arXiv preprint arXiv:1907.11692.
- Nouri, M., Amani, M., Zohrabi, R., & Asgari, E. (2023). The language model, resources, and computational pipelines for the under-resourced Iranian Azerbaijani. Findings of ACL: AACL-IJCNLP 2023.
- Salimov, R., & Mammadova, G. (2023). Evaluation of neural machine translation for the Azerbaijani–English language pair. Proceedings of the Workshop on Language Resources for Turkic Languages.
- Sorokin, A., Shavrina, T., & Lyashevskaya, O. (2021). A large-scale study of machine translation in the Turkic languages. Proceedings of ACL 2021.
- Karimova, S., & Pecina, P. (2023). Integrated approach to adapting open-source AI models for machine translation of low-resource Turkic languages. Computers, 15(2), 73.
- Kartal, J., et al. (2024). Enhancing language learning through technology: A new English–Azerbaijani parallel corpus. arXiv preprint arXiv:2407.05189.
Artificial intelligence (AI) systems, such as large-scale speech recognition systems, are vital for acquiring
linguistic competence and achieving this is essential for their use. Thus, for large-scale speech recognition systems,
acquisition of linguistic competence is essential for usage of the system. Even for the widely spoken language like English,
Mandarin and Spanish, these corpora are vast and have hundreds of billions of tokens, which can be used to power up
virtually all natural-language processing (NLP) tasks. In the online world, however, the Azerbaijani language is far less
well resourced: publicly available language corpora are limited to a small number of high-resource languages. In the
article, the main problems related to corpus scarcity, such as poor quality of machine translation, incorrect morphological
processing, recognition dialects of the Azerbaijan language and failure of chatbots to understand the Azerbaijan language
are discussed.
Keywords :
Azerbaijani Language, Natural Language Processing, Language Corpus, Low-Resource Languages, Machine Translation, Aİ Language Models, Turkic Languages.