A Comprehensive Part-of-Speech Tagging to Standardize Central-Kurdish Language: A Research Guide for Kurdish Natural Language Processing Tasks

Abstract
Abstract\r\nThe field of natural language processing (NLP) has undergone significant expansion over the last decade. Many human-being applications are conducted daily via NLP tasks, starting from machine translation, speech recognition, text generation and recommendations, Part-of-Speech tagging (POS), and Named-Entity Recognition. However, low-resourced languages, such as the Central Kurdish language (CKL), remain largely unexamined due to the richness of their language and the shortage of necessary resources to support their development. The POS tagging task is the foundation of other NLP tasks; for example, the POS tag set has been used to standardize languages, providing the relationship between words within sentences, which is then followed by machine translation and text recommendation. Specifically, for the CKL, most of the utilized or provided POS tagsets are neither standardized nor comprehensive. To this end, this study presents an accurate and comprehensive POS tagset for the CKL, aiming to improve the performance of Kurdish NLP tasks. The article also collected most of the POS tags from various studies, as well as from Kurdish linguistic experts, to standardize part-of-speech tags. The developed POS tagset is designed to annotate a large CKL corpus, supporting Kurdish NLP tasks. The initial investigations of this study, conducted by comparing the listed POS tagset with the Universal Dependencies framework for standard languages, indicate that it can streamline or correct sentences more accurately for Kurdish NLP tasks

Author
Safar Maghdid Asaad

DOI
https://doi.org/10.53898/josse2025531

ISSN
2789-634X

Publish Date: 30-Dec-2025

اتصل بنا

التسجيل: +964 750 3000 600
التسجيل: +964 750 3000 700
الرئاسة: +964 750 3000 800

ارسل لنا عبر البريد الإلكتروني

[email protected]