🤖 AI Summary
To address challenges in part-of-speech (POS) tagging for low-resource languages—including difficult cross-lingual transfer, high adaptation cost, and label imbalance—this paper proposes a language-agnostic, modular, lightweight Transformer framework. The architecture decouples language-specific preprocessing, label-space mapping, and the model backbone, enabling rapid adaptation to new languages with minimal code changes. It employs a unified cross-lingual tagset and targeted data augmentation to enhance robustness under data scarcity and partial language overlap. Evaluated on Bangla and Hindi, the framework achieves 96.85% and 97.00% token-level accuracy, respectively, with stable F1 scores. This work significantly lowers the barrier to model design and hyperparameter tuning for low-resource POS tagging, providing a reusable, extensible infrastructure for cross-lingual NLP.
📝 Abstract
This study proposes a language-agnostic transformer-based POS tagging framework designed for low-resource languages, using Bangla and Hindi as case studies. With only three lines of framework-specific code, the model was adapted from Bangla to Hindi, demonstrating effective portability with minimal modification. The framework achieves 96.85 percent and 97 percent token-level accuracy across POS categories in Bangla and Hindi while sustaining strong F1 scores despite dataset imbalance and linguistic overlap. A performance discrepancy in a specific POS category underscores ongoing challenges in dataset curation. The strong results stem from the underlying transformer architecture, which can be replaced with limited code adjustments. Its modular and open-source design enables rapid cross-lingual adaptation while reducing model design and tuning overhead, allowing researchers to focus on linguistic preprocessing and dataset refinement, which are essential for advancing NLP in underrepresented languages.