Unsupervised Cross-Lingual Part-of-Speech Tagging with Monolingual Corpora Only

📅 2026-02-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes an unsupervised cross-lingual part-of-speech (POS) tagging framework that operates without parallel corpora, addressing the challenge of scarce labeled data in low-resource languages. By leveraging unsupervised neural machine translation (UNMT), the method generates pseudo-parallel sentence pairs from monolingual corpora alone. It then integrates word alignment with a multi-source cross-lingual label projection mechanism to accurately transfer POS tags to the target language. Notably, this approach achieves effective cross-lingual POS tagging for the first time without relying on genuine parallel data. Evaluated across 28 language pairs, the framework matches or surpasses the performance of baseline methods that depend on authentic parallel corpora, yielding an average accuracy improvement of 1.3%.

Technology Category

Natural Language Processing: Machine Translation, Multilinguality, Cross-Lingual NLPMachine Learning: Multimodal LearningComputer Vision: Language and Vision

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataSemantics and Knowledge: Methods, algorithms and applications for the development of semantic models, knowledge graphs and other forms of structured data models with machine-interpretable semantics
📝 Abstract
Due to the scarcity of part-of-speech annotated data, existing studies on low-resource languages typically adopt unsupervised approaches for POS tagging. Among these, POS tag projection with word alignment method transfers POS tags from a high-resource source language to a low-resource target language based on parallel corpora, making it particularly suitable for low-resource language settings. However, this approach relies heavily on parallel corpora, which are often unavailable for many low-resource languages. To overcome this limitation, we propose a fully unsupervised cross-lingual part-of-speech(POS) tagging framework that relies solely on monolingual corpora by leveraging unsupervised neural machine translation(UNMT) system. This UNMT system first translates sentences from a high-resource language into a low-resource one, thereby constructing pseudo-parallel sentence pairs. Then, we train a POS tagger for the target language following the standard projection procedure based on word alignments. Moreover, we propose a multi-source projection technique to calibrate the projected POS tags on the target side, enhancing to train a more effective POS tagger. We evaluate our framework on 28 language pairs, covering four source languages (English, German, Spanish and French) and seven target languages (Afrikaans, Basque, Finnis, Indonesian, Lithuanian, Portuguese and Turkish). Experimental results show that our method can achieve performance comparable to the baseline cross-lingual POS tagger with parallel sentence pairs, and even exceeds it for certain target languages. Furthermore, our proposed multi-source projection technique further boosts performance, yielding an average improvement of 1.3% over previous methods.
Problem

Research questions and friction points this paper is trying to address.

cross-lingual POS tagging
low-resource languages
parallel corpora scarcity
unsupervised POS tagging
monolingual corpora
Innovation

Methods, ideas, or system contributions that make the work stand out.

unsupervised cross-lingual POS tagging
monolingual corpora
unsupervised neural machine translation
pseudo-parallel corpus
multi-source projection
🔎 Similar Papers
2017-08-30Conference on Empirical Methods in Natural Language ProcessingCitations: 73
💼 Related Jobs
No related jobs found.
J
Jianyu Zheng
School of Foreign Languages, University of Electronic Science and Technology of China, Chengdu, Sichuan Province, 611731 CN, China; School of Computer Science and Engineering, University of Electronic Science and Technology of China, Chengdu, Sichuan Province, 611731 CN, China