Cross-Lingual IPA Contrastive Learning for Zero-Shot NER

📅 2025-03-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses zero-shot named entity recognition (NER) for low-resource languages. To bridge the phonological representation gap between high- and low-resource languages, we propose IPAC—a phoneme-level alignment method grounded in the International Phonetic Alphabet (IPA). We introduce CONLIPA, the first cross-lingual IPA contrastive learning dataset covering ten major language families, constructed without reliance on translations or parallel corpora. IPAC aligns phonological representations across languages via IPA phoneme embeddings and contrastive learning, then integrates these with multilingual pre-trained language models for fine-tuning. Experiments demonstrate that IPAC achieves statistically significant average improvements over state-of-the-art baselines on zero-shot NER benchmarks. Results confirm that IPA-based representations effectively mitigate cross-lingual phonological divergence, establishing a transferable, translation-free paradigm for low-resource NER.

Technology Category

Natural Language Processing: Language Grounding & Multi-modal NLPMachine Learning: Large Multimodal Models (LMMs)Computer Vision: Language and Vision

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Existing approaches to zero-shot Named Entity Recognition (NER) for low-resource languages have primarily relied on machine translation, whereas more recent methods have shifted focus to phonemic representation. Building upon this, we investigate how reducing the phonemic representation gap in IPA transcription between languages with similar phonetic characteristics enables models trained on high-resource languages to perform effectively on low-resource languages. In this work, we propose CONtrastive Learning with IPA (CONLIPA) dataset containing 10 English and high resource languages IPA pairs from 10 frequently used language families. We also propose a cross-lingual IPA Contrastive learning method (IPAC) using the CONLIPA dataset. Furthermore, our proposed dataset and methodology demonstrate a substantial average gain when compared to the best performing baseline.
Problem

Research questions and friction points this paper is trying to address.

Reduces phonemic representation gap for zero-shot NER.
Uses IPA transcription to improve cross-lingual model performance.
Introduces CONLIPA dataset and IPAC method for low-resource languages.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-lingual IPA contrastive learning method
CONLIPA dataset with IPA pairs
Reduces phonemic representation gap
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Jimin Sohn
Jimin Sohn
LG Innotek
OCRNERComputer VisionSemantic Segmentation
D
David R. Mortensen
Carnegie Mellon University, USA