UniBERTs: Adversarial Training for Language-Universal Representations

📅 2025-03-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the high computational cost and weak cross-lingual generalization of multilingual models, this paper proposes UniBERT, a lightweight multilingual language model. Methodologically, it innovatively integrates gradient-projection-based adversarial training (FGSM/PGD) with teacher-student knowledge distillation into a unified masked language modeling framework, jointly optimized across Wikipedia corpora in 107 languages to strengthen language-agnostic representation learning. UniBERT employs shared subword tokenization and multilingual tokenization, significantly reducing both pretraining and inference overhead. Empirically, it achieves an average relative performance gain of 7.72% (p = 0.0181) over strong baselines on four cross-lingual tasks—named entity recognition, natural language inference, question answering, and semantic textual similarity—demonstrating the efficacy of synergistically enhancing adversarial robustness and knowledge transfer for multilingual representation learning.

Technology Category

Natural Language Processing: Machine Translation, Multilinguality, Cross-Lingual NLPMachine Learning: Adversarial Learning & RobustnessMultiagent Systems: Adversarial Agents

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Large pretrained models with web dataSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
This paper presents UniBERT, a compact multilingual language model that leverages an innovative training framework integrating three components: masked language modeling, adversarial training, and knowledge distillation. Pre-trained on a meticulously curated Wikipedia corpus spanning 107 languages, UniBERT is designed to reduce the computational demands of large-scale models while maintaining competitive performance across various natural language processing tasks. Comprehensive evaluations on four tasks -- named entity recognition, natural language inference, question answering, and semantic textual similarity -- demonstrate that our multilingual training strategy enhanced by an adversarial objective significantly improves cross-lingual generalization. Specifically, UniBERT models show an average relative improvement of 7.72% over traditional baselines, which achieved an average relative improvement of only 1.17%, with statistical analysis confirming the significance of these gains (p-value = 0.0181). This work highlights the benefits of combining adversarial training and knowledge distillation to build scalable and robust language models, thereby advancing the field of multilingual and cross-lingual natural language processing.
Problem

Research questions and friction points this paper is trying to address.

Develops a compact multilingual language model for NLP tasks.
Reduces computational demands while maintaining competitive performance.
Improves cross-lingual generalization using adversarial training.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adversarial training enhances multilingual generalization.
Knowledge distillation reduces computational demands effectively.
Masked language modeling supports diverse NLP tasks.
🔎 Similar Papers
No similar papers found.
Andrei-Marius Avram
Andrei-Marius Avram
Adobe
machine learningnatural language processingspeech recognition
M
Marian Lupacscu
University of Bucharest, Bucharest, Romania
Dumitru-Clementin Cercel
Dumitru-Clementin Cercel
Teaching Assistant of Computer Science, University Politehnica of Bucharest
Social Network AnalysisNatural Language ProcessingInformation RetrievalMachine Learning
I
Ionuct Mironicua
Adobe Research, Bucharest, Romania
S
Stefan Truaucsan-Matu
National University of Science and Technology POLITEHNICA Bucharest, Bucharest, Romania, Academy of Romanian Scientists (AOSR), Bucharest, Romania