SEA-CLIP-Tiny: Efficient Multilingual Text-Vision Embedding for Southeast Asian Languages

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of data scarcity and limited model support in multimodal retrieval for Southeast Asian languages by proposing a compact cross-lingual image-text retrieval model with fewer than 50 million parameters. Methodologically, it introduces a region-aware training strategy alongside a curated localized dataset, and employs an enhanced CLIP-KD framework that leverages knowledge distillation from a multilingual teacher model to achieve efficient cross-lingual alignment. Experimental results demonstrate that the proposed model attains state-of-the-art average retrieval performance across seven Southeast Asian languages, outperforming MobileCLIP2 by 12.1 percentage points in R@10 while maintaining lower inference latency. These findings establish the model as a high-performance, lightweight solution well-suited for resource-constrained deployment scenarios.
📝 Abstract
Multilingual text-vision embedding models are essential for cross-lingual image-text retrieval, but Southeast Asian languages remain poorly supported due to the region's linguistic diversity and limited data and computing resources. In this paper, we introduce SEA-CLIP-Tiny, a compact multilingual text-vision embedding model for Southeast Asia with fewer than 50M parameters. Our model adapts a CLIP-KD-style framework to Southeast Asian multilingual settings through regional data curation and multilingual teacher guidance. Experiments across seven Southeast Asian languages show that SEA-CLIP-Tiny achieves the strongest average retrieval performance among the evaluated student models, reaching 12.9%, 31.5%, and 42.2% at R@1, R@5, and R@10, respectively. Compared with MobileCLIP2, it improves average R@10 by 12.1 points while using 38.4% fewer parameters and lower measured CPU latency. These results highlight the importance of region-aware training for efficient multilingual text-vision models in Southeast Asia.
Problem

Research questions and friction points this paper is trying to address.

Multilingual text-vision embedding
Southeast Asian languages
Cross-lingual image-text retrieval
Resource-constrained
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multilingual Text-Vision Embedding
Knowledge Distillation
Southeast Asian Languages
Lightweight Model
Cross-modal Retrieval
💼 Related Jobs
No related jobs found.
P
Puja Ahmad Habibi
SEACrowd
F
Faiz Assabil Firdaus
SEACrowd, University of Indonesia
A
Ashvanth S
SEACrowd, Cohere Labs Community, Technical University of Denmark
Ekapol Chuangsuwanich
Ekapol Chuangsuwanich
Chulalongkorn University
Speech ProcessingNatural Language ProcessingMedical AI
P
Pume Tuchinda
SEACrowd, Vidyasirimedhi Institute of Science and Technology, AI Singapore
Peerat Limkonchotiwat
Peerat Limkonchotiwat
Research Fellow, AI Singapore, National University of Singapore
Evaluation and BenchmarkRepresentation LearningLarge Language ModelMultilingual Learning