LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators

📅 2025-07-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Multilingual content moderation systems exhibit insufficient generalization on localized and low-resource language variants—such as Singaporean Chinese dialects, Malay variants, and Tamil—posing critical security vulnerabilities. Method: We propose a lightweight, fine-tuning-free large language model (LLM)-enhanced architecture that integrates OpenAI’s multilingual pretrained embeddings with a multi-head ordinal classifier, augmented by few-shot learning and localization-oriented data augmentation to support multilingual mixed inputs and fine-grained risk scoring. Contribution/Results: Without LLM fine-tuning, our approach significantly improves detection accuracy and inference efficiency in low-resource settings. It outperforms leading commercial and open-source systems across all 17 cross-lingual safety benchmarks. Deployed at scale in Singapore’s national content governance platform, the system demonstrates real-world operational efficacy. Model weights and a subset of annotated data are publicly released to advance research in LLM safety and localized content moderation.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Safety and RobustnessComputer Vision: Large Vision Models

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Large pretrained models with web dataUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Modern moderation systems increasingly support multiple languages, but often fail to address localisation and low-resource variants - creating safety gaps in real-world deployments. Small models offer a potential alternative to large LLMs, yet still demand considerable data and compute. We present LionGuard 2, a lightweight, multilingual moderation classifier tailored to the Singapore context, supporting English, Chinese, Malay, and partial Tamil. Built on pre-trained OpenAI embeddings and a multi-head ordinal classifier, LionGuard 2 outperforms several commercial and open-source systems across 17 benchmarks, including both Singapore-specific and public English datasets. The system is actively deployed within the Singapore Government, demonstrating practical efficacy at scale. Our findings show that high-quality local data and robust multilingual embeddings can achieve strong moderation performance, without fine-tuning large models. We release our model weights and part of our training data to support future work on LLM safety.
Problem

Research questions and friction points this paper is trying to address.

Addressing localization gaps in multilingual content moderation systems
Reducing data and compute demands for lightweight moderation models
Improving performance for low-resource languages in moderation tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Lightweight multilingual moderation classifier
Pre-trained OpenAI embeddings usage
Multi-head ordinal classifier architecture
🔎 Similar Papers
No similar papers found.