🤖 AI Summary
Multilingual content moderation systems exhibit insufficient generalization on localized and low-resource language variants—such as Singaporean Chinese dialects, Malay variants, and Tamil—posing critical security vulnerabilities. Method: We propose a lightweight, fine-tuning-free large language model (LLM)-enhanced architecture that integrates OpenAI’s multilingual pretrained embeddings with a multi-head ordinal classifier, augmented by few-shot learning and localization-oriented data augmentation to support multilingual mixed inputs and fine-grained risk scoring. Contribution/Results: Without LLM fine-tuning, our approach significantly improves detection accuracy and inference efficiency in low-resource settings. It outperforms leading commercial and open-source systems across all 17 cross-lingual safety benchmarks. Deployed at scale in Singapore’s national content governance platform, the system demonstrates real-world operational efficacy. Model weights and a subset of annotated data are publicly released to advance research in LLM safety and localized content moderation.
📝 Abstract
Modern moderation systems increasingly support multiple languages, but often fail to address localisation and low-resource variants - creating safety gaps in real-world deployments. Small models offer a potential alternative to large LLMs, yet still demand considerable data and compute. We present LionGuard 2, a lightweight, multilingual moderation classifier tailored to the Singapore context, supporting English, Chinese, Malay, and partial Tamil. Built on pre-trained OpenAI embeddings and a multi-head ordinal classifier, LionGuard 2 outperforms several commercial and open-source systems across 17 benchmarks, including both Singapore-specific and public English datasets. The system is actively deployed within the Singapore Government, demonstrating practical efficacy at scale. Our findings show that high-quality local data and robust multilingual embeddings can achieve strong moderation performance, without fine-tuning large models. We release our model weights and part of our training data to support future work on LLM safety.