🤖 AI Summary
Existing large language models and safety classifiers lack rigorous evaluation on low-resource languages—particularly Singaporean English, Mandarin, Malay, and Tamil—due to the absence of localized, multilingual safety benchmarks. Method: This paper introduces RabakBench, the first safety evaluation benchmark explicitly designed for Singapore’s multilingual sociolinguistic context. It employs a scalable, three-stage data construction pipeline: (1) LLM-driven red-teaming for adversarial prompt generation, (2) semi-automatic multi-label annotation via majority-voting LLM annotators, and (3) high-fidelity cross-lingual transfer that preserves both toxicity semantics and language-specific nuances. Contributions/Results: We release an open-source dataset comprising 5,000+ samples across four languages and six fine-grained risk categories. Empirical evaluation demonstrates substantial performance degradation of mainstream safety classifiers in this localized setting. Additionally, we establish a reusable, principled framework for constructing low-resource multilingual safety data.
📝 Abstract
Large language models (LLMs) and their safety classifiers often perform poorly on low-resource languages due to limited training data and evaluation benchmarks. This paper introduces RabakBench, a new multilingual safety benchmark localized to Singapore's unique linguistic context, covering Singlish, Chinese, Malay, and Tamil. RabakBench is constructed through a scalable three-stage pipeline: (i) Generate - adversarial example generation by augmenting real Singlish web content with LLM-driven red teaming; (ii) Label - semi-automated multi-label safety annotation using majority-voted LLM labelers aligned with human judgments; and (iii) Translate - high-fidelity translation preserving linguistic nuance and toxicity across languages. The final dataset comprises over 5,000 safety-labeled examples across four languages and six fine-grained safety categories with severity levels. Evaluations of 11 popular open-source and closed-source guardrail classifiers reveal significant performance degradation. RabakBench not only enables robust safety evaluation in Southeast Asian multilingual settings but also offers a reproducible framework for building localized safety datasets in low-resource environments. The benchmark dataset, including the human-verified translations, and evaluation code are publicly available.