Score
Designs and implements benchmarks, audits, and evaluation pipelines that measure and compare model safety and robustness across languages, scripts, and cultural regions, including multilingual and vision–language models, code‑switched input, and translation‑based workflows. Builds and analyzes cross‑lingual toxicity detectors and mitigation methods, cross‑lingual adversarial and jailbreak tests and language‑aware obfuscation benchmarks, and quantifies metrics such as attack success rates, detector transferability, heterogeneous safety drift, and region‑specific vulnerabilities.
研究评估了现有AI安全基准对小型语言模型的有效性和可靠性,发现这些基准在评估小型语言模型时存在局限性,特别是在处理复杂提示和模型架构时。
This study addresses the pronounced vulnerability of large language models in non-English and low-resource languages, where harmful prompts rejected in English are frequently executed incorrectly in other languages. The authors construct a multilingual jailbreaking benchmark spanning 18 languages across four resource tiers and incorporating four perturbation types—including script variation, code-switching, and translationese—and combine geometric mechanism analysis with subspace projection to systematically investigate how such perturbations affect safety alignment. Their findings reveal a sharp performance cliff between resource tiers 2 and 3 across all models and identify geometrically misaligned subspaces in low-resource languages that evade safety rejection mechanisms. These results demonstrate that reliance on English-only safety evaluations is insufficient and underscore the necessity of language-specific alignment strategies tailored to script families and perturbation categories.
Existing evaluation benchmarks struggle to comprehensively assess safety risks of vision-language models in multilingual and multimodal settings and lack semantically aligned harmful image-text pairs. This work proposes the first safety evaluation framework that disentangles language and modality dimensions, introducing a benchmark dataset comprising 100,440 semantically aligned harmful image-text pairs across 10 languages. The framework distinguishes between image-dominant and text-dominant risk types and employs red-teaming attacks alongside human and automated evaluations to systematically test 11 open-source vision-language models. The study reveals that high-resource languages are more susceptible to image-dominant attacks, whereas low-resource languages exhibit greater vulnerability under text-dominant risks. Although model scaling reduces overall attack success rates, it exacerbates disparities in safety performance across languages.
研究审计了八个网络安全基准在不同语言模型上的表现,揭示了评分依赖于评估流程配置的问题,并提出标准化评估流程以提高模型评价可靠性。
The capability of existing large language models (LLMs) in detecting Chinese illegal content—particularly politically sensitive material, pornography, and phonetic/orthographic variants (e.g., homophone-based evasion)—remains poorly characterized. Method: We introduce ChineseSafe, the first LLM safety benchmark tailored to Chinese regulatory requirements, comprising 205K samples across four high-level risk categories and ten fine-grained subcategories. It features a rule-enhanced multi-stage annotation framework, human-expert validation for data curation, and a zero-/few-shot evaluation paradigm supporting both open-source models and commercial APIs. Crucially, it enables fine-grained safety classification and constructs localized adversarial examples (e.g., homophone perturbations) for the first time. Contribution/Results: Experiments expose critical vulnerabilities in mainstream LLMs—especially in identifying political sensitivities and phonetic evasion—some exceeding legally acceptable risk thresholds. ChineseSafe is publicly released and officially endorsed by Hugging Face.
This study addresses coverage gaps and quality deficiencies in monolingual dimensions of multilingual AI safety benchmarks by proposing a reusable slice-level auditing methodology. Through cross-tier empirical comparisons and controlled experiments across 21 resources, we systematically reveal structural deficits in non-English data regarding provenance, annotation, and harm taxonomy. The research demonstrates that these disparities are quantifiably remediable and establishes a causal link between data-sparse regions and asymmetric model robustness against jailbreaking. By filling critical coverage voids in categories such as self-harm, this work renders multilingual safety claims empirically verifiable and provides methodological foundations for constructing high-quality multilingual safety datasets.
This study addresses the limitations of existing safety evaluation benchmarks for large language models, which are predominantly English-centric and insufficient for assessing cultural sensitivity and localized harms. The authors construct a cross-cultural safety benchmark encompassing 10 country–language pairs and 5,500 test cases, introducing two novel metrics—Neutral-Safe Rate and Cultural Sensitivity Rate—to distinguish between universal harms and culturally embedded sensitive content. Through a multi-stage construction pipeline involving model-assisted discovery, automated validation, and dual-native annotation, along with a unified evaluation framework, they assess 10 frontier models and 27 localized models. The evaluation reveals a decoupling between jailbreak robustness and cultural awareness in frontier models and demonstrates that the apparent safety of many localized models often stems from generation failures rather than genuine alignment.
This work addresses the inconsistent detection and mitigation of toxic content by multilingual large language models across diverse linguistic and cultural contexts. It presents the first systematic synthesis of research on multilingual toxicity handling, proposing a comprehensive framework that encompasses threat modeling, task formulation, detection strategies—such as cross-lingual encoders, translation pipelines, and representation probing—and mitigation approaches, including data filtering, alignment tuning, decoding controls, and multilingual safeguards. The study identifies core challenges such as uneven language coverage and culturally contingent definitions of harm, while highlighting critical issues like fragmented evaluation protocols and the unintended suppression of legitimate expression. By elucidating these dimensions, the paper establishes a theoretical foundation and practical roadmap for achieving cross-lingual safety alignment in multilingual language models.
This study addresses the limitations of current safety evaluations for multilingual large language models, which often rely on literal translations of English benchmarks and overlook cultural contextual differences, leading to inaccurate risk assessments. The authors present the first systematic construction of paired red-teaming datasets in Korean, Japanese, Thai, and Khmer, featuring both direct translation (DT) and culturally adapted (CA) prompts matched one-to-one by seed. Evaluation using attack success rate (ASR) and a novel Cultural Contextualization Score (C3) reveals that CA prompts increase ASR by an average of 9.3 percentage points across all 16 language–model combinations. Furthermore, DT significantly underestimates local threats in 44 out of 48 risk categories, while C3 scores rise from a mean of 0.17 to as high as 2.51, demonstrating that cultural adaptation is essential for accurately capturing localized safety risks.
This work addresses the vulnerability of instruction-tuned large language models to task-level backdoor poisoning attacks when trained on unverified data, proposing PoisonForge—a systematic benchmark for evaluating such threats. By injecting only a minimal number (e.g., 10 samples, or 1% of the training set) of carefully crafted instruction-response pairs, the attack reliably induces the model to output attacker-specified content on a targeted task while preserving near-original performance on others (leakage rate < 0.5%). The study introduces the first four-dimensional parameterization of task-level poisoning threats, demonstrating that attack success hinges primarily on poison design rather than model scale, and develops a generalizable risk prediction model. Evaluated across 12 mainstream open-source models, the approach achieves over 70% attack success rates in 11 models under the most vulnerable configurations.