Language Shapes Instruction Hierarchy Compliance in Multilingual LLMs

πŸ“… 2026-07-26
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the lack of systematic evaluation of multilingual large language models (LLMs) regarding instruction hierarchy (IH) compliance, particularly their ability to consistently follow high-priority instructions in both cross-lingual and intra-lingual conflict scenarios. The authors introduce XIH-Bench, a multilingual IH evaluation benchmark spanning six languages, four domains, and three IH configurations. Their analysis reveals a previously undocumented β€œlanguage boundary effect,” wherein cross-lingual conflicts paradoxically enhance IH compliance. Furthermore, models exhibit greater difficulty overriding low-priority instructions in their preferred languages, posing potential safety risks. Extensive experiments across multiple mainstream LLMs demonstrate that language choice significantly influences IH adherence, offering critical empirical insights for improving the controllability and safety of multilingual AI systems.
πŸ“ Abstract
Instruction hierarchy (IH) requires models to prioritize instructions by source, ensuring that higher-priority instructions override lower-priority ones. Despite its importance for safe and controllable deployment, existing evaluations have focused almost exclusively on English, leaving it unclear whether IH compliance remains stable in multilingual settings. We introduce XIH-Bench, a benchmark for multilingual IH evaluation with both same-language and cross-language conflicts across six languages, four domains, and three IH settings. Across models, we find two consistent patterns. First, IH compliance exhibits a clear language-dependent asymmetry: a language that strengthens compliance in the higher-priority position can become disruptive in the lower-priority position. Second, cross-language conflicts yield higher compliance than same-language conflicts, a phenomenon we term the Language Boundary Effect. We further show that language specialization can make lower-priority instructions in model-favored languages harder to override, creating multilingual reliability and security risks.
Problem

Research questions and friction points this paper is trying to address.

Instruction Hierarchy
Multilingual LLMs
Language Dependency
Compliance Asymmetry
Cross-language Conflicts
Innovation

Methods, ideas, or system contributions that make the work stand out.

instruction hierarchy
multilingual LLMs
language boundary effect
XIH-Bench
compliance asymmetry
πŸ”Ž Similar Papers
No similar papers found.