actr: aligning thoughts and responses for multilingual safety in reasoning llms

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the thought-action misalignment in reasoning large language models under multilingual jailbreak attacks, wherein models identify risks yet generate unsafe responses. To align cross-lingual chains-of-thought with safe outputs, we propose the ACTR framework. Methodologically, we introduce a novel Thought Gap Score (TGS) metric to precisely localize safety-reasoning neurons, alongside Neuron Selective Consistency Optimization (NSCO), which requires no manual annotation, and a frozen judge-model reward mechanism. Experimental results demonstrate that ACTR significantly reduces attack success rates across multiple languages and generalizes effectively to unseen languages. Furthermore, the proposed approach preserves knowledge reasoning capabilities while mitigating over-refusal rates.
📝 Abstract
Ensuring the safety of reasoning large language models (LLMs) across languages is essential for their reliable deployment. However, when exposed to jailbreak attacks in non-high-resource languages, these models may generate unsafe responses even when their reasoning traces identify safety risks. To address this issue, we propose aligning cross-lingual thoughts and responses (ACTR), a framework that improves multilingual safety alignment by strengthening the use of existing safety reasoning. Specifically, we first present the think gap score (TGS) to compare the normalized contributions of reasoning traces to attention outputs during response generation across languages, and use reasoning- trace substitution to measure the cross-lingual safety gap. Next, using a corpus of jailbreak queries, we assess neuron importance through changes in response representations caused by neuron masking and compare the high-importance neuron sets obtained with reasoning enabled and disabled to identify safety think neurons that support the use of safety reasoning. Finally, we devise neuron-selective consistency optimization (NSCO), which uses a frozen judge model to reward agreement between the safety categories of reasoning traces and responses while updating only the parameters associated with the selected neurons, requiring no human-annotated responses or preference data. Across two reasoning models, ACTR achieves lower average attack success rates than the evaluated state-of-the-art methods on AdvBench-X and MultiJail, with safety gains extending to unseen languages, while preserving or improving average performance on multilingual knowledge and mathematical reasoning tasks and limiting false refusals of benign requests. Warning: this paper contains examples with unsafe content.
Problem

Research questions and friction points this paper is trying to address.

reasoning LLMs
multilingual safety
jailbreak attacks
cross-lingual alignment
safety reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multilingual Safety Alignment
Think Gap Score
Safety Think Neurons
Neuron-Selective Consistency Optimization
Reasoning LLMs
🔎 Similar Papers
No similar papers found.