How Alignment Routes: Localizing, Scaling, and Controlling Policy Circuits in Language Models

📅 2026-04-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
The internal mechanisms by which aligned language models implement strategic refusal—such as safety filtering—remain poorly understood. This work identifies a sparse routing mechanism across nine models through natural experiments: specific gating attention heads detect sensitive content and activate downstream amplification heads to enhance refusal signals. This architecture reveals, for the first time, a structural separation between intent recognition and policy execution, enabling continuous modulation of refusal strength. The robustness and intervenability of this routing pathway are validated via necessity-sufficiency tests, signal modulation, cipher probes, and cross-model ablations, demonstrating high reproducibility (Jaccard similarity 0.92–1.0). Notably, even when the routing fails under ciphered inputs, deep representations retain detectable signals of harmful content.

Technology Category

Natural Language Processing: Safety and RobustnessPlanning, Routing, and Scheduling: Planning with Language ModelsMachine Learning: Large Multimodal Models (LMMs)

Application Category

Responsible Web: Social and technical mechanisms of refusal for web technologies and applicationsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
We identify a recurring sparse routing mechanism in alignment-trained language models: a gate attention head reads detected content and triggers downstream amplifier heads that boost the signal toward refusal. Using political censorship and safety refusal as natural experiments, we trace this mechanism across 9 models from 6 labs, all validated on corpora of 120 prompt pairs. The gate head passes necessity and sufficiency interchange tests (p < 0.001, permutation null), and core amplifier heads are stable under bootstrap resampling (Jaccard 0.92-1.0). Three same-generation scaling pairs show that routing distributes at scale (ablation up to 17x weaker) while remaining detectable by interchange. By modulating the detection-layer signal, we continuously control policy strength from hard refusal through steering to factual compliance, with routing thresholds that vary by topic. The circuit also reveals a structural separation between intent recognition and policy routing: under cipher encoding, the gate head's routing contribution collapses (78% in Phi-4 at n=120) while the model responds with puzzle-solving rather than refusal. The routing mechanism never fires, even though probe scores at deeper layers indicate the model begins to represent the harmful content. This asymmetry is consistent with different robustness properties of pretraining and post-training: broad semantic understanding versus narrower policy binding that generalizes less well under input transformation.
Problem

Research questions and friction points this paper is trying to address.

alignment
policy circuits
language models
refusal mechanism
routing
Innovation

Methods, ideas, or system contributions that make the work stand out.

sparse routing mechanism
gate attention head
policy circuits
interchange intervention
alignment robustness
🔎 Similar Papers
No similar papers found.