Jailbreaking the Matrix: Nullspace Steering for Controlled Model Subversion

📅 2026-04-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Despite alignment and instruction tuning, large language models remain vulnerable to jailbreaking attacks. This work proposes Head-Masked Nullspace Steering (HMNS), a novel approach that uniquely integrates geometric awareness with interpretability. By leveraging causal analysis to identify critical attention heads, HMNS masks their write paths and injects nullspace-constrained perturbations into the orthogonal complement of the suppressed subspace. The method further incorporates residual norm scaling and iterative re-identification to establish a closed-loop detection-and-intervention mechanism. Evaluated across multiple mainstream models and jailbreaking benchmarks, HMNS achieves state-of-the-art attack success rates while demonstrating significantly higher query efficiency compared to existing techniques.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Safety and RobustnessComputer Vision: Adversarial Attacks & Robustness

Application Category

User Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systemsSearch and Retrieval-Augmented AI: Large language models for searchSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Large language models remain vulnerable to jailbreak attacks -- inputs designed to bypass safety mechanisms and elicit harmful responses -- despite advances in alignment and instruction tuning. We propose Head-Masked Nullspace Steering (HMNS), a circuit-level intervention that (i) identifies attention heads most causally responsible for a model's default behavior, (ii) suppresses their write paths via targeted column masking, and (iii) injects a perturbation constrained to the orthogonal complement of the muted subspace. HMNS operates in a closed-loop detection-intervention cycle, re-identifying causal heads and reapplying interventions across multiple decoding attempts. Across multiple jailbreak benchmarks, strong safety defenses, and widely used language models, HMNS attains state-of-the-art attack success rates with fewer queries than prior methods. Ablations confirm that nullspace-constrained injection, residual norm scaling, and iterative re-identification are key to its effectiveness. To our knowledge, this is the first jailbreak method to leverage geometry-aware, interpretability-informed interventions, highlighting a new paradigm for controlled model steering and adversarial safety circumvention.
Problem

Research questions and friction points this paper is trying to address.

jailbreak attacks
large language models
safety mechanisms
adversarial safety circumvention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Nullspace Steering
Jailbreak Attack
Attention Head Intervention
Closed-loop Detection
Geometry-aware Perturbation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
V
Vishal Pramanik
Department of Computer & Information Science & Engineering, University of Florida, Gainesville, FL 32611, USA
M
Maisha Maliha
School of Computer Science, University of Oklahoma, Norman, OK 73019, USA
Susmit Jha
Susmit Jha
Director, Neurosymbolic Computing and Intelligence, SRI International
Aritificial IntelligenceAutonomyFormal MethodsMachine Learning
Sumit Kumar Jha
Sumit Kumar Jha
University of Florida