Probing Latent Subspaces in LLM for AI Security: Identifying and Manipulating Adversarial States

📅 2025-03-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the security risk of jailbreaking—i.e., unintended, harmful outputs—in large language models (LLMs) induced by prompt injection attacks. We propose a novel *representation-level proactive defense* paradigm grounded in neuroscience-inspired attractor dynamics: we model LLM hidden states as *semi-stable attractors*, identify latent subspaces corresponding to safe versus jailbroken behavioral regimes via hidden-layer activation analysis, and construct cross-layer perturbation vectors that locally induce or suppress state transitions at targeted layers. Experiments across multiple LLMs demonstrate statistically significant triggering or suppression of jailbroken responses. Our approach provides the first empirical evidence for *causal intervention at the representation level*, shifting the defensive paradigm from reactive output filtering to proactive internal-state control.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Safety and RobustnessComputer Vision: Large Vision Models

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks, yet they remain vulnerable to adversarial manipulations such as jailbreaking via prompt injection attacks. These attacks bypass safety mechanisms to generate restricted or harmful content. In this study, we investigated the underlying latent subspaces of safe and jailbroken states by extracting hidden activations from a LLM. Inspired by attractor dynamics in neuroscience, we hypothesized that LLM activations settle into semi stable states that can be identified and perturbed to induce state transitions. Using dimensionality reduction techniques, we projected activations from safe and jailbroken responses to reveal latent subspaces in lower dimensional spaces. We then derived a perturbation vector that when applied to safe representations, shifted the model towards a jailbreak state. Our results demonstrate that this causal intervention results in statistically significant jailbreak responses in a subset of prompts. Next, we probed how these perturbations propagate through the model's layers, testing whether the induced state change remains localized or cascades throughout the network. Our findings indicate that targeted perturbations induced distinct shifts in activations and model responses. Our approach paves the way for potential proactive defenses, shifting from traditional guardrail based methods to preemptive, model agnostic techniques that neutralize adversarial states at the representation level.
Problem

Research questions and friction points this paper is trying to address.

Identifying latent subspaces in LLMs for AI security
Manipulating adversarial states to induce jailbreak transitions
Developing proactive defenses against adversarial prompt injections
Innovation

Methods, ideas, or system contributions that make the work stand out.

Extracted hidden activations to identify latent subspaces.
Used dimensionality reduction to reveal lower dimensional spaces.
Derived perturbation vector to shift model states.
💼 Related Jobs
No related jobs found.
X
Xin Wei Chia
Home Team Science and Technology Agency, Singapore
J
Jonathan Pan
Home Team Science and Technology Agency, Singapore