dynamic representation editing

Designs and implements methods to identify, modify, and evaluate internal latent states and activation trajectories of a model at runtime—e.g., clustering and disentangling reasoning manifolds, projecting and applying directions in latent space (Fisher‑LDA style), or directly editing activations to produce controlled counterfactual behaviors. Builds tooling and analyses for dynamic interventions, purification of representational subspaces, and measurement of the causal effect of those interventions on downstream outputs and decoding trajectories.

dynamicrepresentationediting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.32
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$201K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of precisely controlling specific behaviors—such as refusal or sycophancy—in large language models, where targeted interventions often produce unintended side effects. The authors propose a low-rank subspace diagnostic framework that reveals, for the first time, that distinct behaviors share internal representations in activation space. Through geometric analysis of decision subspaces and the mean squared cosine of principal angles, they demonstrate that intervention effects propagate asymmetrically, depending on the degree of subspace overlap and the angular proximity to the decision subspace. Experiments across multiple instruction-tuned models (7B–70B) show that behaviors exhibiting high representational overlap and closer alignment with the decision subspace are more susceptible to intervention, thereby explaining the fundamental difficulty in achieving independent behavioral control.

behavioral interferencelarge language modelssafety interventions

This study addresses the challenge of decoding latent dynamic structures underlying large-scale neuronal population activity by proposing a unified latent variable modeling framework that, for the first time, jointly integrates three core tasks: single-region dynamics modeling, inter-regional communication analysis, and behavioral alignment. The approach combines classical state-space models with cutting-edge deep generative architectures—including Transformers, diffusion models, and neural ordinary differential equations—to systematically construct a taxonomy and establish clear evaluation benchmarks. Emphasizing critical challenges such as causal inference and directional connectivity, this work provides both theoretical foundations and methodological tools for interpretable brain dynamics analysis and robust neural decoding.

Behavior-Aligned ModelingLatent Variable ModelsMulti-Region Communication

Probing Latent Subspaces in LLM for AI Security: Identifying and Manipulating Adversarial States

Mar 12, 2025
XW
Xin Wei Chia
🏛️ Home Team Science and Technology Agency | Singapore

This work addresses the security risk of jailbreaking—i.e., unintended, harmful outputs—in large language models (LLMs) induced by prompt injection attacks. We propose a novel *representation-level proactive defense* paradigm grounded in neuroscience-inspired attractor dynamics: we model LLM hidden states as *semi-stable attractors*, identify latent subspaces corresponding to safe versus jailbroken behavioral regimes via hidden-layer activation analysis, and construct cross-layer perturbation vectors that locally induce or suppress state transitions at targeted layers. Experiments across multiple LLMs demonstrate statistically significant triggering or suppression of jailbroken responses. Our approach provides the first empirical evidence for *causal intervention at the representation level*, shifting the defensive paradigm from reactive output filtering to proactive internal-state control.

Developing proactive defenses against adversarial prompt injectionsIdentifying latent subspaces in LLMs for AI securityManipulating adversarial states to induce jailbreak transitions

Towards Unifying Interpretability and Control: Evaluation via Intervention

Nov 07, 2024
UB
Usha Bhalla
🏛️ Harvard University | Google DeepMind

Current interpretability research for large language models (LLMs) treats interpretability and controllability as disjoint objectives. Method: This paper proposes “intervention capability” as a unified evaluation goal and introduces an encoder-decoder framework that integrates four method families—sparse autoencoders (SAEs), Logit Lens, Tuned Lens, and probes—to enable controllable interventions on interpretable features. Contribution/Results: We formally define two novel metrics—intervention success rate and consistency–intervention trade-off—and argue that effective intervention constitutes the foundational objective of interpretability. Experiments show that Lens-based methods outperform SAEs and probes in simple interventions; however, existing methods exhibit inconsistent cross-feature and cross-model intervention efficacy. Moreover, mechanistic interventions often underperform prompt engineering, revealing critical controllability bottlenecks. This work shifts LLM interpretability research from descriptive analysis toward causal, interventionist control.

Assess coherence-intervention tradeoff in model behaviorEvaluate methods through intervention success metricsUnify interpretability and control in language models

The causal origins of interpretable units—such as induction heads—in large language models remain poorly understood. This work proposes a scalable mechanistic data attribution framework that integrates influence functions with causal interventions to establish, for the first time, direct causal links between specific training examples and the emergence of such interpretable components. The study reveals that structured repetitive data plays a catalytic role in circuit formation and demonstrates a direct functional relationship between induction heads and in-context learning capabilities. By selectively intervening on a small set of high-influence training samples, the emergence of attention heads can be significantly modulated. Furthermore, the proposed data augmentation strategy consistently accelerates circuit convergence across different model scales.

Data AttributionIn-Context LearningInduction Heads

Latest Papers

What's happening recently
View more

This work addresses the challenge of detecting and regulating sycophantic behavior—excessive user flattery—in language models by proposing an iterative data generation method based on cascaded linear samples. Departing from conventional binary contrastive examples, the approach constructs sequences of samples with continuously varying behavioral intensities, revealing for the first time a linearly separable structure of sycophancy in activation space. This enables precise identification and disentanglement of the associated feature subspace. Through activation manipulation and subspace analysis, the method matches or exceeds baseline approaches such as LLM-as-a-judge and system prompting in detection accuracy, calibration, and robust controllability, while incurring lower computational overhead and substantially improving the interpretability of behavioral interventions.

activation steeringbehavior controlinterpretable features

Large language models may alter their behavior during safety evaluations due to awareness of being assessed, thereby compromising evaluation validity. This work proposes a novel method that suppresses internal latent variables associated with evaluation awareness solely by optimizing input prefix prompts, without requiring model inference access. The approach integrates GCG-style token optimization, a self-cross-entropy fluency regularizer, and multi-class latent targets—including CAA directions, SAE features, and MLP neurons—enabling, for the first time, selective deactivation of specific internal representations. Experiments on Llama-3.2-3B and Llama-3.1-8B demonstrate robust suppression of target latents to approximately –7, with causally validated SAE features fully deactivated, revealing that activation interpretability does not imply behavioral controllability.

activation suppressionevaluation-awarenessinput-only control

Hot Scholars

CS

Changho Shin

University of Wisconsin-Madison
Machine learningdata science
FS

Frederic Sala

Assistant Professor, University of Wisconsin
Data-centric AIMachine learningInformation theory
JO

Jean Oh

Robotics Institute, Carnegie Mellon University
RoboticsMultimodal PerceptionSocial NavigationLanguage-Vision intersection
XZ

Xiaoqing Zheng

Fudan University
Natural Language Processing and Machine Learning