Activation Flow: Manufacturing Activations for Steering

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the failure of traditional mean-difference steering methods when models deliberately conceal their capabilities, thereby withholding authentic activations. To overcome this limitation, this work proposes Activation Flow, a framework that introduces the first ordinary differential equation (ODE)-based activation synthesis technique to eliminate reliance on genuine activations. By integrating residual stream single-vector injection, Jacobian singular direction truncation, and ODE solving, the method generates targeted activation vectors from minimal labeled data to steer model behavior without fine-tuning. On capability-locked models, the proposed approach significantly improves task accuracy from 0.05 to 0.85, achieving performance comparable to full fine-tuning while comprehensively outperforming existing baselines.
📝 Abstract
Difference-in-means steering requires activations recorded while a model shows the desired behavior, which a sandbagging model withholds by deliberately underperforming. We introduce Activation Flow (ActFlow), which manufactures these activations from $k$ correct labels without fine-tuning. ActFlow sets target logits that rank each labeled item's correct answer first, and moves the logits toward them by adding one vector $x$ to all $k$ residual streams at one layer. ActFlow is a family of ordinary differential equations for $x$, one for each rule that maps the required logit change to the velocity of $x$. The smallest-norm rule lands exactly on the targets, while the others keep only the top singular directions of the Jacobian. We test ActFlow on three instruction-tuned models, each locked by a sandbagging prompt and by a password-locked LoRA. At $k=40$, ActFlow keeping five singular directions raises the mean held-out ARC-Easy accuracy over the six locked models from $0.05$ to $0.85$, against $0.88$ for fine-tuning and $0.92$ for the honest models. Furthermore, it scores higher than the smallest-norm rule in 16 of the 18 combinations of locked model and $k$, and its steering direction is nearly orthogonal to the honest difference-in-means direction. It also unlocks two LoRA locks where the honest direction fails.
Problem

Research questions and friction points this paper is trying to address.

sandbagging
activation steering
capability elicitation
locked models
difference-in-means
Innovation

Methods, ideas, or system contributions that make the work stand out.

Activation Flow
Sandbagging
Activation Steering
Ordinary Differential Equations
Jacobian Singular Directions
🔎 Similar Papers
No similar papers found.