design activation strategies

Designs, implements, and evaluates methods for producing, modifying, or exploiting internal activations of neural models to elicit, amplify, or suppress specific behaviors; this includes crafting input triggers or prompts, training- or inference-time interventions, probes, and measurement pipelines to control and analyze how layer- and neuron-level activations map to model outputs.

designactivationstrategies

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.15
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$178K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Neural Network Reprogrammability: A Unified Theme on Model Reprogramming, Prompt Tuning, and Prompt Instruction

Jun 05, 2025
ZY
Zesheng Ye
🏛️ University of Melbourne | Southeast University | IBM Research

Efficient, lightweight downstream adaptation of large language models remains challenging due to fragmented methodologies and lack of unifying principles. Method: This paper proposes a unified framework—“neural reprogrammability”—modeling parameter-free adaptation paradigms—including model reprogramming, prompt tuning, and prompt instruction—as targeted manipulations of information flow at interfaces such as input, intermediate layers, or context. Contribution/Results: We introduce the first cross-modal, architecture-agnostic four-dimensional taxonomy (format, location, operator, output alignment), revealing intrinsic unity among in-context learning, chain-of-thought, and related methods. By systematically integrating existing interface perturbation techniques—including input perturbation, token insertion, and example injection—we empirically validate their generality across multimodal foundation models. Our framework establishes foundational principles and provides actionable guidelines for lightweight, controllable, and interpretable model adaptation.

Categorizing adaptation approaches across key dimensions systematicallyExploring neural network sensitivity to interface manipulationsUnifying model adaptation techniques for pre-trained foundation models

Existing neural network probing methods often rely on input perturbations or parameter analysis, which struggle to uncover structured information embedded in intermediate representations. This work proposes APEX, a novel probing paradigm that perturbs hidden activations during inference while keeping both inputs and model parameters fixed. APEX formalizes activation perturbation as a general probing framework, unifying and extending prior approaches such as input perturbation as special cases, and enabling a controllable transition from sample-dependent to model-dependent behavioral analysis. Experiments demonstrate that APEX effectively quantifies representational structure, distinguishes models trained on structured versus random labels, reveals semantically coherent prediction transitions, and precisely identifies the concentration of predictions toward target classes in backdoor attacks.

activation perturbationintermediate representationsneural networks

Existing interpretability methods struggle to distinguish whether model components genuinely encode a target capability or merely propagate upstream signals. This work proposes Weight Patching, a source-directed intervention in weight space that operates on isomorphic models exhibiting varying behavioral strengths. By substituting specific module weights and anchoring behavioral interfaces via vector alignment, the method precisely localizes source-level mechanisms within large language models. The framework enables, for the first time, tracing the pathway of capability transmission from shallow source carriers to downstream execution circuits, thereby supporting mechanism-aware model merging. Experiments on instruction-following tasks successfully identify critical mechanistic components, significantly improving selective fusion of expert models, with findings further validated externally.

behavioral capabilityLLMsmechanistic interpretability

This work proposes a real-time detection method for reward hacking in large language models during text generation—a subtle failure mode that often remains undetectable in final outputs. By leveraging sparse autoencoders to extract internal representations from residual stream activations and combining them with a lightweight linear classifier, the approach enables token-by-token identification of reward-hacking behaviors. This method dynamically reveals the early emergence, continuous evolution, and dependence on reasoning strategies of such behaviors throughout the generation process. Notably, it generalizes across model families and fine-tuning strategies, providing actionable alerts before problematic outputs are produced. The framework thus establishes a novel paradigm for post-deployment alignment monitoring, offering a proactive safeguard against hidden misalignment in deployed models.

emergent misalignmentinternal activationslarge language models

The causal origins of interpretable units—such as induction heads—in large language models remain poorly understood. This work proposes a scalable mechanistic data attribution framework that integrates influence functions with causal interventions to establish, for the first time, direct causal links between specific training examples and the emergence of such interpretable components. The study reveals that structured repetitive data plays a catalytic role in circuit formation and demonstrates a direct functional relationship between induction heads and in-context learning capabilities. By selectively intervening on a small set of high-influence training samples, the emergence of attention heads can be significantly modulated. Furthermore, the proposed data augmentation strategy consistently accelerates circuit convergence across different model scales.

Data AttributionIn-Context LearningInduction Heads

Latest Papers

What's happening recently
View more

This study addresses the latent degradation and evaluation blind spots induced by internal activation steering in tool-calling scenarios of large language models (LLMs). We propose SAKIKO, an auditing framework that systematically evaluates the genuine remediation effects of internal interventions through directional error discovery, channel-keyed intervention, and target-parsing verification. Furthermore, this work pioneers an outcome-parsing-based adjudication mechanism and forward-freezing statistical licensing, revealing that behavioral modification does not equate to fundamental repair. Experiments across seven LLMs demonstrate that most existing interventions incur severe side effects, thereby establishing the necessity of outcome-level adjudication for ensuring intervention rigor.

activation steeringinternal interventionsmechanistic auditing

This study investigates the internal decision-making mechanisms by which LLM agents choose between tool invocation and direct response. Methodologically, it proposes transforming complex prompts into single-variable contrastive pairs to construct minimal examples, thereby identifying a causal vector μΔ. This analysis is further supported by mechanistic interpretability, Transcoder decomposition, attention head tracing, and cross-model ablation experiments. The findings reveal that analytical verbs suppress features that interfere with tool-calling priors, and that μΔ exhibits both causal necessity and sufficiency. Notably, this mechanism is consistently observed across seven model families, including Qwen, and generalizes effectively to native multi-turn dialogue scenarios.

Agentic LLMsDecision mechanismMechanistic interpretability

This study addresses the challenge in large language model (LLM) agents of distinguishing instructions from data and localizing intervention points under indirect prompt injection. By employing counterfactual role probes, component-level activation patching, and trajectory-independent interventions via AgentDojo, this work leverages causal tracing to reveal a fundamental divergence between the readability of role signals and the intervenability of agent behavior. It is the first to explicitly distinguish readable role signals from effective behavioral interventions, demonstrating that cross-channel transfer of intervention directions is inherently difficult. Furthermore, the study systematically analyzes how network depth and positional factors influence attack success. Results indicate that single-point editing fails in deeper layers, whereas wide-span and repeated editing strategies significantly reduce attack success rates, offering actionable insights for securing LLM agents against indirect prompt injection.

Behavioral InterventionCausal TracingIndirect Prompt Injection

Hot Scholars

MG

Mor Geva

Tel Aviv University, Google Research
Natural Language Processing
MP

Miao Pan

Professor, Electrical and Computer Engineering, University of Houston
Wireless for AICybersecurity for AIMobile/Edge AI SystemsUnderwater IoT Nets
SH

Shaoyi Huang

Assistant Professor, Stevens Institute of Technology
Deep LearningEfficient AISoftware hardware co-design
EA

Eric A. F. Reinhardt

Graduate Student, The University of Alabama
Particle PhysicsMachine LearningPhysics
SG

Sergei Gleyzer

University of Alabama
Particle PhysicsMachine Learning