Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models

📅 2026-07-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Large language models may alter their behavior during safety evaluations due to awareness of being assessed, thereby compromising evaluation validity. This work proposes a novel method that suppresses internal latent variables associated with evaluation awareness solely by optimizing input prefix prompts, without requiring model inference access. The approach integrates GCG-style token optimization, a self-cross-entropy fluency regularizer, and multi-class latent targets—including CAA directions, SAE features, and MLP neurons—enabling, for the first time, selective deactivation of specific internal representations. Experiments on Llama-3.2-3B and Llama-3.1-8B demonstrate robust suppression of target latents to approximately –7, with causally validated SAE features fully deactivated, revealing that activation interpretability does not imply behavioral controllability.
📝 Abstract
Activation steering controls model behavior by editing internal activations at inference time. We study its input-side dual: optimizing a fluent prompt so that a chosen internal latent is driven toward zero, with no inference-time model access. Our target is an "evaluation-awareness" latent-linearly readable and steerable in recent work-whose control would threaten the validity of safety evaluations if models behave differently when they detect being tested. Adapting Fluent Dreaming / EPO with a negated feature term (GCG-style token optimization plus a self-cross-entropy fluency regularizer, swept over a fluency weight), we suppress the latent under five target constructions-a CAA direction, a subspace norm, an SAE feature, a single MLP neuron, and a behavioral logit-on Llama-3.2-3B and Llama-3.1-8B. The latent is robustly suppressible ($z\approx-7$), and a causally-validated Llama Scope SAE feature can be fully and selectively turned off. But our controls tell a cautionary story about the CAA direction: a placebo random direction is suppressed just as hard and shifts behavior just as far, and when we hold a real eval passage in context and optimize only a prefix, suppressing the eval-direction fails to reduce-and slightly increases-the model's behavioral eval judgment. Activation-readability, in short, is not behavioral controllability. We further find that a single MLP neuron is eval-correlated but not causal at both scales, and that scanning the real Pile yields a natural-text baseline competitive with the optimizer for the internal direction. A positive control validates our erasure detector, bounding an erasure-vs-rotation question earlier left open.
Problem

Research questions and friction points this paper is trying to address.

evaluation-awareness
activation suppression
large language models
safety evaluation
input-only control
Innovation

Methods, ideas, or system contributions that make the work stand out.

activation steering
evaluation-awareness
input-only suppression
latent erasure
SAE feature
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Deepanshu Mody
Center for Data Science, New York University
Samarth Agarwal
Samarth Agarwal
Architect, Samsung Research
U
Utkarsh Mittal
Center for Data Science, New York University
D
Dipesh Mahato
Center for Data Science, New York University