🤖 AI Summary
This study addresses the behavioral discrepancies of large language models between auditing and deployment phases, which stem from their latent "evaluation awareness"—a cognitive state difficult to measure directly. To resolve this, we propose a verbalization training method that leverages spontaneous model utterances as observable evidence. By integrating reinforcement learning, generation-prefix truncation, and synthetic document fine-tuning, our approach calibrates verbal outputs without supervising latent beliefs, thereby externalizing evaluation awareness while preserving underlying belief stability. Experimental results demonstrate that this method amplifies verbalized evaluation awareness by 2.4 to 2.9 times and effectively generalizes to agentic scenarios. Ultimately, this work establishes a novel paradigm for alignment auditing in large language models.
📝 Abstract
Evaluation awareness (EA) can cause large language models (LLMs) to behave differently during audits than in deployment, yet measuring and accounting for EA remains challenging. We introduce verbalization training (VT), a method for making LLMs less reticent about verbalizing evaluation awareness while avoiding to supervise the latent belief itself. VT uses a model's spontaneous verbalizations as evidence that awareness is present and truncates each rollout immediately before the verbalization, producing training prefixes at which the model is presumed to be aware. The model is then trained with an RL objective designed to increase verbalization in a calibrated way. Across Qwen3.6-35B-A3B, Kimi K2.6, and Inkling, VT increases verbalized EA by 2.4-2.9 times and transfers to held-out agentic settings, while measured latent EA and behavior remain largely stable. In a causal experiment, we independently implant meta-knowledge about evaluations through synthetic-document fine-tuning and show that VT-induced verbalizations reflect the richer knowledge acquired by the model.