COMPASS: Finding Where Reasoning Lives in Language Models

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing reasoning methods largely rely on predefined features. This work proposes COMPASS, a method that locates and steers internal reasoning mechanisms in language models using only answer correctness signals. Specifically, it constructs latent directions from correctness feedback, identifies critical attention heads via logit-space attribution, and applies activation interventions during inference, eliminating the need for complex prompting or sparse autoencoders. COMPASS outperforms baselines across multiple mathematical benchmarks, achieving an average accuracy improvement of 16% on GSM8K while reducing generated tokens by 20–70%, yielding performance comparable to chain-of-thought reasoning. Furthermore, the learned interventions demonstrate strong cross-benchmark transferability.
📝 Abstract
Explicitly eliciting reasoning substantially improves LLM performance. Existing approaches require a predefined characterization of reasoning, whether through CoT prompt design, contrastive CoT directions, or via SAE derived reasoning features. For mathematical reasoning with verifiable answers, we show that a much simpler signal suffices, which is the correctness of the model's own direct answer attempts. This signal yields a latent direction that elicits reasoning. This direction is decodable within the activations of most attention heads, but only a small subset of them can be effectively intervened. We introduce COMPASS, an inference-time steering method that identifies these heads using a logit-space attribution score and steers their activations along the correctness direction, requiring only per-head activation statistics. Across three model families and multiple math benchmarks, COMPASS outperforms the activation-steering baselines we compare against, improves GSM8K accuracy by 16 percentage points on average, and approaches CoT accuracy with 20-70\% fewer generated tokens. Interventions transfer without re-fitting to unseen benchmarks, and ablations show that both the correctness direction and the small set of heads carrying it are necessary, with the effect concentrated in remarkably few heads.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Inference-time steering
Activation intervention
Logit-space attribution
Correctness direction
Attention heads
P
Pratyay Dutta
University of California, Riverside, CA 92521. Work performed during a summer internship at Lawrence Livermore National Laboratory
Kowshik Thopalli
Kowshik Thopalli
Ph.D. Student, Arizona State University
computer visionmachine learningdeep learningartificial intelligencedifferential geometry
V
Vivek Narayanaswamy
Lawrence Livermore National Laboratory, Livermore, CA 94550.