Targeted Retrieval, Compact Representations: How CoT Reasoning Improves Long-Context Counting

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the internal mechanisms by which Chain-of-Thought (CoT) prompting enhances long-context counting accuracy in large language models. Leveraging the Needle-in-a-Haystack task, we integrate mechanistic interpretability analysis, causal intervention experiments, and autoregressive training simulations. Our findings reveal that CoT induces target retrieval and compact representation strategies, optimizing attention distributions and representational geometry while enabling efficient counting through internal counter state tracking. We propose a theoretical framework grounded in target retrieval and state tracking, demonstrating substantial performance gains of CoT in large-scale counting scenarios and confirming the existence of internal counter mechanisms. This work offers novel insights into the underlying principles governing CoT reasoning.
📝 Abstract
Large language models (LLMs) have been rapidly improving in long-context tasks, powered by Chain-of-Thought (CoT) reasoning. However, the internal mechanisms underlying this improvement remain unclear. We investigate these mechanisms through a needle-in-a-haystack (NIAH) counting task, where an LLM is asked to count the number of records dispersed in a long text. Across twelve model comparison groups, Thinking (or reasoning) improves counting accuracy over Non-thinking, with pronounced gains at larger counts. This motivates our mechanistic analysis, which identifies two contrasting mechanisms: (i) broad retrieval, where Non-thinking models broadly attend to multiple needles; (ii) targeted retrieval, where Thinking models use enumeration in CoT traces to successively retrieve needles. Targeted retrieval concentrates attention on individual needles and is accompanied by more compact internal representations. Moreover, causal intervention analysis suggests that Thinking models use the CoT trace to maintain and update an internal counter as needles are successively retrieved, even without explicit numbering. In small controlled experiments, both retrieval mechanisms and counter states emerge under standard autoregressive training. Together, our results connect long-context retrieval with representation geometry of counting, supporting a state-tracking account of CoT reasoning.
Problem

Research questions and friction points this paper is trying to address.

Chain-of-Thought reasoning
long-context counting
mechanistic interpretability
needle-in-a-haystack
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Chain-of-Thought Reasoning
Targeted Retrieval
Compact Representations
Causal Intervention
State-tracking
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Liang Twist Shan
Department of Statistics, University of Wisconsin-Madison
Tianyu Hu
Tianyu Hu
Peking University
nlp
H
Hao Yan
Department of Statistics, University of Wisconsin-Madison
Yiqiao Zhong
Yiqiao Zhong
Assistant Professor, University of Wisconsin--Madison
Interpretability of LLMsDeep Learning TheoryMachine LearningStatistics