Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models

πŸ“… 2026-09-22
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
ι€šθΏ‡APIθ―±ε―Όε‰ζ²Ώζ¨‘εž‹ε€–ιƒ¨εŒ–δΈ­ι—΄ζŽ¨η†θΏ‡η¨‹οΌŒιͺŒθ―ε…ΆζŽ¨η†θƒ½εŠ›οΌŒεΉΆεˆ†ζžδΈεŒζ¨‘εž‹ηš„ζŽ¨η†η»“ζž„ε’Œζ•ˆηŽ‡γ€‚
πŸ“ Abstract
The rapid capability gains of frontier language models are widely attributed to improved reasoning abilities, yet this cannot be verified as raw CoT traces in closed-source systems are hidden. By registering a simple custom tool through a standard API feature, we induce frontier models to externalize intermediate reasoning. Because these traces may reflect post-hoc rationalization rather than genuine reasoning, we first evaluate against native CoT on open-source models and extend to closed-source frontier models including GPT-6 Astra. We find that the extracted reasoning matches native reasoning performance and substantially outperforms no-reasoning baselines, across competition mathematics, science, and code generation. We then characterize how frontier models structure their intermediate reasoning. Across token efficiency, reasoning-step types, and induced reasoning trees, we identify systematic differences in how models externalize, compress, and organize reasoning. We find that Astra exhibits token-efficient directed reasoning, selecting a correct trajectory earlier, while resolving elementary steps internally and externalizing only crucial reasoning. These findings provide a behavioral lens on frontier-model reasoning beyond benchmark scores.
Problem

Research questions and friction points this paper is trying to address.

frontier language models
reasoning abilities
chain-of-thought
closed-source systems
Innovation

Methods, ideas, or system contributions that make the work stand out.

Custom Tool Induction
Externalized Reasoning
Token Efficiency
Intermediate Reasoning Structure
GPT-6 Astra
πŸ”Ž Similar Papers
No similar papers found.
X
Xiaoyu Luo
Department of Computer Science, Aalborg University, Copenhagen, Denmark
Tao Ren
Tao Ren
Peking University
Foundation modelOptimizationReinforcement Learning
W
Wenrui Yu
Department of Electronic Systems, Aalborg University, Copenhagen, Denmark
X
Xiao Li
Seafill
Q
Qiongxiu Li
Department of Electronic Systems, Aalborg University, Copenhagen, Denmark; Seafill
Johannes Bjerva
Johannes Bjerva
Full Professor, Department of Computer Science, Aalborg University
Natural Language ProcessingNLP SecurityLLM SecurityComputational LinguisticsTypology