🤖 AI Summary
This work investigates the temporal mechanism underlying answer generation in multi-step arithmetic reasoning by large language models (LLMs): specifically, whether answers are formed prior to chain-of-thought (CoT) activation (“think-to-talk”) or incrementally constructed during CoT execution (“talk-to-think”).
Method: We design controlled arithmetic tasks and employ causal probing combined with latent state intervention to isolate and perturb reasoning dynamics across model layers and timesteps.
Contribution/Results: Our analysis reveals, for the first time, a consistent cross-model hierarchical timing pattern: single-step subproblems are resolved before CoT initiation, whereas multi-step composite computations dynamically depend on the unfolding CoT process. This finding challenges the oversimplified assumption that CoT merely verbalizes precomputed answers, establishing instead that CoT serves a dual function—performing internal computation *and* externalizing reasoning steps. The results provide critical empirical evidence for understanding the computational architecture of LLM reasoning.
📝 Abstract
This study investigates the internal reasoning process of language models during arithmetic multi-step reasoning, motivated by the question of when they internally form their answers during reasoning. Particularly, we inspect whether the answer is determined before or after chain-of-thought (CoT) begins to determine whether models follow a post-hoc Think-to-Talk mode or a step-by-step Talk-to-Think mode of explanation. Through causal probing experiments in controlled arithmetic reasoning tasks, we found systematic internal reasoning patterns across models in our case study; for example, single-step subproblems are solved before CoT begins, and more complicated multi-step calculations are performed during CoT.