Score
Designs and builds models and algorithms that internalize multi-step chain-of-thought reasoning into a single forward pass, so intermediate reasoning states are represented implicitly and can be restored or reconstructed from the final representation. Analyzes how these single-pass architectures model interactions among sequential reasoning factors, recover or compensate for missing or degraded intermediate steps, and trade off accuracy, interpretability, and multi-step computational overhead.
Explicit chain-of-thought (CoT) reasoning is constrained by the expressive bandwidth of natural language and its reliance on token-level supervision. Method: This work systematically investigates implicit reasoning—multi-step, language-free inference conducted directly within the continuous hidden states of large language models (LLMs). We propose a unified analytical framework and introduce an infinite-depth implicit reasoning mechanism based on masked diffusion models, enabling globally consistent, invertible, and supervision-free inference. Furthermore, we model hierarchical reasoning at the neural level via activation recurrence, latent-state propagation, and trajectory-internalized fine-tuning. Contribution/Results: Our work establishes the first comprehensive survey and knowledge framework for implicit reasoning techniques and releases an open-source GitHub repository to advance the field.
This work investigates the temporal mechanism underlying answer generation in multi-step arithmetic reasoning by large language models (LLMs): specifically, whether answers are formed prior to chain-of-thought (CoT) activation (“think-to-talk”) or incrementally constructed during CoT execution (“talk-to-think”). Method: We design controlled arithmetic tasks and employ causal probing combined with latent state intervention to isolate and perturb reasoning dynamics across model layers and timesteps. Contribution/Results: Our analysis reveals, for the first time, a consistent cross-model hierarchical timing pattern: single-step subproblems are resolved before CoT initiation, whereas multi-step composite computations dynamically depend on the unfolding CoT process. This finding challenges the oversimplified assumption that CoT merely verbalizes precomputed answers, establishing instead that CoT serves a dual function—performing internal computation *and* externalizing reasoning steps. The results provide critical empirical evidence for understanding the computational architecture of LLM reasoning.
This work addresses key limitations of existing long chain-of-thought (Long CoT) methods in mathematical reasoning—namely, excessively lengthy inference sequences, high computational overhead, and context loss due to unidirectional reasoning. To overcome these challenges, the authors propose the Cognitive Loop of Thought (CLoT) framework, which introduces a reversible hierarchical Markov chain to decompose problems into structured subproblems and incorporates a human cognition-inspired, layer-wise backward verification mechanism. Coupled with a dynamic KV cache pruning strategy that eliminates redundant low-level reasoning paths after high-level validation, CLoT effectively transcends the memory and unidirectional constraints of conventional CoT approaches. The method achieves state-of-the-art performance across four mathematical reasoning benchmarks, attaining 99.0% accuracy on the AddSub dataset with GPT-4o-mini—outperforming standard CoT and CoT-SC by 4.1% and 2.9%, respectively.
Current evaluations of large language model reasoning overly rely on final-answer accuracy, making it difficult to diagnose the reasoning process itself. This work proposes a process-oriented evaluation framework centered on adaptive, multi-step search, modeling reasoning as an input-dependent, variable-depth search procedure. The approach emphasizes assessing the faithfulness and effectiveness of intermediate reasoning trajectories rather than just end results. By leveraging intermediate decoding and explicit reasoning traces, the method analyzes model behavior in step selection and termination mechanisms, revealing structural limitations of single-pass forward architectures in achieving variable-depth computation. This shift enables the development of more interpretable and debuggable evaluation standards that capture the dynamics of reasoning beyond static correctness.
Chain-of-thought (CoT) reasoning in mathematical problem solving suffers from excessive token consumption, high KV cache overhead, and low inference efficiency due to long dependency chains. Method: We propose Markovian Chain-of-Thought (MCoT), modeling each reasoning step as a text state augmented with executable Python code; an integrated code interpreter enables automatic verification and dynamic history compression, reducing redundant intermediate steps to equivalent problem representations. MCoT formally recasts multi-step CoT as a Markov process, eliminating reliance on full-history KV caching. We construct the MCoTInstruct dataset—grounded in symbolic reasoning, code execution, and instruction tuning—and adapt it to mainstream LLM inference pipelines. Results: Experiments show MCoT matches baseline accuracy on mathematical reasoning while significantly reducing latency (−38% on average) and KV cache memory usage (−62%), validating a novel paradigm for efficient long-horizon reasoning.
Large language models (LLMs) lack intrinsic self-refinement capability within a single forward pass, limiting their ability to correct reasoning errors without external feedback or parallel generation. Method: We propose Diversified Chain-of-Thought (DCoT) fine-tuning—a supervised fine-tuning paradigm that constructs structured DCoT datasets by integrating diversity-aware sampling, inter-chain quality comparison, and prompt engineering. This enables the model to generate multiple complementary reasoning chains in one forward pass and perform intra-chain self-correction without external signals or post-hoc aggregation. Contribution/Results: DCoT is the first method to achieve CoT self-refinement *within* a single forward pass, departing from conventional multi-chain parallel generation followed by post-processing. Evaluated across 1.3B–70B models, DCoT consistently outperforms standard CoT baselines—especially on numerically intensive tasks with large state spaces. Human evaluation confirms an intra-chain improvement rate of 68.3%.
This work addresses critical reliability limitations of chain-of-thought (CoT) reasoning in large language models for AI safety monitoring, identifying three distinct pathological failure modes: post-hoc rationalization, encoded reasoning, and internalized reasoning. The study presents the first systematic characterization and differentiation of these CoT pathologies and introduces a lightweight, task-agnostic, and computationally efficient diagnostic toolkit capable of real-time monitoring during model training. By leveraging behavior-based diagnostic metrics and purpose-built model organisms, the proposed method accurately identifies and distinguishes among the three pathological patterns. This approach offers a practical, low-cost solution to enhance the monitorability and safety of large language models without requiring extensive architectural modifications or computational overhead.
This work addresses the high computational and memory costs of chain-of-thought (CoT) reasoning in large language models, which existing compression methods exacerbate by discarding fine-grained information and degrading accuracy. To overcome this trade-off, the paper proposes HybridThinker, a novel approach that integrates temporary retention of critical reasoning steps with memory token compression. It employs a hybrid training strategy wherein some intermediate reasoning steps remain visible to subsequent layers while others are masked, compelling the model to learn efficient compression and retrieval of essential information. This method achieves comparable performance to uncompressed baselines across four reasoning benchmarks while significantly reducing resource overhead. On average, HybridThinker outperforms current CoT compression techniques by 5.8 percentage points in accuracy, with similar inference latency.
This study investigates whether chain-of-thought (CoT) reasoning traces faithfully reflect a model’s actual internal decision-making process, thereby questioning their reliability as a supervisory and auditing mechanism. To this end, the authors propose a step-level Detect-Classify-Compare framework, integrating multidimensional validation techniques—including answer-commitment agents, Patchscopes, tuned-lens probes, causal ablation, truncation experiments, and donor contamination tests. Experiments across nine models and seven reasoning benchmarks reveal that, on average, only 61.9% of CoT steps align with the model’s internal computations; in 58% of misaligned cases, models generate redundant “reasoning” after the answer has already been determined—a phenomenon termed “hallucinated continuation.” Notably, stronger CoT performance correlates with lower temporal fidelity. This work provides the first systematic evidence of a fundamental disconnect between CoT traces and genuine reasoning dynamics, challenging the core assumption that CoT serves as a faithful reasoning log.