Score
Designs, builds, or analyzes model architectures, training curricula, supervision signals, and inference procedures that represent and execute multi-step reasoning as continuous, vector-valued latent trajectories instead of explicit token chains. This work covers methods for encoding intermediate thoughts and numeric/algorithmic steps in latent space, reflective or iterative updating, latent-chain training and supervision, multimodal integration, and evaluation or interpretation of the implicit reasoning trajectory during inference.
Explicit chain-of-thought (CoT) reasoning is constrained by the expressive bandwidth of natural language and its reliance on token-level supervision. Method: This work systematically investigates implicit reasoning—multi-step, language-free inference conducted directly within the continuous hidden states of large language models (LLMs). We propose a unified analytical framework and introduce an infinite-depth implicit reasoning mechanism based on masked diffusion models, enabling globally consistent, invertible, and supervision-free inference. Furthermore, we model hierarchical reasoning at the neural level via activation recurrence, latent-state propagation, and trajectory-internalized fine-tuning. Contribution/Results: Our work establishes the first comprehensive survey and knowledge framework for implicit reasoning techniques and releases an open-source GitHub repository to advance the field.
Existing research lacks a mechanistic, system-level analysis of implicit reasoning in large language models (LLMs). This paper introduces the first taxonomy of implicit reasoning centered on *execution paradigms*, departing from conventional representation-based analyses and instead characterizing reasoning strategies computationally. We categorize implicit reasoning into three types: latent optimization, signal-guided control, and layer-wise recurrent execution. Leveraging intermediate-layer activation analysis, dynamic signal modulation, intra-layer recursive modeling, and behavioral interpretability experiments, we integrate structural, behavioral, and representational evidence to uncover how reasoning “silently unfolds” within internal representations. Furthermore, we establish a systematic evaluation framework covering mainstream benchmarks and low-latency inference metrics. All code and resources are publicly released and actively maintained.
Traditional chain-of-thought (CoT) prompting in language models relies on explicit natural-language reasoning, which is computationally inefficient and ill-suited for abstract, non-linguistic inference. Method: This work systematically investigates the latent CoT paradigm—where large language models perform implicit, language-free, and efficient abstract reasoning within their internal latent spaces. We propose a novel four-dimensional unified taxonomy covering token-level strategies, architectural mechanisms, analytical methodologies, and application scenarios. Contribution/Results: Our study establishes the first structured knowledge framework for latent CoT, comprehensively surveys over 100 state-of-the-art works—including advances in training paradigms, model architecture innovations, latent-space interpretability analysis, and multi-task empirical validation—and clarifies technical trajectories, recurring design patterns, and fundamental challenges. To foster community progress, we open-source implementation code and a curated resource repository.
This study investigates whether latent reasoning methods genuinely perform multi-step inference under weak and strong supervision, and whether their internal mechanisms implement structured search. Through comparative experiments, latent space representation analysis, and behavioral diagnostics, the work systematically evaluates the reasoning processes of various models in continuous latent spaces. The findings reveal that existing approaches predominantly rely on shortcut learning rather than genuine implicit reasoning; while the latent space can encode multiple hypotheses, the reasoning process manifests as implicit pruning rather than structured exploration. A key contribution is the identification of a trade-off between supervision strength and reasoning behavior: strong supervision suppresses shortcuts but constrains hypothesis diversity, whereas weak supervision preserves richer representations yet exacerbates shortcut reliance. These results challenge the prevailing assumption that latent reasoning equates to implicit breadth-first search.
This work aims to evaluate the causal role of intermediate steps in implicit chain-of-thought (CoT) reasoning on final answer correctness. To this end, we model implicit CoT as a structural causal model (SCM) in representation space and employ do-intervention analysis to assess the causal necessity of latent reasoning steps, trace influence propagation pathways, and examine answer commitment mechanisms. For the first time, we reveal—through the lens of causal intervention—the stage-wise functionality and non-local routing properties inherent in implicit CoT, proposing an analytical framework that integrates modality conditioning with stability awareness. Our experiments uncover that the budget of latent steps is allocated in a stage-wise rather than uniform manner, and that output bias emerges prior to representational commitment, resulting in a persistent gap. These findings establish new objectives for improving the training and decoding of implicit reasoning systems.
Implicit chain-of-thought reasoning relying solely on outcome supervision is prone to semantic drift and gradient vanishing, hindering robust inference. This work addresses these limitations by reframing process supervision through an information-theoretic lens, decoupling it into trajectory and spatial supervision. Rather than enforcing geometric compression, the proposed approach preserves informational fidelity in the reasoning space via mutual information maximization. The authors introduce a “dual collapse” mechanism to elucidate the root causes of failure and develop a dual-dimensional supervision framework that jointly governs trajectory and spatial aspects. Semantic structure is retained through generative reconstruction instead of rigid geometric constraints. Experiments demonstrate a strong correlation between reasoning accuracy and the fidelity of latent trajectory information, establishing an “information–performance binding” principle that offers principled guidance for supervising implicit reasoning systems.
This study investigates the evolution of internal representations during chain-of-thought reasoning in large language models, aiming to understand, predict, and intervene in the correctness of their outputs. By modeling reasoning as structured trajectories in representation space, the work reveals for the first time that reasoning steps progress along specific subspaces in an ordered manner, with correct and incorrect solutions systematically diverging in later stages. Leveraging the geometric properties of these trajectories, the authors propose a novel inference-time trajectory-guidance paradigm. Integrating representational geometry analysis, trajectory clustering, and ROC-AUC evaluation, this approach enables high-accuracy prediction of final answer correctness as early as the midpoint of reasoning (achieving an AUC of 0.87) and supports dynamic correction and control over reasoning length.
This study investigates whether Latent-CoT models, exemplified by CODI, genuinely perform implicit step-by-step reasoning. Employing interpretability techniques—including logit-lens decoding, linear probing, attention analysis, and activation patching—the work systematically examines how intermediate states are represented and propagated in polynomial iteration tasks. The analysis reveals, for the first time, that CODI constructs complete reasoning paths in short-hop tasks but shifts to relying on compressed shortcuts in long-hop settings, retaining only late-stage intermediate representations. This strategy proves highly fragile under distributional shifts or increased optimization difficulty, exposing a fundamental vulnerability in its reasoning process.
This work addresses a critical limitation in current large language model (LLM) reasoning research, which overly relies on surface-level chain-of-thought (CoT) traces while neglecting the role of latent state trajectories, leading to biased interpretability analyses, evaluations, and intervention strategies. The paper reconceptualizes LLM reasoning as the emergence of latent state trajectories and formally distinguishes and articulates three competing hypotheses about the nature of reasoning, advocating for latent state trajectories as the default object of study. Through a combination of matched computational budget scaling, latent interventions, and surface trace decomposition—supported by empirical analysis, mechanistic investigation, and computational auditing—the study demonstrates the superiority of the latent state trajectory hypothesis (H1), establishing a new paradigm for evaluating and intervening in LLM reasoning processes.
Existing continuous chain-of-thought (Continuous CoT) methods rely on slow autoregressive generation and suffer significant performance degradation on tasks requiring long reasoning trajectories. This work proposes C-MTP, a novel approach that, for the first time, directly supervises hidden states using the mean of corresponding chain-of-thought embeddings, thereby employing embedding averages as supervision signals to simplify training and eliminate the need for autoregressive decoding. The method outperforms existing direct supervision approaches on short reasoning tasks and matches the performance of indirect supervision methods. However, on long reasoning trajectories spanning hundreds of tokens, all current methods—including C-MTP—experience a performance drop of approximately 65%, revealing a fundamental limitation of contemporary Continuous CoT frameworks in long-horizon reasoning.
This work addresses the limitations of traditional multimodal reasoning methods that rely on textual chain-of-thought (CoT), which suffer from slow inference and constraints imposed by language expression. The authors propose CoLT, a novel framework that enables efficient reasoning through a small number of implicit latent states, eliminating the need for verbose intermediate text generation. CoLT introduces a lightweight external decoder trained under dual supervision—forward (predicting the next step) and backward (aligning with context)—alongside a coherence constraint on internal latent states to ensure stable training and semantically meaningful reasoning chains. During inference, the supervisory modules are removed to maximize efficiency. Experiments demonstrate that CoLT outperforms existing implicit reasoning and image-augmented visual reasoning approaches across eight benchmarks, achieving a 10.1× reduction in overall reasoning time and a 22.6× decrease in text decoding latency compared to textual CoT.
This work proposes Abstract Chain-of-Thought (ACoT), a novel approach that replaces natural language reasoning chains with sequences of discrete latent variables to achieve efficient and effective reasoning. Addressing the high computational cost of traditional explicit chain-of-thought methods and the performance degradation of existing non-linguistic compression techniques, ACoT integrates vocabulary-preserving discrete latents, mask-supervised fine-tuning, constrained decoding with self-distillation, and warm-start reinforcement learning. The model learns to reason using compact abstract symbols while maintaining strong task performance. Experiments demonstrate up to 11.6× reduction in reasoning tokens on mathematical reasoning, instruction following, and multi-hop tasks, matching the accuracy of explicit chain-of-thought methods and exhibiting strong cross-model generalization. Notably, the learned abstract symbols follow a power-law distribution reminiscent of natural language.
This work addresses the instability of existing implicit reasoning methods across diverse scenarios, where low-confidence reasoning paths often introduce noise and lead to high-confidence errors. To mitigate this, the authors propose a confidence-aware dynamic routing mechanism that adaptively selects between continuous latent-space reasoning for high-confidence contexts and discrete symbolic reasoning for low-confidence ones, enabling the first synergistic switching between these two paradigms. By integrating large language model–based confidence estimation, latent-space inference, and discrete chain-of-thought generation, the approach effectively suppresses error propagation and corrects flawed reasoning. Experimental results demonstrate consistent improvements across multiple STEM and programming benchmarks, with an average 19.70-point gain in Pass@1 accuracy, up to a 15.55% reduction in output length, and significantly enhanced calibration of reasoning confidence.