latent diffusion reasoning

Designs and evaluates models that represent reasoning as a temporal diffusion process in latent spaces—often visual latent spaces—by evolving latent states over time, generating progressive decoded “thought” steps, and predicting end-task outcomes such as final skill quality. Work includes specifying latent diffusion dynamics, training denoising/generative trajectories, constructing decoders to produce interpretable visual reasoning sequences, and measuring trajectory fidelity and outcome-prediction performance.

latentdiffusionreasoning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.39
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work challenges the prevailing assumption that video generation models perform temporal modeling through frame-to-frame sequential reasoning. Instead, it demonstrates that inference unfolds along the denoising steps of the diffusion process, revealing a “Chain-of-Steps” mechanism: the model explores multiple plausible solutions in early denoising stages and gradually converges in later steps, exhibiting reasoning-like behaviors such as working memory, self-correction, and perceptual anticipation. Through training-free analyses—including qualitative inspection, targeted probing, and latent trajectory integration—the study uncovers functional stratification within Diffusion Transformers. Leveraging these insights, a latent-space ensemble strategy substantially enhances inference performance, offering a novel perspective on the dynamic reasoning mechanisms underlying video generative models.

Chain-of-Framesdenoising stepsdiffusion models

Tracing the Traces: Latent Temporal Signals for Efficient and Accurate Reasoning

Oct 12, 2025
MG
Martina G. Vilas
🏛️ Goethe University Frankfurt | Microsoft Research | NVIDIA

This work addresses computational inefficiency in large language model (LLM) inference, where redundant token generation wastes resources. We propose a latent-trajectory–based early path screening method that dynamically predicts the success probability of candidate reasoning paths during autoregressive decoding. Our core innovation is the Latent-Trajectory signal—a lightweight metric quantifying three aspects of hidden-state evolution: (i) initial-to-final representation divergence, (ii) cumulative intermediate variation, and (iii) convergence toward the final state. Unlike conventional confidence scores or majority voting, this signal enables robust path pruning and answer selection across multiple parallel sampling trajectories. Experiments on standard reasoning benchmarks demonstrate a 2.6% absolute accuracy gain while reducing total token consumption by 70%, significantly improving both inference efficiency and effectiveness under test-time scaling.

Improving efficiency through early selection of promising reasoning pathsPredicting reasoning path success to reduce wasted computationUsing latent trajectory signals for accurate solution prediction

This work investigates interpretable visual reasoning without linguistic supervision and proposes a rectified flow–based diffusion Transformer that models image-to-image reasoning end-to-end purely in pixel space. By integrating noise-free flow matching, Euler integration, and a Transformer architecture, the method efficiently solves tasks such as maze navigation in as few as ten iterative steps. Trained on million-scale synthetic data, it achieves a 192-fold reduction in training loss and a 22.7-fold improvement in L2 error. Notably, the study reports the first observation of an “aha!”-like phase transition analogous to human insight: in 68% of cases, reasoning progresses minimally for most of the trajectory before the correct solution emerges globally and synchronously within the final 2% of steps, challenging conventional assumptions of sequential, incremental reasoning.

diffusion modelsimplicit thoughtinsight phenomenon

In multimodal implicit reasoning, lightweight student models often rely excessively on linguistic priors while neglecting genuine visual perception, leading to significant divergence in visual attention from their teacher counterparts. To address this, this work proposes a novel paradigm that aligns the "latent visual thinking" of student and teacher models. Specifically, it employs autoregressive reconstruction of the teacher’s visual semantics and attention trajectories to align their dynamic visual reasoning processes prior to text generation. A curriculum-based sensory gating mechanism is further introduced to suppress shortcut learning. This approach represents the first explicit modeling and transfer of the teacher’s dynamic visual attention, achieving up to a 16.9% performance gain on complex reasoning tasks and enabling a 3B-parameter model to surpass both larger open-source models and closed-source systems such as GPT-4o.

knowledge distillationlatent reasoningmultimodal reasoning

Latest Papers

What's happening recently
View more

This work addresses the underutilization of visually grounded latent variables in current vision-language models, which—despite their rich semantic content—are systematically suppressed during inference. The study identifies this phenomenon as “silent visual latent variables” and introduces a novel inference-time optimization mechanism that requires no updates to the backbone parameters. By leveraging query-guided contrastive alignment between latent variables and visual features, coupled with a confidence-progressive reward scheme, the method enhances latent semantic quality and steers the prediction pathway in two stages while keeping the backbone frozen. Evaluated across eight benchmarks and four model backbones, the approach consistently achieves significant gains in multimodal reasoning performance, effectively unlocking the previously suppressed inferential capacity of visual latent variables.

latent reasoningmultimodal modelsoptimization pathology

This work addresses the limited interpretability of existing action quality assessment models, which often fail to reveal the visual reasoning processes underlying their judgments. To this end, the paper introduces, for the first time, a unified framework that integrates keypoint-guided Monte Carlo tree search into a latent visual diffusion model, enabling simultaneous high-accuracy skill evaluation and interpretable step-by-step visual reasoning. The proposed method produces clear reasoning trajectories that explicitly highlight the critical visual evidence supporting each prediction. Extensive experiments across four cross-domain datasets demonstrate that the model not only achieves competitive quantitative performance but also effectively visualizes its decision rationale, thereby substantially enhancing the transparency and trustworthiness of the assessment process.

action quality assessmentblack-box modelsfine-grained skill activities

This work addresses the challenge of discerning whether lengthy reasoning traces generated by large language models genuinely reflect their internal reasoning processes or merely constitute redundant output. To this end, the authors propose StALT, a novel metric that, for the first time, quantifies dynamic patterns in model hidden states across both temporal and layer dimensions without requiring additional training. StALT constructs a spatiotemporal magnitude statistic by analyzing hidden state trajectories, applying inter-layer saliency weighting, and measuring state transitions between adjacent tokens. Experimental results demonstrate that StALT reliably distinguishes between correct and incorrect reasoning traces across diverse models and reasoning-intensive tasks, while remaining sensitive to variations in reasoning demands, thereby validating its effectiveness and generality as an internal reasoning probe.

correctnesshidden-state dynamicsinternal computation

This work addresses the instability and ambiguity of reasoning trajectories in language model hidden states, which are highly sensitive to input rephrasing, model versions, and perturbations, lacking consistent and transferable reasoning directions. To tackle this, the authors propose TILR, a training-free intervention framework that, for the first time, reveals the existence of low-dimensional invariant subspaces within implicit reasoning trajectories. TILR extracts these subspaces by contrasting strong and weak reasoning paths and applies adaptive alignment gating to intervene in hidden states. Experiments demonstrate that TILR significantly improves answer consistency under rephrasing (by approximately 10%) and reduces trajectory variance by up to 50% across six reasoning benchmarks, all while preserving original reasoning accuracy—thereby enabling effective identification and manipulation of stable reasoning structures.

hidden-state spaceinvariant directionslanguage models

Hot Scholars

HW

Haoru Wang

Peking University, undergraduate
3d-visionCG
LB

Lei Bai

Shanghai AI Laboratory
Foundation ModelScience IntelligenceMulti-Agent SystemAutonomous Discovery
JH

Junjie Huang

College of Computer and Information Science, Southwest University, China
Social Network AnalysisGraph Neural NetworksComputational Social Science
MX

Meilong Xu

Stony Brook University
Machine LearningComputer VisionTopological Data Analysis