Score
Designs and evaluates models that represent reasoning as a temporal diffusion process in latent spaces—often visual latent spaces—by evolving latent states over time, generating progressive decoded “thought” steps, and predicting end-task outcomes such as final skill quality. Work includes specifying latent diffusion dynamics, training denoising/generative trajectories, constructing decoders to produce interpretable visual reasoning sequences, and measuring trajectory fidelity and outcome-prediction performance.
This work challenges the prevailing assumption that video generation models perform temporal modeling through frame-to-frame sequential reasoning. Instead, it demonstrates that inference unfolds along the denoising steps of the diffusion process, revealing a “Chain-of-Steps” mechanism: the model explores multiple plausible solutions in early denoising stages and gradually converges in later steps, exhibiting reasoning-like behaviors such as working memory, self-correction, and perceptual anticipation. Through training-free analyses—including qualitative inspection, targeted probing, and latent trajectory integration—the study uncovers functional stratification within Diffusion Transformers. Leveraging these insights, a latent-space ensemble strategy substantially enhances inference performance, offering a novel perspective on the dynamic reasoning mechanisms underlying video generative models.
This work addresses computational inefficiency in large language model (LLM) inference, where redundant token generation wastes resources. We propose a latent-trajectory–based early path screening method that dynamically predicts the success probability of candidate reasoning paths during autoregressive decoding. Our core innovation is the Latent-Trajectory signal—a lightweight metric quantifying three aspects of hidden-state evolution: (i) initial-to-final representation divergence, (ii) cumulative intermediate variation, and (iii) convergence toward the final state. Unlike conventional confidence scores or majority voting, this signal enables robust path pruning and answer selection across multiple parallel sampling trajectories. Experiments on standard reasoning benchmarks demonstrate a 2.6% absolute accuracy gain while reducing total token consumption by 70%, significantly improving both inference efficiency and effectiveness under test-time scaling.
This work investigates interpretable visual reasoning without linguistic supervision and proposes a rectified flow–based diffusion Transformer that models image-to-image reasoning end-to-end purely in pixel space. By integrating noise-free flow matching, Euler integration, and a Transformer architecture, the method efficiently solves tasks such as maze navigation in as few as ten iterative steps. Trained on million-scale synthetic data, it achieves a 192-fold reduction in training loss and a 22.7-fold improvement in L2 error. Notably, the study reports the first observation of an “aha!”-like phase transition analogous to human insight: in 68% of cases, reasoning progresses minimally for most of the trajectory before the correct solution emerges globally and synchronously within the final 2% of steps, challenging conventional assumptions of sequential, incremental reasoning.
In multimodal implicit reasoning, lightweight student models often rely excessively on linguistic priors while neglecting genuine visual perception, leading to significant divergence in visual attention from their teacher counterparts. To address this, this work proposes a novel paradigm that aligns the "latent visual thinking" of student and teacher models. Specifically, it employs autoregressive reconstruction of the teacher’s visual semantics and attention trajectories to align their dynamic visual reasoning processes prior to text generation. A curriculum-based sensory gating mechanism is further introduced to suppress shortcut learning. This approach represents the first explicit modeling and transfer of the teacher’s dynamic visual attention, achieving up to a 16.9% performance gain on complex reasoning tasks and enabling a 3B-parameter model to surpass both larger open-source models and closed-source systems such as GPT-4o.
This work addresses the underutilization of visually grounded latent variables in current vision-language models, which—despite their rich semantic content—are systematically suppressed during inference. The study identifies this phenomenon as “silent visual latent variables” and introduces a novel inference-time optimization mechanism that requires no updates to the backbone parameters. By leveraging query-guided contrastive alignment between latent variables and visual features, coupled with a confidence-progressive reward scheme, the method enhances latent semantic quality and steers the prediction pathway in two stages while keeping the backbone frozen. Evaluated across eight benchmarks and four model backbones, the approach consistently achieves significant gains in multimodal reasoning performance, effectively unlocking the previously suppressed inferential capacity of visual latent variables.
This work addresses the limited interpretability of existing action quality assessment models, which often fail to reveal the visual reasoning processes underlying their judgments. To this end, the paper introduces, for the first time, a unified framework that integrates keypoint-guided Monte Carlo tree search into a latent visual diffusion model, enabling simultaneous high-accuracy skill evaluation and interpretable step-by-step visual reasoning. The proposed method produces clear reasoning trajectories that explicitly highlight the critical visual evidence supporting each prediction. Extensive experiments across four cross-domain datasets demonstrate that the model not only achieves competitive quantitative performance but also effectively visualizes its decision rationale, thereby substantially enhancing the transparency and trustworthiness of the assessment process.
This work addresses the challenge of discerning whether lengthy reasoning traces generated by large language models genuinely reflect their internal reasoning processes or merely constitute redundant output. To this end, the authors propose StALT, a novel metric that, for the first time, quantifies dynamic patterns in model hidden states across both temporal and layer dimensions without requiring additional training. StALT constructs a spatiotemporal magnitude statistic by analyzing hidden state trajectories, applying inter-layer saliency weighting, and measuring state transitions between adjacent tokens. Experimental results demonstrate that StALT reliably distinguishes between correct and incorrect reasoning traces across diverse models and reasoning-intensive tasks, while remaining sensitive to variations in reasoning demands, thereby validating its effectiveness and generality as an internal reasoning probe.
This work addresses the instability and ambiguity of reasoning trajectories in language model hidden states, which are highly sensitive to input rephrasing, model versions, and perturbations, lacking consistent and transferable reasoning directions. To tackle this, the authors propose TILR, a training-free intervention framework that, for the first time, reveals the existence of low-dimensional invariant subspaces within implicit reasoning trajectories. TILR extracts these subspaces by contrasting strong and weak reasoning paths and applies adaptive alignment gating to intervene in hidden states. Experiments demonstrate that TILR significantly improves answer consistency under rephrasing (by approximately 10%) and reduces trajectory variance by up to 50% across six reasoning benchmarks, all while preserving original reasoning accuracy—thereby enabling effective identification and manipulation of stable reasoning structures.