anticipatory state prediction

Designs and builds models and modules that predict future internal (latent) states from current observations and refine those predicted latents to improve their semantic distinctness and usefulness for downstream inference. This includes learning future-latent predictors and anticipatory latent‑sharpening techniques to increase predictive accuracy and separability, particularly at short lead times.

anticipatorystateprediction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.51
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work proposes EIDOS, a novel time series foundation model that shifts the pretraining objective from direct observation-space forecasting to latent-space dynamics modeling. Existing approaches predict future values directly in the observation space, making them susceptible to noise and resulting in unstructured, inconsistent latent representations. In contrast, EIDOS employs a causal Transformer to forecast the evolution of latent representations, complemented by a lightweight aggregation branch to construct stable training targets. The model is trained via multi-task joint optimization, integrating latent alignment, observation anchoring, and direct prediction supervision to learn predictable and well-structured latent dynamics. Evaluated on the GIFT-Eval benchmark, EIDOS achieves state-of-the-art performance, significantly enhancing representation consistency and model robustness.

latent representationspredictive learningrepresentation structure

This work addresses a critical yet previously unarticulated issue in time series forecasting—termed “latent chaos”—where conventional methods operating directly in the observation space learn representations that are temporally inconsistent and lack continuity, thereby failing to capture the true underlying dynamics of the system. To overcome this limitation, the paper introduces LatentTSF, a novel paradigm that leverages an autoencoder to construct a high-dimensional latent state space in which prediction is performed. By implicitly maximizing the mutual information among latent states, ground-truth system states, and observations, LatentTSF enforces temporal coherence and dynamical fidelity. Theoretical analysis and extensive experiments demonstrate that this approach substantially mitigates latent chaos and achieves state-of-the-art forecasting performance across multiple established benchmarks.

Latent ChaosLatent RepresentationObservation-space Forecasting

This work addresses the challenge of maintaining stable and coherent internal representations in long-sequence modeling, particularly under out-of-distribution (OOD) generalization settings where existing methods often underperform. To this end, the authors propose an alternating fast-slow recurrent mechanism that interleaves slow observational updates with rapid, self-organizing latent state updates. This design enables the model to internally “reason” while processing inputs, dynamically constructing clustered yet temporally coherent long-range representations. By integrating self-organizing representation learning with sequential modeling, the approach significantly outperforms established baselines—including LSTMs, state space models, and Transformers—in both reinforcement learning and algorithmic tasks, demonstrating markedly improved OOD generalization over extended time horizons.

internal representation stabilitylatent recurrencelong-horizon sequential modeling

Standard next-token prediction is limited in long-horizon reasoning by its narrow receptive field and the accumulation of multi-step errors. This work proposes a hierarchical latent variable prediction mechanism that introduces high-level abstract latent variables for the first time, integrating self-supervised latent space prediction, multi-step rollout, and hierarchical representation learning. This approach effectively mitigates error propagation during latent space rollouts and enhances the model’s capacity to maintain coherent, long-range belief states. Evaluated on code generation and multi-step reasoning benchmarks, the method significantly outperforms existing approaches and substantially improves speculative decoding efficiency.

error accumulationlanguage model pre-traininglatent-space rollout

Latest Papers

What's happening recently
View more

This work addresses a critical limitation in existing latent-variable world models, which rely on average prediction error over training data for training and selection—a metric that fails to reflect actual controller performance due to a mismatch between the evaluation distribution and the distribution queried by the planner. The authors propose instead to center model assessment on the discrepancy between predicted and true costs over states reachable by the planner. They establish, for the first time, a rigorous theoretical link between this discrepancy and control suboptimality, proving it provides a valid upper bound on performance loss, whereas conventional prediction errors neither bound nor track performance. Leveraging control theory, spectral analysis, and non-normal operator theory, they decompose the discrepancy into an intrinsic manifold residual and an off-manifold divergence term, and introduce a fidelity score to quantify alignment of the planner’s reachable distribution. Experiments on synthetic systems and model predictive control confirm that the proposed metric reliably tracks control performance, while single-step prediction error shows virtually no correlation.

latent world modelsmodel-based controloff-manifold divergence

This work addresses the tendency of large language models to rely on time-specific knowledge in prediction tasks, which induces look-ahead bias. By employing sparse autoencoders to dissect internal model representations, the authors identify interpretable features associated with temporal awareness and, for the first time, apply targeted causal interventions to modulate these features. Experimental results demonstrate that enhancing temporal-awareness features significantly mitigates look-ahead bias, whereas direct intervention on bias-related features proves ineffective, underscoring the pivotal role of temporal perception in improving predictive generalization. The proposed approach effectively steers the model toward greater reliance on historical information while preserving its general reasoning capabilities.

forecastinggeneralizationlarge language models

Although large language models achieve high prediction accuracy, they often exhibit poor calibration and their chain-of-thought reasoning may not faithfully reflect the true basis of their predictions. This work proposes a representation-based probing method that pools intermediate-layer activations to uncover the model’s internal decision rationale. It demonstrates, for the first time, that internal representations provide a more reliable indicator of prediction grounds than chain-of-thought outputs, serving effectively as both a calibration tool and a “lie detector.” By integrating evidence ablation, perturbation injection, and pre-inference answer distribution analysis, the approach substantially improves calibration performance, correctly predicting the direction of perturbation effects in 84% of cases and reducing generated tokens by 30–47% through answer-aware routing without compromising accuracy.

calibrationchain-of-thoughtfaithfulness

Hot Scholars

YW

Yusong Wu

University of Montreal, Mila
Deep LearningGenerative ModelMusic Generation
IS

Ian Simon

Google Brain
Music GenerationMachine LearningMusic Information RetrievalComputer Vision
HC

Hongkai Chen

The Chinese Unversity of Hong Kong
Cyber-Physical SystemsInternet of ThingsSmart HealthSmart Cities
YF

Yuang Fan

PhD Student, Columbia University
XJ

Xiaofan Jiang

Associate Professor of Electrical Engineering, Columbia University
Mobile and Embedded SystemsArtificial Intelligence of ThingsSmart Health and FitnessCPHS