Score
Designs and builds predictive forward-dynamics models that estimate how a system's state evolves when particular actions are taken, i.e., functions that map current state and candidate action(s) to next state(s), observations, or distributions over trajectories. Work includes modeling the causal effects of actions, simulating state transitions under candidate behaviors, and quantifying uncertainty and multi-step rollout performance for planning or control.
This study addresses the conceptual ambiguity surrounding World Action Models (WAMs) by clarifying their distinctions from related paradigms such as world models, video generation, and vision-language action policies. Through two complementary lenses—generated content type and methodological composition—the work proposes the first unified taxonomy to systematically deconstruct WAM design paradigms. It reveals a fundamental trade-off between representational richness and computational, memory, latency, and action annotation costs, highlighting a growing trend toward generating minimal yet control-critical future information. The paper characterizes WAMs as inherently predictive-action synergistic mechanisms, distills common design patterns, and provides a systematic overview of current advances and open challenges across key dimensions including interactivity, causality, persistence, physical plausibility, and generalization capability.
How can piecewise-linear trajectory tracking be achieved for unknown nonlinear systems with time-varying dynamics—without prior system knowledge—while ensuring robustness against abrupt dynamic changes? Method: We propose a data-driven online control framework that (i) identifies local system dynamics in real time via small-perturbation excitation and local system identification; (ii) analytically derives the state reachable set without model assumptions, leveraging an upper bound on the local growth rate; and (iii) synthesizes a receding-horizon closed-loop controller based on reachable-set prediction. Contribution/Results: The approach avoids global modeling and parametric structural assumptions, enabling rapid adaptation to sudden dynamic shifts. Experimental validation across multiple unknown nonlinear systems demonstrates stable waypoint-sequence tracking, confirming strong robustness and generalization capability under unmodeled dynamics and disturbances.
This work addresses the challenge of generating dynamically feasible control trajectories in complex environments. We propose a novel framework that implicitly embeds system dynamics into the denoising process of diffusion models. Methodologically, we design a temporal-aware physical projection mechanism that aligns denoising steps with the noise schedule and enforces kinematic or dynamic constraints at each iteration—without requiring explicit dynamical priors—enabling recovery of linear feedback controller trajectories directly from expert demonstrations. Our key contributions are: (i) the first diffusion-based trajectory generation method ensuring dynamical consistency under maximum-likelihood estimation; and (ii) the integration of implicit dynamics learning with projection-based constraint optimization. Evaluated on standard control benchmarks and non-convex optimal control tasks involving obstacle avoidance and path tracking, our approach significantly improves trajectory feasibility and task success rates, demonstrating strong potential for real-world deployment.
This work proposes a unified mathematical framework grounded in dynamic information flow for constructing structurally rigorous models of future prediction. By integrating filtering theory, regular conditional probabilities, Markov semigroups, infinitesimal generators, and multiple information geometries—including Hilbert, Fisher–Rao, and Wasserstein—the approach conceptualizes prediction as the construction of conditional distributions governed by informational, geometric, and modeling constraints. The framework elucidates deep connections among classical results such as the tower property and semigroup laws, as well as Itô’s formula and backward equations. Explicit transition laws, spectral decompositions, term structures, and asymptotic behaviors are derived within canonical models like Ornstein–Uhlenbeck and Cox–Ingersoll–Ross, thereby establishing a compact mathematical mapping from idealized theoretical constructs to empirical forecasting.
This paper addresses the problem of next-state prediction for model-free dynamical systems—i.e., predicting future states without knowledge of the underlying evolution function and without parametric assumptions. Methodologically, it introduces novel combinatorial metrics and dimensions to characterize the fundamental statistical complexity of model-free time-series prediction for the first time. Leveraging tools from combinatorial learning theory, dynamical systems analysis, and online learning regret analysis, the paper rigorously derives tight optimal mistake bounds (in the realizable setting) and regret bounds (in the agnostic setting). The results establish the first systematic theoretical foundation for model-free time-series prediction, precisely quantifying its intrinsic statistical difficulty and delineating fundamental algorithmic limits.
Designing effective policies for non-stationary Markov decision processes (MDPs), where time-varying transition and reward functions undermine conventional reinforcement learning assumptions, remains challenging. Method: This paper proposes a low-regret online control framework that integrates forward-looking predictions—e.g., renewable generation and load forecasts—within a robust model predictive control (MPC) architecture. It introduces an online policy update mechanism resilient to model uncertainty and explicitly accounts for prediction inaccuracies. Contribution/Results: We establish, for the first time, an exponential decay relationship between lookahead horizon length and dynamic regret. Theoretically, we prove that dynamic regret remains strictly bounded under sub-exponential prediction error growth. Empirical evaluation in representative non-stationary environments demonstrates substantial performance gains over state-of-the-art baselines, validating both theoretical guarantees and practical efficacy.
This work addresses the limited human-like risk perception of autonomous vehicles in complex environments by proposing a learning framework that integrates multi-step belief-space distribution dynamics with model predictive control (MPC). The approach employs a structured neural network to learn the evolution of uncertainty, enabling natural speed modulation and intelligent decision-making without excessive conservatism. Innovatively embedding distributional predictions directly into real-time MPC optimization significantly enhances the system’s adaptability to dynamic risks. Evaluated on a large-scale real-world off-road driving dataset and validated on a full-scale vehicle, the method demonstrates robust performance and effectiveness across diverse and challenging terrains.
This work addresses a critical limitation in existing latent-variable world models, which rely on average prediction error over training data for training and selection—a metric that fails to reflect actual controller performance due to a mismatch between the evaluation distribution and the distribution queried by the planner. The authors propose instead to center model assessment on the discrepancy between predicted and true costs over states reachable by the planner. They establish, for the first time, a rigorous theoretical link between this discrepancy and control suboptimality, proving it provides a valid upper bound on performance loss, whereas conventional prediction errors neither bound nor track performance. Leveraging control theory, spectral analysis, and non-normal operator theory, they decompose the discrepancy into an intrinsic manifold residual and an off-manifold divergence term, and introduce a fidelity score to quantify alignment of the planner’s reachable distribution. Experiments on synthetic systems and model predictive control confirm that the proposed metric reliably tracks control performance, while single-step prediction error shows virtually no correlation.
This work investigates how to achieve optimal policies in Markov decision processes (MDPs) that incorporate future information—such as reference trajectories or predictions—by leveraging model predictive control (MPC). The authors formulate MPC as a class of parameterized policies and train them end-to-end via reinforcement learning. Their key contribution lies in establishing, for the first time, the precise structural conditions under which MPC can exactly represent the optimal value function and policy, thereby providing a theoretical foundation for MPC as a structured function approximator with formal guarantees. Empirical validation on a point-mass racing task with future reference trajectories demonstrates that the proposed approach learns policies approaching optimality, confirming its effectiveness.
Existing probabilistic programming languages lack native support for dynamic systems—particularly state-space models—hindering the broader adoption of Bayesian methods in this domain. This work introduces dynestyx, a library that provides first-class, unified, and user-friendly support for state-space models within a probabilistic programming framework. dynestyx enables flexible specification of priors, accommodates both discrete- and continuous-time dynamics, handles mixed-effects data, and facilitates joint Bayesian inference over latent states and model parameters with full uncertainty quantification. By doing so, this contribution substantially enhances the accessibility, flexibility, and practical utility of dynamic system modeling across statistics, signal processing, and machine learning.
Real-world time series are often highly irregular and severely missing due to sensor dormancy or transmission delays. Existing methods typically assume future observation times are known, overlooking the critical question of whether future values will even be observable. This work proposes Timeflies, a novel framework that reframes time series forecasting as a joint task of inferring future observability and estimating numerical values. Timeflies employs a dual-stream architecture—comprising an observation stream and a value stream—augmented with reliability-aware embeddings, observation-guided dependency modeling, and continuous-time dynamics to explicitly co-learn the presence of observations and the evolution of underlying states. Evaluated on the newly introduced Shadow benchmark and OVJE metric, Timeflies significantly outperforms existing approaches, demonstrating that jointly modeling future observability is essential for improving forecasting performance under substantial missingness.