Score
Designs and evaluates models and representations that capture how a system’s state evolves under actions or time, including predictive dynamics models, dynamics-aware architectures, and low-dimensional dynamics embeddings. These artifacts include methods for learning transferable dynamics (for example via spectral MDP decomposition) to enable zero-shot policy generalization and rapid few-shot adaptation across tasks.
Behavioral foundation models (BFMs) exhibit severely limited zero-shot generalization under dynamics shifts, hindering their deployment in real-world robotic applications. To address this, we propose a Forward-Backward (FB) adaptive representation framework. Our method introduces a novel Transformer-based belief estimator that implicitly models unknown dynamics, coupled with unsupervised clustering of dynamics-aware policy embeddings—effectively decoupling policy representations from environmental dynamics to enhance cross-dynamics zero-shot transfer. Crucially, the approach requires no fine-tuning. Evaluated on both discrete and continuous control benchmarks, it achieves zero-shot returns twice those of state-of-the-art baselines. This significantly improves BFMs’ robustness and generalization to unseen dynamics at test time.
This work addresses the challenge of efficiently estimating the spectra of Koopman and transfer operators for nonlinear dynamical systems from data. We propose a dynamic adaptive orthogonal local basis learning framework that, for the first time, deeply integrates adaptive basis function learning with invariant subspace construction for the target operator. Within an end-to-end differentiable framework, we jointly optimize deep neural network parameters, orthogonality constraints, and kernel-based reproducing structure—yielding basis functions that are locally supported, strictly orthogonal, and dynamically sensitive. Compared to standard DMD and EDMD, our method achieves significantly improved accuracy in recovering dominant spectral components (e.g., decay rates and frequencies) across diverse chaotic systems and high-dimensional manifolds; operator approximation error is reduced by over 40%, and the framework demonstrates strong generalization capability.
This work addresses the limited generalization of reinforcement learning policies under unmodeled or time-varying dynamics by proposing a trajectory-outcome-driven implicit dynamics representation that eschews reliance on predefined physical parameters. A task-specific smooth latent space is constructed via semi-supervised contrastive learning, and the authors theoretically establish a monotonic relationship between the regret bound in the target domain and the Lipschitz constant of the trajectory encoder. Leveraging this insight, they enforce Lipschitz constraints to optimize the geometry of the latent space, thereby enhancing robustness. Experiments on MuJoCo benchmarks demonstrate that the proposed method substantially outperforms parameter-centric baselines, effectively handling complex dynamics shifts while improving in-domain stability and interpretability of the latent representation.
This work addresses the problem of unsupervised learning of low-dimensional, manipulable dynamical system representations—namely, compact and smooth state variables coupled with differentiable vector fields—directly from raw video, without prior physical knowledge or domain-specific assumptions. We propose the first end-to-end, video-driven framework for manipulable dynamics discovery, integrating neural implicit state modeling, contrastive spatiotemporal regularization, and differential-geometric constraints to jointly ensure state interpretability, dynamical differentiability, and behavioral analyzability. Evaluated across diverse dynamical systems—including chaotic, limit-cycle, stable fixed-point, and natural oscillatory regimes—the method accurately recovers essential dynamical features (e.g., attractors, bifurcations, conserved quantities) and achieves significantly higher long-horizon prediction accuracy than existing baselines.
Multi-source short time-series data suffer from limited per-sequence length, hindering accurate modeling of complex dynamical mechanisms. Method: We propose the first hierarchical unsupervised generative framework that jointly learns population-level shared priors and domain-specific dynamics. Our approach integrates variational inference, multi-domain dynamical system reconstruction (DSR), and interpretable latent-space learning to construct a linearly controllable, low-dimensional feature space—enabling cross-parameter-domain transfer and fundamental dynamical modeling. Contributions/Results: (1) First automatic discovery of interpretable dynamical features under a hierarchical structure; (2) High-fidelity single-domain reconstruction on standard DSR benchmarks and real-world neuroscience/clinical datasets; (3) Significantly improved generalization to unseen parameter regimes and modeling robustness in few-shot settings.
This work addresses the lack of interpretability in existing foundation models for zero-shot dynamical system reconstruction, which often obscure their prediction mechanisms. The authors propose DynaBase, a minimal interpretable architecture comprising only two parameters, that predicts future states via a linear combination of the current latent state and the nearest neighbors—along with their successors—in a contextual memory bank. By integrating model parsimony, nearest-neighbor retrieval, and analytical optimization, DynaBase achieves high-performance zero-shot reconstruction for the first time and yields a one-parameter family of maps that unifies chaotic and periodic dynamics, reconciling conflicting views in the literature. Evaluated across diverse systems, DynaBase outperforms existing models while using orders of magnitude fewer parameters and admits a closed-form MSE solution, enabling direct optimization toward reconstruction metrics.
This work addresses the problem of efficiently learning compact world model representations for reinforcement learning from interaction data. The proposed method models the environment transition kernel as a third-order tensor over state–action–next-state triples and introduces, for the first time, CP decomposition to obtain factorized feature maps for each modality. A unified spectral representation is constructed by jointly optimizing encoders via a noise-contrastive objective. This approach substantially reduces the hypothesis space, leading to improved sample efficiency and strong performance on high-dimensional control tasks. Moreover, the learned state encoder exhibits cross-actuator transferability, requiring only fine-tuning of the action encoder to adapt to new dynamics.
This work addresses the challenge of modeling and forecasting stochastic nonlinear dynamical systems under noisy and partially observable conditions by proposing a deep spectral learning framework. The method employs a learnable neural encoder to construct Markovian latent states in a feature space, whose dynamics and observations are governed by learned transition and observation operators. It uniquely unifies spectral learning, Bayesian filtering, and Koopman mode decomposition within an end-to-end trainable architecture. Efficient state estimation is achieved through functional canonical correlation analysis, Galerkin projection, and a closed-form ridge-regularized solution. Experimental results demonstrate that the proposed approach significantly outperforms baseline methods—including sequential Bayesian filters and dynamic mode decomposition—across diverse scenarios, exhibiting strong robustness to both observation noise and partial observability.
Existing dynamical system reconstruction models exhibit limited out-of-distribution generalization, particularly when extrapolating across critical points. This work identifies three fundamental structural deficiencies underlying this limitation and introduces an improved framework based on topological feature disentanglement and hierarchical modeling. For the first time, the study derives a closed-form theoretical bound characterizing the reliable extrapolation range of such models. The proposed approach enables high-accuracy, zero-shot predictions in unseen dynamical regimes—such as regions straddling bifurcation points—without requiring additional training, thereby substantially enhancing out-of-distribution generalization performance.
This work addresses the limited generalization of reinforcement learning agents to new tasks, which often necessitates training from scratch. To overcome this, the authors propose Outcome-Predictive State Representations (OPSRs) and the OPSR Skill framework, which construct compact, task-agnostic state abstractions and define reusable abstract actions—referred to as skills—on top of these representations. This approach is the first to jointly abstract both states and actions, enabling cross-task skill transfer without requiring task-specific preprocessing, while preserving policy optimality. Empirical results demonstrate that OPSR-based skills significantly accelerate learning across multiple unseen tasks, confirming their strong generalization capability and effectiveness.