Score
Designs, implements, and evaluates models and algorithms that infer and iteratively refine hidden (latent) states from observed data, producing time‑indexed state trajectories and weighting vectors. Analyzes and quantifies within‑ and between‑unit variability, maps latent states to decision probabilities or other downstream variables, and develops procedures for latent‑state modeling, refinement, and convergence assessment.
Existing latent variable models often suffer from under-constrained objectives, leading to non-identifiable, ambiguous, and poorly interpretable representations. This work proposes the Constrained Latent State Modeling (CLSM) framework, which systematically integrates six core constraints—namely predictive sufficiency, minimality, temporal consistency, and others—for the first time. Grounded in information theory and dynamical systems theory, CLSM formally characterizes the intrinsic couplings and trade-offs among these constraints. By reframing representation learning as a constrained optimization problem, the framework unifies diverse approaches such as variational autoencoders and state-space models, revealing that non-identifiability stems from insufficient constraints rather than technical shortcomings. CLSM thus provides a principled foundation for designing latent variable models that are interpretable, robust, and aligned with downstream tasks.
Accurate estimation of latent object states and motion dynamics from image observations remains challenging due to occlusions, clutter, and nonlinear observation processes. Method: This paper proposes an end-to-end joint learning framework that simultaneously performs latent-state filtering and identifies linear dynamical equations directly from raw images—marking the first such approach. It integrates deep generative modeling, variational Bayesian inference, and latent-space linear system identification within a fully differentiable architecture, circumventing error propagation inherent in conventional staged modeling. Contribution/Results: The framework unifies nonlinear observation mapping, latent-state dynamics modeling, and linear dynamical structure discovery, balancing estimation accuracy with model interpretability. Evaluated on simulated visual dynamical environments, it reduces filtering error by 32% compared to baseline methods; the identified dynamical parameters retain clear physical meaning. Consequently, the method significantly enhances both robustness and interpretability of state estimation under complex, realistic visual observations.
To address the cyclic dependency problem arising from the coupling of latent states and nonlinear dynamics in time-series modeling, this paper proposes LaNoLem: a method that models the system as a time-varying dynamical process in a latent space and decouples latent-state inference from dynamics learning via an alternating minimization algorithm. It introduces a fully automated, human-in-the-loop-free complexity regularization criterion to enable adaptive control of model capacity. By jointly optimizing latent-state representation, nonlinear differential equation learning, and dynamics estimation, LaNoLem achieves state-of-the-art accuracy in dynamical system identification. Moreover, it significantly outperforms existing methods on multi-step long-horizon forecasting tasks—particularly for systems exhibiting intricate hidden mechanisms and long-range temporal dependencies.
MuZero’s dynamics network relies on unobservable latent states for planning, undermining decision interpretability. Method: We propose an interpretability analysis framework that systematically evaluates latent state representation quality and dynamic modeling mechanisms—across Go, Gomoku, and Atari—using observation reconstruction loss and state consistency regularization. Contribution/Results: First, we find latent states achieve higher fidelity in board games but degrade in long-horizon Atari simulations. Second, the planning process itself exhibits intrinsic error-correction capability, mitigating distortions in latent state representations. Third, this self-correcting mechanism substantially enhances behavioral interpretability and robustness. Our findings provide novel insights into the internal operation of model-based reinforcement learning and establish a reproducible analytical paradigm for probing latent dynamics.
This work addresses a critical yet previously unarticulated issue in time series forecasting—termed “latent chaos”—where conventional methods operating directly in the observation space learn representations that are temporally inconsistent and lack continuity, thereby failing to capture the true underlying dynamics of the system. To overcome this limitation, the paper introduces LatentTSF, a novel paradigm that leverages an autoencoder to construct a high-dimensional latent state space in which prediction is performed. By implicitly maximizing the mutual information among latent states, ground-truth system states, and observations, LatentTSF enforces temporal coherence and dynamical fidelity. Theoretical analysis and extensive experiments demonstrate that this approach substantially mitigates latent chaos and achieves state-of-the-art forecasting performance across multiple established benchmarks.
This study addresses the problem of recovering latent discrete states from the evolving weights of models trained on time-varying data streams to characterize non-stationary distributional shifts. The proposed approach trains classifiers over sliding time windows, aligns their weight trajectories, and fits a hidden Markov model (HMM) to these trajectories—enabling, for the first time, the identification of semantically coherent temporal phases solely from weight dynamics. Experiments on the Fakeddit and Yelp datasets demonstrate that transfer performance within the same inferred state significantly outperforms cross-state transfer, and this advantage persists independently of temporal proximity and shifts in class distribution. These findings confirm that model weights encode structural information about data distributions that extends beyond local temporal correlations.