Score
Designs and implements control policies or controller modules that consume inferred latent variables (shared, residual, or latent-informed representations) to modulate actuation in real time; this includes building mechanisms to interpret latents, adjust synchronization strength, execution asymmetry and smoothness, and enforce joint-level safety constraints. These controllers are constructed to integrate with standard control pipelines so the latent information directly influences closed-loop behavior and safety checks.
本文通过条件流匹配方法将控制潜变量解码为四旋翼模型分布,以实现固定策略的在线预测调优和鲁棒性分析。
In partially observable environments relying solely on RGB observations, existing latent-space safety filters struggle to capture unseen safety constraints, leading to myopic safety behaviors. Method: We propose a mutual information-based metric to quantify observation completeness; design a multimodal supervised training strategy to enhance the latent space’s capacity to model implicit safety states; and integrate world-model learning of latent dynamics with Hamilton–Jacobi reachability analysis and a classification-based mechanism to construct a robust safety controller. Contribution: This work is the first to systematically expose the fundamental limitations of latent-space safety modeling under RGB-only observation. Evaluated in simulation and on a Franka Research 3 robotic arm, our approach effectively prevents latent hazards—such as overheating of a wax pot—demonstrating substantial improvements in both robustness and generalization of safety-critical control.
This study addresses the challenge of ensuring safety when embedding industrial control models into deployed systems, particularly regarding safety constraints in pressurized water reactor (PWR) load-following operations. To this end, a physics-decomposition-based structured representation learning method is proposed. This approach constructs a multi-timescale separation embedding architecture that maps variables across different timescales into independent latent spaces to emulate expert policies. It further enables hybrid deployment by integrating behavioral cloning with nonlinear model predictive control (NMPC). Experimental results demonstrate that the proposed method significantly improves the accuracy and feasibility of long-horizon trajectories, achieving fully feasible solutions with near-optimal costs while reducing computation time by approximately 15% compared to the expert controller.
To address safety and efficiency challenges in autonomous systems executing complex tasks under latent risks, this paper proposes a multi-timescale safety-critical control framework integrating large language models (LLMs), numerical optimization, and model predictive control (MPC). The method features a three-layer synergistic architecture: (i) a high-level semantic-driven subtask decomposition module leveraging in-context learning; (ii) a mid-level risk-adaptive parameter synthesis module guided by chain-of-thought reasoning to steer numerical optimization; and (iii) a low-level MPC-based closed-loop controller augmented with physics-informed simulation for enhanced dynamic robustness. Evaluated on robotic manipulation and autonomous driving simulations, the framework significantly improves task completion rates and behavioral safety in risk-sensitive scenarios, while enabling efficient online learning and adaptive behavior refinement.
This paper addresses the joint learning of state representations and controllers for unknown partially observable linear systems under the LQG control paradigm. Unlike conventional representation learning approaches that require observation reconstruction, we propose a cost-driven latent dynamics modeling framework that directly optimizes multi-step control costs—bypassing observation reconstruction and enabling end-to-end joint learning of representations and controllers. Theoretically, we establish the first finite-sample, provably guaranteed analysis for cost-driven latent model learning, revealing that accurate multi-step cost prediction is both necessary and sufficient for representation identifiability and near-optimal control performance. Methodologically, our approach integrates empirical risk minimization, system identification, and robust control analysis. Under finite-sample conditions, the learned representation and controller converge to the optimal solution, thereby closing a long-standing gap in provable guarantees for this paradigm.
This work addresses a critical gap in existing approaches that typically treat tool selection or action sequences as control variables while overlooking explicit regulation of prompt context construction. For the first time, we formalize context assembly—encompassing prompt templates, example selection, and the amount of retrieved content—as controllable variables and introduce a novel dual-layer policy architecture. The outer layer employs either a contextual multi-armed bandit or REINFORCE algorithm for online context modulation, while the inner layer leverages a frozen large language model to execute downstream tasks. We establish a theoretical framework analyzing stability and uncertainty, proving that expected reward is non-decreasing under bounded policy updates. Empirical results further demonstrate that the controller’s confidence is well-calibrated with task performance, validating the efficacy of our approach.
This study addresses the challenge of ensuring physical quantity recovery and command response consistency in latent world models for vehicle control. We propose a decoupled evaluation framework for action-conditioned latent predictors based on the temporal Joint Embedding Predictive Architecture (JEPA). Leveraging IPG CarMaker simulation data and a physics readout mechanism, this framework independently assesses state retention, geometric organization, prediction accuracy, and local responsiveness using untrained encoder baselines, thereby precisely localizing error sources within either the representation or the predictor. Our analysis reveals an inherent trade-off between retention and responsiveness, demonstrating that prediction error alone is insufficient for evaluating control suitability. Ultimately, this work establishes a rigorous foundation for closed-loop control assessment in latent space.
This work addresses the challenge of achieving general and efficient control of partial differential equation (PDE) systems without task-specific objectives or reward signals. To this end, it proposes a goal-agnostic PDE control framework that combines an offline-trained, frozen Vision Transformer (ViT) encoder with an action-conditional latent dynamics model based on the Joint Embedding Predictive Architecture (JEPA), integrated with Model Predictive Path Integral (MPPI) control for online planning. Innovatively, the approach couples goal-agnostic latent space modeling with probes on physically meaningful observables—such as kinetic energy—enabling a single frozen world model to support diverse control tasks. Evaluated on Navier–Stokes benchmarks, the method substantially improves performance: kinetic-energy-probe-based planning raises the 50-episode average reward from −12.08 to −10.90 and reduces late-stage velocity field RMSE by 9.5%; across three unseen non-periodic targets, it cuts late-field RMSE by 53% and wins all 30 head-to-head trials; steady-state control achieves a 2.7% average relative error.
This work addresses a critical limitation in existing latent-variable world models, which rely on average prediction error over training data for training and selection—a metric that fails to reflect actual controller performance due to a mismatch between the evaluation distribution and the distribution queried by the planner. The authors propose instead to center model assessment on the discrepancy between predicted and true costs over states reachable by the planner. They establish, for the first time, a rigorous theoretical link between this discrepancy and control suboptimality, proving it provides a valid upper bound on performance loss, whereas conventional prediction errors neither bound nor track performance. Leveraging control theory, spectral analysis, and non-normal operator theory, they decompose the discrepancy into an intrinsic manifold residual and an off-manifold divergence term, and introduce a fidelity score to quantify alignment of the planner’s reachable distribution. Experiments on synthetic systems and model predictive control confirm that the proposed metric reliably tracks control performance, while single-step prediction error shows virtually no correlation.
This study addresses the challenge of robust decision-making arising from the difficulty of quantifying latent-space uncertainty in world models. We propose a conformal prediction-based latent perturbation modeling approach that constructs trustworthy latent dynamic sets within a game-theoretic optimization framework. By integrating dynamics-aware similarity metrics with out-of-distribution detection, the method effectively balances pessimism and plausibility in decision-making. Both simulation and hardware experiments demonstrate that our approach reduces safety filter failure rates by 70% and sampling strategy overhead by 54%, achieving data-efficient robust decision-making and policy guidance in latent spaces.