Score
Designs and implements models, loss terms, and training procedures that infer per-step actions from pairs of states (inverse dynamics) and enforce cycle or action-consistency between those inferred actions and the actions predicted by a forward-transition model. Builds regularizers that compute and adaptively scale per-step action residuals (the difference between inferred and predicted actions), integrate them into optimization, and analyze their effect on learned transition and action predictors.
This work addresses the practical bottleneck in imitation learning—scarce access to expert action labels—by systematically studying “Learning from Observation” (LfO), which relies solely on expert state sequences. We propose the first unified taxonomy for LfO, organizing methods along two orthogonal dimensions: modeling objectives (state reconstruction, latent variable inference, policy alignment) and algorithmic mechanisms (integration with RL, model-based prediction, or hierarchical architectures). We establish, for the first time, rigorous theoretical connections between LfO and offline RL, model-based RL, and hierarchical RL. Our analysis precisely characterizes key assumptions, fundamental performance limits, and domain-specific applicability conditions for each method class. Furthermore, we identify major open challenges—including ambiguity in inverse dynamics and distributional shift—and propose empirically verifiable pathways toward resolution. This work lays a structured conceptual foundation for LfO, advancing its development toward greater robustness, interpretability, and real-world applicability.
In learning-based control, single-step autoregressive prediction suffers from error accumulation in closed-loop operation, leading to performance degradation—especially under model misspecification (e.g., partial observability). This work focuses on linear dynamical systems and provides the first rigorous characterization of the bias–complexity trade-off between single-step and multi-step prediction. We theoretically establish that multi-step prediction significantly reduces asymptotic bias even under imperfect models, and its benefits outweigh the associated increase in modeling complexity. Methodologically, we propose training single-step models using a multi-step loss, complemented by closed-loop performance evaluation and tight error bound analysis. Numerical experiments demonstrate that our approach outperforms standard single-step methods both in open-loop prediction accuracy and closed-loop robustness. This work furnishes both theoretical foundations and a practical design paradigm for learning-based controllers.
This work addresses the insufficient joint modeling of action steps and scene state changes in procedural activity understanding. We propose a process-aware video representation learning framework that, for the first time, leverages explicit state-change descriptions generated by large language models (LLMs) as supervisory signals—and constructs their counterfactual variants—to jointly model bidirectional causal relationships between actions and states, thereby enhancing “if–then” reasoning. Our key contributions are: (1) the first use of LLM-generated state descriptions and their counterfactual counterparts for video representation learning; and (2) unified causal modeling of both normative procedures and anomalous/erroneous steps. Extensive experiments demonstrate significant improvements over state-of-the-art methods on temporal action segmentation and procedural error detection, validating the effectiveness of explicit state supervision and counterfactual reasoning for procedural understanding.
Machine learning training is significantly more time-consuming than inference, and the design of input or parameter perturbations has long relied on empirical trial-and-error. Method: This paper models training dynamics as a first-passage process and introduces a statistical mechanics framework to analyze model responses to input/parameter perturbations. It proposes, for the first time, a single-frequency perturbation response theory grounded in the quasi-stationary assumption, and rigorously proves its generalizability to multi-frequency perturbation regimes—enabling rational optimization of perturbation protocols. Contribution/Results: Evaluated on ResNet-18 trained for CIFAR-10 classification, the method precisely identifies the optimal perturbation type and frequency, reducing training iterations by 23% and improving test accuracy by 1.4 percentage points, thereby substantially enhancing both training efficiency and generalization performance.
This work addresses the problem of automatic learning of planning domain models without action parameter annotations: given only action names and fully observable state trajectories, the task is to infer the number and types of action parameters, along with their preconditions and effects. We propose a hybrid method integrating state-transition analysis, logical induction, and constraint solving, augmented by IPC benchmark-driven heuristic pruning and pattern matching. To our knowledge, this is the first end-to-end approach that learns complete action models without any prior assumptions—neither on parameter count, type, nor domain constraints. We prove its computational complexity is at least as hard as graph isomorphism, yet demonstrate practical feasibility. Evaluated on multiple IPC benchmarks, our method achieves significantly higher model similarity than SAM and Extended SAM, while requiring fewer input assumptions and imposing looser trajectory constraints.
This paper addresses the fundamental question of whether latent action models (LAMs) genuinely learn action-driven inter-frame dynamics or merely capture exogenous noise. Method: We develop an analytically tractable linear system model to theoretically characterize the learning mechanism of LAMs, uncovering their intrinsic relationship with principal component analysis (PCA) and rigorously analyzing how structural coupling among observations, actions, and noise governs model performance. Leveraging controllability theory, we derive principled guidelines for designing data generation strategies. Contribution/Results: These guidelines inform video data augmentation, noise denoising, and auxiliary action prediction. Numerical simulations demonstrate that our strategy significantly enhances learning of action-relevant features, thereby advancing the interpretability and reliability of unsupervised action representation learning.
This study addresses the issue of unreliable policy updates in continuous control caused by errors in the action derivatives of the critic. To overcome this limitation, it proposes a forward entropy-regularized policy optimization algorithm that directly optimizes the policy by constructing a target distribution and minimizing the forward Kullback-Leibler (KL) divergence, thereby circumventing the need for critic differentiation. By leveraging the mode-covering property of the forward KL divergence, the method effectively explores multimodal high-value regions. Furthermore, training stability is maintained through a combination of KL regularization constraints and self-normalized importance sampling. Experimental evaluations on the MuJoCo and ManiSkill benchmarks demonstrate that the proposed approach achieves competitive performance and sample efficiency, while delivering faster actor update speeds compared to REPPO.
This study addresses the limitation of Joint Embedding Predictive Architecture (JEPA) models, which achieve accurate predictions yet struggle to recover latent causal states. To this end, we investigate the theoretical foundations and conditions under which causal states can be recovered from observations. Methodologically, we propose an information-theoretic objective function alongside identifiability conditions, and introduce the Action-Modulated JEPA (A-JEPA). This framework integrates latent variable modeling, conditional likelihood estimation, entropy maximization, and Gaussian additive noise to facilitate effective optimization. Experimental results on visual benchmarks demonstrate that the proposed approach significantly enhances the recovery of causal states while exhibiting strong transfer generalizability to previously unseen mechanisms.
This study addresses the problem of recursive prediction error accumulation in visual planning by proposing the SALT model. Its core innovation lies in introducing state-affine properties into latent-space dynamics modeling for the first time, ensuring that the error propagation operator depends solely on the action sequence and thereby eliminating nonlinear error propagation. Furthermore, SALT integrates a joint-embedding world model with a recursive multi-step rollout supervision mechanism to optimize training. Experimental results demonstrate that SALT improves closed-loop success rates by an average of 10% and significantly reduces the failure rate on the OGBench-Cube task from 23.3% to 2.0%, achieving highly reliable visual planning.
Although JEPA-based world models perform prediction in latent space, they remain susceptible to visual perturbations, leading to distorted state representations and biased action-conditioned predictions. This work proposes Action-Conditioned Prediction Consistency (ACPC), a diagnostic framework that, for the first time, operationalizes the twin-simulation concept into a computable metric by quantifying the divergence between multi-step forward rollouts from clean and perturbed observations under identical action sequences. To enable robustness evaluation across tasks and architectures, we introduce two complementary metrics: Invariance Radius (IR) and Separation Rate (SR). Empirical results demonstrate that ACPC effectively predicts perturbation-induced errors in prediction and planning costs, and the IR–SR pair exhibits strong generalization and diagnostic capability across models such as LeWM and PLDM.
本文解决了重建与动作不匹配的问题,提出ACT-LAM框架,通过改进动作提取和利用方法,提高潜行动态模型的性能。