Score
Designs and implements models that represent and learn probability distributions over actions or policies, capturing multimodal, diverse candidate actions and enabling conditional sampling based on observed context. Builds training and inference procedures to generate, score, and sample from these action distributions while avoiding averaging effects that produce suboptimal actions.
This study addresses the conceptual ambiguity surrounding World Action Models (WAMs) by clarifying their distinctions from related paradigms such as world models, video generation, and vision-language action policies. Through two complementary lenses—generated content type and methodological composition—the work proposes the first unified taxonomy to systematically deconstruct WAM design paradigms. It reveals a fundamental trade-off between representational richness and computational, memory, latency, and action annotation costs, highlighting a growing trend toward generating minimal yet control-critical future information. The paper characterizes WAMs as inherently predictive-action synergistic mechanisms, distills common design patterns, and provides a systematic overview of current advances and open challenges across key dimensions including interactivity, causality, persistence, physical plausibility, and generalization capability.
This work addresses the inefficiency in imitation learning inference caused by repeated sampling of failed actions. We propose a recovery-oriented action generation framework based on conditional diffusion models. Methodologically, we pioneer the decomposition of long-horizon failure recovery into composable sub-policies and introduce a guided resampling mechanism that dynamically corrects the sampling distribution solely using successful demonstration data—without requiring additional exploration or high-level controllers. Our core contributions are: (1) policy decomposition modeling via diffusion models, and (2) distribution-guided resampling grounded in successful trajectories. Evaluated on tasks including door opening in unknown directions, object manipulation, and button searching, our approach significantly improves action success rates and execution efficiency while demonstrating strong robustness to varying numbers of prior failures.
To address the reliance of procedural action planning in instructional videos on frame-level annotations or speech commands—rendering methods susceptible to error accumulation—this paper proposes an end-to-end diffusion model supervised solely by start/end-frame visual observations and task-level labels. The method frames action sequence generation as a conditional distribution fitting problem without intermediate supervision. Its core contributions include: (i) the first formulation of procedural action generation as an unsupervised latent sequence modeling task under coarse-grained visual and semantic constraints; and (ii) a differentiable projection-guidance mechanism that injects both visual and semantic priors seamlessly during both training and sampling, drastically reducing dependence on fine-grained annotations. Built upon a U-Net backbone, the model fuses start/end-frame visual embeddings with an efficient, differentiable projection operator. Evaluated on three diverse-scale instructional video datasets, it achieves state-of-the-art performance across multiple metrics, outperforming prior approaches while requiring no intermediate action annotations.
This work addresses the challenge of fine-grained conditional control in pretrained diffusion models. We propose CTRL, a framework that formulates conditional generation as a reinforcement learning (RL) policy optimization problem. Specifically, CTRL employs a PPO variant with KL regularization to jointly optimize a policy network using offline annotated data; the composite reward comprises classifier outputs and KL divergence from the target conditional distribution. To our knowledge, this is the first approach to cast conditional diffusion generation as end-to-end RL policy learning—eliminating the need for intermediate state classifiers, thereby significantly reducing data requirements and training complexity. Moreover, CTRL enables soft optimal conditional sampling and enjoys theoretical convergence guarantees to the desired conditional distribution. Experiments on image generation demonstrate that CTRL outperforms classifier guidance and classifier-free guidance (CFG) using fewer labeled examples, achieving superior trade-offs between control fidelity and sample diversity.
In distributed reinforcement learning, parameterized distribution approximations of the true return distribution introduce substantial inductive bias, degrading generalization and undermining the reliability of uncertainty estimation. To address this, we propose the Diverse Projection Ensemble (DPE), a framework that integrates multiple Wasserstein-distance-compatible projection operators with parameterized distribution representations. We theoretically characterize how projection bias fundamentally affects generalization. Furthermore, DPE couples ensemble disagreement with an exploration reward derived from the 1-Wasserstein distance, yielding an uncertainty-aware deep exploration mechanism. Evaluated on the Behavior Suite and VizDoom benchmarks, DPE significantly outperforms state-of-the-art methods—particularly excelling in directed exploration tasks. These results empirically validate that diversity across projections is critical for robust exploration and reliable uncertainty estimation.
To address low sample efficiency in offline policy evaluation (OPE) and offline policy learning (OPL) under large action spaces, this paper proposes sDM—a unified Bayesian framework that introduces structured prior modeling to OPE/OPL for the first time, explicitly capturing inter-action correlations. Methodologically, sDM designs a correlation-aware Bayesian metric, replacing conventional worst-case analysis to enable instance-averaged performance assessment; it further integrates online Bayesian multi-armed bandit heuristics to jointly optimize OPE and OPL. Theoretically, we prove that modeling action correlations significantly reduces estimation variance. Empirically, sDM achieves superior evaluation accuracy and policy performance over state-of-the-art baselines across multiple benchmark tasks, while maintaining linear time complexity.
本文提出OHCAM方法,通过减少假设间的不确定性来学习包含条件和量化效果的动作模型,解决了现有方法在处理此类问题时的计算不可行性。
Current single-pass inference paradigms constrain the performance of non-deterministic generative models in robotic manipulation. To address this limitation, this work proposes TapSampling—a plug-and-play, inference-time sampling framework that enables policy-agnostic execution refinement by efficiently exploring the action latent space and incorporating a semantically interpretable task-progress prediction verifier. Built upon an action variational autoencoder (Action-VAE), TapSampling is compatible with both diffusion and autoregressive models and requires no fine-tuning. Experiments demonstrate that it significantly enhances task success rates and robustness across diverse general-purpose policies in both simulated and real-world environments.
This work addresses a fundamental limitation in conventional Bayesian experimental design, which relies on prior-to-posterior uncertainty reduction and yields an intractable objective that is doubly hard to evaluate and poorly aligned with downstream tasks. By reframing the problem through decision theory, the authors formulate it as optimizing the expected future loss (EFL) of downstream actions, thereby reducing the objective to a singly intractable form that obviates explicit posterior or marginal likelihood computation. They introduce a stochastic gradient method that jointly optimizes both the experimental design and the action policy, requiring only samples from the joint parameter–data model and evaluations of the loss function. This approach naturally accommodates implicit modeling and task-specific customization, demonstrating marked improvements over existing methods in both optimization efficiency and task adaptability.
研究测试时使用世界动作模型进行规划,通过比较预测结果选择最优动作,探索如何利用想象的未来指导动作选择。
This work addresses the challenge that diffusion-based policies often fail to discover rare yet effective behavioral patterns under scarce demonstration data, frequently converging to suboptimal solutions or generating infeasible trajectories. To overcome this limitation, the authors propose a novel framework integrating a Feynman–Kac corrector with a learnable guidance potential, which systematically steers the diffusion process toward under-explored feasible regions of the trajectory space. By coupling this guided exploration with sample-based trajectory optimization and an iterative retraining mechanism, the method continuously refines and reuses newly discovered behaviors. Empirical results demonstrate that the approach substantially enhances policy diversity, feasibility, and generalization, consistently uncovering novel and effective strategies across diverse manipulation tasks—surpassing the capabilities of conventional sampling and reinforcement learning methods.