distributional action modeling

Designs and implements models that represent and learn probability distributions over actions or policies, capturing multimodal, diverse candidate actions and enabling conditional sampling based on observed context. Builds training and inference procedures to generate, score, and sample from these action distributions while avoiding averaging effects that produce suboptimal actions.

distributionalactionmodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.03
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$213K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

CCDP: Composition of Conditional Diffusion Policies with Guided Sampling

Mar 19, 2025
AR
Amirreza Razmjoo
🏛️ Ecole Polytechnique FedExceptionale de Lausanne (EPFL) | Idiap Research Institute | Honda Research Institute Europe GmbH

This work addresses the inefficiency in imitation learning inference caused by repeated sampling of failed actions. We propose a recovery-oriented action generation framework based on conditional diffusion models. Methodologically, we pioneer the decomposition of long-horizon failure recovery into composable sub-policies and introduce a guided resampling mechanism that dynamically corrects the sampling distribution solely using successful demonstration data—without requiring additional exploration or high-level controllers. Our core contributions are: (1) policy decomposition modeling via diffusion models, and (2) distribution-guided resampling grounded in successful trajectories. Evaluated on tasks including door opening in unknown directions, object manipulation, and button searching, our approach significantly improves action success rates and execution efficiency while demonstrating strong robustness to varying numbers of prior failures.

Decomposes primary problem into manageable sub-problemsEnhances sampling strategy to avoid unsuccessful actionsInfers recovery actions using successful demonstration data

PDPP: Projected Diffusion for Procedure Planning in Instructional Videos

Mar 26, 2023
HW
Hanlin Wang
🏛️ Nanjing University | MYbank | Ant Group | Shanghai AI Laboratory

To address the reliance of procedural action planning in instructional videos on frame-level annotations or speech commands—rendering methods susceptible to error accumulation—this paper proposes an end-to-end diffusion model supervised solely by start/end-frame visual observations and task-level labels. The method frames action sequence generation as a conditional distribution fitting problem without intermediate supervision. Its core contributions include: (i) the first formulation of procedural action generation as an unsupervised latent sequence modeling task under coarse-grained visual and semantic constraints; and (ii) a differentiable projection-guidance mechanism that injects both visual and semantic priors seamlessly during both training and sampling, drastically reducing dependence on fine-grained annotations. Built upon a U-Net backbone, the model fuses start/end-frame visual embeddings with an efficient, differentiable projection operator. Evaluated on three diverse-scale instructional video datasets, it achieves state-of-the-art performance across multiple metrics, outperforming prior approaches while requiring no intermediate action annotations.

Automatic Action PlanningError AccumulationTeaching Videos

Adding Conditional Control to Diffusion Models with Reinforcement Learning

Jun 17, 2024
YZ
Yulai Zhao
🏛️ Princeton University | Genentech | University of California, Berkeley

This work addresses the challenge of fine-grained conditional control in pretrained diffusion models. We propose CTRL, a framework that formulates conditional generation as a reinforcement learning (RL) policy optimization problem. Specifically, CTRL employs a PPO variant with KL regularization to jointly optimize a policy network using offline annotated data; the composite reward comprises classifier outputs and KL divergence from the target conditional distribution. To our knowledge, this is the first approach to cast conditional diffusion generation as end-to-end RL policy learning—eliminating the need for intermediate state classifiers, thereby significantly reducing data requirements and training complexity. Moreover, CTRL enables soft optimal conditional sampling and enjoys theoretical convergence guarantees to the desired conditional distribution. Experiments on image generation demonstrate that CTRL outperforms classifier guidance and classifier-free guidance (CFG) using fewer labeled examples, achieving superior trade-offs between control fidelity and sample diversity.

Enhancing control in diffusion modelsImproving sample efficiency and dataset constructionUsing reinforcement learning for fine-tuning

Diverse Projection Ensembles for Distributional Reinforcement Learning

Jun 12, 2023
MA
Moritz A. Zanger
🏛️ Delft University of Technology

In distributed reinforcement learning, parameterized distribution approximations of the true return distribution introduce substantial inductive bias, degrading generalization and undermining the reliability of uncertainty estimation. To address this, we propose the Diverse Projection Ensemble (DPE), a framework that integrates multiple Wasserstein-distance-compatible projection operators with parameterized distribution representations. We theoretically characterize how projection bias fundamentally affects generalization. Furthermore, DPE couples ensemble disagreement with an exploration reward derived from the 1-Wasserstein distance, yielding an uncertainty-aware deep exploration mechanism. Evaluated on the Behavior Suite and VizDoom benchmarks, DPE significantly outperforms state-of-the-art methods—particularly excelling in directed exploration tasks. These results empirically validate that diversity across projections is critical for robust exploration and reliable uncertainty estimation.

Distributional RL aims to learn return distributions, not just expected values.Diverse projection ensembles improve uncertainty estimation and exploration performance.Projection in distributional RL introduces bias, affecting model generalization.

Bayesian Off-Policy Evaluation and Learning for Large Action Spaces

Feb 22, 2024
IA
Imad Aouali
🏛️ Criteo AI Lab | CREST | ENSAE | IP Paris

To address low sample efficiency in offline policy evaluation (OPE) and offline policy learning (OPL) under large action spaces, this paper proposes sDM—a unified Bayesian framework that introduces structured prior modeling to OPE/OPL for the first time, explicitly capturing inter-action correlations. Methodologically, sDM designs a correlation-aware Bayesian metric, replacing conventional worst-case analysis to enable instance-averaged performance assessment; it further integrates online Bayesian multi-armed bandit heuristics to jointly optimize OPE and OPL. Theoretically, we prove that modeling action correlations significantly reduces estimation variance. Empirically, sDM achieves superior evaluation accuracy and policy performance over state-of-the-art baselines across multiple benchmark tasks, while maintaining linear time complexity.

Enhances off-policy evaluation in large action spacesIntroduces Bayesian metrics for average performance assessmentLeverages action correlations for efficient learning

Latest Papers

What's happening recently
View more

Current single-pass inference paradigms constrain the performance of non-deterministic generative models in robotic manipulation. To address this limitation, this work proposes TapSampling—a plug-and-play, inference-time sampling framework that enables policy-agnostic execution refinement by efficiently exploring the action latent space and incorporating a semantically interpretable task-progress prediction verifier. Built upon an action variational autoencoder (Action-VAE), TapSampling is compatible with both diffusion and autoregressive models and requires no fine-tuning. Experiments demonstrate that it significantly enhances task success rates and robustness across diverse general-purpose policies in both simulated and real-world environments.

action selectionembodied controlinference-time sampling

This work addresses a fundamental limitation in conventional Bayesian experimental design, which relies on prior-to-posterior uncertainty reduction and yields an intractable objective that is doubly hard to evaluate and poorly aligned with downstream tasks. By reframing the problem through decision theory, the authors formulate it as optimizing the expected future loss (EFL) of downstream actions, thereby reducing the objective to a singly intractable form that obviates explicit posterior or marginal likelihood computation. They introduce a stochastic gradient method that jointly optimizes both the experimental design and the action policy, requiring only samples from the joint parameter–data model and evaluations of the loss function. This approach naturally accommodates implicit modeling and task-specific customization, demonstrating marked improvements over existing methods in both optimization efficiency and task adaptability.

Bayesian experimental designdoubly intractable objectivesdownstream tasks

This work addresses the challenge that diffusion-based policies often fail to discover rare yet effective behavioral patterns under scarce demonstration data, frequently converging to suboptimal solutions or generating infeasible trajectories. To overcome this limitation, the authors propose a novel framework integrating a Feynman–Kac corrector with a learnable guidance potential, which systematically steers the diffusion process toward under-explored feasible regions of the trajectory space. By coupling this guided exploration with sample-based trajectory optimization and an iterative retraining mechanism, the method continuously refines and reuses newly discovered behaviors. Empirical results demonstrate that the approach substantially enhances policy diversity, feasibility, and generalization, consistently uncovering novel and effective strategies across diverse manipulation tasks—surpassing the capabilities of conventional sampling and reinforcement learning methods.

behavior discoverydiffusion policiesdiverse behaviors

Hot Scholars

MN

Mehwish Nasim

Senior Lecturer UWA Australia, Network Analysis & Social Influence Modelling (NASIM)Lab
Information WarfareNetwork ScienceNLPComplex Systems
AD

Amitava Datta

Professor of Computer Science, University of Western Australia
parallel and distributed computingwireless networksbioinformaticsdata mining
NF

Nicolas Fay

Associate Professor, University of Western Australia
Social InfluenceCollective BehaviourCultural EvolutionLanguage Evolution