amortised policy learning

Design, train, and evaluate policies that map belief representations (probability distributions or state posteriors) to actions that select measurements, viewpoints, or other acquisition decisions, with the policy amortised across tasks or datasets to enable fast, real-time inference. This includes formulating and optimizing belief-space reinforcement-learning or information-gain objectives, implementing belief-conditioned policy architectures and amortisation strategies, and applying offline/online training procedures (e.g., PPO, offline RL, supervised amortisation) to reduce sample requirements and deploy belief-aware acquisition/viewpoint-selection behaviors.

amortisedpolicylearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.01
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing adaptive data acquisition methods often suffer from inefficient policy learning due to reliance on biased posterior approximations or inadequate exploitation of model representations. This work proposes POLAR, a novel framework that decouples representation learning from policy learning by leveraging pretrained predictive foundation models to encode belief states. POLAR unifies Bayesian experimental design, Bayesian optimization, and active learning within a single coherent framework. By integrating amortized policy learning with task-specific utility functions, the approach substantially reduces the required number of training samples and consistently outperforms state-of-the-art amortized methods across diverse tasks, significantly enhancing the scalability and efficiency of data acquisition.

adaptive data acquisitionamortised inferenceBayesian experimental design

This work addresses the challenge of partial observability in partially observable Markov decision processes (POMDPs) under high-dimensional visual observations, where traditional solvers struggle to scale effectively. The authors propose the Perception-Driven Belief Propagation (PBP) framework, which integrates an image classifier with a classical POMDP solver for the first time. The classifier maps raw images to state probability distributions, which are then incorporated into the belief update process, thereby circumventing direct handling of the high-dimensional observation space. To account for inaccuracies in the perception model, PBP further introduces an uncertainty quantification mechanism. Experimental results demonstrate that PBP outperforms end-to-end deep reinforcement learning approaches across multiple tasks, and that the explicit modeling of perceptual uncertainty significantly enhances robustness against visual disturbances.

belief updatehigh-dimensional observationsperception

Adaptive teachers for amortized samplers

Oct 02, 2024
MK
Minsu Kim
🏛️ KAIST | Recursion | Université de Montréal | POSTECH | University of Edinburgh

To address the low sample efficiency and inadequate multimodal coverage in approximate inference for hard-to-sample unnormalized density distributions, this paper reformulates sampling as a sequential decision-making process and introduces a novel adaptive teacher-guided framework: dynamically identifying high-loss regions where the student sampler underperforms and actively constructing a progressive training curriculum. The method integrates reinforcement learning–based normalizing flows, off-policy training, auxiliary behavioral modeling, and amortized inference. Evaluated across synthetic exploration environments, two diffusion-based sampling tasks, and four biochemical discovery benchmarks, it achieves substantial improvements—averaging +37% in sample efficiency and +52% in mode coverage—while notably enhancing discovery of low-probability, high-reward modes.

Enhancing mode coverage through adaptive training distributionImproving exploration efficiency in reinforcement learning methodsTraining parametric models for intractable distribution approximation

This work addresses the challenge of efficiently sampling combinatorial discrete objects from unnormalized posterior distributions, where existing amortized inference methods based on Markov decision processes suffer from state aliasing, leading to impaired signal propagation and limited expressivity. To overcome these limitations, the authors propose a path-dependent amortized sampling framework that introduces a learnable implicit dynamical system, enabling the policy to model the full generation trajectory rather than relying solely on the current state. This approach effectively relaxes the Markov assumption and allows for conditional modeling over entire trajectories. Theoretically, the framework preserves the scalability of existing discrete amortized algorithms under this extended setting. Empirical results demonstrate that the proposed method significantly accelerates training convergence and enhances exploration in the state space, outperforming current approaches on standard benchmark tasks.

compositional objectsdiscrete samplingMarkov assumption

Latest Papers

What's happening recently
View more

This work addresses the challenge in offline reinforcement learning of jointly quantifying and optimizing epistemic uncertainty arising from limited data coverage and ambiguity in dynamics model identification. The authors propose the Posterior Hybrid Bayesian (PhyB) framework, which treats the dynamics model as a random variable and maintains a posterior belief over it. By constructing convex combinations over subsets of models to approximate the expected objective, PhyB enables efficient policy optimization without resorting to computationally intractable search or strong posterior assumptions. The approach preserves the adaptability of Bayesian RL while providing a metric-agnostic guarantee of monotonic performance improvement. Empirical results demonstrate that PhyB achieves state-of-the-art performance across multiple offline RL benchmarks, confirming its effectiveness and robustness.

Bayesian RLepistemic uncertaintyoffline reinforcement learning

This work integrates the minimization of expected free energy from active inference into an optimizable decision-making framework by formally casting it as a Convex Markov Decision Process (Convex MDP) for the first time, unifying epistemic exploration and pragmatic goals in the space of state marginal distributions. By revealing that expected free energy corresponds to a policy-dependent instrumental reward, the study establishes compatibility with dynamic programming and actor-critic methods. Leveraging convex optimization and mirror descent, it derives policy optimization algorithms applicable to finite-horizon, discounted, and average-reward settings. This approach provides theoretical guarantees for policy improvement and bridges active inference with modern reinforcement learning through a rigorous theoretical foundation.

Active InferenceConvex MDPEpistemic Value

This work addresses the challenge of constructing reliable and compact belief representations that support near-optimal decision-making under perceptual and actuation noise. It introduces a “reliability cell covering” approach that replaces traditional equivalence-class partitioning by defining cells in belief space wherein the optimal action-value function varies by no more than a tolerance ε. The method employs a reliability entropy measure to quantify decision-relevant belief complexity and distinguishes between representational sufficiency and performance limits imposed by noise. By leveraging a fixed observation filtering map, predictive observation laws, and a controlled belief transition kernel, the construction yields an ε-cover under Lipschitz continuity assumptions, applicable to finite POMDPs, linear-Gaussian systems, and particle-filter-based models. The resulting piecewise-constant policies achieve a suboptimality bound of 2ε/(1−γ) and admit analytical or empirical verification across diverse filtering frameworks.

actuation noisebelief-space representationdecision-making under uncertainty

This work addresses the challenge of integrating information seeking with closed-loop control in active inference. It proposes a cognitive-prior variational free energy framework that enables closed-loop planning by jointly modeling the posterior over states and actions, explicitly incorporating the dependence of future actions on anticipated states into the inference process. The key insight is that the advantage of complex reasoning stems from the closed-loop architecture itself rather than tree search per se, a mechanism unified under the notion of cognitive priors. Experiments on the Reactivity Maze benchmark demonstrate that neither purely cognition-driven nor open-loop strategies succeed in completing the task, whereas the proposed method achieves robust goal-directed behavior.

active inferenceclosed-loop controlepistemic priors

This study addresses the challenge of effectively balancing intrinsic motivation-driven exploration and goal-directed decision-making for agents operating in partially observable environments. To this end, it extends the maximum occupancy principle to partially observable settings by introducing a belief-based reasoning mechanism. Furthermore, this work proposes a Bellman reformulation of expected free energy, enabling offline value iteration solutions across the entire belief state space. This methodological advance significantly improves algorithmic tractability under uncertainty. Experimental results demonstrate that the proposed agent dynamically switches between exploration and exploitation based on energy and belief states, thereby overcoming the limitation of conventional active inference strategies that tend to converge on single sources. Ultimately, the approach achieves superior adaptive behavior generation.

Active InferenceBelief StateIntrinsic Motivation

Hot Scholars

ZP

Zhuoyang Pan

University of Pennsylvania
Computer VisionComputer Graphics
MT

Marc Toussaint

Professor of Computer Science, TU Berlin, Germany
RoboticsRobot LearningArtificial Intelligence
XW

Xuhong Wang

Shanghai Artificial Intelligence Laboratory
LLMKnowledge SystemAI Simulation
SG

Stephan Günnemann

Professor of Computer Science, Technical University of Munich
Machine LearningGraphsGraph Neural NetworksRobustness
GL

Guoyu Lu

SUNY Binghamton
RoboticsComputer VisionMachine Learning