Dynamical Priors as a Training Objective in Reinforcement Learning

📅 2026-04-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Reinforcement learning policies often exhibit abrupt shifts, oscillations, or degenerate stagnation due to a lack of temporal coherence. To address this, this work proposes Dynamic Prior Reinforcement Learning (DP-RL), which introduces an auxiliary loss derived from external state dynamics—without altering the reward function, environment, or policy architecture—to explicitly incorporate a dynamic prior that embodies evidence accumulation and hysteresis effects as part of the training objective. By solely modifying the optimization target, DP-RL enables control over the temporal geometric properties of an agent’s decision trajectory. Implemented within a policy gradient framework, DP-RL is validated in three minimal environments, demonstrating its ability to systematically shape behaviors with structured temporal characteristics in a task-dependent manner, surpassing the capabilities of generic smoothing approaches.

Technology Category

Machine Learning: Reinforcement LearningHumans and AI: Human-Aware Planning and Behavior PredictionSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingResponsible Web: Machine-in-the-loop, human agency and autonomyGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphs
📝 Abstract
Standard reinforcement learning (RL) optimizes policies for reward but imposes few constraints on how decisions evolve over time. As a result, policies may achieve high performance while exhibiting temporally incoherent behavior such as abrupt confidence shifts, oscillations, or degenerate inactivity. We introduce Dynamical Prior Reinforcement Learning (DP-RL), a training framework that augments policy gradient learning with an auxiliary loss derived from external state dynamics that implement evidence accumulation and hysteresis. Without modifying the reward, environment, or policy architecture, this prior shapes the temporal evolution of action probabilities during learning. Across three minimal environments, we show that dynamical priors systematically alter decision trajectories in task-dependent ways, promoting temporally structured behavior that cannot be explained by generic smoothing. These results demonstrate that training objectives alone can control the temporal geometry of decision-making in RL agents.
Problem

Research questions and friction points this paper is trying to address.

reinforcement learning
temporal coherence
dynamical priors
decision trajectories
policy training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamical Priors
Reinforcement Learning
Temporal Coherence
Policy Gradient
Evidence Accumulation
🔎 Similar Papers
No similar papers found.