FLAC: Maximum Entropy RL via Kinetic Energy Regularized Bridge Matching

📅 2026-02-13
📈 Citations: 0
✨ Influential: 0
📄 PDF

Technology Category

Machine Learning: Reinforcement LearningSearch and Optimization: Learning to SearchReasoning under Uncertainty: Stochastic Optimization

Application Category

Search and Retrieval-Augmented AI: Agentic searchEconomics, Online Markets and Human Computation: Economic ramifications for generative AI infrastructure and applicationsResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
Iterative generative policies, such as diffusion models and flow matching, offer superior expressivity for continuous control but complicate Maximum Entropy Reinforcement Learning because their action log-densities are not directly accessible. To address this, we propose Field Least-Energy Actor-Critic (FLAC), a likelihood-free framework that regulates policy stochasticity by penalizing the kinetic energy of the velocity field. Our key insight is to formulate policy optimization as a Generalized Schr\"odinger Bridge (GSB) problem relative to a high-entropy reference process (e.g., uniform). Under this view, the maximum-entropy principle emerges naturally as staying close to a high-entropy reference while optimizing return, without requiring explicit action densities. In this framework, kinetic energy serves as a physically grounded proxy for divergence from the reference: minimizing path-space energy bounds the deviation of the induced terminal action distribution. Building on this view, we derive an energy-regularized policy iteration scheme and a practical off-policy algorithm that automatically tunes the kinetic energy via a Lagrangian dual mechanism. Empirically, FLAC achieves superior or comparable performance on high-dimensional benchmarks relative to strong baselines, while avoiding explicit density estimation.
Problem

Research questions and friction points this paper is trying to address.

Maximum Entropy Reinforcement Learning
Iterative generative policies
Action log-densities
Continuous control
Likelihood-free
Innovation

Methods, ideas, or system contributions that make the work stand out.

Maximum Entropy Reinforcement Learning
Kinetic Energy Regularization
Generalized Schrödinger Bridge
Likelihood-Free Policy Optimization
Flow Matching
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Lei Lv
Shanghai Research Institute for Intelligent Autonomous Systems
Yunfei Li
Yunfei Li
ByteDance Seed
Reinforcement LearningRobotics
Y
Yu Luo
Tsinghua University
F
Fuchun Sun
Tsinghua University
Xiao Ma
Xiao Ma
ByteDance Seed
Robot LearningReinforcement LearningRobotics