Data-Driven Exploration for a Class of Continuous-Time Linear--Quadratic Reinforcement Learning Problems

📅 2025-06-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the model-free reinforcement learning problem in continuous-time stochastic linear-quadratic (LQ) control, where the system volatility depends on both state and control, and no running control reward is present. To overcome inefficiencies and hyperparameter sensitivity inherent in conventional fixed-exploration strategies, we propose an adaptive exploration mechanism within an actor-critic framework: the critic dynamically adjusts the entropy regularization strength, while the actor real-time modulates policy variance, enabling online balancing of exploration and exploitation. The method requires no prior model knowledge and—first among continuous-time LQ settings—achieves a sublinear regret bound matching the best known optimal rate. It breaks the limitations of static exploration scheduling, substantially reducing hyperparameter tuning effort and accelerating convergence. Numerical experiments demonstrate superior regret performance and faster convergence compared to both non-adaptive model-free and model-based baselines.

Technology Category

Search and Optimization: Learning to SearchMachine Learning: Online Learning & BanditsReasoning under Uncertainty: Stochastic Optimization

Application Category

Economics, Online Markets and Human Computation: LLM based quality controls for crowd workSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
We study reinforcement learning (RL) for the same class of continuous-time stochastic linear--quadratic (LQ) control problems as in cite{huang2024sublinear}, where volatilities depend on both states and controls while states are scalar-valued and running control rewards are absent. We propose a model-free, data-driven exploration mechanism that adaptively adjusts entropy regularization by the critic and policy variance by the actor. Unlike the constant or deterministic exploration schedules employed in cite{huang2024sublinear}, which require extensive tuning for implementations and ignore learning progresses during iterations, our adaptive exploratory approach boosts learning efficiency with minimal tuning. Despite its flexibility, our method achieves a sublinear regret bound that matches the best-known model-free results for this class of LQ problems, which were previously derived only with fixed exploration schedules. Numerical experiments demonstrate that adaptive explorations accelerate convergence and improve regret performance compared to the non-adaptive model-free and model-based counterparts.
Problem

Research questions and friction points this paper is trying to address.

Develop adaptive exploration for continuous-time LQ reinforcement learning
Improve learning efficiency with minimal tuning in RL
Achieve sublinear regret bound matching best-known results
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive entropy regularization by critic
Dynamic policy variance adjustment by actor
Model-free data-driven exploration mechanism
🔎 Similar Papers
No similar papers found.
Y
Yilie Huang
Department of Industrial Engineering and Operations Research, Columbia University, New York, NY 10027, USA
X
Xun Yu Zhou
Department of Industrial Engineering and Operations Research & Data Science Institute, Columbia University, New York, NY 10027, USA