CARE-VI: Conservative Adaptive Reliability Estimation for Value Improvement in Off-Policy Actor-Critic Learning

📅 2026-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出CARE-VI框架,通过CARS、SEVA和DARE方法解决离线策略学习中目标可靠性问题,提高价值估计的准确性和稳定性。
📝 Abstract
Reliable temporal-difference targets are central to off-policy actor-critic learning. Direct value improvement refines the next-state target with alternative actions, but the reliability of this refinement depends on how candidate actions are ranked, reviewed, and weighted. Noisy rankings may force premature candidate commitment, reusing selection scores may bias target valuation, and fixed enhancement weights may amplify weak evidence. To address these risks, we develop Conservative Adaptive Ranking and Screening (CARS), which retains an ordered candidate prefix within a preset budget and narrows it only when the observed boundary gap exceeds a disagreement-scaled uncertainty radius. Selector-Evaluator Value Assessment (SEVA) uses selector critics to order candidates and a separately parameterized evaluator critic to review the selected value, then caps the reviewed value at the selector reference. Dynamic Adaptive Risk-aware Enhancement (DARE) then regulates each residual correction using candidate reliability, the gap between selector and evaluator signals, and a finite stage factor. Together, CARS, SEVA, and DARE form CARE-VI, an evidence-regulated target construction framework that preserves the backbone interfaces for critic regression and actor updates. The analysis bounds the CARS boundary error, the SEVA selected-value overestimation, and the one-sided deviation of the DARE residual displacement from its population counterpart, and establishes fixed-policy recovery after the finite-stage perturbation ends. Experiments with SAC, TD3, and TD7 on four MuJoCo tasks show that CARE-VI achieves the highest mean return in all twelve settings. Grouped ablations and scalar diagnostics support the roles of the three components in improving target reliability.
Problem

Research questions and friction points this paper is trying to address.

off-policy actor-critic learning
reliable temporal-difference targets
value improvement
Innovation

Methods, ideas, or system contributions that make the work stand out.

Conservative Adaptive Ranking and Screening (CARS)
Selector-Evaluator Value Assessment (SEVA)
Dynamic Adaptive Risk-aware Enhancement (DARE)
💼 Related Jobs
No related jobs found.
X
Xiang Zou
School of Mathematics, Harbin Institute of Technology, Harbin 150001, China
Shengzhu Shi
Shengzhu Shi
Lecturer, Harbin Institute of Technology
uncertainty quantificationdeep learningimage processingoptimal controlpreconditioning
Junqi Gao
Junqi Gao
Shanghai AI Lab, 哈尔滨工业大学
Deep LearningGenerative ModelsContinual Learning
Z
Zhichang Guo
School of Mathematics, Harbin Institute of Technology, Harbin 150001, China