FSPO: Policy-Consistent Risk and Pareto-Feasible Control for Budgeted LLM RL Post-Training

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of risk estimation bias, calibration failure, and multi-resource feasibility certification in adaptive controllers for reinforcement learning-based post-training of large language models. To this end, we propose FSPO, a feedback-state controller that introduces a policy-consistent risk model alongside a decision-conditioned trajectory calibration mechanism. By integrating Bellman objective alignment, cross-fitted trajectory generation, and Pareto resource continuation certificates, FSPO achieves joint policy optimization under budget constraints. Experimental results demonstrate that, within the GRPO resource envelope, FSPO attains held-out and out-of-distribution accuracies of 66.11% and 59.43%, respectively. These outcomes significantly surpass baseline methods while effectively reducing calibration error, highlighting the efficacy of the proposed framework for robust and resource-aware policy optimization.
📝 Abstract
Adaptive LLM reinforcement-learning post-training changes multiple training actuators online, including rollout temperature, group size, clipping, KL regularization, verifier allocation, and update budget. Three coupled issues remain unresolved. A future-risk model trained from behavior trajectories need not estimate the risk induced by the controller that will be deployed; a score calibrated on logged state-action pairs can become miscalibrated after selective action choice; and independent per-resource minimum costs do not in general certify a feasible multi-resource continuation. We introduce FSPO, a feedback-state controller for budgeted LLM RL post-training that addresses these issues jointly. FSPO learns a policy-consistent risk-to-go model whose Bellman target follows the same frozen controller used for future decisions, together with a long-horizon utility model. Decision-conditioned trajectory calibration (DCTC) calibrates risk on cross-fitted trajectories generated by actions selected by provisional controllers. A Pareto resource continuation certificate (PRCC) admits an action only when a non-dominated cumulative reservation remains feasible over the residual horizon. Under a matched GRPO resource envelope, FSPO reaches 66.11% held-out and 59.43% OOD accuracy, compared with 64.47% and 57.03% for PB2, the strongest evaluated adaptive baseline. Three paired training seeds give gains of +2.42 and +3.19 percentage points over the contextual bandit on held-out and OOD evaluation. Under high behavior-deployment mismatch, policy-consistent risk lowers selected-decision ECE from 0.108 to 0.053; DCTC lowers it from 0.039 to 0.022 at matched acceptance; PRCC removes false-feasible admissions on an 18-action catalog ($0.197\rightarrow0.000$); and enabling all three components reduces trajectory failure from 0.181 to 0.083 in a factorial ablation.
Problem

Research questions and friction points this paper is trying to address.

LLM RL post-training
budgeted reinforcement learning
policy-consistent risk
score calibration
multi-resource feasibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

Policy-Consistent Risk
Decision-Conditioned Trajectory Calibration
Pareto Resource Continuation Certificate
Budgeted RL Post-Training
Feedback-State Controller
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Miaobo Hu
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
S
Shuhao Hu
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
Xiaobo Guo
Xiaobo Guo
Dartmouth College
machine learningdeep learningnatural language processingsocia mediapropagantion
Xin Wang
Xin Wang
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences
Biomedical Engineering
Bokun Wang
Bokun Wang
Texas A&M University
Machine LearningArtificial IntelligenceMultimodal Machine Learning
D
Daren Zha
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
J
Jun Xiao
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China