SIPO: Selective-Inference Policy Optimization for Tree-Structured Agentic RL

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the statistical asymmetry and selection history bias arising from adaptive expansion in tree-structured reinforcement learning. To mitigate these issues, we propose Selective Reasoning Policy Optimization, a method incorporating scale-invariant branching criteria, exchangeable sampling, and order-statistic correction mechanisms. These components effectively eliminate selection bias in credit estimation, ensuring that leaf-node budget allocation remains consistent with the optimization objective while enhancing fairness in tree search. Extensive experiments conducted on the Qwen model series across seven question-answering benchmarks demonstrate that our approach achieves superior average performance, yielding significant improvements in both single-hop and multi-hop question-answering accuracy.
📝 Abstract
Tree-structured reinforcement learning trains search agents by comparing alternative continuations and propagating terminal rewards to intermediate decisions. Adaptive expansion, however, creates a statistical asymmetry: an incumbent is selected using its own generation statistic, whereas fresh siblings are sampled after selection. When that statistic is associated with return, branch values can reflect selection history as well as continuation quality, even for a shared parent. We propose Selective-Inference Policy Optimization (\SIPO{}), which incorporates this distinction into tree-based credit estimation. Its scale-free branch criterion keeps generation scores and sibling penalties on a consistent relative scale; exchangeable branching supplies multiple fresh continuations from each selected parent; and order-statistic correction adjusts retained incumbent values using selection rank and the estimated score--outcome association. These mechanisms preserve the leaf budget and the host policy optimisation objective. Across seven QA benchmarks using Qwen3-4B, Qwen3-8B, and Qwen2.5-7B, \SIPO{} achieves the highest reported multi-hop and single-hop averages among the compared methods. On Qwen3-8B, it improves these averages over AT\textsuperscript{2}PO by $1.31$ and $1.07$ percentage points, respectively, and ranks first on six of seven benchmarks. Component ablations evaluate the individual and combined changes, while early-training paired diagnostics show a selected--fresh value gap alongside a near-zero fresh--fresh reference. Together, these results support accounting for selection history when constructing and evaluating search-agent rollouts. Our code is available at https://github.com/Zenghuang-Fu/SIPO
Problem

Research questions and friction points this paper is trying to address.

Tree-structured reinforcement learning
Statistical asymmetry
Selection bias
Credit estimation
Agentic RL
Innovation

Methods, ideas, or system contributions that make the work stand out.

Selective-Inference Policy Optimization
Tree-Structured Reinforcement Learning
Order-Statistic Correction
Scale-Free Branch Criterion
Agentic RL
🔎 Similar Papers
Z
Zenghuang Fu
University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences
N
Ningqi Chen
The University of Hong Kong
M
Mingda Jia
University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences
X
Xiaofeng Han
University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences
Zhaoyang Li
Zhaoyang Li
Ph.D student, University of Science and Technology of China
Computer Vision
Q
Qiuyuan Ai
Peking University
Z
Zelong Zheng
University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences
H
Haoyu Wu
Mininglamp Technology
Tianyu Fu
Tianyu Fu
Ph.D at Tsinghua University
efficient AILLMsparse computation
Chenxu Zhao
Chenxu Zhao
Research Director, Mininglamp Technology.
Multi-modal Large Language ModelGenerative AIMeta-learningComputer Vision
Minghui Wu
Minghui Wu
Zhejiang University City College
Mobile ComputingBig DataMachine LearningSoftware Engineering
Guannan He
Guannan He
Peking University
Energy SystemMobilityEnergy StorageOptimization
Changwei Wang
Changwei Wang
Shandong Computer Science Center
Multimodal LearningEmbodied AIEdge Intelligent ComputingAI for HealthcareSafety Alignment