OPTS-TTPO: Enhancing Finite-Sample Policy-Gradient Learning with Tree Search

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the gradient bias in finite-sample policy gradient methods caused by overlooking rare high-reward trajectories. To mitigate this, we propose online parallel tree search and tree trajectory optimization. Leveraging a branch aggregation lemma, our approach enhances trajectory coverage while controlling bias under a fixed computational budget. Furthermore, we introduce an online tree sampling mechanism that eliminates the need for action distribution correction, and formally prove that the expected return increases monotonically with the budget under deterministic dynamics. Empirical evaluations demonstrate that our method improves tail returns by 28.6% over PPO in MuJoCo environments and significantly outperforms strong baselines on both Atari games and large language model reasoning tasks.
📝 Abstract
The policy-gradient theorem gives the exact gradient under the current policy, but finite on-policy samples may miss rare high-return trajectories. We study whether tree search improves their coverage within a fixed budget while controlling gradient bias. We introduce On-Policy Parallel Tree Search (OPTS) and Tree Trajectory Policy Optimization (TTPO) using on-policy tree trajectories, which sample new suffixes from the current policy at visited states. This needs no action-distribution correction, although branching changes state visitation. Our Branch Aggregation Lemma shows that branch-weighted tree statistics recover chain expectations when branch choices and weights are fixed before outgoing transitions are sampled. OPTS selects expansion states using estimated performance differences. Under deterministic dynamics, exact values, and max-backup advantages, the induced search policy's expected return improves monotonically with the budget. We bound the gradient bias from adaptive expansion and show that max backup assigns prefix credit to actions leading to better discovered suffixes. Against a finite chain reference, TTPG's measured bias stays near its no-branching level, while NaivePG's bias grows from 0.1251 to 0.4884. At matched budgets, reward- and value-guided OPTS improve correct-answer coverage and majority-vote accuracy over independent sampling. At matched branch counts, OPTS + TTPG gains coverage with a modest bias increase relative to Fixed-branch + TTPG. Under matched interaction or rollout budgets, OPTS-TTPO improves MuJoCo tail returns over PPO by up to 28.6%, achieves a 34-22-1 win-loss-tie record against PPO on Atari-57 under the last-100-log mean-return metric, and improves micro-averaged avg@32 and pass@32 over PPO across all four Qwen3 models.
Problem

Research questions and friction points this paper is trying to address.

policy gradient
finite-sample coverage
tree search
gradient bias
high-return trajectories
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Tree Search
Policy Gradient
Branch Aggregation Lemma
Gradient Bias Control
Trajectory Coverage
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Junyu Lu
Dobot Robotics
S
Shichao Weng
Dobot Robotics
Z
Zhiqiang Wang
Dobot Robotics
H
Haojie Luo
Dobot Robotics
Jingfan Zhang
Jingfan Zhang
Unknown affiliation
NLP
Y
Yuhua Zhou
Zhejiang University
C
Cheng Du
Independent Researcher
Y
Yuzhuo Zhang
Fudan University
X
Xi Li
Dobot Robotics
J
Jinwei Du
Independent Researcher
T
Tiancheng Feng
Dobot Robotics
Chuan Xiao
Chuan Xiao
Associate Professor, Osaka University
Agent-Based ModelingComputer SimulationData PreprocessingData ManagementData Science
Shuyuan Zheng
Shuyuan Zheng
The University of Osaka
Data ValuationData SecurityLegal AI