T2SPO: Trajectory-to-Step Policy Optimization for Agentic Reinforcement Learning

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of sparse terminal rewards in reinforcement learning for LLM agents, which leaves intermediate steps unsupervised. To tackle this issue, the authors propose T2SPO, a method that leverages historical trajectories to derive remaining-distance targets. T2SPO innovatively incorporates a pretrained TabPFN regressor to estimate state values through in-context dynamic updating rather than gradient-based training, thereby providing step-level auxiliary credit assignment for policy optimization. Experimental evaluations on the ALFWorld and WebShop benchmarks demonstrate that T2SPO significantly improves task success rates for both 1.5B and 7B parameter models, consistently outperforming the GRPO baseline.
📝 Abstract
Reinforcement learning enables large language model (LLM) agents to learn multi-step behaviors through interaction with their environments. However, rewards in many interactive tasks reflect only the final outcome, providing limited guidance on which intermediate decisions advance the task. Successful training trajectories contain intermediate states that can provide supervision for subsequent interactions. We introduce Trajectory-to-Step Policy Optimization (T2SPO), a method that uses past interaction trajectories to provide step-level feedback for policy learning. T2SPO derives remaining-distance targets from successful trajectories and pairs them with representations of the states visited along the way. Conditioned on these examples, a pretrained TabPFN regressor estimates the remaining distance to success at each state of a new rollout. Changes in this distance estimate across consecutive states yield auxiliary credit for agent steps alongside task-level supervision. As training proceeds, newly completed trajectories refresh the estimator's context, incorporating new experience without updating its parameters. Experiments with 1.5B and 7B language models on ALFWorld and WebShop show that T2SPO consistently improves overall task success over GRPO.
Problem

Research questions and friction points this paper is trying to address.

reinforcement learning
LLM agents
sparse reward
credit assignment
multi-step decision making
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Reinforcement Learning
Step-level Credit Assignment
TabPFN
Trajectory-to-Step Policy Optimization
Remaining Distance Estimation
B
Bo-Wen Zhang
State Key Laboratory of Novel Software Technology, Nanjing University; School of Intelligence Science and Technology, Nanjing University; ByteDance
Junwei He
Junwei He
Institute of Computing Technology, Chinese Academy of Sciences
LLM ReasoningGraph Learning
M
Maoqi Liu
ByteDance
F
Feiran Li
ByteDance
S
Song-Lin Lv
State Key Laboratory of Novel Software Technology, Nanjing University; School of Intelligence Science and Technology, Nanjing University
Wentao Ma
Wentao Ma
Alibaba DAMO Academy
Dialog SystemLarge Language ModelQuestion Answering
R
Rongyi Lin
ByteDance
S
Shuhan Zhong
ByteDance
Lan-Zhe Guo
Lan-Zhe Guo
LAMDA Group, Nanjing University
Machine Learning