SEEA-R1: Tree-Structured Reinforcement Fine-Tuning for Self-Evolving Embodied Agents

📅 2025-06-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Embodied intelligence’s self-evolution is hindered by two key bottlenecks: sparse rewards in multi-step reasoning and poor generalizability of handcrafted reward functions. To address these, we propose Tree-Structured Reinforcement Fine-Tuning (Tree-RFT): (1) Tree-GRPO, the first algorithm integrating Monte Carlo Tree Search (MCTS) into Group Relative Policy Optimization (GRPO), automatically generating dense intermediate rewards; (2) a Multimodal Generative Reward Model (MGRM) enabling cross-task, human-free reward modeling. Tree-RFT eliminates reliance on environment-provided rewards and supports end-to-end autonomous optimization for long-horizon, multimodal embodied tasks. On ALFWorld, it achieves state-of-the-art performance—85.07% on text-based tasks and 36.19% on multimodal tasks—outperforming GPT-4o and all open-source baselines. Remarkably, it retains 80.3% of its performance even under zero environment reward, demonstrating robust reward autonomy.

Technology Category

Intelligent Robots: Embodied AIMachine Learning: Reinforcement LearningMultiagent Systems: Multiagent Learning

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGResponsible Web: Machine-in-the-loop, human agency and autonomyEconomics, Online Markets and Human Computation: LLM based quality controls for crowd work
📝 Abstract
Self-evolution, the ability of agents to autonomously improve their reasoning and behavior, is essential for the embodied domain with long-horizon, real-world tasks. Despite current advancements in reinforcement fine-tuning (RFT) showing strong performance in enhancing reasoning in LLMs, its potential to enable self-evolving embodied intelligence with multi-modal interactions remains largely unexplored. Specifically, reinforcement fine-tuning faces two fundamental obstacles in embodied settings: (i) the lack of accessible intermediate rewards in multi-step reasoning tasks limits effective learning signals, and (ii) reliance on hand-crafted reward functions restricts generalization to novel tasks and environments. To address these challenges, we present Self-Evolving Embodied Agents-R1, SEEA-R1, the first RFT framework designed for enabling the self-evolving capabilities of embodied agents. Specifically, to convert sparse delayed rewards into denser intermediate signals that improve multi-step reasoning, we propose Tree-based group relative policy optimization (Tree-GRPO), which integrates Monte Carlo Tree Search into GRPO. To generalize reward estimation across tasks and scenes, supporting autonomous adaptation and reward-driven self-evolution, we further introduce Multi-modal Generative Reward Model (MGRM). To holistically evaluate the effectiveness of SEEA-R1, we evaluate on the ALFWorld benchmark, surpassing state-of-the-art methods with scores of 85.07% (textual) and 36.19% (multi-modal), outperforming prior models including GPT-4o. SEEA-R1 also achieves scores of 80.3% without environmental reward, surpassing all open-source baselines and highlighting its scalability as a self-evolving embodied agent. Additional experiments and qualitative analysis further support the potential of SEEA-R1 for future research in scalable embodied intelligence.
Problem

Research questions and friction points this paper is trying to address.

Enables self-evolving embodied agents with multi-modal interactions
Addresses sparse rewards in multi-step reasoning tasks
Generalizes reward estimation across novel tasks and environments
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tree-GRPO integrates Monte Carlo Tree Search
Multi-modal Generative Reward Model generalizes rewards
SEEA-R1 framework enables self-evolving embodied agents
🔎 Similar Papers
W
Wanxin Tian
Beijing Innovation Center of Humanoid Robotics
S
Shijie Zhang
Beijing Innovation Center of Humanoid Robotics
K
Kevin Zhang
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
Xiaowei Chi
Xiaowei Chi
The Hong Kong University of Science and Technology
Multimodal GenerationRoboticsComputer Vision
Yulin Luo
Yulin Luo
Peking University
Data-centric AILLMVLMEmbodied AI
J
Junyu Lu
Beijing Innovation Center of Humanoid Robotics
C
Chunkai Fan
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
Q
Qiang Zhou
Beijing Innovation Center of Humanoid Robotics
Y
Yiming Zhao
Beijing Innovation Center of Humanoid Robotics
N
Ning Liu
Beijing Innovation Center of Humanoid Robotics
Siyu Lin
Siyu Lin
Beijing Jiaotong University
Wireless communicaions
Z
Zhiyuan Qin
Beijing Innovation Center of Humanoid Robotics
X
Xiaozhu Ju
Beijing Innovation Center of Humanoid Robotics
Shanghang Zhang
Shanghang Zhang
Peking University
Embodied AIFoundation Models
J
Jian Tang
Beijing Innovation Center of Humanoid Robotics