EmbodiedMind: Adaptive Data Curation and Prefix-Tree Reinforcement Learning for Efficient Embodied Intelligence

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对具身智能训练中的样本利用率低、任务间梯度贡献不平衡及长期规划信用分配问题,提出了一种结合数据筛选与层次策略优化的新训练范式。
📝 Abstract
Training embodied foundation models typically requires massive-scale datasets and extensive computational resources, yet often suffers from three critical limitations: (1) inefficient sample utilization due to low-informative samples; (2) imbalanced gradient contributions across heterogeneous tasks; and (3) severe credit assignment problem in long-horizon planning, where trajectory-level rewards indiscriminately penalize all tokens. To address these issues, we propose an efficient training paradigm that achieves state-of-the-art average performance through strategic data selection and hierarchical policy optimization. Our approach consists of three synergistic stages. First, Rejection Sampling-based Fine-Tuning (RSFT) filters out low-informative samples to establish robust behavioral priors while preventing distributional collapse. Second, Iterative Rejection GRPO (IR-GRPO) employs task-specific queues stratified by difficulty to keep datasets balanced across reinforcement learning iterations, coupled with a hybrid reward mechanism for precise cross-task feedback. Third, to enhance long-horizon task planning, we introduce Trie-GRPO, a novel reinforcement learning algorithm based on action prefix trees, which enables step-level advantage estimation. This resolves the credit assignment problem by isolating intermediate correct decisions from downstream errors, while effectively balancing exploration efficiency and depth compared to conventional search trees. As a result, EmbodiedMind achieves a state-of-the-art average performance of 70.02% across 18 benchmarks, and significantly outperforms other embodied foundation models in long-horizon task planning accuracy. Our project will be released for reproducibility.
Problem

Research questions and friction points this paper is trying to address.

Embodied Intelligence
Data Curation
Reinforcement Learning
Credit Assignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Rejection Sampling-based Fine-Tuning (RSFT)
Iterative Rejection GRPO (IR-GRPO)
Trie-GRPO
Action Prefix Trees
Hierarchical Policy Optimization
🔎 Similar Papers
2024-10-04International Conference on Learning RepresentationsCitations: 0
Feifan Wang
Feifan Wang
Southeast university
Computer visionAffective computing
Z
Zongbing Zhang
ZTE Corporation, Shenzhen, China
Y
Yu Zhang
ZTE Corporation, Shenzhen, China
L
Lingfeng Wang
ZTE Corporation, Shenzhen, China
Yurui Zhu
Yurui Zhu
University of Science and Technology of China
J
Jin Deng
ZTE Corporation, Shenzhen, China
M
Mingliang Zhang
ZTE Corporation, Shenzhen, China
Z
Zhengguang Gao
ZTE Corporation, Shenzhen, China
Y
Yongcheng Wang
ZTE Corporation, Shenzhen, China
J
Jin Xu
ZTE Corporation, Shenzhen, China
R
Ri Yang
ZTE Corporation, Shenzhen, China