ROOT: Discovering Rewards for User-Specified Embodied Behaviors

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of precisely specifying fine-grained behavioral rewards using natural language in embodied intelligence by proposing the ROOT framework. This method reformulates reward design as a continuous search process over observation trees, leveraging video-language models to diagnose policy failure causes and guide reward optimization, thereby overcoming the limitations of traditional approaches that rely on scalar statistics. By integrating reinforcement learning with large language models, ROOT implements an observation-guided tree search algorithm. Experimental results demonstrate that this framework improves behavioral alignment by 16.5% and achieves human preference rates of 51%–63%, significantly outperforming existing baseline methods.
📝 Abstract
Reinforcement learning for embodied control remains constrained by the difficulty of reward specification. Although recent large language model (LLM)-based methods can synthesize reward functions from natural-language descriptions, they often fail to capture subtle behavioral properties that humans care about, such as natural gait, posture, and movement style. This limitation arises because many desired behaviors are easier to recognize visually than to encode in a reward function. We introduce Reward Optimization via Observable Trees (ROOT), a framework for discovering reward functions that align learned policies with user-specified embodied behaviors. Rather than relying solely on scalar training statistics, ROOT casts reward design as an observation-guided search over a persistent experiment tree that stores reward programs, trained policies, and rollout observations, together with behavioral insights distilled by a video-language model, to diagnose behavioral failures and guide subsequent reward refinements. We evaluate ROOT on seven tasks across four embodiments: simulated Hopper, HalfCheetah, Ant, Unitree Go2, and as well as the real-world Unitree Go2. ROOT produces behaviors that better align with user intent than those generated by existing LLM-based reward-generation methods, achieving up to 86.8% locomotion-completeness accuracy and improving Vid-LLM behavioral alignment from 3.56/5 to 4.14/5, a 16.5% improvement over baselines. Human evaluations further support these results, with ROOT preferred in 51-63% of pairwise comparisons.
Problem

Research questions and friction points this paper is trying to address.

reward specification
embodied behaviors
reinforcement learning
behavioral alignment
reward design
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reward Optimization
Embodied AI
Video-Language Model
Experiment Tree
Reinforcement Learning
E
Eren Sadikoglu
Robotics and Autonomous Systems, Arizona State University, Tempe, Arizona, USA
Aditya Taparia
Aditya Taparia
Ph.D. student at Arizona State University
Generative VisionXAIReinforcement LearningDeep Learning
X
Xinyuan Liu
School of Computing and Augmented Intelligence, Arizona State University, Tempe, Arizona, USA
Ransalu Senanayake
Ransalu Senanayake
ASU | Stanford University
Machine LearningRoboticsHealthcare ML