RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文针对具身智能体指令跟随能力虚高问题,提出RoboFollow基准,通过增加场景熵、多层次诊断协议及控制混淆因素等方法来准确评估真实指令跟随性能。
📝 Abstract
Modern embodied agents achieve impressive success rates, yet their actual instruction-following ability is far weaker than these numbers suggest. We trace this illusion to a structural property we term low scene entropy: when a visual scene admits only one valid task, language becomes redundant and a policy can score highly while barely using it. We introduce RoboFollow, a diagnostic benchmark with three principles: (1) High Scene Entropy: each training scene supports multiple kinematically distinct task branches, making vision alone insufficient and forcing reliance on language. (2) Hierarchical Diagnostic Protocol: a four-level protocol (L0--L3) progressively perturbs visual layout and semantics, probing whether equivalent instructions yield consistent behavior and distinct ones yield discriminable behavior across spatial relations, attributes, trajectory constraints, and logic. (3) Confound-Controlled Diagnosis: we simplify interaction objects, restrict actions to the trained repertoire and report stage-wise Intent and Execution scores, isolating comprehension from motor execution. Evaluation of nine VLA and WAM policies shows that strong L0 performance, where attained, does not reliably transfer to L1--L3 under our fine-tuning setup. Representative mitigations, including stronger VLM backbones, QA co-training, LangForce, and Classifier-Free Guidance, all fail to close this gap. RoboFollow exposes genuine instruction following as a critical, overlooked bottleneck. Code and dataset are available at https://github.com/AutoLab-SAI-SJTU/RoboFollow and https://huggingface.co/datasets/AutoLab-SJTU/robofollow-data.
Problem

Research questions and friction points this paper is trying to address.

Instruction Following
Scene Entropy
Embodied Agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

High Scene Entropy
Hierarchical Diagnostic Protocol
Confound-Controlled Diagnosis
🔎 Similar Papers
2024-10-04International Conference on Learning RepresentationsCitations: 0
C
Chang Guo
AutoLab, School of Artificial Intelligence, Shanghai Jiao Tong University
Y
Yukun Xie
AutoLab, School of Artificial Intelligence, Shanghai Jiao Tong University
B
Bohan Tan
AutoLab, School of Artificial Intelligence, Shanghai Jiao Tong University
Z
Zheng Chang
AutoLab, School of Artificial Intelligence, Shanghai Jiao Tong University
Z
Zhaokai Yin
AutoLab, School of Artificial Intelligence, Shanghai Jiao Tong University
Qianli Ma
Qianli Ma
Shanghai Jiao Tong University
Deep LearningGenerative AILLMsMLLMs
Y
Yingqiao Wang
AutoLab, School of Artificial Intelligence, Shanghai Jiao Tong University
C
Chao Liang
Research Lab, Anyverse Dynamics
Zhipeng Zhang
Zhipeng Zhang
School of Artificial Intelligence, Shanghai Jiao Tong University
Computer Vision,Object Tracking and Segmentation