An Unexpected Robot Policy: Early Evaluations of GPT-6 Astra on RoboDojo and Beyond

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探索了大型语言模型作为机器人操作策略的可行性,无需特定任务微调。GPT-6 Astra在RoboDojo任务中表现出色,但高精度控制仍有限制。
📝 Abstract
Embodied AI systems are often organized into System 1 and System 2. System 1 is typically a pretrained policy that generates actions at high frequency, whereas System 2 is often instantiated as a vision-enabled language model for high-level planning. We ask whether a large language model (LLM) can act as the policy for robot manipulation without task-specific finetuning. We call this setting LLM as policy. We evaluate three LLMs on all 42 RoboDojo tasks and compare their scores with 40 public policies. Astra and GPT-5.5 use the official 50-episode-per-task protocol; DeepSeek-Flash uses 10 episodes per task. GPT-6 Astra achieves 22.48% average success rate and 28.97 Score over 2,100 trials, ranking above every public entry. Yet GPT-5.5 and DeepSeek-Flash reach only 0.88% and 1.92% average success rate with the same post-processing. We find that Astra exhibits a sharply polarized capability profile. It generalizes well to tasks that require semantic understanding but not high-precision control. In contrast, it performs poorly on tasks that require precision, dynamic control, or complex bimanual coordination. In-context experiments show no aggregate benefit from one-shot demonstrations, while selected interaction traces show within-episode corrections under perturbations. Overall, the evaluated LLMs vary substantially in manipulation performance. Astra stands out and provides initial evidence for the potential of a general-purpose manipulation model, although reliable precision and dynamic control remain limitations in the evaluated setting.
Problem

Research questions and friction points this paper is trying to address.

large language model
robot manipulation
task-specific finetuning
semantic understanding
high-precision control
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM as policy
general-purpose manipulation model
semantic understanding
high-precision control
in-context learning
🔎 Similar Papers
No similar papers found.
W
Wenbo Zhang
RoboProbe
K
Kaixuan Wang
RoboDojo, The University of Hong Kong
Y
Yutao Ouyang
RoboProbe, Tsinghua University
X
Xiaoyu Huang
RoboProbe, University of California, Berkeley
L
Liyang Li
RoboProbe
K
Kailun Su
RoboDojo, Tsinghua University
W
Weiyang Jin
The University of Hong Kong
Wenhao Chai
Wenhao Chai
Princeton University
Machine LearningComputer Vision
H
Haotian Liang
The University of Hong Kong
Z
Zhiyang Dou
Massachusetts Institute of Technology
Yue Chen
Yue Chen
Peking University
RoboticsLarge Language Model
T
Tianxing Chen
RoboDojo, The University of Hong Kong