X2Real: an eXtensive simulation benchmark for real-world generalist policies

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为了解决现有仿真基准中的模拟与现实差距、任务覆盖范围窄及评估不公平问题,X2Real基于Nvidia Isaac Lab-Arena开发了一个包含多样性和公平性的可进化仿真基准。
📝 Abstract
Generalist robot manipulation policies have developed rapidly, yet their reliable evaluation remains challenging due to fundamental flaws in existing simulation benchmarks: prominent sim-to-real gaps, narrow task coverage, and unfair evaluation caused by ambiguous training-test pipelines. Prior works only partially resolve these issues and lack simultaneous faithfulness, diversity, and fairness, while static benchmark designs fail to sustain long-term policy development. We present X2Real, an evolvable simulation benchmark for faithfully evaluating the real-world performance of robotic manipulation policies based on Nvidia Isaac Lab-Arena. Following three core principles (faithfulness, diversity, and fairness), X2Real calibrates simulation visual and physical properties to align with real hardware, achieving a 0.84 linear correlation between simulated and real-robot evaluation results. It features a comprehensive taxonomy with 10 capability dimensions and 44 hierarchical long-horizon tasks, covering basic manipulation skills and advanced capacities such as visual grounding, language understanding, and bimanual control. We further adopt multi-axis domain randomization and strictly disjoint training-evaluation pipelines to mitigate benchmark exploitation and ensure credible evaluation. Powered by a custom physical domain-specific language, the Mana simulation ecosystem supports modular task design and iterative performance analysis, alongside a nearly 300-hour annotated simulation trajectory dataset. X2Real offers a faithful, diverse, and fair evolving evaluation infrastructure, effectively bridging the sim-to-real evaluation gap and supporting the advancement of generalist robotic manipulation policies.
Problem

Research questions and friction points this paper is trying to address.

sim-to-real gap
task coverage
unfair evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

evolvable simulation benchmark
multi-axis domain randomization
disjoint training-evaluation pipelines
🔎 Similar Papers
No similar papers found.
L
Lian Ruan
J
Jade Yang
S
Sherphylan Gao
F
Felix Gao
K
Kyson Liang
G
Galen Liu
L
Ligo Wu
L
Lane Jin
G
Guu Gu
B
Bevan Xie
C
Cloud Yan
Z
Zongzi Yuan
K
Kino Luo
Emma Chen
Emma Chen
Harvard University
AI for Healthcare
S
Shuwen Chen
Y
Yang Ping
M
Miles Guo
R
Rain Sun
K
Kayden Zhang
A
Alex Du
Ruihai Wu
Ruihai Wu
Peking University
computer visionrobotics
L
Liang Hao
Z
Zhaoshuo Li
R
Roy Gan
H
Hao Wang