HappyWorld-Bench

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为评估生成世界的质量和一致性,HappyWorld-Bench通过三个独立的评测轨道和自动化度量方法,对世界模型在交互、修改下的表现进行了全面评价。
📝 Abstract
Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification. We introduce HappyWorld-Bench, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them. Our design is built on a hierarchical capability framework of six world capabilities (W1-W6), from generative construction to unified world modeling, instantiated across three independent evaluation tracks: video world models, spatial world models, and embodied world models. HappyWorld-Bench comprises 1,138 video prompts, 300 spatial scenes, and 254 embodied test cases. Across all three tracks, we build and operate HappyWorld-Arena to organize human A/B comparisons and derive model-level Elo ratings, which complement newly designed automated metrics that capture behavioral correctness. We evaluate 14 video world models, 9 spatial systems, and 8 embodied candidates under this unified framework. Results reveal remaining reliability gaps across all three tracks: video models exhibit reduced consistency during extended rollouts and revisits, spatial models achieve at best 70.14% placement accuracy and 73.33% edit execution, and embodied models struggle to preserve state across multi-step actions and respond precisely to altered action conditions and physical rules. These findings highlight the need to evaluate world models not only by visual quality, but also by state consistency and the correctness of their responses to actions and interventions.
Problem

Research questions and friction points this paper is trying to address.

world models
reliability
consistency
interaction
evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

HappyWorld-Bench
world models evaluation
hierarchical capability framework
behavioral correctness
Elo ratings
🔎 Similar Papers
No similar papers found.
Z
Zhiqi Bai
Alibaba Token Hub, Alibaba Group
J
Junai Cai
Alibaba Token Hub, Alibaba Group
Y
Yixin Chen
State Key Laboratory of General Artificial Intelligence, BIGAI
J
Jingrun Du
Alibaba Token Hub, Alibaba Group
T
Tao Feng
Alibaba Token Hub, Alibaba Group
W
Wei Gong
Alibaba Token Hub, Alibaba Group
Siyuan Huang
Siyuan Huang
Beijing Institute for General Artificial Intelligence (BIGAI)
Embodied AI3D VisionRobotics3D Scene Understanding
X
Xiao Lin
Alibaba Token Hub, Alibaba Group
J
Jiaheng Liu
Nanjing University
J
Jun Luo
State Key Laboratory of General Artificial Intelligence, BIGAI
Y
Yongzhe Lyu
State Key Laboratory of General Artificial Intelligence, BIGAI
Liya Ma
Liya Ma
University of Malaya
RF-MEMSPrintable electronicsMicroelectronics
Z
Zenan Meng
Alibaba Token Hub, Alibaba Group
L
Lin Qu
Alibaba Token Hub, Alibaba Group
W
Wenbo Su
Alibaba Token Hub, Alibaba Group
Jiaming Wang
Jiaming Wang
National University of Singapore
Generative AIRobotics
Qinghe Wang
Qinghe Wang
Dalian University of Technology
Video/Image GenerationDiffusion ModelsGANs
S
Shaofei Wang
State Key Laboratory of General Artificial Intelligence, BIGAI
Y
Yanghai Wang
Nanjing University
Zequn Wang
Zequn Wang
UESTC
reliability-based design
Z
Ziming Wang
Alibaba Token Hub, Alibaba Group
H
Hu Wei
Alibaba Token Hub, Alibaba Group
Jiangtao Wu
Jiangtao Wu
PhD Student of Solid Mechanics,Georgia Institute of Technology
Solid mechanics3D printingShape memory polymerMolecular dynamicsDensity functional theory
R
Ruiqi Wu
Alibaba Token Hub, Alibaba Group
Jiaxin Xie
Jiaxin Xie
HKUST