WorldMark: A Unified Benchmark Suite for Interactive Video World Models

πŸ“… 2026-04-23
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the lack of a unified evaluation framework for interactive video generation models, which hinders fair and systematic comparison. To this end, we propose World Model Arenaβ€”a standardized benchmark that establishes consistent WASD-style action mappings, control interfaces, and scene configurations. It features a hierarchical test suite and a modular evaluation toolkit to support multidimensional analysis across visual quality, control alignment, and world consistency. The benchmark encompasses 500 diverse evaluation cases spanning multiple viewpoints, artistic styles, and difficulty levels, and includes a comprehensive comparison of six state-of-the-art image-to-video world models. All data, code, model outputs, and an online battle platform are publicly released to foster reproducible and comparable research in the field.

Technology Category

Computer Vision: Large Vision ModelsHumans and AI: Human-Aware Planning and Behavior PredictionMachine Learning: Large Multimodal Models (LMMs)

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsUser Modeling, Personalization and Recommendation: User modeling and simulation for interactive and conversational systemsResponsible Web: Machine-in-the-loop, human agency and autonomy
πŸ“ Abstract
Interactive video generation models such as Genie, YUME, HY-World, and Matrix-Game are advancing rapidly, yet every model is evaluated on its own benchmark with private scenes and trajectories, making fair cross-model comparison impossible. Existing public benchmarks offer useful metrics such as trajectory error, aesthetic scores, and VLM-based judgments, but none supplies the standardized test conditions -- identical scenes, identical action sequences, and a unified control interface -- needed to make those metrics comparable across models with heterogeneous inputs. We introduce WorldMark, the first benchmark that provides such a common playing field for interactive Image-to-Video world models. WorldMark contributes: (1) a unified action-mapping layer that translates a shared WASD-style action vocabulary into each model's native control format, enabling apples-to-apples comparison across six major models on identical scenes and trajectories; (2) a hierarchical test suite of 500 evaluation cases covering first- and third-person viewpoints, photorealistic and stylized scenes, and three difficulty tiers from Easy to Hard spanning 20-60s; and (3) a modular evaluation toolkit for Visual Quality, Control Alignment, and World Consistency, designed so that researchers can reuse our standardized inputs while plugging in their own metrics as the field evolves. We will release all data, evaluation code, and model outputs to facilitate future research. Beyond offline metrics, we launch World Model Arena (warena.ai), an online platform where anyone can pit leading world models against each other in side-by-side battles and watch the live leaderboard.
Problem

Research questions and friction points this paper is trying to address.

interactive video generation
benchmark
world models
model evaluation
standardized testing
Innovation

Methods, ideas, or system contributions that make the work stand out.

interactive video generation
unified benchmark
action mapping
world models
model evaluation
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
X
Xiaojie Xu
Alaya Studio, Shanda AI Research Tokyo
Z
Zhengyuan Lin
Alaya Studio, Shanda AI Research Tokyo; The University of Tokyo
Kang He
Kang He
Purdue University
Large Language ModelsReasoningAgentic AI
Y
Yukang Feng
Alaya Studio, Shanda AI Research Tokyo; Shanghai Innovation Institute
Xiaofeng Mao
Xiaofeng Mao
Alibaba Group
Computer VisionAdversarial Machine Learning
Y
Yuanyang Yin
Alaya Studio, Shanda AI Research Tokyo; Shanghai Innovation Institute
Kaipeng Zhang
Kaipeng Zhang
Shanghai AI Laboratory
LLMMultimodal LLMsAIGC
Yongtao Ge
Yongtao Ge
The University of Adelaide
Computer VisionHuman DigitizationPerception3D Reconstruction