PlaySuite: A Large-Scale Benchmark for Interactive Visual Intelligence

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of long-term interaction evaluation for multimodal models in dynamic environments by constructing an interactive visual intelligence benchmark based on 5,000 open-source games. Methodologically, it proposes a unified evaluation framework for heterogeneous games alongside a closed-loop interaction system, and introduces a Video-LLM-as-a-Judge mechanism that effectively mitigates data leakage while quantifying game milestones. The research reveals a significant perception-action gap in current models, demonstrating their persistent limitations in spatial localization, action execution, and self-correction. Ultimately, this work establishes a novel paradigm for evaluating embodied intelligence.
📝 Abstract
Recent advances in multimodal foundation models yield strong performance on static perception and reasoning benchmarks, yet such evaluations largely overlook a central aspect of intelligence: acting competently in dynamic environments over extended time horizons. We introduce PlaySuite, a large-scale benchmark for evaluating interactive visual intelligence across more than 5K open-source video games curated from PyWeek and itch.io. Spanning diverse genres and engines, including Pygame, HTML5, Godot, and Unity, these independent games are largely out-of-distribution for current models, reducing the likelihood that success can be achieved by retrieving memorized walkthroughs or web-scale training artifacts. To enable scalable evaluation across heterogeneous titles, we develop a unified closed-loop interaction framework optimized for HPC clusters alongside a Video-LLM-as-a-judge protocol that maps observable gameplay milestones to standardized progress levels. We evaluate fourteen recent open models spanning vision-language models, computer-use agents, and vision-language-action models. Our results yield strong evidence of a perception-action gap: despite strong reasoning capabilities, current models struggle to make sustained progress and exhibit recurring failures in spatial grounding, action execution, and self-correction. PlaySuite provides a reproducible and extensible testbed for measuring progress from visual perception to goal-directed interaction, and a foundation for developing models that can act, adapt, and generalize in dynamic visual environments.
Problem

Research questions and friction points this paper is trying to address.

interactive visual intelligence
multimodal foundation models
perception-action gap
dynamic environments
benchmark evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Interactive Visual Intelligence
Large-Scale Benchmark
Closed-loop Interaction Framework
Video-LLM-as-a-Judge
Perception-Action Gap
🔎 Similar Papers