WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of effectively coupling navigation with visual reasoning for anomaly detection by multimodal agents in 3D interactive environments. To this end, we construct a 3D world auditing benchmark built upon Unreal Engine 5 and Three.js, comprising 213 tasks across 13 distinct environments. We propose an end-to-end vision-language model (VLM) agent paradigm and present the first systematic evaluation of VLM and vision-language-action (VLA) models regarding their exploration, evidence gathering, and anomaly detection capabilities under fixed budgets. Experimental results reveal that model success rates range from only 6.6% to 42.3%, falling substantially short of the human baseline of 83.4%. These findings highlight significant technical bottlenecks in the action-reasoning coordination mechanisms of current multimodal agents.
📝 Abstract
As interactive 3D worlds are increasingly used to study intelligent behavior, it becomes important to develop efficient pipelines for identifying anomalies in these simulated environments, such as floating objects, traversable walls, or objects inconsistent with the surrounding scene. Multimodal AI systems, including vision-language models (VLMs) and vision-language-action models (VLAs), have shown potential for automating this task. However, 3D world auditing is complex, requiring the close coupling of two distinct capabilities: action, to navigate the 3D world and search for anomalies systematically and efficiently; and visual reasoning, to understand the environment and identify anomalies from multimodal observations. It remains largely unexplored whether multimodal agents can effectively couple these two capabilities, using visual reasoning to identify potential anomalies while taking actions to validate them. In this paper, we introduce WorldAuditBench, a benchmark for 3D world auditing comprising 213 anomaly tasks across 13 environments built with Unreal Engine 5 and Three.js, spanning five anomaly families. We evaluate five frontier models under a fixed exploration budget using two auditing paradigms: VLA-based exploration followed by VLM-based anomaly identification, and an end-to-end VLM agent in which visual reasoning directly guides action selection. Across the evaluated models and two paradigms, success rates range from 6.6% to 42.3%, substantially below human performance (83.4%). Through the task of world auditing, WorldAuditBench provides a testbed for studying how multimodal agents couple action and visual reasoning in interactive 3D environments, while highlighting current limitations in their ability to gather and interpret evidence during exploration.
Problem

Research questions and friction points this paper is trying to address.

3D world auditing
anomaly detection
multimodal agents
visual reasoning
interactive 3D environments
Innovation

Methods, ideas, or system contributions that make the work stand out.

3D World Auditing
Multimodal Agents
Vision-Language-Action Models
Benchmark
Visual Reasoning
🔎 Similar Papers
No similar papers found.