🤖 AI Summary
This study addresses the limited spatial intelligence of multimodal large language models caused by the scarcity of complex 3D question-answering data. To this end, we propose an exemplar-driven, multi-agent collaborative generation paradigm. By leveraging a multi-agent coding framework that invokes geometric toolkits to execute deterministic programs, our approach expands static templates into a large-scale, high-fidelity synthetic dataset, effectively circumventing the inherent spatial reasoning deficiencies of LLMs. Fine-tuning Qwen2.5-VL on this synthesized corpus yields substantial performance improvements across indoor, outdoor, and hybrid scene benchmarks, successfully narrowing the sim-to-real gap.
📝 Abstract
Advancing spatial intelligence in Multimodal Large Language Models (MLLMs) is bottlenecked by the scarcity of complex, scalable 3D question-answer (QA) data. While manual annotation is labor-intensive, directly utilizing LLMs to synthesize these QA pairs often fails due to their inherent deficiencies in spatial and geometric computation. We introduce Exemplar2VQA, a scalable exemplar-driven visual question answering generation framework that rapidly synthesizes large-scale spatial QA pairs in simulated environments via multi-agent coding. By equipping collaborative agents with a meticulously designed library of geometric utilities, Exemplar2VQA bypasses LLMs' spatial reasoning flaws through deterministic code execution. Crucially, the framework exhibits remarkable versatility: taking diverse static object-centric spatial query templates as exemplars, it seamlessly and autonomously scales them into massive, high-fidelity synthetic datasets. Fine-tuning Qwen2.5-VL (3B/7B) exclusively on Exemplar2VQA-generated synthetic indoor data yields significant performance improvements across various diverse benchmarks. Furthermore, its effectiveness is not limited to in-domain indoor datasets but also robustly extends to outdoor and mixed-scene benchmarks. These results establish Exemplar2VQA as a scalable and powerful paradigm for bridging the sim-to-real gap in Embodied AI. Our code is at https://github.com/yingjiayu12/Exemplar2VQA