Exemplar2VQA: A Scalable Exemplar-Driven Visual Question Answering Generation Framework via Multi-Agent Coding

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited spatial intelligence of multimodal large language models caused by the scarcity of complex 3D question-answering data. To this end, we propose an exemplar-driven, multi-agent collaborative generation paradigm. By leveraging a multi-agent coding framework that invokes geometric toolkits to execute deterministic programs, our approach expands static templates into a large-scale, high-fidelity synthetic dataset, effectively circumventing the inherent spatial reasoning deficiencies of LLMs. Fine-tuning Qwen2.5-VL on this synthesized corpus yields substantial performance improvements across indoor, outdoor, and hybrid scene benchmarks, successfully narrowing the sim-to-real gap.
📝 Abstract
Advancing spatial intelligence in Multimodal Large Language Models (MLLMs) is bottlenecked by the scarcity of complex, scalable 3D question-answer (QA) data. While manual annotation is labor-intensive, directly utilizing LLMs to synthesize these QA pairs often fails due to their inherent deficiencies in spatial and geometric computation. We introduce Exemplar2VQA, a scalable exemplar-driven visual question answering generation framework that rapidly synthesizes large-scale spatial QA pairs in simulated environments via multi-agent coding. By equipping collaborative agents with a meticulously designed library of geometric utilities, Exemplar2VQA bypasses LLMs' spatial reasoning flaws through deterministic code execution. Crucially, the framework exhibits remarkable versatility: taking diverse static object-centric spatial query templates as exemplars, it seamlessly and autonomously scales them into massive, high-fidelity synthetic datasets. Fine-tuning Qwen2.5-VL (3B/7B) exclusively on Exemplar2VQA-generated synthetic indoor data yields significant performance improvements across various diverse benchmarks. Furthermore, its effectiveness is not limited to in-domain indoor datasets but also robustly extends to outdoor and mixed-scene benchmarks. These results establish Exemplar2VQA as a scalable and powerful paradigm for bridging the sim-to-real gap in Embodied AI. Our code is at https://github.com/yingjiayu12/Exemplar2VQA
Problem

Research questions and friction points this paper is trying to address.

Spatial Intelligence
3D Question Answering
Multimodal Large Language Models
Embodied AI
Sim-to-Real Gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Agent Coding
Visual Question Answering
Spatial Intelligence
Sim-to-Real
Synthetic Data Generation
🔎 Similar Papers
No similar papers found.