Score
Designs, implements, and evaluates systems that take visual inputs (images or video) together with natural-language questions and produce answers, covering model architectures, multi-modal fusion, temporal and spatial reasoning, training and inference pipelines, and evaluation metrics. This competence also includes creating datasets and annotation tools, grounding and attention mechanisms, and deployment or API components for interactive visual question answering.
This survey addresses core challenges in visual question answering (VQA): insufficient deep image–language understanding, inconsistent evaluation protocols, and practical deployment difficulties. Methodologically, it introduces the first unified VQA architecture taxonomy grounded in information flow and cross-modal interaction; integrates large vision-language models (LVLMs) as an emerging paradigm; and comprehensively reviews mainstream datasets, evaluation metrics, multimodal alignment techniques, and representative applications. Key contributions include: (1) the first comparable four-dimensional analytical framework spanning architectures, datasets, evaluation, and applications; (2) a distilled synthesis of six open research questions and persistent technical bottlenecks; and (3) concrete, scenario-driven implementation pathways for education, healthcare, and accessibility domains. By establishing a rigorous theoretical benchmark and a reproducible technical roadmap, this work provides a foundational reference for advancing VQA research and real-world adoption.
This study investigates the potential of vision-language models (VLMs) to enhance learner engagement and knowledge retention in automated question generation for educational videos. Methodologically, it systematically evaluates zero-shot capabilities of representative VLMs—including LLaVA and Qwen-VL—integrated with video frame sampling, cross-modal alignment, and supervised fine-tuning; an expert-designed human evaluation framework quantifies question relevance, diversity, and answer solvability. Key contributions include: (1) the first comprehensive benchmark of VLMs for educational video QA generation; (2) empirical validation that zero-shot generation is feasible yet exhibits substantial bias; (3) after fine-tuning, question relevance improves by 32% and answer solvability reaches 86%; and (4) identification of modality redundancy and difficulty calibration as critical bottlenecks, leading to a novel multimodal data curation paradigm tailored to educational contexts and concrete directions for future research.
Current large multimodal models (LMMs) exhibit strong visual perception capabilities but lack high-level, task-specific compositional reasoning—hindering progress toward general visual intelligence. Method: We propose a human-inspired “Perceive–Reason–Answer” single-pass reasoning paradigm, enabling end-to-end visual reasoning within a single forward pass—without iterative inference, multi-step API calls, or external tools. Our approach extends large vision-language model architectures with integrated visual grounding, multi-granularity representation learning, and instruction tuning, trained on a newly curated 334K-sample high-quality visual instruction dataset. Contribution/Results: The resulting model, Griffon-R, achieves state-of-the-art performance on complex visual reasoning benchmarks (e.g., VSR, CLEVR) and substantially improves results across mainstream multimodal evaluations (e.g., MMBench, ScienceQA). It further demonstrates enhanced interpretability and response faithfulness, advancing both capability and reliability in visual language understanding.
This study systematically evaluates five representative VQA models—ABC-CNN, KICNLE, MVLM, BLIP-2, and OFA—addressing three core challenges: dataset bias, weak commonsense reasoning, and poor real-world generalization. Methodologically, it conducts the first unified, cross-paradigm, multi-dimensional analysis across bias robustness, cross-domain generalization, and implicit reasoning capability. The evaluation framework introduces novel, realistic-scenario-oriented assessment protocols. Key findings reveal substantial performance disparities: BLIP-2 achieves superior open-domain generalization, whereas OFA excels in fine-grained visual understanding; critically, all models attain <38% average accuracy on counterfactual questions, exposing a fundamental commonsense reasoning bottleneck. These results provide empirical grounding and methodological insights for future VQA model design and evaluation, emphasizing the need for bias-mitigated training, enhanced causal and counterfactual reasoning, and more ecologically valid benchmarks.
Existing visual question answering (VQA) methods struggle with complex spatial relational reasoning due to weak cross-modal interaction and insufficient modeling of entity-level spatial configurations. To address this, we propose the Dynamic Bidirectional Spatial Tower (DBST), a paradigm shift from passive “seeing” to active “perceiving–organizing” grounded in Gestalt principles. DBST decomposes images hierarchically across four layers—dynamic spatial decomposition, bidirectional hierarchical perception, Gestalt-driven multi-scale spatial organization, and lightweight cross-modal fusion—introducing strong structural priors for spatial relation modeling and overcoming the limitations of pixel-level exhaustive search. Its modular design enables plug-and-play integration. Built upon DBST, the July model (3B parameters) achieves state-of-the-art performance on spatial relational VQA benchmarks, demonstrating significant gains in both accuracy and interpretability.
Embodied AI urgently requires real-time audio-visual interaction capabilities, yet no standardized benchmark exists for synchronous camera-microphone input in dynamic scenes. Method: We introduce IVD—the first Interactive Video Dialogue benchmark for real-time multimodal question answering—featuring a naturalistic face-to-face interaction evaluation paradigm. IVD incorporates core challenges including fine-grained audio-visual temporal alignment and cross-modal coreference resolution, addressed via end-to-end audio-visual QA fine-tuning and dynamic coreference modeling. Results: State-of-the-art vision-language models underperform humans by over 35% on average on IVD; after fine-grained audio-visual joint fine-tuning, key perceptual capabilities improve by up to 42.6%, significantly narrowing the human–machine performance gap. This work provides the first systematic characterization of real-time audio-visual interaction bottlenecks, establishing both a practical technical pathway and a standardized evaluation foundation for embodied AI.
Existing vision-language models (VLMs) predominantly rely on text-dominated reasoning paradigms, limiting their capacity for fine-grained, truly image-centric multimodal reasoning—i.e., “thinking in images.” Method: We propose an image-interactive reasoning paradigm, introducing the first large-scale, interleaved image-text dataset with 31K chain-of-thought trajectories, and design an endogenous visual thinking generation mechanism. This mechanism performs reasoning directly within the visual embedding space—without external tool invocation—enabling more flexible and image-faithful thought evolution. Contribution/Results: Experiments demonstrate substantial improvements over strong baselines across multiple multimodal reasoning benchmarks. Our results validate both the efficacy of the curated dataset construction strategy and the superiority of visual-space reasoning modeling. The proposed framework advances VLMs toward deeper visual understanding by shifting reasoning from textual abstraction to intrinsic visual representation, offering a novel pathway for next-generation multimodal intelligence.
This work addresses a critical gap in existing visual question answering (VQA) benchmarks: their inability to evaluate models’ understanding of the underlying data distributions behind scientific charts, particularly the complex, non-bijective relationships between visual representations and their source data. To this end, the authors propose the first data-distribution-centered VQA paradigm, introducing a novel benchmark constructed from synthetically generated histograms grounded in real underlying data. The dataset includes annotations for distribution parameters, raw data points, and bounding boxes of visual elements. Questions are carefully designed to probe distributional reasoning, combining human-authored and large language model–generated queries to comprehensively assess models’ grasp of the data generation process. The released open-source dataset not only exposes significant limitations of current VQA models in distributional reasoning but also establishes a robust foundation for future research.
Multimodal large language models (MLLMs) suffer from visual hallucinations and over-reliance on textual priors during visual reasoning. Method: We propose a tool-augmented agent architecture that decouples the LLM from a lightweight, specialized vision module, enabling fine-grained visual analysis and iterative reasoning via chain-of-thought guidance. We introduce a three-stage diagnostic evaluation framework to systematically uncover failure modes of mainstream MLLMs and design a modular, interpretable agent workflow supporting dynamic visual tool invocation and result verification. Contribution/Results: Our approach achieves +10.3 and +6.0 absolute improvements on MMMU and MathVista, respectively—surpassing same-parameter-scale models and approaching the performance of significantly larger ones. The code and evaluation framework are publicly released.
本文针对大型视频模型在视觉理解上的不足,通过引入三个新的以视觉为中心的评估基准来解决这一问题。
研究提出了一种视觉问答模型,用于非破坏性检测图像分析,通过深度学习和自然语言处理技术提高检测效率和准确性。