Visual question answering: from early developments to recent advances -- a survey

๐Ÿ“… 2025-01-07
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This survey addresses core challenges in visual question answering (VQA): insufficient deep imageโ€“language understanding, inconsistent evaluation protocols, and practical deployment difficulties. Methodologically, it introduces the first unified VQA architecture taxonomy grounded in information flow and cross-modal interaction; integrates large vision-language models (LVLMs) as an emerging paradigm; and comprehensively reviews mainstream datasets, evaluation metrics, multimodal alignment techniques, and representative applications. Key contributions include: (1) the first comparable four-dimensional analytical framework spanning architectures, datasets, evaluation, and applications; (2) a distilled synthesis of six open research questions and persistent technical bottlenecks; and (3) concrete, scenario-driven implementation pathways for education, healthcare, and accessibility domains. By establishing a rigorous theoretical benchmark and a reproducible technical roadmap, this work provides a foundational reference for advancing VQA research and real-world adoption.

Technology Category

Computer Vision: Multi-modal VisionNatural Language Processing: Question AnsweringMachine Learning: Deep Neural Architectures and Foundation Models

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsWeb Mining and Content Analysis: Community question answering
๐Ÿ“ Abstract
Visual Question Answering (VQA) is an evolving research field aimed at enabling machines to answer questions about visual content by integrating image and language processing techniques such as feature extraction, object detection, text embedding, natural language understanding, and language generation. With the growth of multimodal data research, VQA has gained significant attention due to its broad applications, including interactive educational tools, medical image diagnosis, customer service, entertainment, and social media captioning. Additionally, VQA plays a vital role in assisting visually impaired individuals by generating descriptive content from images. This survey introduces a taxonomy of VQA architectures, categorizing them based on design choices and key components to facilitate comparative analysis and evaluation. We review major VQA approaches, focusing on deep learning-based methods, and explore the emerging field of Large Visual Language Models (LVLMs) that have demonstrated success in multimodal tasks like VQA. The paper further examines available datasets and evaluation metrics essential for measuring VQA system performance, followed by an exploration of real-world VQA applications. Finally, we highlight ongoing challenges and future directions in VQA research, presenting open questions and potential areas for further development. This survey serves as a comprehensive resource for researchers and practitioners interested in the latest advancements and future
Problem

Research questions and friction points this paper is trying to address.

Visual Question Answering
Performance Improvement
Practical Applications
Innovation

Methods, ideas, or system contributions that make the work stand out.

Deep Learning
Visual Question Answering
Real-world Applications
๐Ÿ”Ž Similar Papers
No similar papers found.