🤖 AI Summary
Existing visual reasoning research is fragmented across subdomains—including relational, symbolic, temporal, causal, and commonsense reasoning—lacking a unified taxonomy, comparable evaluation protocols, and systematic analysis. Method: We propose the first cross-paradigmatic unified classification framework for visual reasoning, integrating graph neural networks, memory-augmented architectures, attention mechanisms, and neuro-symbolic methods into a cohesive perception-reasoning architecture. We further design a multidimensional evaluation protocol assessing functional correctness, structural consistency, and causal validity. Contribution/Results: Our analysis reveals shared bottlenecks across state-of-the-art methods—particularly in out-of-distribution generalization, interpretability, and performance under weak supervision. The framework establishes a theoretical benchmark and practical technical guidelines for visual reasoning, enabling robust, trustworthy AI applications in domains such as autonomous driving and medical diagnosis.
📝 Abstract
Visual reasoning is critical for a wide range of computer vision tasks that go beyond surface-level object detection and classification. Despite notable advances in relational, symbolic, temporal, causal, and commonsense reasoning, existing surveys often address these directions in isolation, lacking a unified analysis and comparison across reasoning types, methodologies, and evaluation protocols. This survey aims to address this gap by categorizing visual reasoning into five major types (relational, symbolic, temporal, causal, and commonsense) and systematically examining their implementation through architectures such as graph-based models, memory networks, attention mechanisms, and neuro-symbolic systems. We review evaluation protocols designed to assess functional correctness, structural consistency, and causal validity, and critically analyze their limitations in terms of generalizability, reproducibility, and explanatory power. Beyond evaluation, we identify key open challenges in visual reasoning, including scalability to complex scenes, deeper integration of symbolic and neural paradigms, the lack of comprehensive benchmark datasets, and reasoning under weak supervision. Finally, we outline a forward-looking research agenda for next-generation vision systems, emphasizing that bridging perception and reasoning is essential for building transparent, trustworthy, and cross-domain adaptive AI systems, particularly in critical domains such as autonomous driving and medical diagnostics.