Dynamic Double Space Tower

📅 2025-06-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing visual question answering (VQA) methods struggle with complex spatial relational reasoning due to weak cross-modal interaction and insufficient modeling of entity-level spatial configurations. To address this, we propose the Dynamic Bidirectional Spatial Tower (DBST), a paradigm shift from passive “seeing” to active “perceiving–organizing” grounded in Gestalt principles. DBST decomposes images hierarchically across four layers—dynamic spatial decomposition, bidirectional hierarchical perception, Gestalt-driven multi-scale spatial organization, and lightweight cross-modal fusion—introducing strong structural priors for spatial relation modeling and overcoming the limitations of pixel-level exhaustive search. Its modular design enables plug-and-play integration. Built upon DBST, the July model (3B parameters) achieves state-of-the-art performance on spatial relational VQA benchmarks, demonstrating significant gains in both accuracy and interpretability.

Technology Category

Knowledge Representation and Reasoning: Geometric, Spatial, and Temporal ReasoningReasoning under Uncertainty: Relational Probabilistic ModelsComputer Vision: Visual Reasoning & Symbolic Representations

Application Category

Search and Retrieval-Augmented AI: Vertical and domain-specific searchSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphs
📝 Abstract
The Visual Question Answering (VQA) task requires the simultaneous understanding of image content and question semantics. However, existing methods often have difficulty handling complex reasoning scenarios due to insufficient cross-modal interaction and capturing the entity spatial relationships in the image.cite{huang2023adaptive}cite{liu2021comparing}cite{guibas2021adaptive}cite{zhang2022vsa}We studied a brand-new approach to replace the attention mechanism in order to enhance the reasoning ability of the model and its understanding of spatial relationships.Specifically, we propose a dynamic bidirectional spatial tower, which is divided into four layers to observe the image according to the principle of human gestalt vision. This naturally provides a powerful structural prior for the spatial organization between entities, enabling the model to no longer blindly search for relationships between pixels but make judgments based on more meaningful perceptual units. Change from"seeing images"to"perceiving and organizing image content".A large number of experiments have shown that our module can be used in any other multimodal model and achieve advanced results, demonstrating its potential in spatial relationship processing.Meanwhile, the multimodal visual question-answering model July trained by our method has achieved state-of-the-art results with only 3B parameters, especially on the question-answering dataset of spatial relations.
Problem

Research questions and friction points this paper is trying to address.

Enhancing VQA model reasoning via cross-modal interaction improvement
Capturing entity spatial relationships in images effectively
Replacing attention mechanism with dynamic bidirectional spatial tower
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic bidirectional spatial tower enhances reasoning
Four-layer gestalt vision principle improves perception
Achieves state-of-the-art with minimal parameters
🔎 Similar Papers
No similar papers found.
W
Weikai Sun
S
Shijie Song
H
Han Wang