🤖 AI Summary
This work addresses the substantial computational overhead in knowledge-intensive multimodal question answering (KI-MMQA), where full visual processing, dense retrieval, and cross-modal fusion are typically performed for every query despite only a small fraction of content being relevant. To tackle this inefficiency, we propose SKIP, the first architecture enabling joint sparse routing conditioned on the question, image, and external knowledge. SKIP dynamically allocates computation through question-guided visual token pruning, region-conditioned sparse retrieval, bipartite sparse cross-attention, and speculative knowledge verification. We further introduce an adaptive computation budget controller and derive an information-theoretic bound on visual sparsity via the information bottleneck principle. Evaluated on five KI-MMQA benchmarks, SKIP matches or exceeds the accuracy of dense baselines while reducing FLOPs by 3.4–6.8× and end-to-end latency by 2.7×.
📝 Abstract
Knowledge-intensive multimodal question answering (KI-MMQA) sits at the intersection of three expensive primitives: long visual token sequences, dense retrieval over large external corpora, and full cross-modal fusion. Existing systems pay all three costs uniformly per query, even though only a small fraction of visual content and retrieved knowledge is actually relevant to any given question. We introduce SKIP (Salient Knowledge-Injected Pathways), a unified inference architecture that routes computation along sparse pathways jointly conditioned on the question, the image, and a difficulty estimate. SKIP combines question-guided visual token pruning, region-conditional sparse retrieval, bipartite sparse cross-attention, and speculative knowledge verification with an adaptive budget controller that allocates compute proportional to predicted question difficulty. We derive an information-bottleneck bound showing that the optimal visual sparsity rate scales as $O(1/\sqrt{N})$ under realistic question-image mutual-information assumptions, with retained accuracy guarantees. Across five KI-MMQA benchmarks (OK-VQA, A-OKVQA, InfoSeek, Encyclopedic-VQA, and ViQuAE), SKIP matches or exceeds the accuracy of strong dense baselines while using $3.4$--$6.8\times$ fewer FLOPs and $2.7\times$ less end-to-end latency. Code available at: https://pmlrbd.github.io/skip/