NeuronEye: Query-Guided Visual Concept Activation for Vision-Language Reasoning

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the implicit entanglement of visual features in vision-language models, which hinders the isolation of query-specific evidence. Inspired by biological vision, we propose a plug-and-play sparse concept activation framework that constructs an overcomplete sparse concept vocabulary. Through query-guided cluster activation and patch-level localization, the method focuses on local evidence while introducing a complementary inhibition mechanism to preserve weak cues, thereby unifying interpretability and reasoning capability within a single forward pass. Evaluated on Qwen2.5-VL, the proposed approach improves accuracy by 3.1% on CV-Bench, 9.5% on distance-related tasks, and 8.3% on BLINK multi-view benchmarks, demonstrating its effectiveness in enhancing both transparent and robust visual reasoning.
📝 Abstract
Current vision-language models (VLMs) encode visual information in dense hidden states where object identity, spatial layout, and local attributes are implicitly entangled rather than explicitly disentangled, limiting their ability to isolate and modulate the specific visual evidence required by a given language query. Inspired by sparse population coding and top-down modulation in biological vision, we introduce NeuronEye, a plug-in framework that constructs a sparse, concept-level neuron vocabulary from intermediate VLM representations and selectively activates query-relevant visual concepts during inference. NeuronEye decomposes vision-token states into an overcomplete sparse basis organized by concept-level clusters, uses the language query to activate relevant clusters and localize the patches where selected concepts are expressed, and injects the focused evidence back into vision tokens. A complementary suppression mechanism attenuates dominant perceptual directions to preserve weaker but relevant cues. All operations run in a single forward pass over a frozen VLM backbone. On Qwen2.5-VL-7B, NeuronEye raises CV-Bench overall accuracy by +3.1 with gains of +9.5 on Distance, and improves BLINK Multi-view by +8.3, with similar trends on LLaVA-1.6-7B. These results suggest that sparse neuron vocabularies can serve not only as post-hoc interpretability tools but also as active interfaces for concept-level visual reasoning.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Visual Reasoning
Dense Representations
Concept Disentanglement
Query-Guided Activation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Neuron Vocabulary
Query-Guided Activation
Concept-Level Reasoning
Top-Down Modulation
Plug-and-Play Framework
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
R
Ruiyu Yan
New York University
B
Bowen Chen
New Jersey Institute of Technology
S
Shaowen Wan
New Jersey Institute of Technology
Lin Zhao
Lin Zhao
New Jersey Institute of Technology
Brain-inspired AIMedical Image AnalysisArtificial General Intelligence