Score
Designs and implements decoding modules that adaptively fuse, calibrate, and hierarchically process multi-scale feature representations to produce dense outputs such as segmentation masks or boundary maps. This work includes building scale-aware attention and two-stage/cortical-style decoders that align local texture with global shape, preserve fine-grained details, refine ambiguous boundaries, and evaluate robustness across object scales.
Vision Transformers (ViTs) suffer from weak hierarchical modeling capability in conventional Patch Merging, struggling to jointly capture global dependencies and local details. To address this, we propose a brain-inspired Stepwise Patch Merging (SPM) paradigm. SPM decouples global integration from local refinement: a learnable Multi-Scale Aggregation (MSA) module formalizes the brain’s multi-scale fusion mechanism to enhance long-range modeling; a Guided Local Enhancement (GLE) module dynamically strengthens salient local structures. Integrated into hierarchical ViT backbones, SPM supports end-to-end joint training. Extensive experiments on ImageNet-1K, COCO, and ADE20K demonstrate that SPM significantly improves performance on dense prediction tasks—achieving state-of-the-art accuracy and robustness in object detection and semantic segmentation.
This study decodes visual information from high-density neural recordings in the primate cortex to investigate how neural activity underpins perception. We systematically evaluate the impact of model architecture, training objectives, and data scale on decoding performance, proposing an efficient decoder that combines a lightweight temporal attention module with a shallow multilayer perceptron. Furthermore, we introduce a generative framework integrating low-resolution image reconstruction with semantic-conditioned diffusion. Experiments demonstrate that our approach achieves 70% Top-1 accuracy on image retrieval tasks, substantially outperforming existing methods. Our findings also reveal diminishing returns with increasing input dimensionality and dataset size, underscoring the critical role of temporal dynamics modeling in visual neural decoding.
This work addresses the challenge of simultaneously modeling global context and preserving local details in image understanding by proposing ConvNeur, a novel architecture that explicitly decouples global reasoning from local representation for the first time. ConvNeur employs a dual-branch design: a lightweight neural memory branch efficiently captures global context, while a local preservation branch leverages convolutions to retain fine-grained structural details. A learnable gating mechanism adaptively modulates local features using global information. Combined with a compact token aggregation strategy, the model achieves sub-quadratic computational complexity while maintaining local inductive biases. Extensive experiments demonstrate that ConvNeur outperforms existing methods across image classification, object detection, and semantic segmentation tasks, achieving superior accuracy-latency trade-offs at comparable or lower computational costs.
Single-image super-resolution (SISR) models typically exhibit poor generalization across diverse scaling factors. Method: We propose a plug-and-play Scale-Aware Attention Module (SAAM)—a lightweight, parameter-free design that introduces the first scale-adaptive attention mechanism. SAAM is integrated with a gradient variance loss to enhance texture sharpness and is compatible with mainstream SISR backbones (e.g., SCNet, HiT-SR, OverNet). Crucially, it enables fixed-scale models to perform arbitrary integer or non-integer upscaling without retraining. Results: Extensive evaluations on standard benchmarks demonstrate that our approach significantly improves cross-scale robustness and real-world applicability while maintaining low computational overhead. It achieves state-of-the-art performance across multiple metrics and scales, outperforming existing methods in both quantitative accuracy and qualitative fidelity.
Vision-language models (VLMs) remain plagued by object hallucination, undermining visual understanding accuracy. To address this, we propose a selective and contrastive decoding framework centered on objects: it employs attention gating to progressively select multi-scale visual features and introduces a contrastive decoding loss, an object-aware fusion module, and a theoretically grounded scale-consistency constraint. Crucially, our approach is the first to formalize human perceptual alignment as a cross-scale priority ranking and discrepancy suppression process. Evaluated on mainstream hallucination benchmarks—including POPE and HallusionBench—our method achieves an average improvement of 12.7%, significantly outperforming strong baselines such as LLaVA and Qwen-VL. Theoretical analysis establishes convergence guarantees and demonstrates superior generalization properties.
This work addresses functional mismatch and redundancy in the attention mechanisms of current large vision-language models, which fail to efficiently exploit visual context. By establishing a unified framework grounded in information theory and information geometry, the study quantifies the geometric structure and entropy characteristics of residual updates, revealing a functional decoupling between attention mechanisms and feed-forward networks (FFNs) in subspace operations. For the first time from an information-geometric perspective, it clarifies their distinct intrinsic roles and demonstrates that attention can be replaced by predefined weights—such as those derived from Gaussian noise—without performance degradation. Empirical results show that this simplified model matches or even surpasses the original architecture across multiple benchmarks, challenging the prevailing design paradigm reliant on dynamic attention and confirming its substantial redundancy.
Existing image super-resolution methods struggle to balance reconstruction quality and model complexity. To address this challenge, this work proposes a lightweight Multi-scale Spatially Adaptive Attention Network (MSAAN), whose core component is the Multi-scale Spatially Adaptive Attention (MSAA) module. The MSAA module integrates Global Feature Modulation (GFM) and Multi-scale Feature Aggregation (MFA), complemented by a Local Enhancement Block (LEB) and a Feature Interaction Gated Feed-Forward (FIGFF) module to efficiently model both local details and long-range dependencies. Extensive experiments demonstrate that MSAAN and its lightweight variant achieve state-of-the-art or competitive performance in terms of PSNR and SSIM across multiple benchmark datasets, while significantly reducing model parameters and computational overhead.
This work addresses the significant computational redundancy in existing lightweight semantic segmentation methods, which independently compute attention at each level of multi-scale decoders. To overcome this inefficiency, the authors propose a Cross-Stage Attention Propagation (CSAP) mechanism that computes attention only once at the deepest feature layer and efficiently propagates it to shallower layers, thereby eliminating redundant query-key operations while preserving multi-scale contextual modeling capability. Built upon CSAP, the lightweight model CSAP-Tiny achieves 42.9% mIoU on ADE20K with only 5.5 GFLOPs, outperforming SegNeXt-Tiny by 1.8% mIoU while reducing computational cost by 16.8%. This approach is the first to enable cross-stage sharing of attention distributions, effectively balancing efficiency and performance.
This work addresses the underutilized potential of large-scale vision foundation models in salient object detection and the high cost and overfitting risks associated with full fine-tuning. To this end, the authors propose GLASSNet, a framework that freezes the SAMv2 encoder and introduces a lightweight spatial-aware convolutional adapter comprising less than 3% learnable parameters. GLASSNet employs dual decoders to separately model global semantics and local details, fusing their features to produce high-precision saliency maps. Extensive experiments demonstrate that the proposed method surpasses state-of-the-art approaches across multiple benchmarks for both salient and camouflaged object detection, validating its efficiency and effectiveness.
Existing multimodal salient and camouflaged object detection methods suffer from structural complexity and large parameter counts, making it challenging to balance accuracy and efficiency. Inspired by the human visual system, this work proposes a lightweight, unified architecture that integrates a Retinal Integration Module (RIM) for hierarchical, multi-stage cross-modal feature fusion and a Cortical Decoder (CD) that mimics visual cortical mechanisms for layered decoding. This approach establishes a biologically inspired, simplified modeling paradigm capable of supporting diverse modalities and tasks within a single framework. Evaluated across four modalities, seven tasks, and 22 datasets, the model achieves an excellent trade-off between accuracy and efficiency with a compact structure, demonstrating strong generalization capability.