adaptive multi-scale decoding

Designs and implements decoding modules that adaptively fuse, calibrate, and hierarchically process multi-scale feature representations to produce dense outputs such as segmentation masks or boundary maps. This work includes building scale-aware attention and two-stage/cortical-style decoders that align local texture with global shape, preserve fine-grained details, refine ambiguous boundaries, and evaluate robustness across object scales.

adaptivemulti-scaledecoding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.56
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Brain-Inspired Stepwise Patch Merging for Vision Transformers

Sep 11, 2024
YY
Yong Yu
🏛️ University of Chinese Academy of Sciences | Institute of Automation, Chinese Academy of Sciences | Center for Long-term Artificial Intelligence

Vision Transformers (ViTs) suffer from weak hierarchical modeling capability in conventional Patch Merging, struggling to jointly capture global dependencies and local details. To address this, we propose a brain-inspired Stepwise Patch Merging (SPM) paradigm. SPM decouples global integration from local refinement: a learnable Multi-Scale Aggregation (MSA) module formalizes the brain’s multi-scale fusion mechanism to enhance long-range modeling; a Guided Local Enhancement (GLE) module dynamically strengthens salient local structures. Integrated into hierarchical ViT backbones, SPM supports end-to-end joint training. Extensive experiments on ImageNet-1K, COCO, and ADE20K demonstrate that SPM significantly improves performance on dense prediction tasks—achieving state-of-the-art accuracy and robustness in object detection and semantic segmentation.

Balances long-range dependency modeling and local feature enhancementEnhances Vision Transformers' hierarchical architecture via brain-inspired patch mergingImproves performance in dense prediction tasks like object detection

This study decodes visual information from high-density neural recordings in the primate cortex to investigate how neural activity underpins perception. We systematically evaluate the impact of model architecture, training objectives, and data scale on decoding performance, proposing an efficient decoder that combines a lightweight temporal attention module with a shallow multilayer perceptron. Furthermore, we introduce a generative framework integrating low-resolution image reconstruction with semantic-conditioned diffusion. Experiments demonstrate that our approach achieves 70% Top-1 accuracy on image retrieval tasks, substantially outperforming existing methods. Our findings also reveal diminishing returns with increasing input dimensionality and dataset size, underscoring the critical role of temporal dynamics modeling in visual neural decoding.

brain-computer interfaceintracortical recordingsneural decoding

This work addresses the challenge of simultaneously modeling global context and preserving local details in image understanding by proposing ConvNeur, a novel architecture that explicitly decouples global reasoning from local representation for the first time. ConvNeur employs a dual-branch design: a lightweight neural memory branch efficiently captures global context, while a local preservation branch leverages convolutions to retain fine-grained structural details. A learnable gating mechanism adaptively modulates local features using global information. Combined with a compact token aggregation strategy, the model achieves sub-quadratic computational complexity while maintaining local inductive biases. Extensive experiments demonstrate that ConvNeur outperforms existing methods across image classification, object detection, and semantic segmentation tasks, achieving superior accuracy-latency trade-offs at comparable or lower computational costs.

computational efficiencyglobal contextglobal-local decoupling

Single-image super-resolution (SISR) models typically exhibit poor generalization across diverse scaling factors. Method: We propose a plug-and-play Scale-Aware Attention Module (SAAM)—a lightweight, parameter-free design that introduces the first scale-adaptive attention mechanism. SAAM is integrated with a gradient variance loss to enhance texture sharpness and is compatible with mainstream SISR backbones (e.g., SCNet, HiT-SR, OverNet). Crucially, it enables fixed-scale models to perform arbitrary integer or non-integer upscaling without retraining. Results: Extensive evaluations on standard benchmarks demonstrate that our approach significantly improves cross-scale robustness and real-world applicability while maintaining low computational overhead. It achieves state-of-the-art performance across multiple metrics and scales, outperforming existing methods in both quantitative accuracy and qualitative fidelity.

Enabling arbitrary-scale super-resolution in fixed modelsEnhancing generalization across varying scale factorsImproving real-world applicability with minimal overhead

Vision-language models (VLMs) remain plagued by object hallucination, undermining visual understanding accuracy. To address this, we propose a selective and contrastive decoding framework centered on objects: it employs attention gating to progressively select multi-scale visual features and introduces a contrastive decoding loss, an object-aware fusion module, and a theoretically grounded scale-consistency constraint. Crucially, our approach is the first to formalize human perceptual alignment as a cross-scale priority ranking and discrepancy suppression process. Evaluated on mainstream hallucination benchmarks—including POPE and HallusionBench—our method achieves an average improvement of 12.7%, significantly outperforming strong baselines such as LLaVA and Qwen-VL. Theoretical analysis establishes convergence guarantees and demonstrates superior generalization properties.

Addressing object hallucination in Vision-Language ModelsLeveraging multi-scale visual information effectivelyReducing perceptual hallucinations via selective contrastive decoding

Latest Papers

What's happening recently
View more

This work addresses functional mismatch and redundancy in the attention mechanisms of current large vision-language models, which fail to efficiently exploit visual context. By establishing a unified framework grounded in information theory and information geometry, the study quantifies the geometric structure and entropy characteristics of residual updates, revealing a functional decoupling between attention mechanisms and feed-forward networks (FFNs) in subspace operations. For the first time from an information-geometric perspective, it clarifies their distinct intrinsic roles and demonstrates that attention can be replaced by predefined weights—such as those derived from Gaussian noise—without performance degradation. Empirical results show that this simplified model matches or even surpasses the original architecture across multiple benchmarks, challenging the prevailing design paradigm reliant on dynamic attention and confirming its substantial redundancy.

Attention MechanismLarge Vision-Language ModelsModel Redundancy

Existing image super-resolution methods struggle to balance reconstruction quality and model complexity. To address this challenge, this work proposes a lightweight Multi-scale Spatially Adaptive Attention Network (MSAAN), whose core component is the Multi-scale Spatially Adaptive Attention (MSAA) module. The MSAA module integrates Global Feature Modulation (GFM) and Multi-scale Feature Aggregation (MFA), complemented by a Local Enhancement Block (LEB) and a Feature Interaction Gated Feed-Forward (FIGFF) module to efficiently model both local details and long-range dependencies. Extensive experiments demonstrate that MSAAN and its lightweight variant achieve state-of-the-art or competitive performance in terms of PSNR and SSIM across multiple benchmark datasets, while significantly reducing model parameters and computational overhead.

Image Super-ResolutionLightweight NetworkModel Complexity

This work addresses the significant computational redundancy in existing lightweight semantic segmentation methods, which independently compute attention at each level of multi-scale decoders. To overcome this inefficiency, the authors propose a Cross-Stage Attention Propagation (CSAP) mechanism that computes attention only once at the deepest feature layer and efficiently propagates it to shallower layers, thereby eliminating redundant query-key operations while preserving multi-scale contextual modeling capability. Built upon CSAP, the lightweight model CSAP-Tiny achieves 42.9% mIoU on ADE20K with only 5.5 GFLOPs, outperforming SegNeXt-Tiny by 1.8% mIoU while reducing computational cost by 16.8%. This approach is the first to enable cross-stage sharing of attention distributions, effectively balancing efficiency and performance.

attention redundancycomputational efficiencylightweight models

This work addresses the underutilized potential of large-scale vision foundation models in salient object detection and the high cost and overfitting risks associated with full fine-tuning. To this end, the authors propose GLASSNet, a framework that freezes the SAMv2 encoder and introduces a lightweight spatial-aware convolutional adapter comprising less than 3% learnable parameters. GLASSNet employs dual decoders to separately model global semantics and local details, fusing their features to produce high-precision saliency maps. Extensive experiments demonstrate that the proposed method surpasses state-of-the-art approaches across multiple benchmarks for both salient and camouflaged object detection, validating its efficiency and effectiveness.

Computational CostFoundation ModelsOverfitting

Existing multimodal salient and camouflaged object detection methods suffer from structural complexity and large parameter counts, making it challenging to balance accuracy and efficiency. Inspired by the human visual system, this work proposes a lightweight, unified architecture that integrates a Retinal Integration Module (RIM) for hierarchical, multi-stage cross-modal feature fusion and a Cortical Decoder (CD) that mimics visual cortical mechanisms for layered decoding. This approach establishes a biologically inspired, simplified modeling paradigm capable of supporting diverse modalities and tasks within a single framework. Evaluated across four modalities, seven tasks, and 22 datasets, the model achieves an excellent trade-off between accuracy and efficiency with a compact structure, demonstrating strong generalization capability.

camouflaged object detectionmodel efficiencymultimodal fusion

Hot Scholars

MA

Mahdi Abdelguerfi

Professor of Computer Science, University of New Orleans
Geospatial IntelligenceBig DataAI
MM

Md Meftahul Ferdaus

University of New Orleans, Postdoctoral Research Scientist
MLOpsLightweight Neural NetworksComputer Vision and RobotsMR Materials
CL

Chenxin Li

The Chinese University of Hong Kong
Multimodal LLMAgentWorld Model
TF

Teng Fei

School of Resources and Environmental Science, Wuhan University
Remote SensingGISSocial SensingPlanning