Score
Designs and implements modules and algorithms that compute and distribute input-wide contextual representations—using global pooling, cross-attention, or learnable context nodes—and integrate those representations into local feature processing. These components aggregate long-range cues from spatial or structured inputs (e.g., scene elements, keypoints, or graph nodes) to propagate global information, resolve local ambiguities beyond receptive fields, and improve semantic and geometric coherence of predictions.
This work addresses the challenge of simultaneously modeling global context and preserving local details in image understanding by proposing ConvNeur, a novel architecture that explicitly decouples global reasoning from local representation for the first time. ConvNeur employs a dual-branch design: a lightweight neural memory branch efficiently captures global context, while a local preservation branch leverages convolutions to retain fine-grained structural details. A learnable gating mechanism adaptively modulates local features using global information. Combined with a compact token aggregation strategy, the model achieves sub-quadratic computational complexity while maintaining local inductive biases. Extensive experiments demonstrate that ConvNeur outperforms existing methods across image classification, object detection, and semantic segmentation tasks, achieving superior accuracy-latency trade-offs at comparable or lower computational costs.
This work addresses the limitation of existing linear attention methods, which rely on predefined spatial layouts to compress image tokens, thereby constraining information aggregation to coordinate positions rather than semantic content. To overcome this, the paper introduces Representative Attention (RPAttention), a novel representation-driven token compression mechanism that dynamically generates semantic representative tokens to enable spatially agnostic global interactions. RPAttention adopts a lightweight Gather-Interact-Distribute paradigm, integrating competitive similarity-based routing, interaction among representative tokens in a compact latent space, and query-driven cross-attention. This design maintains linear computational complexity while substantially enhancing semantic alignment. Experimental results demonstrate consistent and significant performance gains across image classification, object detection, and semantic segmentation tasks.
This study investigates whether self-supervised vision models spontaneously develop human-like Gestalt perception—specifically illusory contour completion, convexity preference, and dynamic figure-ground segregation—and examines the necessity of global spatial structure modeling. We introduce DiSRT (Distorted Spatial Relationship Testbench), the first diagnostic benchmark to systematically evaluate model sensitivity to core Gestalt principles—including closure, proximity, and figure-ground assignment. Our experiments reveal that self-supervised pretraining (e.g., MAE) induces robust Gestalt perception, whereas subsequent supervised fine-tuning degrades it; reintroducing Top-K sparse activation effectively restores global spatial sensitivity. Notably, ViT and ConvNeXt models evaluated on DiSRT outperform supervised baselines, with some metrics exceeding human performance. These results demonstrate that Gestalt organization does not require attention mechanisms per se and is instead modulated by training paradigms—highlighting the critical role of objective design in shaping perceptual priors.
This paper addresses the poor generalizability, training instability, and weak interpretability of conventional spatial attention mechanisms in convolutional neural networks (CNNs), which stem from irregular, pixel-level attention regions. To this end, we propose a parametric Rectangular Spatial Attention Module (RSAM) that explicitly defines a rectangular attention region using only five learnable parameters. RSAM is fully differentiable and enables end-to-end joint optimization, serving as a plug-and-play component compatible with arbitrary CNN architectures. Our key contribution is the first explicit geometric constraint of spatial attention to a rectangle—enhancing boundary regularity, training stability, and cross-sample generalization, while improving semantic interpretability of attended locations. Extensive experiments on multiple benchmarks demonstrate that RSAM consistently outperforms pixel-wise attention methods, achieving significant gains in classification accuracy, robustness to input perturbations, and visual localization consistency.
This paper addresses two fundamental challenges in foundation model research: the opaque nature of representation mechanisms and diminishing returns from scaling. To resolve these, we propose the “contexture” theory—a unified characterization of representation learning wherein optimal representations maximize mutual information between inputs and contextual variables, with peak generalization achieved at moderate contextual strength. We establish the first unified mathematical framework proving that scaling bottlenecks stem primarily from contextual *quality*, not scale. We introduce two general-purpose context-aware learning objectives—SVME and KISE—and a multi-context fusion strategy. Leveraging information theory and statistical learning theory, we derive a generalization bound for representation learning, unifying theoretical explanations across supervised, self-supervised, and generative pretraining paradigms. Empirical validation confirms that mainstream pretraining objectives implicitly optimize contexture. Our work provides both theoretical foundations and practical guidelines for designing efficient, context-driven pretraining frameworks.
Current visual models lack a unified framework, leading to fragmented local, global, and mechanistic interpretability. This work proposes a unified interpretability framework centered on the instantiated effective receptive field (iERF), which leverages Shared Ratio Decomposition (SRD) to generate high-fidelity saliency maps, integrates Concept-Anchored Feature Explanation (CAFE) to link abstract features with pixel-level evidence, and employs Inter-layer Concept Attribution and Trajectory (ICAT) to uncover concept evolution pathways within deep networks. For the first time, this approach unifies the three interpretability paradigms, significantly outperforming baselines on ResNet50, VGG16, and ViT in both fidelity and robustness. It successfully dissects sparse autoencoder features and clearly delineates dominant concept trajectories underlying correct predictions, misclassifications, and adversarial examples.
This work addresses the unclear mechanisms by which current vision-language models associate spatial relationships with object attributes, particularly the lack of understanding regarding how these models internally process spatial information. Through representational analysis, disentanglement of spatial relations, and enhancement of global visual tokens, the study systematically evaluates the contribution of individual components to spatial reasoning. It reveals, for the first time, that the visual encoder plays a dominant role in spatial reasoning: its output encodes global spatial signals distributed broadly across all image tokens—including background regions—rather than being confined to object-centric areas. Leveraging this insight, augmenting the visual encoder’s global spatial representations substantially improves spatial reasoning performance on natural images, challenging the conventional paradigm that focuses exclusively on object regions.
This study addresses the challenge of constructing neurally plausible, coherent continuous-space representations that unify physical and perceptual information. Building upon the spatial semantic pointer framework, we propose a radial basis kernel implementation tailored for distributed representations and systematically analyze the role of grid-cell-like codes within this architecture. Both theoretical analysis and empirical results demonstrate that grid-cell-like representations not only naturally facilitate the construction of radial basis kernels but also exhibit optimality in terms of neural implementability. Our work is the first to integrate spatial semantic pointers, distributed Fourier embeddings, and grid-cell-like mechanisms into a unified model, yielding an efficient and biologically plausible spatial representation system that demonstrates significant advantages in representational capacity and neural implementation efficiency.
This study investigates how Large Vision-Language Models (LVLMs) construct linear representations of contextual images. By analyzing the principles of linear self-attention projections through both gradient descent analysis and empirical validation, this work elucidates the dimensionality reduction mechanisms operating in early network layers. The primary contribution is the first characterization of a novel mechanism by which LVLMs exploit visual modality properties to achieve dimensional reduction of contextual images. Specifically, the findings reveal that early layers perform a principal component analysis-like process, compressing images into linearly separable representations while discovering shared discriminative geometric structures. This mechanism effectively facilitates downstream classification tasks, offering new insights into the representational dynamics of vision-language architectures.
This work proposes a novel perspective on sequence modeling by reinterpreting the attention mechanism as a dynamic parameter prediction process within a multilayer perceptron (MLP). Traditional Transformers rely on explicit attention for global context modeling, yet their quadratic computational complexity hinders scalability. The proposed approach entirely eliminates explicit attention, instead implicitly compressing global contextual information through dynamically generated MLP parameters, thereby achieving linear computational complexity. This is the first method to fully replace the attention mechanism with a dynamic parameterization strategy while preserving strong global modeling capabilities. Empirical results demonstrate that the model significantly reduces computational overhead in vision tasks without sacrificing performance, matching or closely approaching that of standard Transformers.