Score
Design and analyze neural encoders for visual data (images or video) that produce hierarchical, multi-stage representations by applying cascaded or grouped attention and progressive token interactions to refine features coarse-to-fine; such systems integrate spatial and temporal information and model fine-grained temporal dependencies across stages.
Existing attention mechanisms struggle to unify modeling across multimodal and multiscale data due to ad hoc heuristic designs lacking theoretical grounding. Method: This paper proposes a hierarchical self-attention framework grounded in entropy minimization—a first-principles derivation of the optimal attention form that intrinsically incorporates hierarchical geometric priors. We mathematically model multiscale structure via explicit construction and enable efficient hierarchical attention computation through dynamic programming, ensuring full compatibility with standard Transformer architectures. Contribution/Results: The method seamlessly integrates into pretrained models, supporting both zero-shot transfer and end-to-end fine-tuning. Experiments demonstrate significant improvements over state-of-the-art heuristic approaches on multimodal benchmarks, achieving superior inference efficiency and stronger generalization—without architectural modifications or additional parameters.
Existing methods struggle to model the hierarchical organization and temporal dynamics of visual cognition, resulting in a significant modality gap between neural responses and visual inputs. To address this, this work proposes NeuroAlign, a novel framework that, for the first time, incorporates the hierarchical structure and temporal dynamics of the biological visual pathway into cross-modal alignment. NeuroAlign achieves fine-grained matching between fMRI signals and video through a two-stage mechanism: it first captures global semantics via Neural-Temporal Contrastive Learning (NTCL), then aligns local patterns using enhanced vector quantization. The framework further introduces a dynamic multimodal fusion module, DynaSyncMM-EMA, along with bidirectional cross-modal prediction. Experiments demonstrate that the proposed method substantially outperforms existing approaches on cross-modal retrieval tasks, offering a new paradigm for understanding the mechanisms underlying visual cognition.
This study addresses the neglect of dynamic, cross-regional neural generation processes in current brain–vision decoding. We propose the Adaptive Topological Vision Transformer (AT-ViT). Methodologically, leveraging the Allen Neuropixel dataset, AT-ViT integrates hierarchical intra-regional neural representations with dynamic inter-regional information gradient modeling—enabling, for the first time, quantitative characterization of hierarchical information flow across visual brain regions. Crucially, we innovatively incorporate hippocampal neural stochasticity as a decoding feature and jointly employ deep neural networks with neuroinformation gradient analysis to unify fine-grained and coarse-grained decoding paradigms. Experiments demonstrate that regional information hierarchy is strongly correlated with decoding performance; visual stimulus reconstruction accuracy improves significantly; and robust, interpretable decoding is achieved across multiple tasks.
Conventional CNNs struggle to model high-order pixel-wise correlations in images. Method: Inspired by nonlinear processing mechanisms in biological vision, we propose a learnable high-order Volterra convolution module that explicitly models multiplicative interactions among pixels. This is the first incorporation of biologically plausible high-order nonlinear convolution into deep visual models, supporting dynamic learning of the optimal expansion order—empirically found to be 3–4, aligning with statistical properties of natural images. Contribution/Results: Through representational similarity analysis (RSA), systematic perturbation studies, and evaluation across multiple datasets (MNIST–Imagenette), our method achieves significant performance gains over standard CNNs on CIFAR-10/100. It reveals order-specific encoding of distinct visual information subdimensions and characterizes hierarchical differences in representational geometry across network layers—establishing a novel paradigm for interpretable, biologically grounded visual modeling.
This work addresses the challenge of enhancing neural networks’ ability to focus on salient information in long-sequence and multimodal tasks. By establishing a unified theoretical framework for attention mechanisms, the study systematically analyzes their mathematical foundations, computational properties, and cross-task generalizability. The framework is instantiated across diverse architectures—including autoregressive Transformers, bidirectional encoders, Vision Transformers, and cross-modal attention models—demonstrating consistent performance gains. The research further uncovers an intrinsic relationship between attention structure and model interpretability, validates empirical scaling laws governing training dynamics and performance, and achieves state-of-the-art results on multiple benchmark datasets. Attention visualization techniques are employed to enhance model transparency, offering insights into the decision-making process of these architectures.
This work addresses the challenge of effectively modeling complex temporal dependencies in existing implicit neural representation (INR)-based video compression methods. To this end, we propose TeNeRV, a hierarchical temporal neural representation framework that enhances local temporal consistency through an inter-frame feature fusion module and jointly captures both short- and long-term dependencies via a Group-of-Pictures (GoP)-adaptive modulation mechanism. TeNeRV implicitly models video content using continuous functions and dynamically adjusts its neural representation parameters according to the GoP structure. Experimental results demonstrate that TeNeRV achieves significantly superior rate-distortion performance compared to current INR-based video compression approaches.
This study decodes visual information from high-density neural recordings in the primate cortex to investigate how neural activity underpins perception. We systematically evaluate the impact of model architecture, training objectives, and data scale on decoding performance, proposing an efficient decoder that combines a lightweight temporal attention module with a shallow multilayer perceptron. Furthermore, we introduce a generative framework integrating low-resolution image reconstruction with semantic-conditioned diffusion. Experiments demonstrate that our approach achieves 70% Top-1 accuracy on image retrieval tasks, substantially outperforming existing methods. Our findings also reveal diminishing returns with increasing input dimensionality and dataset size, underscoring the critical role of temporal dynamics modeling in visual neural decoding.
This work addresses the puzzle of how biological visual systems achieve efficient learning from limited experience without relying on extensive labeled data. It proposes a fully hierarchical unsupervised efficient coding framework that progressively compresses natural images by exploiting their local statistical regularities, layer by layer, without requiring labels or backpropagation. The model constructs human-interpretable features—such as edges, color, texture, and shape—in a bottom-up manner. For the first time, the principle of efficient coding is consistently applied throughout an entire deep network architecture. The resulting representations exhibit strong alignment with human visual cortical responses measured via fMRI and significantly enhance both category learning efficiency and neural alignment under few-shot conditions.
This work aims to establish a correspondence between deep visual models and the hierarchical organization of the human visual cortex to enable high-fidelity reconstruction of fMRI brain activity from images. To this end, we propose CHASMBrain, the first framework to introduce Mamba into brain decoding, featuring a dual-stream Mamba architecture that separately models global semantics and local spatial information. Integrating a Mamba-based variational autoencoder (Mamba-VAE) with a two-stage coarse-to-fine strategy, our approach first predicts region-of-interest activation and subsequently refines predictions to the voxel level. Through causal branch ablation and cross-subject transfer learning, we uncover a causal mapping between the dual-stream design and functional subdivisions of the visual cortex while learning shared visual representations. Evaluated on the Natural Scenes Dataset (NSD), our method achieves a Pearson correlation coefficient of 0.429 and a mean squared error of 0.261, significantly outperforming existing approaches.
The quadratic computational complexity of attention mechanisms severely hinders their scalability in video understanding and generation. Existing sparse attention methods rely on binary block masks, leading to substantial information loss under high sparsity. To address this, we propose Pyramid Sparse Attention (PSA), which constructs hierarchical key-value (KV) representations via multi-level pooling and dynamically assigns queries to appropriate pooling levels based on their importance—enabling fine-grained, continuous sparsity that preserves critical contextual information while drastically reducing computation. PSA employs a hardware-efficient decoupled block-tile kernel and an importance-aware KV block hierarchy. Experiments across diverse video tasks demonstrate that PSA significantly outperforms state-of-the-art sparse attention methods under tight computational budgets, achieving superior long-range context modeling and higher visual fidelity—thereby delivering a better efficiency–quality trade-off.