Score
Designs and implements neural encoder or backbone modules that convert convolutional feature maps into token sequences and apply transformer-based encoders (e.g., multi-head attention) to capture global, cross-region dependencies while preserving local convolutional inductive biases. Builds and analyzes hybrid architectures (for example Swin-style patching or integrated transformer–CNN blocks) that combine local CNN processing with transformer token encoding and evaluates their representational behavior, computational trade-offs, and downstream performance.
This work investigates the fundamental differences between spatial token mixers (STMs)—the spatial feature aggregation mechanisms—in Vision Transformers and convolutional networks. To enable a fair, architecture-agnostic comparison, we propose a unified STM modeling paradigm that decouples network-level design from the spatial aggregation module, implementing both convolutional and attention-based STMs on a neutral backbone. Our methodology includes: (1) designing a modular, swappable STM interface; (2) systematically analyzing inductive biases—including receptive field size, translation invariance, and adversarial robustness; and (3) conducting multi-task performance benchmarking. Results show that while modern network-level designs yield substantial gains, intrinsic performance gaps among STMs persist. Crucially, we quantitatively demonstrate for the first time that convolutions exhibit superior translation invariance and local robustness, whereas attention achieves larger effective receptive fields but is more vulnerable to input perturbations.
To address the high computational cost and heavy data dependency of Vision Transformers (ViTs) in visual tasks, this paper proposes NiNformer: a lightweight architecture that replaces standard self-attention layers with Network-in-Network (NiN) blocks. It introduces a learnable, element-wise dynamic gating mechanism driven by token mixing to enable efficient feature transformation. Unlike conventional ViTs, NiNformer abandons global attention and static MLP-based fusion, instead pioneering the integration of NiN-style hierarchical convolutional abstraction with token-mixing–driven gating. This design preserves strong representational capacity while drastically reducing FLOPs. Extensive experiments demonstrate that NiNformer consistently outperforms ViT, MLP-Mixer, and Conv-Mixer on mainstream image classification benchmarks—including ImageNet—achieving higher accuracy with significantly lower computational cost. The work establishes a novel, efficient paradigm for vision modeling grounded in architectural innovation rather than scale.
Scaling Transformer models is prohibitively expensive due to fixed-parameter linear projection layers; architectural modifications necessitate full retraining. Method: We propose TokenFormer, the first architecture introducing *parameter tokenization*, which models model parameters as learnable tokens and replaces all linear layers with token-parameter self-attention—unifying parameter and input token representations in a shared latent space. Contribution/Results: Our method enables zero-shot, progressive parameter expansion without retraining, overcoming classical scaling bottlenecks. Without altering network topology, we scale model parameters from 124M to 1.4B while matching the performance of fully trained baselines, achieving substantial training cost reduction. The code and models are publicly released.
This work challenges the prevailing consensus that locality-inductive bias is indispensable in vision Transformers. It investigates whether pixel-level tokenization—bypassing conventional patch-based partitioning (e.g., 16×16) and convolutional priors—is both feasible and effective. Method: The authors directly serialize raw images into pixel-level tokens and adopt a standard Transformer architecture, trained via masked autoencoding and diffusion-model paradigms. They comprehensively evaluate performance across image classification, dense prediction, self-supervised reconstruction, and generative modeling. Contribution/Results: Experiments demonstrate that pure pixel-level Transformers achieve competitive or superior performance to state-of-the-art ViTs on multiple benchmarks—including ImageNet-1K, COCO, and ADE20K—without any explicit locality bias. This is the first empirical evidence that locality-inductive bias is not strictly necessary for high visual representation learning. The findings establish a new architectural paradigm and provide theoretical grounding for next-generation vision models grounded in sequence-based, bias-free design principles.
To address the limited representational capacity of conventional static convolutions in CNN-Transformer hybrid architectures, this paper proposes the input-adaptive Dual-Dynamic Token Mixer (D-Mixer)—the first to jointly integrate input-driven depthwise separable convolution with lightweight global attention for synergistic modeling of local details and long-range dependencies. D-Mixer enables dynamic cross-module feature fusion and adaptive receptive field expansion, overcoming the inherent limitations of static convolution. Built upon D-Mixer, TransXNet-T achieves a 0.3% top-1 accuracy gain on ImageNet-1K over Swin-T while consuming less than 50% of its FLOPs; its small and base variants attain 83.8% and 84.6% top-1 accuracy, respectively. Moreover, TransXNet demonstrates superior performance on dense prediction tasks—e.g., semantic segmentation and object detection—at significantly lower computational cost, surpassing state-of-the-art methods.
This work addresses the feature learning conflict in Vision Transformers arising from the shared computational pathway for both the global [CLS] token and local patch tokens, which hinders performance in dense prediction tasks. The study reveals, for the first time, that normalization layers implicitly differentiate between these two token types. Building on this insight, the authors propose a lightweight token-specific processing mechanism that decouples their computational flows within the normalization layers and early QKV projections. This approach incurs no additional computational overhead and increases model parameters by only 8%, yet consistently improves performance by over 2 mIoU on standard segmentation benchmarks while preserving strong image classification accuracy.
This work addresses the high computational cost of Transformer-based models in 3D medical image segmentation by proposing the Token-UNet family of lightweight architectures. Building upon a standard 3D convolutional encoder, the approach integrates TokenLearner and TokenFuser modules to efficiently tokenize feature maps, enabling synergistic modeling of both local and global structural information. The resulting framework substantially reduces computational overhead while producing interpretable attention maps. Experimental results demonstrate that the best-performing variant achieves a Dice score of 87.21% ± 0.35%, outperforming SwinUNETR (86.75% ± 0.19%) while reducing memory consumption, inference time, and parameter count to 33%, 10%, and 35% of those of SwinUNETR, respectively.
This work challenges the prevailing reliance of Chinese language models on discrete token embeddings by investigating whether effective language modeling can be achieved using only glyph images. To this end, the authors construct a dual-branch controlled framework: one branch rasterizes character sequences into images processed by a visual encoder composed of ResNet and a shallow Vision Transformer (ViT), while the other employs conventional index-based embeddings as a baseline; both branches share an identical decoder to ensure strict variable control. Experiments demonstrate for the first time that pure glyph-based input is not only viable but consistently outperforms the baseline across all decoders—achieving up to 0.429 accuracy (a 21% relative improvement), converging nearly twice as fast, and showing advantages with only 21% of the training data. The approach also exhibits greater robustness to character perturbations, revealing both the modality-agnostic capacity of Transformers and the information-rich structural properties inherent in Chinese characters.