hybrid cnn-transformer encoding

Designs and implements neural encoder or backbone modules that convert convolutional feature maps into token sequences and apply transformer-based encoders (e.g., multi-head attention) to capture global, cross-region dependencies while preserving local convolutional inductive biases. Builds and analyzes hybrid architectures (for example Swin-style patching or integrated transformer–CNN blocks) that combine local CNN processing with transformer token encoding and evaluates their representational behavior, computational trade-offs, and downstream performance.

hybridcnn-transformerencoding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.55
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Demystify Transformers & Convolutions in Modern Image Deep Networks

Nov 10, 2022
JD
Jifeng Dai
🏛️ Tsinghua University | Shanghai Artificial Intelligence Laboratory | Huazhong University of Science and Technology | Fudan University | The Chinese University of Hong Kong | SenseTime Research | South China University of Technology

This work investigates the fundamental differences between spatial token mixers (STMs)—the spatial feature aggregation mechanisms—in Vision Transformers and convolutional networks. To enable a fair, architecture-agnostic comparison, we propose a unified STM modeling paradigm that decouples network-level design from the spatial aggregation module, implementing both convolutional and attention-based STMs on a neutral backbone. Our methodology includes: (1) designing a modular, swappable STM interface; (2) systematically analyzing inductive biases—including receptive field size, translation invariance, and adversarial robustness; and (3) conducting multi-task performance benchmarking. Results show that while modern network-level designs yield substantial gains, intrinsic performance gaps among STMs persist. Crucially, we quantitatively demonstrate for the first time that convolutions exhibit superior translation invariance and local robustness, whereas attention achieves larger effective receptive fields but is more vulnerable to input perturbations.

Analyze performance differences in attention vs convolutionCompare spatial token mixers in vision backbonesUnify architecture to isolate feature transformation effects

NiNformer: A Network in Network Transformer with Token Mixing Generated Gating Function

Mar 04, 2024
AN
Abdullah Nazhat Abdullah
🏛️ Bahcesehir University

To address the high computational cost and heavy data dependency of Vision Transformers (ViTs) in visual tasks, this paper proposes NiNformer: a lightweight architecture that replaces standard self-attention layers with Network-in-Network (NiN) blocks. It introduces a learnable, element-wise dynamic gating mechanism driven by token mixing to enable efficient feature transformation. Unlike conventional ViTs, NiNformer abandons global attention and static MLP-based fusion, instead pioneering the integration of NiN-style hierarchical convolutional abstraction with token-mixing–driven gating. This design preserves strong representational capacity while drastically reducing FLOPs. Extensive experiments demonstrate that NiNformer consistently outperforms ViT, MLP-Mixer, and Conv-Mixer on mainstream image classification benchmarks—including ImageNet—achieving higher accuracy with significantly lower computational cost. The work establishes a novel, efficient paradigm for vision modeling grounded in architectural innovation rather than scale.

Improves image classification performance over baseline architecturesReduces computational cost of transformer attention mechanismsReplaces ViT attention with dynamic Network in Network structure

TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters

Oct 30, 2024
HW
Haiyang Wang
🏛️ Max Planck Institute for Informatics | Peking University | Google

Scaling Transformer models is prohibitively expensive due to fixed-parameter linear projection layers; architectural modifications necessitate full retraining. Method: We propose TokenFormer, the first architecture introducing *parameter tokenization*, which models model parameters as learnable tokens and replaces all linear layers with token-parameter self-attention—unifying parameter and input token representations in a shared latent space. Contribution/Results: Our method enables zero-shot, progressive parameter expansion without retraining, overcoming classical scaling bottlenecks. Without altering network topology, we scale model parameters from 124M to 1.4B while matching the performance of fully trained baselines, achieving substantial training cost reduction. The code and models are publicly released.

Dependence on fixed parameters requiring full retrainingHigh computational cost of scaling Transformer modelsLack of efficient progressive scaling for large models

An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels

Jun 13, 2024
DN
Duy-Kien Nguyen
🏛️ University of Amsterdam | FAIR | Meta AI

This work challenges the prevailing consensus that locality-inductive bias is indispensable in vision Transformers. It investigates whether pixel-level tokenization—bypassing conventional patch-based partitioning (e.g., 16×16) and convolutional priors—is both feasible and effective. Method: The authors directly serialize raw images into pixel-level tokens and adopt a standard Transformer architecture, trained via masked autoencoding and diffusion-model paradigms. They comprehensively evaluate performance across image classification, dense prediction, self-supervised reconstruction, and generative modeling. Contribution/Results: Experiments demonstrate that pure pixel-level Transformers achieve competitive or superior performance to state-of-the-art ViTs on multiple benchmarks—including ImageNet-1K, COCO, and ADE20K—without any explicit locality bias. This is the first empirical evidence that locality-inductive bias is not strictly necessary for high visual representation learning. The findings establish a new architectural paradigm and provide theoretical grounding for next-generation vision models grounded in sequence-based, bias-free design principles.

Challenges the necessity of locality bias in vision Transformers.Demonstrates Transformers can perform well using individual pixels as tokens.Explores pixel-level Transformers in classification, autoencoding, and image generation.

TransXNet: Learning Both Global and Local Dynamics with a Dual Dynamic Token Mixer for Visual Recognition

Oct 30, 2023
ML
Meng Lou
🏛️ Deepwise Healthcare | The University of Hong Kong | ShanghaiTech University

To address the limited representational capacity of conventional static convolutions in CNN-Transformer hybrid architectures, this paper proposes the input-adaptive Dual-Dynamic Token Mixer (D-Mixer)—the first to jointly integrate input-driven depthwise separable convolution with lightweight global attention for synergistic modeling of local details and long-range dependencies. D-Mixer enables dynamic cross-module feature fusion and adaptive receptive field expansion, overcoming the inherent limitations of static convolution. Built upon D-Mixer, TransXNet-T achieves a 0.3% top-1 accuracy gain on ImageNet-1K over Swin-T while consuming less than 50% of its FLOPs; its small and base variants attain 83.8% and 84.6% top-1 accuracy, respectively. Moreover, TransXNet demonstrates superior performance on dense prediction tasks—e.g., semantic segmentation and object detection—at significantly lower computational cost, surpassing state-of-the-art methods.

Addresses static convolution limitations in hybrid CNN-Transformer networks.Enhances network performance with reduced computational costs.Proposes Dual Dynamic Token Mixer for global and local dynamics learning.

Latest Papers

What's happening recently
View more

This work addresses the feature learning conflict in Vision Transformers arising from the shared computational pathway for both the global [CLS] token and local patch tokens, which hinders performance in dense prediction tasks. The study reveals, for the first time, that normalization layers implicitly differentiate between these two token types. Building on this insight, the authors propose a lightweight token-specific processing mechanism that decouples their computational flows within the normalization layers and early QKV projections. This approach incurs no additional computational overhead and increases model parameters by only 8%, yet consistently improves performance by over 2 mIoU on standard segmentation benchmarks while preserving strong image classification accuracy.

[CLS] tokendense predictionfeature learning

This work addresses the high computational cost of Transformer-based models in 3D medical image segmentation by proposing the Token-UNet family of lightweight architectures. Building upon a standard 3D convolutional encoder, the approach integrates TokenLearner and TokenFuser modules to efficiently tokenize feature maps, enabling synergistic modeling of both local and global structural information. The resulting framework substantially reduces computational overhead while producing interpretable attention maps. Experimental results demonstrate that the best-performing variant achieves a Dice score of 87.21% ± 0.35%, outperforming SwinUNETR (86.75% ± 0.19%) while reducing memory consumption, inference time, and parameter count to 33%, 10%, and 35% of those of SwinUNETR, respectively.

3D medical imagingbrain segmentationcomputational efficiency

This work challenges the prevailing reliance of Chinese language models on discrete token embeddings by investigating whether effective language modeling can be achieved using only glyph images. To this end, the authors construct a dual-branch controlled framework: one branch rasterizes character sequences into images processed by a visual encoder composed of ResNet and a shallow Vision Transformer (ViT), while the other employs conventional index-based embeddings as a baseline; both branches share an identical decoder to ensure strict variable control. Experiments demonstrate for the first time that pure glyph-based input is not only viable but consistently outperforms the baseline across all decoders—achieving up to 0.429 accuracy (a 21% relative improvement), converging nearly twice as fast, and showing advantages with only 21% of the training data. The approach also exhibits greater robustness to character perturbations, revealing both the modality-agnostic capacity of Transformers and the information-rich structural properties inherent in Chinese characters.

Chinese language modelingglyph imagesinput representation

Hot Scholars

GG

Guangwei Gao

Professor of PCALab@NJUST, IEEE/CCF/CSIG/CAAI/CAA Senior Member
Pattern RecognitionImage UnderstandingMachine Learning
CW

Chaoli Wang

Professor of Computer Science and Engineering, University of Notre Dame
Scientific VisualizationVisual AnalyticsVisualization
HH

Hongyang He

University of Warwick,University of Birmingham,Imperial College London
Machine LearningComputer VisionOptimisation Theory
JZ

Jun Zeng

University of California, Berkeley
Robotics
JL

Jinpeng Lu

University of Science and Technology of China
Biomedical Image ProcessingMultimodal Learning