hybrid transformer-unet

Design and build encoder–decoder neural architectures that integrate transformer-based attention modules (e.g., Swin) with U‑Net style convolutional encoders/decoders to fuse attention and convolutional representations for spatial prediction. These hybrid models address image and optionally spatio‑temporal inputs (e.g., video) by augmenting or replacing U‑Net blocks with transformers and are trained end-to-end to encode spatio‑temporal features and decode spatial outputs.

hybridtransformer-unet

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.58
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Convolutional Rectangular Attention Module

Mar 13, 2025
HN
Hai-Vy Nguyen
🏛️ Ampere Software Technology | Institut de mathématiques de Toulouse | Institut de Recherche en Informatique de Toulouse | Université Côte d'Azur

This paper addresses the poor generalizability, training instability, and weak interpretability of conventional spatial attention mechanisms in convolutional neural networks (CNNs), which stem from irregular, pixel-level attention regions. To this end, we propose a parametric Rectangular Spatial Attention Module (RSAM) that explicitly defines a rectangular attention region using only five learnable parameters. RSAM is fully differentiable and enables end-to-end joint optimization, serving as a plug-and-play component compatible with arbitrary CNN architectures. Our key contribution is the first explicit geometric constraint of spatial attention to a rectangle—enhancing boundary regularity, training stability, and cross-sample generalization, while improving semantic interpretability of attended locations. Extensive experiments on multiple benchmarks demonstrate that RSAM consistently outperforms pixel-wise attention methods, achieving significant gains in classification accuracy, robustness to input perturbations, and visual localization consistency.

Enhances generalization with rectangular attention regions using 5 parameters.Improves model performance by focusing on discriminative image parts.Introduces a spatial attention module for convolutional networks.

Demystify Transformers & Convolutions in Modern Image Deep Networks

Nov 10, 2022
JD
Jifeng Dai
🏛️ Tsinghua University | Shanghai Artificial Intelligence Laboratory | Huazhong University of Science and Technology | Fudan University | The Chinese University of Hong Kong | SenseTime Research | South China University of Technology

This work investigates the fundamental differences between spatial token mixers (STMs)—the spatial feature aggregation mechanisms—in Vision Transformers and convolutional networks. To enable a fair, architecture-agnostic comparison, we propose a unified STM modeling paradigm that decouples network-level design from the spatial aggregation module, implementing both convolutional and attention-based STMs on a neutral backbone. Our methodology includes: (1) designing a modular, swappable STM interface; (2) systematically analyzing inductive biases—including receptive field size, translation invariance, and adversarial robustness; and (3) conducting multi-task performance benchmarking. Results show that while modern network-level designs yield substantial gains, intrinsic performance gaps among STMs persist. Crucially, we quantitatively demonstrate for the first time that convolutions exhibit superior translation invariance and local robustness, whereas attention achieves larger effective receptive fields but is more vulnerable to input perturbations.

Analyze performance differences in attention vs convolutionCompare spatial token mixers in vision backbonesUnify architecture to isolate feature transformation effects

Spatio-Temporal Graph Convolutional Networks: Optimised Temporal Architecture

Jan 14, 2025
ET
Edward Turner
🏛️ University of Oxford

Traditional ST-GCNs employ单一 temporal modules—either CNNs or LSTMs—leading to insufficient capture of dynamic spatiotemporal patterns. To address this, we propose a plug-and-play hybrid temporal module that, for the first time, synergistically integrates CNNs and LSTMs within a unified co-temporal block. This design jointly models local temporal features and long-range dependencies. Through theoretical analysis and cross-dataset ablation studies, we systematically characterize the intrinsic relationship between temporal module architecture and representational capacity. Evaluated on standard spatiotemporal graph benchmarks—including NTU-RGB+D and PeMSD7—our method achieves significant improvements in prediction accuracy and cross-domain generalization. It consistently outperforms pure-CNN and pure-LSTM baselines in temporal representation learning. The proposed module establishes a reusable, principled design paradigm for temporal modeling in ST-GCNs, advancing both expressiveness and architectural flexibility.

EfficiencySpatial Graph Convolutional NetworksTime Series Data

Spatiotemporal Tile-based Attention-guided LSTMs for Traffic Video Prediction

Oct 24, 2019
TN
Tu Nguyen
🏛️ Daimler Autonomous Services | Daimler AG

Addressing the challenge of jointly modeling fine-grained (pixel-level) and coarse-grained (regional-level) spatial structures while preserving long-range temporal dependencies in traffic video forecasting, this paper proposes a tile-based spatial attention-enhanced encoder-decoder LSTM framework. Our key contributions are: (1) a tile-aware spatial attention mechanism that explicitly captures multi-scale spatial hierarchies; (2) an attention-guided LSTM cell integrating convolutional features to improve trajectory modeling fidelity; and (3) an adaptive frame-sampling strategy coupled with an end-to-end trainable encoder-decoder architecture, balancing computational efficiency and robustness. Evaluated on the Traffic4Cast 2019 dataset, our method significantly outperforms both 2D/3D CNNs and standard ConvLSTM baselines, achieving superior prediction accuracy and enhanced spatiotemporal consistency.

Achieve scalable memory-accuracy tradeoffs for large mapsModel fine-grained and coarse spatial traffic structuresPreserve temporal relationships across long video sequences

Continuum Attention for Neural Operators

Jun 10, 2024
EC
E. Calvello
🏛️ California Institute of Technology | NVIDIA | The Broad Institute of MIT and Harvard

Existing attention mechanisms operate on discrete sequences, limiting their applicability to continuous function spaces essential for scientific machine learning tasks such as PDE solving and physical simulation. Method: This work generalizes attention to continuous function spaces by introducing the Transformer Neural Operator (TNO), the first rigorously defined attention mechanism on functions. It establishes a mathematically sound formulation of functional attention and proposes a patching-based continuous attention mechanism coupled with an efficient discretization strategy to mitigate computational complexity in high dimensions. Contribution/Results: TNO is proven to be a universal approximator for arbitrary continuous operators. Experiments demonstrate that it significantly outperforms state-of-the-art neural operators across diverse PDE benchmarks and physics-informed simulation tasks, validating its effectiveness, scalability, and generalization capability in scientific machine learning.

Developing efficient attention for multidimensional function domainsExtending attention mechanism to function space mappingsProving universal approximation for transformer neural operators

Latest Papers

What's happening recently
View more

This work addresses functional mismatch and redundancy in the attention mechanisms of current large vision-language models, which fail to efficiently exploit visual context. By establishing a unified framework grounded in information theory and information geometry, the study quantifies the geometric structure and entropy characteristics of residual updates, revealing a functional decoupling between attention mechanisms and feed-forward networks (FFNs) in subspace operations. For the first time from an information-geometric perspective, it clarifies their distinct intrinsic roles and demonstrates that attention can be replaced by predefined weights—such as those derived from Gaussian noise—without performance degradation. Empirical results show that this simplified model matches or even surpasses the original architecture across multiple benchmarks, challenging the prevailing design paradigm reliant on dynamic attention and confirming its substantial redundancy.

Attention MechanismLarge Vision-Language ModelsModel Redundancy

This work addresses a key limitation in existing feed-forward neural view synthesis (NVS) Transformers, where semantic and spatial information are entangled within a shared feature space, causing spatial bias to interfere with appearance representation and degrade rendering fidelity. To resolve this, the authors propose a semantics-spatial disentangled architecture that explicitly separates feature representations into independent branches while enabling efficient cross-branch interaction through shared attention routing. Additionally, they introduce optional classification supervision and a bidirectional modulation mechanism to enhance representational capacity with negligible impact on inference latency. The proposed approach consistently improves performance across both decoder-only and encoder-decoder variants of feed-forward NVS models, yielding significantly higher rendering quality.

novel view synthesisrendering fidelityrepresentation ambiguity

This work addresses the challenge of detecting AI-generated videos from high-fidelity models such as Sora2 and Veo3, which current methods struggle to handle due to their reliance on shallow features or computationally expensive multimodal architectures. We propose EA-Swin, an embedding-agnostic Swin Transformer that leverages a factorized window attention mechanism to directly model spatiotemporal dependencies in pretrained video embeddings, offering compatibility with any Vision Transformer (ViT)-style encoder. To support comprehensive evaluation, we introduce EA-Video, a new benchmark comprising 130,000 videos, enabling the first unified assessment of generalization across diverse ViT-based embeddings and unseen generators. EA-Swin achieves state-of-the-art accuracy of 0.97–0.99 on mainstream generators, outperforming existing methods by 5%–20%, and demonstrates strong generalization to previously unseen video synthesis models.

AI-generated video detectiondeepfake detectionfoundation video generators

This work addresses the high computational cost of existing hybrid CNN-Transformer super-resolution models when scaling receptive fields, which hinders deployment on resource-constrained devices. The authors propose UCAN, a lightweight architecture that unifies convolution and attention mechanisms to efficiently model both local textures and long-range dependencies. Key innovations include an integrated window-based spatial attention combined with Hedgehog Attention, a distilled large-kernel convolution module, and a cross-layer parameter sharing strategy, collectively reducing computational complexity. UCAN achieves a PSNR of 31.63 dB on Manga109 (4×) with only 48.4G MACs and 27.79 dB on BSDS100, demonstrating superior accuracy-efficiency trade-offs compared to most existing methods while using a smaller model footprint.

computational costlightweightreceptive field

This work addresses the high computational cost of Transformer-based models in 3D medical image segmentation by proposing the Token-UNet family of lightweight architectures. Building upon a standard 3D convolutional encoder, the approach integrates TokenLearner and TokenFuser modules to efficiently tokenize feature maps, enabling synergistic modeling of both local and global structural information. The resulting framework substantially reduces computational overhead while producing interpretable attention maps. Experimental results demonstrate that the best-performing variant achieves a Dice score of 87.21% ± 0.35%, outperforming SwinUNETR (86.75% ± 0.19%) while reducing memory consumption, inference time, and parameter count to 33%, 10%, and 35% of those of SwinUNETR, respectively.

3D medical imagingbrain segmentationcomputational efficiency

Hot Scholars

ZH

Zhihai He

Southern University of Science and Technology
Deep learningcomputer visionmachine learningsmart cyber-physical systems
YY

Yongquan Yang

Dr of Computer Science, Ocean University of China
cloud computingpervasive computingmulti-touch
XX

Xiao Xue

Tianjin University
Computational ExperimentService EcosystemAgent Based SimulationCrowd Intelligence
IH

Ilker Hacihaliloglu

Department of Radiology, Department of Medicine, University of British Columbia
Biomedical EngineeringMedical Image ProcessingUltrasound Image ProcessingImage Guided Surgery and Therapy