transformer-based segmentation

Designs, implements, and evaluates image-based segmentation systems that use transformer architectures (e.g., ViT, Swin, SegFormer) to produce pixel- or region-level labels; develops model architectures, training pipelines, loss functions, and evaluation metrics to optimize boundary delineation, class-wise accuracy, and generalization across visual domains.

transformer-basedsegmentation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

To address the high computational overhead and memory bottlenecks hindering Vision Transformer (ViT) deployment on edge devices, this paper presents a systematic survey of lightweighting and acceleration techniques tailored for edge scenarios—spanning model compression (e.g., pruning, quantization, knowledge distillation, attention simplification), software optimization (e.g., compiler frameworks such as TVM), and hardware adaptation (e.g., GPU/TPU/FPGA mapping). Its key contributions include: (1) proposing the first unified taxonomy for ViT edge deployment, explicitly characterizing trade-offs among accuracy, latency, power consumption, and hardware platforms; (2) establishing a structured evaluation framework covering 120+ works to identify real-world deployment bottlenecks; and (3) delivering a reproducible, cross-platform technical selection guide to advance co-optimization of accuracy, latency, and power efficiency.

High computational complexity of vision transformers.Lack of comprehensive review on model compression.Memory demands for edge device deployment.

Your ViT is Secretly an Image Segmentation Model

Mar 24, 2025
TK
Tommie Kerssies
🏛️ Eindhoven University of Technology | Polytechnic of Turin | RWTH Aachen University

This work addresses the challenge of achieving efficient image segmentation using a pure Vision Transformer (ViT) encoder—without task-specific modules such as convolutional adapters, pixel decoders, or Transformer decoders. We propose the Encoder-only Mask Transformer (EoMT), which leverages large-scale ViTs (e.g., ViT-L) and employs joint self-supervised and supervised pretraining, coupled with only a lightweight mask prediction head and simple upsampling decoder. Our key finding is the first empirical demonstration that a sufficiently pretrained ViT encoder inherently acquires multi-scale inductive biases essential for segmentation. Evaluated on benchmarks including ADE20K, EoMT achieves state-of-the-art accuracy while accelerating inference by up to 4× over prior methods. It thus establishes a new Pareto-optimal trade-off between segmentation accuracy and computational efficiency.

Achieve high segmentation accuracy with large-scale models and pre-trainingOptimize balance between segmentation accuracy and prediction speedRepurpose plain ViT for image segmentation without task-specific components

To address the high computational cost and memory overhead of Transformer models in medical image semantic segmentation, this work systematically investigates the application of key-value (KV)-only attention mechanisms in both pure Transformer and CNN-Transformer hybrid architectures. By eliminating the query (Q) branch and retaining only the Key-Value pathways, self-attention computation is significantly simplified. Experiments across multiple medical segmentation benchmarks demonstrate that the KV variant reduces model parameters by over 40% and multiply-accumulate operations (MACs) by approximately 35%, while maintaining mIoU performance comparable to standard QKV-based models. This study not only validates the effectiveness of KV attention for pixel-wise dense prediction tasks but also reveals its practical advantages in the accuracy-efficiency trade-off. The proposed approach establishes a novel paradigm for lightweight medical image segmentation, offering a principled path toward resource-efficient yet accurate deep models.

Comparing KV and QKV Transformers' performance and complexityEvaluating KV Transformers for medical image segmentationReducing model complexity while maintaining segmentation accuracy

On Efficient Variants of Segment Anything Model: A Survey

Oct 07, 2024
XS
Xiaorui Sun
🏛️ UESTC | Lancaster University | Tongji University

While the Segment Anything Model (SAM) exhibits strong generalization capability, its substantial computational overhead hinders deployment on resource-constrained edge devices. This work presents a systematic survey of efficient SAM variants tailored for edge deployment. We introduce the first unified evaluation framework spanning diverse hardware platforms—including CPU, GPU, and Edge TPU—and conduct joint accuracy–latency–memory benchmarking on COCO and SA-1B. Our analysis categorizes acceleration techniques along six technical axes: model pruning, knowledge distillation, lightweight attention mechanisms, quantization, module substitution, and hardware-aware compilation—characterizing their Pareto-optimal trade-offs. The core contributions are: (1) an open-source, fully reproducible edge-SAM benchmark; and (2) empirical insights into the applicability domains and fundamental accuracy-efficiency trade-offs of each acceleration strategy—providing both theoretical foundations and practical guidelines for designing lightweight vision foundation models.

Addressing high computational demands of Segment Anything ModelEnhancing SAM efficiency for resource-limited environmentsSurveying acceleration techniques for SAM variants

Deep Learning and Machine Learning - Object Detection and Semantic Segmentation: From Theory to Applications

Oct 21, 2024
JR
Jintao Ren
🏛️ Aarhus University | Indiana University | Kyoto University | AppCubic | Rutgers University | University of Wisconsin-Madison | Georgia Institute of Technology | National Taiwan Normal University | University of Hawaii | Xi'an Jiaotong-Liverpool University | Zhejiang University | Purdue University

To address performance limitations in object detection and semantic segmentation under complex scenarios—including occlusion, small objects, and cross-domain generalization—this paper proposes a novel multimodal detection paradigm synergizing large language models (LLMs). Methodologically, it systematically integrates CNNs, YOLOv5/v8, and DETR architectures into an LLM-augmented inference framework, augmented by scalable data pipelines, model pruning, and quantization, and evaluated via a multi-dimensional metric system based on mAP and mIoU. Key contributions include: (1) bridging the gap between traditional feature engineering and end-to-end deep learning; (2) introducing a dynamic context enhancement mechanism tailored for challenging environments; and (3) achieving state-of-the-art accuracy-efficiency trade-offs on COCO and ADE20K. The fully open-sourced, reproducible framework significantly improves model generalizability and robustness across diverse real-world conditions.

Bridging traditional methods with modern AI for large-scale detection tasksExploring object detection and semantic segmentation from theory to applicationsReviewing state-of-the-art deep learning architectures for computer vision

Latest Papers

What's happening recently
View more

This study addresses the challenge that the effects of module-level features and their interactions on generalization in Vision Transformer (ViT) architectures remain insufficiently isolated and quantified. By analyzing ViT representational structures, we identify feature collapse during initialization and propose quantitative metrics based on feature entropy and minimum eigenvalues. Our analysis reveals the critical roles of token spaces and linear submodules in generalization, systematically validated through multi-scale architectural comparison experiments. Results demonstrate that the proposed surrogate metrics improve correlation rankings with generalization performance by 18%–48%, enabling precise identification of low-compute, high-accuracy ViT architectures. This work provides both theoretical foundations and practical guidance for efficient visual model design.

Architecture DesignFeature InformationGeneralization Behavior

Weakly Supervised Food Image Segmentation using Vision Transformers and Segment Anything Model

Sep 23, 2025
IS
Ioannis Sarafis
🏛️ Aristotle University of Thessaloniki

To address the high cost of pixel-level annotations in food image semantic segmentation, this paper proposes a weakly supervised approach that generates high-quality food region masks using only image-level labels. The method leverages a Swin Transformer to produce Class Activation Maps (CAMs), which are automatically converted into point and bounding-box prompts for the Segment Anything Model (SAM). Additionally, an image-adaptive preprocessing pipeline and a multi-mask fusion strategy are introduced to enhance segmentation robustness. Evaluated on the FoodSeg103 dataset, the method yields an average of 2.4 valid masks per image, achieving 0.54 mIoU with the multi-mask scheme—substantially outperforming existing weakly supervised methods. This work is the first to synergistically integrate ViT-based CAMs with SAM’s prompting mechanism for weakly supervised food segmentation, enabling high-accuracy, scalable segmentation without any pixel-level annotations and establishing a novel paradigm for real-world food analysis.

Improving food mask quality through preprocessing and multi-mask strategiesLeveraging Vision Transformers and SAM for zero-shot food segmentationSegmenting food images using weak supervision without pixel-level annotations

This work investigates whether medical image segmentation still necessitates U-Net–style decoders when built upon modern Vision Transformer (ViT) backbones. To this end, we propose EoSeg, the first pure encoder-based segmentation architecture tailored for multimodal medical imaging, which departs from the conventional U-Net paradigm. EoSeg achieves end-to-end dense prediction through a multi-level query mechanism and a learnable feature fusion module. Extensive experiments across seven medical imaging benchmarks demonstrate that EoSeg significantly outperforms existing methods, achieving state-of-the-art results on datasets such as Synapse (mDice 85.50%), ACDC (91.73%), and GlaS (93.27%).

dense predictionencoder-only architecturemedical image segmentation

Hands-on Evaluation of Visual Transformers for Object Recognition and Detection

Dec 10, 2025
DN
Dimitrios N. Vlachogiannis
🏛️ University of Patras

Convolutional neural networks (CNNs) struggle to model global contextual dependencies in images. Method: This work systematically evaluates pure vision transformers (ViTs), hierarchical ViTs (e.g., Swin, CvT), and hybrid architectures across image classification (ImageNet), object detection (COCO), and medical image classification (ChestX-ray14), benchmarking all against unified CNN baselines and quantifying accuracy–efficiency trade-offs. Contribution/Results: It presents the first cross-task, cross-domain comparative evaluation of diverse ViT families and introduces a medical imaging–specific data augmentation strategy. Hierarchical ViTs consistently outperform CNNs: achieving +3.2% average AUC on ChestX-ray14 and 51.7% mAP on COCO. Results empirically validate the structural advantage of self-attention for modeling long-range dependencies, providing evidence-based guidance for deploying ViTs in both clinical and general-purpose vision applications.

Assesses performance trade-offs between accuracy and computational efficiencyCompares Vision Transformers and CNNs for object recognition and detection tasksEvaluates transformer models on standard and medical imaging datasets

Low-light conditions and occlusions severely degrade thermal imaging-based weapon segmentation accuracy. Method: This paper pioneers the systematic application of Vision Transformers (ViTs) to this task, introducing the first large-scale thermal weapon segmentation dataset—comprising 9,711 real-world surveillance images—and conducting a rigorous comparative evaluation of SegFormer, Swin Transformer, SegNeXt, and DeepLabV3+ within the MMSegmentation framework. To enhance data quality and generalization, SAM2-assisted annotation and standardized augmentation strategies are integrated. Results: SegFormer-b5 achieves state-of-the-art performance with 94.15% mIoU and 97.04% pixel accuracy; SegFormer-b0 attains real-time inference at 98.32 FPS. ViT-based architectures consistently outperform CNNs, demonstrating superior capability in modeling long-range dependencies and capturing fine-grained structural details. This work establishes a new high-accuracy, high-efficiency segmentation paradigm for thermal-imaging-based security applications.

Adapting Vision Transformers for thermal weapon segmentation tasksAddressing limited long-range dependency capture in thermal CNNsEvaluating transformer architectures on real-world thermal surveillance data

Hot Scholars

HK

Hitoshi Kiya

Professor Emeritus, Tokyo Metropolitan University
Signal ProcessingComputer VisionMachine LearningInformation Security
JH

Jiacheng Hu

Tulane University
Deep LearningComputer VisionNatural language processingAI in Medicine
AK

Anil K. Jain

Michigan State University
BiometricsComputer visionPattern recognitionMachine learning
RK

Ranjay Krishna

University of Washington, Allen Institute for AI
Computer VisionNatural Language ProcessingMachine LearningHuman Computer Interaction