Score
Designs, implements, and evaluates image-based segmentation systems that use transformer architectures (e.g., ViT, Swin, SegFormer) to produce pixel- or region-level labels; develops model architectures, training pipelines, loss functions, and evaluation metrics to optimize boundary delineation, class-wise accuracy, and generalization across visual domains.
Convolutional Neural Networks (CNNs) exhibit limitations in modeling long-range dependencies, multi-scale objects, and contextual information for image segmentation. Method: This paper presents a systematic survey of segmentation-specific Transformer architectures. It introduces a unified architectural taxonomy covering encoder-decoder designs, multi-scale feature fusion strategies, mask prediction heads, and attention variants—including windowed and axial attention. A “challenge–solution” mapping framework is established to identify three key bottlenecks: computational overhead, poor generalization under low-data regimes, and constraints on real-time deployment. Contribution/Results: The survey proposes two principal evolutionary directions—lightweight design and data-efficient learning—and establishes a unified evaluation benchmark to delineate state-of-the-art performance boundaries. Collectively, this work provides a principled, industrially viable roadmap for deploying Transformer-based segmentation systems.
To address the high computational overhead and memory bottlenecks hindering Vision Transformer (ViT) deployment on edge devices, this paper presents a systematic survey of lightweighting and acceleration techniques tailored for edge scenarios—spanning model compression (e.g., pruning, quantization, knowledge distillation, attention simplification), software optimization (e.g., compiler frameworks such as TVM), and hardware adaptation (e.g., GPU/TPU/FPGA mapping). Its key contributions include: (1) proposing the first unified taxonomy for ViT edge deployment, explicitly characterizing trade-offs among accuracy, latency, power consumption, and hardware platforms; (2) establishing a structured evaluation framework covering 120+ works to identify real-world deployment bottlenecks; and (3) delivering a reproducible, cross-platform technical selection guide to advance co-optimization of accuracy, latency, and power efficiency.
This work addresses the challenge of achieving efficient image segmentation using a pure Vision Transformer (ViT) encoder—without task-specific modules such as convolutional adapters, pixel decoders, or Transformer decoders. We propose the Encoder-only Mask Transformer (EoMT), which leverages large-scale ViTs (e.g., ViT-L) and employs joint self-supervised and supervised pretraining, coupled with only a lightweight mask prediction head and simple upsampling decoder. Our key finding is the first empirical demonstration that a sufficiently pretrained ViT encoder inherently acquires multi-scale inductive biases essential for segmentation. Evaluated on benchmarks including ADE20K, EoMT achieves state-of-the-art accuracy while accelerating inference by up to 4× over prior methods. It thus establishes a new Pareto-optimal trade-off between segmentation accuracy and computational efficiency.
To address the high computational cost and memory overhead of Transformer models in medical image semantic segmentation, this work systematically investigates the application of key-value (KV)-only attention mechanisms in both pure Transformer and CNN-Transformer hybrid architectures. By eliminating the query (Q) branch and retaining only the Key-Value pathways, self-attention computation is significantly simplified. Experiments across multiple medical segmentation benchmarks demonstrate that the KV variant reduces model parameters by over 40% and multiply-accumulate operations (MACs) by approximately 35%, while maintaining mIoU performance comparable to standard QKV-based models. This study not only validates the effectiveness of KV attention for pixel-wise dense prediction tasks but also reveals its practical advantages in the accuracy-efficiency trade-off. The proposed approach establishes a novel paradigm for lightweight medical image segmentation, offering a principled path toward resource-efficient yet accurate deep models.
While the Segment Anything Model (SAM) exhibits strong generalization capability, its substantial computational overhead hinders deployment on resource-constrained edge devices. This work presents a systematic survey of efficient SAM variants tailored for edge deployment. We introduce the first unified evaluation framework spanning diverse hardware platforms—including CPU, GPU, and Edge TPU—and conduct joint accuracy–latency–memory benchmarking on COCO and SA-1B. Our analysis categorizes acceleration techniques along six technical axes: model pruning, knowledge distillation, lightweight attention mechanisms, quantization, module substitution, and hardware-aware compilation—characterizing their Pareto-optimal trade-offs. The core contributions are: (1) an open-source, fully reproducible edge-SAM benchmark; and (2) empirical insights into the applicability domains and fundamental accuracy-efficiency trade-offs of each acceleration strategy—providing both theoretical foundations and practical guidelines for designing lightweight vision foundation models.
To address performance limitations in object detection and semantic segmentation under complex scenarios—including occlusion, small objects, and cross-domain generalization—this paper proposes a novel multimodal detection paradigm synergizing large language models (LLMs). Methodologically, it systematically integrates CNNs, YOLOv5/v8, and DETR architectures into an LLM-augmented inference framework, augmented by scalable data pipelines, model pruning, and quantization, and evaluated via a multi-dimensional metric system based on mAP and mIoU. Key contributions include: (1) bridging the gap between traditional feature engineering and end-to-end deep learning; (2) introducing a dynamic context enhancement mechanism tailored for challenging environments; and (3) achieving state-of-the-art accuracy-efficiency trade-offs on COCO and ADE20K. The fully open-sourced, reproducible framework significantly improves model generalizability and robustness across diverse real-world conditions.
This study addresses the challenge that the effects of module-level features and their interactions on generalization in Vision Transformer (ViT) architectures remain insufficiently isolated and quantified. By analyzing ViT representational structures, we identify feature collapse during initialization and propose quantitative metrics based on feature entropy and minimum eigenvalues. Our analysis reveals the critical roles of token spaces and linear submodules in generalization, systematically validated through multi-scale architectural comparison experiments. Results demonstrate that the proposed surrogate metrics improve correlation rankings with generalization performance by 18%–48%, enabling precise identification of low-compute, high-accuracy ViT architectures. This work provides both theoretical foundations and practical guidance for efficient visual model design.
To address the high cost of pixel-level annotations in food image semantic segmentation, this paper proposes a weakly supervised approach that generates high-quality food region masks using only image-level labels. The method leverages a Swin Transformer to produce Class Activation Maps (CAMs), which are automatically converted into point and bounding-box prompts for the Segment Anything Model (SAM). Additionally, an image-adaptive preprocessing pipeline and a multi-mask fusion strategy are introduced to enhance segmentation robustness. Evaluated on the FoodSeg103 dataset, the method yields an average of 2.4 valid masks per image, achieving 0.54 mIoU with the multi-mask scheme—substantially outperforming existing weakly supervised methods. This work is the first to synergistically integrate ViT-based CAMs with SAM’s prompting mechanism for weakly supervised food segmentation, enabling high-accuracy, scalable segmentation without any pixel-level annotations and establishing a novel paradigm for real-world food analysis.
This work investigates whether medical image segmentation still necessitates U-Net–style decoders when built upon modern Vision Transformer (ViT) backbones. To this end, we propose EoSeg, the first pure encoder-based segmentation architecture tailored for multimodal medical imaging, which departs from the conventional U-Net paradigm. EoSeg achieves end-to-end dense prediction through a multi-level query mechanism and a learnable feature fusion module. Extensive experiments across seven medical imaging benchmarks demonstrate that EoSeg significantly outperforms existing methods, achieving state-of-the-art results on datasets such as Synapse (mDice 85.50%), ACDC (91.73%), and GlaS (93.27%).
Convolutional neural networks (CNNs) struggle to model global contextual dependencies in images. Method: This work systematically evaluates pure vision transformers (ViTs), hierarchical ViTs (e.g., Swin, CvT), and hybrid architectures across image classification (ImageNet), object detection (COCO), and medical image classification (ChestX-ray14), benchmarking all against unified CNN baselines and quantifying accuracy–efficiency trade-offs. Contribution/Results: It presents the first cross-task, cross-domain comparative evaluation of diverse ViT families and introduces a medical imaging–specific data augmentation strategy. Hierarchical ViTs consistently outperform CNNs: achieving +3.2% average AUC on ChestX-ray14 and 51.7% mAP on COCO. Results empirically validate the structural advantage of self-attention for modeling long-range dependencies, providing evidence-based guidance for deploying ViTs in both clinical and general-purpose vision applications.
Low-light conditions and occlusions severely degrade thermal imaging-based weapon segmentation accuracy. Method: This paper pioneers the systematic application of Vision Transformers (ViTs) to this task, introducing the first large-scale thermal weapon segmentation dataset—comprising 9,711 real-world surveillance images—and conducting a rigorous comparative evaluation of SegFormer, Swin Transformer, SegNeXt, and DeepLabV3+ within the MMSegmentation framework. To enhance data quality and generalization, SAM2-assisted annotation and standardized augmentation strategies are integrated. Results: SegFormer-b5 achieves state-of-the-art performance with 94.15% mIoU and 97.04% pixel accuracy; SegFormer-b0 attains real-time inference at 98.32 FPS. ViT-based architectures consistently outperform CNNs, demonstrating superior capability in modeling long-range dependencies and capturing fine-grained structural details. This work establishes a new high-accuracy, high-efficiency segmentation paradigm for thermal-imaging-based security applications.