🤖 AI Summary
Convolutional Neural Networks (CNNs) exhibit limitations in modeling long-range dependencies, multi-scale objects, and contextual information for image segmentation. Method: This paper presents a systematic survey of segmentation-specific Transformer architectures. It introduces a unified architectural taxonomy covering encoder-decoder designs, multi-scale feature fusion strategies, mask prediction heads, and attention variants—including windowed and axial attention. A “challenge–solution” mapping framework is established to identify three key bottlenecks: computational overhead, poor generalization under low-data regimes, and constraints on real-time deployment. Contribution/Results: The survey proposes two principal evolutionary directions—lightweight design and data-efficient learning—and establishes a unified evaluation benchmark to delineate state-of-the-art performance boundaries. Collectively, this work provides a principled, industrially viable roadmap for deploying Transformer-based segmentation systems.
📝 Abstract
Image segmentation, a key task in computer vision, has traditionally relied on convolutional neural networks (CNNs), yet these models struggle with capturing complex spatial dependencies, objects with varying scales, need for manually crafted architecture components and contextual information. This paper explores the shortcomings of CNN-based models and the shift towards transformer architectures -to overcome those limitations. This work reviews state-of-the-art transformer-based segmentation models, addressing segmentation-specific challenges and their solutions. The paper discusses current challenges in transformer-based segmentation and outlines promising future trends, such as lightweight architectures and enhanced data efficiency. This survey serves as a guide for understanding the impact of transformers in advancing segmentation capabilities and overcoming the limitations of traditional models.