🤖 AI Summary
While SAM and SAM 2 excel at segmenting context-agnostic objects (e.g., persons, vehicles), they exhibit significant limitations on context-dependent (CD) concepts—such as visual saliency, camouflaged objects, industrial defects, and medical lesions—primarily due to insufficient global-local semantic co-modeling. Method: We introduce the first comprehensive CD benchmark spanning 11 concept categories across natural, medical, and industrial domains, incorporating 2D/3D images and videos. We propose a unified evaluation framework supporting human annotation, automated metrics, and self-prompted interaction, augmented with prompt robustness testing and SAM 2’s in-context learning analysis. Our method further incorporates multi-granularity prompt generation, cross-modal self-prompting, and context-aware evaluation metrics. Contribution/Results: Experiments reveal fundamental architectural bottlenecks in SAM-series models, delivering the first quantitative analysis of CD segmentation performance and establishing an empirical foundation for designing SAM 3.
📝 Abstract
As a foundational model, SAM has significantly influenced multiple fields within computer vision, and its upgraded version, SAM 2, enhances capabilities in video segmentation, poised to make a substantial impact once again. While SAMs (SAM and SAM 2) have demonstrated excellent performance in segmenting context-independent concepts like people, cars, and roads, they overlook more challenging context-dependent (CD) concepts, such as visual saliency, camouflage, product defects, and medical lesions. CD concepts rely heavily on global and local contextual information, making them susceptible to shifts in different contexts, which requires strong discriminative capabilities from the model. The lack of comprehensive evaluation of SAMs limits understanding of their performance boundaries, which may hinder the design of future models. In this paper, we conduct a thorough quantitative evaluation of SAMs on 11 CD concepts across 2D and 3D images and videos in various visual modalities within natural, medical, and industrial scenes. We develop a unified evaluation framework for SAM and SAM 2 that supports manual, automatic, and intermediate self-prompting, aided by our specific prompt generation and interaction strategies. We further explore the potential of SAM 2 for in-context learning and introduce prompt robustness testing to simulate real-world imperfect prompts. Finally, we analyze the benefits and limitations of SAMs in understanding CD concepts and discuss their future development in segmentation tasks. This work aims to provide valuable insights to guide future research in both context-independent and context-dependent concepts segmentation, potentially informing the development of the next version -- SAM 3.