Score
Designs and trains deep semantic segmentation systems by applying transfer learning from pretrained encoder/backbone networks, adapting and fine-tuning feature extractors, architectures, and training regimes to produce accurate pixelwise labels. This work builds and evaluates encoder–decoder style models, backbone adaptations, loss and augmentation strategies, and optimization choices to improve boundary/seam continuity and to reduce catastrophic failures such as severe zero-IoU predictions.
Addressing the few-shot semantic segmentation (FSS) challenge in the era of foundation models, this work introduces the first benchmark specifically designed for adapting large-scale vision models to FSS. We systematically evaluate five representative models—DINO v2, SAM, CLIP, MAE, and ResNet50-COCO—alongside five adaptation strategies: linear probing, LoRA, feature distillation, prompt tuning, and full fine-tuning. Notably, this is the first comprehensive evaluation of both multimodal and unimodal vision foundation models in FSS. Our results reveal that DINO v2 substantially outperforms all others (achieving an average mIoU 8.2 percentage points higher than the second-best on Pascal-5i and COCO-20i), and linear probing alone attains 97.3% of full fine-tuning performance—drastically reducing computational overhead. These findings challenge prevailing assumptions about the necessity of complex adaptation mechanisms and establish a new empirical baseline and practical guidance for integrating foundation models with FSS.
To address performance limitations in object detection and semantic segmentation under complex scenarios—including occlusion, small objects, and cross-domain generalization—this paper proposes a novel multimodal detection paradigm synergizing large language models (LLMs). Methodologically, it systematically integrates CNNs, YOLOv5/v8, and DETR architectures into an LLM-augmented inference framework, augmented by scalable data pipelines, model pruning, and quantization, and evaluated via a multi-dimensional metric system based on mAP and mIoU. Key contributions include: (1) bridging the gap between traditional feature engineering and end-to-end deep learning; (2) introducing a dynamic context enhancement mechanism tailored for challenging environments; and (3) achieving state-of-the-art accuracy-efficiency trade-offs on COCO and ADE20K. The fully open-sourced, reproducible framework significantly improves model generalizability and robustness across diverse real-world conditions.
To address the challenges of detail loss and high computational overhead in few-shot semantic segmentation, this paper proposes an efficient and accurate Transformer-based approach. Methodologically, it introduces (1) a novel spatial Transformer decoder coupled with a context-aware mask generation module; (2) a multi-scale hierarchical decoding mechanism that fuses intermediate-layer global features to enhance fine-grained localization; and (3) prototype-guided multi-scale feature pyramid decoding, integrated with spatial-attention-driven support-query relational modeling. With only 1.5M parameters, the model achieves state-of-the-art performance on both PASCAL-5ⁱ and COCO-20ⁱ benchmarks under 1-shot and 5-shot settings. It simultaneously delivers superior accuracy, strong generalization across diverse base/novel classes, and high inference efficiency—outperforming all existing methods while maintaining architectural compactness and computational tractability.
Deep models for image classification heavily rely on large-scale labeled data, yet real-world scenarios often suffer from data scarcity, making transfer learning a critical solution—yet systematic surveys and theoretical frameworks remain lacking. This paper proposes the first unified taxonomy for deep transfer learning in image classification, formally defining the problem, identifying core challenges (e.g., source/target domain shift, limited target-sample size), and characterizing failure boundaries. It systematically integrates both CNN- and Transformer-based architectures, covering major paradigms: feature extraction, fine-tuning, domain adaptation, and meta-transfer learning. Through cross-paradigm comparative analysis, it uncovers intrinsic relationships between model performance, data characteristics (e.g., domain gap, sample scale), and architectural/algorithmic choices. The study clarifies current research gaps, establishes necessary conditions for effective transfer, and delivers a reusable methodological framework for few-shot and cross-domain image classification.
To address the degradation of semantic segmentation model generalization under visual condition shifts (e.g., day/night, weather), this paper proposes a feature-level domain adaptation method grounded in feature invariance. Leveraging image style transfer as an intermediary, the approach aligns feature representations between stylized and original images at the encoder level, thereby disentangling style variation from semantic structure and enabling condition-agnostic semantic understanding. Its core innovation is the first feature-level invariance loss function explicitly designed for semantic segmentation. Built upon a state-of-the-art unsupervised domain adaptation framework, the method requires no target-domain annotations. Experiments demonstrate that it achieves state-of-the-art performance on Cityscapes→Dark Zurich and ranks second on Cityscapes→ACDC. Moreover, it exhibits strong zero-shot transfer capability to unseen domains—including BDD100K-night and ACDC-night—without fine-tuning.
Existing methods struggle to provide clear and generalizable explanations of self-supervised vision models, particularly in distinguishing behavioral differences between contrastive learning and masked image modeling. This work proposes a visualization protocol based on unsupervised semantic segmentation, introduced for the first time as an interpretability tool, to intuitively reveal positional bias, locality bias, and scaling properties across different layers and representations of such models. The approach effectively disentangles positional effects from local structural patterns and uncovers novel phenomena—such as boundary artifacts in DINOv3-Large—offering a unified, intuitive, and reproducible visual framework for understanding diverse self-supervised pretraining paradigms.
This work addresses the challenge of insufficient accuracy in segmenting fine structures and object boundaries—particularly in resource-constrained settings—by proposing FoR-Net, a lightweight architecture for semantic segmentation. FoR-Net incorporates a region-focused inductive bias that selectively enhances information-rich and challenging regions through importance map prediction and a Top-K activation mechanism. It further aggregates multi-scale spatial context via parallel convolutional branches. Evaluated on the Cityscapes benchmark under standard training protocols and limited computational resources, FoR-Net achieves competitive overall performance while significantly improving segmentation consistency in difficult regions.
This work addresses the high computational cost and overfitting issues inherent in training-dependent approaches for cross-domain few-shot segmentation, as well as the limited or even degraded performance observed when integrating current vision foundation models. To overcome these challenges, the paper introduces the first fully training-free segmentation framework. Built upon the DINOv3 self-supervised encoder, the method leverages three key components—Semantic-Aware Feature Refusion (SAFR), Adaptive Support Enhancement (ASE), and Hybrid Prototype Matching (HPM)—to enable effective segmentation of unseen categories. Evaluated on four target-domain datasets, the proposed approach achieves state-of-the-art performance, substantially improving cross-domain generalization and demonstrating the efficacy and robustness of a training-free paradigm in few-shot segmentation.
This study addresses the challenge of spatial misalignment between remote sensing imagery and external labels—such as those from OpenStreetMap—which hinders effective training of building segmentation models. To overcome this, the authors propose Align and Segment (AnS), an unsupervised framework that jointly performs high-quality building segmentation and automatic image-to-label alignment without requiring precisely registered annotations. The method employs a differentiable spatial transformation module to affinely align labels and incorporates a self-supervised regularization loss to prevent shortcut learning. This work is the first to simultaneously achieve accurate segmentation and alignment under misaligned supervision, thereby relaxing the conventional reliance on pixel-perfect ground truth. Extensive experiments on real and synthetic datasets across multiple cities demonstrate the superiority of AnS in both segmentation accuracy and alignment fidelity.
This work addresses catastrophic forgetting in real-time semantic segmentation models during incremental learning when new classes are introduced. To tackle this challenge, the authors propose PILOT, a lightweight continual learning framework tailored for PIDNet. PILOT freezes the original network parameters and introduces a parallel derivative branch (D-branch) to capture high-frequency boundary information of new classes, enabling knowledge increment without access to historical data. By leveraging a boundary-guided mechanism and a data-agnostic learning strategy, PILOT effectively mitigates forgetting and substantially reduces training overhead while incurring negligible inference latency. Experimental results demonstrate that PILOT achieves accurate segmentation on newly added classes while preserving high mIoU for original classes, outperforming current state-of-the-art continual learning approaches.