Score
Designs, implements, and evaluates cascaded coarse‑to‑fine pipelines that first detect candidate foreground/object regions and then classify or segment those regions, including architectures that crop or mask regions to pass to a second-stage segmenter or classifier. Builds and analyzes background modeling, background removal/suppression, and foreground‑background separation components to reduce sensitivity to background noise, handle out‑of‑distribution backgrounds, and speed processing of large datasets by limiting computation to detected regions.
To address the inefficiency, time consumption, and high resource demands of manual mask generation in visual effects (VFX) production, this paper proposes a text-prompt-driven automated video segmentation pipeline. The method integrates text-guided instance segmentation for flexible object localization, fine-grained frame-wise semantic segmentation, and temporally consistent video object tracking, deployed via lightweight containerization to ensure seamless integration with artist workflows. Key contributions include: (i) the first end-to-end incorporation of text prompting into VFX-grade video segmentation; (ii) a multi-stage collaborative architecture ensuring inter-frame segmentation stability; and (iii) structured, editable outputs—including mask sequences, confidence maps, and trajectory metadata. Experiments demonstrate that the system reduces pre-compositing preparation time by 72% on average and decreases manual intervention by 89%, significantly enhancing automation and productivity across VFX pipelines.
Small-object detection performance is hindered by fragmented optimization across stages in conventional pipeline-based detectors. To address this, we propose PLUSNet, an end-to-end co-optimization framework introducing the novel “Purify–Label–Utilize” paradigm. Specifically, we design a hierarchical feature purifier to suppress noise; develop a multi-criterion dynamic label assignment mechanism to improve positive/negative sample quality; and introduce a frequency-domain decoupled detection head for fine-grained feature modeling. All modules are lightweight, modular, and seamlessly integrate with mainstream detectors. Extensive experiments on MS COCO, VisDrone, and other benchmarks demonstrate consistent and significant gains in small-object AP (+3.2–5.8 points), validating the effectiveness of joint upstream-downstream optimization and strong generalizability across diverse scenarios and architectures.
This work addresses the challenges of high early-stage uncertainty in manufacturing monitoring system development—leading to redundant modeling and substantial training costs—and the limited transferability of filtering pipelines in cross-domain image segmentation tasks. To tackle these issues, the authors propose a problem-centric design paradigm that constructs an abstract system model to continuously accumulate and retrieve historical segmentation tasks along with their associated filtering pipelines, enabling solution reuse and incremental optimization. The approach integrates similarity-based problem retrieval, abstract modeling, pipeline reuse, and a retrieval-augmented evolutionary learning mechanism. Experimental results demonstrate that the method significantly reduces training costs and late-stage revision risks, provides the first systematic validation of filtering pipeline transferability across similar segmentation tasks, and achieves a favorable balance among complexity, technical requirements, and reliability under lightweight model constraints.
To address performance limitations in object detection and semantic segmentation under complex scenarios—including occlusion, small objects, and cross-domain generalization—this paper proposes a novel multimodal detection paradigm synergizing large language models (LLMs). Methodologically, it systematically integrates CNNs, YOLOv5/v8, and DETR architectures into an LLM-augmented inference framework, augmented by scalable data pipelines, model pruning, and quantization, and evaluated via a multi-dimensional metric system based on mAP and mIoU. Key contributions include: (1) bridging the gap between traditional feature engineering and end-to-end deep learning; (2) introducing a dynamic context enhancement mechanism tailored for challenging environments; and (3) achieving state-of-the-art accuracy-efficiency trade-offs on COCO and ADE20K. The fully open-sourced, reproducible framework significantly improves model generalizability and robustness across diverse real-world conditions.
While the Segment Anything Model (SAM) exhibits strong generalization capability, its substantial computational overhead hinders deployment on resource-constrained edge devices. This work presents a systematic survey of efficient SAM variants tailored for edge deployment. We introduce the first unified evaluation framework spanning diverse hardware platforms—including CPU, GPU, and Edge TPU—and conduct joint accuracy–latency–memory benchmarking on COCO and SA-1B. Our analysis categorizes acceleration techniques along six technical axes: model pruning, knowledge distillation, lightweight attention mechanisms, quantization, module substitution, and hardware-aware compilation—characterizing their Pareto-optimal trade-offs. The core contributions are: (1) an open-source, fully reproducible edge-SAM benchmark; and (2) empirical insights into the applicability domains and fundamental accuracy-efficiency trade-offs of each acceleration strategy—providing both theoretical foundations and practical guidelines for designing lightweight vision foundation models.
Existing CLIP-based foreground-background decomposition methods for few-shot out-of-distribution (OOD) detection employ a uniform background suppression strategy and overlook semantically ambiguous local regions within the foreground, thereby limiting performance. To address these limitations, this work proposes a plug-and-play framework that refines semantic control over image regions through three key components: foreground-background decomposition, entropy-weighted adaptive background suppression, and confusion-aware foreground correction based on local semantic similarity analysis. By moving beyond the conventional practice of treating all background regions equally and ignoring intra-foreground semantic ambiguities, the proposed method achieves significant performance gains over existing foreground-background approaches across multiple benchmarks, demonstrating its effectiveness and generalizability.
Existing segmentation evaluation metrics often lack transparency and modularity, making them ill-suited for diverse tasks such as transparent object, specular surface, or lesion segmentation. This work proposes a unified evaluation framework that decomposes metrics into five modular components: prediction representation, target extraction, target matching, score computation, and metric reporting. For the first time, it systematically analyzes the implicit assumptions and design limitations of mainstream binary segmentation metrics through this modular lens. The framework enables task-aware customization of evaluation protocols, reveals evolutionary trajectories among existing metrics, and is accompanied by an open-source toolkit. By offering a principled and interpretable foundation, this approach paves the way for developing more rational and adaptable segmentation evaluation methodologies.
This study addresses the challenges of detail loss and dataset discrepancies in satellite image time series crop segmentation by proposing PAtteRNS, a hybrid Transformer-convolutional model. Its core innovation lies in the first parallel dimension self-attention architecture tailored for Sentinel-2 multispectral data, which independently decouples attention computation across temporal, spectral, and spatial dimensions, achieving fully factorized attention while significantly reducing computational complexity. Experimental results demonstrate that PAtteRNS surpasses existing state-of-the-art methods on the PASTIS and MTLCC datasets, exhibiting particularly notable advantages in parcel boundary delineation quality. Furthermore, this work reveals the potential impact of dataset grouping deficiencies on model performance.
为解决半导体制造中缺陷分析自动化不足的问题,提出了一种基于弱监督的晶圆缺陷分割方法SePArate,通过三阶段训练实现精确分割。
Existing vision Mixture-of-Experts (MoE) approaches perform routing at the image or patch level, which struggles to align with the instance-centric nature of object detection. This work proposes a hierarchical instance-conditioned MoE architecture that introduces a two-stage routing mechanism—operating at both scene and instance levels—within a DETR-style detector, achieving fine-grained expert assignment aligned with instance queries for the first time. The method employs a lightweight scene router and an instance router that jointly balance sparse computation with the heterogeneity of individual instances. Experiments demonstrate that the model outperforms dense DINO baselines and simplified routing variants on COCO, significantly enhancing small object detection performance and offering preliminary evidence of functional specialization among experts.