Score
Designs and implements training objectives, auxiliary losses, and pretraining tasks that explicitly target region boundaries and edges in dense-prediction models, including masked-boundary prediction, boundary-targeted masking, and edge/auxiliary boundary supervision. Builds components and representations (e.g., dynamic identification of boundary-bearing tokens and sub-pixel boundary encodings) and evaluation practices to sharpen region-aligned segmentation and to encourage learning of dense visual tokens that capture precise boundary geometry.
Current vision foundation models emphasize semantic invariance, which falls short of meeting the metric-level dense spatial understanding required by embodied intelligence. This work proposes Mask Boundary Modeling—a self-supervised approach that leverages a boundary-centric perspective to dynamically learn sub-pixel accurate boundary representations and guide the learning of dense visual tokens. By elevating boundary modeling from simple line segments to a scalable pretraining paradigm, the method effectively drives dense spatial representation learning. Built upon this framework, LingBot-Vision significantly outperforms the DINOv3 baseline on downstream tasks such as depth estimation, enabling the upgrade of LingBot-Depth from version 1.0 to 2.0 and substantially improving performance in depth completion and estimation.
This study addresses the issues of error fixation caused by hard labels and insufficient boundary precision in weakly supervised semantic segmentation by proposing the DS-CRF framework. Inspired by Conditional Random Fields, this method decouples unary and pairwise potential supervision signals. Specifically, it leverages DINO text CAMs to provide soft classification supervision and incorporates SAM to extract boundary information, optimizing training via a custom CRF loss function. This design effectively prevents the error amplification typically induced by conventional pseudo-label fusion while preserving prediction uncertainty to enhance boundary learning. Experimental results demonstrate that the proposed framework achieves 56.5% mIoU on the MS COCO dataset, establishing a new state-of-the-art performance.
Existing open-vocabulary segmentation models rely on predefined category prompts and produce only sparse predictions, limiting free-text-driven pixel-level mask generation and automatic discovery of unseen categories. To address this, we propose the first patch-wise perception paradigm for open-set dense segmentation, enabling joint dense and sparse mask prediction. We introduce an instruction-response dialogue fine-tuning mechanism to transcend closed-set category constraints. Our approach integrates multimodal large model (LMM) vision-language alignment, patch-wise attention modeling, instruction tuning, and iterative text-guided mask refinement. Evaluated on comprehensive multi-task open-set segmentation benchmarks, our method achieves state-of-the-art performance. Notably, it is the first framework to unify zero-shot category generation, high-precision dense segmentation, and fine-grained semantic correction within a single architecture—enabling both open-vocabulary expressivity and dense spatial reasoning.
In masked image modeling (MIM) pretraining of Vision Transformers (ViTs), tokenization and local masking induce spatially inconsistent reconstruction supervision, degrading representation discriminability. This work is the first to systematically identify and address this spatial inconsistency issue, proposing Dynamic Token Morphing (DTM): a context-aware, dynamic token aggregation mechanism that generates spatially coherent reconstruction targets. DTM introduces no additional parameters or computational overhead and is plug-and-play across diverse MIM frameworks. On ImageNet-1K and ADE20K, DTM achieves significant gains over state-of-the-art MIM methods—yielding lower training loss and more stable convergence. When transferred to downstream tasks such as iNaturalist, it delivers consistent performance improvements. The core contribution is the first lightweight, parameter-free, and framework-agnostic solution specifically designed to resolve the spatial inconsistency problem in MIM.
In knowledge distillation (KD) for dense image prediction tasks, student models often suffer from fragmented object boundaries and poor region connectivity. To address this, we propose Boundary-and-Context Distillation (BCD), a complementary distillation framework that jointly integrates explicit semantic boundary extraction and implicit pixel-level context transfer into KD. BCD explicitly distills boundary-aware representations by leveraging multi-level backbone features, while implicitly transferring contextual information via self-relational modeling—enabling unsupervised, end-to-end differentiable context alignment without auxiliary annotations or inference overhead. Evaluated on semantic segmentation, instance segmentation, and object detection benchmarks, BCD consistently outperforms state-of-the-art KD methods, yielding predictions with sharper object boundaries and significantly improved intra-object region continuity.
This work addresses the challenge of instance segmentation in low-texture, uniformly colored industrial scenes, where existing methods struggle due to insufficient texture and contextual cues, and where instance definitions often vary by task. The authors propose a boundary-aware few-shot instance segmentation framework that leverages a pretrained foundation model to extract visual features and employs a lightweight signed distance function (SDF) head to predict boundary-sensitive distance maps. Instance masks are then reconstructed from these SDF predictions. By shifting supervision from internal appearance to explicit boundary modeling, the approach enables flexible instance definition—such as whole objects or subparts—via mask conditioning. A pixel-wise shallow MLP facilitates rapid training and fine-grained control. Experiments demonstrate superior few-shot generalization, robustness, and precise control over segmentation granularity on low-texture industrial components and food imagery.
This study addresses the problem of recovering unknown object boundaries from noisy, unlabeled images in an unsupervised and nonparametric setting. To this end, the authors propose a boundary detection method that integrates a continuous hinge-type surrogate loss with deep neural networks, embedded within a robust Gibbs posterior framework based on a thresholded misclassification loss. Theoretical analysis demonstrates that the resulting estimator achieves minimax optimal convergence rates—up to logarithmic factors—for piecewise smooth boundaries that may include corners and kinks, while also enjoying Fisher consistency and calibration properties. Extensive experiments confirm the method’s stability and superior performance across varying noise levels and complex boundary shapes, significantly outperforming existing unsupervised boundary detection approaches.
Existing object detectors often learn task-driven features that rely on shortcut correlations, failing to adequately capture the underlying annotation structure, which limits their generalization, interpretability, and robustness under task shifts or sparse supervision. To address this, this work proposes an annotation-guided feature enhancement framework that explicitly integrates geometric annotation priors into feature learning for the first time. By constructing a dense spatial feature grid and injecting it into the backbone network—where it fuses with the feature pyramid—the method steers region proposal and detection heads toward representations better aligned with annotation structure. Evaluated on wildlife and remote sensing datasets, the approach significantly improves object focus, reduces background sensitivity, and demonstrates superior generalization and data efficiency in weakly supervised and unseen-task settings.
This work addresses the challenge of insufficient accuracy in segmenting fine structures and object boundaries—particularly in resource-constrained settings—by proposing FoR-Net, a lightweight architecture for semantic segmentation. FoR-Net incorporates a region-focused inductive bias that selectively enhances information-rich and challenging regions through importance map prediction and a Top-K activation mechanism. It further aggregates multi-scale spatial context via parallel convolutional branches. Evaluated on the Cityscapes benchmark under standard training protocols and limited computational resources, FoR-Net achieves competitive overall performance while significantly improving segmentation consistency in difficult regions.
Existing temporal action segmentation methods face deployment challenges due to architectural complexity, imprecise boundary localization, and weak intra-segment consistency. This work proposes a lightweight dual-loss training framework that enhances performance without modifying the backbone architecture—requiring only an additional output channel and two auxiliary loss terms. The first is a single-channel boundary regression loss designed to improve temporal boundary accuracy, and the second is a segment-level regularization based on the cumulative distribution function (CDF) to strengthen intra-segment consistency. The approach is architecture-agnostic and seamlessly integrates with mainstream models such as MS-TCN, C2F-TCN, and FACT. Evaluated on three benchmark datasets, it consistently achieves significant gains in F1 and Edit scores while maintaining stable frame-wise accuracy, demonstrating the efficacy of a minimalist yet well-designed loss formulation.