Score
Designs and implements algorithms and modules that identify correspondences between image patches or blocks using semantic cues, including encoders that produce semantic tokens and matchers that find best-match tokens across images or views. Builds decoders or reconstruction components that use those matched tokens to complete or refine regions and analyzes methods for maintaining appearance and semantic consistency across matched views.
Image stitching often suffers from misalignment and visual discontinuities due to viewpoint differences and dynamic foreground objects; conventional seam-cutting methods ignore semantic structure, frequently severing salient foreground objects. To address this, we propose SemanticStitch—a novel framework that, for the first time, integrates foreground-aware semantic priors into seam optimization. It jointly leverages a semantic segmentation network and an enhanced seam-cutting algorithm, complemented by a real-scenario-oriented evaluation dataset. Crucially, we design a semantic consistency loss that preserves foreground object integrity while improving structural and appearance coherence across stitched regions. Extensive experiments demonstrate that SemanticStitch significantly outperforms state-of-the-art methods in visual quality, structural similarity (SSIM), and user-perceived quality. Moreover, it exhibits strong robustness to scene dynamics and viewpoint variation, confirming its practical viability for real-world deployment.
This work identifies that vision encoders (e.g., CLIP) implicitly encode imperceptible device- and algorithm-specific artifacts—such as camera model, compression parameters, and post-processing pipelines—introduced during image acquisition and processing. It systematically evaluates how these non-semantic signals interfere with or enhance downstream semantic predictions. Using feature interpretability analysis, linear probing, controlled ablation experiments, and distributional correlation modeling, the study quantitatively demonstrates: (1) acquisition and processing parameters are highly recoverable from CLIP visual representations (mean accuracy >92%); (2) such artifacts significantly degrade classification robustness under distribution shifts, inducing up to 15.3% error fluctuation; and (3) their statistical correlation with semantic labels modulates prediction confidence—either positively or negatively. This is the first systematic investigation revealing “implicit metadata contamination” in vision representations—a previously overlooked source of spurious correlation. To foster reproducibility, the authors release all code and benchmark datasets.
Semantic correspondence aims to match semantically consistent keypoints across images, yet has long suffered from the absence of systematic surveys and unified evaluation protocols. This paper introduces the first comprehensive, reproducible semantic correspondence benchmark. Our contributions are threefold: (1) We propose the first hierarchical taxonomy of methodologies, clarifying the technical evolution and conceptual relationships among existing approaches; (2) We design a lightweight, efficient baseline model that achieves state-of-the-art performance on major benchmarks including PF-PASCAL and SPair-71k; (3) We establish cross-dataset standardized evaluation protocols and perform component-wise ablation studies, substantially improving experimental comparability and reproducibility. All code, configuration files, and evaluation tools are publicly released under an open-source license, enabling plug-and-play reproduction.
In computational pathology, large tissue sections are segmented into multiple fragments and stitched into whole-mount slide (WMS) images; however, existing boundary-based stitching methods suffer from poor robustness due to tissue defects, deformations, staining heterogeneity, and edge abrasion. To address this, we propose the first semantic-driven stitching framework leveraging Vision Foundation Models (VFMs): it extracts high-dimensional latent semantic features from pretrained pathological VFMs, constructs cross-fragment semantic correspondence candidate sets, and enables robust pose estimation and precise spatial registration. By decoupling stitching from explicit geometric boundary constraints, our method significantly improves tolerance to nonrigid deformations and staining variations. Evaluated on three public histopathological datasets, it consistently outperforms state-of-the-art methods in boundary matching accuracy, yielding more stable and higher-fidelity WMS reconstructions.
To address SAM’s lack of semantic awareness and its inability to support open-vocabulary and multi-granularity semantic segmentation, this paper proposes a dual-type composable prompting framework: Type-I prompts align textual class labels with SAM’s base segmentation tokens semantically; Type-II prompts model instance consistency by unifying affinity modeling between semantic/instance queries and SAM tokens. The method requires no fine-tuning, integrating zero-shot SAM segmentation, CLIP-based text–image matching, affinity graph construction, and hierarchical token merging. It supports semantic, instance, and panoptic segmentation in both open- and closed-vocabulary settings. On open-vocabulary segmentation benchmarks, it achieves state-of-the-art performance, significantly outperforming existing adaptation methods across multiple datasets. Notably, it is the first framework to enable single-model, zero-shot, multi-granularity, open-vocabulary, semantic-aware segmentation.
This study addresses the performance stagnation of existing semantic correspondence methods under fine-grained thresholds, which stems from quantization errors induced by the discrete grids of Vision Transformers (ViTs). To overcome this limitation, this work proposes the concept of continuous feature fields, constructing a continuous feature space via implicit neural representations to support queries at arbitrary coordinates. Specifically, a FiLM-conditioned decoder is employed to embed sub-pixel positional information into ViT features, thereby transcending discrete grid constraints and enabling continuous coordinate sampling. The proposed method achieves a 6.2 percentage point improvement in PCK@0.01 on the SPair-71k benchmark, substantially enhancing fine-grained semantic matching accuracy.
Current text-to-image generation models struggle to achieve smooth transitions between semantically similar prompts due to substantial differences in token sequences—particularly in wording, ordering, and conceptual positioning—which hinders effective image blending and continuous editing. This work proposes a Token-to-Token Alignment framework that, without modifying the underlying model, employs a two-stage strategy: first aligning the semantic structures of prompts and then aligning their token embedding representations. By reconstructing diverse prompts into a shared structural form, the method reveals that the latent continuous semantic structure within the text embedding space can be effectively leveraged through representation alignment. Consequently, linear interpolation in this aligned space yields coherent semantic transitions, significantly enhancing the quality of image semantic mixing and continuous editing.
This work proposes SemTok, a semantic-driven one-dimensional tokenizer that addresses the limitations of existing vision tokenizers, which often rely on fixed 2D grids and prioritize pixel-level reconstruction at the expense of compact global semantics. SemTok compresses images into high-level discrete semantic tokens and introduces a masked autoregressive generative framework. Its key innovations include a 2D-to-1D semantic tokenization strategy, a semantic alignment constraint mechanism, and a two-stage generative training paradigm. Experimental results demonstrate that SemTok achieves state-of-the-art performance in image reconstruction, delivering higher fidelity under extremely compact token representations and significantly enhancing downstream generative capabilities.
This work addresses the lack of explicit modeling of co-visible regions in image correspondence estimation under large viewpoint and scale variations by proposing a structured feature matching method grounded in co-visibility modeling. It extends the Segment Anything Model (SAM) to multi-view correspondence inference for the first time, leveraging predicted cross-view co-visible masks and bounding boxes as structured priors. A symmetric cross-view interaction mechanism is introduced to enable bidirectional feature exchange and semantic alignment. By integrating mask–box consistency constraints with a unified supervision strategy, the approach shifts the matching paradigm from pixel-level to region-level. The method achieves significant performance gains over existing techniques across multiple challenging benchmarks, demonstrating notably enhanced robustness under extreme viewpoint and scale changes.
This study addresses the challenges of semantic-geometric inconsistency and the absence of global guidance in cross-modal matching by proposing the CDPM framework. This work is the first to introduce geometric consistency constraints into DINOv3 feature adaptation, constructing a DINO-centric multi-scale feature pyramid through geometry-aware patch pair mining. By integrating a lightweight CNN for local structure refinement, the method establishes a semantics-led, detail-assisted matching architecture that overcomes the conventional stability-accuracy trade-off. Evaluated on VIS-IR datasets, CDPM achieves an AUC improvement exceeding 10 percentage points and reduces the mean angular error to 2.78 pixels. Furthermore, it outperforms RoMa v2 while reducing computational cost by 45.6%.