semantic-guided block matching

Designs and implements algorithms and modules that identify correspondences between image patches or blocks using semantic cues, including encoders that produce semantic tokens and matchers that find best-match tokens across images or views. Builds decoders or reconstruction components that use those matched tokens to complete or refine regions and analyzes methods for maintaining appearance and semantic consistency across matched views.

semantic-guidedblockmatching

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

SemanticStitch: Enhancing Image Coherence through Foreground-Aware Seam Carving

Nov 15, 2025
JJ
Ji-Ping Jin
🏛️ ShanghaiTech University | University of Macau

Image stitching often suffers from misalignment and visual discontinuities due to viewpoint differences and dynamic foreground objects; conventional seam-cutting methods ignore semantic structure, frequently severing salient foreground objects. To address this, we propose SemanticStitch—a novel framework that, for the first time, integrates foreground-aware semantic priors into seam optimization. It jointly leverages a semantic segmentation network and an enhanced seam-cutting algorithm, complemented by a real-scenario-oriented evaluation dataset. Crucially, we design a semantic consistency loss that preserves foreground object integrity while improving structural and appearance coherence across stitched regions. Extensive experiments demonstrate that SemanticStitch significantly outperforms state-of-the-art methods in visual quality, structural similarity (SSIM), and user-perceived quality. Moreover, it exhibits strong robustness to scene dynamics and viewpoint variation, confirming its practical viability for real-world deployment.

Addresses misalignments in image stitching from varying angles and movementsImproves visual coherence through deep learning with semantic priorsPreserves foreground object integrity using semantic-aware seam carving

Processing and acquisition traces in visual encoders: What does CLIP know about your camera?

Aug 14, 2025
RR
Ryan Ramos
🏛️ The University of Osaka | VRG | FEE | Czech Technical University in Prague

This work identifies that vision encoders (e.g., CLIP) implicitly encode imperceptible device- and algorithm-specific artifacts—such as camera model, compression parameters, and post-processing pipelines—introduced during image acquisition and processing. It systematically evaluates how these non-semantic signals interfere with or enhance downstream semantic predictions. Using feature interpretability analysis, linear probing, controlled ablation experiments, and distributional correlation modeling, the study quantitatively demonstrates: (1) acquisition and processing parameters are highly recoverable from CLIP visual representations (mean accuracy >92%); (2) such artifacts significantly degrade classification robustness under distribution shifts, inducing up to 15.3% error fluctuation; and (3) their statistical correlation with semantic labels modulates prediction confidence—either positively or negatively. This is the first systematic investigation revealing “implicit metadata contamination” in vision representations—a previously overlooked source of spurious correlation. To foster reproducibility, the authors release all code and benchmark datasets.

Analyzing subtle image acquisition parameters in CLIPExploring correlation between acquisition traces and model performanceInvestigating impact of imperceptible traces on semantic predictions

Semantic Correspondence: Unified Benchmarking and a Strong Baseline

May 23, 2025
KZ
Kaiyan Zhang
🏛️ The University of Hong Kong | University of Oxford

Semantic correspondence aims to match semantically consistent keypoints across images, yet has long suffered from the absence of systematic surveys and unified evaluation protocols. This paper introduces the first comprehensive, reproducible semantic correspondence benchmark. Our contributions are threefold: (1) We propose the first hierarchical taxonomy of methodologies, clarifying the technical evolution and conceptual relationships among existing approaches; (2) We design a lightweight, efficient baseline model that achieves state-of-the-art performance on major benchmarks including PF-PASCAL and SPair-71k; (3) We establish cross-dataset standardized evaluation protocols and perform component-wise ablation studies, substantially improving experimental comparability and reproducibility. All code, configuration files, and evaluation tools are publicly released under an open-source license, enabling plug-and-play reproduction.

Establish semantic correspondence across different imagesPropose a strong baseline for future researchReview and analyze existing semantic correspondence methods

Semantic Mosaicing of Histo-Pathology Image Fragments using Visual Foundation Models

Aug 05, 2025
SB
Stefan Brandstätter
🏛️ Medical University Vienna

In computational pathology, large tissue sections are segmented into multiple fragments and stitched into whole-mount slide (WMS) images; however, existing boundary-based stitching methods suffer from poor robustness due to tissue defects, deformations, staining heterogeneity, and edge abrasion. To address this, we propose the first semantic-driven stitching framework leveraging Vision Foundation Models (VFMs): it extracts high-dimensional latent semantic features from pretrained pathological VFMs, constructs cross-fragment semantic correspondence candidate sets, and enables robust pose estimation and precise spatial registration. By decoupling stitching from explicit geometric boundary constraints, our method significantly improves tolerance to nonrigid deformations and staining variations. Evaluated on three public histopathological datasets, it consistently outperforms state-of-the-art methods in boundary matching accuracy, yielding more stable and higher-fidelity WMS reconstructions.

Automated stitching of large histopathology tissue fragmentsImproving accuracy in whole mount slide reconstructionOvercoming challenges like tissue loss and distortion

SAM-CP: Marrying SAM with Composable Prompts for Versatile Segmentation

Jul 23, 2024
PC
Pengfei Chen
🏛️ University of Chinese Academy of Sciences | Huawei Inc. | University of Science and Technology of China

To address SAM’s lack of semantic awareness and its inability to support open-vocabulary and multi-granularity semantic segmentation, this paper proposes a dual-type composable prompting framework: Type-I prompts align textual class labels with SAM’s base segmentation tokens semantically; Type-II prompts model instance consistency by unifying affinity modeling between semantic/instance queries and SAM tokens. The method requires no fine-tuning, integrating zero-shot SAM segmentation, CLIP-based text–image matching, affinity graph construction, and hierarchical token merging. It supports semantic, instance, and panoptic segmentation in both open- and closed-vocabulary settings. On open-vocabulary segmentation benchmarks, it achieves state-of-the-art performance, significantly outperforming existing adaptation methods across multiple datasets. Notably, it is the first framework to enable single-model, zero-shot, multi-granularity, open-vocabulary, semantic-aware segmentation.

Achieving versatile segmentation in open and closed domainsEnhancing SAM for semantic-aware segmentation with composable promptsReducing complexity in handling multiple semantic classes and patches

Latest Papers

What's happening recently
View more

This study addresses the performance stagnation of existing semantic correspondence methods under fine-grained thresholds, which stems from quantization errors induced by the discrete grids of Vision Transformers (ViTs). To overcome this limitation, this work proposes the concept of continuous feature fields, constructing a continuous feature space via implicit neural representations to support queries at arbitrary coordinates. Specifically, a FiLM-conditioned decoder is employed to embed sub-pixel positional information into ViT features, thereby transcending discrete grid constraints and enabling continuous coordinate sampling. The proposed method achieves a 6.2 percentage point improvement in PCK@0.01 on the SPair-71k benchmark, substantially enhancing fine-grained semantic matching accuracy.

Fine-grained MatchingQuantization ErrorSemantic Correspondence

Current text-to-image generation models struggle to achieve smooth transitions between semantically similar prompts due to substantial differences in token sequences—particularly in wording, ordering, and conceptual positioning—which hinders effective image blending and continuous editing. This work proposes a Token-to-Token Alignment framework that, without modifying the underlying model, employs a two-stage strategy: first aligning the semantic structures of prompts and then aligning their token embedding representations. By reconstructing diverse prompts into a shared structural form, the method reveals that the latent continuous semantic structure within the text embedding space can be effectively leveraged through representation alignment. Consequently, linear interpolation in this aligned space yields coherent semantic transitions, significantly enhancing the quality of image semantic mixing and continuous editing.

embedding spacesemantic blendingsemantic structure

This work proposes SemTok, a semantic-driven one-dimensional tokenizer that addresses the limitations of existing vision tokenizers, which often rely on fixed 2D grids and prioritize pixel-level reconstruction at the expense of compact global semantics. SemTok compresses images into high-level discrete semantic tokens and introduces a masked autoregressive generative framework. Its key innovations include a 2D-to-1D semantic tokenization strategy, a semantic alignment constraint mechanism, and a two-stage generative training paradigm. Experimental results demonstrate that SemTok achieves state-of-the-art performance in image reconstruction, delivering higher fidelity under extremely compact token representations and significantly enhancing downstream generative capabilities.

2D-to-1D mappingimage reconstructionlatent space

This work addresses the lack of explicit modeling of co-visible regions in image correspondence estimation under large viewpoint and scale variations by proposing a structured feature matching method grounded in co-visibility modeling. It extends the Segment Anything Model (SAM) to multi-view correspondence inference for the first time, leveraging predicted cross-view co-visible masks and bounding boxes as structured priors. A symmetric cross-view interaction mechanism is introduced to enable bidirectional feature exchange and semantic alignment. By integrating mask–box consistency constraints with a unified supervision strategy, the approach shifts the matching paradigm from pixel-level to region-level. The method achieves significant performance gains over existing techniques across multiple challenging benchmarks, demonstrating notably enhanced robustness under extreme viewpoint and scale changes.

co-visibility modelingcorrespondence estimationfeature matching

This study addresses the challenges of semantic-geometric inconsistency and the absence of global guidance in cross-modal matching by proposing the CDPM framework. This work is the first to introduce geometric consistency constraints into DINOv3 feature adaptation, constructing a DINO-centric multi-scale feature pyramid through geometry-aware patch pair mining. By integrating a lightweight CNN for local structure refinement, the method establishes a semantics-led, detail-assisted matching architecture that overcomes the conventional stability-accuracy trade-off. Evaluated on VIS-IR datasets, CDPM achieves an AUC improvement exceeding 10 percentage points and reduces the mean angular error to 2.78 pixels. Furthermore, it outperforms RoMa v2 while reducing computational cost by 45.6%.

Cross-modal image registrationFeature alignmentGeometric correspondence

Hot Scholars

SM

Siwei Ma

Peking University
Video Coding and Processing
CJ

Chuanmin Jia

Peking University
Video CodingMultimediaData Compression
YH

Yuxing Han

Tsinghua University
Smart AgricultureArtificial IntelligenceVideoCommunication
LP

Ling Pei

Shanghai Jiao Tong University
NavigationPositioningSLAMSensor Fusion
XZ

Xinfeng Zhang

Fuxi AI Lab, NetEase Inc.
Vision-Language ModelsMultimodal