Geometric Similarity in VLM Low-Level Vision Representations

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited understanding of the geometric organization of low-level visual representations in vision-language models (VLMs), which constrains the development of unified image restoration models. To this end, we propose GeoSim, a framework that introduces a pioneering four-tier analytical system encompassing global, local, sparse, and topological perspectives. By integrating techniques such as representational similarity analysis, sparse feature decomposition, and multi-task conditional extraction, GeoSim systematically evaluates the representational properties of autoregressive (AR) and diffusion transformer (DiT) architectures across 24 low-level tasks. Our findings reveal the organizational principles and latent transferability of low-level visual representations, delineate consistency boundaries across tasks and models, and establish an interpretable analytical paradigm for diagnosing model deficiencies.
📝 Abstract
Vision-language models (VLMs) have emerged as powerful candidates for universal vision backbones, with representative architectures including autoregressive (AR) models and diffusion transformers (DiTs). Yet, adapting them efficiently for all-in-one low-level image restoration remains a challenge. Crucially, the field lacks an understanding of how VLMs organize hidden-layer representations and whether these structurally distinct paradigms share a common geometric organization for pixel-level perception. Such shared organization is a prerequisite for building highly transferable, unified restoration VLMs and adapters. In this paper, we systematically investigate representational similarity across 24 low-level tasks spanning 5 categories. We propose GeoSim, a unified four-level framework that analyzes task-conditioned representations from global similarity, local geometry, sparse feature decomposition, and topological verification perspectives. Our formulation applies to the analysis of hidden states in AR models and feature maps in DiTs across same- and cross-task/model settings. Our results reveal the organizing principles of low-level visual representations while exposing their limits in cross-task and cross-model agreement. Ultimately, GeoSim provides an interpretability lens for probing latent transferability in low-level vision and diagnosing model limitations in task- or model-specific scenarios.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Low-Level Vision
Representational Similarity
Image Restoration
Geometric Organization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Geometric Similarity
Low-Level Vision
Representation Analysis
GeoSim Framework
🔎 Similar Papers
2024-08-29arXiv.orgCitations: 7