Are In-Context Images Worth 10 Dimensions?

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how Large Vision-Language Models (LVLMs) construct linear representations of contextual images. By analyzing the principles of linear self-attention projections through both gradient descent analysis and empirical validation, this work elucidates the dimensionality reduction mechanisms operating in early network layers. The primary contribution is the first characterization of a novel mechanism by which LVLMs exploit visual modality properties to achieve dimensional reduction of contextual images. Specifically, the findings reveal that early layers perform a principal component analysis-like process, compressing images into linearly separable representations while discovering shared discriminative geometric structures. This mechanism effectively facilitates downstream classification tasks, offering new insights into the representational dynamics of vision-language architectures.
📝 Abstract
There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit. For a few-shot classification task, the induction circuit leverages linear representations of each labeled example in-context in order to classify an unlabeled query. However, few works focus on how those linear representations are built in the first place. Leveraging the expressivity of the vision modality compared to text, we uncover a Shared Discriminative Geometry (SDG) inside Large Vision Language Models (LVLMs). It is a low-dimensional space, shared across all image classification tasks, in which in-context images are compressed into linearly separable representations later used to perform classification. We observe that this is the result of the model performing a dimensionality reduction of vision representations in early layers. In order to explain this phenomenon: (1) We show analytically that linear self-attention can perform a dimensionality reduction by projecting in-context data onto its principal components, with each layer implementing one gradient descent step toward this objective. (2) We provide evidence that trained LVLMs reduce the dimensionality of vision representations in early layers via a similar mechanism.
Problem

Research questions and friction points this paper is trying to address.

In-Context Learning
Large Vision Language Models
Linear Representations
Dimensionality Reduction
Shared Discriminative Geometry
Innovation

Methods, ideas, or system contributions that make the work stand out.

In-Context Learning
Large Vision Language Models
Shared Discriminative Geometry
Dimensionality Reduction
Linear Self-Attention
🔎 Similar Papers
No similar papers found.
A
Adhemar de Senneville
Université Paris-Saclay, CNRS, ENS Paris-Saclay, Centre Borelli, Paris, France
Xavier Bou
Xavier Bou
Centre Borelli, ENS Paris-Saclay
Computer Vision
J
Jérémy Anger
Université Paris-Saclay, CNRS, ENS Paris-Saclay, Centre Borelli, Paris, France
R
Rafael Grompone
Université Paris-Saclay, CNRS, ENS Paris-Saclay, Centre Borelli, Paris, France
Gabriele Facciolo
Gabriele Facciolo
Professor of Mathematics, Centre Borelli, ENS Paris-Saclay
Image ProcessingComputer VisionRemote Sensing