structured visual representation learning

Designs, builds, and evaluates models and representation spaces that encode visual entities and their relationships, producing per-object (unary) embeddings and pairwise or higher-order relational embeddings; and analyzes methods to explicitly represent, reason about, and manipulate object–object relations within a structured visual representation.

structuredvisualrepresentationlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.61
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Interpreting the structure of multi-object representations in vision encoders

Jun 13, 2024
TK
Tarun Khajuria
🏛️ University of Tartu

This work investigates the representational structure of visual encoders in multi-object scenes, formalizing “structured representation” as the binding of object-specific information to discrete tokens with cross-object disentanglement. It introduces the first quantitative measures for two core properties: binding fidelity and isolation degree. Leveraging a novel object decoding task built on COCO, the study employs token-level attribution, inter-layer representation decomposition, and structuralness evaluation to systematically benchmark ViT, CLIP, and DINOv2. Results show that DINOv2 and deeper ViT layers exhibit superior structured representation capability; the [CLS] token exhibits significant task bias; and pretraining objectives critically influence object separation quality. Crucially, the proposed metrics strongly correlate with downstream multi-object task performance—enabling interpretable model selection and task-aware architecture adaptation. (149 words)

Evaluate structured object-binding properties across different encodersInterpret multi-object scene representations in vision encodersPropose measures to quantify structured representation properties

Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models

Mar 21, 2025
DB
Davide Berasi
🏛️ Fondazione Bruno Kessler | University of Trento | University of Twente

This work investigates whether human-like semantic compositionality—i.e., interpretable decomposition and recomposition of image representations along semantic parts—exists in the visual embedding space of vision-language models (VLMs). Conventional linear compositional analysis fails on visual embeddings due to high noise and sparsity in image data. To address this, we propose Geodesically Decomposable Embeddings (GDE), a geometry-aware framework that replaces linear assumptions with manifold geodesic structure to model nonlinear semantic composition. We provide the first systematic empirical validation that mainstream VLMs exhibit significant, interpretable compositionality in their visual embeddings. GDE outperforms linear baselines on compositional classification and surpasses specialized methods in group robustness. Moreover, it reveals that VLMs implicitly possess automatic compositional reasoning capabilities. Our findings establish a new paradigm for interpretable and structured visual understanding in VLMs.

Evaluates compositional classification and group robustness performanceInvestigates compositionality in visual embeddings of VLMsProposes GDE framework for geometry-aware compositional structures

Achieving compositional generalization in vision models under unseen compositional scenarios remains a fundamental challenge for foundation models. This work theoretically establishes, for the first time, that effective compositional generalization necessitates representations that are disentangled, transferable, and stable—properties that require conceptual components in the embedding space to be linearly decomposable and mutually orthogonal. We further derive a geometric constraint linking the number of concepts to the embedding dimensionality. Guided by this theory, we conduct empirical analyses on prominent models including CLIP, SigLIP, and DINO, revealing that their representations consistently exhibit partial linear factorization and low-rank near-orthogonal structures. Crucially, the degree of such structure strongly correlates with compositional generalization performance, thereby validating the proposed theoretical framework.

compositional generalizationlinear representationsorthogonal representations

How Can Objects Help Video-Language Understanding?

Apr 10, 2025
ZT
Zitian Tang
🏛️ Brown University | Samsung Electronics

This study investigates the impact of explicitly incorporating object representations on video-language understanding in multimodal large language models (MLLMs), addressing the necessity and efficient implementation of object-centric modeling. Method: We propose a lightweight symbolic object representation scheme, seamlessly integrated into MLLM architectures via learnable adapters; we systematically compare symbolic versus distributed object representations and validate that explicit perception module fusion outperforms implicit inductive biases. Results: Experiments across five mainstream video question answering benchmarks demonstrate that our approach achieves state-of-the-art performance with higher data efficiency. It is the first work to empirically establish that symbolic object representations simultaneously preserve strong generalization capability and exhibit superior integration friendliness—offering a concise, effective, and scalable object modeling paradigm for visual grounding in MLLMs.

Explicit object-centric representation improves video question answeringHow objects enhance video-language understanding in MLLMsTrade-off between object representation expressiveness and integration difficulty

Exploring the Effectiveness of Object-Centric Representations in Visual Question Answering: Comparative Insights with Foundation Models

Jul 22, 2024
AM
Amir Mohammad Karimi Mamaghan
🏛️ KTH Royal Institute of Technology | University of Amsterdam | TU Munich | MPI for Intelligent Systems

This study investigates how object-centric (OC) representations enhance compositional generalization and structured reasoning in visual question answering (VQA), and analyzes their complementarity with large vision-language foundation models (e.g., ViT, CLIP). We introduce the first large-scale empirical framework, evaluating over 600 downstream VQA models across 15 upstream representation types—including OC models (Slot Attention, IODINE)—and incorporating multi-stage fine-tuning and prompting strategies. Our key contributions are: (1) the first empirical validation that OC representations substantially improve compositional generalization on both synthetic (CLEVR) and real-world (GQA) benchmarks; (2) a hybrid paradigm integrating OC representations with foundation models; and (3) experimental results demonstrating an average accuracy gain of 3.2% and a 21% improvement in robustness, revealing a synergistic division of labor—OC representations excel at structured, part-based reasoning, while foundation models support open-domain semantic understanding.

Compares OC models with foundation models for VQA tasks.Evaluates object-centric representations in Visual Question Answering.Identifies optimal strategies combining OC and foundation models.

Latest Papers

What's happening recently
View more

This work addresses the limitations of existing representation alignment methods, which predominantly rely on geometric properties and struggle to capture the global structural organization of model representations. To overcome this, the study introduces topological data analysis into the field for the first time, proposing a Mapper-based visual analytics framework. By integrating force-directed layout, Bubble Sets, motif querying, and membrane-inspired heuristics, the framework enables a unified analytical pipeline spanning global structure alignment, local region matching, and fine-grained pattern exploration. Case studies on language and multimodal models, complemented by expert evaluations, demonstrate that the approach effectively reveals and compares the topological organization of representations across different models or layers, offering deep structural insights.

global structuremodel comparisonneural representations

Existing layout-to-image generation methods suffer from fragmented representations under few-shot, atypical scenarios, leading to image distortion and loss of detail. This work proposes a representation-driven framework that explicitly decouples semantic identity from visual primitives for the first time. Semantic anchoring aggregates category-level semantics to stabilize object identity, while primitive injection models recombinable local primitives to enhance fine-grained details. Furthermore, a concept-guided mechanism incorporating saliency-aware optimization is introduced to improve foreground semantic consistency. Evaluated under a strict 5-shot setting, the proposed method consistently outperforms state-of-the-art approaches across multiple atypical domains, achieving notable improvements in both visual fidelity and layout alignment.

atypicalfew-shotlayout-to-image

The field of 3D vision suffers from fragmented data representations, learning paradigms, and benchmarking protocols, leading to a lack of unified understanding regarding efficiency, fidelity, and scalability. This work proposes the first cohesive conceptual framework that integrates geometric representations—such as point clouds, meshes, voxels, and 3D Gaussians—with diverse learning paradigms—including 2D-supervised learning, implicit neural representations, and 4D modeling—and connects them to real-world application scenarios. By constructing a structured knowledge graph of 3D vision, the study systematically relates dataset design, supervision mechanisms, and task requirements, clarifying the trade-offs between efficiency and fidelity and charting pathways for multimodal geometric grounding. This framework offers systematic guidance for reconstruction, generation, and dynamic scene modeling, advancing the field toward a unified and efficient paradigm.

3D visionbenchmark fragmentationdata representation

This study addresses the entanglement of objects and attributes in unsupervised image representation learning and the suboptimality of existing uniform decomposition strategies. To this end, it proposes a probabilistic generative model grounded in the Linear Representation Hypothesis (LRH). Methodologically, this work pioneers the introduction of LRH into slot space and designs a block attention architecture. Through variational inference, it jointly optimizes the evidence lower bound (ELBO) across object and attribute spaces, thereby achieving their decoupled learning. Experimental results demonstrate that the proposed method surpasses state-of-the-art approaches in terms of the Disentanglement, Completeness, and Informativeness (DCI) metric across multiple benchmark datasets. These findings effectively validate the plausibility of applying LRH within slot space and further support semantically interpretable image editing.

disentangled representationLinear Representation Hypothesisslot-based representation

研究使用Qwen3-VL-4B模型和一个合成数据集来探究视觉-语言模型是否真正理解视觉关系,发现这些模型结合了真正的视觉推理和基于语言线索的捷径策略。

Language CuesRelational ReasoningVision-Language Models

Hot Scholars

RM

Rui Mao

Nanyang Technological University
Computational LinguisticsCognitive ComputingMetaphorQuantitative Finance
QL

Quanyu Long

Nanyang Technological University
TLNLP
WW

Wenya Wang

Nanyang Technological University
Deep LearningKnowledge ReasoningNatural Language ProcessingSentiment Analysis
FW

Fang Wang

Postdoc, Stanford University
Reading acquisitiondyslexiacross-linguistic researchbilingualism