Score
Designs and trains models and training pipelines that produce continuous, object-centered semantic embedding fields and the associated object-centric representations — i.e., learnable functions over an object’s spatial domain that map spatial queries or input samples to embeddings encoding part-level semantics. Work includes specifying architectures, input conditioning (e.g., on object samples), loss functions and supervision strategies to produce part-aware semantic embeddings and to evaluate or analyze the learned object-centric representations.
This study investigates how object-centric (OC) representations enhance compositional generalization and structured reasoning in visual question answering (VQA), and analyzes their complementarity with large vision-language foundation models (e.g., ViT, CLIP). We introduce the first large-scale empirical framework, evaluating over 600 downstream VQA models across 15 upstream representation types—including OC models (Slot Attention, IODINE)—and incorporating multi-stage fine-tuning and prompting strategies. Our key contributions are: (1) the first empirical validation that OC representations substantially improve compositional generalization on both synthetic (CLEVR) and real-world (GQA) benchmarks; (2) a hybrid paradigm integrating OC representations with foundation models; and (3) experimental results demonstrating an average accuracy gain of 3.2% and a 21% improvement in robustness, revealing a synergistic division of labor—OC representations excel at structured, part-based reasoning, while foundation models support open-domain semantic understanding.
Existing object-centric representation learning models can unsupervisedly discover scene objects but lack language controllability, hindering targeted extraction of specific object instances via natural language instructions. To address this, we propose the first language-driven object-centric representation framework. Our method introduces slot-language cross-modulation and CLIP feature-guided differentiable attention routing to achieve end-to-end self-supervised semantic binding and cross-modal alignment. Crucially, it operates without pixel-level mask supervision, enabling text-directed object localization and instance-level representation generation. Evaluated on real-world complex scenes, our approach significantly improves language-guided object extraction accuracy, enhances instance fidelity in text-to-image generation, and achieves state-of-the-art performance on visual question answering. The framework bridges the gap between object-centric learning and grounded language understanding, enabling precise, interpretable, and controllable scene decomposition through natural language.
Existing unsupervised 3D object discovery methods exhibit poor generalization in real-world scenes due to entangled modeling of intrinsic object properties (e.g., shape, appearance) and observer-dependent pose (i.e., 6DoF position and orientation). Method: We propose the first object-centric Neural Radiance Fields (NeRF) framework for single real-world images. It employs explicit camera parameter decoupling (intrinsic/extrinsic), a differentiable spatial segmentation module, and self-supervised consistency regularization to achieve unsupervised disentangled 3D reconstruction—separating object shape/appearance from 6DoF pose. Contribution/Results: Evaluated on a newly constructed real-world kitchen dataset, our method achieves high-fidelity object-level 3D segmentation and editing. It discovers and reconstructs multiple objects from a single input image and demonstrates strong zero-shot generalization to unseen objects, significantly outperforming baselines in reconstruction accuracy.
This work addresses unsupervised object-centric representation learning by explicitly disentangling shape and texture factors of objects in images, thereby enhancing model robustness to structural variations and cross-object generalization. Methodologically, it is the first to impose a predefined dimensional partitioning of the latent space within an object-centric framework, enforcing strict separation between shape and texture subspaces. Building upon Invariant Slot Attention, the approach introduces prior-driven structural constraints and a dual-branch feature disentanglement architecture, enabling controllable texture generation and cross-shape–texture transfer. Evaluated on multiple standard benchmarks, the method achieves significant improvements in disentanglement quality—measured via established metrics—and consistently outperforms existing baselines across diverse downstream tasks, including segmentation, reconstruction, and compositional generalization. These results empirically validate the effectiveness and practicality of explicit latent-space structural design for object-centric learning.
This work addresses unsupervised object decomposition to enhance object-centric scene understanding, particularly for set-wise attribute prediction. We propose a clustering slot module based on a learnable Gaussian Mixture Model (GMM), wherein each slot is modeled as a cluster centroid jointly embedded with distance-aware representations relative to other clusters—replacing conventional Slot Attention. To our knowledge, this is the first incorporation of GMM into slot-based learning, explicitly encoding inter-cluster structural relationships to improve representation discriminability and geometric consistency. Our method integrates soft assignment, end-to-end optimization, and slot-wise distance-aware embedding. Evaluated on CLEVR6 and Multi-dSprites, it significantly outperforms Slot Attention and other baselines: set-wise attribute prediction accuracy improves by 3.2–5.7 percentage points, achieving state-of-the-art performance.
This work addresses the challenge of extracting object-centric structured representations from complex real-world scenes—characterized by multiple objects and low contrast—under unsupervised conditions. We propose a novel approach based on Cycle-Consistent Generative Adversarial Networks (Cycle-Consistent GANs), which, for the first time, introduces cycle consistency into object-centric representation learning, thereby overcoming limitations of conventional autoencoder architectures. By jointly performing unsupervised object segmentation and modeling a low-dimensional latent space, our method decomposes input images into independent object representations and reconstructs them faithfully. Experiments demonstrate that the proposed method achieves state-of-the-art performance on synthetic data and is currently the only approach capable of effectively handling multi-object, low-contrast real-world images. The learned representations enable object-level manipulation and exhibit strong scalability with respect to both object count and image resolution.
This work addresses the lack of stable, viewpoint-invariant 3D semantic understanding of functional object parts—such as handles or openings—in existing robotic manipulation approaches. The authors propose an object-centric continuous semantic field that learns part-aware semantic embeddings from point cloud inputs, enabling queryable representations at arbitrary 3D locations and serving as conditional input for manipulation policies. This formulation overcomes the discreteness and observation dependence of conventional point-level semantics, yielding a continuous, object-level, and viewpoint-invariant semantic representation. Evaluated in both RoboTwin simulation and real-world bimanual robot tasks, the method significantly outperforms baselines using raw point clouds, 2D feature augmentation, or 3D point-level semantics in terms of policy performance and stability in identifying functional parts.
Existing evaluations of object-centric models are largely confined to object discovery and simple reasoning tasks, which inadequately assess their representational capacity under compositional generalization and out-of-distribution robustness. To address this limitation, this work proposes a novel evaluation framework that leverages instruction-tuned vision-language models (VLMs) to probe the support provided by object-centric representations for complex reasoning across diverse visual question answering (VQA) tasks. The framework introduces a unified task design that simultaneously evaluates localization accuracy and representational effectiveness. By employing VLMs as scalable evaluators and integrating multi-feature reconstruction baselines, the approach overcomes the fragmentation inherent in conventional metrics, enabling a more comprehensive, consistent, and scalable assessment of object-centric models’ representational capabilities in complex scenarios.
This study investigates whether object-centric representations confer advantages for compositional generalization—reasoning about novel combinations of familiar concepts—in visually rich scenes. To this end, the authors introduce three controlled visual question answering benchmarks (CLEVRTex, Super-CLEVR, and MOVi-C) and systematically compare visual encoders with and without object-centric biases, built upon DINOv2 and SigLIP, under rigorously matched conditions regarding data diversity, sample size, representation scale, and downstream computational resources. The work provides the first empirical validation of object-centric representations in diverse yet controlled settings, demonstrating their superior performance on challenging compositional tasks and higher sample efficiency, particularly outperforming dense representation approaches when data or computational budgets are limited.
This study investigates how the quality of slot representations in object-centric world models influences planning performance and out-of-distribution generalization, clarifying the role of their inductive biases. Through controlled visual model-predictive control experiments, the authors evaluate object-centric models against scene-centric baselines along two axes: representation quality and robustness to distribution shifts, introducing unsupervised slot-quality metrics (FG-ARI and mBO). Their findings reveal a positive correlation—subject to saturation—between slot quality and planning success; high-quality slots eliminate reliance on proprioceptive or mask-based biases. Notably, an object-centric model leveraging frozen DINO pretrained features (DINO-WM) achieves substantially improved planning performance and out-of-distribution robustness, outperforming the end-to-end scene-centric LeWM model when slot binding is effective.