Score
Designs and implements computational world models that represent entities as collections of parts, capturing each part's geometry, appearance, and kinematic relationships so the system encodes joint visual and motion/kinematic distributions. Builds methods to infer part-level structure and dynamics from observations, to reconstruct or generate articulated objects (e.g., in 3D) from sensory input, and to enable generalization to novel or out-of-distribution objects and configurations.
This work addresses the lack of a unified predictive framework for world models in robotic manipulation, which has led to fragmented research and ambiguous design choices. Focusing on three core questions—what to predict, how to relate predictions to actions, and when to use predictions—the paper proposes a functional taxonomy that distinguishes between integrated prediction-action models and explicit predictive planners, positioning world models as foundational predictive infrastructure for robot learning. The study presents a systematic review covering latent dynamics models, action-conditioned video generation, 3D/4D scene prediction, physics simulators, and prediction modules in vision–language–action systems. It also consolidates evaluation protocols across 34 manipulation datasets, highlighting open challenges such as contact modeling and hallucination control through the lenses of prediction fidelity, task performance, and simulation reliability.
Autonomous robotic manipulation critically requires world models capable of understanding the physical mechanisms and dynamics of the environment; however, existing definitions remain ambiguous and capability boundaries ill-defined, hindering both generality and practical deployment. Method: This work adopts a functional perspective to systematically characterize the core capabilities of world models for robotic manipulation, proposing a task-driven, unified framework that integrates state representation learning, dynamic modeling, sequence prediction, closed-loop control, and model-based reinforcement learning. Contribution: We formally identify the essential components and functional roles of world models within the perception–prediction–control loop for the first time. Moving beyond conventional static representations, we emphasize dynamic modeling, causal reasoning, and online planning as critical capabilities. Furthermore, we provide a clear capability taxonomy and a systematic construction methodology, enabling the development of generalizable, scalable world models for robotics.
This paper presents a systematic survey of state-of-the-art 3D modeling techniques for articulated objects, addressing the fundamental challenge of jointly modeling geometry—i.e., part structure and shape—and motion—i.e., dynamics and kinematic constraints. It establishes, for the first time, a unified taxonomy of co-design paradigms that integrate geometric and motion reasoning, rigorously defines task boundaries, and identifies core bottlenecks in generalization, physical plausibility, and cross-domain transfer. The survey comprehensively covers major technical approaches, including optimization-based methods, multi-view and single-image reconstruction, NeRFs, implicit neural representations, graph neural networks, and physics-based simulation. Building on this analysis, the authors propose the first structured classification framework for articulated object modeling, distilling seven key challenges and four concrete future research directions. This work serves as an authoritative benchmark and roadmap for researchers in computer vision, computer graphics, and robotics.
This work addresses the challenging problem of generating articulated 3D furniture models from a single input image. We propose the first end-to-end, single-image-driven method, overcoming key limitations of prior approaches—namely, their reliance on multi-view or multi-state inputs and coarse-grained joint modeling. Methodologically, we introduce a geometry-motion joint diffusion model that unifies part-level fine-grained articulation synthesis with physically plausible motion constraints. We further design a cross-domain coarse-to-fine generation framework, explicit part connectivity graph modeling, and an abstraction-aware representation mechanism. Quantitatively and qualitatively, our method achieves state-of-the-art performance in realism, image fidelity, and structural reconstruction accuracy. It significantly advances practical utility and generalization capability for single-image articulated object modeling, establishing new benchmarks in this emerging domain.
This work addresses the problem of physically plausible completion of missing parts in interactive objects. We propose a diffusion-based generative method that, for the first time, explicitly incorporates physical constraints—such as structural stability and kinematic mobility—as differentiable loss terms within the diffusion sampling process. Our approach integrates classifier-free geometric guidance with a motion success rate evaluation mechanism, enabling part dependency modeling and hierarchical sequential generation. Quantitative evaluation demonstrates that our method achieves state-of-the-art performance across both geometric fidelity (e.g., Chamfer distance, F-Score) and physical plausibility metrics. Crucially, the newly introduced motion success rate metric empirically validates the high physical credibility of generated parts. Experiments confirm end-to-end applicability to downstream tasks including 3D printing, robotic manipulation, and modeling of complex assemblies.
This study addresses the limitations of traditional physics simulators in robotics—such as restricted expressiveness due to simplifying assumptions, high data costs, and difficulties in modeling complex physical interactions—by systematically reviewing video generation models as embodied world models. Integrating high-fidelity, multimodal-conditioned video synthesis with imitation learning, reinforcement learning, and visual planning frameworks, this work provides the first comprehensive analysis of their potential and limitations in tasks including action prediction, dynamics modeling, and policy evaluation. The review highlights breakthroughs in high-fidelity modeling of physical interactions while identifying key challenges in instruction following, physical consistency, and safety. These insights lay a theoretical foundation and outline future directions for replacing conventional simulators and enabling deployment in safety-critical scenarios.
This paper addresses the challenging problem of reconstructing multi-instance, structurally and articulationally variable man-made articulated objects from a single RGB-D image. We propose the first detection-driven, part-level joint modeling framework. Our method follows a “detect-group-fuse” paradigm: (1) detecting individual part instances; (2) performing kinematics-aware test-time part grouping to enable structure-adaptive assembly; (3) introducing anisotropic scale normalization to mitigate inter-part size variation; and (4) designing a cross-space optimization mechanism that jointly refines shape reconstruction, pose estimation, and kinematic parameters—including joint type, axis orientation, and motion range—for mutual consistency. The end-to-end trainable framework significantly outperforms prior methods on both synthetic and real-world benchmarks. It robustly handles complex topologies and false detections, while achieving state-of-the-art accuracy in part-level shape reconstruction and kinematic parameter estimation.
This work proposes a Geometric Primary Structure (GPS) representation to enhance robotic manipulation of articulated objects, enabling the modeling of both geometric and kinematic properties of movable parts from a single RGB-D image. To support scalable and high-quality data annotation, the authors develop a VR-GPS system that leverages portable virtual reality devices, effectively balancing annotation efficiency with data fidelity. Building upon GPS predictions, a heuristic manipulation policy is designed to guide robotic task execution. The approach is evaluated on a large-scale dataset comprising 234 articulated objects and 41K frames; without any in-domain fine-tuning, it achieves a 73% success rate across 270 initial configurations spanning nine object categories.
Existing methods struggle to generalize to unseen categories or AI-generated articulated 3D objects due to scarce annotated data. This work proposes integrating instructional kinematic specifications—comprising part descriptions, connectivity relationships, and joint types—into the reconstruction pipeline, enabling, for the first time, specification-guided end-to-end motion structure prediction. By constructing a heterogeneous 3D dataset of over 150,000 instances and leveraging vision-language models to automatically generate specifications during inference, the approach supports multi-granularity annotations and effectively unifies heterogeneous data sources. Experiments demonstrate that the method substantially improves generalization across object categories and AI-generated meshes, achieving high-quality reconstruction of articulated 3D assets from a single input image.
Existing world models struggle to infer the complete physical structure of scenes and interactions among objects from partially observed videos. This work proposes a novel probabilistic world model based on autoregressive sequence modeling, which enables efficient training and supports conditional estimation over arbitrary visual variables—such as appearance and dynamics. The model generates multimodal future states over multiple steps, automatically discovers objects and their subparts, and facilitates 3D manipulation and physical reasoning. Experiments demonstrate that the model successfully extracts hierarchical object structures in tasks such as Visual Jenga, significantly enhancing the understanding of complex physical interactions.
This work addresses the challenge of generating articulated 3D objects from a single image, where accurately inferring kinematic structures remains difficult due to insufficient static visual cues, error propagation in two-stage pipelines, and scarcity of motion-labeled data. To overcome these limitations, we propose PWM-ArtGen, a unified part-based world model that, for the first time, formulates articulated objects as dynamic systems. By coupling image diffusion with action diffusion, our approach enables joint training of visual and action branches without requiring explicit kinematic annotations. We introduce a dual-diffusion architecture with independent timesteps and construct a large-scale dataset of part-level image pairs to support this framework. Experiments demonstrate that our method significantly outperforms existing baselines in static pose generation and exhibits strong zero-shot generalization, effectively handling complex out-of-distribution real-world objects.
The field of 3D vision suffers from fragmented data representations, learning paradigms, and benchmarking protocols, leading to a lack of unified understanding regarding efficiency, fidelity, and scalability. This work proposes the first cohesive conceptual framework that integrates geometric representations—such as point clouds, meshes, voxels, and 3D Gaussians—with diverse learning paradigms—including 2D-supervised learning, implicit neural representations, and 4D modeling—and connects them to real-world application scenarios. By constructing a structured knowledge graph of 3D vision, the study systematically relates dataset design, supervision mechanisms, and task requirements, clarifying the trade-offs between efficiency and fidelity and charting pathways for multimodal geometric grounding. This framework offers systematic guidance for reconstruction, generation, and dynamic scene modeling, advancing the field toward a unified and efficient paradigm.