Score
Designs and implements models and algorithms that evaluate and enforce the physical plausibility of scenes, objects, and motions by encoding kinematic constraints, dynamics, collisions, and contact geometry. Builds simulation and validation tools that predict scene‑specific object interactions and generate realizable target configurations and trajectories.
Current AIGC approaches for 3D/4D generation suffer from pervasive physical distortions—unrealistic deformations, unstable motion, and implausible interactions—stemming from an overreliance on appearance consistency while neglecting physical priors. This paper presents the first systematic survey of physics-driven AIGC methods, proposing a unified classification framework spanning multiple representations (e.g., NeRF, 3D Gaussian Splatting, multi-view geometry) and dimensions (3D/4D). We integrate rigid- and soft-body dynamics simulation, differentiable rendering, and physics-constrained modeling to delineate method applicability across material properties and dynamical regimes. Our analysis identifies shared limitations in structural stability, deformation plausibility, and interaction fidelity across existing works, and prescribes concrete optimization pathways. The study establishes theoretical foundations and practical guidelines for developing physically consistent, predictable, and editable generative models.
Reconstructing physically valid 3D scenes from single-view observations is a prerequisite for bridging the gap between visual perception and robotic control. However, in scenarios requiring precise contact reasoning, such as robotic manipulation in highly cluttered environments, geometric fidelity alone is insufficient. Standard perception pipelines often neglect physical constraints, resulting in invalid states, e.g., floating objects or severe inter-penetration, rendering downstream simulation unreliable. To address these limitations, we propose a novel physics-constrained Real-to-Sim pipeline that reconstructs physically consistent 3D scenes from single-view RGB-D data. Central to our approach is a differentiable optimization pipeline that explicitly models spatial dependencies via a contact graph, jointly refining object poses and physical properties through differentiable rigid-body simulation. Extensive evaluations in both simulation and real-world settings demonstrate that our reconstructed scenes achieve high physical fidelity and faithfully replicate real-world contact dynamics, enabling stable and reliable contact-rich manipulation.
Manual measurement of geometric and physical properties of real-world objects hinders scalable simulation asset creation. Method: This paper proposes an end-to-end, fully automatic Real2Sim pipeline that generates simulation-ready assets solely from unlabeled real-world robotic grasping interactions, using only robot joint torque sensors and external multi-view cameras. Contribution/Results: Its core innovation is an object-centric transparent alpha training strategy—first enabling direct foreground object decoupling and high-fidelity estimation of collision geometry, mass, and inertia tensor directly from photometric reconstructions (e.g., NeRF or Gaussian Splatting), without environmental modification or human intervention. The method integrates force sensing, multi-view visual reconstruction, physical identification, and foreground-background segmentation. Evaluated on diverse physical objects, the generated assets enable plug-and-play, high-fidelity dynamic simulation, significantly improving scalability and automation in robotic simulation dataset construction.
This work addresses the challenge of achieving explicitly physics-driven controllable dynamics in image-to-video generation, particularly the confusion of physical attributes in multi-object interactions. To this end, we propose PhyParam, a novel framework that integrates object-level forces, mass, friction, and scene gravity into the diffusion process via a lightweight physics-aware attention mechanism, further enhanced by semantic-structural feature supervision to improve dynamic modeling. We introduce PhyParam-Dataset, comprising 130,000 meticulously annotated video clips, and present the first approach enabling explicit physical control over rigid-body motion in image-to-video synthesis. Additionally, we establish PhyParam-Bench, a new evaluation benchmark assessing physical consistency across temporal dynamics, spatial stability, and semantic-physical alignment. Experiments demonstrate that our method significantly enhances physical plausibility while preserving high visual fidelity. Code, dataset, and benchmark are publicly released.
This work addresses the challenge of constructing high-fidelity digital twins from real-robot trajectories, where severe occlusions, noisy camera poses, and strong dynamic disturbances impede accurate modeling. We propose an end-to-end physics–vision co-optimization framework. Methodologically, we introduce the first hybrid scene representation integrating 3D Gaussian splatting (for appearance modeling) with explicit physical object meshes (for geometry and physics modeling). Our approach jointly optimizes geometry, appearance, robot pose, and rigid-body physical parameters—enabling unsupervised pose calibration and high-fidelity reconstruction. By unifying differentiable rendering with the differentiable MuJoCo physics engine, we achieve tight coupling between visual and physical optimization. Evaluated on the ALOHA 2 bimanual platform, our method achieves millimeter-accurate object mesh reconstruction, high-quality novel-view synthesis, and zero-shot pose calibration—significantly improving geometric fidelity, physical simulatability, and visual realism in real-to-simulation transfer.
This work addresses the lack of physical plausibility in single-image 3D reconstruction. We propose the first physics-compatible reconstruction framework that enforces static equilibrium as a hard constraint. Methodologically, we explicitly decouple and jointly optimize material stiffness, external loading forces, and the static equilibrium geometry; deformation responses are modeled via differentiable physics simulation, enabling gradient-based joint optimization of all variables. Our approach breaks from conventional simplifications—such as rigid-body assumptions or neglect of external forces—by embedding real-world physical constraints directly into the single-image reconstruction pipeline. Evaluated on Objaverse, our method yields reconstructions with significantly improved mechanical stability, suitable for downstream dynamic simulation and 3D printing. Physical validation via real-world force testing further confirms the structural robustness of the generated models.
Existing 3D generation methods often neglect physical properties or are confined to a single object category, failing to meet the demand for diverse and physically plausible assets in downstream simulation tasks. This work proposes PhysX-Omni, a unified framework that achieves joint generation of rigid, deformable, and articulated 3D objects with physical fidelity for the first time. Key innovations include a compression-free, high-resolution geometric representation tailored for vision-language models, the first universal simulation-ready 3D dataset—PhysXVerse—and PhysX-Bench, a comprehensive evaluation benchmark encompassing six-dimensional physical attributes. Experiments demonstrate that the proposed method excels on both conventional metrics and PhysX-Bench, significantly enhancing performance in downstream applications such as simulation scene generation and robot policy learning.
This work proposes Dream.exe, a framework that, for the first time, validates the executability of outputs from video generation models through real-world physical execution. Addressing the question of whether generated manipulation videos adhere to physical laws and can be realized by robots, the method integrates video generation models, trajectory extraction algorithms, and a physics simulator to establish an end-to-end video-to-execution evaluation pipeline. Evaluation across 101 manipulation tasks on eight model families reveals that certain models achieve notably high execution success rates, indicating their acquisition of effective physical priors from large-scale training data. The study further uncovers a significant disconnect between visual fidelity and executability, thereby advocating for new evaluation dimensions that go beyond purely perceptual metrics.
Manually editing heterogeneous collision meshes in bulk is time-consuming and poorly scalable, while existing automatic methods often fail to accurately capture user intent. This work introduces neural symbolic program synthesis to 3D collision mesh editing for the first time, formulating the task as a programming-by-example problem: users provide only a few edited examples, and the system automatically synthesizes a reusable program that generalizes to similar meshes. Evaluated on 24 tasks involving 600 meshes, the approach successfully completes 23 tasks, requiring an average of just 2.2 examples per task and synthesizing programs in approximately 3.5 seconds, thereby significantly improving both editing efficiency and scalability.
Existing physics-guided video generation methods rely on single-pass predictions of physical parameters, which struggle to accurately capture user intent—particularly in modeling fine-grained dynamics, complex trajectories, and temporally coherent interactions. To address this limitation, this work proposes a reflective agent framework that treats physical programs as executable hypotheses and iteratively refines motion and interaction through a closed-loop “generate–simulate–verify–repair” process. The framework integrates a vision-language model, a physics simulation engine, and a dedicated control API to enable multi-stage user interaction and progressively realize precise event outcomes. Experimental results demonstrate that the generated videos significantly outperform existing approaches in terms of physical plausibility, alignment with input prompts, and generalization across diverse scenarios.