Score
Designs and implements physics-grounded simulation environments and device-aware simulators that model real-world physical processes to generate synthetic data and to embed known ground-truth treatment effects. Builds simulation-grounded validation and sim-in-the-loop frameworks that introduce realistic adoption biases and confounding, support benchmarking of estimators and decision systems, detect physically or physiologically implausible outputs, and close the loop between simulation and model or content generation.
To address performance degradation in sim-to-real transfer for embodied intelligence—caused by modeling discrepancies in physics simulators—this paper presents the first systematic, three-dimensional evaluation of mainstream engines (e.g., PyBullet, MuJoCo, Isaac Gym) along physical fidelity, task adaptability, and hardware constraints. We integrate cutting-edge techniques—including world models and geometrically equivariant networks—to establish a comprehensive benchmark featuring multi-task datasets, unified evaluation metrics, and an open-source platform. Furthermore, we propose a task-aware simulator selection framework that quantifies trade-offs among accuracy, real-time capability, differentiability, and deployment compatibility for navigation and manipulation tasks. Our contributions include an open-source evaluation repository and practical guidelines, providing both theoretical foundations and engineering evidence to reduce real-world training costs and enhance transfer robustness.
This work addresses the frequent mismatch between user-specified physical intent and the actual behavior of multiphysics simulation code generated by large language models, often due to erroneous implementations of partial differential equations (PDEs). To bridge this gap, we propose a PDE-structure-based intent verification method that deterministically reconstructs the governing equations implicitly encoded in the generated code and compares them against the user’s intended PDEs, enabling semantic correctness validation and iterative refinement. We introduce, for the first time, a formal metric termed the Intent Fidelity Score (IFS) to quantify alignment with physical intent, establish a PDE-driven feedback loop, and demonstrate compatibility with major PDE frameworks including MOOSE, FEniCS, and FreeFEM. Evaluated on 220 cases in MooseBench, our approach substantially improves IFS—by 0.22–0.41 on challenging instances with initial IFS < 0.7—while audits reveal that execution-only repair strategies still yield physically incorrect results in 39–40% of cases.
This paper addresses the lack of theoretical foundations for Simulation-Grounded Neural Networks (SGNNs), particularly in scientific modeling scenarios where ground-truth labels are unavailable or scarce. Methodologically, it introduces a simulator-driven learning framework that generates synthetic data from mechanistic simulators and integrates Bayesian inference approximation with neural network amortization. Theoretically, it rigorously proves that SGNNs converge to the Bayes-optimal predictor under mild regularity conditions. The approach enables identification and prediction of unobserved latent variables and ensures posterior-consistent scientific interpretability via simulator-based mechanistic attribution. Experiments demonstrate that the model accurately recovers latent parameters, achieves half the error rate of AIC on dynamical mechanism discrimination tasks, and exhibits strong robustness under model misspecification—significantly outperforming conventional statistical methods.
This work addresses the limitation of existing world model evaluations, which overly rely on visual fidelity and fail to reliably assess whether generated multi-agent dynamics adhere to fundamental physical laws. To this end, we propose CrashTwin, a novel framework that, for the first time, enables metric-scale physical quantity recovery from video replays without requiring camera calibration. We introduce a large-scale multi-agent dataset encompassing both synthetic and real-world collision scenarios. Leveraging uncalibrated 3D reconstruction and verification against physical conservation laws—such as momentum and kinetic energy—CrashTwin establishes a multidimensional diagnostic benchmark centered on physical plausibility. Experiments reveal that state-of-the-art world models, despite achieving high perceptual quality, frequently violate basic physics; CrashTwin effectively uncovers these failure modes, offering a crucial tool for developing physically reliable simulation models.
This work addresses the absence of a unified, measurement-based evaluation framework for assessing how accurately existing physics simulators and video world models reproduce real-world physical dynamics. To this end, we introduce GAUGE—a diagnostic benchmark comprising 22 controlled tasks that, for the first time, integrates real-world trajectories, uncertainty annotations, and task-specific observables to enable fine-grained, interpretable evaluation of core physical processes such as collisions, friction, and deformation. Through analyses of generalized trajectory error, consistency with physical laws, and temporal parameter stability, GAUGE reveals significant discrepancies in mainstream simulators during impulsive contacts, rapid cloth motion, and volumetric deformation. Moreover, while video world models can often fit trajectory shapes, they frequently misestimate acceleration, momentum transfer, and oscillation timing.
This work addresses the ambiguity in physics-based simulations described in natural language, which often leads to modeling inconsistencies that compromise correctness and reproducibility. To tackle this challenge, the authors propose a closed-loop modeling framework centered on simulator validation, employing an “interpret–execute–verify” cycle that integrates document retrieval, code generation, static analysis, and solver diagnostics to explicitly identify and resolve modeling uncertainties. Built upon the differentiable Julia simulator JutulDarcy, the intelligent agent JutulGPT enables explicit logging and interactive refinement of modeling assumptions. The study further introduces a novel approach to reconstruct reference models directly from textual descriptions to audit reproducibility. All code, prompts, and logs are publicly released, and experiments demonstrate the framework’s effectiveness in uncovering hidden degrees of freedom introduced by default parameters, establishing a verifiable and traceable paradigm for scientific modeling.
This study addresses the unclear impact of world fidelity and human behavioral similarity in simulated data on policy performance. To this end, we propose a Real2Sim2Real co-training framework that decouples and independently modulates world grounding and behavior grounding during simulation data generation. By incorporating dynamic dexterous manipulation tasks and latent space analysis, this work reveals the complementary mechanisms through which both forms of grounding enhance policy performance. Experimental results demonstrate that full grounding improves the success rate from 52% to 86%, validating the critical value of highly grounded simulated data for foundation model co-training.
研究针对赛博物理系统设计中的现实差距问题,提出一种识别、分类、描述并利用设计错觉的方法,将其转化为可操作的设计知识。
This study addresses the labor-intensive nature of designing assets, layouts, and parameters for physics simulation scenes, which hinders the rapid generation of executable dynamic scenarios from text. To overcome this, we propose a hierarchical agent pipeline built upon the Genesis engine, featuring a Planner-Writer-Critic architecture that constitutes the first agent framework specifically tailored for simulation. Furthermore, we introduce compact Debug Cards distilled from graphical demonstrations to enable character-specific physical guidance and automated execution repair. Across 42 evaluation tasks, our method consistently surpasses existing baselines in both physical fidelity and visual quality. User preference studies under blind testing conditions demonstrate significant improvements, while the framework effectively supports multimodal downstream applications such as dataset construction.
This study addresses the catastrophic forgetting and prediction imbalance arising from simulation-to-experiment fine-tuning by proposing a multi-objective joint training paradigm. Specifically, simulation and experimental predictions are formulated as a multi-objective optimization task, where domain adaptation and spatiotemporal physical system modeling are leveraged to jointly optimize risks across both domains, effectively overcoming the limitations of conventional sequential fine-tuning. Evaluated on tasks such as fluid dynamics, the proposed method achieves optimal balanced performance under diverse weight configurations. It significantly enhances cross-domain retention capabilities while substantially improving simulation-domain performance without compromising features inherent to the scarce experimental data.
This study addresses the "correct reasoning yet suboptimal action" problem in autonomous driving by redefining the groundedness of physical intelligence and proposing the GroundAct framework. GroundAct establishes the concept of grounded planning, which treats entities as fundamental units and employs lightweight reference tokens to associate states with interactions. This design enables an explicit mapping from symbolic reasoning to planning, effectively bridging the gap between reasoning and action. By integrating closed-loop simulation with open-loop planning algorithms, experimental results demonstrate that the proposed method exhibits strong open-loop planning capabilities across routine, out-of-distribution, and safety-critical scenarios. Furthermore, the framework's effectiveness is validated through closed-loop driving evaluations, confirming its potential for robust and reliable deployment in real-world autonomous driving applications.