Score
Designs, builds, and evaluates AI agents and systems that are physically situated and interact with the real world or realistic simulations via sensors and actuators. This includes developing perception, control, planning and learning algorithms, integrating and calibrating sensor–actuator hardware and software, creating embodied simulation and sim‑to‑real pipelines, and measuring performance on tasks such as navigation, manipulation, locomotion, and human–robot interaction.
This paper addresses the weak generalization and poor adaptability of embodied intelligence in sim-to-real transfer. We propose a unified framework that synergistically integrates high-fidelity physics simulators—with their accurate modeling of external environmental dynamics—and world models—which learn compact, predictive internal representations. Our method jointly incorporates predictive coding, representation learning, and reinforcement learning to enable predictive planning and adaptive decision-making within the perception–reasoning–action loop. Key contributions include: (1) a systematic characterization of the complementary synergy between physics simulators and world models, establishing a novel paradigm bridging simulation-based training and real-world deployment; (2) an open-source, continuously updated literature repository comprehensively surveying technical approaches, state-of-the-art advances, and core challenges; and (3) a scalable theoretical and practical framework to enhance autonomy, cross-task generalization, and environmental adaptability of embodied AI systems.
Current AI systems exhibit a fundamental disconnect between physics-aware perception and symbolic physical reasoning, hindering the deep integration of physical laws into artificial intelligence. To address this, we propose the first unified framework that bridges perception, reasoning, modeling, and interaction—thereby realizing a next-generation world model grounded in both physical priors and embodied causal inference. Our approach formally unifies *theoretical physics reasoning* (axiomatic, formal deduction) with *practical physics understanding* (causal modeling grounded in embodied interaction), systematically integrating symbolic reasoning, generative modeling, multimodal perception, and first-principles physics modeling. The resulting structured physics-AI knowledge system is released as an open-source resource library. Empirical evaluation demonstrates substantial improvements in model interpretability, cross-scenario generalizability, and physical consistency—establishing foundational progress toward safe, generalizable, and verifiable physics-integrated intelligence.
Existing world models for embodied AI agents exhibit fragmentation in environmental prediction, intention recognition, and social context modeling, hindering coherent physical–social interaction. Method: We propose a unified “physical–mental” two-layer world modeling framework: a lower layer integrates multimodal perception, memory, and causal reasoning to model dynamic physical environments; an upper layer employs mental state inference to jointly represent user beliefs, goals, and social norms. The architecture enables cross-modal reasoning, long-horizon planning, and collaborative decision-making. Contribution/Results: Experiments demonstrate significant improvements in task completion rate, intention recognition accuracy, and human–robot interaction naturalness across both simulated and real-world settings. Our framework provides a scalable, human-aligned paradigm for advancing embodied intelligence toward human-like interactive capabilities.
To address performance degradation in sim-to-real transfer for embodied intelligence—caused by modeling discrepancies in physics simulators—this paper presents the first systematic, three-dimensional evaluation of mainstream engines (e.g., PyBullet, MuJoCo, Isaac Gym) along physical fidelity, task adaptability, and hardware constraints. We integrate cutting-edge techniques—including world models and geometrically equivariant networks—to establish a comprehensive benchmark featuring multi-task datasets, unified evaluation metrics, and an open-source platform. Furthermore, we propose a task-aware simulator selection framework that quantifies trade-offs among accuracy, real-time capability, differentiability, and deployment compatibility for navigation and manipulation tasks. Our contributions include an open-source evaluation repository and practical guidelines, providing both theoretical foundations and engineering evidence to reduce real-world training costs and enhance transfer robustness.
本文提出物理代理AI框架,通过将任务分解为阶段并分配给机器人技能对,结合语义规划与执行间的接口验证每个计划动作,以解决机器人团队协作中的不可行、时机不当或不安全行为问题。
This study addresses the deep coupling of perception, communication, and action in embodied intelligence, as well as how 6G networks can support distributed task collaboration and secure control. We propose a Perception-Communication-Action (PCA) architecture oriented toward the 6G orchestration plane. By designing mechanisms such as task state interfaces, semantic freshness metrics, and predictive digital twins, the architecture exposes task states and safety boundaries to the network layer, enabling cross-agent semantic communication and secure coordination. Multi-robot simulations demonstrate that this framework effectively delineates the functional boundaries between 5G and 6G, confirming that network orchestration significantly enhances task utility while underscoring the necessity of local autonomous control in ensuring system safety.
Embodied agents operating in physical environments face a fundamental “action poverty” issue: existing simulators provide only limited, domain-specific APIs, hindering cross-task generalization due to insufficiently expressive primitive action definitions. Method: We propose the first cross-domain API universe bootstrapped from human-authored tutorials (wikiHow), systematically inducing minimal, executable primitives by mapping tutorial instructions to grounded policies. Our iterative API induction framework integrates few-shot prompting with GPT-4, Python program generation, instruction-action alignment, and contextual policy grounding. Contribution/Results: Applied to just 0.5% of wikiHow, our method induces over 300 high-frequency, semantically grounded APIs. Human evaluation reveals that leading embodied simulators cover only nine of the top-50 induced APIs—quantifying for the first time the severity of action-space sparsity. This work establishes a foundational benchmark and design paradigm for principled primitive action space construction in embodied AI.
This work addresses a critical gap in the evaluation of action-conditioned world models (ACWMs), which has predominantly emphasized visual fidelity or task performance while neglecting their core function as physical simulators—namely, the causal fidelity between actions and environmental responses. To this end, we formalize the notion of an "observable simulator contract" and introduce WorldSimProbe, a fine-grained diagnostic framework that assesses ACWMs across five dimensions: local control sensitivity, global trajectory variation, multi-source action consistency, interaction grounding, and dynamics. Built upon controlled testing protocols, our framework leverages calibration analysis, dense action-motion correspondence, and spurious interaction detection. Evaluated on RoboTwin, ManiSkill, and LIBERO across six open-source models (>18,000 instances), it reveals systematic deficiencies in action execution, interaction grounding, and dynamics, with results strongly aligned with human judgment and downstream task performance.
This study addresses the challenges of manual reward engineering and the difficulty of integrating perception with control in robotic skill development by proposing the RPG framework. This method leverages offline data to construct simulated practice tasks and introduces a novel self-improvement mechanism that requires no model weight updates. Specifically, it employs multimodal large language models, combining privileged state information with video feedback to diagnose failures, dynamically reconstruct a reusable symbolic skill library, and iteratively optimize prompts. Experimental results demonstrate that this framework increases the success rate from 28.6% to 95.0% across 22 manipulation tasks, significantly outperforming existing baselines. Furthermore, it achieves a 100% success rate over 30 real-world physical trials, validating its superior generalization capability and practical utility.
This study addresses the lack of systematic synthesis at the intersection of artificial intelligence (AI) and modeling and simulation (M&S) by proposing, for the first time, a structured framework based on the full M&S lifecycle—encompassing model construction, input modeling, execution, experimentation, validation, and output analysis. It elucidates the bidirectional integration mechanisms between AI and simulation: how AI enhances or substitutes traditional simulation components, and how simulation supports AI training and evaluation. Incorporating generative AI technologies such as large language models, the paper identifies representative application paradigms and integration approaches across each phase, synthesizes key achievements, and presents a conceptual roadmap tailored to the rapidly evolving ecosystem, while highlighting current limitations and open research challenges.
论文探讨了通过融合物理-数字AI代理来解决无人机理解人类意图和飞行动态的问题,提出了5+5框架以实现适应性和持续进化的无人机自主性。
本文提出Robot Data Factory,通过持续生成、验证和重用机器人经验来解决物理AI中的知识获取问题,采用基础设施和方法论实现部署-测量-学习-重复的闭环。