GPUSimBench: Towards Scalable and Reliable GPU-Accelerated Simulators in Embodied AI

📅 2026-07-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the implicit trade-offs among physical fidelity, scaling efficiency, and execution determinism in GPU-accelerated robotic simulators, which undermine learning reliability and reproducibility in parallel training. The authors introduce GPUSimBench, a benchmark that evaluates sim-to-real dynamic consistency through an inclined-plane task and quantifies throughput, memory footprint, and inter-run as well as inter-environment non-determinism induced by GPU batching across varying scales. For the first time, the study systematically uncovers hidden limitations of mainstream platforms—such as Isaac Lab and Genesis—in scalability, physical consistency, and determinism, identifies four classes of stochastic mechanisms, and demonstrates that unconstrained parallelization significantly degrades reproducibility. These findings provide empirical guidance for building high-fidelity, scalable, and deterministic simulation infrastructures for embodied AI.
📝 Abstract
Data-driven embodied AI is rapidly transitioning into a paradigm that scales training through massively parallel simulation, where GPU-accelerated simulators serve as the foundational data infrastructure. However, as computational throughput scales, the underlying trade-offs between parallel efficiency, physical fidelity, and execution determinism remain largely unexamined, hindering the development of reliable robot learning. In this paper, we expose the hidden limits of mainstream GPU-based robotic simulators (e.g., Isaac Lab, Genesis) by introducing GPUSimBench, which focuses on scalability, physical consistency, and computational determinism. First, GPUSimBench establishes a physical grounding evaluation with a controlled inclined-plane task, quantifying the distributional alignment between simulated dynamics and their real-world counterparts. Second, we benchmark parallel scalability by measuring throughput and memory footprints across scaling environment counts. Crucially, beyond standard performance metrics, we unveil and quantify the inherent non-determinism introduced by GPU-batched execution, characterized by significant run-to-run and inter-environment variability even under identical initial conditions. Finally, we identify four empirical regimes of stochasticity within current simulator stacks, highlighting that unbounded scaling can compromise reproducibility without explicit constraints.
Problem

Research questions and friction points this paper is trying to address.

GPU-accelerated simulators
embodied AI
physical fidelity
execution determinism
parallel scalability
Innovation

Methods, ideas, or system contributions that make the work stand out.

GPU-accelerated simulation
scalability
physical fidelity
execution determinism
embodied AI
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Huzhenyu Zhang
Shanghai AI Laboratory, China; MIRAI, Russia
S
Shenghai Yuan
Nanyang Technological University, Singapore
W
Wenrui Yan
Shanghai AI Laboratory, China
L
Li Ma
Shanghai AI Laboratory, China
H
Hengjie Li
Shanghai AI Laboratory, China
J
Jingcheng Pang
Nanjing University, China
D
Dmitry Yudin
MIRAI, Russia; AXXX, Moscow, Russia