Score
Designs and implements GPU-accelerated, massively-parallel and SIMD/vectorized simulators and agent-based models—including batched environments and rollouts—by authoring GPU kernels, SIMD-friendly data layouts, memory and synchronization schemes, and parallel solvers to scale simulations to many concurrent instances and millions of agents and to support parallel deformable-physics computation. Builds and integrates GPU-based rendering pipelines (real-time, physically‑based, photorealistic and point‑based rendering) to provide high-fidelity 3D visual and sensor outputs for simulation, enabling high-throughput training and analysis.
This work addresses the implicit trade-offs among physical fidelity, scaling efficiency, and execution determinism in GPU-accelerated robotic simulators, which undermine learning reliability and reproducibility in parallel training. The authors introduce GPUSimBench, a benchmark that evaluates sim-to-real dynamic consistency through an inclined-plane task and quantifies throughput, memory footprint, and inter-run as well as inter-environment non-determinism induced by GPU batching across varying scales. For the first time, the study systematically uncovers hidden limitations of mainstream platforms—such as Isaac Lab and Genesis—in scalability, physical consistency, and determinism, identifies four classes of stochastic mechanisms, and demonstrates that unconstrained parallelization significantly degrades reproducibility. These findings provide empirical guidance for building high-fidelity, scalable, and deterministic simulation infrastructures for embodied AI.
This work addresses the inability of CUDA and Vulkan to execute compute and graphics tasks concurrently on GPUs due to scheduling isolation, which severely limits hardware utilization. To overcome this limitation, the authors present the first cross-ecosystem spatial sharing solution between CUDA and Vulkan. Their approach introduces driver-level mechanisms—including channel redirection, virtual address space merging, and page table grafting—to unify scheduling and memory address spaces without requiring data copies. A lightweight developer annotation API is also provided to facilitate integration. Evaluated on representative embodied AI workloads, the system achieves up to 85% higher throughput compared to a time-multiplexing baseline, while significantly reducing end-to-end latency and improving overall GPU utilization.
Modern AI models are scaling rapidly, while interconnect bandwidth growth lags, making multi-GPU communication a critical performance bottleneck. Existing overlap optimization techniques struggle to approach theoretical peak throughput under heterogeneous workloads and on emerging accelerators. This paper introduces the first systematic, general-purpose multi-GPU kernel design paradigm: it defines eight fundamental communication primitives and establishes a unified programming template, reducing complex kernel development to reusable, principle-based abstractions. Built upon CUDA extensions, the ThunderKittens framework integrates transmission mechanism modeling, resource-aware scheduling optimization, and overhead control. Evaluated on Hopper and Blackwell architectures, it achieves 2.33×–4.08× speedup over state-of-the-art baselines using fewer than 50 lines of device code—significantly improving both cross-architecture and cross-workload communication efficiency as well as developer productivity.
Multi-agent planning research is hindered by reliance on billions of simulation steps and low computational efficiency. Method: This paper introduces GPUDrive, a GPU-accelerated closed-loop multi-agent driving simulator. It pioneers the integration of heterogeneous agent behavioral modeling with low-level CUDA optimizations, achieving over one million frames per second throughput while retaining Python usability. Built upon the Madrona engine and custom CUDA C++ kernels, GPUDrive natively supports reinforcement learning (RL) frameworks and is fully compatible with the Waymo Open Motion Dataset. Contribution/Results: It enables RL training for single tasks in minutes and scales to thousands of scenarios within hours. Empirical evaluation on the Waymo dataset demonstrates efficient goal-directed driving performance. The codebase and pre-trained models are publicly released.
This work addresses the challenge that existing GPU performance models struggle to accurately simulate highly optimized large language model (LLM) kernels, which rely on fine-grained scheduling and compute–memory overlap, as traditional simulators are computationally expensive while analytical models are overly coarse. To bridge this gap, the authors propose a tile-centric simulation framework that, for the first time, models LLM kernels as tile-level dependency graphs. By combining an automated frontend for graph construction with a graph-driven backend simulator, the approach efficiently captures execution dependencies and overlapping behaviors. The framework supports GEMM, attention mechanisms, and end-to-end LLM inference, achieving average absolute percentage errors of 1.22%–8.71% on A100/H100 GPUs and successfully generalizing to the Blackwell architecture. This advancement significantly enhances simulation accuracy and scalability, facilitating effective hardware–software co-design for LLMs.
This work addresses the inefficiency of traditional CPU-based simulators, which are constrained by sequential execution, and existing GPU approaches, which suffer from kernel launch overhead and redundant global memory accesses in large-scale multi-agent financial market simulations. The authors propose a state-persistent, block-level parallel design pattern that caches mutable agent states in shared memory within GPU thread blocks, aggregates agent actions via shared-memory atomic operations, and co-executes matching logic. This approach substantially reduces critical-path depth and global memory traffic. Notably, it introduces state-persistent reduction to limit order book simulation for the first time, achieving step-count-independent memory access complexity while preserving high accuracy and reproducibility. Implemented on a lightweight custom CUDA engine, the system attains a peak throughput of 54.7 billion events per second—3,406× faster than CPU and 27.8× faster than PyTorch GPU—with one-tenth the memory footprint and statistical error below 0.1%.
This work addresses the tension between performance gains and scientific validity when porting large legacy scientific codes to GPUs by proposing a verification-centric, AI-assisted migration workflow. The approach integrates a large language model–driven CLI agent, OpenACC-based automated code transformation, and physics-informed kernel benchmark generation, ensuring consistency through both element-wise numerical comparison and application-level meteorological simulations. For the first time, scientific validation is deeply embedded into an AI-assisted porting pipeline, enabling automatic detection of floating-point semantic discrepancies and branch sensitivity, while highlighting the critical roles of conversational context management and runtime state reconstruction. Applied to the 250K-line Fortran weather model CReSS, the method successfully produced verified GPU implementations for 162 core kernels, achieving a 5.1× speedup in real typhoon simulations and uncovering five instances of numerical divergence, substantially reducing migration costs.