gpu-accelerated simulation

Designs and implements GPU-accelerated, massively-parallel and SIMD/vectorized simulators and agent-based models—including batched environments and rollouts—by authoring GPU kernels, SIMD-friendly data layouts, memory and synchronization schemes, and parallel solvers to scale simulations to many concurrent instances and millions of agents and to support parallel deformable-physics computation. Builds and integrates GPU-based rendering pipelines (real-time, physically‑based, photorealistic and point‑based rendering) to provide high-fidelity 3D visual and sensor outputs for simulation, enabling high-throughput training and analysis.

gpu-acceleratedsimulation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.59
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the implicit trade-offs among physical fidelity, scaling efficiency, and execution determinism in GPU-accelerated robotic simulators, which undermine learning reliability and reproducibility in parallel training. The authors introduce GPUSimBench, a benchmark that evaluates sim-to-real dynamic consistency through an inclined-plane task and quantifies throughput, memory footprint, and inter-run as well as inter-environment non-determinism induced by GPU batching across varying scales. For the first time, the study systematically uncovers hidden limitations of mainstream platforms—such as Isaac Lab and Genesis—in scalability, physical consistency, and determinism, identifies four classes of stochastic mechanisms, and demonstrates that unconstrained parallelization significantly degrades reproducibility. These findings provide empirical guidance for building high-fidelity, scalable, and deterministic simulation infrastructures for embodied AI.

embodied AIexecution determinismGPU-accelerated simulators

This work addresses the inability of CUDA and Vulkan to execute compute and graphics tasks concurrently on GPUs due to scheduling isolation, which severely limits hardware utilization. To overcome this limitation, the authors present the first cross-ecosystem spatial sharing solution between CUDA and Vulkan. Their approach introduces driver-level mechanisms—including channel redirection, virtual address space merging, and page table grafting—to unify scheduling and memory address spaces without requiring data copies. A lightweight developer annotation API is also provided to facilitate integration. Evaluated on representative embodied AI workloads, the system achieves up to 85% higher throughput compared to a time-multiplexing baseline, while significantly reducing end-to-end latency and improving overall GPU utilization.

concurrent executionCUDA-Vulkan interoperabilityexecution isolation

ParallelKittens: Systematic and Practical Simplification of Multi-GPU AI Kernels

Nov 17, 2025
SH
Stuart H. Sul
🏛️ Stanford University

Modern AI models are scaling rapidly, while interconnect bandwidth growth lags, making multi-GPU communication a critical performance bottleneck. Existing overlap optimization techniques struggle to approach theoretical peak throughput under heterogeneous workloads and on emerging accelerators. This paper introduces the first systematic, general-purpose multi-GPU kernel design paradigm: it defines eight fundamental communication primitives and establishes a unified programming template, reducing complex kernel development to reusable, principle-based abstractions. Built upon CUDA extensions, the ThunderKittens framework integrates transmission mechanism modeling, resource-aware scheduling optimization, and overhead control. Evaluated on Hopper and Blackwell architectures, it achieves 2.33×–4.08× speedup over state-of-the-art baselines using fewer than 50 lines of device code—significantly improving both cross-architecture and cross-workload communication efficiency as well as developer productivity.

Addressing inter-GPU communication bottlenecks in modern AI workloadsOptimizing performance across heterogeneous workloads and new accelerator architecturesSimplifying development of overlapped multi-GPU kernels through reusable principles

GPUDrive: Data-driven, multi-agent driving simulation at 1 million FPS

Aug 02, 2024
SK
Saman Kazemkhani
🏛️ New York University | Stanford University

Multi-agent planning research is hindered by reliance on billions of simulation steps and low computational efficiency. Method: This paper introduces GPUDrive, a GPU-accelerated closed-loop multi-agent driving simulator. It pioneers the integration of heterogeneous agent behavioral modeling with low-level CUDA optimizations, achieving over one million frames per second throughput while retaining Python usability. Built upon the Madrona engine and custom CUDA C++ kernels, GPUDrive natively supports reinforcement learning (RL) frameworks and is fully compatible with the Waymo Open Motion Dataset. Contribution/Results: It enables RL training for single tasks in minutes and scales to thousands of scenarios within hours. Empirical evaluation on the Waymo dataset demonstrates efficient goal-directed driving performance. The codebase and pre-trained models are publicly released.

Accelerates training with GPU-optimized, high-speed simulations.Enables large-scale multi-agent driving simulations.Facilitates efficient reinforcement learning for complex agent behaviors.

Latest Papers

What's happening recently
View more

This work addresses the challenge that existing GPU performance models struggle to accurately simulate highly optimized large language model (LLM) kernels, which rely on fine-grained scheduling and compute–memory overlap, as traditional simulators are computationally expensive while analytical models are overly coarse. To bridge this gap, the authors propose a tile-centric simulation framework that, for the first time, models LLM kernels as tile-level dependency graphs. By combining an automated frontend for graph construction with a graph-driven backend simulator, the approach efficiently captures execution dependencies and overlapping behaviors. The framework supports GEMM, attention mechanisms, and end-to-end LLM inference, achieving average absolute percentage errors of 1.22%–8.71% on A100/H100 GPUs and successfully generalizing to the Blackwell architecture. This advancement significantly enhances simulation accuracy and scalability, facilitating effective hardware–software co-design for LLMs.

GPU simulationhardware-software co-designlarge language models

This work addresses the inefficiency of traditional CPU-based simulators, which are constrained by sequential execution, and existing GPU approaches, which suffer from kernel launch overhead and redundant global memory accesses in large-scale multi-agent financial market simulations. The authors propose a state-persistent, block-level parallel design pattern that caches mutable agent states in shared memory within GPU thread blocks, aggregates agent actions via shared-memory atomic operations, and co-executes matching logic. This approach substantially reduces critical-path depth and global memory traffic. Notably, it introduces state-persistent reduction to limit order book simulation for the first time, achieving step-count-independent memory access complexity while preserving high accuracy and reproducibility. Implemented on a lightweight custom CUDA engine, the system attains a peak throughput of 54.7 billion events per second—3,406× faster than CPU and 27.8× faster than PyTorch GPU—with one-tenth the memory footprint and statistical error below 0.1%.

GPU accelerationmarket simulationmulti-agent systems

This work addresses the tension between performance gains and scientific validity when porting large legacy scientific codes to GPUs by proposing a verification-centric, AI-assisted migration workflow. The approach integrates a large language model–driven CLI agent, OpenACC-based automated code transformation, and physics-informed kernel benchmark generation, ensuring consistency through both element-wise numerical comparison and application-level meteorological simulations. For the first time, scientific validation is deeply embedded into an AI-assisted porting pipeline, enabling automatic detection of floating-point semantic discrepancies and branch sensitivity, while highlighting the critical roles of conversational context management and runtime state reconstruction. Applied to the 250K-line Fortran weather model CReSS, the method successfully produced verified GPU implementations for 162 core kernels, achieving a 5.1× speedup in real typhoon simulations and uncovering five instances of numerical divergence, substantially reducing migration costs.

GPU portinglegacy codenumerical validation

Hot Scholars

AT

Andrea Tagliasacchi

Associate Prof, SFU; Research Scientist, Google DeepMind
3D Deep Learning
YL

Yu-Lun Liu

Assistant Professor, National Yang Ming Chiao Tung University
Computer VisionImage ProcessingMachine LearningDeep Learning
GK

George Kopanas

RunwayML, Member of Technical Staff - Team Lead
Gaussian SplattingNeRFNeural RenderingView Synthesis
YS

Yujun Shen

Ant Group
Generative ModelingComputer VisionDeep Learning
BS

Boxin Shi

Peking University
Computer VisionComputational Photography