๐ค AI Summary
This work addresses the inefficiency of traditional CPU-based simulators, which are constrained by sequential execution, and existing GPU approaches, which suffer from kernel launch overhead and redundant global memory accesses in large-scale multi-agent financial market simulations. The authors propose a state-persistent, block-level parallel design pattern that caches mutable agent states in shared memory within GPU thread blocks, aggregates agent actions via shared-memory atomic operations, and co-executes matching logic. This approach substantially reduces critical-path depth and global memory traffic. Notably, it introduces state-persistent reduction to limit order book simulation for the first time, achieving step-count-independent memory access complexity while preserving high accuracy and reproducibility. Implemented on a lightweight custom CUDA engine, the system attains a peak throughput of 54.7 billion events per secondโ3,406ร faster than CPU and 27.8ร faster than PyTorch GPUโwith one-tenth the memory footprint and statistical error below 0.1%.
๐ Abstract
Simulating financial markets at scale with multi-agent (Agent-Based) models is critical for market design, regulatory stress-testing, and reinforcement learning, but traditional CPU simulators are bottlenecked by sequential processing while vectorized GPU frameworks suffer from kernel-launch overhead and redundant global-memory round-trips. We formalize, analyze, and evaluate a reusable parallel design pattern: persistent, state-carrying clearing for iterative multi-agent reductions. By caching mutable simulation state in thread-block shared memory across step boundaries, aggregating agent actions via shared-memory atomics, and resolving the clearing function cooperatively, the pattern reduces the per-step critical-path depth from Theta(L+A) for sequential clearing (L price-grid ticks, A agents) to Theta(log L + ceil(A/L)) and makes global-memory traffic independent of the step count. We implement this in KineticSim, a lightweight GPU execution engine that simulates massive ensembles of limit-order books in parallel, reaching a peak throughput of over 54.7 billion agent-events per second. On a fixed workload it delivers speedups of 3406x over CPU (NumPy), 27.8x over PyTorch GPU, 42.8x over JAX GPU, and 8.4x over a naive custom CUDA baseline, while using roughly an order of magnitude less GPU memory than PyTorch. Across 53 configurations the two custom CUDA engines produce bitwise-identical order books, and aggregate statistics match the CPU reference to within 0.1%. The pattern generalizes to other iterative multi-agent workloads requiring state-persistent, block-localized reductions.