KineticSim: A Lightweight, High-Performance Execution Engine for Real-Time Market Simulators

๐Ÿ“… 2026-06-19
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the inefficiency of traditional CPU-based simulators, which are constrained by sequential execution, and existing GPU approaches, which suffer from kernel launch overhead and redundant global memory accesses in large-scale multi-agent financial market simulations. The authors propose a state-persistent, block-level parallel design pattern that caches mutable agent states in shared memory within GPU thread blocks, aggregates agent actions via shared-memory atomic operations, and co-executes matching logic. This approach substantially reduces critical-path depth and global memory traffic. Notably, it introduces state-persistent reduction to limit order book simulation for the first time, achieving step-count-independent memory access complexity while preserving high accuracy and reproducibility. Implemented on a lightweight custom CUDA engine, the system attains a peak throughput of 54.7 billion events per secondโ€”3,406ร— faster than CPU and 27.8ร— faster than PyTorch GPUโ€”with one-tenth the memory footprint and statistical error below 0.1%.
๐Ÿ“ Abstract
Simulating financial markets at scale with multi-agent (Agent-Based) models is critical for market design, regulatory stress-testing, and reinforcement learning, but traditional CPU simulators are bottlenecked by sequential processing while vectorized GPU frameworks suffer from kernel-launch overhead and redundant global-memory round-trips. We formalize, analyze, and evaluate a reusable parallel design pattern: persistent, state-carrying clearing for iterative multi-agent reductions. By caching mutable simulation state in thread-block shared memory across step boundaries, aggregating agent actions via shared-memory atomics, and resolving the clearing function cooperatively, the pattern reduces the per-step critical-path depth from Theta(L+A) for sequential clearing (L price-grid ticks, A agents) to Theta(log L + ceil(A/L)) and makes global-memory traffic independent of the step count. We implement this in KineticSim, a lightweight GPU execution engine that simulates massive ensembles of limit-order books in parallel, reaching a peak throughput of over 54.7 billion agent-events per second. On a fixed workload it delivers speedups of 3406x over CPU (NumPy), 27.8x over PyTorch GPU, 42.8x over JAX GPU, and 8.4x over a naive custom CUDA baseline, while using roughly an order of magnitude less GPU memory than PyTorch. Across 53 configurations the two custom CUDA engines produce bitwise-identical order books, and aggregate statistics match the CPU reference to within 0.1%. The pattern generalizes to other iterative multi-agent workloads requiring state-persistent, block-localized reductions.
Problem

Research questions and friction points this paper is trying to address.

market simulation
multi-agent systems
GPU acceleration
performance bottleneck
real-time simulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

persistent clearing
shared-memory atomics
multi-agent simulation
GPU acceleration
order book simulation
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
S
Shakya Jayakody
Independent Researcher, Orlando, FL USA
P
Prarthinie Jayakody
Independent Researcher, Colombo, Sri Lanka