Hybrid-Adaptive Thread Tuning to Mitigate Simulation Execution Bottlenecks in High-Performance Reinforcement Learning Inference

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses execution bottlenecks in reinforcement learning simulations caused by dynamic workloads and suboptimal thread configurations. To this end, the authors propose AutoThread, a novel approach that, for the first time, integrates physics-informed neural operators (PINO) with a finite-source M/M/1 queueing model to enable online prediction and dynamic tuning of thread counts. AutoThread further incorporates a workload-aware adaptive fine-tuning mechanism to enhance responsiveness to runtime variations. Experimental results demonstrate that, compared to static threading strategies, AutoThread achieves an average speedup of 18.4%, delivers 1.7× the throughput of XGBoost and 1.8× that of Reinforcer, and reduces maximum execution time by up to 83.8%.
📝 Abstract
In simulation-in-the-loop decision-making systems, reinforcement learning (RL) inference is often constrained by simulator-side execution overhead, where workloads are highly dynamic and sensitive to runtime thread configurations. Existing multithreaded strategies struggle to match thread resources before or during execution, causing resource contention, scheduling overhead, and reduced throughput. Through empirical analysis, we identify the ratio of task execution time to scheduling time as the key factor determining the optimal thread count. Building on this insight, we propose AutoThread, a hybrid adaptive thread-tuning method for mitigating simulation bottlenecks in RL inference. AutoThread employs a Physics-Informed Neural Operator (PINO) as a thread-count predictor and incorporates a finite-source M/M/1 queueing model to constrain and guide prediction, enabling fast and accurate estimation under dynamic workloads. It further performs load-aware online fine-tuning to compensate for prediction errors and refine resource allocation. Experiments show that AutoThread improves average speedup by 18.4\% over static strategies, achieves average throughput of 1.7x and 1.8x that of XGBoost and Reinforcer, respectively, and reduces execution time by up to 83.8\% compared with state-of-the-art methods. Our code and dataset are publicly available at https://github.com/suchenjm/AutoThread.
Problem

Research questions and friction points this paper is trying to address.

simulation bottlenecks
reinforcement learning inference
thread tuning
resource contention
dynamic workloads
Innovation

Methods, ideas, or system contributions that make the work stand out.

AutoThread
hybrid adaptive thread tuning
Physics-Informed Neural Operator
queueing model
reinforcement learning inference
🔎 Similar Papers
No similar papers found.
J
Jiming Su
College of Systems Engineering, National University of Defense Technology
H
Hantao Hua
College of Computer Science and Technology, National University of Defense Technology
L
Lujia Yin
College of Systems Engineering, National University of Defense Technology
Y
Yiping Yao
College of Systems Engineering, National University of Defense Technology
Feng Zhu
Feng Zhu
University of Technology Sydney
Computer VisionEfficient Computing