Institution profile

BespokeLabs

Research institutionnorthamerica · us
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

ThunderSyncRL: Lossless Acceleration of Agentic Reinforcement Learning

Oct 05, 2026

This study addresses the inherent trade-off in agent reinforcement learning between training engine idleness under synchronous execution and policy staleness under asynchronous execution. To reconcile this dilemma, we propose a fine-grained gradient flow mechanism that guarantees zero policy staleness. By processing single-trajectory or single-round data on the fly within Group Relative Policy Optimization (GRPO) and Online Policy Distillation (OPD), our method is theoretically proven to yield updates mathematically identical to those of batch synchronous training. Experimental results demonstrate that the proposed mechanism accelerates training by up to 1.9× compared to synchronous baselines while improving performance by up to 2.47 percentage points over asynchronous methods under a fixed computational budget, thereby successfully unifying training efficiency with optimization accuracy.

0 citationsRead paper
Recent publications

Latest Papers

ThunderSyncRL: Lossless Acceleration of Agentic Reinforcement Learning

Oct 05, 2026

This study addresses the inherent trade-off in agent reinforcement learning between training engine idleness under synchronous execution and policy staleness under asynchronous execution. To reconcile this dilemma, we propose a fine-grained gradient flow mechanism that guarantees zero policy staleness. By processing single-trajectory or single-round data on the fly within Group Relative Policy Optimization (GRPO) and Online Policy Distillation (OPD), our method is theoretically proven to yield updates mathematically identical to those of batch synchronous training. Experimental results demonstrate that the proposed mechanism accelerates training by up to 1.9× compared to synchronous baselines while improving performance by up to 2.47 percentage points over asynchronous methods under a fixed computational budget, thereby successfully unifying training efficiency with optimization accuracy.

0 citationsRead paper