SpecRoll: Fast-Slow Verifier-Feedback Adaptation for Speculative Reinforcement Learning Rollouts

๐Ÿ“… 2026-08-05
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the inefficiency of autoregressive rollout generation in reinforcement learning post-training and the inability of existing speculative decoding methods to adapt to continuously evolving policies. To overcome these limitations, we propose SpecRollโ€”a dual-timescale adaptive speculative rollout engine that significantly accelerates generation while strictly preserving the target policyโ€™s sampling distribution. SpecRoll employs a lightweight future-token head for parallel proposal generation and integrates a backpropagation-free Reflex module for local hidden-state correction along trajectories. It further incorporates concurrency-aware sparse-tree verification, exact target validation, adaptive fast-slow path routing, and delayed validator feedback. Evaluated across models ranging from 1.5B to 14B parameters and three mathematical reasoning benchmarks, SpecRoll achieves 1.26โ€“2.15ร— faster rollout generation and 1.21โ€“2.04ร— end-to-end speedup, consistently outperforming FastGRPO.
๐Ÿ“ Abstract
Reinforcement learning (RL) post-training improves the reasoning capabilities of large language models, but autoregressive rollout generation remains a major efficiency bottleneck. Speculative decoding can accelerate generation, yet applying it during RL is difficult because the target policy continually evolves: static proposers become stale, while frequent drafter updates add substantial overhead. We introduce SpecRoll, a speculative rollout engine that preserves the target model's sampling distribution while adapting at two timescales. Lightweight future-token heads generate parallel proposals, while our proposed Reflex module uses delayed verifier feedback to perform bounded, trajectory-local hidden-state corrections without backpropagation. A complementary slow path updates the head parameters only when sustained degradation is detected. SpecRoll combines these mechanisms with concurrency-aware sparse-tree verification and exact target verification, leaving the target rollout distribution and GRPO objective unchanged. Across five models ranging from 1.5B to 14B and three mathematical reasoning datasets, SpecRoll achieves 1.26-2.15x generation speedup and 1.21-2.04x end-to-end speedup over vanilla GRPO. It also outperforms FastGRPO in both generation and end-to-end time across all 15 matched settings, with an average pairwise end-to-end gain of 1.18x. Controlled ablations show that the fast and slow adaptation paths provide complementary benefits. Our source code is available at https://anonymous.4open.science/r/SpecRoll-26062006.
Problem

Research questions and friction points this paper is trying to address.

speculative decoding
reinforcement learning
rollout generation
efficiency bottleneck
policy adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

speculative decoding
reinforcement learning
dual-timescale adaptation
verifier feedback
fast-slow adaptation
๐Ÿ”Ž Similar Papers
No similar papers found.