Reward-rate Policy Gradient for Efficient Machine Learning Engineering Agents

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of traditional reinforcement learning, which assumes constant action durations and thus struggles with machine learning engineering agents that perform variable-length, costly actions. Grounded in Semi-Markov Decision Processes (SMDPs), this work introduces continuous-time reinforcement learning into agent training for the first time by proposing a reward rate policy gradient algorithm. Rather than maximizing cumulative rewards, the method directly optimizes long-term rewards per unit time, thereby circumventing the challenge of enumerating the policy space. It further achieves efficient training by integrating off-policy sampling estimation with a self-improvement loop driven by a small language model (Qwen3.5-4B). Experimental results demonstrate that, under fixed time budgets, the proposed approach improves reward performance over baselines by 19.2% on MLE-Bench and 85.7% on NanoGPT.
📝 Abstract
Traditional reinforcement learning (RL) techniques focus on maximizing expected cumulative reward, where each action assumes to take a constant unit of time. However, this assumption does not hold for agentic RL tasks such as machine learning engineering (MLE) agents, where actions involve data loading, feature engineering, and model training that take variable durations. Efficiency matters in modern agentic RL where actions are costly. To address this limitation, we adapt from continuous-time RL and Semi-Markov Decision Process (SMDP) formulation and propose Reward-rate Policy Gradient (RPG), where we focus on optimizing the reward rate -- the long-term reward per unit of time. RPG estimates the reward rate from off-policy samples, then charges each action for the time it consumes at that rate. We first conduct theoretical analysis in the bandit setting to establish that RPG approximates the optimal reward rate and empirically demonstrate it outperforms baselines while avoiding enumeration over the policy space, a known issue for an existing method. We then further apply RPG on a small language model (Qwen3.5-4B) with self-improvement loops and empirically show it obtains higher rewards within a fixed time budget than vanilla RL on MLE-Bench and NanoGPT, with a 19.2% and 85.7% margin, respectively. Our method provides a practical solution for optimizing performance under wait time considerations in modern agentic RL tasks, where actions interact with external environments and cost time.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Machine Learning Engineering Agents
Reward Rate Optimization
Variable-duration Actions
Semi-Markov Decision Process
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reward-rate Policy Gradient
Semi-Markov Decision Process
Continuous-time Reinforcement Learning
Machine Learning Engineering Agents
Off-policy Estimation