Technical Program Manager, RL Scaling, DeepMind

Google DeepMind
Mountain View, CA, USA

About the job

The Gemini Reinforcement learning (RL) Scaling team is at the frontier of reinforcement learning research for large language models, driving the reasoning, multimodal, and agentic capabilities that define the next generation of Gemini models. As a Technical Program Manager, you will independently drive program execution for critical RL research workstreams. You will sit at the intersection of empirical research and large-scale distributed systems, partnering directly with research scientists and research engineers to operationalize scaling experiments, manage RL training pipelines including SFT initialization and data workflows optimize compute utilization, and accelerate the progress of research breakthroughs into frontier Gemini releases.

Responsibilities

Scope, plan, and lead execution for RL research workstreams in collaboration with tech leads, turning research hypotheses into structured roadmaps, experiment plans, and deliverable model milestones.

Partner with engineering and infrastructure leads to manage and track experiments and compute allocations, enabling prioritization, monitoring training efficiency, and unblocking runs.

Identify and resolve cross-functional dependencies across the RL research ecosystem (data pipelines, evaluations, distributed infrastructure) and partner teams.

Establish reliable operational rhythms including experiment status dashboards, launch criteria, retrospectives, and milestone reviews, synthesizing complex training dynamics into actionable updates for tech leads and leadership.

Qualifications

Minimum

Bachelor's degree in Computer Science, a related technical field or equivalent practical experience.

5 years of experience in technical program management.

Preferred

Master's degree or PhD in Computer Science or a closely related technical field.

Over 5 years leading cross-functional AI model programs, with expertise in large-scale distributed training pipelines and reinforcement learning.

Proven ability to design lightweight, high-impact processes that structure fast-moving research environments without hindering team velocity.

Highly comfortable with ambiguity; a strong communicator who builds trust and drives alignment across engineering, research, and leadership stakeholders.