About the job
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. The Predictive Planning team (PrePlan) develops and deploys state-of-the-art machine learning solutions that predict the future state of the world and plan the Waymo Driver’s behavior. We have an exciting opportunity for a Staff Technical Lead Manager to lead our ML Evaluation team. In this role, you will define the strategic vision for our evaluation platforms, scaling the critical infrastructure and metrics required, and partner closely with the modeling teams to rigorously validate our next-generation deep neural networks and accelerate ML developer velocity across PrePlan.
Responsibilities
Influence the strategic direction of foundational infrastructure and evaluation platforms to robustly support next-generation ML model evaluation use cases
Collaborate cross-functionally with ML engineers, data scientists, and infrastructure teams to identify, define, and surface critical signals on model, component, and system-level performance
Leverage and scale evaluation and infrastructure platforms to significantly enhance the ML developer experience, enabling faster iteration through earlier, more reliable, and trusted model evaluation
Manage and mentor a focused team of engineers, aligning their career growth and aspirations with critical organizational needs
Drive best practices and leverage deep technical awareness of the Alphabet ML stack (e.g., TensorFlow, JAX, Flax, Apache Beam) to optimize evaluation workflows
Stay at the forefront of emerging technologies, industry trends, and research in ML evaluation methodologies and advanced metrics design
Qualifications
Minimum
M.S. in Computer Science, Mathematics, or equivalent industry experience in Robotics or large-scale ML systems with critical evaluation needs
5+ years of experience building and maintaining large-scale distributed infrastructure, ML inference systems, or evaluation platforms, including 3+ years of engineering management experience
Strong coding and testing proficiency, specifically in Python and C++
Strong foundational knowledge of model evaluation and core data science principles (e.g., confidence intervals, outlier identification, curve fitting, and causality analysis)
Familiarity with large-scale ML deployment and orchestration tools (e.g., TF Serving, TorchServe, Kubeflow, SageMaker Pipelines, or Vertex AI Pipelines)
Understanding of machine learning fundamentals and experience with popular ML frameworks such as JAX, PyTorch, or TensorFlow
Preferred
Experience developing and maintaining evaluation pipelines for ML models
Experience deploying and supporting machine learning models for computer vision, natural language processing, robotics/motion planning, or recommendation systems
Experience supporting a small team of MLEs developing high-capacity, production-grade models and components
Strong understanding of metrics computation and regression detection at scale