About the job
Google Kubernetes Engine (GKE) is the industry standard for container orchestration and the core of Google Cloud’s modernization strategy. We are now embarking on a mission to reinvent GKE and Kubernetes as the premier substrate for the next generation of computing: AI inference at massive scale. As the Principal Engineer, you will lead the technical and architectural reinvention of GKE to become the inference engine for the world.
Responsibilities
Lead the architectural direction for llm-d, ensuring a highly optimized, scalable foundation for distributed LLM and Reinforcement Learning (RL) serving across the GKE fleet.\nDefine GKE's evolution to support massive-scale inference and RL, solving novel orchestration problems in dynamic resource allocation, multi-host TPU/GPU scheduling, and high-throughput networking.\nPartner with strategic AI model builders, DeepMind, and Vertex AI to co-develop an AI-first roadmap, leveraging Google's custom silicon to optimize throughput and compute density.\nLead the broader Kubernetes ecosystem and Open Source Software (OSS) community, driving key upstream initiatives to establish industry standards for AI, RL, and accelerator orchestration.
Qualifications
Minimum
Bachelor's degree in Computer Science, a related technical field, or equivalent practical experience.\n15 years of experience in software engineering, or 15 years of experience with an advanced degree.\nExperience building distributed systems and driving technical strategy for platform-level infrastructure.\nExperience with Kubernetes, container runtimes, and AI/ML infrastructure (e.g., inference serving, LLM, hardware accelerators).
Preferred
Master's degree or PhD in Computer Science or related technical field.\nExperience interacting with senior customer stakeholders (CTOs, chief architects) to represent the technical vision of the organization.\nDemonstrated track record of significant technical contributions to the Kubernetes open-source project or related CNCF AI/ML projects (e.g., Kueue).\nDemonstrated track record of influencing cross-functional teams (product, engineering, research) to deliver complex technical outcomes.\nDeep technical understanding of high-performance networking (RDMA, NCCL), storage/caching architectures for massive model weights, and accelerator virtualization/sharing mechanisms.