About the job
Manages team delivering scalable distributed systems and components on a 2–4 quarter horizon. Standardizes engineering practices and scalability requirements across teams; oversees optimization for high‑throughput, hyper‑scale workloads; and ensures effective use of distributed state tools and data plane platforms. Guides teams to design fault‑tolerant, in‑service‑upgradable systems, set SLO‑aligned durability/availability targets, and implement resiliency mechanisms (load‑shedding, throttling, rate‑limiting).
Responsibilities
Manages team delivering scalable distributed systems and components on a 2–4 quarter horizon
Own and build solutions to scale and optimize AI compute infrastructure components like GPU control plane and GPU data plane
Standardizes engineering practices and scalability requirements across teams
Oversees optimization for high‑throughput, hyper‑scale workloads
Guides teams to design fault‑tolerant, in‑service‑upgradable systems and implement resiliency mechanisms
Provides oversight for KPIs, telemetry, and moderately complex dashboards
Qualifications
Minimum
10+ years' experience in software development with programming languages including, but not limited to, C, C++, C#, Java, Go, Rust
5+ years' experience in people management or leadership role while working on cross-functional projects
5+ years' experience designing and developing large-scale distributed systems, services, and infrastructure
BS (or equivalent experience) in Computer Science, Engineering, or related field
Strong communication, collaboration, and project management skills
Ability to adapt to a fast-paced, dynamic environment and manage multiple tasks and priorities effectively
Preferred
Experience managing cloud infrastructure with hundreds of thousands of servers
Experience with containerization technologies such as Docker and Kubernetes
Experience scheduling high-performance workloads on Kubernetes or Slurm