About the job
In Annapurna Labs we are at the forefront of hardware/software co-design not just in Amazon Web Services (AWS) but across the industry. The Machine Learning Acceleration Fleet Operations Team is looking for a technical leader to manage a team of 5-10 engineers and own operations across multiple ML server platforms spanning tens of thousands of hosts globally.
Responsibilities
- Build, hire, mentor, and grow a team of platform development engineers responsible for ML fleet operations across multiple accelerator platforms
- Define team roadmap and technical strategy for fleet health, automation, and data infrastructure — balancing near-term operational demands against long-term engineering investments
- Drive operational excellence by establishing metrics, SLAs, and processes that maximize platform sellability and customer experience
- Partner with hardware engineering, software engineering, and product teams to prioritize debug efforts and translate fleet learnings into permanent design fixes
- Own escalation paths for critical fleet incidents and lead cross-functional war rooms to resolution
- Influence org-level priorities by surfacing fleet-wide patterns and advocating for systemic improvements across the ML hardware portfolio
- Raise the bar on team software practices — ensuring automation is maintainable, tested, documented, and reusable at scale
- Represent fleet operations in executive reviews, providing data-driven narratives on platform health and roadmap
Qualifications
Minimum
- Bachelor's degree in computer science, electrical engineering, or related field
- 2+ years of engineering team management experience
- Knowledge of and proficiency in the use of Python scripting language
- Experience with general troubleshooting/debugging of hardware
- Experience designing, building, operating, and managing large-scale distributed systems or web services
- 7+ years of experience in systems engineering, platform engineering, SRE, or hardware operations
Preferred
- Experience in automating, deploying, and supporting large-scale infrastructure
- Experience in server technologies such as, thermal, mechanical, power, and signal integrity
- Experience working cross-functionally across several teams both technical and non-technical
- Experience with GPU, ML accelerator, or high-performance computing hardware
- Experience managing teams through ambiguity on new or unreleased products