AI Hardware Systems Manager, Annapurna Labs, Trainium Machine Learning Fleet Operations

Amazon
USA, TX, Austin2026-06-15ONSITE

About the job

In Annapurna Labs we are at the forefront of hardware/software co-design not just in Amazon Web Services (AWS) but across the industry. The Machine Learning Acceleration Fleet Operations Team is looking for a technical leader to manage a team of 5-10 engineers and own operations across multiple ML server platforms spanning tens of thousands of hosts globally.

Responsibilities

- Build, hire, mentor, and grow a team of platform development engineers responsible for ML fleet operations across multiple accelerator platforms

- Define team roadmap and technical strategy for fleet health, automation, and data infrastructure — balancing near-term operational demands against long-term engineering investments

- Drive operational excellence by establishing metrics, SLAs, and processes that maximize platform sellability and customer experience

- Partner with hardware engineering, software engineering, and product teams to prioritize debug efforts and translate fleet learnings into permanent design fixes

- Own escalation paths for critical fleet incidents and lead cross-functional war rooms to resolution

- Influence org-level priorities by surfacing fleet-wide patterns and advocating for systemic improvements across the ML hardware portfolio

- Raise the bar on team software practices — ensuring automation is maintainable, tested, documented, and reusable at scale

- Represent fleet operations in executive reviews, providing data-driven narratives on platform health and roadmap

Qualifications

Minimum

- Bachelor's degree in computer science, electrical engineering, or related field

- 2+ years of engineering team management experience

- Knowledge of and proficiency in the use of Python scripting language

- Experience with general troubleshooting/debugging of hardware

- Experience designing, building, operating, and managing large-scale distributed systems or web services

- 7+ years of experience in systems engineering, platform engineering, SRE, or hardware operations

Preferred

- Experience in automating, deploying, and supporting large-scale infrastructure

- Experience in server technologies such as, thermal, mechanical, power, and signal integrity

- Experience working cross-functionally across several teams both technical and non-technical

- Experience with GPU, ML accelerator, or high-performance computing hardware

- Experience managing teams through ambiguity on new or unreleased products