Sr Technical Program Manager, AWS Generative AI & ML Servers

Amazon
Cupertino, CA, USA / Seattle, WA, USA2026-06-15ONSITE

About the job

AWS Infrastructure Services is seeking a Senior Technical Program Manager to drive end-to-end delivery of GPU-accelerated servers powering AI/ML workloads across our global fleet. You will facilitate requirements gathering with internal customers, coordinate cross-functional engineering teams spanning hardware and software disciplines, manage ODM partnerships across multiple continents, and establish closed-loop quality feedback systems connecting operational data to design improvements.

Responsibilities

Facilitate requirements gathering sessions with internal customers and stakeholders. Work with engineering teams to validate feasibility, identify dependencies, and establish clear success criteria. Build program timelines with milestone tracking, critical path analysis, and proactive risk assessment. Translate business objectives into program deliverables that engineering teams can execute against.

Drive cross-functional alignment across hardware, firmware, software, and operations teams to maintain development schedules. Manage ODM partnerships by tracking design reviews, manufacturing readiness gates, and quality checkpoints. Identify blockers early and escalate dependencies before they impact critical path. Facilitate decision-making when teams are stuck on technical trade-offs. Communicate program status to stakeholders with clear visibility into progress against milestones.

Identify program risks by challenging technical workstreams and connecting technical constraints to schedule impact. Escalate risks proactively with quantified business impact and proposed mitigation plans. Lead root cause analysis of fleet-wide hardware failures, driving engineering teams to implement corrective actions that improve operational metrics. Establish feedback loops connecting operational telemetry back to validation criteria for future designs.

Define acceptance criteria and coordinate qualification testing across teams. Manage deployment readiness by tracking open issues, validating documentation completeness, and ensuring operations teams are prepared for production handoff. Facilitate go/no-go decisions with clear risk posture and mitigation status for mission-critical AI/ML workloads.

Conduct knowledge transfer sessions with receiving teams. Document lessons learned with actionable recommendations that improve future program execution. Ensure all program artifacts are complete and accessible.

Qualifications

Minimum

5+ years of technical product or program management experience

7+ years of working directly with engineering teams experience

Experience managing programs across cross functional teams, building processes and coordinating release schedules

Experience leading engineering design projects and interacting with cross-functional teams

Preferred

8+ years of project management disciplines including scope, schedule, budget, quality, along with risk and critical path management experience

Knowledge of the electrical and mechanical systems involved in critical data center operations including systems such as feeders, transformers, generators, switchgear, UPS systems, ATS units, PDU units, chillers, pumps, air handling units, and CRAC units

Experience in machine learning, data mining, information retrieval, statistics or natural language processing, or experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware

Experience developing and driving projects with tight schedules and multiple dependencies

Experience using data and metrics to drive improvements

Experience coordinating programs with ODM partners across multiple geographies, managing design reviews, manufacturing readiness gates, and quality checkpoints

Ability communicating program status to executive stakeholders by connecting technical constraints to business impact