Senior Accelerator Engineer, Cloud AI/ML server team

Amazon
Cupertino, CA, USA / Seattle, WA, USA / Denver, CO, USA2026-08-27ONSITE

About the job

AWS Hardware Engineering is looking for a Senior Hardware Development Engineer to own the technical roadmap across all AI accelerator platforms in the organization. You will be the single-threaded owner of GPU lifecycle from defining hardware, firmware and diagnostics requirements, qualification through field operations and translating fleet-scale failure data into vendor action.

Responsibilities

- Own the technical relationship with vendors across all platforms in the portfolio: roadmap alignment, escalations, partnerships.

- Own qualification of new GPU SKUs and baseboard assemblies during NPI bring-up -- define test plans, acceptance criteria, and production readiness gates

- Define and maintain GPU firmware qualification criteria across the org -- pass/fail gates, staged rollout policy, regression detection methodology

- Drive RMA strategy: build failure evidence packages, negotiate acceptance criteria with vendors, manage submission quotas and pipeline velocity

- Lead root-cause analysis on fleet-wide GPU failure modes (component errors, PCIE interface errors, thermal events, link degradation, manufacturing escapes) using telemetry, event log data, and vendor diagnostics

- Set GPU health standards: define the metrics, thresholds, and alerting that platform teams execute against

- Define GPU fleet health dashboards: identify relevant telemetry, failure rate trends, replacement pipeline status, firmware version distribution, qualification status

- Represent the organization in technical discussions with leading vendor’s engineering -- translate fleet-scale patterns into prioritized vendor action items

- Partner with server teams to ensure consistent GPU operational practices; provide expertise without owning their execution

- Present GPU fleet health, replacement pipeline status, and qualification progress to senior leadership (VP-level) regularly

- Mentor engineers on GPU failure analysis methodology

Qualifications

Minimum

- Bachelor's degree in electrical engineering, computer engineering, or equivalent

- Experience in developing functional specifications, design verification plans and functional test procedures

- 7+ years of hardware design and development experience for server, compute, or large-scale infrastructure platforms

Preferred

- 7+ years hardware engineering experience

- Direct experience with data center GPUs and associated tooling

- Experience developing or influencing server roadmap

- Track record of influencing vendor engineering priorities through failure evidence

- Familiarity with GPU thermal management, power delivery, and PCIe and interconnect architectures

- Experience with firmware lifecycle management at scale (qualification, staged rollout, regression detection, rollback)

- Comfortable presenting to VP-level audiences

- Experience working horizontally across multiple platform teams without direct authority

- Strong data analysis skills at fleet scale (statistical failure modeling, trend detection, threshold setting)