About the job
AWS Hardware Engineering is looking for a Senior Hardware Development Engineer to own the technical roadmap across all AI accelerator platforms in the organization. You will be the single-threaded owner of GPU lifecycle from defining hardware, firmware and diagnostics requirements, qualification through field operations and translating fleet-scale failure data into vendor action.
Responsibilities
- Own the technical relationship with vendors across all platforms in the portfolio: roadmap alignment, escalations, partnerships.
- Own qualification of new GPU SKUs and baseboard assemblies during NPI bring-up -- define test plans, acceptance criteria, and production readiness gates
- Define and maintain GPU firmware qualification criteria across the org -- pass/fail gates, staged rollout policy, regression detection methodology
- Drive RMA strategy: build failure evidence packages, negotiate acceptance criteria with vendors, manage submission quotas and pipeline velocity
- Lead root-cause analysis on fleet-wide GPU failure modes (component errors, PCIE interface errors, thermal events, link degradation, manufacturing escapes) using telemetry, event log data, and vendor diagnostics
- Set GPU health standards: define the metrics, thresholds, and alerting that platform teams execute against
- Define GPU fleet health dashboards: identify relevant telemetry, failure rate trends, replacement pipeline status, firmware version distribution, qualification status
- Represent the organization in technical discussions with leading vendor’s engineering -- translate fleet-scale patterns into prioritized vendor action items
- Partner with server teams to ensure consistent GPU operational practices; provide expertise without owning their execution
- Present GPU fleet health, replacement pipeline status, and qualification progress to senior leadership (VP-level) regularly
- Mentor engineers on GPU failure analysis methodology
Qualifications
Minimum
- Bachelor's degree in electrical engineering, computer engineering, or equivalent
- Experience in developing functional specifications, design verification plans and functional test procedures
- 7+ years of hardware design and development experience for server, compute, or large-scale infrastructure platforms
Preferred
- 7+ years hardware engineering experience
- Direct experience with data center GPUs and associated tooling
- Experience developing or influencing server roadmap
- Track record of influencing vendor engineering priorities through failure evidence
- Familiarity with GPU thermal management, power delivery, and PCIe and interconnect architectures
- Experience with firmware lifecycle management at scale (qualification, staged rollout, regression detection, rollback)
- Comfortable presenting to VP-level audiences
- Experience working horizontally across multiple platform teams without direct authority
- Strong data analysis skills at fleet scale (statistical failure modeling, trend detection, threshold setting)