Senior Hardware Development Engineer, Cloud AI/ML Server Team

Amazon
Cupertino, CA, USA / Denver, CO, USA / Seattle, WA, USA2026-08-27ONSITE

About the job

AWS operates the world's largest fleet of GPU-accelerated servers powering AI/ML training and inference at cloud scale. Our team defines the server architectures, drives the hardware designs, and owns the fleet quality for these platforms — from component selection through datacenter operations. If you want to shape the physical hardware that frontier models train on, this is the role.

We are seeking a Cloud Hardware Development Engineer to define server architectures based on workload demand, translate them into detailed component specifications, and drive validation from PCBA bring-up through rack integration. You will lead ODM design partners through development and production, triage hardware issues across manufacturing and datacenters, and own fleet quality metrics post-launch.

Responsibilities

- Define server architectures based on workload demand and customer requirements, translating them into detailed designs and component specifications that enable high-performance AI training and inference at scale

- Work with interdisciplinary teams of component, firmware, test, qualification, and integration engineers to deliver cohesive designs

- Drive design reviews with ODM/JDM partners covering schematic, layout, BOM, and manufacturing DFx (Design for Test, Design for Manufacturing)

- Define and execute validation strategies from PCBA bring-up through server and rack integration — covering power sequencing, signal integrity, thermal characterization, and accelerator interconnect performance

- Own hardware debug during EVT/DVT/PVT builds, correlating failures across PCIe, power rails, memory channels, and GPU subsystems

- Triage hardware issues at both ODM facilities and datacenters, conduct root cause analysis, and implement corrective actions

- Own fleet quality metrics post-launch: server-level annualized failure rates and component-level failure modes

- Monitor operational telemetry to identify systemic issues and drive design or process changes for current and future platforms

- Partner with test and automation teams to improve manufacturing yield and reduce test dwell times

- Work with EC2 architecture teams to align on instance definitions, workload requirements, and platform trade-offs

- Drive ODM/JDM design partners through development milestones and production ramp

- Collaborate with firmware, software, and operations teams to ensure designs are debuggable, serviceable, and automation-ready

Qualifications

Minimum

- Bachelor's degree in electrical engineering, computer engineering, or equivalent

- Experience in developing functional specifications, design verification plans and functional test procedures

- 7+ years of hardware design and development experience for server, compute, or large-scale infrastructure platforms

- Experience in one or more server technologies: thermal/mechanical design, power delivery, high-speed signal integrity, or accelerator subsystems

- Experience leading hardware development through full product lifecycle (concept through production ramp)

Preferred

- Master's degree or PhD in Electrical Engineering, Computer Engineering, or a related field

- 5+ years of experience working with ODMs through the product development and manufacturing lifecycle (EVT, DVT, PVT)

- In-depth expertise in high-speed bus design, signal integrity analysis, or power delivery for GPU/accelerator platforms

- 5+ years of experience with hardware bring-up, debug, and root cause analysis across PCIe, NVMe, memory, and accelerator interconnects

- Experience owning fleet quality metrics and driving design improvements based on operational failure data

- Experience with thermal/mechanical design for high-power-density compute platforms (liquid cooling, air cooling, or hybrid)