About the job
AWS operates the world's largest fleet of GPU-accelerated servers powering AI/ML training and inference at cloud scale. Our team defines the server architectures, drives the hardware designs, and owns the fleet quality for these platforms — from component selection through datacenter operations. If you want to shape the physical hardware that frontier models train on, this is the role.
We are seeking a Cloud Hardware Development Engineer to define server architectures based on workload demand, translate them into detailed component specifications, and drive validation from PCBA bring-up through rack integration. You will lead ODM design partners through development and production, triage hardware issues across manufacturing and datacenters, and own fleet quality metrics post-launch.
Responsibilities
- Define server architectures based on workload demand and customer requirements, translating them into detailed designs and component specifications that enable high-performance AI training and inference at scale
- Work with interdisciplinary teams of component, firmware, test, qualification, and integration engineers to deliver cohesive designs
- Drive design reviews with ODM/JDM partners covering schematic, layout, BOM, and manufacturing DFx (Design for Test, Design for Manufacturing)
- Define and execute validation strategies from PCBA bring-up through server and rack integration — covering power sequencing, signal integrity, thermal characterization, and accelerator interconnect performance
- Own hardware debug during EVT/DVT/PVT builds, correlating failures across PCIe, power rails, memory channels, and GPU subsystems
- Triage hardware issues at both ODM facilities and datacenters, conduct root cause analysis, and implement corrective actions
- Own fleet quality metrics post-launch: server-level annualized failure rates and component-level failure modes
- Monitor operational telemetry to identify systemic issues and drive design or process changes for current and future platforms
- Partner with test and automation teams to improve manufacturing yield and reduce test dwell times
- Work with EC2 architecture teams to align on instance definitions, workload requirements, and platform trade-offs
- Drive ODM/JDM design partners through development milestones and production ramp
- Collaborate with firmware, software, and operations teams to ensure designs are debuggable, serviceable, and automation-ready
Qualifications
Minimum
- Bachelor's degree in electrical engineering, computer engineering, or equivalent
- Experience in developing functional specifications, design verification plans and functional test procedures
- 7+ years of hardware design and development experience for server, compute, or large-scale infrastructure platforms
- Experience in one or more server technologies: thermal/mechanical design, power delivery, high-speed signal integrity, or accelerator subsystems
- Experience leading hardware development through full product lifecycle (concept through production ramp)
Preferred
- Master's degree or PhD in Electrical Engineering, Computer Engineering, or a related field
- 5+ years of experience working with ODMs through the product development and manufacturing lifecycle (EVT, DVT, PVT)
- In-depth expertise in high-speed bus design, signal integrity analysis, or power delivery for GPU/accelerator platforms
- 5+ years of experience with hardware bring-up, debug, and root cause analysis across PCIe, NVMe, memory, and accelerator interconnects
- Experience owning fleet quality metrics and driving design improvements based on operational failure data
- Experience with thermal/mechanical design for high-power-density compute platforms (liquid cooling, air cooling, or hybrid)