About the job
Join the Systems Pathfinding and Architecture (SPARC) team within Microsoft’s Azure Hardware Systems and Infrastructure (AHSI) organization. In this role, you will own and define the system architecture requirements, making a direct impact on Microsoft AI accelerator systems.
Responsibilities
Define AI rack- and cluster-scale architecture across accelerator, host, network, memory, storage, power, cooling, and datacenter constraints. Make quantitative tradeoffs across performance and TCO.\u000aDevelop functional proof-of-concepts to evaluate and de-risk new capabilities.\u000aWrite and own high-level product requirements and specifications.\u000aDrive alignment across cross-functional engineering and leadership teams. Own the outcomes.\u000aUse performance simulations and hardware models for what-if studies to evaluate and choose between architecture options.\u000aParticipate in industry consortiums to shape standards, and influence vendor roadmaps.
Qualifications
Minimum
Master's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 9+ years technical engineering experience OR Bachelor's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 11+ years technical engineering experience OR equivalent experience
Preferred
Doctorate in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 8+ years technical engineering experience OR Master's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 12+ years technical engineering experience OR Bachelor's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 15+ years technical engineering experience OR equivalent experience.\u000aExperience in large scale AI datacenter systems and their network.\u000aDemonstrated expertise in one or more of these domains: networking (RoCE, Infiniband transports, congestion control, network topology design), reliability and serviceability, datacenter architecture for large scale AI systems.\u000aExperience owning hardware system architecture at scale – machine, rack row and cluster, authoring specifications and requirements.\u000aDeep expertise in AI systems or a key domain such as networking, reliability, power management.\u000aSkilled in partnering and influencing architects, hardware engineers, and software leads.\u000aAbility to manage through ambiguity, bringing clarity to establish direction and driving decisions.\u000aCollaboration skills, teamwork, and sense of presumed responsibility.