About the job
The Azure HPC & AI Infrastructure team defines and delivers the systems that power large-scale AI training and inference. We are seeking a Senior Product Manager to own the product strategy and execution for back-end networking across GPU infrastructure, spanning across clusters and datacenters to power the next generation of AI Factories.
In this role, you will translate workload and customer needs into product requirements, align architecture and investment choices, and drive products from concept through qualification, launch, and fleet deployment. You will work across networking, systems engineering, software, datacenter infrastructure, operations, supply chain, finance, and strategic suppliers to improve performance, reliability, time to market, and total cost of ownership.
Responsibilities
Define the product strategy, and multi-year roadmap for GPU back-end networking capabilities.
Engage strategic customers and internal workload teams to understand requirements, manage competing product priorities to deliver infrastructure for AI workloads at scale.
Translate customer AI training and inference workload needs into clear product requirements for bandwidth, latency, topology, congestion management, resiliency, security, observability, and serviceability.
Own the end-to-end product lifecycle from concept and plan of record through new product introduction, qualification, launch, scale deployment, and lifecycle improvement.
Drive alignment and execution across hardware and software engineering, networking, architecture, datacenter infrastructure, operations, supply chain, finance, and customer-facing teams.
Influence supplier roadmaps and product definitions across GPU systems, network interface devices, switches, optics, cables, and interconnect technologies; resolve technical and business tradeoffs to improve platform outcomes.
Establish success metrics and operating mechanisms for performance, reliability, availability, deployment readiness, cost efficiency, and launch health; use customer signals, telemetry, and incident learnings to prioritize improvements.
Qualifications
Minimum
Bachelor's Degree AND 5+ years experience in product/service/program management or software development OR equivalent experience.
1+ years experience with GPU or AI-accelerated hardware systems and infrastructure, including the ability to translate complex system concepts into product requirements, roadmaps, and execution plans.
1+ years experience with datacenter networking or accelerated computing infrastructure, including technologies such as Ethernet, InfiniBand, RDMA, network interface devices, switches, optics, NVLink etc.
Ability to meet Microsoft, customer and/or government security screening requirements are required for this role.
Preferred
Bachelor's Degree AND 8+ years experience in product/service/program management or software development OR equivalent experience.
2+ years experience taking a product, feature, or experience to market (e.g., design, addressing product market fit, and launch, internal tool/framework).
4+ years experience improving product metrics for a product, feature, or experience in a market (e.g., growing customer base, expanding customer usage, avoiding customer churn).
4+ years experience disrupting a market for a product, feature, or experience (e.g., competitive disruption, taking the place of an established competing product).