About the job
NVIDIA is seeking an extraordinary innovator in networking and system architecture to join our Research team. In this role, you will focus on designing networks optimized for AI systems, as well as leveraging AI and machine learning to advance network architecture and enable intelligent, real-time decision-making for tasks such as control, routing, congestion management, and scheduling.
Responsibilities
Develop innovative network architectures, algorithms, and hardware/software co-design approaches for high-performance interconnects and large-scale distributed AI systems, including AI-assisted methods where appropriate, to enable efficient, scalable, and robust communication and extend the state of the art in networking, distributed computing, and system architecture.\nCreate and evaluate mechanisms for network and system decision-making in large-scale GPU/accelerator clusters, including AI/ML-based approaches for routing, traffic engineering, congestion control, scheduling, topology design, and telemetry-driven control loops."Invent new techniques, technologies, methodologies, processes, and devices, to enable new products or types of products. Deliverable results include prototypes, patents, publications, and product impact.\
Qualifications
Minimum
Pursuing or recently completed a PhD in relevant discipline(s) (CS, CE, EE, Physics, Math) or equivalent experience.",
Preferred
2+ years of relevant industrial and academic experience preferred. Relevant experience includes systems and network design for AI or HPC infrastructure and/or applying AI/ML to systems or networking problems (e.g., optimization, reinforcement learning, graph learning, telemetry-driven control).\Background and publication record in systems, networking, computer architecture, and/or ML for networking/systems. Publication at venues such as ISCA, HPCA, MICRO, SIGCOMM, NSDI (and related venues) is a plus.\Experience with AI/ML methods for systems and networks supporting AI workloads (e.g., PyTorch/TensorFlow/JAX) is valuable. It includes applying these methods to system or network building, simulation, optimization, and control/decision loops.\Strong programming and prototyping ability, with experience building research artifacts, simulators, or system prototypes; C++ and Python preferred, and experience with hardware description languages or HLS is desirable.