About the job
We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network.
Responsibilities
- Own the operational health, availability, and reliability of production AI network infrastructure across 1P and 3P data centers.
- Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR).
- Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks.
- Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact.
- Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance.
- Support new AI cluster deployments, data center expansions, and infrastructure migrations in partnership with deployment and engineering teams.
- Partner with cloud service providers (CSPs), colocation providers, Smart Hands teams, and hardware vendors to maintain production infrastructure.
- Perform root-cause analysis (RCA) for production incidents and drive permanent corrective actions that eliminate recurring issues.
- Build and maintain monitoring, telemetry, dashboards, and alerting to improve network observability and proactive issue detection.
- Develop and improve operational runbooks, playbooks, troubleshooting documentation, and standard operating procedures.
- Automate repetitive operational tasks using Python and infrastructure automation frameworks to reduce toil and improve efficiency.
- Continuously identify opportunities to improve service reliability, scalability, operational maturity, and engineering efficiency.
Qualifications
Minimum
- Bachelor’s degree in Computer Science, Network Engineering, or a related discipline, or equivalent practical experience.
- 5+ years of experience operating large-scale data center, cloud, AI, or HPC network infrastructure.
- Experience supporting production network environments with high-availability requirements.
- Hands-on experience with one or more of the following platforms: Cisco NX-OS, Arista EOS, NVIDIA Spectrum / Cumulus Linux, or Juniper JunOS.
- Strong knowledge of Layer 2 and Layer 3 networking, BGP, OSPF, ECMP, MLAG, LACP, VRFs, and VLANs.
- Experience troubleshooting physical infrastructure, including fiber optics, transceivers, DAC/AOC cables, and high-speed Ethernet links.
- Experience performing software upgrades, hardware maintenance, and production change management.
- Excellent analytical and troubleshooting skills, with the ability to communicate technical risk clearly across teams.
Preferred
- Experience operating AI or High-Performance Computing (HPC) network environments.
- Experience with NVIDIA AI networking technologies and GPU infrastructure.
- Experience supporting RoCE v2 or RDMA-based Ethernet fabrics, with a strong understanding of Priority Flow Control (PFC), Explicit Congestion Notification (ECN), Data Center Quantized Congestion Notification (DCQCN), Quality of Service (QoS), and lossless Ethernet networking.
- Experience supporting 100G, 200G, 400G, and 800G Ethernet networks.
- Experience with GPU platforms including NVIDIA HGX, DGX, GB200, or equivalent AI infrastructure.
- Experience supporting distributed storage environments such as VAST, DDN, or similar technologies.
- Experience working with cloud service providers such as AWS, Azure, or Google Cloud, and with third-party colocation providers.
- Experience with network monitoring and telemetry technologies, including Prometheus, Grafana, gNMI, streaming telemetry, SNMP, or similar tools.
- Experience developing automation using Python, Git, REST APIs, Terraform, or similar automation frameworks.