Network Operations Engineer, AI Networking

OpenAI
San Francisco2026-07-07

About the job

We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network.

Responsibilities

- Own the operational health, availability, and reliability of production AI network infrastructure across 1P and 3P data centers.

- Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR).

- Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks.

- Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact.

- Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance.

- Support new AI cluster deployments, data center expansions, and infrastructure migrations in partnership with deployment and engineering teams.

- Partner with cloud service providers (CSPs), colocation providers, Smart Hands teams, and hardware vendors to maintain production infrastructure.

- Perform root-cause analysis (RCA) for production incidents and drive permanent corrective actions that eliminate recurring issues.

- Build and maintain monitoring, telemetry, dashboards, and alerting to improve network observability and proactive issue detection.

- Develop and improve operational runbooks, playbooks, troubleshooting documentation, and standard operating procedures.

- Automate repetitive operational tasks using Python and infrastructure automation frameworks to reduce toil and improve efficiency.

- Continuously identify opportunities to improve service reliability, scalability, operational maturity, and engineering efficiency.

Qualifications

Minimum

- Bachelor’s degree in Computer Science, Network Engineering, or a related discipline, or equivalent practical experience.

- 5+ years of experience operating large-scale data center, cloud, AI, or HPC network infrastructure.

- Experience supporting production network environments with high-availability requirements.

- Hands-on experience with one or more of the following platforms: Cisco NX-OS, Arista EOS, NVIDIA Spectrum / Cumulus Linux, or Juniper JunOS.

- Strong knowledge of Layer 2 and Layer 3 networking, BGP, OSPF, ECMP, MLAG, LACP, VRFs, and VLANs.

- Experience troubleshooting physical infrastructure, including fiber optics, transceivers, DAC/AOC cables, and high-speed Ethernet links.

- Experience performing software upgrades, hardware maintenance, and production change management.

- Excellent analytical and troubleshooting skills, with the ability to communicate technical risk clearly across teams.

Preferred

- Experience operating AI or High-Performance Computing (HPC) network environments.

- Experience with NVIDIA AI networking technologies and GPU infrastructure.

- Experience supporting RoCE v2 or RDMA-based Ethernet fabrics, with a strong understanding of Priority Flow Control (PFC), Explicit Congestion Notification (ECN), Data Center Quantized Congestion Notification (DCQCN), Quality of Service (QoS), and lossless Ethernet networking.

- Experience supporting 100G, 200G, 400G, and 800G Ethernet networks.

- Experience with GPU platforms including NVIDIA HGX, DGX, GB200, or equivalent AI infrastructure.

- Experience supporting distributed storage environments such as VAST, DDN, or similar technologies.

- Experience working with cloud service providers such as AWS, Azure, or Google Cloud, and with third-party colocation providers.

- Experience with network monitoring and telemetry technologies, including Prometheus, Grafana, gNMI, streaming telemetry, SNMP, or similar tools.

- Experience developing automation using Python, Git, REST APIs, Terraform, or similar automation frameworks.