Senior MLOps Engineer - DSX Enablement

Nvidia
US, CA, Santa Clara / US, WA, Seattle / Remote - US2026-08-20remote_local

About the job

NVIDIA is seeking a Senior MLOps Engineer to join our DSX Enablement team, collaborating closely with strategic customers to implement and enhance groundbreaking AI workloads. We partner with the world's most innovative AI companies and open-source communities to address their most challenging technical problems.

Responsibilities

Build and deploy custom AI solutions on NeoCloud platforms and NVIDIA Cloud Partners (NCPs), including distributed training, inference optimization, and MLOps pipelines

Act as a primary technical contact for internal and external customers and partners, guiding joint engagements, ensuring the success of initiatives on DGX Cloud, and solving complex problems in production

Work closely with the teams building the infrastructure software and accelerated frameworks that support today’s most compelling AI applications

Profile and tune large-scale training and inference workloads on NCP platforms, leading efforts to reduce latency, cost, and operational risk

Develop open-source tools and reference architectures to make it easier to build and manage machine learning and AI workloads, pipelines, and systems at scale

Qualifications

Minimum

BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field, or equivalent experience

8+ years of experience in technical roles such as data science, data engineering, or ML engineering, ideally targeting large‑scale production systems

Demonstrated AI/ML experience across multiple phases of the machine learning lifecycle, from exploratory analysis to production systems

Facility with systems topics including Linux, batch schedulers, Kubernetes, distributed filesystems, and advanced networking at datacenter scale

Solid scripting and programming skills in languages like bash and Python and solid systems programming skills in a language like C++, Go, or Rust

Experience using machine learning or deep learning frameworks for training and inference

Excellent communication and technical presentation skills, with the ability to clearly articulate architectures, trade‑offs, and recommendations to both engineering and leadership audiences

A clear record of engineering discipline and execution on interesting projects, whether you’re working alone or collaborating on a team

Preferred

Experience contributing to and working in open-source communities

Experience with the NVIDIA ecosystem, including DGX systems, CUDA, NeMo, RAPIDS, Triton, NIM, and NVIDIA networking technologies such as InfiniBand, NVLink, and RoCE

Experience and familiarity building machine learning systems in a security-critical environment and distributed training and inference frameworks

Familiarity with MLOps practices in a cloud‑native context: containerization, CI/CD pipelines, workflow automation, observability stacks, and GitOps workflows

Direct experience drawing on deep systems knowledge to diagnose and fix performance or correctness problems that span multiple layers of the application stack, like hardware, networking, accelerator, hypervisor or OS, compilers or runtimes, application code, and libraries