About the job
NVIDIA is seeking a Senior MLOps Engineer to join our DSX Enablement team, collaborating closely with strategic customers to implement and enhance groundbreaking AI workloads. We partner with the world's most innovative AI companies and open-source communities to address their most challenging technical problems.
Responsibilities
Build and deploy custom AI solutions on NeoCloud platforms and NVIDIA Cloud Partners (NCPs), including distributed training, inference optimization, and MLOps pipelines
Act as a primary technical contact for internal and external customers and partners, guiding joint engagements, ensuring the success of initiatives on DGX Cloud, and solving complex problems in production
Work closely with the teams building the infrastructure software and accelerated frameworks that support today’s most compelling AI applications
Profile and tune large-scale training and inference workloads on NCP platforms, leading efforts to reduce latency, cost, and operational risk
Develop open-source tools and reference architectures to make it easier to build and manage machine learning and AI workloads, pipelines, and systems at scale
Qualifications
Minimum
BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field, or equivalent experience
8+ years of experience in technical roles such as data science, data engineering, or ML engineering, ideally targeting large‑scale production systems
Demonstrated AI/ML experience across multiple phases of the machine learning lifecycle, from exploratory analysis to production systems
Facility with systems topics including Linux, batch schedulers, Kubernetes, distributed filesystems, and advanced networking at datacenter scale
Solid scripting and programming skills in languages like bash and Python and solid systems programming skills in a language like C++, Go, or Rust
Experience using machine learning or deep learning frameworks for training and inference
Excellent communication and technical presentation skills, with the ability to clearly articulate architectures, trade‑offs, and recommendations to both engineering and leadership audiences
A clear record of engineering discipline and execution on interesting projects, whether you’re working alone or collaborating on a team
Preferred
Experience contributing to and working in open-source communities
Experience with the NVIDIA ecosystem, including DGX systems, CUDA, NeMo, RAPIDS, Triton, NIM, and NVIDIA networking technologies such as InfiniBand, NVLink, and RoCE
Experience and familiarity building machine learning systems in a security-critical environment and distributed training and inference frameworks
Familiarity with MLOps practices in a cloud‑native context: containerization, CI/CD pipelines, workflow automation, observability stacks, and GitOps workflows
Direct experience drawing on deep systems knowledge to diagnose and fix performance or correctness problems that span multiple layers of the application stack, like hardware, networking, accelerator, hypervisor or OS, compilers or runtimes, application code, and libraries