Senior Solutions Architect, NVIDIA Cloud Partner Operations

Nvidia
US, CA, Santa Clara2026-08-12onsite

About the job

NVIDIA is looking for a hands-on Solutions Architect to raise the Day 2 operations bar across our NVIDIA Cloud Partner ecosystem. Day 2 starts when a cluster is installed and validated: keeping the service healthy, adapting it as technology and customer demand change, and improving performance, stability, efficiency and economics over time. You will work with engineers running AI clouds at scale on the problems that decide whether customers stay and whether the next generation of NVIDIA technology lands successfully.

Responsibilities

Solve hard Day 2 operations problems at scale alongside partner engineers to find causes, prototype approaches, validate under load, and leave behind operable practices.

Make new technology Day 2 ready by helping partners prepare operating models for new NVIDIA platforms, capacity, services, and use cases before customer dependence.

Improve reliability, performance, and economics using measures such as incident frequency, recovery time, utilization, and cost per token to identify and fix performance or margin losses.

Raise each partner's Day 2 maturity by identifying and closing gaps across people, process, tooling, telemetry, security, and incident response.

Turn validated solutions into ecosystem capability by converting work into operating procedures, reference architectures, assessments, automation, and agentic workflows.

Create feedback loops by spotting patterns across partners and bringing field evidence to account teams, support, product, and engineering to fix repeated problems.

Qualifications

Minimum

BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or a related field - or equivalent experience.

12+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or a similar technical role; alternatively, 5+ years of exceptional specialist-level work in large-scale GPU or AI infrastructure.

Experience building, operating, or improving distributed infrastructure under real production load.

Deep expertise in at least one part of the Day 2 stack, backed by hands-on work with large-scale GPU, HPC, or cloud infrastructure.

Working experience across the broader operating platform, including Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry, and automation with Terraform, Ansible, Argo CD, or similar tooling.

Strong Linux knowledge and enough Python, Bash, or similar experience to automate measurement, diagnosis, validation, or remediation.

A detailed evidence-led approach to troubleshooting across system boundaries.

The ability to lead sophisticated work with partner engineers and cross-functional teams without direct authority.

Strong communication, prioritization, and time-management skills across multiple partner engagements.

Preferred

Real world experience operating a GPU cloud, HPC environment, or large-scale AI platform under customer load.

Built or matured a 24/7 operations function, including observability, incident and problem management, coverage, and on-call design.

Hands on experience with NVIDIA rack-scale platforms such as GB200 or GB300 NVL72 into production, or have hands-on experience with NVIDIA operations technologies such as Spectrum-X, UFM, Base Command Manager, Mission Control, and the GPU or Network Operators.

Driven improved fleet health or unit economics through benchmarking, infrastructure as code, GitOps, automated diagnosis, or agent-based remediation.