Senior Technical Marketing Engineer - DSX AI Infrastructure Software

Nvidia
US, CA, Santa Clara / US, CA, Remote2026-08-26remote_local

About the job

NVIDIA DSX brings together facilities infrastructure, hardware, software, simulation, and partner technologies to build and run efficient AI factories. We are looking for a Senior Technical Marketing Engineer to show and educate our AI factory ecosystem how to bring up and operate the entire stack, ranging from facilities and multi-node GPU infrastructure to provisioning, networking, storage, cluster orchestration, security, observability, and workload enablement.

Responsibilities

Stand up and validate complete DSX-aligned software stacks on multi-node GPU systems, capturing dependencies, configuration order, validation steps, and operational handoffs.

Turn working deployments into useful technical content including reference architectures, quick-starts, installation and upgrade guides, troubleshooting runbooks, code examples, blogs, whitepapers, and demo videos.

Build reusable examples and automation with APIs, Python or shell scripting, infrastructure-as-code, containers, Kubernetes, Slurm, Helm, GitOps, and CI/CD.

Build demos, labs, and training addressing practical aspects of operating an AI factory, from deployment and tenant setup to upgrades, monitoring, scheduling, fault isolation, remediation, capacity management, and security.

Test pre-release software using representative training and inference workloads to identify rough edges, assess interoperability and resiliency, and provide feedback to Product and Engineering.

Help solution architects, field teams, cloud and OEM partners, ISVs, and system integrators use the stack successfully through repeatable assets, train-the-trainer sessions, live demos, and direct support.

Qualifications

Minimum

BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or another technical field, or equivalent experience.

8+ years of experience in infrastructure engineering, systems engineering, solutions architecture, software engineering, technical marketing engineering, site reliability engineering, or a related role.

Hands-on experience deploying and operating Linux-based data center, cloud, HPC, or AI infrastructure, including multi-node GPU systems and production operational practices.

Strong working knowledge of Kubernetes and/or Slurm, including containers, operators, Helm charts, cluster lifecycle, and workload scheduling.

Experience in several core infrastructure domains, such as bare-metal provisioning, firmware and drivers, compute, Ethernet or InfiniBand networking, storage, identity, multi-tenancy, secrets or certificate management, telemetry, observability, and fleet health.

Ability to automate deployments and operations through scripting, APIs, configuration management, infrastructure-as-code, Git-based workflows, and CI/CD.

Examples of technical work for practitioner audiences, such as deployment guides, documentation, reference architectures, code repositories, demos, workshops, blog posts, conference talks, or training.

Excellent written, verbal, and visual communication skills.

Ability to balance multiple projects and constituents, prioritize under tight deadlines, and work well across Engineering, Product, Field, Marketing, and partner teams.

Preferred

Experience with NVIDIA DSX, DGX systems, DGX Cloud, NVIDIA AI Enterprise, BlueField DPUs, DOCA, or related NVIDIA infrastructure software.

Experience operating large GPU clusters and diagnosing distributed performance, networking, storage, scheduling, or hardware-health issues.

Experience with AI training and inference workloads and the requirements for operating them dependably on accelerated infrastructure.

Experience connecting infrastructure software to facilities or operational technology systems, including power, cooling, building management systems.

Active participation in cloud-native, HPC, infrastructure automation, or open-source communities, including published examples or project contributions.