Senior Machine Learning Engineer, AI Platform

Adobe
San Jose, California, United States of America2026-09-03Full time

About the job

We are currently hiring a Senior Machine Learning Engineer, AI Platform to lead the architecture and operations of Adobe's AI and generative AI infrastructure. Drive the transition of models from experimentation to production, ensuring high latency and throughput. Ideal for experienced engineers with deep expertise in distributed systems and cloud infrastructure.

Responsibilities

● Own the architecture and roadmap for major components of the ML compute and inference platform, such as training orchestration, GPU scheduling and utilization, model serving, or the developer-facing surfaces ML teams build on.

● Design and operate distributed systems that run large-scale training and low-latency, high-throughput inference reliably across thousands of accelerators.

● Drive multi-tenancy, elasticity, and cost/utilization efficiency across a shared GPU fleet serving many teams with competing demands.

● Build the paths that move a model from experiment to production without re-implementation, from packaging and registry through deployment and safe rollout.

● Set engineering standards for reliability, observability, and performance, and raise the bar for how the platform is built and operated.

● Partner with ML researchers and product teams to turn emerging workloads into first-class platform capabilities, and inform capacity and hardware strategy.

● Provide technical leadership and mentorship across the platform organization.

Qualifications

Minimum

● 7+ years building and operating large-scale platform, infrastructure, or distributed systems in production, with direct ownership of performance, scalability, and reliability.

● Deep expertise in distributed systems and cloud infrastructure, including Kubernetes, containerized workloads, and operating large multi-node and multi-region clusters.

● Strong programming ability in Python and at least one systems language (Go, C++, Rust, or Java).

● A track record of designing systems that other engineers build on, making deliberate architectural tradeoffs and taking them from design to production at scale.

● A bias for measurable outcomes (latency, throughput, utilization, reliability) and the collaboration skills to drive them across teams and partners.

Preferred

● Experience with GPU or accelerator scheduling, performance tuning, or fleet management.

● Familiarity with ML framework internals or distributed training (PyTorch, FSDP, DeepSpeed) or modern inference stacks (vLLM, TensorRT-LLM, Triton, Ray Serve).

● Experience operating ML or data infrastructure at the scale of a major ML-driven product organization