Principal Machine Learning Engineer

Adobe
California / Washington2026-08-17Full time

About the job

Firefly Foundry is Adobe's enterprise managed-service offering for custom multimedia generative AI — deep-tuned image, video, and 3D models built on each customer's IP, paired with creative production workflows and a media-intelligence layer, and deployed across new and existing Adobe surfaces and products, including Firefly, Photoshop, Illustrator, Express, Stock, and Premiere. We are hiring a Principal Machine Learning Engineer to serve as the technical lead for our GenAI Services area. This is not a model-training or research role — it is the senior-most hands-on engineering authority over how our generative models are architected, optimized, and served at enterprise scale. You will set the inference architecture and technical standards that a growing organization of engineers builds against, co-develop and optimize the inference code that makes those systems fast and cost-efficient, and architect the APIs and product backend that let Adobe's first-party and third-party models reach both internal applications and external plugin integrations.

Responsibilities

Lead the development of core GenAI services and APIs that integrate a wide range of first-party and third-party generative models into Adobe's flagship products.

Architect ML serving workflows for enterprise-scale model customization, deployment, and ecosystem integration — including externalizable, self-serve fine-tuning flows.

Co-develop and optimize GPU-accelerated inference pipelines — prioritizing latency, throughput, scalability, and reliability — using tools such as PyTorch, CUDA, Triton, and TensorRT.

Design and architect the product backend and plugin ecosystem that lets internal applications and external integrations consume Firefly Foundry's model services.

Provide hands-on technical leadership: guide engineers through architecture, design, implementation, and best practices, and mentor a growing organization of ML engineers.

Research and evaluate emerging inference and MLOps technologies — serving runtimes, quantization, GPU scheduling — to improve engineering velocity and system performance.

Qualifications

Minimum

MS or PhD in Computer Science, Machine Learning, or a related field — or equivalent industry experience.

8+ years of experience in machine learning engineering, including production-scale deployment and serving — not training or research experimentation.

3+ years leading the technical direction of large-scale, GPU-intensive GenAI inference systems — serving, architecture, and optimization.

Deep experience with inference frameworks and tools such as PyTorch, CUDA, Triton, TensorRT, Nvidia Dynamo, and Python.

Strong understanding of generative model architectures — diffusion models, transformers, GANs, LLMs — sufficient to make architecture and optimization calls and reason about output quality, in partnership with Applied Science.

Proven experience architecting multi-model pipelines and serving them behind APIs at enterprise scale.

Experience designing product backend systems and plugin architectures consumed by internal applications and external integrations.

Proven success leading cross-functional teams through complex, high-stakes technical initiatives, with a track record of driving alignment in matrixed organizations.

Excellent communication and technical leadership skills.

Preferred

Experience with model serving, orchestration, and GPU resource management in large-scale environments.

Hands-on expertise in Kubernetes, distributed systems, and MLOps platforms.

Experience with RAG architectures and multi-turn, agentic conversational systems.

Experience with quantization, distillation, or other model-optimization techniques for inference.