About the job
Join our team as a Senior Machine Learning Engineer and help build enterprise-grade AI solutions at Firefly Foundry. Drive the development of multi-model pipelines, deploy production services, and ensure high availability and quality for cutting-edge generative AI models. Shape the future of media and entertainment technology with Adobe.
Responsibilities
Compose and build production pipelines from a heterogeneous set of models — prompt-rewrite LLMs, image and video generation, 3D mesh reconstruction, upsampling, NSFW and safety checkers, and IP guardrails.
Deploy these pipelines as services, scale them to enterprise traffic, and hold them to SLAs for availability, latency, and throughput.
Ensure served quality matches the training and reference environment — closing train/serve gaps across precision, preprocessing, and model versions.
Engineer for enterprise from the ground up: tenancy boundaries, data isolation, and the controls that let us honor customer IP contracts under audit.
Build the platform underneath it all — rapid pipeline deployment, observability, monitoring, and alerting.
Qualifications
Minimum
5+ years in machine learning engineering, with significant ownership of production ML or inference services at scale.
Strong Python and deep-learning engineering skills (PyTorch), with hands-on experience deploying and scaling model-backed services.
Experience composing multi-model pipelines and serving them behind APIs — orchestration, batching, autoscaling, and version management.
A track record owning production SLAs — availability, latency, and throughput — backed by real observability, monitoring, and alerting.
Comfort working across multiple, distinct generative model architectures (LLMs and VLMs, diffusion and transformer models, 3D/mesh) — enough to integrate, optimize, and reason about output quality, in partnership with Applied Science.
Experience with multi-tenant systems and data isolation in an enterprise or regulated context.
Fluency with containers and orchestration (Docker, Kubernetes), CI/CD for ML, and a major cloud (AWS or Azure).
GPU inference optimization for latency and cost — quantization, batching, and serving runtimes; custom CUDA a plus.
Preferred
No preferred qualifications listed.