Staff Machine Learning Engineer - ML Frameworks

Adobe
California / Washington2026-09-02Full time

About the job

Exciting opportunity for a Staff Machine Learning Engineer to design and maintain AI/ML infrastructure supporting large-scale generative AI models. Lead the development of distributed training frameworks and collaborate with data scientists. Ideal for candidates with significant industry experience and expertise in Python, Kubernetes, and GPU-based machine learning.

Responsibilities

Design, develop, and maintain robust AI/ML infrastructure solutions to support the training and deployment of large-scale AI models, using Kubernetes and Python on AWS cloud

Implement and improve distributed training frameworks leveraging GPUs to improve performance and scalability

Improve resiliency, elasticity, data loading and provide out-of-the-box support for FSDP and model parallelism

Help train better models by improving orchestration and scheduling, scaling the number of jobs, faster experimentation with AutoML and similar

Collaborate with data scientists and ML researchers to streamline the model training pipeline and ensuring efficient resource utilization

Drive innovation in infrastructure practices to support pioneering machine learning research and development

Qualifications

Minimum

PhD or Master’s in computer science or related field and 5+ years of hands-on industry experience

Proven proficiency with Python and developing systems, frameworks and SDKs

Experience with infrastructure and understanding of model serving, training, orchestration, and management of GPU resources

Experience with machine learning and distributed Pytorch

Strong critical thinking, analytical and quantitative problem-solving ability

Excellent communication, relationship skills and a strong teammate

Preferred

Experience with KubeFlow, MLFlow, Ray, SageMaker, or similar

Experience with Pytorch distributed, MPI, Megatron, Horovod and other AI training frameworks