About the job
Exciting opportunity for a Staff Machine Learning Engineer to design and maintain AI/ML infrastructure supporting large-scale generative AI models. Lead the development of distributed training frameworks and collaborate with data scientists. Ideal for candidates with significant industry experience and expertise in Python, Kubernetes, and GPU-based machine learning.
Responsibilities
Design, develop, and maintain robust AI/ML infrastructure solutions to support the training and deployment of large-scale AI models, using Kubernetes and Python on AWS cloud
Implement and improve distributed training frameworks leveraging GPUs to improve performance and scalability
Improve resiliency, elasticity, data loading and provide out-of-the-box support for FSDP and model parallelism
Help train better models by improving orchestration and scheduling, scaling the number of jobs, faster experimentation with AutoML and similar
Collaborate with data scientists and ML researchers to streamline the model training pipeline and ensuring efficient resource utilization
Drive innovation in infrastructure practices to support pioneering machine learning research and development
Qualifications
Minimum
PhD or Master’s in computer science or related field and 5+ years of hands-on industry experience
Proven proficiency with Python and developing systems, frameworks and SDKs
Experience with infrastructure and understanding of model serving, training, orchestration, and management of GPU resources
Experience with machine learning and distributed Pytorch
Strong critical thinking, analytical and quantitative problem-solving ability
Excellent communication, relationship skills and a strong teammate
Preferred
Experience with KubeFlow, MLFlow, Ray, SageMaker, or similar
Experience with Pytorch distributed, MPI, Megatron, Horovod and other AI training frameworks