Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training

Amazon
Cupertino, CA, USA2026-08-26ONSITE

About the job

This role is for a Senior Machine Learning Engineer in the Distribute Training team for AWS Neuron, responsible for development, enablement and performance tuning of a wide variety of ML model families, including massive-scale Large Language Models (LLM) such as GPT and Llama, as well as Stable Diffusion, Vision Transformers (ViT) and many more.

Responsibilities

Lead efforts to build distributed training support into PyTorch and JAX using XLA, the Neuron compiler, and runtime stacks

Optimize models to achieve peak performance and maximize efficiency on AWS custom silicon, including Trainium and Inferentia, as well as Trn2, Trn1, Inf1, and Inf2 servers

Work side by side with chip architects, compiler engineers and runtime engineers to create, build and tune distributed training solutions with Trainium instances

Extend distributed training libraries such as FSDP, Deepspeed, and Nemo for the Neuron based system

Qualifications

Minimum

Bachelor's degree in computer science or equivalent

5+ years of non-internship professional software development experience

5+ years of programming with at least one software programming language experience

5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience

5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience

Experience as a mentor, tech lead or leading an engineering team

Experience in machine learning, data mining, information retrieval, statistics or natural language processing

Preferred

Master's degree in computer science or equivalent

Experience in computer architecture

Previous software engineering expertise with Pytorch/Jax/Tensorflow, Distributed libraries and Frameworks, End-to-end Model Training