Senior Software Engineer - AI/ML, AWS Neuron Inference

Amazon
USA, WA, Seattle2026-05-18ONSITE

About the job

AWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud-scale machine learning accelerators. This role is for a senior software engineer in the Machine Learning Inference Applications team. This role is responsible for development and performance optimization of core building blocks of LLM Inference - Attention, MLP, Quantization, Speculative Decoding, Mixture of Experts, etc.

Responsibilities

- Adapting latest research in LLM optimization to Neuron chips to extract best performance from both open source as well as internally developed models.

- Working across teams and organizations is key.

Qualifications

Minimum

- 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience

- Bachelor's degree in computer science or equivalent

- 5+ years of programming using a modern programming language such as Java, C++, or C#, including object-oriented design experience

- Fundamentals of Machine learning models, their architecture, training and inference lifecycles along with work experience on some optimizations for improving the model performance.

Preferred

- Master's degree in computer science or equivalent

- Hands-on experience with PyTorch or Jax - preferably involving developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware.