About the job
Shape the Future of AI Accelerators at AWS Neuron
We build Amazon Neuron, the software development kit used to accelerate deep learning and GenAI workloads on Amazon’s custom machine learning accelerators, Inferentia and Trainium.
As a Senior Software Engineer on our Machine Learning Applications team, you will optimize the world's most demanding AI models at a scale few engineers ever get to work on. This role offers a unique opportunity to work at the intersection of machine learning, high-performance computing, and distributed architectures, where you'll help shape the future of AI acceleration technology.
Responsibilities
Design, develop, and optimize machine learning models including GPT, Kimi, and Qwen on custom AI accelerators.
Participate in all stages of the ML system development lifecycle including distributed computing based architecture design, implementation, performance profiling, low level optimizations, and production deployment.
Build infrastructure to systematically analyze and onboard multiple models with diverse architecture.
Design and implement high-performance kernels and features for ML operations, leveraging the Neuron architecture and programming models
Analyze and optimize system-level performance across multiple generations of Neuron hardware
Conduct detailed performance analysis using profiling tools to identify and resolve bottlenecks
Implement optimizations such as fusion, sharding, tiling, and scheduling
Work directly with customers to enable and optimize their ML models on AWS accelerators
Collaborate across teams to develop innovative optimization techniques
Develop kernels to improve model efficiency on Amazon AI Accelerators
Transform complex tensor operations into highly optimized graph implementations
Optimize state-of-the-art language, vision, and multimodal generative AI models for Neuron hardware
Qualifications
Minimum
Bachelor's degree
5+ years of non-internship professional software development experience
Knowledge of Python and/or C++ programming
5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience
Experience in debugging, profiling, and implementing software engineering best practices in large-scale systems
Knowledge of system performance, memory management, and parallel computing principles
Experience owning a performance optimization roadmap and mentoring engineers on optimization
Preferred
Master's degree in computer science or equivalent
Knowledge of machine learning model architecture and inference
Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques
Hands-on development with PyTorch is preferred
Experience scaling workloads across multi-GPU and multi-node topologies with NCCL and tensor, pipeline, or expert parallelism
Experience writing and optimizing custom CUDA/Triton kernels for tensor operations