About the job
The Annapurna Labs team at Amazon Web Services (AWS) builds AWS Neuron, the software development kit used to accelerate deep learning and GenAI workloads on AWS Trainium, Amazon's custom machine learning accelerator. Neuron includes an ML compiler, runtime, collectives library, and application framework that integrate with PyTorch and JAX, so customers can train frontier-scale models on Trainium without rewriting their stack.
Responsibilities
Implement and tune components of our distributed training stack for large-scale training, post-training, and reinforcement learning workloads on the latest Trainium instances, working across PyTorch and the Neuron software stack; Contribute to parallelism strategies such as data, tensor, and pipeline parallelism and apply reduced-precision formats under the guidance of senior team members; Profile workloads to help determine whether a bottleneck sits in compute, memory, collectives, or host overhead, and work with compiler, runtime, and collectives engineers to help land the fix; Build and maintain internal tooling, benchmarks, and tests that keep the team's performance work reproducible, and take on increasing ownership as you grow in the role.
Qualifications
Minimum
Bachelor's degree or above in computer science or equivalent; 3+ years of experience with full software development life cycle in production; 3+ years of experience with at least one programming language such as Python, C/C++, or a similar language; 2+ years of experience with system design (design patterns, reliability and scaling) of new or existing systems; Familiarity with LLM/transformer fundamentals such as attention mechanisms, autoregressive decoding, KV-cache behavior, and various forms of parallelism
Preferred
Master's degree or above in computer science or equivalent; Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience in computer architecture; Experience with ML frameworks such as Pytorch/Jax, Distributed libraries and Frameworks, RL frameworks or End-to-end Model Training; Experience with performance engineering: workload profiling, characterization (compute bound, memory bound, network bound), and optimization; Experience with open source projects