About the job
AWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud-scale machine learning accelerators and the trn* and inf* servers that use them. This position is for a Software Engineer that will lead the development of various services that will aid in optimization, analysis and release of machine learning workloads and artifacts. This candidate must have had experience leading distributed systems and machine learning related projects, preferably starting from architecture through several generations of delivery to customers. Deep knowledge of optimization, resource management, scheduling are needed. The ideal candidate will have experience working on services like EC2, EKS, Lambda in AWS or similar services on other cloud providers.
Responsibilities
* Lead the design and implementation of new tools, pipelines and automation, will work with developers, system architects, hardware engineers and users both within and external to Amazon to ensure compatibility of this new toolset with existing and next-generation AI accelerators.
* Design, implement, and maintain CI/CD pipelines to automate the software release process.
* Collaborate with development teams to integrate new software releases.
* Manage and automate infrastructure provisioning.
* Ensure high availability and scalability of systems through effective infrastructure management.
* Implement monitoring solutions to track system performance. Identify bottlenecks and optimize system performance.
* Implement security best practices in the DevOps pipeline. Conduct regular vulnerability assessments and risk management.
Qualifications
Minimum
- 3+ years of non-internship professional software development experience
- 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience
- Experience programming with at least one software programming language
- Knowledge of system performance, memory management, and parallel computing principles
- Experience in debugging, profiling, and implementing software engineering best practices in large-scale systems, or experience debugging, profiling, and implementing best software engineering practices in large-scale systems
- Experience with AWS or cloud technologies
Preferred
- 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
- Bachelor's degree in computer science or equivalent
- Knowledge of fundamentals of networking, security, databases (relational or NoSQL), operating systems (Unix, Linux, and/or Windows)
- Fundamentals of Machine learning and LLMs, their architecture along with work experience on certain LLM models.