About the job
The Fleet Performance Optimization (FPO) organization builds the intelligence layer for Amazon Robotics manipulation workcells. We own the systems that select what work a robot should attempt, monitor how it performs, and close the loop by improving models with every induct. Our platforms serve Sparrow, Cardinal, FlexCell, and Robin workcells that handle millions of packages across Amazon's fulfillment network.
Responsibilities
- Building high-throughput, event-based scoring services that process real-time inventory signals across 10+ warehouses and hundreds of workcells
- Designing config-based annotation orchestration systems that onboard new ML data pipelines without code changes
- Developing model deployment infrastructure spanning edge devices, cloud inference services, and planning systems
- Building fleet monitoring systems that use VLMs and statistical methods to detect performance anomalies and surface root causes
- Creating unified data exploration and visualization tools used daily by engineers, scientists, and operations
- Implementing ongoing learning systems that intelligently select which data to collect based on inference outcomes
- Extending our foundation models, which are the backbone for multiple prediction tasks
Qualifications
Minimum
- 3+ years of non-internship professional software development experience
- 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience
- Experience programming with at least one software programming language
- Bachelor's degree in computer science or equivalent
- Experience designing and building distributed systems or data-intensive applications
- Experience with the full software development lifecycle: design, implementation, testing, deployment, and operations
Preferred
- 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
- Experience with ML infrastructure (training pipelines, model serving, feature stores, monitoring)
- Experience with event-driven architectures and stream processing at scale (Kinesis, Kafka, SQS)
- Experience with AWS services (ECS, Lambda, SageMaker, DynamoDB, S3, Athena)
- Experience building developer tools, data platforms, or observability systems
- Familiarity with ML concepts (model evaluation, data drift, active learning, annotation pipelines)
- Experience working in robotics, or computer vision
- Experience with infrastructure-as-code (CDK, CloudFormation)
- Strong written communication skills, ability to author design documents and influence technical decisions