AI/ML Engineer, Amazon Global Data Center Ops Central Insight and Analytics Team

Amazon
Seattle, WA, USA2026-08-25ONSITE

About the job

We are looking for an AI/ML Engineer to build, deploy, and operate the ML/AI systems that power the agentic decision intelligence workflow we are building. You are the person who takes a model from a notebook to production, builds the LLM integration layer, implements RAG pipelines, creates evaluation frameworks, and ensures our AI systems are reliable, observable, and continuously improving.

This is a hands-on engineering role with deep ML/AI focus — you write production code that runs AI systems, not research papers. If you love the intersection of ML infrastructure, LLM applications, and production engineering, this role is for you.

Responsibilities

- Build and maintain LLM-powered components: structured reasoning chains, narrative generation, recommendation rationale

- Implement and optimize prompt engineering pipelines with version control, A/B testing, and regression detection

- Build RAG (Retrieval-Augmented Generation) systems that ground LLM outputs in operational data, historical playbooks, and domain knowledge

- Build guardrails, validation layers, and output parsing for LLM responses. Optimize latency, cost, and quality trade-offs across LLM providers

- Deploy ML models to production. Implement model monitoring: drift detection, performance degradation alerts, automated retraining triggers

- Build A/B testing infrastructure for model experiments. Manage model versioning, rollback, and canary deployment. Ensure SLA compliance for inference latency and availability

- Own the operational health of AI/ML services: monitoring, alarming, on-call, incident response, observability across the AI stack (prompt traces, latency histograms, token usage, error rates)

- Write comprehensive tests (unit, integration, end-to-end) for ML pipelines

Qualifications

Minimum

- 3+ years of non-internship professional software development experience

- Bachelor's degree in Computer Science, Machine Learning, or related field (or equivalent experience)

- 2+ years deploying ML models to production environments

- Strong Python proficiency + experience with ML frameworks

- Experience with LLM APIs and prompt engineering

- Experience with cloud ML services

- Experience building data pipelines for ML (feature engineering, preprocessing, training data management)

- Solid software engineering fundamentals (testing, CI/CD, code review, production operations)

Preferred

- 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience

- Experience building RAG systems (vector databases, embedding models, retrieval pipelines)

- Experience with agent/orchestration frameworks (LangChain, LangGraph, CrewAI, Bedrock Agents, or custom)

- Experience with ML evaluation frameworks (especially for generative AI / LLM outputs)

- Experience with time-series ML (forecasting, anomaly detection)

- Experience with MLOps tooling (MLflow, SageMaker Pipelines, Step Functions, feature stores)

- Experience with infrastructure-as-code (CDK, CloudFormation, Terraform)

- Background in operational/infrastructure environments