About the job
We are looking for an AI/ML Engineer to build, deploy, and operate the ML/AI systems that power the agentic decision intelligence workflow we are building. You are the person who takes a model from a notebook to production, builds the LLM integration layer, implements RAG pipelines, creates evaluation frameworks, and ensures our AI systems are reliable, observable, and continuously improving.
This is a hands-on engineering role with deep ML/AI focus — you write production code that runs AI systems, not research papers. If you love the intersection of ML infrastructure, LLM applications, and production engineering, this role is for you.
Responsibilities
- Build and maintain LLM-powered components: structured reasoning chains, narrative generation, recommendation rationale
- Implement and optimize prompt engineering pipelines with version control, A/B testing, and regression detection
- Build RAG (Retrieval-Augmented Generation) systems that ground LLM outputs in operational data, historical playbooks, and domain knowledge
- Build guardrails, validation layers, and output parsing for LLM responses. Optimize latency, cost, and quality trade-offs across LLM providers
- Deploy ML models to production. Implement model monitoring: drift detection, performance degradation alerts, automated retraining triggers
- Build A/B testing infrastructure for model experiments. Manage model versioning, rollback, and canary deployment. Ensure SLA compliance for inference latency and availability
- Own the operational health of AI/ML services: monitoring, alarming, on-call, incident response, observability across the AI stack (prompt traces, latency histograms, token usage, error rates)
- Write comprehensive tests (unit, integration, end-to-end) for ML pipelines
Qualifications
Minimum
- 3+ years of non-internship professional software development experience
- Bachelor's degree in Computer Science, Machine Learning, or related field (or equivalent experience)
- 2+ years deploying ML models to production environments
- Strong Python proficiency + experience with ML frameworks
- Experience with LLM APIs and prompt engineering
- Experience with cloud ML services
- Experience building data pipelines for ML (feature engineering, preprocessing, training data management)
- Solid software engineering fundamentals (testing, CI/CD, code review, production operations)
Preferred
- 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
- Experience building RAG systems (vector databases, embedding models, retrieval pipelines)
- Experience with agent/orchestration frameworks (LangChain, LangGraph, CrewAI, Bedrock Agents, or custom)
- Experience with ML evaluation frameworks (especially for generative AI / LLM outputs)
- Experience with time-series ML (forecasting, anomaly detection)
- Experience with MLOps tooling (MLflow, SageMaker Pipelines, Step Functions, feature stores)
- Experience with infrastructure-as-code (CDK, CloudFormation, Terraform)
- Background in operational/infrastructure environments