About the job
As a Research Engineer focused on Safety, you’ll be responsible for making Decagon’s AI agents safe, reliable, and controllable, from evaluation through production. You’ll identify real-world failure modes and build the models, evaluations, and safeguards that prevent them.
We’re looking for strong engineers who want to advance applied AI safety in production. People here own their work end-to-end, ship real improvements, and are trusted to make high-impact technical decisions.
Responsibilities
Research and build safeguards against prompt injection, unsafe tool use, sensitive-data disclosure, policy violations, and hallucinated commitments
Build adversarial evaluations, simulations, red-team datasets, and regression suites informed by production failures
Develop and deploy classifiers, judges, reward signals, post-training methods, and runtime safeguards for safer agent behavior
Analyze production traces and incidents to identify root causes, test mitigations, and measure their impact
Partner with Security, Product, Infrastructure, Legal, and customer-facing teams to turn enterprise requirements into scalable safeguards and rollout practices
Qualifications
Minimum
4+ years of experience in AI/ML engineering, research, or AI safety
Hands-on experience evaluating, post-training, or deploying language models or agentic systems
Experience with modern post-training techniques, such as reinforcement learning, preference optimization, distillation, model routing, and synthetic-data generation
Experience with adversarial testing, model red teaming, prompt injection, policy enforcement, privacy, or safe tool use
Fluency in Python and modern ML tooling, with strong experimental judgment and the engineering depth to ship production systems
Comfort owning ambiguous, high-stakes technical problems and making clear risk and product tradeoffs
Preferred
Experience building safeguards for high-stakes or regulated enterprise workflows
Familiarity with human-in-the-loop review, incident response, or responsible rollout frameworks for ML systems