Research Engineer, Safety

Decagon
San Francisco / New York City2026-09-04

About the job

As a Research Engineer focused on Safety, you’ll be responsible for making Decagon’s AI agents safe, reliable, and controllable, from evaluation through production. You’ll identify real-world failure modes and build the models, evaluations, and safeguards that prevent them.

We’re looking for strong engineers who want to advance applied AI safety in production. People here own their work end-to-end, ship real improvements, and are trusted to make high-impact technical decisions.

Responsibilities

Research and build safeguards against prompt injection, unsafe tool use, sensitive-data disclosure, policy violations, and hallucinated commitments

Build adversarial evaluations, simulations, red-team datasets, and regression suites informed by production failures

Develop and deploy classifiers, judges, reward signals, post-training methods, and runtime safeguards for safer agent behavior

Analyze production traces and incidents to identify root causes, test mitigations, and measure their impact

Partner with Security, Product, Infrastructure, Legal, and customer-facing teams to turn enterprise requirements into scalable safeguards and rollout practices

Qualifications

Minimum

2+ years of experience in AI/ML engineering, research, or AI safety

Hands-on experience evaluating, post-training, or deploying language models or agentic systems

Experience with modern post-training techniques, such as reinforcement learning, preference optimization, distillation, model routing, and synthetic-data generation

Experience with adversarial testing, model red teaming, prompt injection, policy enforcement, privacy, or safe tool use

Fluency in Python and modern ML tooling, with strong experimental judgment and the engineering depth to ship production systems

Comfort owning ambiguous, high-stakes technical problems and making clear risk and product tradeoffs

Preferred

Experience building safeguards for high-stakes or regulated enterprise workflows

Familiarity with human-in-the-loop review, incident response, or responsible rollout frameworks for ML systems