Research Engineer, Audio and Speech

Decagon
San Francisco / New York City2026-09-04

About the job

As a Research Engineer focused on Audio and Speech, you’ll be responsible for building the models and agent harnesses that power Decagon’s real-time voice agents and taking them all the way from idea to production. Your work will advance multimodal and full-duplex systems that can listen, reason, speak, and respond naturally in real time.

We’re looking for strong engineers who want to build the next generation of AI voice agents. People here own their work end-to-end, ship real improvements, and are trusted to make high-impact technical decisions.

Responsibilities

Design and build next-generation agent harnesses optimized for streaming speech, turn-taking, interruptions, overlapping speech, and continuous interaction

Research and train multimodal and full-duplex models that jointly understand audio, reason, and generate speech

Improve speech recognition, voice activity detection, endpointing, and speech generation across diverse speakers, environments, domains, and languages

Build evaluations and use production calls to ship measurable improvements in accuracy, latency, naturalness, and task outcomes

Optimize end-to-end inference for responsiveness, throughput, stability, and cost, partnering with Voice Platform and Infrastructure teams to deploy at scale

Qualifications

Minimum

2+ years of experience in speech, audio ML, multimodal ML, or production machine learning

Experience developing or adapting autoregressive, diffusion, flow-matching, or codec-based speech models

Hands-on experience with streaming agent systems, low-latency inference, production model serving, and evaluation on real-world audio

Fluency in Python and a modern deep-learning framework such as PyTorch, with strong foundations in machine learning and signal processing

A track record of taking research ideas from prototype to reliable, measurable production impact

Preferred

Familiarity with speech-to-speech or full-duplex models

Experience with telephony, multilingual speech, noisy-channel robustness, speaker adaptation, or expressive speech generation