About the job
As a Research Engineer focused on Audio and Speech, you’ll be responsible for building the models and agent harnesses that power Decagon’s real-time voice agents and taking them all the way from idea to production. Your work will advance multimodal and full-duplex systems that can listen, reason, speak, and respond naturally in real time.
We’re looking for strong engineers who want to build the next generation of AI voice agents. People here own their work end-to-end, ship real improvements, and are trusted to make high-impact technical decisions.
Responsibilities
Design and build next-generation agent harnesses optimized for streaming speech, turn-taking, interruptions, overlapping speech, and continuous interaction
Research and train multimodal and full-duplex models that jointly understand audio, reason, and generate speech
Improve speech recognition, voice activity detection, endpointing, and speech generation across diverse speakers, environments, domains, and languages
Build evaluations and use production calls to ship measurable improvements in accuracy, latency, naturalness, and task outcomes
Optimize end-to-end inference for responsiveness, throughput, stability, and cost, partnering with Voice Platform and Infrastructure teams to deploy at scale
Qualifications
Minimum
2+ years of experience in speech, audio ML, multimodal ML, or production machine learning
Experience developing or adapting autoregressive, diffusion, flow-matching, or codec-based speech models
Hands-on experience with streaming agent systems, low-latency inference, production model serving, and evaluation on real-world audio
Fluency in Python and a modern deep-learning framework such as PyTorch, with strong foundations in machine learning and signal processing
A track record of taking research ideas from prototype to reliable, measurable production impact
Preferred
Familiarity with speech-to-speech or full-duplex models
Experience with telephony, multilingual speech, noisy-channel robustness, speaker adaptation, or expressive speech generation