Batched Speech Decisions Without Decoding: Single-Token Supervision Lets a Frozen LLM Hear Beyond the Transcript

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high response latency caused by conventional autoregressive decoding in full-duplex voice agents. We propose DuplexJev, an architecture that feeds hidden states from an ASR encoder into a frozen large language model via a cross-attention connector. By employing a single-token classification mechanism to directly read decisions, it enables decoding-free, ultra-fast batched inference. Furthermore, we replace transcription distillation with cross-entropy supervision on answer tokens, allowing the model to perceive gender and emotion information beyond textual content. Experimental results demonstrate that on an eight-GPU node, DuplexJev completes eighty decisions within 0.1 seconds, achieving nearly 90% accuracy across spoken question answering, gender recognition, and emotion recognition tasks.
📝 Abstract
Full-duplex voice agents make many small, closed decisions, which current systems answer by slow autoregressive decoding. We propose DuplexJev, which feeds ASR-encoder hidden states through a small connector into a frozen LLM and reads each question as a single-token distribution over its options. Nothing is decoded, and an 8-GPU node answers 80 decisions about eight utterances in about 0.1 s. With a last-layer connector, spoken QA stays close to reading the transcript (90% vs. 91%). DuplexJev also hears the speaker: gender and emotion accuracy both reach 90% (from 55% and 28%) with a cross-attention connector, whose spoken QA drops by only 1 point (83% to 82%). We train decisions with cross-entropy on the read-out answer token, instead of the usual transcript distillation, whose teacher never hears the voice, and keep distillation for content. Encoders and LLMs are interchangeable; we release weights, training recipe, a batched-inference pipeline for full-duplex serving and a bilingual spoken-QA set.
Problem

Research questions and friction points this paper is trying to address.

full-duplex voice agents
autoregressive decoding
spoken QA
paralinguistic information
transcript distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Single-Token Supervision
Decoding-Free
Frozen LLM
Cross-Attention Connector
Full-Duplex Voice Agent
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jie Jin
Adventists.ai, Melbourne, Australia
Z
Ziyin Ma
College of Computer Science and Technology, Zhejiang University, Hangzhou, China
M
Min Yin
Future Design Lab, Innovation Center of Yangtze River Delta, Zhejiang University, Jiaxing, China
Jinyu Chen
Jinyu Chen
The Hong Kong Polytechnic University
Edge/cloud computingVideo transmission.
H
Haigang Song
Adventists.ai, Melbourne, Australia
Z
Zhikun Pang
Adventists.ai, Melbourne, Australia
X
Xiaowen Zhang
Adventists.ai, Melbourne, Australia