SteerablePlex: Can We Steer Full-Duplex Models?

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the tendency of existing full-duplex speech models, when employed as user simulators, to deviate from predefined scenarios as dialogue history grows, thereby compromising evaluation reliability. To mitigate this, we propose a controllable full-duplex user simulation framework that leverages an asynchronous backend large language model to dynamically monitor conversations and inject instructions, ensuring strict adherence to multi-stage constraints. Our core contributions include the construction of SimIF-Bench, a dedicated evaluation benchmark, and the introduction of Group-reward Decoupled Policy Optimization (GDPO), a novel training method. Experimental results demonstrate that our approach significantly outperforms existing open-source models and GPT-Realtime while preserving natural full-duplex interactions, substantially enhancing both instruction-following reliability and simulation fidelity.
📝 Abstract
Full-duplex speech models can listen and speak simultaneously, enabling natural interaction, but become increasingly difficult to control as the conversation history grows. When used as user simulators, this lack of control can cause them to deviate from prescribed scenarios and produce unreliable evaluation outcomes. We introduce SimIF-Bench (Simulator Instruction-Following Benchmark), which evaluates whether a conversational model stays within a prescribed scenario and completes multiple goals in the required order. The benchmark reveals that current open-source full-duplex models struggle to follow such constraints. We then introduce a Group Reward-Decoupled Normalization Policy Optimization (GDPO)-based training recipe that enables a full-duplex model to follow textual instructions during an ongoing conversation while maintaining its turn-taking ability. By connecting the resulting SteerablePlex to an asynchronous backend language model that monitors the conversation and provides instructions when needed, we build a more controllable full-duplex user simulator that follows multi-stage constraints more reliably than existing open-source models and GPT-Realtime.
Problem

Research questions and friction points this paper is trying to address.

Full-duplex speech models
Controllability
User simulator
Instruction following
Multi-stage constraints
Innovation

Methods, ideas, or system contributions that make the work stand out.

Full-Duplex Speech Models
GDPO
SimIF-Bench
User Simulator
Instruction Following
🔎 Similar Papers
No similar papers found.