Prompts versus Rules: Auditing and Controlling Speech Naturalness Behaviors in Voice User Simulators

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the disconnect between configured naturalness behaviors and actual outputs in spoken user simulators, which compromises evaluation validity. To this end, we propose a linguistics-based deterministic rule injection algorithm alongside a naturalness behavior auditing framework. By systematically comparing large language model (LLM) prompt engineering with rule injection in simulating phenomena such as disfluencies and interruptions, we expose the unreliability of LLM prompting. Our findings demonstrate that the proposed rule-based approach significantly outperforms prompt-based methods in both behavioral authenticity and controllability. Furthermore, this work emphasizes that simulator evaluations must audit actual generated behaviors rather than relying solely on configuration parameters, thereby establishing a robust paradigm for constructing high-fidelity spoken interaction evaluation systems.
📝 Abstract
As voice agents gain more popularity commercially, the user simulators used to evaluate the deployed agents are also being developed to include more realistic, variable, and diverse speech naturalness behaviors -- disfluency, interruption and backchanneling. The quality of the user simulator directly affects the validity of agent evaluation results. However, we find that most studies so far have not examined in detail whether the intended configuration for these behaviors is realized in the simulation. In this study, we audit the realized naturalness behaviors of tau-Voice, our own LLM-based prompting approach across three models, and our rule-based injection algorithm for disfluency, interruption, and backchanneling. We find that prompting for these behaviors is unreliable and produces speech inconsistent with the instructions, placed and distributed less naturally than the instruction implies. In contrast, our rule-based, model-free algorithm produces controllable and diverse naturalness behaviors more aligned with natural speech. Our results suggest that LLMs not purpose-trained for user simulation are not sufficient on their own to represent authentic user behavior, and that linguistically informed deterministic approaches or specialized models are needed to close the gap; auditing and reporting realized naturalness behaviors, rather than configured settings, is what makes that gap visible.
Problem

Research questions and friction points this paper is trying to address.

voice user simulator
speech naturalness
disfluency
interruption
backchanneling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Voice User Simulators
Speech Naturalness Behaviors
Rule-based Injection
Prompt Auditing
Large Language Models
🔎 Similar Papers
No similar papers found.