InstCharVoice: Grounding Natural-Language Instructions for Character-Level Control in Text-to-Speech

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of fine-grained control in instruction-based text-to-speech (TTS) and the absence of natural language interaction in character-level TTS by proposing a unified framework that maps natural language instructions to character-level acoustic control. Methodologically, it introduces the first automatic alignment mechanism from instructions to character-level acoustics, incorporating keyword prediction and grounding-aware loss weighting strategies. Leveraging Qwen3-Omni for annotated data construction, an autoregressive model is trained to identify target characters, predict acoustic attributes, and generate speech tokens. Experimental results demonstrate that this approach significantly enhances instruction adherence and keyword-level acoustic control capabilities, achieving explicit character-level controllability while preserving speech naturalness.
📝 Abstract
Instruction-based text-to-speech (ITTS) systems enable natural-language control of expressive speech generation, but often offer limited transparency and fine-grained control over individual text units. Character-level controllable TTS systems provide explicit acoustic control, yet typically rely on user-specified acoustic attributes. To bridge this gap, we propose InstCharVoice, a unified framework that grounds natural-language instructions in character-level acoustic control. We first construct grounded instruction annotations on the WordVoice-5A-zh corpus using Qwen3-Omni. With this supervision, we train an autoregressive model to identify instruction-relevant characters and predict their acoustic attributes before generating the corresponding speech tokens. Keyword prediction and grounding-aware loss weighting help the model focus on instruction-relevant characters and attributes. Experiments show improved instruction following and keyword-level acoustic control over representative ITTS systems, with competitive speech naturalness and explicit character-level controllability. Audio samples are available at https://xxh333.github.io/instcharvoice-demo/.
Problem

Research questions and friction points this paper is trying to address.

Text-to-Speech
Instruction-based TTS
Character-level control
Natural-language instructions
Acoustic attributes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Instruction-based TTS
Character-level control
Grounding
Autoregressive model
Acoustic attributes
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Sihang Nie
South China University of Technology
X
Xueru Li
South China University of Technology
Xiaofen Xing
Xiaofen Xing
South China University of Technology
D
Deyi Tuo
Huya Inc.
Cheng-Bin Jin
Cheng-Bin Jin
Huya Inc.
J
Jingyuan Xing
South China University of Technology
J
Jinxin Ji
The Hongkong Polytechnic University