🤖 AI Summary
This study addresses the lack of fine-grained control in instruction-based text-to-speech (TTS) and the absence of natural language interaction in character-level TTS by proposing a unified framework that maps natural language instructions to character-level acoustic control. Methodologically, it introduces the first automatic alignment mechanism from instructions to character-level acoustics, incorporating keyword prediction and grounding-aware loss weighting strategies. Leveraging Qwen3-Omni for annotated data construction, an autoregressive model is trained to identify target characters, predict acoustic attributes, and generate speech tokens. Experimental results demonstrate that this approach significantly enhances instruction adherence and keyword-level acoustic control capabilities, achieving explicit character-level controllability while preserving speech naturalness.
📝 Abstract
Instruction-based text-to-speech (ITTS) systems enable natural-language control of expressive speech generation, but often offer limited transparency and fine-grained control over individual text units. Character-level controllable TTS systems provide explicit acoustic control, yet typically rely on user-specified acoustic attributes. To bridge this gap, we propose InstCharVoice, a unified framework that grounds natural-language instructions in character-level acoustic control. We first construct grounded instruction annotations on the WordVoice-5A-zh corpus using Qwen3-Omni. With this supervision, we train an autoregressive model to identify instruction-relevant characters and predict their acoustic attributes before generating the corresponding speech tokens. Keyword prediction and grounding-aware loss weighting help the model focus on instruction-relevant characters and attributes. Experiments show improved instruction following and keyword-level acoustic control over representative ITTS systems, with competitive speech naturalness and explicit character-level controllability. Audio samples are available at https://xxh333.github.io/instcharvoice-demo/.