Training-Free Instruction TTS Gender Bias Calibration Using Model-Adaptive Steering

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Instruction-based text-to-speech (ITTS) systems exhibit implicit gender biases that are difficult to calibrate without retraining. This work proposes a model-adaptive steering method that directly adjusts post-encoder representations for training-free bias correction through group bias vectors, coarse-to-fine intensity search, and deterministic lexical gating. To our knowledge, this mechanism is the first to operate without retraining while generalizing across heterogeneous architectures, eliminating implicit biases while strictly preserving explicit prompt semantics. Experimental results demonstrate that the proposed approach reduces aggregate calibration error by 0.8 to 5.9 percentage points across four mainstream ITTS models, with negligible degradation in audio quality and intelligibility.
📝 Abstract
Instruction text-to-speech (ITTS) systems systematically encode acoustic gender skews from descriptive style prompts, such as occupations or personas, even when demographic attributes are left unspecified. Calibrating the implicit gender distribution of the synthesized voices, across heterogeneous architectures and without model retraining or overriding explicit user prompts, is an open problem. In this work, we propose \textit{model-adaptive steering}, a training-free bias calibration method that steers post-encoder conditioning representations using a group deviation vector paired with a coarse-to-fine strength search on a development set. A deterministic lexical gate bypasses intervention whenever explicit gender keywords are detected, preserving intended prompt semantics. Evaluated on a 12,800-prompt held-out benchmark across four ITTS models (three architectures), the proposed method reduces aggregate calibration error from 11.5--28.1 to 0.8--5.9 percentage points (e.g., shifting Parler-TTS Mini from 78.1\% and PromptTTS++ from 27.3\% female to 50.8--52.4\%). Output rates stay within 0.9 pp across anchor set sizes at fixed operating points, while UTMOS decreases by at most 0.08 and WER degrades by at most 3.3 pp.
Problem

Research questions and friction points this paper is trying to address.

Instruction Text-to-Speech
Gender Bias
Bias Calibration
Training-Free
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-Free Bias Calibration
Model-Adaptive Steering
Instruction Text-to-Speech
Gender Bias
Deterministic Lexical Gate
🔎 Similar Papers
No similar papers found.