SPEAR-Gen: Generation-Aware Pre-training for Unified Speech Representations

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent trade-off in existing speech models, where representation capabilities for understanding and generation remain disjointed. To bridge this gap, we propose a unified speech representation framework that, for the first time, enables a single representation to simultaneously support both tasks through task-aligned feature aggregation and a coarse-to-fine optimization objective. Specifically, our method synergistically trains a frozen encoder with discrete masked prediction, Log-Mel spectrogram reconstruction, and residual flow matching. Experimental results demonstrate that the proposed model maintains strong understanding performance on the SUPERB benchmark while significantly improving speech resynthesis quality and speaker preservation. By effectively closing the performance gap between understanding and generation, this work establishes a robust foundation for versatile speech modeling within a unified representational paradigm.
📝 Abstract
Speech understanding and generation place different demands on speech representations, and existing models are typically optimised towards one capability or the other. To reduce this gap, we introduce SPEAR-Gen, a speech representation model that learns a single representation for both capabilities. Task-aligned feature aggregation consolidates complementary linguistic and paralinguistic information across a frozen encoder into discrete targets for masked prediction, while a coarse-to-fine objective combines log-Mel reconstruction with residual flow matching to preserve spectral structure and fine-grained acoustic variation. Experiments on SUPERB and speech resynthesis show that SPEAR-Gen maintains strong understanding performance while substantially improving resynthesis quality and speaker preservation. These results demonstrate that a single speech representation can effectively support both understanding and generation.
Problem

Research questions and friction points this paper is trying to address.

Speech Representation
Speech Understanding
Speech Generation
Unified Model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified Speech Representations
Task-aligned Feature Aggregation
Coarse-to-fine Objective
Residual Flow Matching
Masked Prediction
🔎 Similar Papers
No similar papers found.