Unifying Speech Recognition, Synthesis and Conversion with Autoregressive Transformers

📅 2026-01-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes GPA, a general-purpose audio model that unifies speech synthesis, automatic speech recognition, and voice conversion within a single autoregressive Transformer architecture—addressing the fragmentation, poor scalability, and limited generalization of traditional task-specific speech systems. By leveraging a shared discrete speech token space and an instruction-driven mechanism, GPA enables zero-architecture-modification task switching. The model employs multi-task joint training and a high-throughput inference pipeline, facilitating lightweight deployment. Experimental results demonstrate that GPA achieves competitive performance across multiple tasks, with its 0.3B-parameter variant particularly well-suited for low-latency, resource-constrained edge scenarios.

Technology Category

Natural Language Processing: SpeechMachine Learning: Large Multimodal Models (LMMs)Computer Vision: Generative Adversarial Networks (GANs) for Vision

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGEconomics, Online Markets and Human Computation: Economic ramifications for generative AI infrastructure and applicationsSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applications
📝 Abstract
Traditional speech systems typically rely on separate, task-specific models for text-to-speech (TTS), automatic speech recognition (ASR), and voice conversion (VC), resulting in fragmented pipelines that limit scalability, efficiency, and cross-task generalization. In this paper, we present General-Purpose Audio (GPA), a unified audio foundation model that integrates multiple core speech tasks within a single large language model (LLM) architecture. GPA operates on a shared discrete audio token space and supports instruction-driven task induction, enabling a single autoregressive model to flexibly perform TTS, ASR, and VC without architectural modifications. This unified design combines a fully autoregressive formulation over discrete speech tokens, joint multi-task training across speech domains, and a scalable inference pipeline that achieves high concurrency and throughput. The resulting model family supports efficient multi-scale deployment, including a lightweight 0.3B-parameter variant optimized for edge and resource-constrained environments. Together, these design choices demonstrate that a unified autoregressive architecture can achieve competitive performance across diverse speech tasks while remaining viable for low-latency, practical deployment.
Problem

Research questions and friction points this paper is trying to address.

speech recognition
speech synthesis
voice conversion
unified model
autoregressive transformers
Innovation

Methods, ideas, or system contributions that make the work stand out.

unified speech model
autoregressive transformer
discrete audio tokens
instruction-driven task induction
multi-task speech processing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Runyuan Cai
AutoArk-AI
Y
Yu Lin
AutoArk-AI
Y
Yiming Wang
AutoArk-AI
C
Chunlin Fu
AutoArk-AI
X
Xiaodong Zeng
AutoArk-AI