Toward Complex-Valued Neural Networks for Waveform Generation

📅 2026-03-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing iSTFT-based neural vocoders, which employ real-valued networks to process the real and imaginary components of complex spectra separately, thereby failing to effectively capture their intrinsic structural relationships and constraining synthesis quality. To overcome this, the authors propose ComVo—the first neural vocoder built upon native complex-valued neural networks—featuring a complex-valued generator and discriminator trained within an adversarial framework. ComVo further introduces a novel phase quantization mechanism to structurally guide phase transformation and incorporates a block-matrix computation strategy to reduce redundancy and enhance efficiency. Experimental results demonstrate that ComVo achieves superior audio synthesis quality compared to real-valued baselines while reducing training time by 25%.

Technology Category

Computer Vision: Generative Adversarial Networks (GANs) for VisionMachine Learning: Deep Generative Models & AutoencodersCognitive Modeling & Cognitive Systems: Neural Spike Coding

Application Category

Social Networks and Social Media: Generative AI / large language models and their impact on social systemsEconomics, Online Markets and Human Computation: Economic ramifications for generative AI infrastructure and applicationsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Neural vocoders have recently advanced waveform generation, yielding natural and expressive audio. Among these approaches, iSTFT-based vocoders have recently gained attention. They predict a complex-valued spectrogram and then synthesize the waveform via iSTFT, thereby avoiding learned upsampling stages that can increase computational cost. However, current approaches use real-valued networks that process the real and imaginary parts independently. This separation limits their ability to capture the inherent structure of complex spectrograms. We present ComVo, a Complex-valued neural Vocoder whose generator and discriminator use native complex arithmetic. This enables an adversarial training framework that provides structured feedback in complex-valued representations. To guide phase transformations in a structured manner, we introduce phase quantization, which discretizes phase values and regularizes the training process. Finally, we propose a block-matrix computation scheme to improve training efficiency by reducing redundant operations. Experiments demonstrate that ComVo achieves higher synthesis quality than comparable real-valued baselines, and that its block-matrix scheme reduces training time by 25%. Audio samples and code are available at https://hs-oh-prml.github.io/ComVo/.
Problem

Research questions and friction points this paper is trying to address.

complex-valued neural networks
neural vocoder
waveform generation
complex spectrogram
iSTFT-based vocoder
Innovation

Methods, ideas, or system contributions that make the work stand out.

Complex-valued neural networks
Neural vocoder
Phase quantization
Block-matrix computation
Adversarial training
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Hyung-Seok Oh
Hyung-Seok Oh
Korea Unviersity
Speech synthesis Deep Learning
D
Deok-Hyeon Cho
Department of Artificial Intelligence, Korea University, Seoul, Republic of Korea
Seung-Bin Kim
Seung-Bin Kim
Department of Artificial Intelligence, Korea University, Seoul, Korea
Speech Synthesis
S
Seong-Whan Lee
Department of Artificial Intelligence, Korea University, Seoul, Republic of Korea