FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow Matching

๐Ÿ“… 2025-01-09
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Audio super-resolution is an ill-posed inverse problem, where existing methods rely on multi-step sampling for high-fidelity reconstruction, resulting in low inference efficiency. This paper pioneers the application of flow matching to audio super-resolution, designing a spectrogram-aware customized probability path to enable single-step, end-to-end, high-fidelity reconstruction. The proposed method supports multiple input sampling rates and achieves state-of-the-art performance on the VCTK dataset with only one sampling stepโ€”attaining optimal Log-Spectral Distance and ViSQOL scores. Moreover, it accelerates inference by over an order of magnitude compared to mainstream diffusion-based approaches. The core contribution lies in establishing the first audio-adapted flow matching paradigm, which eliminates the iterative sampling bottleneck without compromising reconstruction quality.

Technology Category

Computer Vision: Diffusion Models for VisionSearch and Optimization: Sampling/Simulation-based SearchMachine Learning: Multimodal Learning

Application Category

Search and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesSystems and Infrastructure for Web, Mobile and WoT: Data management and stream processing for Web, mobile and wireless applicationsUser Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendation
๐Ÿ“ Abstract
Audio super-resolution is challenging owing to its ill-posed nature. Recently, the application of diffusion models in audio super-resolution has shown promising results in alleviating this challenge. However, diffusion-based models have limitations, primarily the necessity for numerous sampling steps, which causes significantly increased latency when synthesizing high-quality audio samples. In this paper, we propose FLowHigh, a novel approach that integrates flow matching, a highly efficient generative model, into audio super-resolution. We also explore probability paths specially tailored for audio super-resolution, which effectively capture high-resolution audio distributions, thereby enhancing reconstruction quality. The proposed method generates high-fidelity, high-resolution audio through a single-step sampling process across various input sampling rates. The experimental results on the VCTK benchmark dataset demonstrate that FLowHigh achieves state-of-the-art performance in audio super-resolution, as evaluated by log-spectral distance and ViSQOL while maintaining computational efficiency with only a single-step sampling process.
Problem

Research questions and friction points this paper is trying to address.

Audio Clarity Enhancement
High-Quality Audio Processing
Long Processing Time
Innovation

Methods, ideas, or system contributions that make the work stand out.

FLowHigh
Audio Clarity Enhancement
Efficient Sampling