Comparative study of state-based neural networks for virtual analog audio effects modeling

📅 2024-05-07
🏛️ EURASIP Journal on Audio, Speech, and Music Processing
📈 Citations: 3
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of dynamic parameter control and real-time performance in virtual analog audio effect modeling. We systematically compare three stateful neural architectures—State-Space models, Linear Recurrent Units (LRUs), and LSTMs (including encoder-decoder variants)—for modeling distortion, equalization, saturation, and compression. To enable fine-grained, differentiable parameter conditioning, we innovatively integrate Feature-wise Linear Modulation (FiLM). Furthermore, we propose a joint time-frequency evaluation metric that jointly assesses transient response, energy consistency, and spectral fidelity. Experimental results show that LSTMs achieve superior performance in distortion and EQ modeling; their encoder-decoder variants and State-Space models better capture the nonlinear dynamics of saturation and compression; all models fail to accurately model low-pass filter characteristics; and LRUs exhibit poor generalization stability.

Technology Category

Machine Learning: Deep Generative Models & AutoencodersCognitive Modeling & Cognitive Systems: Neural Spike CodingComputer Vision: Diffusion Models for Vision

Application Category

User Modeling, Personalization and Recommendation: User modeling and simulation for interactive and conversational systemsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Artificial neural networks are a promising technique for virtual analog modeling, having shown particular success in emulating distortion circuits. Despite their potential, enhancements are needed to enable effect parameters to influence the network’s response and to achieve a low-latency output. While hybrid solutions, which incorporate both analytical and black-box techniques, offer certain advantages, black-box approaches, such as neural networks, can be preferable in contexts where rapid deployment, simplicity, or adaptability are required, and where understanding the internal mechanisms of the system is less critical. In this article, we explore the application of recent machine learning advancements for virtual analog modeling. In particular, we compare State-Space models and Linear Recurrent Units against the more common Long Short-Term Memory networks, with a variety of audio effects. We evaluate the performance and limitations of these models using multiple metrics, providing insights for future research and development. Our metrics aim to assess the models’ ability to accurately replicate the signal’s energy and frequency contents, with a particular focus on transients. The Feature-wise Linear Modulation method is employed to incorporate effect parameters that influence the network’s response, enabling dynamic adaptability based on specified conditions. Experimental results suggest that Long Short-Term Memory networks offer an advantage in emulating distortions and equalizers, although performance differences are sometimes subtle yet statistically significant. On the other hand, encoder-decoder configurations of Long Short-Term Memory networks and State-Space models excel in modeling saturation and compression, effectively managing the dynamic aspects inherent in these effects. However, no models effectively emulate the low-pass filter, and Linear Recurrent Units show inconsistent performance across various audio effects.
Problem

Research questions and friction points this paper is trying to address.

Enhancing neural networks for dynamic parameter control in audio effects
Comparing State-Space, Linear Recurrent Units, and LSTM for low-latency modeling
Evaluating accuracy in replicating signal energy and frequency content
Innovation

Methods, ideas, or system contributions that make the work stand out.

State-Space models for dynamic audio effects
Feature-wise Linear Modulation for adaptability
LSTM networks excel in distortion emulation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
University of Oslo
R
Riccardo Simionato
Department of Musicology, University of Oslo, Problemveien 11, Oslo, 0313, Norway
S
Stefano Fasciani
Department of Musicology, University of Oslo, Problemveien 11, Oslo, 0313, Norway