🤖 AI Summary
This study addresses the challenges of audio synthesizer sound matching, where target representations are difficult to optimize and direct search requires frequent, computationally expensive rendering. To overcome these bottlenecks, this work proposes a render-free search framework based on the Joint Embedding Predictive Architecture (JEPA). By learning mutually predictive representations of audio and parameters, it constructs an audio geometry shaped by parameter mappings, enabling candidate parameter evaluation directly within the embedding space rather than through costly rendering. The proposed method surpasses existing baselines in in-domain performance while demonstrating strong out-of-domain generalization. Furthermore, subjective listening tests reveal an 85% preference rate for the approach, which also supports flexible scaling that trades computational resources for improved synthesis quality.
📝 Abstract
Sound matching can be formulated as optimizing synthesizer parameters against an audio-domain objective. However, objectives derived from generic audio representations are often difficult to optimize, while direct search requires rendering every candidate. We introduce Synth-JEPA, which learns mutually predictive audio and parameter representations from paired synthesizer data. At inference, candidate parameters are scored directly in this learned space, yielding a renderer-free objective whose audio geometry is shaped by parameter correspondences rather than generic audio similarity. We evaluate Synth-JEPA on Surge XT using held-out synthesizer sounds and out-of-domain NSynth and FSD50K targets, against inverse models, direct search, and learned proxy objectives. Synth-JEPA outperforms all baselines in-domain and remains competitive out-of-domain. Its matching quality continues to improve with additional test-time search, allowing compute to be traded for match quality. In pairwise listening tests, listeners preferred Synth-JEPA in 85% of trials overall. Together, these results show that an audio representation with a parameter-induced geometry allows synthesizer sound matching to be approached as an effective renderer-free search problem.