Don't CLAP: Are Music-Text Models Bag-of-Words?

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether audio-text matching scores, such as those produced by CLAP, can accurately capture fine-grained attribute binding between music and text. To this end, it proposes an attribute-swapping perturbation test that systematically permutes musical attributes to evaluate the representational capacity of both contrastive learning models and large audio language models. The findings reveal that existing metrics are insensitive to semantic variations, with their representations degenerating into bag-of-words patterns that over-rely on linguistic priors rather than capturing nuanced musical semantics. By exposing these fundamental limitations, this work challenges the validity of prevailing evaluation paradigms in the field and provides critical empirical evidence for developing more robust and reliable audio-language assessment frameworks.
📝 Abstract
Text-to-music systems are assessed on audio quality and on how faithfully the music follows its prompt, and the CLAP score, the cosine similarity between a music-text model's audio and text embeddings, is the standard objective metric of faithfulness. We ask how accurately that score reflects the text: when an attribute is linked to an instrument (e.g., distorted guitar), does the text embedding capture that binding? To find out, we introduce an attribute swap perturbation: the caption of a real recording is edited by exchanging exactly one property, timbre, lead versus accompaniment, or order of first appearance, between two instruments. We then test four contrastive music-text models and one large audio-language model on whether the audio scores higher against the original caption than against the perturbed one. No contrastive model distinguishes the two captions reliably. The audio-language model does better, but further experiments show that its advantage rests largely on audio-agnostic language priors. Our results thus provide compelling evidence that the CLAP score and related metrics do not capture fine-grained musical meaning or attribute bindings; their representation is closer to a bag-of-words that leaves them insensitive to meaning-changing perturbations of the caption.
Problem

Research questions and friction points this paper is trying to address.

CLAP score
music-text models
attribute binding
bag-of-words
text-to-music evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Attribute Swap Perturbation
CLAP Score
Music-Text Models
Bag-of-Words
Audio-Language Model
🔎 Similar Papers
No similar papers found.