🤖 AI Summary
This study addresses a critical gap in the evaluation of AI-generated music detection methods, which are typically assessed on synthetic data and thus fail to reflect real-world broadcast conditions. To bridge this gap, the authors introduce BAMM, a novel dataset comprising 40 hours of authentic television broadcast recordings. They systematically evaluate a CNN-based detector’s ability to distinguish between AI-generated and human-composed music across three scenarios: Clean Foreground Music, Synthetic TV Broadcast, and Real TV Broadcast. Experimental results reveal a substantial performance drop in real broadcast settings, with significant score overlap between the two music types. This highlights a pronounced domain shift and underscores the inadequacy of current approaches for practical broadcast monitoring applications.
📝 Abstract
The proliferation of AI-generated music in broadcast media raises concerns about transparency and fair compensation, but reliable detection under real broadcast conditions remains unresolved. Existing studies report substantial performance degradation in this domain, yet their evaluations are limited to synthetic broadcast data. To address this gap, we introduce BAMM (Broadcast AI-Music Monitoring), a 40-hour dataset of real-world television recordings containing AI-generated and human-made music. We compare clean-trained and broadcast-trained CNN variants across three progressively more challenging scenarios: Clean Foreground Music (CFM), Synthetic TV Broadcast (STB), and Real TV Broadcast (RTB). Both models achieve near-perfect performance on CFM but degrade substantially under synthetic broadcast conditions. Broadcast-oriented training improves robustness compared with clean training, although performance remains limited. On RTB, evaluated using BAMM, both models degrade further and show substantial score overlap between AI-generated and human-made music. These results expose a critical domain gap and show that current training approaches on CNN-based detectors remain insufficient for reliable AI-generated music detection in broadcast monitoring.