🤖 AI Summary
This study addresses the challenge of plagiarism detection in AI-generated music, where conventional threshold-based methods fail due to generation-resynthesis pipelines. To overcome this limitation, this work pioneers a musical version identification paradigm tailored for AI plagiarism detection. By adapting state-of-the-art version identification architectures to human-AI plagiarism scenarios, it proposes a coordinate-shift-based supervision framework that effectively mitigates dispersed noise introduced during resynthesis through embedding space shift analysis. Furthermore, this research constructs COPYCAT, a benchmark comprising 350,000 sample pairs. Experimental results demonstrate a substantial improvement in the overall F0.5 score from 0.612 to 0.803, validating the effectiveness and robustness of the proposed supervised learning framework.
📝 Abstract
The rapid expansion of text-to-music generative models challenges traditional paradigms of music creation and intellectual property. Plagiarism in this context is rarely an absolute mathematical binary, but an ambiguous threshold negotiated over harmonic structure, melodic contours, or overall perceived stylistic character. In this work, we test the transferability of state-of-the-art music version identification architectures from the human-to-human cover domain to the human-to-AI plagiarism setting. To evaluate this task, we introduce COPYCAT, a benchmark derived from real-world plagiarism cases and extended through generative re-synthesis and digital signal processing obfuscations, yielding 350,654 evaluation pairs. We show that scalar distance thresholding collapses under generative re-synthesis, while a supervised framework leveraging coordinate-wise embedding shifts recovers the dispersed plagiarism signal, raising overall $F_{0.5}$ from $0.612$ to $0.803$.