A Stem-Agnostic Approach to Hybrid AI Music Detection
为解决混合音乐中AI生成音频检测问题,提出了一种基于inspectrogram和Wiener滤波器的框架,通过CNN模型评估音轨是否由AI生成。
为解决混合音乐中AI生成音频检测问题,提出了一种基于inspectrogram和Wiener滤波器的框架,通过CNN模型评估音轨是否由AI生成。
本文提出CoJEPA方法,结合对比学习和JEPA解决音乐表示中的全局-局部问题,通过共享骨干网络联合训练以获得更丰富的音乐表示。
This study presents the first exploration into the detectability of AI-generated tracks within music co-created by humans and artificial intelligence. Addressing the challenge that general-purpose source separation methods often fail to reliably recover AI-specific artifacts, the authors propose a parallel detection architecture that operates without requiring full track separation. The approach leverages a neural audio codec to simulate the mixing process and combines short-time audio block analysis, relative energy estimation, and a binary classifier to directly identify AI-generated components within the mixed audio signal. Experiments on the MUSDB18-HQ dataset demonstrate promising track-level detection performance, confirming the feasibility of effectively discerning AI-generated content even in complex musical mixtures.
This work addresses the regulatory challenges posed by the proliferation of AI-generated music from unknown models by proposing an unsupervised, zero-shot detection framework that effectively distinguishes authentic from synthetic audio without relying on prior knowledge of generative models. The method integrates artifact-based feature extraction, non-negative matrix factorization (NMF), and a combination of one-class classification with unsupervised clustering strategies, thereby introducing zero-shot learning to AI music detection for the first time in a systematic manner. Experimental results demonstrate that the approach achieves strong performance in both binary authenticity discrimination and multi-class clustering of unseen AI-generated music sources, making it well-suited for monitoring high-purity synthetic content at scale within large music repositories.
This study addresses the lack of efficient, scalable, and human-aligned automatic evaluation methods for natural language responses in conversational music recommendation systems. It presents the first empirical investigation into the alignment between large language models employed as judges (LLM-as-a-Judge) and human experts, specifically along the dimensions of personalization and explanation quality. Through multi-turn dialogue sampling, generation of responses via four types of instruction-tuned models, expert human ratings, and bootstrap-based correlation analysis, the work demonstrates that LLM judges exhibit moderate positive correlation with human assessments—significantly outperforming conventional baselines. The findings further reveal the impact of model scale and contextual information on judging performance, offering practical guidance for model selection and deployment conditions in real-world applications.
为解决混合音乐中AI生成音频检测问题,提出了一种基于inspectrogram和Wiener滤波器的框架,通过CNN模型评估音轨是否由AI生成。
本文提出CoJEPA方法,结合对比学习和JEPA解决音乐表示中的全局-局部问题,通过共享骨干网络联合训练以获得更丰富的音乐表示。
This study presents the first exploration into the detectability of AI-generated tracks within music co-created by humans and artificial intelligence. Addressing the challenge that general-purpose source separation methods often fail to reliably recover AI-specific artifacts, the authors propose a parallel detection architecture that operates without requiring full track separation. The approach leverages a neural audio codec to simulate the mixing process and combines short-time audio block analysis, relative energy estimation, and a binary classifier to directly identify AI-generated components within the mixed audio signal. Experiments on the MUSDB18-HQ dataset demonstrate promising track-level detection performance, confirming the feasibility of effectively discerning AI-generated content even in complex musical mixtures.
This work addresses the regulatory challenges posed by the proliferation of AI-generated music from unknown models by proposing an unsupervised, zero-shot detection framework that effectively distinguishes authentic from synthetic audio without relying on prior knowledge of generative models. The method integrates artifact-based feature extraction, non-negative matrix factorization (NMF), and a combination of one-class classification with unsupervised clustering strategies, thereby introducing zero-shot learning to AI music detection for the first time in a systematic manner. Experimental results demonstrate that the approach achieves strong performance in both binary authenticity discrimination and multi-class clustering of unseen AI-generated music sources, making it well-suited for monitoring high-purity synthetic content at scale within large music repositories.
This study addresses the lack of efficient, scalable, and human-aligned automatic evaluation methods for natural language responses in conversational music recommendation systems. It presents the first empirical investigation into the alignment between large language models employed as judges (LLM-as-a-Judge) and human experts, specifically along the dimensions of personalization and explanation quality. Through multi-turn dialogue sampling, generation of responses via four types of instruction-tuned models, expert human ratings, and bootstrap-based correlation analysis, the work demonstrates that LLM judges exhibit moderate positive correlation with human assessments—significantly outperforming conventional baselines. The findings further reveal the impact of model scale and contextual information on judging performance, offering practical guidance for model selection and deployment conditions in real-world applications.