Sample-Conditioned Representation Selection for Audio Few-Shot Learning

📅 2026-09-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文针对少样本音频分类中的前景-背景共现依赖问题,提出SAMPLESELECT方法,通过预测特征掩码来提高跨背景的泛化能力。
📝 Abstract
Few-shot audio classifiers may rely on foreground-background co-occurrences and fail when those correlations shift. On SpurAudio, the resulting representation shift is concentrated and class dependent: for ResNet12, the top 10 percent of channels explain 82.80 percent of the null-corrected shift contribution. We propose SAMPLESELECT, which predicts a fixed-budget feature mask independently for each input while keeping the encoder and source classifier frozen. Training uses differentiable Gumbel Top-k selection with foreground classification and cross-background contrastive losses; inference uses deterministic Top-k masks and support-only linear adaptation. Across ResNet12 and Conv64 in 5-way 1-shot and 5-shot evaluation, SAMPLESELECT gives the best OOD accuracy among the compared methods and improves the matched full-representation control by 4.90-8.38 percentage points. Ablations and representation analyses further support the learned selection mechanism. Code is available at https://github.com/Cross-Innovation-Lab/SAMPLESELECT/
Problem

Research questions and friction points this paper is trying to address.

few-shot learning
audio classification
representation shift
foreground-background co-occurrences
feature selection
Innovation

Methods, ideas, or system contributions that make the work stand out.

SAMPLESELECT
few-shot learning
audio classification
representation selection
Gumbel Top-k
🔎 Similar Papers
No similar papers found.