🤖 AI Summary
This work addresses the susceptibility of existing open-vocabulary keyword spotting systems to false triggers caused by over-sensitivity to shared prefixes in spoken commands. The authors identify this issue as stemming from training data bias and positional bias in cross-modal scoring. To systematically evaluate prefix-induced confusion, they introduce the Partial Overlap Benchmark, comprising POB-Spark and POB-LibriPhrase datasets. They further propose a lightweight Equal-weighting Position Scoring (EPS) mechanism that uniformly attends to the entire speech sequence. Within an audio-text joint embedding framework, integrating EPS alone reduces the equal error rate (EER) on POB-Spark from 64.4% to 29.3% and achieves 96.8% accuracy on POB-LibriPhrase, while preserving performance on original benchmarks. Combining EPS with POB training data yields the best overall results.
📝 Abstract
Open-vocabulary keyword spotting (OV-KWS) enables personalized device control via arbitrary voice commands. Recently, researchers have explored using audio-text joint embeddings, allowing users to enroll phrases with text, and proposed techniques to disambiguate similar utterances. We find that existing OV-KWS solutions often overly bias the beginning phonemes of an enrollment, causing false triggers when negative enrollment-query-pairs share a prefix (``turn the volume up''vs. ``turn the volume down''). We trace this to two factors: training data bias and position-biased cross-modal scoring. To address these limitations, we introduce the Partial Overlap Benchmark (POB) with two datasets, POB-Spark and POB-LibriPhrase (POB-LP), containing mismatched audio-text pairs with shared prefixes, and propose Equal-weighting Position Scoring (EPS), a lightweight decision layer. Using EPS alone reduces EER on POB-Spark from 64.4\% to 29.3\% and improves POB-LP accuracy from 87.6\% to 96.8\%, while maintaining performance on LibriPhrase and Google Speech Commands (GSC). With POB data added in training, our work achieves the best POB benchmark results while incurring the least amount of degradation on prior metrics among baselines. This degradation is most pronounced in GSC, which contains only one-word commands. We surface mitigating this trade-off as future work.