Improving the performance of an ASV system using hybrid speech features

📅 2026-07-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the significant performance degradation of automatic speaker verification (ASV) systems under noisy conditions and adversarial attacks. To enhance robustness, the study proposes a novel multimodal feature fusion approach that integrates a newly designed RAB descriptor with conventional acoustic features—namely MFCC, CQCC, and PNCC—to construct a hybrid feature set. Experimental evaluation on the Google Speech Commands dataset demonstrates that the proposed PNCC+RAB combination substantially reduces the equal error rate (EER) in noisy environments, thereby markedly improving both the accuracy and robustness of ASV systems.
📝 Abstract
The growing need for secure and convenient authentication methods has led to the increasing popularity of biometric solutions. In addition to traditional and popular methods, such as fingerprint or iris scanning, voice-based approaches are also employed. User identity verification based on voice is conducted using Automatic Speaker Verification (ASV) systems. Despite their many advantages, these systems are sensitive to various types of attacks and acoustic noises, which can reduce verification accuracy. This work examines the potential to improve the performance of ASV systems by using hybrid feature sets that combine different signal representations, starting with widely-used Mel-Frequency Cepstral Coefficients (MFCC), through Constant Q Cepstral Coefficients (CQCC) and ending with the innovative RAB descriptor. Experiments were conducted on recordings from the Google Speech Commands dataset under two scenarios: in clean conditions and in the presence of acoustic noise. Finally, the systems' performance was compared using the EER metric to determine whether hybrid feature sets decrease verification error. The results show that using a hybrid feature set (PNCC+RAB) improves speaker verification performance under noisy conditions.
Problem

Research questions and friction points this paper is trying to address.

Automatic Speaker Verification
acoustic noise
verification accuracy
biometric authentication
speaker verification performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

hybrid speech features
Automatic Speaker Verification
RAB descriptor
noise robustness
EER
🔎 Similar Papers
No similar papers found.