🤖 AI Summary
This study addresses the absence of unified attackability evaluation standards for speech quality metrics used as rewards, alongside vulnerabilities in existing methods to processing chain shifts and single-attack randomness. We propose a revised evaluation protocol referencing unperturbed round trips and introduce a multi-random-seed worst-case assessment mechanism. Benchmarking is conducted by integrating neural audio codecs, adversarial training, and closed-loop attack-defense auditing. Results reveal substantial disparities in attack success rates across metrics such as NISQA (6%–90%). Furthermore, we demonstrate that patch-based defenses are effective only within specific attack spaces and incur out-of-domain performance degradation, underperforming random perturbation baselines. Nevertheless, augmenters can effectively enhance the robustness of patched metrics.
📝 Abstract
Speech quality predictors are increasingly used as rewards, yet no agreed measure of their hackability exists. The usual measurement has two flaws. First, the perturbation reaches the predictor through a processing chain -- here a neural codec -- that shifts the score on its own, which scoring against the raw input charges to the attack. Referencing the unperturbed round trip instead changes measured hackability by up to a factor of four (0.31 to 0.08 for one defence). Second, one trained attacker is a sample, not a measurement: five attackers differing only in random seed reach success rates from 0.00 to 0.38 against one fixed predictor, so a defence claim needs the worst case over several. Under this protocol, four published predictors differ widely: NISQA is hacked on 90% of utterances, SSL-MOS on 21%, DNSMOS on 14% and UTMOS on 6%. We then audit a closed attack-detect-patch loop. It hardens the predictor only in its own attack space, by less than the spread between attackers; a random-perturbation baseline matches it; and it costs up to 0.30 system SRCC out of domain. Enhancers post-trained against patched predictors hack them far less (PESQ -0.03 versus -0.23). Code, preregistration and run outputs are released.