🤖 AI Summary
This study addresses the temporal collapse in traditional audio watermarking caused by temporal evidence aggregation, which constrains both robustness and multi-payload capacity. We propose an optimization framework that eliminates the need to retrain foundation models: a pretrained encoder and detector remain frozen while only a low-latency Conformer decoder is fine-tuned to optimize temporal soft outputs, complemented by system-specific adapters to overcome performance bottlenecks. Experimental results demonstrate that this approach substantially enhances resilience against adversarial attacks, yielding improvements of 6.9 to 48.0 percentage points in detection AUROC and dual-payload joint recovery rates. Ultimately, this work achieves synergistic gains in robustness and payload capacity through highly parameter-efficient fine-tuning.
📝 Abstract
Modern neural audio watermarking systems typically embed a message repeatedly across time and then collapse the resulting temporal evidence into a single payload using averaging, voting, or another fixed aggregation rule. We argue that this temporal collapse limits both robustness and the recovery of multiple payloads, and that the limitation can be addressed without retraining the underlying watermarker. We freeze a pretrained watermarker's encoder and detector and train only a low-latency Conformer-based decoder. The decoder consumes the detector's temporal soft outputs, which a system-specific adapter pools into a sequence of window-level representations, and predicts the embedded message. On three frozen watermarkers (AURA, AudioSeal, and WavMark), the learned decoder improves recovery of attacked messages and yields higher detection AUROC point estimates on all three. Under controlled full- and partial-coverage multiplexing, it improves joint-exact recovery of two alternating payload words by 9.7-48.0, 6.9-17.3, and 8.2-14.2 percentage points, respectively, under one to three chained attacks on feasible clips.