π€ AI Summary
This study addresses the vulnerability of watermarks to codec attacks and the challenge of individual-level tracing in autoregressive audio generation. We propose MARC, a multi-bit watermarking method whose core innovation lies in constructing a token clustering space that integrates intrinsic relationships with retokenization confusion patterns. This establishes a perceptual domain supporting multi-bit payload embedding, thereby overcoming the limitations of conventional zero-bit detection. Furthermore, MARC enhances robustness through payload-driven cluster scheduling and multi-codec adversarial training. Experimental results demonstrate that the proposed method achieves an average bit extraction accuracy of 97.3% on unmodified audio while maintaining resilience against diverse codec and overlay attacks.
π Abstract
Generated audio is now used in a range of applications, creating a need to verify its origin after distribution and signal processing. This task is particularly challenging for autoregressive audio generation because codec processing can alter the token sequence recovered from the waveform. Such changes reduce the reliability of watermark detection and payload decoding. Existing methods construct token groups using either intrinsic token representations or substitution patterns caused by transformations. As a result, intrinsic token relationships and codec induced substitutions are modeled separately. In addition, most methods support only zero bit detection. They can determine whether a watermark is present but cannot distinguish individual generated outputs. We propose \textbf{MARC}, a multi-bit generative watermarking method for autoregressive audio generation. MARC integrates intrinsic token representations with confusion patterns obtained through retokenization and multiple codecs, forming a codec-aware token-cluster space. Within this space, payload-driven cluster scheduling is used to embed a multi-bit watermark, while detection and payload decoding are performed on retokenized observations. Experiments on speech, dialogue, and music generation show that MARC achieves an average of 97.3\% bit extraction accuracy on unmodified watermarked audio and the watermark can still be extracted under diverse codec attacks. MARC also demonstrates robustness to overwriting attacks.