🤖 AI Summary
This study addresses the loss of parent-child topological relationships among scattering paths caused by coefficient flattening in wavelet scattering-based deepfake speech detection, which limits the exploitation of forensic cues. To overcome this, we propose explicitly restoring such topology within the wavelet scattering front-end for the first time by restructuring scattering paths into a sparse modulated-carrier grid. Combined with length-aware adaptive local attention pooling, this constructs a waveform-to-graph interface that preserves acoustic-axis characteristics, enabling the effective integration of parameter-free fixed interfaces with physical features. The proposed method maintains competitive performance with AASIST while reducing trainable parameters by approximately 60%, and achieves significant performance improvements on cross-domain benchmarks.
📝 Abstract
The acoustic front-end determines which forensic cues a speech deepfake detector can exploit. The wavelet scattering transform (WST) provides stable multiscale coefficients with explicit coordinates, yet direct flattening obscures the parent relation between paths. We introduce WST-Graph, reconstructing these paths as a sparse modulation-carrier grid for an AASIST graph backend. Modulation-level normalization and length-aware adaptive local attention pooling produce fixed relative-time representations while retaining the acoustic axes before learned adaptation. This yields a waveform-to-graph interface with a fixed, parameter-free WST. Our configurations remain competitive with AASIST while using approximately 60% fewer trainable parameters and show clear gains on selected out-of-domain benchmarks. These results underscore the value of preserving parent-child relations within the carrier-modulation topology when constructing a compact, physically grounded interface for graph-based speech deepfake detection. Code will be released at https://github.com/saki-ciallo/wst-graph.