Camera-Noise Residuals for Face-Swap Detection: Redundant, Not Complementary, and Why

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether camera noise residuals and RGB features exhibit complementarity in face forgery detection. Leveraging the FaceForensics++ dataset and an Xception backbone integrated with a Noiseprint++ noise branch, we conduct multi-model ablation experiments to identify performance bottlenecks. Our analysis reveals that Instance Normalization layers suppress the discriminative capacity of noise features by eliminating their statistical properties. Consequently, affected by this bottleneck, noise and RGB features prove redundant rather than complementary. Although removing these normalization layers recovers the noise signal, the resulting fusion approach still fails to surpass the pure RGB baseline. These findings clearly delineate the application limitations of noise features within such architectures for deepfake detection.
📝 Abstract
Fusing a learned camera-noise fingerprint with an RGB appearance backbone is an appealing route to generator-independent deepfake detection, because the noise residual is grounded in image-formation physics rather than in the texture statistics of a particular generator. We test, on FaceForensics++, whether a Noiseprint++ residual channel carries information \emph{complementary} to an RGB Xception backbone for face-swap detection. A three-model ablation (RGB-only, residual-only, late-fusion) shows that fusion does not improve over RGB alone and that the residual branch alone is near chance. A seven-level bottleneck diagnostic localizes the cause: the noise maps do carry a discriminative signal, but it is statistical---carried by the per-sample first and second moments (mean, variance, energy) of the residual---and the per-sample \texttt{InstanceNorm} layer placed at the noise-branch input, following the TruFor template, standardizes exactly those moments away (five-fold cross-validated AUC drops from $0.747$ to $0.554$). A context-crop control rules out cropping geometry, and two fixed-fusion variants that remove the bottleneck recover the statistical signal yet still fail to beat RGB on every dataset. We conclude that, on this manipulation distribution, the noise residual is redundant with RGB rather than complementary, and we give concrete guidance for practitioners adopting noise-residual fusion for face-swap detection.
Problem

Research questions and friction points this paper is trying to address.

face-swap detection
camera-noise residuals
deepfake detection
feature fusion
InstanceNorm
Innovation

Methods, ideas, or system contributions that make the work stand out.

Face-Swap Detection
Camera-Noise Residuals
Information Redundancy
Bottleneck Diagnostics
InstanceNorm