🤖 AI Summary
This study addresses the limitations of conventional 3D Morphable Model (3DMM) approaches in person-of-interest (POI) deepfake detection, which typically rely on full coefficient sets while neglecting their inherent redundancy and surface-level information. To this end, we propose MOTIF, a novel detection framework that, for the first time, quantifies the contributions of individual 3DMM coefficient groups, revealing the dominant role of shape parameters and the complementary value of dense surfaces. Based on these insights, MOTIF extracts purely visual features fused with surface information to construct a detector trained exclusively on real videos, eliminating the need for forged samples or POI-specific data. Extensive experiments demonstrate that MOTIF comprehensively outperforms existing state-of-the-art methods across multiple datasets and varying quality benchmarks, significantly enhancing generalization against diverse manipulation techniques.
📝 Abstract
Video deepfakes targeting a specific individual, the Person-of-Interest (POI), are the most harmful ones, and, since a public figure is abundantly recorded, a detector can be built from genuine footage of that individual. Such detectors commonly describe a subject through a 3D Morphable Model (3DMM) and adopt its coefficients as a whole, so which part of that description carries the signal has never been measured. We dissect it, holding the encoder, the training corpus and the enrollment protocol fixed and varying only what the encoder observes. The groups of coefficients prove largely redundant, since the shape block alone recovers almost all the accuracy of the full vector, and their temporal evolution contributes a real but bounded amount. We further show that the dense surface the same fit returns, which these detectors discard, carries identity information that the coefficients do not, and that it helps precisely where they are weakest. We assemble the best configuration into MOTIF, a visual-only detector trained on real videos only, with no manipulated video and no POI-specific data. It improves on both state-of-the-art POI detectors in every dataset and manipulation of our benchmark and at two quality levels. Our experimental code will be released at https://github.com/polimi-ispl/MOTIF.