🤖 AI Summary
Existing deepfake detection methods rely heavily on artifacts left by generative models, leading to a significant drop in generalization when confronted with emerging generative architectures and interactive deception scenarios—such as video or voice impersonation—where the core threat lies in deceptive behavior rather than signal-level anomalies. This work breaks from conventional signal-centric paradigms by systematically integrating Speech Act Theory, Grice’s Cooperative Principle, and Cialdini’s Principles of Influence to construct a novel three-tiered analytical framework encompassing speech acts, dialogic interaction, and audience response. By introducing foundational social theories into media forensics, this framework not only exposes the “generalization illusion” inherent in current approaches but also establishes a new pathway for detecting deception in interactive deepfake contexts, while highlighting critical open challenges in the field.
📝 Abstract
For nearly a decade, deepfake detection has been framed as a classification task: given an audio or video clip, decide whether it is real or synthetic. Top detectors often report high accuracy on standard benchmarks; however, performance drops sharply on content from newer or unseen generators. We argue that better classifiers of synthetic media alone will not solve this problem, especially for interactive deepfakes such as impersonation in video and voice calls, where the harm lies not in the artifact (manipulated media signal) but in the act of deception. Deepfake detection therefore requires a complementary analytical layer focused on communicative interaction, not just media realism. We identify five assumptions that artifact-based detection (the forensic analysis of low-level signal traces) relies on and show that all five are eroding as generative models improve, producing what we call the Generalization Illusion. To address this, we draw on three well-established frameworks from philosophy of language and social psychology, namely, Speech Act Theory, Grice's Cooperative Principle, and Cialdini's principles of influence, to examine forensic signals at three levels: the utterance, the conversation, and the listener response. The result is a unified framework that complements existing forensic methods. We close with open problems for future work. https://jesseeho.github.io/deepfake-deception/