"What I See is What I Hear": Deepfake Detection Across Diverse Hearing Abilities

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the overlooked disparities in audiovisual perception and asymmetric risks faced by individuals with hearing impairments in existing deepfake detection research. Employing a mixed-methods approach, we recruited 80 participants across varying hearing levels to systematically evaluate their ability to detect manipulated videos using multimodal stimuli, including speech synthesis and voice conversion. This work is the first to quantify differential sensitivities to deepfakes among distinct hearing groups. Results reveal that hard-of-hearing participants exhibit significantly lower overall detection accuracy than normal-hearing individuals, primarily driven by elevated false positive rates induced by audio-channel manipulations, with deaf participants performing poorest under purely audio-based manipulations. These findings underscore the urgent need for accessible defense mechanisms and provide empirical evidence for developing inclusive deepfake detection frameworks.
📝 Abstract
The proliferation of audiovisual deepfakes has lowered the cost of fraud, impersonation, and misinformation, but their success ultimately depends on human perception. Detection requires integrating auditory and visual cues, yet security and privacy research has largely overlooked d/Deaf and hard-of-hearing (DHH) populations. We address this gap with an in-person, mixed-methods study of 80 participants: 31 hearing persons (HPs), 15 hard-of-hearing (HoH) participants, 17 d/Deaf participants, and 17 cochlear implant (CI) users. Each participant judged the authenticity of 30 clips, where manipulations spanned text-to-speech, voice conversion, lip-sync, or face-swap. DHH participants were less accurate than HPs overall (76.4% vs. 88.0%, p<.001), primarily because they more often classified authentic clips as manipulated (FPR: 29.7% vs. 11.2%). Differences depended strongly on the manipulated channel. For audio-only manipulations, HoH participants matched HPs (90.0% vs. 90.3%), followed by CI users (79.4%) and d/Deaf participants (41.2%). When clips contained an audiovisual manipulation, accuracy clustered between 84% and 87%, although performance still varied by manipulation method. Our work systematically characterizes how deepfakes affect DHH populations, highlighting the asymmetric risks audiovisual manipulations may pose to groups with different hearing abilities and the need for accessible, tailored defenses that support all users.
Problem

Research questions and friction points this paper is trying to address.

Deepfake detection
Hearing impairment
Audiovisual manipulation
Accessibility
DHH population
Innovation

Methods, ideas, or system contributions that make the work stand out.

Deepfake Detection
Audiovisual Manipulation
Deaf and Hard-of-Hearing (DHH)
Human Perception
Accessible Security
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Magdalena Pasternak
University of Florida
M
Malvika Jadhav
University of Florida
P
Palavi V. Bhole
Rochester Institute of Technology
A
Aviva Smith
University of Florida
E
Elaina Trapatsos
Rochester Institute of Technology
Vincent Bindschaedler
Vincent Bindschaedler
Assistant Professor at the University of Florida
PrivacySecurityApplied Cryptography
R
Roshan Peiris
Rochester Institute of Technology
Ersin Uzun
Ersin Uzun
Rochester Institute of Technology
Patrick Traynor
Patrick Traynor
University of Florida
Network SecurityMobile Networks
M
Matthew Wright
Rochester Institute of Technology
K
Kevin R. B. Butler
University of Florida