🤖 AI Summary
This study reveals a practical privacy threat: user-typed text can be accurately reconstructed solely from keystroke acoustics, even without labeled data from the target device. To address this, the authors propose a self-supervised acoustic eavesdropping method that integrates unsupervised acoustic clustering, Transformer-based language model inference, and an iterative self-training mechanism, leveraging audio captured via smartphones or contact microphones. This approach achieves high-accuracy, cross-platform text reconstruction across diverse real-world scenarios—including remote meetings and through-wall eavesdropping—without requiring device-specific annotations. In close proximity, it attains 99% accuracy with only 100–150 keystrokes; under realistic noisy conditions such as at a distance of 3 meters, through walls, or during online calls, it consistently achieves over 90% reconstruction accuracy using 150–250 keystrokes.
📝 Abstract
We present a self-supervised acoustic eavesdropping attack that reconstructs typed text solely from keystroke sounds, without requiring labeled data for the target device. The proposed attack enables stealthy eavesdropping in two real-world scenarios-physical spaces (public and semi-public) and online meetings. Our method combines unsupervised acoustic clustering with Transformer-based language model inference and iterative self-training, enabling stable character inference under highly uncertain acoustic-to-character mappings. We demonstrate that the proposed method achieves over 99% reconstruction accuracy with only 100-150 observed keystrokes under a close-proximity recording setup using a smartphone placed near the target device, significantly outperforming prior unsupervised baselines in low-data regimes. We further evaluate robustness across multiple laptop platforms and in realistic acquisition channels, including distance recording from approximately 3 meters away on the same desk, through-the-wall eavesdropping with a contact microphone, and background keyboard noise in online conferencing systems. Across these scenarios, the proposed method achieves high reconstruction accuracy (often exceeding 90%) with approximately 150-250 observed keystrokes. These results indicate that accurate text reconstruction from keystroke sounds is feasible in practice under an audio-only setting, even with limited observed keystrokes and without requiring device-specific labeled data, highlighting a realistic and previously underestimated privacy risk.