🤖 AI Summary
This work addresses the challenge of achieving high-accuracy, training-free inference of contact-based keypad inputs in real-world settings, where pure WiFi-based sensing suffers from temporal ambiguity due to overlapping Channel State Information (CSI) waveforms caused by rapid keystrokes and network permission constraints. To overcome this limitation, the paper proposes AirKey, a cross-modal sensing framework that uniquely integrates acoustic signals with CSI extracted from WiFi ACK frames. By leveraging acoustic cues as precise temporal anchors to segment overlapped CSI traces, AirKey enables passive, zero-training PIN code inference without requiring network infrastructure modifications. Experimental results demonstrate that AirKey achieves over four times higher accuracy than existing zero-training, single-modality approaches in realistic environments, successfully recovering target PIN codes within six attempts.
📝 Abstract
Contactless keystroke inference via WiFi sensing highlights severe privacy threats, yet its real-world feasibility is hindered by two fundamental physical and deployment bottlenecks: the strict requirement for network privileges to acquire stable sensing streams, and the inherent "waveform fusion" ambiguity of pure WiFi signals during rapid, muscle-memory typing. To overcome these limitations, we propose AirKey, a novel cross-modal sensing framework that achieves highly stealthy, zero-training PIN eavesdropping. First, to bypass network deployment barriers, AirKey exploits fundamental IEEE 802.11 mechanisms to predictably elicit Acknowledgment (ACK) responses from unmodified target devices. By passively harvesting Channel State Information (CSI) from these ACKs using a low-cost microcontroller, AirKey secures a continuous spatial sensing stream entirely without network association. Crucially, to resolve the WiFi waveform fusion bottleneck, AirKey introduces a cross-modal complementarity mechanism. By utilizing lightweight acoustic signals as precise temporal anchors, the system robustly guides the segmentation of overlapping CSI trajectories. This joint spatiotemporal fusion strictly intersects CSI-derived spatial similarities with acoustic-guided inter-keystroke timing. Extensive real-world evaluations demonstrate that AirKey achieves over 4x higher accuracy than state-of-the-art unimodal zero-training schemes, successfully recovering device-unlock PINs within 6 attempts. Ultimately, this work exposes a critical vulnerability in contemporary smart interfaces, underscoring the severe privacy implications of ubiquitous multimodal sensing.