🤖 AI Summary
This study addresses the significant degradation in activity recognition performance caused by acoustic cue loss in low-resolution audio due to reduced sampling rates, compounded by sensor resolution mismatches between training and deployment. To overcome this, we propose a resolution-aware transfer framework that leverages high-resolution audio as privileged information to facilitate model learning. Specifically, the method achieves efficient representation compression and transfer from high- to low-resolution domains through privileged knowledge distillation, neighborhood structure preservation, and token-level local feature alignment. Experimental evaluations on the SAMoSA and AudioIMU datasets demonstrate that the proposed approach improves recognition accuracy by approximately 7.8% over baseline methods. Notably, the framework relies exclusively on low-resolution audio during inference, effectively resolving the challenge of cross-resolution deployment in practical applications.
📝 Abstract
Audio is increasingly used for human activity recognition (HAR) because it captures object interactions, environmental events, and contextual cues in everyday environments. High-resolution (HR) audio provides rich acoustic information for model development but incurs substantial energy and storage costs and may expose sensitive speech content. Low-resolution (LR) audio offers a more privacy-preserving and resource-efficient alternative for deployment, but reduced sampling rates can remove acoustic cues essential for activity recognition, leading to significant performance degradation. We formulate this training-deployment mismatch as sensor-resolution privileged learning, in which HR audio is available during training, while inference relies exclusively on LR audio. We propose RAST, a resolution-aware transfer framework that compresses HR teacher representations by preserving token-level information and neighborhood structure before performing localized HR-LR alignment. Experiments on the SAMoSA and AudioIMU datasets show that RAST consistently outperforms LR-only training and direct teacher-transfer baselines, improving LR-only recognition by up to approximately 7.8% while requiring only LR audio at inference.