RAST: Resolution-Aware Privileged Structure Transfer for Low-Resolution Audio Activity Recognition

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant degradation in activity recognition performance caused by acoustic cue loss in low-resolution audio due to reduced sampling rates, compounded by sensor resolution mismatches between training and deployment. To overcome this, we propose a resolution-aware transfer framework that leverages high-resolution audio as privileged information to facilitate model learning. Specifically, the method achieves efficient representation compression and transfer from high- to low-resolution domains through privileged knowledge distillation, neighborhood structure preservation, and token-level local feature alignment. Experimental evaluations on the SAMoSA and AudioIMU datasets demonstrate that the proposed approach improves recognition accuracy by approximately 7.8% over baseline methods. Notably, the framework relies exclusively on low-resolution audio during inference, effectively resolving the challenge of cross-resolution deployment in practical applications.
📝 Abstract
Audio is increasingly used for human activity recognition (HAR) because it captures object interactions, environmental events, and contextual cues in everyday environments. High-resolution (HR) audio provides rich acoustic information for model development but incurs substantial energy and storage costs and may expose sensitive speech content. Low-resolution (LR) audio offers a more privacy-preserving and resource-efficient alternative for deployment, but reduced sampling rates can remove acoustic cues essential for activity recognition, leading to significant performance degradation. We formulate this training-deployment mismatch as sensor-resolution privileged learning, in which HR audio is available during training, while inference relies exclusively on LR audio. We propose RAST, a resolution-aware transfer framework that compresses HR teacher representations by preserving token-level information and neighborhood structure before performing localized HR-LR alignment. Experiments on the SAMoSA and AudioIMU datasets show that RAST consistently outperforms LR-only training and direct teacher-transfer baselines, improving LR-only recognition by up to approximately 7.8% while requiring only LR audio at inference.
Problem

Research questions and friction points this paper is trying to address.

Human Activity Recognition
Low-Resolution Audio
Privileged Learning
Training-Deployment Mismatch
Innovation

Methods, ideas, or system contributions that make the work stand out.

Privileged Learning
Resolution-Aware Transfer
Low-Resolution Audio
Activity Recognition
Knowledge Distillation
🔎 Similar Papers
No similar papers found.