🤖 AI Summary
This study addresses the significant degradation in target perception capabilities of large audio models under noisy interference. Inspired by human long-term memory, it proposes LTM-AE, a training-free inference-time optimization framework. This method extracts hidden states from clean reference audio to construct a long-term memory representation, which is integrated with token reconstruction and interpolation techniques alongside an optional gating mechanism. By keeping all model parameters frozen, the framework effectively guides the enhancement of specific target sound representations. To our knowledge, this work represents the first introduction of long-term memory principles into audio enhancement. Experimental results demonstrate substantial improvements, increasing multi-class classification accuracy by 29.53%–46.15% and reducing the word error rate of Qwen2-Audio for speech recognition from 23.07% to 14.77%.
📝 Abstract
Audio large language models (ALLMs) can reason about the content of audio recordings to perform complex tasks. However, these capabilities usually collapse in real-world environments when background noise and competing sources mix the target sound. Inspired by long-term memory in human listening, we propose Long-Term Memory-Guided Audio Enhancement (LTM-AE) to improve selective target perception by refining the audio representations of ALLMs without training. LTM-AE extracts representations in hidden states from separate clean reference recordings as long-term memory for each category, guiding enhancement toward a user-specified listening target. We reconstruct incoming audio tokens in the selected category long-term memory and interpolate the reconstructions with the original tokens before language backbone decoding. This interpolation controls the influence of stored auditory experience while keeping all ALLM parameters fixed. Diagnostic readouts across twenty sound categories and three ALLMs show that LTM-AE strengthens responses to a specified target amid three interfering sources. Averaged over constrained and free-form classification, accuracy gains over raw mixtures range from 29.53 to 46.15 percentage points across multiple open source models. For speech content recovery, LTM-AE with an additional learned token-level gate reduces Qwen2-Audio's word error rate from 23.07% to 14.77%. This work takes an initial step toward using principles of human long-term memory to enhance ALLMs for real-world listening. Our code is available at https://github.com/aynlp/ltm-audio-code