REVE: Efficient Hallucination Correction for Large Audio-Language Models via Reused Encoder States

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出REVE方法,利用已计算的编码器状态来校正大音频-语言模型中的幻觉事件,无需二次音频编码,有效减少了92.9%的标签不支持的提及。
📝 Abstract
Large audio-language models may mention acoustic events that are absent from the input. A separate audio event detector can verify these mentions, but doing so requires a second audio encoder and a separate forward pass. We propose Reused Encoder States for Verifying Events (REVE), a lightweight method that uses states already computed by the target model. One readout summarizes class scores across audio frames, while another uses pooled states from four consecutive frame intervals. Class-aware score fusion combines their outputs to verify generated event mentions without encoding the audio again. On AudioSet, REVE removes 92.9% of label-unsupported mentions under a faithful-mention recall constraint. With fewer added parameters and no second audio-encoding pass, REVE achieves a reduction comparable to those of CED-Tiny and CED-Base. Its complete verification latency is about 1/18 of the CED-Base path. Results on controlled DESED mixtures and different target-model architectures further confirm the effectiveness of encoder-state reuse.
Problem

Research questions and friction points this paper is trying to address.

audio-language models
hallucination correction
acoustic events
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reused Encoder States
Class-aware Score Fusion
Lightweight Method
Efficient Hallucination Correction
🔎 Similar Papers
No similar papers found.
H
Hongjin Song
Beijing Institute of Technology, Zhuhai, Guangdong, China
J
Jiasheng Kuang
Beijing Institute of Technology, Zhuhai, Guangdong, China
X
Xinyu Yang
Beijing Institute of Technology, Zhuhai, Guangdong, China
Q
Qiuyu Fang
Beijing Institute of Technology, Zhuhai, Guangdong, China
Ziyu Wu
Ziyu Wu
USTC
smart textile applicationspressure sensing mattresshuman pose and shape estimation
G
Guowu Tan
Beijing Institute of Technology, Zhuhai, Guangdong, China
Xiang Xie
Xiang Xie
PADO Labs
CryptographyPrivacy-Preserving Machine Learning