AudioJev: Direct Audio Decisions with Order-Calibrated Probabilities

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the susceptibility of existing audio decision models to option-order interference and acoustic information loss during transcription. We propose an end-to-end, fully fine-tuned multimodal framework that directly maps waveforms, questions, and options to candidate distributions. The core innovation is an "order calibration" mechanism employing random derangement and symmetric KL-divergence training to ensure stable probability outputs under semantically invariant permutations, eliminating the need for additional calibration heads. Experiments demonstrate that this approach achieves 68.88% and 55.33% accuracy on the MMAU and MMAR benchmarks, respectively, while reducing performance variance caused by randomized ordering by 43.8% and 60.5%. These results confirm that the proposed method effectively balances decision accuracy with robustness against positional bias.
📝 Abstract
Audio decisions often depend on evidence that a transcript does not preserve, while their probability estimates can depend on how answer options are ordered. AudioJev maps a waveform, question and supplied alternatives directly to a candidate distribution through one shared full-parameter model. We define order calibration as preserving an answer's probability under meaning-preserving option permutations. Random-derangement SKL training pairs each question with a reordered view in which every alternative changes position, supervises both answers, and aligns the two distributions before applying a symmetric KL penalty. Inference retains a single candidate-scoring forward, with no calibration head or order ensemble. Across three training seeds, AudioJev reaches 68.88%/55.33% mean accuracy on complete MMAU/MMAR and reduces random-order SKL by 43.8%/60.5% relative to the single-view removal ablation. The same model handles intent, environmental sound, note properties, speech activity and conversational transitions. Paired removal ablations and multi-order evaluation measure predictive accuracy and probability stability together, establishing a direct audio interface whose calibration objective acts on candidate meaning rather than presentation position. Inference code and model weights are available at https://github.com/SihanLv/AudioJev-Inference and https://huggingface.co/shlv/AudioJev.
Problem

Research questions and friction points this paper is trying to address.

audio decision-making
order calibration
probability estimation
multiple-choice audio reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Direct Audio Decisions
Order Calibration
Random-Derangement SKL
Symmetric KL Penalty
Shared Full-Parameter Model
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Sihan Lv
Zhejiang University, Hangzhou, China
Zhen Li
Zhen Li
China Mobile Information Technology Center; Peking University
Z
Zhiqi Cao
Innovation and Management Center of the School of Software (Ningbo), Zhejiang University, Ningbo, China
J
Jinshan Zhang
Zhejiang University, Hangzhou, China
Y
Ying Li
Zhejiang University, Hangzhou, China
Meng Xi
Meng Xi
College of Computer Science and Technology, Zhejiang University
service computingservice patterndata miningartificial intelligence
Jianwei Yin
Jianwei Yin
Professor of Computer Science and Technology, Zhejiang University
Service ComputingComputer ArchitectureDistributed ComputingAI