Vox-Infinity: Benchmarking the Limits of Long-Context Spoken Language Models

📅 2026-09-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入Vox-Infinity基准来解决长上下文语音模型理解能力不足的问题,该基准通过增加对话轮数和时长评估模型处理长音频输入的能力。
📝 Abstract
Long-context understanding remains a fundamental challenge for large language models, as excessively long inputs often lead models to forget salient information. This issue is even more pronounced in the speech domain, where audio, as a low-compression modality, requires substantially more embeddings than text to preserve both semantic content and acoustic cues. To address this challenge, we introduce \textbf{Vox-Infinity}, the first benchmark specifically designed to evaluate long-context understanding in spoken language models. Vox-Infinity systematically extends audio history along two dimensions: turn count and turn duration. It covers a diverse range of representative scenarios with varying interaction structures and semantic complexity. Crucially, Vox-Infinity provides explicit answer-provenance annotations and organizes samples according to the amount of historical context required to resolve each query, enabling precise and length-aware evaluation. Extensive evaluations of seven representative spoken language models reveal a clear overall recency effect: models generally achieve higher accuracy when answer-supporting evidence is closer to the query, but struggle to retrieve and use evidence located farther back in the dialogue history. Cases and datasets are available at https://vox-infinity.github.io.
Problem

Research questions and friction points this paper is trying to address.

long-context understanding
spoken language models
audio embeddings
semantic content
acoustic cues
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vox-Infinity
long-context understanding
spoken language models
audio history
answer-provenance annotations
🔎 Similar Papers
No similar papers found.
X
Xize Cheng
Zhejiang University
W
Wenxu Jia
Zhejiang University
C
Chenyuhao Wen
Zhejiang University
Dongjie Fu
Dongjie Fu
LREIS, Institute of Geographic Sciences and Natural Resources Research, Chinese Academy of Sciences
Remote Sensing
Zehan Wang
Zehan Wang
Zhejiang University
Multi-modal LearningComputer Vision
X
Xinyu Zhang
Zhejiang University
T
Tao Jin
Zhejiang University