Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding

📅 2026-07-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing Chinese multi-turn dialogue benchmarks, which inadequately assess models’ deep comprehension in long conversations and lack fine-grained attribution of failure modes. To this end, the authors propose the first six-dimensional evaluation framework—encompassing constrained memory, precise execution, constraint synthesis, object grounding, action inhibition, and coreference resolution—derived from real-world chatbot failure cases. They construct a challenging Chinese multi-turn dialogue benchmark comprising 209 controlled tasks spanning 12 to 76 turns, incorporating off-topic distractions and colloquial expressions. Through carefully designed tasks, capability disentanglement, and human validation, the benchmark enables interpretable assessment of models’ deep understanding abilities. Experiments reveal that even the strongest model, GPT-5.5, achieves only 41.1% fully correct responses, with no model excelling across all dimensions, highlighting significant gaps in current systems’ capacity for deep multi-turn reasoning.
📝 Abstract
Long-running multi-turn interactions with chatbots and agents are now common, and a correct response often depends on remembering earlier details, tracking later revisions, identifying intended objects or referents, and withholding action when required conditions are unmet. Existing multi-turn benchmarks typically cover short exchanges and do not fully evaluate these capabilities in long multi-turn interactions, particularly in Chinese, while offering limited insight into how and why models fail. To address these limitations, we analyze real chatbot failures to identify six recurring mechanisms and use them to define six controlled evaluation modes in Hy-MultiTurn, a Chinese benchmark for deep multi-turn dialogue understanding. The six modes evaluate constraint memory, precise execution, constraint synthesis, object localization, action suppression, and reference resolution. Across the six modes, we construct 209 controlled tasks spanning 12-76 turns, with dialogue length, irrelevant-topic distraction, and colloquial phrasing adding further difficulty. Evaluation of 22 frontier model configurations shows that Hy-MultiTurn is broadly challenging, as even GPT-5.5, the strongest overall configuration, satisfies all requirements in only 41.1 percent of responses and no model performs best in all six modes.
Problem

Research questions and friction points this paper is trying to address.

multi-turn dialogue
dialogue understanding
benchmark
long-context interaction
Chinese dialogue
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-turn dialogue understanding
dialogue benchmark
constraint memory
action suppression
reference resolution
E
Eileen Ye
Hunyuan Team, Tencent
J
Jiawen Tao
Hunyuan Team, Tencent; Peking University
Y
Yaoming Li
Peking University
C
Chenxu Liu
Hunyuan Team, Tencent
W
Wenhan Yu
Peking University
Y
Yaxin Fan
Hunyuan Team, Tencent
X
Xiaokun Yuan
Hunyuan Team, Tencent; Peking University
Mengzhou Wu
Mengzhou Wu
Peking University
Software EngineeringLarge Language Model
Y
Yanbing Jiang
Peking University
M
Maxm Pan
Hunyuan Team, Tencent