TPBench: A Turning-Point Benchmark for Dialogue Compression

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue of dialogue compression discarding critical turning points, which leads to the loss of user intent. We introduce the concept of "turning point eviction" and construct TPBench, a benchmark that evaluates information retention under varying budget constraints through initial goals, current values, and joint probing metrics, revealing how single aggregate scores obscure the loss of key updates. Using human-annotated data from MultiWOZ and SGD, and validated with Llama and Mistral across multilingual and long-memory scenarios, our experiments demonstrate that removing turns containing updates significantly degrades accuracy. Full-context processing yields optimal performance, while recency-based strategies prove most effective among compression methods.
📝 Abstract
A compressor can keep the facts of a dialogue and still drop the turn that changed them. A user corrects a price, reverses a choice, or adds a constraint. We call this failure turning-point eviction. One overall retention score hides it, because that score mixes what the user first wanted with what the user wants now. We introduce TPBench, which evaluates three complementary information targets at shared nominal retention budgets. P1 asks for the user's initial goal. P2 asks for the current value of a slot the user revised. P3 asks for both, in dialogues with a late annotated slot update. The current-value answers come from the human dialogue-state annotations of MultiWOZ and SGD. The initial-goal answer is the first sentence of the first user turn. Neither requires new crowdsourcing. The probe-specific evaluations rank compression methods differently. On the joint probe at a retained fraction of 0.30, every tested compressed method remains below full context with the main Llama reader. Deleting the turn that carries the update sharply lowers current-value accuracy, while deleting one matched irrelevant turn leaves it unchanged. A Mistral reader repeats the P2/P3 rankings and the joint-probe gap. Current-value recovery is tested on an additional corpus, LongMemEval-KU, and on Chinese RiSAWOZ: full context has the highest accuracy, and recency has the highest compressed-method mean in both evaluations.
Problem

Research questions and friction points this paper is trying to address.

Dialogue Compression
Turning-Point Eviction
Benchmark Evaluation
Information Retention
Slot Update
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dialogue Compression
Turning-Point Benchmark
Information Retention
Dialogue State Tracking
LLM Evaluation