Preserving What Matters: Semantic Scaffolds Beyond Saturation in Summarization Evaluation

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决现有摘要评价指标无法有效区分模型的问题,提出基于语义支架的新框架,通过提取事实、问题和实体属性等信息,生成新的评价指标。
📝 Abstract
Summarization ships in countless production systems, making model selection a routine decision that depends on measuring summary quality. Existing metrics struggle to support this: ROUGE captures only surface overlap, while LLM-as-judge scores saturate to near-identical values that fail to rank models effectively. We observe this saturation across three public datasets, two proprietary datasets, and multilingual settings. Motivated by this, we introduce Semantic Scaffold, an evaluation framework that extracts a hierarchical representation of facts, questions, and entity attributes from a source text, labeling each as a main point or supporting detail, and reusing this structure as a fixed reference for scoring summaries. From this representation, we derive three diagnostic metrics: Fact Preservation Score (FPS), Question Preservation Score (QPS), and Entity Preservation Score (EPS), designed to reward the preservation of essential information while penalizing detail overload, and position them as interpretable diagnostics that remain informative where holistic axes collapse. Finally, we analyze four recurring failure modes of ROUGE and LLM-as-judge scores, demonstrating that scaffold-based evaluation remains informative where conventional metrics collapse.
Problem

Research questions and friction points this paper is trying to address.

Summarization Evaluation
Metric Saturation
Summary Quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic Scaffold
Fact Preservation Score (FPS)
Question Preservation Score (QPS)
Entity Preservation Score (EPS)
N
Nikhil Reddy Pottanigari
ServiceNow Canada
Ramin Fahimi
Ramin Fahimi
ServiceNow Canada
N
Noah Bolger
ServiceNow Canada
S
Sepideh Kharaghani
ServiceNow Canada
Ying Zhang
Ying Zhang
ServiceNow Canada