Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics

📅 2026-08-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing text evaluation metrics, despite exhibiting high correlation with human judgments, are vulnerable to strategic manipulation and lack robustness against irrelevant perturbations. This work introduces dual criteria—statistical alignment and strategic alignment—and formally defines strategic alignment for the first time, establishing a principled framework that encompasses human correlation, degradation sensitivity, and robustness to manipulation. Grounded in mutual information theory, the authors propose a unified metric design framework composed of four components: information measures, estimation methods, text representations, and prediction mechanisms. Empirical results demonstrate that strong correlation with human scores does not imply strategic robustness; the proposed metrics significantly enhance resistance to manipulation across peer review, summarization, and question-answering tasks while maintaining high agreement with human evaluations.
📝 Abstract
Reference-based text evaluation metrics, which are widely used to assess natural language generation systems, score a candidate response by comparing it with a reference response. The reliability of an evaluation metric is usually judged by its statistical correlation with human ratings. However, as these metrics are increasingly used as optimization objectives, correlation alone is no longer sufficient: agents may strategically game the evaluation metric. We study this issue through two complementary notions of alignment. A metric is statistically aligned if it correlates with human ratings and strategically aligned if it resists perturbations that do not add task-relevant information. We make two contributions. First, we propose test principles for reference-based metrics consisting of human-rating correlation, degradation sensitivity, and manipulation robustness. These principles evaluate whether a metric agrees with human judgments, penalizes low-effort information loss, and resists strategic score inflation. Second, we develop a unified design framework for mutual-information-based metrics that decomposes existing and new metrics into four choices: information measure, estimation method, text representation, and prediction mechanism. Across peer review, summarization, and question answering, we find that strong human-rating correlation does not imply strategic alignment: LLM-as-a-Judge achieves high correlation but is susceptible to manipulation. In contrast, mutual-information-based metrics substantially improve manipulation robustness. Our framework also uncovers a new metric that achieves the strongest overall robustness in our experiments while remaining competitive on human-rating correlation.
Problem

Research questions and friction points this paper is trying to address.

evaluation metrics
strategic manipulation
statistical alignment
natural language generation
reference-based scoring
Innovation

Methods, ideas, or system contributions that make the work stand out.

strategic alignment
mutual information
evaluation robustness
text generation metrics
manipulation resistance
🔎 Similar Papers
No similar papers found.