🤖 AI Summary
This study addresses the limitation of existing text generation metrics in detecting global coherence deficiencies such as logical contradictions and topic drift. To this end, it proposes CHORD, a novel evaluation metric that identifies the hidden representations of large language models (LLMs) as a critical bottleneck for coherence detection. Specifically, the method extracts frozen LLM hidden states via coherence-prompt encoding and measures distributional shifts using Radial Basis Function Maximum Mean Discrepancy (MMD). Experimental results demonstrate that CHORD precisely distinguishes benign rewrites from genuine errors. Furthermore, its model rankings exhibit strong alignment with human judgments regarding logicality and anthropomorphism, significantly outperforming conventional baselines such as perplexity and MAUVE.
📝 Abstract
Existing metrics for open-ended text generation measure likelihood, lexical diversity, or distributional similarity in generic representation space, yet they can miss fundamental dimensions of quality. A prominent blind spot is global coherence: a generated passage may be locally fluent while remaining globally contradictory, causally inconsistent, or topically disconnected. Such failures can still preserve the token-level and lexical statistics that existing metrics rely on. We identify representation as a central bottleneck in detecting these failures and introduce CHORD (Coherence-aware Hidden-state Open-generation Reference Distance), a coherence-sensitive distributional metric. CHORD encodes generated and human-written corpora in the hidden-state space of a frozen LLM using a coherence-eliciting prompt, and compares the resulting distributions using MMD with an RBF kernel. To validate that the metric responds to coherence degradation but not generic textual change, we construct a counterfactual evaluation suite that pairs graded coherence-degrading perturbations with meaning-preserving controls. CHORD selectively detects relation, discourse, structural, and mixture failures that perplexity, entropy, MAUVE, FBD, and MMD-based baselines either miss or cannot separate from benign rewriting. Factorial ablations show that representation is the primary source of coherence sensitivity,while RBF-MMD improves sample efficiency once the relevant distinctions become visible. Larger backbones capture finer-grained distinctions, but coherence prompting improves selectivity only when the backbone can follow the prompt.On unconditional generation and prefix continuation, CHORD yields model rankings that strongly align with human judgments of whether outputs make sense and appear human-written. Together, these results establish representation design as central to reliable distributional evaluation.