Score
Designs, builds, and evaluates metrics, scoring functions, and models that compute scalar semantic similarity or divergence between texts or their representations, including methods for estimating, modeling, and calibrating similarity from embeddings, token alignments, or model outputs. Uses those scores to detect semantic mismatches and to convert divergence into uncertainty or decision signals (e.g., flags or calibrated probabilities) for unreliable retrieval, generation, or reasoning steps.
This study addresses the fundamental question of whether semantic similarity measures genuinely comprehend semantic relationships. We propose the first evaluation framework based on controlled, small-scale semantic transformations to systematically assess the semantic discrimination capability of 18 state-of-the-art methods—including bag-of-words, embedding-based, LLM-based, and structure-aware models—on software engineering texts and code. Experiments reveal that mainstream embedding methods exhibit up to 99.9% misclassification rates in semantic opposition scenarios, exposing their reliance on superficial surface patterns. Substituting cosine similarity for Euclidean distance improves performance by 24–66%. LLM-based methods demonstrate superior fine-grained semantic distinction. Critically, our framework uncovers a foundational limitation in existing measures: their failure to capture semantic essence. It establishes the first reproducible, scalable benchmark paradigm for trustworthy semantic computation in software engineering contexts.
This paper addresses the two-sample testing problem for determining whether two document collections are drawn from the same generative distribution in large-scale text data. We propose the first language-model-based entropy estimation framework for such tests. Methodologically, we (1) embed neural language model–estimated document entropy—e.g., from Transformer-based LMs—into a unified estimation-inference pipeline; and (2) construct an asymptotically normal test statistic based on entropy difference, augmented by a multi-split p-value aggregation strategy to enhance statistical power. The method rigorously controls Type-I error under both synthetic and real-world text benchmarks, while achieving substantially higher statistical power than existing text similarity–based two-sample tests. By grounding distributional comparison in information-theoretic principles, our approach establishes a verifiable, entropy-driven paradigm for assessing generative equivalence of textual sources.
Existing text evaluation metrics, despite exhibiting high correlation with human judgments, are vulnerable to strategic manipulation and lack robustness against irrelevant perturbations. This work introduces dual criteria—statistical alignment and strategic alignment—and formally defines strategic alignment for the first time, establishing a principled framework that encompasses human correlation, degradation sensitivity, and robustness to manipulation. Grounded in mutual information theory, the authors propose a unified metric design framework composed of four components: information measures, estimation methods, text representations, and prediction mechanisms. Empirical results demonstrate that strong correlation with human scores does not imply strategic robustness; the proposed metrics significantly enhance resistance to manipulation across peer review, summarization, and question-answering tasks while maintaining high agreement with human evaluations.
A lack of standardized, reproducible methods for quantifying textual diversity in large language models (LLMs) hinders rigorous evaluation of generation quality and cross-model or cross-corpus comparisons. Method: We propose the first systematic framework for text diversity evaluation, empirically validating convergent validity of diversity metrics and identifying a minimal, complete metric set—comprising compression ratio (zlib/lz4), long n-gram self-repetition rate, Self-BLEU, and BERTScore—that exhibits low inter-metric correlation and complementary multidimensional coverage. Contribution/Results: We release *diversity*, an open-source Python library enabling efficient computation and interactive visualization. Empirical analysis demonstrates that lightweight compression-based metrics robustly substitute for computationally expensive n-gram homogeneity scores. The framework substantially enhances interpretability, comparability, and practical utility of diversity assessment in LLM research.
Existing reviewer assignment research lacks publicly available, fine-grained gold-standard data, hindering empirical evaluation of similarity algorithms. Method: We introduce the first open, human-annotated gold-standard dataset for reviewer–paper matching, comprising self-assessed expertise scores from 58 researchers across 477 papers—providing realistic, granular similarity ground truth. Using this benchmark, we systematically evaluate TF-IDF, SPECTER2, and multiple large language models (LLMs) under varying input modalities (e.g., full text, title+abstract). Results: State-of-the-art methods exhibit high misranking rates—12% to 43% on easy versus hard cases. TF-IDF matches SPECTER2’s accuracy when using full-text inputs, whereas LLMs significantly underperform; SPECTER2 achieves optimal performance with title+abstract inputs. This work fills a critical data gap in the field and establishes a reproducible, empirically grounded evaluation framework for algorithm selection and improvement.
This work addresses the challenge that similarity scores from different embedding models are often incomparable due to geometric discrepancies, which hinders the transferability of fixed similarity thresholds across models. To overcome this limitation without requiring real queries, the authors propose a synthetic query probing method that generates controllable query–text pairs to enable large-scale analysis of cross-model similarity distributions. By learning mappings between score spaces, the approach aligns outputs using calibration strategies including linear regression, isotonic regression, and quantile mapping. Experimental results reveal that while models exhibit consistent ranking behavior, their similarity scores suffer from systematic offsets. The learned mappings substantially improve threshold portability across models, with isotonic regression yielding the best performance.
本文通过将TF-IDF和BM25解释为两种概率模型之间的KL散度,解决了这两种查询文档相关性评分方法缺乏统一统计理论基础的问题。
This work addresses a critical limitation of existing automatic text evaluation metrics—such as ROUGE and BERTScore—which often fail to distinguish between semantically consistent and contradictory content, frequently assigning inflated scores to factually inconsistent outputs. To remedy this, the authors propose MATCHA, a reference-driven evaluation method that requires no training data and introduces, for the first time, a dual-perspective contrastive mechanism. MATCHA generates counterfactual contradictory samples adversarially and leverages semantic embedding-based contrastive learning to simultaneously reward alignment with the reference text and penalize similarity to the generated contradictions. Evaluated across eight benchmarks, MATCHA substantially outperforms prevailing metrics, achieving relative improvements of 18.38% and 20.82% over ROUGE-L and BERTScore, respectively, on TruthfulQA, and demonstrating superior performance across 23 embedding models—thereby exposing a fundamental flaw in current evaluation approaches regarding semantic contradiction detection.
This study investigates whether Herman Melville’s reading influenced his writing at the semantic level. By leveraging BERTScore to compute sentence-level and 5-gram-level semantic similarity between Melville’s works and texts from his personal library, the research replaces conventional fixed thresholds with a semantic alignment metric to enable finer-grained analysis of literary influence. The approach integrates precision, recall, and F1 score for comprehensive evaluation, successfully reproducing established cases of influence documented in the scholarly literature while also uncovering several novel candidate passages suggestive of previously unrecognized connections. This work thus offers a verifiable and scalable computational framework for tracing literary sources and influences.
This work addresses the limitations of existing reference-free summarization evaluation methods, which often suffer from inadequate calibration and reliance on human annotations or large language models, thereby failing to reliably reflect true summary quality. The authors propose a general framework that requires neither reference summaries nor human labels, capable of producing proxy scores for both individual and average summary quality. Central to this approach is Group Isotonic Regression Binning (GIRB), a novel calibration technique designed for continuous-valued tasks, which—used for the first time in a reference-free setting—enables high-quality proxy scoring. Experiments across seven datasets demonstrate that the proposed method significantly outperforms current baselines, substantially improving the reliability and generalizability of evaluation metrics, with straightforward extension to discrete tasks such as question answering.