semantic divergence scoring

Designs, builds, and evaluates metrics, scoring functions, and models that compute scalar semantic similarity or divergence between texts or their representations, including methods for estimating, modeling, and calibrating similarity from embeddings, token alignments, or model outputs. Uses those scores to detect semantic mismatches and to convert divergence into uncertainty or decision signals (e.g., flags or calibrated probabilities) for unreliable retrieval, generation, or reasoning steps.

semanticdivergencescoring

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$201K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

How Small Transformation Expose the Weakness of Semantic Similarity Measures

Sep 08, 2025
SL
Serge Lionel Nikiema
🏛️ University of Luxembourg

This study addresses the fundamental question of whether semantic similarity measures genuinely comprehend semantic relationships. We propose the first evaluation framework based on controlled, small-scale semantic transformations to systematically assess the semantic discrimination capability of 18 state-of-the-art methods—including bag-of-words, embedding-based, LLM-based, and structure-aware models—on software engineering texts and code. Experiments reveal that mainstream embedding methods exhibit up to 99.9% misclassification rates in semantic opposition scenarios, exposing their reliance on superficial surface patterns. Substituting cosine similarity for Euclidean distance improves performance by 24–66%. LLM-based methods demonstrate superior fine-grained semantic distinction. Critically, our framework uncovers a foundational limitation in existing measures: their failure to capture semantic essence. It establishes the first reproducible, scalable benchmark paradigm for trustworthy semantic computation in software engineering contexts.

Evaluating semantic similarity measures for software engineering tasksIdentifying flaws where methods confuse opposites and synonymsTesting 18 methods including embeddings and LLMs on semantic understanding

A Two-Sample Test of Text Generation Similarity

May 08, 2025
JX
Jingbin Xu
🏛️ Dalian University of Technology | Virginia Tech

This paper addresses the two-sample testing problem for determining whether two document collections are drawn from the same generative distribution in large-scale text data. We propose the first language-model-based entropy estimation framework for such tests. Methodologically, we (1) embed neural language model–estimated document entropy—e.g., from Transformer-based LMs—into a unified estimation-inference pipeline; and (2) construct an asymptotically normal test statistic based on entropy difference, augmented by a multi-split p-value aggregation strategy to enhance statistical power. The method rigorously controls Type-I error under both synthetic and real-world text benchmarks, while achieving substantially higher statistical power than existing text similarity–based two-sample tests. By grounding distributional comparison in information-theoretic principles, our approach establishes a verifiable, entropy-driven paradigm for assessing generative equivalence of textual sources.

Assesses text similarity via entropy using neural networksDevelops a two-sample test for comparing document similarityEnhances test power with multiple data-splitting strategy

Existing text evaluation metrics, despite exhibiting high correlation with human judgments, are vulnerable to strategic manipulation and lack robustness against irrelevant perturbations. This work introduces dual criteria—statistical alignment and strategic alignment—and formally defines strategic alignment for the first time, establishing a principled framework that encompasses human correlation, degradation sensitivity, and robustness to manipulation. Grounded in mutual information theory, the authors propose a unified metric design framework composed of four components: information measures, estimation methods, text representations, and prediction mechanisms. Empirical results demonstrate that strong correlation with human scores does not imply strategic robustness; the proposed metrics significantly enhance resistance to manipulation across peer review, summarization, and question-answering tasks while maintaining high agreement with human evaluations.

evaluation metricsnatural language generationreference-based scoring

Standardizing the Measurement of Text Diversity: A Tool and a Comparative Analysis of Scores

Mar 01, 2024
CS
Chantal Shaib
🏛️ Northeastern University | Adobe

A lack of standardized, reproducible methods for quantifying textual diversity in large language models (LLMs) hinders rigorous evaluation of generation quality and cross-model or cross-corpus comparisons. Method: We propose the first systematic framework for text diversity evaluation, empirically validating convergent validity of diversity metrics and identifying a minimal, complete metric set—comprising compression ratio (zlib/lz4), long n-gram self-repetition rate, Self-BLEU, and BERTScore—that exhibits low inter-metric correlation and complementary multidimensional coverage. Contribution/Results: We release *diversity*, an open-source Python library enabling efficient computation and interactive visualization. Empirical analysis demonstrates that lightweight compression-based metrics robustly substitute for computationally expensive n-gram homogeneity scores. The framework substantially enhances interpretability, comparability, and practical utility of diversity assessment in LLM research.

Evaluating convergent validity of existing diversity scores is neededIdentifying repetitive structures in large text corpora is challengingStandardizing text diversity measurement lacks a universal method

Existing reviewer assignment research lacks publicly available, fine-grained gold-standard data, hindering empirical evaluation of similarity algorithms. Method: We introduce the first open, human-annotated gold-standard dataset for reviewer–paper matching, comprising self-assessed expertise scores from 58 researchers across 477 papers—providing realistic, granular similarity ground truth. Using this benchmark, we systematically evaluate TF-IDF, SPECTER2, and multiple large language models (LLMs) under varying input modalities (e.g., full text, title+abstract). Results: State-of-the-art methods exhibit high misranking rates—12% to 43% on easy versus hard cases. TF-IDF matches SPECTER2’s accuracy when using full-text inputs, whereas LLMs significantly underperform; SPECTER2 achieves optimal performance with title+abstract inputs. This work fills a critical data gap in the field and establishes a reproducible, empirically grounded evaluation framework for algorithm selection and improvement.

High error rates in existing similarity score algorithmsLack of gold-standard data for reviewer assignment algorithms comparisonNeed for evidence-based algorithm selection in peer-review systems

Latest Papers

What's happening recently
View more

This work addresses the challenge that similarity scores from different embedding models are often incomparable due to geometric discrepancies, which hinders the transferability of fixed similarity thresholds across models. To overcome this limitation without requiring real queries, the authors propose a synthetic query probing method that generates controllable query–text pairs to enable large-scale analysis of cross-model similarity distributions. By learning mappings between score spaces, the approach aligns outputs using calibration strategies including linear regression, isotonic regression, and quantile mapping. Experimental results reveal that while models exhibit consistent ranking behavior, their similarity scores suffer from systematic offsets. The learned mappings substantially improve threshold portability across models, with isotonic regression yielding the best performance.

cross-model comparabilityembedding modelsretrieval-augmented generation

本文通过将TF-IDF和BM25解释为两种概率模型之间的KL散度,解决了这两种查询文档相关性评分方法缺乏统一统计理论基础的问题。

BM25Information RetrievalKullback-Leibler Divergences

This work addresses a critical limitation of existing automatic text evaluation metrics—such as ROUGE and BERTScore—which often fail to distinguish between semantically consistent and contradictory content, frequently assigning inflated scores to factually inconsistent outputs. To remedy this, the authors propose MATCHA, a reference-driven evaluation method that requires no training data and introduces, for the first time, a dual-perspective contrastive mechanism. MATCHA generates counterfactual contradictory samples adversarially and leverages semantic embedding-based contrastive learning to simultaneously reward alignment with the reference text and penalize similarity to the generated contradictions. Evaluated across eight benchmarks, MATCHA substantially outperforms prevailing metrics, achieving relative improvements of 18.38% and 20.82% over ROUGE-L and BERTScore, respectively, on TruthfulQA, and demonstrating superior performance across 23 embedding models—thereby exposing a fundamental flaw in current evaluation approaches regarding semantic contradiction detection.

automatic scoringcontradiction detectionevaluation metrics

This study investigates whether Herman Melville’s reading influenced his writing at the semantic level. By leveraging BERTScore to compute sentence-level and 5-gram-level semantic similarity between Melville’s works and texts from his personal library, the research replaces conventional fixed thresholds with a semantic alignment metric to enable finer-grained analysis of literary influence. The approach integrates precision, recall, and F1 score for comprehensive evaluation, successfully reproducing established cases of influence documented in the scholarly literature while also uncovering several novel candidate passages suggestive of previously unrecognized connections. This work thus offers a verifiable and scalable computational framework for tracing literary sources and influences.

computational analysisHerman Melvilleliterary influence

This work addresses the limitations of existing reference-free summarization evaluation methods, which often suffer from inadequate calibration and reliance on human annotations or large language models, thereby failing to reliably reflect true summary quality. The authors propose a general framework that requires neither reference summaries nor human labels, capable of producing proxy scores for both individual and average summary quality. Central to this approach is Group Isotonic Regression Binning (GIRB), a novel calibration technique designed for continuous-valued tasks, which—used for the first time in a reference-free setting—enables high-quality proxy scoring. Experiments across seven datasets demonstrate that the proposed method significantly outperforms current baselines, substantially improving the reliability and generalizability of evaluation metrics, with straightforward extension to discrete tasks such as question answering.

calibrationmiscalibrationmodel-based metrics

Hot Scholars

JH

Jiawei Han

Abel Bliss Professor of Computer Science, University of Illinois
data miningdatabase systemsdata warehousinginformation networks
PS

Philip S. Yu

Professor of Computer Science, University of Illinons at Chicago
Data miningDatabasePrivacy
DY

Dawei Yin

Senior Director, Head of Search Science at Baidu
Machine LearningWeb MiningData Mining
SW

Shuaiqiang Wang

Principal Architect of Search Strategy, Baidu Inc.
Large language modelsInformation retrieval
SW

Seung-won Hwang

Seoul National University
language/data understanding