diversity metrics evaluation

Designs, implements, and applies quantitative metrics and software to measure and analyze diversity in datasets and model outputs—covering lexical, temporal, and distributional aspects—by computing statistics such as type–token ratios, entropy, and information‑theoretic divergences. Uses those computations to evaluate and compare metrics, validate their behavior, and detect phenomena like diversity collapse and trade‑offs between measures.

diversitymetricsevaluation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.12
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Entropy and type-token ratio in gigaword corpora

Nov 15, 2024
PR
P. Rosillo-Rodes
🏛️ University of the Balearic Islands | Spanish National Research Council

This study investigates the quantitative relationship between lexical diversity—measured by word entropy and type-token ratio (TTR)—across large-scale corpora (≈1 billion tokens each) of English, Spanish, and Turkish, spanning books, news, and tweets. Using word-frequency distributions, information-theoretic entropy computation, and power-law fitting, we empirically establish, for the first time across multiple languages and genres at scale, a highly consistent negative functional relationship between word entropy and TTR (R² > 0.99). By integrating Zipf’s law and Heaps’ law, we derive an analytical expression for this relationship in the asymptotic limit of large texts. The resulting function proves cross-linguistically invariant, revealing a universal scaling law governing lexical diversity. This work provides a unified theoretical framework for modeling linguistic diversity and establishes a testable, mathematically grounded metric foundation for quantitative language analysis.

Analyze entropy and type-token ratioExamine large text corporaMeasure lexical diversity in languages

This study addresses the lack of a unified definition and quantification of training dataset diversity, which hinders accurate assessment of its impact on model robustness. Focusing on medical imaging, the work presents the first systematic comparison between reference-free and semantic diversity metrics. Leveraging the MorphoMNIST and PadChest datasets, the authors conduct a multidimensional analysis—integrating Fréchet Inception Distance (FID), AUC, semantic diversity measures, controlled perturbations, and clinical expert evaluations—to examine how diversity correlates with expert intuition, downstream performance, and training dynamics. The findings reveal that FID and semantic diversity more effectively predict model performance, whereas merely increasing the number of imaging device sources can inadvertently encourage models to rely on non-robust shortcut features, thereby exposing a critical pitfall in data diversity design.

classification modelsdataset diversitydiversity metrics

We Need to Measure Data Diversity in NLP -- Better and Broader

May 26, 2025
DN
Dong Nguyen
🏛️ Utrecht University | Aalborg University

This study addresses fundamental challenges in quantifying dataset diversity for NLP—including conceptual ambiguity, granularity mismatch, and poor cross-domain generalizability—by proposing the first interdisciplinary diversity evaluation framework integrating linguistics, sociology, and information theory. Through conceptual analysis and methodological critique, it rigorously defines three core dimensions: semantic distance, distributional shift, and demographic representation balance, thereby identifying three principal measurement challenges: definitional vagueness, scale misalignment, and the value-neutrality dilemma. The framework enables fine-grained modeling, empirically verifiable assessment, and task-adaptive calibration. It establishes a theoretical foundation for developing fairness-aware, reproducible, and interpretable diversity metrics, advancing dataset quality evaluation from heuristic judgment toward principled, scientific quantification.

Addressing underexplored methodological challengesIncorporating interdisciplinary perspectives for better measuresMeasuring data diversity in NLP effectively

Standardizing the Measurement of Text Diversity: A Tool and a Comparative Analysis of Scores

Mar 01, 2024
CS
Chantal Shaib
🏛️ Northeastern University | Adobe

A lack of standardized, reproducible methods for quantifying textual diversity in large language models (LLMs) hinders rigorous evaluation of generation quality and cross-model or cross-corpus comparisons. Method: We propose the first systematic framework for text diversity evaluation, empirically validating convergent validity of diversity metrics and identifying a minimal, complete metric set—comprising compression ratio (zlib/lz4), long n-gram self-repetition rate, Self-BLEU, and BERTScore—that exhibits low inter-metric correlation and complementary multidimensional coverage. Contribution/Results: We release *diversity*, an open-source Python library enabling efficient computation and interactive visualization. Empirical analysis demonstrates that lightweight compression-based metrics robustly substitute for computationally expensive n-gram homogeneity scores. The framework substantially enhances interpretability, comparability, and practical utility of diversity assessment in LLM research.

Evaluating convergent validity of existing diversity scores is neededIdentifying repetitive structures in large text corpora is challengingStandardizing text diversity measurement lacks a universal method

Measuring Diversity: Axioms and Challenges

Oct 18, 2024
MM
Mikhail Mironov
🏛️ Yandex Research

Quantifying set diversity lacks a rigorous foundation, as mainstream metrics—e.g., distance-based, entropy-based, or coverage-based measures—fail to satisfy three desirable axioms: monotonicity, uniqueness, and continuity, undermining their reliability. Method: We establish the first axiomatized framework for diversity measurement, formally defining and proving the joint satisfiability of these axioms. Through systematic counterexample analysis, we demonstrate that existing approaches violate at least one axiom. We then construct an axiomatically complete metric and analyze its computational properties. Contribution/Results: We prove that any metric satisfying all three axioms is inherently NP-hard to compute, thereby establishing the “axiomatically complete yet computationally feasible” diversity measure as an open problem. This work provides the first rigorous benchmark for diversity evaluation and precisely characterizes the fundamental tension between theoretical soundness and practical computability in diversity quantification.

Constructing practical diversity measures with proven propertiesExisting measures lack desirable axiomatic propertiesQuantifying diversity for sets of objects

Latest Papers

What's happening recently
View more

Existing approaches to evaluating data diversity are often confined to lexical-level metrics and lack standardization, hindering unified cross-dimensional and cross-data-type analysis. This work proposes a general-purpose embedding-based diversity measurement framework that is compatible with any embeddable data and arbitrary embedding models, enabling, for the first time, the unified quantification of diversity across multiple dimensions—including style, semantics, language, and speaker characteristics. By operating directly in embedding spaces, the framework addresses the longstanding gap in systematic diversity assessment and demonstrates strong generality and practical utility across diverse datasets.

data diversitydataset diversityembedding-based measurement

This study investigates whether prevailing diversity metrics genuinely capture model disagreement in large language model (LLM) ensembles or merely reflect individual model capabilities. Through controlled experiments on 31,900 subsets of 30 LLMs evaluated on MMLU-Pro and TruthfulQA, the authors systematically assess the predictive power of five diversity measures for majority-vote performance gains. Using Spearman correlation, linear regression, and determinant-based algebraic analysis, they find that most diversity metrics are highly collinear with model capability. After rigorously controlling for capability, only “strict diversity,” “disagreement,” and “double-failure” retain weak yet consistent and directionally aligned associations with ensemble failure rates. The results further reveal that simple majority voting surpasses the best individual model in only a small minority of subsets, highlighting the limited utility of current diversity metrics in LLM ensembling.

capability controldiversity metricsLLM ensembles

Existing data quality evaluation methods struggle to disentangle and separately measure the fidelity and diversity of textual data, hindering the construction of high-quality training sets. This work addresses this limitation by introducing optimal transport theory into discrete text evaluation for the first time, proposing a pair of metrics based on optimal transport divergences to quantify, respectively, the fidelity (similarity) and diversity (coverage breadth) of candidate texts relative to reference data. The resulting two-dimensional framework effectively reveals the distinct impacts of these qualities on downstream model performance. Experiments on the M2D2 and GSM8K mathematical datasets demonstrate that the proposed metrics accurately identify diversity deficiencies in synthetic data and uncover a significant correlation between such deficiencies and degraded fine-tuned model accuracy.

data qualitydiversityfidelity

This study addresses a critical gap in biodiversity assessment: the absence of site-specific benchmarks for potential diversity that enable quantification of the disparity between observed communities and their ecological capacity. The authors propose a theoretical framework grounded in information geometry, which integrates Hill diversity and Rao’s quadratic entropy within a constrained variational principle on the probability simplex to construct a continuous, abundance-weighted baseline of potential diversity. This approach unifies evenness, functional traits, and phylogenetic dissimilarity into a single coherent measure, representing the first synthesis of information geometry with ecological capacity and linking explicitly to the concept of “dark diversity.” The resulting model cleanly distinguishes current diversity from attainable potential and supports dynamic extensions under scenarios of species migration and climate change, thereby providing a robust theoretical foundation for regional biodiversity conservation.

biodiversity benchmarkdiversity gapinformation geometry

This work addresses the challenge of effectively evaluating textual diversity in AI-generated and human-written outputs to detect phenomena such as mode collapse, differences in decoding strategies, and degradation in creativity. The authors propose Decan ($D_{Ca_n}$), an unsupervised, context-learning-based diversity metric that requires only a single forward pass through a base language model. Leveraging per-token log-probabilities and information-theoretic principles, Decan operates without embeddings, reference corpora, or human annotations. It is the first method to integrate information theory with in-context learning for universal, training-free diversity assessment, enabling similarity analysis across response sets of arbitrary size. Evaluated on the McDiv benchmark, Decan achieves an OCA of 0.846 and exhibits a monotonic decline across successive post-training stages of OLMo-2-7B (SFT → DPO → RLVR), effectively capturing diversity loss in creative writing.

creative writingdecoding strategiesdiversity measurement

Hot Scholars

YC

Yejin Choi

Stanford University / NVIDIA
Natural Language ProcessingDeep LearningArtificial IntelligenceCommonsense Reasoning
JZ

James Zou

Stanford University
Machine learningcomputational biologycomputational healthstatistics
YD

Yilun Du

Harvard University
Artificial IntelligenceMachine LearningRoboticsComputer Vision
PY

Pierre-Yves Oudeyer

Research director, Inria
Artificial intelligencecognitive sciencedevelopmental AIcuriosity