evaluation metric design

Defining and validating quantitative and human-evaluated metrics (including user studies) that capture outcome accuracy, evidence/reasoning quality, and agreement signals; used to run and interpret experiments such as spoken language ID, forecasting evaluation, and embedding-agreement comparisons.

evaluationmetricdesign

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing automatic evaluation metrics, such as BLEU and TER, commonly overlook communicative context in interpreting quality assessment, thereby failing to accurately reflect the real-world effectiveness of interpretation. This study addresses this gap by incorporating communicative context as a core dimension, integrating theoretical insights from interpreting studies with empirical analyses of automatic metrics to systematically examine their applicability and limitations in authentic interpreting scenarios. The findings reveal that context-independent automatic metrics cannot reliably evaluate interpreting quality in isolation and must be complemented by contextual factors for a valid assessment. By foregrounding the situated nature of interpreting, this work provides both theoretical grounding and methodological innovation for developing evaluation frameworks that better align with the communicative essence of interpreting practice.

automated evaluation metricscommunicative contextinterpreting quality

A Measure of the System Dependence of Automated Metrics

Dec 04, 2024
PV
Pius von Däniken
🏛️ ZHAW School of Engineering

This paper addresses the pervasive system dependency problem in machine translation (MT) automatic evaluation metrics—where metric scores are biased by idiosyncratic characteristics of specific MT systems, leading to unreliable and unfair cross-system comparisons. We propose the first formal, quantitative framework for measuring system dependency. Our methodology integrates rank stability analysis, cross-system score distribution comparison, Monte Carlo perturbation experiments, and statistical significance testing to enable reproducible, diagnostic assessment of metric bias. Empirical validation on WMT benchmarks reveals statistically significant system dependency in BLEU, COMET, and BERTScore. Beyond exposing latent fairness deficiencies in widely adopted metrics, this work establishes a new benchmark for fair evaluation and releases an open-source diagnostic toolkit. It provides both theoretical foundations and practical guidance for metric design, selection, and refinement—advancing rigor and equity in MT evaluation.

Automated ScoringFairness EvaluationMachine Translation

HCI scale development has long suffered from nonstandardized processes, poor construct-theory alignment, and low item reuse rates. This paper introduces the first interactive support system integrating large language models (LLMs) with a structured, empirically grounded measurement knowledge base, enabling a closed-loop workflow: construct identification → theory-informed custom definition → context-aware item generation. The system retrieves theoretically appropriate constructs from a literature-anchored database and leverages LLMs to generate semantically coherent, domain-specific items, supporting human-AI co-refinement. Its key innovation lies in the deep coupling of LLMs with an evidence-validated construct–item relational database, shifting scale development from experience-driven practice toward evidence-enhanced collaborative measurement. Experiments show a 62% reduction in design time, a 3.1× increase in item reuse, and significantly improved theoretical fidelity; expert evaluations across multiple rounds confirm ≥92% contextual appropriateness. The system has been integrated into a prototype HCI research workflow.

Improving rigor and efficiency in HCI measurement designLeveraging LLMs and prior literature for construct developmentStandardizing measurement item design process for researchers

Amid the growing diversity of natural language processing tasks, existing inter-annotator agreement (IAA) metrics often suffer from limited applicability and interpretability when confronted with heterogeneous task types, label imbalance, and missing data. This work systematically reviews the theoretical foundations and practical methodologies of IAA, offering the first structured integration of mainstream metrics—such as Cohen’s Kappa and Krippendorff’s Alpha—organized by task type. It clarifies their underlying assumptions and delineates their boundaries of applicability. Furthermore, the study proposes a reliability assessment strategy that combines confidence intervals with analysis of disagreement patterns. By providing a clear, principled guide for selecting IAA metrics, this research significantly enhances the transparency, reproducibility, and scientific rigor of human annotation and evaluation practices in the NLP community.

agreement metricsannotation reliabilityhuman annotation

Measuring Hypothesis Testing Errors in the Evaluation of Retrieval Systems

Jul 10, 2025
JM
Jack McKechnie
🏛️ University of Glasgow

This study addresses the quantification of discriminative power in query-document relevance judgments (qrels) for information retrieval (IR) evaluation, with particular emphasis on the historically underexamined Type II error (false negatives) and its joint analysis with Type I error (false positives). Method: We systematically introduce Type II error modeling into IR evaluation for the first time and propose replacing conventional significance testing with balanced classification metrics—such as balanced accuracy—as a more principled basis for assessing qrels’ discriminative capability. A unified, comparable measurement framework is thereby established. Results: Empirical hypothesis testing and statistical significance analysis across multiple qrels generation strategies demonstrate that jointly evaluating both error types exposes latent quality deficiencies in qrels more comprehensively than traditional approaches. Balanced classification metrics robustly aggregate discriminative performance, substantially enhancing the reliability and interpretability of IR system evaluation.

Assessing discriminative power of relevance assessments (qrels)Quantifying Type II errors in IR system evaluationsUsing balanced metrics to summarize qrels' effectiveness

Latest Papers

What's happening recently
View more

This study addresses the lack of empirical validation for existing retrieval-augmented generation (RAG) evaluation metrics in real-world scenarios. Leveraging a human-annotated commercial question-answering dataset, it presents the first systematic comparison of prominent metrics from four major evaluation frameworks—Ragas, DeepEval, RAGChecker, and Opik. Through correlation analyses against human judgments and traditional metrics such as recall, the work reveals that most automatic metrics exhibit weak alignment with human assessments. These findings not only highlight significant limitations in current RAG evaluation methodologies when applied to practical settings but also provide empirical grounding and actionable directions for developing more reliable, real-world-oriented evaluation approaches.

empirical studyevaluation metricsquestion answering

This study addresses the inconsistent metric selection and inadequate reporting practices in LLM-as-judge research, which hinder reproducibility and cross-study comparison. Through a systematic analysis of agreement measures between large language models and human evaluators, the work reveals mathematical equivalences among multiple correlation coefficients—such as Pearson, Spearman, phi, and Matthews—under binary scoring, clarifies the distinct utility of Cohen’s κ, and elucidates how handling abstentions fundamentally affects evaluation outcomes. Building on these statistical insights, the paper proposes a standardized reporting checklist that encompasses rating scales, treatment of abstentions and ties, coverage, confusion matrices, and aggregation levels. This framework substantially enhances the transparency, comparability, and reproducibility of LLM-as-judge evaluations.

abstention handlingagreement metricsbinary evaluation

Current AI research tools lack evaluation benchmarks that simultaneously account for usability, interpretability, and integration into scientific workflows, making it difficult to assess their practical reliability in academic settings. This work proposes a comprehensive evaluation framework that integrates human-centered dimensions—such as usability and interpretability—with computational metrics. Through a human-AI collaborative approach—including explainable AI (xAI) analysis, source tracing validation, task-oriented testing, and workflow integration observation—the study systematically evaluates AI-powered question-answering and literature review tools on both exploratory and precision-oriented tasks. Findings reveal a core tension: while these tools effectively support initial exploration by providing useful overviews, they exhibit unreliable precision in factual extraction, with xAI highlights often misaligned with actual answers. Similarly, literature tools aid preliminary discovery but suffer from poor reproducibility and low transparency, necessitating rigorous human verification.

academic researchAI toolsbenchmarking

It remains unclear whether individual scores derived from large language models (LLMs) for inferring user states in operational settings exhibit psychometric stability and interpretability. This study introduces the first reproducible psychometric evaluation framework, integrating test–retest reliability and aggregate reliability analyses to systematically assess 213 user state indicators generated by multimodal LLMs—including GPT-4o audio, Gemini 2.0 Flash, and Gemini 2.5 Flash. Findings reveal that only 31 indicators meet established reliability criteria. Most individual scores demonstrate insufficient stability for real-time adaptation; however, they remain valuable in post-hoc analyses for uncovering patterns of user interaction and their associations with satisfaction, trust, and engagement.

AI trustworthinesslarge language modelsmetric stability

This study addresses the lack of transparent and reproducible protocols in human evaluation for long-form text generation, which hinders interpretability and cross-study comparison. To remedy this gap, the work proposes the first set of 20 reportable standards tailored to this task and conducts a large-scale systematic review of human evaluation practices. The analysis integrates manual annotation of 284 papers from *CL conferences (2023–2025) with large language model–assisted examination of over 1,800 additional publications, yielding a structured framework for assessment. The findings reveal that most studies omit critical methodological details, prompting the authors to formulate actionable recommendations to enhance transparency and reproducibility. Accompanying this contribution, the code and annotated dataset are publicly released.

evaluation protocolshuman evaluationlong-form text generation

Hot Scholars

HQ

Huamin Qu

Chair Professor, Hong Kong University of Science and Technology
Data visualizationHuman-Computer InteractionExplainable AIE-Learning
RD

Ronnie de Souza Santos

Assistant Professor, University of Calgary
Human Aspects of Software EngineeringSoftware TestingSoftware FairnessSoftware Development
MK

Marcos Kalinowski

Professor, Pontifical Catholic University of Rio de Janeiro (PUC-Rio)
Empirical Software EngineeringAI EngineeringAI4SEHuman Aspects in Software Engineering
CT

Christoph Treude

Associate Professor of Computer Science, Singapore Management University
Software EngineeringEmpirical Software EngineeringHuman-AI InteractionAI for Science
PH

Pan Hui

Chair Professor, Nokia Chair in Data Science, FREng & IEEE Fellow (HKUST & University of Helsinki)
Ubiquitous ComputingMobile ComputingAugmented RealityData Science