llm-based evaluation

Designs and implements evaluation systems that use large language models as automated judges and scorers, including role‑ and layer‑based, multimodal, and author‑perspective or role‑playing protocols to compare generated outputs to references and to measure fine‑grained semantics, relevance, planning, and classification. Builds evaluation metrics and benchmark tests, orchestrates inference pipelines (e.g., oracle–retriever–generator flows), produces interpretable and contestable rationales, filters and ranks retrieval pools, and analyzes robustness to prompt injection and alignment with human judgment.

llm-basedevaluation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.31
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$204K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks

Apr 26, 2025
YC
Yixin Cao
🏛️ Fudan University | Nanyang Technological University | Singapore Management University | Tsinghua University | Singapore University of Technology and Design | University of California Davis | National University of Singapore | University of Illinois Urbana-Champaign | Australian National University

Existing evaluation methodologies for large language models (LLMs) suffer from insufficient generalization assessment, as static benchmarks fail to capture the continuously expanding capability boundaries of evolving LLMs. Method: We formally define “evaluation generalizability” and propose a four-dimensional analytical framework encompassing evaluation methodologies, datasets, evaluators, and metrics. Our approach innovatively integrates LLM-as-a-judge, dynamically updated datasets, capability-decoupled benchmark design, and a multidimensional meta-evaluation framework. Contribution: We establish a novel, capability-oriented, automated, and sustainably evolvable evaluation paradigm covering critical dimensions—including knowledge, reasoning, instruction following, multimodal understanding, and safety. Concurrently, we release an open-source, extensible GitHub “living review” repository—a community-maintained, versioned resource—to advance evaluation practice from static benchmarking toward dynamic, collaborative co-evolution.

Addressing evaluation challenges posed by advancing Large Language ModelsOvercoming generalization issues in bounded test sets for LLMsTransitioning from task-specific to capability-based model evaluation

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of existing automatic evaluation methods for large language model (LLM) outputs, which often rely on reference texts and exhibit limited generalizability, thereby struggling to accurately assess the quality and relevance of generated content. The authors propose a reference-free, domain-agnostic automated evaluation framework that leverages pairwise comparisons among multiple LLMs, integrated with an Elo rating system to produce stable and interpretable rankings. A tunable consistency threshold is introduced to balance evaluation confidence against coverage. Evaluated on scientific abstract quality assessment, the method yields rankings that align closely with expert judgments, significantly reducing the need for manual evaluation while demonstrating near-expert assessment capability.

automated evaluationLarge Language Modelsquality assessment

Role-Playing Evaluation for Large Language Models

May 19, 2025
YE
Yassine El Boudouri
🏛️ Univ. Lille | CNRS | Centrale Lille

Current evaluation of large language models’ (LLMs) role-playing capabilities suffers from high manual annotation costs and significant biases in automated metrics. To address this, we propose RPEval—the first multidimensional benchmark specifically designed for role-playing evaluation—systematically defining and quantifying four core dimensions: emotional understanding, decision-making reasoning, moral alignment, and role consistency. Methodologically, RPEval constructs multi-turn dialogue tasks grounded in real-world scenarios and integrates expert annotation, adversarial testing, and consistency verification to enable dual-track assessment via automated scoring and human arbitration. The framework ensures reproducibility, extensibility, and human-AI collaborative validation. We publicly release the dataset and implementation code, and conduct baseline evaluations across mainstream LLMs. Results reveal substantial deficiencies in moral alignment and long-term role consistency—insights previously unattainable due to the absence of standardized benchmarks—thereby establishing the first community-wide evaluation standard for LLM role-playing.

Evaluating LLM role-playing ability is challengingHuman assessments are resource-intensive and automated ones biasedProposes RPEval benchmark for multi-dimensional role-playing evaluation

Adaptively evaluating models with task elicitation

Mar 03, 2025
DB
Davis Brown
🏛️ University of Pennsylvania

The rapid advancement of large language models (LLMs) outpaces the availability of up-to-date, human-annotated evaluation data, hindering timely and reliable assessment of model capabilities. Method: We propose a task-elicited adaptive evaluation framework grounded in scaffolded evaluation agents. It performs behavioral space search and dynamic task generation over domain-specific corpora to automatically discover high-difficulty, high-discriminative failure cases. The method supports cross-model transfer of challenging instances and incorporates human verification to ensure validity. Contribution/Results: Applied to legal reasoning, predictive modeling, and online harassment detection, the framework uncovers systematic inconsistency flaws in state-of-the-art LLMs. The resulting benchmark exhibits strong generalizability across models with diverse capability profiles, enabling high-quality, sustainable, domain-specific evaluation—a novel paradigm for robust LLM assessment.

Creating domain-specific datasets for diverse tasksIdentifying model failure modes through adaptive probingScalable evaluation of rapidly evolving language models

This study investigates the mechanisms influencing human–LLM judgment alignment in human–AI collaborative evaluation, focusing on how task characteristics and AI assistance strategies shape users’ construction and dynamic refinement of evaluation criteria, as well as their model selection behavior. Method: We conducted a controlled human–AI interaction study involving 15 ML practitioners performing 131 real-world evaluation tasks, comparing direct assessment versus pairwise comparison paradigms, augmented by multi-round LLM-assisted judgments and qualitative behavioral analysis. Contribution/Results: We present the first empirical evidence that direct assessment significantly enhances user engagement and criterion-task alignment, facilitating personalized criterion customization, dynamic judgment adjustment, and adaptive model switching. Based on these findings, we propose design principles for front-end evaluation tools tailored to human–AI collaboration. Our work advances low-overhead, interpretable, and task-adaptive AI-assisted evaluation frameworks.

Aligning human and LLM judgments for evaluationsImproving AI-assisted evaluation strategies and toolsReducing cost and time in LLM output assessments

Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks

Oct 30, 2023
QL
Qintong Li
🏛️ The University of Hong Kong | Tencent AI lab

This study systematically evaluates the reliability and human-expert consistency of large language models (LLMs) as automated evaluators for NLP tasks. Method: Leveraging a multi-task human-annotated benchmark, we assess zero-shot and few-shot LLM-based evaluation across dimensions—including fluency, factual consistency, and logical coherence—in story generation and mathematical reasoning, employing consistency metrics and fine-grained error attribution. Contribution/Results: We uncover, for the first time, systematic biases in LLMs’ generation of evaluation criteria. To address this, we propose “pre-writing human–AI collaborative evaluation”: LLMs first generate structured, criterion-grounded rationales, which humans then calibrate. Experiments show this paradigm improves human evaluation objectivity by 27% and substantially mitigates subjectivity and outlier annotations. While LLMs achieve near-human performance on general dimensions (e.g., fluency), they remain significantly lagging on complex, quantitative reasoning criteria.

Evaluation ReliabilityLanguage ModelsText Processing Tasks

Latest Papers

What's happening recently
View more

This work addresses the inefficiency and poor reproducibility of large language model (LLM) evaluation, which often relies on labor-intensive manual processes such as benchmark selection, code reproduction, and metric interpretation. To overcome these limitations, we propose the first agent-driven automated evaluation system that translates natural language evaluation requests into end-to-end executable, traceable, and customizable evaluation workflows. The system leverages NL2Bench for intent parsing and benchmark planning, BenchResolve for standardized data acquisition, and integrates task-aware metric selection with decision-oriented report generation. Human-in-the-loop checkpoints and a sample evidence chain mechanism are introduced to significantly enhance transparency, controllability, and reproducibility. In industrial settings, the framework enables efficient execution of diverse evaluation tasks with minimal human intervention.

benchmark selectionevaluation automationLLM evaluation

Existing NLP evaluation metrics struggle to effectively assess role-playing large language models (LLMs) in terms of character consistency, logical coherence, and long-term narrative stability. To address this gap, this work proposes RPA-Check, a four-stage automated evaluation framework that decomposes evaluation dimensions, generates Boolean checklists, performs semantic deduplication and isolation, and integrates a chain-of-thought-enhanced LLM-as-a-Judge mechanism. This framework establishes the first structured, reproducible benchmark specifically designed for role-playing agents. Experimental results from the LLM Court forensic training game reveal that instruction-finetuned small models (8–9B parameters) outperform larger counterparts in procedural consistency, suggesting an inverse relationship between model scale and role consistency—thereby challenging the prevailing “bigger is better” paradigm in LLM development.

Agent AssessmentAutomated EvaluationLarge Language Models

This study addresses the lack of systematic evaluation regarding the reliability and alignment with human judgment of large language models (LLMs) when deployed as automated evaluators. The authors construct a human-annotated gold-standard dataset spanning eight distinct tasks and conduct the first large-scale empirical analysis of 37 open- and closed-source conversational LLMs under five裁判 prompting strategies, a two-stage judging mechanism, and task-specific fine-tuning. Results demonstrate that GPT-4o, open-source models with at least 32 billion parameters, and Qwen2.5-14B achieve high agreement with human judgments when paired with appropriate prompts, thereby validating the feasibility of using LLMs as reliable automated evaluators. The findings offer empirical guidance for prompt design, model selection, and architectural optimization in automated assessment systems.

Automated JudgmentFidelityHuman Agreement

This work proposes a fine-grained evaluation framework to assess the capability of large language models (LLMs) as relevance judges in information retrieval, extending beyond holistic document-level judgments to identify the specific textual spans that support those judgments. Leveraging a Wikipedia test collection derived from INEX, the study employs prompt engineering to guide LLMs in simultaneously performing document-level relevance assessment and span-level annotation, followed by comparative analysis against human annotations. By introducing fine-grained relevance evaluation into the LLMs-as-Judges paradigm, this research is the first to examine whether models are “right for the right reasons,” thereby substantially enhancing the credibility of automated evaluation. Experimental results demonstrate that, under human supervision, LLMs can accurately identify both relevant documents and the key evidence spans within them.

Fine-grained EvaluationInformation RetrievalLLMs-as-Judges

AutoBench: Automating LLM Evaluation through Reciprocal Peer Assessment

Oct 26, 2025
DL
Dario Loi
🏛️ Sapienza University of Rome | eZecute S.R.L.

To address severe test-set contamination, poor adaptability of static benchmarks, and limited capacity for dynamic task generation in LLM evaluation, this paper proposes a fully automated, contamination-resistant distributed evaluation framework. Methodologically, it introduces a multi-agent mutual-evaluation mechanism wherein models alternately assume the roles of “task generator” and “evaluator”; task generation and closed-loop assessment are enabled via cyclic weighting, consensus-based aggregation, and iterative reliability calibration. Crucially, the framework eliminates reliance on fixed test sets and enhances robustness and human alignment through collaborative judgment by multiple evaluators. Empirical results show Pearson correlations of 78% with human ratings on MMLU-Pro and 63% on GPQA—substantially outperforming single-evaluator baselines—demonstrating both effectiveness and strong generalization across diverse reasoning-intensive benchmarks.

Addressing test-set contamination in static benchmarksAutomating LLM evaluation via reciprocal peer assessmentGenerating dynamic tasks for robust model comparison

Hot Scholars

AC

Arman Cohan

Yale University; Allen Institute for AI
Natural Language ProcessingMachine LearningArtificial Intelligence
DT

Dacheng Tao

Nanyang Technological University
artificial intelligencemachine learningcomputer visionimage processing
SJ

Shafiq Joty

Sr. Research Director at Salesforce Research, Assoc. Prof. at NTU (on leave)
Natural Language ProcessingMachine Learning
GD

Greg Durrett

Associate Professor of Computer Science, New York University
Natural Language Processing
MT

Md Tahmid Rahman Laskar

Senior Applied Scientist, Dialpad
Large Language ModelsNatural Language ProcessingDeep LearningQuestion Answering