Score
Designs and implements evaluation systems that use large language models as automated judges and scorers, including role‑ and layer‑based, multimodal, and author‑perspective or role‑playing protocols to compare generated outputs to references and to measure fine‑grained semantics, relevance, planning, and classification. Builds evaluation metrics and benchmark tests, orchestrates inference pipelines (e.g., oracle–retriever–generator flows), produces interpretable and contestable rationales, filters and ranks retrieval pools, and analyzes robustness to prompt injection and alignment with human judgment.
The reliability of LLM-as-a-Judge—particularly concerning evaluation accuracy, inter-annotator consistency, and fairness—remains a critical bottleneck. Method: This paper introduces the first systematic reliability assessment framework, featuring a dedicated benchmark (ReliBench) and synthesizing three core strategies: consistency enhancement, bias mitigation, and scenario adaptation. Our methodology integrates multi-model cross-validation, adversarial testing, interpretability analysis, and structured prompt engineering, underpinned by a standardized evaluation protocol. Contributions: (1) The first comprehensive, domain-wide survey of LLM-as-a-Judge reliability; (2) An open-source, reproducible evaluation toolkit and benchmark; (3) A paradigm shift from heuristic, experience-driven LLM evaluation toward rigorous, scientific, and standardized assessment practices.
Existing evaluation methodologies for large language models (LLMs) suffer from insufficient generalization assessment, as static benchmarks fail to capture the continuously expanding capability boundaries of evolving LLMs. Method: We formally define “evaluation generalizability” and propose a four-dimensional analytical framework encompassing evaluation methodologies, datasets, evaluators, and metrics. Our approach innovatively integrates LLM-as-a-judge, dynamically updated datasets, capability-decoupled benchmark design, and a multidimensional meta-evaluation framework. Contribution: We establish a novel, capability-oriented, automated, and sustainably evolvable evaluation paradigm covering critical dimensions—including knowledge, reasoning, instruction following, multimodal understanding, and safety. Concurrently, we release an open-source, extensible GitHub “living review” repository—a community-maintained, versioned resource—to advance evaluation practice from static benchmarking toward dynamic, collaborative co-evolution.
This work addresses the limitations of existing automatic evaluation methods for large language model (LLM) outputs, which often rely on reference texts and exhibit limited generalizability, thereby struggling to accurately assess the quality and relevance of generated content. The authors propose a reference-free, domain-agnostic automated evaluation framework that leverages pairwise comparisons among multiple LLMs, integrated with an Elo rating system to produce stable and interpretable rankings. A tunable consistency threshold is introduced to balance evaluation confidence against coverage. Evaluated on scientific abstract quality assessment, the method yields rankings that align closely with expert judgments, significantly reducing the need for manual evaluation while demonstrating near-expert assessment capability.
Current evaluation of large language models’ (LLMs) role-playing capabilities suffers from high manual annotation costs and significant biases in automated metrics. To address this, we propose RPEval—the first multidimensional benchmark specifically designed for role-playing evaluation—systematically defining and quantifying four core dimensions: emotional understanding, decision-making reasoning, moral alignment, and role consistency. Methodologically, RPEval constructs multi-turn dialogue tasks grounded in real-world scenarios and integrates expert annotation, adversarial testing, and consistency verification to enable dual-track assessment via automated scoring and human arbitration. The framework ensures reproducibility, extensibility, and human-AI collaborative validation. We publicly release the dataset and implementation code, and conduct baseline evaluations across mainstream LLMs. Results reveal substantial deficiencies in moral alignment and long-term role consistency—insights previously unattainable due to the absence of standardized benchmarks—thereby establishing the first community-wide evaluation standard for LLM role-playing.
The rapid advancement of large language models (LLMs) outpaces the availability of up-to-date, human-annotated evaluation data, hindering timely and reliable assessment of model capabilities. Method: We propose a task-elicited adaptive evaluation framework grounded in scaffolded evaluation agents. It performs behavioral space search and dynamic task generation over domain-specific corpora to automatically discover high-difficulty, high-discriminative failure cases. The method supports cross-model transfer of challenging instances and incorporates human verification to ensure validity. Contribution/Results: Applied to legal reasoning, predictive modeling, and online harassment detection, the framework uncovers systematic inconsistency flaws in state-of-the-art LLMs. The resulting benchmark exhibits strong generalizability across models with diverse capability profiles, enabling high-quality, sustainable, domain-specific evaluation—a novel paradigm for robust LLM assessment.
This study investigates the mechanisms influencing human–LLM judgment alignment in human–AI collaborative evaluation, focusing on how task characteristics and AI assistance strategies shape users’ construction and dynamic refinement of evaluation criteria, as well as their model selection behavior. Method: We conducted a controlled human–AI interaction study involving 15 ML practitioners performing 131 real-world evaluation tasks, comparing direct assessment versus pairwise comparison paradigms, augmented by multi-round LLM-assisted judgments and qualitative behavioral analysis. Contribution/Results: We present the first empirical evidence that direct assessment significantly enhances user engagement and criterion-task alignment, facilitating personalized criterion customization, dynamic judgment adjustment, and adaptive model switching. Based on these findings, we propose design principles for front-end evaluation tools tailored to human–AI collaboration. Our work advances low-overhead, interpretable, and task-adaptive AI-assisted evaluation frameworks.
This study systematically evaluates the reliability and human-expert consistency of large language models (LLMs) as automated evaluators for NLP tasks. Method: Leveraging a multi-task human-annotated benchmark, we assess zero-shot and few-shot LLM-based evaluation across dimensions—including fluency, factual consistency, and logical coherence—in story generation and mathematical reasoning, employing consistency metrics and fine-grained error attribution. Contribution/Results: We uncover, for the first time, systematic biases in LLMs’ generation of evaluation criteria. To address this, we propose “pre-writing human–AI collaborative evaluation”: LLMs first generate structured, criterion-grounded rationales, which humans then calibrate. Experiments show this paradigm improves human evaluation objectivity by 27% and substantially mitigates subjectivity and outlier annotations. While LLMs achieve near-human performance on general dimensions (e.g., fluency), they remain significantly lagging on complex, quantitative reasoning criteria.
This work addresses the inefficiency and poor reproducibility of large language model (LLM) evaluation, which often relies on labor-intensive manual processes such as benchmark selection, code reproduction, and metric interpretation. To overcome these limitations, we propose the first agent-driven automated evaluation system that translates natural language evaluation requests into end-to-end executable, traceable, and customizable evaluation workflows. The system leverages NL2Bench for intent parsing and benchmark planning, BenchResolve for standardized data acquisition, and integrates task-aware metric selection with decision-oriented report generation. Human-in-the-loop checkpoints and a sample evidence chain mechanism are introduced to significantly enhance transparency, controllability, and reproducibility. In industrial settings, the framework enables efficient execution of diverse evaluation tasks with minimal human intervention.
Existing NLP evaluation metrics struggle to effectively assess role-playing large language models (LLMs) in terms of character consistency, logical coherence, and long-term narrative stability. To address this gap, this work proposes RPA-Check, a four-stage automated evaluation framework that decomposes evaluation dimensions, generates Boolean checklists, performs semantic deduplication and isolation, and integrates a chain-of-thought-enhanced LLM-as-a-Judge mechanism. This framework establishes the first structured, reproducible benchmark specifically designed for role-playing agents. Experimental results from the LLM Court forensic training game reveal that instruction-finetuned small models (8–9B parameters) outperform larger counterparts in procedural consistency, suggesting an inverse relationship between model scale and role consistency—thereby challenging the prevailing “bigger is better” paradigm in LLM development.
This study addresses the lack of systematic evaluation regarding the reliability and alignment with human judgment of large language models (LLMs) when deployed as automated evaluators. The authors construct a human-annotated gold-standard dataset spanning eight distinct tasks and conduct the first large-scale empirical analysis of 37 open- and closed-source conversational LLMs under five裁判 prompting strategies, a two-stage judging mechanism, and task-specific fine-tuning. Results demonstrate that GPT-4o, open-source models with at least 32 billion parameters, and Qwen2.5-14B achieve high agreement with human judgments when paired with appropriate prompts, thereby validating the feasibility of using LLMs as reliable automated evaluators. The findings offer empirical guidance for prompt design, model selection, and architectural optimization in automated assessment systems.
This work proposes a fine-grained evaluation framework to assess the capability of large language models (LLMs) as relevance judges in information retrieval, extending beyond holistic document-level judgments to identify the specific textual spans that support those judgments. Leveraging a Wikipedia test collection derived from INEX, the study employs prompt engineering to guide LLMs in simultaneously performing document-level relevance assessment and span-level annotation, followed by comparative analysis against human annotations. By introducing fine-grained relevance evaluation into the LLMs-as-Judges paradigm, this research is the first to examine whether models are “right for the right reasons,” thereby substantially enhancing the credibility of automated evaluation. Experimental results demonstrate that, under human supervision, LLMs can accurately identify both relevant documents and the key evidence spans within them.
To address severe test-set contamination, poor adaptability of static benchmarks, and limited capacity for dynamic task generation in LLM evaluation, this paper proposes a fully automated, contamination-resistant distributed evaluation framework. Methodologically, it introduces a multi-agent mutual-evaluation mechanism wherein models alternately assume the roles of “task generator” and “evaluator”; task generation and closed-loop assessment are enabled via cyclic weighting, consensus-based aggregation, and iterative reliability calibration. Crucially, the framework eliminates reliance on fixed test sets and enhances robustness and human alignment through collaborative judgment by multiple evaluators. Empirical results show Pearson correlations of 78% with human ratings on MMLU-Pro and 63% on GPQA—substantially outperforming single-evaluator baselines—demonstrating both effectiveness and strong generalization across diverse reasoning-intensive benchmarks.