llm evidence scoring

Design and implement functions or modules that convert LLM-generated text and related language signals into quantitative evidence or confidence scores. This includes algorithms to map responses to numeric scores, calibrate and fuse language-derived scores with behavioral or temporal indicators, and adjust scores for anomalies or temporal context.

llmevidencescoring

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Accurately extracting UK Research Excellence Framework (REF) ratings (1*–4*) from noisy, unstructured text containing missing or invalid values presents a significant challenge, requiring large language models (LLMs) to output only normalized integers (1–4) or a designated missing-value indicator (−1). To address this, this work introduces the first standardized prompt engineering benchmark for complex numerical extraction tasks, accompanied by a publicly available dataset of 1,446 short texts with gold-standard annotations. By integrating semantic understanding with explicit rule-based constraints, an initial prompting strategy achieves 72.6% accuracy. The study clarifies the definition of valid ratings and formalizes a mechanism for handling missing data, thereby advancing research into LLMs’ numerical reasoning and instruction-following capabilities, with the aim of fostering community-driven improvements in structured information extraction from noisy textual sources.

information extractionlarge language modelsmessy text

This study addresses the challenge of error-prone manual verification of tables, figures, and listings (TFLs) in clinical trial reports, which often fails to detect structural or logical inconsistencies. The authors propose PROVE, a novel framework that leverages large language models (LLMs) for semantic parsing and evidence tracing of TFL content, integrated with a programmable rule engine to perform deterministic numerical and logical validation against SDTM/ADaM standards. Designed as a multi-agent architecture, PROVE combines LLM-driven semantic understanding with rule-based checks to enable auditable, configurable automated cross-verification. The approach achieves 100% accuracy under exact label matching; when confronted with linguistic variations, LLM assistance boosts recall from 0.588 to 0.993 and F1 score from 0.735 to 0.996.

clinical trial reportingcross-output consistencyregulatory compliance

Detecting LLM-Generated Short Answers and Effects on Learner Performance

Jun 20, 2025
SB
Shambhavi Bhushan
🏛️ Carnegie Mellon University | Learning Engineering Virtual Institute

Current LLM misuse detection tools in online learning suffer from low reliability, lack of standardized evaluation criteria, and insufficient understanding of educational implications. Method: This study proposes actionable, interpretable criteria for identifying LLM-generated text and introduces a novel multidimensional detection paradigm based on fine-tuned GPT-4o, integrating statistical indicators—including anomalously high scores, readability metrics, and response latency. Results: The method achieves 80% accuracy (F1 = 0.78) on short-answer detection, substantially outperforming GPTZero (70%, F1 = 0.50); robustness is confirmed via human coding and comparative evaluation. Empirical analysis further reveals that students misusing LLMs exhibit abnormally elevated post-test accuracy—indicating superficial engagement and bypassing of deep learning processes. This work constitutes the first systematic integration of detection methodology, interpretable decision criteria, and pedagogical impact analysis, providing both a methodological framework and empirical foundation for AI governance in education.

Assess learning performance impact from LLM misuseCompare accuracy of existing LLM detection methodsDetect LLM-generated short answers in online learning

How to Correctly Report LLM-as-a-Judge Evaluations

Nov 26, 2025
CL
Chungpa Lee
🏛️ Yonsei University | University of Wisconsin–Madison

Large language models (LLMs) used as evaluators suffer from estimation bias and noise due to insufficient specificity and sensitivity, undermining the statistical reliability of automated evaluation. Method: We propose the first practical framework that corrects such bias and constructs statistically rigorous confidence intervals. It introduces a plug-in bias-correction mechanism that jointly models uncertainty over both test and calibration sets, coupled with an adaptive sampling algorithm to optimize calibration sample allocation. Leveraging estimated specificity and sensitivity, the framework employs statistical inference to derive bias-corrected confidence intervals. Contribution/Results: Our approach significantly reduces both bias and variance in accuracy estimation. Extensive evaluation across multiple benchmarks demonstrates its robustness, reliability, and generalizability. The framework establishes a new paradigm for LLM-based automated evaluation—reproducible, interpretable, and statistically trustworthy.

Constructing confidence intervals with imperfect specificity and sensitivity estimatesCorrecting biased accuracy estimates in LLM-as-a-judge evaluationsReducing uncertainty through adaptive calibration sample allocation

Enhancing LLM Evaluations: The Garbling Trick

Nov 03, 2024
WF
William F. Bradley
🏛️ Mirabolic Consulting

Conventional LLM evaluation metrics exhibit saturation effects, failing to discern subtle differences—particularly in reasoning capabilities—among state-of-the-art models. Method: We propose the “Gibberish Technique”, a scalable evaluation augmentation paradigm that transforms original tasks into a family of progressively challenging, reasoning-oriented multiple-choice questions. This is achieved through semantics-preserving input perturbations and multi-level difficulty construction. Contribution/Results: The technique uncovers previously masked capability gradients between base models and specialized reasoning models—revealing distinctions invisible under standard benchmarks. Experiments across multiple mainstream LLMs demonstrate that our augmented evaluation significantly improves discriminative power, accurately characterizing hierarchical reasoning competencies. It establishes a more sensitive and diagnostically informative benchmark for LLM assessment, enabling fine-grained differentiation where traditional metrics fall short.

Comparing base LLMs and reasoning models effectivelyRevealing performance differences not seen in original assessmentsTransforming LLM evaluations into progressively harder tasks

Latest Papers

What's happening recently
View more

This study addresses the threat posed by the stochasticity of large language model (LLM) outputs to research reproducibility, demonstrating that randomness arises not only from sampling strategies but also from non-sampling factors such as silent model updates, numerical rounding, and expert routing in mixture-of-experts architectures. The work is the first to explicitly model LLM outputs as draws from a probability distribution and systematically evaluates the impact of various sources of randomness on downstream outcomes through regression analysis in a sentiment classification task. Empirical comparisons across multiple software and hardware environments—using both API-accessible and locally deployed open-source models—reveal substantial output variability even at zero temperature. Building on these findings, the paper proposes standardized reporting guidelines for research papers and reproduction packages to encourage the adoption of stricter LLM usage protocols within the community.

large language modelsLLM outputsrandomness

This work addresses the challenges of behavioral drift, miscalibrated uncertainty, and declining trustworthiness in large language models (LLMs) under continuous deployment, primarily caused by inadequate modeling and monitoring of temporal interactions and dynamic feedback. To tackle this, the paper introduces, for the first time, a sequential statistical inference framework tailored for trustworthy LLM deployment. Built upon dependent stochastic processes, the proposed framework integrates sequential hypothesis testing, change-point detection, calibration assessment, and fairness monitoring. This paradigm provides rigorous uncertainty guarantees in settings involving dependent interactions, repeated usage, and adaptive behavior, while enabling real-time monitoring and early warning for critical attributes such as hallucination rates, refusal patterns, and fairness violations. The approach significantly enhances the stability and reliability of LLMs operating in dynamic environments.

behavioral shiftslarge language modelssequential inference

This study addresses the questionable diagnostic validity of automated scoring by large language models (LLMs), which often appears superficially comprehensive. Leveraging the JorGPT dataset, this work pioneers a diagnostic-validity perspective to deconstruct LLMs’ structured outputs. By integrating statistical analyses, including variance inflation factor (VIF) assessment, with natural language processing techniques, it systematically compares human and machine scoring across multidimensional score redundancy, misconception detection rates, and tone bias. The findings reveal systematic deficiencies in LLM scoring, notably high sub-dimension collinearity, low misconception detection rates, and an absence of severity moderation. Furthermore, the analysis demonstrates that LLMs remain reliable primarily within procedural knowledge domains. Ultimately, this research delineates critical boundaries requiring human oversight for human–AI collaborative assessment in higher education.

Automated GradingDiagnostic QualityFeedback

This study addresses the instability of large language model (LLM) annotations under varying task designs, revealing substantial “instrument uncertainty” that confidence scores fail to mitigate. By evaluating seven LLMs across twelve task designs on a sample of 3,000 tweets, this work systematically quantifies the impact of prompt engineering on annotation outcomes using Fleiss’ and Cohen’s Kappa coefficients. The findings demonstrate that design variations inflate the variance of prevalence estimates by 76- to 110-fold, far exceeding typical inter-annotator disagreement among humans. Furthermore, this research establishes that comparing multiple task designs is the only effective approach for measuring such uncertainty, exposing critical limitations in existing evaluation methodologies and highlighting systematic biases introduced by both model selection and prompt design.

Annotation ReliabilityInstrument UncertaintyLabel Variance

Hot Scholars

SS

Shuzheng Si

Tsinghua University
Natural Language ProcessingLarge Language Models
MS

Maosong Sun

Professor of Computer Science and Technology, Tsinghua University
Natural Language ProcessingArtificial IntelligenceSocial Computing
JW

Jiancan Wu

University of Science and Technology of China
LLMsRecommendationGraph Neural Network
XL

Xiaohan Li

Walmart Inc.
Data MiningRecommender systemMedical AI