build grounded explanation dataset

Design and build datasets of examples annotated with grounded, natural-language explanations that are explicitly aligned to model evidence or attribution traces; produce structured annotations (including partial-spoof or perturbed grounding labels) and metadata to support human evaluation, faithfulness checks, and inside-accuracy benchmarking of explanations.

buildgroundedexplanationdataset

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Large Language Model Sourcing: A Survey

Oct 11, 2025
LP
Liang Pang
🏛️ State Key Laboratory of AI Safety | Institute of Computing Technology, Chinese Academy of Sciences | University of Chinese Academy of Sciences | Gaoling School of Artificial Intelligence | Renmin University of China

To address the challenges of tracing provenance and ensuring credibility of large language model (LLM)–generated content, this paper proposes the first holistic four-dimensional provenance framework integrating both model- and data-centric perspectives: model origin identification, architectural and mechanistic analysis, training data attribution, and external information verification. We introduce a novel “prior–posterior” dual-paradigm classification system and unify techniques including model fingerprinting, response-level verification, and traceability-aware embedding to support both proactive and reactive reasoning. The framework systematically consolidates fragmented provenance research efforts, significantly enhancing the explainability, verifiability, and transparency of AI-generated content. It establishes a theoretical foundation and scalable technical methodology for detecting AI-generated content (AIGC), identifying model identities, and ensuring information reliability.

Addressing hallucinations and bias through multi-perspective sourcingDeveloping traceability methods for model structure and training dataTracking provenance of LLM-generated content to enhance transparency

Must-Read Papers

Most classic and influential ideas
View more

Evaluating Evidence Attribution in Generated Fact Checking Explanations

Jun 18, 2024
RX
Rui Xing
🏛️ The University of Melbourne | MBZUAI

Evidence misattribution—frequent in fact-checking explanation generation—leads to hallucinations and low credibility. To address this, we propose the Citation Masking and Recovery (CMR) evaluation protocol, the first quantifiable framework for assessing evidence attribution quality. Leveraging large language models (LLMs) for automated annotation, crowdsourced human evaluation, and controlled comparative experiments, we find that state-of-the-art LLMs still exhibit substantial attribution error rates. However, LLM-generated attributions align strongly with human annotations (Spearman ρ > 0.85), validating their efficacy as scalable proxies for human assessment. Crucially, our experiments demonstrate that human-curated evidence selection significantly improves both explanation accuracy and interpretability. This work establishes a novel, empirically grounded evaluation standard for trustworthy explanation generation in fact-checking systems.

Assessing attribution quality using citation masking.Evaluating evidence attribution in fact-checking explanations.Identifying inaccuracies in LLM-generated explanations.

This study addresses the high variability in human annotations for subjective natural language processing tasks, which often arises from semantic ambiguity, and investigates the unclear impact of large language model (LLM)-generated reasoning on human annotation behavior. The authors propose ReasonAlign, a two-round Delphi-style annotation protocol that exposes annotators solely to LLM-generated rationales while withholding predicted labels, thereby isolating the effect of explanatory reasoning on inter-annotator agreement and label revision. Introducing AEP (Annotator Effort Proxy) as a novel metric to quantify the extent of annotation revisions, experiments on sentiment classification and opinion detection demonstrate that exposure to LLM rationales significantly improves annotation consistency with only minimal label changes, suggesting that such reasoning primarily aids in resolving ambiguous cases.

annotation variabilityhuman annotationhuman-AI co-annotation

This work addresses a critical limitation in conventional large language model (LLM) evaluation, which treats benchmark datasets as homogeneous aggregates and overlooks the heterogeneity among samples in cognitive, linguistic, and task-related attributes. The authors propose a dataset-centric meta-evaluation framework that introduces fine-grained, sample-level annotations across five dimensions: cognitive demand, language quality, task characteristics, contextual dependency, and ethical safety. For the first time, this approach enables multidimensional auditing of widely used benchmarks such as MMLU and ARC. By allowing dynamic subset composition aligned with specific evaluation objectives, the framework uncovers the diversity obscured by aggregate accuracy metrics and establishes a composable evaluation paradigm tailored to targeted capabilities—such as reasoning depth or ethical sensitivity—thereby substantially enhancing the precision and interpretability of LLM assessments.

benchmark heterogeneitydataset introspectionevaluation bias

Manual annotation of textual explanations for interpretable NLP is costly and inherently unscalable. Method: We propose a multi-LLM collaborative framework for automated explanation generation to enhance natural language inference (NLI) classifiers. It integrates outputs from multiple state-of-the-art large language models to produce high-quality, faithful reasoning rationales; employs NLG evaluation metrics to assess explanation quality; and fine-tunes downstream NLI classifiers—specifically on the SNLI and MNLI benchmarks—using these generated explanations as auxiliary supervision. Contribution/Results: Explanations automatically generated by LLMs significantly improve pre-trained NLI model performance, matching the efficacy of human-annotated explanations. This work provides the first empirical validation of the effectiveness and scalability of *automatically generated* explanations for model enhancement. By eliminating reliance on manual annotation, it establishes a novel, scalable paradigm for interpretable NLP.

Assessing impact of automated explanations on model task performanceAutomating textual explanations to replace costly human annotationsEvaluating if LLM-generated explanations improve classification performance

LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Jun 26, 2024
AB
Anna Bavaresco
🏛️ University of Amsterdam | University of Trento | University of Copenhagen | Utrecht University | ETH Zürich | Saarland University | Universidade de Lisboa | LMU Munich | University of Potsdam | Heriot-Watt University | Unbabel | MCML

This study investigates whether large language models (LLMs) can reliably replace human annotators for evaluating NLP models. Method: We introduce JUDGE-BENCH—the first large-scale, multi-task, multi-dimensional automatic evaluation benchmark with high-quality human annotations—and systematically assess the effectiveness and consistency of 11 state-of-the-art LLMs as automatic evaluators across 20 NLP tasks. Our methodology integrates human annotation quality analysis, statistical significance testing, and cross-model correlation metrics (Kendall’s τ and Spearman’s ρ). Contribution/Results: LLM-based evaluation performance is highly contingent on evaluation attributes, annotator expertise level, and text source; while LLMs approximate human judgments in certain tasks, they lack universal reliability. Human annotations remain indispensable as the gold standard for pre-validation. We publicly release JUDGE-BENCH—including all human annotations, model outputs, and evaluation scripts—to advance standardized, reproducible research on LLM-based evaluation.

Assessing validity of LLMs replacing human judges in NLP evaluationsEvaluating reproducibility of proprietary vs open-weight LLM modelsMeasuring variance in LLM performance across diverse NLP tasks

Latest Papers

What's happening recently
View more

This work addresses the lack of traceability in existing domain-specific fine-tuning approaches, which often leads to blind and inefficient data augmentation. The authors propose a “programming with data” paradigm that treats structured knowledge representations as a unified foundation for both training and evaluation, drawing an analogy to software development: training data serve as source code, model training as compilation, evaluation as unit testing, and data refinement as debugging. This framework enables precise, concept- and reasoning-chain–oriented model repair through structured knowledge extraction, test-driven data engineering, concept-level gap analysis, and diagnosis of broken reasoning chains. Validated across 16 disciplines, the approach significantly enhances model performance without compromising general capabilities, and the authors release an open-source knowledge base, evaluation suite, and training corpora to support reproducibility and further research.

data engineeringdomain specializationknowledge transfer

This study addresses the widespread problem of incomplete reporting of annotation practices in natural language processing (NLP) research, which undermines reproducibility and quality assessment. Analyzing 1,603 papers from major NLP conferences between 2018 and 2025, the work introduces a unified taxonomy for annotation reporting that spans tasks, time, and domains, along with a minimal reporting standard. Leveraging a gold-standard dataset—Annotated-gold—curated through a combination of large language models and human adjudication, the authors construct Annotated-llm, achieving human-level inter-annotator agreement (Krippendorff’s α = 0.606) on structured information extraction. Despite gradual improvements in reporting over time, critical details—such as annotator training, linguistic competence, and compensation—remain frequently omitted. These findings advance the push toward more transparent and reliable annotation practices in NLP.

annotation reportingannotation validityhuman annotation

This study addresses the limitations of prevailing binary support/refutation frameworks in evaluating AI-generated text, which fail to capture the nuanced semantic relationships between generated content and source documents. Moving beyond conventional groundedness paradigms, the work proposes a reader-centered, fine-grained taxonomy of evidential relations by integrating insights from linguistics and philosophy of language, encompassing diverse linkage types such as syntactic rephrasing and inferential strategies. Through theoretical analysis, a human annotation protocol, and benchmark evaluations, the authors systematically demonstrate the feasibility and efficacy of this framework. The resulting approach offers a more transparent and interpretable provenance mechanism for AI outputs, establishing both theoretical foundations and practical pathways for fine-grained evaluation and explainable interfaces in natural language generation systems.

generative AIgroundednesshallucination

This work proposes a graph neural network–based approach to support compliance and safety verification by analyzing the structural soundness and source authenticity of assurance cases. For the first time, assurance cases are systematically modeled as textual attributed graphs, enabling a joint learning framework that performs link prediction—assessing structural coherence—and human–machine generated case classification—discriminating origin authenticity. The method reveals significant differences in hierarchical linkage patterns between large language model–generated cases and those authored by humans. Experimental results on real-world datasets demonstrate the effectiveness and novelty of the proposed framework, achieving a ROC-AUC of 0.760 for link prediction and an F1 score of 0.94 for human–machine case classification.

assurance casesgraph classificationlink prediction

Hot Scholars

XS

X. Sean Wang

School of Computer Science, Fudan University
Database SystemsInformation Security and PrivacyWireless Sensor NetworksStreaming Data Processing Time Series Queries
MV

Michalis Vlachos

Professor, HEC Lausanne, University of Lausanne
Recommender SystemsAI in EducationTime Series MiningDigital Watermarking
MN

Mohammad Noorchenarboo

Ph.D. Candidate in Department of Electrical and Computer Engineering, Western University
Artificial intelligenceExplainabilityBiostatisticsBioinformatics