grounding validation

Designs and implements evaluation methods, metrics, and test suites that verify system outputs are supported by specified evidence or ground-truth representations, including extracting and matching closed-set evidence units and performing representation-aware checks that answers are supported by those units. This work covers evidence-based validation across modalities (e.g., natural language and visual), checks that queries are self-contained, and measures coverage for multi-hop or procedural grounding and alignment.

groundingvalidation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.3
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$187K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the lack of systematic understanding of reasoning capabilities evaluated in existing fact-checking benchmarks. The authors propose the first decomposition of reasoning components in fact-checking, generating 24K structured reasoning traces using GPT-4o-mini and employing a lightweight 1B-parameter verifier to classify and analyze errors. Their analysis reveals that prevailing benchmarks predominantly assess direct evidence extraction, while multi-sentence synthesis and numerical reasoning remain severely underrepresented. High model performance largely reflects retrieval and textual entailment abilities rather than complex reasoning. Distinct error patterns emerge across domains: general-domain errors stem from lexical overlap bias, scientific-domain errors arise from excessive caution, and mathematical-domain failures result from inadequate arithmetic reasoning. These findings expose critical limitations in current evaluation protocols and offer concrete directions for developing more challenging and comprehensive fact-checking benchmarks.

claim verificationdataset biasevaluation benchmarks

This work investigates large language models’ (LLMs) capabilities in evidence-based claim verification, specifically evaluating deductive versus abductive reasoning. To this end, we introduce RECV—the first benchmark featuring real-world claims with fine-grained, atomic-level annotations of reasoning types—and propose a reasoning-type decomposition evaluation framework. Through systematic assessment of mainstream closed-source LLMs across multiple difficulty levels and prompting strategies, complemented by semantic similarity analysis, we find that: (1) LLMs exhibit robust performance on deductive reasoning but suffer from systematic failures in abductive reasoning; (2) generated explanations achieve high semantic similarity to human-written ones—especially for deductive tasks—but rationalization does not consistently improve verification accuracy. This study provides the first empirical evidence of LLMs’ fundamental limitations in abductive reasoning, establishing a novel, trustworthy benchmark and methodology for rigorous reasoning evaluation.

Assessing deductive vs abductive reasoningLLMs' reasoning in claim verificationRationale generation impact on LLMs

This work addresses the challenge of providing verifiable safety assurance for robotic systems in safety-critical domains, where traditional assurance cases rely on manually generated evidence that is costly, error-prone, and difficult to maintain. The paper proposes a model-based automated approach that deeply integrates formal verification into the assurance workflow. It employs RoboChart—a domain-specific modeling language with formal semantics—to capture system designs, and introduces a template-driven mechanism to automatically translate natural-language requirements into formal assertions. These assertions are then discharged through a combination of model checking and theorem proving tools, yielding formally verified evidence that can be seamlessly integrated into assurance cases. Case studies demonstrate that the proposed method significantly enhances the reliability, maintainability, and degree of automation in safety argumentation.

Assurance CasesFormal EvidenceRobotic Software

Latest Papers

What's happening recently
View more

This work addresses the susceptibility of large language models to hallucinations in empirical reasoning—outputs lacking verifiable evidence or formal guarantees. The authors propose EG-VAR, an architecture that uniquely leverages the Lean 4 formal proof kernel as the sole trusted generator of claims. By integrating tool-certified axioms and a source elevation mechanism, EG-VAR ensures every output is bound to a kernel-verified chain of reasoning and tool invocation; otherwise, it abstains and provides a fully traceable audit trail. Evaluated on a TableBench subset, EG-VAR achieves perfect accuracy (120/120), substantially outperforming a 95% baseline. In counterfactual tests, it maintains 100% source fidelity—significantly higher than competing methods (80–90%)—and exhibits remarkably low semantic formalization error rates of 1.7% (Opus) and 3.3% (Sonnet).

empirical inferenceevidence groundingformal verification

This work addresses the limitation of conventional execution coverage in UI component testing, which fails to verify whether tests adequately capture behavioral relationships implied by APIs and documentation. The paper proposes the first evaluation framework based on inferred metamorphic relations (MRs): it automatically derives MRs using a UI-specific taxonomy from source code and documentation, aligns test executions to these MRs through deterministic and semantic analysis, and introduces relation-level MR coverage as a novel metric. By treating inferred MRs as empirical benchmarks for behavioral validation, the approach exposes verification gaps invisible to traditional coverage metrics—particularly in weak-oracle scenarios. Empirical results across three LLM configurations show MR coverage ranging only from 42.5% to 47.6%, substantially lower than MR reachability; uncovered MRs are predominantly of the weak-oracle type, demonstrating that MR coverage meaningfully complements conventional metrics and offers practical utility in fault detection and issue mapping.

behavioral validationmetamorphic relationstest coverage

This work addresses a critical limitation in existing cloud skill testing, which focuses solely on task success rates and fails to reveal uncovered behaviors, leaving test adequacy unquantifiable. To bridge this gap, the study introduces the first formal definition of test coverage units, coverage relationships, and a computational pipeline for cloud skills. It proposes a coverage evaluation method grounded in natural language skill packages: by parsing user prompts and initial resource states, the approach reconstructs operational obligations and models workflow context to establish an end-to-end measurement pipeline. Integrating model-assisted candidate generation with expert review, the framework produces auditable coverage reports and closes the loop by mapping coverage gaps to source-level test improvement recommendations. Empirical results demonstrate that this methodology enables quantitative assessment of test coverage for production-grade cloud skills, substantially enhancing their reliability and observability.

AI AgentsCloud SkillsSkill Evaluation

This work addresses the challenge of providing auditable, high-assurance verification for AI-generated informal reasoning while maintaining broad coverage. It proposes a structured approach that reformulates solutions as typed state-transition sequences, where every step must be explicitly justified—via citations, computations, or given premises—and enforces a “change completeness” invariant to surface hidden assumptions and fabricated references. By integrating explicit justification licensing, human-readable proof traces, and an adversarial validation framework, the method enables transparent, traceable, and contestable reasoning. Empirical evaluation demonstrates strict certification accuracies of 91.4% on HLE-Verified Gold and 97.1% on GPQA Diamond, significantly outperforming existing monolithic LLM-based evaluators in detecting concealed premises and hallucinated citations.

AI reasoningauditabilityinformal reasoning

Hot Scholars

HR

Hongliang Ren

Chinese University of Hong Kong | National University of Singapore | JHU/Harvard(RF) | CUHK(PhD)
Biorobotics & intelligent systemsmedical mechatronicscontinuumsoft flexible robots/sensors
LB

Long Bai

Research Assistant, Institute of Computing Technology, Chinese Academy of Sciences
Event-Centric AnalysisKnowledge GraphNatural Language Processing
ML

Manling Li

Assistant Professor at Northwestern University
Natural Language ProcessingVision-LanguageEmbodied Agents
PT

Philip Torr

Professor, University of Oxford
Department of Engineering
LK

Lingdong Kong

National University of Singapore
Computer VisionDeep Learning