faithfulness evaluation

Designs and implements methods, metrics, and procedures to measure and verify how faithfully a system's outputs reflect a defined source of truth or formal specification. This includes building automatic faithfulness metrics and scoring systems, statistical and logical consistency-checking and analysis tools, protocol- and proof-based verification techniques, and diagnostics that detect contradictions, quantify agreement, and prioritize corrective updates.

faithfulnessevaluation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.43
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$185K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Current approaches to generating Lean formal statements from natural language rely solely on compilation success rates, which often fail to ensure semantic faithfulness, leading to issues such as omitted premises, incorrect domain specifications, or trivial propositions. This work proposes a novel evaluation paradigm centered on “consensus faithfulness,” introducing a benchmark dataset of 400 graduate-level mathematical problems. Through an integrated methodology combining Lean compilation checks, multi-model semantic evaluations, expert calibration, and a 2³ factorial experiment, the study systematically identifies key factors affecting faithfulness. Experiments reveal that even the best tool-augmented agent achieves only 60.5% consensus faithfulness despite an 89.5% compilation rate, and manual auditing confirms this metric’s high conservativeness, highlighting a substantial gap between syntactic compilability and genuine semantic accuracy.

faithful formalizationformal statement generationnatural-language-to-Lean

Despite its efficacy in isolated projects, deductive verification has yet to achieve broad industrial adoption. To identify root barriers and key enablers, this paper conducts semi-structured interviews with 30 practitioners, followed by thematic analysis. We systematically uncover fundamental obstacles—including high proof maintenance overhead, limited automation, poor tool usability, and lack of workflow integration—as well as critical enabling factors. Diverging from prior work, we empirically establish *usability* and *workflow adaptability* as core dimensions governing adoption. Based on these findings, we propose three actionable improvement principles: (1) enhancing automation support for proof construction and evolution; (2) reducing proof maintenance burden through modularization and abstraction; and (3) deepening integration with IDEs and CI/CD pipelines. Our empirically grounded insights provide concrete, evidence-based guidance for tool developers, practitioners, and researchers—bridging the gap between academic verification techniques and engineering practice.

Developing recommendations for practitioners and tool builders to improve adoptionIdentifying underexplored obstacles like proof maintenance and usability issuesInvestigating barriers to mainstream adoption of deductive verification methods

This work addresses the challenge that counterexamples generated by formal verification often consist of numerous low-level Boolean variables, rendering them difficult for developers to interpret at the application-domain level. To bridge this gap, the paper proposes a novel hierarchical explanation method that integrates predicate relevance metrics with dependency graph analysis—a first-time fusion of these two techniques—to automatically extract human-readable, domain-oriented explanations from logical formulas. By leveraging formal modeling and a dedicated explanation-generation algorithm, the approach produces concise and semantically clear descriptions of failure causes across multiple case studies. Empirical results demonstrate that the method significantly outperforms existing techniques, offering effective support for fault localization in practical verification tasks.

application domain modelcounterexample interpretationformal verification

Tool-Assisted Conformance Checking to Reference Process Models

Aug 01, 2025
BR
Bernhard Rumpe
🏛️ RWTH Aachen University

Existing conformance checking approaches between process models and reference models suffer from limited semantic expressiveness and insufficient automation, hindering fine-grained compliance verification. This paper proposes a semantic consistency checking method grounded in causal dependency analysis of tasks and events, transcending traditional trajectory-based dependency modeling by formally encoding causal constraints at the semantic level. We establish a unified framework integrating causal dependency modeling, semantic representation, and formal verification, and design an automated conformance checking algorithm implemented in a prototype tool. Empirical evaluation demonstrates that our approach significantly outperforms state-of-the-art techniques in both accuracy and flexibility, achieving— for the first time—the fully automated, high-expressivity semantic conformance verification of process models against reference models.

Automated conformance checks for process models against reference modelsEnhancing accuracy and flexibility in process model conformance verificationLack of expressiveness and automation in semantic model comparison

Validating Network Protocol Parsers with Traceable RFC Document Interpretation

Apr 25, 2025
MZ
Mingwei Zheng
🏛️ Purdue University | Nanjing University

Addressing the “oracle absence” and “error attribution difficulty” challenges in network protocol parser verification, this paper proposes an LLM-driven framework for RFC semantic parsing and feedback-based oracle refinement. First, large language models automatically translate unstructured RFC text into formal message specifications. Second, an iterative, quasi-oracle is constructed to support specification-guided fuzz testing and cross-language (C/Python/Go) protocol implementation verification. Finally, vulnerabilities are precisely traced back to their originating RFC clauses. This work is the first to integrate LLM-based semantic understanding with dynamic oracle refinement. Evaluated on nine mainstream protocols, it discovers 69 vulnerabilities—36 of which have been confirmed—surpassing state-of-the-art approaches in both effectiveness and efficiency. It also demonstrates, for the first time, the feasibility of fully automated derivation of test oracles directly from natural-language protocol specifications.

Addressing oracle and traceability issues in protocol validationAutomating software validation via LLM-based specification translationValidating network protocol parsers using RFC documents

Latest Papers

What's happening recently
View more

Current AI-assisted scientific writing lacks auditable generation processes and mechanisms for accountability, undermining the verifiability of research credibility and compliance. This work proposes a novel auditing paradigm embedded directly within the production workflow, enforcing end-to-end traceability, immutability, and third-party reproducibility of AI involvement through preregistered blind-spot indicator cards, sealed execution environments, and automated gatekeeping intercepts. Core technical components include Git-sealed lineage anchoring, hash-bound provenance tracking, red-flag interception protocols, cross-model role isolation, and programmatic assembly. In experimental validation, one project was automatically terminated when preregistered confirmatory tests triggered a No-Go decision. An open-source toolkit is released to enable independent recomputation of all core audit metrics by third parties.

AI AccountabilityAuditable AIProvenance

This study addresses the growing challenge posed by the widespread involvement of AI agents in software development, which undermines the long-standing assumption that development artifacts are exclusively produced by human professionals—an assumption underpinning traditional software metrics. The work systematically exposes how AI-generated traces compromise the foundational premises of established software measurement practices, thereby threatening the validity of prior empirical conclusions. To confront this issue, the authors propose an AI-augmented, systematic replication methodology that integrates modern data analytics with empirical software engineering techniques to rigorously re-evaluate key findings. The project advances a dynamic, reproducible, and sustainable measurement paradigm capable of adapting to evolving data ecosystems, offering a robust and timely framework for software metrics in the AI era.

AI agentsfoundational assumptionsreplication

This work addresses critical reliability issues in existing Lean theorem-proving benchmarks, where inconsistencies between formal statements and informal problem descriptions, along with susceptibility to trivial or adversarial solutions, undermine evaluation validity. To tackle this, the study introduces the first fault taxonomy for formal mathematical datasets and develops an automated auditing toolkit integrating static program analysis, formal verification, semantic auditing via prompt engineering, and manual validation. Applying this framework, the authors conduct a large-scale audit of five prominent benchmarks, uncovering 4,833 issues—including 398 severe defects—and demonstrate that uncorrected flaws significantly distort prover rankings. The paper releases both the auditing tools and corrected dataset snapshots to foster reproducible and trustworthy evaluations in theorem proving.

dataset defectsevaluation reliabilityformal verification

This study addresses the challenges of applying formal verification to production-grade software, where high modeling costs and consistency risks in fault handling hinder adoption. By integrating runtime execution traces with formal specifications, the authors verify a real-world restaurant point-of-sale (POS) payment workflow and leverage large language models (LLMs) to automatically generate these specifications. Their analysis reveals that the structural form—not the natural language phrasing—of specifications primarily governs LLM-generated correctness, and uncovers a shared “relevant oracle failure” issue between code and simulators. Extending fault models to include crash-recovery, stale reads, and retries, the team conducts simulation-based audits, verifying core protocol correctness, identifying and reproducing seven fault-handling vulnerabilities, and revalidating after fixes. They also expose how deviations in API response structures render recovery paths unreachable—a finding consistently replicated across seven LLMs from two vendors.

failure handlingformal verificationpayment workflow

This work addresses the problem of global inconsistency in multi-component intelligent agent releases, where local validation passes but cross-component relational integrity fails due to the absence of holistic consistency guarantees. To tackle this, we propose the Schema-SIP Relational Consistency (SIP-RC) framework—the first systematic approach to formally define and mitigate relational inconsistency faults in multi-component deployments. SIP-RC models release packages as graph structures and integrates schema documentation with product contract principles to enable cross-component relational verification. Key mechanisms include declarative–evidential linkage, decision authority scoping, provenance tracking of derived components, and byte-level consistency checks. Preliminary experiments demonstrate the feasibility of the proposed framework, offering a practical and actionable paradigm for ensuring relational consistency in intelligent agent releases.

Agent SystemsMulti-Artifact ReleasesPackage Consistency

Hot Scholars

YZ

Yutao Zhong

Department of Mathematics, Courant Institute of Mathematical Sciences
Machine LearningMathematics
MM

Mehryar Mohri

Head, ML Theory, Google Research; Professor, Courant Institute of Mathematical Sciences.
Machine Learning
WZ

Wentao Zhang

Institute of Physics, Chinese Academy of Sciences
photoemissionsuperconductivitycupratehtsc
LZ

Linfeng Zhang

DP Technology; AI for Science Institute
AI for Sciencemulti-scale modelingmolecular simulationdrug/materials design
YL

Yuyu Luo

Assistant Professor, HKUST(GZ) / HKUST
Data AgentsLLM AgentsDatabaseText-to-SQL