claim-evidence alignment

Designs and builds methods, representations, and tools that decompose assertions into atomic claims and align each claim to supporting or contradicting evidence — including passages, table cells, charts, and attribute-structured records — across modalities. Produces structured outputs and visual analytics (e.g., claim-evidence matrices and evidence visualizations) that summarize support composition and confidence, expose coverage gaps, contradictions, and modality imbalances, and enable coordinated claim-level inspection and auditing.

claim-evidencealignment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.7
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$207K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the instability of existing claim decomposition approaches in verifying complex, multifaceted claims, which stems primarily from insufficient evidence alignment and inadequate modeling of error patterns in sub-claims. The authors introduce a new dataset featuring temporally constrained evidence and human-annotated evidence spans for sub-claims, enabling a systematic evaluation of claim decomposition under two evidence alignment settings. They reveal, for the first time, the critical impact of fine-grained evidence alignment and label bias in sub-models on verification performance, and propose a structured decomposition framework that contrasts sub-claim aligned evidence (SAE) with repeated claim-level evidence (SRE). Experiments demonstrate that decomposition significantly improves performance only under strictly aligned, fine-grained evidence conditions, and that incorporating an “abstain” strategy effectively mitigates error propagation, with consistent results across multiple datasets.

claim verificationerror propagationevidence alignment

Large language models frequently generate assertions in financial question answering that blend well-supported claims, weakly reasoned statements, and unsupported conclusions, thereby compromising the reliability of high-stakes decisions. This work addresses the issue by framing it as an assertion–evidence alignment task and introduces a multimodal assertion–evidence matrix that systematically links textual, tabular, and graphical evidence through a lightweight alignment pipeline and standardized JSON evidence artifacts. Integrating a deterministic audit-prioritization algorithm with an interactive visualization interface, the proposed approach significantly enhances analysts’ ability to discern reliable assertions from overconfident or unsubstantiated content in financial statement audits, outperforming conventional chat-based information presentation methods.

auditabilityclaim-evidence alignmentfinancial question answering

Format Matters: The Robustness of Multimodal LLMs in Reviewing Evidence from Tables and Charts

Nov 13, 2025
XH
Xanh Ho
🏛️ National Institute of Informatics | National Taiwan University | The University of Tokyo | Inria | LS2N | Nantes Université

This study systematically evaluates the robustness of multimodal large language models (MLLMs) in verifying scientific claims using tabular and graphical evidence. We conduct cross-format experiments on an adapted scientific paper dataset with 12 state-of-the-art MLLMs and include human expert benchmarking for rigorous comparison. Results reveal that current MLLMs achieve strong performance on table understanding but exhibit substantial deficiencies in chart comprehension, coupled with poor cross-modal generalization; in contrast, human experts maintain high accuracy across both modalities. To our knowledge, this work provides the first quantitative characterization of structural weaknesses in MLLMs’ chart semantic parsing—highlighting the critical need to enhance chart understanding for trustworthy, AI-assisted scientific review. Our findings deliver empirical grounding for future model development, identifying concrete limitations and actionable directions for improving multimodal reasoning in scientific domains.

Assessing performance gap between table-based and chart-based evidence verificationEvaluating robustness of multimodal LLMs in verifying scientific claims using tables and chartsInvestigating limited cross-modal generalization in smaller multimodal LLMs

This work addresses the limitations of existing fact-checking approaches, which are often confined to unimodal text or lack interpretability, thereby struggling to verify claims requiring joint visual and textual evidence. The authors propose a two-tier multimodal graph architecture that enables fine-grained evidence retrieval through a bidirectional image-text reasoning mechanism. Multimodal information is fused at both token and evidence levels to support claim verification, while a dedicated fusion decoder generates natural language explanations—realizing, for the first time, an integrated pipeline for retrieval, verification, and explanation. Key contributions include a novel multi-granularity fusion strategy, the bidirectional reasoning mechanism, and AIChartClaim, the first multimodal claim dataset centered on scientific figures in the AI domain. Experiments demonstrate that the proposed method significantly outperforms current baselines on multimodal claim verification tasks.

claim verificationevidence retrievalexplainability

This work investigates large language models’ (LLMs) capabilities in evidence-based claim verification, specifically evaluating deductive versus abductive reasoning. To this end, we introduce RECV—the first benchmark featuring real-world claims with fine-grained, atomic-level annotations of reasoning types—and propose a reasoning-type decomposition evaluation framework. Through systematic assessment of mainstream closed-source LLMs across multiple difficulty levels and prompting strategies, complemented by semantic similarity analysis, we find that: (1) LLMs exhibit robust performance on deductive reasoning but suffer from systematic failures in abductive reasoning; (2) generated explanations achieve high semantic similarity to human-written ones—especially for deductive tasks—but rationalization does not consistently improve verification accuracy. This study provides the first empirical evidence of LLMs’ fundamental limitations in abductive reasoning, establishing a novel, trustworthy benchmark and methodology for rigorous reasoning evaluation.

Assessing deductive vs abductive reasoningLLMs' reasoning in claim verificationRationale generation impact on LLMs

Latest Papers

What's happening recently
View more

Debugging in data-intensive programming faces significant challenges, including fragmented evidence, difficulty in discerning discrepancies between expected and observed behaviors, and the complexity of tracking state evolution across components. Through semi-structured interviews and thematic analysis, this study systematically characterizes practitioners’ debugging practices and, for the first time, identifies three core requirements: cross-artifact evidence alignment, expectation-based comparison mechanisms, and traceable state evolution. Building on these insights, the work constructs a visualization-driven design space tailored to debugging in data-intensive contexts, exposing critical gaps in existing tools and providing a theoretical foundation and clear direction for the development of future debugging aids.

data-intensive programmingdebugging challengesevidence-driven reasoning

Industrial research agents often generate experimental trajectories containing invalid or incomplete information, rendering them unreliable for direct decision-making. This work proposes an evidence-oriented framework that automatically transforms such trajectories into structured evidence through a context-isolated generate–verify–repair pipeline. The approach introduces intervention-level claim categorization—distinguishing actionable repairs, diagnostic safeguards, and retained discoveries—and incorporates end-to-end provenance tracking to enable claim scoping and auditability. Experimental results demonstrate that the resulting candidate solutions outperform existing baselines. Audits further reveal that trajectory evolution is non-monotonic, and that applicability assessment constitutes a key performance bottleneck for the controller.

auditable recordsevidence validationindustrial machine learning

This work addresses the challenges of verifying factual consistency in multi-document summarization, where evidence is often scattered, conflicting, or absent. To tackle this, the authors propose the Summary Verification Space (SVS) system, which introduces two novel two-dimensional spatial layouts—SUMMARY-GUIDED and SOURCE-GUIDED—combined with a coordinated provenance visualization mechanism to explicitly reveal relationships among summaries, source documents, and supporting evidence. The system integrates semantic similarity computation, spatial layout algorithms, and provenance tracking into a scalable visual analytics framework. User studies demonstrate that both layouts significantly outperform a linear baseline in verification accuracy, efficiency, and cognitive load, with the SUMMARY-GUIDED layout achieving the best overall performance.

evidence groundingmulti-document summarizationsource provenance

Hot Scholars

YH

Yuyang Huang

University of Chicago
system for mlcomputer systemoperating system
TM

Tom Marchant

Clinical Scientist, Christie Medical Physics and Engineering, The Christie NHS Foundation Trust
Image guided radiotherapycone beam CTmotion tracking and correction
GN

Guoshun Nan

Professor of Beijing University of Posts and Telecommunications
Multimodal LearningVideo LLM6G SecuritySemantic Communications
JC

Jienan Chen

University of Electronic Science and Technology of China
Electrical Engineering