Score
Designs and builds methods, representations, and tools that decompose assertions into atomic claims and align each claim to supporting or contradicting evidence — including passages, table cells, charts, and attribute-structured records — across modalities. Produces structured outputs and visual analytics (e.g., claim-evidence matrices and evidence visualizations) that summarize support composition and confidence, expose coverage gaps, contradictions, and modality imbalances, and enable coordinated claim-level inspection and auditing.
This study addresses the instability of existing claim decomposition approaches in verifying complex, multifaceted claims, which stems primarily from insufficient evidence alignment and inadequate modeling of error patterns in sub-claims. The authors introduce a new dataset featuring temporally constrained evidence and human-annotated evidence spans for sub-claims, enabling a systematic evaluation of claim decomposition under two evidence alignment settings. They reveal, for the first time, the critical impact of fine-grained evidence alignment and label bias in sub-models on verification performance, and propose a structured decomposition framework that contrasts sub-claim aligned evidence (SAE) with repeated claim-level evidence (SRE). Experiments demonstrate that decomposition significantly improves performance only under strictly aligned, fine-grained evidence conditions, and that incorporating an “abstain” strategy effectively mitigates error propagation, with consistent results across multiple datasets.
Large language models frequently generate assertions in financial question answering that blend well-supported claims, weakly reasoned statements, and unsupported conclusions, thereby compromising the reliability of high-stakes decisions. This work addresses the issue by framing it as an assertion–evidence alignment task and introduces a multimodal assertion–evidence matrix that systematically links textual, tabular, and graphical evidence through a lightweight alignment pipeline and standardized JSON evidence artifacts. Integrating a deterministic audit-prioritization algorithm with an interactive visualization interface, the proposed approach significantly enhances analysts’ ability to discern reliable assertions from overconfident or unsubstantiated content in financial statement audits, outperforming conventional chat-based information presentation methods.
This study systematically evaluates the robustness of multimodal large language models (MLLMs) in verifying scientific claims using tabular and graphical evidence. We conduct cross-format experiments on an adapted scientific paper dataset with 12 state-of-the-art MLLMs and include human expert benchmarking for rigorous comparison. Results reveal that current MLLMs achieve strong performance on table understanding but exhibit substantial deficiencies in chart comprehension, coupled with poor cross-modal generalization; in contrast, human experts maintain high accuracy across both modalities. To our knowledge, this work provides the first quantitative characterization of structural weaknesses in MLLMs’ chart semantic parsing—highlighting the critical need to enhance chart understanding for trustworthy, AI-assisted scientific review. Our findings deliver empirical grounding for future model development, identifying concrete limitations and actionable directions for improving multimodal reasoning in scientific domains.
This work addresses the limitations of existing fact-checking approaches, which are often confined to unimodal text or lack interpretability, thereby struggling to verify claims requiring joint visual and textual evidence. The authors propose a two-tier multimodal graph architecture that enables fine-grained evidence retrieval through a bidirectional image-text reasoning mechanism. Multimodal information is fused at both token and evidence levels to support claim verification, while a dedicated fusion decoder generates natural language explanations—realizing, for the first time, an integrated pipeline for retrieval, verification, and explanation. Key contributions include a novel multi-granularity fusion strategy, the bidirectional reasoning mechanism, and AIChartClaim, the first multimodal claim dataset centered on scientific figures in the AI domain. Experiments demonstrate that the proposed method significantly outperforms current baselines on multimodal claim verification tasks.
This work investigates large language models’ (LLMs) capabilities in evidence-based claim verification, specifically evaluating deductive versus abductive reasoning. To this end, we introduce RECV—the first benchmark featuring real-world claims with fine-grained, atomic-level annotations of reasoning types—and propose a reasoning-type decomposition evaluation framework. Through systematic assessment of mainstream closed-source LLMs across multiple difficulty levels and prompting strategies, complemented by semantic similarity analysis, we find that: (1) LLMs exhibit robust performance on deductive reasoning but suffer from systematic failures in abductive reasoning; (2) generated explanations achieve high semantic similarity to human-written ones—especially for deductive tasks—but rationalization does not consistently improve verification accuracy. This study provides the first empirical evidence of LLMs’ fundamental limitations in abductive reasoning, establishing a novel, trustworthy benchmark and methodology for rigorous reasoning evaluation.
Debugging in data-intensive programming faces significant challenges, including fragmented evidence, difficulty in discerning discrepancies between expected and observed behaviors, and the complexity of tracking state evolution across components. Through semi-structured interviews and thematic analysis, this study systematically characterizes practitioners’ debugging practices and, for the first time, identifies three core requirements: cross-artifact evidence alignment, expectation-based comparison mechanisms, and traceable state evolution. Building on these insights, the work constructs a visualization-driven design space tailored to debugging in data-intensive contexts, exposing critical gaps in existing tools and providing a theoretical foundation and clear direction for the development of future debugging aids.
Industrial research agents often generate experimental trajectories containing invalid or incomplete information, rendering them unreliable for direct decision-making. This work proposes an evidence-oriented framework that automatically transforms such trajectories into structured evidence through a context-isolated generate–verify–repair pipeline. The approach introduces intervention-level claim categorization—distinguishing actionable repairs, diagnostic safeguards, and retained discoveries—and incorporates end-to-end provenance tracking to enable claim scoping and auditability. Experimental results demonstrate that the resulting candidate solutions outperform existing baselines. Audits further reveal that trajectory evolution is non-monotonic, and that applicability assessment constitutes a key performance bottleneck for the controller.
This work addresses the challenges of verifying factual consistency in multi-document summarization, where evidence is often scattered, conflicting, or absent. To tackle this, the authors propose the Summary Verification Space (SVS) system, which introduces two novel two-dimensional spatial layouts—SUMMARY-GUIDED and SOURCE-GUIDED—combined with a coordinated provenance visualization mechanism to explicitly reveal relationships among summaries, source documents, and supporting evidence. The system integrates semantic similarity computation, spatial layout algorithms, and provenance tracking into a scalable visual analytics framework. User studies demonstrate that both layouts significantly outperform a linear baseline in verification accuracy, efficiency, and cognitive load, with the SUMMARY-GUIDED layout achieving the best overall performance.