decompositional evidence grounding

Designs, builds, or evaluates methods that decompose complex queries into explicit sub-questions and produce intermediate sub-answers each grounded in localized evidence. The work identifies and links the specific evidence items (e.g., text spans or image regions) that support each sub-answer, verifies them across the chain of evidence, and composes a final, evidence-supported response.

decompositionalevidencegrounding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.55
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the lack of scientifically grounded evidence supporting large language model (LLM) answers in biomedical question answering. Methodologically, we propose a multi-source retrieval-augmented generation (RAG) framework that integrates heterogeneous biomedical literature sources—including the novel observational evidence repository Alexandria (formerly Atropos Evidence Library), PubMed, and Perplexity—to conduct empirical analysis on real-world physician questions. Our key contributions are: (1) the first incorporation of dynamic observational study evidence into LLM evaluation, substantially broadening the scope of evidence coverage; and (2) cross-source validation revealing that PubMed alone supports answers to ~44% of questions, Alexandria supports ~50%, and their combination enables evidence-based, reliable answers for over 70% of questions. This establishes a reproducible benchmark and technical pathway for trustworthy biomedical AI.

Comparing retrieval systems for scientific literature groundingEvaluating evidence support in LLM biomedical answersImproving reliability of LLM responses to medical questions

Evaluating large language models’ (LLMs) mathematical reasoning solely via final-answer accuracy risks overlooking spurious correctness—correct outputs arising from flawed or redundant intermediate reasoning steps. Method: We propose a training-free, zero-parameter sub-thought–driven evaluation framework: (1) automatically segmenting chain-of-thought (CoT) rationales into semantically coherent sub-thoughts; (2) prompting multiple independent continuations per sub-thought; and (3) aggregating results via majority voting and consistency confidence scoring to yield robust predictions. Contribution/Results: This is the first method to systematically uncover and exploit latent redundancy in sub-thought–level correctness, moving beyond conventional single-path CoT evaluation. On AIME2024 and AIME2025, it achieves +13% and +10% absolute accuracy gains over strong baselines, respectively. Crucially, its consistency metric reliably distinguishes erroneous reasoning from superficially correct outputs, enhancing diagnostic interpretability.

Can alternative reasoning paths yield different results?Does the final answer reliably represent the model's optimal conclusion?How to improve accuracy by analyzing intermediate reasoning steps?

This work addresses the limitations of large language models in tackling highly complex questions that demand long-horizon planning and integration of massive, heterogeneous evidence. To overcome these challenges, the authors propose an autonomous complex question-answering framework that synergistically combines “ultra-wide” and “ultra-deep” research paradigms. The approach employs structured task decomposition, large-scale multi-source retrieval, iterative deep querying, and a graph-anchored auditing protocol to produce verifiable research reports enriched with fine-grained citations and intermediate reasoning artifacts. Evaluated on a benchmark of 300 expert-level complex questions, the system demonstrates the capability to analyze thousands of web pages and perform hundred-step reasoning chains, substantially enhancing answer traceability and multidimensional credibility.

autonomous researchcomplex question answeringevidence synthesis

This work addresses the persistent hallucination problem in retrieval-augmented generation (RAG) systems, which often occurs even when relevant documents are available, and highlights the limitations of conventional evaluation methods in identifying fine-grained issues in evidence utilization. The authors propose a diagnostic framework for question-answering tasks that decomposes questions into atomic reasoning facets and constructs a facet–text chunk matrix. By integrating retrieval relevance with natural language inference–based faithfulness scores, the framework analyzes evidence usage across three reasoning paradigms: Strict RAG, Soft RAG, and LLM-only. This approach enables, for the first time, facet-level diagnosis of RAG behavior, uncovering systematic failure modes—such as missing, misaligned, or prior-dominated evidence—that remain invisible at the answer level. Experimental results demonstrate that hallucinations primarily stem from flawed evidence integration rather than retrieval inaccuracies, offering an interpretable foundation for improving RAG systems.

evidence groundingfacet-level analysishallucination

Latest Papers

What's happening recently
View more

This work addresses the challenge of disjointed page-level and region-level evidence localization in multi-page document visual question answering, which hinders effective retrieval of sparse information. The authors propose a hierarchical evidence routing framework that models long-document evidence acquisition as a two-stage set prediction process: a page-level policy first identifies key pages, followed by a region-level policy that extracts semantic elements for the answer generation model. This approach uniquely decouples page and region selection into sequential, independently optimizable stages, enabling coordinated localization through granularity-specific structured set rewards. Integrating a two-tier policy architecture, staged GRPO optimization, fused OCR and table text representations, and joint global-local inputs, the method achieves state-of-the-art or competitive performance among open-source systems across multiple benchmarks, outperforming the strongest open-source baseline on LongDocURL by 16.87%, with region-level evidence boosting accuracy and F1 by 5.51% and 4.82%, respectively.

evidence routinglong-documentpage-level selection

Current AI-generated hypotheses in materials science, while linguistically fluent, lack verifiable scientific grounding in their internal reasoning mechanisms. This work proposes a visual diagnostic framework integrating semantic backtracking, graph-structure perturbation, and activation recovery metrics to scrutinize the “graph-to-answer” hypothesis generation pathway—the first such approach to focus mechanism interpretability on this specific process. Leveraging the Graph-PrefLexOR-8B model, residual stream scanning and layer–token grid visualizations across 100 open-ended materials science questions reveal that mechanistic recovery is highly concentrated in late-stage synthesis layers—particularly around layers 30 and 36—and that the final answers align most closely with the model’s own synthesized representations, thereby identifying critical loci where graph-based reasoning mechanisms are predominantly encoded.

graph-to-answerhypothesis generationmaterials science

Hot Scholars

ZX

Zhen Xing

Alibaba Tongyi Lab | Zhejiang University | Fudan University
Computer VisionVideo GenerationAIGCVideo Diffusion
DG

Diandian Guo

The Chinese University of Hong Kong
Deep learning
KJ

Kaixun Jiang

Fudan University
Computer VisionAdversarial Examples
DL

Disheng Liu

Case Western Reserve University
Computer Vision