Score
Designs and implements end-to-end systems and processes that locate, prioritize, and filter external evidence to support downstream tasks. This includes hypothesis- or intent-driven retrieval, source-level selection and weighting, retrieval-based grounding, and auditable pipelines and criteria to exclude weak or unverifiable sources so outputs are grounded in diverse, authoritative, verifiable references.
Current information retrieval systems struggle with complex queries requiring logical constraints, multi-step reasoning, and evidence synthesis, primarily due to a lack of structured reasoning capabilities. This work proposes the first unified framework for structured reasoning in information retrieval, systematically integrating interdisciplinary approaches—including large language model reasoning strategies, neuro-symbolic systems, probabilistic and Bayesian methods, geometric representations, and energy-based models—to elucidate their inherent trade-offs and complementary mechanisms. By bridging disciplinary boundaries, the framework clarifies the central role of retrieval within general-purpose reasoning systems and provides researchers with conceptual tools and practical guidance to advance the development of verifiable, structured reasoning architectures.
Current information retrieval paradigms struggle to support complex analytical tasks such as trend analysis and causal inference, lacking end-to-end problem-solving capabilities, controllable reasoning processes, and verifiable results. This work proposes a novel paradigm termed “analytical search,” formally defining it as a distinct search type separate from traditional retrieval and retrieval-augmented generation (RAG). By explicitly modeling analytical intent, the approach constructs an evidence-driven, process-oriented, multi-step structured reasoning workflow. The study introduces a unified framework that integrates query understanding, recall-oriented retrieval, reasoning-aware fusion, and adaptive verification mechanisms. This framework lays the theoretical foundation and outlines future research directions for next-generation analytical search engines that are highly accountable and capable of supporting multi-objective analytical tasks.
Industrial research agents often generate experimental trajectories containing invalid or incomplete information, rendering them unreliable for direct decision-making. This work proposes an evidence-oriented framework that automatically transforms such trajectories into structured evidence through a context-isolated generate–verify–repair pipeline. The approach introduces intervention-level claim categorization—distinguishing actionable repairs, diagnostic safeguards, and retained discoveries—and incorporates end-to-end provenance tracking to enable claim scoping and auditability. Experimental results demonstrate that the resulting candidate solutions outperform existing baselines. Audits further reveal that trajectory evolution is non-monotonic, and that applicability assessment constitutes a key performance bottleneck for the controller.
Existing fact-checking retrieval models primarily rely on relevance ranking, neglecting the actual discriminative utility of retrieved evidence for verifiers. Method: We propose a “utility-driven” evidence retrieval paradigm that directly enhances the support of retrieved evidence for claim verification. To this end, we design a Feedback-based Evidence Retriever (FER), which employs the KL divergence between the verifier’s (e.g., BERT-based) prediction distributions over retrieved versus gold evidence as a differentiable feedback signal, enabling end-to-end joint optimization of retrieval and verification. Our approach integrates the retriever, verifier, and KL-based feedback mechanism without requiring human-annotated evidence. Contribution/Results: On benchmarks such as FEVER, FER significantly outperforms prevailing relevance-driven baselines. It is the first work to systematically demonstrate that utility-oriented retrieval yields substantial improvements in fact verification performance, offering a novel, interpretable, and efficient pathway for automated fact-checking.
Current AI-assisted scientific writing lacks auditable generation processes and mechanisms for accountability, undermining the verifiability of research credibility and compliance. This work proposes a novel auditing paradigm embedded directly within the production workflow, enforcing end-to-end traceability, immutability, and third-party reproducibility of AI involvement through preregistered blind-spot indicator cards, sealed execution environments, and automated gatekeeping intercepts. Core technical components include Git-sealed lineage anchoring, hash-bound provenance tracking, red-flag interception protocols, cross-model role isolation, and programmatic assembly. In experimental validation, one project was automatically terminated when preregistered confirmatory tests triggered a No-Go decision. An open-source toolkit is released to enable independent recomputation of all core audit metrics by third parties.
This study addresses the challenges of synthesizing multi-source heterogeneous evidence—such as academic papers, reports, policies, and media content—which vary widely in quality and structure and entail high manual effort. The reliability of current large language models (LLMs) across individual synthesis subtasks remains unclear. To tackle this, the authors propose the Knowledge Synthesis Review (KSR) framework, decomposing the review process into four stages: screening, extraction, analysis, and synthesis. Using a high-agreement expert gold standard (92.2% agreement, κ=0.80), they conduct task-level evaluations of leading LLMs—including GPT-5, Claude Sonnet 4, and Gemini 2.5 Pro—and introduce a dynamic routing mechanism that automatically selects the best-performing model under human supervision. This model-agnostic, auditable, and transparent approach significantly enhances review efficiency and coverage. Experiments reveal no single model dominates all tasks: Claude Sonnet 4 achieves the highest screening accuracy (82.8%), while GPT-5 attains the best recall (91.8%). Moreover, multi-source synthesis uncovers critical themes—such as worker well-being, small and medium enterprises, and Global South perspectives—often missed in single-source analyses.