local evidence search

Designs and implements algorithms and decision procedures that identify, select, and prioritize localized regions of data (‘‘local evidence’’) for inspection or verification under explicit resource or budget constraints. This competence covers building budget-aware selection strategies and streaming, fine-grained action planners (e.g., sequential or tree-search methods) that produce prioritized action sequences to efficiently acquire checkable evidence.

localevidencesearch

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.58
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of designing an auditable, evidence-driven authorization mechanism for decisions with heterogeneous persistence—such as transient routing versus permanent exclusion—in budget-constrained multi-source learning. We propose ALIVE, the first framework to decouple persistent exclusion actions from non-persistent routing decisions, introducing an auditable exclusion protocol grounded in evidence thresholds and capacity feasibility. Under a shared audit-and-learning budget, source exclusion is permitted only when a strict majority inconsistency condition is met and dual-certificate verification succeeds. By integrating sampling without replacement via random prefix selection with Serfling/FPC finite-population corrections, ALIVE provides theoretical family-wise error rate guarantees at any time. Experiments show that ALIVE improves accuracy–AUBC by 0.1935 percentage points over pure routing on CIFAR, substantially reduces median evidence volume in the PPR engine (e40 drops from 304 to 96), and covers 88% of samples with shorter prefixes on a fixed natural panel.

auditingbudgeted learningmulti-source learning

This work addresses the challenge of dynamically deciding whether a multi-agent large language model system should execute an answer or defer to human review under a limited error budget. The problem is formalized as a budget-constrained “act-or-defer” decision framework, where debate prefixes are mapped to low-dimensional states, and a k-nearest-neighbor lower confidence bound on state-conditional correctness is computed using calibration data. An action is executed only when this bound exceeds a user-specified threshold. The authors introduce a novel conditionally safe assurance framework that decomposes the total error budget into three interpretable components: calibration failure, residual risk, and representation gap, enabling falsifiable diagnostics and task-difficulty-normalized budget allocation. Evaluated across six benchmarks, the method achieves 84% automation rate and 96% execution accuracy using only 9–12% of the error budget, substantially outperforming nine baselines.

act-or-deferbudgeted decision makingLLM reasoning

This work addresses the limitations of large language models (LLMs) in open-ended investigative tasks, where constrained context windows and complex evidence dependencies hinder the generation of coherent and reliable global explanations. To overcome this, the authors propose the Explanation-on-Graph (EoG) framework, which formulates the task as abductive reasoning over a dependency graph. Within EoG, an LLM performs local evidence extraction and annotation, while a deterministic controller orchestrates graph traversal, state maintenance, and belief propagation to collaboratively construct a minimal explanation frontier. By decoupling semantic reasoning from control logic and integrating a graph-structured belief propagation mechanism, the framework enables dynamic evidence aggregation and autonomous hypothesis revision. Evaluated on the ITBench diagnostic benchmark, EoG achieves a seven-fold average improvement in Majority-at-k entity F1 over ReAct baselines, substantially enhancing explanation consistency and reliability.

belief propagationcontext window limitationevidence mining

Latest Papers

What's happening recently
View more

Existing approaches to video misinformation detection typically model entire videos holistically, overlooking the fact that deceptive content often relies on sparse, critical cues. This leads to computational redundancy and dilutes discriminative evidence. To address this, this work proposes SIEVE, a novel framework that introduces, for the first time, an agent-based paradigm for actively searching multimodal sparse evidence. SIEVE decouples evidence acquisition from verification: an agent employs evidence-aware reinforcement learning to efficiently extract a minimal set of key cues, forming a compact evidence package, which a verifier then uses to make interpretable judgments. The method significantly outperforms current state-of-the-art techniques across multiple benchmarks, achieving high detection accuracy with highly distilled evidence while providing transparent and traceable decision rationales.

decision-relevant cluesevidence structuremultimodal video misinformation detection

This work addresses the vulnerability of large language model (LLM) agents to unauthorized evidence in mixed-trust contexts, which can lead to policy-violating decisions. The authors propose a goal-oriented, fine-grained authorization auditing framework that isolates the influence of evidence source authority by fixing task specifications, propositions, stances, and policies, while systematically varying only the trustworthiness of evidence sources. Through contextual subset ablation and multi-model comparison, they quantify how unauthorized information alters action selection. The study introduces a novel auditing mechanism that separately annotates contextual factors with respect to tool-use and parameter-setting goals, alongside a controlled stress-testing protocol. In 450 controlled tasks, 5.4% of actions changed due to source differences, with 2.4% retaining conflicting unauthorized evidence—indicating that while LLMs perceive source cues, they remain susceptible to their influence.

authorization auditevidence groundingLLM agents

This work addresses the inadequacy of current AI runtime logs in providing the structured evidence necessary for legal fact-finding—such as data boundary violations or human interventions. It formalizes, for the first time, the binary factual requirements of regulatory compliance into a criterion of evidentiary sufficiency for runtime records, mandating that logs explicitly encode the legal category of events and their determinative relationships (e.g., provenance, authorization, temporal validity). By integrating legal ontologies, event-type systems, provenance semantics, and temporal validity constraints—and drawing on the law of requisite variety and the Good Regulator theorem from cybernetics—the approach exposes limitations in tamper-proof logging and generic provenance mechanisms. Validation against selected obligations of the EU AI Act demonstrates that this criterion precisely delineates the boundary between traces and hyperproperties in runtime verification, thereby establishing a verifiable foundation for compliance.

Agentic AIevidentiary adequacylegal findings

This work addresses the challenge in vision-language tasks where full responses are difficult to verify, yet only partial information—such as regions, relations, or numerical values—is verifiable, rendering traditional Best-of-N (BoN) selection ineffective. The authors propose the Best-of-Evidence (BoE) framework, which formalizes candidate selection under partial verification for the first time. BoE models reusable claims via a signed candidate-factor graph and dynamically selects the most decision-influential evidence within a limited query budget. Theoretically, shared factor queries reduce query complexity from Θ(K) to O(log K), with BoE naturally degenerating to BoN under zero budget. Experiments demonstrate that BoE significantly improves selection performance across four medical VQA datasets, rectifies failure cases of BoN, and reveals that channel quality and candidate generation capacity critically constrain achievable performance.

Best-of-Ncandidate selectionevidence-based reasoning

Hot Scholars

CG

Cathal Gurrin

School of Computing, Dublin City University & Adapt Centre
LifeloggingQuantified SelfPersonal Data AnalyticsMobile HCI
WB

Werner Bailer

JOANNEUM RESEARCH
multimedia analysismachine learningmetadata models
LR

Luca Rossetto

Dublin City University
Multimedia RetrievalMultimedia AnalysisData-centric AIMulti-modal Graphs
KS

Klaus Schoeffmann

Associate Professor, Klagenfurt University, Austria
Medical Video AnalysisDeep LearningComputer VisionVideo Retrieval