multi-agent evidence fusion

Designs and implements algorithms and systems that ingest heterogeneous, uncertain evidence produced by multiple agents, calibrate and weight those inputs, reconcile or quantify inter-agent conflict, and produce consolidated belief distributions or fused risk scores with propagated uncertainty. It also includes analysis and evaluation of fusion performance, conflict metrics, and uncertainty propagation to improve predictive accuracy and decision quality.

multi-agentevidencefusion

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of existing evidence fusion methods, which struggle to simultaneously capture inter-evidence conflict and intra-evidence uncertainty while neglecting the long-term reliability of evidence sources. To overcome these issues, the authors propose a unified evidential reasoning framework. It introduces a chaos-conflict joint measure satisfying five axioms to coherently quantify both conflict and nonspecificity. Furthermore, it incorporates a context-aware reliability assessment mechanism derived from historical fusion outcomes, leveraging spectral clustering and regret theory. This reliability estimate drives an adaptive combination rule and a belief-interval-based decision strategy. Evaluated on 16 real-world datasets, the method achieves an F1 score of 85.78% and an AUC of 93.30%, significantly outperforming eight Dempster–Shafer theory baselines and three gradient boosting approaches. Ablation studies confirm the contribution of each component.

conflict measurementDempster-Shafer theoryevidence fusion

This paper addresses the modeling challenge of “unconceivable uncertainty”—events unforeseeable yet possible—arising in social and life sciences, where classical probability theory fails due to its inherent reliance on known, enumerable possibilities. We propose an extended evidence-theoretic framework that formally distinguishes between uncertainties within and beyond the agent’s cognitive boundary, integrating imprecise probabilities, subadditive measures, and non-standard information-theoretic approaches. Crucially, we establish a novel interface between this framework and multi-agent systems, rigorously differentiating representable from unrepresentable uncertainty sources. The resulting formalism provides a new mathematical foundation and analytical paradigm for studying complex socio-biological systems, particularly in risk perception and cultural information diffusion. (128 words)

Compares extended Evidence Theory with advanced Probability Theory variantsExplores multi-agent applications of enhanced uncertainty reasoningExtends Evidence Theory to handle unforeseen event uncertainties

This work addresses the challenge of selecting uncertainty representations that align with decision objectives to achieve optimal and trustworthy decisions under state-variable uncertainty. Drawing on decision theory, it systematically analyzes the optimal forms of uncertainty representation for both risk-neutral and risk-averse agents in known and unknown environments, revealing the minimal uncertainty information required under distinct risk preferences. The study innovatively unifies three approaches to epistemic uncertainty—calibrated prediction, confidence-set robust optimization, and Bayesian inference—establishing a theoretical link between uncertainty representation and decision goals. This integration yields a reliable decision-making framework that provides agents with verifiable utility guarantees.

decision makingepistemic uncertaintyposterior distribution

This work addresses the problem of “debate collapse” in multi-agent debate systems, where erroneous reasoning often dominates due to the absence of effective detection and intervention mechanisms. The authors propose a hierarchical uncertainty quantification framework that measures behavioral uncertainty at the individual, interaction, and system levels, introducing it for the first time as a diagnostic indicator of system failure. Building on this, they develop an uncertainty-driven strategy optimization method that dynamically penalizes self-contradictory statements, inter-agent conflicts, and low-confidence outputs. Experimental results demonstrate that the proposed approach significantly improves decision accuracy, reduces internal inconsistency, and achieves reliable calibration of multi-agent systems across multiple benchmarks.

debate collapseLLM reasoningmulti-agent systems

This study addresses the limitations of existing audit risk assessment approaches, which predominantly rely on point predictions and fail to explicitly capture the consistency and uncertainty inherent in heterogeneous evidence. To overcome this, the authors propose the UMAR framework, which introduces a multi-agent collaborative mechanism for the first time: three specialized agents are constructed based on MD&A text, financial ratios, and key audit matters, respectively, to generate calibrated, uncertainty-aware risk scores. These scores are then fused using Dempster–Shafer evidence theory, which also quantifies inter-agent evidence conflict. Evaluated on a sample of 3,200 U.S. public firms, the method achieves an AUROC of 0.782 and a PR-AUC of 0.341, with the lowest expected calibration error (ECE = 0.052), significantly outperforming baseline models and delivering interpretable, actionable risk signals.

audit risk assessmentcalibrated predictionsevidence conflict

Latest Papers

What's happening recently
View more

This work addresses the lack of reliable uncertainty estimation for reasoning failures in agentic retrieval-augmented generation (RAG) systems during multi-hop question answering. It proposes the first uncertainty-aware agentic RAG framework, which integrates semantic disagreement metrics with a generator self-evaluation mechanism to produce stage-wise uncertainty signals. For the first time, a Bayesian network is employed to propagate uncertainty across the system and identify failure-prone components. Experimental results on StrategyQA and HotpotQA demonstrate that the proposed framework significantly improves AUROC and AUARC while reducing ECE and Brier Score, effectively modeling uncertainty accumulation in multi-hop reasoning and enabling both node-level failure warnings and system-level confidence assessment.

Agentic RAGFailure DetectionMulti-Hop Question Answering

This work addresses the absence of a unified confidence assessment mechanism for collective outputs in existing multi-agent systems. The authors propose three protocols that standardize individual agents’ raw confidence signals and integrate soft voting with Bayesian fusion strategies to produce a single, comparable confidence estimate for the final answer. The approach innovatively combines sequential probabilities with self-reported confidence estimators and unifies parametric and non-parametric calibration techniques. Extensive experiments across five benchmarks and four task categories demonstrate that the proposed aggregated confidence significantly outperforms both single-agent and standard debate baselines—evidenced by notable gains in AUARC—while maintaining stable F1 scores. Particularly in ambiguous tasks, the method effectively compensates for performance degradation typically observed in multi-agent debates.

confidence aggregationmultiagent debatemultiagent systems

This study investigates how localized misinformation propagates through interactions in large language model (LLM) multi-agent systems and undermines collective fact-recovery capabilities. To this end, the authors propose the Hi-Agreement evaluation framework, which compares scenarios of honest collaboration against those where key evidence holders introduce deceptive statements within a controlled environment. By integrating multi-stage voting, testimony adoption tracking, and evidence provenance analysis, the framework elucidates the mechanisms of distributed information aggregation. Experiments reveal, for the first time, that a single false testimony is more readily adopted and propagates more widely than truthful statements, exerting persistent influence even after the deceiving agent withdraws. Across 120 five-agent scenarios, collective fact-recovery rates plummeted from 72.50% to 14.17%, highlighting the system’s acute vulnerability to deception; while observers without direct evidence can mitigate erroneous consensus, they fail to significantly enhance truth recovery.

distributed reasoningfact recoveryLLM-based collaboration

This work addresses a fundamental limitation in distributed Bayesian experimental design, where local nodes evaluate experiments using only local information, leaving the fusion center without access to global likelihoods or information gains. The paper proposes the first Bayesian optimal fusion rule: local agents submit design decisions based on expected information gain as their utility function, and the fusion center selects the experiment maximizing the conditional expectation of centralized information gain. Departing from conventional classification error criteria, this approach defines loss via information gain regret and establishes corresponding bounds on information loss along with conditions for asymptotic equivalence to the centralized optimum. Numerical experiments demonstrate that the proposed rule closely approximates the performance of an ideal centralized design and significantly outperforms heuristic strategies such as majority voting.

Bayesian experimental designdecision fusiondistributed experimental design

Hot Scholars

RX

Runsheng Xu

Waymo Research, UCLA
Computer VisionDeep LearningAutonomous Driving
ZT

Zhengzhong Tu

Texas A&M University, Google Research, University of Texas at Austin
Agentic AITrustworthy AIEmbodied AI
SZ

Seth Z. Zhao

University of California, Los Angeles
SimulationMulti-agent LearningAutonomous DrivingCooperative Driving
EF

Emanuel Figetakis

PhD Candidate, University of Guelph
Machine LearningIoTReinforcement learningComputer Vision