🤖 AI Summary
This study addresses the evidence selection bias inherent in deep research agents during adaptive search, which frequently yields accurate citations yet misleading conclusions. To tackle this issue, we formalize it as an adaptive evidence sampling problem and propose the CESS framework for the first time. Specifically, CESS leverages causal inference and log-probability weighting to correct the average evidence direction within the candidate pool, incorporates shrinkage estimation to stabilize short searches, and decouples pool estimation from policy effects to audit the influence of search strategies on final conclusions. Empirical evaluations demonstrate that CESS reduces error by 9.2% and improves robustness by 39.4% on the MS2 benchmark. Furthermore, when evaluated on real-world trajectories, the framework achieves a 60.1% reduction in error alongside an 87.2% improvement in robustness.
📝 Abstract
Deep Research agents synthesize evidence into cited reports, yet a well-cited report can still reach a misleading conclusion. Citation correctness checks whether cited sources support individual claims. It does not show whether adaptive search exposed a representative view of all documents made available for evaluation, which we call the candidate pool. Early findings redirect later queries, document choices, and stopping, so the documents an agent reads form a selective sample. Existing evaluations rarely account for this selection. We formulate the problem as adaptive evidence sampling and introduce Causal Evidence Selection Correction (CESS). CESS predicts each candidate document's evidence direction and corrects the candidate-pool average using the logged probabilities of selecting each document and reaching each search round. Shrinkage stabilizes short searches, while intervals replace point estimates when some documents cannot be sampled. We also prove that estimating the average evidence direction of a common pool differs from measuring how a change in search policy alters the evidence read. The latter requires intervention. On questions from the MS2 systematic-review benchmark, CESS reduces mean absolute error against the candidate-pool average by $9.2\%$ and reduces the estimate's change under opposing document rankings by $39.4\%$ relative to averaging the evidence scores of documents read. Across trajectories from a public Open Deep Research agent, the corresponding reductions reach $60.1\%$ and $87.2\%$. A further 4,800 trajectories under paired interventions confirm that correcting a pool estimate and measuring a policy effect are different tasks. CESS therefore audits whether the evidence direction underlying a report reflects the documents available for evaluation, while a separate intervention analysis measures the effect of search decisions.