🤖 AI Summary
Existing evaluations of scientific search agents rely solely on final accuracy, making it difficult to pinpoint specific failure causes across target exposure, examination attempts, and answer acceptance. This work proposes a label-free decision checkpoint protocol that logs observations and tool actions throughout the reasoning process, decomposing aggregate accuracy into question-level discrepancies to precisely distinguish stage-specific behaviors. Experiments conducted on 540 questions from the AutoResearchBench Deep dataset demonstrate that keyword-based search achieves an accuracy of 24.6%, outperforming vanilla search. Furthermore, the analysis reveals significant differences among strategies regarding evidence invocation volume and unexamined targets.
📝 Abstract
Final-answer accuracy does not reveal whether a scientific-search agent failed to encounter a target paper, attempt to inspect it, or return an accepted answer after inspection. We introduce decision checkpoints that record observations and tool actions without benchmark labels during inference, then join target identities and evaluator labels to assign outcome categories from recorded events. Across five conditions on 540 answerable AutoResearchBench Deep questions in a fixed, target-enriched environment, keyword search achieves 24.6\% accuracy, compared with 17.8\% for raw search. The keyword condition has fewer incorrect answers with neither target exposure nor inspection, but more incorrect answers after the target is exposed and left uninspected. Compared with keyword search, read-first has 27.4\% more recorded evidence-search calls. Target inspection attempts occur on 199 questions under read-first and 191 under keyword search; both conditions achieve 24.6\% accuracy. The checkpoint protocol makes these question-level differences explicit, distinguishing target exposure and inspection from aggregate accuracy and total tool use.