Not Until the Evidence Says So: Teaching LLM Investigators When to Close a Case

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the tendency of large language models (LLMs) to erroneously close cases with insufficient evidence due to overconfidence during incident investigations. To tackle this, the authors construct Nautil, a multi-domain dataset incorporating counterfactual evidence, and propose an evidence-dependency evaluation framework. LLM investigators are trained via supervised fine-tuning and reinforcement learning to determine whether to close a case or identify missing information based on evidence sufficiency. Experimental results demonstrate that the overstatement rate of a 9B-parameter model decreases from 97% to 35%, while the balanced accuracy for case closure reaches 83.3%. These findings indicate a successful paradigm shift from indiscriminate case closure to evidence-driven decision-making in automated incident investigation.
📝 Abstract
Accident, defect and outage investigations end with a decision that ordinary question answering never faces: whether the evidence gathered so far is enough to close the case. We study this decision for LLM investigators, which request evidence from a case file, revise their hypotheses, and either close the case with a conclusion grounded in what they read or leave it open and name what is missing. This judgment does not come with capability: an untrained 9B model overstates its evidence in 97% of its answers, and a frontier model that identifies the right cause in 84% of cases still overstates in 91% and closes 17 of the 41 cases whose official finding is "cause undetermined". Measuring it is also non-trivial: the source of a case largely predicts its label, and a rule that reads only the source reaches 83.0 balanced accuracy on our test cases. We therefore evaluate closure with three tests: closure accuracy, reported against this rule and within each source; evidence dependence, which removes the grounds of a conclusion and checks whether the model stops closing; and conclusion and gap quality, a judged checklist of what the model asserts and what it says is missing. We build Nautil, 731 audited cases from aviation, rail, maritime, chemical-safety and vehicle-defect reports and production server incidents, with teacher trajectories, an out-of-distribution test set and counterfactual evidence versions. Fine-tuning a 9B model on these trajectories makes its closures follow the evidence: removing the grounds lowers its closure rate by 26 points relative to a matched control, overstatement falls from 97% to 35%, and correct, non-overstated conclusions rise from 3% to 43%. Reinforcement learning that rewards only the closure decision then raises balanced accuracy from 69.2 to 83.3, on par with the teacher, and within-source accuracy from 60.4 to 74.1, at some cost in evidence dependence.
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM investigators
evidence-based closure
Nautil benchmark
supervised fine-tuning
reinforcement learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Tingzhu Bi
Peking University
Ping Wang
Ping Wang
Peking University
M
Meng Ma
Peking University