Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts

๐Ÿ“… 2026-09-28
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study investigates whether scientific explanations can substantively enhance the experimental prediction accuracy of research agents. To this end, it proposes a โ€œprediction creditโ€ quantification framework that decouples descriptive, explanatory, and contextual effects, alongside five standardized validation protocols. Paired predictions are executed using the DeepSeek V4 model, evaluated through ROC AUC interval scoring, mean absolute error, and coverage metrics across controlled learning settings and benchmarks such as Tox21. Empirical findings reveal no supported predictive gain from natural language explanations; although human-mechanism control groups significantly reduce errors, general explanations fail to improve point prediction accuracy effectively. This work provides rigorous, counterintuitive evidence regarding the practical predictive value of scientific explanations and establishes a systematic evaluation paradigm for future research.
๐Ÿ“ Abstract
Research agents explain planned experiments. We measure predictive credit with paired forecasts sharing an intervention, forecaster, and outcome while varying description, matched explanation, and donor context. Five checks track commitment, delivery, predictive gain, alignment, and known-signal uptake. Across 336 prospective states in controlled learning, 12 Tox21 endpoints, and 24 OpenML tasks, v5's frozen credit decision was inconclusive. Tox21's preregistered ROC AUC interval-score harm test was unmet ($D-M=-.0026$, 95 percent interval [$-.0174$, .0104]); OpenML's joint formation, point-equivalence, and repeatability rule was unmet. Matched point-accuracy gains over description remained unconfirmed, and Tox21/OpenML seed-donor intervals spanned zero. Under requested DeepSeek V4 Pro, matched and donor cards reduced secondary Tox21 drift by 64.5 and 59.1 percent. A DeepSeek V4 Flash replay raised matched point MAE from .01823 to .02020 and missed matched-donor interval-score equivalence. OpenML full-card assignment widened nominal 80 percent intervals by 21 percent, with 49.3 percent coverage versus 51.4 percent for description and content in 66/144 cards. Direct-text Flash delivered all 144 notes without detectable matched point-accuracy gain. A researcher-authored mechanism positive control lowered point MAE by 2.60 percentage points versus description. The protocol measures predictive credit for research-agent benchmarks and scientific forecasting; natural-explanation credit remained unconfirmed at the tested donor resolutions.
Problem

Research questions and friction points this paper is trying to address.

Predictive Credit
Scientific Explanations
Experimental Forecasts
Research Agents
Prediction Accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Predictive Credit
Research Agents
Scientific Forecasting
Paired Forecasts
Evaluation Protocol
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
J
Jingjie Ning
Carnegie Mellon University
Xueqi Li
Xueqi Li
Shenzhen University
Y
Yibo Kong
Carnegie Mellon University
D
Dongting Li
Tsinghua University