Beyond Overlap: Estimating the Causal Effect of Benchmark Exposure

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出LeakScale框架,通过创建新任务并控制信息访问来量化训练材料对评估性能的影响,解决了基准分数难以解释的问题。
📝 Abstract
Evidence that evaluation material entered training does not reveal how much it affected evaluation. This distinction leaves a contaminated benchmark score difficult to interpret: provenance can establish contact, but only a counterfactual can quantify the performance attributable to that contact. We present LeakScale, an interventional framework for estimating this missing quantity. LeakScale creates fresh executable tasks that require private, family-specific information absent from and non-derivable from the public task, controls access to that information, and estimates the resulting control-adjusted change in executable accuracy. Across 2,048 unique families, two model families, two executable domains, and 262,144 generations, exposure improves accuracy in every model-by-domain combination, with gains ranging from +7.17 to +27.31 percentage points. These findings separate two empirical questions that are often conflated: whether benchmark contact occurred and how strongly a reported score depends on it. LeakScale makes the latter directly measurable.
Problem

Research questions and friction points this paper is trying to address.

benchmark exposure
causal effect
performance attribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

LeakScale
interventional framework
benchmark exposure
counterfactual
executable accuracy
🔎 Similar Papers
No similar papers found.