When benchmark inferences do not compose: Projectibility in AI evaluation

📅 2026-07-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current AI benchmarks lack guarantees regarding the overall validity of reasoning chains when extrapolating from limited evidence to real-world deployment. This work introduces the concept of “projectibility,” emphasizing the coherent transmission of premises, assumptions, and uncertainties across reasoning steps, and articulates a “non-compositional principle” demonstrating that locally valid inferences may collectively fail. Drawing on philosophical epistemology—particularly the problem of projectibility—and integrating frameworks of argumentative validity with reanalysis of case studies and simulation experiments, the authors develop a projectibility auditing methodology. Applying this approach to legal AI reveals a critical disconnect: while benchmark evaluations and deployment studies may each appear sound in isolation, they often fail to align in practice. Simulations further show that aggregate stability can obscure underlying discrepancies, thereby validating the proposed auditing framework.
📝 Abstract
An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and combine it with assumptions about human review and downstream consequences. Validity-centred approaches require evidence for each claim. This paper identifies a further epistemic problem: warranted links don't automatically make a warranted chain. The target of one study may not be the source of the next; system, population, outcome, or conditions may change at the interface; and shared data or model lineage may make apparently independent support dependent. Projectibility concerns whether a bounded extension from observed to unobserved cases is warranted. Goodman supplies the problem of rival extensions; argument-based validity supplies an architecture for testing them. The paper's distinctive claim is a non-composition principle: support for adjacent projections warrants their composition only when endpoints and assumptions align and dependence and uncertainty are carried through. A legal-research case shows how benchmark evidence and a deployment study can each be sound while remaining parallel. A reanalysis and simulation show why aggregate stability can erase distinctions a later projection requires. The resulting projectibility audit diagnoses unsupported joins in benchmark-to-use arguments.
Problem

Research questions and friction points this paper is trying to address.

projectibility
AI evaluation
benchmarking
non-composition
epistemic validity
Innovation

Methods, ideas, or system contributions that make the work stand out.

projectibility
non-composition principle
AI evaluation
argument-based validity
benchmark generalization
🔎 Similar Papers