Recipe-Matching, Not Equivalence

๐Ÿ“… 2026-09-25
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the issue that existing mathematical retrieval benchmarks yield inflated scores due to their reliance on large language model (LLM)-generated training data, thereby obscuring deficiencies in genuine semantic understanding. To mitigate this, we introduce the concept of โ€œrecipe matchingโ€ and employ computer algebra system verification, cross-lingual comparisons, and non-LLM paraphrasing control groups to disentangle prompt biases from structural features, quantifying their interference with retrieval performance. Our analysis reveals that 45% of the R@1 gains stem from LLM generation artifacts rather than true semantic comprehension, uncovering a paradox where benchmark scores improve while actual retention rates decline. Furthermore, we release a generator-free evaluation dataset to facilitate more reliable assessments of generalization capabilities.
๐Ÿ“ Abstract
MathNet-Retrieve asks a retriever to find, for a math problem, a document stating the same problem. An LLM under one fixed prompt writes each gold document and its near-miss distractors; LLM judges filter them. We call this procedure the"recipe", training on pairs built the same way"recipe-matching", and ask how much score it buys beyond the ability the benchmark claims to test. Two models from one base, matched in rows and settings, differ only in the training file: pairs written under the benchmark's published prompt by another vendor's LLM and judge, or computer-algebra-verified pairs with no LLM anywhere. The first leads by 45 R@1 points on the easy tier. By a non-LLM paraphrase control, half to two thirds of that gap comes from the pairs being LLM-written at all: LLM rewrites under two unrelated prompts, with the verified model's negatives, recover 30 and 22 of the 45 points; back-translations with the same negatives recover almost none. The remaining 15 to 25 points appear only under the benchmark's own prompt and vanish on real duplicates no generator wrote, the same problem in two languages. The hard tier rewards the recipe's pair structure, a deep rewrite against a minimal-edit near-miss: LLM rewrites alone score zero on it, attaching negatives unlocks it, and every negative that does so costs cross-language points; the sets scoring highest on it separate near-misses no LLM wrote worse than LLM rewrites with verified negatives. MELD also moves when a model trains on pairs built its way, without losing retention; on SABER-Math the registered attack fails, and the one gain, from its LLM-written summaries, is small but holds at a matched budget. Only on MathNet-Retrieve could we pin an inversion, benchmark score up and real retention down, to one edit of a training file. We release the generator-free duplicate evaluations, the near-miss test and three trained models.
Problem

Research questions and friction points this paper is trying to address.

benchmark evaluation
recipe-matching
math retrieval
LLM-generated data
evaluation validity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Recipe-Matching
Retrieval Benchmark Evaluation
LLM-generated Artifacts
Near-miss Distractors
Performance Inversion
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
A
Ali Habibullah
KAUST Academy, Computer, Electrical & Mathematical Sciences & Engineering (CEMSE), King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia
M
Mohammad Alshiekh
KAUST Academy, Computer, Electrical & Mathematical Sciences & Engineering (CEMSE), King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia
Y
Yazan Alshoibi
KAUST Academy, Computer, Electrical & Mathematical Sciences & Engineering (CEMSE), King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia
Salman Khan
Salman Khan
Research Fellow, Oxford Brookes University
Computer VisionDeep learningFire / Smoke detectionAction RecognitionMedical Image Analysis
N
Naeemullah Khan
KAUST Academy & CEMSE, KAUST; Lady Margaret Hall, University of Oxford