🤖 AI Summary
This study addresses the vulnerability of large language models (LLMs) acting as judges to subtle prompt formatting cues, such as punctuation, which can lead to the erroneous acceptance of incorrect answers. Through paired intervention experiments and statistical calibration on the DROP and GSM8K datasets, the authors quantify the impact of presentation formats on evaluation accuracy, incorporating GPT-6 Sol for comparative analysis. The findings reveal a paradoxical coexistence of fundamental judging competence and heightened sensitivity to editorial cues. Specifically, experimental results demonstrate that format variations cause the false acceptance rate of Jev models to surge from 1.0% to 26.0%, whereas GPT-6 Sol remains unaffected by such perturbations. These observations confirm the critical role of prompt structure in LLM-based evaluation and provide empirical evidence for enhancing the robustness of automated assessment frameworks.
📝 Abstract
An answer judge instructed to grade the final commitment should reject an explicitly wrong final value even when an earlier value matches the reference. We show that adding one colon to a candidate can violate this requirement depending on the presentation of the structured judging request. Numeric references certify the error, and paired interventions distinguish the candidate edit from the integration's presentation choices. On 200 previously unused DROP and GSM8K source clusters, the edit increased Jev's false acceptance from 1.0% to 26.0% with three output labels and from 3.0% to 26.5% with the published four-label grading instruction under sorted JSON keys. Both candidate variants were rejected under insertion presentation. These interactions passed the prespecified statistical correction even though Jev met the control thresholds under both grading configurations and presentations. Most excess acceptances occurred among candidates assigned larger numerical errors. GPT-6 Sol produced no observed cue-condition false acceptances, with missing responses unresolved. The result shows that basic judging competence can coexist with sharply different vulnerability to a fixed candidate edit across logically equivalent request presentations. The reference-aware grammar and compound ordering change limit the finding's operational scope and leave its internal cause unmeasured.