When Verifiable Counts Depend on Wording: Auditing Wording Robustness in Instruction Following
This study addresses the scoring instability in instruction-following benchmarks caused by phrasing variations by proposing the WISE evaluation suite. This suite systematically examines how different phrasings affect model compliance under fixed constraints and introduces a novel matching evaluation protocol to quantify phrasing robustness. A large-scale audit is conducted through multi-model repeated generation, rigorous JSON probing, and human verification. Results demonstrate that negative sentence structures significantly reduce compliance rates and induce ranking reversals. To address this, new metrics—average and worst-case compliance rates—are introduced to supplement traditional scoring systems, revealing critical linguistic sensitivity deficiencies in large language models.