When Verifiable Counts Depend on Wording: Auditing Wording Robustness in Instruction Following

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scoring instability in instruction-following benchmarks caused by phrasing variations by proposing the WISE evaluation suite. This suite systematically examines how different phrasings affect model compliance under fixed constraints and introduces a novel matching evaluation protocol to quantify phrasing robustness. A large-scale audit is conducted through multi-model repeated generation, rigorous JSON probing, and human verification. Results demonstrate that negative sentence structures significantly reduce compliance rates and induce ranking reversals. To address this, new metrics—average and worst-case compliance rates—are introduced to supplement traditional scoring systems, revealing critical linguistic sensitivity deficiencies in large language models.
📝 Abstract
Verifiable instruction-following benchmarks often express each constraint through one fixed template. We test whether scores remain stable when the operational requirement is unchanged but its wording varies. We introduce WISE, a matched evaluation suite and reporting protocol instantiated on exact word count, keyword inclusion exactly once, and an inclusive 8--12 word range. Across 100 matched tasks, up to thirteen models from seven providers, and repeated generations scored over the complete visible output, wording alone produces substantial compliance shifts. In an avoidance-family panel, five avoidance and exclusion forms fall below the positive baseline, while constructional controls also shift compliance substantially: in the nine-model control panel, compliance is 54.9% for the original positive form, 48.2% for a longer positive form, 36.7% when the target appears later, and 33.8% for AVOID1. A strict JSON-structure probe shows wording sensitivity beyond counting, with a different direction of effect. Effect sizes, failure directions, weakest forms, and model rankings vary across realizations. Under the most disruptive exclusion form, the top-ranked model changes and 24.1% of strictly ordered model pairs reverse. Human validation further shows that unanimous agreement on an exact-count interpretation can coexist with substantially different model behavior. WISE supplements conventional scores with mean and worst-form compliance, wording gaps, failure profiles, and ranking stability.
Problem

Research questions and friction points this paper is trying to address.

instruction following
wording robustness
verifiable benchmarks
evaluation reliability
compliance stability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Wording Robustness
Instruction Following
Evaluation Benchmark
Large Language Models
WISE Protocol
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Q
Qishi Zhan
Marquette University
Seoyeon Jang
Seoyeon Jang
KAIST Urban Robotics Lab
LiDAR SLAMMoving Object SegmentationStatic Mapping
Z
Zihan Dong
Georgia Tech
M
Minxuan Hu
Cornell University
Z
Ziheng Chen
The University of Texas at Austin
T
Tonghui Qu
Hikvision