🤖 AI Summary
This study addresses the "structure tax" problem, wherein enforcing structured outputs in large language models degrades accuracy, by conducting a systematic evaluation across multiple models and datasets. Methodologically, it employs Centered Kernel Alignment (CKA) to analyze representation separability in intermediate Transformer layers, alongside confidence calibration and hidden-layer geometric measurements. The findings reveal that accuracy degradation stems from schema design rather than structural constraints per se, proposing a paradigm shift from "whether to structure" to "how to structure." Experiments demonstrate that a reasoning-prioritized field ordering strategy matches or surpasses free-text performance while significantly improving model calibration.
📝 Abstract
Deploying large language models in production often requires constraining outputs to structured formats such as JSON or XML, and prior work treats the resulting accuracy loss as an inherent `structure tax'. We re-examine this claim by evaluating a battery of models, datasets and schemas, measuring task accuracy, confidence calibration, and hidden-state geometry. The tax turns out to depend on schema design rather than on structure per se: reasoning-first field ordering matches or exceeds free-form accuracy, while answer-first ordering causes steep drops, particularly in smaller models. Format sensitivity scales inversely with a task's own structural constraints, and schemas that preserve reasoning order also improve calibration with CKA showing greater separability between correct and incorrect representations in middle transformer layers. Our findings indicate that properly designed structured formats can match or exceed free-form performance, reframing the critical question from `whether to structure' to `how to structure' for optimal reasoning preservation.