🤖 AI Summary
This study addresses the limitation of existing evaluation frameworks for large language model (LLM) generation, which fail to disentangle coverage improvements arising from repeated execution from genuine task-specific specialization. To resolve this, we propose a novel evaluation criterion that decouples coverage from specialization. Using the MATH-500 benchmark, we compare eight generation frameworks against nine baseline replicas, employing multi-execution mechanisms to trace failure modes and establishing diversity evaluation protocols centered on cross-execution persistence and budget alignment. Our analysis reveals that identical programs yield merely a 2.16% coverage margin, while generated programs predominantly expose persistent weaknesses without performance gains from frozen selectors. These findings demonstrate that raw coverage metrics alone are insufficient to substantiate effective task-specific specialization in LLM generation systems.
📝 Abstract
Automated generation of LLM harnesses promises to improve inference through task specialization. Yet additional answer coverage can arise from repeated execution of the same program, making specialization difficult to identify. We introduce a controlled evaluation that separates answer coverage, repeatable task advantages, and gains from pre-execution selection. On 386 MATH-500 tasks, we compare eight generated harnesses plus a baseline with nine byte-identical baseline copies, using three executions per member. Identical programs yield 2.16 percentage points of repeat-averaged oracle headroom. Generated programs exhibit substantially more repeatable score patterns, but these chiefly reveal persistent weaknesses: losses relative to the baseline persist across all three repeats on 100 tasks, while persistent wins occur on only one task and are sensitive to answer extraction. The frozen selector gains 0.00 percentage points, and both populations reach 98.70% oracle coverage at 27 harness executions. Stable complementarity remains unresolved at three repeats. Supporting BIRD traces locate failures in mechanism implementation, activation, and output validity. Together, these findings establish why coverage and repeatability alone cannot justify claims of useful specialization. They motivate an evaluation standard for harness diversity: task advantages should persist across executions, guide usable decisions, and improve on additional fixed-program executions under matched inference budgets.