More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses
This study addresses the limitation of existing evaluation frameworks for large language model (LLM) generation, which fail to disentangle coverage improvements arising from repeated execution from genuine task-specific specialization. To resolve this, we propose a novel evaluation criterion that decouples coverage from specialization. Using the MATH-500 benchmark, we compare eight generation frameworks against nine baseline replicas, employing multi-execution mechanisms to trace failure modes and establishing diversity evaluation protocols centered on cross-execution persistence and budget alignment. Our analysis reveals that identical programs yield merely a 2.16% coverage margin, while generated programs predominantly expose persistent weaknesses without performance gains from frozen selectors. These findings demonstrate that raw coverage metrics alone are insufficient to substantiate effective task-specific specialization in LLM generation systems.