Institution profile

SINOPEC Research Institute of Petroleum Processing

Academic institutionasia · cn
Official website
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses

Sep 26, 2026

This study addresses the limitation of existing evaluation frameworks for large language model (LLM) generation, which fail to disentangle coverage improvements arising from repeated execution from genuine task-specific specialization. To resolve this, we propose a novel evaluation criterion that decouples coverage from specialization. Using the MATH-500 benchmark, we compare eight generation frameworks against nine baseline replicas, employing multi-execution mechanisms to trace failure modes and establishing diversity evaluation protocols centered on cross-execution persistence and budget alignment. Our analysis reveals that identical programs yield merely a 2.16% coverage margin, while generated programs predominantly expose persistent weaknesses without performance gains from frozen selectors. These findings demonstrate that raw coverage metrics alone are insufficient to substantiate effective task-specific specialization in LLM generation systems.

0 citationsRead paper
Recent publications

Latest Papers

More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses

Sep 26, 2026

This study addresses the limitation of existing evaluation frameworks for large language model (LLM) generation, which fail to disentangle coverage improvements arising from repeated execution from genuine task-specific specialization. To resolve this, we propose a novel evaluation criterion that decouples coverage from specialization. Using the MATH-500 benchmark, we compare eight generation frameworks against nine baseline replicas, employing multi-execution mechanisms to trace failure modes and establishing diversity evaluation protocols centered on cross-execution persistence and budget alignment. Our analysis reveals that identical programs yield merely a 2.16% coverage margin, while generated programs predominantly expose persistent weaknesses without performance gains from frozen selectors. These findings demonstrate that raw coverage metrics alone are insufficient to substantiate effective task-specific specialization in LLM generation systems.

0 citationsRead paper