🤖 AI Summary
This study addresses the fragmentation in existing code generation benchmarks regarding language coverage, task granularity, and evaluation protocols by constructing the first unified cross-granularity evaluation framework. Integrating five programming languages and 58 real-world open-source repositories, the framework spans function- to repository-level tasks and introduces an execution-based integrated evaluation protocol to systematically assess the generative capabilities of large language models, specialized methods, and general-purpose coding agents. The findings reveal significant performance disparities across methods and granularities, as well as notable context augmentation effects. Specifically, generating complete code snippets remains challenging for current approaches, achieving only 31.0% accuracy at the repository level, whereas providing implementation context substantially improves the executable correctness of function-level generation.
📝 Abstract
As large language models increasingly move toward repository-level software engineering, existing code-generation benchmarks remain fragmented across language coverage, task granularity, and evaluation protocols, impeding systematic comparison. To address this gap, we present PolyCodeEval, a unified multilingual and multi-granularity benchmark for code generation. It comprises 2,590 code generation tasks spanning functions to repositories, derived from 58 real, executable open-source repositories in five programming languages. All tasks are evaluated under a unified execution-based protocol with integration procedures tailored to their generation targets. Building on this benchmark, we evaluate frontier large language models, state-of-the-art specialized methods, and general coding agents. Our results show that existing approaches still struggle to correctly generate complete code fragments across granularities and languages. Specifically, the studied methods generate at most 71.7%, 76.7%, and 31.0% correct functions, files, and repositories, respectively, with performance varying widely across languages. Paired experiments further show that implementation context from related functions in the same file improves the executable correctness of function generation. Method rankings also vary across task granularities and programming languages, highlighting the importance of multilingual, multi-granularity evaluation for comprehensively assessing code generation capabilities.