🤖 AI Summary
Existing code generation benchmarks are confined to single programming languages, impeding the unified cross-ecosystem evaluation of large language models (LLMs). This work proposes the first multilingual reproduction test generation framework, constructing a unified evaluation benchmark spanning eight programming languages and conducting a systematic empirical study of LLMs via an agent-based architecture. The study quantifies performance disparities across languages, such as from Python to C++, and identifies critical failure modes including systematic language gaps and implicit context configuration issues. Ultimately, this research delineates the core bottlenecks and improvement directions for generalizing the coding capabilities of LLMs.
📝 Abstract
Reproduction test generation translates a natural-language issue description into executable tests that fail on the original code and pass after the issue is resolved, providing executable evidence for verifying candidate patches. Existing benchmarks are constructed for individual programming languages, preventing a unified evaluation across diverse programming ecosystems. To address this limitation, we introduce MULTI-SWT-BENCH, a multilingual benchmark for reproduction test generation consisting of 1,963 instances across eight programming languages: Python, Java, TypeScript, JavaScript, Go, Rust, C, and C++. Using this benchmark, we conduct an empirical study of state-of-the-art LLMs with four representative methods (MSWE-agent, MOpenHands, Codex, and Claude Code) and perform a failure analysis across programming languages. Our evaluation reveals a systematic language gap. Across every evaluated method and LLM, the success rate on Python exceeds the aggregate success rate across all languages, while C++ exhibits particularly low success rates. Our failure analysis identifies both language-specific challenges arising from repository testing conventions and cross-language challenges in inferring implicit setup requirements and preserving the target behavior through iterative revisions. These findings demonstrate the importance of multilingual evaluation and provide actionable directions for developing reproduction test generation methods that generalize across software ecosystems and reliably capture issue-specific behavior.