🤖 AI Summary
This study addresses the challenges of latent semantic errors and repetitive repair failures in code generation using open-source large language models. To this end, it proposes an execution-guided, memory-augmented Monte Carlo Tree Search (MCTS) framework. Specifically, the method employs an LLM to guide MCTS in organizing candidate programs while leveraging long-term memory retrieval to share cross-branch failure experiences, thereby preventing redundant trial-and-error. Furthermore, a reflection mechanism combined with branch-local debugging contexts is introduced to precisely distinguish failed assertions from missing evidence, enabling stable and efficient code search. Extensive evaluations on the HumanEval and MBPP benchmarks demonstrate that the proposed framework significantly outperforms direct generation methods under most configurations, while also confirming its compatibility across different compiler backends.
📝 Abstract
Open-weight large language models (LLMs) can generate function-level programs from natural-language prompts, but plausible candidates still fail on hidden semantics and repeat mistakes across repair attempts. We present ReMCTS, an execution-grounded, memory-augmented, LLM-guided MCTS-style search framework. It organizes program candidates as tree states, retains branch-local debugging context, retrieves failure experience across branches, and distinguishes failed checks from unavailable evidence. On HumanEval and MBPP-Sanitized, visible-test ReMCTS improves over direct generation in 8 of 10 model-dataset pairs under held-out evaluation, whereas proxy-only search is less stable. Controlled tree-search, sampling, repair, and memory ablations characterize the source and limits of these gains. A 30-task HumanEval-X C++ pilot further demonstrates compatibility with compiler-backed execution, but does not constitute a broad multilingual evaluation.