ARIADNE: Agentic Reward-Informed Adaptive Decision Exploration via Blackboard-Driven MCTS for Competitive Program Generation

📅 2026-05-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of large language models in generating programming competition problems—specifically, their lack of explicit algorithmic planning, difficulty in robustly handling boundary conditions, and inefficient use of execution feedback—by proposing a blackboard-driven Monte Carlo Tree Search (MCTS) framework. The approach formulates program synthesis as a sequential decision-making process that jointly optimizes five stages: strategy selection, code generation, test case generation, quality evaluation, and repair. By integrating a blackboard architecture with MCTS for the first time, the method enables continuous accumulation of structured evidence and coordinated multi-stage decision-making, substantially enhancing the reliability of program generation under constrained settings. Evaluated on four benchmarks—including APPS and CodeContests—the framework achieves state-of-the-art Pass@1 performance, outperforming the strongest baseline, CodeSim, by 26.06 percentage points when using GPT-4o.
📝 Abstract
Competitive program generation aims to automatically produce correct and efficient solutions for programming-contest problems under strict time and memory constraints. Existing LLM-based approaches often fail to perform explicit algorithmic planning and to handle edge cases robustly, leading to unreliable one-shot generation. Moreover, although execution feedback is essential for iterative debugging and refinement, incorporating such feedback effectively within limited computational budgets remains difficult. To overcome these limitations, we propose {\tool}, a blackboard-driven Monte Carlo Tree Search (MCTS) framework that models program generation as a sequential decision process. {\tool} organizes the generation workflow into five coordinated stages (i.e., strategy selection, code generation, test generation, quality evaluation, and code repair) while maintaining a shared blackboard that accumulates structured evidence to guide subsequent decisions. Experiments on four benchmarks (APPS, CodeContests, CodeContests+, and LiveCodeBench) show that {\tool} consistently achieves the best Pass@1 performance across multiple LLM backends. With GPT-4o, {\tool} attains Pass@1 scores of 41.30, 46.67, 27.27, and 20.91, surpassing the strongest baseline CodeSim by up to 26.06 points, while further improvements are observed with DeepSeek-V3.2. These results indicate that combining global search through MCTS with persistent evidence accumulation on a shared blackboard enables systematic exploration and effective feedback utilization, substantially enhancing the capability of LLMs in competitive program generation.
Problem

Research questions and friction points this paper is trying to address.

competitive program generation
algorithmic planning
edge cases
execution feedback
iterative debugging
Innovation

Methods, ideas, or system contributions that make the work stand out.

Blackboard-driven MCTS
Agentic Reward-Informed Exploration
Competitive Program Generation
Structured Evidence Accumulation
Sequential Decision Process
🔎 Similar Papers
No similar papers found.
M
Minnan Wei
School of Artificial Intelligence and Computer Science, Nantong University, Nantong, China
X
Xiang Chen
School of Artificial Intelligence and Computer Science, Nantong University, Nantong, China
X
Xiaoshuai Niu
School of Artificial Intelligence and Computer Science, Nantong University, Nantong, China
S
Siyu Chen
School of Artificial Intelligence and Computer Science, Nantong University, Nantong, China