Are you Synthesizing or Recalling? Evaluating LLMs on Algorithmic Code Retrieval

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the opacity of existing LLM code generation processes, which makes it difficult to disentangle algorithmic knowledge recall from logical reasoning capabilities. To this end, this work proposes the concept of parameterized code retrieval, formalizing classical algorithm code generation as a retrieval task, and constructs AlgoREval, a large-scale isolated evaluation benchmark encompassing multiple programming languages and diverse graph inputs. Through systematic experiments involving zero-shot evaluation, prompt enhancement, supervised fine-tuning (SFT), and GRPO-based reinforcement learning, the research reveals cross-lingual retrieval discrepancies among models. It further demonstrates that prompt enhancement improves accuracy on complex algorithms, while SFT and GRPO optimize retrieval breadth and depth, respectively. Ultimately, this work establishes a novel paradigm for precisely evaluating the algorithmic reproduction capabilities of large language models.
📝 Abstract
Large language models (LLMs) have demonstrated strong performance in code generation, where success depends on both recalling relevant algorithmic knowledge and reasoning about how to apply it. However, existing LLM pipelines are opaque, with no explicit separation between these two components. We argue that for well-known algorithms whose canonical implementations are widely accessible in pretraining corpora, code generation is better measured as \textit{parametric code retrieval}: reproducing a named algorithm from internalised knowledge rather than synthesizing a novel one. We introduce AlgoREval, a benchmark of 599 problems spanning classical 77 algorithms across 14 domains, 7 programming languages, and 4 graph-input representations to evaluate this capability in isolation, and assess 15 models (7B--34B parameters) in a zero-shot setting. We find substantial variation in retrieval accuracy across languages and input representations, even for widely documented algorithms and show that prompt augmentation with retrieved code snippets or structured algorithmic hints improve accuracy on complex algorithms, while SFT achieves broader language gains and GRPO achieves larger per-language gains on specific languages. Together, our results establish parametric code retrieval as a distinct, measurable capability and caution against deploying AI-generated algorithmic code without systematic validation.\footnote{Code and dataset are available at https://github.com/Nickil21/AlgoREval
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Code Generation
Algorithmic Code Retrieval
Parametric Knowledge
Evaluation Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Parametric Code Retrieval
AlgoREval Benchmark
Algorithmic Code Generation
Prompt Augmentation
GRPO
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
N
Nickil Maveli
School of Informatics, University of Edinburgh
Antonio Vergari
Antonio Vergari
Reader (Associate Professor), University of Edinburgh, UK
Artificial IntelligenceProbabilistic Machine LearningProbabilistic CircuitsNeuro-Symbolic AI
S
Shay B. Cohen
School of Informatics, University of Edinburgh