🤖 AI Summary
This study addresses the high interaction costs and inefficiencies that arise when AI agents use existing proof assistant interfaces designed primarily for humans. We propose an LLM-driven evolutionary tool design paradigm that incrementally generates and filters features using frontier models, integrated with the Model Context Protocol (MCP) architecture and automated benchmarking, to construct AI-agent-optimized interfaces for Rocq and Lean. This approach pioneers a transition from human-centric to machine-centric interface optimization while demonstrating cross-system transferability. Experimental results show that our method significantly improves success rates and reduces solving overhead on miniF2F-Rocq. Furthermore, it successfully transfers to Lean, yielding enhanced performance on PutnamBench.
📝 Abstract
Recent achievements in AI-assisted mathematics require intensive interaction of agents with proof assistants to generate machine-checked proof certificates. Agents interact with proof assistants such as Rocq or Lean through an interface that controls what the agent receives from the prover and the cost of these interactions. Today, these interfaces are adapted from tools designed for humans and not optimized for agents. We propose an evolutionary method where a frontier model incrementally proposes new features and only keeps the ones that improve the overall performance of smaller models. We demonstrate the effectiveness of our method by growing, on a curated set of mathematical problems, \rme, a new MCP server for the Rocq prover. On the held-out \texttt{test} split of miniF2F-Rocq, an agent equipped with \rme outperforms both the baseline that only exposes the Rocq compiler and an established MCP server, across four models from two families, in success rate, cost per solve, and time per solve. Although evolved for Rocq, the resulting server transfers to Lean, improving cost and time per solve on a subset of PutnamBench. We release \rme and its port to Lean.