TCSAlgBench: Benchmarking Automated Proving for Research-Level Theoretical Computer Science

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of systematic evaluation of large language models' reasoning capabilities in research-level theoretical computer science by proposing a refreshable benchmark comprising 398 theorem challenges. The benchmark preserves computational assumptions through expert rules, supports the discovery and verification of natural language proofs, and introduces a reusable pipeline for automatically generating versioned test sets from newly published papers, thereby bridging the evaluation gap in research-level automated proving. Methodologically, it integrates large language models with a prover-verifier discussion mechanism, agent-based decomposition planning, and multi-round sampling for automated proof search. Experimental results demonstrate that GPT-5.6 Sol max achieves a coverage rate of 23.6%, which improves to 25.4% when combined with agent-based planning, significantly outperforming direct reasoning baselines.
📝 Abstract
Large language models perform strongly on competition mathematics, but their research-level reasoning remains difficult to evaluate systematically. Theoretical computer science (TCS) connects algorithm design to explicit guarantees and fundamental limits, providing a setting for evaluating whether models can justify computational improvements with arguments humans can inspect. We introduce TCSAlgBench, a benchmark and reusable pipeline for natural-language proof discovery, comprising 398 theorem-level challenges from 138 STOC and COLT 2026 papers. Expert-designed rules complete paper-specific context, preserve computational assumptions and quantitative guarantees, and withhold constructions when discovering an algorithm is part of the task. For each task, prover systems receive theorem statements and access to cited prior work. The pipeline supports fresh, versioned challenge batches from newly released papers. We evaluate ten model configurations from four families under direct inference and prover-verifier discussion, and compare four agent workflows under matched model-call opportunities. All evaluations use the full benchmark. In the model comparison, GPT-5.6 Sol max achieves the highest five-run verifier-accepted coverage at 23.6% after 10-round discussion. Discussion and repeated sampling improve coverage. In the separate agent comparison using GPT-5.5 xhigh, decomposition improves coverage over discussion, and agentic planning achieves the highest five-run verifier-accepted coverage at 25.4%. TCSAlgBench provides a refreshable testbed for measuring progress in model reasoning and studying how agent workflows support research-level proof discovery.
Problem

Research questions and friction points this paper is trying to address.

automated proving
theoretical computer science
benchmarking
large language models
research-level reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Automated Theorem Proving
Benchmark Pipeline
Theoretical Computer Science
Prover-Verifier Discussion
Agentic Planning
🔎 Similar Papers
No similar papers found.