agentic proof planning

Design, build, and analyze multi‑agent systems and engineering workflows that plan, synthesize, and automate the construction of formal proofs or theorems. Work includes decomposing proof tasks into subtasks, coordinating agent actions to construct and iteratively refine proof attempts, using model-based guidance and task decomposition to drive search, and engineering pipelines for automated proof synthesis and verification.

agenticproofplanning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work proposes a novel paradigm termed “agent-based proof automation” to address the high cost of manually crafting lengthy formal proof scripts. In this approach, human experts supply key mathematical insights, while large language model (LLM) agents autonomously generate and iteratively refine proof scripts within the Lean 4 environment. Relying solely on off-the-shelf LLMs and lightweight verification tools, the method demonstrates— for the first time—the capacity for efficient, large-scale collaborative formal verification by LLM agents. Evaluated on the 14,000-line System Capless type safety proof, the system successfully completed 189 out of 217 tasks (87% success rate), with only 16% of the tasks requiring human intervention.

automationformal verificationlarge language models

This work addresses the vulnerability of existing large-scale multi-agent systems to failure in complex tasks due to error propagation and insufficient verification mechanisms. The authors propose a two-stage framework that automatically constructs and executes task-specific multi-agent systems from natural language instructions, incorporating dual verification mechanisms—during both construction and runtime. The approach decomposes tasks into directed acyclic graphs, defines input/output contracts, grounds knowledge via web search, and auto-generates prompts and tools. It further introduces a three-level error attribution scheme and intermediate output validation gating to enable targeted recovery strategies. Experimental results demonstrate that the method significantly outperforms strong baselines across programming, in-context learning, and open-ended reasoning tasks, consistently improving task success rates, error recovery capability, and workflow stability.

error propagationmulti-agent systemstask decomposition

This work addresses the lack of structural correctness verification during the design phase in existing AI agent workflow platforms, which typically rely on runtime safeguards. The authors propose a workflow modeling approach centered on reusable building blocks and introduce, for the first time, a set of twelve structural rules. By leveraging graph-based representations and a rule engine, the method enables static, formal checks for compatibility and logical consistency at design time. Experimental evaluation demonstrates that the prototype system efficiently detects design violations on a dataset comprising 48 defective workflows and 168 structural variants, maintaining high detection accuracy even when tasks are split across multiple agents. This significantly enhances the reliability and maintainability of workflow designs.

Agentic AIBuilding BlocksConceptual Models

This work addresses the degradation of reasoning quality and unreliable verification in LLM-based multi-agent systems for scientific computing—specifically linear-elastic finite element analysis—caused by collaborative dynamics. We systematically identify three systemic failure modes: confirmation bias, premature consensus, and verification–validation decoupling, leading to undetected physics-inconsistent code. Building upon the AutoGen framework, we design a role-specialized tri-agent system (Coder/Executor/Critic) and evaluate it via controlled dialogue experiments under a dual-criteria assessment paradigm: physical consistency and executable correctness. Results show that functional complementarity outweighs team size; Critic involvement achieves 100% correctness in both physics and visualization; and confirmation bias is detected with 85–92% accuracy. Based on these findings, we propose three actionable design principles—role differentiation, multi-level verification, and anti-premature-convergence interaction—to establish a foundation for engineering-grade trustworthy multi-agent systems.

Addressing systematic failure modes in automated computational workflows for engineeringDeveloping design principles for reliable multi-agent collaboration in finite element analysisExamining how inter-agent dynamics affect reasoning quality in multi-agent LLM systems

Autonomous Task Planning for Heterogeneous Multi-Agent Systems

Sep 18, 2022
AA
Anatoli A. Tziola
🏛️ Cyprus University of Technology

This paper addresses automatic task planning for heterogeneous multi-agent systems operating under dynamic fault conditions. Method: We propose a fault-aware two-layer task synthesis framework. Agent capabilities, operational constraints, and fault models are uniformly encoded using ε-transition-enabled nondeterministic finite automata (NFAs). A decoupled design separates the system model, initial state, and task specification—enabling online injection of fault modes. Formal verification is integrated with heuristic search to ensure solution completeness and optimality while substantially reducing computational overhead. Results: Experiments demonstrate that the approach maintains effectiveness and robustness across diverse fault scenarios. It provides a scalable, formal planning paradigm for high-reliability autonomous decision-making in multi-agent systems.

Automatic task planning for multi-agent systemsGenerating optimal solutions with system constraintsIncorporating agent capabilities and failure modes

Latest Papers

What's happening recently
View more

This work proposes Mechanic, a novel automated theorem proving system that addresses the inefficiency and contextual redundancy plaguing existing approaches when tackling complex mathematical problems. Traditional methods often suffer from costly full-proof regeneration or iterative error correction. Mechanic introduces a formal decomposition strategy centered on Lean’s `sorry` placeholders, precisely isolating unresolved subgoals while preserving the verified proof structure. By extracting failed subproblems into independent contexts, the system delegates them to large language model agents for targeted resolution. This approach simultaneously ensures proof reuse and maintains contextual conciseness, overcoming the limitations of conventional regeneration or repair paradigms. Evaluated on challenging mathematical competition benchmarks—including IMO 2025 and Putnam 2025—Mechanic demonstrates significant gains in proof efficiency.

automated theorem provingcontext lengthmathematical reasoning

This study addresses the disruptive impact of large language models and AI agent systems—capable of generating vast volumes of code—on traditional software engineering paradigms. The work proposes a new paradigm centered on agent orchestration, verification of AI-generated code, and structured human-AI collaboration. Through a structured synthesis of literature review and industry practices, it constructs a comprehensive framework encompassing education, toolchains, lifecycle management, and governance. The research reveals a fundamental shift in the nature of code—from a scarce craft artifact to a consumable commodity—and identifies the evolving role of software engineers toward system design, semantic validation, and accountability oversight. It further establishes key directions such as a verification-first software development lifecycle, offering both theoretical grounding and practical pathways for software engineering transformation in the AI era.

Agentic AI SystemsAI-generated CodeHuman-AI Collaboration

This work addresses the lack of concise and reproducible baseline systems in AI-driven automated theorem proving, which hinders fair architectural comparisons. To this end, we propose a minimalist yet competitive proof agent that integrates three core mechanisms: iterative proof refinement, theorem library retrieval, and context management. The system enables systematic evaluation of diverse large language models and design choices, achieving performance on par with state-of-the-art methods across multiple heterogeneous benchmarks. Our experiments demonstrate that iterative proof generation significantly outperforms single-pass generation, offering superior sample efficiency and reduced inference cost. The codebase is publicly released to provide the community with a standardized reference implementation.

agentautomated theorem provingbaseline

Formal methods remain underutilized due to their high mathematical barrier, and existing large language model (LLM) approaches are typically confined to isolated tasks, lacking mechanisms for the co-evolution of modeling and verification. This work proposes the first LLM-based agent framework that supports joint iterative refinement of formal models and proofs: it generates an initial Event-B model from natural language requirements and leverages feedback from formal verification—such as theorem proving and model checking—to iteratively refine and automatically repair the model, thereby emulating the intertwined nature of modeling and verification in real-world software design. Experimental results demonstrate that the approach significantly outperforms baseline methods across systems of varying complexity, achieving higher success rates in end-to-end formal model synthesis and repair while maintaining reasonable computational efficiency.

autoformalizationcorrect-by-constructionformal methods

Hot Scholars

CT

Chaofan Tao

The University of Hong Kong
Efficient MLNatural Language ProcessingMultimodal
JX

Jing Xiong

The University of Hong Kong
Natural Language ProcessingAutomated Theorem Proving
DM

Debmalya Mandal

Assistant Professor, University of Warwick
Computational Social ChoiceAlgorithmic FairnessReinforcement Learning
LL

Laurentiu Leustean

University of Bucharest & IMAR & Institute for Logic and Data Science
Mathematical logicProof theoryOptimizationErgodic theory