agentic proof planning

Design, build, and analyze multi‑agent systems and engineering workflows that plan, synthesize, and automate the construction of formal proofs or theorems. Work includes decomposing proof tasks into subtasks, coordinating agent actions to construct and iteratively refine proof attempts, using model-based guidance and task decomposition to drive search, and engineering pipelines for automated proof synthesis and verification.

agenticproofplanning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.08
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work proposes a novel paradigm termed “agent-based proof automation” to address the high cost of manually crafting lengthy formal proof scripts. In this approach, human experts supply key mathematical insights, while large language model (LLM) agents autonomously generate and iteratively refine proof scripts within the Lean 4 environment. Relying solely on off-the-shelf LLMs and lightweight verification tools, the method demonstrates— for the first time—the capacity for efficient, large-scale collaborative formal verification by LLM agents. Evaluated on the 14,000-line System Capless type safety proof, the system successfully completed 189 out of 217 tasks (87% success rate), with only 16% of the tasks requiring human intervention.

automationformal verificationlarge language models

This work addresses the vulnerability of existing large-scale multi-agent systems to failure in complex tasks due to error propagation and insufficient verification mechanisms. The authors propose a two-stage framework that automatically constructs and executes task-specific multi-agent systems from natural language instructions, incorporating dual verification mechanisms—during both construction and runtime. The approach decomposes tasks into directed acyclic graphs, defines input/output contracts, grounds knowledge via web search, and auto-generates prompts and tools. It further introduces a three-level error attribution scheme and intermediate output validation gating to enable targeted recovery strategies. Experimental results demonstrate that the method significantly outperforms strong baselines across programming, in-context learning, and open-ended reasoning tasks, consistently improving task success rates, error recovery capability, and workflow stability.

error propagationmulti-agent systemstask decomposition

This study addresses the challenges of correctness verification and cross-domain comprehension for AI-generated code in scientific computing by proposing a deterministic, orchestrator-based human-AI collaborative workflow. In this framework, domain experts define specifications while AI agents perform coding, with quality ensured through numerical consistency checks and manual review. The core innovation lies in a lightweight "author-reviewer" loop that replaces complex multi-agent systems to enable efficient Fortran-to-C++ code migration. High-energy physics experiments demonstrate that this streamlined iterative architecture reduces costs to approximately one-third of those incurred by comparable approaches on equivalent tasks, achieving both reliability and cost-effectiveness.

Agentic AICode TranslationCode Verification

This work addresses the degradation of reasoning quality and unreliable verification in LLM-based multi-agent systems for scientific computing—specifically linear-elastic finite element analysis—caused by collaborative dynamics. We systematically identify three systemic failure modes: confirmation bias, premature consensus, and verification–validation decoupling, leading to undetected physics-inconsistent code. Building upon the AutoGen framework, we design a role-specialized tri-agent system (Coder/Executor/Critic) and evaluate it via controlled dialogue experiments under a dual-criteria assessment paradigm: physical consistency and executable correctness. Results show that functional complementarity outweighs team size; Critic involvement achieves 100% correctness in both physics and visualization; and confirmation bias is detected with 85–92% accuracy. Based on these findings, we propose three actionable design principles—role differentiation, multi-level verification, and anti-premature-convergence interaction—to establish a foundation for engineering-grade trustworthy multi-agent systems.

Addressing systematic failure modes in automated computational workflows for engineeringDeveloping design principles for reliable multi-agent collaboration in finite element analysisExamining how inter-agent dynamics affect reasoning quality in multi-agent LLM systems

Autonomous Task Planning for Heterogeneous Multi-Agent Systems

Sep 18, 2022
AA
Anatoli A. Tziola
🏛️ Cyprus University of Technology

This paper addresses automatic task planning for heterogeneous multi-agent systems operating under dynamic fault conditions. Method: We propose a fault-aware two-layer task synthesis framework. Agent capabilities, operational constraints, and fault models are uniformly encoded using ε-transition-enabled nondeterministic finite automata (NFAs). A decoupled design separates the system model, initial state, and task specification—enabling online injection of fault modes. Formal verification is integrated with heuristic search to ensure solution completeness and optimality while substantially reducing computational overhead. Results: Experiments demonstrate that the approach maintains effectiveness and robustness across diverse fault scenarios. It provides a scalable, formal planning paradigm for high-reliability autonomous decision-making in multi-agent systems.

Automatic task planning for multi-agent systemsGenerating optimal solutions with system constraintsIncorporating agent capabilities and failure modes

Latest Papers

What's happening recently
View more

This work addresses the lack of structural correctness verification during the design phase in existing AI agent workflow platforms, which typically rely on runtime safeguards. The authors propose a workflow modeling approach centered on reusable building blocks and introduce, for the first time, a set of twelve structural rules. By leveraging graph-based representations and a rule engine, the method enables static, formal checks for compatibility and logical consistency at design time. Experimental evaluation demonstrates that the prototype system efficiently detects design violations on a dataset comprising 48 defective workflows and 168 structural variants, maintaining high detection accuracy even when tasks are split across multiple agents. This significantly enhances the reliability and maintainability of workflow designs.

Agentic AIBuilding BlocksConceptual Models

Recent multi-agent LLM systems increasingly rely on graph-structured communication to coordinate specialized agents. We revisit multi-agent orchestration from a graph-engineering perspective: rather than optimizing a static topology, we synthesize a task-conditioned temporal workflow graph that jointly specifies agent connectivity and edge-level communication semantics. We introduce ReActNet, a training-free framework that compiles a query and a set of role-specialized agents into a sequence of directed communication graphs. Each graph snapshot corresponds to one reasoning stage, and each edge carries a natural-language instruction specifying the message that a source agent should provide to a target agent. The compiled temporal graph is then executed through structured message passing: agents update their reasoning states by integrating their previous states with messages from controller-assigned neighbors, and a final aggregator synthesizes the resulting states into the answer. This design separates graph compilation from graph execution, making multi-agent coordination explicit, inspectable, and task-conditioned without requiring reinforcement learning or gradient-based topology optimization. Across knowledge reasoning, mathematical problem solving, code generation, and GAIA-style assistant tasks, ReActNet consistently improves over fixed-topology and learned-topology baselines while maintaining competitive inference cost. These results suggest that effective multi-agent orchestration depends not only on which agents communicate, but also on engineering executable workflow graphs that encode when, why, and how information should flow during reasoning.

agent connectivitycommunication semanticsgraph-structured communication

This study addresses the limitation of single-pass generation in large language models when tackling open mathematical problems that require long-horizon exploration, overcoming technical obstacles, and preserving intermediate progress. To this end, it proposes a Gemini-based multi-agent orchestration framework in which an orchestrator assigns competing conjectural directions and coordinates independent provers through iterative proving and adversarial verification cycles. A persistent ledger is maintained to accumulate intermediate reasoning steps, thereby overcoming bottlenecks associated with long-horizon reasoning. The proposed approach yields novel mathematical results on five open problems spanning online learning, auction theory, and other domains. These findings have been independently verified by domain experts, demonstrating the framework’s effectiveness in addressing the challenges of long-range collaborative reasoning for complex mathematical problem-solving.

automated proof discoverylanguage modelsmathematical reasoning

This study addresses the disruptive impact of large language models and AI agent systems—capable of generating vast volumes of code—on traditional software engineering paradigms. The work proposes a new paradigm centered on agent orchestration, verification of AI-generated code, and structured human-AI collaboration. Through a structured synthesis of literature review and industry practices, it constructs a comprehensive framework encompassing education, toolchains, lifecycle management, and governance. The research reveals a fundamental shift in the nature of code—from a scarce craft artifact to a consumable commodity—and identifies the evolving role of software engineers toward system design, semantic validation, and accountability oversight. It further establishes key directions such as a verification-first software development lifecycle, offering both theoretical grounding and practical pathways for software engineering transformation in the AI era.

Agentic AI SystemsAI-generated CodeHuman-AI Collaboration

This study addresses the theoretical drift caused by model modifications and the challenges of automated construction in the formalization of stochastic optimization algorithms. We propose a fully automated formalization framework driven by large language model (LLM) agents. This framework employs proof obligations to guide the automatic construction of Lean models and supporting theories, introduces signature contracts alongside independent auditing mechanisms to prevent assumption weakening, and establishes a reusable verification library, SOptLib, to enable cumulative verification cycles. Experimental results demonstrate that the system achieves an average score of 6.3 out of 7 across 15 tasks, generates 490,000 lines of `sorry`-free code, and identifies 28 formula errors and proof gaps in published literature, thereby realizing highly reliable automated formalization of research-grade algorithms.

autoformalizationdomain theoryLean

Hot Scholars

CT

Chaofan Tao

The University of Hong Kong
Efficient MLNatural Language ProcessingMultimodal
JX

Jing Xiong

The University of Hong Kong
Natural Language ProcessingAutomated Theorem Proving
DM

Debmalya Mandal

Assistant Professor, University of Warwick
Computational Social ChoiceAlgorithmic FairnessReinforcement Learning
LL

Laurentiu Leustean

University of Bucharest & IMAR & Institute for Logic and Data Science
Mathematical logicProof theoryOptimizationErgodic theory