divide-and-conquer repair

Designs and builds automated program-repair workflows that partition faults or violations into clusters, decompose each cluster by root cause, and apply divide-and-conquer strategies to generate coordinated, focused patches (including edits across multiple files) for each subproblem.

divide-and-conquerrepair

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the unreliability of large language model (LLM)-driven agent workflows, which stems from output nondeterminism, complex node dependencies, and tool heterogeneity, and proposes FlowFixer—a novel framework that introduces symbolic reasoning into automated workflow repair. FlowFixer models execution traces symbolically to generate behavioral specifications, enabling precise fault localization and root cause identification, and dynamically synthesizes targeted repair patches. To reduce verification overhead, it incorporates a multidimensional pre-evaluation mechanism. Experimental evaluation on Dify, Coze, and n8n platforms demonstrates that FlowFixer achieves a repair success rate of 71.3%, outperforming existing methods by 11.9%–27.6%, and improves root cause analysis accuracy by 15.3%–38.8%.

agentic workflowautomatic repairfailure root cause

Towards Reliable Generation of Executable Workflows by Foundation Models

Sep 29, 2025
SM
Sogol Masoumzadeh
🏛️ McGill University | Huawei Canada | Queen’s University

Foundation models (FMs) exhibit low accuracy and high defect rates when generating domain-specific language (DSL) workflows, necessitating extensive manual debugging. Method: We propose a static-analysis–driven closed-loop repair framework. First, we establish a taxonomy of 18 defect classes specific to DSL workflows. Second, we develop Timon—the first static analysis tool tailored for DSL workflows—and integrate it into the FM generation pipeline to automate defect detection, localization, and repair. Contribution/Results: Our analysis reveals that 87.27% of FM-generated workflows contain detectable defects, nine of which are precisely identifiable by Timon. Empirical evaluation demonstrates that our framework significantly improves workflow executability and reliability, advancing end-to-end automation from natural language requirements to executable DSL workflows.

Automating workflow generation from natural language using foundation modelsDetecting and repairing defects in domain-specific language workflowsImproving reliability of executable workflow creation through static analysis

This work addresses the limited diversity in repair strategies generated by current large language models for automated program repair, which often stems from redundant execution traces and repetitive sampling. To overcome this, the authors propose CT-Repair, a novel framework that integrates static and dynamic evidence by combining Code Property Graphs (CPGs) with Temporal Execution Graphs (TEGs). CT-Repair introduces a finite state machine–guided multi-perspective agent collaboration mechanism, enabling independent generation and optimization of diverse repair strategies. Coupled with a three-stage filtering pipeline and validation feedback, the approach substantially enhances both repair diversity and accuracy. Evaluated on 854 Java bugs from Defects4J v3.0, CT-Repair successfully repairs 489, outperforming ReinFix and RepairAgent; the joint use of three perspectives yields 99 more fixes than the strongest single perspective, while execution-based filtering reduces the search space by an average of 94.85%.

Automated Program RepairExecution TracesPatch Sampling

Latest Papers

What's happening recently
View more

Existing tools struggle to effectively repair complex web accessibility (A11Y) issues that span multiple files and exhibit strong structural interdependencies. This work proposes a divide-and-conquer repair framework powered by large language models, which clusters related violations, decomposes them according to root causes, and integrates WCAG domain knowledge to enable coordinated and precise multi-file patch generation. Notably, it is the first approach to embed WCAG guidelines directly into the divide-and-conquer repair pipeline, substantially improving both consistency and efficiency. Evaluated on a newly constructed A11YBench benchmark, the method outperforms current state-of-the-art techniques, and its generated patches have been adopted by real-world open-source projects from Google, Microsoft, and Facebook, demonstrating its practical efficacy and applicability.

accessibility violationsautomated program repairmulti-fault repair

This work addresses the challenge that existing large language model–based program repair approaches struggle to effectively handle multi-hunk bugs requiring coordinated modifications across multiple code locations. To overcome this limitation, we propose MultiFixer, the first multi-agent framework featuring a coordinator–proposer architecture. MultiFixer integrates tool-augmented bug analysis, fine-grained contextual modeling, iterative patch generation, and a two-stage (syntactic–semantic) refinement mechanism to enable coordinated, cross-method and cross-file repairs with explicit repair sequencing. Evaluated on Defects4J, MultiFixer successfully fixes 326 bugs—including 62 multi-method and 27 multi-file defects—and achieves a new state of the art by repairing 420 bugs when combined with Claude-3.5-Sonnet, significantly outperforming current methods across multiple benchmark datasets.

Automated Program RepairLarge Language ModelsMulti-Hunk Bugs

This work addresses a critical yet overlooked reliability issue in code generated by large language models (LLMs): despite passing compilation and unit tests, such code often fails in deployment due to structural inconsistencies—such as missing configurations, invalid imports, or omitted security controls—that evade detection by conventional CI/SAST tools. The paper introduces the “patchwork problem” to characterize these cross-module global defects, proposes an eight-category taxonomy specific to LLM-generated code, and formalizes structural consistency via invariants derived from a multidimensional code graph encompassing imports, calls, dependencies, configurations, and routing. Building on this foundation, the authors design a hybrid verification framework that integrates traditional static analysis with custom graph-based invariant checkers to precisely identify structural flaws invisible to existing tools. Empirical evaluation reveals that such defects are pervasive across major LLMs under diverse prompting strategies and exhibit distinct model-specific patterns.

global inconsistencyLLM-generated codepatchwork problem

This study addresses the lack of systematic understanding regarding the impact of repair loop iteration counts in large language model (LLM)-based software engineering tasks, where prior work often relies on arbitrarily defined repair budgets. Through a cross-task (code generation, test generation, code translation) and cross-model empirical analysis, this work reveals—for the first time—a pronounced diminishing marginal returns phenomenon in iterative repair: performance gains are concentrated within the first 3–4 iterations, with negligible improvements thereafter. The findings underscore that the design of the repair workflow and feedback mechanisms exerts a far greater influence on repair efficacy than the choice of LLM itself. The authors advocate for treating repair budget as a critical experimental variable to ensure reliable, computationally efficient, and reproducible evaluation outcomes.

diminishing returnsiteration limitsLLM-based software engineering

Existing coding agents are largely confined to code generation and lack support for the full workflow lifecycle, including composition, iteration, deployment, and sharing. This work proposes CURATE, a novel system that integrates modular cataloging and FAIR principles into a large language model–driven multi-agent framework to enable human-in-the-loop, end-to-end workflow development and automated execution. Built upon Claude Opus 4.8, CURATE incorporates user-in-the-loop mechanisms and a module registry to facilitate cross-workflow sharing of reusable components. The system successfully reproduces four SeBS-Flow benchmark workflows and automatically constructs a complex anaerobic digestion simulation pipeline, demonstrating its feasibility and effectiveness in supporting comprehensive, collaborative scientific workflow automation.

code generationdeploymentmodule reuse

Hot Scholars

BS

Barna Saha

Harry E. Gruber Endowed Chair Professor, University of California San Diego
AlgorithmsProbabilistic MethodsData Management
SH

Sariel Har-Peled

Professor of Computer Science, UIUC
Computational Geometry
HB

Hadley Black

Postdoc in Computer Science, UC San Diego
theoretical computer sciencecombinatoricssublinear algorithmslearning theory
AM

Arya Mazumdar

HDSI Endowed Chair Professor in AI, University of California, San Diego
Information TheoryCoding TheoryLearning TheoryMathematical Statistics