automated pipeline synthesis

Designs and builds systems that automatically assemble, configure, and validate sequences of interoperable algorithmic components into executable pipelines; produces and enumerates candidate pipeline configurations and context-specific optimization workflows for evaluation and deployment.

automatedpipelinesynthesis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the heavy reliance on expert knowledge in designing and debugging scientific workflows, a challenge exacerbated by existing large language model approaches that directly generate code without ensuring transparency, reproducibility, or seamless system integration. To overcome these limitations, we propose an AI-assisted scientific workflow management framework that decouples user intent from implementation through a structured specification phase, enabling specification-driven workflow generation and validation. We further introduce a multi-layer debugging agent powered by large language models to automate error diagnosis and correction. By deeply integrating with the Pegasus workflow system via the Model Context Protocol (MCP), our approach supports end-to-end workflow lifecycle management. Empirical evaluation demonstrates successful generation and execution of federated learning medical imaging workflows comprising thousands of tasks, substantially reducing debugging effort and empowering non-expert users to construct complex workflows adhering to expert-level design patterns.

debugginglarge language modelsreproducibility

This work addresses the challenge developers face in efficiently authoring CI/CD configurations due to limited DevOps expertise by proposing a large language model (LLM)-based, context-aware generation approach. The method leverages both natural language descriptions and repository structure to automatically produce accurate and executable pipeline configurations for platforms such as GitHub Actions and GitLab CI/CD. Integrated with automated validation and human-in-the-loop feedback mechanisms, this framework is the first to combine repository context understanding with natural language-driven configuration synthesis. Experimental results demonstrate that the approach significantly lowers the barrier to DevOps adoption, markedly improves the accuracy and validity of generated configurations, and substantially reduces manual configuration effort.

CI/CD pipeline configurationconfiguration errorsdeveloper productivity

Existing coding agents are largely confined to code generation and lack support for the full workflow lifecycle, including composition, iteration, deployment, and sharing. This work proposes CURATE, a novel system that integrates modular cataloging and FAIR principles into a large language model–driven multi-agent framework to enable human-in-the-loop, end-to-end workflow development and automated execution. Built upon Claude Opus 4.8, CURATE incorporates user-in-the-loop mechanisms and a module registry to facilitate cross-workflow sharing of reusable components. The system successfully reproduces four SeBS-Flow benchmark workflows and automatically constructs a complex anaerobic digestion simulation pipeline, demonstrating its feasibility and effectiveness in supporting comprehensive, collaborative scientific workflow automation.

code generationdeploymentmodule reuse

Automatic Pipeline Provisioning

Nov 18, 2025
AL
Alexandre-Xavier Labonté-Lamoureux
🏛️ École de Technologie Supérieure

To address the challenges of prolonged CI pipeline deployment cycles, error-prone manual configuration, and poor cross-project consistency, this paper proposes an automated pipeline configuration framework grounded in Infrastructure-as-Code (IaC) principles and templated configuration. The framework enables declarative definition and one-click generation of CI/CD pipelines via reusable YAML templates, a parameterized pipeline engine, and an integrated automation toolchain. Compared to conventional manual approaches, our method reduces average pipeline deployment time by 72% and decreases human configuration errors by 91%, while substantially improving consistency in build logic and execution environments across projects. Empirical validation across six open-source projects demonstrates the framework’s engineering practicality and methodological generality. It provides a reusable implementation model and actionable methodology for CI/CD automation, advancing scalable, maintainable, and reproducible software delivery practices.

Applying automatic deployment for software engineering projectsExploring benefits of automatic pipeline provisioningFocusing on CI pipelines with similar CD implications

This work addresses the challenge of efficiently balancing accuracy and latency in structured large language model (LLM) workflows, where the combinatorial design space—spanning model selection, inference budgets, and pipeline architecture—is prohibitively vast. To tackle this, the study introduces, for the first time, machine learning compilation principles into LLM workflow optimization. It performs global design space exploration prior to deployment by leveraging sub-agent decomposition, multi-configuration performance profiling, and a structure-aware surrogate model to accurately estimate and jointly optimize workflow-level accuracy and latency. The proposed approach generates a reusable set of Pareto-optimal configurations that span diverse accuracy–latency trade-offs, without requiring online adaptation or retraining. Experiments across multiple complex workflows and benchmarks demonstrate substantial improvements over heuristic and routing baselines, achieving up to 6.4× speedup while enabling flexible deployment and downstream scheduling.

accuracy-latency trade-offcompile-time optimizationdesign space exploration

Latest Papers

What's happening recently
View more

This work addresses the challenges of low reliability, poor auditability, high cost, and security risks associated with runtime invocation of large language models (LLMs) in high-stakes enterprise workflows. To overcome these issues, we propose a novel “compiled AI” paradigm, wherein LLMs generate executable code during compilation, eliminating the need for model calls at runtime and thereby ensuring deterministic execution. We present the first systematic application of this paradigm to high-risk scenarios, integrating constrained code generation, a four-stage verification pipeline, template-embedded business logic functions, and an operation-oriented evaluation framework to jointly achieve reliability, auditability, and security. Experiments demonstrate a 96% success rate on function-calling tasks with zero runtime token consumption; 80.0% and 80.4% accuracy on critical field extraction and line-item recognition in document intelligence tasks; and strong security performance, with 96.7% prompt injection detection accuracy and 87.5% static analysis precision without false positives.

auditabilitycompiled AIdeterministic execution

This study addresses the heavy reliance of scientific software optimization on domain experts, which impedes large-scale data processing. To overcome this limitation, this work proposes a framework in which large language model (LLM) agents autonomously optimize mature scientific software. Within this paradigm, human involvement is restricted to defining objectives and verification mechanisms, while automated acceleration is achieved through algorithmic restructuring and low-level code optimization. The contributions demonstrate that, for verifiable problems, artificial intelligence can attain automated optimization surpassing manual efforts, thereby reshaping human–machine collaboration paradigms. Empirically, the proposed approach yields up to two orders of magnitude speedup in tasks such as t-SNE and discovers novel graph counting algorithms.

autonomous agentslarge language modelsperformance improvement

Hot Scholars

DW

Dingmin Wang

Applied Scientist@AWS AI Lab
Natural Language ProcessingLLMs for Code
QC

Quanquan C. Liu

Yale University
Graph AlgorithmsParallel/Distributed AlgorithmsHPCGraph Differential Privacy
XY

Xu-Yao Zhang

Institute of Automation, Chinese Academy of Sciences
Pattern RecognitionMachine LearningOCR