Score
Designs and builds feedback-driven, closed-loop workflow systems that collect execution and experimental readouts, analyze failures and performance metrics, and automatically or semi-automatically update protocol definitions, control scripts, and automation code. Maintains and version-controls runtime skill libraries, prioritizes and validates corrective revisions using expert-grounded rubrics, and ensures workflows continuously adapt and improve.
This work addresses the lack of a unified framework in current LLM agent workflows, which hinders method comparison and reproducibility. To resolve this, we propose the Agent Computation Graph (ACG) framework, which models workflows as computation graphs and adopts “structure determines timing” as a core principle. The framework explicitly distinguishes between reusable templates, runtime instance graphs, and execution traces, enabling a systematic categorization of static and dynamic optimization approaches. Through a comprehensive literature review and conceptual modeling, we develop a multidimensional evaluation framework that integrates structural properties, establishes precise terminology, and defines standardized evaluation criteria. This foundation supports a reproducible and highly comparable research paradigm for optimizing LLM agent workflows.
This work addresses the limitations of existing agent workflows, which predominantly rely on abstract structures from large language models and lack genuine tool integration, resulting in poor usability and stability. To overcome this, we propose FlowScout, a novel framework that explicitly models real-world tool invocations as nodes in a directed graph. FlowScout integrates tool coordination skeleton mining with a Monte Carlo Tree Search mechanism guided by execution feedback to automatically optimize workflow topology. Experimental results across four task domains demonstrate that FlowScout significantly outperforms baseline methods—including PM4Py, ReAct, and AFlow—with at least a 92.69% improvement in tool invocation accuracy, a minimum 17.66% gain in execution quality, and enhanced runtime stability.
This work addresses the lack of closed-loop control in traditional software development lifecycles, which often fails to simultaneously ensure security, auditability, and highly reliable automation. The authors propose a deterministic autonomous control framework that models the lifecycle as a seven-stage automated pipeline, integrating Jira-based task orchestration, structured context, resource constraints, and human-review gating mechanisms to establish a secure closed loop. Key innovations include a state-contract-based collision locking mechanism, a degradation protocol for fallback operation, and a traceable control architecture. Implemented with 12,661 lines of Python code and 6,907 lines of versioned prompt specifications—including 101 exception handlers and 12 centralized locks—the system achieved a 100% success rate (95% CI [97.6%, 100%]) across 152 initial runs, producing over 795 artifacts. All 51 issues identified through adversarial review were fully resolved, with 60% of security tickets autonomously completed.
This study addresses the significant burden developers face in authoring and maintaining GitHub Actions workflows, stemming from a lack of systematic understanding of real-world automation and reuse practices. Through a mixed-methods approach combining a survey of 419 practitioners with qualitative and quantitative analysis, this work presents the first developer-centric characterization of common automation tasks, patterns of reuse mechanism adoption, and maintenance pain points in workflow development. The findings reveal that while developers heavily rely on reusable Actions, they seldom adopt reusable workflows; version management challenges lead to rampant copy-pasting; and critical aspects such as security and performance monitoring remain under-automated. These insights provide empirical foundations for improving CI/CD toolchains and reuse mechanisms.
This study addresses the challenges of control design in complex industrial processes characterized by multivariable coupled dynamics by proposing an automated control strategy generation framework that integrates large language models (LLMs) with Bayesian optimization. The approach decomposes control design into structured code generation steps, ensuring physical consistency through execution-based validation and feedback-driven repair. It pioneers the automatic synthesis of decentralized PI controller architectures and their tuning environments directly from dynamic process models. Evaluated on a nonlinear gas preheater benchmark, the generated control schemes—subsequently refined via Bayesian optimization—achieve a 26.5% improvement in closed-loop performance and significantly enhance the transient response of pressure loops, thereby demonstrating the method’s effectiveness and novelty.
This work proposes a fine-tuning-free, large language model (LLM)-driven approach to address the need for high-quality, context-aware, and goal-directed automated code review comments in enterprise settings. By leveraging prompt engineering, contextual retrieval, and a comment quality filtering mechanism, the authors developed and deployed RovoDev Code Reviewer—an integrated system within Atlassian Bitbucket. Evaluation over a one-year period in a real-world industrial environment demonstrates that 38.7% of the system’s automatically generated comments led developers to modify their code, resulting in a 30.8% reduction in average pull request (PR) cycle time and a 35.6% decrease in manual reviewer comments. The system also effectively identified actionable code defects, confirming its practicality and effectiveness without requiring model fine-tuning.
This study investigates the evolutionary mechanisms of self-evolving agent skills under multi-round feedback, focusing on how feedback type—success, failure, or both—affects skill refinement and whether test-time computation can reproduce the observed evolutionary gains. By constructing a controlled evaluation framework across five benchmarks and three models, the authors employ a multi-round feedback design with fixed executors, optimizers, and validation rules, complemented by byte-level difference detection and a validation-based selection mechanism. Their findings reveal that skill self-evolution is fundamentally a sparse search process critically dependent on validation filtering. Across 14 experimental settings, 11 successfully selected evolved skills, with 9 demonstrating improved test performance; notably, all effective evolutions required feedback incorporating failure trajectories. Moreover, test-time scaling with GPT-5.5 fails to fully recover these gains, highlighting a misalignment between validation criteria and downstream evaluation preferences.
Existing coding agents are largely confined to code generation and lack support for the full workflow lifecycle, including composition, iteration, deployment, and sharing. This work proposes CURATE, a novel system that integrates modular cataloging and FAIR principles into a large language model–driven multi-agent framework to enable human-in-the-loop, end-to-end workflow development and automated execution. Built upon Claude Opus 4.8, CURATE incorporates user-in-the-loop mechanisms and a module registry to facilitate cross-workflow sharing of reusable components. The system successfully reproduces four SeBS-Flow benchmark workflows and automatically constructs a complex anaerobic digestion simulation pipeline, demonstrating its feasibility and effectiveness in supporting comprehensive, collaborative scientific workflow automation.
This work addresses the challenges faced by large language model (LLM) agents in policy-constrained enterprise workflows—such as document auditing—including sparse feedback, performance degradation due to frequent rule updates, and the need to balance accuracy, reasoning cost, and auditability. To tackle these issues, the authors propose FRAMES, a framework that bootstraps deployable skills during cold-start and enables continuous skill evolution through consensus-based skill mutation, Pareto-optimal trade-offs between precision and computational cost, and a degradation-aware validation mechanism within a closed-loop system. FRAMES is the first approach to achieve auditable, efficient, and robust skill iteration in policy-intensive settings, demonstrating state-of-the-art precision–cost trade-offs on both an internal production system and the tau-bench benchmark, significantly outperforming existing baselines.