Score
Designs, implements, and operates the processes, tooling, and automations that keep production services running—including monitoring, alerting, deployment and rollback mechanisms, runbooks, and on‑call rotations. Builds and executes incident response and post‑incident review practices and analyzes alerts, logs, and metrics to reduce time‑to‑detect/repair and prevent recurrence.
本文介绍了Meta为解决大规模系统持续部署中速度与可靠性之间的矛盾,通过构建服务健康检查器进行自动回滚等方法保障部署安全。
To address the challenges of standardizing Site Reliability Engineering (SRE) practices in heterogeneous environments and balancing system reliability with development agility, this paper proposes a customizable SRE process framework. The framework integrates automated operations, multidimensional observability (metrics, logs, traces), error-budget-driven governance, standardized incident response, and progressive delivery (canary and blue-green deployments). It is designed for cross-technology-stack adaptability, enabling contextual implementation of core SRE principles. Evaluated in production systems, the framework reduced mean time to recovery by 42%, decreased unplanned outages by 67%, lowered operational staffing requirements by 35%, and achieved 99.99% service availability. Its primary contribution is the first methodology for customizing SRE processes specifically for heterogeneous environments, empirically demonstrating synergistic improvements in both system reliability and operational efficiency.
This study addresses the challenge of effectively monitoring early-stage agent systems, where structural flaws often obscure task-level errors. The authors propose a three-dimensional (quality, suitability, efficiency) and three-granularity (intra-run, inter-run, structural) monitoring and triaging framework tailored for low-maturity agent systems. They introduce a novel system maturity staging model based on the coefficient of variation and monitoring granularity, integrated with a severity classification adapted from FMEA to guide human review. The resulting transferable monitoring architecture supports document-driven, multi-stage workflows, enhanced by a synthetic testbed with controlled error injection. Experimental results demonstrate that structural defects significantly mask task-level signals; 97% of issues can be automatically traced, with only 2% requiring human intervention, and each granularity level precisely identifies its corresponding defect type (coefficients of variation: 0.02, 1.25, and 0.00, respectively).
This work addresses the lack of closed-loop control in traditional software development lifecycles, which often fails to simultaneously ensure security, auditability, and highly reliable automation. The authors propose a deterministic autonomous control framework that models the lifecycle as a seven-stage automated pipeline, integrating Jira-based task orchestration, structured context, resource constraints, and human-review gating mechanisms to establish a secure closed loop. Key innovations include a state-contract-based collision locking mechanism, a degradation protocol for fallback operation, and a traceable control architecture. Implemented with 12,661 lines of Python code and 6,907 lines of versioned prompt specifications—including 101 exception handlers and 12 centralized locks—the system achieved a 100% success rate (95% CI [97.6%, 100%]) across 152 initial runs, producing over 795 artifacts. All 51 issues identified through adversarial review were fully resolved, with 60% of security tickets autonomously completed.
Industrial applications heavily rely on open-source libraries, yet stalled community maintenance frequently leaves vulnerabilities unpatched for extended periods, posing critical software supply chain security risks. Existing approaches suffer from label scarcity, sparse feature representations, and incomplete modeling of transitive dependency relationships, hindering practical deployment in industrial settings. This paper proposes the first maintenance-activity monitoring framework that jointly models direct and transitive dependencies. It constructs fine-grained maintenance metrics from multi-source repository metadata—including commits, releases, issues, and pull requests—and introduces a graph propagation model to quantify the cross-dependency transmission of maintenance decay. Crucially, the method operates without manual labeling. Evaluated across multiple enterprise projects, it achieves early warning of high-risk stagnant libraries 3–6 months in advance, substantially reducing manual auditing effort and significantly enhancing the security and maintainability of open-source dependency ecosystems.
为解决ERP系统中数据集成和流程监控的碎片化问题,本文提出一种企业流程控制塔,通过集成状态观测、语义翻译、机器学习诊断等方法提升IT团队的工作效率。
This study addresses the challenge of automating workflows in complex industries—such as logistics, healthcare, and construction—where processes are fragmented across heterogeneous tools and involve multi-party collaboration. The work proposes orchestration as a core abstraction to enable effective automation by dynamically coordinating multi-step tasks, enforcing domain-specific constraints, managing human approvals, and integrating legacy systems. It introduces the novel concept of “orchestration bottlenecks” and develops a theoretical framework that unifies multi-agent systems, workflow modeling, constraint reasoning, and human–AI collaboration, while exposing critical gaps in current multi-agent approaches at the orchestration level. Based on distinct sources of operational friction across domains, the paper advocates for targeted architectural safeguards—such as constraint enforcement or explainability—and phased implementation strategies to provide actionable pathways for automation in complex operational environments.
研究通过分析GitHub Agentic Workflows的结构和维护方式,探讨了开发者如何定义和维护由AI代理执行的工作流程,并建议增加防御措施。
Enterprise operational workflows are notoriously difficult to automate end-to-end due to their heavy reliance on human intervention and limited adaptability to change. This work proposes the first action-centric workflow graph framework, which achieves automated construction, execution, and evolution through a three-stage pipeline: structured workflow graphs are extracted from human operation traces, executed via multi-agent online traversal, and continuously optimized in a closed loop using an Adaptive Traversal Reinforcement (ATR) mechanism. Integrating large-scale offline graph construction, graph-guided retrieval, and large language model reasoning, the approach was deployed across four cloud database services. It substantially outperforms the Trace-RAG baseline in coverage breadth, factual accuracy, and diagnostic throughput, achieving an expert blind-review score of 4.95 out of 5.
This work addresses the limitations of traditional expert-manual-based cybersecurity response methods, which struggle to adapt to dynamic attack scenarios and evolving recovery objectives, as well as the instability of existing large-model approaches in long-horizon tasks. The authors propose an end-to-end agent planning framework that innovatively models event states using a graph structure (Graph-as-State), incorporates a phase-aware agent routing mechanism, and establishes a verifiable experience reuse loop to guide action selection and state updates. The system integrates multi-agent large language models with experience retrieval augmentation and execution feedback verification, enabling dynamic, stable, and evolvable response planning within a Docker-based network range simulation environment. Experimental results demonstrate that the proposed method achieves a normalized defense score of 0.94 across 100 simulated scenarios, representing a 9.5% improvement over the strongest baseline.