Score
Designs, implements, and evaluates systems or processes that detect and localize operational issues and then close the remediation loop by coordinating fixes, tracking actions, and verifying resolution through feedback and monitoring. Covers end-to-end work across detection, root-cause analysis, remediation orchestration, and validation to ensure incidents are resolved and do not recur.
Existing root cause analysis (RCA) research lacks a goal-oriented, systematic taxonomy, leading to task ambiguity and hindered progress assessment. Method: This paper proposes the first RCA classification framework centered on fundamental objectives—departing from conventional data-type–based taxonomies—and systematically categorizes 135 studies (2014–2025) according to core goals such as fault localization and defect remediation. Guided by a systematic literature review, we construct a multi-level RCA objective hierarchy that characterizes the state of the art, recurrent challenges, and critical technical gaps per task. Contribution/Results: We present the first RCA objective-method mapping atlas tailored to cloud service scenarios, establishing a theoretical foundation for academic research and a practical technology roadmap for industrial deployment.
This study addresses the challenge of effectively monitoring early-stage agent systems, where structural flaws often obscure task-level errors. The authors propose a three-dimensional (quality, suitability, efficiency) and three-granularity (intra-run, inter-run, structural) monitoring and triaging framework tailored for low-maturity agent systems. They introduce a novel system maturity staging model based on the coefficient of variation and monitoring granularity, integrated with a severity classification adapted from FMEA to guide human review. The resulting transferable monitoring architecture supports document-driven, multi-stage workflows, enhanced by a synthetic testbed with controlled error injection. Experimental results demonstrate that structural defects significantly mask task-level signals; 97% of issues can be automatically traced, with only 2% requiring human intervention, and each granularity level precisely identifies its corresponding defect type (coefficients of variation: 0.02, 1.25, and 0.00, respectively).
Prior work lacks empirical characterization of problem-solving processes in software development. Method: Integrating grounded theory coding, sequential pattern mining, and multidimensional statistical analysis on 356 Mozilla Firefox issue reports, this study extracts fine-grained, reusable problem-solving process patterns from collaborative textual artifacts. Contribution/Results: We identify 47 empirically grounded process patterns—challenging the traditional linear assumption by revealing pervasive nonlinearity: 73% of fixes involve iterative backtracking or parallel activities. The resulting process landscape and pattern catalog systematically characterize distributional regularities across issue types, defect categories, and repair durations. This advances understanding of real-world engineering complexity and provides an evidence-based foundation for process optimization, collaborative tool design, and developer support.
Existing visualization tools for compliance checking lack systematic characterization of analytical tasks, hindering rigorous effectiveness evaluation. This paper introduces the first multidimensional task taxonomy specifically designed for compliance checking, modeling core trace-to-model alignment tasks in process mining along six dimensions: objective, method, constraint type, data characteristics, data target, and cardinality. Crucially, this taxonomy explicitly links the semantic requirements of compliance checking with established visual analytics design principles—thereby bridging the semantic gap between process mining and visual analytics. It provides a reusable theoretical framework to rigorously define visualization purposes, evaluate tool effectiveness, and support co-design of analysis systems. As a result, the interpretability and practical utility of complex compliance analysis outcomes are significantly enhanced.
This work addresses the challenge of costly erroneous repairs in existing automated remediation systems, which often lack the ability to assess intervention necessity and thus rely on manual approval for safety. The authors formulate safe repair as an intervention decision problem under risk constraints and introduce a three-dimensional risk decomposition framework encompassing impact scope, reversibility, and epistemic uncertainty. They further design a context-adaptive human-in-the-loop gating strategy that enables interpretable, workload-aware safety interventions. Built upon constrained Markov decision processes (CMDPs), offline policy learning, Chaos Mesh fault injection, and the RCAEval classification framework, the proposed approach reduces erroneous repair rates by 39% and improves repair success rates by 2.5 percentage points on the Train Ticket benchmark, while decreasing on-call escalation burden by 17% compared to fixed-threshold baselines.
This study addresses the lack of systematic comparative analysis in business process compliance monitoring, particularly for non-conformance checking techniques. Through a systematic literature review (SLR), process mining, compliance modeling, and qualitative comparative analysis, it maps real-world applications across domains, operational workflows, technical foundations, and result representations. The analysis identifies key implementation barriers—especially pervasive human dependence and the absence of standardized evaluation criteria. As the first structured survey framework dedicated to non-conformance checking, the study introduces a standardized, multi-dimensional evaluation framework that clarifies commonalities and distinctions across the technical landscape. It further proposes an extensible theoretical pathway and practical guidelines for automated compliance monitoring. This work provides a methodological foundation and strategic direction for both academic research and industrial deployment. (149 words)
研究通过轨迹级分析方法评估LLM代理在微服务根因分析中的表现,提出DiagGuard框架以提高诊断准确性。
This study addresses the challenges in assessing the completeness of multi-patch vulnerability fixes and the lack of systematic understanding of their root causes and characteristics. Through manual analysis of 1,646 multi-patch repair records, this work proposes the first three-tier classification framework grounded in root causes, revealing the evolutionary patterns of such repairs. By contrasting key features, it clarifies the distinctions between multi-patch and single-patch fixes and evaluates the effectiveness of mainstream vulnerability detection tools in verifying repair completeness. The findings delineate predominant multi-patch repair patterns and associated challenges, expose limitations of current tools, and provide a novel perspective along with an empirical foundation for future research on repair validation.
为解决ERP系统中数据集成和流程监控的碎片化问题,本文提出一种企业流程控制塔,通过集成状态观测、语义翻译、机器学习诊断等方法提升IT团队的工作效率。
为解决工具调用超时导致的静默失败问题,引入Outcome Monitors方法检测结果合同违规,并提供恢复工具建议,提高任务完成率。
This work addresses the limitations of existing large language models, which are typically confined to isolated tasks and struggle to integrate into industrial-scale, multi-stage security workflows. To bridge this gap, the authors propose the first role-based multi-agent framework tailored to the entire vulnerability lifecycle, incorporating specialized agents—Planner, Analyzer, Fixer, and Verifier—augmented with CodeQL static analysis for enhanced precision. By introducing a role-oriented multi-agent architecture into end-to-end vulnerability management, this approach effectively aligns the capabilities of large models with real-world security engineering demands. Evaluated on 25 real-world C/C++ vulnerabilities, the system achieves a detection accuracy of 44%—comparable to GPT-5.5—and a repair accuracy of 19%, offering a practical and collaborative paradigm for intelligent security operations.