negative-case analysis

Designs and conducts analyses and tooling to identify, collect, categorize, and monitor negative cases — bad examples, failure modes, or harmful/low‑quality instances — in datasets, model outputs, or operational traces. Builds attribution, triage, and management workflows that assign root causes, prioritize cases for remediation, and produce metrics and alerts to track regressions and the effect of fixes.

negative-caseanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.13
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the lack of empirical guidance on tool design and composition for large language model (LLM) agents in microservice root cause analysis (RCA) by constructing the first systematic empirical benchmark dedicated to agentic RCA tool abstraction and composition. We propose a hierarchical tool architecture spanning levels L0 through L3 and conduct multi-model comparative experiments alongside trajectory analysis to quantitatively evaluate how different tool configurations affect diagnostic performance. Results demonstrate that higher-level tools (L3) halve fault localization time while improving fault type identification, revealing inherent accuracy-efficiency trade-offs across tool hierarchy levels. These findings provide data-driven decision-making foundations for agent tool selection and design in automated microservice diagnostics.

Empirical StudyLLM AgentsMicroservice Systems

DrP: Meta's Efficient Investigations Platform at Scale

Dec 03, 2025
SS
Shubham Somani
🏛️ Meta

In large-scale systems, on-call engineers rely on manual procedures or ad-hoc scripts for incident investigation, resulting in high mean time to resolution (MTTR), elevated operational overhead, and diminished productivity. This paper introduces DrP—the first end-to-end automated investigation framework designed for heterogeneous domains including services, AI/ML, and mobile systems. DrP’s key contributions are: (1) a declarative SDK enabling low-code development of reusable, domain-agnostic analysis logic; (2) a distributed execution engine with a plugin-based architecture supporting high-concurrency diagnostics and deep integration with alerting, event management, and remediation systems; and (3) a unified abstraction layer that transparently insulates users from infrastructure heterogeneity. Deployed at scale within Meta, DrP executes ~50,000 analyses daily across 300+ engineering teams, reducing average MTTR by 20% overall and up to 80% in specific scenarios—significantly enhancing SRE responsiveness and system observability.

Automates manual investigation processes to reduce incident resolution timeProvides an end-to-end framework for scalable, automated incident analysis and mitigationReduces on-call toil and improves productivity in large-scale systems

The root cause analysis (RCA) community suffers from a critical shortage of large-scale, open-source, multimodal benchmark datasets, hindering rigorous method evaluation and advancement. To address this, we introduce LEMMA-RCA—the first cross-domain, multimodal, open-source RCA dataset tailored for IT/OT systems, encompassing realistic failure scenarios from microservices, water supply, and wastewater treatment. It comprises hundreds of system entities and fine-grained causal relationship annotations. Uniquely integrating four heterogeneous modalities—time-series metrics, logs, topology graphs, and alerts—it supports both offline/online and unimodal/multimodal RCA evaluation. Built via distributed monitoring, multi-source alignment, controllable fault injection, and causal graph annotation, the dataset ensures high fidelity and representativeness. Extensive evaluation across eight baseline methods demonstrates that multimodal joint modeling improves average F1-score by 23.6%. LEMMA-RCA is publicly released, establishing a new community benchmark for RCA research.

Absence of real-world fault scenarios from IT and OT systemsLack of large-scale open-source datasets for root cause analysisNeed for diverse RCA tasks across multiple domains and modalities

What-if Analysis for Business Professionals: Current Practices and Future Opportunities

Dec 27, 2022
SG
Sneha Gathani
🏛️ University of Maryland | University of Massachusetts | AWS AI Labs | MIT CSAIL

Business professionals—non-technical domain experts—lack appropriate tools and methodologies for effective what-if analysis (WIA), hindering data-informed decision-making. Method: We conducted a two-phase mixed-methods user study—comprising contextual interviews and in-situ task-based evaluations—to systematically characterize their analytical behaviors for the first time. Contribution/Results: Based on empirical findings, we propose three domain-grounded design principles: business-contextual data preparation, risk-aware assessment, and domain-knowledge integration. We implemented and validated these principles in an interactive visual analytics prototype. The study identifies three critical support gaps, empirically confirms that six classes of what-if techniques significantly improve decision efficiency and confidence, and yields eight actionable design guidelines for commercial business intelligence systems. This work bridges a key theoretical and practical gap in WIA research concerning non-technical users.

Addresses lack of WIA support for business professionalsExplores non-technical WIA practices and challengesProposes design improvements for business analytics systems

Latest Papers

What's happening recently
View more

This work addresses the vulnerability of existing LLM-driven microservice root cause analysis (RCA) agents to early reasoning errors that lead to diagnostic failure and their inability to localize or correct such mistakes. The study introduces a novel formulation of RCA errors as stage-localizable reasoning flaws and proposes a structured four-stage framework—comprising evidence bundles, hypothesis sets, analytical structures, and decision reports. To enhance robustness, it integrates stage-level auditing, budget-aware fast-slow path routing, counterfactual candidate evaluation, and stage-specific patch replay mechanisms. Implemented via LangGraph, the system demonstrates significant improvements in root cause localization and fault classification accuracy on both public benchmarks and real-world production data, precisely identifying erroneous stages and successfully repairing most execution trajectories within one to two iterations, thereby substantially improving agent debuggability and self-repair capability.

AIOpsLLM-based AgentsMicroservices

This study addresses the lack of standardized and automated case planning processes in medical social work, which currently relies heavily on individual practitioner experience and suffers from inefficiency. The authors propose a model-agnostic, open-source large language model (LLM) workflow that systematically integrates established social work practice frameworks into LLM prompt design for the first time. The approach decomposes case planning into six sequential stages—assessment, problem analysis, goal setting, intervention planning, risk anticipation, and outcome evaluation—and combines structured client profiling with staged prompt engineering to generate professional, reviewable draft assessment forms and service plans. Designed to be compatible across multiple LLM platforms, the framework ensures cross-model reproducibility, and its code has been publicly released to provide a standardized tool for advancing intelligent support in medical social work.

case planningLLM workflowmedical social work

Existing root cause analysis (RCA) research is constrained by Top@k metrics, making it difficult to distinguish between failures originating in the retrieval and reranking stages. This work proposes a retrieval-reranking decoupled framework that constructs a two-stage pipeline comprising a multi-signal fusion retriever and a large language model-based reranker. The proposed approach achieves high-precision root cause localization without requiring causal graphs or annotated data. Experimental results demonstrate that this method comprehensively outperforms the strongest baselines across six benchmarks, improving Top@1 accuracy by up to 18 percentage points.

Anomaly DetectionComplex Monitored SystemsEvaluation Metric

Hot Scholars

SK

Sinan Kalkan

Dept. of Computer Eng., Middle East Technical University
Computer VisionDeep LearningRobotics
LB

Luca Barletta

Politecnico di Milano
Communication TheoryInformation Theory