llm fault classification

Designs and evaluates classification pipelines that use large language models (often kept frozen) with in-context, zero‑ or few‑shot prompting to map textual or structured failure descriptions to fault categories. These systems produce a predicted class (and alternative candidates), quantitative confidence scores, and evidence-based explanations supporting each prediction.

llmfaultclassification

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.11
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the long-standing reliance on manual intervention in CI/CD pipeline failure management, which suffers from low efficiency and limited automation capabilities, particularly when handling unstructured failure information. To overcome these challenges, the authors propose an automated approach that integrates large language models (LLMs) with domain-specific knowledge—including historical failure data, pipeline context, and repair instructions—enabling end-to-end failure localization and repair generation for the first time in the large-scale industrial project SAP HANA. Experimental results demonstrate that, when augmented with domain knowledge, the method achieves a 97.4% accuracy in failure localization and generates fully correct repair solutions in 92.1% of cases, substantially outperforming knowledge-agnostic baselines. Ablation studies further confirm the critical contribution of historical failure data to overall performance gains.

automationCI/CD pipelinefailure management

A System-Level Taxonomy of Failure Modes in Large Language Model Applications

Nov 25, 2025
VV
Vaishali Vinay
🏛️ Independent Researcher

This paper addresses the under-recognized reliability challenges of large language models (LLMs) in real-world system deployments. Adopting a systems engineering perspective, it establishes the first fault taxonomy for LLM-based applications. Through systematic analysis and multi-case root-cause investigation, the study identifies 15 classes of latent failures—including multi-step reasoning drift, context boundary degradation, erroneous tool invocation, and latent inconsistency—exposing fundamental limitations of current evaluation benchmarks in stability, reproducibility, and workflow integration. The work introduces high-level design principles centered on observability, cost sensitivity, and version evolution, shifting LLM reliability research from a model-centric to a system-integration paradigm. It delivers the first structured fault classification framework and practical guidance for building reliable, maintainable, and auditable LLM software systems.

Addressing system-level reliability challenges in LLM deploymentAnalyzing gaps in evaluation methods for stability and reproducibilityClassifying fifteen hidden failure modes in real-world LLM applications

In text classification, manual verification of predictions is costly and ill-suited for continuous retraining under data drift. This work pioneers a systematic investigation into leveraging large language models (LLMs) as trustworthy automated validators—replacing human annotation to ensure classifier quality and enable efficient incremental updates. Our method integrates prompt engineering, zero- and few-shot inference, consistency checking, task-specific semantic constraints, and model confidence analysis. Evaluated across multiple benchmark datasets, LLM-based validation achieves over 92% agreement with expert annotations, substantially reducing verification cost while improving pipeline timeliness and scalability. The core contribution is the first LLM-based trustworthy validation framework specifically designed for classifier prediction verification—establishing a new paradigm for low-cost, robust continual learning.

Addressing high costs and limited availability of human annotatorsAutomating text classifier validation using LLMs to reduce human effortEnsuring model quality and enabling efficient incremental learning

What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Jun 18, 2024
FE
Federico Errica
🏛️ NEC Italia | NEC Laboratories Europe

Large language models (LLMs) exhibit unstable outputs in software applications when prompts undergo minor rephrasings, hindering reliable deployment. Method: This paper introduces two label-free, quantifiable metrics—sensitivity (cross-prompt prediction variance) and consistency (prediction stability across semantically equivalent prompts)—to formally decouple and evaluate LLM robustness to prompt perturbations. Leveraging text classification tasks, we conduct systematic, multi-round prompt rewriting and statistical analysis of prediction distributions. Contribution/Results: Empirical evaluation reveals that mainstream LLMs consistently exhibit high sensitivity and low consistency, exposing a critical robustness gap. Our framework provides a reproducible, ground-truth-label-free diagnostic paradigm for prompt engineering, enabling joint optimization of accuracy and robustness. This work establishes the first formal, measurement-driven approach to assessing and improving LLM resilience against prompt variations.

Large Language ModelsPrediction InstabilitySemantic Sensitivity

Prompt Stability Scoring for Text Annotation with Large Language Models

Jul 02, 2024
CB
Christopher Barrie
🏛️ University of Edinburgh | Independent Researcher | University of Amsterdam

Large language models (LLMs) exhibit poor reproducibility in text annotation tasks due to sensitivity to minor prompt perturbations, yet no standardized metric exists for quantifying prompt stability. To address this, we systematically adapt inter-annotator agreement principles from coding reliability research to prompt engineering, introducing the Prompt Stability Score (PSS)—a unified, computationally tractable metric for stability assessment. Our method integrates multi-prompt sampling, batched LLM inference, consistency analysis via Cohen’s and Fleiss’ Kappa, and an automated Python evaluation framework (open-sourced as PromptStability). Empirical validation across six benchmark datasets and twelve annotation task types—encompassing over 150,000 samples—demonstrates PSS’s effectiveness in precisely identifying low-stability prompting configurations. This work establishes the first standardized diagnostic paradigm for evaluating prompt robustness, thereby enabling reproducible, interpretable, and empirically grounded prompt engineering practices.

Addresses reproducibility issues in text annotationMeasures prompt stability in large language modelsProvides framework for reliable classification routines

Latest Papers

What's happening recently
View more

This study addresses the limitations of existing approaches for automatically extracting machine learning (ML) pipeline structures, which often rely on manual annotations or suffer from insufficient generalization to keep pace with the rapid evolution of the ML ecosystem. The work presents the first systematic evaluation of small language models (SLMs) for reverse-engineering ML pipelines and proposes an SLM-based method for their automatic identification and reconstruction. Through comprehensive comparative experiments across multiple SLMs and rigorous statistical validation using Cochran’s Q, McNemar, and Pearson’s chi-squared tests, the authors demonstrate that the best-performing SLM significantly outperforms current methods and exhibits robustness across diverse classification schemes. This approach uncovers finer-grained patterns in data science practices and overcomes longstanding bottlenecks in scalability and domain adaptability inherent in traditional techniques.

Code UnderstandingData Science PracticesMachine Learning Pipelines

Current evaluations of large language models (LLMs) on ill-defined tasks—such as complex instruction following and natural language-to-Mermaid sequence diagram generation—suffer from insufficient coverage, sensitivity to phrasing, incomparable metrics, and instability in LLM-based judging, thereby failing to yield reliable or diagnostic assessment signals. This work presents the first systematic analysis of confounding failure modes in such tasks, integrating case studies, failure mode categorization, and a multidimensional evaluation framework to demonstrate how existing benchmarks often conflate distinct error types, leading to distorted scores. Moving beyond monolithic aggregate metrics, the proposed approach delivers actionable, fine-grained insights that lay both theoretical and practical foundations for building more robust and interpretable evaluation systems.

diagnostic evaluationevaluation benchmarksill-defined tasks

Intermittent failures in continuous integration (CI) pipelines are notoriously difficult to diagnose, leading to wasted resources and reduced development efficiency. This work proposes FlaXifyer, a few-shot learning approach that integrates the interpretable AI technique LogSift to fine-tune pretrained language models on pipeline logs using only 12 labeled examples per failure class. The method simultaneously predicts failure categories and pinpoints critical log entries indicative of root causes. Evaluated on 2,458 real-world CI failures, FlaXifyer achieves a Macro F1 score of 84.3% and a Top-2 accuracy of 92.0%, reducing the required log inspection effort by 74.4%. Furthermore, it successfully identifies the underlying fault in 87% of cases, demonstrating its effectiveness in accelerating failure diagnosis with minimal labeled data.

automated triagecontinuous integrationfailure diagnosis

Teaching People LLM's Errors and Getting it Right

Dec 24, 2025
NS
Nathan Stringham
🏛️ University of Utah

Users often over-rely on large language models (LLMs) in simple tasks (e.g., arithmetic) due to their strong performance on complex ones (e.g., poetry generation), misjudging reliability. Existing methods—using embedding clustering to identify LLM failure modes and teach users—show limited effectiveness. Method: We conduct the first empirical validation of the groupability and teachability of systematic LLM failure patterns. Introducing a novel paradigm for instructional efficacy—user accuracy in *anticipating* LLM errors—we replace traditional human-AI collaboration accuracy metrics. Using meta-label grouping, embedding clustering, prompt engineering, and controlled user studies, we evaluate current automated failure detection and instruction approaches. Contribution/Results: We find that state-of-the-art automatic failure discovery lacks stability; critically, our new teaching paradigm significantly improves users’ error anticipation accuracy (p < 0.01), providing both theoretical grounding and practical pathways for reliable human-LLM collaboration.

Investigates why teaching LLM failure patterns fails to reduce user overreliance.Proposes a new metric to assess teaching effectiveness in mitigating overreliance.Tests if automated methods can effectively surface LLM failure patterns for users.

Hot Scholars

JS

Jieke Shi

PhD Candidate & Research Engineer, Singapore Management University
Software EngineeringAI Software Testing
SR

Shawn Rasheed

UCOL | Te Pūkenga
program analysisprogramming languagessecurity
SL

Shuai Liang

College of Environmental Science and Engineering, Beijing Forestry University
Membrane TechnologyNanotechnology
PC

Pengfei Chen

Sun Yat-sen University, Associated Professor
Distributed computingCloud computing and Blockchain