Score
Designs and evaluates classification pipelines that use large language models (often kept frozen) with in-context, zero‑ or few‑shot prompting to map textual or structured failure descriptions to fault categories. These systems produce a predicted class (and alternative candidates), quantitative confidence scores, and evidence-based explanations supporting each prediction.
This work addresses the long-standing reliance on manual intervention in CI/CD pipeline failure management, which suffers from low efficiency and limited automation capabilities, particularly when handling unstructured failure information. To overcome these challenges, the authors propose an automated approach that integrates large language models (LLMs) with domain-specific knowledge—including historical failure data, pipeline context, and repair instructions—enabling end-to-end failure localization and repair generation for the first time in the large-scale industrial project SAP HANA. Experimental results demonstrate that, when augmented with domain knowledge, the method achieves a 97.4% accuracy in failure localization and generates fully correct repair solutions in 92.1% of cases, substantially outperforming knowledge-agnostic baselines. Ablation studies further confirm the critical contribution of historical failure data to overall performance gains.
This paper addresses the under-recognized reliability challenges of large language models (LLMs) in real-world system deployments. Adopting a systems engineering perspective, it establishes the first fault taxonomy for LLM-based applications. Through systematic analysis and multi-case root-cause investigation, the study identifies 15 classes of latent failures—including multi-step reasoning drift, context boundary degradation, erroneous tool invocation, and latent inconsistency—exposing fundamental limitations of current evaluation benchmarks in stability, reproducibility, and workflow integration. The work introduces high-level design principles centered on observability, cost sensitivity, and version evolution, shifting LLM reliability research from a model-centric to a system-integration paradigm. It delivers the first structured fault classification framework and practical guidance for building reliable, maintainable, and auditable LLM software systems.
In text classification, manual verification of predictions is costly and ill-suited for continuous retraining under data drift. This work pioneers a systematic investigation into leveraging large language models (LLMs) as trustworthy automated validators—replacing human annotation to ensure classifier quality and enable efficient incremental updates. Our method integrates prompt engineering, zero- and few-shot inference, consistency checking, task-specific semantic constraints, and model confidence analysis. Evaluated across multiple benchmark datasets, LLM-based validation achieves over 92% agreement with expert annotations, substantially reducing verification cost while improving pipeline timeliness and scalability. The core contribution is the first LLM-based trustworthy validation framework specifically designed for classifier prediction verification—establishing a new paradigm for low-cost, robust continual learning.
Large language models (LLMs) exhibit unstable outputs in software applications when prompts undergo minor rephrasings, hindering reliable deployment. Method: This paper introduces two label-free, quantifiable metrics—sensitivity (cross-prompt prediction variance) and consistency (prediction stability across semantically equivalent prompts)—to formally decouple and evaluate LLM robustness to prompt perturbations. Leveraging text classification tasks, we conduct systematic, multi-round prompt rewriting and statistical analysis of prediction distributions. Contribution/Results: Empirical evaluation reveals that mainstream LLMs consistently exhibit high sensitivity and low consistency, exposing a critical robustness gap. Our framework provides a reproducible, ground-truth-label-free diagnostic paradigm for prompt engineering, enabling joint optimization of accuracy and robustness. This work establishes the first formal, measurement-driven approach to assessing and improving LLM resilience against prompt variations.
Large language models (LLMs) exhibit poor reproducibility in text annotation tasks due to sensitivity to minor prompt perturbations, yet no standardized metric exists for quantifying prompt stability. To address this, we systematically adapt inter-annotator agreement principles from coding reliability research to prompt engineering, introducing the Prompt Stability Score (PSS)—a unified, computationally tractable metric for stability assessment. Our method integrates multi-prompt sampling, batched LLM inference, consistency analysis via Cohen’s and Fleiss’ Kappa, and an automated Python evaluation framework (open-sourced as PromptStability). Empirical validation across six benchmark datasets and twelve annotation task types—encompassing over 150,000 samples—demonstrates PSS’s effectiveness in precisely identifying low-stability prompting configurations. This work establishes the first standardized diagnostic paradigm for evaluating prompt robustness, thereby enabling reproducible, interpretable, and empirically grounded prompt engineering practices.
This study addresses the limitations of existing approaches for automatically extracting machine learning (ML) pipeline structures, which often rely on manual annotations or suffer from insufficient generalization to keep pace with the rapid evolution of the ML ecosystem. The work presents the first systematic evaluation of small language models (SLMs) for reverse-engineering ML pipelines and proposes an SLM-based method for their automatic identification and reconstruction. Through comprehensive comparative experiments across multiple SLMs and rigorous statistical validation using Cochran’s Q, McNemar, and Pearson’s chi-squared tests, the authors demonstrate that the best-performing SLM significantly outperforms current methods and exhibits robustness across diverse classification schemes. This approach uncovers finer-grained patterns in data science practices and overcomes longstanding bottlenecks in scalability and domain adaptability inherent in traditional techniques.
Current evaluations of large language models (LLMs) on ill-defined tasks—such as complex instruction following and natural language-to-Mermaid sequence diagram generation—suffer from insufficient coverage, sensitivity to phrasing, incomparable metrics, and instability in LLM-based judging, thereby failing to yield reliable or diagnostic assessment signals. This work presents the first systematic analysis of confounding failure modes in such tasks, integrating case studies, failure mode categorization, and a multidimensional evaluation framework to demonstrate how existing benchmarks often conflate distinct error types, leading to distorted scores. Moving beyond monolithic aggregate metrics, the proposed approach delivers actionable, fine-grained insights that lay both theoretical and practical foundations for building more robust and interpretable evaluation systems.
Intermittent failures in continuous integration (CI) pipelines are notoriously difficult to diagnose, leading to wasted resources and reduced development efficiency. This work proposes FlaXifyer, a few-shot learning approach that integrates the interpretable AI technique LogSift to fine-tune pretrained language models on pipeline logs using only 12 labeled examples per failure class. The method simultaneously predicts failure categories and pinpoints critical log entries indicative of root causes. Evaluated on 2,458 real-world CI failures, FlaXifyer achieves a Macro F1 score of 84.3% and a Top-2 accuracy of 92.0%, reducing the required log inspection effort by 74.4%. Furthermore, it successfully identifies the underlying fault in 87% of cases, demonstrating its effectiveness in accelerating failure diagnosis with minimal labeled data.
Users often over-rely on large language models (LLMs) in simple tasks (e.g., arithmetic) due to their strong performance on complex ones (e.g., poetry generation), misjudging reliability. Existing methods—using embedding clustering to identify LLM failure modes and teach users—show limited effectiveness. Method: We conduct the first empirical validation of the groupability and teachability of systematic LLM failure patterns. Introducing a novel paradigm for instructional efficacy—user accuracy in *anticipating* LLM errors—we replace traditional human-AI collaboration accuracy metrics. Using meta-label grouping, embedding clustering, prompt engineering, and controlled user studies, we evaluate current automated failure detection and instruction approaches. Contribution/Results: We find that state-of-the-art automatic failure discovery lacks stability; critically, our new teaching paradigm significantly improves users’ error anticipation accuracy (p < 0.01), providing both theoretical grounding and practical pathways for reliable human-LLM collaboration.