Score
Designs, implements, and evaluates methods and tools to detect, analyze, and mitigate nondeterministic or intermittently failing software tests; work includes automated detection (e.g., rerun-based and statistical heuristics), root-cause analysis of flakiness (timing, concurrency, environment or resource dependencies), and interventions to reduce recurrence. Also involves building fixes and processes—such as test isolation, deterministic seeding, environment stabilization, CI configuration, and monitoring/metrics—to prevent, track, and regress-test flakiness in test pipelines.
To address result instability in automated testing caused by inter-test ordering dependencies, this paper proposes a fragility-aware test prioritization method based on shared static field characteristics. The approach employs static code analysis to identify shared static state within test classes that is prone to induce ordering dependencies, then constructs a lightweight fragility scoring model to prioritize high-risk tests—thereby accelerating defect exposure. This work is the first to use shared static fields as the core heuristic feature for ordering-dependent test prioritization. Evaluated across 27 modules, it achieves complete and accurate prioritized coverage of all ordering-dependent tests in 23 modules. On average, it reduces total test execution count by 65.92% and eliminates 72.19% of redundant re-executions, significantly improving detection efficiency while lowering computational resource overhead.
In software testing, flaky tests—those exhibiting non-deterministic pass/fail behavior—often fail in clusters due to shared root causes (e.g., network instability, transient external dependencies), forming “systemic flakiness,” which substantially increases debugging and repair costs. This phenomenon challenges the conventional assumption that flakiness is isolated and test-specific. Method: We introduce the concept of systemic flakiness and empirically analyze 24 Java open-source projects, revealing that 75% of flaky tests belong to co-occurring clusters (average size: 13.5 tests per cluster). To efficiently identify such clusters, we propose a lightweight ML model leveraging static test-distance features, combined with bottom-up clustering and manual stack-trace validation. Contribution/Results: Evaluated on 810 flaky tests, our approach achieves high detection accuracy and enables elimination of thousands of redundant test executions. This work establishes a root-cause–driven paradigm for batch diagnosis and remediation of flakiness, shifting focus from individual test fixes to systemic root-cause resolution.
Flaky tests—non-deterministic in pass/fail outcomes—pose a critical reliability threat to quantum software, yet their prevalence, characteristics, and detectability lack empirical grounding. To address this gap, we conduct the first large-scale dynamic empirical study, executing 27,026 test cases across 23 versions of Qiskit Terra for 10,000 repetitions each, identifying 290 flaky tests. We propose a Wilson confidence interval–based method to quantify rerun budgets for flakiness detection, revealing extremely low occurrence probabilities (~10⁻⁴), high sporadicity, and severe distribution skew. Our work introduces the first cross-version volatility tracking and subsystem mapping analysis for quantum flaky tests, and releases the first publicly available quantum flaky test dataset. Results show that although flakiness rates are low (0–0.4%), detecting them requires tens of thousands of executions—highlighting a severe challenge to quantum testing stability.
This study addresses the prevalent issue of non-deterministic test failures in JavaScript projects caused by environmental variations—such as operating systems, Node.js versions, or browsers—which significantly undermine continuous integration (CI) efficiency. The authors present the first systematic quantification of test flakiness induced by three categories of environmental factors and introduce js-env-sanitizer, a lightweight tool that automatically identifies and skips environment-sensitive tests without requiring re-execution. By leveraging cross-environment execution and dynamic interception, the tool generates detailed reports while maintaining compatibility with popular testing frameworks including Jest, Mocha, and Vitest. Evaluated on 65 real-world projects, js-env-sanitizer demonstrates high accuracy in detecting fragile tests, substantially reducing CI pipeline interruptions with minimal performance and configuration overhead.
This study addresses the lack of effective methods for identifying and diagnosing flaky tests in industrial-scale software systems. It presents the first validation and enhancement of lexical-based flaky test detection within the large-scale industrial environment of SAP HANA. By integrating TF-IDF and TF-IDFC-RF feature extraction with CodeBERT and XGBoost classification models, the approach achieves F1 scores of 0.96 and 0.99 on the original benchmark and SAP HANA datasets, respectively. The research identifies external dependencies as the primary root cause of flakiness in SAP HANA and systematically evaluates the effectiveness of various features and models in real-world industrial settings. These findings provide a robust empirical foundation for the automated diagnosis of flaky tests, offering practical insights for improving test reliability in complex software systems.
Existing code-level test flakiness detection methods are limited by their neglect of execution environments and dynamic behaviors. This work constructs a controlled dataset, C-IDoFT, alongside a CI log-mined dataset, systematically uncovering label shortcuts and evaluation protocol biases in current benchmarks. It proposes a novel execution-environment-centered judgment paradigm that explicitly distinguishes between “whether a test is flaky” and “whether a failure is caused by flakiness.” Through replication of the CodeBERT detector and comprehensive validation—including cross-project evaluation, repeated execution, CI log analysis, and causal attribution—the study reveals that existing models degrade to baseline performance under project-separated protocols despite strong results on FlakeBench. Furthermore, CI logs indicate that 58% of flaky tests require execution evidence rather than static code features for accurate identification.
This study addresses the critical yet underexplored issue of flakiness in REST API testing, which severely undermines the reliability of automated test suites. Through an empirical analysis of nearly 3,000 failing test cases across 36 real-world APIs, the work systematically identifies and categorizes the root causes of flakiness in REST API fuzzing. Building on these insights, the authors propose FlakyCatch, a general-purpose approach capable of effectively detecting and mitigating flakiness in tests generated by both white-box and black-box fuzzers, including EvoMaster. Experimental evaluation demonstrates that FlakyCatch substantially enhances test stability and reliability, establishing a new paradigm for robust REST API testing.
Quantum software is prone to “flaky tests” due to its probabilistic outputs, leading to unstable test results that hinder defect diagnosis and development efficiency. This work proposes the first automated pipeline integrating large language models (LLMs) with cosine similarity to detect flaky tests in quantum repositories, link them to associated pull requests, and support root cause analysis. We present the first systematic evaluation of mainstream LLMs—including GPT, LLaMA, Gemini, and Claude—for this task, expanding the existing dataset by 54% through the identification of 25 previously unknown flaky tests. Experimental results demonstrate that Gemini achieves the best performance, attaining F1 scores of 0.9420 for flakiness detection and 0.9643 for root cause identification, thereby validating the practical potential of LLMs in quantum software testing.
Reproducing and repairing flaky tests remains highly challenging due to the absence of reproducible environments and standardized validation mechanisms. This work introduces ReproFlake, a dataset comprising 1,115 reproducible flaky tests spanning four canonical categories, and presents the first comprehensive ecosystem for flaky test reproducibility, including standardized build environments, automated reproduction and repair validation scripts, detailed execution logs, and community contribution guidelines. By integrating developer-reported cases with existing datasets, the project achieves end-to-end automation in test reproduction, repair application, and log collection. Empirical analysis reveals that error messages aid in identifying flaky test categories, repair locations strongly correlate with test types, and build failures in legacy projects constitute a primary obstacle—collectively establishing a solid empirical foundation for future research.
This study addresses the challenge of reproducing failures in cyber-physical system (CPS) simulation testing, where non-deterministic behaviors often hinder consistent fault manifestation. To tackle this issue, the work extends delta debugging to stochastic CPS scenarios for the first time, introducing three novel delta debugging algorithms tailored for randomized environments. The proposed approach integrates statistical failure analysis, repeated execution, and environment-aware input minimization to identify a minimal triggering input while preserving fault semantics. Empirical evaluation on case studies involving elevator scheduling and autonomous mobile robots demonstrates that the method significantly enhances the stability of failure reproduction and substantially reduces debugging time, effectively mitigating execution flakiness inherent in such systems.