multi-round execution verification

Design and build execution-based test harnesses and validation procedures that repeatedly run target queries or programs across multiple temporally separated rounds to detect flakiness and confirm result stability. Analyze outputs and execution traces to filter or flag unstable instances and produce validated results or reports before release.

multi-roundexecutionverification

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.39
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing code-level test flakiness detection methods are limited by their neglect of execution environments and dynamic behaviors. This work constructs a controlled dataset, C-IDoFT, alongside a CI log-mined dataset, systematically uncovering label shortcuts and evaluation protocol biases in current benchmarks. It proposes a novel execution-environment-centered judgment paradigm that explicitly distinguishes between “whether a test is flaky” and “whether a failure is caused by flakiness.” Through replication of the CodeBERT detector and comprehensive validation—including cross-project evaluation, repeated execution, CI log analysis, and causal attribution—the study reveals that existing models degrade to baseline performance under project-separated protocols despite strong results on FlakeBench. Furthermore, CI logs indicate that 58% of flaky tests require execution evidence rather than static code features for accurate identification.

CI pipelinescode-based detectionevaluation bias

This study addresses the reliability challenges posed by flakiness in test cases generated by large language models (LLMs), which can compromise defect detection. For the first time, it systematically quantifies the prevalence and root causes of flakiness in tests produced by GPT-4o and Mistral-Large-Instruct-2407 across four database systems: SAP HANA, DuckDB, MySQL, and SQLite. Through manual analysis and cross-system comparative experiments, the authors find that LLM-generated tests exhibit slightly higher flakiness than existing human-written tests. Notably, 63% of flaky cases stem from implicit assumptions about unordered collections—termed “non-guaranteed order dependencies.” Furthermore, flakiness can propagate from existing tests to newly generated ones via contextual prompts, a phenomenon particularly pronounced in closed-source systems.

database management systemsflaky testsLLM-generated tests

This study addresses the challenge of flaky tests undermining code quality assessment in the industrial-scale database system SAP HANA. To enable systematic root cause analysis, the authors introduce a large language model (LLM) as an automated annotator—deployed for the first time in a large database system—and combine it with internal and external consistency validation strategies to classify the root causes of 559 resolved flaky test reports. Their analysis reveals that 23% (130 cases) of flakiness stems from concurrency-related issues and uncovers distinct instability challenges across different test types. The proposed approach establishes a scalable new paradigm for empirical analysis and mitigation of flaky tests, offering actionable guidance for improving reliability in complex software systems.

database management systemflaky testsroot cause analysis

This study addresses the lack of effective methods for identifying and diagnosing flaky tests in industrial-scale software systems. It presents the first validation and enhancement of lexical-based flaky test detection within the large-scale industrial environment of SAP HANA. By integrating TF-IDF and TF-IDFC-RF feature extraction with CodeBERT and XGBoost classification models, the approach achieves F1 scores of 0.96 and 0.99 on the original benchmark and SAP HANA datasets, respectively. The research identifies external dependencies as the primary root cause of flakiness in SAP HANA and systematically evaluates the effectiveness of various features and models in real-world industrial settings. These findings provide a robust empirical foundation for the automated diagnosis of flaky tests, offering practical insights for improving test reliability in complex software systems.

automated testingflaky testsindustrial software

Detecting Flaky Tests in Quantum Software: A Dynamic Approach

Dec 19, 2025
DK
Dongchan Kim
🏛️ University of Maryland, Baltimore County | Toronto Metropolitan University

Flaky tests—non-deterministic in pass/fail outcomes—pose a critical reliability threat to quantum software, yet their prevalence, characteristics, and detectability lack empirical grounding. To address this gap, we conduct the first large-scale dynamic empirical study, executing 27,026 test cases across 23 versions of Qiskit Terra for 10,000 repetitions each, identifying 290 flaky tests. We propose a Wilson confidence interval–based method to quantify rerun budgets for flakiness detection, revealing extremely low occurrence probabilities (~10⁻⁴), high sporadicity, and severe distribution skew. Our work introduces the first cross-version volatility tracking and subsystem mapping analysis for quantum flaky tests, and releases the first publicly available quantum flaky test dataset. Results show that although flakiness rates are low (0–0.4%), detecting them requires tens of thousands of executions—highlighting a severe challenge to quantum testing stability.

Characterizes prevalence and patterns of flakiness across Qiskit releasesDetects flaky tests in quantum software using dynamic analysisEstimates execution budgets needed for reliable flaky test detection

Latest Papers

What's happening recently
View more

Quantum software is prone to “flaky tests” due to its probabilistic outputs, leading to unstable test results that hinder defect diagnosis and development efficiency. This work proposes the first automated pipeline integrating large language models (LLMs) with cosine similarity to detect flaky tests in quantum repositories, link them to associated pull requests, and support root cause analysis. We present the first systematic evaluation of mainstream LLMs—including GPT, LLaMA, Gemini, and Claude—for this task, expanding the existing dataset by 54% through the identification of 25 previously unknown flaky tests. Experimental results demonstrate that Gemini achieves the best performance, attaining F1 scores of 0.9420 for flakiness detection and 0.9643 for root cause identification, thereby validating the practical potential of LLMs in quantum software testing.

automated testingflaky testsquantum software

In enterprise microservice regression testing, QA engineers often lack up-to-date documentation and must rely on real user traffic to reconstruct business scenarios; however, transforming such traffic into replayable test cases with stable assertions is labor-intensive and error-prone. This work proposes NL2Test, a method that combines the semantic understanding of large language models with deterministic algorithms to generate executable API test cases end-to-end from natural language scenario descriptions and captured execution traces. NL2Test automatically slices request sequences, reconstructs data dependencies, masks non-deterministic fields, and produces business-aligned, reliable assertions. Evaluated on 51 industrial scenarios, it achieves an exact match rate of 82.4%, with 98.0% of generated test cases becoming functional after minor tuning. During a nine-month production deployment, it produced 3,196 test cases, of which 85.4% were accepted and integrated into the codebase.

assertion generationmicroservice systemsregression testing

Test flakiness severely undermines developer trust, CI reliability, and resource efficiency in large-scale software ecosystems, yet its cross-project manifestations remain underexplored. This study presents the first empirical analysis of test behavior across 649 projects in the OpenStack ecosystem, leveraging extensive CI logs, execution records, and a mixed-methods approach combining quantitative statistics with qualitative root-cause analysis. The work systematically identifies and quantifies two novel forms of flakiness: cross-project and inconsistent flaky tests. Findings reveal 1,535 cross-project and 1,105 inconsistent flaky tests, affecting 55% of the projects. Notably, 70% of unit tests are impacted by cross-project flakiness, significantly prolonging code review cycles and increasing resource consumption—challenging the long-held assumption that unit tests are inherently isolated.

Continuous Integrationcross-project flakinessinconsistent flakiness

This study addresses the critical yet underexplored issue of flakiness in REST API testing, which severely undermines the reliability of automated test suites. Through an empirical analysis of nearly 3,000 failing test cases across 36 real-world APIs, the work systematically identifies and categorizes the root causes of flakiness in REST API fuzzing. Building on these insights, the authors propose FlakyCatch, a general-purpose approach capable of effectively detecting and mitigating flakiness in tests generated by both white-box and black-box fuzzers, including EvoMaster. Experimental evaluation demonstrates that FlakyCatch substantially enhances test stability and reliability, establishing a new paradigm for robust REST API testing.

automated testingfuzzingREST API

Hot Scholars

CT

Chongyang Tao

Associate Professor of Computer Science, Beihang University
Natural Language ProcessingDialogue SystemsInformation RetrievalData Intelligence
WZ

Wenhui Zhu

Arizona State University
Computer VisionArtificial intelligenceVision Language ModelLarge Language Model
JS

Jinyan Su

Cornell university
LLM reasoningLLM agentretrieval augmented generation
VS

Veselin Stoyanov

Tome AI
Natural Language ProcessingMachine LearningStructured PredictionInformation Extraction
SL

Salem Lahlou

MBZUAI
probabilistic modelinguncertainty estimationgflownetsLLM reasoning