Score
Designs, implements, and executes tests and test suites (manual and automated) to verify system correctness, behavior, performance, and reliability. Builds test plans, test cases, test harnesses and CI test pipelines, and analyzes test outcomes, coverage, and failure causes to assess and improve software quality.
Real-time testing in cloud service production environments risks interfering with live traffic, compromising service reliability and SLA compliance. Method: This paper proposes an automated test planning framework that jointly models test configuration selection, deployment planning, and execution scheduling as a risk-constrained optimization problem. It integrates lightweight online traffic forecasting with dynamic risk mitigation strategies to coordinate test execution with production workload fluctuations. Contribution/Results: Compared to manual approaches, the method significantly reduces configuration errors and service disruption risks. In real-world cloud service deployments, it decreases SLA violations induced by testing by 62% and shortens average test preparation time by 57%, while maintaining full test coverage and effectiveness. The framework provides a scalable, resource-aware automation solution for safe and efficient online testing of large-scale distributed systems.
This study investigates the practical prioritization of key quality attributes—such as fault detection capability, usability, and maintainability—in test cases and test suites, along with associated implementation challenges in industrial software testing. We designed a structured questionnaire grounded in a systematic literature review and deployed it across a large-scale, heterogeneous cohort of software testing practitioners on LinkedIn, yielding 354 valid responses. Mixed-method analysis (qualitative and quantitative) revealed significant contextual variations in attribute prioritization across domains—including agile, embedded, and web development—and identified three pervasive barriers: ambiguous attribute definitions, absence of actionable measurement metrics, and lack of formal review mechanisms. To our knowledge, this is the first empirical study to systematically characterize such cross-domain perceptual differences and practical impediments. The findings provide evidence-based guidance for refining test quality assessment frameworks and prioritizing engineering improvements in real-world testing practice.
Continuous Integration (CI) practices suffer from severe monitoring deficiencies: developers largely neglect critical metrics such as “build health” and “time-to-fix failed builds,” while mainstream CI services offer only weak native monitoring capabilities, forcing reliance on fragmented and often redundant third-party tools. Method: We conducted a triangulated investigation—including documentation analysis, developer surveys, functional audits of CI platforms, and case studies of open-source projects—to systematically identify cognitive gaps and practical monitoring needs. Contribution/Results: Our study provides the first empirical evidence that although over 80% of developers track test coverage, only a minority monitor build health or timeliness; further, all major CI services lack built-in multidimensional monitoring support. These findings establish an evidence-based foundation for designing next-generation CI monitoring frameworks and prioritizing tooling enhancements.
To address the challenges of automated test execution across projects, programming languages, build systems, and testing frameworks, this paper proposes ExecutionAgent—a large language model (LLM)-based autonomous agent. It employs a meta-prompt-driven system interaction paradigm to parse source code, perform environment-aware configuration, and autonomously generate test scripts for arbitrary open-source projects, supporting feedback-guided iterative debugging. Its novel zero-shot adaptation mechanism requires no predefined rules or human intervention, ensuring compatibility with 14 programming languages and mainstream ecosystems. Evaluated on 50 heterogeneous projects, ExecutionAgent successfully executed 33 test suites with only 7.5% result deviation from ground truth, achieving 6.6× higher performance than state-of-the-art methods. The average per-project execution time is 74 minutes, with an LLM inference cost of just $0.16.
This study addresses the imbalance in the test pyramid—characterized by an overreliance on coarse-grained integration and system tests, which leads to difficulties in fault localization and slow execution—by proposing, for the first time, a method to automatically generate unit tests from existing integration tests. The approach combines static and dynamic analysis to automatically isolate component dependencies and enhance coverage at the unit level. Implemented as a Node.js tool and evaluated on twelve open-source JavaScript projects, the technique produces high-quality unit tests that significantly improve test suite structure, thereby increasing both testing efficiency and maintainability.
This work addresses a critical limitation in existing cloud skill testing, which focuses solely on task success rates and fails to reveal uncovered behaviors, leaving test adequacy unquantifiable. To bridge this gap, the study introduces the first formal definition of test coverage units, coverage relationships, and a computational pipeline for cloud skills. It proposes a coverage evaluation method grounded in natural language skill packages: by parsing user prompts and initial resource states, the approach reconstructs operational obligations and models workflow context to establish an end-to-end measurement pipeline. Integrating model-assisted candidate generation with expert review, the framework produces auditable coverage reports and closes the loop by mapping coverage gaps to source-level test improvement recommendations. Empirical results demonstrate that this methodology enables quantitative assessment of test coverage for production-grade cloud skills, substantially enhancing their reliability and observability.
Traditional test adequacy metrics, such as code coverage and mutation testing, focus on implementation details and struggle to capture discrepancies between expected and actual program behavior. This work proposes an automated approach that extracts method-level expected behaviors from natural language documentation and source code, then maps them to existing test cases, thereby formalizing and empirically evaluating “behavioral gaps”—a dimension of test adequacy independent of structural metrics. By integrating natural language processing, static analysis, and behavioral mapping techniques, our method identifies 20,729 behaviors across ten Java open-source libraries with 93.1% precision, revealing that 17.5% of expected behaviors remain entirely untested. Notably, these gaps persist even in methods exhibiting high code coverage or high mutation kill rates, exposing a systematic deficiency in current testing practices—including automatically generated tests—in validating intended program behavior.
This work addresses the high cost of regression testing in continuous integration by proposing a unified learning model that incorporates the structural semantics of code diffs into test prioritization—a dimension overlooked by existing approaches. The model integrates diff structural features, test coverage relationships, and historical execution behavior to predict the likelihood of test cases revealing faults. Evaluated through cross-project experiments on five projects from Defects4J, the approach demonstrates significantly superior fault detection effectiveness and model generalizability compared to baseline methods that do not account for commit-aware information.
This work addresses the growing complexity of CI/CD pipelines and the lack of structured analysis capabilities in existing tools for understanding their behavior, failures, and version evolution. The authors propose an innovative approach that uniquely integrates digital twin technology with BPMN-based modeling in DevOps contexts. By automatically parsing raw CI configurations and execution logs, the method constructs structured, high-level process models that enable pipeline visualization, failure traceability, and cross-version comparison. Evaluated across multiple open-source projects, the approach demonstrates effectiveness in monitoring, evolutionary analysis, and fault diagnosis, offering a modular and extensible foundational framework for the analysis and optimization of CI/CD pipelines.