Score
Designs, implements, and evaluates software tools, utilities, and automation that enable development, testing, deployment, and maintenance—for example IDE plugins, command-line interfaces, build/test pipelines, debuggers, profilers, and deployment toolchains. Builds integrations, APIs, instrumentation, and automation scripts, and analyzes tool reliability, performance, usability, and workflow compatibility.
This work addresses the limited accessibility of large language model (LLM) and agent workflow development for engineers without machine learning expertise, primarily due to the absence of integrated testing, debugging, and reproducibility capabilities. To bridge this gap, the authors propose a novel IDE-native AI observability workflow, implemented as the AI Toolkit plugin for JetBrains IDEs. This approach seamlessly embeds trace capture and evaluation into standard run/debug cycles, enabling automatic hierarchical trace logging during execution, one-click dataset persistence, and a pluggable, unit-test-like evaluation framework. By minimizing environment setup and context-switching overhead, the solution facilitates routine evaluation and immediate trace visualization. Empirical data from the initial PyCharm release demonstrates high adoption, sustained usage, and low churn, confirming that IDE-integrated tooling effectively lowers the barrier to entry for non-ML developers.
This study addresses the lack of empirical evidence on the real-world impact of UI testing frameworks in CI/CD pipelines. Using GitHub API data collection, YAML configuration parsing, CI log metric extraction, and controlled time-series analysis across open-source repositories, we systematically quantify the integration patterns and effects of Selenium, Playwright, and Cypress within GitHub Actions workflows. Results show that UI testing significantly improves test pass-rate stability but increases mean build duration by 12% initially. Highly active repositories prefer Playwright—its built-in retry mechanism reduces flaky-test-induced pipeline interruptions by 35%. This work fills a critical gap in understanding UI testing’s practical implications in production CI/CD environments, providing data-driven insights for quality assurance strategy design and framework selection.
Implementation discrepancies across software repository mining tools severely threaten the validity of empirical findings. Method: We conduct a dual-tool comparative analysis of 10 large-scale open-source projects, systematically identifying how minor implementation differences—such as commit parsing logic and author deduplication rules—induce up to 500% deviation in key metrics (e.g., commit count, developer count). We propose a “tool-level configuration + post-hoc normalization” co-optimization framework to mitigate metric divergence and perform multi-tool experiments, quantitative consistency assessment, and code-level root-cause analysis. Contribution/Results: We identify six technical sources undermining data validity and establish the first validity assessment paradigm for Mining Software Projects Research (MSPR) explicitly addressing tool heterogeneity—thereby enabling rigorous, reproducible, and comparable empirical software engineering studies.
This paper addresses the conceptual ambiguity, ill-defined boundaries, and lack of implementation standards between Infrastructure-as-Code (IaC) and Pipeline-as-Code in DevOps practice. To resolve these issues, we systematically delineate their respective roles and synergistic mechanisms within the DevOps ecosystem and propose a reusable, standardized IaC-driven CI/CD implementation framework. Our approach integrates Terraform for infrastructure provisioning, Ansible for configuration management, GitLab CI for pipeline orchestration, and Docker/Kubernetes for containerized deployment—enabling an end-to-end automated delivery pipeline. Empirical evaluation demonstrates 99.8% configuration change accuracy, reduces environment provisioning time from hours to minutes, and significantly improves deployment consistency and delivery efficiency.
Current AI assistant features in IDEs exhibit a significant misalignment with developers’ authentic needs, necessitating a systematic understanding of heterogeneous user requirements. Method: We conducted semi-structured interviews with 35 practitioners—comprising AI adopters, attriters, and non-users—to empirically construct the first human-AI interaction design space for IDE-integrated AI assistants. Through thematic coding and cross-cohort comparative analysis, we identified fundamental divergences across user groups along five dimensions: reliability, privacy, personalization, proactivity, and ethical concerns. Contribution/Results: We propose a role-driven, five-dimensional design framework—encompassing technical robustness, interaction modality, goal alignment, skill abstraction, and cognitive offloading—alongside 12 actionable design guidelines. This work advances IDE AI tools toward greater reliability, contextual awareness, privacy-by-design, and seamless workflow integration.
Bug fixing is a complex and time-consuming task in software development. Bug localization research tends to focus on the accuracy of automated tools that suggest source code files for developers to look at. However, little is known about how developers use these tools in practice. This paper reports on an ongoing qualitative user study. Eleven participants worked through four realistic bug localization tasks in a controlled environment and were given varying levels of support information offered by a specialized tool. Participants were asked to think aloud in a semi-structured interview session. The preliminary findings provide insight into three aspects of practice: how developers interact with tools, the role social and contextual information plays, and problem solving. The study demonstrates that bug localization is complex and suggests that the adoption of effective tools depends on more than their accuracy.
Software testing is a fundamental process of software development, and prior work has shown that visualizations of test results support testers' decision-making. However, Human-Computer Interaction research on software testing has yet to explore and understand the shared interface elements and patterns in visualization of testing outputs. To address this, we conducted a visual comparative analysis of the output of 50 software testing tools and harnesses (44 with CLI output, 6 with GUI output) across four popular programming languages. Our analysis reveals the common interface elements in software testing tools, how these tools display and visualize test results, as well as the specific make-up of the output. Our findings provide insight on how visual testing output is formatted and how colour is used across both CLI and GUI environments, identifying trends that can be applied by developers of testing tools.
This work addresses the challenge of balancing software quality, testability, and maintainability under rapid iteration and frequent requirement changes. It proposes Algorithm-Driven Development (ADD), a novel approach that unifies requirements specification and technical design by using algorithm flowcharts as a single, coherent artifact. This integration enables end-to-end modeling of requirements, architecture, and testing. Leveraging this model, the system automatically generates high-coverage acceptance tests and incorporates continuous integration with code coverage feedback. Industrial adoption at Dassault Systèmes demonstrates that ADD achieves over 95% code coverage, substantially reduces defect density, and ensures a stable delivery cadence, outperforming conventional test-driven development and test-after approaches.
This study addresses the lack of systematic understanding regarding how GitHub Actions workflows are used in real-world scenarios, how developers respond to workflow failures, and how these practices relate to project characteristics. Combining large-scale quantitative analysis of 258,300 workflow runs with qualitative case studies across 21 diverse repositories, this work identifies three typical patterns developers employ to handle workflow failures and uncovers a “configuration–usage gap”—where YAML configurations exist but workflows remain effectively unused. Furthermore, the study empirically validates five hypotheses linking project features to workflow usage intensity, revealing a significant positive correlation between high usage intensity and low failure rates. These findings provide actionable empirical evidence for improving CI/CD practices.
This study addresses the challenge of systematically integrating generative AI into the entire software development lifecycle to enhance productivity while ensuring quality and governance. The authors propose a progressive integration framework centered on an innovative “AI harness” that unifies management of project context, access control, validation, logging, and human approval workflows. This architecture enables seamless co-evolution of technical capabilities, organizational processes, and quality assurance mechanisms. The framework supports a transition from informal AI assistance toward controlled, agent-based development and is empirically validated through a case study in a mid-sized software enterprise, offering both a practical roadmap and evidence-based foundation for AI-driven transformation in software engineering.