Score
Designs, configures, and maintains build toolchains for Apple platforms, including installing and customizing Xcode toolchains and Apple Silicon/Intel compiler, SDK, linker, and runtime sets; manages multiple toolchain versions and cross‑architecture compilation/ABI compatibility. Integrates and deploys toolchains into CI/CD and build systems to ensure reproducible builds, correct debugging artefacts, and platform-specific packaging and deployment workflows.
This study addresses core challenges in CI/CD practices for open-source Android applications—namely, configuration complexity, high maintenance overhead, and lack of cross-platform standardization. We conduct an empirical analysis of 2,557 GitHub projects utilizing GitHub Actions, Travis CI, CircleCI, and GitLab CI/CD, employing YAML configuration parsing, commit log topic modeling, and cross-platform comparative analysis. Key findings include: only 9% of projects achieve automated deployment; configurations are updated on average every two months, with over one-third of changes addressing build failures; 50% perform only basic testing; and maintenance intensity strongly correlates with project activity, size, and community engagement. We identify 11 recurring maintenance themes and highlight urgent needs for AI-driven automation tools and portable open-source solutions. This work provides empirical foundations and practical guidelines for standardizing CI/CD in mobile development.
This work addresses the frequent failures in continuous integration (CI) builds of embedded open-source software, which often stem from cross-compilation complexities, board-specific configurations, and toolchain constraints. These failures are compounded by heterogeneous, ephemeral build logs that are difficult to reuse. To tackle this challenge, the authors propose PhantomRun, a framework that enables standardized reproduction of historical failed builds through a build log abstraction layer, metadata standardization, containerized replay environments, and heterogeneous log parsing techniques. PhantomRun is the first system to support large-scale, controllable replay of failed embedded CI builds, offering a unified, machine-readable interface for build artifacts and metadata. Evaluated on 4,628 failed runs, PhantomRun successfully reconstructed 91.8% of the builds, with 98% preserving the original execution outcomes, demonstrating high reproducibility fidelity.
ASIC development faces challenges in IP reuse and lacks integrated hardware-software co-verification and unified build infrastructure. Method: This paper introduces SoCMake—the first unified SoC build system supporting cross-compilation of Chisel/SystemRDL hardware descriptions with C/C++/assembly code. It integrates RTL generation, simulation, firmware compilation, and SoC configuration into a single workflow, overcoming the limited software compilation support of conventional hardware build tools. By deeply embedding SystemC, the RISC-V toolchain, and CMake’s extensibility framework, SoCMake enables automated, abstraction-level–aware co-building across hardware description → RTL → firmware. Contribution/Results: SoCMake has successfully accelerated iterative deployment of radiation-tolerant RISC-V SoCs in high-energy physics applications. After open-sourcing, it has become a de facto standard for generic SoC generation, reducing overall SoC development time by over 40% in empirical evaluations.
This work proposes PhantomRun, a novel framework that leverages large language models (LLMs) to automatically repair compilation failures in continuous integration (CI) pipelines for embedded open-source software—a domain often plagued by hardware dependencies, syntax errors, and build script issues that incur substantial debugging overhead. PhantomRun integrates build logs, source code, historical fixes, and error diagnostics to generate and validate repair patches. The framework incorporates an adapter layer to ensure compatibility with diverse CI platforms such as GitHub Actions and GitLab CI, as well as multiple build systems. Experimental evaluation on four widely used embedded software projects demonstrates that PhantomRun successfully resolves 45% of CI compilation failures, thereby establishing the effectiveness and practicality of LLMs in this challenging context.
This study presents the first empirical investigation into the evolution of CI/CD configurations in machine learning (ML) projects. Addressing the lack of understanding regarding how CI/CD configurations co-evolve with ML components, the authors analyze 508 open-source ML projects, 343 manually annotated commits, and 15,634 automated CI/CD commits. They propose a novel 14-category taxonomy capturing synergistic changes between CI/CD and ML components, develop a dedicated clustering tool to identify recurrent evolutionary patterns, and establish an empirically grounded model linking developer experience to CI/CD configuration modification behavior. Results show that 61.8% of CI/CD-related commits involve build strategy modifications; common anti-patterns—including dependency hardcoding and missing test frameworks—are identified; and senior developers modify CI/CD configurations more frequently and effectively than juniors, confirming the critical role of experience in CI/CD maintenance.
This study addresses the impact of continuous integration (CI) build duration on developer productivity and investigates the underexplored adoption and efficacy of caching as a build acceleration technique. Through a large-scale empirical analysis of 513,384 builds across 1,279 GitHub projects on Travis CI—combining log analysis, pull request intervention experiments, and developer feedback—the work reveals that CI caching is adopted by only 30% of projects, largely due to developers’ limited awareness and the perceived maintenance complexity. Notably, nearly half of previously non-adopting projects accepted caching configurations when proposed via pull requests, with approximately one-third achieving significant speedups. However, widespread issues of cache redundancy and staleness were also observed, underscoring both the feasibility and necessity of optimizing CI caching practices.
This work addresses the challenge of detecting and repairing compilation errors caused by feature interactions in configurable C systems—errors that traditional compilers and existing variability-aware tools struggle to handle effectively. We present the first systematic exploration of leveraging foundation models for this task, proposing a variability-aware error detection and repair approach based on GPT-OSS-20B and Gemini 3 Pro. Our method is evaluated across synthetic systems, real-world GitHub commits, and mutation testing scenarios. Experimental results demonstrate that GPT-OSS-20B achieves 0.97 precision, 0.90 recall, and 0.94 accuracy on small-scale systems, successfully repairing over 70% of the errors. Notably, it also uncovers potential compilation defects in real Linux commits, offering a low-overhead, high-coverage alternative for variability-aware compilation.
This work addresses the challenge of automated configuration migration across CI platforms—particularly from Travis CI to GitHub Actions—where manual translation is error-prone and labor-intensive. We propose an LLM-based translation framework grounded in empirical analysis of 811 real-world migration cases, enabling the first quantitative characterization of configuration conversion effort. We introduce a four-category taxonomy of translation problems and identify recurring developer pain points. Methodologically, we design a composite prompting strategy integrating documentation-guided instruction, iterative refinement, and in-context learning to enhance LLM robustness and fidelity. Evaluation on GPT-4o shows our approach achieves 75.5% end-to-end build success rate—nearly tripling the performance of baseline prompting—while substantially reducing manual intervention. The framework provides a reproducible, quantitatively evaluable pathway for intelligent CI/CD configuration migration.
This work addresses the challenges in edge and embedded application development—namely, heterogeneous software stacks, multi-language runtimes, and difficult debugging—which lead to rigid deployment workflows and complex fault diagnosis. To overcome these limitations, the paper proposes a novel architecture enabling unified end-edge-cloud development. Its core components include a single programming language, a retargetable runtime system, a local recording and replay mechanism for distributed events, and a cross-platform deployment framework. This design breaks down traditional debugging barriers in edge–cloud collaborative development, facilitating seamless scalability, consistent testing, and flexible deployment across heterogeneous environments. Evaluation of the prototype system demonstrates that the proposed approach significantly simplifies deployment procedures and enhances fault diagnosis efficiency.
This work addresses the challenge of automatically repairing software build failures during cross-Instruction-Set-Architecture (ISA) migration. To this end, we introduce Build-bench—the first end-to-end evaluation benchmark specifically designed for this scenario. Build-bench innovatively integrates architecture-aware reasoning, tool-augmented inference, and executable validation, enabling multi-turn autonomous repair via structure extraction, content modification, build execution, and log-driven feedback. We systematically evaluate six state-of-the-art large language models (LLMs) on 268 real-world build-failing packages; the best-performing model achieves a 63% build repair success rate. Our analysis reveals, for the first time, substantial disparities among models in tool-calling strategies and iterative repair behaviors. This work establishes a reproducible, executable evaluation paradigm and provides empirical foundations for LLM-driven cross-architecture software migration.