Score
Designs, builds, and measures the technical systems, processes, and experiments that increase adoption of a software platform — for example onboarding flows, migration and integration tooling, SDKs, telemetry and analytics, and feedback/notification infrastructure. Analyzes adoption funnels, retention cohorts, A/B tests, and technical barriers to recommend and implement engineering changes that improve activation, retention, and long‑term platform usage.
New software engineers often struggle to comprehend large legacy systems, leading to prolonged onboarding periods. Method: This study introduces a systems-thinking training program grounded in Labelled Transition System (LTS) modeling and a structured understanding template—the first application of LTS modeling in software engineering onboarding education—featuring differentiated learning pathways across five sessions, integrating pedagogical best practices and pre-/post-assessment design. Contribution/Results: While overall comprehension gains were not statistically significant, learners with low initial proficiency showed a robust 15-percentage-point improvement (p < 0.05). Qualitative feedback indicated high engagement and perceived practical relevance. The framework offers a scalable, low-cost, reusable instructional model for cultivating software comprehension skills, addressing a critical gap in industry onboarding programs by introducing formal modeling techniques into foundational training.
Industrial adoption of machine learning–enhanced system trace analysis tools remains limited due to the “excellence paradox”: advanced technical capabilities often compromise usability, transparency, and user trust. Method: We propose an adoption-oriented design paradigm grounded in three principles—cognitive compatibility, expert knowledge embedding, and transparency-driven trust building—and implement TMLL, a tool library integrating system tracing, interpretable ML, and human-in-the-loop mechanisms to support semi-automated analysis, incremental workflow integration, and verifiable outputs. Contribution/Results: Validated by Ericsson experts and integrated into the Eclipse Foundation ecosystem, TMLL was evaluated via a survey of 40 industry and academic experts. Results show 77.5% prioritize result credibility and 67.5% prefer controllable semi-automation. This work establishes both a theoretical framework and practical pathway for deploying trustworthy AI tools in industrial settings.
Contemporary technical hiring practices suffer from stress-induced bias and evidentiary gaps, resulting in distorted competency assessments and compromised fairness. Method: This paper proposes an evidence-driven paradigm for software engineer competency evaluation. It systematically identifies and bridges evidentiary gaps in technical hiring through (1) multi-source behavioral data integration, (2) low-stress, authentic task design, and (3) a verifiable fairness framework grounded in educational measurement, human-computer interaction evaluation, algorithmic fairness auditing, and structured competency modeling. Contribution/Results: The approach yields a scalable, empirically validated hiring effectiveness metric suite. Empirical evaluation demonstrates significant improvements in employer hiring accuracy. Crucially, it establishes a reproducible, auditable foundation for equitable assessment—enhancing both validity and procedural fairness for candidates while enabling rigorous, transparent evaluation of hiring systems.
本文提出了一种通过生产使用信号直接计算变更管理阶段进展的方法来衡量企业AI采用情况,并提供了一个统一框架、具体操作化模型及开源实现。
Existing RTI policy monitoring tools suffer from inefficient information acquisition and inadequate dynamic tracking capabilities. This study proposes a methodology for developing an open-source web-based system dedicated to Research, Technology, and Innovation (RTI) policy monitoring. It employs role-driven requirements engineering to precisely elicit heterogeneous stakeholder needs and introduces a novel modular architectural paradigm centered on a user-configurable dashboard, with strict separation across presentation, service, and data layers. The approach integrates open-data interoperability standards and interactive visualization techniques. As key contributions, the work delivers a reusable RTI monitoring system architecture specification and a standardized dashboard requirements template. These artifacts were empirically validated through deployment in the Austrian RTI Monitor—a national-scale platform enabling cross-departmental, real-time indicator tracking and evidence-informed policy coordination—thereby substantially enhancing the timeliness, accessibility, and scalability of RTI policy monitoring.
Frameworks such as SPACE, DevEx, and DORA established that developer productivity is inherently multidimensional, but left practitioners with a practical question: what should we measure, and how should we use it to improve? This paper introduces Engineering Thrive (EngThrive), a measurement and improvement system developed and deployed across Microsoft's engineering organization. EngThrive organizes productivity around three dimensions - Speed, Ease, and Quality - with Thriving as a guardrail to ensure developer wellbeing improves alongside performance. Within each dimension, outcome-oriented North Star metrics are paired with diagnostic submetrics, combining system telemetry with developer surveys to provide both scale and context. We describe the design principles that guide metric selection, including an approach in which well-chosen metrics align "gaming" behavior with genuine improvement. We also outline the data platform, survey program, and dashboard ecosystem required to operationalize this approach in practice, and present case studies demonstrating how outcome-oriented measurement enables sustained, system-level improvements. Finally, we show that EngThrive functions as a general-purpose evaluation language, applicable not only to developer tools and AI, but to organizational policies, work environments, and other factors that shape how developers experience their work. We offer EngThrive as a concrete model for organizations seeking to move beyond measuring activity toward improving outcomes.
This study addresses the lack of empirical evidence guiding developers’ choices among continuous integration (CI) services, which obscures whether adoption decisions stem from genuine project requirements or social influence, and leaves unclear their receptiveness to CI recommendation systems. By conducting an online survey with approximately 5,000 active GitHub developers and integrating their behavioral data, this work systematically disentangles demand-driven factors from social influence in CI adoption and investigates developers’ perceptions of, and barriers to adopting, automated CI recommendation systems. The findings provide an empirical foundation for designing effective CI recommendation tools, thereby supporting open-source projects in more successfully promoting CI practices.
This study investigates the drivers of continued intention to use generative artificial intelligence (GenAI) tools among software developers in Italian small and medium-sized enterprises following initial adoption. Building on an extended UTAUT2 framework, the research employs a two-stage mixed-methods approach, integrating a longitudinal pilot study with a cross-sectional survey, and analyzes questionnaire data and semi-structured interviews using partial least squares structural equation modeling (PLS-SEM). The findings reveal that individual perceptions—such as enhanced productivity, perceived ease of use, and hedonic motivation—significantly influence sustained usage intentions, whereas social and organizational factors exhibit negligible effects, thereby challenging the applicability of traditional technology acceptance theories. The proposed model accounts for 64.7% of the variance in continuance intention, underscoring the pivotal role of positive individual experiences in the ongoing adoption of GenAI tools.
This study addresses the challenge faced by open-source software (OSS) contributors in selecting suitable projects, which often hampers their participation efficiency. Through a survey of 208 contributors—both newcomers and experienced participants—the work presents the first systematic comparison of differences in project selection preferences, motivational drivers, and demographic characteristics. Integrating quantitative statistical analysis with qualitative feedback, the findings reveal that age, gender, and contributor role significantly shape participation motives. Furthermore, preferences regarding project age, development stage, and documentation quality vary systematically with these underlying motivations. These insights offer empirical grounding and actionable directions for improving contributor retention and designing motivation-aware recommendation mechanisms for OSS projects.
This work proposes a methodology for customizing large language models (LLMs) tailored to enterprise software engineering contexts, aiming to enhance developer productivity and code quality. By constructing a trillion-token-scale dataset comprising proprietary enterprise code and engineering artifacts, the approach integrates continual pretraining, post-training optimization, and an intermediate training strategy designed to mitigate catastrophic forgetting, thereby enabling deep adaptation to enterprise development environments. The end-to-end pipeline encompasses high-value signal extraction, data curation, and deployment. In a blind evaluation involving 29,000 developers, the customized model significantly reduced interactive iteration counts by 23% and improved code survival rates by approximately 17%, demonstrating its effectiveness and generalizability.