Score
Designs and implements measurement instruments, metrics, and empirical analyses to evaluate how AI tools are integrated into the software development lifecycle and affect development practices. This includes profiling repository-level AI usage patterns, analyzing chat and AI-generated content, and measuring impacts on contributor activity, code quality, and merge/acceptance signals.
This study addresses critical challenges in human-AI collaboration for software development—including low co-development efficiency, insufficient trust, and weak perceived control. To tackle these issues, we propose the first taxonomy of developer-AI interactions spanning the entire software engineering lifecycle. Grounded in empirical analysis and consensus among domain experts, the taxonomy systematically classifies eleven distinct interaction patterns, including auto-completion, instruction-driven programming, and conversational assistance. Unlike prior fragmented characterizations, this structured framework establishes a foundational paradigm for AI tool design, human factors evaluation, and research on trustworthy collaborative mechanisms. The taxonomy enables principled, theory-guided optimization of adaptive and reliable programming assistants—shifting human-AI software development from empirically driven practice toward rigorous, evidence-based methodology.
This study addresses the lack of a systematic understanding of generative artificial intelligence’s role across the full software development lifecycle. Through a systematic literature review complemented by structured surveys of 65 developers, this work integrates empirical data with existing research to comprehensively evaluate the real-world impact and adoption patterns of large language models (LLMs) in each development phase. Findings indicate that over 70% of developers save more than 50% of their time on boilerplate code generation and documentation tasks, and 79% use browser-based LLMs daily. While nascent governance mechanisms are emerging, benefits in early-stage activities—such as requirements elicitation and architectural design—remain limited. The results suggest that generative AI is shifting the locus of development value from coding toward upstream design activities.
This study addresses the growing challenge posed by the widespread involvement of AI agents in software development, which undermines the long-standing assumption that development artifacts are exclusively produced by human professionals—an assumption underpinning traditional software metrics. The work systematically exposes how AI-generated traces compromise the foundational premises of established software measurement practices, thereby threatening the validity of prior empirical conclusions. To confront this issue, the authors propose an AI-augmented, systematic replication methodology that integrates modern data analytics with empirical software engineering techniques to rigorously re-evaluate key findings. The project advances a dynamic, reproducible, and sustainable measurement paradigm capable of adapting to evolving data ecosystems, offering a robust and timely framework for software metrics in the AI era.
This study addresses the challenge of systematically integrating generative AI into the entire software development lifecycle to enhance productivity while ensuring quality and governance. The authors propose a progressive integration framework centered on an innovative “AI harness” that unifies management of project context, access control, validation, logging, and human approval workflows. This architecture enables seamless co-evolution of technical capabilities, organizational processes, and quality assurance mechanisms. The framework supports a transition from informal AI assistance toward controlled, agent-based development and is empirically validated through a case study in a mid-sized software enterprise, offering both a practical roadmap and evidence-based foundation for AI-driven transformation in software engineering.
This study addresses the lack of large-scale empirical analysis on AI-generated code in real-world software repositories. By constructing a large dataset of authentic code repositories and combining heuristic filtering with large language model–based classification, this work systematically identifies and measures multidimensional characteristics of AI-generated code in realistic development contexts for the first time. It evaluates such code comprehensively from both code-level and commit-level perspectives, examining complexity, structural properties, defect rates, and developer behavior. The findings reveal that AI-generated code significantly differs from human-written code in structural simplicity, complexity distribution, and commit stability, offering critical empirical evidence for understanding the practical impact of AI-assisted programming.
This study investigates the mechanisms through which AI-generated code affects software engineers’ productivity and long-term software quality—specifically maintainability and extensibility. Using a mixed-methods approach, it combines practitioner surveys with multi-project empirical codebase analysis and statistical modeling across tasks of varying complexity. Results show that AI significantly improves developer efficiency for small-scale tasks without compromising code quality; however, in high-complexity systems, direct integration of AI-generated code increases architectural coupling and maintenance risks. Such challenges necessitate architect-led problem decomposition and manual integration. Based on these findings, the study proposes a “human–AI layered collaboration” model: AI handles module-level implementation, while humans retain responsibility for architectural design and cross-module integration. This model preserves productivity gains while mitigating AI’s adverse effects on long-term software quality, offering a practical governance framework for AI-augmented software engineering.
Current AI assistant features in IDEs exhibit a significant misalignment with developers’ authentic needs, necessitating a systematic understanding of heterogeneous user requirements. Method: We conducted semi-structured interviews with 35 practitioners—comprising AI adopters, attriters, and non-users—to empirically construct the first human-AI interaction design space for IDE-integrated AI assistants. Through thematic coding and cross-cohort comparative analysis, we identified fundamental divergences across user groups along five dimensions: reliability, privacy, personalization, proactivity, and ethical concerns. Contribution/Results: We propose a role-driven, five-dimensional design framework—encompassing technical robustness, interaction modality, goal alignment, skill abstraction, and cognitive offloading—alongside 12 actionable design guidelines. This work advances IDE AI tools toward greater reliability, contextual awareness, privacy-by-design, and seamless workflow integration.
This study investigates the impact of generative artificial intelligence (GenAI) tools on team creativity and collaboration in software development—a domain inherently collaborative yet increasingly mediated by individually oriented AI technologies. Through semi-structured interviews with 13 software engineers across four companies, the research employs qualitative thematic analysis to uncover an emergent triadic collaboration paradigm among developers, colleagues, and AI. Findings reveal that while GenAI expands avenues for idea generation, it may concurrently suppress independent ideation; developers often prioritize consulting AI over seeking input from teammates, thereby diminishing peer learning opportunities. Crucially, affective dynamics serve as a pivotal cross-cutting dimension, permeating both creative and collaborative processes. These insights offer theoretical and practical implications for understanding how AI reshapes collaborative mechanisms within real-world software engineering teams.
This work addresses the limited accessibility of large language model (LLM) and agent workflow development for engineers without machine learning expertise, primarily due to the absence of integrated testing, debugging, and reproducibility capabilities. To bridge this gap, the authors propose a novel IDE-native AI observability workflow, implemented as the AI Toolkit plugin for JetBrains IDEs. This approach seamlessly embeds trace capture and evaluation into standard run/debug cycles, enabling automatic hierarchical trace logging during execution, one-click dataset persistence, and a pluggable, unit-test-like evaluation framework. By minimizing environment setup and context-switching overhead, the solution facilitates routine evaluation and immediate trace visualization. Empirical data from the initial PyCharm release demonstrates high adoption, sustained usage, and low churn, confirming that IDE-integrated tooling effectively lowers the barrier to entry for non-ML developers.
This study addresses the current lack of a systematic theoretical framework explaining how software professionals evaluate AI-generated code and the underlying cognitive processes and preferences involved. Employing a constructivist grounded theory approach, the research integrates questionnaire surveys, semi-structured interviews, and laddering interviews to iteratively collect data from 20 to 50 practitioners until theoretical saturation is achieved. The work presents the first empirically grounded theoretical framework for assessing AI-generated code, elucidating the evaluation mechanisms developers employ in human-AI collaborative programming contexts. By doing so, it fills a critical gap in the literature concerning both behavioral and cognitive dimensions of code evaluation in AI-assisted software development.
This study addresses the growing concerns regarding the quality, reliability, and security of AI-generated code by systematically identifying its key influencing factors. Through a rigorous systematic literature review—combining AI-assisted screening with manual validation—the authors synthesize evidence from 24 empirical studies to integrate, for the first time, three critical dimensions: human factors (e.g., developer expertise), system characteristics (e.g., prompt design), and human-AI interaction aspects (e.g., task specification). The findings reveal that AI-assisted programming functions as a socio-technical system, wherein different quality attributes—such as correctness and security—exhibit markedly divergent behaviors. While the approach demonstrates considerable potential for enhancing software development, it also introduces non-trivial risks, thereby offering both theoretical insights and practical guidance for the future design and evaluation of AI-powered programming tools.
This study addresses the challenges of frequent requirement changes, quality assurance, and delivery efficiency in agile software development by systematically investigating the application of artificial intelligence—particularly machine learning and natural language processing—in critical phases such as requirements management, code generation, and testing. Through a comprehensive literature review and empirical research involving industry practitioners, the work demonstrates that AI not only enhances existing agile practices but also fundamentally reshapes software development paradigms by introducing automation and intelligent decision-making capabilities. This transformation significantly improves development efficiency, product quality, and team responsiveness, thereby enabling a synergistic advancement in quality, speed, and innovation.