evaluate ai-assisted development

Designs and implements measurement instruments, metrics, and empirical analyses to evaluate how AI tools are integrated into the software development lifecycle and affect development practices. This includes profiling repository-level AI usage patterns, analyzing chat and AI-generated content, and measuring impacts on contributor activity, code quality, and merge/acceptance signals.

evaluateai-assisteddevelopment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

This study addresses the lack of a systematic understanding of generative artificial intelligence’s role across the full software development lifecycle. Through a systematic literature review complemented by structured surveys of 65 developers, this work integrates empirical data with existing research to comprehensively evaluate the real-world impact and adoption patterns of large language models (LLMs) in each development phase. Findings indicate that over 70% of developers save more than 50% of their time on boilerplate code generation and documentation tasks, and 79% use browser-based LLMs daily. While nascent governance mechanisms are emerging, benefits in early-stage activities—such as requirements elicitation and architectural design—remain limited. The results suggest that generative AI is shifting the locus of development value from coding toward upstream design activities.

AI GovernanceDeveloper SurveyGenerative AI

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the growing challenge posed by the widespread involvement of AI agents in software development, which undermines the long-standing assumption that development artifacts are exclusively produced by human professionals—an assumption underpinning traditional software metrics. The work systematically exposes how AI-generated traces compromise the foundational premises of established software measurement practices, thereby threatening the validity of prior empirical conclusions. To confront this issue, the authors propose an AI-augmented, systematic replication methodology that integrates modern data analytics with empirical software engineering techniques to rigorously re-evaluate key findings. The project advances a dynamic, reproducible, and sustainable measurement paradigm capable of adapting to evolving data ecosystems, offering a robust and timely framework for software metrics in the AI era.

AI agentsfoundational assumptionsreplication

This study addresses the challenge of systematically integrating generative AI into the entire software development lifecycle to enhance productivity while ensuring quality and governance. The authors propose a progressive integration framework centered on an innovative “AI harness” that unifies management of project context, access control, validation, logging, and human approval workflows. This architecture enables seamless co-evolution of technical capabilities, organizational processes, and quality assurance mechanisms. The framework supports a transition from informal AI assistance toward controlled, agent-based development and is empirically validated through a case study in a mid-sized software enterprise, offering both a practical roadmap and evidence-based foundation for AI-driven transformation in software engineering.

Agentic DevelopmentAI-driven Software DevelopmentDevelopment Process Governance

This study addresses the lack of large-scale empirical analysis on AI-generated code in real-world software repositories. By constructing a large dataset of authentic code repositories and combining heuristic filtering with large language model–based classification, this work systematically identifies and measures multidimensional characteristics of AI-generated code in realistic development contexts for the first time. It evaluates such code comprehensively from both code-level and commit-level perspectives, examining complexity, structural properties, defect rates, and developer behavior. The findings reveal that AI-generated code significantly differs from human-written code in structural simplicity, complexity distribution, and commit stability, offering critical empirical evidence for understanding the practical impact of AI-assisted programming.

AI-generated codeempirical studylarge-scale measurement

This study investigates the mechanisms through which AI-generated code affects software engineers’ productivity and long-term software quality—specifically maintainability and extensibility. Using a mixed-methods approach, it combines practitioner surveys with multi-project empirical codebase analysis and statistical modeling across tasks of varying complexity. Results show that AI significantly improves developer efficiency for small-scale tasks without compromising code quality; however, in high-complexity systems, direct integration of AI-generated code increases architectural coupling and maintenance risks. Such challenges necessitate architect-led problem decomposition and manual integration. Based on these findings, the study proposes a “human–AI layered collaboration” model: AI handles module-level implementation, while humans retain responsibility for architectural design and cross-module integration. This model preserves productivity gains while mitigating AI’s adverse effects on long-term software quality, offering a practical governance framework for AI-augmented software engineering.

Analyzing AI-generated solution quality in complex projectsAssessing AI tools' impact on software engineer productivityEvaluating long-term effects of AI solutions on software quality

The Design Space of in-IDE Human-AI Experience

Oct 11, 2024
AS
Agnia Sergeyuk
🏛️ JetBrains Research | Delft University of Technology

Current AI assistant features in IDEs exhibit a significant misalignment with developers’ authentic needs, necessitating a systematic understanding of heterogeneous user requirements. Method: We conducted semi-structured interviews with 35 practitioners—comprising AI adopters, attriters, and non-users—to empirically construct the first human-AI interaction design space for IDE-integrated AI assistants. Through thematic coding and cross-cohort comparative analysis, we identified fundamental divergences across user groups along five dimensions: reliability, privacy, personalization, proactivity, and ethical concerns. Contribution/Results: We propose a role-driven, five-dimensional design framework—encompassing technical robustness, interaction modality, goal alignment, skill abstraction, and cognitive offloading—alongside 12 actionable design guidelines. This work advances IDE AI tools toward greater reliability, contextual awareness, privacy-by-design, and seamless workflow integration.

Address gaps in proactive and maintenance AI supportAssess feasibility of implementing requested AI featuresIdentify developers' unmet needs for AI assistants in IDEs

Latest Papers

What's happening recently
View more

This study investigates the impact of generative artificial intelligence (GenAI) tools on team creativity and collaboration in software development—a domain inherently collaborative yet increasingly mediated by individually oriented AI technologies. Through semi-structured interviews with 13 software engineers across four companies, the research employs qualitative thematic analysis to uncover an emergent triadic collaboration paradigm among developers, colleagues, and AI. Findings reveal that while GenAI expands avenues for idea generation, it may concurrently suppress independent ideation; developers often prioritize consulting AI over seeking input from teammates, thereby diminishing peer learning opportunities. Crucially, affective dynamics serve as a pivotal cross-cutting dimension, permeating both creative and collaborative processes. These insights offer theoretical and practical implications for understanding how AI reshapes collaborative mechanisms within real-world software engineering teams.

collaborationcreativityGenerative AI

This work addresses the limited accessibility of large language model (LLM) and agent workflow development for engineers without machine learning expertise, primarily due to the absence of integrated testing, debugging, and reproducibility capabilities. To bridge this gap, the authors propose a novel IDE-native AI observability workflow, implemented as the AI Toolkit plugin for JetBrains IDEs. This approach seamlessly embeds trace capture and evaluation into standard run/debug cycles, enabling automatic hierarchical trace logging during execution, one-click dataset persistence, and a pluggable, unit-test-like evaluation framework. By minimizing environment setup and context-switching overhead, the solution facilitates routine evaluation and immediate trace visualization. Empirical data from the initial PyCharm release demonstrates high adoption, sustained usage, and low churn, confirming that IDE-integrated tooling effectively lowers the barrier to entry for non-ML developers.

AI debuggingAI evaluationIDE integration

This study addresses the current lack of a systematic theoretical framework explaining how software professionals evaluate AI-generated code and the underlying cognitive processes and preferences involved. Employing a constructivist grounded theory approach, the research integrates questionnaire surveys, semi-structured interviews, and laddering interviews to iteratively collect data from 20 to 50 practitioners until theoretical saturation is achieved. The work presents the first empirically grounded theoretical framework for assessing AI-generated code, elucidating the evaluation mechanisms developers employ in human-AI collaborative programming contexts. By doing so, it fills a critical gap in the literature concerning both behavioral and cognitive dimensions of code evaluation in AI-assisted software development.

AI-generated codecode evaluationgenerative AI

This study addresses the growing concerns regarding the quality, reliability, and security of AI-generated code by systematically identifying its key influencing factors. Through a rigorous systematic literature review—combining AI-assisted screening with manual validation—the authors synthesize evidence from 24 empirical studies to integrate, for the first time, three critical dimensions: human factors (e.g., developer expertise), system characteristics (e.g., prompt design), and human-AI interaction aspects (e.g., task specification). The findings reveal that AI-assisted programming functions as a socio-technical system, wherein different quality attributes—such as correctness and security—exhibit markedly divergent behaviors. While the approach demonstrates considerable potential for enhancing software development, it also introduces non-trivial risks, thereby offering both theoretical insights and practical guidance for the future design and evaluation of AI-powered programming tools.

AI-generated codecode qualityempirical evidence

This study addresses the challenges of frequent requirement changes, quality assurance, and delivery efficiency in agile software development by systematically investigating the application of artificial intelligence—particularly machine learning and natural language processing—in critical phases such as requirements management, code generation, and testing. Through a comprehensive literature review and empirical research involving industry practitioners, the work demonstrates that AI not only enhances existing agile practices but also fundamentally reshapes software development paradigms by introducing automation and intelligent decision-making capabilities. This transformation significantly improves development efficiency, product quality, and team responsiveness, thereby enabling a synergistic advancement in quality, speed, and innovation.

Agile DevelopmentDevelopment EfficiencyProduct Quality

Hot Scholars

LA

Lennart Ante

Constructor University
blockchaincryptocurrencystablecoinsdecentralized finance
JB

James Brusseau

Pace University, AI Ethics Site
History of philosophy and ethicsArtificial intelligenceDecadence in philosophy
AJ

Aditya Johri

George Mason University, Professor & Endowed Research Fellow
Computing EducationEngineering Education ResearchAI Ethics EducationSocial Computing