Score
Designs, builds, and evaluates tools and workflows that integrate machine learning and generative models to assist software development tasks; this includes creating IDE plugins, APIs, datasets, benchmarks, and user interfaces for features such as code generation, completion, refactoring, testing, debugging, documentation, and code review. Work also covers implementing and measuring human–AI interaction, correctness and safety checks, automation of developer tasks, and infrastructure to deploy and iterate on these assistive capabilities.
This study addresses critical challenges in human-AI collaboration for software development—including low co-development efficiency, insufficient trust, and weak perceived control. To tackle these issues, we propose the first taxonomy of developer-AI interactions spanning the entire software engineering lifecycle. Grounded in empirical analysis and consensus among domain experts, the taxonomy systematically classifies eleven distinct interaction patterns, including auto-completion, instruction-driven programming, and conversational assistance. Unlike prior fragmented characterizations, this structured framework establishes a foundational paradigm for AI tool design, human factors evaluation, and research on trustworthy collaborative mechanisms. The taxonomy enables principled, theory-guided optimization of adaptive and reliable programming assistants—shifting human-AI software development from empirically driven practice toward rigorous, evidence-based methodology.
This study addresses the lack of a systematic understanding of generative artificial intelligence’s role across the full software development lifecycle. Through a systematic literature review complemented by structured surveys of 65 developers, this work integrates empirical data with existing research to comprehensively evaluate the real-world impact and adoption patterns of large language models (LLMs) in each development phase. Findings indicate that over 70% of developers save more than 50% of their time on boilerplate code generation and documentation tasks, and 79% use browser-based LLMs daily. While nascent governance mechanisms are emerging, benefits in early-stage activities—such as requirements elicitation and architectural design—remain limited. The results suggest that generative AI is shifting the locus of development value from coding toward upstream design activities.
This work addresses the limited accessibility of large language model (LLM) and agent workflow development for engineers without machine learning expertise, primarily due to the absence of integrated testing, debugging, and reproducibility capabilities. To bridge this gap, the authors propose a novel IDE-native AI observability workflow, implemented as the AI Toolkit plugin for JetBrains IDEs. This approach seamlessly embeds trace capture and evaluation into standard run/debug cycles, enabling automatic hierarchical trace logging during execution, one-click dataset persistence, and a pluggable, unit-test-like evaluation framework. By minimizing environment setup and context-switching overhead, the solution facilitates routine evaluation and immediate trace visualization. Empirical data from the initial PyCharm release demonstrates high adoption, sustained usage, and low churn, confirming that IDE-integrated tooling effectively lowers the barrier to entry for non-ML developers.
Despite growing adoption of generative AI in software development, its integration into professional practice—particularly across diverse engineering tasks—remains poorly understood. Method: We employed a mixed-methods approach, combining a large-scale survey of 91 professional software engineers with in-depth qualitative analysis, grounded in a software engineering task taxonomy to systematically examine prompting strategies, multi-turn interaction patterns, and reliability assessment behaviors. Contribution/Results: We find that while code generation is widely adopted, proficiency differentiation emerges more clearly in advanced tasks such as debugging and code review; documentation generation achieves the highest reliability, whereas complex logic implementation remains challenging. We propose a novel, empirically grounded “progressive workflow integration” paradigm—from isolated code generation toward deep, context-aware toolchain integration—and provide the first evidence-based characterization of developers’ iterative, multi-turn interaction preferences with generative AI in real-world settings, offering concrete empirical benchmarks and actionable design guidelines for AI-assisted development tools.
This study addresses the challenge of systematically integrating generative AI into the entire software development lifecycle to enhance productivity while ensuring quality and governance. The authors propose a progressive integration framework centered on an innovative “AI harness” that unifies management of project context, access control, validation, logging, and human approval workflows. This architecture enables seamless co-evolution of technical capabilities, organizational processes, and quality assurance mechanisms. The framework supports a transition from informal AI assistance toward controlled, agent-based development and is empirically validated through a case study in a mid-sized software enterprise, offering both a practical roadmap and evidence-based foundation for AI-driven transformation in software engineering.
In the AI era, developers face skill atrophy and diminished professional adaptability due to overreliance on generative AI. To address this, we conducted in-depth interviews with 21 cutting-edge AI practitioners and employed qualitative thematic coding alongside workplace empirical analysis. We systematically identified 12 categories of development objectives, 75 granular tasks, and their associated competency requirements. This study introduces—first in the literature—the four-dimensional “AI-Augmented Developer” competency framework (encompassing generative AI application, core software engineering, adjacent engineering domains, and non-engineering literacies) and a six-step task adaptation model. We further propose evidence-based educational interventions to mitigate skill degradation, yielding five key insights and actionable recommendations. These findings provide empirical grounding for university curriculum reform and enterprise AI literacy programs, advancing the systematic mapping and evolution of developer competencies.
Current AI assistant features in IDEs exhibit a significant misalignment with developers’ authentic needs, necessitating a systematic understanding of heterogeneous user requirements. Method: We conducted semi-structured interviews with 35 practitioners—comprising AI adopters, attriters, and non-users—to empirically construct the first human-AI interaction design space for IDE-integrated AI assistants. Through thematic coding and cross-cohort comparative analysis, we identified fundamental divergences across user groups along five dimensions: reliability, privacy, personalization, proactivity, and ethical concerns. Contribution/Results: We propose a role-driven, five-dimensional design framework—encompassing technical robustness, interaction modality, goal alignment, skill abstraction, and cognitive offloading—alongside 12 actionable design guidelines. This work advances IDE AI tools toward greater reliability, contextual awareness, privacy-by-design, and seamless workflow integration.
This study addresses the current lack of a systematic theoretical framework explaining how software professionals evaluate AI-generated code and the underlying cognitive processes and preferences involved. Employing a constructivist grounded theory approach, the research integrates questionnaire surveys, semi-structured interviews, and laddering interviews to iteratively collect data from 20 to 50 practitioners until theoretical saturation is achieved. The work presents the first empirically grounded theoretical framework for assessing AI-generated code, elucidating the evaluation mechanisms developers employ in human-AI collaborative programming contexts. By doing so, it fills a critical gap in the literature concerning both behavioral and cognitive dimensions of code evaluation in AI-assisted software development.
This study addresses the lack of systematic understanding regarding how developers actually employ generative AI tools—such as ChatGPT and GitHub Copilot—in open-source projects and the resulting implications for development workflows, licensing, and trustworthiness. By mining self-reported traces in GitHub commits, issues, and pull requests and complementing them with manual coding, the authors construct the first comprehensive taxonomy of generative AI usage, encompassing seven high-level categories and 64 specific tasks. The analysis further tracks the evolution of usage patterns over a two-year period, revealing a decline in early concerns, a substantial broadening of application scenarios, and consistently positive developer feedback on tool utility. These findings offer empirical grounding for both software engineering practices and the automated identification of AI-assisted development activities.
This work addresses the lack of systematic evaluation of large language models as intelligent IDE agents in realistic multilingual, full-stack development environments. We propose a Docker-based, IDE-native benchmarking framework that simulates authentic development workflows through structured tool interfaces—including code search, structured editing, and full-stack testing—and constructs 80 real-world tasks across eight private, unpublished codebases spanning C/C++, Java, and MERN stacks. These tasks encompass feature implementation, bug fixing, refactoring, and performance optimization. Our framework establishes, for the first time under contamination-free conditions, a systematic linkage between agent intent and project-level modification outcomes, thereby introducing an evaluation paradigm that closely mirrors real-world engineering collaboration and enables comprehensive, reliable assessment of AI-powered IDE agents.
This study investigates the transformative impact of generative artificial intelligence (GenAI) on human-computer interaction paradigms within integrated development environments (IDEs), along with its associated challenges and opportunities. Through a four-day interdisciplinary workshop at the Shonan Meeting 222, involving 33 experts from software engineering, artificial intelligence, and human-computer interaction, the research systematically identifies four core themes shaping the future evolution of IDEs under GenAI influence. The findings highlight GenAI’s potential to enhance developer productivity in tasks such as code generation, testing, review, and repair, while articulating key research challenges and practical opportunities. This work provides a foundational roadmap for advancing collaborative programming research in human-AI co-development contexts.
This study addresses a critical gap in existing research by shifting focus from the impact of generative AI on development efficiency and code quality to developers’ lived interaction experiences and subjective well-being in real-world settings. Employing a mixed-methods approach—integrating controlled experiments, naturalistic observation, log analysis, and in-depth interviews—the work systematically investigates how professional developers interact with tools such as GitHub Copilot. It proposes empirically grounded heuristics for selecting AI interaction modes based on task characteristics, revealing that concurrently using code suggestions and chat prompts can diminish effectiveness. The findings identify cognitive load and output quality as pivotal factors shaping developer experience, report high overall satisfaction—particularly for repetitive tasks—and demonstrate that interaction strategy efficacy varies by task type, while participation in the study heightened developers’ intentional use of AI tools.