Score
Designs, builds, and evaluates systems and tools that use AI/ML models to assist software development tasks—including code completion, code generation, debugging, automated code review, editing and documentation, developer-facing coding agents, and automation of build/test workflows. Also develops and analyzes integrations with IDEs and CI pipelines, model and prompt pipelines, security and correctness mitigations for model outputs, and evaluation methods for functional accuracy, safety, and developer experience.
This study addresses critical challenges in human-AI collaboration for software development—including low co-development efficiency, insufficient trust, and weak perceived control. To tackle these issues, we propose the first taxonomy of developer-AI interactions spanning the entire software engineering lifecycle. Grounded in empirical analysis and consensus among domain experts, the taxonomy systematically classifies eleven distinct interaction patterns, including auto-completion, instruction-driven programming, and conversational assistance. Unlike prior fragmented characterizations, this structured framework establishes a foundational paradigm for AI tool design, human factors evaluation, and research on trustworthy collaborative mechanisms. The taxonomy enables principled, theory-guided optimization of adaptive and reliable programming assistants—shifting human-AI software development from empirically driven practice toward rigorous, evidence-based methodology.
This study addresses the lack of a systematic understanding of generative artificial intelligence’s role across the full software development lifecycle. Through a systematic literature review complemented by structured surveys of 65 developers, this work integrates empirical data with existing research to comprehensively evaluate the real-world impact and adoption patterns of large language models (LLMs) in each development phase. Findings indicate that over 70% of developers save more than 50% of their time on boilerplate code generation and documentation tasks, and 79% use browser-based LLMs daily. While nascent governance mechanisms are emerging, benefits in early-stage activities—such as requirements elicitation and architectural design—remain limited. The results suggest that generative AI is shifting the locus of development value from coding toward upstream design activities.
Current IDEs lack intelligent, end-to-end support for the machine learning (ML) lifecycle, while MLOps platforms remain decoupled from coding environments. To bridge this gap, we propose a novel large language model (LLM)-enhanced intelligent IDE paradigm that deeply integrates LLMs into the development environment. This enables synergistic, closed-loop automation across code-level intelligent programming—such as code generation, debugging, and completion—and full-stack MLOps pipeline orchestration—including data validation, feature store management, data drift detection, retraining triggers, and CI/CD deployment. The system unifies development, experimentation, validation, and monitoring phases, significantly improving engineering efficiency and reproducibility. Empirical evaluation on the UCI Adult and M5 datasets demonstrates a 61% reduction in pipeline configuration time, a 45% improvement in experimental reproducibility, and a 14% increase in data drift detection accuracy.
This study investigates the applicability of AI-powered programming assistants in real-world enterprise software projects and their impact on software engineering workflows and developer experience. Drawing on a survey of 57 developers with diverse backgrounds, a systematic review of 35 existing user studies, and semi-structured interviews complemented by literature analysis, this work presents the first empirical investigation—combining qualitative and quantitative methods—specifically focused on enterprise-level deployment contexts. The research formulates a core requirements framework for AI programming assistants grounded in practical engineering needs, uncovering critical challenges in current deployments and articulating key user expectations. These findings offer empirically grounded insights and clear guidance for the design of developer tools and the optimization of underlying AI models.
This work addresses the limited accessibility of large language model (LLM) and agent workflow development for engineers without machine learning expertise, primarily due to the absence of integrated testing, debugging, and reproducibility capabilities. To bridge this gap, the authors propose a novel IDE-native AI observability workflow, implemented as the AI Toolkit plugin for JetBrains IDEs. This approach seamlessly embeds trace capture and evaluation into standard run/debug cycles, enabling automatic hierarchical trace logging during execution, one-click dataset persistence, and a pluggable, unit-test-like evaluation framework. By minimizing environment setup and context-switching overhead, the solution facilitates routine evaluation and immediate trace visualization. Empirical data from the initial PyCharm release demonstrates high adoption, sustained usage, and low churn, confirming that IDE-integrated tooling effectively lowers the barrier to entry for non-ML developers.
Traditional code review relies heavily on static rule-based checks, lacking proactive risk prediction and pedagogical support. To address this, we propose an intelligent code review agent powered by large language models (LLMs). Our method fine-tunes LLMs on heterogeneous code-related data—including vast codebases, historical review comments, defect reports, and best-practice documentation—and integrates deep code semantic understanding with developer sentiment analysis. The agent performs code smell detection, anticipatory defect identification, actionable improvement suggestions, and educational feedback. Our key contributions are twofold: (1) the first application of LLMs to *proactive* code risk forecasting—moving beyond retrospective rule matching—and (2) an empirically grounded human-AI collaborative evaluation framework. Experimental results demonstrate a statistically significant reduction in post-release defect density, improved review throughput, and high developer acceptance, validating the agent’s dual efficacy in enhancing software quality assurance and fostering engineering skill development.
Addressing the acute shortage of AI/ML expertise in software engineering (SE), this study investigates the effectiveness and adoption barriers of AutoML for SE decision-making. Method: We systematically benchmark 12 state-of-the-art AutoML tools (e.g., H2O, Auto-sklearn, TPOT) on SE datasets and complement quantitative evaluation with surveys and expert interviews. Contribution/Results: Our empirical analysis reveals that AutoML-generated models achieve significantly higher average accuracy than manually tuned models on SE classification tasks. However, 83% of the tools lack automated feature engineering and deployment capabilities, and provide insufficient workflow support for non-ML experts—exposing a critical “pseudo end-to-end” limitation. The study identifies structural gaps in full-lifecycle automation and cross-role collaboration within current AutoML systems, thereby providing evidence-based insights and concrete design directions for next-generation AutoML tailored to SE contexts.
This study addresses the challenge of systematically integrating generative AI into the entire software development lifecycle to enhance productivity while ensuring quality and governance. The authors propose a progressive integration framework centered on an innovative “AI harness” that unifies management of project context, access control, validation, logging, and human approval workflows. This architecture enables seamless co-evolution of technical capabilities, organizational processes, and quality assurance mechanisms. The framework supports a transition from informal AI assistance toward controlled, agent-based development and is empirically validated through a case study in a mid-sized software enterprise, offering both a practical roadmap and evidence-based foundation for AI-driven transformation in software engineering.
This study addresses the challenges of frequent requirement changes, quality assurance, and delivery efficiency in agile software development by systematically investigating the application of artificial intelligence—particularly machine learning and natural language processing—in critical phases such as requirements management, code generation, and testing. Through a comprehensive literature review and empirical research involving industry practitioners, the work demonstrates that AI not only enhances existing agile practices but also fundamentally reshapes software development paradigms by introducing automation and intelligent decision-making capabilities. This transformation significantly improves development efficiency, product quality, and team responsiveness, thereby enabling a synergistic advancement in quality, speed, and innovation.
This study addresses the persistent challenges of inefficiency and inconsistent design fidelity that developers encounter when translating high-fidelity mockups into production-grade user interfaces. Through controlled experiments conducted across Angular, iOS, and Android platforms in an industrial setting, the work presents the first empirical evaluation of an AI-assisted development tool integrated with a design system. The findings demonstrate that this approach substantially enhances both development efficiency and design consistency: delivery time was reduced by 46.7%–69.4%, task completion rates improved, performance variability decreased, and workflow friction was markedly alleviated. These results validate the synergistic value of design-system-aware AI tools in enabling automation and standardization across multi-platform front-end development workflows.
This study addresses the lack of systematic understanding regarding how developers continuously use and evolve AI-generated code in real-world projects. By analyzing 35,361 GitHub code comments referencing AI and their associated 12,996 subsequent commits, the authors construct the first taxonomy of AI-assisted development activities. Integrating open coding, LLM-based dual-classifier annotation, Dawid-Skene aggregation, and longitudinal temporal analysis, they reveal a long-term evolutionary trend wherein AI tools shift from initial code generation toward knowledge support and code enhancement. The findings indicate that developers primarily employ large language models (LLMs) for implementation, debugging, and code augmentation, while subsequent commits predominantly involve refactoring and feature extension. Moreover, AI references increasingly reflect conceptual collaboration, suggesting that AI is becoming an embedded development partner.