Score
Designs, implements, and evolves software tools and toolchains—including CLIs, language-specific libraries (e.g., Python), plugins, compiler toolchains, and AI-assisted or internal tooling—by specifying tool interfaces and APIs and building processing pipelines and automation. Integrates and validates these tools into developer workflows, deploys and maintains them, and engineers toolchain interoperability and tooling automation.
This study addresses the challenge of systematically integrating generative AI into the entire software development lifecycle to enhance productivity while ensuring quality and governance. The authors propose a progressive integration framework centered on an innovative “AI harness” that unifies management of project context, access control, validation, logging, and human approval workflows. This architecture enables seamless co-evolution of technical capabilities, organizational processes, and quality assurance mechanisms. The framework supports a transition from informal AI assistance toward controlled, agent-based development and is empirically validated through a case study in a mid-sized software enterprise, offering both a practical roadmap and evidence-based foundation for AI-driven transformation in software engineering.
This work addresses the limited accessibility of large language model (LLM) and agent workflow development for engineers without machine learning expertise, primarily due to the absence of integrated testing, debugging, and reproducibility capabilities. To bridge this gap, the authors propose a novel IDE-native AI observability workflow, implemented as the AI Toolkit plugin for JetBrains IDEs. This approach seamlessly embeds trace capture and evaluation into standard run/debug cycles, enabling automatic hierarchical trace logging during execution, one-click dataset persistence, and a pluggable, unit-test-like evaluation framework. By minimizing environment setup and context-switching overhead, the solution facilitates routine evaluation and immediate trace visualization. Empirical data from the initial PyCharm release demonstrates high adoption, sustained usage, and low churn, confirming that IDE-integrated tooling effectively lowers the barrier to entry for non-ML developers.
Implementation discrepancies across software repository mining tools severely threaten the validity of empirical findings. Method: We conduct a dual-tool comparative analysis of 10 large-scale open-source projects, systematically identifying how minor implementation differences—such as commit parsing logic and author deduplication rules—induce up to 500% deviation in key metrics (e.g., commit count, developer count). We propose a “tool-level configuration + post-hoc normalization” co-optimization framework to mitigate metric divergence and perform multi-tool experiments, quantitative consistency assessment, and code-level root-cause analysis. Contribution/Results: We identify six technical sources undermining data validity and establish the first validity assessment paradigm for Mining Software Projects Research (MSPR) explicitly addressing tool heterogeneity—thereby enabling rigorous, reproducible, and comparable empirical software engineering studies.
This study addresses the underexplored reliability issues of AI-powered coding tools in real-world engineering practice. Through a large-scale empirical analysis of over 3,800 public bug reports on GitHub concerning Claude-Code, Codex, and Gemini CLI, the authors employ open coding to synthesize issue descriptions, user discussions, and developer responses into a multidimensional defect taxonomy. The findings reveal that more than 67% of reported defects are functional in nature, with 36.9% stemming from API misuse, integration flaws, or configuration errors. The most prevalent symptoms include API errors (18.3%), terminal anomalies (14%), and command failures (12.7%), highlighting critical fragility points during tool invocation and execution. These insights provide empirical grounding and actionable design guidance for developing more robust AI coding assistants.
Current AI assistant features in IDEs exhibit a significant misalignment with developers’ authentic needs, necessitating a systematic understanding of heterogeneous user requirements. Method: We conducted semi-structured interviews with 35 practitioners—comprising AI adopters, attriters, and non-users—to empirically construct the first human-AI interaction design space for IDE-integrated AI assistants. Through thematic coding and cross-cohort comparative analysis, we identified fundamental divergences across user groups along five dimensions: reliability, privacy, personalization, proactivity, and ethical concerns. Contribution/Results: We propose a role-driven, five-dimensional design framework—encompassing technical robustness, interaction modality, goal alignment, skill abstraction, and cognitive offloading—alongside 12 actionable design guidelines. This work advances IDE AI tools toward greater reliability, contextual awareness, privacy-by-design, and seamless workflow integration.
This study addresses the underexplored nature of rules in AI-powered integrated development environments (IDEs) as an emerging class of software artifacts, whose taxonomy, evolution patterns, and practical impact remain poorly understood. Through a mixed-methods approach analyzing 7,310 rules from 83 open-source projects alongside survey data from 99 developers, this work proposes the first systematic classification framework comprising five top-level categories and 25 subcategories. It reveals a significant discrepancy between developer intent and actual rule configurations, demonstrates that rule evolution is frequent and primarily driven by contextual expansion, and quantifies a substantial improvement in software compliance—increasing from an average of 49.14% to 72.13% (+22.99%)—following rule updates.
This study addresses the lack of systematic understanding regarding how developers continuously use and evolve AI-generated code in real-world projects. By analyzing 35,361 GitHub code comments referencing AI and their associated 12,996 subsequent commits, the authors construct the first taxonomy of AI-assisted development activities. Integrating open coding, LLM-based dual-classifier annotation, Dawid-Skene aggregation, and longitudinal temporal analysis, they reveal a long-term evolutionary trend wherein AI tools shift from initial code generation toward knowledge support and code enhancement. The findings indicate that developers primarily employ large language models (LLMs) for implementation, debugging, and code augmentation, while subsequent commits predominantly involve refactoring and feature extension. Moreover, AI references increasingly reflect conceptual collaboration, suggesting that AI is becoming an embedded development partner.
This work addresses the challenges of IDE development posed by the rapid evolution of smart contract languages such as Move by presenting a high-performance IDE support system built atop the Move compiler and adhering to the Language Server Protocol (LSP). Through deep integration with existing language toolchains and the application of incremental parsing and optimized semantic analysis techniques, the system efficiently delivers rich IDE features even as the language undergoes continuous iteration. Deployed successfully within the Sui platform’s Move ecosystem, it significantly enhances developer experience and yields a reusable, evolution-aware IDE construction strategy applicable to other emerging programming language ecosystems.
Existing approaches rely on tool-level graph representations of historical trajectories, which struggle to generalize to new tool sets and thereby limit the planning capabilities of large language models. To address this, this work proposes a Functional-level Workflow Graph (FWG) that abstracts tool-specific behaviors into functional-level workflows through trajectory uplifting, effectively decoupling workflow planning from tool selection. The framework incorporates a source-gating mechanism and skill-specific rewards, combined with reinforcement learning, to ensure reliable and traceable data flows. Evaluated on two in-distribution and three out-of-distribution benchmarks, the method significantly outperforms current state-of-the-art approaches and demonstrates strong cross-domain generalization to unseen tool sets.