Score
Designs and implements the infrastructure, integrations, and processes required to operationalize AI and generative‑AI models and end‑to‑end AI workflows—including deployment pipelines, monitoring, logging, scaling, model/version management, retraining, and incident response. Builds cross‑team rollout and adoption plans, automation and tooling to integrate AI capabilities with existing systems, and instrumentation to measure runtime performance, reliability, cost, and compliance for ongoing operations.
This paper identifies a core dilemma in organizational responsible AI governance: ambiguous responsibility boundaries across AI lifecycle stages and a lack of role- and stage-appropriate operational tools. Methodologically, the study systematically reviews over 220 responsible AI tools and proposes a novel two-dimensional (Actor, Stage) classification framework, integrating systematic review, meta-analysis, and qualitative coding. It identifies three critical governance gaps: (1) unclear accountability attribution, (2) absence of empirical validation for most tools, and (3) severe coverage imbalance across actors and stages. Results show that >80% of tools target developers during data and modeling phases; tools for leadership, deployers, end users, and stages such as value proposition definition and deployment are virtually absent. Moreover, >90% of tools lack empirical evidence. The study establishes a theoretically grounded, empirically benchmarked framework to advance actor–stage–aligned AI governance tool ecosystems.
Production-grade autonomous AI workflows face significant engineering challenges in reliability, observability, maintainability, and security governance. Method: We propose a structured, full-lifecycle methodology comprising a multi-agent architecture with collaborative reasoning, tool augmentation, and dynamic orchestration—integrated with the Model Context Protocol (MCP), deterministic orchestration, pure function invocation, containerized deployment, and modular tool integration. We further define nine core engineering practices, including tool-first design, single-responsibility agents, externalized prompt management, and model-federation-driven responsible AI design. Contribution/Results: This work establishes the first systematic engineering paradigm for Agentic AI productionization, markedly improving system simplicity, observability, and governability. Empirical validation via a multimodal news analysis–media generation use case demonstrates robustness and scalability. The methodology provides a reusable framework and practical benchmark for industrial-scale autonomous AI systems.
This work addresses the challenge that existing AI systems struggle to dynamically observe, intervene in, and optimize agent behavior at runtime, making it difficult to simultaneously achieve high task success rates, low latency, token efficiency, reliability, and safety. To overcome this limitation, the paper proposes a novel runtime infrastructure layer situated between the model and the application, which treats the AI execution process itself as an optimizable object—departing from conventional approaches that restrict optimization to static model or log-level adjustments. This layer enables proactive intervention and multi-dimensional performance co-optimization through mechanisms such as runtime monitoring, real-time inference, adaptive memory management, fault recovery, and policy enforcement. Experimental results demonstrate that the proposed approach significantly enhances the holistic performance of long-horizon agent workflows across task success rate, response latency, token efficiency, system reliability, and safety compliance.
This study addresses the high manual overhead faced by DevOps teams in managing multi-interface cloud infrastructures. We propose and systematically evaluate an LLM-driven AI agent framework for automation. Methodologically, the agent unifies heterogeneous interfaces—including SDKs, CLIs, Infrastructure-as-Code (IaC) tools, and web portals—to support core tasks such as configuration deployment, monitoring/alerting, and incident remediation. Key contributions include: (1) the first evaluation framework specifically designed for AI agents in cloud infrastructure management; (2) identification and systematic mitigation of three critical bottlenecks—interface semantic gaps, action execution reliability, and security constraint compliance; and (3) domain-specific optimization strategies validated in real-world deployments, demonstrating both task feasibility and cross-scenario generalizability. Our work establishes a reusable methodology and empirical benchmark for AI-native cloud operations.
This study addresses the challenge of systematically integrating generative AI into the entire software development lifecycle to enhance productivity while ensuring quality and governance. The authors propose a progressive integration framework centered on an innovative “AI harness” that unifies management of project context, access control, validation, logging, and human approval workflows. This architecture enables seamless co-evolution of technical capabilities, organizational processes, and quality assurance mechanisms. The framework supports a transition from informal AI assistance toward controlled, agent-based development and is empirically validated through a case study in a mid-sized software enterprise, offering both a practical roadmap and evidence-based foundation for AI-driven transformation in software engineering.
To address the escalating computational demands, high costs, and cross-environment coordination challenges in large language model (LLM) training, this project develops an end-to-end hybrid-cloud AI infrastructure comprising the cloud-based multi-tenant supercomputing platform Vela and the on-premises ultra-large-scale training system Blue Vela. It introduces a novel cloud-edge collaborative dynamic resource scheduling paradigm, integrating AI-optimized hardware clusters, a full-stack software–hardware co-designed training framework, and a unified telemetry and elastic orchestration system. Compared to conventional approaches, the infrastructure achieves a 35% improvement in training efficiency at the thousand-GPU scale and reduces fault recovery time by 60%. It has successfully accelerated iterative development of IBM’s third-generation and beyond generative AI models, while delivering commercial inference services with millisecond-level latency and 99.99% system availability.
This study addresses the pervasive “capability–deployment validation gap” that impedes the real-world adoption of advanced AI agent systems in industry. Through in-depth interviews with 16 practitioners across 12 enterprises of varying scales and domains, and by applying a six-level AI maturity framework, the research systematically assesses current agent adoption practices. It reveals, for the first time, that this gap stems primarily from information asymmetry and a lack of organizational readiness. Key technical barriers include large language models’ context limitations, non-deterministic behavior, insufficient support for proprietary languages, and data confidentiality constraints. Findings indicate most organizations operate at Level 1 (AI assistant) or Level 2 (AI compensator), with only one reaching Level 3 (multi-agent orchestration); notably, four firms could not achieve production deployment due to the absence of output validation mechanisms.
This study addresses the prevailing gap in AI education, which emphasizes model development while neglecting system engineering practices, leaving students ill-equipped to handle real-world challenges such as architectural design, deployment, and monitoring. To bridge this gap, the authors implemented a master’s-level course in which students built a movie recommendation system under realistic constraints, with a focus on integrating AI components into robust software systems, adopting data-driven machine learning practices, and cultivating systems-level thinking. Using a mixed-methods approach—combining analysis of student project artifacts with survey data—the research evaluates learners’ performance in architectural decision-making, integration of heterogeneous models, and adaptation to evolving requirements. Findings reveal common difficulties students encounter in AI system engineering and demonstrate the course’s effectiveness in addressing critical deficiencies in AI engineering education and enhancing systems-aware competencies.
This study addresses the inadequacy of the current U.S. Department of Defense software acquisition pathways in effectively managing the unique challenges posed by artificial intelligence systems—particularly their data dynamism, model evolution, and governance requirements. Through scenario-based policy analysis, the authors embed a hypothetical AI-enabled project into critical junctures of the existing acquisition process to systematically evaluate how policies translate into practice. The analysis reveals that core guidance documents lack operational specificity, while AI-related controls are fragmented across supplementary materials, leading programs to rely on inconsistent local interpretations. To bridge this gap, the paper proposes a dedicated AI acquisition sub-pathway alongside targeted documentation enhancements, substantially aligning policy with practice in areas such as data provenance, lifecycle management, and human oversight.
This work addresses the challenges of constructing and managing generative AI agent systems for long-horizon, stateful, multi-step business processes by proposing a graph-structured workflow design methodology. Leveraging the LangGraph framework, it explicitly models core mechanisms such as state management, conditional routing, and human-in-the-loop interventions. The approach is instantiated in three representative applications: SQL analysis with repair loops, retrieval-augmented generation gated by evidential validation, and human-AI collaborative policy review supporting interruption and checkpoint-based recovery. By treating behaviors like routing, pausing, and audit trails as explicit product features rather than implicit prompt logic, this study not only delineates the applicability boundaries of LangGraph in high-complexity workflows but also substantially enhances system controllability, reliability, and auditability in real-world operational settings, establishing a reusable engineering paradigm.