Twelve quick tips for designing AI-driven HPC workflows

📅 2026-06-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges AI-driven scientific workflows encounter in high-performance computing (HPC) environments—namely data intensity, resource heterogeneity, frequent iteration cycles, and I/O bottlenecks—which hinder their compatibility with conventional linear pipeline architectures. The work presents the first systematic integration of AI workflow characteristics with HPC system design, proposing a transformative framework tailored for adaptive intelligent computing environments and articulating twelve practical design principles. This framework incorporates key technologies including containerization, job array scheduling, explicit feedback mechanisms, heterogeneous resource management, and small-file I/O optimization, making it particularly well-suited for high-throughput domains such as computational biology. The resulting guidelines offer researchers actionable strategies to substantially enhance the efficiency, portability, and scalability of AI-HPC workflows.
📝 Abstract
High-performance computing (HPC) clusters remain the backbone of large-scale scientific computation, traditionally executing deterministic, linear pipelines optimised for predictable performance. However, the pervasive integration of artificial intelligence (AI) and foundation models into scientific research has introduced a fundamentally new computational paradigm. AI-driven workflows are characteristically iterative, data-driven, and probabilistic, introducing unique challenges regarding data gravity, heterogeneous resource management, and complex workflow orchestration. This guide provides twelve practical tips designed to help researchers design efficient, scalable, and reproducible AI-driven HPC workflows. By addressing critical system-level bottlenecks - such as containerisation for environment portability, strategic deployment of job arrays, explicit feedback loop mechanics, and I/O optimisation for small files - this article offers a framework for transitioning from rigid execution pipelines to adaptive, intelligent computational environments. While these architectural principles are broadly applicable across distributed environments, they are particularly tailored to the resource-intensive throughput demands of modern computational biology.
Problem

Research questions and friction points this paper is trying to address.

AI-driven workflows
High-performance computing
foundation models
workflow orchestration
heterogeneous resource management
Innovation

Methods, ideas, or system contributions that make the work stand out.

AI-driven workflows
HPC
containerisation
job arrays
I/O optimisation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jamie J. Alnasir
Department of Computer Science, Royal Holloway University of London