ASCEND: Personal AI Agents for Autonomous Scientific Computing Across HPC Clusters and GPU Workstations

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of cumbersome environment configuration and difficult fault diagnosis in traditional scientific computing by proposing a local AI agent framework based on large language models. Without requiring facility-level service support, the method autonomously schedules jobs across clusters via multiplexed connections using only account authentication. A local policy validator is integrated to ensure execution safety, thereby achieving an end-to-end automated closed loop. Experimental results demonstrate that the proposed agent can automatically recover from tensor device failures, reproduce meteorological model evaluations with errors below 2.5%, and reduce the runtime of geological solvers from twelve hours to two hours. These findings indicate that the framework significantly enhances both the efficiency and reliability of scientific computing workflows.
📝 Abstract
Traditional scientific computing requires researchers to translate computational intent into environment configuration, resource requests, and executable jobs, then diagnose failures from scheduler state and application logs. We present ASCEND (Autonomous Scientific Computing Engine and Novel Discovery), an AI-powered agent interface that runs the agent on the researcher's own laptop, reaching Slurm-managed clusters and a GPU workstation over a multiplexed authenticated connection, with site-specific execution policies checked by locally executed tools; the language model is hosted remotely and holds no credentials. No facility-scale service is required: an account on each resource is sufficient, and the public installer lets users link additional Slurm clusters or workstations of their own. We report four recorded cases: (1) the agent closed a failure-recovery loop on a planted tensor-device fault, submitting, diagnosing, repairing and resubmitting with job-level artifacts preserved; (2) it reproduced the published evaluation of a weather-forecasting model from the author's released forecasts, agreeing with the published curves to 2.1% (z500) and 2.4% (t850) while identifying a unit discrepancy in the paper's prose and an initialization-field discrepancy in its released data; (3) it parallelized a released 12,693-line geophysical solver under a bit-for-bit identity requirement, reducing wall-clock runtime from about twelve hours to about two; (4) that requirement exposed two instances of undefined behaviour in the published solver, both repaired and reported upstream. Separately, a pre-specified evaluation of the policy layer found the deployed validator rejected 29 of 30 constructed violations and held the remaining one for approval, while denying 3 of 14 legitimate requests. Autonomy was exercised under author supervision; an end-to-end recovery benchmark remains outstanding.
Problem

Research questions and friction points this paper is trying to address.

scientific computing
autonomous agents
HPC clusters
failure diagnosis
job management
Innovation

Methods, ideas, or system contributions that make the work stand out.

Autonomous Scientific Computing
AI Agents
HPC Clusters
Failure Recovery
Policy Validation
💼 Related Jobs
No related jobs found.