LLMoxie: Exploring Agentic AI for Scientific Software Development

📅 2026-07-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
General-purpose AI coding agents often fail to meet the core requirements of scientific software—namely, citability, reproducibility, auditability, and extensibility—and frequently overlook scientific Python conventions, sensitive data handling, and traceable decision-making. This work proposes the first institutional-scale platform that systematically embeds Research Software Engineering (RSE) principles into an AI agent architecture through a three-tier design: a multi-cloud/on-premises inference layer; a control plane built on LiteLLM and MLflow integrating authentication, budgeting, PII redaction, and observability; and a plugin-agent-skill framework powered by an RSE-Plugins ecosystem. Deployed over 20 months at a university RSE center, the platform supports a six-stage scientific workflow and has successfully enabled research software development across astronomy, climate science, agriculture, and health domains, effectively addressing infrastructure, governance, and process challenges while establishing a reusable platform paradigm.
📝 Abstract
In this paper, we describe LLMoxie, an institutional AI platform whose three-tiered architecture supports multi-cloud and on-premise inference, a LiteLLM/MLflow control plane for authentication, budgeting, PII masking, and observability, and an application augmentation layer for AI coding agents. Layered on top, an open-source RSE-Plugins ecosystem encodes accumulated RSE knowledge as a Plugin-Agent-Skill hierarchy spanning scientific Python practice, domain-specific knowledge, a six-phase research-and-implement workflow, and project lifecycle management. Scientific software is judged less by raw code quality than by whether it can be cited, audited, reproduced, and extended. Off-the-shelf AI coding agents, optimized against commercial software benchmarks, are poorly calibrated for this setting: they ignore the conventions of the scientific Python libraries they invoke, mishandle sensitive or embargoed data, and leave decision trails that are difficult to reconstruct after the fact. We report on twenty months of practice at a university-based research software engineering (RSE) center, where RSEs embedded across astronomy, earth and climate science, agriculture, and health projects worked to close this gap. We characterize the recurring infrastructure, governance, and process challenges of adopting Agentic AI inside a multi-domain RSE center, describe the platform and plugin design, and distill operational lessons from real scientific software deployments. Together, the platform and plugins shift AI coding agents from generic code generators into domain-aware collaborators that respect community norms and produce auditable provenance of technical reasoning.
Problem

Research questions and friction points this paper is trying to address.

Agentic AI
Scientific Software Development
Research Software Engineering
AI Coding Agents
Reproducibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic AI
Scientific Software Development
RSE-Plugins
Provenance-aware AI
Domain-aware Coding Agents
L
Landung Setiawan
eScience Institute, University of Washington
A
Anant Mittal
eScience Institute, University of Washington
C
Cordero Core
eScience Institute, University of Washington
A
Anshul Tambay
eScience Institute, University of Washington
C
Carlos Garcia Jurado Suarez
eScience Institute, University of Washington
David A. C. Beck
David A. C. Beck
Assoc. Res. Prof. (Chem. Eng.) Dir. of Res. (eScience Institute), UW
A
Andrew J. Connolly
eScience Institute, Dept of Astronomy, University of Washington
V
Vani Mandava
eScience Institute, University of Washington