π€ AI Summary
This study addresses the unreliability of large language model (LLM) agents in invoking scientific software and the absence of natural language interfaces for such tasks. To this end, we propose Epydemix, a framework that introduces the first auditable and reproducible interaction layer for epidemiological modeling tailored to AI agents. By integrating LLM agent techniques with open-source Python libraries, Epydemix employs a four-tier mechanism encompassing model discovery, validation, execution, and result verification, alongside declarative specifications, to achieve end-to-end automation from natural language to complex simulations without requiring custom code. Experimental evaluations across 50 sessions demonstrate that the proposed approach significantly reduces interaction rounds, token consumption, and computational costs while ensuring full-process traceability and result reproducibility.
π Abstract
Artificial Intelligence agents based on large language models provide convenient natural language interfaces to scientific software, but reliability is not automatic. Here we introduce the Epydemix Agent Framework, an additive layer over Epydemix, an open-source Python library for stochastic compartmental epidemic modeling. The framework extends the library with four capabilities to facilitate interaction with an AI agent: discovery of available models and parameters, preventive validation of a declarative scenario specification, execution through tested library code, and inspectability of results. These capabilities let an agent handle the entire modeling process, from the natural-language description of the scenario to quantitative results, figures, and interpretation of findings without writing custom code. Each step reads input files and saves results in a separate output bundle, making the process auditable and reproducible. First, we show the end-to-end workflow with a case study comparing vaccination strategies for a novel respiratory virus. Second, we assessed the framework across 50 agent sessions and five modeling tasks by comparing the agent use of the framework against the direct use of the Python interface. The framework reduced turns, output tokens, and cost on most tasks, unless it trades resources for per-point reproducibility.