EpiWorld: Grounding LLM Policy Agents in Epidemiological World Models

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of large language models (LLMs) in accurately evaluating intervention outcomes due to their lack of knowledge regarding epidemic dynamics and institutional constraints. To overcome this, we propose a closed-loop framework that embeds LLM agents within an epidemic world model and a hierarchical skill library. Specifically, this work introduces action-conditioned world models and reusable experience repositories, integrating post-hoc adaptive learning for counterfactual reasoning to optimize public health policies. The proposed end-to-end framework preserves decision interpretability while achieving the lowest peak error in the world model. Furthermore, it reduces cumulative hospitalization rates by up to 59%, significantly outperforming baseline approaches such as reinforcement learning.
📝 Abstract
Epidemic intervention policies are textual artefacts that human decision-makers interpret, justify, and revise through natural language, making large language models a natural candidate for epidemic policy reasoning. A naive LLM, however, lacks the epidemic dynamics needed to project intervention consequences, the quantitative surveillance signals required to assess severity, and the institutional constraints that define admissible actions. We present EpiWorld, a closed-loop framework that grounds an LLM policy actor in a learned action-conditioned epidemiological world model and a tiered skill library of public-health protocols, surveillance tools, and adaptive lessons accumulated through after-action analysis. Given a candidate intervention, the world model predicts regional epidemic evolution and enables fast counterfactual rollouts that provide feedback for policy selection and refinement. Outcomes of simulated futures are distilled into reusable lessons while protocol constraints remain fixed, allowing the decision process to improve without sacrificing interpretability or controllability. We evaluate both the world model and the end-to-end framework on retrospective COVID-19 and Influenza datasets: the world model achieves the best out-of-distribution Peak-MAE among all forecasting baselines, and the closed-loop framework reduces cumulative hospitalisation by up to 59% across datasets and by an average of ~16% across six LLM backbones, outperforming reinforcement-learning and optimal-control policy baselines.
Problem

Research questions and friction points this paper is trying to address.

Epidemiological policy reasoning
Large language models
Epidemic dynamics
Intervention consequences
Public health constraints
Innovation

Methods, ideas, or system contributions that make the work stand out.

Epidemiological World Models
LLM Policy Agents
Closed-loop Framework
Counterfactual Rollouts
Tiered Skill Library
🔎 Similar Papers
2024-10-06Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations)Citations: 13