Grounding Large Language Models in DSGE Simulators for Policy Generation and Forecasting

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of dynamic consistency verification and long-horizon credit assignment mechanisms when large language models (LLMs) formulate economic policies. To overcome these limitations, this work proposes a closed-loop interactive framework that embeds instruction-tuned LLMs within a Dynamic Stochastic General Equilibrium (DSGE) simulator. Policy actions are optimized using the Proximal Policy Optimization (PPO) reinforcement learning algorithm, thereby establishing an evaluation paradigm grounded in economic consequences rather than textual plausibility. This approach effectively resolves the challenge of delayed reward propagation, enabling dynamic simulation and rigorous assessment of policy effects under historical shocks. Ultimately, this research provides a quantifiable and verifiable pathway for AI-driven macroeconomic decision-making.
📝 Abstract
Large language models can produce economic policy responses that sound reasonable, but this does not show that their actions are consistent with economic dynamics. We test this by placing an instruction-tuned language model inside six Snowdrop-backed dynamic stochastic general equilibrium (DSGE) simulators. At each turn, the model observes the economy and a change in economic discourse, selects a bounded policy action, and receives the next simulated state and an economic reward. We implement a common Python interface for repeated rollouts, persistent shocks, state cloning, and rolling-horizon simulation. This setting creates a long-horizon credit-assignment problem. Policy effects may appear several quarters after an action is taken. PPO has a learned value function that can propagate delayed reward to earlier tokens through generalized advantage estimation. GRPO has no learned value function and instead assigns a group-relative advantage from complete rollout returns. It therefore cannot distinguish which earlier turn caused the outcome; if every rollout receives the same return, the normalized advantage is zero. We use PPO as the primary method and GRPO as a matched critic-free baseline. The experiments also test directional semantic signals, reward horizon, trajectory warm starts, cross-simulator transfer, and historically anchored pandemic and monetary-policy shocks. The objective is to judge policy actions by their simulated economic consequences rather than by plausible language alone.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
DSGE Simulators
Economic Policy Generation
Long-horizon Credit Assignment
Policy Forecasting
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
DSGE Simulators
Reinforcement Learning
Credit Assignment
Policy Generation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Aditya Dubey
Birla Institute of Technology and Science, Pilani, India
N
Namah Gupta
Birla Institute of Technology and Science, Pilani, India
Vinti Agarwal
Vinti Agarwal
Birla Institute of Science and Technology, Pilani, India
Machine LearningSemi-supervised learningGraph deep learningSocial Recommender SystemsData