🤖 AI Summary
Generative agent-based modeling (ABM) powered by large language models (LLMs) promises to address ABM’s longstanding challenges—limited realism, weak empirical validation, and opaque causal mechanisms—but risks exacerbating scientific rigor deficits. Method: Through critical literature analysis, methodological reflection, and integrative paradigm diagnosis, this study systematically evaluates the epistemic suitability of LLM-ABM for social simulation. Contribution/Results: We find that current LLM-ABM implementations routinely neglect core ABM methodological principles and lack rigorous validation protocols. LLMs’ inherent opacity intensifies the explanatory and verifiability crisis, rendering subjective “trustworthiness” assessments inadequate substitutes for operational validity testing. Crucially, the paper demonstrates—not merely identifies—that LLM integration *amplifies*, rather than alleviates, ABM’s scientific credibility gap. It thereby challenges LLM-ABM’s capacity to advance social-scientific theory construction and provides foundational methodological guardrails and normative design principles for future generative ABM research.
📝 Abstract
Recent advancements in AI have reinvigorated Agent-Based Models (ABMs), as the integration of Large Language Models (LLMs) has led to the emergence of ``generative ABMs'' as a novel approach to simulating social systems. While ABMs offer means to bridge micro-level interactions with macro-level patterns, they have long faced criticisms from social scientists, pointing to e.g., lack of realism, computational complexity, and challenges of calibrating and validating against empirical data. This paper reviews the generative ABM literature to assess how this new approach adequately addresses these long-standing criticisms. Our findings show that studies show limited awareness of historical debates. Validation remains poorly addressed, with many studies relying solely on subjective assessments of model `believability', and even the most rigorous validation failing to adequately evidence operational validity. We argue that there are reasons to believe that LLMs will exacerbate rather than resolve the long-standing challenges of ABMs. The black-box nature of LLMs moreover limit their usefulness for disentangling complex emergent causal mechanisms. While generative ABMs are still in a stage of early experimentation, these findings question of whether and how the field can transition to the type of rigorous modeling needed to contribute to social scientific theory.