Score
Design and implement agent-based simulations where individual agents are implemented as or controlled by large language models (and optionally hybrid LLM + rule-based controllers), including the environment, interaction protocols, and agent behavior modules to simulate communication, decision-making, and resource interactions. Build measurement and analysis pipelines to validate agent behavior and quantify emergent system-level outcomes such as cooperation, resource distribution, reputation dynamics, and conflict.
Generative agent-based modeling (ABM) powered by large language models (LLMs) promises to address ABM’s longstanding challenges—limited realism, weak empirical validation, and opaque causal mechanisms—but risks exacerbating scientific rigor deficits. Method: Through critical literature analysis, methodological reflection, and integrative paradigm diagnosis, this study systematically evaluates the epistemic suitability of LLM-ABM for social simulation. Contribution/Results: We find that current LLM-ABM implementations routinely neglect core ABM methodological principles and lack rigorous validation protocols. LLMs’ inherent opacity intensifies the explanatory and verifiability crisis, rendering subjective “trustworthiness” assessments inadequate substitutes for operational validity testing. Crucially, the paper demonstrates—not merely identifies—that LLM integration *amplifies*, rather than alleviates, ABM’s scientific credibility gap. It thereby challenges LLM-ABM’s capacity to advance social-scientific theory construction and provides foundational methodological guardrails and normative design principles for future generative ABM research.
Non-technical users struggle to intuitively operate complex simulation systems, while existing large language models (LLMs) lack grounding in real-world dynamic constraints, leading to physically implausible or causally inconsistent outputs. Method: We propose the first bidirectional collaborative framework integrating simulation systems and LLMs—enabling natural-language-driven simulation execution while constraining LLM reasoning with causally accurate, structured, real-time simulation states. Our approach innovatively combines prompt-engineering-driven LLM-Simulation API orchestration, dynamic knowledge grounding, and a causal-aware state-mapping interface. Contribution/Results: Evaluated across multi-domain decision-making tasks, the framework significantly improves answer accuracy (+32%) and operational success rate (+41%). It enables zero-code invocation of high-fidelity simulations, achieving a principled balance among interpretability, usability, and physical consistency.
This paper addresses the core challenges and integration pathways for incorporating large language models (LLMs) into agent-based social simulation (ABSS). Recognizing LLMs’ limitations in theory of mind and social inference—key cognitive modeling capabilities—it proposes a “rules + LLM” hybrid architecture: embedding LLMs within established simulation platforms (e.g., GAMA, NetLogo) to enhance agent behavioral expressivity and social interaction fidelity, while retaining rule-based modules to ensure transparency, interpretability, and reproducibility. The study systematically evaluates state-of-the-art implementations—including Smallville and AgentSociety—to delineate LLMs’ appropriate use cases in generative social interaction and their constraints in predictive modeling. Its primary contribution is a methodological framework for LLM-augmented ABSS that jointly optimizes behavioral fidelity, explainability, and reproducibility. Furthermore, it provides empirical evidence and design principles for robust LLM integration in social computing.
本文提出一个多代理框架,使大型语言模型能够通过科学模拟模型进行受控实验,以优化制药过程设计,提高输出的具体性和实用性。
Current large language model (LLM)-based agents in social simulation are often mistakenly assumed to spontaneously reproduce authentic human collective behaviors, despite lacking a scientific foundation in behavioral validity and environmental interaction mechanisms. This work critically exposes the fundamental gap between role-playing plausibility and genuine human behavioral efficacy, reframing social simulation as a Markov game that explicitly incorporates environmental participation. It introduces explicit scheduling and information exposure mechanisms, emphasizing the pivotal roles of initial conditions, interaction protocols, and environmental dynamics in shaping emergent collective behavior. By doing so, the proposed framework systematically enhances the scientific rigor, auditability, and reproducibility of LLM-based multi-agent simulations, offering actionable design, evaluation, and interpretability guidelines for AI-driven social modeling.
This study investigates the integration of large language models (LLMs) into agent-based models (ABMs) to assess their impact on simulation fidelity, computational overhead, and agent behavior. Focusing on the classic Schelling segregation model, the authors design a hybrid agent architecture in which an LLM processes natural-language descriptions of neighborhood information via tool calls and integrates this interpretation into the agent’s original decision logic. The work presents the first controlled ABM framework that combines LLMs with statistical model checking (MultiVeStA), enabling quantitative evaluation of LLM-induced effects. Experimental results demonstrate that smaller LLMs exhibit insufficient reliability in semantic classification and tool invocation, whereas larger models successfully pass basic functional validation, thereby confirming the feasibility and potential of this approach.
Social simulation in policy-making suffers from limited credibility and trust across domains. Method: This study introduces a verifiable human-AI co-modeling paradigm grounded in large language model (LLM) agents, developed through a year-long iterative collaboration with a university emergency response team. The system comprises 13,000 LLM agents simulating crowd mobility and information diffusion dynamics during large-scale event emergencies. Contribution/Results: We propose three design principles: (1) initiating modeling with high-fidelity, verifiable scenarios to establish cross-disciplinary trust; (2) using preliminary simulations to elicit domain experts’ tacit knowledge; and (3) treating simulation development and policy refinement as a co-evolutionary process. The framework has directly informed operational improvements—including volunteer training optimization, evacuation protocol revision, and infrastructure layout adjustment—demonstrating end-to-end validation from simulation to policy implementation.
Current research on large language model (LLM) multi-agent systems lacks a unified theoretical framework, making it difficult to systematically characterize their social interactions and strategic behaviors. This work addresses this gap by introducing game theory as a foundational lens, constructing an interpretable, comparable, and scalable analytical and design framework centered on four core elements: players, strategies, payoffs, and information. Through game-theoretic modeling, comprehensive literature review, and integrative framework synthesis, the study presents the first unified taxonomy specifically tailored for LLM multi-agent systems. The proposed framework not only offers a structured basis for classifying existing approaches but also provides essential theoretical grounding for the future design, evaluation, and development of such systems, effectively bridging the fragmented landscape in this emerging field.
Although large language models exhibit strong linguistic capabilities, they lack a foundational understanding of social dynamics, making it difficult to generate socially intelligible behaviors that align with roles, norms, and contextual constraints. This work proposes, for the first time, a systematic framework in which character-driven role modeling serves as the core mechanism for building socially intelligent agents. By formalizing roles as structured persona descriptions, the approach enables a principled transformation from linguistic competence to contextually appropriate social behavior. The study introduces a hybrid control architecture that integrates large language models with theories of social roles and persona-based modeling, and outlines a research roadmap encompassing representation, control, and evaluation. This framework establishes a foundational theoretical and methodological baseline for advancing research on socially capable artificial agents.
This study presents the first systematic evaluation of the scientific validity of large language models (LLMs) in automatically generating reproducible and verifiable agent-based model code from ODD protocol specifications. Using the PPHPC predator–prey model as a benchmark, the authors assess Python implementations produced by 17 LLMs across four dimensions: executability, behavioral consistency, computational efficiency, and code maintainability. Results indicate that GPT-4.1 consistently generates statistically valid and efficient code, with Claude 3.7 Sonnet performing comparably but exhibiting lower stability. Crucially, the study demonstrates that executability does not guarantee behavioral fidelity, underscoring the essential role of formal verification in scientific modeling. While these findings reveal the potential of LLMs in specification-driven modeling, they also affirm that current models cannot yet replace human expertise in rigorous scientific model development.
Existing social simulation approaches face a dual bottleneck: rule-based agents lack semantic understanding, while LLM-based agents incur prohibitive computational overhead. This paper proposes a hybrid agent framework integrating large language models (LLMs) and diffusion models, featuring a modular dual-agent architecture—where the LLM agent performs semantic content parsing and reasoning, and the diffusion-model agent efficiently captures user-state evolution and social influence propagation. The framework jointly incorporates personalized user modeling, quantified social influence, and content-aware mechanisms to enable scalable, multi-factor-coupled simulation. Experiments on three real-world social datasets demonstrate substantial improvements in information diffusion prediction accuracy, achieving a favorable trade-off between precision and efficiency. Results validate the effectiveness and generalizability of the hybrid architecture for large-scale social simulation.