Score
Designs, implements, and analyzes computational simulations of interacting autonomous agents by specifying agent rules, heterogeneous and hybrid agent types, interaction protocols, and network or spatial environments to produce emergent population-level behavior. Builds experiment pipelines (Monte Carlo realizations, parameter sweeps, matched reference populations, and counterfactual interventions) to estimate outcome distributions, variances, and the effects of policy or design changes.
This study addresses a key limitation of traditional agent-based modeling (ABM)—its emphasis on simulation at the expense of experimental rigor, which hinders the identification of causal mechanisms in complex systems. To overcome this, the authors propose an integrated framework that combines computational experimentation with ABM, enabling systematic manipulation of input variables and counterfactual simulations. This approach constructs a “parallel world” capable of exploring multiple evolutionary trajectories, thereby transcending the constraints of conventional scenario analysis that relies heavily on subjective reasoning. The framework facilitates causal inference regarding the dynamic evolution of complex social systems, offers interpretable causal pathways to understand emergent phenomena, and establishes a theoretical foundation for computational experimentation in complex systems research.
Conventional physics-based simulation methods struggle to capture historical evolution, heterogeneity, and emergent phenomena inherent in social systems. Method: This study systematically reviews the development trajectory and design principles of agent-based modeling (ABM) in the social sciences. It introduces— for the first time—a tripartite classification of ABM social simulation paradigms: thought experiments, mechanism exploration, and parallel optimization. A generic three-layer modeling framework is proposed, integrating agents, environments, and interaction rules, alongside a taxonomy tailored for social simulator development. Contribution/Results: By synthesizing canonical case studies and core methodological challenges, the work establishes a theoretical foundation and practical guidelines for the standardization of ABM methodology and the systematic engineering of social simulators, thereby advancing rigorous, interpretable, and scalable computational social science.
Multi-agent systems (MAS) dynamically evolve under feedback, adaptation, and non-stationarity, yet lack a general adaptive modeling framework. Method: We propose an interpretable, learning-control-integrated MAS modeling framework that unifies multi-agent reinforcement learning, entropy rate and statistical complexity analysis, predictive information metrics, structural causal models (SCMs), and clustering-based emergent behavior identification; it partitions system dynamics into four regimes—stationary, oscillatory, drifting, and chaotic. Contribution/Results: This work is the first to embed information-theoretic diagnostics and causal inference directly into the adaptive learning pipeline, enabling prior-free agent initialization, unsupervised behavioral pattern discovery, and falsifiable policy evaluation. The framework provides formal definitions, computable operators, and standardized experimental templates, facilitating systematic cross-environment comparison across stability, performance, and interpretability in non-stationary settings.
This study investigates the spontaneous emergence of coordination, communication dynamics, and role differentiation within large-scale decentralized populations of AI agents. Deploying over 770,000 large language model–based agents in the MoltBook environment, we conducted a three-week longitudinal observation of more than 90,000 active individuals. Integrating network clustering, information cascade modeling, Cox survival analysis, and collaborative event detection, our analysis reveals that 93.5% of agents occupy homogeneous peripheral roles, information propagation follows a power-law distribution (α = 2.57), and the success rate of collaborative tasks is only 6.7%—significantly lower than that of a single-agent baseline (Cohen’s d = −0.88). These findings provide the first empirical evidence of emergent collective behavior in massive AI agent systems and establish a critical foundation for understanding both the mechanisms and limitations of artificial collective intelligence.
High-dimensional stochastic agent-based models (ABMs) are notoriously difficult to analyze systematically due to the curse of dimensionality and inherent stochasticity. This work proposes a multi-stage automated exploration framework that first employs model-driven experimental design to identify key variables and partition the parameter space, then leverages machine learning surrogate models to efficiently capture residual nonlinear interaction effects. The approach operates without human intervention, automatically detecting unstable regions within the simulator and enabling robust sensitivity analysis and policy testing. Applied to a predator–prey case study, the framework successfully isolates dominant variables and highly sensitive nonlinear regimes, substantially enhancing the efficiency and reliability of ABM exploration.
Current large language model (LLM)-based agents in social simulation are often mistakenly assumed to spontaneously reproduce authentic human collective behaviors, despite lacking a scientific foundation in behavioral validity and environmental interaction mechanisms. This work critically exposes the fundamental gap between role-playing plausibility and genuine human behavioral efficacy, reframing social simulation as a Markov game that explicitly incorporates environmental participation. It introduces explicit scheduling and information exposure mechanisms, emphasizing the pivotal roles of initial conditions, interaction protocols, and environmental dynamics in shaping emergent collective behavior. By doing so, the proposed framework systematically enhances the scientific rigor, auditability, and reproducibility of LLM-based multi-agent simulations, offering actionable design, evaluation, and interpretability guidelines for AI-driven social modeling.
Efficient simulation and analysis of emergent behaviors in large-scale multi-agent resource foraging remain challenging due to computational bottlenecks and lack of differentiability. Method: We propose the first JAX-based, fully vectorized, end-to-end differentiable multi-agent foraging framework. It enables parallel simulation of thousands of agents in a shared environment, unifying customizable dynamics, perception models, policies, and boundary conditions, while supporting runtime agent addition and removal. Contribution/Results: By tightly integrating vectorized programming, automatic differentiation, and hardware acceleration (e.g., GPUs/TPUs), our framework achieves real-time, differentiable simulation at the thousand-agent scale—unprecedented in prior work. This significantly improves modeling efficiency and enables gradient-driven analysis (e.g., policy optimization, sensitivity analysis). Beyond foraging applications, our framework advances the generality and scalability of differentiable multi-agent simulation, establishing new benchmarks for performance and flexibility in learned and physics-informed agent-based modeling.
This study addresses the challenge of effectively evaluating the impact of AI systems in knowledge work, which is hindered by traditional experimental methods that rely on unstructured textual descriptions lacking comparability, reusability, and auditability. To overcome this limitation, the authors propose the SEED framework, which formalizes human–AI collaborative experimental designs as typed participant–process graphs. This approach enables explicit representation of interaction structures, assessment of design novelty, and generation of feasible configurations under specified constraints. Integrating structured encoding, graph-guided generation, and lightweight validation, SEED significantly enhances process clarity, hypothesis specificity, and regulatory compliance in a medical triage task. The results demonstrate its effectiveness as a traceable, comparable, and generative tool for supporting rigorous experimental design in human–AI collaboration.
This study addresses the lack of systematic investigation into the design space of large language model (LLM)-based social simulations, which hinders the assessment of simulation fidelity. It reveals for the first time that this design space exhibits a nontrivial geometric structure. Through systematic analysis of key design choices—including base LLM type and agent connectivity patterns—and their interaction effects, the work identifies the base LLM as the dominant factor shaping simulation outcomes. While some parameters exert additive effects, others engage in complex interactions. Integrating LLM-based agent modeling, social network topology, and survey-validated opinion alignment among agents, this research provides a reproducible design framework for constructing realistic and trustworthy silicon-based societies.
This study investigates the integration of large language models (LLMs) into agent-based models (ABMs) to assess their impact on simulation fidelity, computational overhead, and agent behavior. Focusing on the classic Schelling segregation model, the authors design a hybrid agent architecture in which an LLM processes natural-language descriptions of neighborhood information via tool calls and integrates this interpretation into the agent’s original decision logic. The work presents the first controlled ABM framework that combines LLMs with statistical model checking (MultiVeStA), enabling quantitative evaluation of LLM-induced effects. Experimental results demonstrate that smaller LLMs exhibit insufficient reliability in semantic classification and tool invocation, whereas larger models successfully pass basic functional validation, thereby confirming the feasibility and potential of this approach.
Current AI agents lack a unified and effective capability for modeling environmental dynamics, hindering their reliable performance in complex physical, digital, social, and scientific settings. This work proposes a “Hierarchy × Laws” framework that structures a three-level world model system—comprising a predictor (L1), simulator (L2), and evolver (L3)—and integrates four categories of domain-specific laws. Synthesizing over 400 studies, the project introduces a novel taxonomy to unify world modeling concepts across disciplines, establishes decision-oriented evaluation principles, and provides a reproducible benchmark suite. By surveying more than 100 systems spanning reinforcement learning, video generation, GUI/Web agents, multi-agent simulation, and AI-driven scientific discovery, the study delineates methodological approaches, failure modes, and evaluation criteria for each configuration, offering both a theoretical foundation and a practical roadmap toward building advanced agents capable of simulating and reshaping their environments.
This study addresses the unclear causal mechanisms through which individual behavioral rules give rise to emergent collective phenomena. To this end, it proposes RePair, a framework that translates natural language rules into quantifiable generative agent behaviors, constructs simulated worlds via simulation calibration, and implements matching interventions to trace the association between behavioral trajectories and causal processes. This work contributes a systematic mapping from qualitative rules to quantitative collective effects while achieving convergence in cross-rule comparisons. Furthermore, it validates the efficacy of natural language-driven causal inference, offering a reliable and interpretable methodological guide for exploring complex social phenomena.