Operational Validation of Large-Language-Model Agent Social Simulation: Evidence from Voat v/technology

šŸ“… 2025-08-29
šŸ“ˆ Citations: 0
✨ Influential: 0
šŸ“„ PDF
šŸ¤– AI Summary
This study addresses the challenge of empirically validating the operational fidelity of LLM-driven generative social simulations within extremist online ecosystems—specifically, the alt-right platform Voat. We introduce YSocial, the first fine-grained simulation framework tailored to far-right technical communities, built upon Dolphin 3.0 (Llama 3.1 8B). YSocial integrates agent-level user profiles—including political orientation, educational background—and platform-specific norms to model posting, replying, reacting, and diurnal behavioral rhythms. Over a 30-day simulation, it successfully reproduces key empirical dynamics of the target community: heavy-tailed participation distributions, core-periphery network topology, AI/Big Tech–centric topic concentration, and elevated toxicity levels. Critically, this work provides the first empirical demonstration that LLM-based agents can faithfully emulate complex cultural norms, interaction tempo, and toxicity propagation mechanisms. It establishes a novel, controllable experimental paradigm for evaluating content governance interventions in ideologically charged digital environments.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Cognitive Modeling & Cognitive Systems: Social Cognition And InteractionMultiagent Systems: Agent-Based Simulation and Emergent Behavior

Application Category

Social Networks and Social Media: Generative AI / large language models and their impact on social systemsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
šŸ“ Abstract
Large Language Models (LLMs) enable generative social simulations that can capture culturally informed, norm-guided interaction on online social platforms. We build a technology community simulation modeled on Voat, a Reddit-like alt-right news aggregator and discussion platform active from 2014 to 2020. Using the YSocial framework, we seed the simulation with a fixed catalog of technology links sampled from Voat's shared URLs (covering 30+ domains) and calibrate parameters to Voat's v/technology using samples from the MADOC dataset. Agents use a base, uncensored model (Dolphin 3.0, based on Llama 3.1 8B) and concise personas (demographics, political leaning, interests, education, toxicity propensity) to generate posts, replies, and reactions under platform rules for link and text submissions, threaded replies and daily activity cycles. We run a 30-day simulation and evaluate operational validity by comparing distributions and structures with matched Voat data: activity patterns, interaction networks, toxicity, and topic coverage. Results indicate familiar online regularities: similar activity rhythms, heavy-tailed participation, sparse low-clustering interaction networks, core-periphery structure, topical alignment with Voat, and elevated toxicity. Limitations of the current study include the stateless agent design and evaluation based on a single 30-day run, which constrains external validity and variance estimates. The simulation generates realistic discussions, often featuring toxic language, primarily centered on technology topics such as Big Tech and AI. This approach offers a valuable method for examining toxicity dynamics and testing moderation strategies within a controlled environment.
Problem

Research questions and friction points this paper is trying to address.

Validating LLM agent social simulation for online communities
Assessing operational validity of simulated toxic discussions
Testing moderation strategies in controlled social environments
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-based generative social simulation
Uncensored model with concise personas
Controlled environment for toxicity analysis
šŸ”Ž Similar Papers
2024-10-06Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations)Citations: 13
A
Aleksandar TomaÅ”ević
Institute of Physics Belgrade, University of Belgrade, Belgrade, Serbia
D
Darja Cvetković
Institute of Physics Belgrade, University of Belgrade, Belgrade, Serbia
S
Sara Major
Faculty of Philosophy, University of Novi Sad, Novi Sad, Serbia
S
Slobodan Maletić
Vinča Institute of Nuclear Sciences, University of Belgrade, Belgrade, Serbia
M
Miroslav Anđelković
Vinča Institute of Nuclear Sciences, University of Belgrade, Belgrade, Serbia
Ana Vranić
Ana Vranić
Social Physics and Complexity (SPAC), LIP, Portugal
statistical physicscomplex systemscomplex networks
B
Boris Stupovski
Institute of Physics Belgrade, University of Belgrade, Belgrade, Serbia
D
DuÅ”an Vudragović
Institute of Physics Belgrade, University of Belgrade, Belgrade, Serbia
A
Aleksandar Bogojević
Institute of Physics Belgrade, University of Belgrade, Belgrade, Serbia
M
Marija Mitrović Dankulov
Institute of Physics Belgrade, University of Belgrade, Belgrade, Serbia