🤖 AI Summary
This study addresses the systemic risks—such as knowledge dilution, skill degradation, and model collapse—that may arise from feedback loops in human–large language model (LLM) collaboration for constructing online knowledge bases. The authors propose a novel, parsimonious, and interpretable dynamical systems model that captures the co-evolution of knowledge base scale and quality, human and AI capabilities, and query volume. Through numerical simulations and empirical validation using real-world data from Wikipedia, PubMed, and GitHub/Copilot, the model successfully reproduces the observed “inversion” in knowledge contribution patterns before and after the release of ChatGPT. It further uncovers multiple evolutionary trajectories—including healthy growth, oscillation, and inverted flow—and identifies key control levers. These findings offer both theoretical grounding and actionable strategies for designing sustainable human–AI collective knowledge ecosystems.
📝 Abstract
Humans and large language models (LLMs) now co-produce and co-consume the web's shared knowledge archives. Such human-AI collective knowledge ecosystems contain feedback loops with both benefits (e.g., faster growth, easier learning) and systemic risks (e.g., quality dilution, skill reduction, model collapse). To understand such phenomena, we propose a minimal, interpretable dynamical model of the co-evolution of archive size, archive quality, model (LLM) skill, aggregate human skill, and query volume. The model captures two content inflows (human, LLM) controlled by a gate on LLM-content admissions, two learning pathways for humans (archive study vs. LLM assistance), and two LLM-training modalities (corpus-driven scaling vs. learning from human feedback). Through numerical experiments, we identify different growth regimes (e.g., healthy growth, inverted flow, inverted learning, oscillations), and show how platform and policy levers (gate strictness, LLM training, human learning pathways) shift the system across regime boundaries. Two domain configurations (PubMed, GitHub and Copilot) illustrate contrasting steady states under different growth rates and moderation norms. We also fit the model to Wikipedia's knowledge flow during pre-ChatGPT and post-ChatGPT eras separately. We find a rise in LLM additions with a concurrent decline in human inflow, consistent with a regime identified by the model. Our model and analysis yield actionable insights for sustainable growth of human-AI collective knowledge on the Web.