Experimental evidence of progressive ChatGPT models self-convergence

📅 2026-03-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the potential risk of “model self-convergence” in large language models trained recursively on synthetic data—a phenomenon characterized by increasing output similarity across model versions, declining diversity, and semantic degradation. While hypothesized, long-term empirical evidence has been lacking. To investigate this, the authors conduct the first longitudinal analysis of multiple ChatGPT versions, employing controlled text generation experiments with temperature set to 1 to maximize output diversity. Using established text similarity metrics, they systematically evaluate the evolution of output variability over time. Their findings reveal a significant trend toward convergence in recent model versions, even under high-diversity settings, providing the first empirical validation of model self-convergence and demonstrating its strong association with the rising proportion of synthetic data in training corpora.

Technology Category

Natural Language Processing: (Large) Language ModelsMachine Learning: Large Multimodal Models (LMMs)Cognitive Modeling & Cognitive Systems: Computational Creativity

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Large language models for searchWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Large Language Models (LLMs) that undergo recursive training on synthetically generated data are susceptible to model collapse, a phenomenon marked by the generation of meaningless output. Existing research has examined this issue from either theoretical or empirical perspectives, often focusing on a single model trained recursively on its own outputs. While prior studies have cautioned against the potential degradation of LLM output quality under such conditions, no longitudinal investigation has yet been conducted to assess this effect over time. In this study, we employ a text similarity metric to evaluate different ChatGPT models' capacity to generate diverse textual outputs. Our findings indicate a measurable decline of recent ChatGPT releases' ability to produce varied text, even when explicitly prompted to do so, by setting the temperature parameter to one. The observed reduction in output diversity may be attributed to the influence of the amounts of synthetic data incorporated within their training datasets as the result of internet infiltration by LLM generated data. The phenomenon is defined as model self-convergence because of the gradual increase of similarities of produced texts among different ChatGPT versions.
Problem

Research questions and friction points this paper is trying to address.

model collapse
self-convergence
synthetic data
output diversity
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

model self-convergence
synthetic data
model collapse
text diversity
large language models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Konstantinos F. Xylogiannopoulos
University of Calgary, Department of Computer Science, Calgary, AB, Canada; Stetson University, School of Business Administration, DeLand, FL, USA
Petros Xanthopoulos
Petros Xanthopoulos
Associate professor, Executive Director of Graduate Programs, Stetson University
analyticsmachine learningoperations research
Panagiotis Karampelas
Panagiotis Karampelas
Hellenic Air Force Academy
Software EngineeringData MiningCyber SecurityDigital ForensicsSocial Network Analysis
G
Georgios A. Bakamitsos
Stetson University, Marketing Department, School of Business Administration, DeLand, FL, USA