One Model Is Not a Crowd: Multi-LLM and Aspect-Conditioned Diverse Comment Generation

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the homogenization of AI-generated content and the associated risk of model collapse by proposing a multi-LLM collaboration framework with aspect-conditioned generation. By integrating semantic, linguistic, and sociopragmatic features, the approach simulates diverse human perspectives. We introduce a novel three-dimensional diversity evaluation system grounded in dispersion, coverage, and alignment metrics. The framework is validated using an ensemble of large language models from multiple providers alongside large-scale YouTube comment data. Experimental results demonstrate that multi-model collaboration significantly outperforms single-model approaches, yielding generated data that more closely approximates human distributions and exhibits strong performance in pre-training data filtering and downstream tasks. However, the generated outputs have not yet fully attained the level of diversity observed in human-produced text.
📝 Abstract
Human communication on the internet is shaped by diverse perspectives, most visibly expressed in online comment spaces. As large language model (LLM)based AI agents begin to inhabit these spaces, a key question arises: whether synthetic comment threads can capture the diversity inherent in human discourse. This concern is increasingly important, as the growing presence of homogenized AI-generated content risks reducing diversity over time, potentially leading to model collapse and degrading the richness of digital communication. Inspired by the plurality of human crowds and the aspect-driven nature of discourse, we hypothesize that comment diversity is better approximated by combining multiple LLMs with aspect-conditioned generation. We formalize and evaluate this approach using models from different providers and introduce a framework that characterizes diversity across semantic, linguistic, and socio-pragmatic features along three axes: dispersion, coverage, and alignment. Using this framework, we conduct a large-scale study on over 2 million YouTube comments across multiple domains. Our results reveal that multi-LLM and aspect-conditioned generation better align with human comment distributions and such data remains viable under pretraining style curation and is effective for downstream tasks. Yet, human diversity remains unmatched. Overall, our findings provide a practical foundation for generating more diverse and socially grounded discourse in AI-mediated environments.
Problem

Research questions and friction points this paper is trying to address.

diverse comment generation
multi-LLM
homogenized AI content
model collapse
human discourse diversity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-LLM
Aspect-Conditioned Generation
Diversity Evaluation Framework
Comment Generation
Model Collapse
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Nafis Irtiza Tripto
Nafis Irtiza Tripto
PhD Student
Natural Language ProcessingMachine LearningData Mining
Delvin Ce Zhang
Delvin Ce Zhang
Assistant Professor, University of Sheffield
Multimodal LLMAI for Science
M
Mahjabin Nahar
College of Information Sciences and Technology, Pennsylvania State University, PA, USA
D
Dongwon Lee
College of Information Sciences and Technology, Pennsylvania State University, PA, USA