No Model Required: Text Entropy Rate Filtering Mitigates Iterative Fine-Tuning Collapse

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue of model collapse induced by iterative fine-tuning on synthetic data, noting that existing mitigation strategies typically rely on external models or authentic data. To overcome these limitations, this work proposes a model-free information-theoretic filtering scheme based on Kontoyiannis entropy rate estimation. By leveraging only the statistical properties of raw text to curate training corpora, the approach effectively suppresses diversity degradation without requiring probabilistic models. Experimental evaluations across six generations of QLoRA iterative fine-tuning on Llama-3.1 demonstrate that the proposed method increases the number of unique triplets by 42% and vocabulary size by 30%, while reducing repetition rates by 19%. These results indicate significant improvements over baseline methods, highlighting the efficacy of entropy-based filtering in preserving linguistic diversity during iterative synthetic training.
📝 Abstract
Iterative fine-tuning on synthetic data causes \emph{model collapse}: output diversity narrows as rare patterns are progressively lost, a signature most visible as phrase-level repetition. Existing mitigations either require model log-probabilities, an external oracle, or continued access to real human data. Here we develop a new approach grounded in mathematical information theory: the non-parametric Kontoyiannis entropy rate estimator $h_k$, computed entirely from raw text via match-length statistics, with no model of any kind. We show that this is in fact a \emph{superior} training-data filter on text-diversity metrics in a fully-synthetic, single-lineage fine-tuning setting. In a six-generation QLoRA collapse experiment on Llama-3.1-8B, logprob-based filtering (the most established model-access-requiring baseline) provides no significant text-diversity benefit on any metric ($p > 0.23$), whereas $h_k$-filtering yields $+42\%$ unique trigrams, $+30\%$ vocabulary, and $-19\%$ repetition (all $p < 0.001$). We validate $h_k$ as a cross-domain entropy proxy ($β= 0.924$, $R^2 = 0.746$) and collapse detector ($ρ= +0.454$, $p < 0.0001$) across 4~domains, 2~temperatures, 2~generator--scorer model pairs, and 1{,}520 generated documents. Our results demonstrate that information theoretic approaches to collapse mitigation are efficient, and suggest new approaches for maintaining multi-agent diversity.
Problem

Research questions and friction points this paper is trying to address.

model collapse
iterative fine-tuning
synthetic data
text diversity
entropy rate
Innovation

Methods, ideas, or system contributions that make the work stand out.

Entropy Rate Estimation
Model Collapse Mitigation
Iterative Fine-Tuning
Synthetic Data Filtering
Information Theory
🔎 Similar Papers
No similar papers found.