🤖 AI Summary
This study addresses the vulnerability of synthetic data provenance to text paraphrasing and its limited utility in guiding data selection for recursive training. Focusing on financial text generation, we conduct comparative experiments integrating style rewriting with recursive model training to systematically evaluate how filtering strategies—based on source identifiers versus reference model scores—affect model degradation. Our findings reveal that data identity recognition and training value assessment constitute distinct problems, challenging the prevailing assumption that traceability inherently implies superiority. Experiments demonstrate that source attribution accuracy drops precipitously after rewriting and fails to consistently mitigate model collapse, establishing that provenance is not a sufficient condition for ensuring recursive training quality. This work thereby provides a new paradigm for synthetic data curation.
📝 Abstract
Repeated training on model-generated data can degrade later models. One possible response is to use provenance when deciding which generated examples to reuse. We test both how reliably that provenance can be recovered and whether it helps identify better training data. Using financial-risk text, we first identify the source of generated passages and then repeat the test after rewriting them. Generator attribution is 98.7% accurate on the original passages but falls to 53.1% after paraphrasing and 29.0% after style rewriting. Generated-versus-human detection remains close to perfect against the tested human comparison set. We then compare two ways of selecting generated examples over three rounds of generation and retraining. One uses source information. The other uses a score from a separate reference model. The two rules select different examples, but the planned comparison does not detect a stable difference in the degradation of the resulting models. The results show that identifying where data came from and identifying which data are useful for training are separate problems. The experiment therefore separates source identity, criterion-facing selection, and recursive training outcome: neither the provenance score nor the tested criterion-facing proxy is established as sufficient for future recursive behaviour.