🤖 AI Summary
This study addresses the challenge of interpreting the intrinsic properties of large language model (LLM)-generated data, whose utility and limitations in learning tasks remain unclear. To this end, it proposes a synthetic data modeling framework grounded in sample-level learnability representations. By dynamically inferring empirical data distributions through encoder training, the approach systematically compares various LLM families, model scales, and human-written data, while validating cross-encoder robustness and data selection strategies across single- and multi-label classification tasks. The findings reveal fundamental differences in learnability between machine-generated and human-authored data, demonstrating that learnability-driven data selection strategies are effective across diverse data sources. Ultimately, this work establishes a novel paradigm for evaluating synthetic data quality, offering actionable insights into the principled utilization of LLM-generated content for downstream learning applications.
📝 Abstract
Interpreting properties of LLM-generated data is important for understanding its utility and limitations across learning tasks. In this work, we characterize synthetic data through sample-level learnability, studying variation among LLM families and scales, alongside human-written data as a reference. We first generate synthetic datasets spanning single- and multi-label classification, labeling, and tree prediction tasks. We then derive empirical data distributions from encoder training dynamics for both machine and organic data, and estimate the robustness of these distributions across encoders. Finally, we evaluate how data selection strategies based on these learnability signals affect both data sources differently.