Score
Designs and runs systematic benchmarks that measure how well dataset distillation and condensation methods compress full datasets into small synthetic or summarized training sets, including accuracy and generalization of models trained on the distilled data. Builds experimental comparisons across training protocols and baselines (e.g., coreset selection), and quantifies construction cost, training cost, and scalability to evaluate practical trade‑offs.
Existing dataset distillation methods suffer from poor generalization on small-scale data and prohibitively high computational overhead on large-scale datasets, hindering training efficiency and practical deployment. This paper proposes a unified distillation framework that systematically models and optimizes the entire distillation design space for the first time. Key contributions include: (i) a soft class-aware matching mechanism to enhance semantic consistency between synthetic and real data; (ii) gradient-matching-based meta-optimization coupled with theory-guided architecture selection; and (iii) a dynamic, self-adaptive learning rate scheduling strategy. On ImageNet-1k, our method achieves 48.6% Top-1 accuracy with ResNet-18 using only 10 synthetic images per class (IPC), corresponding to a compression ratio of 0.78%. It significantly outperforms state-of-the-art approaches such as SRe2L, while simultaneously ensuring diversity, fidelity, and training efficiency.
In knowledge distillation, the unavailability of the teacher model’s original training data—due to constraints such as continual learning or data privacy—poses a critical practical bottleneck. Method: This paper systematically investigates the efficacy of substitute datasets for data-free distillation. It proposes and validates non-natural images (e.g., StyleGAN-generated samples) as effective distillation sources, challenging the conventional assumption that original data is indispensable. A multi-dimensional evaluation framework is introduced to quantify distillation data quality along axes of diversity, discriminability, and feature alignment with the teacher. Contribution/Results: Through cross-domain data assessment, teacher–student feature alignment analysis, and ablation studies, the work demonstrates that diverse real and synthetic substitutes achieve distillation performance on par with original data on benchmarks like CIFAR-100—yielding up to a 3.2% accuracy gain in student models. This establishes a novel paradigm and practical guidelines for data-free knowledge distillation.
To address scalability bottlenecks in large-scale dataset distillation—including prohibitive computational cost, high memory consumption, synthetic image homogenization, and poor generalization—this paper proposes Curricular Distillation, a novel curriculum-based framework. It progressively synthesizes training images according to sample difficulty, introduces a first-of-its-kind curriculum-aware evaluation mechanism to mitigate homogenization, and incorporates an adversarial optimization module to enhance image representativeness and cross-architecture robustness. The method integrates curriculum learning, gradient-matching distillation, multi-scale evaluation, and adversarial training. Extensive experiments demonstrate state-of-the-art performance: +11.1%, +9.0%, and +7.3% top-1 accuracy gains on Tiny-ImageNet, ImageNet-1K, and ImageNet-21K, respectively, significantly outperforming prior approaches and establishing a new benchmark for large-scale dataset distillation.
This paper investigates the mechanistic role and dominance of soft labels in dataset distillation. Addressing the limitation of existing methods—overemphasis on synthetic image generation while neglecting label information—we establish soft labels as the primary determinant of distillation performance, validated for the first time via systematic ablation studies and theoretical analysis. Our contributions are threefold: (1) We demonstrate the necessity of structured soft labels, which encode inter-class semantic relationships beyond hard labels; (2) We discover an image–label scaling law, revealing their joint optimization principle; and (3) We construct a Pareto frontier for data-efficient learning, enabling quantitative characterization of the accuracy–data-size trade-off. Extensive experiments on CIFAR-10/100 and ImageNet confirm significant improvements in few-shot generalization. The code is publicly released to ensure reproducibility.
To address the high training cost and substantial data redundancy in deep learning—hindering simultaneous optimization of generalization and efficiency—this paper proposes a dataset distillation method based on cross-layer attention matching. Our approach leverages multi-layer spatial attention maps from randomly initialized neural network ensembles as discriminative supervision signals, guiding gradient-based synthesis of distilled images. To ensure distribution consistency and class separability, we further introduce multi-scale feature alignment. Crucially, our novel cross-layer attention matching mechanism significantly enhances the fidelity of synthesized data compared to prior methods. The proposed method achieves state-of-the-art performance on CIFAR-10/100, TinyImageNet, and ImageNet-1K: classification accuracy improves by 6.5% on CIFAR-100 and 4.1% on ImageNet-1K. Moreover, we demonstrate its broad applicability in continual learning and neural architecture search, confirming strong generalization beyond standard supervised settings.
This study addresses the lack of a standardized evaluation protocol in dataset distillation research, which has hindered objective comparisons between distilled datasets and real-data baselines such as coreset methods. Under a unified experimental setup, the authors conduct the first systematic comparison of seven state-of-the-art distillation techniques against three coreset selection strategies across ImageNet-1K, ImageNet100, and ImageNette, employing both standard empirical risk minimization (ERM) and single/multi-teacher training protocols. Comprehensive evaluations along dimensions of accuracy, representativeness, diversity, and distributional coverage reveal that current distillation approaches do not consistently outperform—and often underperform—coreset methods on large-scale datasets, despite incurring substantially higher computational costs. Notably, coresets demonstrate superior coverage of the original data distribution.
Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.
Dataset distillation (DD) lacks a unified theoretical foundation; existing methods pursue heterogeneous objectives, exhibit unclear generalization bounds, and lack well-characterized conditions for effectiveness under varying training configurations (e.g., optimizers, architectures, data augmentations). Method: We propose the first unified analytical framework encompassing mainstream DD approaches, establishing a “configuration–dynamics–error” theory. Leveraging generalization error analysis, we derive scaling laws for performance saturation and coverage laws for configuration robustness. By integrating gradient matching, distribution matching, and trajectory matching—rigorously grounded in theoretical analysis and validated across diverse training configurations—we provide a principled characterization of DD behavior. Contribution/Results: We prove, for the first time, an optimal linear lower bound on distilled dataset size with respect to configuration diversity. This yields interpretable, verifiable theoretical guarantees for robust DD design, bridging theory and practice in data-efficient learning.
This study addresses the significant performance degradation observed when merging multi-source independently distilled datasets in federated learning. By integrating dataset distillation, Hessian variance analysis, and trajectory compression theory, this work derives, for the first time, an exact composition error formula under quadratic objectives, decoupling the error into local bias and residual components. The analysis reveals fundamental nonlinear discrepancies between independent and joint distillation, demonstrating that training fidelity and downstream accuracy requirements are inherently distinct. Furthermore, it proves that single-dataset evaluation cannot guarantee joint performance and identifies that residuals dominate the error under endpoint matching conditions. These findings provide rigorous theoretical justification for the advantages of joint distillation in federated settings.
Existing dataset distillation methods struggle to significantly outperform random baselines under soft-label settings, often leading to misleading performance evaluations. This work systematically investigates the impact of various labeling strategies—including soft labels, hard labels, and knowledge distillation—on subset quality and model performance, uncovering a performance saturation phenomenon inherent to soft-label distillation. To address this, the authors propose CA2D, an efficient distillation framework tailored for hard-label evaluation. CA2D integrates a computation-aware pruning metric (CAD-Prune), difficulty-aware sampling, and a compute-budget alignment strategy. Evaluated on ImageNet-1K under hard-label constraints, CA2D consistently surpasses both random baselines and state-of-the-art distillation and coreset methods, thereby demonstrating the critical role of data quality in hard-label assessment scenarios.