Score
Designs and implements algorithms and compact synthetic training sets that compress a large dataset into a small set of representative samples (condensed data points) and optimizes those samples or the condensation procedure to preserve downstream model accuracy. Evaluates and analyzes condensation methods, the generated condensed samples, and associated training pipelines to enable competitive model training with substantially fewer examples.
Existing dataset distillation methods suffer from poor generalization on small-scale data and prohibitively high computational overhead on large-scale datasets, hindering training efficiency and practical deployment. This paper proposes a unified distillation framework that systematically models and optimizes the entire distillation design space for the first time. Key contributions include: (i) a soft class-aware matching mechanism to enhance semantic consistency between synthetic and real data; (ii) gradient-matching-based meta-optimization coupled with theory-guided architecture selection; and (iii) a dynamic, self-adaptive learning rate scheduling strategy. On ImageNet-1k, our method achieves 48.6% Top-1 accuracy with ResNet-18 using only 10 synthetic images per class (IPC), corresponding to a compression ratio of 0.78%. It significantly outperforms state-of-the-art approaches such as SRe2L, while simultaneously ensuring diversity, fidelity, and training efficiency.
In knowledge distillation, performance degradation of student models arises from information scarcity in compressed datasets. To address this, we propose a teacher-model inversion-based method that generates complementary synthetic data to strategically augment the compressed dataset, thereby better approximating the original data distribution. This work is the first to integrate model inversion techniques into compressed-data expansion, enabling effective single-sample-per-class and few-shot distillation—scenarios where conventional direct distillation fails. By jointly optimizing synthetic-data distribution alignment and knowledge distillation, our approach achieves substantial accuracy gains across multiple datasets and model architectures: it significantly outperforms distillation using only compressed data, and improves upon standard model-inversion distillation by up to 11.4%. The method bridges a critical gap between data efficiency and knowledge transfer fidelity, advancing the state of the art in compact yet high-fidelity distillation.
This work investigates whether dataset compression can effectively support adversarial training. It identifies a critical limitation: existing compression methods produce synthetic datasets that fail to transfer adversarial robustness. To address this, the authors propose the first robustness-aware dataset compression framework, grounded in Minimum Finite Cover (MFC), which explicitly incorporates adversarial robustness into the compression objective. Theoretically, the MFC-based method is provably robust, computed in a single pass, model-agnostic, and generalizable. Through rigorous comparison—leveraging generalized adversarial loss minimization and distribution matching—the framework significantly outperforms state-of-the-art dataset compression (DC) methods across three benchmark datasets, achieving superior trade-offs between clean accuracy and adversarial robustness. Moreover, adversarial training on MFC-compressed datasets converges faster and yields higher final robustness than training on datasets compressed via competing DC approaches.
The conventional “fine-tune-then-compress” paradigm for lightweighting large language models (LLMs) during post-training incurs significant performance degradation and introduces redundant intermediate models. Method: This paper proposes the first end-to-end framework that jointly optimizes fine-tuning and structured compression—integrating progressive knowledge distillation, dynamic structured pruning, and low-rank parameter constraints directly into the downstream fine-tuning process to cooperatively shrink the parameter space. Contribution/Results: By eliminating the need to store and compute full-sized intermediate models, our approach reduces memory and computational overhead. On multiple benchmark tasks, it achieves an average accuracy gain of 2.1% at equivalent parameter counts and compresses model size by up to 4.3×, substantially mitigating performance decay inherent in conventional lightweighting pipelines.
This paper addresses the core legal question of whether large language model weights constitute copyright-infringing copies or derivative works of training data. Methodologically, it introduces the novel “training-as-compression” theoretical framework, modeling weights as information-theoretic compressions of training data and integrating entropy, reconstruction fidelity, deep learning training dynamics, and copyright law principles. It rigorously demonstrates— for the first time—that model weights may satisfy the legal criteria for copyright infringement under specific conditions, particularly when they enable high-fidelity reconstruction of copyrighted material. Based on this insight, the paper proposes a dual-dimensional copyright risk assessment methodology grounded in information entropy and reconstruction fidelity. The resulting paradigm provides a rigorous, actionable foundation for adjudicating AI-generated content ownership, designing copyright-compliant training protocols, and informing evidence-based copyright policy development in the age of foundation models.
This work addresses the high computational cost and neglect of heterogeneous features and class imbalance in existing tabular data condensation methods. We propose the first training-free framework for tabular data condensation, formulating the condensation objective as a class-adaptive clustering assignment problem that jointly optimizes class allocation and feature representation. To efficiently solve this NP-hard problem, we introduce a Hybrid Categorical Feature Encoding (HCFE) scheme coupled with a heuristic local search algorithm (HFILS), leveraging soft assignment and intra-class clustering strategies. Extensive experiments on ten real-world datasets demonstrate that our method achieves at least two orders of magnitude speedup over state-of-the-art approaches while delivering superior performance on downstream tasks.
This work addresses the challenge of preserving representational quality while achieving high storage efficiency in dataset condensation under extremely low-bit quantization. To this end, we propose a plug-and-play compression framework that, for the first time, introduces post-training quantization to this task. The method mitigates accuracy degradation through patch-level local quantization, reduces parameter overhead via quantization-aware clustering, and compensates for quantization-induced distribution shifts with a dedicated alignment module. Notably, our approach requires no additional training and consistently outperforms existing methods across CIFAR-10/100, Tiny ImageNet, and ImageNet subsets. Under extreme 2-bit compression with an image-per-class (IPC) budget of 1, it nearly doubles test accuracy—from 26.0% to 54.1%—demonstrating substantial gains in both efficiency and fidelity.
Current climate emulation training datasets suffer from limited scenario diversity, which constrains the generalization capability of machine learning surrogate models. This work addresses this limitation by treating the training scenarios themselves as optimization variables and introduces an iterative data optimization method based on a differentiable simple climate model. By leveraging sensitivity analysis to compute the impact of scenario perturbations on model loss, the approach dynamically generates compact yet highly informative training scenarios. Remarkably, a surrogate model trained on just a single optimized scenario outperforms those trained on six standard ScenarioMIP pathways, effectively disentangling physical responses to distinct external forcings. The method’s superior generalization and data efficiency are further validated on an intermediate-complexity climate model, demonstrating significant improvements in both model skill and sample efficiency.
This study addresses the lack of a standardized evaluation protocol in dataset distillation research, which has hindered objective comparisons between distilled datasets and real-data baselines such as coreset methods. Under a unified experimental setup, the authors conduct the first systematic comparison of seven state-of-the-art distillation techniques against three coreset selection strategies across ImageNet-1K, ImageNet100, and ImageNette, employing both standard empirical risk minimization (ERM) and single/multi-teacher training protocols. Comprehensive evaluations along dimensions of accuracy, representativeness, diversity, and distributional coverage reveal that current distillation approaches do not consistently outperform—and often underperform—coreset methods on large-scale datasets, despite incurring substantially higher computational costs. Notably, coresets demonstrate superior coverage of the original data distribution.
Existing graph neural network (GNN) compression methods typically rely on full-graph training and model-specific architectures, resulting in high computational overhead, limited generalization, and deployment challenges. This work systematically analyzes the fundamental shortcomings of current compression paradigms in methodology, evaluation protocols, and practical applicability. We propose a novel direction that entirely dispenses with full-graph training and model dependency, advocating for lightweight, architecture-agnostic graph compression techniques. Furthermore, we advocate a restructured evaluation framework centered on genuine resource savings and deployability, thereby establishing a theoretical foundation for efficient and scalable GNN training.