Score
Design, implement and evaluate augmentation pipelines that expand labeled training sets by creating synthetic examples through input-space transforms, auxiliary-dataset mixing and filtering, and latent-space operations (e.g., interpolation, anisotropic or ellipsoidal perturbations, and latent diffusion), as well as by text-oriented techniques such as back-translation/round-trip translation and iterative segmentation augmentation. Analyze and tune how these augmentations affect model generalization, robustness to site- or dataset-shift and label noise, overfitting, and training-data efficiency by selecting augmentation types, mixing proportions, and quality-control or filtering procedures.
Existing surveys are limited to unimodal or operation-centric taxonomies, lacking a unified conceptualization of cross-modal data augmentation. To address this, we propose the first modality-agnostic classification framework centered on *intrinsic data relationships*, systematically organizing augmentation techniques across five modalities—images, text, speech, time series, and graph-structured data—according to sample granularity (single-sample, sample-pair, and population-level). Moving beyond conventional modality- or transformation-based categorizations, our framework adopts a *data-centric perspective* to uncover the fundamental principle of *relationship-driven augmentation*. Through inductive analysis, cross-modal comparison, and empirical validation, we construct a comprehensive, three-tiered taxonomy grounded in relational dimensions—semantic, structural, and distributional—which spans all modalities. This taxonomy significantly enhances method comparability, cross-modal transferability, and theoretical coherence in data augmentation research.
To address the weak generalization and severe prediction bias of large language models (LLMs) on few-shot and class-imbalanced text data, this paper proposes an embedding-space synthetic feature augmentation method. Unlike conventional approaches, it operates directly in the language model’s latent embedding space—bypassing raw text generation—and jointly synthesizes minority-class features via embedding interpolation, noise perturbation, and adversarial generation to optimize semantic representation distributions. The method integrates seamlessly into standard fine-tuning pipelines and is compatible with mainstream open-source text classification benchmarks. Experiments across multiple benchmarks demonstrate up to a 12.3% improvement in minority-class F1 score, alongside consistent gains in overall accuracy and robustness. The core innovation lies in migrating synthetic data generation from the input space to the embedding space, enabling efficient, lossless, and fair representation calibration.
Traditional data augmentation techniques (e.g., rotation, flipping) only perturb low-level geometric attributes of images and cannot control high-level semantics (e.g., animal species, plant categories), resulting in insufficient semantic diversity in few-shot learning scenarios. To address this, we propose the first fine-tuning-free, semantic-level image augmentation framework leveraging frozen pre-trained text-to-image diffusion models (e.g., Stable Diffusion). Our method integrates CLIP-guided latent-space editing, prompt-driven semantic redrawing, and conditional inversion to enable zero-shot cross-species and cross-category semantic editing. Crucially, it requires no additional training and generalizes to unseen concepts. Evaluated on few-shot classification and real-world agricultural weed recognition tasks, our approach improves average accuracy by 4.2–9.7%, demonstrating the critical role of semantic diversity in enhancing downstream task performance.
Manual data augmentation design is labor-intensive and suboptimal, limiting model generalization and robustness. Method: We propose the first systematic framework for AutoML-driven data augmentation, unifying three paradigms—data transformation, ensemble-based augmentation, and synthetic-data generation—via integrated Bayesian optimization, reinforcement learning, and meta-learning. The framework supports end-to-end differentiable search across multimodal domains (images and text). Contribution/Results: We establish a standardized evaluation protocol and empirically demonstrate, on CIFAR-10/100, ImageNet, and NLP benchmarks, an average test accuracy gain of 1.2–2.7% over state-of-the-art hand-crafted augmentations, alongside substantial reduction in manual hyperparameter tuning effort. Further experiments confirm superior generalization across unseen domains and enhanced robustness to distributional shifts and adversarial perturbations.
Deep learning models often suffer from overfitting and limited generalization due to their reliance on large-scale labeled datasets. To address this, we propose a semantics-oriented data augmentation method explicitly designed to enhance generalization. Our approach is the first to systematically harness the semantic generation capabilities of pre-trained text-to-image diffusion models (e.g., Stable Diffusion), leveraging prompt engineering, semantic consistency constraints, and class-aware sampling to synthesize augmented images—without requiring additional annotations or model fine-tuning. Critically, the generated samples exhibit high semantic fidelity and robustness to out-of-distribution shifts, surpassing conventional pixel-level augmentation techniques. Empirical evaluation across multiple benchmarks demonstrates substantial improvements in cross-domain generalization, achieving an average 5.2% gain in cross-domain accuracy. The method effectively mitigates overfitting while preserving label semantics and distributional coherence.
Existing research lacks systematic analysis of how hybrid-sample data augmentation techniques—such as CutMix and SaliencyMix—affect the interpretability of deep neural networks. Method: To address this gap, we propose the first three-dimensional interpretability evaluation framework integrating human alignment, model faithfulness, and number of identifiable concepts. We validate it through multi-faceted analysis: gradient- and mask-based attribution, human cognitive experiments, and concept activation vector detection. Contribution/Results: Our experiments reveal—for the first time—that CutMix and SaliencyMix significantly degrade model interpretability, reducing attribution map quality by 23–37%. This work fills a critical void in the joint analysis of data augmentation and interpretability, providing both theoretical foundations and empirical evidence to guide the selection of augmentation strategies under interpretability constraints—particularly in high-stakes applications.
This work addresses the severe overfitting that autoregressive language models exhibit under data-constrained yet compute-rich pretraining regimes, where repeated training epochs on a fixed corpus degrade generalization. To mitigate this, the authors propose data augmentation as a regularization mechanism enabling efficient pretraining for hundreds of epochs on static datasets. Three orthogonal augmentation strategies are introduced: token-level noise (e.g., random token replacement), sequence reordering (e.g., right-to-left prediction and infilling), and target-shifted prediction (e.g., forecasting future tokens). Empirical results demonstrate that each strategy effectively reduces validation loss, with random token replacement yielding the strongest individual gains. Combining these augmentations further lowers validation loss, substantially delaying overfitting and enhancing training efficiency.
This work addresses the challenge that conventional data augmentation methods often fail to simultaneously preserve task relevance and introduce highly diverse, realistic synthetic data, frequently leading to performance degradation due to mismatched augmentations. To overcome this limitation, the authors propose EvoAug, a novel framework that integrates conditional diffusion models and few-shot NeRF-based generative models with evolutionary algorithms to automatically discover task-specific structured stochastic augmentation trees. This approach enables adaptive, learnable data augmentation strategies tailored to the downstream task. Extensive experiments on fine-grained classification and few-shot learning benchmarks demonstrate that EvoAug significantly improves model performance, validating both the effectiveness and generalization capability of the learned augmentation policies.
Traditional text augmentation methods (e.g., back-translation) are limited to lexical substitution and yield semantically homogeneous variants; while direct LLM generation offers knowledge emergence potential, it often compromises semantic fidelity and stylistic controllability. To address this, we propose LMTransplant, the first framework introducing a “transplant-and-regenerate” paradigm: it first embeds a seed text into an expanded semantic context derived from the LLM’s internal knowledge, then regenerates linguistically richer and structurally more diverse variants preserving the core semantics. Its key innovation is a context expansion mechanism that automatically activates the LLM’s knowledge integration capability—without human annotation—thereby jointly optimizing semantic diversity and controllability. Experiments demonstrate that LMTransplant significantly outperforms state-of-the-art augmentation baselines across multiple NLP tasks. Moreover, its performance scales consistently with increasing augmented data volume, exhibiting strong scalability.
To address the insufficient robustness of deep vision models under common image corruptions, this paper proposes a data augmentation pipeline integrating neural style transfer with controllable synthetic image generation. We first observe that stylized degradation—though increasing Fréchet Inception Distance (FID)—significantly improves corruption robustness. We further uncover the complementary mechanisms between style transfer and synthetic data augmentation, and formally characterize their compatibility boundary with rule-based methods such as TrivialAugment. Through systematic hyperparameter analysis and cross-benchmark evaluation, our method achieves state-of-the-art robust accuracy on CIFAR-10-C (93.54%), CIFAR-100-C (74.90%), and TinyImageNet-C (50.86%), establishing new SOTA results on small-scale corruption benchmarks.
This work addresses the inconsistency in existing data augmentation methods when jointly transforming images and their associated multimodal annotations—such as masks, bounding boxes, and keypoints—where mismatched random transformations often lead to misaligned training samples and degraded data quality. To resolve this, the authors propose a unified augmentation framework that encapsulates the augmentation pipeline into composable Compose objects, rigorously synchronizing transformation parameters and random seeds across all modalities. The framework supports diverse data types including images, masks, bounding boxes, keypoints, stereo views, video frames, and volumetric data. Furthermore, it incorporates an augmentation history logging and replay mechanism, ensuring fully reproducible and traceable augmentation processes. This approach significantly enhances the reliability of training data and improves model robustness.