data augmentation

Design, implement and evaluate augmentation pipelines that expand labeled training sets by creating synthetic examples through input-space transforms, auxiliary-dataset mixing and filtering, and latent-space operations (e.g., interpolation, anisotropic or ellipsoidal perturbations, and latent diffusion), as well as by text-oriented techniques such as back-translation/round-trip translation and iterative segmentation augmentation. Analyze and tune how these augmentations affect model generalization, robustness to site- or dataset-shift and label noise, overfitting, and training-data efficiency by selecting augmentation types, mixing proportions, and quality-control or filtering procedures.

dataaugmentation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.72
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

To address the weak generalization and severe prediction bias of large language models (LLMs) on few-shot and class-imbalanced text data, this paper proposes an embedding-space synthetic feature augmentation method. Unlike conventional approaches, it operates directly in the language model’s latent embedding space—bypassing raw text generation—and jointly synthesizes minority-class features via embedding interpolation, noise perturbation, and adversarial generation to optimize semantic representation distributions. The method integrates seamlessly into standard fine-tuning pipelines and is compatible with mainstream open-source text classification benchmarks. Experiments across multiple benchmarks demonstrate up to a 12.3% improvement in minority-class F1 score, alongside consistent gains in overall accuracy and robustness. The core innovation lies in migrating synthetic data generation from the input space to the embedding space, enabling efficient, lossless, and fair representation calibration.

Imbalanced DataLarge Language ModelsPredictive Inequity

Effective Data Augmentation With Diffusion Models

Feb 07, 2023
BT
Brandon Trabucco
🏛️ Carnegie Mellon University | MPG Ranch | University Of Chicago Laboratory Schools

Traditional data augmentation techniques (e.g., rotation, flipping) only perturb low-level geometric attributes of images and cannot control high-level semantics (e.g., animal species, plant categories), resulting in insufficient semantic diversity in few-shot learning scenarios. To address this, we propose the first fine-tuning-free, semantic-level image augmentation framework leveraging frozen pre-trained text-to-image diffusion models (e.g., Stable Diffusion). Our method integrates CLIP-guided latent-space editing, prompt-driven semantic redrawing, and conditional inversion to enable zero-shot cross-species and cross-category semantic editing. Crucially, it requires no additional training and generalizes to unseen concepts. Evaluated on few-shot classification and real-world agricultural weed recognition tasks, our approach improves average accuracy by 4.2–9.7%, demonstrating the critical role of semantic diversity in enhancing downstream task performance.

Enhance data augmentation diversity using diffusion modelsImprove few-shot image classification accuracyModify high-level semantic attributes in images

Data augmentation with automated machine learning: approaches and performance comparison with classical data augmentation methods

Mar 13, 2024
AM
A. Mumuni
🏛️ Cape Coast Technical University | University of Mines and Technology

Manual data augmentation design is labor-intensive and suboptimal, limiting model generalization and robustness. Method: We propose the first systematic framework for AutoML-driven data augmentation, unifying three paradigms—data transformation, ensemble-based augmentation, and synthetic-data generation—via integrated Bayesian optimization, reinforcement learning, and meta-learning. The framework supports end-to-end differentiable search across multimodal domains (images and text). Contribution/Results: We establish a standardized evaluation protocol and empirically demonstrate, on CIFAR-10/100, ImageNet, and NLP benchmarks, an average test accuracy gain of 1.2–2.7% over state-of-the-art hand-crafted augmentations, alongside substantial reduction in manual hyperparameter tuning effort. Further experiments confirm superior generalization across unseen domains and enhanced robustness to distributional shifts and adversarial perturbations.

Automatic Data AugmentationComparison with Traditional MethodsMachine Learning Performance

Semantic Augmentation in Images using Language

Apr 02, 2024
SY
Sahiti Yerramilli
🏛️ Carnegie Mellon University

Deep learning models often suffer from overfitting and limited generalization due to their reliance on large-scale labeled datasets. To address this, we propose a semantics-oriented data augmentation method explicitly designed to enhance generalization. Our approach is the first to systematically harness the semantic generation capabilities of pre-trained text-to-image diffusion models (e.g., Stable Diffusion), leveraging prompt engineering, semantic consistency constraints, and class-aware sampling to synthesize augmented images—without requiring additional annotations or model fine-tuning. Critically, the generated samples exhibit high semantic fidelity and robustness to out-of-distribution shifts, surpassing conventional pixel-level augmentation techniques. Empirical evaluation across multiple benchmarks demonstrates substantial improvements in cross-domain generalization, achieving an average 5.2% gain in cross-domain accuracy. The method effectively mitigates overfitting while preserving label semantics and distributional coherence.

Addressing data scarcity in deep learning with semantic augmentationImproving out-of-domain generalization via dataset augmentationReducing overfitting by leveraging diffusion-generated images

Analyzing Effects of Mixed Sample Data Augmentation on Model Interpretability

Mar 26, 2023
SW
Soyoun Won
🏛️ Kyung Hee University

Existing research lacks systematic analysis of how hybrid-sample data augmentation techniques—such as CutMix and SaliencyMix—affect the interpretability of deep neural networks. Method: To address this gap, we propose the first three-dimensional interpretability evaluation framework integrating human alignment, model faithfulness, and number of identifiable concepts. We validate it through multi-faceted analysis: gradient- and mask-based attribution, human cognitive experiments, and concept activation vector detection. Contribution/Results: Our experiments reveal—for the first time—that CutMix and SaliencyMix significantly degrade model interpretability, reducing attribution map quality by 23–37%. This work fills a critical void in the joint analysis of data augmentation and interpretability, providing both theoretical foundations and empirical evidence to guide the selection of augmentation strategies under interpretability constraints—particularly in high-stakes applications.

Effect of label mixing on interpretability degradationImpact of mixed sample augmentation on model interpretabilityMeasuring interpretability via feature attribution maps

Latest Papers

What's happening recently
View more

This work addresses the severe overfitting that autoregressive language models exhibit under data-constrained yet compute-rich pretraining regimes, where repeated training epochs on a fixed corpus degrade generalization. To mitigate this, the authors propose data augmentation as a regularization mechanism enabling efficient pretraining for hundreds of epochs on static datasets. Three orthogonal augmentation strategies are introduced: token-level noise (e.g., random token replacement), sequence reordering (e.g., right-to-left prediction and infilling), and target-shifted prediction (e.g., forecasting future tokens). Empirical results demonstrate that each strategy effectively reduces validation loss, with random token replacement yielding the strongest individual gains. Combining these augmentations further lowers validation loss, substantially delaying overfitting and enhancing training efficiency.

autoregressive pretrainingdata augmentationdata-constrained pretraining

This work addresses the challenge that conventional data augmentation methods often fail to simultaneously preserve task relevance and introduce highly diverse, realistic synthetic data, frequently leading to performance degradation due to mismatched augmentations. To overcome this limitation, the authors propose EvoAug, a novel framework that integrates conditional diffusion models and few-shot NeRF-based generative models with evolutionary algorithms to automatically discover task-specific structured stochastic augmentation trees. This approach enables adaptive, learnable data augmentation strategies tailored to the downstream task. Extensive experiments on fine-grained classification and few-shot learning benchmarks demonstrate that EvoAug significantly improves model performance, validating both the effectiveness and generalization capability of the learned augmentation policies.

data augmentationgenerative modelsoverfitting

Transplant Then Regenerate: A New Paradigm for Text Data Augmentation

Aug 20, 2025
GW
Guangzhan Wang
🏛️ Shanghai Jiao Tong University | Chongqing University

Traditional text augmentation methods (e.g., back-translation) are limited to lexical substitution and yield semantically homogeneous variants; while direct LLM generation offers knowledge emergence potential, it often compromises semantic fidelity and stylistic controllability. To address this, we propose LMTransplant, the first framework introducing a “transplant-and-regenerate” paradigm: it first embeds a seed text into an expanded semantic context derived from the LLM’s internal knowledge, then regenerates linguistically richer and structurally more diverse variants preserving the core semantics. Its key innovation is a context expansion mechanism that automatically activates the LLM’s knowledge integration capability—without human annotation—thereby jointly optimizing semantic diversity and controllability. Experiments demonstrate that LMTransplant significantly outperforms state-of-the-art augmentation baselines across multiple NLP tasks. Moreover, its performance scales consistently with increasing augmented data volume, exhibiting strong scalability.

Controlling style and structure in LLM-generated outputsEnhancing text data augmentation diversity via LLMsOvercoming limitations of traditional semantic-preserving augmentation methods

Stylized Synthetic Augmentation further improves Corruption Robustness

Dec 17, 2025
GS
Georg Siedel
🏛️ University of Stuttgart | Federal Institute for Occupational Safety and Health (BAuA)

To address the insufficient robustness of deep vision models under common image corruptions, this paper proposes a data augmentation pipeline integrating neural style transfer with controllable synthetic image generation. We first observe that stylized degradation—though increasing Fréchet Inception Distance (FID)—significantly improves corruption robustness. We further uncover the complementary mechanisms between style transfer and synthetic data augmentation, and formally characterize their compatibility boundary with rule-based methods such as TrivialAugment. Through systematic hyperparameter analysis and cross-benchmark evaluation, our method achieves state-of-the-art robust accuracy on CIFAR-10-C (93.54%), CIFAR-100-C (74.90%), and TinyImageNet-C (50.86%), establishing new SOTA results on small-scale corruption benchmarks.

Combining synthetic data with style transfer for training augmentationEnhancing deep vision models' robustness to image corruptionsImproving classifier performance on corrupted image benchmarks

This work addresses the inconsistency in existing data augmentation methods when jointly transforming images and their associated multimodal annotations—such as masks, bounding boxes, and keypoints—where mismatched random transformations often lead to misaligned training samples and degraded data quality. To resolve this, the authors propose a unified augmentation framework that encapsulates the augmentation pipeline into composable Compose objects, rigorously synchronizing transformation parameters and random seeds across all modalities. The framework supports diverse data types including images, masks, bounding boxes, keypoints, stereo views, video frames, and volumetric data. Furthermore, it incorporates an augmentation history logging and replay mechanism, ensuring fully reproducible and traceable augmentation processes. This approach significantly enhances the reliability of training data and improves model robustness.

annotation alignmentdata augmentationimage transformation

Hot Scholars

TN

Trung-Nghia Le

University of Science, VNU-HCM
Applied Deep LearningApplied Computer VisionMultimedia Security
MT

Minh-Triet Tran

University of Science & John von Neumann Institute, VNU-HCM
Cryptography and SecurityMultimedia and InteractionComputer Vision and Machine LearningSoftware Engineering
SM

Sébastien Marcel

Senior researcher ( Idiap research institute ) and Professor ( University of Lausanne )
AIbiometricssecurity and privacyanti-spoofing and deepfakes
SP

Symeon Papadopoulos

Information Technologies Institute (ITI)
Artificial IntelligenceMedia VerificationAI FairnessWeb Mining
SY

Suorong Yang

Nanjing University
Computer VisionDeep LearningMultimodal Learning