Text Data Augmentation for Large Language Models: A Comprehensive Survey of Methods, Challenges, and Opportunities

📅 2025-01-31
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In low-resource settings, large language models (LLMs) suffer from insufficient training data, while existing text data augmentation methods struggle to simultaneously ensure enhanced text quality and factual fidelity. Method: This paper systematically surveys LLM-driven text data augmentation techniques and proposes a unified taxonomy—comprising simple, prompt-based, retrieval-augmented, and hybrid augmentation—alongside a novel “generate–retrieve–verify” collaborative augmentation paradigm that emphasizes post-hoc verification for improving data trustworthiness. It integrates prompt engineering, external knowledge retrieval, generation-quality filtering, and multi-dimensional evaluation to enable task-controllable, high-fidelity text generation. Contribution/Results: The work delivers a comprehensive technical landscape covering methodologies, open challenges, and emerging opportunities, yielding a reusable, empirically verifiable augmentation framework specifically designed to advance low-resource LLM training.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: GenerationData Mining & Knowledge Management: Conversational Systems for Recommendation & Retrieval

Application Category

Search and Retrieval-Augmented AI: Large language models for searchSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
The increasing size and complexity of pre-trained language models have demonstrated superior performance in many applications, but they usually require large training datasets to be adequately trained. Insufficient training sets could unexpectedly make the model overfit and fail to cope with complex tasks. Large language models (LLMs) trained on extensive corpora have prominent text generation capabilities, which improve the quality and quantity of data and play a crucial role in data augmentation. Specifically, distinctive prompt templates are given in personalised tasks to guide LLMs in generating the required content. Recent promising retrieval-based techniques further improve the expressive performance of LLMs in data augmentation by introducing external knowledge to enable them to produce more grounded-truth data. This survey provides an in-depth analysis of data augmentation in LLMs, classifying the techniques into Simple Augmentation, Prompt-based Augmentation, Retrieval-based Augmentation and Hybrid Augmentation. We summarise the post-processing approaches in data augmentation, which contributes significantly to refining the augmented data and enabling the model to filter out unfaithful content. Then, we provide the common tasks and evaluation metrics. Finally, we introduce existing challenges and future opportunities that could bring further improvement to data augmentation.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Data Augmentation
Text Quality and Authenticity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Data Augmentation
Large Language Models
Knowledge Injection
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yaping Chai
School of Data Science, Lingnan University, Hong Kong
H
Haoran Xie
School of Data Science, Lingnan University, Hong Kong
J
Joe S. Qin
School of Data Science, Lingnan University, Hong Kong