Transplant Then Regenerate: A New Paradigm for Text Data Augmentation

📅 2025-08-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Traditional text augmentation methods (e.g., back-translation) are limited to lexical substitution and yield semantically homogeneous variants; while direct LLM generation offers knowledge emergence potential, it often compromises semantic fidelity and stylistic controllability. To address this, we propose LMTransplant, the first framework introducing a “transplant-and-regenerate” paradigm: it first embeds a seed text into an expanded semantic context derived from the LLM’s internal knowledge, then regenerates linguistically richer and structurally more diverse variants preserving the core semantics. Its key innovation is a context expansion mechanism that automatically activates the LLM’s knowledge integration capability—without human annotation—thereby jointly optimizing semantic diversity and controllability. Experiments demonstrate that LMTransplant significantly outperforms state-of-the-art augmentation baselines across multiple NLP tasks. Moreover, its performance scales consistently with increasing augmented data volume, exhibiting strong scalability.

Technology Category

Natural Language Processing: GenerationMachine Learning: Large Multimodal Models (LMMs)Planning, Routing, and Scheduling: Planning with Language Models

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved information
📝 Abstract
Data augmentation is a critical technique in deep learning. Traditional methods like Back-translation typically focus on lexical-level rephrasing, which primarily produces variations with the same semantics. While large language models (LLMs) have enhanced text augmentation by their "knowledge emergence" capability, controlling the style and structure of these outputs remains challenging and requires meticulous prompt engineering. In this paper, we propose LMTransplant, a novel text augmentation paradigm leveraging LLMs. The core idea of LMTransplant is transplant-then-regenerate: incorporating seed text into a context expanded by LLM, and asking the LLM to regenerate a variant based on the expanded context. This strategy allows the model to create more diverse and creative content-level variants by fully leveraging the knowledge embedded in LLMs, while preserving the core attributes of the original text. We evaluate LMTransplant across various text-related tasks, demonstrating its superior performance over existing text augmentation methods. Moreover, LMTransplant demonstrates exceptional scalability as the size of augmented data grows.
Problem

Research questions and friction points this paper is trying to address.

Enhancing text data augmentation diversity via LLMs
Controlling style and structure in LLM-generated outputs
Overcoming limitations of traditional semantic-preserving augmentation methods
Innovation

Methods, ideas, or system contributions that make the work stand out.

Transplant-then-regenerate paradigm for text augmentation
Leverages LLM-expanded context to generate variants
Preserves core attributes while enhancing diversity
🔎 Similar Papers
No similar papers found.