generate paraphrases

Design and build models, algorithms, or pipelines that automatically rewrite input text into alternative phrasings that preserve the original semantics (paraphrase generation / semantic paraphrasing), optionally producing multiple variants and controlling stylistic attributes such as tone or normalization of wording.

generateparaphrases

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.39
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work proposes a fine-grained paradigm for modeling semantic equivalence by decomposing paraphrasing into specific linguistic operations—such as lexical substitution and syntactic transformation—to construct structured, cognitively plausible representations of meaning. Rather than relying on coarse binary classification or single rewrites, the approach explicitly captures the nuanced mechanisms underlying paraphrase generation. A deep semantic model trained on annotated data achieves 89.6% accuracy on a Wikipedia plagiarism detection task and 66.5% on an arXiv dataset, substantially outperforming human baselines. Furthermore, the method demonstrates consistent performance gains on downstream tasks such as Quora duplicate question detection, validating its effectiveness in both paraphrase understanding and controllable generation.

language modelinglinguistic variationmeaning preservation

This study addresses the challenges of semantic distortion and hallucination commonly encountered by small language models when rewriting short texts with high semantic density. The authors present the first systematic exploration of optimization strategies for this task, constructing a supervised dataset generated by GPT-4 and applying prompt distillation combined with parameter-efficient fine-tuning to adapt the Phi Silica model. Model performance is evaluated using both LLM-as-a-judge metrics and human preference assessments. Results demonstrate that the optimized model surpasses GPT-4-generated outputs in semantic fidelity, hallucination suppression, and win rates in human preference comparisons, substantially narrowing the performance gap with large cloud-based language models.

hallucination robustnessparaphrasingsemantic fidelity

Say It Another Way: A Framework for User-Grounded Paraphrasing

May 06, 2025
CC
Cléa Chataigner
🏛️ Mila | McGill University | University of Waterloo | Vector Institute | Université de Montréal | Quebec AI Institute

Large language models (LLMs) exhibit high sensitivity to minor lexical or syntactic variations in prompts; however, existing evaluation methods often rely on hand-crafted or unnatural perturbations, failing to reflect robustness under authentic linguistic usage. Method: We propose the first linguistics-driven minimal transformation classification framework for prompt rewriting—characterized by fine-grained, controllable, and interpretable transformations grounded in user context. Our approach integrates BBQ benchmark adaptation, dual-verification via human annotation and automated consistency checking, and quantitative stability analysis. Contribution/Results: Experiments reveal that natural paraphrasing induces accuracy fluctuations exceeding 20%, exposing a widespread lack of paraphrase robustness in current LLM evaluations. This work establishes a foundational paradigm for paraphrase-aware LLM assessment, advancing evaluation standards toward linguistic realism and contextual fidelity.

Assess LLM sensitivity to subtle paraphrasing changesDevelop framework for natural prompt variation generationStudy how prompt wording affects LLM behavior stability

ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data

Apr 20, 2025
TC
Tong Chen
🏛️ University of Washington | Allen Institute for Artificial Intelligence

Language models frequently verbatim reproduce pretraining data in non-adversarial settings, posing risks to copyright, privacy, and originality. To address this, we propose ParaPO—a post-training reinforcement alignment method grounded in paraphrasing preference optimization. ParaPO jointly leverages preference optimization and contrastive learning to steer models toward paraphrasing rather than literal repetition, while incorporating a plug-and-play system prompt mechanism that suppresses unnecessary verbatim copying without compromising legitimate quotation (e.g., canonical aphorisms). Experiments on Llama3.1-8B and Tulu3-8B demonstrate that ParaPO significantly reduces verbatim repetition rates in creative writing (e.g., from 17.3% to 12.9%), outperforming existing unlearning techniques, while preserving recall of well-known quotations. To our knowledge, ParaPO is the first alignment framework that jointly models paraphrasing preferences and enables controllable, prompt-based suppression of unwanted memorization—achieving both fidelity to source meaning and respect for intellectual property norms.

Controlling regurgitation behavior using system promptsPreserving model utility while minimizing unintentional memorizationReducing verbatim reproduction of pretraining data in LMs

Linearly Controlled Language Generation with Performative Guarantees

May 24, 2024
EC
Emily Cheng
🏛️ Universitat Pompeu Fabra | ICREA | ETH Zürich

To address the need for controllable text generation by large language models (LLMs) in safety-critical applications, this paper tackles the challenge of efficiently ensuring semantic safety of generated outputs. Methodologically, it formalizes semantic constraints as linear structures in the model’s latent space and models the generation process as a trajectory evolution; it then introduces the first gradient-free, closed-form geometric intervention strategy for latent-space control. Theoretically, it establishes the first probabilistic guarantee that generated texts provably reside within a pre-specified safe semantic region. Empirical evaluation on toxicity mitigation demonstrates substantial reduction in harmful content generation while preserving linguistic fluency and lexical diversity—achieving a balanced optimization between controllability and generation quality.

Control language generation for performance guaranteesEnsure activations enter predefined allowed semantics spaceSteer trajectories away from undesired semantic regions

Latest Papers

What's happening recently
View more

This study addresses the underexplored impact of differentially private text rewriting on stylistic register, despite its aim to preserve semantic content. By conducting multidimensional register analysis and employing both autoregressive paraphrasing and bidirectional substitution methods for sentence-level differential privacy under varying privacy budgets, the work reveals that privacy constraints systematically reduce interactional markers, contextual references, and complex subordinate clauses. Consequently, privatized texts gravitate toward a homogeneous, non-interactive, and non-persuasive register, thereby diminishing stylistic nuance. This paper is the first to demonstrate the systematic influence of differential privacy mechanisms on textual communicative functions, highlighting the inherent tension between privacy preservation and stylistic fidelity.

Differential PrivacyLinguistic StyleRegister Identity

This work proposes RewriteNets, a neural architecture grounded in explicit parallel string rewriting to address the limitations of conventional sequence models such as Transformers, which suffer from quadratic computational complexity due to implicit structural representations and exhibit poor systematic generalization. RewriteNets incorporate learnable rewriting rules at each layer, enabling efficient sequence modeling through fuzzy pattern matching, conflict resolution, and non-overlapping rule selection. The model integrates a straight-through Gumbel-Sinkhorn estimator to facilitate end-to-end differentiable training, effectively casting sequence modeling as a differentiable symbolic rewriting process with explicit structural inductive bias. Evaluated on the SCAN length-split task, RewriteNets achieve 98.7% accuracy—substantially outperforming LSTM and Transformer baselines—while demonstrating superior computational efficiency and enhanced systematic generalization capabilities.

computational efficiencysequence modelingstring rewriting

This study investigates the dependence of synthetic rewriting on data quality in continual pretraining for Portuguese, examining whether it can serve as a substitute for high-quality data curation. Using the ClassiCC-PT corpus—annotated with STEM content and educational quality scores—the authors construct 10B-token subsets of high- and low-quality text. They generate approximately 80B tokens of synthetically rewritten data in four styles using a 7B instruction-tuned model, then train 1.1B and 7B models evaluated on the PoETa V2 benchmark. The work provides the first systematic evidence in a non-English setting that synthetic rewriting acts as an amplifier—not a replacement—for data quality, with this effect intensifying at larger scales: the 7B model gains +3.4 NPM with rewritten high-quality data versus only +0.5 NPM with low-quality data, while the 1.1B model shows no significant amplification.

data qualitylanguage model pretrainingPortuguese

This work addresses the challenge of detecting high-fidelity paraphrased text generated by large language models (LLMs), which preserves semantic content so effectively that conventional detection methods often fail. To tackle this issue, the authors propose SearchLLM, a novel approach that integrates a search engine with a text regeneration mechanism. Specifically, SearchLLM retrieves candidate source texts, regenerates their LLM-paraphrased versions, and compares these against the input text to assess similarity. Designed as a plug-in proxy layer, SearchLLM seamlessly enhances existing detectors without requiring architectural modifications. Experimental results demonstrate that this method significantly improves detection accuracy across diverse LLM-generated paraphrasing datasets, thereby substantially increasing robustness against semantically faithful paraphrasing attacks.

AI-generated textLLM paraphrasingoriginal source identification

Scientific literature faces growing threats from semantic distortion plagiarism induced by automated rewriting tools (e.g., substituting “artificial intelligence” with “counterfeit consciousness”), against which existing detection methods suffer from high false-negative rates and lack traceability. This paper proposes the first context-aware, two-stage semantic reconstruction framework: (1) a domain-adapted SciBERT-based pseudo-perplexity anomaly detection module identifies distorted phrases; (2) a hybrid retrieval and alignment stage integrates FAISS-enabled dense retrieval with SBERT-based sentence-level alignment to mathematically reconstruct original terms and locate source documents. To address the high lexical variance of scientific terminology, we introduce a static thresholding strategy. Evaluated on adversarial parallel corpora, our method achieves 23.67% original-term recovery accuracy—surpassing the zero-shot baseline (0%)—demonstrating substantial improvements in detection robustness and provenance tracing capability.

Addresses limitations of static blocklists and general language modelsDetects adversarial plagiarism using 'tortured phrases' in scientific literatureRecovers original terminology from obfuscated text via semantic reconstruction

Hot Scholars

YC

Yejin Choi

Stanford University / NVIDIA
Natural Language ProcessingDeep LearningArtificial IntelligenceCommonsense Reasoning
NA

Nicholas Andrews

Johns Hopkins University
natural language processingmachine learning
BS

Binesh Sadanandan

Ph.D Student, University of New Haven
Machine LearningHealthcareMultimodal LLM
VB

Vahid Behzadan

Assistant Professor - University of New Haven
AI SafetySecurityWireless CommunicationsGame Theory
DT

Dacheng Tao

Nanyang Technological University
artificial intelligencemachine learningcomputer visionimage processing