The power of text similarity in identifying AI-LLM paraphrased documents: The case of BBC news articles and ChatGPT

📅 2025-05-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Generative AI systems (e.g., ChatGPT) rewriting news articles raise serious copyright infringement concerns and erode creators’ revenue. Method: This paper proposes a lightweight, non-deep-learning provenance detection method that models n-gram overlap, word-order preservation, and syntactic structure similarity, augmented by handcrafted features explicitly designed to be sensitive to semantic-preserving perturbations typical of ChatGPT paraphrasing. Contribution/Results: (1) First fine-grained provenance detector specifically tailored for ChatGPT-generated paraphrases; (2) Introduction of the first publicly available BBC News–ChatGPT paraphrase paired benchmark dataset (4,448 instances); (3) A paradigm shift from conventional NLP-based detection—relying solely on pattern matching rather than neural models. On the curated dataset, the method achieves ≥96.2% accuracy, precision, recall, specificity, and F1-score.

Technology Category

Natural Language Processing: GenerationMachine Learning: Large Multimodal Models (LMMs)Cognitive Modeling & Cognitive Systems: Computational Creativity

Application Category

Web Mining and Content Analysis: Web data provenance, reliability, and authenticitySocial Networks and Social Media: Generative AI / large language models and their impact on social systemsSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web search
📝 Abstract
Generative AI paraphrased text can be used for copyright infringement and the AI paraphrased content can deprive substantial revenue from original content creators. Despite this recent surge of malicious use of generative AI, there are few academic publications that research this threat. In this article, we demonstrate the ability of pattern-based similarity detection for AI paraphrased news recognition. We propose an algorithmic scheme, which is not limited to detect whether an article is an AI paraphrase, but, more importantly, to identify that the source of infringement is the ChatGPT. The proposed method is tested with a benchmark dataset specifically created for this task that incorporates real articles from BBC, incorporating a total of 2,224 articles across five different news categories, as well as 2,224 paraphrased articles created with ChatGPT. Results show that our pattern similarity-based method, that makes no use of deep learning, can detect ChatGPT assisted paraphrased articles at percentages 96.23% for accuracy, 96.25% for precision, 96.21% for sensitivity, 96.25% for specificity and 96.23% for F1 score.
Problem

Research questions and friction points this paper is trying to address.

Detect AI-paraphrased news articles for copyright protection
Identify ChatGPT as the source of text infringement
Evaluate pattern-based similarity without deep learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pattern-based similarity detection for AI paraphrased news
Algorithmic scheme to identify ChatGPT as infringement source
High accuracy without deep learning, achieving 96.23% F1 score
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Konstantinos Xylogiannopoulos
Stetson University, School of Business Administration, DeLand, FL, USA
K
Konstantinos Xylogiannopoulos
University of Calgary, Calgary, AB, Canada
Petros Xanthopoulos
Petros Xanthopoulos
Associate professor, Executive Director of Graduate Programs, Stetson University
analyticsmachine learningoperations research
Panagiotis Karampelas
Panagiotis Karampelas
Hellenic Air Force Academy
Software EngineeringData MiningCyber SecurityDigital ForensicsSocial Network Analysis
G
Georgios Bakamitsos
Stetson University, Marketing Department, School of Business Administration, DeLand, FL, USA