Semantic Reconstruction of Adversarial Plagiarism: A Context-Aware Framework for Detecting and Restoring "Tortured Phrases" in Scientific Literature

📅 2025-12-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Scientific literature faces growing threats from semantic distortion plagiarism induced by automated rewriting tools (e.g., substituting “artificial intelligence” with “counterfeit consciousness”), against which existing detection methods suffer from high false-negative rates and lack traceability. This paper proposes the first context-aware, two-stage semantic reconstruction framework: (1) a domain-adapted SciBERT-based pseudo-perplexity anomaly detection module identifies distorted phrases; (2) a hybrid retrieval and alignment stage integrates FAISS-enabled dense retrieval with SBERT-based sentence-level alignment to mathematically reconstruct original terms and locate source documents. To address the high lexical variance of scientific terminology, we introduce a static thresholding strategy. Evaluated on adversarial parallel corpora, our method achieves 23.67% original-term recovery accuracy—surpassing the zero-shot baseline (0%)—demonstrating substantial improvements in detection robustness and provenance tracing capability.

Technology Category

Natural Language Processing: Fact-Checking / Misinformation Detection (NLP Focus)Computer Vision: Adversarial Attacks & RobustnessMachine Learning: Adversarial Learning & Robustness

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsSemantics and Knowledge: Scalable techniques for the creation, curation, publication, maintenance, and consumption of large, Web-based, structured, reusable, knowledge graphs and ontologiesWeb Mining and Content Analysis: Robustness and generalizability of Web mining methods
📝 Abstract
The integrity and reliability of scientific literature is facing a serious threat by adversarial text generation techniques, specifically from the use of automated paraphrasing tools to mask plagiarism. These tools generate "tortured phrases", statistically improbable synonyms (e.g. "counterfeit consciousness" for "artificial intelligence"), that preserve the local grammar while obscuring the original source. Most existing detection methods depend heavily on static blocklists or general-domain language models, which suffer from high false-negative rates for novel obfuscations and cannot determine the source of the plagiarized content. In this paper, we propose Semantic Reconstruction of Adversarial Plagiarism (SRAP), a framework designed not only to detect these anomalies but to mathematically recover the original terminology. We use a two-stage architecture: (1) statistical anomaly detection with a domain-specific masked language model (SciBERT) using token-level pseudo-perplexity, and (2) source-based semantic reconstruction using dense vector retrieval (FAISS) and sentence-level alignment (SBERT). Experiments on a parallel corpus of adversarial scientific text show that while zero-shot baselines fail completely (0.00 percent restoration accuracy), our retrieval-augmented approach achieves 23.67 percent restoration accuracy, significantly outperforming baseline methods. We also show that static decision boundaries are necessary for robust detection in jargon-heavy scientific text, since dynamic thresholding fails under high variance. SRAP enables forensic analysis by linking obfuscated expressions back to their most probable source documents.
Problem

Research questions and friction points this paper is trying to address.

Detects adversarial plagiarism using 'tortured phrases' in scientific literature
Recovers original terminology from obfuscated text via semantic reconstruction
Addresses limitations of static blocklists and general language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Two-stage detection with domain-specific language model
Source-based semantic reconstruction using dense retrieval
Static decision boundaries for robust jargon detection
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Agniva Maiti
School of Computer Engineering, KIIT University, Bhubaneswar, India.
P
Prajwal Panth
School of Computer Engineering, KIIT University, Bhubaneswar, India.
S
Suresh Chandra Satapathy
School of Computer Engineering, KIIT University, Bhubaneswar, India.