TREND: A Whitespace Replacement Information Hiding Method

📅 2025-02-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the challenge of tracing text generated by large language models (LLMs), this paper proposes a semantic-lossless steganographic method leveraging Unicode whitespace characters. Without altering lexical meaning or modifying character count, it covertly embeds arbitrary byte sequences into plain text. The approach innovatively employs visually imperceptible Unicode whitespace code points (e.g., U+200B–U+200F) to construct configurable ciphertext structures, integrating LZ4 compression, AES-256 encryption, SHA-256 hashing, and Reed–Solomon error correction. We implement a cross-platform Kotlin library, a command-line interface tool, and a web-based interface. Evaluated on a million-document Wikipedia benchmark, the method exhibits zero perceptibility to human readers and robustness against common text-processing operations—including encoding conversion, formatting, and copy-paste. It achieves an embedding capacity of 1 bit per character, outperforming ten state-of-the-art steganographic schemes.

Technology Category

Natural Language Processing: (Large) Language ModelsMachine Learning: Large Multimodal Models (LMMs)Planning, Routing, and Scheduling: Planning with Language Models

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Large Language Models (LLMs) have gained significant popularity in recent years. Differentiating between a text written by a human and a text generated by an LLM has become almost impossible. Information hiding techniques such as digital watermarking or steganography can help by embedding information inside text without being noticed. However, existing techniques, such as linguistic-based or format-based methods, change the semantics or do not work on pure, unformatted text. In this paper, we introduce a novel method for information hiding termed TREND, which is able to conceal any byte-encoded sequence within a cover text. The proposed method is implemented as a multi-platform library using the Kotlin programming language, accompanied by a command-line tool and a web interface provided as examples of usage. By substituting conventional whitespace characters with visually similar Unicode whitespace characters, our proposed scheme preserves the semantics of the cover text without increasing the number of characters. Furthermore, we propose a specified structure for secret messages that enables configurable compression, encryption, hashing, and error correction. Our experimental benchmark comparison on a dataset of one million Wikipedia articles compares ten algorithms from literature and practice. It proves the robustness of our proposed method in various applications while remaining imperceptible to humans. We discuss the limitations of limited embedding capacity and further robustness, which guide implications for future work.
Problem

Research questions and friction points this paper is trying to address.

Differentiate human vs LLM-generated text.
Hide information without altering text semantics.
Enhance robustness and capacity of information hiding.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses Unicode whitespace for hiding
Preserves text semantics completely
Configures compression, encryption, hashing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Malte Hellmeier
Fraunhofer ISST, Speicherstr. 6, 44147 Dortmund, Germany
H
Hendrik Norkowski
Radix Security GmbH, Universitaetsstr. 136, 44799 Bochum, Germany
E
Ernst-Christoph Schrewe
Fraunhofer ISST, Speicherstr. 6, 44147 Dortmund, Germany
Haydar Qarawlus
Haydar Qarawlus
Fraunhofer ISST, Speicherstr. 6, 44147 Dortmund, Germany
Falk Howar
Falk Howar
TU Dortmund
computer scienceformal methodsautomata learning