"Nothing to See Here'': Unintended Disclosure through Revision Traces of LLM Deliverables

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the risk of inadvertent privacy leakage in large language models (LLMs) during sensitive information redaction, where residual revision traces often enable unintended disclosure. To systematically quantify this vulnerability, we introduce RevLeakBench, a novel benchmark encompassing both conversational and agent-based scenarios, and evaluate six mainstream LLMs. Furthermore, we propose an output-side filtering mechanism designed to balance privacy protection with content integrity. Experimental results reveal that approximately 13% of model outputs permit the recovery of sensitive information. The proposed filter significantly mitigates this leakage rate while preserving essential content with negligible degradation. Collectively, this work provides an effective safeguard for secure LLM deployment by highlighting and addressing the overlooked risks associated with revision trace retention.
📝 Abstract
Large language model (LLM) assistants increasingly help users draft content for third-party recipients. During private drafting, the user or the model may introduce an item and later remove or replace it. The model may remove the item from the intended content but reveal it again when stating the edit. We call such statements revision traces. For example, after a user removes the password before sharing a configuration file, the model may delete it but leave a comment saying,"Removed the password'No****4!'as requested."A third-party recipient who sees only the delivered file can therefore recover the withdrawn password from the comment. In an in-the-wild analysis of three public conversation corpora, we identify 26,753 revision requests, of which 2,363 (8.8%) leave revision traces. We study them in greater depth under controlled conditions by introducing RevLeakBench, a benchmark of 100 tasks across five scenarios with a conversation track and an agent track. We measure trace occurrence, withdrawn-item recovery, trace position, and required-content retention. Across six models, about half of the deliverables in both tracks state the edit after a revocation, and a reader that sees only the deliverable can recover the withdrawn item from about 13% of them. Telling the model that its entire reply will be forwarded to the recipient still leaves revision traces in 36.4% of the deliverables. We compare prompt defenses and a delivery boundary, and propose an output-side filter that sharply reduces recovery with little loss of required content. We believe our work can benefit efforts to understand and mitigate unintended disclosure in LLM interactions.
Problem

Research questions and friction points this paper is trying to address.

unintended disclosure
revision traces
large language models
information leakage
privacy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Revision Traces
Unintended Disclosure
RevLeakBench
Output-side Filter
Large Language Models
🔎 Similar Papers
No similar papers found.