The Poisoned Conversation: Privacy-Leaking Watermarks in Unified Multimodal Models

๐Ÿ“… 2026-10-04
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the cross-modal privacy leakage risks in locally deployed unified multimodal models, where sensitive information within conversational histories can be covertly exfiltrated through image generation. To investigate this vulnerability, we propose the first context-conditioned invisible watermarking attack paradigm that leverages malicious triggers to map early textual cues into subsequently generated images. By integrating adversarial intervention with trigger-dependent watermark embedding, this approach enables imperceptible data exfiltration. Our work reveals the failure mechanisms of inter-modal privacy isolation inherent in unified architectures. Evaluated across two mainstream model families, the proposed method achieves a 100% detection rate alongside a 1% false positive rate, demonstrating severe security vulnerabilities in local deployment scenarios.
๐Ÿ“ Abstract
Multimodal models are increasingly shifting toward unified architectures that understand and generate text, images, and other modalities within a shared conversational context. This design enables fluid interaction across modalities, but it also changes the privacy threat model: Information revealed in one part of a conversation may remain accessible when the model later generates content in another modality. This risk is particularly concerning in settings where users rely on locally deployed models for privacy, assuming that sensitive interactions remain confined to their device. We introduce Privacy-Leaking Watermarks (PLWs): invisible, trigger-dependent watermarks that a malicious model provider can condition on prior chat history. With this adversarial intervention, the usual separation breaks: a sensitive keyword or semantic cue mentioned earlier in the conversation can cause a later, unrelated image to carry a hidden yet detectable watermark. PLWs pose a novel threat to users of unified multimodal models: A poisoned model can retain utility while covertly turning image generation into a channel for privacy leakage, even when deployed locally. Across 13 sensitive-attribute triggers and two model families, PLWs reach up to 100.0% TPR at 1% FPR. For example, across all tested conversational separations, OmniGen2 detects every prior disclosure of depression while falsely flagging only 1% of images generated without such a disclosure.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Models
Privacy Leakage
Watermarking
Adversarial Attack
Unified Architecture
Innovation

Methods, ideas, or system contributions that make the work stand out.

Privacy-Leaking Watermarks
Unified Multimodal Models
Cross-modal Privacy Leakage
Adversarial Watermarking
Trigger-dependent Watermarks
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
T
Tobias Braun
TU Darmstadt, Germany; hessian.AI, Germany
J
Jonas Henry Grebe
TU Darmstadt, Germany; hessian.AI, Germany
E
Emil Sivic
TU Darmstadt, Germany
P
Patrick Mohr Gordillo
TU Darmstadt, Germany
H
Hossein Shakibania
TU Darmstadt, Germany; Zuse School ELIZA
Marcus Rohrbach
Marcus Rohrbach
Professor for Multimodal Reliable AI, TU Darmstadt, Germany
Machine LearningComputer VisionAI
Anna Rohrbach
Anna Rohrbach
Professor, TU Darmstadt, Germany
Vision and LanguageArtificial IntelligenceMultimodal Grounded Learning