๐ค AI Summary
This study addresses the cross-modal privacy leakage risks in locally deployed unified multimodal models, where sensitive information within conversational histories can be covertly exfiltrated through image generation. To investigate this vulnerability, we propose the first context-conditioned invisible watermarking attack paradigm that leverages malicious triggers to map early textual cues into subsequently generated images. By integrating adversarial intervention with trigger-dependent watermark embedding, this approach enables imperceptible data exfiltration. Our work reveals the failure mechanisms of inter-modal privacy isolation inherent in unified architectures. Evaluated across two mainstream model families, the proposed method achieves a 100% detection rate alongside a 1% false positive rate, demonstrating severe security vulnerabilities in local deployment scenarios.
๐ Abstract
Multimodal models are increasingly shifting toward unified architectures that understand and generate text, images, and other modalities within a shared conversational context. This design enables fluid interaction across modalities, but it also changes the privacy threat model: Information revealed in one part of a conversation may remain accessible when the model later generates content in another modality. This risk is particularly concerning in settings where users rely on locally deployed models for privacy, assuming that sensitive interactions remain confined to their device. We introduce Privacy-Leaking Watermarks (PLWs): invisible, trigger-dependent watermarks that a malicious model provider can condition on prior chat history. With this adversarial intervention, the usual separation breaks: a sensitive keyword or semantic cue mentioned earlier in the conversation can cause a later, unrelated image to carry a hidden yet detectable watermark. PLWs pose a novel threat to users of unified multimodal models: A poisoned model can retain utility while covertly turning image generation into a channel for privacy leakage, even when deployed locally. Across 13 sensitive-attribute triggers and two model families, PLWs reach up to 100.0% TPR at 1% FPR. For example, across all tested conversational separations, OmniGen2 detects every prior disclosure of depression while falsely flagging only 1% of images generated without such a disclosure.