Hiding in Plain Sight: A Diffusion-based Mitigation of Geolocation Privacy Leakage in Vision-Language Models

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对多模态大推理模型导致的地理位置隐私泄露问题,提出了一种基于扩散模型的方法,在图像的潜在空间中注入扰动以有效防御。
📝 Abstract
Multimodal large reasoning models (MLRMs) have demonstrated remarkable capabilities in complex visual understanding. However, this very power introduces a critical yet underexplored privacy threat: adversaries can exploit MLRMs to precisely infer users' geographic locations from casually shared photographs, by performing structured reasoning over subtle visual cues such as architectural styles, vegetation, and lighting conditions. In this work, we present a systematic study of MLRM-driven geolocation privacy leakage. We first reveal that refusal-based safeguards are critically insufficient, as carefully crafted jailbreak prompts can raise model response rates to 100%. We further identify that existing defenses, which inject imperceptible perturbations into shared images, suffer from structural limitations intrinsic to their pixel-space optimization, resulting in degraded black-box transferability and pronounced visual artifacts. Motivated by these findings, we propose a diffusion-based framework that provides targeted, proactive defense against geolocation privacy leakage. By injecting perturbations into the latent space of a diffusion model during reverse sampling, our method operates directly on high-level semantic representations, thereby resolving the effectiveness-utility bottlenecks by construction. We further ground our optimization with GeoCLIP, a model explicitly aligned with GPS coordinates, as a surrogate to pinpoint and disrupt the geographic signals that MLRMs exploit for location inference. This targeted semantic disruption yields significantly stronger black-box transferability while preserving perceptual image quality, offering a seamless integration on social media platforms.
Problem

Research questions and friction points this paper is trying to address.

Geolocation Privacy
Vision-Language Models
Multimodal Large Reasoning Models
Privacy Threat
Innovation

Methods, ideas, or system contributions that make the work stand out.

diffusion-based framework
latent space perturbation
geolocation privacy leakage
semantic disruption
black-box transferability
🔎 Similar Papers
No similar papers found.
Yining Wang
Yining Wang
Fudan University
X
Xi Li
Fudan University
Mi Zhang
Mi Zhang
Fudan University
AI Security
Xiaohan Zhang
Xiaohan Zhang
Fudan University
systems and securityAI security
X
Xiaoyu You
East China University of Science and Technology
Z
Zhenxing Qian
Fudan University
M
Mi Wen
Shanghai University of Electric Power