Privacy Leakage Through AI-mediated Analysis of Smartphone Data

πŸ“… 2026-09-22
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the privacy risks arising from AI-driven analysis on smartphones, where conventional permission prompts fail to mitigate deep inference threats. To investigate this, we develop Priva-See, a system that provides the first empirical demonstration of large language models (LLMs) inferring sensitive user profiles from single-permission multimodal data. Furthermore, we conduct an IRB-compliant user study involving 465 participants to evaluate the perceived intrusiveness of such inferences. Our findings reveal that significant privacy violations substantially diminish users’ willingness to share data, thereby exposing the inadequacy of traditional authorization mechanisms. Based on these insights, this work proposes recommendations for operating-system-level permission reforms to better safeguard user privacy against advanced inferential capabilities.
πŸ“ Abstract
Over the past thirty years, the online advertising industry built a large-scale data collection ecosystem, with the goal of tracking a user's online activity to infer their demographics and interests. Traditionally, the ecosystem relied upon the collation and analysis of highly-structured text data like user IP addresses, GPS coordinates, e-commerce purchase histories, and visited URLs. However, recent ML models can parse not only structured text, but also multimedia files and unstructured text inputs---meaning a user's photos, videos, inboxes, and calendars are now ripe for automated analysis. The privacy risks are particularly acute in the context of smartphone apps. A user's phone already acts as a natural collation point for sensitive user information, but users may not understand that permitting an app to, for example, access a user's photo does not just give the app access to the bytes in the photo: the app also receives access to inferences about the user that are enabled by the photo. To explore these privacy risks, we built Priva-See, an LLM-based inference system for app-collected user data; Priva-See reflects our best understanding of how real-life adtech companies would leverage machine learning to build user profiles. Through an IRB-approved user study, 465 participants deployed Priva-See on their phones; Priva-See made privacy-invasive inferences despite having access to only a subset of a user's data. We see the experience significantly impacted participant willingness to share permissions data moving forward. Based on the observed privacy violations, we suggest changes to how smartphone OSes should gather user consent for data access, to better inform users about downstream data usage capability.
Problem

Research questions and friction points this paper is trying to address.

privacy leakage
smartphone data
AI-mediated inference
user profiling
data consent
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Privacy Leakage
Smartphone Data
User Profiling
Multimodal Analysis
πŸ”Ž Similar Papers
No similar papers found.
Sarah Radway
Sarah Radway
Harvard University
Z
Zoe Robert
Unaffiliated
M
Matthew Soto
Carnegie Mellon University
J
Julianna Cimillo
Unaffiliated
S
Sebastian Diaz
Harvard University
M
Meg Marco
Harvard University
James Mickens
James Mickens
Harvard University
SystemsOperating SystemsDistributed SystemsWeb ServicesDatacenters