WebPII: Benchmarking Visual PII Detection for Computer-Use Agents

📅 2026-03-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the risk of visual personally identifiable information (PII) leakage when AI agents process web screenshots—a vulnerability overlooked by existing research due to the absence of dedicated benchmarks. To bridge this gap, we introduce WebPII, the first fine-grained, agent-oriented benchmark for synthetic visual PII detection, comprising 44,865 annotated e-commerce interface images. WebPII features an expanded PII taxonomy and forward-looking detection categories such as transaction identifiers and partially filled forms. Leveraging vision-language models to synthesize UIs and auto-annotate PII, we develop WebRedact, a real-time PII recognition and redaction model that achieves a mAP@50 of 0.753—substantially outperforming text-based baselines (0.357)—and runs at 20 ms inference latency on CPU. Both the dataset and model are publicly released.

Technology Category

Humans and AI: AI for AccessibilityComputer Vision: Large Vision ModelsPhilosophy and Ethics of AI: Privacy & Security

Application Category

Responsible Web: Data and user privacy-enhancing technologies for the WebGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Computer use agents create new privacy risks: training data collected from real websites inevitably contains sensitive information, and cloud-hosted inference exposes user screenshots. Detecting personally identifiable information in web screenshots is critical for privacy-preserving deployment, but no public benchmark exists for this task. We introduce WebPII, a fine-grained synthetic benchmark of 44,865 annotated e-commerce UI images designed with three key properties: extended PII taxonomy including transaction-level identifiers that enable reidentification, anticipatory detection for partially-filled forms where users are actively entering data, and scalable generation through VLM-based UI reproduction. Experiments validate that these design choices improve layout-invariant detection across diverse interfaces and generalization to held-out page types. We train WebRedact to demonstrate practical utility, more than doubling text-extraction baseline accuracy (0.753 vs 0.357 mAP@50) at real-time CPU latency (20ms). We release the dataset and model to support privacy-preserving computer use research.
Problem

Research questions and friction points this paper is trying to address.

PII detection
privacy preservation
computer-use agents
web screenshots
visual privacy
Innovation

Methods, ideas, or system contributions that make the work stand out.

visual PII detection
synthetic benchmark
anticipatory detection
VLM-based UI generation
privacy-preserving agents
N
Nathan Zhao
Stanford University