Privacy-Friendly Cohort Determination: Sealed, CSP-Independent In-Browser ML Inference of Professional Segments for Identity-Less Advertising

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of identifying professional attributes for B2B advertising under third-party cookie restrictions and tightening privacy regulations. We propose SIF, a sealed inference framework that executes model inference client-side via WebAssembly, employs cross-origin iframes to bypass Content Security Policy constraints, and isolates network access through nested Workers to preserve user privacy. Furthermore, SIF integrates local differential privacy, randomized response, and k-anonymity mechanisms to generate privacy-preserving labels, which are seamlessly injected into the OpenRTB protocol for identifier-free precision targeting. Empirical evaluations demonstrate that the proposed approach effectively predicts company size and type while strictly bounding information leakage from adversarial models to five bits per site per week.
📝 Abstract
B2B advertising targets a viewer's professional attributes (employer size and industry, function, seniority) and has obtained them by matching identities across sites. Safari and Firefox block third-party cookies, Google retired the Privacy Sandbox cohort APIs in 2025, and reverse-IP firmographics decay under remote work. We present SIF (Sealed Inference Frame), which infers coarse professional cohorts on the device and emits only a locally differentially private, taxonomy-coded label into the OpenRTB bid stream, with no cross-site identifier. It rests on a property of the web platform we make precise: a navigated cross-origin iframe is the only way third-party code obtains a policy it controls, so inference runs in WebAssembly even where the publisher's CSP forbids it, and a nested worker served with default-src 'none' gives the model no network. Even a malicious model leaks at most about 5 bits per site per week. Labels pass through a memoised k-ary randomised response keyed to the publisher's first-party identifier, which gives $\varepsilon$-local differential privacy, defeats averaging, and links requests no better than the identifier already sent. An org-conditional k-anonymity rule suppresses cells, more strictly on corporate networks than at home. Cohorts ride OpenRTB user.data in a LinkedIn-aligned taxonomy, and attribution uses LinkedIn's click-scoped li_fat_id without bridging identities. We report a crawl of CSP deployment on 7,969 top sites and 431 B2B publishers, Heavy-Ad budgets, closed-form privacy-utility trade-offs, a re-identification simulation, and an assessment of which attributes are predictable at all: company type and size are, seniority largely is not. On-device is a design property, not a consent exemption.
Problem

Research questions and friction points this paper is trying to address.

B2B advertising
privacy-preserving inference
third-party cookie deprecation
cross-site tracking
professional cohort determination
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sealed Inference Frame
WebAssembly
Local Differential Privacy
On-device ML Inference
Content Security Policy
🔎 Similar Papers
No similar papers found.
O
Om Shankar Tiwari
Applied AI Technical Lead, Google
N
Navnit Shukla
Principal AI Architect, Snowflake Inc.
G
Guanyu Wang
Staff Software Engineer, TikTok
A
Akshay Jain
Staff Software Engineer, LinkedIn