PrivCert: Certifying Statement Support under Differential Privacy

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the "evidence gap" in differentially private (DP) text generation, where generated claims lack verifiable evidential support. To this end, we propose PrivCert, a framework that explicitly quantifies statement support through privacy certificates and a "release-or-abstain" decision rule. By decoupling candidate discovery from private certification, PrivCert reformulates privacy reporting as an evidence design problem. Technically, it integrates histogram, sparse vector, and Gaussian mechanisms to certify support under differential privacy, while establishing a privacy-honesty frontier and quantifying the cost of multi-statement certification. Experiments across multiple datasets demonstrate that PrivCert significantly reduces the publication rate of unsupported statements, outperforming existing free-text DP baselines.
📝 Abstract
Differentially private (DP) text generation can protect individual records, but privacy alone does not specify what evidence a released statement carries about the underlying data. We identify this as an evidence gap: a private report may contain plausible claims without indicating whether they are strongly supported by the private dataset. We introduce PrivCert, a framework for privacy-preserving reporting that makes statement support explicit through privacy-preserving certificates and emit-or-abstain decisions. As a canonical instantiation, PrivCert-PF (Proposal-and-Filter) separates data-independent candidate discovery from private support certification, emitting only statements whose support passes a private evidence test. We provide theoretical grounding for this framework by characterizing the limits of implicit evidence under DP, deriving a sharp privacy--honesty frontier for single-statement certification, and establishing a worst-case cost for fine-grained multi-statement certification. Experiments on synthetic tasks and TAB, WildChat, and Yelp show that explicit certification maintains low unsupported emission, while free-text DP baselines frequently produce low-support claims under the same declared support semantics. We further show that the PrivCert contract can be realized with histogram, sparse-vector, and Gaussian mechanisms, and use DP synthetic data to illustrate an important boundary: support in a private proxy does not automatically certify support in the original data. Together, these results position privacy-preserving reporting as an evidence-design problem: not only how to generate private text, but what a private report can substantiate about its underlying data.
Problem

Research questions and friction points this paper is trying to address.

Differential Privacy
Text Generation
Evidence Gap
Statement Support
Privacy-Preserving Reporting
Innovation

Methods, ideas, or system contributions that make the work stand out.

Differential Privacy
PrivCert
Emit-or-Abstain
Privacy-Honesty Frontier
Text Generation