DP-ES: Differentially Private Evolution Strategies for Prompt Optimization

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the instability and noise sensitivity of greedy construction in differentially private prompt optimization under strict privacy budgets. To overcome these limitations, we propose DP-ES, a method that replaces word-by-word greedy search with a population-based evolutionary strategy. By maintaining a complete population of prompts for mutation and evaluation, DP-ES strictly confines privacy consumption to the Gaussian-sampled evaluation phase and incorporates Gumbel-smoothed selection to achieve efficient and secure prompt optimization. Experimental results demonstrate that DP-ES attains 88.1% accuracy on GSM8K, representing a 38.6 percentage point improvement over baselines. Furthermore, it reduces variance by a factor of nine, accelerates inference speed by 2.5 times, and requires fewer privacy queries, highlighting its effectiveness for privacy-preserving prompt engineering.
📝 Abstract
Token-level differentially private (DP) prompt optimization methods such as DP-OPT can become unstable under tight privacy budgets: on GSM8K, DP-OPT obtains $49.5\pm28.5\%$ across 30 runs, and a logged search trajectory reveals prompt-template drift and noise-sensitive irreversible choices. We diagnose these as structural consequences of greedy token-by-token construction over privately aggregated counts. We then propose DP-ES (Differentially Private Evolution Strategies), a structurally cleaner alternative that maintains a population of full prompts, mutates them via LLM calls that never access the private dataset, and spends privacy only on sampled-Gaussian evaluation; deterministic or Gumbel-smoothed selection is post-processing. Under a conservative $(\varepsilon\leq1.0,\delta=10^{-5})$ guarantee, DP-ES achieves 88.1% on GSM8K (+38.6 pp over DP-OPT, approximately 9 times lower standard deviation), 99.7% on MedQA, 73.5% on BANKING77, and 86.8% on Alpaca. It is also 2.5 times faster in wall-clock time and uses 3.3 times fewer logged private-data call groups than DP-OPT. Selection and population ablations, implementation-level noise checks, and a 200-profile exact-match memorization stress test complement the formal guarantee. Scope: Our experiments establish optimization robustness under DP noise, especially where prompt structure is critical; end-to-end validation on genuinely sensitive, non-saturated deployment data remains future work.
Problem

Research questions and friction points this paper is trying to address.

differential privacy
prompt optimization
optimization instability
privacy budget
Innovation

Methods, ideas, or system contributions that make the work stand out.

Differential Privacy
Evolution Strategies
Prompt Optimization
Population-based Search
Privacy Budget
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Ziniu Liu
National University of Defense Technology
A
Aiping Li
National University of Defense Technology
Y
Yue Han
National University of Defense Technology
H
Han Yu
National University of Defense Technology
J
Junjian Zhang
National University of Defense Technology
D
Dong Zhu
National University of Defense Technology
Changjian Li
Changjian Li
Assistant Professor at University of Edinburgh
Computer Graphics3D VisionGeometry Analysis and ProcessingMedical Image Analysis
S
Shiqiang Zhang
CRRC Zhuzhou Electric Locomotive Research Institute Co., Ltd., China Academy of Railway Sciences