selection-guided synthesis

Designs and implements algorithms and pipelines that synthesize data records by iteratively selecting and evolving candidate records using progressive selection strategies while enforcing differential privacy. This work includes mechanisms to concentrate and allocate privacy budget on selection steps, perform multi-batch top-1 private selection, and coordinate progressive selection across batches to produce synthetic outputs under strict DP guarantees.

selection-guidedsynthesis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of preserving high-order variable correlations in synthetic tabular data under differential privacy. The authors propose Tab-PE, the first algorithm to extend the Private Evolution framework to tabular data generation. Tab-PE introduces low-overhead heuristic evolutionary operators that efficiently optimize candidate datasets within differential privacy constraints, combined with a private scoring and selection mechanism to simultaneously capture complex correlations and ensure scalability. Experimental results demonstrate that Tab-PE significantly outperforms existing methods on both real-world and synthetic datasets, achieving up to a 10% improvement in classification accuracy over the AIM baseline while running 28 times faster.

data sharingdifferential privacyhigh-order correlations

This work addresses the challenge of generating high-quality synthetic structured text under strict privacy constraints, where real data are scarce and heavily regulated. Existing differentially private synthesis methods struggle to simultaneously preserve semantic fidelity and structural validity. To overcome this, the authors propose SelPE, a novel framework that concentrates the privacy budget on a multi-batch top-1 progressive selection process and decouples generation into two stages—semantic abstraction followed by structural realization. Candidate samples are evaluated using a multi-channel distance kernel, and diversity is enhanced through a non-private contrastive expansion mechanism that incurs no additional privacy cost. Experiments demonstrate that, under stringent differential privacy guarantees and limited sample sizes, SelPE significantly improves the structural validity, semantic fidelity, and downstream utility of the synthesized data.

data synthesisdifferential privacyprivacy-preserving

Differentially Private Synthetic Data Generation for Relational Databases

May 29, 2024
KA
Kaveh Alimohammadi
🏛️ MIT | RedHat | IBM

Generating differentially private synthetic data for multi-table relational databases (e.g., academic transcript systems) remains challenging due to structural constraints and utility degradation from flattening. Method: We propose the first end-to-end differentially private (DP) relational synthetic data generation algorithm. It avoids flattening by iteratively calibrating low-order marginals while explicitly enforcing relational constraints—ensuring referential integrity and approximating low-dimensional joint distributions. Contribution/Results: Our approach is the first to decouple arbitrary DP mechanisms from relational structure, achieving both computational efficiency and high-dimensional scalability. It provides rigorous DP guarantees and theoretical utility bounds. Experiments on real-world datasets demonstrate substantial improvements over flattened baselines in query accuracy, cardinality consistency, and JOIN pattern preservation—yielding high-fidelity synthetic relational data.

Complex Structured DataPrivacy ProtectionSynthetic Data

Differentially Private Data Generation with Missing Data

Oct 17, 2023
SM
Shubhankar Mohapatra
🏛️ University of Waterloo

Existing differential privacy (DP) synthetic data methods suffer substantial utility degradation when input data contain missing values. Method: This paper formally defines the DP synthetic data generation problem under missingness, characterizing how missingness mechanisms—Missing Completely at Random (MCAR), Missing at Random (MAR), and Missing Not at Random (MNAR)—propagate privacy loss and affect theoretical privacy bounds. Building on this analysis, we propose three adaptive synthesis strategies, each tailored to a specific missingness mechanism and privacy budget. Contribution/Results: Our key innovation is integrating missingness mechanism modeling directly into the DP synthetic data framework, enabling joint optimization of privacy protection and data utility. Evaluated on four real-world datasets with natural missingness, our approach achieves significant utility gains—improving downstream machine learning task accuracy by an average of 8.3%—while yielding tighter analytical privacy upper bounds compared to baseline methods.

Generating differentially private synthetic data with missing valuesImproving synthetic data utility under privacy constraintsModeling missing mechanisms for tighter privacy guarantees

Systematic Assessment of Tabular Data Synthesis Algorithms

Feb 09, 2024
YD
Yuntao Du
🏛️ Purdue University

Existing tabular data synthesis methods lack a unified, comparable evaluation framework due to fragmented metrics and missing standardized benchmarks—particularly concerning privacy guarantees (differential privacy vs. heuristic approaches) and model paradigms (diffusion models, LLMs vs. statistical methods). Method: We introduce the first systematic evaluation framework for privacy-preserving tabular synthesis, featuring a three-dimensional quantitative metric system—fidelity, privacy, and utility—and a differentiable, unified objective function enabling fair cross-paradigm comparison across diffusion models, LLM-based synthesizers, and marginal-distribution methods. The framework integrates formal differential privacy verification, multi-scale statistical utility assessment (e.g., MMD, JS divergence), adversarial privacy attack benchmarks, and downstream task generalization tests. Results: Extensive experiments across 12 real-world datasets and 8 synthesizers reveal fundamental performance boundaries and trade-off patterns, providing empirical guidance and concrete improvement pathways for next-generation privacy-enhanced synthetic data generation.

Addressing limitations in current synthesis evaluation metricsComparing diffusion/LLM-based synthesizers with statistical methodsEvaluating privacy and utility of tabular data synthesizers

Latest Papers

What's happening recently
View more

This work addresses the problem of efficiently generating synthetic data under differential privacy for a given family of queries. By parameterizing the problem with the treewidth of the query family’s associated graph, the authors establish—for the first time—that the problem is fixed-parameter tractable. They propose a unified dynamic programming framework that integrates linear programming duality-based separation, subsampled private multiplicative weights, and Gibbs sampling techniques. This approach achieves theoretically optimal error rates across the full parameter regime, significantly enhancing both the scalability and practical utility of differentially private synthetic data generation.

differential privacyfixed-parameter tractabilityincidence graph

This work addresses the vulnerability of differentially private synthetic data in protecting anomalous individuals—such as patients with rare diseases—who face significantly higher success rates in membership inference attacks. To mitigate this risk, the authors propose a risk-balanced differentially private synthesis framework that first evaluates the anomaly level of each record using a small privacy budget and then inversely weights records by their risk during generative model training to attenuate the influence of high-risk samples. This approach provides stronger privacy guarantees by incorporating record-level risk awareness into differentially private synthesis for the first time, enabling targeted protection for highly anomalous individuals and yielding a closed-form per-record privacy bound. Experiments on both synthetic and real-world datasets (e.g., Breast Cancer, Adult) demonstrate a substantial reduction in membership inference success against high-anomaly records, with performance contingent on the synergy between the risk scorer and the synthesis pipeline.

differential privacymembership inferenceoutliers

This work addresses fairness concerns in differentially private synthetic data, which often arise from spurious associations between sensitive attributes and outcomes. To mitigate this issue, the authors propose PrivCI, a method that incorporates conditional independence (CI) constraints into the differentially private data synthesis process to eliminate such biases. The core innovation lies in a CI-aware greedy minimum spanning tree algorithm that integrates feasibility checks and the exponential mechanism during Kruskal’s construction, enabling graph structure learning that jointly preserves privacy, fairness, and data fidelity. Experimental results demonstrate that PrivCI significantly outperforms existing approaches in data fidelity and predictive accuracy on standard fairness benchmarks while strictly adhering to prescribed CI constraints.

conditional independencedata synthesisdifferential privacy

This study addresses the instability and noise sensitivity of greedy construction in differentially private prompt optimization under strict privacy budgets. To overcome these limitations, we propose DP-ES, a method that replaces word-by-word greedy search with a population-based evolutionary strategy. By maintaining a complete population of prompts for mutation and evaluation, DP-ES strictly confines privacy consumption to the Gaussian-sampled evaluation phase and incorporates Gumbel-smoothed selection to achieve efficient and secure prompt optimization. Experimental results demonstrate that DP-ES attains 88.1% accuracy on GSM8K, representing a 38.6 percentage point improvement over baselines. Furthermore, it reduces variance by a factor of nine, accelerates inference speed by 2.5 times, and requires fewer privacy queries, highlighting its effectiveness for privacy-preserving prompt engineering.

differential privacyoptimization instabilityprivacy budget

Hot Scholars

SW

Shuai Wang

The Hong Kong University of Science and Technology
Computer SecuritySoftware Engineering
XM

Xing Ma

Meituan, NLP engineer
Dialog SystemLarge Language ModelConversation Analysis
ZC

Zehua Chen

PostDoc at Tsinghua University | Ph.D. from Imperial College
Generative ModelsMulti-modal GenerationHealth Monitoring
KD

Karan Dua

Senior Applied Scientist
Computer VisionNLPSynthetic Data GenerationMLOps
MO

Michael Orshansky

University of Texas at Austin
Hardware SecurityML/HW Co-DesignApproximate Computing