evaluate synthetic data

Designs and implements evaluation frameworks, benchmarks, and diagnostic analyses for synthetic datasets that quantify fidelity, diversity, realism, and downstream utility by building and applying metrics, statistical-comparison tests, and retraining experiments (e.g., train-on-synthetic test-on-real, retrain-on-synthetic) to measure performance gaps and protocol conformity. Performs privacy-aware validation and privacy–utility–fairness assessments, property-based validation, and failure-mode detection to compare synthesizers, assess augmentation or restoration effects, and evaluate synthetic-data detection and benchmarking tools.

evaluatesyntheticdata

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.63
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$257K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing approaches struggle to effectively quantify the similarity and quality between synthetic and real data in evaluating tool-augmented agents. To address this gap, this work proposes SynAE, a novel framework that establishes the first multi-axis evaluation system tailored for multi-turn tool-use scenarios. SynAE introduces four fine-grained metric categories—assessing task instructions, tool invocations, final outputs, and downstream evaluation performance—to systematically measure synthetic data across dimensions of validity, fidelity, and diversity. Integrating natural language processing, trajectory modeling, and controllable generation techniques, the framework enables a reproducible evaluation pipeline and successfully identifies several representative failure modes in synthetic data generation. Empirical results demonstrate that such multidimensional assessment is essential for enhancing the reliability of agent evaluations.

benchmarkingdata qualityevaluation framework

To address the challenge of balancing privacy preservation and utility in synthetic tabular data evaluation, this paper proposes the first standardized, multidimensional assessment framework. Methodologically, it integrates low- and high-dimensional distributional comparisons, deep embedding similarity metrics, k-nearest-neighbor distance statistics, and statistical distribution tests—establishing a unified benchmark paradigm covering statistical utility, privacy security, and interpretable diagnostics, applicable to both sequential and context-structured data. Contributions include: (1) a quantifiable, interpretable, and cross-method comparable quality evaluation system; (2) adoption of a hold-out benchmarking strategy with standardized metrics, significantly enhancing reproducibility and assessment consistency; and (3) open-sourcing of the implementation to foster methodological unification in the field.

Developing a multi-dimensional benchmarking framework for synthetic dataEnsuring reproducibility in synthetic data generation techniquesEvaluating synthetic data quality for privacy and utility

Empirical Evaluation of Structured Synthetic Data Privacy Metrics: Novel experimental framework

Dec 18, 2025
MN
Milton Nicolás Plasencia Palacios
🏛️ Aindo SpA | LAST-JD | Alma AI | Alma Mater Studiorum, University of Bologna | DIKE research group | PREC department | Vrije Universiteit Brussel | Luiss University | University of Trieste

Current evaluations of synthetic data privacy lack quantifiable, comparable metrics due to ambiguous privacy definitions and existing measures’ inability to reflect real-world disclosure risks. Method: We propose the first benchmark framework based on deliberate risk insertion—integrating legal theory with a black-box threat model—to enable reproducible, cross-method assessment of privacy-utility trade-offs. Our approach systematically controls perturbations, models diverse black-box attacks, maps outputs to regulatory compliance criteria, and validates findings on public datasets. Contribution/Results: Empirical evaluation reveals substantial discrepancies between mainstream privacy metrics (e.g., k-anonymity, differential privacy estimates) and actual re-identification risks under realistic attack scenarios. This work establishes the first evaluation paradigm for privacy-enhancing technologies (PETs) that is simultaneously interpretable, empirically grounded, and aligned with regulatory requirements—thereby bridging theoretical guarantees, practical security, and legal accountability.

Assessing privacy quantification methods through risk insertionBenchmarking tabular synthetic data privacy metricsEvaluating synthetic data privacy protection efficacy

This work proposes a GAN-inspired privacy-preserving synthetic data generation method that avoids direct access to original data during training. Instead, it leverages fuzz testing to produce candidate samples and iteratively refines them through a discriminator-guided feedback loop combined with statistical distribution constraints to approximate the original data distribution. By innovatively integrating fuzz testing, adversarial discrimination, and indirect constraint mechanisms, the approach achieves strong privacy guarantees—effectively resisting membership inference and data reconstruction attacks—while preserving high data utility. Extensive experiments on four benchmark datasets demonstrate that the proposed method strikes a superior balance between privacy protection and data fidelity compared to existing techniques.

data confidentialityprivacy preservationstatistical distribution

A Multi-Faceted Evaluation Framework for Assessing Synthetic Data Generated by Large Language Models

Apr 20, 2024
YY
Yefeng Yuan
🏛️ Santa Clara University | eBay Inc.

Existing research lacks a unified, quantitative framework to simultaneously evaluate the quality, downstream utility, and privacy preservation capabilities of large language model (LLM)-generated structured synthetic data—e.g., product reviews. Method: We propose SynEval, an open-source evaluation framework that introduces the first tri-dimensional quantitative paradigm integrating fidelity (statistical similarity), utility (performance on downstream ML tasks), and privacy (robustness against membership and attribute inference attacks). Contribution/Results: SynEval systematically characterizes inherent trade-offs among these dimensions through rigorous statistical analysis, task-based benchmarking, and adversarial privacy testing. Empirically validated on synthetic reviews generated by ChatGPT, Claude, and Llama, it delivers reproducible, interpretable insights for synthetic data selection and deployment—enabling principled, evidence-based decision-making in real-world applications.

Lack of comprehensive framework for synthetic data evaluationNeed to assess fidelity, utility, and privacy of synthetic dataPrivacy concerns in synthetic data generation using LLMs

Latest Papers

What's happening recently
View more

Current evaluation of synthetic tabular data lacks a unified framework, resulting in heterogeneous metrics and fragmented practices. To address this, we propose the first three-dimensional fidelity assessment paradigm integrating statistical distribution alignment, variable dependency preservation, and graph-structured representation learning. We implement this as SynthEval—a modular Python library supporting automated data-type inference, cross-domain benchmarking, and end-to-end interactive visualization report generation. SynthEval integrates principled methods including Wasserstein distance for marginal distribution comparison, Hilbert–Schmidt Independence Criterion (HSIC) for dependency quantification, GNN-based embedding similarity for relational structure fidelity, and structural graph modeling. Built with Plotly and Seaborn, it enables interactive exploratory analysis. Empirical validation across three heterogeneous domains—medical diagnosis, socioeconomic modeling, and cybersecurity—demonstrates significant improvements in evaluation consistency, interpretability, and robustness, particularly for high-cardinality categorical variables and high-dimensional time-series signals.

Addresses fragmented evaluation with a modular library for consistent benchmarking.Assesses data across diverse domains like healthcare, finance, and cybersecurity.Evaluates synthetic tabular data fidelity using statistical and structural metrics.

This study systematically evaluates the suitability and effectiveness of synthetic data across three canonical scenarios: data sharing, model training augmentation, and variance reduction in statistical estimation. By integrating formal modeling, theoretical analysis of generative models, and empirical case studies, the work presents the first comprehensive taxonomy of synthetic data applications and delineates their boundaries of applicability. The research elucidates both the potential and fundamental limitations of synthetic data in enhancing privacy preservation, model performance, and statistical stability. It further demonstrates that many existing or proposed use cases are misaligned with the intrinsic properties of synthetic data, thereby providing decision-makers with a principled theoretical framework to assess whether synthetic data is appropriate for addressing specific data availability challenges.

data augmentationdata sharingprivacy

This study addresses the challenge of sharing real-world educational data under strict privacy constraints, where existing differentially private synthetic data methods—often reliant on deep learning—are hindered by engineering complexity and limited practicality in small-sample, high-dimensional settings. The authors propose a training-free, two-stage framework: first, leveraging large language models to generate differentially private synthetic data for broad sharing and exploratory analysis; second, enabling on-demand validation of research findings on the original data through a secure remote code submission mechanism. Evaluated on three years of real educational data, the approach achieves synthetic data quality comparable to deep learning baselines while substantially reducing implementation overhead. Case studies show that approximately 36% of findings are reproducible on the真实 data, with the validation process introducing only negligible additional privacy loss.

differentially private synthetic data generationeducational data sharingprivacy constraints

This work addresses the risk that generative AI models may inadvertently leak private user data from their training sets through synthetic outputs, a concern exacerbated by the difficulty of distinguishing genuine privacy leakage from coincidental matches. To tackle this challenge, the paper introduces the first model-agnostic causal auditing framework that requires only synthetic outputs and a held-out reference dataset. By integrating statistical hypothesis testing with causal inference, the method rigorously differentiates true data leakage from spurious (“phantom”) matches. Applicable to any generative mechanism, the approach operates without shadow models or honeypot data, incurs computational costs orders of magnitude lower than existing techniques, and yields interpretable, tight lower bounds on privacy leakage.

data disclosuregenerative AImembership inference

This study addresses a critical gap in synthetic data generation (SDG) research, which has predominantly focused on privacy attacks initiated by data recipients while overlooking internal adversaries—such as data owners or generators—who may degrade data quality by perturbing real data. The work formally introduces this internal threat model and proposes targeted perturbation strategies based on label flipping and feature importance manipulation. Through systematic evaluation across multiple mainstream SDG frameworks, the experiments demonstrate that even minimal perturbations can substantially impair downstream task performance and amplify statistical distributional biases. These findings reveal a pronounced vulnerability in current SDG pipelines regarding data integrity and underscore the urgent need for robustness and integrity verification mechanisms in synthetic data workflows.

Adversarial ManipulationData IntegrityPrivacy-Preserving Data Sharing

Hot Scholars

HK

Hirokatsu Kataoka

AIST / University of Oxford
Computer VisionAction RecognitionAction PredictionVisual Pre-training
XX

Xiangyang Xue

Professor of Computer Science, Fudan University
Computer VisionPattern RecognitionMachine Learning
WC

Wuyang Chen

Assistant Professor, CS@Simon Fraser University
Scientific Machine LearningComputer VisionLarge Language ModelsReasoning
TA

Tejumade Afonja

CISPA Helmholtz Center for Information Security | AI Saturdays Lagos | BioRamp | Saarland University
Intersection of Privacy and Security in Machine Learning | African Accented Speech Researcher
MF

Mario Fritz

Faculty CISPA Helmholtz Center for Information Security; Professor Saarland University
Computer VisionMachine LearningTrustworthy AISecurity