Score
Designs and builds synthetic diagnostic testbeds and evaluation protocols that use generative models or procedurally created data to produce controlled attribute variations and perturbations for assessing model behavior. Analyzes the resulting performance trends to identify failure modes, quantify robustness and likely real-world performance gaps, and to prioritize or guide targeted data collection and remediation.
Accurately estimating model test error under scarce labeled data remains challenging. Method: This paper proposes a novel error estimation paradigm leveraging high-quality synthetic data. Theoretically, we derive a new generalization error upper bound incorporating generator quality constraints, quantifying for the first time how generative model fidelity critically affects estimation bias. Methodologically, we design an interpretable and optimization-friendly synthetic sample construction strategy that jointly leverages generative modeling and generalization theory to enhance assessment reliability. Results: Extensive experiments on both synthetic and real-world tabular datasets demonstrate that our approach consistently outperforms existing baselines, achieving significant and robust improvements in both accuracy and stability of error estimation.
This work addresses the risks of bias in statistical inference when using synthetic data generated by modern generative AI models—such as diffusion models, GANs, and large language models—due to model misspecification, underestimation of uncertainty, and insufficient generalization. It presents the first systematic integration of generative modeling with statistical inference theory, clarifying the assumptions and conditions under which synthetic data can reliably support scientific discovery. By unifying uncertainty quantification, model diagnostics, and downstream task analysis, the study proposes a principled framework that delineates effective usage guidelines, identifies critical failure modes, and offers practical recommendations for researchers and developers, along with directions for future research.
This study addresses the limitations of traditional Monte Carlo simulations, which rely on ad hoc assumptions and struggle to generate data reflecting realistic multilevel structures, thereby compromising the validity of quantitative method evaluations. To overcome this, the authors propose the first six-stage workflow integrating generative AI with multilevel data simulation. They innovatively adapt diffusion models and generative adversarial networks (GANs) to accommodate hierarchical data structures and introduce a comprehensive synthetic data quality assessment framework that ensures both within-table and cross-table consistency. Empirical experiments on real-world social science datasets demonstrate that the proposed approach substantially enhances the realism and reliability of Monte Carlo simulations, outperforming conventional strategies and providing a more empirically grounded benchmark for evaluating predictive performance and parameter recovery in quantitative methods.
Do generative models inevitably suffer “model collapse” during large-scale pretraining with early-stage synthetic data? This paper systematically compares three synthetic-data training paradigms—replacement, accumulation, and constrained subset iteration—across Gaussian estimation, kernel density estimation, and language model fine-tuning. Methodologically, it introduces a generational iterative training framework, a multi-task benchmark suite, and dynamic test-loss modeling. Results demonstrate that “accumulation + full-dataset training” completely avoids collapse (test loss remains stable), whereas “constrained subset iteration” induces progressive performance degradation, and pure replacement inevitably collapses. These findings refute the monolithic assumption that synthetic data inherently causes collapse, instead establishing “data-evolution path dependence” as a new paradigm and empirically delineating safe operational boundaries for synthetic-data utilization.
Medical causal inference is hindered by the scarcity of real-world clinical data, while existing synthetic data generation methods fail to preserve—specifically for treatment effect estimation—the essential characteristics of covariate distributions, treatment assignment mechanisms, and outcome generation mechanisms. To address this, we propose STEAM, the first framework that systematically models the triple generative mechanism underlying therapeutic data and introduces a dedicated evaluation metric suite tailored for causal inference tasks. STEAM integrates generative modeling with structural causal models, explicitly optimizing fidelity in both treatment and outcome mechanisms. Extensive experiments demonstrate that STEAM significantly outperforms state-of-the-art baselines across multiple causal evaluation metrics—particularly under challenging settings involving high-dimensional covariates, nonlinear relationships, and strong confounding dependencies.
Current evaluations of synthetic medical data predominantly rely on statistical similarity and predictive performance, which often fail to capture clinical validity. This work proposes an epidemiology-informed, multidimensional evaluation framework that systematically assesses generative models on structured electronic health records across three dimensions: descriptive fidelity, clinical utility, and structural validity. Applying this framework to a real-world PRIME-CVD cohort comprising 50,000 individuals, the authors empirically compare four model families—GANs, VAE-based approaches, diffusion models, and masked modeling—and find that even models achieving high distributional fidelity exhibit miscalibration and distorted inter-variable relationships. These findings reveal that conventional evaluation metrics tend to overestimate synthetic data quality and underscore the necessity of shifting toward domain-driven validation of clinical effectiveness.
This work addresses the lack of fine-grained diagnostic tools for evaluating the performance limitations of existing aerial object detectors in complex scenes. It introduces, for the first time, large-scale text-to-image generative models into the diagnostic pipeline of aerial detection systems, establishing a controllable synthetic testing platform. Through text-guided image generation, attribute-controllable editing, and automated validation, the framework enables systematic evaluation of pretrained vehicle detectors. The approach accurately predicts real-world performance deficiencies and effectively guides targeted data collection: augmenting the training set with only a small amount of carefully selected real data improves AP50 by up to 13%, substantially outperforming non-directed augmentation strategies. The modular and extensible design of the framework establishes a new paradigm for robustness analysis in aerial vision systems.
While generic generative synthetic data often performs well in predictive tasks, it struggles to preserve causal estimands such as the average treatment effect (ATE). This work proposes a hybrid synthesis framework that decouples covariate generation from the treatment–outcome mechanism, constructing (W, A, Y) triplets by integrating nearest-neighbor distance diagnostics with an independently learned interference model, and further introduces targeted synthetic augmentation to mitigate positivity violations. The study is the first to systematically uncover the failure mechanisms of generative models in causal estimation, advocates for a synthesis strategy that separates covariates from causal mechanisms, and develops a synthetic simulation engine tailored for evaluating causal estimators under limited sample settings. Experiments demonstrate that the proposed approach substantially improves ATE fidelity over fully generative baselines and provides practical diagnostic tools, with consistent effectiveness validated across diverse configurations.
This work addresses the challenge of efficiently and legally auditing high-risk language models at scale for harmful specialization—such as generation of child sexual abuse material (CSAM)—without producing illicit content. The authors propose a novel non-generative evaluation paradigm that detects harmful specialization by analyzing internal model states, specifically perturbations in intermediate representations induced by LoRA adapters. Leveraging Gaussian probing techniques, the method quantifies changes in internal representations through Gaussian latent ensembles. Experimental results demonstrate that this approach reliably distinguishes between benign and harmful models in CSAM-related tasks and exhibits robustness against adversarial interventions such as weight scaling.
Existing approaches struggle to effectively quantify the similarity and quality between synthetic and real data in evaluating tool-augmented agents. To address this gap, this work proposes SynAE, a novel framework that establishes the first multi-axis evaluation system tailored for multi-turn tool-use scenarios. SynAE introduces four fine-grained metric categories—assessing task instructions, tool invocations, final outputs, and downstream evaluation performance—to systematically measure synthetic data across dimensions of validity, fidelity, and diversity. Integrating natural language processing, trajectory modeling, and controllable generation techniques, the framework enables a reproducible evaluation pipeline and successfully identifies several representative failure modes in synthetic data generation. Empirical results demonstrate that such multidimensional assessment is essential for enhancing the reliability of agent evaluations.