Score
Designs and implements methods and tools to estimate, model, compare, and visualize probability distributions from data, including parametric and nonparametric density estimators, empirical distribution estimators, and procedures for sampling from high-density regions. Builds and evaluates distributional metrics and divergence measures, constructs distribution-matching and regularization techniques, and develops detectors and estimators for distribution shift plus tests and visualizations to quantify and monitor distributional similarity and change.
This study addresses the fundamental statistical problem of conditional density estimation by systematically comparing classical nonparametric approaches—such as single-index models, basis expansion methods (e.g., FlexCode), and DeepCDE—with modern generative models, including conditional GANs and conditional denoising diffusion probabilistic models. For the first time, these methods are evaluated within a unified and reproducible framework using metrics like mean squared error and Wasserstein distance to assess their accuracy, flexibility, and computational cost in estimating conditional means and standard deviations. The analysis clarifies the performance boundaries and practical applicability of each approach, offering both theoretical insights and actionable guidance for selecting appropriate methods in predictive modeling, uncertainty quantification, and probabilistic inference tasks.
Distribution shift in tabular data lacks empirical grounding, and existing robust learning methods—such as distributionally robust optimization (DRO)—frequently fail in real-world settings. Prior work predominantly assumes covariate shift (X-shift), whereas empirical analysis reveals that label-conditional shift (Y|X-shift) is both more prevalent and dominant in practice. Method: The authors construct a comprehensive benchmark platform spanning five real-world tabular datasets and 60,000 experimental configurations, enabling large-scale ablation studies on robustness. Contribution/Results: (1) Theory-driven methods like DRO offer no consistent advantage over empirical risk minimization (ERM); (2) implementation details—including model selection, hyperparameter tuning, and imbalance handling—exert far greater influence on robustness than the design of ambiguity sets; (3) a data-driven, inductive paradigm should supplant prior structural assumptions. These findings challenge the conventional paradigm in distributionally robust learning and provide reproducible, empirically grounded guidance for co-optimizing algorithms and data.
Existing methods struggle to align and interpret distribution shifts across heterogeneous, domain-consistent datasets—such as tabular, textual, visual, and time-series data—especially when scale and modality disparities are pronounced, resulting in poor interpretability. This paper introduces the first human-centric, cross-modal distribution discrepancy explanation framework, implemented as an interpretable dataset comparison toolbox. It integrates statistical hypothesis testing, feature importance decomposition, class activation mapping (CAM), contrastive representation learning, and interpretable generative modeling to enable fine-grained, semantically readable attribution and visualization of distributional shifts. Evaluated across diverse real-world scenarios, the framework significantly improves users’ efficiency in understanding shift causes and enhances the accuracy of intervention decisions—thereby overcoming the limitations of conventional black-box shift detection approaches.
Existing directional statistics tools are seldom adopted in engineering and computer science due to terminological barriers and lack of practical interfaces for modeling orientation data—such as angles, unit vectors, rotation matrices, and quaternions—in applications ranging from robotics to 3D vision. Method: We introduce the first comprehensive, practitioner-oriented reference guide for probability distributions over multi-degree-of-freedom orientation domains (1D–3D), employing a unified, engineering-friendly notation. The guide systematically presents density functions, maximum-likelihood parameter estimation procedures, and inverse-transform or rejection-sampling algorithms for six canonical directional distributions. Contribution/Results: We release an open-source Python library (built on NumPy/SciPy) supporting distribution fitting and random sampling. Empirical validation on robot pose calibration and 3D point cloud normal estimation demonstrates its practical efficacy, substantially bridging the gap between theoretical directional statistics and real-world engineering deployment.
Diffusion models unexpectedly generate cartoonized or blurry images—nonexistent in training data—within high-density regions of the learned distribution. Method: We propose Mode Tracking Theory to precisely localize modes in the diffusion denoising distribution; design a zero-overhead SDE likelihood tracking method that estimates and optimizes sample likelihood without additional computation; and develop an efficient high-density sampler that targets atypical, high-likelihood samples overlooked by conventional samplers. Results: Experiments demonstrate substantial improvement in sampling likelihood, stable generation of cartoon/blurry high-density images, and faithful reproduction of this phenomenon on purely real-image datasets—revealing an intrinsic, implicit structural bias inherent to diffusion models.
This work addresses the challenge that traditional discrete probability distributions rely on manually derived analytical forms, hindering the automatic discovery of interpretable models. We propose Symbolic Density Estimation (SDE), a novel framework that, for the first time, integrates structural priors, evolutionary search, and validity-aware parameter inference to automatically discover closed-form probability mass functions within a structured symbolic space composed of elementary mathematical operations. SDE accommodates complex distributional features such as zero-inflation and finite mixtures. We introduce the first systematic benchmark dataset for this task and demonstrate that SDE accurately recovers all target distribution families. On real-world data, SDE discovers concise, interpretable mixture models that achieve superior goodness-of-fit compared to standard methods.
This work addresses the challenge of achieving statistically sound machine unlearning over entire data domains—such as those containing biased or copyrighted content—while preserving performance on the target task. The authors propose a distribution-level unlearning framework that models forgettable and retainable domains as probability distributions. By leveraging hypothesis testing, the method identifies an optimal subset for removal and characterizes both the fundamental regions of editable distributions and the removal–retention Pareto frontier. Combining analyses across parametric and nonparametric distribution families—including Gaussian, Poisson, and log-concave noise models—the approach provides finite-sample guarantees of Pareto optimality, elucidates the information–computation gap, and demonstrates robustness and interpretability in empirical validation.
This study addresses the two-sample testing problem without distributional assumptions. It proposes PReLU-TST, a novel approach based on integral probability metrics (IPMs), which constructs a nonparametric test statistic using a parametrized discriminator consisting of only a single neuron. This design retains the flexibility of nonparametric methods while substantially improving computational efficiency. The proposed test is proven to be consistent and asymptotically equivalent to classical nonparametric IPM-based tests. Empirical evaluations demonstrate that PReLU-TST achieves higher or at least comparable finite-sample testing power against existing methods across a range of synthetic and real-world datasets.