Score
Designs and implements weighted-sample estimators, resampling and acceptance/rejection samplers, and evaluation metrics that reweight observations to correct distribution mismatch or selection bias and to approximate expectations under a specified target distribution. This includes constructing importance-sampling and importance-weighted estimators, deriving and applying weight formulas from likelihood ratios or proxy models, and analyzing bias–variance tradeoffs and variance-reduction methods (e.g., control variates, resampling schemes, constrained weighting) to guarantee unbiasedness or controlled variance under limited data, budget, or sampling constraints.
In observational causal inference, weighting methods mitigate covariate imbalance but often inflate variance estimates and yield overly conservative standard errors. This paper proposes augmenting weighted regression with main effects of covariates and their interactions with the treatment variable, integrated with residualization and parametric model augmentation to form a unified inferential framework. We establish, for the first time under design-based, model-based, and finite-sample-corrected superpopulation sampling assumptions, that this approach yields asymptotically valid and more precise standard errors. Theory, simulations, and multiple empirical applications demonstrate substantially narrower confidence intervals—on average 15–30% shorter—with improved inferential accuracy and robustness to both exact and approximately balanced weights. The key innovation lies in achieving simultaneous gains in statistical efficiency and asymptotic validity at minimal variance cost.
This paper addresses the problem of biased predictions by machine learning models against marginalized groups in real-world data. To jointly optimize predictive accuracy and fairness, we propose a genetic algorithm-based sample weighting method that evolves instance-level weights through multi-objective optimization. Unlike conventional uniform or feature-driven weighting schemes, our approach simultaneously optimizes accuracy, AUC, demographic parity difference, and subgroup false negative rate. Extensive experiments on 11 publicly available datasets—including two healthcare benchmarks—demonstrate that the evolved weights substantially improve the fairness–performance trade-off. The most significant gains are achieved when jointly optimizing for accuracy and demographic parity difference, confirming the method’s effectiveness and generalizability in practical, high-stakes domains.
Slice sampling suffers from low efficiency and heavy reliance on manual tuning when applied to complex target distributions—such as highly skewed or constrained spaces. To address this, we propose Quantile Slice Sampling (QSS), a novel framework that (1) integrates probability integral transformation with quantile mapping to enable automatic initialization and unit-interval standardization; (2) introduces an evaluable pseudo-target importance reweighting mechanism, coupled with dual-metric quality assessment and adaptive parameter optimization; and (3) extends slice sampling to multivariate and constrained state spaces by incorporating elliptical slicing, Neal’s shrinkage, and Gibbs-like coordinate updates. Experiments on benchmark distributions and Bayesian modeling tasks demonstrate that QSS significantly outperforms conventional slice sampling and Metropolis–Hastings: in highly skewed and constrained settings, it reduces rejection rates by over 30%, while delivering enhanced robustness, full automation, and practical usability.
To address nonignorable nonresponse in sample surveys, this paper extends model-assisted estimation to the missing-at-random (MAR) framework. We propose a calibratable inverse-probability weighting (IPW) method that reweights sampled units in a second stage to compensate for nonrespondents, and systematically construct a Horvitz–Thompson-type adjusted estimator. Theoretically, we establish its asymptotic design-unbiasedness and design-consistency, derive a closed-form asymptotic variance expression, and provide a consistent variance estimator. Monte Carlo simulations demonstrate that the proposed estimator significantly outperforms the conventional Horvitz–Thompson estimator under diverse nonresponse mechanisms. Our key contributions are: (i) the first systematic adaptation of model-assisted estimation to the MAR setting; and (ii) a novel IPW weighting scheme that simultaneously satisfies calibration constraints and enjoys rigorous asymptotic properties—namely, design-consistency, asymptotic normality, and consistent variance estimation.
This paper addresses the “weak paradox” of inverse probability weighting (IPW) estimators—highlighted by Basu (1988) and Wasserman (2004)—in survey sampling, causal inference, and Bayesian evidence estimation. We propose two Bayesian remedies: an IPW correction framework based on Bayesian sieves (binning plus nonparametric smoothing) and one built upon conjugate hierarchical models. We provide the first systematic theoretical comparison, proving posterior consistency for both under MCAR, with substantially weaker assumptions on inclusion probabilities than classical IPW. Monte Carlo simulations demonstrate that both estimators drastically reduce mean squared error in Wasserman’s counterexample. Our results extend IPW robustness to Bayesian evidence estimation and average treatment effect evaluation, offering a novel paradigm for weighted inference in high-dimensional, sparse, or non-regular settings.
This study addresses the high computational cost of Markov chain Monte Carlo (MCMC) steps commonly employed in sequential Monte Carlo (SMC) and approximate Bayesian computation (ABC)-SMC to control the variance of importance weights. For the first time, it systematically integrates Pareto smoothed importance sampling (PSIS) into the SMC framework, leveraging a fitted generalized Pareto distribution to adjust the tails of the importance weights and thereby reduce their variance, with the aim of diminishing reliance on MCMC. However, empirical analysis reveals that because SMC inherently mitigates weight degeneracy through its sequence of intermediate target distributions, the additional variance reduction offered by PSIS is limited. This finding challenges the presumed necessity of PSIS in this context and provides new theoretical and empirical insights into strategies for stabilizing importance weights in SMC algorithms.
This study addresses the issue of variance inflation in regression models under complex survey designs, which often arises from unnecessary variability in sampling weights. The authors propose a novel approach that, for the first time, integrates stabilized weights with generalized raking within a two-stage sampling framework, leveraging auxiliary covariate information to effectively reduce extraneous weight variation. This method substantially enhances the efficiency of design-based estimators while remaining compatible with standard statistical software. Simulation studies demonstrate that, under typical two-stage survey designs, the proposed estimator achieves markedly higher precision compared to existing methods. The approach has been successfully applied to a large-scale multinational study of Kaposi’s sarcoma, illustrating its practical utility and robustness in real-world settings.
This work addresses the challenge of risk prediction when acquiring true outcomes is prohibitively expensive, allowing labels for only a subset of samples. The authors propose a surrogate-assisted optimal sampling framework that, under a fixed annotation budget, leverages covariates, surrogate variables, and an initial estimator to construct a sampling strategy minimizing the expected out-of-sample cross-entropy loss. Coupled with an inverse probability weighted cross-entropy estimator for model training, this approach achieves—without requiring access to true responses during design—theoretically optimal sampling. It uniquely guarantees predictive optimality, robustness to surrogate misspecification, and stability in settings with rare outcomes. Both theoretical analysis and empirical experiments demonstrate that the method significantly outperforms existing approaches, particularly when the surrogate is imperfect or the event of interest is rare.