Score
Designs and implements statistical models, estimators, and adjustment procedures to correct for measurement error—including non‑classical errors—using error‑in‑variables techniques and validation or gold‑label data. Analyzes estimator properties (bias, probability limits, and efficiency), quantifies the direction and magnitude of bias, and develops calibration and debiasing procedures that adjust predictions and retain efficiency in large‑scale settings.
Conventional statistical inference suffers from bias under non-Gaussian measurement errors, violating the classical Gaussian error assumption. Method: This paper proposes a novel debiasing framework grounded in hypercomplex algebra—marking the first application of hypercomplex numbers to measurement error modeling. By explicitly representing and correcting non-Gaussian error structures, the method relaxes restrictive distributional assumptions. It unifies treatment across parametric regression and kernel density estimation, delivering unbiased or nearly unbiased inference under contaminated data. Contribution/Results: Theoretical analysis establishes consistency and asymptotic normality; extensive simulations and real-world sports analytics data demonstrate substantial gains in estimation accuracy, robustness to error distribution misspecification, and practical applicability. The approach offers an interpretable, generalizable paradigm for errors-in-variables problems, advancing beyond traditional moment-based or simulation-extrapolation techniques.
Measurement error in heteroscedastic continuous exposures, across multiple time points and calibration settings, substantially biases sample size planning, induces estimation bias, and distorts standard error accuracy in distributed lag models—yet existing methods lack a unified framework to address these challenges. This paper introduces the first analytical approximation framework that jointly accommodates nonlinear exposure–response relationships, diverse measurement error structures (differential vs. nondifferential; additive vs. multiplicative; temporally autocorrelated), and multitemporal calibration with or without validation data. Leveraging error propagation theory and Taylor series expansion, integrated with polynomial effect modeling and heteroscedasticity-robust inference, we derive closed-form expressions for required sample size, bias correction, and standard error adjustment. The proposed method markedly improves estimation accuracy and statistical power, as demonstrated through comprehensive simulations and application to real-world environmental exposure studies.
This study addresses the inconsistency of conventional estimators in mixed-data sampling (MIDAS) regression when both high- and low-frequency variables are subject to measurement error. To resolve this issue, the paper introduces the corrected score method into the MIDAS framework for the first time and combines it with profile likelihood to construct a consistent estimator. This approach effectively overcomes the inconsistency that plagues existing profile likelihood estimators under measurement error. Through comprehensive Monte Carlo simulations, the authors systematically investigate the impacts of sample size, lag order, and nuisance parameters on estimation performance. The results demonstrate that the proposed estimator exhibits strong consistency and favorable finite-sample properties across a range of sample sizes and model specifications.
This paper addresses the bias in linear regression parameter estimation arising from misclassification errors in categorical covariates. We propose an asymptotically bias-corrected estimator that requires neither access to true covariate observations nor data recalculation. The method integrates the least-squares estimator based on noisy covariates, the marginal distribution of the true covariates, and the misclassification transition probability matrix, explicitly modeling and eliminating systematic bias induced by measurement error—particularly improving consistency of the intercept estimate. Theoretical analysis establishes the consistency and asymptotic normality of the corrected estimator. Simulation studies confirm its effectiveness across diverse misclassification structures, significantly reducing parameter bias, enhancing estimation accuracy, and increasing statistical power. The key innovation lies in achieving an analytical correction of classical linear regression estimates using only prior knowledge of the misclassification mechanism—specifically, the conditional misclassification probabilities—without requiring additional data or iterative procedures.
Estimating the functional relationship between a continuous exposure and a binary outcome is challenging when covariates are measured with error. This study presents the first systematic evaluation of Simulation-Extrapolation, Regression Calibration, multiple imputation, and Bayesian correction methods, each coupled with flexible modeling techniques—including B-splines, P-splines, and fractional polynomials—within a multi-team, fully blinded, neutral simulation framework. By generating 155 distinct simulation scenarios and repeated samples, the research quantifies the bias and variance of each approach, revealing their relative strengths and limitations. The findings not only inform method selection under measurement error but also demonstrate the feasibility and value of this neutral comparative paradigm for rigorous methodological assessment.
This study addresses the unreliable estimation of repeatability, between-laboratory, and reproducibility variance components under ISO 5725 standards when sample sizes are small or variance structures are extreme. To overcome this limitation, the authors propose a tailored Bootstrap resampling strategy adapted to a one-way random effects model. The approach refines point estimates by adjusting within-laboratory resampling and constructs confidence intervals via a two-stage resampling scheme integrated with bias-corrected and accelerated (BCa) techniques. Extensive simulations and validation using real data from ISO 5725-4 demonstrate that the proposed method substantially improves estimation accuracy and confidence interval coverage. It yields reliable, near-nominal or conservatively valid inferences for small- to moderate-sized experiments and clearly delineates optimal strategies across different practical scenarios.
This study addresses attenuation bias in scalar-on-density regression when the number of repeated measurements per observational unit is limited, leading to insufficient effective sample size for accurate coefficient function estimation. The work systematically establishes, for the first time, a monotonic decreasing relationship between the number of measurements and the magnitude of attenuation bias. To correct this bias, the authors innovatively integrate the Simulation-Extrapolation (SIMEX) method with bootstrap resampling within a functional data analysis framework, simulating scenarios with fewer measurements and extrapolating estimates to the theoretical limit of infinite replicates. Combining techniques from functional data analysis and density estimation, the proposed approach substantially reduces estimation bias in simulations. Applied to NHANES data, it successfully identifies and corrects finite-measurement bias in the association between physical activity density and all-cause mortality.
This study addresses the lack of optimal experimental design guidance for the standard addition method under non-decreasing measurement error structures. Building on c-optimality theory and integrating linear response modeling with analysis of variance, the authors systematically derive an optimal two-concentration-point design that minimizes estimation variance under constant, linear, or quadratic error growth. This work represents the first application of optimal experimental design theory to the standard addition method and demonstrates that the proposed two-point design achieves universal optimality across all considered non-decreasing error scenarios. Notably, the optimal allocation of replicate measurements deviates from the conventional 50:50 ratio and yields minimum-variance unbiased estimates without requiring weighted regression.
This study addresses the limitations of traditional Gaussian process (GP) calibration methods, which neglect intermediate variables in computer experiments, leading to inadequate bias modeling and non-identifiability between the simulator and the discrepancy term. To resolve this, the authors propose a robust GP calibration framework that explicitly incorporates intermediate variables. The approach systematically selects key intermediate variables, constrains the discrepancy term using a scaled Gaussian stochastic process (S-GaSP), and employs space-filling designs to choose constraint points, thereby enabling identifiable joint modeling of the simulator and bias. This work is the first to systematically integrate intermediate variables into GP calibration, substantially improving predictive accuracy and the reliability of uncertainty quantification. Empirical results on nuclear binding energy prediction demonstrate clear superiority over existing baseline methods.