lms-centile imputation

Designs and applies algorithms to impute missing anthropometric or growth-related measurements by estimating or assigning subject-specific centiles relative to a reference using the LMS (lambda–mu–sigma) transformation, then converting those centiles back to expected measurement values; methods include carrying observed centiles forward or backward within a subject and defaulting to a central centile (e.g., the 50th) when no subject-specific information is available.

lms-centileimputation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.06
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$162K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Longitudinal anthropometric data are often challenging to integrate due to missing values and inconsistencies between multiple growth reference standards, such as those from the WHO and CDC. This study proposes a two-stage imputation method: first, linear interpolation is applied to fill missing values between observed measurements within individuals; second, remaining gaps are imputed using age- and sex-specific LMS growth models, with reference standards assigned according to each datum’s original source, thereby estimating individual percentiles. This approach uniquely embeds growth references explicitly into the imputation pipeline, enabling auditable, self-contained reconstruction that preserves data provenance. Evaluated on synthetic data with 30% missingness, the method achieved mean absolute errors of 1.78 kg (3.5%) for weight and 2.84 cm (2.0%) for height—negligible biases—and restored 100% data completeness.

anthropometric measurementsdata harmonisationgrowth reference standards

Missing data arising from unit and item nonresponse in complex survey designs pose significant challenges for valid statistical inference. Method: This paper proposes a novel multiple imputation framework that integrates the original design weights—rather than synthetic or reweighted ones—with known marginal distributions of auxiliary variables. Specifically, it embeds the true design weights directly into an imputation model for nonignorable unit nonresponse, combining weighted generalized linear models with marginal constraint optimization to ensure calibration to population totals. Contribution/Results: Simulation studies demonstrate that the proposed method strictly satisfies auxiliary marginal constraints while substantially reducing bias and mean squared error in target parameter estimates. It improves estimator consistency and overall representativeness relative to existing approaches relying on artificial weights, thereby enhancing the validity and efficiency of survey inference under complex nonresponse mechanisms.

Extends multiple imputation for nonresponse in surveysIncorporates design weights for all sampled unitsLeverages known auxiliary margins for estimation

Some Simplifications for the Expectation-Maximization (EM) Algorithm: The Linear Regression Model Case

Sep 23, 2025
DA
Daniel A. Griffith
🏛️ University of Texas at Dallas

This paper addresses linear regression under missing-at-random (MAR) data by proposing an ANCOVA-equivalent reformulation of the EM algorithm that simplifies computation. The method recasts EM iterations as standard linear or nonlinear regression procedures, enabling constrained prediction and asymptotic variance estimation. Through rigorous theoretical derivation, we establish six theorems that— for the first time—unify the maximum likelihood estimation (MLE) consistency foundations of diverse imputation strategies, thereby enhancing interpretability and implementation flexibility. The approach is validated within the SAS PROC MI framework and applied to reanalyze 14 canonical datasets; imputation results match the gold-standard reference exactly, confirming its accuracy, robustness, and broad applicability.

Deriving analytical results for prediction and variance estimationExpressing EM solutions via regression models for imputationSimplifying EM algorithm for maximum likelihood with missing data

Current approaches to sample size calculation for clinical prediction models typically neglect the impact of missing data, often resulting in overfitting and poor calibration. This study is the first to integrate missing data mechanisms and handling strategies—such as multiple imputation—into a posterior-distribution-based sample size framework. Through simulation studies and Expected Value of Perfect Information (EVPI) analyses, the research quantifies how missingness affects model performance. Findings reveal that under common missing data scenarios, even when existing minimum sample size criteria are met, calibration slopes frequently fall below 0.9. In certain settings, nearly twice the conventional sample size is required to achieve performance comparable to that with complete data, underscoring both the necessity and feasibility of dynamically adjusting sample size requirements in the presence of missing data.

clinical prediction modelsmissing datamodel calibration

Imputing With Predictive Mean Matching Can Be Severely Biased When Values Are Missing At Random

Jun 28, 2025
PT
Paul T. von Hippel
🏛️ The University of Texas at Austin

This study identifies a severe and systematic estimation bias in predictive mean matching (PMM) under the missing-at-random (MAR) mechanism. Through theoretical analysis and extensive simulation experiments, we demonstrate that when an observed covariate (X) strongly predicts the missingness probability of outcome (Y) and is highly correlated with (Y), PMM yields regression slope estimates biased by up to 80%. This bias persists even as the (X)–(Y) correlation weakens and only approaches zero under large samples ((n = 1{,}000)) and the more restrictive missing-completely-at-random (MCAR) assumption. Compared to alternative imputation methods, PMM exhibits greater sensitivity to the missingness mechanism and requires substantially larger sample sizes to achieve acceptable accuracy. Our work provides the first systematic characterization of how PMM bias varies with missingness patterns and sample size. These findings challenge the widespread recommendation of PMM as a default imputation method and deliver critical methodological warnings for applied missing-data analysis.

PMM causes bias in missing-at-random data imputationPMM fails with correlated predictors of missingnessPMM requires large samples for unbiased results

Latest Papers

What's happening recently
View more

This study addresses the challenge of missing multimodal measurements and low inter-modality redundancy in clinical fetal growth assessment by proposing a linear Gaussian factor model that leverages marginalization rather than imputation. The model jointly represents four classes of fetal and maternal data, naturally encoding observation completeness and yielding uncertainty-aware latent representations. Factor dimensionality is determined via parallel analysis, with K=8 factors subjected to VARIMAX rotation, and anomalous records are flagged using standardized residuals. Empirical results demonstrate near-independence across modalities (cross-modal R² = 0.023), a significant correlation between latent dimensions and birth weight percentile (ρ = 0.55), and an AUC of 0.70 for predicting small-for-gestational-age outcomes along a hemodynamic axis. The model achieves nominal 97% confidence interval coverage and successfully identifies 36 data entry errors.

data incompletenessfetal growth analysismissing data

This study addresses a critical limitation of deterministic imputation methods based on minimizing mean squared error (MSE), which, despite yielding accurate point estimates, systematically underestimate data variability and thereby introduce bias into downstream statistics such as variance, correlation, and regression coefficients. To rectify this, the authors propose a stochastic imputation strategy that augments MSE-optimal predictions with random noise scaled to the residual error variance, thereby restoring the original distributional properties of the data. Through multivariate normal simulations, the work demonstrates for the first time the inadequacy of MSE as a sole imputation quality metric and reveals pervasive bias in widely used predictive imputation methods—including missForest, softImpute, and MICE. The proposed stochastic approach effectively eliminates this bias, ensuring statistical validity in subsequent analyses and advocating a paradigm shift from deterministic to stochastic imputation.

data variabilitydownstream analysisimputation bias

This study systematically investigates how missingness mechanisms, missing rates, missingness locations, and sample sizes jointly affect the accuracy and identifiability of average treatment effect (ATE) estimation under time-varying confounding. Through simulation, it compares the performance of complete-case analysis, stratified hot-deck imputation, single-model imputation, and multiple imputation by chained equations (MICE) combined with propensity score weighting, offering the first comprehensive assessment of their interactive effects. The results demonstrate that multiple imputation substantially reduces bias and improves confidence interval coverage across most scenarios. Crucially, the missingness mechanism emerges as a key determinant of estimator performance: missing not at random (MNAR) conditions, high missing rates, or small sample sizes frequently violate the positivity assumption, thereby undermining estimation validity.

average treatment effectidentifiabilitymissing data

Estimating the functional relationship between a continuous exposure and a binary outcome is challenging when covariates are measured with error. This study presents the first systematic evaluation of Simulation-Extrapolation, Regression Calibration, multiple imputation, and Bayesian correction methods, each coupled with flexible modeling techniques—including B-splines, P-splines, and fractional polynomials—within a multi-team, fully blinded, neutral simulation framework. By generating 155 distinct simulation scenarios and repeated samples, the research quantifies the bias and variance of each approach, revealing their relative strengths and limitations. The findings not only inform method selection under measurement error but also demonstrate the feasibility and value of this neutral comparative paradigm for rigorous methodological assessment.

covariate adjustmentexposure-outcome relationshipfunctional form

Hot Scholars

CX

Chunjing Xiao

Henan University
Anomaly DetectionActivity RecognitionDiffusion Model
KC

Kevin Chetty

Professor, University College London
Wireless SensingRadar and Security Technologies
FG

Florian Grensing

Helmut-Schmidt-Universität
Physiological SignalsVirtual RealityEmotion Recognition
BC

Beyza Cinar

PhD Student in Data Engineering
AI in MedicineData ScienceDigital TwinsDigital Health