Score
Modeling and analyzing how varying doses affect outcomes (including non-linearities like hysteresis), validating and calibrating simulators against biological and clinical dose‑response data, and assessing finite‑sample performance and minimum effective doses.
To address the lack of standardized frameworks for validating tumor growth and treatment response prediction models, this study proposes the first multi-level validation paradigm applicable across both preclinical and clinical settings, integrating patient-specific data, experimental constraints, and prospective cohort validation. Methodologically, it unifies ordinary differential equation (ODE), partial differential equation (PDE), and agent-based modeling approaches, augmented by AIC/BIC-based model selection, Bayesian parameter estimation, uncertainty quantification, and counterfactual simulation analysis. Key contributions include: (i) systematic identification of validation bottlenecks; (ii) establishment of a reproducible validation workflow and standardized evaluation metric suite; and (iii) substantial improvement in predictive reliability and model interpretability. Collectively, this framework provides a rigorous methodological foundation for regulatory review—by agencies such as the FDA and NMPA—of AI-driven oncology prediction tools.
This study addresses the challenge of disentangling sources of inter-laboratory variability—specifically baseline offsets versus differences in sensitivity—in multi-laboratory assessments of linear dose–response relationships. To this end, the authors propose a precision evaluation framework based on linear mixed-effects models, integrating analysis of variance, F-tests, and ISO 5725 standards to define and estimate repeatability and between-laboratory variance components. Overall measurement precision is quantified via average dose-specific variance. Under a fully balanced design, the framework yields an exact decomposition of total sum of squares and closed-form ANOVA estimators, overcoming the limitation of conventional fixed-effects models that detect only the presence of differences without identifying their origin. The approach was successfully applied to bronchoalveolar lavage fluid data from a rat intratracheal instillation study involving nanomaterials, effectively distinguishing the sources of observed variability.
Current sample size calculations for clinical prediction models ensure only that performance metrics meet target values *on average*, neglecting sampling variability—resulting in unstable model performance and low probability of achieving acceptable performance (PrAP) in practice. This paper proposes a novel sample size determination framework centered on PrAP, formally adopting “probability of attaining acceptable performance” as the primary statistical objective—replacing conventional expectation-based approaches. Through simulation studies and analytical derivations, we develop robust methods for estimating calibration slope in binary outcome settings, implemented in the R package `samplesizedev`. Results demonstrate that conventional methods yield PrAPs typically below 60%, whereas our approach consistently achieves PrAPs exceeding 80%, with particularly pronounced gains when fewer predictors are included. This substantially improves model reliability and reproducibility.
In early-phase oncology trials, patient heterogeneity is often unknown a priori, making prespecification of subpopulations infeasible. To address this, we propose a two-stage adaptive dose optimization design: Stage I defines a safe dose set using toxicity data; Stage II jointly models efficacy, baseline covariates, and pharmacokinetic exposure via Bayesian sparse group selection to dynamically identify unknown heterogeneous subpopulations and recommend personalized optimal doses. Our key innovation is the first implementation—within a Phase I trial framework—of fully data-driven, prespecification-free subpopulation identification, integrated with exposure-toxicity modeling and futility assessment to adaptively enrich the target population. Simulation studies demonstrate that the method accurately recovers critical heterogeneity-driving biomarkers, distinguishes multiple biologically distinct subgroups, and assigns differential optimal doses—outperforming conventional designs relying on prespecified subpopulations.
This paper addresses nonparametric estimation and valid inference for causal dose–response curves under continuous interventions: it seeks to avoid bias from parametric model misspecification while overcoming the slow convergence rates of conventional nonparametric methods caused by pathwise non-differentiability. To this end, we propose the first plug-in estimator that embeds the highly adaptive lasso (HAL) maximum likelihood estimator within a marginal structural framework, augmented with undersmoothing and smoothness-adaptive fitting. The resulting estimator for the marginal dose–response curve is consistent and semiparametrically efficient. It imposes no differentiability or linearity assumptions on the dose–response function and permits robust standard error construction. Simulation studies demonstrate that the proposed method achieves substantially higher estimation accuracy and nominal coverage of confidence intervals compared to existing approaches, with excellent finite-sample performance.
This work proposes the shared Keyboard design, a novel model-assisted approach for phase I oncology trials that overcomes the limitation of conventional Keyboard designs by incorporating information from neighboring doses. By introducing a Beta kernel process, the method enables controlled borrowing of strength across doses through kernel-weighted pseudo-counts, while preserving the original decision framework. It further supports asymmetric kernels to enhance overdose control. The approach naturally extends to adaptive dose insertion and time-to-event toxicity outcomes. Simulation studies demonstrate that the shared Keyboard design significantly improves both accuracy and safety in identifying the maximum tolerated dose compared to existing methods, achieves higher efficiency in locating the target dose under dose-insertion scenarios, and maintains comparable modification rates.
Phase II dose selection is often hindered by model misspecification risk—e.g., MCP-Mod requires prespecifying candidate dose–response models—thereby compromising Phase III success. Method: We propose MAP-curvature, a Bayesian framework that avoids prespecifying parametric models. It employs curvature-penalized priors for flexible nonlinear modeling and embeds a sigmoid Emax structure to better capture pharmacological S-shaped responses. The method innovatively integrates model-agnostic flexibility with Bayesian hierarchical modeling, enabling robust borrowing from historical data. Posterior inference is performed via MCMC, balancing curve smoothness and biological plausibility. Contribution/Results: Simulation studies demonstrate that MAP-curvature significantly outperforms MCP-Mod and linear models under S-shaped scenarios, yielding higher statistical power, improved detection of dose–response signals, and more accurate estimation of the minimum effective dose—thereby enhancing the reliability of Phase II-to-Phase III translation.
In A/B testing, control variates and regression adjustment are widely used variance reduction techniques, yet their theoretical relationship remains unclear, their methodological frameworks are disjointed, and both have long been confined to design-driven paradigms. Method: This paper establishes, for the first time, a formal equivalence between these two approaches and proposes a novel grouped coefficient estimation method that unifies design-based and model-based estimation frameworks—enabling a paradigm shift from design-driven to model-driven inference. Contribution/Results: Theoretical analysis demonstrates improved estimation accuracy and statistical power. Empirical validation on millions of real-world experiments at ByteDance confirms efficacy: the proposed method has been fully deployed in its online experimentation platform, yielding an average 12.3% increase in statistical significance and a 19.6% improvement in detection sensitivity.
To address the modeling challenges posed by nonlinear dynamics and genuine lag effects in count time series, this paper proposes the Threshold Lagged Poisson Autoregression (HPART) model. HPART incorporates a piecewise-linear threshold mechanism to capture regime shifts and explicitly introduces scientifically interpretable lagged control factors—thereby relaxing the oversimplified lag-structure assumptions of conventional BPART models and enabling more accurate characterization of lag-driven dynamic evolution. Parameter inference, model selection, and asymptotic theory are systematically established via maximum likelihood estimation, non-nested hypothesis testing, and Monte Carlo simulation. Empirical evaluations across multiple real-world datasets demonstrate that HPART significantly improves out-of-sample forecasting accuracy, while maintaining statistical rigor and practical applicability.
This study addresses the lack of statistical rigor in evaluating medical AI models across patient- and imaging-based stratifications. We propose the first end-to-end subgroup analysis framework. Methodologically, it integrates Bayesian posterior estimation with bootstrap resampling to quantify metric uncertainty, applies the Benjamini–Hochberg procedure for multiplicity correction, and introduces an adaptive subgroup scanning algorithm to automatically detect statistically significant and clinically relevant performance disparities. The framework explicitly handles challenges including small-sample subgroups, base-rate imbalance, and intersectional subgroup testing. Evaluated on ISIC2020 (skin lesion classification) and MIMIC-CXR (chest radiograph diagnosis), it successfully identifies interpretable fairness-critical subgroups—e.g., demographic or anatomical cohorts exhibiting systematic underperformance—thereby substantially enhancing both the statistical reliability and clinical interpretability of model evaluation.