Score
Statistical modeling of dichotomous (go/no-go, incident/no-incident) outcomes that separates latent factors (such as skill and task difficulty), produces interpretable intervals, and supports robust evaluation and application in clinical or decision-making contexts.
This study addresses the lack of systematic preprocessing standards, integrated analytical workflows, and cross-method consistency checks in current computer-based assessment process data. To bridge this gap, the authors propose an end-to-end analytical framework featuring a unified preprocessing pipeline and a dual-path analysis paradigm that synergistically combines feature engineering with model-based inference. The framework incorporates large language models (LLMs) to standardize action sequences and facilitate process-data-driven differential item functioning (DIF) detection. Technically, it integrates timestamp correction, action chunking, n-gram and TF-IDF feature extraction, multidimensional scaling, hidden Markov modeling, and subtask identification. Empirical results demonstrate that n-gram–based behavioral clustering offers diagnostic value for incorrect responders, multidimensional scaling effectively reconstructs behavioral constructs, and process data can identify and mitigate construct-irrelevant group differences.
This study addresses the heterogeneity of treatment effects in clinical research, where conventional subgroup analyses lack individual-level predictive power and purely machine learning–based approaches often lack statistical guarantees. To bridge this gap, the authors propose a two-stage hybrid workflow: first, using formal statistical hypothesis testing to confirm the presence of heterogeneous treatment effects, then constructing an individualized treatment strategy evaluated via cross-fitted doubly robust estimation under a Neyman–Pearson risk constraint. This framework integrates the interpretability of statistical inference with the predictive strength of machine learning, yielding a transparent, auditable, and statistically principled approach to heterogeneity. The method demonstrates efficacy in both simulation studies and the ACTG 175 HIV trial, and is accompanied by a practical implementation checklist along with guidance for alignment with regulatory-oriented heterogeneous treatment effect (HTE) assessment protocols.
This paper addresses the sensitivity of hierarchical bounded count data—such as medication adherence—to outliers and overdispersion in mixed-effects logistic regression. We propose a Bayesian t-distributed latent variable mixed logistic regression model. Unlike conventional approaches—including beta-binomial, binomial-logit-normal, and standard binomial models—our method employs a t-distribution for random effects, directly parameterizes the *median* (rather than the mean) of the response, and thereby achieves both robustness and interpretability. We derive a closed-form analytical expression for the posterior distribution of the median. To our knowledge, this is the first work to incorporate the t-distribution into hierarchical bounded count modeling. Extensive simulation studies and real-data experiments demonstrate that the proposed model significantly enhances robustness against contamination, yields more accurate and stable parameter estimates, and fills a critical gap in robust outlier-resistant modeling for bounded count outcomes.
Traditional approaches struggle to capture the temporal evolution of variability, skewness, and tail behavior in the underlying risk distribution of recurrent binary events such as hospital readmissions. This work proposes two novel frameworks: BLaS-Recurrent, a Bayesian model based on the sinh-arcsinh distribution, and QuaD-Recurrent, a quasi-distributional model leveraging nonparametric surface mapping. Both frameworks uniquely embed time-varying location, scale, skewness, and kurtosis into a flexible distributional family, jointly modeling dynamic shifts in both risk level and distributional shape. Moving beyond the limitation of estimating only mean risk, the proposed methods demonstrate superior calibration, robustness, and clinical interpretability in both simulations and real-world MIMIC-IV readmission data, uncovering previously overlooked patterns such as increasing right-skewness and expanding dispersion over time.
This study addresses the inconsistency in causal effect estimates between observational studies and randomized controlled trials (RCTs) by proposing the first unified framework for decomposing causal effect heterogeneity. The framework systematically identifies and quantifies three sources of heterogeneity: differences in covariate distributions, variation in mediating pathways, and shifts in outcome-generating mechanisms. Methodologically, it formally defines effect decomposition across data types (observational vs. experimental), integrating causal inference, sensitivity analysis, and decomposition modeling, while enabling robust parameter estimation under multiple hypotheses. Evaluated through simulation studies and an empirical analysis of the “Moving to Opportunity” experiment, the framework demonstrates improved interpretability, robustness, and policy generalizability in synthesizing evidence from heterogeneous data sources.
This study addresses the limitations of conventional frequentist approaches in effectively incorporating prior knowledge, which constrains adaptive decision-making and reliability in clinical trials. The authors propose a Bayesian framework tailored for discrete probability distributions—such as binomial, Poisson, and negative binomial—to model binary responses and overdispersed clinical endpoints using Bayesian networks. By continuously integrating accumulating evidence, the framework dynamically optimizes trial design and evaluation. Compared to maximum likelihood estimation, this approach demonstrates greater flexibility and robustness in both inferential behavior and practical performance, substantially enhancing decision quality while mitigating misinterpretation of results and reproducibility challenges.
This paper addresses longitudinal polytomous response data featuring ordinal attributes and individual-level covariates. We propose a Restricted Latent-Class Hidden Markov Model (RLC-HMM) that jointly models the evolution of latent attributes and the response-generating process, accommodating time non-homogeneity and conditional dependencies among states and covariates. To our knowledge, this is the first rigorous proof of model identifiability under such a complex structure. By integrating latent-class analysis with covariate-dependent state transitions, the RLC-HMM supports exploratory longitudinal cognitive diagnosis. Bayesian inference via MCMC is employed; simulations confirm accurate and robust parameter estimation. An empirical application to mathematics assessment data demonstrates substantial improvements over existing confirmatory approaches, effectively uncovering dynamic developmental trajectories of student proficiency.
This study addresses missing data in bivariate longitudinal settings arising from non-random dropout—particularly when the two response variables exhibit asynchronous dropout times and complex distributional features such as skewness and heavy tails. The authors propose an innovative Bayesian nonparametric joint modeling approach that simultaneously characterizes the observation processes and dropout mechanisms for both variables. By incorporating identification constraints based on dropout indicators and sensitivity parameters, the method achieves partial identification of the missing data distribution under nonignorable missingness. Assigning priors to the sensitivity parameters enables systematic sensitivity analyses across a range of plausible missingness scenarios. This work represents the first extension of Bayesian nonparametric modeling to bivariate longitudinal dropout contexts and demonstrates its practical utility through successful application to a cost-effectiveness clinical trial on intellectual disability interventions, yielding robust evidence to inform health policy decisions.
This study addresses the arbitrariness in recidivism risk prediction arising from model multiplicity by leveraging a judicial system with over 15 years of operational history. The authors formalize legal rules into algorithmic labels to construct a high-quality dataset, train interpretable models, and analyze how structural diversity among models influences predictive disagreement. For the first time in a real-world judicial setting, they quantify the relationship between model multiplicity and prediction arbitrariness, establish a theoretical lower bound, and demonstrate that actual inter-model consistency substantially exceeds worst-case expectations. Innovatively adopting a “minimum risk score across multiple models” strategy, the approach simultaneously safeguards individual rights and reduces decision arbitrariness. The resulting models not only achieve superior predictive performance and more equitable error distributions across demographic groups but also effectively capture inmates’ rehabilitation progress.
This study addresses the challenge of latent class modeling for mixed continuous and binary data by proposing a unified joint-likelihood framework that models continuous variables with normal distributions and binary variables with Bernoulli distributions. The latent class structure is efficiently estimated via the EM algorithm. A key contribution is the development of the first frequentist R package tailored to such mixed-data latent class models, which accommodates heteroscedasticity and censoring, eliminates the need for manual likelihood derivation by users, and includes dedicated tools for summarization and visualization. The method has been successfully applied to EQ-5D-5L value set estimation, demonstrating its practical utility and ease of use in health economics and related fields.
Current large language models (LLMs) lack systematic evaluation of interpretive reliability in psychiatric clinical risk assessment and are susceptible to interference from non-clinical information and prompt design. This study introduces the first LLM reliability auditing framework tailored to psychiatry, systematically evaluating the stability of hospitalization risk scores across four leading models—Gemini, LLaMA, Claude, and GPT—using synthetically generated patient profiles that incorporate both clinical and non-clinical features, alongside four prompting strategies: neutral, logical, humanistic influence, and clinical judgment. Results demonstrate that non-clinical information significantly increases both the mean risk scores and output variability across all models, revealing a high sensitivity to contextual noise and underscoring substantial reliability concerns for clinical deployment.
This study addresses a critical pitfall in explainable machine learning when predicting composite mental health constructs such as burnout-depression: overlapping latent constructs can induce spuriously stable feature importance rankings, thereby misleading assessments of model generalizability. To mitigate this, the authors propose a transferable residualization testing protocol that integrates ElasticNet regression, Kendall’s τ rank consistency analysis, and cross-cohort validation to systematically disentangle shared variance between predictors and outcomes. Empirical results reveal a high correlation (r = 0.72) between trait anxiety and depressive symptom subscales; after residualization, model R² plummeted from 0.41 to as low as 0.016, with prediction intervals spanning 35.4 points, indicating unsuitability for individual-level inference. This approach provides a crucial auditing tool for XAI research, demonstrating that apparent stability often stems from construct confounding rather than genuine predictive signal.