Institution profile

Chiang Mai University

Academic institutionasia · th
Official website
Research library11linked papers
Opportunities0open roles
Selected work

Representative Papers

Hierarchical Clustering and Signal Denoising on Digraphs

Sep 28, 2026

This study addresses the challenges of directed graph clustering and signal denoising by proposing a spectral clustering framework that jointly accounts for connectivity and directionality. Methodologically, a Hermitian matrix is constructed to represent the directed graph structure, enabling hierarchical clustering through spectral decomposition combined with recursive K-means. Furthermore, the approach integrates hierarchical filtering with B-spline quasi-interpolation, achieving multiscale denoising and reconstruction of graph signals via adaptive thresholding. Experimental results on both synthetic and real-world datasets demonstrate that the proposed method significantly enhances clustering consistency while effectively improving signal recovery performance in terms of RMSE and SNR metrics.

0 citationsRead paper

Comparing Imputation Methods for Clinical Prediction Model Development under Complex Missingness Scenarios: A Simulation Study Using Real-World Cardiac Data

Jul 08, 2026

This study addresses the challenge of diminished stability and generalizability of clinical prediction models under complex missing data, where the impact of different imputation strategies remains unclear. Leveraging a real-world cardiac disease cohort, we simulated 18 distinct missingness mechanisms to systematically evaluate how multiple imputation, missForest, k-nearest neighbors (kNN) imputation, and complete-case analysis affect logistic regression model performance. Model assessment encompassed internal and external validation metrics including AUC, calibration slope, prediction error, and computational efficiency. Our work provides the first comprehensive comparison of imputation methods across diverse missing data patterns, revealing that kNN imputation demonstrates superior robustness—particularly under high missingness rates and complex missingness structures—while achieving excellent external generalizability and the lowest computational cost, making it especially suitable for large-scale clinical modeling.

0 citationsRead paper

A repeated k-fold cross-validation approach for evaluating the instability of clinical prediction models: an empirical comparison to the bootstrap approach

Jul 03, 2026

This study addresses the lack of systematic comparison between cross-validation and bootstrapping for assessing instability in clinical prediction models. Leveraging a cohort of 19,418 emergency department patients, it presents the first comprehensive evaluation of repeated five-fold cross-validation versus bootstrapping across varying events-per-variable (EPV) scenarios, using logistic regression and random forest models. Performance was assessed via AUC, calibration slope, large-scale calibration, and mean absolute prediction error (MAPE). Results indicate that when EPV ≥ 30, both methods yield comparable discriminative ability; however, cross-validation provides more accurate calibration estimates and significantly lower MAPE. These advantages render cross-validation particularly suitable for evaluating model instability across multiple algorithms, offering a dual benefit of internal validation and quantification of predictive stability.

0 citationsRead paper

Class Imbalance Corrections Failed to Enhance Discrimination, Model Calibration, and Prediction Stability: An Empirical Simulation Study Based on Clinical Dataset

Jun 07, 2026

Whether class imbalance correction improves the performance of clinical prediction models remains controversial. This study leverages data from the GUSTO-I clinical trial to systematically evaluate the impact of various correction strategies—including algorithm-level rebalancing, oversampling, and hybrid sampling—on model discrimination (AUC), calibration (calibration plots and MAPE), and predictive stability (Classification Instability Index, CII) across varying sample sizes. Using penalized logistic regression with 200 bootstrap replications, we find that all correction methods fail to enhance discriminative performance and instead introduce greater calibration bias, risk overestimation, and increased prediction instability. These results challenge the common practice of routinely applying class imbalance corrections in clinical modeling and, for the first time in large-scale simulations, reveal their potential harms.

0 citationsRead paper

Influence of continuous predictor modelling methods on prediction stability in clinical prediction model development: an empirical comparison using real clinical data

Jun 05, 2026

This study addresses the lack of systematic empirical evidence on how modeling approaches for continuous variables affect prediction stability in clinical prediction models, particularly across varying sample sizes. Leveraging real-world emergency department data within a bootstrapping framework, we comprehensively compare six methods—dichotomization, tertile categorization, linear and quadratic terms, multivariable fractional polynomials, and XGBoost—in terms of discrimination, calibration, and predictive stability. Our findings indicate that linear specifications achieve acceptable stability even with small samples, whereas more complex models require substantially larger sample sizes to stabilize. Although XGBoost yields higher AUC values, it exhibits persistent calibration bias. This work is the first to quantitatively characterize the trade-off between sample size and modeling complexity, offering empirical guidance for developing robust clinical prediction models.

0 citationsRead paper
Recent publications

Latest Papers

Hierarchical Clustering and Signal Denoising on Digraphs

Sep 28, 2026

This study addresses the challenges of directed graph clustering and signal denoising by proposing a spectral clustering framework that jointly accounts for connectivity and directionality. Methodologically, a Hermitian matrix is constructed to represent the directed graph structure, enabling hierarchical clustering through spectral decomposition combined with recursive K-means. Furthermore, the approach integrates hierarchical filtering with B-spline quasi-interpolation, achieving multiscale denoising and reconstruction of graph signals via adaptive thresholding. Experimental results on both synthetic and real-world datasets demonstrate that the proposed method significantly enhances clustering consistency while effectively improving signal recovery performance in terms of RMSE and SNR metrics.

0 citationsRead paper

Comparing Imputation Methods for Clinical Prediction Model Development under Complex Missingness Scenarios: A Simulation Study Using Real-World Cardiac Data

Jul 08, 2026

This study addresses the challenge of diminished stability and generalizability of clinical prediction models under complex missing data, where the impact of different imputation strategies remains unclear. Leveraging a real-world cardiac disease cohort, we simulated 18 distinct missingness mechanisms to systematically evaluate how multiple imputation, missForest, k-nearest neighbors (kNN) imputation, and complete-case analysis affect logistic regression model performance. Model assessment encompassed internal and external validation metrics including AUC, calibration slope, prediction error, and computational efficiency. Our work provides the first comprehensive comparison of imputation methods across diverse missing data patterns, revealing that kNN imputation demonstrates superior robustness—particularly under high missingness rates and complex missingness structures—while achieving excellent external generalizability and the lowest computational cost, making it especially suitable for large-scale clinical modeling.

0 citationsRead paper

A repeated k-fold cross-validation approach for evaluating the instability of clinical prediction models: an empirical comparison to the bootstrap approach

Jul 03, 2026

This study addresses the lack of systematic comparison between cross-validation and bootstrapping for assessing instability in clinical prediction models. Leveraging a cohort of 19,418 emergency department patients, it presents the first comprehensive evaluation of repeated five-fold cross-validation versus bootstrapping across varying events-per-variable (EPV) scenarios, using logistic regression and random forest models. Performance was assessed via AUC, calibration slope, large-scale calibration, and mean absolute prediction error (MAPE). Results indicate that when EPV ≥ 30, both methods yield comparable discriminative ability; however, cross-validation provides more accurate calibration estimates and significantly lower MAPE. These advantages render cross-validation particularly suitable for evaluating model instability across multiple algorithms, offering a dual benefit of internal validation and quantification of predictive stability.

0 citationsRead paper

Class Imbalance Corrections Failed to Enhance Discrimination, Model Calibration, and Prediction Stability: An Empirical Simulation Study Based on Clinical Dataset

Jun 07, 2026

Whether class imbalance correction improves the performance of clinical prediction models remains controversial. This study leverages data from the GUSTO-I clinical trial to systematically evaluate the impact of various correction strategies—including algorithm-level rebalancing, oversampling, and hybrid sampling—on model discrimination (AUC), calibration (calibration plots and MAPE), and predictive stability (Classification Instability Index, CII) across varying sample sizes. Using penalized logistic regression with 200 bootstrap replications, we find that all correction methods fail to enhance discriminative performance and instead introduce greater calibration bias, risk overestimation, and increased prediction instability. These results challenge the common practice of routinely applying class imbalance corrections in clinical modeling and, for the first time in large-scale simulations, reveal their potential harms.

0 citationsRead paper

Influence of continuous predictor modelling methods on prediction stability in clinical prediction model development: an empirical comparison using real clinical data

Jun 05, 2026

This study addresses the lack of systematic empirical evidence on how modeling approaches for continuous variables affect prediction stability in clinical prediction models, particularly across varying sample sizes. Leveraging real-world emergency department data within a bootstrapping framework, we comprehensively compare six methods—dichotomization, tertile categorization, linear and quadratic terms, multivariable fractional polynomials, and XGBoost—in terms of discrimination, calibration, and predictive stability. Our findings indicate that linear specifications achieve acceptable stability even with small samples, whereas more complex models require substantially larger sample sizes to stabilize. Although XGBoost yields higher AUC values, it exhibits persistent calibration bias. This work is the first to quantitatively characterize the trade-off between sample size and modeling complexity, offering empirical guidance for developing robust clinical prediction models.

0 citationsRead paper