Preprint
Article

This version is not peer-reviewed.

Leakage-Aware Evaluation of Ensemble Classifiers for Symptom-Based Classification of Diabetes Status: Performance, Calibration, and Explainability

Submitted:

18 September 2026

Posted:

20 September 2026

You are already at the latest version

Abstract
Background/Objectives: Diabetes impose a substantial public health burden. Machine learning models trained on symptom data may help identify individuals whose recorded symptom profiles warrant confirmatory clinical assessment. This study compared ten en-semble models for symptom-based classification of a dataset-provided diabetes-status la-bel. Methods: The UCI Early-Stage Diabetes Risk Prediction Dataset was used, comprising 520 observations, 251 unique predictor profiles, 16 predictors, and a binary positive/negative diabetes-status label. All observations were retained, while identical predictor profiles were kept within the same partition during data splitting, nested cross-validation, hy-perparameter optimization, and calibration to reduce information leakage. The weighted F1-score was the optimization objective and primary model-selection metric throughout the analysis. Results: Bagging Extra Tree achieved the highest observed mean weighted F1-score (94.07% ± 7.99%) and MCC (0.876 ± 0.167) and was selected as the final model under the prespecified criterion. An exploratory Friedman analysis indicated variation in model rankings across folds (p = 0.001859), although no Nemenyi-adjusted pairwise comparison reached the nominal significance level. On the profile-disjoint internal holdout set, the calibrated model achieved a weighted F1-score of 0.9224, an MCC of 0.8207, a ROC-AUC of 0.9844, and a Brier score of 0.0528. Random Forest was selected in the unique-profile sensitivity analysis but yielded lower holdout performance, indicating sensitivity to the analytical unit. Conclusions: These findings provide internal evidence that tree-based ensembles can clas-sify the dataset-provided diabetes-status label under profile-disjoint evaluation. They do not establish prospective risk prediction or clinical diagnostic utility; independent exter-nal validation is required before clinical implementation.
Keywords: 
;  ;  ;  ;  ;  ;  
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.