Preprint
Article

This version is not peer-reviewed.

Selective Prediction and Uncertainty-Aware Referral for Pap Smear Classification

Submitted:

22 July 2026

Posted:

23 July 2026

You are already at the latest version

Abstract
Deep learning models for cervical cytology are almost always evaluated as if every prediction must be acted upon, yet a screening system deployed alongside a cytopathologist need not classify every slide: it can defer the cases it is least certain about. Evaluating such a system requires asking not only how often it is correct, but whether its confidence ranks its errors to the bottom. This paper studies selective prediction and uncertainty-aware referral on the Herlev Pap smear dataset under a binary Normal-versus-Abnormal formulation. Two lightweight transformer backbones (Swin-Tiny, TinyViT-5M) are fine-tuned on Herlev from ImageNet-pretrained weights with weighted random sampling, calibrated by post-hoc temperature scaling fit on a held-out calibration subset, and compared against a soft-voting ensemble of both models. Discrimination is reported alongside expected calibration error (ECE) and, as the primary endpoint, the area under the risk-coverage curve (AURC). No statistically significant difference was detected between the two configurations in accuracy or macro-F1, yet the ensemble halves AURC (0.0022 vs. 0.0045, a 51.8% reduction, lower in all five folds) and extends the coverage at which zero errors are made from 18.3% to 72.8% of the pooled test predictions. The same ensemble is nonetheless worse calibrated in absolute terms (ECE 0.0339 vs. 0.0247) and produces more false negatives (14 vs. 10). These results separate two properties that are frequently conflated: the ability to rank predictions by trustworthiness, and the accuracy of the confidence values themselves. Ensembling improves the former while degrading the latter, and the former directly governs the observed risk-coverage tradeoff, whereas the latter governs the interpretation of the reported confidence values.
Keywords: 
;  ;  ;  ;  ;  ;  ;  

1. Introduction

Cervical cancer remains a substantial global health burden, and its incidence and mortality remain well above the threshold set by the WHO elimination initiative in most countries, with marked inequalities across levels of human development [1]. Screening by Pap smear is the principal tool for early detection, but manual examination is labor-intensive, requires expert cytologists, and is subject to delays and inconsistency between readers [2]. Deep learning has accordingly been applied to cervical cell classification, with convolutional, transformer, and ensemble architectures all reporting high discrimination on public benchmarks [2]. The task itself, however, is one of triage performed under a heavy and uneven workload: the overwhelming majority of slides examined in a screening program are negative, and the cytopathologist's time is most valuable on the small fraction that are not. An automated classifier deployed in this setting is therefore not required to render a verdict on every slide. It may instead abstain on the cases it finds ambiguous, referring them for human review while auto-clearing the remainder. This is the setting of selective prediction, and the quantity that governs it is not accuracy but the relationship between a model's confidence and its errors.
The distinction matters because two models with identical accuracy can behave very differently under referral. If one model's mistakes are concentrated among its least confident predictions, a low confidence threshold will remove nearly all of them while retaining most of the workload. If another model's mistakes are scattered across the confidence range, no threshold removes them without also discarding a large volume of correct predictions. The risk-coverage curve makes this behavior explicit by plotting the error rate among retained predictions against the fraction retained, and the area under that curve (AURC) summarizes it in a single number [3], [4].
AURC is not, however, a calibration metric, and this paper argues that the difference is easy to lose sight of. Expected calibration error asks whether a stated confidence of 0.9 corresponds to being correct 90% of the time. AURC asks only whether predictions the model is more confident about are more often correct than those it is less confident about. The first is a statement about absolute probability values; the second is a statement about their ordering. A model can rank its predictions perfectly while systematically overstating or understating its confidence, and post-hoc temperature scaling, which rescales confidence without reordering predictions [5], changes the second not at all while changing the first substantially.
Uncertainty estimation and abstention have been studied in medical imaging: rejecting low-confidence samples on the basis of predictive uncertainty has been shown to improve the reliability of the predictions that are retained [6], and selective prediction has recently been extended from classification to segmentation, where its authors note that selective classification has received considerable attention while only a few studies have explored the segmentation setting [7]. A related principle argues for reporting reliability at the level of the individual case rather than only in aggregate: converting a continuous quality score into discrete per-case risk categories makes the reliability of each prediction explicit where aggregate metrics alone would obscure it [8]. In cervical cytology specifically, evaluation continues to center on accuracy, F1-score, sensitivity, specificity, and the area under the ROC curve (AUROC). Calibration has received limited attention on this task, few studies have evaluated referral behavior, and the joint effect of ensembling on calibration error and on confidence-based error ranking remains underexplored on Herlev.
This paper reports a case in which the two properties diverge cleanly. On the Herlev Pap smear dataset, a soft-voting ensemble of two lightweight transformer backbones shows no statistically significant difference from the better of its two members in accuracy or macro-F1, and is measurably worse calibrated, yet it halves AURC and quadruples the coverage at which no errors are made. The contribution is threefold. First, the study is conducted entirely on Herlev, with all models trained directly on that dataset rather than transferred from another source, and with temperature scaling fit on a calibration subset carved out of each cross-validation fold's training portion so that no sample used to fit a temperature is ever used to evaluate it. Second, selective prediction is evaluated as a primary endpoint rather than as an afterthought, with AURC reported per fold so that the consistency of the effect can be assessed rather than only its pooled magnitude. Third, the calibration and referral results are reported together, which is what makes the divergence between them visible.

3. Methodology

All experiments were conducted directly on the Herlev dataset. No models or checkpoints were transferred from any other cervical cytology dataset; every model reported here was fine-tuned from ImageNet-pretrained weights on Herlev alone.
Figure 1 summarizes the full experimental pipeline; the remainder of this section describes each stage in turn.

3.1. Dataset

Herlev comprises 917 single-cell Pap smear images distributed across seven native subtypes. Following standard practice, the three normal subtypes (superficial, intermediate, and columnar) are mapped to a Normal class and the four abnormal subtypes (light, moderate, and severe dysplasia, and carcinoma in situ) to an Abnormal class. The resulting binary problem is skewed toward the abnormal category, which is the reverse of the skew found in a screening population and is an artifact of how the dataset was assembled. Table 1 reports the class distribution and the corresponding inverse-frequency sampling weights used to counteract this imbalance during training (Section 3.4). Table 2 reports the underlying seven-subtype composition. Figure 2 shows one representative image from each of the seven native subtypes.

3.2. Class Formulation

The binary Normal-versus-Abnormal formulation is used throughout. This is a deliberate simplification relative to the seven-subtype labels, adopted because it is the convention in the Herlev literature and because the referral question this paper studies, whether a slide can be auto-cleared or must be reviewed, is itself binary. The cost of the simplification is that it discards the distinction between low-grade and high-grade abnormality, which a clinical triage system would need.

3.3. Candidate Models and Ensemble

Two lightweight transformer architectures were selected as the candidate pool: Swin-Tiny and TinyViT-5M. The choice of compact transformer backbones follows evidence that such models remain competitive with heavier convolutional baselines at reduced deployment cost on cervical cell images [16]. Each was fine-tuned directly on Herlev from ImageNet-pretrained weights. A single soft-voting ensemble, denoted Hybrid-2, was constructed by averaging the predicted class probabilities of both models and renormalizing. Because the candidate pool contains exactly two models, Hybrid-2 comprises the entire pool; no model selection is performed, and the comparison reported in this paper is therefore between the better of the two individual models and the average of both, not between a selected subset and a larger pool.

3.4. Class-Imbalance Handling

Class imbalance was addressed during training using a WeightedRandomSampler with inverse class-frequency sampling weights, so that images from the minority Normal class are sampled more frequently per epoch than their raw frequency in the dataset. This changes each sample's sampling probability during training; it does not alter the underlying dataset or its true class distribution, which is reported separately in Table 1 alongside the sampling weights it implies.

3.5. Calibration Protocol

Post-hoc temperature scaling was applied to calibrate each model's predicted probabilities. To avoid the optimistic calibration estimates that result when the same data is used both to fit a temperature and to evaluate it, each cross-validation fold's training portion was further split into a model-training subset (85%) and a calibration subset (15%, stratified by class), with the calibration subset held out from model training entirely. The temperature was fit by minimizing negative log-likelihood on the calibration subset only, over the bounded interval [0.25, 5.0]. Final metrics, both before and after calibration, are computed exclusively on the fold's held-out test partition, which is never used for either training or calibration.
As a specific leakage check: within every cross-validation fold, the temperature-scaling parameter is fit exclusively on the calibration subset described above, and every accuracy, macro-F1, AUROC, ECE, and AURC value reported in this paper is computed exclusively on that fold's held-out test partition. No sample used to fit a temperature is ever also used to evaluate it, and no sample used for either training or calibration is ever included in a reported test metric. This partitioning discipline follows the leakage-aware benchmarking protocol we developed to keep overlapping partitions from inflating reported reliability [30], [31].

3.6. Selective Prediction and the Risk-Coverage Curve

Selective prediction is evaluated using the maximum calibrated class probability as the confidence score, which is the standard baseline selection function [24]. For a coverage level c, the c fraction of test predictions with the highest confidence is retained and the remainder is deferred; the risk at that coverage is the error rate among the retained predictions. Sweeping c from 1/n to 1 traces the risk-coverage curve, where n is the number of test predictions.
The area under this curve (AURC) is computed by trapezoidal integration over the swept coverage grid and normalized by the width of that grid, which spans from 1/n to 1 rather than from 0 to 1. This normalization differs from definitions that integrate over the full unit interval, and yields slightly larger values than an unnormalized integral would; it is stated here explicitly so that the absolute magnitudes reported below are interpreted correctly. All comparisons in this paper are made between configurations evaluated under the identical normalization, so the relative differences are unaffected. AURC is computed on after-calibration probabilities throughout.
AURC is reported per fold, so that a mean and standard deviation across the five folds can be given and the consistency of any difference assessed. A pooled risk-coverage curve, obtained by concatenating the held-out test predictions of all five folds, is also produced for visualization. The pooled curve and the per-fold mean are not the same quantity: pooling mixes predictions whose confidence values were produced under different fold-specific temperatures, so the pooled AURC can differ appreciably from the mean of the per-fold AURCs. Both are reported below, and the per-fold statistics are used for all inferential claims.

3.7. Cross-Validation and Evaluation Metrics

Stratified five-fold cross-validation was used, with training, calibration, and test partitions constructed as described above within each fold. Four metrics were computed on the held-out test partition of every fold: accuracy, macro-averaged F1, AUROC, and expected calibration error (ECE, 15 bins). Each is reported both before calibration (raw softmax outputs) and after calibration (temperature-scaled outputs). AURC is reported on calibrated outputs only. Models were trained for 15 epochs with batch size 32 using AdamW (learning rate 1e-4, weight decay 1e-4), cross-entropy loss with label smoothing 0.1, a two-epoch linear warmup followed by cosine decay, and 224 × 224 inputs.

4. Results

Table 3 reports discrimination metrics for the best individual model (Swin-Tiny) and the two-model ensemble (Hybrid-2), averaged across the five cross-validation folds. Swin-Tiny was the better of the two individual models under the after-calibration ranking criterion and is reported as the best single model throughout.
Accuracy, macro-F1, and AUROC are unchanged to four decimal places before and after calibration for Swin-Tiny, as expected: temperature scaling divides all logits by a single positive scalar and therefore cannot reorder predicted classes. The small changes visible for Hybrid-2 arise because calibration is applied to each member independently before their probabilities are averaged, so the averaging operation combines differently scaled distributions and can alter the ensemble's argmax.
The two configurations are close on the metrics most commonly reported. Across the five folds a paired t-test finds no significant difference in accuracy (0.9727 vs. 0.9662, p = 0.109) or macro-F1 (0.9650 vs. 0.9559, p = 0.094). The ensemble does achieve a significantly higher AUROC (0.9947 vs. 0.9907, p = 0.014). On the basis of Table 3 alone, a reader would reasonably conclude that ensembling buys a marginal and largely non-significant improvement.
Table 4 tells a different story. Temperature scaling substantially improves calibration for both configurations, reducing ECE by 53.0% for Swin-Tiny and 48.0% for Hybrid-2. But after calibration the ensemble is less well calibrated than the best single model, with a mean ECE of 0.0339 against Swin-Tiny's 0.0247, a gap that holds in four of the five folds. Its advantage lies entirely in the selective-prediction column: AURC falls from 0.0045 to 0.0022, a 51.8% reduction, and this difference is significant under a paired t-test across folds (p = 0.028) and consistent in direction in all five folds. Because this comparison is based on only five overlapping cross-validation folds, the inferential result should be interpreted cautiously.
Figure 3 shows the per-fold AURC values that underlie this comparison. The ensemble achieves the lower AURC in every fold, and its fold-to-fold variability is less than half that of the single model (standard deviation 0.0008 vs. 0.0018). The consistency, rather than the magnitude of the pooled difference, is the stronger evidence here: with only five folds, a difference of this size that reverses in no fold is more informative than a large mean gap driven by one or two folds.
Figure 4 shows the pooled risk-coverage curves. The practical consequence of the AURC difference is visible on the left of the plot: Hybrid-2 makes no errors at all until 72.8% of predictions are retained, whereas Swin-Tiny's first error appears once coverage exceeds 18.3%. In a referral setting, this is the difference between a system that can retain 72.8% of pooled cross-validation predictions before the first observed error and one that can safely auto-clear only 18.3% before its first. At full coverage the two curves converge toward their respective overall error rates, 2.73% and 3.38%, and the gap between them narrows to a margin that the accuracy comparison in Table 3 already described as non-significant.
Figure 5 reports pooled confusion matrices. Hybrid-2 makes fewer total errors than Swin-Tiny (25 vs. 31) and is markedly more specific, misclassifying 11 Normal images as Abnormal against Swin-Tiny's 21. It is, however, less sensitive: it misses 14 Abnormal images against Swin-Tiny's 10, giving sensitivities of 0.9793 and 0.9852 respectively. In cervical screening the false negative is the more costly error, since an abnormality passed as normal leaves the screening pathway entirely. The ensemble's superior aggregate error count therefore does not translate straightforwardly into superior clinical behavior, and this qualification applies to its AURC advantage as well: AURC weights all errors equally.
Table 5 disaggregates the two individual models. Both fit temperatures below one (0.655 and 0.587), indicating that both are under-confident on this dataset rather than over-confident, a direction opposite to that typically reported for large networks on natural-image benchmarks. Neither temperature approaches the bounds of the search interval. The two models are close on every metric, which is consistent with the ensemble of the two behaving similarly to either member on accuracy while differing from both on how its confidence is distributed.

5. Discussion

The central finding of this study is that ensembling improved the ranking of predictions by confidence while degrading the accuracy of the confidence values themselves. Hybrid-2 halved AURC relative to the best single model, and did so in every fold, yet its post-calibration ECE was higher than that of the single model in four folds out of five. These are not contradictory results. ECE measures whether a stated confidence of p is correct p of the time; AURC measures only whether predictions with higher confidence are more often correct than predictions with lower confidence. Averaging the probabilities of two models pulls confidence values toward the middle of the range, which harms the first quantity, while suppressing the idiosyncratic high-confidence errors of either member, which helps the second.
For a referral system, confidence-based ranking directly governs the observed risk–coverage tradeoff and therefore how much workload can be retained at a given observed risk. The operating characteristic that matters is where the risk-coverage curve leaves zero, and on that measure the ensemble is not marginally better but qualitatively so: 72.8% coverage at zero pooled risk against 18.3%. Within the pooled cross-validation predictions, the ensemble deferred fewer than one-third of cases before the first observed error among retained predictions. This operating point should not be interpreted as a clinically validated deployment threshold. This is the practical argument for evaluating selective prediction directly rather than inferring it from accuracy, for which no statistically significant difference between the two configurations was detected in this study.
Two qualifications temper the conclusion, and both point the same direction. First, the ensemble is less sensitive than the single model, missing 14 abnormal images against 10. AURC treats a false negative and a false positive as equivalent, so a metric that ranks Hybrid-2 as clearly superior is silent about the error type that matters most in cervical screening. A cost-sensitive analogue of the risk-coverage curve, in which risk is defined by clinical cost rather than raw error, would be the appropriate next instrument, and this study does not provide one. Second, the reversed class balance of Herlev, in which abnormal images outnumber normal ones by nearly three to one, is an artifact of dataset construction and the opposite of a screening population. Sensitivity and specificity estimated on this distribution should not be read as estimates of screening performance.
The findings should be interpreted within the scope of this dataset, model pool, and calibration protocol. Herlev is small (917 single-cell images), and the calibration subset carved out of each fold's training portion is correspondingly small, which is reflected in the fold-to-fold variability of the fitted temperatures. The candidate pool contains only two architectures, so the ensemble is the entire pool and no model selection is exercised; whether the divergence between calibration and ranking persists for larger and more diverse ensembles is an open question that a two-model study cannot answer. Finally, the binary Normal-versus-Abnormal formulation, though conventional for Herlev, discards the low-grade versus high-grade distinction that a clinical triage system would need to make. The small candidate pool also reflects a preference for parsimonious modeling: when data are limited, a model built from a few informative features can generalize better than a larger one, as we found in predicting post-stroke functional outcomes [32]. The authors' broader work spans other machine-learning domains, including speaker identification [33] and underwater sensor networks [34].
Read alongside our recent study on liquid-based cervical cytology, in which ensembling provided no consistent reliability benefit once individual models were properly calibrated [20], the present result suggests that the value of an ensemble depends on which reliability property is being measured. Where that study asked whether ensembling improved calibration and found that it did not, this one asks whether ensembling improves selective prediction and finds that it does, on a different dataset and class formulation. The two findings are compatible, and together they caution against treating ensemble size as a single lever on a single quantity called reliability.

6. Conclusion

This paper evaluated selective prediction and uncertainty-aware referral for Pap smear classification on the Herlev dataset under a binary Normal-versus-Abnormal formulation, fine-tuning two lightweight transformer backbones on Herlev from ImageNet-pretrained weights and comparing the better individual model against a soft-voting ensemble of both. Temperature scaling, fit on a calibration subset held out from both training and test data, reduced expected calibration error by roughly half for both configurations without altering discrimination. No statistically significant difference was detected between the ensemble and the best single model in accuracy or macro-F1, and the ensemble was less well calibrated than the best single model, yet it halved the area under the risk-coverage curve, did so consistently in all five folds, and extended the coverage at which no errors are made from 18.3% to 72.8% of pooled predictions. Its advantage in aggregate error count was accompanied by a loss of sensitivity, the error type most costly in screening. These results indicate that the ability to rank predictions by trustworthiness and the accuracy of the confidence values themselves are distinct properties that respond differently to ensembling, and that a system intended to defer its uncertain cases should be evaluated on the former directly rather than through accuracy or calibration alone.

References

  1. Singh, D.; Vignat, J.; Lorenzoni, V.; Eslahi, M.; Ginsburg, O.; Lauby-Secretan, B.; Arbyn, M.; Basu, P.; Bray, F.; Vaccarella, S. Global estimates of incidence and mortality of cervical cancer in 2020: A baseline analysis of the WHO Global Cervical Cancer Elimination Initiative. Lancet Glob. Health 2023, 11(2), e197–e206. [Google Scholar] [CrossRef] [PubMed]
  2. Sahoo, P.; Saha, S.; Mondal, S.; Seera, M.; Sharma, S. K.; Kumar, M. Enhancing computer-aided cervical cancer detection using a novel fuzzy rank-based fusion. IEEE Access 2023, 11, 145281–145294. [Google Scholar] [CrossRef]
  3. El-Yaniv, R.; Wiener, Y. On the foundations of noise-free selective classification. J. Mach. Learn. Res. 2010, 11, 1605–1641. [Google Scholar]
  4. Zhou, H.; Van Landeghem, J.; Popordanoska, T.; Blaschko, M. B. A novel characterization of the population area under the risk coverage curve (AURC) and rates of finite sample estimators. Int. Conf. Mach. Learn. (ICML) 2025, arXiv:2410.15361. [Google Scholar] [CrossRef]
  5. Guo, C.; Pleiss, G.; Sun, Y.; Weinberger, K. Q. On calibration of modern neural networks. International Conference on Machine Learning, PMLR, 2017; 70, pp. 1321–1330. [Google Scholar]
  6. Ghesu, F. C.; Georgescu, B.; Mansoor, A.; Yoo, Y.; Gibson, E.; Vishwanath, R. S.; Balachandran, A.; Balter, J. M.; Cao, Y.; Singh, R.; Digumarthy, S. R.; Kalra, M. K.; Grbic, S.; Comaniciu, D. Quantifying and leveraging predictive uncertainty for medical image assessment. Med. Image Anal. 2021, 68, 101855. [Google Scholar] [CrossRef] [PubMed]
  7. Sanisoglu, M. A.; Navab, N.; Kim, S. T. Learning to abstain: Reliable medical image segmentation with rejection option. IEEE Access 2026, 14, 32655–32665. [Google Scholar] [CrossRef]
  8. Albzour, N. Reliability-aware CT-MRI registration: A quality engineering framework with stability analysis and risk classification. arXiv 2026, arXiv:2607.02585. [Google Scholar] [CrossRef]
  9. Jantzen, J.; Norup, J.; Dounias, G.; Bjerregaard, B. Pap-smear benchmark data for pattern classification. In Nature Inspired Smart Information Systems (NiSIS); 2005; pp. 1–9. [Google Scholar]
  10. Zhang, L.; Lu, L.; Nogues, I.; Summers, R. M.; Liu, S.; Yao, J. DeepPap: Deep convolutional networks for cervical cell classification. IEEE J. Biomed. Health Inform. 2017, 21(6), 1633–1643. [Google Scholar] [CrossRef] [PubMed]
  11. Lin, H.; Hu, Y.; Chen, S.; Yao, J.; Zhang, L. Fine-grained classification of cervical cells using morphological and appearance based convolutional neural networks. IEEE Access 2019, 7, 71541–71549. [Google Scholar] [CrossRef]
  12. Jain, S.; Jain, A.; Jangid, M.; Shetty, S. Metaheuristic driven framework for classifying cervical cancer on smear images using deep learning approach. IEEE Access 2024, 12, 160805–160821. [Google Scholar] [CrossRef]
  13. Kaur, H.; Sharma, R.; Kaur, J. Comparison of deep transfer learning models for classification of cervical cancer from pap smear images. Sci. Rep. 2025, 15(1), 3945. [Google Scholar] [CrossRef] [PubMed]
  14. Albzour, N.; Lam, S. S. Segmentation and classification of Pap smear images for cervical cancer detection using deep learning. In Proceedings of the IISE Annual Conference & Expo 2025, Also available at. 2025. [Google Scholar] [CrossRef]
  15. Albzour, N.; Lam, S. S. Systematic evaluation of vision transformers for automated cervical cancer classification: Optimization, statistical validation, and clinical interpretability. Cancers 2026, 18(13), 2178. [Google Scholar] [CrossRef] [PubMed]
  16. Albzour, N.; Lam, S. S. A reproducible benchmark of ViT-Tiny against CNN baselines for cervical cell classification: Accuracy, statistical validation, and deployment efficiency. SSRN. 2026. Available online: https://ssrn.com/abstract=6839541.
  17. Liang, G.; Zhang, Y.; Wang, X.; Jacobs, N. Improved trainable calibration method for neural networks on medical imaging classification. British Machine Vision Conference (BMVC), 2020. [Google Scholar]
  18. Carse, J.; Alvarez Olmo, A.; McKenna, S. Calibration of deep medical image classifiers: An empirical comparison using dermatology and histopathology datasets. In Uncertainty for Safe Utilization of Machine Learning in Medical Imaging (UNSURE); MICCAI Workshops, 2022; pp. 89–99. [Google Scholar] [CrossRef]
  19. Penso, C.; Frenkel, L.; Goldberger, J. Confidence calibration of a medical imaging classification system that is robust to label noise. IEEE Trans. Med. Imaging 2024, 43(6), 2050–2060. [Google Scholar] [CrossRef] [PubMed]
  20. Albzour, N.; Lam, S. S. Reliability-aware ensemble classification under class imbalance: A calibration study on liquid-based cervical cytology. arXiv 2026, arXiv:2607.09837. [Google Scholar] [CrossRef]
  21. Chow, C. K. On optimum recognition error and reject tradeoff. IEEE Trans. Inf. Theory 1970, 16(1), 41–46. [Google Scholar] [CrossRef]
  22. Geifman, Y.; El-Yaniv, R. Selective classification for deep neural networks. Adv. Neural Inf. Process. Syst. 2017, 30, 4878–4887. [Google Scholar]
  23. Geifman, Y.; El-Yaniv, R. SelectiveNet: A deep neural network with an integrated reject option. Proceedings of the 36th International Conference on Machine Learning, 2019; 97, pp. 2151–2159. [Google Scholar]
  24. Hendrycks, D.; Gimpel, K. A baseline for detecting misclassified and out-of-distribution examples in neural networks. International Conference on Learning Representations, 2017. [Google Scholar]
  25. Pugnana, A.; Ruggieri, S. AUC-based selective classification. Proceedings of the 26th International Conference on Artificial Intelligence and Statistics 2023, PMLR 206, 2494–2514. [Google Scholar]
  26. Lakshminarayanan, B.; Pritzel, A.; Blundell, C. Simple and scalable predictive uncertainty estimation using deep ensembles. Adv. Neural Inf. Process. Syst. 2017, 30, 6402–6413. [Google Scholar]
  27. Sreelatha, S.; Shivashetty, V. Deep ensemble learning with uncertainty aware prediction ranking for cervical cancer detection using Pap smear images. IAES Int. J. Artif. Intell. 2025, 14(2), 1450–1460. [Google Scholar] [CrossRef]
  28. Albzour, N. External validation and calibration of hybrid deep-boosted ensembles for breast cancer survival prediction: A cross-cohort study on SEER and METABRIC. In Research Square; (preprint, not peer reviewed); 2026. [Google Scholar] [CrossRef] [PubMed]
  29. Albzour, N. Decoupled transferability of calibration and explainability in cross-cohort breast cancer survival prediction; Preprints.org (preprint, not peer reviewed); 2026. [Google Scholar] [CrossRef]
  30. Albzour, N. A leakage-aware comparative benchmark of machine learning, deep learning, and transformer models for reliable leukemia detection. arXiv 2026, arXiv:2606.24944. [Google Scholar]
  31. Al-Zoubi, H.; Al-Bzoor, N. Toward driverless AI: Automating leukemia detection and classification using hyperautomation, a case study. In Research Square; (preprint, not peer reviewed); 2022. [Google Scholar] [CrossRef] [PubMed]
  32. Albzour, N.; Agarwal, S.; Althnaibat, H.; Lu, S. Predicting post-stroke activities of daily living: Enhancing machine learning with feature selection. Proceedings of the IISE Annual Conference & Expo 2025, 2025, 1–6. [Google Scholar] [CrossRef] [PubMed]
  33. Al-Jarrah, M. A.; Al-Jarrah, A.; Jarrah, A.; AlShurbaji, M.; Magableh, S. K.; Al-Tamimi, A.-K.; Albzour, N.; Al-Shamali, M. O. Accurate reader identification for the Arabic holy Quran recitations based on an enhanced VQ algorithm. Rev. d'Intelligence Artif. 2022, 36(6), 815–823. [Google Scholar] [CrossRef]
  34. Albzour, N.; Bawaneh, R. Autonomous underwater vehicle–enabled data collection for underwater wireless sensor networks: A systematic taxonomy, protocol review, and open challenges; Preprints.org (preprint, not peer reviewed); 2026. [Google Scholar] [CrossRef]
Figure 1. Overview of the experimental pipeline: dataset and binary class formulation, the leakage-safe stratified five-fold cross-validation design, training of the two candidate models, post-hoc temperature scaling on the held-out calibration subset, construction of the Hybrid-2 soft-voting ensemble, confidence-based selective prediction, evaluation on held-out test data, and the statistical comparison of the best single model against the ensemble.
Figure 1. Overview of the experimental pipeline: dataset and binary class formulation, the leakage-safe stratified five-fold cross-validation design, training of the two candidate models, post-hoc temperature scaling on the held-out calibration subset, construction of the Hybrid-2 soft-voting ensemble, confidence-based selective prediction, evaluation on held-out test data, and the statistical comparison of the best single model against the ensemble.
Preprints 224543 g001
Figure 2. Representative Herlev images for each of the seven native subtypes, grouped by the binary class to which they are mapped. Each panel shows an equally sized window onto one cell: images are scaled to fill the panel and the overflow is cropped, so aspect ratios are preserved and no cell is stretched. The panels are therefore not to a common scale, and the apparent sizes here do not reflect the true relative sizes of the cells, which differ by more than an order of magnitude in pixel area. Images are upscaled for display.
Figure 2. Representative Herlev images for each of the seven native subtypes, grouped by the binary class to which they are mapped. Each panel shows an equally sized window onto one cell: images are scaled to fill the panel and the overflow is cropped, so aspect ratios are preserved and no cell is stretched. The panels are therefore not to a common scale, and the apparent sizes here do not reflect the true relative sizes of the cells, which differ by more than an order of magnitude in pixel area. Images are upscaled for display.
Preprints 224543 g002
Figure 3. Area under the risk-coverage curve (AURC, lower is better) for the best single model and the Hybrid-2 ensemble, computed independently on each of the five held-out cross-validation test folds. The ensemble achieves a lower AURC in every fold.
Figure 3. Area under the risk-coverage curve (AURC, lower is better) for the best single model and the Hybrid-2 ensemble, computed independently on each of the five held-out cross-validation test folds. The ensemble achieves a lower AURC in every fold.
Preprints 224543 g003
Figure 4. Pooled risk-coverage curves for the best single model and the Hybrid-2 ensemble, obtained by concatenating the held-out test predictions of all five folds. Legend values report the per-fold mean and standard deviation of AURC; the plotted curves are pooled and their areas therefore differ from these means (see Section 3.6).
Figure 4. Pooled risk-coverage curves for the best single model and the Hybrid-2 ensemble, obtained by concatenating the held-out test predictions of all five folds. Legend values report the per-fold mean and standard deviation of AURC; the plotted curves are pooled and their areas therefore differ from these means (see Section 3.6).
Preprints 224543 g004
Figure 5. Pooled confusion matrices across all 5 test folds (917 predictions) for the best single model and the Hybrid-2 ensemble, using calibrated predictions.
Figure 5. Pooled confusion matrices across all 5 test folds (917 predictions) for the best single model and the Hybrid-2 ensemble, using calibrated predictions.
Preprints 224543 g005
Table 1. Class distribution and sampling weights.
Table 1. Class distribution and sampling weights.
Class Count Sampling Weight
Normal 242 0.004132
Abnormal 675 0.001481
Table 2. Native seven-subtype composition of the Herlev dataset, and the binary class to which each subtype is mapped.
Table 2. Native seven-subtype composition of the Herlev dataset, and the binary class to which each subtype is mapped.
Subtype Count Mapped Class
Severe dysplastic 197 Abnormal
Light dysplastic 182 Abnormal
Carcinoma in situ 150 Abnormal
Moderate dysplastic 146 Abnormal
Normal columnar 98 Normal
Normal superficial 74 Normal
Normal intermediate 70 Normal
Table 3. Discrimination metrics for the best single model (Swin-Tiny) and the Hybrid-2 ensemble, before and after temperature scaling, averaged across 5 cross-validation folds. All metrics computed on held-out test partitions only.
Table 3. Discrimination metrics for the best single model (Swin-Tiny) and the Hybrid-2 ensemble, before and after temperature scaling, averaged across 5 cross-validation folds. All metrics computed on held-out test partitions only.
Config Acc. (before) Acc. (after) Macro-F1 (before) Macro-F1 (after) AUROC (before) AUROC (after)
Swin-Tiny 0.9662 0.9662 0.9559 0.9559 0.9907 0.9907
Hybrid-2 0.9706 0.9727 0.9621 0.9650 0.9953 0.9947
Table 4. Reliability and selective-prediction metrics for the same two configurations. ECE is reported before and after calibration; AURC is computed on calibrated outputs and reported as the mean and standard deviation across the five folds, alongside the pooled value.
Table 4. Reliability and selective-prediction metrics for the same two configurations. ECE is reported before and after calibration; AURC is computed on calibrated outputs and reported as the mean and standard deviation across the five folds, alongside the pooled value.
Config ECE (before) ECE (after) AURC (mean) AURC (std) AURC (pooled)
Swin-Tiny 0.0527 0.0247 0.0045 0.0018 0.0072
Hybrid-2 0.0653 0.0339 0.0022 0.0008 0.0020
Table 5. Per-model calibration results for each individual architecture, averaged across the five cross-validation folds. T is the fitted temperature. All metrics computed on held-out test partitions only.
Table 5. Per-model calibration results for each individual architecture, averaged across the five cross-validation folds. T is the fitted temperature. All metrics computed on held-out test partitions only.
Model T Acc. (after) Macro-F1 (after) AUROC (after) ECE (before) ECE (after)
Swin-Tiny 0.655 0.9662 0.9559 0.9907 0.0527 0.0247
TinyViT-5M 0.587 0.9640 0.9545 0.9912 0.0594 0.0303
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.