Submitted:
20 July 2026
Posted:
22 July 2026
You are already at the latest version
Abstract

Keywords:
1. Introduction
1.1. Contributions
- A leakage-controlled ordinal benchmark. To our knowledge, the first head-to-head comparison of multinomial logistic regression (MLR), random forest, XGBoost, and an artificial neural network (ANN) on an ordinal three-tier (Low/Medium/High) formulation of the German, Taiwan, and Lending Club corpora under one identical protocol: the split precedes all fitting, scaling and SMOTE are estimated inside each cross-validation fold, and tier-defining columns are excluded from the predictors. Existing work on these corpora is almost exclusively binary (Table 1).
- A quantification of protocol-driven inflation. By reporting the transition from leaked to honest performance on the same data and models (Section 4.2: from 1.000 accuracy to 0.63–0.86), the study measures how much of the spread in published results on these corpora can be manufactured by preprocessing and resampling design alone, directly explaining why the results reviewed in Table 1 disagree.
- Evidence against a universal winner. Under the controlled protocol the best model is corpus-dependent—MLR on the small, weakly correlated German corpus; tree ensembles on the larger Taiwan and Lending Club corpora—supporting protocol-conditional rather than universal model claims. The quadratic weighted kappa (QWK) is shown to reorder models relative to accuracy when errors concentrate in adjacent tiers.
- A reproducible, deployable pipeline. A fully specified R/tidymodels implementation (pinned versions, fixed seed, serialized reload-and-score objects mapping tiers to graduated lending actions), together with a documented multi-class probability mis-decoding fault that silently drives a trained model to chance, and its diagnostic signature.
2. Literature Review
2.1. Machine Learning for Credit Scoring
2.2. Imbalanced Credit Data
2.3. Leakage and Reproducibility
2.4. Multi-Class and Ordinal Risk
2.5. Explainability and Governance
2.6. Prior Studies on the Three Benchmark Corpora
2.7. Research Gap
3. Data and Methodology
3.1. Datasets
3.2. Three-Tier Target Construction
3.3. Exploratory Analysis
3.4. Leakage-Safe Preprocessing Protocol
- Split first. An 80/20 stratified split (seed 123) is taken on the raw, labelled data before any transformation, isolating the test set.
- Fit preprocessing inside resamples. These steps are applied in a fixed order inside a single recipe: (i) one-hot encoding of nominal variables, (ii) median imputation of missing values, (iii) zero-variance predictor removal, (iv) standardization, and (v) SMOTE oversampling (Chawla et al. 2002). The recipe is re-estimated within each of five cross-validation folds on the training partition only; SMOTE (k = 5) is applied solely within training folds, avoiding synthetic leakage, and the scaler never sees test rows. Evaluation is always performed on the untouched, natural class distribution.
- Exclude tier-defining columns. Because each tier is derived from specific columns (Table 3), those columns are removed from the predictors: the five score inputs for German, the six PAY_* fields for Taiwan, and the grade-deterministic fields (grade, sub-grade, interest rate, installment) for Lending Club. Otherwise, a model trivially recovers the labelling rule.
3.5. Models
- MLR. Penalized multinomial logistic regression (glmnet); class probabilities follow the softmax form in Equation (2), in which each risk tier k has its own coefficient vector βₖ (k = 1, …, K, with K = 3 tiers) and the denominator sums over all K tiers; coefficients are interpretable as log-odds relative to a reference tier.
- Random Forest. 500 trees (ranger) with impurity importance; class assigned by majority vote across decorrelated bootstrap trees (Breiman 2001; Liaw and Wiener 2002).
- XGBoost. Regularised gradient boosting (xgboost), max_depth = 6, η = 0.05, 300 rounds, 0.8 subsampling (Chen and Guestrin 2016). Multi-class probabilities are decoded by the framework, avoiding the manual reshape fault discussed in Section 4.5.
- ANN. A single-hidden-layer feed-forward network (nnet).
3.6. Model Persistence and Inference
3.7. Evaluation Metrics
4. Results
4.1. Comparative Performance
4.2. Why Leakage Matters: An Inflated-Accuracy Cautionary Result
4.3. ROC Analysis: High-Tier Discrimination
4.4. Feature Importance
4.5. A Note on a Common Multi-Class Decoding Fault
5. Discussion
5.1. Dataset-Dependent Winners and the Signal–Model Interaction
5.2. Comparison with Existing Literature
5.3. Tree Ensembles and the Interpretability–Accuracy Trade-Off
5.4. Imbalance Shapes the Metrics
5.5. Genuine Prediction, Not Rule Recovery
5.6. Limitations
6. Implications for AI, Banking, and Academia
7. Conclusions and Future Work
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Akinjole, Abisola; Shobayo, Olamilekan; Popoola, Jumoke; Okoyeigbo, Obinna; Ogunleye, Bayode. Ensemble-Based Machine Learning Algorithm for Loan Default Risk Prediction. Mathematics 2024, 12, 3423. [Google Scholar] [CrossRef]
- Alam, Talha Mahboob; Shaukat, Kamran; Hameed, Ibrahim A.; Luo, Suhuai; Sarwar, Muhammad Umer; Shabbir, Shakir; Li, Jiaming; Khushi, Matloob. An Investigation of Credit Card Default Prediction in the Imbalanced Datasets. IEEE Access 2020, 8, 201173–98. [Google Scholar] [CrossRef]
- Apicella, Andrea; Isgrò, Francesco; Prevete, Roberto. Don’t Push the Button! Exploring Data Leakage Risks in Machine Learning and Transfer Learning. Artif. Intell. Rev. 2025, 58, 339. [Google Scholar] [CrossRef]
- Arik, Sercan Ö.; Pfister, Tomas. TabNet: Attentive Interpretable Tabular Learning. Paper presented at the AAAI Conference on Artificial Intelligence, 2021; vol. 35, pp. 6679–87. [Google Scholar] [CrossRef]
- Ariza-Garzón; Janny, Miller; Arroyo, Javier; Caparrini, Antonio; Segovia-Vargas, María-Jesús. Explainability of a Machine Learning Granting Scoring Model in Peer-to-Peer Lending. IEEE Access 2020, 8, 64873–90. [Google Scholar] [CrossRef]
- Ayari, Helmi; Guetari, Ramzi; Kraïem, Naoufel. Machine Learning Powered Financial Credit Scoring: A Systematic Literature Review. Artif. Intell. Rev. 2025, 59, 13. [Google Scholar] [CrossRef]
- Babaei, Golnoosh; Giudici, Paolo; Raffinetti, Emanuela. Explainable FinTech Lending. J. Econ. Bus. 2023, 125–126, 106126. [Google Scholar] [CrossRef]
- Baesens, Bart; Van Gestel, Tony; Viaene, Stijn; Stepanova, Maria; Suykens, Johan; Vanthienen, Jan. Benchmarking State-of-the-Art Classification Algorithms for Credit Scoring. J. Oper. Res. Soc. 2003, 54, 627–35. [Google Scholar] [CrossRef]
- Ballegeer, Matteo; Bogaert, Matthias; Benoit, Dries F. Evaluating the Stability of Model Explanations in Instance-Dependent Cost-Sensitive Credit Scoring. Eur. J. Oper. Res. 2025, 326, 630–40. [Google Scholar] [CrossRef]
- Bhandary, Rakshith; Ghosh, Bidyut Kumar. Credit Card Default Prediction: An Empirical Analysis on Predictive Performance Using Statistical and Machine Learning Methods. J. Risk Financ. Manag. 2025, 18, 23. [Google Scholar] [CrossRef]
- Borisov, Vadim; Leemann, Tobias; Seßler, Kathrin; Haug, Johannes; Pawelczyk, Martin; Kasneci, Gjergji. Deep Neural Networks and Tabular Data: A Survey. IEEE Trans. Neural Netw. Learn. Syst. 2024, 35, 7499–519. [Google Scholar] [CrossRef] [PubMed]
- Breiman, Leo. Random Forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef]
- Bussmann, Niklas; Giudici, Paolo; Marinelli, Dimitri; Papenbrock, Jochen. Explainable Machine Learning in Credit Risk Management. Comput. Econ. 2021, 57, 203–16. [Google Scholar] [CrossRef]
- Chang, Victor; Sivakulasingam, Sharuga; Wang, Hai; Wong, Siu Tung; Ganatra, Meghana Ashok; Luo, Jiabin. Credit Risk Prediction Using Machine Learning and Deep Learning: A Study on Credit Card Customers. Risks 2024, 12, 174. [Google Scholar] [CrossRef]
- Chawla, Nitesh V.; Bowyer, Kevin W.; Hall, Lawrence O.; Kegelmeyer, W. Philip. SMOTE: Synthetic Minority Over-Sampling Technique. J. Artif. Intell. Res. 2002, 16, 321–57. [Google Scholar] [CrossRef]
- Chen, Tianqi; Guestrin, Carlos. XGBoost: A Scalable Tree Boosting System. Paper presented at the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 2016; pp. 785–94. [Google Scholar] [CrossRef]
- Chen, Yan-Ru; Leu, Jenq-Shiou; Huang, Sheng-An; Wang, Jui-Teng; Takada, Jun-Ichi. Predicting Default Risk on Peer-to-Peer Lending Imbalanced Datasets. IEEE Access 2021, 9, 73103–9. [Google Scholar] [CrossRef]
- Chen, Yujia; Calabrese, Raffaella; Martín-Barragán, Belén. Interpretable Machine Learning for Imbalanced Credit Scoring Datasets. Eur. J. Oper. Res. 2024, 312, 357–72. [Google Scholar] [CrossRef]
- Darwish, Jumanah A. Optimization and Prediction of Corporate Credit Rating through Advanced Feature Selection Based on AI and Deep Learning. Alex. Eng. J. 2025, 127, 586–94. [Google Scholar] [CrossRef]
- Dastile, Xolani; Celik, Turgay; Potsane, Moshe. Statistical and Machine Learning Models in Credit Scoring: A Systematic Literature Survey. Appl. Soft Comput. 2020, 91, 106263. [Google Scholar] [CrossRef]
- de Oliveira, Nazário A.; Basso, Leonardo F. C. Advancing Credit Rating Prediction: The Role of Machine Learning in Corporate Credit Rating Assessment. Risks 2025, 13, 116. [Google Scholar] [CrossRef]
- Giudici, Paolo; Centurelli, Mattia; Turchetta, Stefano. Artificial Intelligence Risk Measurement. Expert Syst. With Appl. 2024, 235, 121220. [Google Scholar] [CrossRef]
- Gorishniy, Yury; Rubachev, Ivan; Khrulkov, Valentin; Babenko, Artem. Revisiting Deep Learning Models for Tabular Data. Paper presented at Advances in Neural Information Processing Systems, 2021; vol. 34, pp. 18932–43. [Google Scholar]
- Grinsztajn, Léo; Oyallon, Edouard; Varoquaux, Gaël. Why Do Tree-Based Models Still Outperform Deep Learning on Typical Tabular Data? Paper presented at Advances in Neural Information Processing Systems, 2022; vol. 35, pp. 507–20. [Google Scholar]
- Hand, David J.; Till, Robert J. A Simple Generalisation of the Area under the ROC Curve for Multiple Class Classification Problems. Mach. Learn. 2001, 45, 171–86. [Google Scholar] [CrossRef]
- Hjelkrem, Lars Ole; de Lange, Petter Eilif. Explaining Deep Learning Models for Credit Scoring with SHAP: A Case Study Using Open Banking Data. J. Risk Financ. Manag. 2023, 16, 221. [Google Scholar] [CrossRef]
- Hlongwane, Rivalani; Ramaboa, Kutlwano K. K. M.; Mongwe, Wilson. Enhancing Credit Scoring Accuracy with a Comprehensive Evaluation of Alternative Data. PLoS ONE 2024, 19, e0303566. [Google Scholar] [CrossRef] [PubMed]
- Hofmann, Hans. Statlog (German Credit Data). In UCI Machine Learning Repository; 1994. [Google Scholar] [CrossRef]
- Khatir, Hussin Adam; Almustfa, Ahmed; Bee, Marco. Machine Learning Models and Data-Balancing Techniques for Credit Scoring: What Is the Best Combination? Risks 2022, 10, 169. [Google Scholar] [CrossRef]
- Kapoor, Sayash; Narayanan, Arvind. Leakage and the Reproducibility Crisis in Machine-Learning-Based Science. Patterns 2023, 4, 100804. [Google Scholar] [CrossRef] [PubMed]
- Kim, Ji-Yoon; Cho, Sung-Bae. Towards Repayment Prediction in Peer-to-Peer Social Lending Using Deep Learning. Mathematics 2019, 7, 1041. [Google Scholar] [CrossRef]
- Kuhn, Max; Wickham, Hadley. Tidymodels: A Collection of Packages for Modeling and Machine Learning Using Tidyverse Principles. 2020. Available online: https://www.tidymodels.org (accessed on 14 July 2026).
- Lending Club. Lending Club Loan Data. Kaggle. 2019. Available online: https://www.kaggle.com/datasets/wordsforthewise/lending-club (accessed on 14 July 2026).
- Lessmann, Stefan; Baesens, Bart; Seow, Hsin-Vonn; Thomas, Lyn C. Benchmarking State-of-the-Art Classification Algorithms for Credit Scoring: An Update of Research. Eur. J. Oper. Res. 2015, 247, 124–36. [Google Scholar] [CrossRef]
- Liaw, Andy; Wiener, Matthew. Classification and Regression by randomForest. R News 2002, 2, 18–22. [Google Scholar]
- Lin, Luyun; Wang, Yiqing. SHAP Stability in Credit Risk Management: A Case Study in Credit Card Default Model. Risks 2025, 13, 238. [Google Scholar] [CrossRef]
- Liu, Weiqi; Li, Meifang. The Impact of AI Washing on Enterprises’ Access to Bank Loans: From the Perspective of External Governance. Financ. Res. Lett. 2026, 98, 109884. [Google Scholar] [CrossRef]
- Liu, Wanan; Fan, Hong; Xia, Meng. Tree-Based Heterogeneous Cascade Ensemble Model for Credit Scoring. Int. J. Forecast. 2023, 39, 1593–614. [Google Scholar] [CrossRef] [PubMed]
- Louzada, Francisco; Ara, Anderson; Fernandes, Guilherme B. Classification Methods Applied to Credit Scoring: Systematic Review and Overall Comparison. Surv. Oper. Res. Manag. Sci. 2016, 21, 117–34. [Google Scholar] [CrossRef]
- Lundberg, Scott M.; Lee, Su-In. A Unified Approach to Interpreting Model Predictions. Paper presented at Advances in Neural Information Processing Systems, 2017; vol. 30, pp. 4765–74. [Google Scholar]
- Malekipirbazari, Milad; Aksakalli, Vural. Risk Assessment in Social Lending via Random Forests. Expert Syst. With Appl. 2015, 42, 4621–31. [Google Scholar] [CrossRef]
- Mapfumo, Irvine; Shongwe, Thokozani. Performance Evaluation of Machine Learning and Deep Learning Models for Credit Risk Prediction. J. Risk Financ. Manag. 2026, 19, 210. [Google Scholar] [CrossRef]
- Markov, Anton; Seleznyova, Zinaida; Lapshin, Victor. Credit Scoring Methods: Latest Trends and Points to Consider. J. Financ. Data Sci. 2022, 8, 180–201. [Google Scholar] [CrossRef]
- Mestiri, Sami. Credit Scoring Using Machine Learning and Deep Learning-Based Models. Data Sci. Financ. Econ. 2024, 4, 236–48. [Google Scholar] [CrossRef]
- Mushava, Jonah; Murray, Michael. Flexible Loss Functions for Binary Classification in Gradient-Boosted Decision Trees: An Application to Credit Scoring. Expert Syst. With Appl. 2024, 238, 121876. [Google Scholar] [CrossRef]
- Nguyen, Thi Hong Thuy; Nguyen, Thi Vinh Ha; Nguyen, Nam Trung; Vu, Thi Thanh Binh; Nguyen, Thu Hang; The Binh Vu. Comparing the Effectiveness of Machine Learning and Deep Learning Models in Student Credit Scoring: A Case Study in Vietnam. Risks 2025, 13, 99. [Google Scholar] [CrossRef]
- Oreski, Goran. Synthesizing Credit Data Using Autoencoders and Generative Adversarial Networks. Knowl.-Based Syst. 2023, 274, 110646. [Google Scholar] [CrossRef]
- Paz, Álex; Crawford, Broderick; Monfroy, Eric; Barrera-García, José; Fritz, Álvaro Peña; Soto, Ricardo; Cisternas-Caneo, Felipe; Yáñez, Andrés. Machine Learning and Metaheuristics Approach for Individual Credit Risk Assessment: A Systematic Literature Review. Biomimetics 2025, 10, 326. [Google Scholar] [CrossRef] [PubMed]
- R Core Team. R: A Language and Environment for Statistical Computing; R Foundation for Statistical Computing: Vienna, 2026; Available online: https://www.R-project.org (accessed on 14 July 2026).
- Ribeiro, Marco Tulio; Singh, Sameer; Guestrin, Carlos. “Why Should I Trust You?”: Explaining the Predictions of Any Classifier. Paper presented at the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 2016; pp. 1135–44. [Google Scholar] [CrossRef]
- Rizinski, Maryan; Trajanov, Dimitar. AI Agents in Finance and Fintech: A Scientific Review of Agent-Based Systems, Applications, and Future Horizons. Comput. Mater. Contin. 2026, 86, 1–34. [Google Scholar] [CrossRef]
- Serrano-Cinca, Carlos; Gutiérrez-Nieto, Begoña; López-Palacios, Luz. Determinants of Default in P2P Lending. PLoS ONE 2015, 10, e0139427. [Google Scholar] [CrossRef] [PubMed]
- Shi, Si; Tse, Rita; Luo, Wuman; D’Addona, Stefano; Pau, Giovanni. Machine Learning-Driven Credit Risk: A Systemic Review. Neural Comput. Appl. 2022, 34, 14327–39. [Google Scholar] [CrossRef]
- Suhadolnik, Nicolas; Ueyama, Jo; Da Silva, Sergio. Machine Learning for Enhanced Credit Risk Assessment: An Empirical Approach. J. Risk Financ. Manag. 2023, 16, 496. [Google Scholar] [CrossRef]
- Talaat, Fatma M.; Aljadani, Abdussalam; Badawy, Mahmoud; Elhosseini, Mostafa. Toward Interpretable Credit Scoring: Integrating Explainable Artificial Intelligence with Deep Learning for Credit Card Default Prediction. Neural Comput. Appl. 2024, 36, 4847–65. [Google Scholar] [CrossRef]
- Trivedi, Shrawan Kumar. A Study on Credit Scoring Modeling with Different Feature Selection and Machine Learning Approaches. Technol. Soc. 2020, 63, 101413. [Google Scholar] [CrossRef]
- Yeh, I-Cheng; Lien, Che-Hui. The Comparisons of Data Mining Techniques for the Predictive Accuracy of Probability of Default of Credit Card Clients. Expert Syst. With Appl. 2009, 36, 2473–80. [Google Scholar] [CrossRef]
- Zhang, Xiaoming; Yu, Lean. Consumer Credit Risk Assessment: A Review from the State-of-the-Art Classification Algorithms, Data Traits, and Learning Methods. Expert Syst. With Appl. 2024, 237, 121484. [Google Scholar] [CrossRef]
- Zheng, Mengmeng; Zhang, Lu; Tripe, David; Zhang, Yuming. Can Artificial Intelligence Mitigate Greenwashed Green Credit? Evidence from Loan Contracts of Chinese Listed Firms. Int. Rev. Financ. Anal. 2026, 110, 104948. [Google Scholar] [CrossRef]
- Zhou, Ying; Shen, Long; Ballester, Laura. A Two-Stage Credit Scoring Model Based on Random Forest: Evidence from Chinese Small Firms. Int. Rev. Financ. Anal. 2023, 89, 102755. [Google Scholar] [CrossRef]
- Zhu, Xu; Chu, Qingyong; Song, Xinchang; Hu, Ping; Peng, Lu. Explainable Prediction of Loan Default Based on Machine Learning Models. Data Sci. Manag. 2023, 6, 123–33. [Google Scholar] [CrossRef]















| Study | Venue/Year | Dataset(s) | Models Compared | Headline Finding (as Published) | Comparability Limitation |
| Baesens et al. | JORS 2003 | 8 credit sets incl. German | 17 classifiers | Small differences among well-tuned classifiers | Binary; pre-leakage-era protocols |
| Lessmann et al. | EJOR 2015 | 8 credit sets incl. German | 41 classifiers | Heterogeneous ensembles best overall | Binary; per-dataset preprocessing varies |
| Yeh and Lien | ESWA 2009 | Taiwan | 6 methods incl. ANN | ANN best for default probability | Binary; no leakage audit |
| Malekipirbazari and Aksakalli | ESWA 2015 | Lending Club | RF, SVM, LR, k-NN | Random forest best | Binary; random split, single vintage |
| Serrano-Cinca et al. | PLoS ONE 2015 | Lending Club | LR, survival analysis | Grade and indebtedness drive default | Explanatory focus, not a benchmark |
| Trivedi | Technol. Soc. 2020 | German | 5 classifiers × 3 FS methods | Ranking shifts with feature selection | Binary; single small corpus |
| Alam et al. | IEEE Access 2020 | Taiwan + others | Classifier × resampling grid | GBDT + oversampling dominate | Resampling outside CV folds |
| Ariza-Garzón et al. | IEEE Access 2020 | Lending Club | XGBoost + SHAP | Explainable granting model | Binary; single corpus |
| Hussin Adam Khatir and Bee | Risks 2022 | German | Classifiers × FS × balancing | RF + RFE + random oversampling best | Binary; single small corpus |
| Suhadolnik et al. | JRFM 2023 | Lending Club (1.3M loans) | 10 statistical and ML algorithms | XGBoost best (80.4% vs. LR 65.7%) | Binary; single corpus |
| Chang et al. | Risks 2024 | Credit-card customers | ML and DL models | Ensembles strongest | Binary; single corpus |
| Nguyen et al. | Risks 2025 | Student loans (Vietnam) | RF, GB, SVM, DNN | DNN best (85.6% accuracy) | Binary; private survey data |
| Lin and Wang | Risks 2025 | Taiwan | 100 XGBoost seeds + SHAP | SHAP rank stability tied to importance level | Explainability focus; binary |
| Mapfumo and Shongwe | JRFM 2026 | German + Taiwan | 10 ML/DL × SMOTE variants | SMOTE-ENN + MLP: F1 0.928 (German) | Evaluation after synthetic resampling |
| This study | — | German + Taiwan + Lending Club | MLR, RF, XGBoost, ANN | Leakage-controlled ordinal three-tier benchmark | Proxy tiers on two corpora (stated) |
| Dataset | Instances | Predictors ¹ | Low/Medium/High |
| German Credit | 1000 | ~15 | 25/436/539 |
| Taiwan Default | 30,000 | ~17 | 19,931/8876/1193 |
| Lending Club | 48,808 | ~60 | 17,370/24,845/6593 |
| Dataset | Three-Tier Rule |
| German | Engineered risk score s ∈ [0,100]; s ≤ 33 Low, 34–66 Medium, ≥ 67 High. |
| Taiwan | Worst recent delay over PAY_0–PAY_6: ≤ 0 Low, 1–2 months Medium, ≥ 3 months High. |
| Lending Club | Lender grade collapsed: A–B Low, C–D Medium, E–G High. |
| Tier | Lending Action |
| Low | Streamlined approval (high creditworthiness) |
| Medium | Conditional approval; monitor/extra documentation |
| High | Decline or require collateral/guarantor |
| Dataset | Model | Accuracy | Precision | Recall | F1 | AUC (Hand–Till) | κ | QWK |
| German | MLR | 0.5572 | 0.4671 | 0.5328 | 0.4337 | 0.6759 | 0.2615 | 0.3066 |
| Random Forest | 0.6318 | 0.4265 | 0.4297 | 0.4273 | 0.6480 | 0.2855 | 0.2756 | |
| XGBoost | 0.6219 | 0.4197 | 0.4228 | 0.4205 | 0.6787 | 0.2662 | 0.2582 | |
| ANN | 0.5075 | 0.3595 | 0.3435 | 0.3502 | 0.5518 | 0.0841 | 0.1683 | |
| Taiwan | MLR | 0.4646 | 0.4132 | 0.5282 | 0.3694 | 0.7190 | 0.1454 | 0.2327 |
| Random Forest | 0.8649 | 0.7902 | 0.7581 | 0.7730 | 0.9049 | 0.7076 | 0.7693 | |
| XGBoost | 0.8604 | 0.7546 | 0.7798 | 0.7652 | 0.9145 | 0.7008 | 0.7688 | |
| ANN | 0.7260 | 0.5773 | 0.6961 | 0.6060 | 0.8610 | 0.4554 | 0.6074 | |
| Lending Club | MLR | 0.6020 | 0.5797 | 0.6452 | 0.5866 | 0.8187 | 0.3864 | 0.5752 |
| Random Forest | 0.6810 | 0.6653 | 0.5973 | 0.6175 | 0.8239 | 0.4406 | 0.5672 | |
| XGBoost | 0.6969 | 0.6819 | 0.6155 | 0.6358 | 0.8404 | 0.4701 | 0.5994 | |
| ANN | 0.6153 | 0.5826 | 0.6320 | 0.5940 | 0.8121 | 0.3916 | 0.5651 |
| Prior Study | Corpus | Reported (Binary, as Published) | This Study (Three-Class, Leakage-Controlled) |
| Mapfumo and Shongwe (2026) | German | Accuracy 95.4%, F1 0.928 (SMOTE-ENN + MLP) | Best accuracy 0.632 (RF); best QWK 0.307 (MLR) |
| Lessmann et al. (2015) | German | Ensembles lead; small margins among tuned classifiers | Qualitative pattern holds only on larger corpora; MLR leads QWK on German |
| Hussin Adam Khatir and Bee (2022) | German | RF + RFE + random oversampling best combination | RF leads accuracy (0.632) but MLR leads ordinal agreement |
| Yeh and Lien (2009) | Taiwan | ANN best for default-probability estimation | RF best (accuracy 0.865, QWK 0.769); ANN trails ensembles |
| Mapfumo and Shongwe (2026) | Taiwan | F1 0.789 (SMOTE-ENN + RF) | RF F1 0.773 on the natural test distribution |
| Malekipirbazari and Aksakalli (2015) | Lending Club | Random forest best | XGBoost best (accuracy 0.697, QWK 0.599); RF second |
| Suhadolnik et al. (2023) | Lending Club | XGBoost best: accuracy 80.4% vs. LR 65.7% | Same ranking: XGBoost first, MLR last on accuracy (0.697 vs. 0.602, three-class) |
| Ariza-Garzón et al. (2020) | Lending Club | Boosting + SHAP; strong binary discrimination | Consistent: boosting leads; SHAP integration deferred to future work |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).