Submitted:
14 August 2025
Posted:
15 August 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Related Work
2.1. Ensemble Learning for Fraud Detection
2.2. Class Imbalance Handling in Fraud Detection
2.3. IEEE-CIS Dataset Studies
3. Methodology
3.1. Dataset Description and Characteristics
3.2. Data Preprocessing and Feature Engineering
3.2.1. Feature Preprocessing Pipeline
3.2.2. Missing Value Analysis and Treatment
- Feature Removal: Features with >95% missing values are removed to prevent sparse representations.
- Strategic Imputation: For categorical features with <95% missing values, we create explicit "missing" categories to capture informational content.
- Numerical Imputation: Numerical features employ median imputation within fraud/legitimate groups separately to preserve class-specific distributions.
- Missingness Indicators: Binary indicators are created for features with <20% missing values to capture missingness patterns as potential fraud signals.
3.2.3. Feature Engineering Strategy
3.3. Handling Class Imbalance
3.4. Ensemble Model Architecture
3.4.1. Base Learner Selection and Diversity Strategy
3.4.2. Meta-Learning and Model Combination Strategy
4. Experimental Setup
4.1. Implementation Environment and Tools
4.2. Hyperparameter Optimization
4.3. Evaluation Metrics and Performance Assessment
4.4. Cross-Validation and Model Selection
4.5. Statistical Significance Testing
5. Results and Analysis
5.1. Overall Performance Comparison
5.2. Individual Algorithm Performance Analysis
5.3. Class Imbalance Handling Effectiveness
5.4. Feature Engineering Impact Assessment
5.5. Ensemble Architecture Analysis
6. Discussion
6.1. Key Findings and Implications
6.2. Practical Implementation Considerations
6.3. Limitations and Future Directions
7. Conclusion
References
- Khalid, A.R.; Owoh, N.; Uthmani, O.; Ashawa, M.; Osamor, J.; Adejoh, J. Enhancing credit card fraud detection: an ensemble machine learning approach. Big Data and Cognitive Computing 2024, 8, 6. [Google Scholar] [CrossRef]
- Homaei, M.H.; Caro Lindo, A.; Sancho Núñez, J.C.; Mogollón Gutiérrez, O.; Alonso Díaz, J. The Role of Artificial Intelligence in Digital Twin’s Cybersecurity. In Proceedings of the XVII Reunión Española sobre Criptología y Seguridad de la Información (RECSI 2022); 2022. [Google Scholar]
- Homaei, M.; Mogollón-Gutiérrez, O.; Sancho, J.C.; Ávila, M.; Caro, A. A review of digital twins and their application in cybersecurity based on artificial intelligence. Artificial Intelligence Review 2024, 57. [Google Scholar] [CrossRef]
- Gandhar, A.; Gupta, K.; Pandey, A.K.; Raj, D. Fraud detection using machine learning and deep learning. SN Computer Science 2024, 5, 453. [Google Scholar] [CrossRef]
- Moradi, F.; Tarif, M.; Homaei, M. A Systematic Review of Machine Learning in Credit Card Fraud Detection. Preprint 2024. [Google Scholar]
- Mienye, I.D.; Jere, N. Deep learning for credit card fraud detection: A review of algorithms, challenges, and solutions. IEEE Access 2024. [Google Scholar] [CrossRef]
- Chen, Y.; Zhao, C.; Xu, Y.; Nie, C. Year-over-Year Developments in Financial Fraud Detection via Deep Learning: A Systematic Literature Review. arXiv 2025, arXiv:2502.00201. [Google Scholar]
- Fernández, A.; García, S.; Galar, M.; Prati, R.C.; Krawczyk, B.; Herrera, F. Learning from imbalanced data sets. 2018, 10. [Google Scholar] [CrossRef]
- Talukder, M.A.; Khalid, M.; Uddin, M.A. An integrated multistage ensemble machine learning model for fraudulent transaction detection. Journal of Big Data 2024, 11. [Google Scholar] [CrossRef]
- Vesta Corporation. IEEE-CIS Fraud Detection Dataset. Kaggle Competition, 2019.
- Suganya, S.S.; Nishanth, S.; Mohanadevi, D. Ensemble Learning Approaches for Fraud Detection in Financial Transactions. 2023 2nd International Conference on Automation, Computing and Renewable Systems (ICACRS); 2023; pp. 805–810. [Google Scholar]
- Almalki, F.; Masud, M. Financial Fraud Detection Using Explainable AI and Stacking Ensemble Methods. arXiv 2025, arXiv:2505.10050. [Google Scholar] [CrossRef]
- Zhao, X.; Zhang, Q.; Zhang, C. Enhancing Transaction Fraud Detection with a Hybrid Machine Learning Model. 2024 IEEE 4th International Conference on Electronic Technology, Communication and Information (ICETCI); 2024; pp. 427–432. [Google Scholar]
- Talukder, M.A.; Khalid, M.; Uddin, M.A. An integrated multistage ensemble machine learning model for fraudulent transaction detection. Journal of Big Data 2024, 11. [Google Scholar] [CrossRef]
- Chawla, N.V.; Bowyer, K.W.; Hall, L.O.; Kegelmeyer, W.P. SMOTE: synthetic minority over-sampling technique. Journal of artificial intelligence research 2002, 16, 321–357. [Google Scholar] [CrossRef]
- Elreedy, D.; Atiya, A.F.; Kamalov, F. A theoretical distribution analysis of synthetic minority oversampling technique (SMOTE) for imbalanced learning. Machine Learning 2024, 113, 4903–4923. [Google Scholar] [CrossRef]
- Salehi, A.R.; Khedmati, M. A cluster-based SMOTE both-sampling (CSBBoost) ensemble algorithm for classifying imbalanced data. Scientific Reports 2024, 14, 5152. [Google Scholar] [CrossRef] [PubMed]
- Li, J.; Wang, H.; Zhang, Y.; Chen, L. Imbalanced Data Classification Based on Improved Random-SMOTE and Feature Standard Deviation. Mathematics 2024, 12, 1709. [Google Scholar] [CrossRef]
- Papers with Code. IEEE CIS Fraud Detection Dataset. Online Repository, 2024.
- Saito, T.; Rehmsmeier, M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PloS one 2015, 10, e0118432. [Google Scholar] [CrossRef] [PubMed]
- Grinsztajn, L.; Oyallon, E.; Varoquaux, G. Why do tree-based models still outperform deep learning on typical tabular data? Advances in Neural Information Processing Systems 2022, 35, 507–520. [Google Scholar]

| (a) Methodology Comparison | |||||
| Study | Ensemble Method | Base Classifiers | Key Contribution | ||
| [1] | Bagging + Boosting | SVM, KNN, RF, Bagging, Boosting | Imbalance-aware ensemble integration | ||
| [11] | Random Forest | Decision Trees | Advanced feature engineering | ||
| [12] | Stacking Ensemble | XGBoost, LightGBM, CatBoost | Explainable AI integration | ||
| [13] | Hybrid ML Model | Multiple Algorithms | Comprehensive feature engineering | ||
| (b) Dataset and Performance Comparison | |||||
| Study | Dataset | Size | Fraud Rate | Features | Best Performance |
| [1] | Custom | Not Specified | Imbalanced | Not Specified | Improved accuracy through ensemble strategies |
| [11] | Financial | Large Dataset | Not Specified | Engineered | Enhanced detection via feature engineering |
| [12] | Financial | Not Specified | Not Specified | Not Specified | 0.99 AUC-ROC with explainable stacking |
| [13] | IEEE-CIS | 590,540 | 3.5% | 431 | Significant improvement via hybrid learning |
| Preprocessing Step | Features Remaining | Features Removed |
|---|---|---|
| Original IEEE-CIS Dataset | 431 | - |
| Remove Features >95% Missing | 298 | 133 |
| Remove Zero-Variance Features | 276 | 22 |
| Remove Highly Correlated (>0.98) | 203 | 73 |
| Remove Low Information Gain (<0.001) | 167 | 36 |
| Baseline Feature Set | 167 | 264 total |
| Feature Engineering Phase | ||
| + Temporal Features | 182 | +15 |
| + Amount Engineering | 194 | +12 |
| + Aggregation Features | 222 | +28 |
| + Interaction Features | 247 | +25 |
| Final Feature Set | 247 | +80 engineered |
| Algorithm | Optimized Parameters |
|---|---|
| XGBoost | n_estimators=500, max_depth=6, learning_rate=0.1, subsample=0.8, colsample_bytree=0.8 |
| LightGBM | n_estimators=400, max_depth=7, learning_rate=0.1, feature_fraction=0.8, bagging_fraction=0.8 |
| Random Forest | n_estimators=300, max_depth=10, min_samples_split=5, min_samples_leaf=2 |
| CatBoost | iterations=400, depth=8, learning_rate=0.1, l2_leaf_reg=3 |
| Neural Network | hidden_layers=(100,50), alpha=0.001, learning_rate_init=0.01 |
| Algorithm | AUC-ROC | AUC-PR | Training Time | Inference Time |
|---|---|---|---|---|
| XGBoost | 0.887±0.004 | 0.834±0.006 | 18.3 min | 45 ms |
| LightGBM | 0.882±0.003 | 0.828±0.005 | 12.7 min | 38 ms |
| CatBoost | 0.873±0.004 | 0.821±0.006 | 24.1 min | 52 ms |
| Random Forest | 0.869±0.005 | 0.802±0.007 | 8.9 min | 28 ms |
| Neural Network | 0.841±0.006 | 0.786±0.008 | 15.6 min | 35 ms |
| K-NN | 0.826±0.007 | 0.771±0.009 | 2.1 min | 125 ms |
| Logistic Regression | 0.829±0.005 | 0.743±0.008 | 1.8 min | 12 ms |
| Technique | AUC-ROC | AUC-PR | F1-Score | Recall@95%P |
|---|---|---|---|---|
| SMOTE + Stacking | 0.918±0.003 | 0.891±0.005 | 0.856±0.004 | 0.847 |
| Borderline-SMOTE + Stacking | 0.912±0.004 | 0.884±0.006 | 0.834±0.005 | 0.823 |
| ADASYN + Stacking | 0.908±0.004 | 0.876±0.006 | 0.825±0.005 | 0.814 |
| SMOTE + Tomek + Stacking | 0.915±0.003 | 0.888±0.005 | 0.845±0.004 | 0.836 |
| No Sampling + Stacking | 0.863±0.005 | 0.812±0.007 | 0.774±0.006 | 0.752 |
| SMOTE + XGBoost | 0.887±0.004 | 0.834±0.006 | 0.798±0.005 | 0.781 |
| No Sampling + XGBoost | 0.821±0.006 | 0.758±0.009 | 0.701±0.007 | 0.679 |
| Feature Set | AUC-ROC | AUC-PR | F1-Score | Features | AUC-PR |
|---|---|---|---|---|---|
| Complete Pipeline | 0.918±0.003 | 0.891±0.005 | 0.856±0.004 | 247 | - |
| - Interaction Features (25) | 0.905±0.004 | 0.873±0.006 | 0.834±0.005 | 222 | -0.018 |
| - Aggregation Features (28) | 0.892±0.004 | 0.856±0.006 | 0.812±0.005 | 219 | -0.035 |
| - Temporal Features (15) | 0.883±0.004 | 0.847±0.006 | 0.801±0.005 | 232 | -0.044 |
| - Amount Engineering (12) | 0.897±0.004 | 0.863±0.006 | 0.825±0.005 | 235 | -0.028 |
| Baseline Features Only | 0.851±0.005 | 0.789±0.007 | 0.743±0.006 | 167 | -0.102 |
| Ensemble Method | AUC-ROC | AUC-PR | Training Time | p-value* |
|---|---|---|---|---|
| Stacking (Proposed) | 0.918±0.003 | 0.891±0.005 | 45.7 min | - |
| Weighted Voting | 0.901±0.003 | 0.847±0.005 | 32.4 min | 0.004 |
| Blending | 0.895±0.004 | 0.842±0.006 | 38.9 min | 0.002 |
| Simple Voting | 0.878±0.004 | 0.823±0.006 | 31.8 min | <0.001 |
| Bagging (RF) | 0.869±0.005 | 0.808±0.007 | 26.3 min | <0.001 |
| AdaBoost | 0.841±0.006 | 0.785±0.008 | 41.2 min | <0.001 |
| *Bonferroni-corrected = 0.0083 | ||||
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).