Submitted:
09 January 2025
Posted:
09 January 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
- Hybrid Resampling Techniques: By combining SMOTE and SMOTEENN, the framework effectively addresses class imbalance while minimizing noise in synthetic samples. This ensures a balanced dataset, which is crucial for improving the sensitivity of minority class predictions.
- Attention-Based Feature Engineering: The attention mechanism adaptively prioritizes significant features, capturing both local and global interactions. This dynamic feature weighting enhances the representation of critical factors contributing to stroke prediction.
- Ensemble and Meta-Learning Integration: By integrating Random Forest and LightGBM as base models and leveraging a deep learning meta-model, the framework optimizes the synergy between diverse classifiers. This approach captures higher-order interactions, improving decision boundaries and overall predictive accuracy.
- Explainable AI: The inclusion of SHAP ensures that the model’s decision-making model is transparent, providing clinicians with actionable insights into feature contributions. This enhances trust in the system, making it more suitable for real-world clinical adoption.
- Extensive Validation: The framework’s performaframework’srously evaluated on three benchmark datasets, DF-1, DF-2, and DF-3, demonstrating consistent superiority in Accuracy, F1-Score, and ROC-AUC metrics. This validates its robustness and highlights its generalizability across diverse datasets.
2. Related Work
3. Datasets
3.1. Dataset Description
3.2. Class Imbalance in DF1, DF2 and Df-3 Datasets



3.3. Feature Distributions
4. Methodology
4.1. Data Preprocessing
4.1.1. Handling Missing Values
4.2. One-Hot Encoding
4.2.1. Standardizing Numerical Features
4.2.2. Feature Correlation and Redundancy Removal
4.3. Imbalance Handling
4.4. Feature Selection
4.5. Model Architecture
4.5.1. Base Models
- Random Forest (RF): A tree-based ensemble method that combines predictions from multiple decision trees. Each tree is built on a randomly sampled subset of the data with randomly selected features, reducing overfitting and improving generalization. The Random Forest prediction is computed as:where represents the prediction of the i-th decision tree, and N is the total number of trees.
- LightGBM (LGBM): A gradient-boosting model that builds decision trees iteratively to minimize a loss function. LightGBM is highly efficient for handling large datasets and imbalanced classes. The model minimizes the loss function :where ℓ is the loss function (e.g., binary cross-entropy), is the prediction at iteration t, and n is the number of data points.
4.5.2. Meta-Learning Model
- Input Layer: Combines the probability outputs and from the Random Forest and LightGBM models:
-
Hidden Layers: Two fully connected dense layers with ReLU activation functions capture non-linear interactions in the input space. Dropout layers are applied for regularization to reduce overfitting:Where is the activation at layer l, and are the weights and biases.
- Output Layer: The final layer is a single neuron with a sigmoid activation function, outputting the final probability prediction:
4.5.3. Meta-Learning Concept and Final Prediction
5. Experimental Results
5.1. Performance Measure over datasets
5.1.1. DF-1 Dataset Results
5.1.2. DF-2 Dataset Results
5.1.3. DF-3 Dataset Results
5.2. Explainable Predictions Using SHAP
5.2.1. Global Feature Importance
5.2.2. Feature Dependency and Interaction
5.2.3. Localized Explanations for Individual Predictions
5.2.4. Cumulative Feature Contributions
5.3. Comparison with State-of-the-Art Methods
5.3.1. DF-1 Dataset
5.3.2. DF-2 Dataset
| Refs | Model Used | Accuracy (%) | F1-Score (%) |
|---|---|---|---|
| [38] | LR, DT, RF, SVM, and NB | 95.5 | 94.5 |
| [41] | Category Boosting Classifier (CBC) | 97 | 96 |
| [39] | DT, SVM, and LR | 95.49 | 96 |
| [40] | Deep neural Networks | 95.49 | 96 |
| [42] | RF | 97.19 | 97.15 |
| [37] | Stacking Algorithms | 97.98 | 98.0 |
| [43] | Boosting Algorithms | 97.97 | 93.0 |
| Proposed Method | Meta + (SMOTE-SMOTEENN) | 98.02 | 98.25 |
5.3.3. DF-3 Dataset
| Refs | Model Used | Accuracy (%) |
|---|---|---|
| [44] | RF | 95.3 |
| [45] | BSPE | 95.3 |
| [46] | LGBM | 94.53 |
| [48] | RF | 98.94 |
| [51] | RF | 97.2 |
| [49] | RF | 99.07 |
| [47] | Voting | 97.0 |
| [50] | DT | 93 |
| Proposed Method | Meta + (SMOTE-SMOTEENN) | 99.34 |
5.3.4. Summary of Comparative Analysis
6. Discussion
6.1. Impact of Resampling Techniques on Model Sensitivity
6.2. Performance Variations with Resampling Strategies
6.3. Significance of High Predictive Metrics in Clinical Applications
6.4. Enhancing Model Interpretability Through SHAP Analysis
6.5. Broader Applicability of the Proposed Framework
7. Conclusions and Future Work
References
- Collaborators, G.S.; et al. Global, regional, and national burden of stroke and its risk factors, 1990–2019: a systematic analysis for the Global Burden of Disease Study 2019. The Lancet. Neurology 2021, 20, 795. [Google Scholar]
- Saini, V.; Guada, L.; Yavagal, D.R. Global epidemiology of stroke and access to acute ischemic stroke interventions. Neurology 2021, 97, S6–S16. [Google Scholar] [CrossRef] [PubMed]
- Saceleanu, V.M.; Toader, C.; Ples, H.; Covache-Busuioc, R.A.; Costin, H.P.; Bratu, B.G.; Dumitrascu, D.I.; Bordeianu, A.; Corlatescu, A.D.; Ciurea, A.V. Integrative approaches in acute ischemic stroke: from symptom recognition to future innovations. Biomedicines 2023, 11, 2617. [Google Scholar] [CrossRef] [PubMed]
- Al Duhayyim, M.; Abbas, S.; Al Hejaili, A.; Kryvinska, N.; Almadhor, A.; Mohammad, U.G. An Ensemble Machine Learning Technique for Stroke Prognosis. Computer Systems Science & Engineering 2023, 47. [Google Scholar]
- Correa, R.; Shaan, M.; Trivedi, H.; Patel, B.; Celi, L.A.G.; Gichoya, J.W.; Banerjee, I. A Systematic review of ‘Fair’AI model development for image classification and prediction. Journal of Medical and Biological Engineering 2022, 42, 816–827. [Google Scholar] [CrossRef]
- Adi Pratama, F.R.; Oktora, S.I. Synthetic Minority Over-sampling Technique (SMOTE) for handling imbalanced data in poverty classification. Statistical Journal of the IAOS 2023, 39, 233–239. [Google Scholar] [CrossRef]
- Muntasir Nishat, M.; Faisal, F.; Jahan Ratul, I.; Al-Monsur, A.; Ar-Rafi, A.M.; Nasrullah, S.M.; Reza, M.T.; Khan, M.R.H. A Comprehensive Investigation of the Performances of Different Machine Learning Classifiers with SMOTE-ENN Oversampling Technique and Hyperparameter Optimization for Imbalanced Heart Failure Dataset. Scientific Programming 2022, 2022, 3649406. [Google Scholar] [CrossRef]
- Yang, Y.; Lv, H.; Chen, N. A survey on ensemble learning under the era of deep learning. Artificial Intelligence Review 2023, 56, 5545–5589. [Google Scholar] [CrossRef]
- Rufo, D.D.; Debelee, T.G.; Ibenthal, A.; Negera, W.G. Diagnosis of diabetes mellitus using gradient boosting machine (LightGBM). Diagnostics 2021, 11, 1714. [Google Scholar] [CrossRef] [PubMed]
- Monteiro, J.P.; Ramos, D.; Carneiro, D.; Duarte, F.; Fernandes, J.M.; Novais, P. Meta-learning and the new challenges of machine learning. International Journal of Intelligent Systems 2021, 36, 6240–6272. [Google Scholar] [CrossRef]
- Chaudhari, S.; Mithal, V.; Polatkan, G.; Ramanath, R. An attentive survey of attention models. ACM Transactions on Intelligent Systems and Technology (TIST) 2021, 12, 1–32. [Google Scholar] [CrossRef]
- Nohara, Y.; Matsumoto, K.; Soejima, H.; Nakashima, N. Explanation of machine learning models using shapley additive explanation and application for real data in hospital. Computer Methods and Programs in Biomedicine 2022, 214, 106584. [Google Scholar] [CrossRef] [PubMed]
- Hassija, V.; Chamola, V.; Mahapatra, A.; Singal, A.; Goel, D.; Huang, K.; Scardapane, S.; Spinelli, I.; Mahmud, M.; Hussain, A. Interpreting black-box models: a review on explainable artificial intelligence. Cognitive Computation 2024, 16, 45–74. [Google Scholar] [CrossRef]
- He, H.; Garcia, E.A. Learning from imbalanced data. IEEE Transactions on Knowledge and Data Engineering 2009, 21, 1263–1284. [Google Scholar] [CrossRef]
- Ning, Q.; Zhao, X.; Ma, Z. A novel method for Identification of Glutarylation sites combining Borderline-SMOTE with Tomek links technique in imbalanced data. IEEE/ACM Transactions on Computational Biology and Bioinformatics 2021, 19, 2632–2641. [Google Scholar] [CrossRef] [PubMed]
- Tripathy, G.; Sharaff, A. AEGA: enhanced feature selection based on ANOVA and extended genetic algorithm for online customer review analysis. The Journal of Supercomputing 2023, 79, 13180–13209. [Google Scholar] [CrossRef]
- Shou, Y.; Liu, H.; Cao, X.; Meng, D.; Dong, B. A low-rank matching attention based cross-modal feature fusion method for conversational emotion recognition. IEEE Transactions on Affective Computing 2024. [Google Scholar] [CrossRef]
- Haleem, A.; Javaid, M.; Singh, R.P.; Suman, R. Medical 4.0 technologies for healthcare: Features, capabilities, and applications. Internet of Things and Cyber-Physical Systems 2022, 2, 12–30. [Google Scholar] [CrossRef]
- Tarsha Kurdi, F.; Amakhchan, W.; Gharineiat, Z. Random forest machine learning technique for automatic vegetation detection and modelling in LiDAR data. International Journal of Environmental Sciences and Natural Resources 2021, 28. [Google Scholar]
- Konstantinov, A.V.; Utkin, L.V. Interpretable machine learning with an ensemble of gradient boosting machines. Knowledge-Based Systems 2021, 222, 106993. [Google Scholar] [CrossRef]
- Rasmy, L.; Xiang, Y.; Xie, Z.; Tao, C.; Zhi, D. Med-BERT: pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction. NPJ digital medicine 2021, 4, 86. [Google Scholar] [CrossRef]
- Khan, K. A Framework for Meta-Learning in Dynamic Adaptive Streaming over HTTP. International Journal of Computing 2023, 12. [Google Scholar]
- Kalusivalingam, A.K.; Sharma, A.; Patel, N.; Singh, V. Leveraging SHAP and LIME for Enhanced Explainability in AI-Driven Diagnostic Systems. International Journal of AI and ML 2021, 2. [Google Scholar]
- Salman, A.H.; Al-Jawher, W.A.M. Performance Comparison of Support Vector Machines, AdaBoost, and Random Forest for Sentiment Text Analysis and Classification. Journal Port Science Research 2024, 7, 300–311. [Google Scholar] [CrossRef]
- Zhang, Y.; Lu, S.; Zhou, X.; Yang, M.; Wu, L.; Liu, B.; Phillips, P.; Wang, S. Comparison of machine learning methods for stationary wavelet entropy-based multiple sclerosis detection: decision tree, k-nearest neighbors, and support vector machine. Simulation 2016, 92, 861–871. [Google Scholar] [CrossRef]
- Mondal, S.; Ghosh, S.; Nag, A. Brain stroke prediction model based on boosting and stacking ensemble approach. International Journal of Information Technology 2024, 16, 437–446. [Google Scholar] [CrossRef]
- Mienye, I.D.; Jere, N. Optimized ensemble learning approach with explainable AI for improved heart disease prediction. Information 2024, 15, 394. [Google Scholar] [CrossRef]
- Khademi, Z.; Ebrahimi, F.; Kordy, H.M. A transfer learning-based CNN and LSTM hybrid deep learning model to classify motor imagery EEG signals. Computers in biology and medicine 2022, 143, 105288. [Google Scholar] [CrossRef] [PubMed]
- S, T. Cerebral Stroke Prediction - Imbalanced Dataset. Available online: https://www.kaggle.com/datasets/shashwatwork/cerebral-stroke-predictionimbalaced-dataset (accessed on 15 December 2022).
- Brain Stroke Dataset. Available online: https://www.kaggle.com/datasets/jillanisofttech/brain-stroke-dataset (accessed on 3 November 2022).
- Stroke Prediction Dataset. Available online: https://www.kaggle.com/datasets/fedesoriano/stroke-prediction-dataset (accessed on 15 December 2022).
- Wang, L.; Jiang, S.; Jiang, S. A feature selection method via analysis of relevance, redundancy, and interaction. Expert Systems with Applications 2021, 183, 115365. [Google Scholar] [CrossRef]
- Li, A.; Mueller, A.; English, B.; Arena, A.; Vera, D.; Kane, A.E.; Sinclair, D.A. Novel feature selection methods for construction of accurate epigenetic clocks. PLoS Computational Biology 2022, 18, e1009938. [Google Scholar] [CrossRef]
- Naseer, A.; Jalal, A. Pixels to precision: features fusion and random forests over labelled-based segmentation. In Proceedings of the 2023 20th International Bhurban Conference on Applied Sciences and Technology (IBCAST); IEEE, 2023; pp. 1–6. [Google Scholar]
- Lokker, C.; Abdelkader, W.; Bagheri, E.; Parrish, R.; Cotoi, C.; Navarro, T.; Germini, F.; Linkins, L.A.; Haynes, R.B.; Chu, L.; et al. Boosting efficiency in a clinical literature surveillance system with LightGBM. PLOS Digital Health 2024, 3, e0000299. [Google Scholar] [CrossRef] [PubMed]
- Xie, H.; Fan, X.; Zhang, Y.; Zhan, Y.; Xu, W.; Huang, L. Predicting the risk of stroke based on imbalanced data set with missing data. In Proceedings of the 2022 IEEE 2nd International Conference on Electronic Technology, Communication and Information (ICETCI); IEEE, 2022; pp. 129–133. [Google Scholar]
- Mondal, S.; Ghosh, S.; Nag, A. Brain stroke prediction model based on boosting and stacking ensemble approach. International Journal of Information Technology 2024, 16, 437–446. [Google Scholar] [CrossRef]
- Ashrafuzzaman, M.; Saha, S.; Nur, K. Prediction of stroke disease using deep CNN based approach. Journal of Advances in Information Technology 2022, 13. [Google Scholar] [CrossRef]
- Geethanjali, T.; Divyashree, M.; Monisha, S.; Sahana, M. Stroke prediction using machine learning. Journal of Emerging Technologies and Innovative Research 2021, 9, 710–717. [Google Scholar]
- Nalini, D. Motyka Similar Feature Selected Softsign Deep Neural Classification For Stroke Disease Prediction. Webology (ISSN: 1735-188X) 2021, 18. [Google Scholar]
- Ahammad, T. Risk factor identification for stroke prognosis using machine-learning algorithms. Jordanian Journal of Computers and Information Technology 2022, 8. [Google Scholar] [CrossRef]
- Bathla, P.; Kumar, R. A hybrid system to predict brain stroke using a combined feature selection and classifier. Intelligent Medicine 2024, 4, 75–82. [Google Scholar] [CrossRef]
- Dubey, Y.; Tarte, Y.; Talatule, N.; Damahe, K.; Palsodkar, P.; Fulzele, P. Explainable and Interpretable Model for the Early Detection of Brain Stroke Using Optimized Boosting Algorithms. Diagnostics 2024, 14, 2514. [Google Scholar] [CrossRef]
- Akter, B.; Rajbongshi, A.; Sazzad, S.; Shakil, R.; Biswas, J.; Sara, U. A machine learning approach to detect the brain stroke disease. In Proceedings of the 2022 4th International Conference on Smart Systems and Inventive Technology (ICSSIT); IEEE, 2022; pp. 897–901. [Google Scholar]
- Devaki, A.; Rao, C.G. An ensemble framework for improving brain stroke prediction performance. In Proceedings of the 2022 First International Conference on Electrical, Electronics, Information and Communication Technologies (ICEEICT); IEEE, 2022; pp. 1–7. [Google Scholar]
- Premisha, P.; Prasanth, S.; Kanagarathnam, M.; Banujan, K. An ensemble machine learning approach for stroke prediction. In Proceedings of the 2022 International Research Conference on Smart Computing and Systems Engineering (SCSE); IEEE, 2022; Volume 5, pp. 165–170. [Google Scholar]
- Emon, M.U.; Keya, M.S.; Meghla, T.I.; Rahman, M.M.; Al Mamun, M.S.; Kaiser, M.S. Performance analysis of machine learning approaches in stroke prediction. In Proceedings of the 2020 4th international conference on electronics, communication and aerospace technology (ICECA); IEEE, 2020; pp. 1464–1469. [Google Scholar]
- Sharma, C.; Sharma, S.; Kumar, M.; Sodhi, A. Early stroke prediction using machine learning. In Proceedings of the 2022 International Conference on Decision Aid Sciences and Applications (DASA); IEEE, 2022; pp. 890–894. [Google Scholar]
- Islam, F.; Ghosh, M. An enhanced stroke prediction scheme using SMOTE and machine learning techniques. In Proceedings of the International conference on computing communication and networking technologies (ICCCNT), Kharagpur, India, 2021. [Google Scholar]
- Hossain, S.; Biswas, P.; Ahmed, P.; Sourov, M.R.; Keya, M.; Khushbu, S.A. Prognostic the risk of stroke using integrated supervised machine learning teachniques. In Proceedings of the 2021 12th International Conference on Computing Communication and Networking Technologies (ICCCNT); IEEE, 2021; pp. 1–5. [Google Scholar]
- Gupta, S.; Raheja, S. Stroke prediction using machine learning methods. In Proceedings of the 2022 12th International Conference on Cloud Computing, Data Science & Engineering (Confluence); IEEE, 2022; pp. 553–558. [Google Scholar]
















| Feature Name | Description | Type |
|---|---|---|
| age | Patient’s age in years | Numeric |
| Gender | Gender of the patient (Male, Female) | Categorical |
| hypertension | Whether the patient has hypertension (0 or 1) | Binary |
| heart_disease | Presence of heart disease (0 or 1) | Binary |
| ever_married | Marital status (Yes, No) | Categorical |
| work_type | Type of work (Private, Self-employed, etc.) | Categorical |
| Residence_type | Area of residence (Urban, Rural) | Categorical |
| avg_glucose_level | Average glucose level in blood | Numeric |
| bmi | Body Mass Index (BMI) | Numeric |
| smoking_status | Smoking status (Never smoked, Smokes, etc.) | Categorical |
| stroke | Stroke occurrence (Target: 0 or 1) | Binary |
| Dataset | Class | Samples | Percentage (%) |
|---|---|---|---|
| DF-1 [29] | Non-Stroke (0) | 42,617 | 98.2% |
| Stroke (1) | 783 | 1.8% | |
| DF-2 [30] | Non-Stroke (0) | 4,733 | 95.0% |
| Stroke (1) | 248 | 5.0% | |
| DF-3 [31] | Non-Stroke (0) | 4,861 | 95.1% |
| Stroke (1) | 249 | 4.9% |
| Feature | Description |
|---|---|
| gender_Female | Indicates if the gender is Female (binary: 0 or 1). |
| gender_Male | Indicates if the gender is Male (binary: 0 or 1). |
| gender_Other | Indicates if the gender is Other (binary: 0 or 1). |
| ever_married_No | Indicates if the individual has never married (binary: 0 or 1). |
| ever_married_Yes | Indicates if the individual has been married (binary: 0 or 1). |
| Residence_type_Rural | Indicates if the residence type is Rural (binary: 0 or 1). |
| Residence_type_Urban | Indicates if the residence type is Urban (binary: 0 or 1). |
| work_type_Govt_job | Indicates if the work type is Government job (binary: 0 or 1). |
| work_type_Never_worked | Indicates if the individual has never worked (binary: 0 or 1). |
| work_type_Private | Indicates if the work type is Private sector (binary: 0 or 1). |
| work_type_Self-employed | Indicates if the work type is Self-employed (binary: 0 or 1). |
| work_type_children | Indicates if the work type is related to children (binary: 0 or 1). |
| smoking_status_Unknown | Indicates if the smoking status is unknown (binary: 0 or 1). |
| smoking_status_formerly smoked | Indicates if the individual formerly smoked (binary: 0 or 1). |
| smoking_status_never smoked | Indicates if the individual never smoked (binary: 0 or 1). |
| smoking_status_smokes | Indicates if the individual currently smokes (binary: 0 or 1). |
| Metric | Description | Equation |
|---|---|---|
| Accuracy | Proportion of correct predictions among all cases. | |
| Precision | Proportion of true positives among all positive predictions. | |
| Recall (Sensitivity) | Proportion of true positives among all actual positives. | |
| F1 Score | Harmonic mean of Precision and Recall. | |
| ROC-AUC | Area under the ROC curve. | |
| Cohen Kappa | Agreement between predicted and actual labels, adjusted for chance. |
| SMOTE | SMOTEENN | SMOTE_SMOTEENN | ||||
|---|---|---|---|---|---|---|
| Mean | Std. | Mean | Std. | Mean | Std. | |
| Accuracy | 0.983739 | 0.001473 | 0.989341 | 0.000959 | 0.992189 | 0.001142 |
| Precision | 0.984304 | 0.003651 | 0.989271 | 0.002561 | 0.993501 | 0.001254 |
| Recall | 0.983176 | 0.003628 | 0.990264 | 0.002680 | 0.991661 | 0.002118 |
| F1-Score | 0.983730 | 0.001471 | 0.989762 | 0.000921 | 0.992579 | 0.001090 |
| ROC AUC | 0.998856 | 0.000226 | 0.999460 | 0.000097 | 0.999728 | 0.000076 |
| Cohen Kappa | 0.967478 | 0.002945 | 0.978645 | 0.001921 | 0.984334 | 0.002290 |
| SMOTE | SMOTEENN | SMOTE_SMOTEENN | ||||
|---|---|---|---|---|---|---|
| Mean | Std. | Mean | Std. | Mean | Std. | |
| Accuracy | 0.956792 | 0.005906 | 0.974961 | 0.0076 | 0.980297 | 0.00248 |
| Precision | 0.953499 | 0.008228 | 0.972186 | 0.011668 | 0.976108 | 0.001943 |
| Recall | 0.960494 | 0.007696 | 0.981709 | 0.00409 | 0.987795 | 0.003084 |
| F1-Score | 0.956956 | 0.005844 | 0.976893 | 0.006891 | 0.981916 | 0.002295 |
| ROC AUC | 0.993853 | 0.001366 | 0.997314 | 0.001116 | 0.998717 | 0.000552 |
| Cohen Kappa | 0.913585 | 0.011811 | 0.94957 | 0.015352 | 0.960277 | 0.004993 |
| SMOTE | SMOTEENN | SMOTE_SMOTEENN | ||||
|---|---|---|---|---|---|---|
| Mean | Std. | Mean | Std. | Mean | Std. | |
| Accuracy | 0.957828 | 0.005484 | 0.974077 | 0.006032 | 0.9934901 | 0.005255 |
| Precision | 0.955037 | 0.006166 | 0.970695 | 0.008552 | 0.97696 | 0.007958 |
| Recall | 0.960912 | 0.006294 | 0.981486 | 0.006761 | 0.989877 | 0.002184 |
| F1-Score | 0.957957 | 0.005471 | 0.976034 | 0.005543 | 0.983365 | 0.004762 |
| ROC AUC | 0.994063 | 0.001678 | 0.997051 | 0.00106 | 0.998943 | 0.000543 |
| Cohen Kappa | 0.915656 | 0.010968 | 0.947807 | 0.012157 | 0.963522 | 0.010617 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).