Results and Discussion
In this section, we embark on an in-depth exploration of the performance exhibited by various machine learning (ML) models concerning their hyperparameters, notably the minimum child weight and learning rate. These parameters play a pivotal role in addressing overfitting and underfitting challenges inherent in ML models. Through a meticulous comparative analysis across multiple ML algorithms, we aim to elucidate their behavior under diverse hyperparameter settings.
Commencing our analysis with XGBoost, as depicted in
Figure 7, we observed a noteworthy reduction in overfitting tendencies upon monitoring the training and validation plots while modulating the minimum child weight and learning rate. Through strategic convergence of these plots, we achieved a commendable accuracy, precision, recall, and F-1 score metrics, culminating in an impressive performance level of 94.84%. Particularly noteworthy is the discernible improvement observed in the confusion matrix plot on unseen datasets, where we attained a prediction accuracy of 96% for patients diagnosed with lung cancer and 94% for non-cancer cases, as illustrated in
Figure 15.
These findings underscore the effectiveness of parameter optimization in enhancing the predictive capabilities of ML models, thereby facilitating more accurate and reliable classification outcomes. Moreover, they highlight the potential of XGBoost as a robust tool for lung cancer diagnosis, capable of delivering clinically relevant insights with high precision and reliability.
Figure 7.
Training and Validation plots under consideration of Min child Weight and Learning Rate in XGBoost.
Figure 7.
Training and Validation plots under consideration of Min child Weight and Learning Rate in XGBoost.
Figure 8.
Training and Validation plots under consideration of Min child Weight and Learning Rate in LGBM.
Figure 8.
Training and Validation plots under consideration of Min child Weight and Learning Rate in LGBM.
AdaBoost emerges as a potent solution for mitigating overfitting concerns, as evidenced by our analysis depicted in
Figure 9. Through meticulous adjustment of hyperparameters, we effectively curtailed overfitting tendencies, thereby achieving commendable accuracy, precision, recall, and F-1 score metrics, culminating in an impressive overall performance rate of 95.87%. Notably, the ensuing evaluation via the confusion matrix plot revealed an outstanding prediction accuracy of 96% for both lung cancer and non-cancer cases, as delineated in
Figure 15.
Similarly, our exploration of Logistic Regression, depicted in
Figure 10, yielded results akin to the AdaBoost model. Although exhibiting slightly lower performance metrics, Logistic Regression still showcased notable accuracy, precision, recall, and F-1 score rates, reaching an overall performance level of 89.69%, as illustrated in
Figure 15.
AdaBoost and Logistic Regression emerge as strong contenders for lung cancer classification. Through meticulous hyperparameter tuning and effective overfitting mitigation strategies, these models demonstrate promising capabilities for accurate and reliable diagnosis. Furthermore, the consistency observed in their performance, as evidenced by the informative confusion matrix plots, strengthens their case as valuable tools for clinical decision-making.
Figure 9.
Training and Validation plots under consideration of Min child Weight and Learning Rate in AdaBoost.
Figure 9.
Training and Validation plots under consideration of Min child Weight and Learning Rate in AdaBoost.
Figure 10.
Training and Validation plots under consideration of Min child Weight and Learning Rate in Logistic Regression.
Figure 10.
Training and Validation plots under consideration of Min child Weight and Learning Rate in Logistic Regression.
Further analysis encompassed the evaluation of additional machine learning models, including Decision Tree (
Figure 11), Random Forest (
Figure 12), CatBoost (
Figure 13), and k-NN (
Figure 14), each offering unique insights into their performance characteristics with regard to overfitting.
The Decision Tree model, as depicted in
Figure 11, exhibited a stable performance, achieving an accuracy rate of approximately 92%. Notably, the absence of discernible overfitting signs between the training and validation sets, as evidenced by the smooth progression per epoch across the range of minimum child weight and learning rate, underscores its reliability in classification tasks.
Similarly, the Random Forest model, illustrated in
Figure 12, showcased remarkable performance with an accuracy rate of approximately 97%. Importantly, the absence of any discernible overfitting between the training and validation sets across the range of minimum child weight and learning rate reaffirms its robustness and efficacy in achieving highly accurate classifications.
Figure 11.
Training and Validation plots under consideration of Min child Weight and Learning Rate in Decision Tree.
Figure 11.
Training and Validation plots under consideration of Min child Weight and Learning Rate in Decision Tree.
Figure 12.
Training and Validation plots under consideration of Min child Weight and Learning Rate in Random forest.
Figure 12.
Training and Validation plots under consideration of Min child Weight and Learning Rate in Random forest.
In the case of the CatBoost model, depicted in
Figure 13, an accuracy rate of approximately 96% was attained, albeit with a slight gap observed between the training and validation sets per epoch, indicating minor overfitting tendencies. Despite this, the model demonstrates impressive performance and offers valuable insights into its adaptability to different hyperparameter settings.
Lastly, the k-NN model, as shown in
Figure 14, achieved an accuracy rate of 92% with a discernible but logical gap observed between the training and validation sets per epoch. This indicates a moderate level of overfitting, albeit within acceptable bounds.
Overall, the comprehensive analysis across these diverse machine learning models underscores their robustness and adaptability in handling variations in hyperparameters. The minimal fluctuations observed in overfitting across the range of learning rates and minimum child weights affirm the effectiveness of these models in achieving stable and reliable performance for lung cancer classification tasks.
Figure 13.
Training and Validation plots under consideration of Min child Weight and Learning Rate in CatBoost.
Figure 13.
Training and Validation plots under consideration of Min child Weight and Learning Rate in CatBoost.
Figure 14.
Training and Validation plots under consideration of Min child Weight and Learning Rate in k-NN.
Figure 14.
Training and Validation plots under consideration of Min child Weight and Learning Rate in k-NN.
In the test analysis of datasets, as depicted in the confusion matrices presented in
Figure 14 and
Figure 15, the performance of various machine learning models in predicting actual datasets is elucidated through the observed errors. Specifically, for XGBoost, 5 errors were noted, while LGBM and AdaBoost exhibited 3 errors each. Logistic Regression, on the other hand, registered 10 errors, followed by Decision Tree with 8 errors, and Random Forest with 3 errors. CatBoost and KNN models demonstrated 4 and 8 errors, respectively. Notably, the DNN model showcased the lowest error count, with only 3 errors observed in prediction.
Figure 15.
Confusion Matrix of XGBoost, LGBM, AdaBoost, Logistic Regression, Decision Tree and Random Forest.
Figure 15.
Confusion Matrix of XGBoost, LGBM, AdaBoost, Logistic Regression, Decision Tree and Random Forest.
Figure 16.
Confusion Matrix of CatBoost, k-NN, and DNN.
Figure 16.
Confusion Matrix of CatBoost, k-NN, and DNN.
Indeed, the Deep Neural Network (DNN) model exhibited remarkable performance, surpassing all other models with an outstanding accuracy, precision, recall, and F-1 score of 96.91%, as evidenced in
Figure 16. This exemplary performance underscores the efficacy of DNN, particularly in scenarios where the correlation between datasets and features is intricate and challenging to discern.
In summary, our analysis delineates DNN as the top-performing model for lung cancer classification tasks. Following closely behind are CatBoost and AdaBoost, which also demonstrated impressive performance metrics. These findings underscore the efficacy of these models in addressing overfitting and achieving high prediction accuracy in lung cancer classification tasks, as illustrated in
Figure 17. Such insights are invaluable for guiding the selection of appropriate models for clinical applications, ultimately facilitating more accurate and reliable diagnoses for improved patient outcomes.
These findings shed light on the performance of different machine learning models in the prediction of lung cancer classifications. The comparative analysis of error rates underscores the varying degrees of accuracy and reliability exhibited by each model. Such insights are invaluable for clinicians and researchers alike, aiding in the selection of appropriate models for real-world applications and informing decision-making processes in clinical settings.