Submitted:
16 July 2026
Posted:
17 July 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
- Development of a reproducible cleanup pipeline that employs format normalization, kNN-based imputation, and robust outlier filtering via winsorization to successfully transform fragmented, multi-stage process records into stable model inputs.
- Implementation of a systematic, multi-stage feature engineering pipeline by utilizing variance thresholding, correlation filtering, and RFFI. While this approach systematically prunes highly correlated and redundant parameters to reduce database dimensionality, it is deliberately designed to accommodate necessary operator intervention.
- Strategic application of the Synthetic Minority Over-sampling Technique (SMOTE) strictly within the training folds of a defect-stratified 10-fold CV. This ensures that critical but infrequent defect types are reliably learned and predicted, rather than being overshadowed by defect-free products.
- Systematic performance comparison of five classifiers (RF, XGBoost, LightGBM, AdaBoost, and kNN) on a real-world foundry dataset. To align with industrial quality standards, hyperparameter tuning and model evaluations were strictly governed by the macro-averaged F3-scores. This is an asymmetric metric intentionally selected to mathematically penalize critical missed defects over FAs.
- Application of a post-processing threshold calibration to enforce strict industrial zero-defect constraints. Utilizing confusion matrices (CM) to map continuous prediction probabilities against asymmetric misclassification costs, the decision support system systematically shifts default decision boundaries to safely act as a pre-filter for manual inspection.
- Implementation of post-calibration decision fusion strategies (e.g., hard voting, rule-based consensus) to synthesize predictions across multiple base models into a final operator assistance system. This stage evaluates whether combining independent classifiers, which have been individually risk-calibrated to be hyper-sensitive to defects, can systematically increase the manual inspection time reduction beyond the capacity of any single algorithm.
2. Related Work
3. Materials and Methods
3.1. Foundry Database Acquisition & Pre-Processing
- Completeness & Distribution: Red cells indicate missing values, while blue cells highlight distributional skew and implausible spikes.
- Consistency: Pink cells denote mixed units within single columns and also potential writing errors.
- Relational Integrity: Orange cells reveal duplicate rows, grey cells mark faulty/non-monotonic timestamps, and light-brown cells indicate misaligned joins in the main database between different product samples.
3.2. Data Cleanup & Feature Engineering
3.2.1. kNN-Based Imputation
3.2.2. Variance Treshold Filtering (VTF)
3.2.3. Winsorization
3.2.4. Pearson Correlation Filtering (PCF)
3.2.5. RFFI-Based Feature Deselection
3.3. Model Development & Tuning
3.3.1. Handling Class Imbalance with SMOTE
3.3.2. Ensemble Classifiers
- RF [32]: An ensemble of bagged decision trees trained with the Gini impurity criterion. It handles mixed data types without scaling and provides feature importance estimates. The number of trees (n_estimators), tree depth (max_depth), minimum samples per leaf (min_samples_leaf), and the number of features considered at each split (max_features) are tuned.
- kNN [33]: A non-parametric, instance-based classifier that assigns a class by majority vote among the kNN closest training samples. Distances are computed with the Manhattan metric (L1-norm) after standardizing the features to ensure equal contribution. The voting scheme dictates how neighbors influence the prediction, i.e., either via uniform weighting, where all kNN neighbors contribute equally to the majority vote, or distance-based weighting, where closer neighbors exert a stronger influence on the final class assignment than more distant ones.
- XGBoost [34]: A regularized gradient boosting method that builds an additive model of shallow decision trees. It uses a second-order Taylor expansion of the loss function and includes L1/L2 regularization on leaf weights. Key tuned parameters are the learning rate (η), tree depth (max_depth), number of trees (n_estimators), subsampling ratios, and regularization strengths ().
- LightGBM [35]: A gradient boosting implementation that grows trees leaf-wise, focusing on the leaf with the highest loss reduction. It uses histogram-based algorithms for efficiency. Tuned hyperparameters include the number of leaves (num_leaves), tree depth (max_depth), learning rate (η), number of trees, minimum child samples, and L1/L2 regularization.
- AdaBoost [36]: AdaBoost was implemented in its Stagewise Additive Modeling using a Multi-class Exponential loss function (SAMME) variant to support the multiclass defect classification setting. The model sequentially fits weak learners, increasing the weight of previously misclassified samples so that later learners focus on harder observations. In this study, the number of estimators (n_estimators) and the learning rate (η) were tuned on the training folds, while the base learner consisted of decision stumps.
3.3.3. k-Fold CV
3.4. Classification Performance Evaluation
3.4.1. Macro-Averaged -Score
3.4.2. CM and Sensitivity-Driven Thresholding
3.4.3. Decision Fusion
- Majority Voting:
- 2.
- Penalized Weighted Voting:
- 3.
- Naïve Bayes Stacking:
3.5. Implementation Details
4. Results & Discussion
4.1. Data Pre-Processing Pipeline Performance
4.1.1. Initial Dataset Characterization
4.1.2. Missing Data Resolution
4.1.3. Outlier Control via Winsorization and VTF
4.1.4. Feature Selection Results
4.2. Defect Classification Performance
4.2.1. Macro-Averaged Classification Metrics
4.2.2. Cross-Validation vs. Test Set Robustness
4.2.3. CM and Risk-Sensitive Evaluation

5. Conclusions
Author Contributions
Data Availability Statement
Conflicts of Interest
Abbreviations
| AdaBoost | Adaptive Boosting |
| AP | Average Precision |
| CART | Classification and Regression Tree |
| CM | Confusion Matrix |
| CV | Cross-Validation |
| DOAJ | Directory of open access journals |
| FA | False Alarm |
| GBM | Gradient Boosting Machine |
| IQR | Interquartile Range |
| kNN | k-Nearest Neighbors |
| LightGBM | Light Gradient Boosting Machine |
| MDPI | Multidisciplinary Digital Publishing Institute |
| ML | Machine Learning |
| NaN | Not a Number |
| NOK / OK | Not ok / ok |
| OOD | Out-of-Distribution |
| PR | Precision-Recall |
| RF | Random Forest |
| RFFI | RF Feature Importance |
| SAMME | Stagewise Additive Modeling using Multiclass Exponential loss fct. |
| SHAP | SHapley Additive exPlanations |
| SME | Small and Medium-sized Enterprise |
| SMOTE | Synthetic Minority Over-sampling Technique |
| TP/FP/FN/TN | True Positive / False Positive / False Negative / True Negative |
| VTF | Variance Threshold Filtering |
| XGBoost | eXtreme Gradient Boosting |
References
- statista GmbH, “Statistik-Report zur Gießerei-Industrie in Deutschland,” statista - Industrien & Märkte, 2025.
- J. Ma, X. Tang, Y. Hou, H. Li, J. Lin and M. W. Fu, “Defects in metal-forming: formation mechanism, prediction and avoidance,” International Journal of Machine Tools and Manufacture, 2025. [CrossRef]
- C. Chelladurai, N. S. Mohan, D. Hariharashayee, S. Manikandan and P. Sivaperumal, “Analyzing the casting defects in small scale casting industry,” in Materials Today: Proceedings, 2021. [CrossRef]
- N. Qosim, A. M. Mufarrih, A. Sai’in, A. H. Firdaus, F. A. Andika, R. Monasari and Z. F. Emzain, “Effect of moisture content of green sand on the casting defects,” Journal of Applied Engineering and Technological Science (JAETS), vol. 2, pp. 1-6, 2020. [CrossRef]
- M. Maheswari and N. Brintha, “A Survey on Detection of Various Casting Defects,” in IEEE - Proceedings of the 2nd International Conference on Intelligent Data Communication Technologies and Internet of Things (IDCIoT), 2024.
- C. W. Chou and Y. T. Hsu, “Robust real-time object detection and counting system for casting foundries,” Applied Soft Computing, no. 176, 2025.
- S. Luo, Z. Z. Liu, R. Chen, Y. L. He, P. C. Zhang and X. Z. Hao, “Analysis of the application of artificial intelligence in the foundry industry,” in 17TH ASIAN FOUNDRY CONGRESS - Part 5: Smart Factory, 2025.
- T. JIANG, T. ALPCAN and K. OTTO, “A data screening framework before engaging machine learning in manufacturing,” Journal of Intelligent Manufacturing, pp. 1-25, 2025.
- D. Pietsch, M. Matthes, U. Wieland, S. Ihlenfeldt and T. Munkelt, “Root cause analysis in industrial manufacturing: A scoping review of current research, challenges and the promises of AI-driven approaches,” Journal of Manufacturing and Materials Processing, vol. 6, no. 8, 2024.
- P. Xu, X. Ji, M. Li and e. al, “Small data machine learning in materials science,” Journal of Computation Materials, vol. 9, no. 1, 2023.
- T. C. Uyan, K. Otto, M. S. Silva, P. Vilaça and E. Armakan, “Industry 4.0 foundry data management and supervised machine learning in low-pressure die casting quality improvement,” International journal of metalcasting, vol. 17, pp. 414-429, 2023.
- C. Bhagyanathan, J. Sarvesh, G. Vishal and R. Rithvik, “Analysis of Automation in Foundries for Transforming Production Efficiency, Quality Control, and Operational Resilience,” in International Conference on Machine Learning and Autonomous Systems (ICMLAS), 2025.
- M. Stubbemann, T. Hille and T. Hanika, “Selecting Features by their Resilience to the Curse of Dimensionality,” arXiv preprint, vol. arXiv:2304.02455, 2023.
- J. Yang, B. Liu and H. Huang, “Research on composition-process-property prediction of die casting Al alloys via combining feature creation and attention mechanisms,” Journal of Materials Research and Technology, vol. 28, pp. 335-346, 2024.
- M. L. Y. Song, A. Sharma and C. S. Chin, “Wafer Region Yield Prediction: Employing Majority Under-Sampling and Output Binarization on Low-Yield Threshold,” IEEE Access, vol. 13, 2025.
- N. Zong, T. Jing and J. C. Gebelin, “Machine learning techniques for the comprehensive analysis of the continuous casting processes: Slab defects,” Ironmaking & Steelmaking, pp. 1-20.
- U. Patwari, S. A. Bhuiyan, K. Noman and W. Ul Navid, “Defects and remedies in casting processes: a combinatorial approach between manual and digital optimization technique for enhanced quality casting,” Discover Mechanical Engineering, vol. 1, no. 3, 2024.
- Y. Han, Z. Wei and G. Huang, “An imbalance data quality monitoring based on SMOTE-XGBOOST supported by edge computing,” Science Report - Nature, no. 14, 2024.
- J. Qian, L. Deng, X. Zhang, S. Pi, Z. Song and X. Zhang, “XAIP: An eXplainable AI-Based Pipeline for Identifying Key Factors of Surface Defects in Strip Steel,” steel research international, vol. 3, no. 96.
- S. Dettori, A. Zaccara, L. Laid, I. Matino, M. Vannucci, V. Colla, G. Bontempi and L. Forlani, “Machine Learning models to forecast defects occurrence on foundry products,” IFAC PapersOnLine, Vols. 58-22, p. 113–118, 2024.
- Y. Z. Çiçek and F. Gürbüz, “Quality improvement in sand casting process: a hybrid approach based on machine learning and metaheuristic optimization methods”.Journal of Intelligent Manufacturing.
- Z. Breznikar, M. Bojinovic and M. Brezocnik, “APPLICATION OF MACHINE LEARNING TO REDUCE CASTING DEFECTS FROM BETONITE SAND MIXTURE,” International journal of simulation modelling, vol. 23, pp. 634-643, 2024.
- R. Shetty, A. Al Majali and L. Wells, “ENHANCED CLASSIFICATION OF REFRACTORY COATINGS IN FOUNDRIES: A VPCA-BASED MACHINE LEARNING APPROACH,” International Journal of Metalcasting, vol. 4, no. 19, 2025.
- D. Blondheim, “Improving manufacturing applications of machine learning by understanding defect classification and the critical error threshold,” International Journal of Metalcasting, vol. 16, no. 2, pp. 502-520, 2022.
- M. R. Mehregan, A. Rezasoltani and A. M. Khani, “A novel hybrid machine learning model for defect prediction in industrial manufacturing processes,” Contrib. Sci. & Tech Eng., vol. 2, 2025.
- Y. Khan, S. F. Shah and S. M. Asim, “A novel ranked k-nearest neighbors algorithm for missing data imputation,” JOURNAL OF APPLIED STATISTICS, vol. 52, p. 1103–1127, 2025.
- P. Duboue, The art of feature engineering: essentials for machine learning, Cambridge University Press, 2020.
- K. Cheng and D. S. Young, “An approach for specifying trimming and winsorization cutoffs,” Journal of Agricultural, Biological and Environmental Statistics, vol. 28, pp. 299-323, 2023.
- H. Gong, Y. Li, J. Zhang, B. Zhang and X. Wang, “A new filter feature selection algorithm for classification task by ensembling pearson correlation coefficient and mutual information,” Engineering Applications of Artificial Intelligence, no. 131, 2024.
- R. Iranzad and X. Liu, “A review of random forest-based feature selection methods for data science education and applications,” International Journal of Data Science and Analytics, vol. 2, no. 20, pp. 197-211, 2025.
- J. Hemmatian, R. Hajizadeh and F. Nazari, “Addressing imbalanced data classification with Cluster-Based reduced noise SMOTE,” Plos one, vol. 2, no. 20, 2025.
- H. A. Salman, A. Kalakech and A. Steiti, “Random forest algorithm overview,” Babylonian Journal of Machine Learning, pp. 69-79, 2024.
- R. K. Halder, M. N. Uddin, M. A. Uddin, S. Aryal and A. Khraisat, “Enhancing K-nearest neighbor algorithm: a comprehensive review and performance analysis of modifications,” Journal of Big Data, vol. 11, 2024.
- H. W. Aung, S. Wongsa and I. Phung-on, “Welding Defect Classification: A Low-Cost Approach Using XGBoost and Simulated Defect Data,” in IEEE International Conference on Information and Communication Technology (ICoICT), 2025.
- H. ZHENG, X. GAO and X. YANG, “Research on complex product quality prediction method based on FA-lightGBM,” in 14th International Conference on Quality, Reliability, Risk, Maintenance, and Safety Engineering (QR2MSE), 2024.
- M. Iqbal and A. K. Madan, “Bearing fault diagnosis in CNC machine using hybrid signal decomposition and gentle AdaBoost learning,” Journal of Vibration Engineering & Technologies, vol. 12, no. 2, pp. 1621-1634, 2024.
- J. M. Gorriz, R. M. Clemente, F. Segovia, J. Ramirez, A. Ortiz and J. Suckling, “Is K-fold cross validation the best model selection method for Machine Learning?,” arXiv.org - Statistics - Machine Learning, 2024.
- M. Grandini, E. Bagli and G. Visani, “Metrics for multi-class classification: an overview,” arXiv, 2020.
- S. Wojciechowski and M. Woźniak, “Evaluation of Multi-and Single-objective Learning Algorithms for Imbalanced Data,” arXiv, 2025.
- Y. Sha, S. Gou, B. Liu, J. Faber, N. Liu, S. Schramm and e. al, “Hierarchical knowledge guided fault intensity diagnosis of complex industrial systems,” in 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2024.
- K. J. Choi and J. T. Seong, “A decision-level fusion hybrid deep learning framework for high-precision automated optical inspection of VCSEL semiconductor devices,” AIMS Mathematics, vol. 11, pp. 14487-14521, 2026.
- S. Alsufyani, M. Forshaw and S. J. Fernstad, “Visualization of missing data: a state-of-the-art survey,” arXiv-PrePrint - Human-Computer Interaction, 2024.
- P. Costa, M. R. Seabra, J. M. C. de Sá and A. D. Santos, “Manufacturing process encoding through natural language processing for prediction of material properties,” Computational Materials Science, p. 112896, 2024.
- J. Obregon and J.-Y. Jung, “Rule-based visualization of faulty process conditions in the die-casting,” Journal of Intelligent Manufacturing, vol. 35, p. 521–537, 2024.
- W. Zasada, P. Guzik, K. B. Kubiak and B. Więckowska, “The Toolbox for Rating Diagnostic Tests: A Guide to Classification Metrics,” Journal of Medical Science, vol. 94, 2025.
- Gomes, 46 D. L. P.; Grégio, A.; Alves, M. A. Z.; P. R. de Almeida, L. Book Title, 2nd ed. Author 1, A., Author 2, B., Editor 1, A., Editor 2, B., Eds.; Publisher Location: Publisher.











| Model | Hyperparameter | Search Space | Chosen Value |
|---|---|---|---|
| RF |
n_estimators max_depth min_samples_split min_samples_leaf max_features class_weight |
200, 300, 500, 1000, 2000 10, 15, 20, 30, 40, 50, None 2, 4, 5, 6, 10 1, 2, 5 ‘sqrt’, ‘log2’, None ‘balanced’, ‘balanced_subsample’, None |
2000 20 4 2 ‘sqrt’ ‘balanced’ |
| kNN | n_neighbors weights metric |
3, 5, 7, 9, 10, 15 ‘uniform’, ‘distance’ ‘euclidean’, ‘manhattan’, ‘chebyshev’ |
15 ‘distance’ ‘manhattan’ |
| XGBoost |
learning_rate max_depth n_estimators subsample colsample_bytree min_child_weight gamma |
0.01, 0.05, 0.1, 0.2 3, 5, 7, 9, 10, 15 100, 200, 300, 500 0.6, 0.8, 1.0 0.6, 0.8, 1.0 1, 3, 5 0, 1, 5 |
0.01 7 500 0.8 1.0 1 5 |
| LightGBM | num_leaves max_depth learning_rate n_estimators min_child_samples class_weight |
15, 31, 63, 127, 255 5, 8, 12, 15, 20, 30, -1 0.01, 0.05, 0.1, 0.2 100, 200, 300, 500, 1000 10, 20, 40, 60 ‘balanced’, None |
63 30 0.01 1000 40 ‘balanced’ |
| AdaBoost | n_estimators learning_rate |
50, 100, 200, 300, 400 0.05, 0.1, 0.5, 0.7, 0.8, 1.0 |
200 0.5 |
| Model | F3-score | Precision | Recall |
|---|---|---|---|
| RF | 0.361 | 0.378 | 0.360 |
| kNN | 0.364 | 0.300 | 0.376 |
| XGBoost | 0.344 | 0.382 | 0.343 |
| LightGBM | 0.397 | 0.375 | 0.400 |
| AdaBoost | 0.309 | 0.253 | 0.332 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.