Submitted:
03 August 2026
Posted:
03 August 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
- How do widely used classifiers compare when they follow the same employee grouping and time rules?
- Can one ensemble use standard fields, organizational hierarchy, and category-aware learning together?
- How well do the models work for employees who took no part in model development?
- Which measures describe risk ranking, the usefulness of the alert list, and the accuracy of probability scores?
- How should uncertain field dates, small job-group counts, and responsible-use requirements limit the conclusions?
2. Related Work and Research Framework
2.1. Turnover Theory and the Observable-Data Boundary
2.1.1. Machine-Learning Development and Employee-Turnover Studies
2.1.2. Imbalance, Ranking, and Probability Reliability
2.1.3. Limited-Data Validation and Model Explanation
2.1.4. People Analytics and Algorithmic Governance
2.2. Conceptual Framework

2.3. Research Gaps
2.4. Methodological Contribution
2.5. From Management Theory to Measurable Evidence
| Theoretical perspective | Core mechanism in the literature | Administrative evidence available here | Valid use in this study | Inference not supported |
|---|---|---|---|---|
| Withdrawal process | dissatisfaction may lead to thoughts, search, and departure | realized voluntary departure and employment dates | define a 90-day prediction target | reconstruct an employee’s unobserved psychological sequence |
| Job embeddedness | links, fit, and sacrifice inhibit leaving | job family, organizational location, and tenure proxies | motivate structured organizational representation | claim that a category code measures embeddedness |
| Job demands–resources | demands deplete resources; resources buffer strain | no validated workload or resource scale | motivate supportive follow-up questions | estimate burnout or workload effects |
| Perceived organizational support | reciprocity follows perceived care and recognition | no direct perception measure | constrain alerts to supportive human review | infer perceived support from a risk score |
| Psychological contract | unmet expectations can trigger withdrawal | no expectation or breach measure | motivate confidential dialogue about expectations | diagnose contract breach from administrative data |
| Unfolding and shock models | events can initiate different departure paths | no event-history or exit-interview sequence | explain why one static rule may miss heterogeneous paths | identify the shock that caused an individual departure |
3. Research Methodology
3.1. Study Design, Prediction Target, and Notation
3.1.1. What the Model Predicts
3.2. General-Purpose Comparison Models
3.2.1. No-information, Linear, and Probabilistic Models
3.2.2. Neighborhood, Margin, and Neural Models
3.2.3. Trees, Bagging, and Imbalance-Aware Ensembles
3.2.4. Boosting Models
3.3. Proposed HRE-XCB Three-Branch Ensemble
3.3.1. Design Rationale
- Raw-XGB learns from filled, standardized, and one-hot-encoded fields.
- HRE-XGB adds category risk summaries that exclude the employee’s own outcome and borrow strength from broader business groups.
- CatBoost keeps the original category labels and calculates category summaries without using the current row’s outcome.
3.3.2. Branch I: Raw-XGB
3.3.3. Business-Derived Categorical Structure
3.3.4. Hierarchical Risk Encoding
3.3.5. Cross-Fitted Category Summaries That Exclude the Employee
3.3.6. Branch II: HRE-XGB
3.3.7. Branch III: CatBoost
3.3.8. Nonnegative Probability Fusion
3.4. Two-Stage Model Search and Alert-Cutoff Selection
3.4.1. Candidate Space Defined before Later Evaluation
3.4.2. Choosing an Alert Cutoff That Protects Recall
3.5. Evaluation Metrics and Probability Interpretation
3.5.1. Measures for the Final Alert List
3.5.2. Ranking Measures Across Possible Cutoffs
3.5.3. Reliability of the Numeric Risk Score
3.6. Uncertainty, Model Explanation, and Job-Group Checks
3.6.1. Bootstrap Uncertainty
3.6.2. Which Variables the Model Uses
3.6.3. Job-Group and Sensitive-Feature Checks
3.7. Data, Employee Separation, and Experimental Protocol

3.8. Models Compared and Their Purpose
| Implemented classifier | Main learning assumption | Limited-data control | Role in the study evidence |
|---|---|---|---|
| DummyPrior | no feature information | none required | exposes majority-class accuracy |
| Logistic regression | regularized linear log-odds | L2 shrinkage and standardized numeric inputs | tuned linear reference |
| SGDLogLoss | linear log-odds learned incrementally | regularization and controlled learning schedule | optimization-sensitive linear check |
| Shrinkage LDA | shared class covariance and linear separation | covariance shrinkage | low-variance discriminant check |
| GaussianNB | conditionally independent Gaussian numeric variables | small parameter count | distributional probabilistic check |
| BernoulliNB | conditionally independent binary indicators | small parameter count | sparse-indicator check |
| KNN | nearby records share outcomes | scaling and bounded neighborhood size | tests local similarity and identity-recurrence risk |
| Linear SVM | maximum-margin linear boundary | margin penalty | linear boundary without calibrated probability output |
| RBF SVM | smooth nonlinear similarity | penalty and kernel-width tuning | flexible margin check |
| Decision tree | recursive axis-aligned partitions | depth and leaf-size limits | interpretable high-variance reference |
| Bagged trees | averaging reduces tree variance | resampling and ensemble averaging | isolates the value of bagging |
| Random forest | bagging plus feature randomization | minimum leaf size and feature subsampling | tuned employee-separated-sample ranking reference |
| Extra Trees | strongly randomized tree partitions | depth, leaf size, and averaging | alternative variance-reduction reference |
| Balanced Random Forest | balanced subsamples per tree | ensemble averaging | imbalance-aware recall check |
| EasyEnsemble | multiple balanced majority subsets | aggregation across subsets | under-sampling ensemble check |
| AdaBoost | sequential emphasis on errors | shallow weak learners and learning rate | reweighting-based boosting check |
| Gradient boosting | additive correction of loss gradients | shallow trees and shrinkage | classical boosting reference |
| Histogram gradient boosting | binned additive trees | leaf and L2 constraints | tuned probability-error reference |
| XGBoost | regularized second-order boosting | depth, sampling, and L1/L2 penalties | tuned general boosting reference |
| LightGBM | histogram-based leaf-wise growth | leaf-count and minimum-data limits | efficient boosting check |
| CatBoost | ordered categorical statistics and symmetric trees | ordered boosting and leaf regularization | category-native high-recall reference |
| MLP | nonlinear interactions in hidden layers | small network, weight decay, and early stopping | flexible neural reference |
3.9. How Hierarchical Risk Encoding Stabilizes Rare Categories
3.10. Model Search, Selection, and Rejection Rules
4. Results
4.1. Model Development Under Employee Separation

4.2. Results for Employees Kept Out of Model Development
| Model | Accuracy | Precision | Recall | F1 | PR-AUC | Brier |
|---|---|---|---|---|---|---|
| RF | .8094 | .2542 | .6250 | .3614 | .3772 | .1134 |
| HistGB | .8165 | .2923 | .7917 | .4270 | .3756 | .0752 |
| Extra Trees | .8489 | .3200 | .6667 | .4324 | .3622 | .1074 |
| XGBoost | .8022 | .2615 | .7083 | .3820 | .3611 | .1305 |
| HRE-XCB | .7842 | .2500 | .7500 | .3750 | .3477 | .1023 |
| LR | .7806 | .2394 | .7083 | .3579 | .3319 | .1306 |
| CatBoost | .6583 | .1743 | .7917 | .2857 | .3127 | .1684 |
| Prior-only | .9137 | .0000 | .0000 | .0000 | .0863 | .1318 |

4.3. HRE-XCB Component Ablation

4.4. Decision-Specific Model Roles

4.5. Paired Uncertainty

4.6. Feature Timing and Probability Calibration

4.7. Integrated Empirical Interpretation

4.8. Full Algorithm Screen and Review Workload
| Tuned baseline | Search trials | Best grouped-CV PR-AUC | Selected complexity control |
|---|---|---|---|
| Logistic regression | 40 | .5770 | L1 penalty, C=.1473 |
| Random forest | 40 | .6199 | depth 8, minimum leaf size 4, square-root feature sampling |
| Extra Trees | 40 | .6195 | unrestricted depth, minimum leaf size 5, square-root feature sampling |
| Histogram gradient boosting | 40 | .6296 | 17 leaf nodes, minimum leaf size 8, learning rate .0239, L2=.2206 |
| XGBoost | 40 | .6142 | depth 5, minimum child weight 3, row/column sampling, L1/L2 penalties |
| Model | True positives | False negatives | Total alerts | False-positive alerts | Alert rate |
|---|---|---|---|---|---|
| Random forest | 15 | 9 | 59 | 44 | 21.22% |
| HistGB | 19 | 5 | 65 | 46 | 23.38% |
| Extra Trees | 16 | 8 | 50 | 34 | 17.99% |
| XGBoost | 17 | 7 | 65 | 48 | 23.38% |
| HRE-XCB | 18 | 6 | 72 | 54 | 25.90% |
| Logistic regression | 17 | 7 | 71 | 54 | 25.54% |
| CatBoost | 19 | 5 | 109 | 90 | 39.21% |
5. Discussion
5.1. Interpretation of the Empirical Evidence
5.2. Credibility, Scope, and Management Implications
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Use of Artificial Intelligence
Abbreviations
| HRE-XCB | Hierarchical Risk Encoding XGBoost–CatBoost Ensemble |
| PR-AUC | Area under the precision–recall curve |
| ROC-AUC | Area under the receiver operating characteristic curve |
| ECE | Expected calibration error |
| HR | Human resources |
References
- Peter W. Hom; Thomas W. Lee; Jason D. Shaw; John P. Hausknecht. One hundred years of employee turnover theory and research. Journal of Applied Psychology, 102(3): 530-545, 2017. [CrossRef]
- Alex L. Rubenstein; Marion B. Eberly; Thomas W. Lee; Terence R. Mitchell. Surveying the forest: A meta-analysis, moderator investigation, and future-oriented discussion of the antecedents of voluntary employee turnover. Personnel Psychology, 71(1): 23-65, 2017. [CrossRef]
- Thomas W. Lee; Peter W. Hom; Marion B. Eberly; Terence R. Mitchell. On the Next Decade of Research in Voluntary Employee Turnover. Academy of Management Perspectives, 31(3): 201-221, 2017. [CrossRef]
- Julie I. Hancock; David G. Allen; Frank A. Bosco; Karen R. McDaniel; Charles A. Pierce. Meta-Analytic Review of Employee Turnover as a Predictor of Firm Performance. Journal of Management, 39(3): 573-603, 2011. [CrossRef]
- Tae-Youn Park; Jason D. Shaw. Turnover rates and organizational performance: A meta-analysis. Journal of Applied Psychology, 98(2): 268-309, 2013. [CrossRef]
- Brooks C. Holtom; Terence R. Mitchell; Thomas W. Lee; Marion B. Eberly. 5 Turnover and Retention Research: A Glance at the Past, a Closer Review of the Present, and a Venture into the Future. Academy of Management Annals, 2(1): 231-274, 2008. [CrossRef]
- T. R. Mitchell; B. C. Holtom; T. W. Lee; C. J. Sablynski; M. Erez. WHY PEOPLE STAY: USING JOB EMBEDDEDNESS TO PREDICT VOLUNTARY TURNOVER. Academy of Management Journal, 44(6): 1102-1121, 2001. [CrossRef]
- William H. Mobley. Intermediate linkages in the relationship between job satisfaction and employee turnover. Journal of Applied Psychology, 62(2): 237-240, 1977. [CrossRef]
- ROBERT P. TETT; JOHN P. MEYER. JOB SATISFACTION, ORGANIZATIONAL COMMITMENT, TURNOVER INTENTION, AND TURNOVER: PATH ANALYSES BASED ON META-ANALYTIC FINDINGS. Personnel Psychology, 46(2): 259-293, 1993. [CrossRef]
- Arnold B. Bakker; Evangelia Demerouti. Job demands–resources theory: Taking stock and looking forward. Journal of Occupational Health Psychology, 22(3): 273-285, 2017. [CrossRef]
- Robert Eisenberger; Robin Huntington; Steven Hutchison; Debora Sowa. Perceived organizational support. Journal of Applied Psychology, 71(3): 500-507, 1986. [CrossRef]
- Denise M. Rousseau. Psychological and implied contracts in organizations. Employee Responsibilities and Rights Journal, 2(2): 121-139, 1989. [CrossRef]
- Leo Breiman. Random Forests. Machine Learning, 45(1): 5-32, 2001. [CrossRef]
- Corinna Cortes; Vladimir Vapnik. Support-vector networks. Machine Learning, 20(3): 273-297, 1995. [CrossRef]
- Jerome H. Friedman. Greedy function approximation: A gradient boosting machine. The Annals of Statistics, 29(5), 2001. [CrossRef]
- Yoav Freund; Robert E Schapire. A Decision-Theoretic Generalization of On-Line Learning and an Application to Boosting. Journal of Computer and System Sciences, 55(1): 119-139, 1997. [CrossRef]
- T. Cover; P. Hart. Nearest neighbor pattern classification. IEEE Transactions on Information Theory, 13(1): 21-27, 1967. [CrossRef]
- David E. Rumelhart; Geoffrey E. Hinton; Ronald J. Williams. Learning representations by back-propagating errors. Nature, 323(6088): 533-536, 1986. [CrossRef]
- Tianqi Chen; Carlos Guestrin. XGBoost. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining: 785-794, 2016. [CrossRef]
- Candice Bentéjac; Anna Csörgo; Gonzalo Martínez-Muñoz. A comparative analysis of gradient boosting algorithms. Artificial Intelligence Review, 54(3): 1937-1967, 2020. [CrossRef]
- Marco Tulio Ribeiro; Sameer Singh; Carlos Guestrin. "Why Should I Trust You?". Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining: 1135-1144, 2016. [CrossRef]
- N. V. Chawla; K. W. Bowyer; L. O. Hall; W. P. Kegelmeyer. SMOTE: Synthetic Minority Over-sampling Technique. Journal of Artificial Intelligence Research, 16: 321-357, 2002. [CrossRef]
- Haibo He; E.A. Garcia. Learning from Imbalanced Data. IEEE Transactions on Knowledge and Data Engineering, 21(9): 1263-1284, 2009. [CrossRef]
- Alexandru Niculescu-Mizil; Rich Caruana. Predicting good probabilities with supervised learning. Proceedings of the 22nd international conference on Machine learning - ICML ’05: 625-632, 2005. [CrossRef]
- Takaya Saito; Marc Rehmsmeier. The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets. PLOS ONE, 10(3): e0118432, 2015. [CrossRef]
- Tom Fawcett. An introduction to ROC analysis. Pattern Recognition Letters, 27(8): 861-874, 2006. [CrossRef]
- Ben Van Calster; David J. McLernon; Maarten van Smeden; Laure Wynants; Ewout W. Steyerberg. Calibration: the Achilles heel of predictive analytics. BMC Medicine, 17(1): 230, 2019. [CrossRef]
- Gary S. Collins; Johannes B. Reitsma; Douglas G. Altman; Karel G.M. Moons. Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis (TRIPOD): The TRIPOD Statement. Annals of Internal Medicine, 162(1): 55-63, 2015. [CrossRef]
- Richard D Riley; Kym IE Snell; Joie Ensor; Danielle L Burke; Frank E Harrell Jr; Karel GM Moons; Gary S Collins. Minimum sample size for developing a multivariable prediction model: PART II - binary and time-to-event outcomes. Statistics in Medicine, 38(7): 1276-1296, 2018. [CrossRef]
- Andrius Vabalas; Emma Gowen; Ellen Poliakoff; Alexander J. Casson. Machine learning algorithm validation with a limited sample size. PLOS ONE, 14(11): e0224365, 2019. [CrossRef]
- Annette M. Molinaro; Richard Simon; Ruth M. Pfeiffer. Prediction error estimation: a comparison of resampling methods. Bioinformatics, 21(15): 3301-3307, 2005. [CrossRef]
- Sudhir Varma; Richard Simon. Bias in error estimation when using cross-validation for model selection. BMC Bioinformatics, 7(1): 91, 2006. [CrossRef]
- David R. Roberts; Volker Bahn; Simone Ciuti; Mark S. Boyce; Jane Elith; Gurutzeta Guillera-Arroita; Severin Hauenstein; José J. Lahoz-Monfort; Boris Schröder; Wilfried Thuiller; David I. Warton; Brendan A. Wintle; Florian Hartig; Carsten F. Dormann. Cross-validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure. Ecography, 40(8): 913-929, 2017. [CrossRef]
- Tal Yarkoni; Jacob Westfall. Choosing Prediction Over Explanation in Psychology: Lessons From Machine Learning. Perspectives on Psychological Science, 12(6): 1100-1122, 2017. [CrossRef]
- Scott M. Lundberg; Gabriel Erion; Hugh Chen; Alex DeGrave; Jordan M. Prutkin; Bala Nair; Ronit Katz; Jonathan Himmelfarb; Nisha Bansal; Su-In Lee. From local explanations to global understanding with explainable AI for trees. Nature Machine Intelligence, 2(1): 56-67, 2020. [CrossRef]
- Alejandro Barredo Arrieta; Natalia Díaz-Rodríguez; Javier Del Ser; Adrien Bennetot; Siham Tabik; Alberto Barbado; Salvador Garcia; Sergio Gil-Lopez; Daniel Molina; Richard Benjamins; Raja Chatila; Francisco Herrera. Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Information Fusion, 58: 82-115, 2020. [CrossRef]
- Janet H. Marler; John W. Boudreau. An evidence-based review of HR Analytics. The International Journal of Human Resource Management, 28(1): 3-26, 2016. [CrossRef]
- Dana B. Minbaeva. Building credible human capital analytics for organizational competitive advantage. Human Resource Management, 57(3): 701-713, 2017. [CrossRef]
- R.H. Hamilton; William A. Sodeman. The questions we ask: Opportunities and challenges for using big data analytics to strategically manage human capital resources. Business Horizons, 63(1): 85-95, 2020. [CrossRef]
- Aizhan Tursunbayeva; Stefano Di Lauro; Claudia Pagliari. People analytics—A scoping review of conceptual boundaries and value propositions. International Journal of Information Management, 43: 224-247, 2018. [CrossRef]
- Aizhan Tursunbayeva; Claudia Pagliari; Stefano Di Lauro; Gilda Antonelli. The ethics of people analytics: risks, opportunities and recommendations. Personnel Review, 51(3): 900-921, 2021. [CrossRef]
- Prasanna Tambe; Peter Cappelli; Valery Yakubovich. Artificial Intelligence in Human Resources Management: Challenges and a Path Forward. California Management Review, 61(4): 15-42, 2019. [CrossRef]
- Katherine C. Kellogg; Melissa A. Valentine; Angéle Christin. Algorithms at Work: The New Contested Terrain of Control. Academy of Management Annals, 14(1): 366-410, 2020. [CrossRef]
- Ninareh Mehrabi; Fred Morstatter; Nripsuta Saxena; Kristina Lerman; Aram Galstyan. A Survey on Bias and Fairness in Machine Learning. ACM Computing Surveys, 54(6): 1-35, 2021. [CrossRef]
- Manish Raghavan; Solon Barocas; Jon Kleinberg; Karen Levy. Mitigating bias in algorithmic hiring. Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency: 469-481, 2020. [CrossRef]
- Cynthia Rudin. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence, 1(5): 206-215, 2019. [CrossRef]
- Jeroen Meijerink; Mark Boons; Anne Keegan; Janet Marler. Algorithmic human resource management: Synthesizing developments and cross-disciplinary insights on digital HRM. The International Journal of Human Resource Management, 32(12): 2545-2562, 2021. [CrossRef]
- Xavier Parent-Rocheleau; Sharon K. Parker. Algorithms as work designers: How algorithmic management influences the design of jobs. Human Resource Management Review, 32(3): 100838, 2022. [CrossRef]
- Sunghoon Kim; Violetta Khoreva; Vlad Vaiman. Strategic Human Resource Management in the Era of Algorithmic Technologies: Key Insights and Future Research Agenda. Human Resource Management, 64(2): 447-464, 2024. [CrossRef]
- Francesca Fallucchi; Marco Coladangelo; Romeo Giuliano; Ernesto William De Luca. Predicting Employee Attrition Using Machine Learning Techniques. Computers, 9(4): 86, 2020. [CrossRef]
- Filippo Guerranti; Giovanna Maria Dimitri. A Comparison of Machine Learning Approaches for Predicting Employee Attrition. Applied Sciences, 13(1): 267, 2022. [CrossRef]
- Xinlei Wang; Jianing Zhi. A machine learning-based analytical framework for employee turnover prediction. Journal of Management Analytics, 8(3): 351-370, 2021. [CrossRef]
- Jungryeol Park; Yituo Feng; Seon-Phil Jeong. Developing an advanced prediction model for new employee turnover intention utilizing machine learning techniques. Scientific Reports, 14(1): 1221, 2024. [CrossRef]
- Aseel Qutub; Asmaa Al-Mehmadi; Munirah Al-Hssan; Ruyan Aljohani; Hanan S. Alghamdi. Prediction of Employee Attrition Using Machine Learning and Ensemble Methods. International Journal of Machine Learning and Computing, 11(2): 110-114, 2021. [CrossRef]
- Kang-Chul Kim; Hai-tong Wei. Development of a Face Detection and Recognition System Using a Raspberry Pi. The Journal of the Korea Institute of Electronic Communication Sciences, 12(5): 859-864, 2017. [CrossRef]
- Haitong Wei; Xinghai Wang. Financial Risk Management Early-Warning Model for Chinese Enterprises. Journal of Risk and Financial Management, 17(7): 255, 2024. [CrossRef]
- Haitong Wei. DecorPGNet: Functional Area Division and Layout Algorithm Model in Living Rooms of Chinese Apartment-Style Family Homes. Civil Engineering Research Journal, 15(1), 2024. [CrossRef]
- Haitong Wei. Exploring and Practicing the Quantification of Interior Design Colors from an IKEA Design Perspective. Journal of Sensor Networks and Data Communications, 4(2): 01-12, 2024. [CrossRef]
- Haitong Wei. Negative Indicators and Ordering Stability in Exploratory Factor Analysis: A Sign-Orientation Theory with Reproducible Simulation Evidence [Preprint]. Preprints.org, 2026. [CrossRef]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).