Submitted:
13 September 2025
Posted:
16 September 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
1.1. Related Work
1.2. Research Gap
- 1.
- Leakage-free, decision-relevant evaluations of LLM-based augmentation on highly imbalanced HR data, with metrics and summaries aligned to asymmetric costs and budget constraints;
- 2.
- Head-to-head evidence on how classical ML versus a compact transformer–MLP exploit synthetic minority samples under identical, leakage-controlled protocols;
- 3.
- Demonstrations that numeric gains translate into interpretable, actionable, and fair decisions (higher minority recall at tolerable false-alarm rates, SHAP levers mapping to HR actions, and bounded subgroup TPR gaps);
- 4.
- Systematic analysis of synthetic scale effects (progressive augmentation) to guide budgeted deployment.
1.3. Our Contributions
- Decision-centric LLM augmentation. We adapt GReaT-style conditional generation [4] to synthesize minority (Attrition=Yes) records strictly within training folds, select models on validation folds, and leave test folds untouched, with the goal of reducing costly false negatives.
- Dual benchmark suite and statistics. We compare logistic and boosted-tree baselines to a compact transformer–MLP; we report ROC-AUC, PR-AUC, balanced accuracy, and decision-facing summaries (alerts per true save), and apply DeLong tests for AUC differences across seeds.
- Progressive augmentation guidance. We trace performance and intervention load across synthetic scales (e.g., to the minority count) and identify robust operating regions for practice.
- Open artifacts. We release code, prompts, and configs to support reproduction consistent with community checklists [3].
1.4. Paper Structure
2. Materials and Methods
2.1. Dataset and Preprocessing
2.2. Synthetic Data Generation
LLM setup.
Scale of augmentation and validity checks.
CTGAN baseline.
2.3. Predictive Models and Training
2.4. Evaluation Metrics
- Discrimination. We report ROC-AUC and Average Precision (AP; area under the Precision–Recall curve[39]) on the test set and plot ROC and PR curves (Figure 3(b)). PR curves emphasize minority-class performance when positives are rare[38,42]. We also show a confusion matrix for the best model (Figure 4(a)) and derive precision and recall for the attrition class. In PR space, the no-skill baseline is a horizontal line at the positive prevalence (here ), and AP measures improvement over this baseline.
- Fairness. Following equal opportunity, we compute the true positive rate (TPR) for females and males and report the TPR gap on the test set[29]. We do not impose fairness constraints; this diagnostic surfaces potential bias. Because base attrition rates differ slightly by gender, we report groupwise TPRs alongside the gap.
- Interpretability. We use SHAP to attribute predictions and summarize global importance for the top model (AutoGluon ensemble)[30]. A SHAP summary plot (Figure 6) ranks features and visualizes value distributions, helping verify alignment with HR factors (e.g., OverTime, Age, MonthlyIncome, JobSatisfaction).
- Scalability and augmentation impact. Starting from the baseline, we evaluate after adding and synthetic “Yes” examples (training size increases of roughly and ). We track ROC-AUC, AP, fairness (TPR gap), and calibration (Brier), and visualize PR curves for AutoGluon and TabTransformer (Figure 5(b)). We also overlay real vs. synthetic projections (PCA/t-SNE) to monitor drift. In our runs, GReaT samples followed the real-data structure and preserved salient correlations, while CTGAN required careful tuning to avoid artifacts (e.g., duplicates or out-of-range values), consistent with prior reports[25].
2.5. Critic-Based Filtering for Quality Enhancement
Design and safeguards.
Relation to prior art.
3. Results
3.1. Overall Model Performance


3.2. Fairness Analysis
3.3. Reliability and Calibration

3.4. Interpretability of Model Predictions
3.5. GReaT vs. GAN: Practical Comparison

4. Discussion
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| MDPI | Multidisciplinary Digital Publishing Institute |
| DOAJ | Directory of open access journals |
| TLA | Three letter acronym |
| LD | Linear dichroism |
References
- Harter, J. & Adkins, A. This Fixable Problem Costs U.S. Businesses $1 Trillion. Gallup Workplace (2019). Available online: https://www.gallup.com/workplace/247391/fixable-problem-costs-businesses-trillion.aspx (accessed on 31 July 2025).
- Qi, Z.; Wang, Y.; Kong, D.; Wang, M. Applying oversampling before cross-validation will lead to high bias in small data sets. Scientific Reports 2024, 14, 62585. [Google Scholar] [CrossRef]
- NeurIPS Paper Checklist Guidelines (accessed 31 July 2025). Available online: https://neurips.cc/public/guides/PaperChecklist.
- Borisov, V.; Seßler, K.; Leemann, T.; Pawelczyk, M.; Kasneci, G. Language Models are Realistic Tabular Data Generators. In International Conference on Learning Representations ICLR. 2023. Available online: https://openreview.net/forum?id=HyeaUS1hXZ (accessed on 31 July 2025).
- Scikit Learn: Preprocessing data. 2025. https://scikit-learn.org/stable/modules/preprocessing.html (accessed on 21 August 2025).
- Scikit Learn: OneHotEncoder. 2025. https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.OneHotEncoder.html (accessed on 21 August 2025).
- Scikit Learn: train_test_split. 2025. https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.train_test_split.html (accessed on 21 August 2025).
- Scikit Learn: StratifiedKFold. 2025. https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedKFold.html (accessed on 21 August 2025).
- Scikit Learn: LogisticRegression. 2025. https://scikit-learn.org/stable/modules/generated/sklearn.linear_model.LogisticRegression.html (accessed on 21 August 2025).
- Ian, T. Jolliffe; Jorge Cadima.Principal component analysis: a review and recent developments. Phil. Trans. R. Soc. A. 2016, 374. [Google Scholar] [CrossRef]
- Scikit Learn: PCA. 2025. https://scikit-learn.org/stable/modules/generated/sklearn.decomposition.PCA.html (accessed on 21 August 2025).
- Laurens van der Maaten; Geoffrey Hinton. Visualizing Data using t-SNE. Journal of Machine Learning Research 2008, 9(86), 2579–2605. http://jmlr.org/papers/v9/vandermaaten08a.html.
- Scikit Learn: TSNE. 2025. https://scikit-learn.org/stable/modules/generated/sklearn.manifold.TSNE.html (accessed on 21 August 2025).
- Samaneh Azadi; Catherine Olsson; Trevor Darrell; Ian J. Goodfellow; Augustus Odena. Discriminator Rejection Sampling. In Proceedings of the 7th International Conference on Learning Representations, ICLR. 2018. https://openreview.net/forum?id=S1GkToR5tm.
- Salimans, Tim; Goodfellow, Ian; Zaremba, Wojciech; Cheung, Vicki; Radford, Alec; Chen, Xi. Improved techniques for training GANs. In Proceedings of the 30th International Conference on Neural Information Processing Systems NeurIPS. 2016. [CrossRef]
- YData-Synthetic: Synthesize tabular data. 2025. https://docs.synthetic.ydata.ai/1.3/synthetic_data/single_table/ctgan_example/ (accessed on 21 August 2025).
- AutoGluon. TabularPredictor.fit. 2025 https://auto.gluon.ai/dev/api/autogluon.tabular.TabularPredictor.fit.html (accessed on 21 August 2025).
- Scikit Learn. BalancedRandomForestClassifier. 2025. https://imbalanced-learn.org/stable/references/generated/imblearn.ensemble.BalancedRandomForestClassifier.html (accessed on 21 August 2025).
- Scikit Learn. roc_auc_score. 2025. https://scikit-learn.org/stable/modules/generated/sklearn.metrics.roc_auc_score.html (accessed on 21 August 2025).
- Chao Chen; Andy Liaw; Leo Breiman. Using Random Forest to Learn Imbalanced Data. University of California, Berkeley, 2004. https://statistics.berkeley.edu/sites/default/files/tech-reports/666.pdf.
- XGBoost. XGBoost Parameters. 2025. https://xgboost.readthedocs.io/en/stable/parameter.html (accessed on 21 August 2025).
- Tianqi Chen; Carlos Guestrin. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data MiningKDD’16. 2016. [CrossRef]
- Scikit Learn. GridSearchCV. 2025. https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GridSearchCV.html (accessed on 21 August 2025).
- Scikit Learn. StratifiedKFold. 2025. https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedKFold.html (accessed on 21 August 2025.
- Xu, Lei; Skoularidou, Maria; Cuesta-Infante, Alfredo; Veeramachaneni, Kalyan. Modeling Tabular Data using Conditional GAN. In Proceedings of the 33rd International Conference on Neural Information Processing Systems NeurIPS. 2019. [CrossRef]
- Austin A. Barr; Robert Rozman; Eddie Guo. Generative adversarial networks vs large language models: a comparative study on synthetic tabular data generation. Available online: https://arxiv.org/abs/2502.14523 (accessed on 19 August 2025).
- Xin Huang; Ashish Khetan; Milan Cvitkovic; Zohar Karnin. TabTransformer: Tabular Data Modeling Using Contextual Embeddings. Available online: https://arxiv.org/abs/2012.06678 (accessed on 19 October 2025).
- Decision Science and Analytics for Management (DSAM) Topic Page. MDPI. Available online: https://www.mdpi.com/topics/9D1436TTND (accessed on 31 July 2025).
- Hardt, M.; Price, E.; Srebro, N. Equality of Opportunity in Supervised Learning. In Proceedings of the 30th International Conference on Neural Information Processing Systems,NeurIPS 2016; Barcelona, Spain, 2016.
- Lundberg, S.M.; Lee, S.-I. A Unified Approach to Interpreting Model Predictions. In Proceedings of the 31th International Conference on Neural Information Processing Systems,NeurIPS 2017; Long Beach, California, USA, 2017.
- IBM Watson Analytics. IBM HR Analytics Employee Attrition & Performance Dataset; 2015. Available online: https://www.kaggle.com/datasets/pavansubhasht/ibm-hr-analytics-attrition-dataset (accessed on 31July 2025).
- Edward J. Hu; Yelong Shen; Phillip Wallis; Zeyuan Allen-Zhu; Yuanzhi Li; Shean Wang; Lu Wang; Weizhu Chen. LoRA: Low-Rank Adaptation of Large Language Models. In The Tenth International Conference on Learning Representations, ICLR, 2022.
- Erickson, N.; Mueller, J.; Zhang, H.; et al. AutoGluon-Tabular: Robust and Accurate AutoML for Structured Data. arXiv 2020, arXiv:2003.06505.
- Zhiqiang; Tang; Haoyang Fang; Su Zhou; Taojiannan Yang; Zihan Zhong; Tony Hu; Katrin Kirchhoff; George Karypis. AutoGluon-Multimodal (AutoMM): Supercharging Multimodal AutoML with Foundation Models. In The International Conference on Automated Machine Learning AutoML, 2024.
- Huang, H.; Zhang, H.; Li, K.; et al. TabTransformer: Tabular Data Modeling Using Contextual Embeddings. arXiv 2020, arXiv:2012.06678.
- Feldman, Michael; Friedler, Sorelle A.; Moeller, John; Scheidegger, Carlos; Venkatasubramanian, Suresh. Certifying and removing disparate impact. In Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; Sydney, NSW, Australia, 2015.
- Katy Molodianovitch; David Faraggi; Benjamin Reiser. Comparing the Areas Under Two Correlated ROC Curves: Parametric and Non-Parametric Approaches. Biometrics 2006, 48(5), 745-757. [CrossRef]
- Saito, T.; Rehmsmeier, M. The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets. PLOS ONE 2015, 10(3), 1-21. [CrossRef]
- Peter A. Flach, Meelis Kull. Precision-Recall-Gain curves: PR analysis done right. In Proceedings of the 29th International Conference on Neural Information Processing Systems NeurIPS; Montreal, Canada. [CrossRef]
- BRIER, G. W. (1950). VERIFICATION OF FORECASTS EXPRESSED IN TERMS OF PROBABILITY. Monthly Weather Review 2023 78(1), 1-3. [CrossRef]
- Telmo Silva Filho; Hao Song; Miquel Perello-Nieto; Raul Santos-Rodriguez; Meelis Kull; Peter Flach. Classifier calibration: a survey on how to assess and improve predicted class probabilities. Mach Learn 2023, 112, 3211–3260. [CrossRef]
- Davis, J.; Goadrich, M. The Relationship Between Precision-Recall and ROC Curves. In Proceedings of the 23rd International Conference on Machine learning ICML; 2006. Pittsburgh, PA, USA. [CrossRef]
- Chawla, N.V.; Bowyer, K.W.; Hall, L.O.; Kegelmeyer, W.P. SMOTE: Synthetic Minority Over-sampling Technique. Journal of Artificial Intelligence ResearchJ. Artif. Int. Res. 2002, 16, 321–357. [Google Scholar] [CrossRef]
- Goodfellow, Ian J.; Pouget-Abadie, Jean; Mirza, Mehdi; Xu, Bing; Warde-Farley, David; Ozair, Sherjil; Courville, Aaron; Bengio, Yoshua. In Proceedings of the 28th International Conference on Neural Information Processing Systems NeurIPS; New York, NY, USA. [CrossRef]
- Ryan D. Turner; Jane Hung; Yunus Saatci; Jason Yosinski. Metropolis-Hastings Generative Adversarial Networks. In Proceedings of the 36rd International Conference on Machine learning ICML; 2019. Pittsburgh, PA, USA. https://proceedings.mlr.press/v97/turner19a/turner19a.pdf.
- Guo, Chuan; Pleiss, Geoff; Sun, Yu; Weinberger, Kilian Q.. On calibration of modern neural networks. In Proceedings of the 34rd International Conference on Machine learning ICML; 2017. Sydney, NSW, Australia. [CrossRef]
- Xin Ding; Yongwei Wang; Z. Jane Wang; William J. Welch. Efficient subsampling of realistic images from GANs conditional on a class or a continuous variable. Neurocomputing 2023, 517, 188-200. [CrossRef]
- Sajjadi, Mehdi S. M.; Bachem, Olivier; Lucic, Mario; Bousquet, Olivier; Gelly, Sylvain. Assessing generative models via precision and recall. In Proceedings of the 32nd International Conference on Neural Information Processing Systems NeurIPS; 2018. [CrossRef]
| 1 | At the GAN optimum , so . |
| 2 | Leaders and scarce roles tend to sit near the upper end. |


| Model | Performance Evaluation | ||||
|---|---|---|---|---|---|
| ROC AUC | Precision | Recall | Avg Precision | ||
| AutoGluon | 0.902 | 0.900 | 0.409 | 0.562 | 0.761 |
| TabTransformer | 0.899 | 0.827 | 0.363 | 0.505 | 0.704 |
| Logistic Regression | 0.894 | 0.833 | 0.303 | 0.444 | 0.698 |
| Random Forest | 0.852 | 0.800 | 0.181 | 0.296 | 0.568 |
| XGBoost | 0.841 | 0.800 | 0.363 | 0.500 | 0.629 |
| Method | Aug. Size | ROC AUC | Acc | Precision | Recall | MCC | |
|---|---|---|---|---|---|---|---|
| LLM-based | 50 | 0.906 | 0.907 | 0.857 | 0.455 | 0.594 | 0.582 |
| 100 | 0.902 | 0.914 | 0.833 | 0.530 | 0.648 | 0.622 | |
| 200 | 0.897 | 0.900 | 0.824 | 0.424 | 0.560 | 0.546 | |
| GAN-based | 50 | 0.904 | 0.912 | 0.966 | 0.424 | 0.589 | 0.607 |
| 100 | 0.901 | 0.912 | 0.865 | 0.485 | 0.621 | 0.607 | |
| 200 | 0.897 | 0.909 | 0.842 | 0.485 | 0.615 | 0.596 |
| Model | AUC | 95% CI (DeLong) |
|---|---|---|
| AutoGluon | ||
| TabTransformer | ||
| Logistic Regression | ||
| Random Forest | ||
| XGBoost | ||
| KNN |
| Model i | Model j | AUC | z | p (raw) | ||
|---|---|---|---|---|---|---|
| AutoGluon | XGBoost | |||||
| AutoGluon | Random Forest | |||||
| AutoGluon | KNN | |||||
| AutoGluon | Logistic Regression | |||||
| AutoGluon | TabTransformer | |||||
| XGBoost | Random Forest | |||||
| XGBoost | KNN | |||||
| XGBoost | Logistic Regression | |||||
| XGBoost | TabTransformer | |||||
| Random Forest | KNN | |||||
| Random Forest | Logistic Regression | |||||
| Random Forest | TabTransformer | |||||
| KNN | Logistic Regression | |||||
| KNN | TabTransformer | |||||
| Logistic Regression | TabTransformer |
| Variant | AUC | 95% CI (DeLong) |
|---|---|---|
| LLM+50 | 0.906 | [0.864, 0.947] |
| GAN+50 | 0.904 | [0.861, 0.948] |
| LLM+100 | 0.902 | [0.860, 0.944] |
| GAN+100 | 0.901 | [0.858, 0.944] |
| LLM+200 | 0.897 | [0.856, 0.939] |
| GAN+200 | 0.897 | [0.856, 0.939] |
| Variant i | Variant j | AUC | z | p (raw) | ||
|---|---|---|---|---|---|---|
| LLM+50 | LLM+100 | |||||
| LLM+50 | LLM+200 | 9.451e-01 | ||||
| LLM+50 | GAN+50 | 9.738e-01 | ||||
| LLM+50 | GAN+100 | 9.451e-01 | ||||
| LLM+50 | GAN+200 | 9.451e-01 | ||||
| LLM+100 | LLM+200 | 9.451e-01 | ||||
| LLM+100 | GAN+50 | 9.738e-01 | ||||
| LLM+100 | GAN+100 | 9.808e-01 | ||||
| LLM+100 | GAN+200 | 9.451e-01 | ||||
| LLM+200 | GAN+50 | 9.451e-01 | ||||
| LLM+200 | GAN+100 | 9.738e-01 | ||||
| LLM+200 | GAN+200 | 9.808e-01 | ||||
| GAN+50 | GAN+100 | 9.451e-01 | ||||
| GAN+50 | GAN+200 | 9.451e-01 | ||||
| GAN+100 | GAN+200 | 9.451e-01 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).