Submitted:
22 July 2026
Posted:
23 July 2026
You are already at the latest version
Abstract

Keywords:
1. Introduction
- Which ML approaches demonstrate the most reliable performance as default-risk classifiers for PIT compliance, the specific risk-modeling task facing revenue authorities?
- Which economic (debt-related) and socio-demographic (debtor-related) attributes most significantly drive taxpayer default risk, and therefore warrant inclusion in a public-sector risk-scoring model?
- How can ML-based risk scoring be operationalized within a risk management framework that meets the accountability and resource-allocation demands specific to public financial management, rather than the private-sector credit risk contexts in which such techniques were first developed?
2. Theoretical and Empirical Background
2.1. Economic and Behavioral Foundations of Tax Compliance
2.2. Machine Learning in Tax Compliance Research
2.3. Explainable AI and Model Interpretability in Risk Modeling
3. Materials and Methods

3.1. Business Understanding
3.2. Data Understanding
- Debt-related factors: the tax amount owed, which is traditionally used by the tax administration for prioritizing tax collection cases(OECD, 2014), and the tax category reflecting the three income-tax chapters defined under Lebanese tax law (MOF, 2008): the first chapter covers profits from industrial, commercial, and non-commercial activity; the second covers salaries, wages, and pension income; and the third covers revenues from movable capital (e.g., interest, dividends, and other investment income). Each chapter is levied at a different rate and is a recognized indicator of compliance risk (Allingham & Sandmo, 1972; Alm et al., 1995; Graetz & Wilde, 1985; Kirchler et al., 2008).
-
Socio-demographic (debtor-related) factors:
- Individual: gender, age, marital status, and collection unit (reflecting the address of the taxpayer)
- Household: number of children and whether the spouse works or not.
3.3. Data Preparation
- Feature engineering: each taxpayer’s date of birth was converted into age.
- Data cleansing: 5,385 exact-duplicate records (33.6% of the raw 16,010 instances) were identified and removed to improve data integrity.
- Encoding: categorical fields were numerically encoded so that all algorithms could process them consistently (Rodríguez et al., 2018).
- Scaling and skew correction: the tax amount field was highly right-skewed, so a standard transformation was applied to compress extreme values while preserving their relative size and sign; this and the other numeric fields (age, number of children) were then standardized to a common scale before training.
3.4. Modeling
- Random Forest (RF): an ensemble of tree-based classifiers combined via majority vote (Breiman, 2001).
- Neural Network (NN): a computational model of interconnected nodes loosely modeled on the brain (Ngai et al., 2011; Gershenson, 2003).
- Decision Tree (DT): a hierarchical, rule-based classifier that recursively partitions the feature space into increasingly homogeneous subsets to maximize class separability at each node (Breiman, Friedman, Olshen, & Stone, 1984).
- XGBoost: a scalable tree-boosting system built on gradient-boosted decision trees (Chen & Guestrin, 2016).
- Support Vector Machine (SVM): separates classes via the decision surface that maximizes the margin between them (Cortes & Vapnik, 1995).
3.5. Evaluation
3.6. Deployment
4. Results
4.1. Predictive Performance
4.3. Interpretability Findings
5. Discussion
6. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| PIT | Personal Income Tax |
| ML | Machine Learning |
| AI | Artificial Intelligence |
| SHAP | Shapley Additive Explanations |
| XGBoost | Extreme Gradient Boosting |
| RF | Random Forest |
| DT | Decision Tree |
| SVM | Support Vector Machine |
| MLP | Multi-Layer Perceptron |
| AUC | Area Under the (ROC) Curve |
References
- Abdul Rahman, R.; Masrom, S.; Omar, N.; Zakaria, M. An application of machine learning on corporate tax avoidance detection model. IAES International Journal of Artificial Intelligence (IJ-AI) 2020, 9(4). [Google Scholar] [CrossRef]
- Allingham, M.; Sandmo, A. Income tax evasion: A theoretical analysis. Journal of Public Economics 1972, 1, 323–338. [Google Scholar] [CrossRef]
- Alm, J.; Sanchez, I.; De Juan, A. Economic and noneconomic factors in tax compliance. Kyklos 1995, 48(1), 1–18. [Google Scholar] [CrossRef]
- Alrasheedi, M. A.; Ijaz, S.; Alrashdi, A. M.; Lee, S.-W. Advanced Tax Fraud Detection: A Soft-Voting Ensemble Based on GAN and Encoder Architecture. Mathematics 2025, 13(4), 642. [Google Scholar] [CrossRef]
- Belhadi, A.; Kamble, S. S.; Mani, V.; Benkhati, I.; Touriki, F. E. An ensemble machine learning approach for forecasting credit risk of agricultural SMEs’ investments in agriculture 4.0 through supply chain finance. Annals of Operations Research 2025, 345(2), 779–807. [Google Scholar] [CrossRef] [PubMed]
- Breiman, L. Random forests. Machine Learning 2001, 45(1), 5–32. [Google Scholar] [CrossRef]
- Breiman, L.; Friedman, J. H.; Olshen, R. A.; Stone, C. J. Classification and regression trees; Chapman & Hall/CRC, 1984. [Google Scholar]
- Ch, R. K.; Meenadevi, K.; Kumar D, D.; Nagaraj, R. Deep Learning in Credit Risk Assessment: A Data-Driven Approach to Transforming Financial Decision-Making and Risk Analytics. Journal of Risk and Financial Management 2026, 19(5), 361. [Google Scholar] [CrossRef]
- Chambi Condori, P. P.; Chambi Vásquez, M.; Saravia Ticona, T. Beyond accuracy: Economic performance of machine learning models in financial fraud detection. Journal of Risk and Financial Management 2026, 19(5), 332. [Google Scholar] [CrossRef]
- Chang, V.; Sivakulasingam, S.; Wang, H.; Wong, S. T.; Ganatra, M. A.; Luo, J. Credit risk prediction using machine learning and deep learning: A study on credit card customers. Risks 2024, 12(11), 174. [Google Scholar] [CrossRef]
- Chapman, P.; Clinton, J.; Kerber, R.; Khabaza, T.; Reinartz, T.; Shearer, C.; Wirth, R. CRISP-DM 1.0: Step-by-step data mining guide. SPSS inc. 2000. [Google Scholar] [CrossRef]
- Chen, T.; Guestrin, C. XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining 785–794; 2016; pp. 785–794. [Google Scholar]
- Cortes, C.; Vapnik, V. Support-vector network. Machine Learning 1995, 20, 1–25. [Google Scholar] [CrossRef]
- Da Silva, L. S.; Carvalho, R. N.; Souza, J. C. F. Predictive models on tax refund claims: Essays of data mining in Brazilian tax administration. International Conference on Electronic Government and the Information Systems Perspective (EGOVIS 2015); 2015; pp. 220–228. [Google Scholar]
- Di Oliveira, V.; Chaim, R. M.; Weigang, L.; Neto, S. A. P. B.; Filho, G. P. R. Towards a Smart Identification of Tax Default Risk with Machine Learning. In International Conference on Web Information Systems and Technologies, WEBIST - Proceedings; 2021; pp. 422–429. [Google Scholar] [CrossRef]
- Feld, L. P.; Frey, B. S. Tax compliance as the result of a psychological tax contract: The role of incentives and responsive regulation. Law & Policy 2007, 29(1), 102–120. [Google Scholar] [CrossRef]
- Galindo, J.; Tamayo, P. Credit risk assessment using statistical and machine learning: Basic methodology and risk modeling applications. Computational Economics 2000, 15(1–2), 107–143. [Google Scholar] [CrossRef]
- Gershenson, C. Artificial neural networks for beginners. arXiv 2003. [Google Scholar]
- Ginting, S. L. B.; Adler, J.; Ginting, Y. R.; Kurniadi, A. H. The development of bank applications for debtors’ selection by using Naïve Bayes classifier technique. IOP Conference Series: Materials Science and Engineering 2018, 407(1), 1–7. [Google Scholar] [CrossRef]
- Graetz, M. J.; Wilde, L. L. The economics of tax compliance: Fact and fantasy. National Tax Journal 1985, 38(3), 355–363. [Google Scholar] [CrossRef]
- Inna, L. The determinants of tax evasion among ukrainian households. Master Thesis, Kyiv School of Economics, 2016. [Google Scholar]
- Khandani, A. E.; Kim, A. J.; Lo, A. W. Consumer credit-risk models via machine-learning algorithms. Journal of Banking & Finance 2010, 34(11), 2767–2787. [Google Scholar] [CrossRef]
- Kirchler, E.; Hoelzl, E.; Wahl, I. Enforced versus voluntary tax compliance: The “slippery slope” framework. Journal of Economic Psychology 2008, 29(2), 210–225. [Google Scholar] [CrossRef]
- Koh, H. C.; Tan, W. C.; Goh, C. P. A two-step method to construct credit scoring models with data mining techniques. International Journal of Business and Information 2006, 1(1), 96–118. [Google Scholar]
- Lee, Y.; Kim, E. Deep Learning-based Delinquent Taxpayer Prediction: A Scientific Administrative Approach. KSII Transactions on Internet & Information Systems 2024, 18(1), 30. [Google Scholar] [CrossRef]
- Lundberg, S. M.; Lee, S.-I. A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems 2017, 30, 4765–4774. [Google Scholar]
- Lundberg, S. M.; Erion, G.; Chen, H.; DeGrave, A.; Prutkin, J. M.; Nair, B.; Katz, R.; Himmelfarb, J.; Bansal, N.; Lee, S.-I. From local explanations to global understanding with explainable AI for trees. Nature Machine Intelligence 2020, 2(1), 56–67. [Google Scholar] [CrossRef] [PubMed]
- Masekoameng, J. L.; Mbona, S. V.; Ananth, A.; Chifurira, R. Bayesian logistic regression for credit risk modelling among South African loan borrowers. Journal of Risk and Financial Management 2026, 19(5), 358. [Google Scholar] [CrossRef]
- MOF. Tax Procedures Law - TPL - and its amendments. 2008.
- Mohamad, A.; Radzuan, N.; Hamid, Z. Tax arrears amongst individual income taxpayers in Malaysia. Journal of Financial Crime 2017, 24(1), 17–34. [Google Scholar] [CrossRef]
- Ngai, E. W. T.; Hu, Y.; Wong, Y. H.; Chen, Y.; Sun, X. The application of data mining techniques in financial fraud detection: A classification framework and an academic review of literature. Decision Support Systems 2011, 50(3), 559–569. [Google Scholar] [CrossRef]
- Nuthalapati, A. Optimizing Lending Risk Analysis & Management with Machine Learning, Big Data, and Cloud Computing. Remittances Review 2022, 7(2), 172–184. [Google Scholar]
- OECD. Working smarter in tax debt management; OECD Publishing: Paris, 2014. [Google Scholar] [CrossRef]
- OECD. Advanced analytics for better tax administration: Putting Data to Work; OECD Publishing: Paris, 2016. [Google Scholar] [CrossRef]
- OECD. Tax Administration 2024: Comparative Information on OECD and other Advanced and Emerging Economies; OECD Publishing: Paris, 2024. [Google Scholar] [CrossRef]
- Rodríguez, P.; Bautista, M. A.; Gonzalez, J.; Escalera, S. Beyond one-hot encoding: Lower dimensional target embedding. Image and Vision Computing 2018, 75, 21–31. [Google Scholar] [CrossRef]
- Saputra, M. P. A.; Sukono; Riaman; Kartiwa, A.; Misiran, M.; Wahid, A. J. Deep learning for credit risk prediction in fintech lending: A systematic literature review on model architectures, imbalanced data handling, and research agenda. Journal of Risk and Financial Management 2026, 19(7), 465. [Google Scholar] [CrossRef]
- Siimon, Õ. R.; Lukason, O. A decision support system for corporate tax arrears prediction. Sustainability (Switzerland) 2021, 13(15). [Google Scholar] [CrossRef]
- Sudhakar, M.; Reddy, C. V. K. Two step credit risk assesment model for retail bank loan applications using decision tree data mining technique. International Journal of Advanced Research in Computer Engineering & Technology (IJARCET,) 2016, 5(3), 705–718. [Google Scholar]
- Strobl, C.; Boulesteix, A. L.; Zeileis, A.; Hothorn, T. Bias in random forest variable importance measures: Illustrations, sources and a solution. BMC Bioinformatics 2007, 8, 25. [Google Scholar] [CrossRef] [PubMed]
- Swenson, C. Using Machine Deep Learning AI to Improve Forecasting of Tax Payments for Corporations. Forecasting 2024, 6(4), 968–984. [Google Scholar] [CrossRef]
- Wirth, R.; Hipp, J. CRISP-DM: Towards a standard process model for data mining. In Proceedings of the 4th International Conference on the Practical Applications of Knowledge Discovery and Data Mining; 2000; pp. 29–39. [Google Scholar]
- Wu, R. C. F. Integrating neurocomputing and auditing expertise. Managerial Auditing Journal 1994, 9(3), 20–26. [Google Scholar] [CrossRef]
- Yan, X.; Zhang, C.; Zhang, S. Towards databases mining: Pre-processing collected data. Applied Artificial Intelligence 2003, 17(5–6), 545–561. [Google Scholar] [CrossRef]
- Zaim, H.; Sahbani, S. Artificial intelligence in tax compliance and evasion mitigation: Trends, mechanisms, and institutional implications. Journal of Risk and Financial Management 2026, 19(7), 513. [Google Scholar] [CrossRef]





| Category | Feature | Description | Raw type | Domain / values |
|---|---|---|---|---|
| Economic | Tax Amount | Total PIT due to the obligation | Float | ≥ 0 |
| Tax Category | PIT chapter/category | Integer (1–3) | {1, 2, 3} | |
| Sociodemographic — Individual | Gender | Taxpayer gender | Categorical (binary) | {F, M} |
| Age | Age at issuance year | Integer | ≥ 18 | |
| Marital Status | Legal marital status | Categorical | {S (Single), M (Married), D (Divorced), W (Widowed)} |
|
| Collection Unit | Taxpayer’s address | Categorical (coded) | 1…K units | |
| Sociodemographic — Household | Number of Children | Count of dependent children | Integer | ≥ 0 |
| Working Spouse | Spouse has declared income | Binary | {0=No, 1=Yes} |
| Model | Accuracy | Precision | Recall | F1 | ROC-AUC |
|---|---|---|---|---|---|
| XGBoost | 0.702 | 0.676 | 0.723 | 0.699 | 0.779 |
| Random Forest | 0.682 | 0.663 | 0.680 | 0.672 | 0.757 |
| Decision Tree | 0.651 | 0.596 | 0.840 | 0.697 | 0.744 |
| Support Vector Machine (RBF) | 0.625 | 0.614 | 0.582 | 0.598 | 0.670 |
| Neural Network (MLP) | 0.620 | 0.636 | 0.477 | 0.545 | 0.668 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).