Submitted:
17 September 2026
Posted:
17 September 2026
You are already at the latest version
Abstract
Machine learning models are increasingly used in credit risk, financial statement fraud detection, financial distress prediction, and related financial classification tasks with substantial consequences. Yet model comparison still often privileges statistical discrimination even though superior predictive performance does not necessarily imply superior financial decisions. This limitation is acute in settings involving rare events, where class imbalance, probability miscalibration, temporal distribution shift, asymmetric error costs, operational capacity constraints, and governance requirements interact. This conceptual and methodological paper develops an integrated framework for decision validity without introducing a new dataset or estimating additional models. The synthesis draws on two complementary empirical streams: evaluating deep learning for financial statement fraud under severe class imbalance and temporally consistent evaluation of credit risk models that accounts for asymmetric costs under distributional shift. These streams integrate established research on precision and recall, probabilistic calibration, concept drift, classification with unequal costs, profitability-based credit scoring, explainable artificial intelligence, and model governance. The proposed framework consists of five sequential gates: validity for rare event detection, probability validity, temporal validity, decision-utility validity, and governance validity. The gates are intentionally not compensatory: strong performance at a later stage should not automatically offset a fundamental failure at an earlier stage. The paper also maps credit default and financial statement fraud, a minimum reporting standard, and a research agenda for dynamic thresholds, temporal calibration, explanation stability, and utility under capacity constraints. The central conclusion is that model superiority in financial machine learning is conditional rather than absolute and should be asserted only relative to an explicit deployment and decision environment.
Keywords:
financial machine learning
; rare event prediction
; credit risk
; financial statement fraud
; class imbalance
; calibration
; temporal distribution shift
; cost-sensitive evaluation
; explainable artificial intelligence
; model governance
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.