Preprint
Article

This version is not peer-reviewed.

When Does a Fraud Model Travel? A Conceptual Framework for Predictive Transportability and External Validation in Financial Statement Fraud Detection

Submitted:

04 September 2026

Posted:

07 September 2026

You are already at the latest version

Abstract
Financial statement fraud detection has progressed from ratio-based screening and statistical classification toward machine learning, ensemble methods, and explainable artificial intelligence. Yet methodological progress has been evaluated predominantly within the datasets and institutional environments in which fraud models are developed. Strong predictive performance in a development sample does not establish that a model will remain reliable when applied to new firms, later periods, different industries, or different institutional environments. This study addresses this problem by developing a conceptual framework for the predictive transportability and external validation of financial statement fraud detection models. Predictive transportability is defined as the degree to which a model developed in a source environment retains acceptable predictive behavior in a specified target environment. The framework proposes four analytical dimensions of transportability: temporal, cross-firm, cross-industry, and cross-institutional. It further integrates dataset-shift concepts with external-validation logic and proposes a lifecycle comprising model development, target-context assessment, shift diagnosis, external validation, recalibration or adaptation, deployment, and continuous monitoring. The study argues that internal predictive performance is evidence of model behavior under observed conditions rather than proof of universal validity. It therefore shifts the central research question from “Does the fraud model work?” to “Where, when, and for whom does the model remain reliable?” The resulting framework provides methodological guidance for evaluating financial statement fraud models beyond their original development samples and establishes testable propositions for future temporal, firm-level, sectoral, and cross-country validation studies.
Keywords: 
;  ;  ;  ;  ;  ;  ;  

1. Introduction

Financial statement fraud is difficult to detect because fraudulent reporting is relatively rare, intentionally concealed, heterogeneous across firms, and typically observed only after an enforcement, audit, restatement, or investigative process identifies the underlying reporting problem. Researchers have therefore long sought to transform accounting and financial information into systematic indicators of elevated manipulation or misstatement risk.
Early quantitative approaches established the foundations of this research stream. Beneish (1999) demonstrated that accounting variables associated with financial statement distortions could be combined into a screening model for earnings manipulation. Dechow et al. (2011) subsequently developed an F-Score based on characteristics associated with material accounting misstatements. These approaches established an important methodological principle: observable accounting information can be converted into probabilistic warning signals that support ex ante screening rather than relying solely on ex post fraud discovery.
The emergence of machine learning expanded this paradigm. Perols (2011) compared statistical and machine-learning algorithms for financial statement fraud detection while explicitly considering fraud prevalence and asymmetric misclassification costs. Bao et al. (2020) later demonstrated the potential of ensemble machine learning using theory-motivated raw accounting numbers. Collectively, such studies shifted the methodological frontier from prespecified statistical relationships toward increasingly flexible algorithms capable of detecting nonlinear and high-dimensional patterns in financial information.
However, improved algorithmic flexibility introduced new evaluation challenges. Importantly, a subsequent erratum to Bao et al. (2020) showed that the treatment of serial fraud spanning training and test periods materially affected estimated model performance (Bao et al., 2022). The correction illustrates a broader methodological problem: observations can be statistically separated into training and test sets while remaining dependent at the firm, fraud-episode, or temporal level. Apparent out-of-sample performance may consequently overstate the ability of a model to operate in genuinely new environments.
More recent financial statement fraud research has moved beyond discrimination alone. Sodnomdavaa and Lkhagvadorj (2026), for example, integrated machine learning, explainable artificial intelligence, calibration, and Decision Curve Analysis using 969 firm-year observations from 132 Mongolian firms over 2013–2024. Their framework shifted attention from classification performance toward interpretability and decision utility. This development raises a further methodological question that precedes practical deployment: whether a model that performs well in its source environment retains its validity when transported elsewhere.
This concern has historical relevance in financial statement fraud research. Enkhbayar and Tsolmon (2015), using Mongolian company data, explicitly observed that fraud-detection models developed in different countries may behave differently because economic, business, taxation, and legal environments differ across jurisdictions. Their observation anticipated a problem that becomes more consequential as fraud analytics moves from relatively transparent statistical models toward increasingly complex machine-learning systems.
A sophisticated model can learn the relationships present in its development environment more accurately without necessarily learning relationships that remain stable outside that environment. Financial structures vary across industries. Fraud prevalence differs across populations. Accounting standards and enforcement practices evolve. Macroeconomic shocks alter the distributions of financial ratios. Ownership and governance structures vary internationally. Fraud strategies themselves may also change.
Strong internal predictive performance ≠ established external validity
Machine-learning research characterizes related problems through the broader concept of dataset shift, which arises when the statistical distribution generating deployment data differs from that underlying model development data. Moreno-Torres et al. (2012) distinguish several forms of distributional change and emphasize that changing training and test populations create distinct classification problems.
The distinction is especially consequential in financial statement fraud detection. A conventional random train-test split may place observations from the same firm, similar reporting periods, common industries, and an unchanged institutional regime in both subsets. The test observations are technically unseen, but the surrounding context may remain highly familiar.
Unseen observation ≠ unseen context
Lin (2024) identifies training-test construction, cross-period financial statement fraud, severe class imbalance, and preprocessing as important methodological considerations in machine-learning-based fraud detection. Cross-period fraud is particularly important because a fraud episode may span multiple reporting periods. If related observations are divided between training and test samples, models may benefit from information structures that will not be available when deployed on genuinely new firms or future periods.
This paper conceptualizes the resulting problem as predictive transportability. For the purposes of this study, predictive transportability is defined as the degree to which a financial statement fraud detection model developed in one source environment retains sufficiently reliable predictive behavior when applied to a specified target environment that differs in time, firms, industry composition, or institutional conditions.
This use of transportability must be distinguished from causal transportability. Pearl and Bareinboim (2014) use transportability in a formal causal-inference sense to describe conditions under which causal effects can be transferred across populations. The present study does not attempt to transport causal effects. Its object is a predictive model, X → P̂(Y = 1 | X). The qualifier predictive is therefore used throughout this paper to avoid conflating the two concepts.
A second distinction concerns external validation. In this study, external validation refers to the empirical process of evaluating a model using data sufficiently distinct from its development data, whereas predictive transportability refers to the extent to which acceptable model performance is retained in the specified target environment. This distinction is consistent with broader prediction-model research in which model performance is evaluated using independent but related populations and development-validation differences are explicitly considered (Debray et al., 2015).
Source data: Dₛ ~ Pₛ(X,Y)
Trained model: Mₛ = A(Dₛ)
Target data: Dₜ ~ Pₜ(X,Y)
Strong source-environment performance does not by itself establish target-environment performance. When Pₛ(X,Y) ≠ Pₜ(X,Y), target performance becomes an empirical question requiring validation. Importantly, distributional change does not necessarily imply model failure.
Dataset shift ≠ model failure; dataset shift⇒ need for target validation
A model may retain discrimination despite moderate covariate shift while losing calibration. Alternatively, both ranking and calibration may deteriorate. In more severe cases, changes in the relationship between predictors and fraud may render the original model unsuitable even after simple recalibration.
The relevant question for fraud analytics is consequently not merely “Does the model work?” but “Where, when, and for whom does the model remain reliable?”
The study addresses that question through three research questions:
  • RQ1. What forms of contextual change threaten the predictive transportability and external validity of financial statement fraud detection models?
  • RQ2. How should predictive transportability be evaluated across firms, industries, time periods, and institutional environments?
  • RQ3. Under what conditions should an existing fraud-detection model be retained, recalibrated, retrained, locally adapted, or rejected in a new target environment?
The study makes four contributions. First, it separates within-environment predictive performance from cross-context predictive transportability. Second, for analytical purposes, it proposes four dimensions of financial statement fraud-model transportability: temporal, cross-firm, cross-industry, and cross-institutional. These dimensions are proposed as an analytical typology rather than claimed as an established taxonomy. Third, the study integrates financial statement fraud research with dataset-shift and external-validation reasoning to develop a structured source-to-target validation framework. Fourth, it reframes model deployment as a lifecycle that explicitly allows for a decision not to deploy a transported model when target validity cannot be demonstrated.

2. From Fraud Prediction to the External Validity Problem

2.1. Statistical Fraud Models and Context Dependence

Early financial statement fraud studies primarily asked whether financial and accounting information could distinguish problematic reporting from normal reporting. Beneish (1999) demonstrated that multiple accounting indicators could be combined into an earnings-manipulation screening mechanism. Dechow et al. (2011) subsequently identified characteristics associated with material accounting misstatements and developed the F-Score as another structured warning mechanism.
These models contributed substantially to fraud analytics because they transformed accounting information into quantitative screening signals. Yet such models were estimated within particular populations, periods, accounting systems, and labeling processes.
Context dependence was explicitly recognized in an early Mongolian application. Enkhbayar and Tsolmon (2015) argued that models developed in different countries may operate differently because economic, business, taxation, and legal conditions vary across jurisdictions. Their study developed a logistic-regression-based financial analysis model using Mongolian company data. The broader methodological implication is that model validity is conditional on context unless evidence demonstrates otherwise.
The question becomes more important rather than less important as predictive models become more flexible. A highly adaptive algorithm may learn local statistical structure very effectively, including structure that is specific to the source environment.
Better source fit ≠ better target validity

2.2. Machine Learning and the Performance Paradigm

Machine learning substantially expanded the technical capacity of financial statement fraud detection. Perols (2011) showed that alternative statistical and machine-learning techniques produced different results and that their usefulness depended on fraud prevalence and misclassification consequences. Bao et al. (2020) subsequently demonstrated the predictive potential of ensemble methods using theoretically motivated accounting information.
Financial statement fraud research has since employed support vector machines, random forests, boosting methods, neural networks, hybrid algorithms, ensemble learning, and increasingly explainable approaches. Shahana et al. (2023) document this methodological expansion and identify continuing challenges involving data, class imbalance, model evaluation, and practical applicability.
Yet predictive sophistication does not automatically establish deployment validity. The Bao et al. case is instructive. Their original 2020 study reported strong results for an ensemble approach. A subsequent erratum corrected the treatment of serial fraud observations spanning training and testing periods (Bao et al., 2022). The correction demonstrates how firm-level and temporal dependence can affect apparently out-of-sample results.
The methodological lesson is broader than any individual study: test-set independence must be evaluated substantively, not only procedurally.

2.3. Prediction, Explainability, and the Remaining Question

Recent research has begun to move beyond classification metrics. Sodnomdavaa and Lkhagvadorj (2026) integrated machine learning with explainable artificial intelligence, calibration, and Decision Curve Analysis. This development recognizes that fraud-model quality cannot be represented adequately by a single discrimination metric.
However, interpretability and decision usefulness remain conditional on the validity of the underlying prediction. An interpretable prediction that is systematically miscalibrated in a new population remains problematic. Likewise, reliable explanation in the source environment does not establish reliable prediction in the target environment.
External validity therefore represents a distinct methodological layer that must be assessed when a model is transferred beyond its development environment.

3. Conceptual Research Design

3.1. Research Approach

This study adopts a conceptual research design rather than an empirical model-comparison design. Following Jaakkola (2020), conceptual research requires an explicit logic of knowledge creation, specification of how existing perspectives are selected and connected, and a defensible conceptual contribution.
The present study combines theory synthesis with conceptual model development. The synthesis connects three research domains: financial statement fraud prediction; machine-learning dataset shift and domain differences; and external validation and predictive model transportability.
The objective is not to conduct a systematic review, and no PRISMA-style completeness claim is made. Instead, the literature is selected according to its relevance to a clearly defined methodological problem: the transfer of a trained financial statement fraud model from a source environment to a materially different target environment.

3.2. Source and Target Environments

Let S denote the source environment and T denote the target environment. The development data satisfy Dₛ ~ Pₛ(X,Y), while deployment data satisfy Dₜ ~ Pₜ(X,Y). A predictive model is trained according to Mₛ = A(Dₛ). Predictive transportability becomes relevant when Pₛ(X,Y) ≠ Pₜ(X,Y).
Three distributional changes are especially useful for conceptualizing the problem.
Covariate shift
Pₛ(X) ≠ Pₜ(X), while Pₛ(Y|X) ≈ Pₜ(Y|X)
In financial statement fraud research, this may occur when firms in the target environment have different leverage, profitability, liquidity, or size distributions.
Prior or prevalence shift
Pₛ(Y) ≠ Pₜ(Y)
This may occur when observed fraud prevalence differs because of enforcement intensity, sample construction, market structure, or fraud-labeling mechanisms.
Concept shift
Pₛ(Y|X) ≠ Pₜ(Y|X)
This represents the more consequential case. A variable that indicates elevated fraud risk in one setting may have a weaker, stronger, or qualitatively different association in another setting. For example, the relation between leverage and fraud risk may differ across regulated financial institutions, manufacturing firms, state-controlled enterprises, or jurisdictions with different creditor-protection systems. Identical predictor values therefore do not necessarily have identical fraud implications across environments.

3.3. Analytical Definition of Predictive Transportability

Predictive transportability is target-specific. It should not be represented as a universal binary characteristic. Conceptually, PT = M(S,T), where PT represents predictive transportability between source environment S and target environment T. The same model may transport successfully to one industry or period but fail in another. This target-specific definition prevents unsupported claims that a model is simply generalizable.

4. Four Dimensions of Predictive Transportability

For analytical purposes, this study proposes four dimensions of fraud-model transportability. The dimensions are intended as a problem-structuring typology rather than an exhaustive or previously established taxonomy.

4.1. Temporal Transportability

Temporal transportability concerns whether a model trained using earlier financial statements remains valid in later periods. Temporal transfer can be threatened by macroeconomic shocks, inflation and interest-rate regimes, accounting-standard changes, enforcement changes, shifts in industry composition, changes in fraud strategies, and structural changes in financing behavior. A ratio associated with elevated fraud risk during a stable macroeconomic environment may behave differently during a crisis. Temporal validation should therefore distinguish random holdout evaluation from future-period holdout evaluation; the latter more closely resembles real-world deployment.

4.2. Cross-Firm Transportability

Cross-firm transportability concerns performance on firms not represented in model development. In panel-type financial data, repeated observations from the same company may share persistent characteristics. If the same firm appears in both training and test sets, the model may exploit entity-specific regularities. Cross-firm validation therefore requires firm separation, for example through grouped or leave-firm-out designs, to provide stronger evidence of performance on genuinely new entities.

4.3. Cross-Industry Transportability

Industries differ structurally in capital structure, working-capital cycles, revenue-recognition patterns, asset composition, regulation, and operating margins. Consequently, both P(X) and potentially P(Y|X) may vary across industries. A financial ratio therefore cannot automatically be assumed to carry an invariant fraud signal across sectors. Cross-industry validation should evaluate both discrimination and calibration deterioration.

4.4. Cross-Institutional Transportability

Cross-institutional transportability represents the broadest form of transfer. Target contexts may differ in accounting standards, legal enforcement, audit quality, tax regulation, disclosure practices, ownership concentration, capital-market development, governance systems, and fraud-label availability. Enkhbayar and Tsolmon (2015) explicitly identified national economic, business, tax, and legal characteristics as reasons why fraud-detection approaches may differ across countries. Cross-country transport therefore cannot be justified merely because predictor definitions appear identical. Mathematically identical ratios can have different institutional interpretations.
Table 1. Proposed dimensions of predictive transportability.
Table 1. Proposed dimensions of predictive transportability.
Dimension Source–target difference Main methodological risk Preferred validation logic
Temporal Earlier vs. later periods Concept/model drift Chronological holdout
Cross-firm Known vs. unseen firms Entity leakage or memorization Grouped or leave-firm-out validation
Cross-industry Different sectors Structural heterogeneity Industry holdout or external industry sample
Cross-institutional Different countries or regimes Institutional domain shift Country/regime external validation
These dimensions are not mutually exclusive. A realistic deployment may simultaneously involve temporal, firm-level, industry, and institutional shift. The larger the combined source-target difference, the weaker the justification for assuming unchanged performance without direct validation.

5. A Source-to-Target External Validation Framework

The proposed framework consists of six linked stages. Figure 1 summarizes the source-to-target logic and emphasizes that successful model development is only the beginning of a deployment decision.

5.1. Stage 1: Source Model Development

A model is developed from Dₛ. At this stage, internal validation evaluates model development quality. Relevant measures may include PR-AUC, ROC-AUC, recall, precision, F1-score, calibration, and decision utility. These metrics characterize source-domain performance and should not be interpreted as universal performance parameters.

5.2. Stage 2: Target-Context Assessment

Before transferring the model, researchers should define the target environment explicitly. At minimum, T should specify the relevant time period, firms, industries, and institutional setting. This step converts external validity from an abstract claim into a target-specific question. The appropriate question is not whether a model is externally valid in general, but whether it is sufficiently valid for target environment T.

5.3. Stage 3: Shift Diagnosis

Source and target data should be evaluated for material differences. The objective is not to prove identical distributions, but to identify where and how S and T differ. Potential diagnostics include predictor-distribution comparisons, fraud prevalence, missingness patterns, accounting-rule differences, sector composition, calibration drift, feature-importance changes, and institutional differences. A detected shift is a warning signal, not automatic proof of model invalidity.

5.4. Stage 4: External Validation

The source model Mₛ is applied without re-estimation to Dₜ. At least three forms of performance should be distinguished: discrimination, calibration, and decision usefulness. Discrimination asks whether the model ranks higher-risk cases above lower-risk cases. Calibration asks whether predicted probabilities correspond sufficiently to observed outcomes. Decision usefulness asks whether model-guided actions remain beneficial under target-context costs and constraints. A model may preserve ranking while losing probability calibration, so target evaluation should not rely on discrimination alone.

5.5. Stage 5: Adaptation Decision

External validation should result in an explicit deployment decision. Retain the model when target performance remains acceptable without material adjustment. Recalibrate when ranking is sufficiently preserved but probability estimates have shifted. Retrain or locally adapt when more substantial target-context differences exist and sufficient local data are available. Reject transport when evidence indicates that the source model cannot be defended in the target context. The final option is essential: the purpose of validation is not to guarantee deployment.

5.6. Stage 6: Deployment and Continuous Monitoring

External validation is not permanent certification. Even after successful deployment, Pₜ(X,Y) can continue to evolve. Monitoring should therefore examine discrimination drift, calibration drift, fraud prevalence, predictor drift, changing explanation patterns, false-positive burden, and false-negative consequences. The lifecycle becomes development → target assessment → shift diagnosis → external validation → adaptation → deployment → monitoring → revalidation.

6. Research Propositions

P1. Internal performance
Strong within-environment predictive performance is insufficient to establish the external validity of a financial statement fraud detection model in a materially different target environment.
P2. Source–target divergence
Greater source–target divergence in temporal, firm-level, industry, or institutional characteristics is associated with greater uncertainty regarding target-context model performance in the absence of local validation.
P3. Temporal validation
When temporal dependence or concept drift is present, chronologically separated validation provides a more realistic estimate of future deployment performance than random observation-level splitting.
P4. Firm separation
When repeated firm observations are present, firm-separated validation provides stronger evidence of cross-firm transportability than validation designs that allow the same firm to appear in both training and test samples.
P5. Calibration vulnerability
Source–target distributional differences can impair probability calibration even when relative discrimination remains comparatively stable.
P6. Adaptation intensity
As source–target differences move from covariate or prevalence shift toward substantial changes in P(Y|X), simple recalibration becomes less sufficient and local retraining, adaptation, or model rejection becomes increasingly appropriate.
Table 2. Testable implications of the framework.
Table 2. Testable implications of the framework.
Proposition Illustrative empirical comparison
P1 Internal holdout vs. independent external sample
P2 Model uncertainty/degradation as source–target distance increases
P3 Random split vs. future-period validation
P4 Random split vs. leave-firm-out validation
P5 Discrimination stability vs. calibration deterioration
P6 No adaptation vs. recalibration vs. retraining/adaptation

7. Discussion

This study reframes financial statement fraud-model evaluation as a source-to-target validity problem. Fraud analytics has progressed through several methodological stages. Statistical models demonstrated that accounting characteristics can be used to screen for manipulation or misstatement. Machine learning increased flexibility and predictive capacity. Explainable AI sought to make complex models more interpretable, while decision-oriented approaches increasingly recognize calibration and practical utility. Yet these developments do not resolve a prior question: does the model retain its validity where it will actually be used?
The distinction is consequential because predictive performance is conditional on an evaluation design. A randomly generated test sample primarily addresses generalization within the source environment. A future-period holdout addresses temporal transportability. A leave-firm-out evaluation addresses cross-firm transportability. An industry holdout addresses cross-industry transportability. A foreign-country sample addresses cross-institutional transportability. These tests answer different questions and should not be treated as interchangeable forms of out-of-sample evaluation.
A second implication concerns performance metrics. Fraud research often emphasizes discrimination metrics because of severe class imbalance. Such metrics remain important, but transportability requires a broader evaluation. A transported model may preserve ranking while losing probability accuracy. If probabilities influence audit prioritization or resource allocation, calibration deterioration may be professionally important even when AUC remains strong.
A third implication concerns context. The framework does not assume that any country-specific or industry-specific model is intrinsically inferior to a globally trained model. A locally developed model can be highly useful within its intended environment. The problem arises when local validity is implicitly interpreted as universal validity. Conversely, pooling countries or industries does not automatically establish transportability because aggregate model performance can conceal subgroup-specific deterioration. The appropriate objective is explicit validation against a defined target context.
A fourth implication concerns model complexity. There is no theoretical reason to assume that a more complex algorithm is inherently more transportable. Greater flexibility can improve source-domain performance, but it can also allow a model to learn more source-specific regularities. Complexity therefore does not imply transportability; it remains an empirical question.
Finally, transportability should be understood as dynamic. A model that successfully transports today may no longer remain appropriate several years later. External validation therefore does not create permanent validity. Continuous monitoring and revalidation are necessary whenever deployment conditions materially change.

8. Implications

8.1. Implications for Fraud-Detection Research

Researchers should state the precise form of generalization their evaluation design supports. Rather than reporting simply that a model demonstrated strong out-of-sample performance, authors should indicate whether the test set represents new observations, new firms, future periods, new sectors, or new institutional environments. Researchers should also avoid generalizing beyond the validation design. For example, chronological validation can support evidence about temporal transfer but does not by itself establish cross-country validity.

8.2. Implications for Model Developers

Fraud-model documentation should identify the intended deployment domain. At minimum, a model record should report the development period, included firms, industries, jurisdiction, outcome definition, and validation domain. A model without an explicit domain of applicability creates a risk that users interpret source-context performance as universal. Developers should also retain the possibility that target validation results in rejection rather than deployment.

8.3. Implications for Auditors and Organizations

Organizations should not adopt fraud models solely because published performance metrics are high. Before use, they should assess how similar the target population is to the development population, whether comparable firms and sectors were represented, whether the development period is economically comparable, whether the model has been externally validated, whether its probabilities are calibrated for the target environment, and whether local recalibration or retraining is required. Model adoption is therefore a validation decision rather than merely a software-procurement decision.

8.4. Implications for Emerging Markets

Predictive transportability may be especially important in emerging markets where locally labeled fraud data are limited and researchers or practitioners may rely on models developed using data from larger capital markets. Direct transfer implicitly assumes that the conditional relationship between predictors and fraud is sufficiently similar across environments. Differences in financial reporting, enforcement, ownership, tax systems, governance, and market discipline may affect both predictor distributions and their association with observed fraud. This does not imply that foreign models cannot be used; it implies that they should be externally validated before reliance.

9. Boundary Conditions, Limitations, and Future Research

This study is conceptual and has several limitations. First, the proposed four-dimensional typology has not yet been empirically validated as an exhaustive representation of financial statement fraud-model transportability. Other forms of transfer may be relevant, including changes in data vendors, reporting formats, label definitions, or feature availability.
Second, the framework does not establish numerical thresholds defining acceptable transportability. A reduction in PR-AUC or calibration performance may be acceptable in one application and unacceptable in another. Such thresholds depend on fraud prevalence, model use, error consequences, available alternatives, and regulatory expectations.
Third, the paper adapts concepts from general machine learning and prediction-model validation to financial statement fraud detection. Although the underlying methodological issues are relevant, future research should determine which shift diagnostics and external-validation measures are most informative specifically for fraud analytics.
Fourth, the framework focuses on supervised predictive models. It may require modification for unsupervised anomaly detection, graph-based fraud systems, large language models, or continuously learning architectures.
Fifth, observed fraud labels themselves can differ between jurisdictions and periods. Apparent dataset shift may therefore partly reflect label-process shift rather than underlying fraud behavior. This issue deserves separate investigation.
Future empirical research should directly compare random splitting with temporal splitting, random splitting with leave-firm-out validation, within-industry with cross-industry validation, and within-country with cross-country validation. A particularly useful design would train identical models in a source environment and evaluate them sequentially across increasingly distant target environments. Such studies could test whether performance degradation is monotonic with contextual distance and whether particular model families are more resistant to distributional change.
A second research stream could examine which performance dimension deteriorates first: calibration, discrimination, or decision utility. A third could test whether explanation patterns themselves transport across environments. If predictive importance changes substantially between source and target populations, explanation drift may itself provide evidence of domain instability.

10. Conclusion

Financial statement fraud detection has developed rapidly from accounting-ratio screening and conventional statistical models toward advanced machine learning, explainable AI, and decision-oriented analytics. However, increasing model sophistication does not answer a fundamental deployment question: where, when, and for whom does the model work?
This study addresses that question through the concept of predictive transportability. The proposed framework distinguishes four analytically defined dimensions: temporal, cross-firm, cross-industry, and cross-institutional transportability. It argues that strong internal predictive performance should be interpreted as evidence of model performance under observed source conditions rather than proof of universal external validity.
The framework further distinguishes dataset shift from model failure. Dataset shift does not automatically imply failure, but it does create a requirement for target-specific validation. Responsible deployment therefore follows a source-to-target lifecycle comprising development, target assessment, shift diagnosis, external validation, adaptation, deployment, monitoring, and revalidation.
The framework explicitly includes model rejection as a legitimate validation outcome. A model should not be transported merely because it is available or because it previously achieved strong predictive performance. The central methodological principle is therefore simple: a fraud model travels only as far as its external validation supports.
By shifting attention from predictive performance within a dataset toward validity across environments, this study provides a conceptual bridge between financial statement fraud analytics, dataset-shift research, and external prediction-model validation. The resulting propositions provide a research agenda for determining how fraud models behave across firms, periods, sectors, and institutional systems and for establishing more defensible boundaries around the use of machine learning in financial statement fraud detection.

Author Contributions

Conceptualization, methodology, investigation, writing—original draft preparation, writing—review and editing, and visualization: L.G. The author has read and agreed to the final version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable. This conceptual study did not involve human participants or animals.

Data Availability Statement

No new data were created or analyzed in this conceptual study.

Use of Artificial Intelligence

Generative artificial intelligence (ChatGPT, OpenAI) was used during manuscript preparation to assist with drafting, restructuring, and language refinement. The author critically reviewed and revised the conceptual arguments, independently verified the cited literature and bibliographic information, and takes full responsibility for the accuracy, integrity, and final content of the manuscript.

Conflicts of Interest

The author declares no conflicts of interest.

References

  1. Bao, Y.; Ke, B.; Li, B.; Yu, Y. J.; Zhang, J. Detecting accounting fraud in publicly traded U.S. firms using a machine learning approach. J. Account. Res. 2020, 58(1), 199–235. [Google Scholar] [CrossRef]
  2. Bao, Y.; Ke, B.; Li, B.; Yu, Y. J.; Zhang, J. Erratum. J. Account. Res. 2022, 60(4), 1635–1646. [Google Scholar] [CrossRef]
  3. Beneish, M. D. The detection of earnings manipulation. Financ. Anal. J. 1999, 55(5), 24–36. [Google Scholar] [CrossRef]
  4. Debray, T. P. A.; Vergouwe, Y.; Koffijberg, H.; Nieboer, D.; Steyerberg, E. W.; Moons, K. G. M. A new framework to enhance the interpretation of external validation studies of clinical prediction models. J. Clin. Epidemiol. 2015, 68(3), 279–289. [Google Scholar] [CrossRef] [PubMed]
  5. Dechow, P. M.; Ge, W.; Larson, C. R.; Sloan, R. G. Predicting material accounting misstatements. Contemp. Account. Res. 2011, 28(1), 17–82. [Google Scholar] [CrossRef]
  6. Enkhbayar, Ch.; Tsolmon, S. Possibility of detecting fraudulent practices in financial statements. Baikal Res. J. 2015, 6(4). [Google Scholar] [CrossRef] [PubMed]
  7. Jaakkola, E. Designing conceptual articles: Four approaches. AMS Rev. 2020, 10, 18–26. [Google Scholar] [CrossRef]
  8. Lin, D. Key considerations to be applied while leveraging machine learning for financial statement fraud detection: A review. IEEE Access 2024, 12, 168213–168228. [Google Scholar] [CrossRef]
  9. Moreno-Torres, J. G.; Raeder, T.; Alaiz-Rodríguez, R.; Chawla, N. V.; Herrera, F. A unifying view on dataset shift in classification. Pattern Recognit. 2012, 45(1), 521–530. [Google Scholar] [CrossRef]
  10. Pearl, J.; Bareinboim, E. External validity: From do-calculus to transportability across populations. Stat. Sci. 2014, 29(4), 579–595. [Google Scholar] [CrossRef]
  11. Perols, J. Financial statement fraud detection: An analysis of statistical and machine learning algorithms. Audit. A J. Pract. Theory 2011, 30(2), 19–50. [Google Scholar] [CrossRef]
  12. Shahana, T.; Lavanya, V.; Bhat, A. R. State of the art in financial statement fraud detection: A systematic review. Technol. Forecast. Soc. Change 2023, 192, 122527. [Google Scholar] [CrossRef]
  13. Sodnomdavaa, T.; Lkhagvadorj, G. Financial statement fraud detection through an integrated machine learning and explainable AI framework. J. Risk Financ. Manag. 2026, 19(1), 13. [Google Scholar] [CrossRef]
Figure 1. Predictive Transportability and External Validation Framework. Proposed source-to-target framework for assessing the predictive transportability and external validation of financial statement fraud detection models. Source–target differences are diagnosed before target-context validation of discrimination, calibration, and decision utility. The validation result supports retention, recalibration, local adaptation/retraining, or rejection, followed by monitoring and revalidation where appropriate.
Figure 1. Predictive Transportability and External Validation Framework. Proposed source-to-target framework for assessing the predictive transportability and external validation of financial statement fraud detection models. Source–target differences are diagnosed before target-context validation of discrimination, calibration, and decision utility. The validation result supports retention, recalibration, local adaptation/retraining, or rejection, followed by monitoring and revalidation where appropriate.
Preprints 231763 g001
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.