Submitted:
04 September 2026
Posted:
07 September 2026
You are already at the latest version
Abstract
Financial statement fraud research has progressed from rule-based and statistical screening toward machine learning and explainable artificial intelligence (XAI). Recent finance and auditing literature has also begun to frame AI as part of human–AI decision and governance systems rather than as a stand-alone predictive tool. Yet a narrower audit-specific problem remains unresolved: how a model-generated fraud-risk signal should be evaluated, challenged, and translated into a proportionate audit response when explanations may be unstable, error consequences are asymmetric, and audit resources are constrained. This conceptual article develops a Human–AI Decision Governance Framework for financial statement fraud risk assessment. Using theory synthesis and conceptual model development, the framework integrates six layers: fraud-risk signaling, explanation assurance, professional judgment, decision utility, audit response, and governance with feedback. It treats AI outputs as decision inputs rather than fraud conclusions and distinguishes model performance from explanation reliability, professional validity, and decision utility. Six propositions specify testable relationships among predictive performance, explanation quality, professional skepticism, resource-sensitive thresholds, human override, and lifecycle governance. The contribution is deliberately domain-specific: it connects fraud analytics and XAI to the auditor’s assessment and response process, including the risks of material misstatement due to fraud under ISA 240 (Revised), rather than claiming a generic theory of human–AI governance. The framework provides a structured research agenda for behavioral, archival, simulation, and field validation.
Keywords:
financial statement fraud
; explainable artificial intelligence
; human–AI decision making
; audit judgment
; professional skepticism
; decision utility
; AI governance
1. Introduction
Financial statement fraud remains a persistent challenge for auditors, regulators, investors, and other users of corporate reporting. Fraudulent reporting is comparatively rare, strategically concealed, and often embedded within otherwise legitimate accounting information. These characteristics make fraud detection a high-stakes, asymmetric decision problem: missing a material fraud can create substantial economic and professional consequences, while investigating false positives consumes scarce audit resources.
Quantitative fraud assessment has consequently developed through several methodological traditions. Early work relied on financial ratios, red-flag indicators, and conventional statistical classification. Beneish (1999) showed that accounting variables could be combined into a screening model for earnings manipulation, while Dechow et al. (2011) developed a probabilistic F-Score for material accounting misstatements. Perols (2011) later demonstrated that fraud-detection performance varies across statistical and machine-learning algorithms and is sensitive to the assumed fraud prevalence and misclassification costs.
This evolution is also visible in emerging-market research. Jadambaa et al. (2018) developed a logistic-regression model using Mongolian company financial statements and subsequently tested the model on a separate set of companies reviewed by public authorities. More recent work has moved from conventional classification toward integrated machine learning and XAI. Sodnomdavaa and Lkhagvadorj (2026) combined multiple machine-learning models with SHAP, LIME, counterfactual explanations, calibration, Decision Curve Analysis (DCA), and audit-cost simulation. That framework represented a shift from asking only whether a model can detect fraud toward asking whether model use improves decision utility.
However, decision-centric model evaluation still does not constitute a complete audit decision architecture. A calibrated risk probability accompanied by a plausible explanation does not, by itself, determine whether an auditor should expand testing, alter the nature or timing of procedures, seek additional corroborating evidence, involve specialists, or escalate concerns to engagement leadership or those charged with governance. Nor does it determine how disagreement between the model and the auditor should be documented or how subsequent outcomes should feed back into model use.
The distinction is increasingly important because explainability itself is not equivalent to explanation reliability. Zhang et al. (2022) established the relevance of XAI to audit documentation and evidence, but a 2026 systematic review by Zafar and Wu identified an evaluation vacuum in fraud-XAI research: approximately 80% of the reviewed studies used predictive performance as a proxy for explanation quality rather than evaluating explanations directly. The review also documented an explainability–imbalance paradox, under which class-imbalance treatments may distort post-hoc explanations. These findings weaken the assumption that a technically interpretable output is automatically suitable for professional reliance.
Human factors create a second unresolved problem. Decision makers may over-rely on automation or, conversely, reject algorithmic recommendations after observing errors. Parasuraman and Riley (1997) distinguish misuse and disuse of automation, while Dietvorst et al. (2015) document algorithm aversion. In auditing, professional skepticism creates an additional obligation to challenge evidence and maintain a questioning mindset rather than defer automatically to either human intuition or machine output. Field evidence further indicates that audit firms view transparency, reliability, bias, auditor overreliance, and insufficient guidance as material barriers to advanced AI adoption (Kokina et al., 2025).
Recent 2026 conceptual work raises the novelty threshold for any additional human–AI governance framework. Kou et al. (2026) reframe finance as a set of human–AI hybrid decision systems centered on delegation, reliance, decision-useful XAI, oversight, and feedback. Engin (2026) conceptualizes human–AI governance through decision authority, process autonomy, and accountability configuration, while Kálmán (2026) translates EU AI Act governance objectives into audit quality-management controls and conditions for evidential reliance. The present study therefore does not claim human oversight, accountability, or lifecycle governance as generic innovations. Its narrower contribution is to connect those concerns to the specific decision structure of financial statement fraud auditing: model-generated fraud risk, explanation reliability, professional skepticism, asymmetric error consequences, constrained audit resources, graded audit response, and feedback under fraud-related auditing responsibilities.
The regulatory context reinforces the importance of this problem. ISA 240 (Revised), issued by the International Auditing and Assurance Standards Board (IAASB) in 2025 and effective for periods beginning on or after 15 December 2026, clarifies auditor responsibilities, strengthens the fraud lens in risk assessment and responses, and reinforces professional skepticism. These requirements define professional obligations but do not prescribe a complete architecture for translating AI-generated fraud signals and explanations into accountable audit actions.
This article therefore addresses the prediction–decision–governance gap by developing a Human–AI Decision Governance Framework for explainable financial statement fraud risk assessment. Rather than proposing another fraud classifier, the study shifts the unit of analysis from model output to the process through which AI-generated evidence is evaluated, challenged, combined with professional judgment, translated into audit responses, and subjected to governance and feedback.
Research questions. The article addresses three questions:
- RQ1. How has financial statement fraud assessment evolved from statistical screening toward explainable and decision-oriented artificial intelligence?
- RQ2. Under what conditions should AI-generated fraud-risk predictions and explanations influence professional audit judgment?
- RQ3. How can prediction, explanation assurance, professional skepticism, decision utility, audit response, and governance be integrated into an accountable human–AI fraud-risk assessment framework?
Contributions. The article makes four domain-specific contributions. First, it separates predictive validity, explanatory validity, professional validity, and decision validity rather than treating model performance as a single notion of usefulness. Second, it positions explanation assurance as an intermediate evaluation layer between XAI output and auditor reliance. Third, it links asymmetric consequences and constrained audit resources to graded audit responses rather than stopping at model-level decision utility. Fourth, it integrates professional skepticism, documented human override, and lifecycle feedback into a fraud-audit decision architecture that is explicitly narrower than general human–AI governance frameworks.
2. From Fraud Detection to AI-Supported Audit Decisions
2.1. Statistical Foundations
Financial statement fraud detection was initially framed as a screening and classification problem. Ratio-based and statistical approaches sought systematic differences between fraudulent and non-fraudulent reporting using observable accounting information. Beneish (1999) demonstrated that manipulation-related accounting characteristics could be assembled into a screening score, while Dechow et al. (2011) linked a broader set of financial and non-financial indicators to the probability of material misstatement.
In Mongolia, Jadambaa et al. (2018) analyzed 370 financial statements using logit, probit, and linear probability models, with the reported logit specification developed from 13 financial variables. The authors then evaluated the model on financial statements of 188 companies reviewed by the Ministry of Finance and the General Department of Taxation. The study is important here not because its statistical architecture remains state of the art, but because it represents an early formalization of fraud screening in a data-constrained emerging-market setting.
Statistical approaches remain valuable when interpretability, small samples, or limited infrastructure constrain more complex models. Their principal limitation is not that they are obsolete, but that fixed functional forms may be less able to capture nonlinear interactions and high-dimensional relationships. Moreover, fraud is typically a rare-event problem; therefore, overall accuracy can conceal weak fraud sensitivity. Perols (2011) showed that comparisons among statistical and machine-learning approaches are materially affected by class proportions and the relative cost of false positives and false negatives.
2.2. Machine Learning and Predictive Flexibility
Machine learning expanded fraud detection by allowing nonlinear relationships and flexible interactions to be learned from financial information. Bao et al. (2020) showed that theory-motivated raw accounting numbers combined with machine learning can improve accounting-fraud prediction relative to established benchmarks. This result is conceptually important because it treats domain theory and machine learning as complements rather than substitutes.
The broader transition toward random forests, boosting, support vector machines, neural networks, ensembles, and related architectures increased predictive flexibility but also increased opacity. In auditing, opacity is consequential because the user must evaluate and document the basis for evidence rather than merely consume a score. This created a transparency problem alongside the accuracy problem.
2.3. Explainable Artificial Intelligence
XAI developed partly in response to the opacity of complex models. Zhang et al. (2022) introduced common XAI techniques to auditing and discussed their potential relevance to audit evidence and documentation. Such methods can identify influential variables, provide local explanations, or show how a prediction might change under counterfactual conditions.
Yet explanation availability and explanation reliability are distinct. Zafar and Wu (2026) synthesize 49 peer-reviewed studies from 2021–2025 and identify two methodological problems that are highly relevant to fraud analytics: the possibility that class-imbalance treatments alter the distributions on which post-hoc explanations depend, and a widespread failure to assess explanation quality directly. Consequently, an intuitive SHAP or LIME output can still be unstable, method-dependent, or insufficiently faithful to the underlying model.
For audit use, the relevant question is therefore not simply whether a model can be explained, but whether an explanation is sufficiently reliable for professional reliance. This article uses the term explanation assurance for the structured evaluation of that reliability. The term denotes a functional layer in the proposed architecture rather than a claim that a new universal measurement construct has already been established.
2.4. From Predictive Accuracy to Decision Utility
Prediction models are commonly evaluated using discrimination and classification metrics, yet such metrics do not fully represent the consequences of acting on a prediction. Decision Curve Analysis was developed precisely to evaluate the net benefit of prediction models across threshold probabilities rather than relying only on accuracy-based measures (Vickers & Elkin, 2006).
Sodnomdavaa and Lkhagvadorj (2026) extended financial statement fraud detection in this direction by combining ML, XAI, calibration, DCA, and audit-cost simulation. The framework demonstrated how model evaluation can move from classification performance toward interpretability and practical decision utility. However, model-level decision utility still leaves open the professional process through which an auditor evaluates the AI signal, reconciles conflicting evidence, determines response intensity, and remains accountable for the final decision.
2.5. Human Judgment, Reliance, and AI Governance
AI adoption in auditing is increasingly recognized as a socio-technical issue. Issa et al. (2016) anticipated workforce supplementation rather than simple human replacement, while Kokina and Davenport (2017) described the changing role of automation in audit work. Fedyk et al. (2022) provide empirical evidence that audit-firm AI investment is associated with improved audit outcomes, but the benefits of AI do not eliminate the need for human capabilities and organizational governance.
Ethical and accountability concerns further complicate adoption. Munoko et al. (2020) identify unintended consequences and ethical issues associated with AI in auditing. Kokina et al. (2025), based on interviews with experienced audit professionals, report concerns about explainability, bias, privacy, robustness, reliability, overreliance, and the need for guidance. Li and Goel (2025) likewise argue that AI auditability depends on technical, process, governance, and competency dimensions rather than model performance alone.
Human reliance is itself unstable. Overreliance can convert automation into a heuristic replacement for vigilant evidence processing (Parasuraman & Riley, 1997; Skitka et al., 2000). Underreliance is also costly: Dietvorst et al. (2015) show that users may avoid algorithms after observing them make errors even when algorithmic performance remains superior on average. Audit design must therefore protect against both excessive acceptance and unjustified rejection of AI outputs.
The adjacent literature has moved quickly toward broader decision-system and governance perspectives. Kou et al. (2026) argue that the relevant unit in finance is increasingly the human–AI hybrid decision system rather than the isolated model. Engin (2026) similarly emphasizes the allocation of authority, autonomy, and accountability across human–AI relationships. In audit practice, Kálmán (2026) develops a governance-to-controls framework linking AI governance requirements to quality management, evidence, documentation, and reviewer challenge. These contributions substantially narrow the remaining gap. What remains insufficiently specified for financial statement fraud assessment is a domain-specific chain linking fraud-risk signals and explanation reliability to professional skepticism, asymmetric audit consequences, resource-sensitive response intensity, and post-decision feedback.
2.6. An Analytical Typology of the Methodological Progression
Table 1 organizes the literature into six analytical stages. These stages are not intended as a strict chronology or as mutually exclusive historical eras. They are a problem-oriented typology: each stage addresses a limitation of the previous methodological focus while introducing a new unresolved problem.
3. Theoretical Foundations
3.1. Decision Theory: From Probability to Action
The first pillar is decision theory. Fraud models generally produce a score, probability, or classification, but a probability does not determine an optimal action without information about consequences. This is particularly important where false negatives and false positives have asymmetric costs and where the cost of additional audit procedures varies across cases. The equations below are an illustrative normative formalization rather than an empirically estimated audit-utility function.
Let pᵢ denote the estimated probability that financial statement i involves fraudulent reporting, conditional on information Xᵢ:
pᵢ = P(Fᵢ = 1 | Xᵢ)
A conventional classifier applies a threshold τ and converts pᵢ into a binary prediction. For professional audit use, however, selecting τ solely to maximize statistical performance is conceptually insufficient. The preferred action should depend on the expected consequences of the available response options.
EUᵢⱼ = pᵢ U(Aⱼ, F=1) + (1 − pᵢ) U(Aⱼ, F=0) − C(Aⱼ)
Aᵢ* = arg max₍Aⱼ₎ EUᵢⱼ
This logic is consistent with consequence-sensitive prediction evaluation and with the motivation underlying DCA (Vickers & Elkin, 2006). The theoretical implication is simple but consequential: the best prediction is not necessarily the best audit decision.
3.2. Professional Judgment and Professional Skepticism
The second pillar is professional judgment, with professional skepticism functioning as a critical control on reliance. Hurtt (2010) conceptualizes professional skepticism as a multidimensional characteristic, while Nelson (2009) and Hurtt et al. (2013) emphasize its importance to skeptical judgment and skeptical action. In an AI-enabled setting, skepticism should operate not only toward management evidence but also toward the reliability, relevance, and limits of algorithmic evidence.
ISA 240 (Revised) reinforces a questioning mindset, consideration of contradictory evidence, and the appropriate design of responses to fraud risk. Accordingly, AI outputs are better understood as evidence inputs that require evaluation rather than as autonomous conclusions.
AI signal + explanation + corroborating evidence professional skepticism professional judgment
Professional skepticism should therefore reduce two opposing failure modes: automatic acceptance of apparently precise AI outputs and reflexive rejection of algorithmic evidence merely because the system is imperfect.
3.3. Socio-Technical Governance and Accountability
The third pillar is socio-technical governance. AI-supported audit decisions emerge from interactions among models, data, explanations, auditors, organizational procedures, standards, and accountability structures. Alkhatib et al. (2026), reviewing 236 AI-and-auditing publications, explicitly frame the field through socio-technical systems theory and identify the need to optimize technical and social subsystems jointly.
This perspective is consistent with the NIST AI Risk Management Framework (AI RMF 1.0), which organizes AI risk management around the Govern, Map, Measure, and Manage functions and identifies validity, reliability, accountability, transparency, explainability, interpretability, privacy, safety, security, resilience, and fairness among trustworthiness characteristics (Tabassi, 2023). Li and Goel (2025) similarly emphasize that AI auditability requires attention to data, system processes, governance, documentation, and multidisciplinary capability.
Applied to financial statement fraud assessment, governance must define who may rely on the model, what validation is required, how human override operates, what documentation is retained, how model and explanation performance are monitored, and how downstream audit outcomes inform continued use. Accountability is therefore an attribute of the decision system, not of the model in isolation.
4. Conceptual Research Design
4.1. Research Approach
This study adopts a conceptual theory-building design rather than an empirical research design. Conceptual research is appropriate when the objective is to integrate partially disconnected literatures, clarify relationships among constructs, and develop propositions that can subsequently be tested. Following Jaakkola (2020), the study combines theory synthesis with conceptual model development.
The article neither estimates a new fraud classifier nor reanalyzes the datasets used in earlier work. Instead, it integrates insights from four domains: financial statement fraud detection, explainable AI, professional audit judgment and skepticism, and AI governance. The unit of analysis is therefore extended from X → P(Fraud) to AI evidence → professional evaluation → audit action → governance.
AI-assisted tools were used during manuscript preparation for language drafting and editing, structural refinement, and preparation/refinement of the conceptual framework figure. The author reviewed and verified the final text, references, equations, and figure and assumes full responsibility for the manuscript.
4.2. Literature-Selection Logic
The synthesis is problem-driven rather than a systematic literature review. Sources were selected purposively for one of four reasons: foundational influence on financial-statement fraud screening; methodological relevance to ML, XAI, calibration, or consequence-sensitive prediction; direct relevance to professional judgment, skepticism, reliance, or audit quality; or authoritative relevance to AI governance and fraud-related auditing requirements.
This design choice avoids an inappropriate PRISMA claim. The purpose is not to estimate the frequency of themes in a defined literature population but to identify mechanisms needed to explain how AI-generated fraud signals can become professionally defensible audit actions. Recent systematic reviews are used as evidence about the state of the literature, not as substitutes for the present conceptual contribution.
4.3. Framework-Development Logic
Framework development proceeded in three analytical steps. First, the literature was decomposed into the six problem-oriented stages shown in Table 1. Second, the unresolved prediction–decision gap was represented as five transitions: prediction → explanation → judgment → action → governance. Third, decision theory, professional skepticism, and socio-technical governance were integrated to specify the mechanisms that govern those transitions.
The resulting architecture contains six interdependent layers: fraud-risk signal, explanation assurance, professional judgment, decision utility, audit response, and governance with feedback. The framework is designed to be model-agnostic: it can accommodate statistical models, tree ensembles, neural networks, language models, or other predictive systems provided that the system supplies risk-relevant information and is used as decision support rather than as an autonomous audit opinion.
4.4. Boundary Conditions
The framework is most relevant when fraud events are relatively rare, investigation resources are constrained, AI outputs influence the allocation or intensity of audit procedures, and the consequences of false-positive and false-negative decisions are asymmetric. It is not intended to imply that every engagement requires AI or that higher model complexity necessarily improves audit evidence.
The framework also assumes meaningful human accountability. Where a system is used merely for administrative automation, or where AI output cannot materially affect professional judgment or audit response, several layers may be unnecessary. Conversely, where AI materially influences risk assessment, testing intensity, or escalation, explanation, documentation, and governance requirements become more important.
Figure 1.
Human–AI Fraud Decision Governance Framework. The vertical sequence represents decision transformation; the feedback path represents lifecycle learning and governance.
Figure 1.
Human–AI Fraud Decision Governance Framework. The vertical sequence represents decision transformation; the feedback path represents lifecycle learning and governance.

5. Proposed Human–AI Decision Governance Framework
5.1. Layer 1: Fraud-Risk Signal
The first layer is the generation of a fraud-risk signal. Contemporary systems may produce a binary label, ranking, anomaly score, or probability estimate. For audit use, a calibrated probability or ranked risk score is generally more informative than a hard classification because it preserves information about uncertainty and supports threshold-based decision analysis.
The framework treats pᵢ as a risk signal rather than a professional conclusion. The estimate is conditional on model specification, training data, class balance, feature construction, and the context in which the model was developed. Strong discrimination does not guarantee calibration, transportability, or robustness under changing economic conditions.
This distinction is also necessary for audit-standard precision. A model label for 'fraudulent financial statement' is not equivalent to the auditor’s assessed risk of material misstatement due to fraud under ISA 240 (Revised). The former is an algorithmic output defined by the model’s labels and data-generating process; the latter is a professional risk assessment that must incorporate the engagement context and sufficient appropriate audit evidence.
pᵢ ≠ fraud conclusion and pᵢ ≠ audit action
5.2. Layer 2: Explanation Assurance
The second layer evaluates whether an explanation is sufficiently reliable to enter professional judgment. Explanation assurance is defined here as a structured evaluation function rather than as a single score that is assumed to exist across all model classes.
EAᵢ = f(Fidelityᵢ, Stabilityᵢ, Robustnessᵢ, Consistencyᵢ, Contextual Plausibilityᵢ)
Fidelity concerns whether an explanation reflects the behavior of the underlying model. Stability concerns whether materially similar cases generate materially similar explanations. Robustness concerns sensitivity to perturbations, resampling, and alternative specifications. Consistency concerns coherence across repeated analytical settings. Contextual plausibility concerns whether the explanation is economically and accounting-wise defensible in the engagement context.
No single XAI method is assumed to be authoritative. SHAP, LIME, counterfactual explanations, model-specific importance, and other methods provide different evidence and have different vulnerabilities. Explanation assurance therefore acts as a gate between technical interpretability and professional reliance.
5.3. Layer 3: Professional Judgment and Skepticism
The third layer converts AI-generated information into professional evaluation. Identical model outputs can justifiably produce different audit judgments because engagement context, materiality, corroborating evidence, internal-control conditions, and management behavior differ across cases.
Jᵢ = g(pᵢ, EAᵢ, Eᵢ, Mᵢ, Cᵢ, Sᵢ)
Here, Eᵢ denotes corroborating evidence, Mᵢ materiality considerations, Cᵢ engagement context, and Sᵢ professional skepticism. When AI output, explanation, and audit evidence converge, the signal may justify a stronger response. When they conflict, the conflict itself should trigger further investigation rather than automatic deference to either the model or prior human expectations.
5.4. Layer 4: Decision Utility and Resource Constraints
The fourth layer recognizes that fraud-risk assessment is also a resource-allocation problem. Audit teams face constraints in hours, specialist capacity, deadlines, and investigation budgets. Consequently, the threshold for additional work should reflect both probability and consequence.
EUᵢⱼ = pᵢ U(Aⱼ, F=1) + (1 − pᵢ) U(Aⱼ, F=0) − C(Aⱼ)
max Σᵢ Σⱼ xᵢⱼ EUᵢⱼ subject to Σᵢ Σⱼ xᵢⱼ cᵢⱼ ≤ B
Σⱼ xᵢⱼ ≤ 1 for each i, xᵢⱼ ∈ {0,1}
Here, Aⱼ is interpreted as a mutually exclusive audit-response package. If individual procedures are modeled as cumulative rather than mutually exclusive, the decision variables and constraints should be reformulated accordingly.
The implication is that prediction ranking and optimal audit allocation need not coincide. A lower-probability case may justify a stronger response if the consequence of a missed fraud is materially higher or if the relevant procedure is relatively inexpensive. Conversely, a high-risk case may not warrant the most resource-intensive response if incremental procedures offer little additional information.
5.5. Layer 5: Audit Response
The fifth layer converts evaluated risk into an operational response. The framework does not prescribe a universal probability-to-procedure mapping. Instead, it proposes graded response intensity, consistent with the principle that fraud-risk assessment should affect the nature, timing, and extent of procedures after professional evaluation.
Table 2.
Illustrative mapping from evaluated fraud risk to audit response. The table is conceptual, not a prescriptive threshold schedule.
Table 2.
Illustrative mapping from evaluated fraud risk to audit response. The table is conceptual, not a prescriptive threshold schedule.
| Decision condition | Illustrative response | Governance emphasis |
|---|---|---|
| Low evaluated risk; stable evidence | No material modification; routine procedures | Retain model and review record |
| Moderate risk or localized concern | Additional analytical or targeted procedures | Document rationale and evidence used |
| Elevated risk with plausible indicators | Expanded substantive testing; targeted journal-entry work | Review explanation and response proportionality |
| High risk with conflicting evidence | Additional corroboration; specialist involvement | Explicit disagreement/override record |
| High risk with strong corroboration | Escalation to engagement leadership / those charged with governance | Enhanced documentation and communication |
5.6. Layer 6: Governance, Documentation, and Feedback
The final layer establishes traceability and continuous control. AI-assisted audit decisions require governance records that conventional standalone classifiers do not necessarily provide. At minimum, the engagement should be able to reconstruct the model version, inputs, output, explanation method, reviewer, decision, override, and relevant downstream outcome.
Gᵢ = {Model, Version, Inputs, Prediction, Explanation, Reviewer, Decision, Override, Outcome}
Human override should remain possible, but it should not be unstructured. Unlimited model authority creates automation risk; unrestricted human override can reintroduce inconsistency, bias, and hindsight rationalization. The preferred design is controlled human authority with documented AI use.
Feedback completes the lifecycle. Subsequent audit findings can be used to evaluate calibration, threshold adequacy, explanation stability, false-positive frequency, missed-risk consequences, override patterns, and model drift. Governance is therefore continuous rather than a one-time model approval.
Prediction → Explanation → Judgment → Action → Outcome → Monitoring → Revision
5.7. Construct Definitions
Table 3.
Construct definitions and candidate operationalizations for future empirical work.
| Construct | Definition in this framework | Primary failure mode | Illustrative future measure |
|---|---|---|---|
| Fraud-risk signal | Model-generated probability, ranking, or anomaly information relevant to fraudulent reporting | Miscalibration; drift; poor transportability | Calibration error; PR-AUC; recall at fixed review capacity |
| Explanation assurance | Evaluation of the reliability of model explanations before professional reliance | Unfaithful or unstable explanations | Fidelity, stability, perturbation sensitivity, agreement across methods |
| Professional judgment | Auditor integration of AI evidence with engagement evidence, materiality, and context | Automation bias; algorithm aversion; confirmation bias | Reliance choice; additional evidence search; judgment accuracy |
| Decision utility | Net value of an audit action given probabilities, consequences, costs, and constraints | Accuracy-optimized but economically inefficient thresholds | Net benefit; expected cost; cases detected per review-hour |
| Audit response | Nature, timing, extent, or escalation of audit procedures | Over-audit; under-response; inconsistent procedure mapping | Procedure intensity; incremental evidence; escalation rate |
| Governance & feedback | Traceability, accountability, monitoring, and learning over the AI lifecycle | Unaccountable override; undetected drift; undocumented reliance | Override logs; drift metrics; monitoring frequency; outcome feedback |
6. Theoretical Propositions
P1 — Predictive Performance and Decision Utility
Predictive performance is positively but imperfectly associated with audit decision utility; therefore, higher predictive performance alone is insufficient to establish superior audit usefulness.
A model can discriminate well yet be poorly calibrated, expensive to act on, or optimized at a threshold that does not reflect audit consequences.
P2 — Explanation Assurance
Conditional on adequate predictive validity, the positive effect of AI-generated fraud-risk signals on audit decision quality increases as explanation fidelity, stability, robustness, and contextual plausibility increase.
Reliable explanations should help auditors distinguish useful risk signals from technically plausible but unstable explanations.
P3 — Professional Skepticism
Professional skepticism moderates the relationship between AI explanations and auditor reliance by reducing both inappropriate algorithmic acceptance and unjustified algorithmic rejection.
Skepticism should produce calibrated reliance rather than simple resistance to or acceptance of AI.
P4 — Resource-Sensitive Thresholds
Fraud-risk thresholds incorporating asymmetric error consequences and resource constraints generate greater audit decision utility than thresholds based solely on predictive classification performance.
Decision thresholds should reflect the value and cost of downstream action, not only model metrics.
P5 — Human Override and Accountability
Human override improves the accountability of AI-assisted fraud assessment when override decisions are explicitly justified, documented, and traceable.
Override is beneficial when governed; unstructured override can reduce consistency and recreate bias.
P6 — Lifecycle Governance
AI-assisted fraud-risk assessment becomes more effective and sustainable when governance jointly monitors predictive performance, explanation reliability, professional overrides, audit outcomes, and model drift.
Monitoring only prediction metrics is insufficient because downstream human and organizational behavior can change decision quality.
Table 4.
Summary of theoretical propositions and suggested validation approaches.
| Proposition | Focal relationship | Expected mechanism | Primary validation approach |
|---|---|---|---|
| P1 | Predictive performance → decision utility | Positive but non-equivalent | Archival or simulation comparison of models/thresholds |
| P2 | Explanation assurance × AI signal → decision quality, conditional on predictive validity | Higher assurance strengthens useful reliance when the underlying signal is adequately valid | Controlled auditor experiment |
| P3 | Professional skepticism × explanation → reliance | Reduces over- and under-reliance | Behavioral experiment with contradictory evidence |
| P4 | Resource-sensitive threshold → utility | Higher expected/net benefit | Decision simulation or field trial |
| P5 | Documented human override → accountability | Higher traceability and consistency | Experiment or audit-file study |
| P6 | Lifecycle governance → sustained effectiveness | Reduces degradation from drift and behavior | Longitudinal field study |
7. Discussion
The proposed framework reframes financial statement fraud detection as a domain-specific professional decision-governance problem rather than a purely predictive modeling problem. This direction is consistent with recent work that shifts finance and auditing from model-centered evaluation toward human–AI decision systems and governance (Engin, 2026; Kálmán, 2026; Kou et al., 2026). The incremental contribution here is narrower: the framework specifies the pathway from a financial statement fraud-risk signal to an audit response by combining explanation assurance, professional skepticism, asymmetric decision consequences, constrained audit resources, documented override, and lifecycle feedback.
This reframing requires four forms of validity. Predictive validity concerns whether a model supplies useful fraud-risk information. Explanatory validity concerns whether the stated reasons for an output are sufficiently faithful and stable. Professional validity concerns whether model information is consistent with corroborating evidence, materiality, and skeptical evaluation. Decision validity concerns whether the selected response appropriately balances detection benefit, resource cost, and professional responsibility. Technical sophistication can therefore coexist with professional inadequacy if the system fails at any one of these layers.
The framework also changes the unit of analysis. Much fraud research evaluates Model → Prediction. The present approach evaluates Model → Explanation → Auditor → Decision → Governance. This broader unit reflects the reality that audit outcomes are produced jointly by technology, professional judgment, organizational procedures, and standards.
A second implication concerns XAI. Explainability should not be treated as the endpoint of trustworthy AI. An explanation creates additional evidence that itself requires evaluation. The explanation-assurance layer therefore prevents the common conceptual shortcut in which the presence of SHAP or LIME is interpreted as sufficient evidence of transparency or trustworthiness.
A third implication concerns human control. Human involvement is neither automatically protective nor automatically harmful. Automation bias can weaken independent challenge, while algorithm aversion can suppress legitimate predictive information. The appropriate design goal is calibrated reliance supported by professional skepticism, documented override, and traceable governance.
Finally, the framework clarifies the role of standards and governance. ISA 240 (Revised) strengthens fraud-focused risk assessment and professional skepticism, while the NIST AI RMF provides a lifecycle-oriented governance vocabulary. Neither source is a substitute for professional judgment. Their value in this framework is to constrain and structure how AI evidence is used, monitored, and documented.
8. Implications
8.1. Implications for Research
For researchers, the framework suggests that model performance should be treated as one component of a larger causal chain. Future studies can test whether explanation reliability changes evidence search, whether skepticism changes reliance, whether resource-sensitive thresholds improve net audit benefit, and whether governance mechanisms reduce harmful override or drift.
The framework also provides a route to cumulative research. Technical studies can supply calibrated risk and explanation measures; behavioral studies can test reliance and skepticism; audit analytics studies can evaluate decision utility; and field studies can examine governance and long-run outcomes. These streams can be linked through common constructs rather than remaining isolated.
8.2. Implications for Auditors and Audit Firms
For practitioners, AI outputs should be treated as structured evidence inputs rather than deterministic conclusions. Audit methodology should define acceptable use cases, validation expectations, explanation review, escalation criteria, override procedures, documentation, and post-engagement monitoring.
Audit firms should also distinguish high-confidence prediction from high-confidence professional conclusion. A technically confident model can still be operating outside its valid context, while a well-explained signal may still be contradicted by stronger engagement evidence. Training therefore needs to include both technical literacy and human-factors risks such as overreliance and algorithm aversion.
Governance responsibility should not sit solely with data-science teams. Audit methodology, quality management, engagement leadership, information security, legal/compliance functions, and technical specialists may all have roles depending on the system's impact.
8.3. Implications for Regulators and Standard Setters
Principles-based guidance could become more explicit about minimum documentation of AI-assisted judgments, expectations for validation and explanation reliability, treatment of disagreement between human and AI assessments, retention of AI-generated evidence, and monitoring of material model changes. Technology-specific rules would risk rapid obsolescence; principles governing evidence, accountability, proportionality, and traceability are more durable.
ISA 240 (Revised) provides a relevant institutional anchor because it strengthens fraud-focused risk assessment, responses, communication, and professional skepticism. The proposed framework translates those responsibilities into questions that can be asked of AI-supported decisions without claiming that the standard itself mandates a particular AI architecture.
8.4. Implications for AI Developers
Developers of audit-oriented fraud systems should optimize for professional usability rather than predictive metrics alone. Relevant design objectives include calibration, uncertainty communication, explanation stability, traceability, configurable thresholds, audit-trail generation, model/version control, and drift monitoring.
Systems should also distinguish prediction confidence from evidence quality. These properties are conceptually different: a model can be confident for the wrong reasons, while a cautious model can still supply valuable directional evidence.
9. Boundary Conditions, Limitations, and Future Validation
The study is conceptual and has not yet empirically validated the proposed relationships. Its value therefore depends on whether the framework generates falsifiable and practically meaningful tests rather than on claims of observed causal effects.
The framework does not prescribe a universal fraud-risk threshold. Thresholds should vary with engagement characteristics, fraud prevalence, materiality, industry, regulatory context, cost of investigation, and the consequences of missed fraud. Likewise, explanation assurance may require different measures across model classes and explanation techniques.
The architecture assumes that AI supports rather than autonomously determines the audit opinion. If regulatory or professional environments permit materially greater autonomy, responsibility allocation and evidence standards would require additional analysis. Governance requirements may also differ across jurisdictions and audit-firm structures.
Future research should therefore move deliberately from conceptual propositions to causal and field validation. Table 5 outlines a practical research agenda.
10. Conclusions
Financial statement fraud detection has progressed from ratio-based screening and statistical classification toward machine learning, XAI, and decision-oriented analytics. In parallel, recent finance and auditing research has begun to articulate broader human–AI decision-system and governance frameworks. This article contributes a narrower fraud-audit architecture focused on the transformation of model-generated risk signals into professionally evaluated and accountable audit responses.
The Human–AI Decision Governance Framework integrates six components: fraud-risk signaling, explanation assurance, professional judgment, decision utility, audit response, and governance with feedback. It distinguishes prediction from decision, interpretability from explanation reliability, and human involvement from accountable human oversight.
The core principle is that AI prediction is an input to judgment, not a substitute for judgment. Useful AI-assisted fraud assessment therefore requires more than accurate models. It requires explanations that can be evaluated, auditors who can challenge algorithmic evidence, thresholds that reflect consequences and resource constraints, responses that remain proportionate to risk, and governance mechanisms that preserve traceability and learning over time.
By shifting attention from fraud classification toward accountable human–AI decision making, the article provides a conceptual bridge among fraud analytics, XAI, professional auditing, and AI governance. The six propositions convert that bridge into a research program that can be tested through behavioral experiments, archival analyses, simulations, and longitudinal field studies.
Author Contributions
Conceptualization, framework development, literature synthesis, writing—original draft, and writing—review and editing: T.A. The sole author reviewed and approved the submitted version.
Funding
This research received no external funding.
Data Availability Statement
No new empirical data were created or analyzed in this conceptual study.
Use of AI-Assisted Tools
AI-assisted tools were used during manuscript preparation for language drafting and editing, structural refinement, and preparation/refinement of the conceptual framework figure. Tamir Ariunsukh, the sole author, reviewed and verified the final text, references, equations, and figure and accepts full responsibility for the submitted content.
Conflicts of Interest
The author declares no conflicts of interest.
References
- Alkhatib, E.; Alkhatib, A.; Jarvis, R. Auditing using artificial intelligence: A systematic literature review with scientometric and topic modeling insights through a socio-technical systems (STS) perspective. Journal of Financial Reporting and Accounting 2026. [Google Scholar] [CrossRef]
- Bao, Y.; Ke, B.; Li, B.; Yu, Y. J.; Zhang, J. Detecting accounting fraud in publicly traded U.S. firms using a machine learning approach. Journal of Accounting Research 2020, 58(1), 199–235. [Google Scholar] [CrossRef]
- Beneish, M. D. The detection of earnings manipulation. Financial Analysts Journal 1999, 55(5), 24–36. [Google Scholar] [CrossRef]
- Dechow, P. M.; Ge, W.; Larson, C. R.; Sloan, R. G. Predicting material accounting misstatements. Contemporary Accounting Research 2011, 28(1), 17–82. [Google Scholar] [CrossRef]
- Dietvorst, B. J.; Simmons, J. P.; Massey, C. Algorithm aversion: People erroneously avoid algorithms after seeing them err. Journal of Experimental Psychology: General 2015, 144(1), 114–126. [Google Scholar] [CrossRef] [PubMed]
- Engin, Z. Human-AI governance (HAIG): A trust-utility approach. Journal of Responsible Technology 2026, 26, 100167. [Google Scholar] [CrossRef]
- Fedyk, A.; Hodson, J.; Khimich, N.; Fedyk, T. Is artificial intelligence improving the audit process? Review of Accounting Studies. 2022, 27, 938–985. [Google Scholar] [CrossRef]
- Hurtt, R. K. Development of a scale to measure professional skepticism. AUDITING: A Journal of Practice & Theory 2010, 29(1), 149–171. [Google Scholar] [CrossRef]
- Hurtt, R. K.; Brown-Liburd, H.; Earley, C. E.; Krishnamoorthy, G. Research on auditor professional skepticism: Literature synthesis and opportunities for future research. AUDITING: A Journal of Practice & Theory 2013, 32 (Supplement 1), 45–97. [Google Scholar] [CrossRef]
- International Auditing and Assurance Standards Board. ISA 240 (Revised), The Auditor’s Responsibilities Relating to Fraud in an Audit of Financial Statements. International Federation of Accountants. 2025. Available online: https://www.iaasb.org/publications/isa-240-revised-auditor-s-responsibilities-relating-fraud-audit-financial-statements.
- Issa, H.; Sun, T.; Vasarhelyi, M. A. Research ideas for artificial intelligence in auditing: The formalization of audit and workforce supplementation. Journal of Emerging Technologies in Accounting 2016, 13(2), 1–20. [Google Scholar] [CrossRef]
- Jaakkola, E. Designing conceptual articles: Four approaches. AMS Review 2020, 10, 18–26. [Google Scholar] [CrossRef]
- Jadambaa, E.; Sodnomdavaa, T.; Purevsukh, N. The probability of detecting false financial statement: Evidence from Mongolian companies. In Sovremennye usloviya vzaimodeystviya nauki i tekhniki: Sbornik statey Mezhdunarodnoy nauchno-prakticheskoy konferentsii; Omega Science, 2018; pp. 69–74. ISBN 978-5-907069-06-0. [Google Scholar]
- Kokina, J.; Davenport, T. H. The emergence of artificial intelligence: How automation is changing auditing. Journal of Emerging Technologies in Accounting 2017, 14(1), 115–122. [Google Scholar] [CrossRef]
- Kokina, J.; Blanchette, S.; Davenport, T. H.; Pachamanova, D. Challenges and opportunities for artificial intelligence in auditing: Evidence from the field. International Journal of Accounting Information Systems 2025, 56, 100734. [Google Scholar] [CrossRef]
- Kou, G.; Li, Y.; Wang, H.; Wang, X. Human–AI hybrid finance: From AI tools to decision systems. Financial Innovation 2026, 12, 125. [Google Scholar] [CrossRef]
- Kálmán, J. From the EU AI Act to audit practice: A governance-to-controls framework for quality management and evidence. Accounting and Auditing 2026, 2(3), 12. [Google Scholar] [CrossRef]
- Li, Y.; Goel, S. Artificial intelligence auditability and auditor readiness for auditing artificial intelligence systems. International Journal of Accounting Information Systems 2025, 56, 100739. [Google Scholar] [CrossRef]
- Munoko, I.; Brown-Liburd, H. L.; Vasarhelyi, M. The ethical implications of using artificial intelligence in auditing. Journal of Business Ethics 2020, 167, 209–234. [Google Scholar] [CrossRef]
- Nelson, M. W. A model and literature review of professional skepticism in auditing. AUDITING: A Journal of Practice & Theory 2009, 28(2), 1–34. [Google Scholar] [CrossRef]
- Parasuraman, R.; Riley, V. Humans and automation: Use, misuse, disuse, abuse. Human Factors 1997, 39(2), 230–253. [Google Scholar] [CrossRef]
- Perols, J. Financial statement fraud detection: An analysis of statistical and machine learning algorithms. AUDITING: A Journal of Practice & Theory 2011, 30(2), 19–50. [Google Scholar] [CrossRef]
- Skitka, L. J.; Mosier, K. L.; Burdick, M.; Rosenblatt, B. Automation bias and errors: Are crews better than individuals? The International Journal of Aviation Psychology 2000, 10(1), 85–97. [Google Scholar] [CrossRef] [PubMed]
- Sodnomdavaa, T.; Lkhagvadorj, G. Financial statement fraud detection through an integrated machine learning and explainable AI framework. Journal of Risk and Financial Management 2026, 19(1), 13. [Google Scholar] [CrossRef]
- Tabassi, E. Artificial Intelligence Risk Management Framework (AI RMF 1.0) (NIST AI 100-1). National Institute of Standards and Technology 2023. [Google Scholar] [CrossRef]
- Vickers, A. J.; Elkin, E. B. Decision curve analysis: A novel method for evaluating prediction models. Medical Decision Making 2006, 26(6), 565–574. [Google Scholar] [CrossRef] [PubMed]
- Zafar, U.; Wu, F. Methodological challenges in explainable AI for fraud detection: A systematic literature review. Artificial Intelligence Review 2026, 59, 115. [Google Scholar] [CrossRef]
- Zhang, C. (A.; Cho, S.; Vasarhelyi, M. Explainable artificial intelligence (XAI) in auditing. International Journal of Accounting Information Systems 2022, 46, 100572. [Google Scholar] [CrossRef]
Table 1.
Analytical typology of financial statement fraud assessment. The stages are problem-oriented and need not be strictly chronological.
Table 1.
Analytical typology of financial statement fraud assessment. The stages are problem-oriented and need not be strictly chronological.
| Stage | Primary objective | Typical methods | Main contribution | Residual limitation |
|---|---|---|---|---|
| I. Rule / red-flag screening | Identify suspicious indicators | Ratios, rules, expert judgment | Transparency and domain alignment | Static; difficult to scale |
| II. Statistical prediction | Estimate fraud probability | Logit, discriminant models, M-/F-Scores | Probabilistic screening | Functional-form and interaction limits |
| III. Machine learning | Improve predictive discrimination | RF, SVM, boosting, neural networks | Nonlinear pattern recognition | Opacity and calibration risk |
| IV. Explainable AI | Interpret model outputs | SHAP, LIME, counterfactuals | Greater transparency | Explanation reliability unresolved |
| V. Decision-centric AI | Evaluate practical net benefit | Calibration, cost-sensitive thresholds, DCA | Links prediction to consequences | Professional judgment architecture incomplete |
| VI. Human–AI decision governance | Convert AI evidence into accountable audit action | Explanation assurance, skepticism, utility, override, monitoring | Integrates technical and professional decision processes | Requires empirical validation |
Table 5.
Empirical research agenda generated by the framework.
| Research question | Suggested design | Manipulation / comparison | Primary outcome | Related proposition |
|---|---|---|---|---|
| Does explanation assurance change auditor reliance? | Controlled experiment | Stable/high-fidelity vs unstable/low-fidelity explanation | Reliance, evidence search, judgment accuracy | P2 |
| Does skepticism reduce over- and under-reliance? | Behavioral experiment | AI recommendation × contradictory evidence × skepticism cue | Verification behavior, override quality | P3 |
| Do consequence-sensitive thresholds improve utility? | Simulation / archival study | F1/AUC-optimal vs expected-utility threshold | Net benefit, review hours, missed fraud cost | P1, P4 |
| Does structured override improve accountability? | Experiment / audit-file study | Documented rationale vs unstructured override | Consistency, traceability, hindsight quality | P5 |
| Does lifecycle monitoring preserve performance? | Longitudinal field study | Governed monitoring vs static deployment | Calibration drift, explanation drift, override patterns | P6 |
| Is Human+AI superior to either alone? | Factorial experiment / field trial | Human-only vs AI-only vs Human+AI | Audit decision quality and efficiency | P1–P6 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.