Preprint
Article

This version is not peer-reviewed.

Professional Judgment and AI Governance in Audit and Sustainability Assurance: Public Evidence from the UK Big Four

Submitted:

22 July 2026

Posted:

22 July 2026

You are already at the latest version

Abstract
Artificial intelligence (AI) is entering audit workflows while sustainability reporting and assurance expand the volume, variety and uncertainty of information subject to professional evaluation. The policy question is whether firms disclose governance arrangements that keep AI-assisted work human-led, reviewable and accountable. This exploratory study analyses the complete cross-section of the 2024 UK transparency reports of Deloitte, EY, KPMG and PwC, coded against seven pre-specified dimensions of AI–judgment governance and aggregated into a transparent, replicable AI–Judgment Governance Disclosure Index (AI-JGDI). The analysis is triangulated with the UK Financial Reporting Council's 2024 inspection results and interpreted against the ISAs, the IESBA Code, the EU Artificial Intelligence Act, the NIST AI Risk Management Framework, CSRD/ESRS, IFRS S1 and S2, and ISSA 5000. All four firms dis-close deployed AI capabilities and explicitly retain human professional responsibility; disclosure is strongest for human oversight, governance ownership and training, and least consistent for AI-specific validation and for explanations that would allow an external reader to reconstruct how an AI output affected an audit judgment. AI-JGDI scores range from 71.4 to 100.0. The study contributes a public-document method, a disclosure index and a governance framework for accountable AI-assisted judgment, and identifies limited explicit integration between AI governance and sustainability-assurance methodology.
Keywords: 
;  ;  ;  ;  ;  ;  ;  

1. Introduction

Financial reporting and assurance are becoming simultaneously more da-ta-intensive and more judgment-intensive. Artificial intelligence can search contracts, identify anomalous transactions, score populations of journal entries, retrieve technical guidance, draft working papers and direct attention toward unusual patterns. Sustainability reporting adds heterogeneous operational, environmental and value-chain data, forward-looking assumptions, double-materiality assessments and in-formation that may not have passed through financial-reporting control systems. These developments enlarge the evidence available to accountants and auditors, but they also increase the need to decide which evidence is relevant, reliable and sufficiently persuasive.
Professional judgment is therefore not a residual activity left after automation. It is the mechanism through which standards, evidence, expertise and ethical responsibility are converted into a defensible conclusion. Accounting judgment research has long shown that performance depends on task knowledge, experience, motivation and the decision environment (Bonner, 1999; Libby & Luft, 1993). Audit research adds professional skepticism: a questioning mind, alertness to contradictory evidence and a critical assessment of audit evidence (Hurtt, 2010; Nelson, 2009). AI changes this environment because a system can structure the evidence presented to the auditor, rank risks, recommend an action or create a first draft that becomes a powerful anchor. The system can improve consistency and coverage, but an apparently precise output can also conceal data limitations, model error, bias or a management-defined objective.
The behavioural evidence is deliberately cautionary. People may avoid algorithms after observing an error (Dietvorst et al., 2015), yet they may also prefer algorithmic advice under other conditions (Logg et al., 2019). In audit settings, reliance depends on task complexity, perceived control and the opportunity to provide input (Commerford et al., 2022, 2024). AI can consequently produce either under-reliance or automation bias. The governance question is how to calibrate reliance: the professional should use the system where it is competent, challenge it where its limits matter, and remain accountable for the conclusion.
The normative environment is also moving quickly. The IESBA technology-related revisions preserve the fundamental principles of integrity, objectivity, professional competence and due care, confidentiality and professional behaviour when technology is used (IESBA, 2023, 2024). ISA 315 (Revised 2019) requires understanding of the entity and its information system and strengthens the exercise of professional skepticism in risk assessment (IAASB, 2019). ISA 540 (Revised) is directly relevant to AI-supported estimates because it emphasises estimation uncertainty, management bias and skeptical evaluation (IAASB, 2018). The EU Artificial Intelligence Act establishes a risk-based regulatory architecture and, for relevant systems, requirements related to governance, technical documentation, logging, transparency, human over-sight, accuracy, robustness and cybersecurity (European Parliament and Council, 2024). The NIST AI Risk Management Framework organises voluntary risk management around Govern, Map, Measure and Manage (NIST, 2023). None of these instruments transfers an auditor's responsibility to an algorithm.
At the same time, sustainability reporting and assurance have moved from voluntary practice toward formal requirements. The Corporate Sustainability Reporting Directive and European Sustainability Reporting Standards require extensive sustainability information and governance disclosures in the European Union (European Commission, 2023; European Parliament and Council, 2022). IFRS S1 and IFRS S2 establish an investor-focused global baseline organised around governance, strategy, risk management, and metrics and targets (ISSB, 2023a, 2023b). ISSA 5000 supplies a glob-al, framework-neutral standard for sustainability assurance engagements (IAASB, 2024). AI may help process these data, but sustainability evidence often contains measurement uncertainty, estimates and value-chain information for which explainability and provenance are indispensable.
Existing research offers important insights but leaves three gaps. First, many studies discuss potential benefits and risks without examining what large audit firms publicly disclose about the governance of AI-assisted judgment. Second, corporate re-ports often combine technology, people, quality and risk narratives; a reproducible coding framework is needed to distinguish a deployed tool from a governance safe-guard. Third, the relationship between AI governance and sustainability assurance is frequently asserted but rarely traced in comparable public documents. Recent work calls for research that examines specific AI configurations and the institutional arrangements surrounding them rather than treating AI adoption as a single binary variable (Kokina et al., 2025; Lehner et al., 2022; Stratopoulos & Wang, 2025).
This study asks three research questions:
RQ1. How do the UK Big Four publicly describe the role of AI in audit work and professional judgment?
RQ2. Which safeguards for validation, documentation, data governance, accountability and learning are disclosed, and how complete are these disclosures across firms?
RQ3. To what extent do the reports connect AI governance with sustainability reporting and assurance?
The empirical setting is the United Kingdom. The Big Four operate under a common regulatory environment, publish transparency reports under comparable requirements and are subject to public inspection by the Financial Reporting Council (FRC). The corpus comprises the complete 2024 transparency-report cross-section for Deloitte, EY, KPMG and PwC UK. A structured content analysis codes seven governance dimensions and constructs an AI–Judgment Governance Disclosure Index (AI-JGDI). The FRC's 2024 inspection results are used only as external context; the study does not infer that a disclosure score causes an inspection outcome.
The paper makes four contributions. It provides public empirical evidence about how leading audit firms frame human–AI responsibility. It introduces a transparent index whose items can be replicated or extended to other jurisdictions and years. It separates disclosure completeness from actual governance effectiveness and from regulatory audit quality. Finally, it develops a practical human–AI judgment protocol relevant to both financial-statement audit and sustainability assurance.

2. Literature Review and Analytical Framework

2.1. Professional Judgment, Discretion and Audit Quality

Professional judgment is a reasoned choice among alternatives under conditions in which rules and evidence do not determine a single answer. It differs from preference because the conclusion must be defensible by reference to applicable requirements, facts, expertise and the objective of the engagement. Judgment is especially important for materiality, risk assessment, accounting estimates, going concern, control evaluation, fraud risk and the sufficiency of evidence. In sustainability assurance, the same logic applies to materiality processes, greenhouse-gas estimates, value-chain information, scenario assumptions and qualitative claims.
Accounting standards create a bounded judgment space. Principles allow economic substance and entity-specific conditions to be reflected, while detailed rules can improve consistency. Neither approach eliminates judgment, and standard precision interacts with preparer incentives, audit-committee strength and auditor challenge (Agoglia et al., 2011; Backof et al., 2016; Bennett et al., 2006; Nelson, 2003). Managerial discretion can communicate private information or can be used opportunistically to select assumptions, timing and disclosures that favour a target (Fields et al., 2001; Healy & Wahlen, 1999). Audit quality therefore depends not only on whether a selected number lies within a plausible range but also on whether evidence was searched neutrally, alternatives were considered and contradictory information was challenged.
Professional skepticism is the behavioural safeguard at this interface. Nelson (2009) distinguishes skeptical judgment from skeptical action: an auditor may recognise risk but still fail to obtain additional evidence or challenge management. AI can support skeptical action by analysing complete populations and surfacing anomalies. It can also weaken it if the system's output becomes a substitute for independent evaluation. Judgment quality consequently contains two dimensions: technical defensibility and process integrity. A technically plausible conclusion produced by a biased or opaque process remains vulnerable.

2.2. AI-Assisted Audit: Augmentation, Reliance and Accountability

AI in audit includes machine learning, natural-language processing, anomaly detection, intelligent search and generative systems. Early research anticipated that automation would shift work from routine procedure execution toward exception analysis and judgment (Abdullah & Almaqtari, 2024; Kokina & Davenport, 2017; Sutton et al., 2016). More recent field evidence shows wider use but also implementation challenges involving data access, integration with methodology, regulation, skills and accountability (Kokina et al., 2025). The potential contribution is strongest when AI and human capabilities are complementary: machines scale search and pattern recognition, whereas professionals frame the problem, assess context, challenge incentives and accept responsibility.
Reliance is not automatically calibrated. Algorithm aversion can lead users to reject a useful model after observing an error (Dietvorst et al., 2015). Algorithm appreciation can produce the opposite tendency (Logg et al., 2019). In complex audit-estimation tasks, the perceived source and competence of an AI system affect reliance (Commerford et al., 2022). Allowing auditors to provide input can increase reliance and perceived control, but greater confidence is beneficial only if it reflects better understanding of system limitations (Commerford et al., 2024). A governance design should therefore support contestability rather than merely encourage adoption.
Ethical risks arise when responsibility is displaced. Training data can reproduce historical bias; model objectives can privilege efficiency over audit quality; generated text can hallucinate; and proprietary systems can be difficult to explain to audit committees or regulators (Lehner et al., 2022; Murikah et al., 2024). Management or engagement teams may also use technological opacity strategically, accepting outputs that support a preferred conclusion and overriding those that do not. This form of algorithmically mediated discretion is particularly problematic because the output appears independent even when objectives, thresholds and data were selected by interested human actors.
Accountable augmentation requires at least seven linked safeguards. First, the use case must be approved for a defined purpose. Second, the system must be tested or validated for that purpose and population. Third, data provenance, confidentiality and permitted use must be controlled. Fourth, users need an explanation adequate to understand why the output is relevant and what uncertainty remains. Fifth, professional review and override must be substantive. Sixth, responsibility and escalation routes must be assigned. Seventh, users require both accounting/audit competence and AI literacy. These safeguards form the coding dimensions used in this study.

2.3. Normative Requirements for AI-Assisted Judgment

No single standard currently governs every use of AI in accounting and audit. The applicable architecture is layered (Table 1). At the engagement level, ISA 200 retains the auditor's responsibility to obtain reasonable assurance and exercise professional judgment and skepticism (IAASB, 2009). ISA 315 (Revised 2019) addresses risk identification, information systems and automated controls. ISA 540 (Revised) requires robust work on accounting estimates and management bias. ISA 220 (Revised) and ISQM 1 allocate engagement and firm-level quality-management responsibilities (IAASB, 2020a, 2020b). These requirements apply regardless of whether evidence or analysis is generated manually or technologically.
Ethical requirements add objectivity, competence, confidentiality and appropriate professional behaviour. The IESBA technology-related revisions explicitly recognise that technology can create threats to compliance with the fundamental principles and that professional accountants must remain alert to information that may be incomplete, biased or misleading (IESBA, 2023). An organisation's approval of a tool does not eliminate the individual professional's responsibility to use it competently and question its output.
General AI-governance instruments supply more specific control concepts. The NIST AI RMF links governance with contextual mapping, measurement and management of risk (NIST, 2023). The EU AI Act uses a risk classification rather than declaring every accounting application high-risk. Nevertheless, its control vocabulary—risk management, data governance, documentation, logging, transparency, human oversight, accuracy, robustness and cybersecurity—provides a useful benchmark for systems that influence consequential professional decisions (European Parliament and Council, 2024). The benchmark is used analytically in this paper, not to claim that every audit tool falls within the Act's high-risk category.
Sustainability standards intensify the need for these controls. CSRD and ESRS broaden the information boundary to impacts, risks and opportunities and require extensive governance and value-chain information (European Commission, 2023; European Parliament and Council, 2022). IFRS S1 and S2 emphasise connected information and decision-useful disclosures (ISSB, 2023a, 2023b). ISSA 5000 applies across sustainability topics, reporting frameworks and both limited- and reasonable-assurance engagements (IAASB, 2024). When AI is used to classify narratives, estimate emissions, screen evidence or draft assurance documentation, provenance and explainability become part of the credibility of the engagement.

2.4. Analytical Model

The study treats AI governance as a conversion system (Figure 1) between technological capability and professional output. AI capability enters an audit task through data, models and interfaces. A professional interprets the output within standards and methodology. Governance controls determine whether that interpretation is informed, reviewable and challengeable. The final judgment then affects audit evidence, documentation and communication. Feedback from review, inspection and incidents should update the system and the organisation's approved-use boundary.
The model predicts neither that more technology necessarily improves quality nor that more disclosure proves effective implementation. It instead identifies observable governance commitments. A public report can provide evidence that a firm recognises a safeguard and describes a mechanism. It cannot demonstrate that the mechanism operated effectively on every engagement. This distinction defines the empirical claims made below.

3. Materials and Methods

3.1. Research Design and Public-Document Corpus

The study uses comparative qualitative content analysis with a structured ordinal coding instrument. Public organisational documents are appropriate for examining how firms construct accountability, identify risks and communicate governance to investors, audit committees, regulators and other stakeholders. They are not neutral descriptions: firms select what to disclose and may use favourable language. For that reason, the unit measured is disclosure completeness, not latent governance quality.
The population was defined as the four largest UK audit firms subject to the FRC's operational-separation principles and commonly described as the Big Four. The sampling frame was each firm's report labelled Transparency Report for the reporting period ending in 2024. All four available reports were included; no observation was sampled within the population. The documents are comparable in jurisdiction, broad regulatory purpose and reporting cycle, although fiscal year-ends and report structures differ. The corpus contains 669 PDF pages in total (Table 2).
The reports were downloaded from the firms' official websites. Searchable text was extracted from the PDF files, and the complete report—not only technology sections—was searched for AI, generative AI, algorithm, professional judgment/judgement, professional skepticism/scepticism, validation, testing, pilot, documentation, explainability, security, privacy, data governance, accountability, training and sustainability. Each located passage was read in page context. The evidence table in Appendix A (Table A1) records the principal pages supporting each score.
The FRC's Annual Review of Audit Quality 2024 was added as a public regulatory source. The FRC reported the percentage of inspected audits assessed as good or requiring no more than limited improvements: Deloitte 94%, EY 76%, KPMG 89% and PwC 76% (FRC, 2024). The inspected audits were a risk-based sample. These percentages are therefore descriptive context and cannot be extrapolated to every audit or treated as a dependent variable in a four-observation causal model.

3.2. Coding Instrument

Seven dimensions were derived deductively from the professional-judgment literature and normative frameworks, then applied consistently to all reports (Table 3). Each dimension received a score of 0, 1 or 2. A score of 0 means that no relevant disclosure was located. A score of 1 means that the report contains a general commitment, adjacent control or limited mechanism. A score of 2 requires an explicit AI- or audit-specific mechanism. The stricter threshold avoids awarding full credit for generic statements about technology or firm-wide security.
The dimensions are intentionally process-oriented. Audit-specific AI integration identifies whether the report moves beyond aspiration to deployed or scaled audit use. Human oversight identifies whether professional judgment remains substantive. Validation and testing identify pre-release or use-case evaluation. Explainability and documentation identify whether an external reviewer can trace how AI affected work. Data governance and security identify controls over data and permitted use. Accountability identifies assigned owners, committees or approval routes. AI-specific learning identifies structured competence development rather than broad technology awareness.

3.3. AI–Judgment Governance Disclosure Index

For firm i, the AI–Judgment Governance Disclosure Index is defined in Equation (1):
A I - J G D I i = 100 14 k = 1 7 s i k ,
where sik is the score (0, 1 or 2) for dimension k. All dimensions have equal weight because there is no validated empirical basis for assigning differential weights. The index ranges from 0 to 100 and should be interpreted as a transparent summary of the coding rules. A score of 100 means that the report contained explicit disclosures for all seven dimensions; it does not mean that AI risk was eliminated or that all engagements complied.
Coding was conducted in two passes. The first pass extracted every potentially relevant passage and recorded its report page. The second pass applied the pre-specified 0–2 rules to the evidence set and revisited the full page when a passage was ambiguous. Because the study has one coder, it does not report inter-coder reliability. Replicability is instead supported by the full codebook, firm-by-dimension matrix and page-level evidence trail. An independent re-coding is an important next step before the index is used in large-sample hypothesis testing.

3.4. Analysis and Limitations of Inference

The analysis combines within-case reading, cross-case comparison and descriptive scoring. No inferential statistics are used because the population contains four firms and the score is ordinal and disclosure-based. The FRC percentages are shown alongside the index to prevent a purely self-reported account, but no correlation or ranking claim is made. Differences may reflect disclosure style, report architecture and reporting choices as well as underlying practice.
All empirical inputs are publicly accessible. No interviews, confidential engagement files, personal data or proprietary datasets were used. The study did not evaluate the source code, model performance or engagement-level operation of any named tool. Assertions about systems are therefore attributed to the firms' public reports.

4. Results

4.1. AI is Disclosed as Deployed Augmentation, Not Autonomous Judgment

All four reports move beyond a generic expectation that AI may be used in the future. Deloitte describes PairD, AI and machine-learning functions in its Omnia platform, fifteen AI/GenAI use cases under development, and pilot applications for technical research, document retrieval, first review and document creation (Deloitte, 2024, pp. 8, 90–92). EY reports globally scaled AI capabilities integrated with EY Canvas to support risk assessment (EY, 2024, pp. 36–37). KPMG states that KPMG Clara AI chat was launched to all UK auditors and that transaction scoring had been deployed to nearly 900 UK audits (KPMG, 2025, pp. 36–37). PwC reports the rollout of ChatPwC, an Audit GenAI Hub and approved engagement use cases (PwC, 2024, pp. 110–112).
The common operational model is augmentation. Deloitte states that expertise, professional skepticism and judgment are used to challenge and assure the reliability of output. EY says that AI-enabled technology supports procedures but does not re-place the professional's experience and judgment. KPMG combines smart technology with curious and inquisitive minds and professional skepticism. PwC describes a human-led, technology-powered audit and requires skeptical review of GenAI outputs. These disclosures support a consistent answer to RQ1: public accountability remains attached to the auditor, even where AI is scaled across the practice.
This framing is important because tools perform different functions. An anomaly score reallocates attention; a technical chatbot retrieves or synthesises guidance; a drafting assistant creates an initial artefact; and a transaction-scoring model analyses a population. None of these functions establishes by itself whether evidence is sufficient or whether an accounting estimate is reasonable. The reports generally recognise this boundary.

4.2. Validation and Traceability Are the Least Consistently Disclosed Safeguards

Validation disclosure varies more than adoption disclosure. EY provides the clearest lifecycle description: technology concepts pass through a global committee; testing with end users, piloting, feedback and certification are prerequisites for release. PwC describes prompt-engineering and validation practices in the Audit GenAI Hub. Deloitte reports pilots, use-case approval, risk thresholding and a clearing-house process, which demonstrate gatekeeping but provide less detail about performance validation. KPMG reports extensive deployment and responsible user challenge, but the 2024 report does not describe an AI-specific validation or certification mechanism at the same level of detail. Under the strict codebook, this produces scores of 2 for EY and PwC, 1 for Deloitte and 0 for KPMG on the validation dimension.
Explainability is also incompletely disclosed. The reports discuss transparent audit services, documentation and the ability of AI tools to support or improve working pa-pers. However, only PwC explicitly states that clear documentation is required where approved GenAI use cases have been used. None of the reports supplies public mod-el-level information such as performance thresholds, error rates, explainability methods, override frequency or post-deployment drift monitoring. Such details may exist internally and may be inappropriate to disclose fully for security or proprietary rea-sons. Yet an external reader cannot determine from most reports what minimum ex-planation must be retained in an engagement file when an AI output materially influences a judgment.
The result identifies a disclosure boundary rather than proving a control deficiency. Public transparency reports are designed for multiple regulatory and stake-holder purposes, not as model cards. Nevertheless, validation and traceability are the dimensions where the difference between a statement of responsible intent and an externally assessable control is greatest.

4.3. Data Governance and Accountability Are More Visible

Deloitte discloses a safe and secure environment for PairD and a firm-wide GenAI risk response involving use-case approval, data use and management, ethical use, cyber risk, a Global Data Council, risk thresholding and a Trustworthy AI framework (Deloitte, 2024, p. 108). EY links responsible technology use with standardised development protocols, a global evaluation committee, information-security policies and privacy-impact assessments for new technology (EY, 2024, pp. 36–37, 60–61, 129–130). KPMG reports Risk Committee deep dives on AI and data risk, secure interaction within Clara, and general information-security governance, but the AI section contains less explicit detail about AI-specific data lineage or permitted-data rules. PwC de-scribes a secure environment, states that ChatPwC does not use prompts or responses to train the underlying model, limits allowable tools and use cases, and requires clear documentation (PwC, 2024, pp. 111–112).
Accountability structures are explicit across the corpus. Deloitte identifies use-case governance and programme ownership. EY identifies a global committee involving Professional Practice, the Assurance Quality Network and Technology. KPMG identifies central technology teams and board/Risk Committee oversight. PwC identifies an Audit GenAI Hub with audit subject-matter experts, data scientists and innovation managers. All four therefore receive the maximum accountability score. The reports differ less on whether someone owns AI governance than on what public evidence is supplied about validation outputs and engagement-level traceability.

4.4. Learning is Treated as a Condition of Responsible Use

Three firms disclose broad AI-specific learning at a level meeting the maximum criterion. Deloitte's principal-risk response includes an AI-fluency workstream. KPMG reports that all auditors were trained in prompt engineering at its 2024 Audit University so that they could engage with and challenge Clara AI chat responsibly. PwC makes training on GenAI fundamentals and audit business rules mandatory before access to ChatPwC. EY discloses AI badges and technology learning, but the 2024 transparency report is less explicit about a mandatory audit-wide AI curriculum; it therefore receives a score of 1 under the strict rule.
The emphasis on learning supports a dual-competence model. An auditor needs domain competence to recognise an implausible output and AI literacy to understand data, uncertainty, permitted use and limitations. Prompt skill alone is not professional competence. Conversely, a technically expert accountant who cannot evaluate model limitations may either reject useful evidence or accept output ceremonially. The re-ports generally treat technology training as complementary to professional skepticism rather than as its replacement.

4.5. Disclosure Scores and External Inspection Context

The cross-firm comparison reveals a common baseline of deployed audit-specific AI, retained human oversight and identifiable accountability structures. The principal differences concern the specificity of validation and testing mechanisms, engagement-level traceability, AI-related data controls and learning arrangements. Table 4 presents the dimension-level coding and the resulting AI–Judgment Governance Disclosure Index (AI-JGDI).
The results indicate that the differences among the firms are not attributable to the absence of deployed AI or human oversight, for which all four reports receive the maximum scores. Instead, the variation arises from the level of detail provided about validation, explainability and documentation, AI-specific data governance and structured learning. PwC provides explicit mechanisms across all seven dimensions, while Deloitte and EY show similarly broad disclosure profiles with more limited information in selected areas. KPMG’s lower index value reflects the absence or limited specificity of public disclosures under the strict coding rules, particularly in relation to AI-specific validation and data governance.
The AI-JGDI should therefore be interpreted as a profile of public disclosure completeness rather than as a ranking of actual governance effectiveness or audit quality. A maximum score indicates that explicit public evidence was identified for every coded dimension; it does not demonstrate that the disclosed controls operated effectively across all engagements. Correspondingly, a lower score indicates that the specified mechanism was not located in the transparency report and does not establish that the firm lacks an equivalent internal control.
The FRC’s risk-based inspection results provide an external context for interpreting the disclosure index. As shown in Table 5, the ordering of the inspection outcomes does not correspond to the AI-JGDI ordering. This descriptive mismatch confirms that public AI-governance disclosure and regulatory audit-quality inspection capture different aspects of accountability and should not be treated as interchangeable measures.
The comparison does not support a firm-level causal inference. The AI-JGDI measures the completeness of governance mechanisms disclosed in public transparency reports, whereas the FRC percentages reflect the outcomes of risk-based inspections of selected audit engagements. The two indicators differ in their objects of measurement, evidence bases and sampling conditions. Consequently, a more complete AI-governance disclosure profile cannot be interpreted as evidence of higher audit quality, just as a lower disclosure score cannot be interpreted as evidence of weaker internal practice.

4.6. AI and Sustainability Assurance Remain Parallel Rather Than Integrated Narratives

All four reports discuss sustainability reporting, climate-related financial report-ing, ESG expertise or the development of sustainability-assurance capacity. Deloitte describes CSRD, ESRS, ISSB Standards, assurance methodology and sustainability training. EY discusses ESG and climate specialists and the effect of emerging risks on audit. KPMG presents sustainability and resilience as part of its purpose and identifies ESG in governance oversight. PwC discusses growth in ESG and assurance capabilities.
Despite this extensive coverage, explicit AI–sustainability integration is limited. The reports do not set out a distinct protocol for using AI to evaluate sustainability evidence, document double-materiality judgments, validate emissions estimates or preserve provenance for value-chain data. Deloitte's principal-risk discussion recognises that GenAI has public-interest and sustainability implications, but this is not the same as an engagement methodology. The answer to RQ3 is therefore that AI and sustainability assurance are both strategic themes, yet they are mainly disclosed in parallel.
This gap matters because sustainability data can be especially susceptible to in-consistent definitions, estimation, missing value-chain evidence and narrative bias. An AI tool may help classify or compare disclosures, but the assurance practitioner still needs to evaluate the reporting criteria, materiality process, source data, internal control and possibility of greenwashing. The governance architecture developed for financial audit should be extended explicitly to sustainability-assurance use cases.

5. Discussion

5.1. AI Redistributes Professional Judgment

The evidence supports a redistribution thesis. AI does not remove judgment; it moves judgment across the workflow. Before use, people decide the purpose, training or reference data, permitted population and performance threshold. During use, the auditor decides whether an output is relevant, whether contradictory evidence exists and whether further procedures are necessary. After use, reviewers decide whether documentation supports the conclusion and whether incidents require remediation. Professional judgment therefore operates both on the accounting or assurance issue and on the technology used to analyse it.
This result extends the accounting judgment literature. The traditional bounded judgment space is created by standards, transactions and uncertainty. AI adds an algorithmic evidence layer that determines what is salient and how alternatives are presented. A high anomaly score, generated summary or suggested conclusion can be-come an anchor. The professional must evaluate both the underlying economic question and the reliability of the mediation. This is why human oversight should be de-fined as effective intervention, not final approval after the system has framed the answer.

5.2. From Human-in-the-Loop to Accountable Human Control

The phrase human-in-the-loop is too weak if the human role is ceremonial. The public reports use stronger language—human-led, professional judgment, challenge and skepticism—but external accountability also requires operational evidence. An accountable human-control design should record: the approved purpose; the tool and version; the source and permissible use of data; relevant validation; the material out-put used; contradictory evidence; the professional's evaluation; any override; the re-viewer; and the final conclusion.
This design aligns engagement quality management with AI risk management (Table 6). ISA 220 (Revised) assigns responsibility for managing and achieving quality at engagement level, while ISQM 1 requires a risk-based system of quality management (IAASB, 2020a, 2020b). NIST's Govern–Map–Measure–Manage sequence provides a compatible technology lens (NIST, 2023). The EU AI Act's concepts of documentation, logging, human oversight, robustness and cybersecurity provide additional design prompts even where a specific audit tool is not legally classified as high-risk (European Parliament and Council, 2024).
The table is not intended as a universal checklist. Controls should be proportion-ate to the influence of the system. A search assistant that retrieves paragraphs from an approved standards library may require different validation from a model that scores journal entries or generates a valuation range. The decisive factor is not whether the tool is labelled AI, but how its output can affect the nature, timing or extent of procedures and the resulting professional judgment.

5.3. Disclosure Completeness is a Governance Outcome in Its Own Right

Transparency reports serve several functions: regulatory compliance, stakeholder communication and reputation. Measuring them cannot reveal every internal control. Yet disclosure completeness matters because audit committees, investors and regulators need a basis for informed dialogue. A firm can reasonably protect proprietary model details while still explaining governance roles, testing categories, permitted-use boundaries, monitoring and the documentation expected when AI influences an engagement.
The AI-JGDI is therefore best understood as a conversation and research instrument. It shows where a public report supplies enough information to identify a mechanism and where it supplies only a general commitment. The equal weighting is deliberately simple. Future research can validate weights through regulator, audit-committee and practitioner surveys or link individual dimensions to engagement outcomes in a larger panel.

5.4. Implications for Sustainability Assurance

Sustainability assurance creates a test of whether AI governance is genuinely integrated. CSRD/ESRS and IFRS S1/S2 require connected, decision-useful information, while ISSA 5000 requires appropriate evidence across diverse sustainability matters. AI can support document comparison, evidence classification, anomaly detection and consistency checks across narrative and quantitative information. It can also scale weak source data or produce fluent but unsupported explanations.
Firms should therefore define sustainability-specific AI use cases and evidence boundaries. A model used to classify value-chain evidence should retain links to source documents. A system used to compare disclosures with ESRS should not be treated as determining materiality. An emissions-estimation model should be evaluated for methodology, data completeness, uncertainty and sensitivity. A generative tool used to draft assurance documentation should not be allowed to convert absence of evidence into confident prose. These controls connect AI governance directly to the professional judgments required by ISSA 5000.

5.5. Implications for Regulators, Firms, Audit Committees and Education

Regulators can improve comparability by developing non-prescriptive disclosure expectations for material AI use in audit. Useful categories include use-case governance, validation, data controls, human oversight, documentation, incident monitoring and competence. This would avoid demanding proprietary source code while enabling stakeholders to distinguish aspiration from an operating governance process.
Audit firms can map approved AI use cases to their system of quality management. The map should identify the quality objective affected, the risk created, the control response, the owner, monitoring evidence and remediation route. Model or tool changes should trigger reassessment. Engagement teams should document material use in the same way they document specialists, data analytics or other sources of evidence.
Audit committees should ask focused questions: Which AI tools affected the audit? What data were used? How were the tools tested for the relevant purpose? What outputs were challenged or overridden? What remains a human judgment? Were any limitations communicated? For sustainability assurance, the committee should ask how source provenance and double-materiality judgments were protected from auto-mated simplification.
Education should integrate accounting judgment and AI literacy rather than teach them separately. Case-based learning can require students to evaluate a model output against accounting standards, identify missing evidence, document an override and explain the conclusion to governance bodies. This approach preserves the professional identity of the accountant while preparing graduates for AI-mediated work.

6. Conclusions

This study examined how the UK Big Four publicly describe AI-assisted professional judgment using a complete 2024 cross-section of transparency reports and only public empirical data. All four firms disclosed audit-specific AI use and retained human professional responsibility. Governance ownership and learning were prominent. Validation, model-level explanation and engagement-file traceability were less consistently described. The AI-JGDI ranged from 71.4 to 100.0, but it measures public disclosure completeness and should not be interpreted as actual audit quality. FRC inspection results were reported separately and did not mirror the disclosure ordering.
The main theoretical conclusion is that AI redistributes rather than eliminates judgment. Accountants and auditors make judgments about the reporting issue, the evidence and the technological system mediating that evidence. The main practical conclusion is that human oversight must be evidenced through purpose approval, validation, data controls, documentation, review, override and assigned accountability. A professional signature alone does not demonstrate meaningful control.
The study also identifies a strategic gap for the special issue theme: AI and sustainability assurance are prominent but largely parallel disclosures. Public reports provide limited detail about how AI governance is adapted to sustainability evidence, materiality processes, emissions estimates or value-chain information. Extending the human–AI judgment protocol to ISSA 5000 engagements is therefore an immediate research and practice priority.
The limitations are material. The sample contains four firms in one jurisdiction and one reporting cycle. Transparency reports are self-reported and vary in structure. The ordinal codebook uses equal weights and one coder. Public disclosures cannot establish engagement-level operation or model performance. FRC inspections are risk-based and cannot be matched to individual AI use cases. These constraints rule out causal inference.
Future research should create a multi-year, multi-jurisdiction panel; use independent coders; validate the index with audit committees and regulators; and examine engagement-level evidence under confidentiality protections. Experiments can test whether the proposed documentation improves calibrated reliance. Field studies can compare financial audit with sustainability assurance. Research should also examine override direction: whether professionals challenge both unfavourable and favourable AI outputs with equal rigor. The public-document method and appendix supplied here provide a reproducible starting point.

Funding

This research was partial funded by Institute for Scientific Research of D. A. Tsenov Academy of Economics, Svishtov, Bulgaria: 1-2026.

Institutional Review Board Statement

Not applicable.

Data Availability Statement

All empirical data are public. The four firm transparency reports and the FRC publication are available at the URLs listed in the References; the page-level evidence supporting each score is documented in Appendix A (Table A1).

Acknowledgments

During the preparation of this manuscript, the author used generative AI tools for language and drafting support. The author has reviewed and edited the output and takes full responsibility for the content of this publication.

Conflicts of Interest

The author declares no conflicts of interest.

Appendix A

Evidence Trail for the AI-JGDI Coding
Appendix A provides the page-level evidence used to support the AI-JGDI coding. The principal disclosures underlying each firm-by-dimension score are presented in Table A1.
Table A1. Page-level evidence trail supporting AI-JGDI scores.
Table A1. Page-level evidence trail supporting AI-JGDI scores.
Dimension Deloitte EY KPMG PwC
AI use 2: PairD; Omnia pilots; AI/ML (pp. 8, 90–92) 2: scaled AI integrated with Canvas (pp. 36–37) 2: Clara AI chat; scoring on nearly 900 audits (pp. 36–37) 2: ChatPwC; approved use cases; GenAI Hub (pp. 110–112)
Human oversight 2: expertise, skepticism and judgment challenge output (p. 8) 2: supports but does not replace experience and judgment (p. 37) 2: inquisitive minds, skepticism and responsible challenge (pp. 36–37) 2: human-led; skepticism and quality review (pp. 110–112)
Validation/testing 1: pilots, use-case approval and risk thresholding (pp. 91, 108) 2: committee evaluation, testing, pilot, feedback and certification (p. 37) 0: no AI-specific validation mechanism located under strict rule 2: prompt-engineering and validation practices (p. 111)
Traceability 1: transparent service and first-review support; no AI-use record rule located 1: consistent documentation; no AI-specific trace rule located 1: AI improves documentation; no AI-specific trace rule located 2: clear documentation required where GenAI is used (p. 112)
Data governance/security 2: secure environment; data-use governance; Data Council (pp. 8, 108) 2: responsible-use policies and privacy assessment for technology (pp. 37, 60–61) 1: secure interaction and general data-risk oversight (pp. 36, 133) 2: secure environment; prompts not used for model training; allowable tools (pp. 111–112)
Accountability 2: use-case governance and clearing-house process (p. 108) 2: global multidisciplinary evaluation committee (p. 37) 2: central team and Risk Committee oversight (pp. 37, 133) 2: dedicated Audit GenAI Hub and business rules (pp. 110–112)
AI learning 2: AI-fluency/adoption workstream (p. 108) 1: AI badges and broader learning; no mandatory audit-wide AI course located 2: all auditors trained in prompt engineering (p. 37) 2: mandatory GenAI fundamentals and business rules (pp. 110–112)
Note: A score of 0 means that the specified disclosure was not located in the public report; it is not evidence that the control does not exist internally.
The page references in Appendix A point to the principal supporting disclosures, not every occurrence of the coded concept. A zero indicates that the specified AI-specific mechanism was not located under the stated rule; it does not prove that the firm lacks an undisclosed internal control.

References

  1. Abdullah; Almaqtari; Abdullah, A. A. H.; Almaqtari, F. A. The impact of artificial intelligence and Industry 4.0 on transforming accounting and auditing practices. Journal of Open Innovation: Technology, Market, and Complexity 2024, 10(1), 100218. [Google Scholar] [CrossRef]
  2. Agoglia; Agoglia, C. P.; Doupnik, T. S.; Tsakumis, G. T.; et al. Principles-based versus rules-based accounting standards: The influence of standard precision and audit committee strength on financial reporting decisions. The Accounting Review 2011, 86(3), 747–767. [Google Scholar] [CrossRef]
  3. Backof; Backof, A. G.; Bamber, E. M.; Carpenter, T. D.; et al. Do auditor judgment frameworks help in constraining aggressive report-ing? Evidence under more precise and less precise accounting standards. Accounting, Organizations and Society 2016, 51, 1–11. [Google Scholar] [CrossRef]
  4. Bennett; Bennett, B.; Bradbury, M.; Prangnell, H.; et al. Rules, principles and judgments in accounting standards. Abacus 2006, 42(2), 189–204. [Google Scholar] [CrossRef]
  5. Bonner; Bonner, S. E. Judgment and decision-making research in accounting. Accounting Horizons 1999, 13(4), 385–398. [Google Scholar] [CrossRef]
  6. Commerford; Commerford, B. P.; Dennis, S. A.; Joe, J. R.; Ulla, J. W.; et al. Man versus machine: Complex estimates and auditor reliance on artificial intelligence. Journal of Accounting Research 2022, 60(1), 171–201. [Google Scholar] [CrossRef]
  7. Commerford; Commerford, B. P.; Eilifsen, A.; Hatfield, R. C.; Holmstrom, K. M.; Kinserdal, F.; et al. Control issues: How providing input affects auditors' reliance on artificial intelligence. Contemporary Accounting Research 2024, 41(4), 2134–2162. [Google Scholar] [CrossRef]
  8. Deloitte, 2024) Deloitte Deloitte LLP and Deloitte Limited 2024 transparency report. Deloitte LLP. 2024. Available online: https://www.deloitte.com/content/dam/assets-zone2/uk/en/docs/about/2024/deloitte-uk-annual-review-2024-audit-transparency-report.pdf (accessed on 12 July 2026).
  9. Dietvorst; Dietvorst, B. J.; Simmons, J. P.; Massey, C.; et al. Algorithm aversion: People erroneously avoid algorithms after seeing them err. Journal of Experimental Psychology: General 2015, 144(1), 114–126. [Google Scholar] [CrossRef] [PubMed]
  10. European Commission; European Commission. Commission Delegated Regulation (EU) 2023/2772 of 31 July 2023 supplementing Directive 2013/34/EU as regards sustainability reporting standards. Official Journal of the European Union. 2023. Available online: https://eur-lex.europa.eu/eli/reg_del/2023/2772/oj/eng (accessed on 12 July 2026).
  11. European Parliament and Council. Directive (EU) 2022/2464 of 14 December 2022 as regards corporate sustainability re-porting. Official Journal of the European Union. 2022. Available online: https://eur-lex.europa.eu/eli/dir/2022/2464/oj/eng (accessed on 12 July 2026).
  12. European Parliament and Council. Regulation (EU) 2024/1689 of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union. 2024. Available online: https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng (accessed on 12 July 2026).
  13. (EY, 2024) EY. (2024). EY UK 2024 transparency report. Ernst & Young LLP. Available online: https://www.ey.com/content/dam/ey-unified-site/ey-com/en-uk/about-us/documents/ey-uk-2024-transparency-report.pdf (accessed on 12 July 2026).
  14. Fields; Fields, T. D.; Lys, T. Z.; Vincent, L.; et al. Empirical research on accounting choice. Journal of Accounting and Economics 2001, 31(1–3), 255–307. [Google Scholar] [CrossRef]
  15. FRC. FRC publishes annual Tier 1 audit firm inspection results. Financial Reporting Council. 2024. Available online: https://www.frc.org.uk/news-and-events/news/2024/07/frc-publishes-annual-tier-1-audit-firm-inspection-results/ (accessed on 12 July 2026).
  16. Healy; Wahlen; Healy, P. M.; Wahlen, J. M. A review of the earnings management literature and its implications for standard setting. Accounting Horizons 1999, 13(4), 365–383. [Google Scholar] [CrossRef]
  17. Hurtt; Hurtt, R. K. Development of a scale to measure professional skepticism. Auditing: A Journal of Practice & Theory 2010, 29(1), 149–171. [Google Scholar] [CrossRef]
  18. IAASB; IAASB. International Standard on Auditing 200: Overall objectives of the independent auditor and the conduct of an audit in ac-cordance with International Standards on Auditing. International Auditing and Assurance Standards Board. 2009. Available online: https://www.iaasb.org/ (accessed on 12 July 2026).
  19. IAASB. International Standard on Auditing 540 (Revised): Auditing accounting estimates and related disclosures; International Auditing and Assurance Standards Board, 2018; Available online: https://www.iaasb.org/focus-areas/embedding-professional-skepticism (accessed on 12 July 2026).
  20. IAASB. International Standard on Auditing 315 (Revised 2019): Identifying and assessing the risks of material misstatement. In-ternational Auditing and Assurance Standards Board. 2019. Available online: https://www.iaasb.org/consultations-projects/isa-315-revised (accessed on 12 July 2026).
  21. IAASB; IAASB. International Standard on Quality Management 1: Quality management for firms that perform audits or reviews of finan-cial statements, or other assurance or related services engagements. International Auditing and Assurance Standards Board. 2020a. Available online: https://www.iaasb.org/publications/international-standard-quality-management-isqm-1-quality-management-firms-perform-audits-or-reviews (accessed on 12 July 2026).
  22. IAASB; IAASB. International Standard on Auditing 220 (Revised): Quality management for an audit of financial statements. Internation-al Auditing and Assurance Standards Board. 2020b. Available online: https://www.iaasb.org/publications/international-standard-auditing-220-revised-quality-management-audit-financial-statements (accessed on 12 July 2026).
  23. IAASB. International Standard on Sustainability Assurance 5000: General requirements for sustainability assurance engagements. International Auditing and Assurance Standards Board. 2024. Available online: https://www.iaasb.org/publications/international-standard-sustainability-assurance-5000-general-requirements-sustainability-assurance (accessed on 12 July 2026).
  24. IESBA. Final pronouncement: Technology-related revisions to the Code. International Ethics Standards Board for Accountants. 2023. Available online: https://www.ethicsboard.org/publications/final-pronouncement-technology-related-revisions-code (accessed on 12 July 2026).
  25. IESBA. 2024 handbook of the International Code of Ethics for Professional Accountants, including International Independence Standards. International Ethics Standards Board for Accountants. 2024. Available online: https://www.ethicsboard.org/publications/2024-handbook-international-code-ethics-professional-accountants (accessed on 12 July 2026).
  26. ISSB. IFRS S1 General requirements for disclosure of sustainability-related financial information. IFRS Foundation. 2023a. Available online: https://www.ifrs.org/issued-standards/ifrs-sustainability-standards-navigator/ifrs-s1-general-requirements/ (accessed on 12 July 2026).
  27. ISSB. IFRS S2 Climate-related disclosures. IFRS Foundation. 2023b. Available online: https://www.ifrs.org/issued-standards/ifrs-sustainability-standards-navigator/ifrs-s2-climate-related-disclosures/ (accessed on 12 July 2026).
  28. Kokina; Kokina, J.; Blanchette, S.; Davenport, T. H.; Pachamanova, D.; et al. Challenges and opportunities for artificial intelligence in auditing: Evidence from the field. International Journal of Accounting Information Systems 2025, 56, 100734. [Google Scholar] [CrossRef]
  29. Kokina; Davenport; Kokina, J.; Davenport, T. H. The emergence of artificial intelligence: How automation is changing auditing. Journal of Emerging Technologies in Accounting 2017, 14(1), 115–122. [Google Scholar] [CrossRef]
  30. KPMG; KPMG. UK transparency report 2024. KPMG LLP. 2025. Available online: https://assets.kpmg.com/content/dam/kpmgsites/uk/pdf/2026/01/uk-transparency-report-2024.pdf (accessed on 12 July 2026).
  31. Lehner; Lehner, O. M.; Ittonen, K.; Silvola, H.; Ström, E.; Wührleitner, A.; et al. Artificial intelligence based decision-making in ac-counting and auditing: Ethical challenges and normative thinking. Accounting, Auditing & Accountability Journal 2022, 35(9), 109–135. [Google Scholar] [CrossRef]
  32. Libby; Luft; Libby, R.; Luft, J. Determinants of judgment performance in accounting settings: Ability, knowledge, motivation, and environment. Accounting, Organizations and Society 1993, 18(5), 425–450. [Google Scholar] [CrossRef]
  33. Logg; Logg, J. M.; Minson, J. A.; Moore, D. A.; et al. Algorithm appreciation: People prefer algorithmic to human judgment. Organi-zational Behavior and Human Decision Processes 2019, 151, 90–103. [Google Scholar] [CrossRef]
  34. Murikah; Murikah, W.; Nthenge, J. K.; Musyoka, F. M.; et al. Bias and ethics of AI systems applied in auditing: A systematic review. Scientific African 2024, 25, e02281. [Google Scholar] [CrossRef]
  35. NIST Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1; National Institute of Standards and Technology, U.S. Department of Commerce, 2023. [CrossRef]
  36. Nelson; Nelson, M. W. Behavioral evidence on the effects of principles- and rules-based standards. Accounting Horizons 2003, 17(1), 91–104. [Google Scholar] [CrossRef]
  37. Nelson; Nelson, M. W. A model and literature review of professional skepticism in auditing. Auditing: A Journal of Practice & Theory 2009, 28(2), 1–34. [Google Scholar] [CrossRef]
  38. PwC. UK transparency report 2024. PricewaterhouseCoopers LLP. 2024. Available online: https://www.pwc.co.uk/transparencyreport/assets/pdf/uk-transparency-report-2024.pdf (accessed on 12 July 2026).
  39. Stratopoulos; Wang; Stratopoulos, T. C.; Wang, V. X. Artificial intelligence and accounting research: A framework and agenda. International Journal of Accounting Information Systems 2025, 100760. [Google Scholar] [CrossRef]
  40. Sutton; Sutton, S. G.; Holt, M.; Arnold, V.; et al. “The reports of my death are greatly exaggerated”–Artificial intelligence research in accounting. International Journal of Accounting Information Systems 2016, 22, 60–73. [Google Scholar] [CrossRef]
Figure 1. Human–AI governance model for accountable professional judgment.
Figure 1. Human–AI governance model for accountable professional judgment.
Preprints 224475 g001
Table 1. Normative layers relevant to AI-assisted professional judgment.
Table 1. Normative layers relevant to AI-assisted professional judgment.
Layer Principal sources Judgment implication AI-governance implication
Engagement and quality management ISA 200; ISA 220 (Revised); ISA 315 (Revised 2019); ISA 540 (Revised); ISQM 1 The auditor retains responsibility for judgment, skepticism, evidence and quality. Tools must be embedded in engagement acceptance, risk assessment, review, documentation and monitoring.
Professional ethics IESBA Code and technology-related revisions Integrity, objectivity, competence, confidentiality and professional behaviour apply to technology use. Technology creates threats that require evaluation, safeguards and appropriate professional action.
General AI risk governance EU AI Act; NIST AI RMF 1.0 Human evaluation is necessary where system output affects a consequential decision. Purpose definition, risk management, data controls, testing, documentation, oversight, robustness and cybersecurity.
Sustainability reporting and assurance CSRD; ESRS; IFRS S1; IFRS S2; ISSA 5000 Materiality, estimates, source reliability and sufficiency of evidence remain professional judgments. AI use should preserve provenance, uncertainty, connected information and reviewability across financial and sustainability data.
Source: Author's synthesis of the cited standards and legislation. The EU AI Act benchmark is applied analytically; the table does not classify every audit tool as a high-risk AI system.
Table 2. Public empirical corpus.
Table 2. Public empirical corpus.
Firm Reporting period PDF pages Principal AI evidence pages Document status
Deloitte UK Year ended 31 May 2024 155 8; 90–92; 108 Official 2024 Transparency Report
EY UK Year ended 28 June 2024 161 36–37; 60–61; 129–130 Official 2024 Transparency Report
KPMG UK Year ended 30 September 2024 177 36–37; 133 Official 2024 Transparency Report
PwC UK Year ended 30 June 2024 176 110–112; 102–103 Official 2024 Transparency Report
Source: Official firm reports. Total corpus: 669 pages. Reports were accessed on 12 July 2026.
Table 3. AI–Judgment Governance Disclosure Index codebook.
Table 3. AI–Judgment Governance Disclosure Index codebook.
Dimension Score 0 Score 1 Score 2: explicit mechanism
AI use No audit AI located Experiment or general aspiration Deployed/scaled audit-specific tool or use case
Human oversight No human role located General professional judgment statement Explicit challenge, review or override of AI output
Validation/testing No relevant disclosure Pilot, approval or general evaluation AI-specific testing/validation and release criteria
Traceability No relevant disclosure General documentation/transparency AI-use documentation, explanation or audit trail required
Data governance/security No relevant disclosure General security/privacy control AI-specific permitted-data, privacy, security or training-data control
Accountability No owner or route located General governance Named committee, owner, hub, approval or escalation route
AI learning No relevant disclosure Optional/general technology learning Structured or mandatory AI-specific learning for auditors
Note: The index measures the completeness of public disclosure, not the operating effectiveness of controls.
Table 4. Firm-by-dimension AI-JGDI scores.
Table 4. Firm-by-dimension AI-JGDI scores.
Dimension Deloitte EY KPMG PwC
Audit-specific AI integration 2 2 2 2
Human judgment and oversight 2 2 2 2
Validation and testing 1 2 0 2
Explainability and documentation 1 1 1 2
Data governance and security 2 2 1 2
Accountability structures 2 2 2 2
AI-specific learning 2 1 2 2
Total (maximum 14) 12 12 10 14
AI-JGDI 85.7 85.7 71.4 100.0
Note: Source: Author's coding of official 2024 UK transparency reports. Scores follow Table 3; evidence is in Appendix A.
Table 5. AI-governance disclosure and FRC inspection context.
Table 5. AI-governance disclosure and FRC inspection context.
Firm AI-JGDI FRC 2024: inspected audits good/limited improvements Permitted interpretation
Deloitte 85.7 94% Disclosure completeness and a risk-based inspection result are separate indicators.
EY 85.7 76% No firm-level causal inference is made.
KPMG 71.4 89% A lower disclosure score does not establish weaker internal practice.
PwC 100.0 76% A complete disclosure score does not establish higher audit quality.
Note: Source: AI-JGDI from Table 4; inspection percentages from FRC (2024). The FRC sample is risk-based and should not be extrapolated to all audits.
Table 6. Accountable human-control protocol for financial and sustainability assurance.
Table 6. Accountable human-control protocol for financial and sustainability assurance.
Control stage Minimum retained evidence Professional judgment question Principal normative anchor
Approve purpose Defined use case, intended user, prohibited use and responsible owner Can this tool appropriately support this task without determining the conclusion? ISQM 1; NIST Govern/Map
Validate Test population, criteria, known limitations, approval and change/version history Is performance sufficient for this purpose and population? NIST Measure; EU AI Act control concepts
Govern data Provenance, permissions, confidentiality, retention and security Is the input complete, lawful, reliable and appropriate? IESBA confidentiality; ISA 315; NIST Govern
Evaluate output Material output, uncertainty, contradictory evidence and additional procedures What does the output not establish, and what evidence could disconfirm it? ISA 200; ISA 540; ISSA 5000
Override and review Acceptance/override rationale, reviewer, escalation and final conclusion Would the same challenge be applied if the output moved the conclusion in the opposite direction? ISA 220; professional skepticism
Monitor and remediate Incidents, drift, user feedback, inspection findings and remediation Should the use case, control or training be changed or withdrawn? ISQM 1; NIST Manage
Note: Source: Author's synthesis. Controls should be proportionate to the influence and risk of the use case.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings