1. Introduction
The rapid integration of artificial intelligence (AI) into financial decision-making has changed the logic, pace, and scale of credit allocation. Nowhere is this more visible than in credit scoring, where algorithmic systems increasingly determine access to loans, mortgages, insurance cover, and other forms of financial capital. Although predictive modeling in consumer credit is not new, the shift from transparent statistical scorecards to complex machine-learning (ML) architectures represents a qualitative leap in both technical capability and in the risks that supervisors and firms must manage (Hurley and Adebayo 2017; Binns 2020).
Credit and insurance scores are foundational to economic participation. They influence not only immediate lending and underwriting decisions but also downstream access to housing, education, entrepreneurship, and health-related products (Barocas et al. 2019). Yet these systems are frequently trained on historical data that reflect entrenched social and economic inequalities. As a result, automated scoring can encode, perpetuate, or amplify disparities across ethnicity, class, gender, and geography (O’Neil 2016; Eubanks 2018). For a regulated institution, those disparities are not only an ethical concern: they are a source of legal, conduct, reputational, and model risk that falls squarely within the remit of enterprise risk management and prudential supervision.
Despite this, the literature has paid far less attention to the distributive consequences of scoring systems than to their accuracy. Mainstream work has emphasized precision and discrimination while downplaying the disparate impacts that emerge across social groups (Mitchell et al. 2021), and the dominant industry mitigation—removing sensitive attributes from the data—is now widely recognized as insufficient, because correlated features act as proxies and reproduce bias through indirect pathways (Kim 2019; Binns 2020). Treating fairness as a first-order risk-management objective, rather than a post hoc compliance formality, remains the exception rather than the rule.
This study addresses that gap. It is positioned at the intersection of computational modeling, financial regulation, and institutional risk governance, and its central premise is that an AI scoring system must be evaluated not only for how well it predicts, but for whom it predicts well, at what social cost, and under which supervisory obligations. The analysis is anchored in credit scoring—the canonical high-risk use case—but is framed throughout for the shared banking-and-insurance perimeter, because the same proxy-discrimination mechanisms, and much of the same regulatory architecture, govern insurance pricing and underwriting.
The remainder of this section reviews the state of the field and the supervisory perimeter, and states the research questions.
Section 2 describes the data, variables, modeling pipeline, and the performance and fairness metrics.
Section 3 reports the empirical findings.
Section 4 discusses the risk-management and governance implications.
Section 5 concludes.
1.1. Machine Learning in Credit Scoring and the Model-Risk Problem
The diffusion of AI across financial services is a structural change in the infrastructure of contemporary finance (Arner et al. 2016). From credit-risk modeling to fraud detection and customer segmentation, AI technologies are reshaping banking and insurance at a scale and speed that often outpaces existing regulatory and ethical frameworks (Pasquale 2015; Bussmann et al. 2021). Each efficiency gain, however, introduces or amplifies a risk that must be governed: model risk from opaque estimators; conduct and legal risk from discriminatory outcomes; operational and data risk from alternative data pipelines; and reputational risk from decisions that customers cannot understand or contest.
Historically, credit scoring quantified default likelihood through logistic regression, discriminant analysis, and decision trees; Thomas et al. (2002) established foundational practices balancing predictive accuracy with interpretability—concerns that remain central to today’s AI-governance debates. Hand and Henley (1997) warned early that reliance on complex classifiers risked sacrificing transparency for marginal accuracy gains, a concern intensified by high-dimensional models that are powerful but difficult to audit (Lipton 2018; Rudin 2019). Since the mid-2000s, supervised ML has captured nonlinearities and high-dimensional interactions and has consistently outperformed traditional techniques across credit datasets (Lessmann et al. 2015; Baesens et al. 2003), with gradient boosting performing especially well under class imbalance and noise (Brown and Mues 2012). Behavioral credit scoring further incorporates transaction histories, mobile usage, and digital footprints (Bazarbash 2019; Berg et al. 2020), which advocates argue extends credit to thin-file applicants (Frost et al. 2019).
From a supervisory standpoint, these gains arrive coupled with model risk—the risk of adverse consequences from decisions based on incorrect or misused model output. Model-risk-management frameworks (Board of Governors of the Federal Reserve System 2011; Basel Committee on Banking Supervision 2005) require firms to validate models across their life cycle and to demonstrate conceptual soundness. Where a scoring model produces systematically different outcomes for structurally disadvantaged groups, that disparity is itself a model-risk and conduct-risk finding, not merely an ethical footnote.
1.2. Algorithmic Bias, Proxy Discrimination, and Fairness as a Risk Category
The promise of algorithmic scoring is tempered by the critique that it can mask systemic bias behind a veneer of objectivity. As O’Neil (2016) argues, systems trained on historical records often encode structural discrimination; seemingly neutral features such as ZIP code or education can proxy for ethnicity, gender, or class (Barocas and Selbst 2016; Eubanks 2018). The risks are not merely theoretical: ML systems in credit and insurance have disproportionately penalized minority and low-income applicants even when risk profiles are comparable (Bartlett et al. 2021; Fuster et al. 2022), while opacity complicates redress and makes adverse decisions difficult to contest (Citron and Pasquale 2014).
In this framing, financial inclusion is more than a design objective—it is a question of economic justice, because exclusion from credit constrains housing, entrepreneurship, education, and healthcare (Hurley and Adebayo 2017; Narayan et al. 2018). Discriminatory algorithms can deepen inequality under the guise of computational neutrality (Binns 2018; Gabor and Brooks 2017). The operative consequence for a regulated institution is that fairness must be measured, monitored, and controlled as an explicit risk category, with quantitative metrics, tolerances, and escalation paths—precisely the logic this paper operationalizes.
1.3. Fairness and Discrimination Risk in Insurance
Although the empirical setting is credit, the same mechanisms and much of the same regulatory logic extend to insurance. Insurance pricing and underwriting are, by construction, acts of risk differentiation, which makes the boundary between legitimate actuarial discrimination and unfair discrimination unusually delicate. The Court of Justice of the European Union’s Test-Achats ruling (C-236/09) prohibited the use of gender as a rating factor in EU insurance, establishing that even a statistically predictive attribute can be legally impermissible on equality grounds. As insurers adopt granular behavioral and telematics data and ML pricing, proxy discrimination becomes a first-order concern: a model can reconstruct protected attributes from correlated features and reproduce prohibited distinctions without ever observing them directly.
Supervisors have responded accordingly. The work of the European Insurance and Occupational Pensions Authority (EIOPA) on the governance of AI and on the ethical use of data in insurance emphasizes fairness, non-discrimination, transparency, and human oversight, and recent attention to differential and behavior-based pricing reflects concern that data-driven personalization can produce outcomes that are unfair even when actuarially defensible. The International Association of Insurance Supervisors (2020) has likewise highlighted fairness, accountability, and explainability as supervisory priorities. The fairness metrics, mitigation techniques, and governance controls examined here are therefore directly transferable from credit to insurance underwriting and pricing.
1.4. The Regulatory and Supervisory Landscape
Regulatory frameworks have tightened around automated decision-making. The European Union’s Artificial Intelligence Act classifies creditworthiness assessment and life- and health-insurance risk pricing as high-risk, imposing obligations on data governance, technical documentation, transparency, human oversight, accuracy, and risk management across the model life cycle (European Commission 2021). The General Data Protection Regulation (GDPR) constrains automated decisions with legal or similarly significant effects and underpins expectations of meaningful information about the logic involved (Voigt and Von dem Bussche 2017; Wachter et al. 2017). In the United States, the Equal Credit Opportunity Act (ECOA) and the Consumer Financial Protection Bureau (CFPB) require that algorithmic credit decisions be explainable enough to support specific adverse-action reasons.
Sector supervisors add a prudential layer. The European Banking Authority (2020) guidelines on loan origination and monitoring set expectations for data quality, documentation, and explainability in credit decisions, while model-risk doctrine (Board of Governors of the Federal Reserve System 2011) requires independent validation and effective challenge. On the insurance side, Solvency II governance and risk-management requirements, together with EIOPA and IAIS guidance, extend analogous expectations to underwriting and pricing models. International bodies have issued principles for responsible AI in finance—transparency, robustness, accountability, and human-in-the-loop control (OECD 2019; Financial Stability Board 2017, 2022).
These instruments have fueled interest in explainable AI (XAI). Tools such as SHapley Additive exPlanations (SHAP) (Lundberg and Lee 2017), Local Interpretable Model-agnostic Explanations (LIME) (Ribeiro et al. 2016), and counterfactual explanations (Wachter et al. 2017) help render complex models interpretable. Yet critics caution that explanation methods can oversimplify or misrepresent model logic (Lipton 2018; Rudin 2019) and do not resolve deeper problems when biased data or misaligned objectives drive behavior. Scholars therefore warn of an implementation gap—firms adopting ethical principles rhetorically while optimizing for efficiency, a dynamic labeled ethics-washing (Mittelstadt 2019; Morley et al. 2020).
1.5. Research Questions
This study asks how AI-driven scoring systems in banking and insurance can be developed, validated, and governed so that they support—rather than undermine—financial inclusion, while remaining defensible under the applicable regulatory frameworks. Three interrelated questions follow.
RQ1. To what extent do AI-based scoring models reproduce or mitigate bias against historically marginalized or underrepresented groups (by ethnicity, gender, socioeconomic status, or geography)? This builds on Fuster et al. (2022) and Bartlett et al. (2021), who show that disparities can persist under ostensibly neutral algorithmic regimes.
RQ2. What role do fairness-aware ML techniques play in mitigating algorithmic discrimination, and at what measurable cost to performance? While explanation tools provide transparency, Lipton (2018) and Rudin (2019) caution that explanation alone does not ensure regulatory compliance if the underlying data or objectives remain misaligned.
RQ3. Which governance frameworks and regulatory mechanisms are most effective in ensuring accountability, transparency, and ethical alignment in financial AI across banking and insurance? A persistent implementation gap remains between principles and organizational practice (Mittelstadt 2019; Morley et al. 2020).
The contribution is threefold: it quantifies the fairness cost of model capacity on a single institutional dataset under a harmonized metric set; it prices the accuracy given up by three distinct mitigation families; and it translates both into a control framework aligned with the supervisory expectations that now govern banking and insurance.