Preprint
Article

This version is not peer-reviewed.

The Augmented Behavioural Capacity (ABC) Model: Large Language Model Assistance and Calibrated Retail-Investor Decision-Making

Submitted:

03 July 2026

Posted:

06 July 2026

You are already at the latest version

Abstract
Retail investors increasingly bring large language models (LLMs) into cognitively demanding trading decisions - not as tools but as conversational reasoning partners that help frame problems, test scenarios, and judge feasibility. This creates a calibration problem with direct bearing on trading behaviour: the same assistance can yield nhancement or miscalibrated confidence, deference, and offloading. Existing adoption and intention theories explain why investors engage such systems but leave unspecified the decision-time mechanism through which dialogic assistance enters evaluative reasoning - a gap diagnosed here as a three-layer (temporal, ontological, phenomenological) insufficiency. The Augmented Behavioural Capacity (ABC) Model addresses it as a domain-bounded extension for LLM-assisted retail-investor decision-making, specifying a moderated process architecture of three elements: Capacity, operationalised through Perceived Cognitive Assistance (PCA); Calibration, a diagnostic overlay separating calibrated enhancement from miscalibrated confidence, deference, or offloading; and Choice, the proximal outcomes spanning complexity intention, strategy evaluation, structural complexity, and risk assertion. The method is theoretical model development - construct roles, boundary conditions, rival explanations, and falsifiable propositions. The contribution is bounded: ABC claims neither improved returns nor full causal validation, but offers a specified architecture, partially anchored by prior PCA validation, and a staged agenda for causal, calibration, longitudinal, and behavioural-footprint testing.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  ;  

1. Introduction

Large language models (LLMs) are entering retail-investor decision-making at a point that earlier digital, financial, and decision-support technologies did not reach. They are no longer encountered only as systems to be adopted, configured, and used (Cohn et al., 2024; Colombatto et al., 2025; Heersmink et al., 2024; Mogaji et al., 2024). They are increasingly encountered as conversational reasoning partners with which an investor can structure a problem, compare alternatives, simulate scenarios, request counterarguments, and check the consistency of a candidate course of action - all at the moment of cognitively demanding evaluation, before a trade is placed (Handler et al., 2024; Li et al., 2025; Si et al., 2024; Wang et al., 2025). The phenomenon this paper takes as its starting point is therefore not whether such systems are accepted or used, but how LLM engagement may reshape the evaluative stage (The term evaluative stage refers to the decision phase in which the actor structures the problem, constructs and weighs alternatives, estimates risk, and forms criteria for action. This corresponds to the Design and Choice phases of Simon’s (1960) classical decision-process taxonomy, to the evaluation phase of Kahneman and Tversky’s (1979) prospect-theoretic account of decisions under risk, and to the constructive evaluation processes characterised by Payne, Bettman, and Johnson (1993) as cognitive-effort-bounded strategy selection) of retail-investor decision-making: the phase in which the investor structures the problem, weighs alternatives, estimates risk, and forms criteria for action, corresponding to the design-and-choice phases of Simon’s (1960) decision-process taxonomy.
The dominant inherited frameworks address adjacent but different problems. The Technology Acceptance Model (TAM) and the Unified Theory of Acceptance and Use of Technology (UTAUT) explain whether users judge a system useful, easy to use, socially supported, or controllable, and whether they intend to use it (Davis, 1989; Venkatesh, Morris, Davis, & Davis, 2003; Venkatesh, Thong, & Xu, 2012). The Theory of Planned Behaviour (TPB) explains how attitudes, subjective norms, and perceived behavioural control combine into behavioural intention (Ajzen, 1991). These families remain necessary, because LLMs are still adopted technologies whose use depends on perceived utility, effort, social influence, and control. They are not, however, sufficient: they were not designed to explain what happens to the evaluative content of a user’s reasoning when a generative, dialogic system enters that reasoning at the moment of choice - in effect, they locate the locus of decision-making within the individual user and treat the system as an external object of use rather than as a participant in the evaluation itself. Section 3 develops this gap as a three-layer (temporal, ontological, phenomenological) insufficiency; here it is stated only as the motivation for a domain-bounded theoretical extension rather than a free-standing new construct.
The paper advances that extension as the Augmented Behavioural Capacity Model, hereafter the ABC Model. The label “capacity” is preferred to “control” because the model is not primarily about classical perceived behavioural control in the sense of the TPB. It concerns tool-contingent perceived cognitive capacity: what an investor feels able to understand, compare, simulate, or execute because an LLM is present as a conversational reasoning partner during the decision episode. The central measurable mechanism is Perceived Cognitive Assistance (PCA), the investor’s felt expansion of cognitive capability at the moment of decision (Gimmelberg & Ludviga, 2025, 2026a, 2026c). The model is summarised as a moderated process: LLM engagement may raise PCA; PCA may shift proximal behavioural choice toward greater complexity, risk assertion, or decision-structure elaboration; and Calibration conditions how that movement should be interpreted.
The model foregrounds two linked but distinct claims, stated here in plain terms and specified formally in Section 4. Claim A, the behavioural-migration claim, proposes that LLM engagement may increase PCA and that elevated PCA may move proximal behavioural choice toward greater strategic complexity, risk assertion, or decision-structure elaboration. Claim B, the calibration overlay claim, proposes that Calibration - the match between perceived capacity and objective comprehension, confidence accuracy, verification discipline, reliance, and offloading - conditions whether PCA-driven movement is best read as calibrated enhancement or behavioural distortion. Enhancement in this sense does not mean better realised results or even better ex ante decision quality. It means that elevated perceived cognitive capacity is not merely a self-reported illusion, because it is matched by objective understanding, appropriate confidence, disciplined verification, and retained independent reasoning. Whether such calibrated enhancement also corresponds to higher ex ante decision quality is a separate criterion-validation question for subsequent testing.
The contribution is deliberately bounded in scope and in evidential status, and the two boundaries should be read together. In scope, ABC is offered as a domain-bounded theoretical model of LLM-assisted retail-investor decision-making, not as a universal theory of artificial-intelligence adoption, financial behaviour, or market efficiency. In evidence, the paper does not claim that LLM use improves realised investment returns, which lie outside the model’s explanatory target, and it does not claim full causal validation of every pathway the architecture implies. What is claimed is the theoretical specification of the ABC Model together with partial empirical anchoring of PCA as its central Capacity-layer mechanism. Controlled causal testing, calibration diagnostics, longitudinal refinement, and behavioural-footprint assessment are stated as future validation tasks rather than completed evidence. The remainder of the article follows the journal’s reporting structure, adapted to a model-development study. Section 2 specifies the model-development method and falsification discipline. Section 3, Section 4 and Section 5 develop and present the model, which is the study’s result: Section 3 establishes the phenomenon and the incumbent-theory gap that motivate it; Section 4 sets out the ABC architecture, construct registry, and claims; and Section 5 states the boundary conditions and falsifiable propositions that delimit it. Section 6 discusses and concludes on the contribution, limitations, and evidence.

2. Materials and Methods: Theoretical Model Development

2.1. Theoretical Model-Development Approach

This is a theoretical model-development paper, and its method is the disciplined construction of a domain-bounded model. A model becomes theoretically useful only when it specifies the phenomenon it explains, the constructs involved, the logical relationships among them, the proposed mechanism, and the boundary conditions under which the explanation is expected to hold (Bacharach, 1989; Dubin, 1978; Wacker, 1998; Whetten, 1989). Theory and model are not treated as interchangeable: a theory is a broader explanatory system of constructs and propositions, while a model is a domain-bounded representation of that explanation in a specified setting, with explicit constructs, links, mechanisms, assumptions, and limits (Bacharach, 1989; Reynolds, 1971; Wacker, 1998). On this basis the ABC Model is advanced as domain-bounded rather than universal.
This distinction sets the standard for model development: theoretical argument integrates intellectual ancestry, propositions, evidence, and relationship logic into a coherent explanatory account of constructs, mechanisms, and boundaries (DiMaggio, 1995; Sutton & Staw, 1995; Weick, 1989; Whetten, 1989).
Three safeguards follow from this logic and discipline the construction throughout. First, the model fixes its level and unit of analysis, because moving between an individual belief, a decision episode, a repeated behavioural pattern, and a market-level outcome without justification produces category errors and ecological or atomistic fallacies (Bacharach, 1989; Klein, Dansereau, & Hall, 1994; Kozlowski & Klein, 2000; Rousseau, 1985). Second, every model element is classified by its theoretical role before propositions are written, because a construct, latent factor, indicator, antecedent, mechanism, moderator, diagnostic variable, outcome, boundary condition, and extension variable are not interchangeable (Edwards & Bagozzi, 2000; MacKenzie, Podsakoff, & Podsakoff, 2011; Podsakoff, MacKenzie, & Podsakoff, 2016; Suddaby, 2010). Third, each link is assigned a formal status - causal, mediating, moderating, sequential, reciprocal, or diagnostic - rather than asserted as mere association, since these alternatives imply different theoretical claims and different empirical tests (Baron & Kenny, 1986; MacKinnon, 2008; Pearl, 2009; Whetten, 1989). These safeguards are necessary precisely because ABC combines inherited adoption constructs, a validated cognitive-assistance construct, calibration diagnostics, behavioural outcomes, and a downstream footprint extension, which could otherwise be misread as equivalent factors.

2.2. Model-Development Sequence Used in This Article

The article follows the same development sequence used to build the model, because each step depends on the ones before it: a model cannot be defensibly tested before its constructs are defined, its level of analysis fixed, its relationship types stated, and its limits made explicit (Bacharach, 1989; Dubin, 1978; Suddaby, 2010; Wacker, 1998). The sequence proceeds from defining the problem and contribution, to stating the development method, establishing the phenomenon and the theory gap, fixing scope and unit of analysis, classifying constructs and model elements, specifying the architecture and relationship logic, stating boundary conditions and rival explanations, deriving falsifiable propositions, and finally stating the contribution, limitations, and evidence status (Corley & Gioia, 2011; Lakatos, 1970; Popper, 1959; Sutton & Staw, 1995; Weick, 1989; Whetten, 1989). Figure 1 presents this sequence as a nine-step model-development flow.
The flow records the logical order in which the model is built and reported, not a claim of empirical completion at each step: the steps that define, classify, and specify the model are completed to the level required for theoretical anchoring, whereas the proposition-development step separates claims already anchored by completed PCA measurement validation from those that remain open to future causal, calibration, longitudinal, and behavioural-footprint testing. This distinction matters because PCA is validated as the model’s central measurement backbone, while full causal validation of the entire ABC architecture is not claimed. The full step-by-step development workflow and its methodological authorities are documented in Appendix A.

2.3. Falsification Logic and Evidence-Status Discipline

The model is disciplined by a refutation-oriented logic adapted from the philosophy of science. Two commitments are operative. First, core claims are distinguished from auxiliary, developmental, and extension claims, so that what would weaken or refute the model’s centre is not hidden among peripheral commitments; refutation conditions are stated at the level of the architecture as well as at the level of individual propositions (Lakatos, 1970; Popper, 1959; Chiambaretto, Fernandez, & Le Roy, 2025; Rubin, 2025). Second, the relationship logic specifies not only that constructs are related but how, why, and within what boundaries, which is a precondition for determinate theoretical testing rather than open-ended association (Pearl, 2009; Whetten, 1989; Busse, Kach, & Wagner, 2017; von Nordenflycht, 2023; Carton, 2025). Auxiliary modifications that merely absorb disconfirming observations without clarifying the core or generating testable refinement are treated as a failure of discipline rather than a defence of the model (Lakatos, 1970; Rubin, 2025).
Applied here, this discipline produces a precise and limited claim. When a phenomenon is theoretically novel and its constructs are still being stabilised, the defensible contribution is to define the phenomenon, specify the model, validate the central measurement mechanism, and state a falsifiable agenda for subsequent tests, rather than to claim a full causal test of every pathway in a single study (Edmondson & McManus, 2007; Lynham, 2002; Van de Ven, 2007). Scale-development and construct-validation guidance makes the same point: content validation, factor-analytic confirmation, discriminant testing, and early predictive evidence can establish a measurement backbone without by themselves validating an entire causal system (Clark & Watson, 2019; Cronbach & Meehl, 1955; MacKenzie, Podsakoff, & Podsakoff, 2011; Messick, 1995). The article therefore claims theoretical specification and partial empirical anchoring, and explicitly not full causal validation; failed or untested elements are reported as such rather than deferred into vague future work.

3. The LLM Phenomenon and the Incumbent-Theory Gap

3.1. LLMs as Evaluative-Stage Participants

LLMs differ from the technologies against which inherited adoption models were calibrated. They are dialogic rather than static, generative rather than retrieval-bound, explanatory rather than merely informative, context-retaining across turns, and capable of participating in the framing of the problem itself (Handler, Larsen, & Hackathorn, 2024; Sabbah & Li, 2025; Tankelevitch et al., 2024; Wang, Lu, & Pan, 2025). Recent work extends the classical extended-mind and scaffolding perspective specifically to LLMs, treating them as interactive cognitive artefacts and scaffolds rather than tools whose relevant properties are fixed before use (Heersmink, de Rooij, Clavel Vázquez, & Colombo, 2024; Kudina, Ballsun-Stanton, & Alfano, 2025; Smart, Clowes, & Clark, 2025). These properties matter because the locus of theoretical interest is the evaluative stage: interpretation, comparison, risk estimation, perceived capability, and perceived manageability of the action. It is at this stage, and not only at the point of adoption, that an LLM can alter the content of an investor’s reasoning (Jakesch, Bhat, Buschek, Zalmanson, & Naaman, 2023; Si et al., 2024).
This locus is consequential in retail finance specifically, because retail investors are documented to trade on overconfidence, to under-diversify, and to be sensitive to information presentation and financial-literacy constraints (Agnew & Szykman, 2005; Barber & Odean, 2000, 2001; Lusardi & Mitchell, 2014). A system that can elaborate, reframe, and appear to validate a strategy at the moment of evaluation engages exactly the cognitive and affective vulnerabilities that this literature identifies, which is why the evaluative stage rather than the adoption decision is the model’s point of entry.
This contrast is summarised visually in Figure 2 across six evaluative-stage dimensions that distinguish LLMs from ordinary adoption-relevant tools.

3.2. Necessary but Incomplete Inherited Theories

The inherited frameworks remain necessary. TAM, UTAUT explain access, perceived usefulness, ease of use, social influence, facilitating conditions, and continued use (Davis, 1989; Venkatesh et al., 2003, 2012), and TPB explains how attitude, subjective norm, and perceived behavioural control combine into intention (Ajzen, 1991). UTAUT2 is the member of this family closest to the present setting: it extends UTAUT to the consumer context with hedonic motivation, price value, and habit (Venkatesh et al., 2012), and retail investors using a consumer technology fall squarely within its scope. A standing critical tradition has nonetheless questioned how far these frameworks travel: commentators have asked whether TAM has reached the limits of its explanatory yield (Bagozzi, 2007; Benbasat & Barki, 2007; Lee, Kozar, & Larsen, 2003; Straub & Burton-Jones, 2007), debated whether TPB should be retired or retained (Ajzen, 2015; Conner, 2015; Sniehotta, Presseau, & Araújo-Soares, 2014), and more recently asked whether generative artificial intelligence (AI) strains the acceptance paradigm itself (Davis & Granić, 2024; Mogaji, Viglia, Srivastava, & Dwivedi, 2024; Schittko, Planing, & Müller, 2026). The point ABC takes from this tradition is narrow: not that the inherited frameworks are wrong, but that they do not specify the tool-contingent perceived cognitive capacity that operates inside the evaluative episode.
What the critique tradition establishes is that the inherited frameworks are reaching the limits of their original formulations; what it does not supply is a positive specification of the decision-time mechanism that LLMs introduce. The published critiques diagnose strain - in TAM’s explanatory yield, in TPB’s coverage of behaviour beyond intention - but they do not name the construct that would carry the missing explanation. ABC’s contribution begins where those critiques stop: rather than retiring or defending the inherited frameworks, it accepts them at the adoption boundary and specifies the evaluative-stage mechanism they leave unaddressed.

3.3. The Three-Layer Gap: Temporal, Ontological, Phenomenological

The insufficiency can be diagnosed at three layers. The temporal layer is a substrate problem: the inherited theories were calibrated against stable, determinate, single-function tools, whereas LLMs are probabilistic, dialogic, and non-stationary across model generations, so constructs designed for fixed-function software do not straightforwardly transfer (Mogaji et al., 2024; Ouyang, Wu, Jiang, Almeida, Wainwright, Mishkin, et al., 2022). The ontological layer concerns where cognition is located: inherited models place reasoning inside the user and treat the technology as an evaluated external object, whereas LLMs partly co-produce the reasoning, a configuration the distributed-cognition and scaffolding tradition anticipates but the acceptance tradition does not encode (Clark & Chalmers, 1998; Hutchins, 1995; Risko & Gilbert, 2016; Wood, Bruner, & Ross, 1976). The phenomenological layer concerns mechanisms that arise in LLM-mediated interaction and have no clean home inside inherited constructs.
Eight such mechanisms can be named in compressed form. Hallucination is the generation of fluent but unfounded content (Ji, Lee, Frieske, Yu, Su, Xu, et al., 2023). Sycophancy is the tendency of dialogue-tuned systems to agree with or flatter the user’s framing, a known by-product of preference-based fine-tuning (Malmqvist, 2024; Ouyang et al., 2022; Sharma, Tong, Korbak, Duvenaud, Askell, Bowman, et al., 2024). Anthropomorphic dialogic interaction elicits social responses to a non-social system (Epley, Waytz, & Cacioppo, 2007; Liew, Lim, Khan, & Tan, 2025; Nass & Moon, 2000). Stochastic non-determinism means the same prompt can yield materially different outputs across runs, so the system the investor evaluates is not fixed. Epistemic opacity with interactional transparency means the system reads as clear and confident while its internal basis is unavailable to the user (Glikson & Woolley, 2020). Asymmetric trust calibration is the mismatch between warranted and actual reliance documented in the automation and advice-taking literatures, where users variously over-rely on or discount machine advice (Dietvorst, Simmons, & Massey, 2015; Lee & See, 2004; Logg, Minson, & Moore, 2019; Parasuraman & Manzey, 2010). Context-window adaptation means the option set shown at a later turn is shaped by framings introduced earlier in the same conversation. Co-construction of the option set is the user-side counterpart: alternatives, criteria, and sometimes attitudes are produced inside the dialogue rather than brought to it, so that an attitude TPB would treat as a stable antecedent may instead be generated within the episode (Jakesch, Bhat, Buschek, Zalmanson, & Naaman, 2023).
The argument here is not that every mechanism lacks intellectual ancestry. Several have clear lineages in metacognition, automation, and distributed cognition (Alter & Oppenheimer, 2009; Lee & See, 2004; Risko & Gilbert, 2016; Rozenblit & Keil, 2002). The argument is that the LLM-mediated configuration activates these mechanisms jointly, inside a single evaluative episode, in a way the inherited models can only patch fragment by fragment by appending one moderator at a time. Perceived ease of use illustrates the cost concretely: calibrated against determinate-output tools, it tracks the friction of obtaining a response but not the difficulty of verifying it, so under LLMs it can register highest in exactly the conditions of greatest over-reliance risk, leaving the user no within-construct signal to separate the two (Buçinca, Malaya, & Gajos, 2021; Heersmink et al., 2024; Vasconcelos et al., 2023). Figure 3 compresses this diagnosis: it aligns the temporal, ontological, and phenomenological layers with the specific feature of LLM-mediated reasoning that inherited adoption and intention theories were not built to carry, and is the basis for treating the gap as structural rather than incremental. The expanded critique tradition and the full mechanism-by-mechanism support are documented in Appendix A.

3.4. Construct Proliferation and the Implication for a Capacity–Calibration–Choice Model

A further symptom supports the diagnosis. The construct-proliferation audit by Gimmelberg and Ludviga (2026b) reinforces this diagnosis by showing that the neighbouring construct space is theoretically active and useful for boundary-setting, while still pointing toward the need for a new theoretical model. The human–AI and GenAI behaviour literature has produced a rapid expansion of partially overlapping constructs - AI self-efficacy and literacy, cognitive absorption, metacognitive demand and workload, task–technology fit, perceived understanding, and dependence or over-reliance, among others (Almeida Lima, Bellei, Ballesteros Martins, & Terlizzi, 2024; Goh, Hartanto, & Majeed, 2025; Huy, Nguyen, Vo-Thanh, Thinh, & Thi Thu Dung, 2024; Sarraf, Kar, & Janssen, 2024; Tankelevitch et al., 2024; Wang, Lu, & Pan, 2025; Wang, Rau, & Yuan, 2023; Wang & Chuang, 2024; Xue, Ghazali, & Mahat, 2025). Read through the construct-clarity lens, this proliferation is itself evidence of theoretical strain: when a phenomenon outruns its inherited constructs, the field tends to multiply neighbouring constructs rather than respecify the mechanism (Shaffer, DeGeest, & Li, 2016; Suddaby, 2010). The implication ABC draws is that the appropriate response is not a further loose construct appended to the inherited frameworks, but a domain-bounded model that names the construct doing the explanatory work, the apparatus that conditions its interpretation, and the proximal outcome family in which the effect appears.

4. Results: The ABC Model: Construct Registry, Architecture, and Claims

4.1. Scope and Unit of Analysis

The unit of analysis is the evaluative decision episode: a self-directed retail investor using an LLM to construct, weigh, or revise a candidate trading strategy, at or near the moment of a cognitively demanding decision. The scope is bounded accordingly. The model addresses retail-investor decisions, not professional, institutional, or fully advised contexts; cognitively demanding tasks, not routine lookup or fully delegated robo-advisory allocation; and proximal choice, not realised returns. Two further bounds are stated as conditions of theoretical applicability rather than as limitations. The model is formulated against the dialogue-tuned, instruction-following LLM capability frontier of 2024–2026; parameters tied to perceived competence - in particular the level and calibration of PCA - are temporally fragile as hallucination rate, sycophancy intensity, reasoning depth, and context-window size move across releases, so findings produced against any one release should be re-anchored as the frontier advances. Parameters may also differ across jurisdictions where LLM-based advisory output is classified or regulated differently, so the present formulation is bounded to its current cultural and regulatory context rather than asserted as cross-jurisdictionally invariant. Empirically, the model is formulated for cross-sectional testing at the episode level, complemented by longitudinal tracking of trust calibration, reliance escalation, and skill atrophy; specific operationalisation choices are a separate empirical question and are not constitutive of the theoretical claim. Fixing the unit at the decision episode locates the model’s elements within a bounded context without collapsing them into a single sequential mediation chain, which matters because the elements occupy different theoretical roles.

4.2. Construct and Element Registry

ABC is organised around five elements, classified by theoretical role rather than treated as five equivalent factors. Adoption and Engagement is the inherited upstream boundary: the conditions of access, uptake, and use carried over from TAM, UTAUT and UTAUT2, and TPB, accepted rather than re-tested here (Ajzen, 1991; Davis, 1989; Venkatesh et al., 2003, 2012). Capacity is the ABC core construct, operationalised as a single latent construct through PCA, the felt expansion of cognitive capability at decision time. PCA measures the user’s perceived expansion of cognitive capability in the presence of an LLM; the extended-mind and scaffolding tradition supplies the motivation for why such a construct should exist as a separate measurement target, but ABC does not claim, at the level PCA measures, that cognition is empirically distributed across the user–system boundary. Calibration is a heterogeneous diagnostic and moderating apparatus, not a unified construct, comprising objective understanding, perceived understanding, overprecision, reliance and deference, and state offloading (Lee & See, 2004; Moore & Healy, 2008; Risko & Gilbert, 2016; Rozenblit & Keil, 2002). Choice is a family of proximal behavioural outcomes - complexity intentions, strategy evaluation, structural complexity of intended trades, and risk assertion - rather than a single dependent variable. Behavioural Footprint is a downstream extension into realised trading behaviour, documented but not part of the defended core.
The structural asymmetry across the three Cs is a substantive feature, not a presentational convenience: Capacity is a construct, Calibration is a diagnostic and moderating apparatus, and Choice is an outcome family. Treating the three as parallel latent factors would misrepresent the architecture and reproduce the construct-confusion the model is meant to resolve. The full operational indicators and the formal treatments are documented in Appendix A.

4.3. The ABC Architecture

The architecture is a five-layer process. The three core stages - Capacity, Calibration, and Choice - are the model’s theoretical claims; the two framing stages - Adoption and Engagement upstream, Behavioural Footprint downstream - are boundary architecture. Defending an asymmetric core follows directly from theory-building practice, which requires a model to mark what it claims, what it inherits, and what it extends, with the boundaries between those zones visible rather than concealed (Bacharach, 1989; Suddaby, 2010; Wacker, 1998; Whetten, 1989). Five layers, rather than three or four, are the minimum that lets the inherited adoption tradition appear upstream and the empirical-extension question appear downstream while keeping both at the boundary: a three-layer core would silently inherit adoption conditions without naming them, and a four-layer version would name one boundary while concealing the other.
Three core concepts, rather than four or five, are retained because each names a distinct kind of theoretical object: Capacity is the construct, Calibration the diagnostic gate, and Choice the manifestation. Other candidate “Cs” that recur in the literature reduce to one of these three or sit outside the core. Confidence and comprehension are not separate stages but components of Calibration’s battery, operationalised as overprecision and as objective understanding (Moore & Healy, 2008; Rozenblit & Keil, 2002). Control is an inherited neighbouring belief from the TPB and UTAUT traditions rather than an ABC core stage; ABC theorises the narrower tool-contingent cognitive-capacity subset that the inherited control construct does not capture (Ajzen, 1991; Venkatesh et al., 2003). Complexity functions as a modulator of Choice and a task-bounding scope condition rather than a stage of its own, and conduct is downstream, at the Footprint boundary. The three Cs are thus the minimum that distinguishes the construct, the gate, and the manifestation; further Cs are not suppressed but located where they belong.
Within the core, the relationship logic is a moderated decision process. PCA captures perceived cognitive capacity; Calibration conditions how that perceived capacity translates into proximal choice. Calibration is specified primarily as a moderator rather than as a second latent construct, because the model’s own language - high PCA matched by objective understanding supports calibrated enhancement, unmatched supports behavioural distortion - is moderation language, and because the moderator form is the most statistically tractable for an early-stage programme and does not foreclose a later latent-class specification if subgroup discontinuities emerge (Aiken & West, 1991; Baron & Kenny, 1986; Hayes, 2018; MacKinnon, 2008; Mathieu & Taylor, 2006; McCutcheon, 1987). Because the five calibration variables are not directionally homogeneous, the apparatus is modelled as a set of sign-aligned miscalibration gaps, with a configural component-specific specification retained as an alternative; a naive additive composite is excluded, because it can mix adequacy and distortion signals and produce an uninterpretable null through cancellation. Figure 4 displays the five-layer architecture, the three-stage core, and the placement of Claim A and Claim B.

4.4. Claim A and Claim B

Claim A (behavioural-migration claim) specifies the theoretically proposed Capacity→ Choice path: LLM engagement may increase PCA, and elevated PCA may shift proximal behavioural choice toward greater strategic complexity, risk assertion, or decision-structure elaboration.
Claim B (calibration overlay claim) specifies the overlay sitting above that path: the match between perceived cognitive capacity and its objective referents across the five calibration diagnostics - objective understanding, perceived understanding, overprecision, reliance and deference, and state offloading - conditions whether the same PCA-driven movement is interpreted as calibrated enhancement or behavioural distortion. Calibrated enhancement means that elevated PCA is matched by objective understanding, confidence accuracy, verification discipline, and retained independent reasoning. Behavioural distortion means that elevated PCA is paired with weak understanding, overprecision, uncritical deference, perceived understanding without objective comprehension, or unverified offloading.
The two claims are linked but distinct. Claim A concerns movement toward complexity; Claim B concerns the calibration-contingent interpretation of that movement. Claim B does not claim that enhancement produces better decisions. It claims that elevated PCA can be classified as calibrated enhancement rather than illusion or distortion when the relevant diagnostic referents are present. Whether calibrated enhancement then corresponds to higher ex ante decision quality is a separate criterion-validation test, not a premise of the classification itself. The model therefore does not posit two separate causal mechanisms so much as a single Capacity → Choice mechanism whose behavioural meaning changes across calibration conditions. Claim B is, at the moderator level, construct-general: whatever raises PCA, its onward effect on Choice is moderated by calibration, so Claim B takes elevated PCA as given rather than tracing its origin.

4.5. Current Evidentiary Status

The two claims are anchored asymmetrically, and the asymmetry is reported rather than smoothed. The Capacity layer is the most developed: PCA has been content-validated against perceived usefulness, perceived ease of use, trust in the LLM, and trading self-efficacy, confirmed as a one-factor scale with adequate reliability and scalar invariance across trading-experience and recency strata, and shown to predict complexity intentions incrementally over inherited proximal comparators, with the PCA–usefulness boundary close but defensible on confirmatory and discriminant criteria (Gimmelberg & Ludviga, 2026a, 2026c). Claim A is therefore partly anchored at the component level - through PCA measurement, one-step incremental prediction, and discriminant validity - but its end-to-end directed path is not confirmed as a mediation chain. Claim B is theoretically specified, with a stated architectural refutation condition, but is not operationally tested, because the matched-domain distortion-diagnostic battery it requires - objective comprehension, overprecision, perceived understanding, reliance and deference, and state offloading at the resolution of the decision episode - is not yet stabilised as a coherent instrument package in retail-investor research. The test of Claim B is therefore a calibration-classification test: whether elevated PCA is matched by objective and behavioural diagnostics of understanding, confidence accuracy, verification, and retained reasoning, or instead by diagnostics of overprecision, deference, and offloading. A later link from calibrated enhancement to higher ex ante decision quality would provide criterion-validity evidence for the classification, but the model does not claim that link in the present article.

5. Boundary Conditions and Falsifiable Propositions

5.1. Boundary Conditions

Boundary specification is part of theoretical precision rather than supplementary limitations prose: a model that holds for all values of all variables is empirically empty, so the conditions under which propositions are asserted must be stated explicitly (Bacharach, 1989; Busse, Kach, & Wagner, 2017; Dubin, 1978; Whetten, 1989). Five boundaries fix the value-ranges across which the ABC propositions are asserted. The number five is model-derived rather than method-prescribed: these are the five points at which ABC’s propositions can be misapplied - phenomenon, population, task context, mechanism, and outcome - while cultural/regulatory context and model-generation era are treated as generalisability questions rather than additional proposition boundaries. The phenomenon boundary restricts the claims to LLM-assisted evaluative reasoning, not technology adoption in general; findings bearing only on adoption, intention, or use fall outside the range. The population boundary restricts them to self-directed retail investors, not professional, institutional, or fully advised actors. The task-context boundary restricts them to cognitively demanding trading decisions, not routine lookup or fully delegated allocation; that low-complexity tasks show weaker effects is itself a proposition rather than a free parameter. The mechanism boundary restricts them to perceived cognitive capacity via PCA, moderated by Calibration, and not to objective ability, skill, or performance; the model’s distinctive predictions concern PCA conditional on calibration, not PCA alone. The outcome boundary restricts them to proximal behavioural choice, not realised execution or return. A contradicting finding outside these ranges is not a refutation; a contradicting finding inside them is.

5.2. Rival-Explanation Logic

The boundary step also asks whether neighbouring accounts could make ABC redundant, mis-scoped, or empirically misleading, following the convergent–discriminant and falsificationist traditions (Campbell & Fiske, 1959; Edwards & Bagozzi, 2000; Lakatos, 1970; Popper, 1959). A construct-proliferation audit (Gimmelberg & Ludviga, 2026b) classified candidate rivals into four threat families - direct substitutes, antecedents and confounders, distortion mechanisms, and background theories - because each requires a different empirical response rather than a single test. Seven external rival families remain on record for boundary treatment: AI self-efficacy, cognitive absorption, cognitive load and workload, illusion of explanatory depth, AI dependence and over-reliance, task–technology fit, and explanation quality or perceived understanding (Agarwal & Karahanna, 2000; Goh, Hartanto, & Majeed, 2025; Goodhue & Thompson, 1995; Leppink et al., 2013; Rozenblit & Keil, 2002; Wang & Chuang, 2024). A substantive result of the audit is that several of these are currently blocked from direct same-level psychometric contest, because matched-domain instruments do not yet exist for retail-investor LLM decision episodes; stating that blockage openly, rather than presenting only the deployable contests, is the discipline the falsification logic requires (Lakatos, 1970). The full audit and rival matrix are documented in Appendix A.

5.3. Architecture-Level Refutation Conditions

Four refutation conditions are stated at the level of the architecture, so that what would refute the model’s core is visible rather than buried in a longer list (Lakatos, 1970; Popper, 1959; Bacharach, 1989; Wacker, 1998). First, construct redundancy: the model fails as a contribution if PCA cannot be distinguished from its proximal comparators - perceived usefulness, ease of use, trust, and trading self-efficacy - on factor-analytic or predictive grounds under stronger designs; closeness is expected, absorption is fatal (Cronbach & Meehl, 1955; Campbell & Fiske, 1959; MacKenzie et al., 2011; Shaffer et al., 2016). Second, misclassification of rivals: the model is weakened both by ignoring the external rival families and by testing any of them as a direct comparator before a matched-domain instrument exists, because rival explanations must be classified by theoretical role, level, and measurement parity before they can adjudicate construct boundaries (Bacharach, 1989; Suddaby, 2010; MacKenzie et al., 2011; Podsakoff et al., 2016). Third, failure of the calibration overlay: the model is weakened at the architectural level if PCA predicts the same proximal-choice pattern regardless of calibration diagnostics - that is, if calibration does not moderate the PCA→Choice relationship - a condition that depends on the miscalibration-gap specification, since a naive additive composite could manufacture a null through cancellation (Aiken & West, 1991; Hayes, 2018; Mathieu & Taylor, 2006; Diamantopoulos & Winklhofer, 2001; Jarvis et al., 2003). This condition is currently programmatic rather than deployable, because the required battery does not yet exist. A later failure to link calibrated enhancement with independently scored decision quality would limit the criterion validity of the classification, but it would not by itself refute the calibration classification claim. Fourth, failure to survive calibration-only alternatives: if calibration diagnostics jointly predict choice while PCA adds no residual role and does not interact with them, PCA is at most a surface appraisal accompanying calibration rather than the mechanism through which engagement becomes capacity; this condition follows from treating calibration-only specification as an architectural rival and testing whether the focal mechanism survives competing structural explanations (Pearl, 2009; MacKinnon, 2008; Whetten, 1989).

5.4. Falsifiable Propositions and Validation Agenda

The audit is translated into seven propositions in three classes - Core (P1–P5), Developmental (P6), and Extension (P7) - each with a determinate refutation condition and a required empirical operation (Lakatos, 1970; Popper, 1959). P1 holds that interactive LLM support raises PCA above static or non-dialogic tools and is the antecedent precondition for Claim A. P2, the component-level test of Claim A, holds that PCA predicts proximal choice incrementally over inherited TAM/UTAUT/TPB comparators. P3 holds that task complexity moderates the PCA→Choice path. P4, a discriminant prerequisite shared by both claims, holds that PCA remains distinguishable from perceived usefulness, dependence, and passive offloading. P5, the direct test of Claim B, holds that elevated PCA is interpreted as calibrated enhancement when matched by objective understanding, confidence accuracy, verification discipline, and retained independent reasoning, and as behavioural distortion when paired with weak understanding, overprecision, uncritical deference, perceived understanding without objective comprehension, or unverified offloading. P6 holds that proficiency and adoption maturity strengthen the productive pathway and reduce distortion. P7 extends the claims into realised trading behaviour at the Footprint boundary (Barber & Odean, 2000, 2001).
The evidence status differs across propositions and is reported separately rather than aggregated. P1 and P3 are open; P2 and P4 are partly anchored by completed PCA validation work; P5 is theoretically specified but instrument-blocked because the matched-domain calibration-diagnostic battery does not yet exist; P6 is developmental; and P7 is an extension. A later decision-quality test may provide criterion-validity evidence for P5, but it is not treated as a current theoretical claim.
The discipline of the set is shown by what it declines to claim: a free-standing mediation proposition of the form LLM use → PCA → choice is deliberately omitted from the core list, because that chain was tested cross-sectionally and not detected, with the bias-corrected indirect-effect interval including zero (Gimmelberg & Ludviga, 2026c). The model treats this non-detection as a constraint on Claim A’s status rather than absorbing it through auxiliary moderation; whether the path is misspecified or merely attenuated in early-stage adoption is left open (MacKinnon, 2008; Mathieu & Taylor, 2006; Quintana, 2024), and P3 and P6 are advanced as independently motivated propositions, not as recovery routes for the unconfirmed mediation. The full proposition-and-refutation matrix is documented in Appendix A.

6. Discussion and Conclusions

The Results sections should be read together as theoretical results of the model-development process rather than as empirical outcome results. The first result is diagnostic: LLM-assisted retail-investor decision-making exposes a gap in inherited adoption and intention frameworks because the system enters the evaluative stage of judgement rather than remaining only an object of use. The second result is architectural: ABC specifies that gap through a Capacity–Calibration–Choice structure, with Perceived Cognitive Assistance as the Capacity-layer mechanism and Calibration as the diagnostic overlay that gives behavioural meaning to PCA-driven movement. The third result is disciplinary: the model is bounded by explicit scope conditions, rival explanations, refutation conditions, and propositions. Taken together, the results position ABC not as proof that LLMs improve investor decisions, but as a specified architecture for studying how dialogic AI assistance may alter perceived decision-time capacity and how that alteration should be interpreted.
The ABC Model contributes a domain-bounded theory of LLM-assisted retail-investor evaluative decision-making. Its theoretical move is to name the mechanism that inherited adoption and intention frameworks reach but do not specify — tool-contingent perceived cognitive capacity, measured as PCA — and to make the interpretation of that capacity conditional on calibration rather than monotonically beneficial. In doing so it reformulates an assumption shared across TAM, UTAUT, UTAUT2 and TPB, that more positive perception of a technology is uniformly beneficial, into a calibration-contingent distinction between calibrated enhancement and behavioural distortion (Ajzen, 1991; Davis, 1989; Venkatesh et al., 2003; Venkatesh et al., 2012). Motivating that theoretical move is a separate diagnostic contribution: a three-layer specification - temporal substrate, ontological cognition-locus, and phenomenological mechanism set - of why inherited adoption theories cannot host LLM-mediated evaluative reasoning. Existing taxonomies of LLM-specific phenomena are risk-organised (Weidinger et al., 2022) or determinant-organised (Eigner & Händler, 2024), and do not group mechanisms by the measurement-boundary crossings of inherited adoption constructs that the eight-mechanism phenomenological layer maps.
This interpretation has two implications for inherited theory. First, ABC does not reject TAM, UTAUT, UTAUT2, or TPB. It relocates them. They remain appropriate for explaining access, adoption, use, usefulness, effort, social influence, control, and intention. Their limitation appears when the theoretical object shifts from adoption of a tool to evaluative reasoning with a dialogic system. Second, ABC treats positive perception of LLM assistance as theoretically ambiguous. Higher perceived cognitive capacity is not assumed to be beneficial by itself. Its interpretation depends on whether perceived capacity is matched by objective understanding, confidence accuracy, verification discipline, and retained independent reasoning. The same apparent movement toward greater strategy complexity can therefore have different behavioural meanings under different calibration conditions.
The practical relevance follows from that distinction. By separating calibrated enhancement from miscalibrated confidence, the model supplies a vocabulary and a falsifiable frame for product, platform, and regulatory audiences deciding how dialogic AI assistance should be embedded in retail-trading interfaces, and PCA offers a deployable instrument for that work. The methodological contribution is to state boundaries, rival explanations, and architecture-level refutation conditions explicitly, rather than to add another loose construct to an already proliferating field (Corley & Gioia, 2011; Shaffer et al., 2016; Suddaby, 2010; Whetten, 1989).
The limitations are intrinsic to the contribution and are stated as such. The model is domain-bounded and does not transfer, without further work, to professional populations, non-financial tasks, or different regulatory and cultural settings (Bacharach, 1989; Busse et al., 2017; Wacker, 1998). It claims no improvement in realised returns or investor welfare, which lie outside its defended core. Its empirical anchoring is moreover a snapshot in a non-stationary landscape, because both the LLM substrate and retail-investor adoption are evolving rapidly (Heersmink et al., 2024; Mogaji et al., 2024; Ouyang et al., 2022); the architecture is intended to be substrate-agnostic, but specific findings will require re-anchoring as the environment matures.
The evidence status of the two claims is deliberately asymmetric. Claim A is partially anchored through PCA measurement and one-step incremental prediction, but is not confirmed as an end-to-end mediation; this non-detection is reported openly rather than explained away. Claim B is theoretically specified with a stated refutation condition, but is not yet operationally tested; its immediate test is whether elevated PCA can be classified as calibrated enhancement or behavioural distortion using matched-domain calibration diagnostics, while any link to ex ante decision quality remains a later criterion-validation question. The validation agenda that follows is determinate: controlled causal tests of the Capacity→Choice path, stabilisation of the calibration battery and a direct test of the moderation claim, longitudinal tests of proficiency effects, and, only if the model is asked to speak to realised behaviour, assessment of the behavioural-footprint extension. The model is offered in a form built to be tested, refined, and where warranted refuted, rather than defended against evidence.

Author Contributions

Conceptualization, D.G. and I.L.; methodology, D.G. and I.L.; formal analysis, D.G.; investigation, D.G.; writing - original draft preparation, D.G.; writing—review and editing, I.L.; supervision, I.L. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Data Availability Statement

The structured literature-retrieval, source-screening, and rival-scoping materials supporting the conclusions of this article are described in Appendix A and will be made available by the corresponding author upon reasonable request.

Conflicts of Interest

The authors declare no conflicts of interest.

Declaration of generative AI in the writing process

During the preparation of this work the authors used the ChatGPT (OpenAI) and Claude (Anthropic) to assist with drafting and language editing, to improve readability, and to organise and structure content. These tools were also used to support source searching and cross-checking; the structured, LLM-assisted literature-retrieval and rival-scoping procedures underlying the 2023–2026 thematic review are reported in full, in a reproducible form, in Appendix A (§A3.1.1, §A3.1.2, and §A6.2). After using these tools, the authors reviewed and edited the content as needed and take full responsibility for the content of the publication.

Appendix A. Methodological and Evidentiary Audit Trail for the ABC Model

A.1. Appendix Content Mapping

Appendix A provides the methodological and evidentiary audit trail for the ABC Model. It documents the model-development workflow and methodological authorities behind the compressed presentation in the main article; reports the literature-search protocol and evidence base supporting the LLM-specific evaluative-stage argument; clarifies the status of the model’s elements, including Capacity, Calibration, Choice, Adoption/Engagement, and Behavioural Footprint; specifies the architecture and relationship logic linking them; and records the boundary conditions, rival explanations, refutation logic, and proposition-level validation agenda. The appendix does not introduce a second model or an additional empirical test; it explains how the ABC architecture was constructed, bounded, and translated into falsifiable claims.

A.2. Model-Development Safeguards and the Nine-Step Flow

A.2.1. Level and Unit of Analysis

First, a social-science behavioural model must state its level and unit of analysis. A model may explain an individual belief, a decision episode, a repeated behavioural pattern, a dyadic interaction, an organisational process, or a market-level outcome. Moving between these levels without justification creates category errors, weakens causal interpretation, and may produce ecological or atomistic fallacies (Rousseau, 1985; Klein, Dansereau, & Hall, 1994; Kozlowski & Klein, 2000). Level specification is therefore not a narrow organisational-research concern; it is part of general theoretical precision in behavioural and social-science modelling (Bacharach, 1989; Whetten, 1989; Johns, 2006).
The primary unit of analysis in the core of the ABC Model is the individual self-directed retail investor at or near the moment of a cognitively demanding trading decision. The model is concerned with how LLM engagement may change the investor’s perceived cognitive capacity during the decision episode, and how calibration determines whether that perceived capacity supports empowered or distorted choice. Market-level traces, longitudinal behavioural patterns, and trading-log footprints are downstream extensions, not the primary level at which the model is first defended.

A.2.2. Classification of Model Elements

Second, every model element must be classified before propositions are written. In this article, a model element means any named unit that appears in the architecture: a layer, stage, construct, latent factor, indicator, antecedent, mechanism, moderator, diagnostic variable, outcome, boundary condition, or extension variable. These roles are not interchangeable. A construct is a theoretical attribute that may be measured by indicators. A latent factor is a statistical measurement dimension estimated from multiple indicators. An indicator is an observed item or behavioural marker. An antecedent is an upstream condition. A mechanism explains how an effect occurs. A moderator specifies when or for whom a relationship changes. A diagnostic variable helps distinguish alternative process states. An outcome is the dependent domain. A boundary condition specifies where the explanation is expected to hold (Edwards & Bagozzi, 2000; MacKenzie et al., 2011; Podsakoff, MacKenzie, & Podsakoff, 2016; Suddaby, 2010).
This classification is especially important for ABC because the model is not a five-factor model. The five ABC layers are positions in a process architecture, not psychometric factors. PCA is the validated factor-like construct inside the cognitive-capacity layer. Calibration is a moderating and diagnostic apparatus, not yet a validated latent factor. Choice is an outcome domain, not a latent factor by default. Adoption and engagement are inherited upstream conditions from the technology-adoption and behavioural-intention traditions. Behavioural footprint is a downstream realised-behaviour extension, not a psychological construct (Bollen, 1989; Fabrigar, Wegener, MacCallum, & Strahan, 1999; Brown, 2015). Treating all of these as generic “factors” would collapse layers into constructs, constructs into indicators, mechanisms into outcomes, and diagnostics into causal variables. Such collapse is exactly what construct-clarity guidance warns against (Suddaby, 2010; MacKenzie et al., 2011; Podsakoff et al., 2016).

A.2.3. Specification of Relationship Types

Third, relationship-building must specify the formal status of each link. Some links may be causal, some mediating, some moderating, some sequential, some reciprocal, and some diagnostic. These alternatives imply different theoretical claims and different empirical tests (Baron & Kenny, 1986; MacKinnon, 2008; Mathieu & Taylor, 2006; Shadish, Cook, & Campbell, 2002; Pearl, 2009). A model that states that two elements are “related” without specifying the relationship type is not yet theoretically precise. It may be useful as an early map, but it is not sufficient as a defensible scientific model (Whetten, 1989; Wacker, 1998; Bacharach, 1989).
In the ABC Model, the core is a moderated decision-process architecture. LLM engagement may produce perceived cognitive capacity through PCA, and PCA may shift proximal behavioural choice. Calibration does not sit between PCA and Choice as a sequential mediator. It is a diagnostic and moderating apparatus that conditions the interpretation of the Capacity → Choice relationship. This relationship logic justifies the separation between ABC layers, validated constructs, calibration diagnostics, behavioural outcomes, and downstream footprint indicators.

A.2.4. Why the Article Uses a Nine-Step Model-Development Flow

The principles introduced above do not, by themselves, dictate a single sequence in which a theory-building paper must be written. They do, however, generate a small number of recurring requirements that any disciplined model-development paper must discharge. A model-development paper must identify the phenomenon to be explained and the contribution claimed (Whetten, 1989; Corley & Gioia, 2011); state an explicit method for building the model, so that the model is not read as a rearrangement of familiar variables (Sutton & Staw, 1995; Weick, 1989; Wacker, 1998); demonstrate the explanatory gap left by existing frameworks (Whetten, 1989; Dubin, 1978); specify scope, level of analysis, and unit of analysis (Bacharach, 1989; Rousseau, 1985; Klein, Dansereau, & Hall, 1994; Kozlowski & Klein, 2000); clarify constructs and classify the status of model elements before propositions are written (Suddaby, 2010; MacKenzie, Podsakoff, & Podsakoff, 2011; Podsakoff, MacKenzie, & Podsakoff, 2016); present a relationship architecture explaining how the constructs are connected and why those connections should exist (Bacharach, 1989; Whetten, 1989; Wacker, 1998); state boundary conditions and rival explanations (Bacharach, 1989; Busse, Kach, & Wagner, 2017; Campbell & Fiske, 1959); formulate falsifiable propositions linked to future empirical operations (Popper, 1959; Lakatos, 1970; Wacker, 1998); and close by specifying contribution, evidence status, limitations, and communication of what is and is not claimed (Whetten, 1989; Sutton & Staw, 1995; Corley & Gioia, 2011; Edmondson & McManus, 2007).
The nine-step flow used in this article is therefore not proposed as a universal theory-building template. It is a paper-specific writing and model-construction sequence that maps these recurring requirements onto a single linear order suitable for the ABC Model. For this article, the number of steps follows from the decision not to collapse analytically distinct requirements into one another. The research problem is separated from the method; the LLM-specific phenomenon is separated from the inherited TAM/UTAUT/TPB gap; scope is separated from construct classification; construct classification is separated from relationship architecture; relationship architecture is separated from boundary and rival-explanation analysis; and proposition development is separated from final contribution and limitation claims. This separation matters because the ABC Model combines inherited adoption constructs, the validated PCA measurement mechanism, calibration diagnostics, behavioural-choice outcomes, and downstream footprint indicators. Treating these as one undifferentiated set of “factors” would reproduce the conceptual confusion that construct-clarity and theory-development guidance explicitly warns against (Suddaby, 2010; MacKenzie et al., 2011; Podsakoff et al., 2016).
The order of the steps is also not arbitrary. Later steps presuppose earlier ones. The phenomenon must be stated before the gap can be evaluated; the model-development method must be stated before the model can be judged; scope and unit of analysis must be fixed before construct status can be classified; constructs and indicators must be classified before relationships can be drawn; relationship types must be specified before boundary conditions and rival explanations can be assessed; and boundaries and rivals must be considered before propositions can be sharpened into falsifiable form (Bacharach, 1989; Whetten, 1989; Wacker, 1998; Popper, 1959). The flow therefore plays a dual role. It is the logical sequence through which the ABC Model is constructed, and it is the rhetorical sequence through which the article presents that construction.
Because Figure 1 in the main article already visualises the nine-step sequence, Appendix A does not repeat the diagram. Table A1 reports the methodological function and main authorities for each step.
Table A1. Model-development workflow and methodological authorities.
Table A1. Model-development workflow and methodological authorities.
Step Model-development function Main methodological support
1. Research problem and contribution Defines the problem and contribution target (Corley & Gioia, 2011; Van de Ven, 2007; Whetten, 1989)
2. Model-development method States what counts as model development and what is not theory (Bacharach, 1989; Sutton & Staw, 1995; Wacker, 1998)
3. Phenomenon and theory gap Establishes why LLM-mediated evaluative reasoning requires extension beyond TAM/UTAUT/TPB (Ajzen, 1991; Davis, 1989; Dubin, 1978; Venkatesh et al., 2003; Venkatesh et al., 2012; Whetten, 1989)
4. Scope and unit of analysis Locks population, decision context, task domain, and level of inference (Johns, 2006; Klein et al., 1994; Kozlowski & Klein, 2000; Rousseau, 1985)
5. Construct and element registry Separates layers, constructs, factors, indicators, diagnostics, outcomes, and extensions (Edwards & Bagozzi, 2000; MacKenzie et al., 2011; Podsakoff et al., 2016; Suddaby, 2010)
6. ABC architecture and relationship logic Specifies how Capacity, Calibration, and Choice relate inside the five-layer architecture (Dubin, 1978; Pearl, 2009; Wacker, 1998; Whetten, 1989)
7. Boundary conditions and rival explanations States where the model applies and what alternative explanations must be considered (Bacharach, 1989; Busse et al., 2017; Campbell & Fiske, 1959; Whetten, 1989)
8. Falsifiable propositions and validation agenda Converts architecture into claims that could be weakened or refuted (Bacharach, 1989; Lakatos, 1970; Popper, 1959; Wacker, 1998)
9. Contribution, limitations, and communication States what is claimed, what is not claimed, and what remains open (Corley & Gioia, 2011; Edmondson & McManus, 2007; Sutton & Staw, 1995; Whetten, 1989)

A.3. Phenomenon and Theory Gap: Why LLMs Affect Evaluative Reasoning Differently

A.3.1 Thematic Literature Search: Protocol and Evidence Base

The phenomenon-and-theory-gap argument developed in this section rests on a corpus of scholarly literature assembled through a deliberate, replicable search protocol rather than through convenience reading. The protocol is specified first; the resulting evidence base is reported afterwards. This ordering - protocol before output - is followed for each model-development step in this paper, in line with the methodological logic stated in §A2.
A.3.1.1. Protocol for an LLM-Mediated Thematic Literature Search
The phenomenon-and-theory-gap argument in §A3.2–§A3.3 rests on a corpus assembled through a replicable, version-controlled protocol. A thematic literature search rather than a pooled-effect synthesis is the appropriate method because the purpose is to map an emerging theoretical domain, identify recurrent mechanisms, and organise fragmented evidence into an explanatory argument, not to estimate effect sizes (Webster & Watson, 2002; Torraco, 2005; Braun & Clarke, 2006). Saturation was assessed at the level of thematic contribution rather than raw source count: additional sources were retained only if they introduced a new mechanism or materially changed an existing dimension of the argument (Saunders et al., 2018).
An earlier publication by the present author and collaborator (Gimmelberg & Ludviga, 2025) used a retrieval-augmented review of more than 120 post-2023 papers to map the broader literature on LLMs in retail-investor contexts; that evidence is not duplicated here. The present analytic task is different - how LLMs differ from ordinary adoption-relevant technologies in entering the user’s evaluative reasoning process - and requires foundational cognitive-science, decision-theory, and human–computer-interaction literature falling outside the post-2023 window of the prior review. A separately executed, structured retrieval was therefore required.
A single, version-controlled research prompt specified the focal phenomenon, the analytic question, seven scoping literature domains, peer-reviewed inclusion criteria, and an APA-citation output format. Holding the prompt constant across runs is the precondition for treating the runs as comparable retrieval trials. The prompt was executed three times in fresh sessions by architecturally and provider-distinct large language models operating in deep-research mode: Run 1 returned 48 references, Run 2 returned 62, and Run 3 returned 35.
The combined retrieval surface was then put through a four-operation screen: intra-run de-duplication; a hosting-class filter retaining canonical scholarly venues and excluding blogs, Medium posts, Scribd documents, and similar non-canonical sources; a topical-relevance filter against the six specificity dimensions used for downstream matrix mapping; and cross-run de-duplication. The per-operation counts are reported in Table A2. Manual bibliographic and content validation - metadata against publisher records and DOI registries, confirmation of peer-reviewed status, and confirmation that the source substantively addresses one or more matrix dimensions - was concentrated on the screened pool of 60 references rather than on the 145 raw retrievals. Each verified reference was then tagged for the dimension(s) it supports on the basis of its primary argument or finding, not tangential association; per-dimension counts in Table A3 are supporting-paper counts, not effect-size aggregates.
Table A2. Evidence-base counts from the thematic literature search.
Table A2. Evidence-base counts from the thematic literature search.
Step Run 1 Run 2 Run 3 Combined Treatment
Raw retrievals returned 48 62 35 145 Combined raw retrieval surface across three independent LLM-driven runs.
After intra-run de-duplication 48 58 35 141 Four redundant entries collapsed within Run 2.
Stage A (hosting / venue class) 48 34 35 117 24 Run 2 items excluded as non-canonical hosting.
Stage B (topical relevance to six dimensions) 48 21 35 104 13 further Run 2 items excluded as topically off-target.
After cross-run de-duplication - - - 73 23 Run 1–Run 3 overlaps and 8 Run 2 overlaps with Runs 1 / 3 collapsed.
Final formal-APA support set - - - 60 13 verified-relevance Run 2 items documented but not promoted; reasons in Appendix A. Manual bibliographic and content validation applied to the 60-reference pool.
Note. Stage A retains references whose hosting class indicates a canonical scholarly venue. Stage B additionally retains only references that are topically on target for at least one of the six specificity dimensions used for downstream matrix mapping. Cross-run de-duplication collapses (i) the 23 references appearing in both Run 1 and Run 3 and (ii) the 8 Run 2 verified-relevance items overlapping with references already cited from Runs 1 or 3. The 13 unique verified-relevance Run 2 items not overlapping with Runs 1 or 3 are documented in Appendix A but not promoted to the formal-APA support set. Manual bibliographic and content validation (Step 5) is applied to the final pool only.
The protocol is subject to the LLM-specific risks identified theoretically in §A3.3.4, and the verification operations are mapped to those risks rather than to a generic notion of AI risk. Hallucinated or fabricated references (Ji et al., 2023; Walters & Wilder, 2023) are caught by bibliographic verification and the hosting filter. Sycophancy and prompt-anchoring effects (Sharma et al., 2024; Malmqvist, 2024) are mitigated by holding the prompt constant across runs and withholding the six-dimension specificity matrix from the retrieval models, so that downstream mapping is not produced by the systems being audited. Fluency-induced false confidence is constrained by restricting support to references whose primary finding substantively addresses the claimed dimension; LLM-generated summaries are not used as evidence. Substrate overlap between the three models is recorded directly in the cross-run de-duplication counts rather than dissolved. The retrieval is therefore reported as LLM-assisted search-surface expansion with a verification spine, not as triangulated retrieval whose triangulation is presumed independent. Every author–year reference cited in Table A3 also appears in the References section, every per-dimension count equals the number of distinct references in the corresponding cell.
A.3.1.2. Resulting Evidence Base
Application of the protocol to the present analytic question produced the corpus summarised in Table A2.
The resulting set of 60 unique scholarly references constitutes the empirical anchor for the academic narrative in §A3.2, the LLM specificity matrix in Table A3, and the conceptual heatmap in Table A4. As noted in Step 6, the per-dimension counts shown in Table A3 are counts of supporting references for each dimension; they are not effect sizes, not meta-analytic aggregates, and not additive across the matrix, because a single reference may legitimately support more than one dimension. The mapped n total across the six dimensions is 71, against the unique-source pool of 60; the implied average of approximately 1.18 dimension-mappings per reference reflects the fact that most references support exactly one dimension and a minority support two.

A.3.2. Phenomenon and Gap Analysis

Large language models (LLMs) affect the evaluative stage of decision-making differently from ordinary adoption tools because they are not only objects of adoption or external aids. Classical technology-adoption theories explain whether users judge a system as useful, easy to use, socially supported, or controllable (Ajzen, 1991; Davis, 1989; Venkatesh et al., 2003, 2012). Decision-support and fit theories add that technologies can improve task performance when system capabilities and information representations fit the task (Goodhue & Thompson, 1995; Todd & Benbasat, 1999; Vessey, 1991). Explanation research further shows that intelligent systems can justify recommendations and change user evaluations of system output (Gedikli et al., 2014; Gregor & Benbasat, 1999). These literatures remain necessary, but they are incomplete for LLM-assisted judgement because they mostly preserve a separation between the human evaluator and the technological instrument: the tool retrieves, displays, filters, recommends, or automates, while the user remains the main site of interpretation.
LLMs weaken that separation. Their distinctive interface is natural-language, multi-turn, generative, and context-sensitive. The user can externalise a partial thought, receive a structured answer, challenge the answer, request alternatives, ask for counterarguments, and continue the same reasoning thread. Recent decision-support scholarship therefore treats LLMs as a general-purpose decision-support technology rather than another narrow information system, and the naturalness of conversational AI affects users’ cognitive effort and ambiguity (Handler et al., 2024; Wang et al., 2025). Instruction-following and chain-of-thought prompting make the interaction appear reasoning-like, even when the underlying system is not reasoning in a human sense (Ouyang et al., 2022; Wei et al., 2022). Sundar’s account of machine agency, work on social responses to computers, and anthropomorphism research explain why this modality is psychologically consequential: the user increasingly communicates with the machine rather than merely through the machine (Epley et al., 2007; Glikson & Woolley, 2020; Nass & Moon, 2000; Sundar, 2020). Heersmink et al. (2024) sharpen this point for LLMs by arguing that they are sufficiently transparent at the interactional level to support fluent use, yet sufficiently opaque internally to make trust calibration difficult.
The relevant decision point is the evaluative stage, not the adoption stage. Evaluation is where the actor interprets an ambiguous situation, forms a mental representation of what is at stake, compares alternatives, estimates risk, and asks whether action is intelligible and feasible. Bounded rationality and heuristics mean that this stage is always simplified and cue-driven rather than fully optimal (Kahneman & Tversky, 1974; Simon, 1955; Slovic, 1987). It is also self-regulatory: people form beliefs about whether they understand the issue and whether they are capable of acting, and contemporary metacognition research emphasises that confidence judgements are inferential and can diverge from objective performance (Bandura, 1977; Fleming, 2024). In retail investing, the problem is intensified by information overload, uneven and partly subjective financial literacy, attention-driven trading, and recurrent overconfidence (Agnew & Szykman, 2005; Barber & Odean, 2000, 2001; Bellofatto et al., 2018; Lusardi & Mitchell, 2014). Therefore, a technology that changes how the investor builds the decision frame can affect more than adoption intention; it can change the investor’s judgement coordinate itself.
The first pathway is empowerment through cognitive scaffolding. LLMs can structure an analytical path, translate a broad view into candidate tactics, compare alternatives side by side, produce what-if scenarios, and highlight inconsistencies. This is why the strongest theoretical anchor is not ordinary usefulness alone but scaffolding, distributed cognition, extended cognition, and cognitive offloading (Clark & Chalmers, 1998; Hollan et al., 2000; Hutchins, 1995; Risko & Gilbert, 2016; Wood et al., 1976). In this interpretation, the LLM functions as a conversational cognitive artefact: it does not simply provide an answer, but helps organise the user’s reasoning activity. Human-AI decision research supports the importance of this distinction. Cognitive forcing functions can reduce overreliance, explanations can shape reliance, and LLM explanations can increase verification efficiency even when accuracy gains are limited or overreliance appears when the model is wrong (Buçinca et al., 2021; Schemmer et al., 2023; Si et al., 2024). Naturalness of conversational interaction further matters because it changes felt cognitive effort and the experience of ambiguity (Wang et al., 2025).
The second pathway is distortion through metacognitive miscalibration. The same features that make LLMs useful as scaffolds can also inflate perceived understanding, confidence, and control. People often infer competence from fluency, coherence, and processing ease (Alter & Oppenheimer, 2009; Koriat, 1997). They can overestimate what they understand, become overprecise in belief, or mistake a plausible explanation for genuine comprehension (Moore & Healy, 2008; Rozenblit & Keil, 2002). LLMs intensify this risk because they present coherent, human-like, testimony-style output; anthropomorphic cues, mental-state attributions, AI confidence, and verbalised uncertainty can all affect trust, perceived accuracy, or user self-confidence (Cohn et al., 2024; Colombatto et al., 2025; Epley et al., 2007; Li et al., 2025; Xu et al., 2025). Explanations are not a guaranteed cure: in clinical decision support, fuller explanations increased trust but also produced over-reliance (Bussone et al., 2015), and AI-generated explanations can be more persuasive than bare classifications and can amplify belief in misinformation when the explanations are deceptive (Danry et al., 2024). Co-writing with opinionated language models can shift users’ expressed views and even subsequent attitudes (Jakesch et al., 2023). Trust-in-automation research further warns that users can misuse, disuse, or overuse automation when reliance is poorly calibrated (Dzindolet et al., 2003; Glikson & Woolley, 2020; Lee & See, 2004; Parasuraman & Manzey, 2010; Parasuraman & Riley, 1997). Hallucination and sycophancy risks add a specifically generative-AI version of the same problem (Ji et al., 2023; Malmqvist, 2024).
Table A3 reports the resulting LLM specificity matrix. Each row pairs one of the six analytic dimensions specified a priori in §A3.1 with the supporting references mapped to that dimension under the protocol described in Step 6, and the n column gives the count of distinct sources tagged to the dimension.
Table A3. LLM specificity matrix.
Table A3. LLM specificity matrix.
Specificity dimension Key finding Supporting references (each entry corresponds to one back-matter reference) n Relevance for PCA and ABC model
Dialogic and contextual interaction LLMs make the interface conversational and adaptive. The user can ask, challenge, refine, and continue the reasoning thread instead of only retrieving or accepting a fixed output. Sundar (2020); Nass and Moon (2000); Heersmink et al. (2024); Ouyang et al. (2022); Wei et al. (2022); Cohn et al. (2024); Colombatto et al. (2025); Xu et al. (2025); Logg et al. (2019); Epley et al. (2007); Glikson and Woolley (2020); Wang et al. (2025) 12 Places the system inside the decision episode as a machine interlocutor rather than as a static adopted tool.
Generative and explanatory reasoning LLMs generate explanations, comparisons, counterarguments, and stepwise reasoning-like text, but the same generativity can also create plausible but wrong, sycophantic, or persuasively misleading explanations. Gregor and Benbasat (1999); Gedikli et al. (2014); Ji et al. (2023); Ouyang et al. (2022); Wei et al. (2022); Si et al. (2024); Malmqvist (2024); Handler et al. (2024); Marchionini (2006) 9 Explains why LLM assistance is not just retrieval: it supplies interpretive structure and justification.
Evaluative co-construction and problem framing Evaluation is the stage in which the actor interprets ambiguity, compares alternatives, judges risk, and assesses whether action is manageable. LLMs can participate directly in this construction of the decision frame, including by shifting users’ expressed views. Simon (1955); Kahneman and Tversky (1974); Slovic (1987); Ajzen (1991); Davis (1989); Goodhue and Thompson (1995); Vessey (1991); Todd and Benbasat (1999); Bandura (1977); Fleming (2024); Jakesch et al. (2023) 11 Bridges technology-adoption logic with behavioural decision theory and metacognitive monitoring.
Cognitive scaffolding and offloading LLMs can reduce felt cognitive load, externalise partial thoughts, structure multi-step reasoning, and support comparison and scenario simulation in ways consistent with distributed and extended cognition. Wood et al. (1976); Hutchins (1995); Clark and Chalmers (1998); Risko and Gilbert (2016); Buçinca et al. (2021); Schemmer et al. (2023); Ma et al. (2024); Li et al. (2025); Hollan et al. (2000); Wang et al. (2025) 10 Provides the theoretical anchor for Perceived Cognitive Assistance as felt expansion of decision-time capability.
Metacognitive calibration and distortion risk Fluent, confident, human-like explanations can inflate perceived understanding, certainty, and reliance even when the underlying reasoning remains weak or unverified; explanations themselves can amplify acceptance of incorrect or deceptive claims. Alter and Oppenheimer (2009); Koriat (1997); Rozenblit and Keil (2002); Moore and Healy (2008); Dzindolet et al. (2003); Lee and See (2004); Parasuraman and Riley (1997); Parasuraman and Manzey (2010); Dietvorst et al. (2015); Logg et al. (2019); Cohn et al. (2024); Colombatto et al. (2025); Li et al. (2025); Xu et al. (2025); Bussone et al. (2015); Danry et al. (2024); Fleming (2024); Glikson and Woolley (2020) 18 Requires calibration diagnostics alongside PCA so empowerment is not confused with distorted confidence.
Ecosystem maturity and user familiarity Retail investors face overload, uneven literacy (including subjective literacy), attention-driven trading, and overconfidence. Ordinary tools mainly improve access, filtering, automation, or adoption intention; they do not usually co-construct evaluative reasoning. Agnew and Szykman (2005); Apesteguia et al. (2020); Barber and Odean (2000); Barber and Odean (2001); Brenner and Meyll (2020); Compeau and Higgins (1995); Davis (1989); Lusardi and Mitchell (2014); Venkatesh et al. (2003); Venkatesh et al. (2012); Bellofatto et al. (2018) 11 Explains why LLM effects are especially important in self-directed trading and investing contexts.
Note. The n column reports the count of supporting references mapped to each dimension and equals the number of distinct author-year entries listed in the adjacent supporting-papers cell; counts are not additive across rows because a single reference may support more than one dimension. Every reference cited in any cell of this table also appears in the References section at the end of this document. The mapped n total across the six dimensions is 71 (12 + 9 + 11 + 10 + 18 + 11), against a unique-source pool of 60. The implied average of approximately 1.18 dimension-mappings per reference reflects the fact that most references support exactly one dimension and a minority support two. This is by construction: a reference contributes to a dimension only when its primary argument or finding is substantively about that dimension, not whenever it could be tangentially associated with one.
Table A4 makes the contrast with ordinary adoption tools explicit. The cells record conceptual synthesis scores assigned by the author against a five-point capability rubric: 1 indicates that the dimension is absent from or outside the technology’s design space; 2 that it is present only incidentally or as a side-effect of other features; 3 that it is supported but not a central design property; 4 that it is supported and well-developed; and 5 that it is a defining design property of the technology class. Scoring is single-rater and is presented as a synthesis aid rather than an empirical measurement, but each cell value is corroborated by two independent evidence streams: the 60-reference base anchoring Table A3, where the per-dimension reference counts give the LLM-side scores their footing in the literature search; and prior published work by the author and collaborator on LLM-versus-baseline cognitive asymmetries in retail-investor decision-making (Gimmelberg & Ludviga, 2025), which provides parallel evidence for the ordinary-tools-versus-LLMs contrast across the six dimensions. The synthesis is reported here to anchor the visual contrast in Figure 2 (reported in the manuscript body) rather than as an independent empirical contribution.
Table A4. Conceptual heatmap of ordinary tools versus LLMs across the six specificity dimensions.
Table A4. Conceptual heatmap of ordinary tools versus LLMs across the six specificity dimensions.
Dimension Ordinary tools LLMs Interpretation
Dialogic and contextual interaction 1 5 LLMs sustain multi-turn natural-language exchange and retain context across turns; ordinary tools present fixed interfaces.
Generative and explanatory reasoning 2 5 LLMs generate explanations, alternatives, and stepwise rationales rather than presenting only fixed outputs.
Evaluative co-construction and problem framing 1 5 LLMs participate directly in defining the problem, comparing alternatives, and shaping the decision frame.
Cognitive scaffolding and offloading 2 5 LLMs can structure reasoning steps and reduce felt complexity; this is the design property that anchors Perceived Cognitive Assistance (PCA).
Metacognitive calibration and distortion risk 2 5 LLMs can inflate perceived understanding, certainty, and reliance through fluency and human-like output; higher score indicates greater exposure to distortion, not a better outcome.
Ecosystem maturity and user familiarity 5 3 Ordinary tools are the established, well-tested retail-investor toolkit - search, dashboards, automated execution, robo-advice - embedded in the everyday decision workflow and operated with native, low-friction familiarity. LLMs are converging on this territory from two directions: they are absorbing tool-like features (retrieval, structured and visual output, dashboard-style summaries) and are increasingly embedded inside the incumbent platforms themselves. But they are newer to the role and users have not yet built comparable fluency, so the established toolkit retains supremacy on this axis. This is also the most temporally unstable of the six dimensions: the gap is expected to narrow as feature-integration and embedding advance.
Figure 2 (reported in the manuscript body) visualises the same scores as a radar plot. Plotting the two profiles on shared axes makes the dominant pattern visible at a glance: ordinary adoption tools dominate on the ecosystem-maturity-and-familiarity axis, where they constitute the established retail-investor toolkit; LLMs dominate on the dialogic, generative, evaluative, scaffolding, and calibration-risk axes, which together define the LLM-specific decision-time mechanism the ABC Model is designed to capture.
A.3.3. Limits of Existing Theories: Why TAM, UTAUT, and TPB Are Necessary but Incomplete
The case advanced here is that the established acceptance and intention-based models - Davis’s (1989) Technology Acceptance Model (TAM), Venkatesh, Morris, Davis, and Davis’s (2003) Unified Theory of Acceptance and Use of Technology (UTAUT), Venkatesh, Thong, and Xu’s (2012) UTAUT2, and Ajzen’s (1991) Theory of Planned Behaviour (TPB) - are necessary but incomplete for the analysis of large-language-model-mediated decision-making. The frameworks examined here - TAM, UTAUT, UTAUT2, and TPB - are selected on three convergent criteria. First, they are the dominant inherited frameworks in the AI-acceptance and behavioural-intention literatures, accounting for the bulk of the construct vocabulary that LLM-acceptance studies currently re-use. Second, together they span the adoption → intention → behaviour pipeline that LLM-mediated decision-making appears to disrupt. Third, each has been independently named by the published critique tradition (§A3.3.1) as facing strain under generative AI. The selection is therefore not contestable on grounds of arbitrary target choice: these are the frameworks the field itself has identified as the locus of difficulty.
The case is built in three layers. The temporal layer concerns the empirical substrate against which these theories were calibrated and the substrate on which they are now being applied. The ontological layer concerns the implicit assumptions the theories make about where cognition lives, what counts as a tool, and how intention forms. The phenomenological layer concerns the appearance, in current AI-mediated activity, of mechanisms with no principled location in the inherited construct space. None of the three layers refutes the inherited theories within their original scope. Together they specify why their application to LLM-mediated behaviour at the evaluative stage of decision-making is an extrapolation rather than a deployment, and why a domain-bounded extension is required.
The three layers are organised from least to most demanding. The temporal layer asks whether the inherited theories were ever calibrated against the empirical substrate to which they are now being applied; this is the weakest form of objection and does not refute the theories within their original scope. The ontological layer asks whether the constructs the theories carry can in principle host the new phenomena; this is a stronger objection because it concerns the theories’ implicit assumptions rather than their domain of application. The phenomenological layer asks whether the empirical phenomena that LLM-mediated activity actually produces have a principled location in the inherited construct space; this is the strongest form of objection because it identifies named mechanisms the inherited frameworks cannot host. The three layers are exhaustive in the sense that any disciplined critique of an inherited theory under a new technology must operate at one of these three levels.
The argument is positioned within an established critical tradition rather than offered against an unanimous orthodoxy. A series of published critiques has already documented limitations of each of the three frameworks; some of those critiques explicitly anticipate the type of difficulty the present analysis develops. The first task of this section is therefore to show what the existing critique tradition has already established, what it has not yet specified, and where the present argument adds precision.
A.3.3.1. The Published Critique Tradition: What Is Already on the Record
A first body of critique is internal to the technology-acceptance research programme. The 2007 special issue of the Journal of the Association for Information Systems titled “Quo Vadis TAM” brought together several senior commentators on the trajectory of TAM-based research. Benbasat and Barki (2007) argued that the intense focus on TAM had diverted research attention from other important issues in information-systems use, had produced an illusion of cumulative progress, and had - through repeated, uncoordinated extension - generated a state of theoretical confusion in which it was no longer clear which version of TAM was the canonical one. Bagozzi (2007), in the same issue, offered a more direct theoretical critique: TAM treats the user as a rational evaluator of stable belief inputs; it omits emotional, group, social, and cultural processes; it does not distinguish goal pursuit from action pursuit; and it provides no apparatus for the self-regulation that is known to mediate the gap between intention and action. Bagozzi extended the same critique to UTAUT, observing that its expansion to a model with around forty independent variables for predicting intention and at least eight for predicting behaviour had pushed the acceptance literature towards a “stage of chaos” rather than towards explanatory consolidation. Goodhue (2007), Straub and Burton-Jones (2007), and Lee, Kozar, and Larsen (2003) added complementary points concerning the “more use is better” assumption, the risk of common-method artefacts in TAM-style measurement, and the historical drift of the literature towards an explanatory monoculture. Williams, Rana, and Dwivedi (2015) and Dwivedi, Rana, Jeyaraj, Clement, and Williams (2019) consolidated these concerns in a systematic review and a revised UTAUT proposal that explicitly trimmed and reorganised the model in response to the critique tradition. The position these critiques converge on is not that TAM, UTAUT, or TPB cannot be extended; manifestly they can, and have been, many times. The position is that repeated extension does not on its own produce explanation: it can equally well attach further variables to a frame whose original mechanism no longer carries the phenomenon being studied.
A second body of critique is external to the information-systems literature. In health psychology, Sniehotta, Presseau, and Araújo-Soares (2014) argued that TPB had outlived its usefulness as a complete model of behavioural prediction. The decisive point in their argument is that TPB rests on a sufficiency hypothesis - the claim that all theory-external influences on behaviour are fully mediated by attitudes, subjective norms, and perceived behavioural control - and that this hypothesis has been falsified by evidence that beliefs and contextual factors predict behaviour over and above intentions, and that age, socio-economic status, environmental features, and physical and mental health continue to predict behaviour after TPB constructs are controlled. Ajzen (2015), Conner (2015), Schwarzer (2015), and Gollwitzer and Oettingen (2015) responded to defend, extend, or partially endorse the critique. Ajzen’s defence acknowledged that the theory does not preclude the addition of further predictors and was developed in part by adding perceived behavioural control to the prior theory of reasoned action; that defence is itself consistent with the position that the original three constructs do not exhaust the relevant influences on behaviour, even on a charitable reading of TPB. The relevance of this exchange for the present analysis is that LLMs are precisely the kind of theory-external influence that TPB’s mediation structure does not absorb cleanly: an LLM contributes to the formation of attitudes, the option set on which they are formed, and the perception of control, all simultaneously and inside the deliberative episode rather than as a stable upstream input.
A third body of critique is specific to generative artificial intelligence. Mogaji, Viglia, Srivastava, and Dwivedi (2024) argued that TAM, in its classical and extended forms, faces five concrete limitations in the era of generative AI: an individual-centric perspective that does not capture the inherently relational character of conversational AI use; a limited scope that ties acceptance to a single act of adoption rather than to ongoing co-production with the system; a static nature that does not accommodate evolving and bidirectional engagement over time; cultural applicability constraints that are difficult to handle within a small set of moderators; and a reliance on self-reported measures whose validity is strained when users are uncertain about what they are accepting and what the system is doing on their behalf. The authors recommend embedding TAM within domain context, integrating industry-specific factors, and combining it with alternative methodologies, while remaining open to the possibility that other models will be needed.
The editorial framing of a recent special issue in the Journal of University Teaching and Learning Practice - under the title “Are Technology Acceptance Models still fit for purpose?” - places the same question on a journal-level agenda. The diagnosis has continued past Mogaji et al. (2024) across several venues. Davis and Granić (2024), in a SpringerBriefs retrospective on three decades of TAM, propose a NeuroIS-based path as the future development of the model; the gesture is consequential because Davis was the originator of TAM, and the book reads as a programmatic concession from inside the tradition rather than an external critique. Scherer (2024) presses the case for stronger longitudinal causal-inference designs, a critique to which the acceptance literature’s predominant cross-sectional survey methodology is directly exposed. None of these contributions, however, identifies the cognitive-capacity mechanism that operates inside the deliberative episode itself; the response remains methodological - biometric measurement, longitudinal designs, more careful causal identification - rather than mechanistic.
Taken together, this published critique tradition establishes three points that the present analysis can take as common ground rather than as contested claims. First, the inherited acceptance and intention-based theories have known limitations and have been the object of serious internal and external critique for at least two decades. Second, the limitations are particularly visible when the theories are applied to AI and, more sharply, to generative AI. Third, the standard response in the literature - repeated extension, addition of moderators, and patch-style construct accumulation - has itself been flagged as problematic, both in 2007 (Benbasat & Barki, 2007; Bagozzi, 2007) and in 2024 (Mogaji et al., 2024). What this tradition has not yet specified, however, is the locus of the LLM-specific effect, the cognitive-capacity mechanism that operates at that locus, or the gating logic by which that mechanism produces empowered or distorted behavioural outcomes. That is the gap the present model is designed to occupy. The remainder of this section makes the three-layer argument that defines the gap.
A.3.3.2 Temporal Layer: The Substrate Problem
TAM was developed against a substrate of deterministic, single-purpose enterprise software. Its outputs were authored, in the strict sense that any given screen or report could in principle be traced to a known specification produced by an identifiable human team; its behaviour for any given input was repeatable; and its evaluation was meaningfully ex ante, in the sense that a prospective user could form a belief about “whether this system will help me” before committing to use. UTAUT and UTAUT2 synthesised eight prior adoption models and extended their reach into consumer technology, but the validation samples that anchored the synthesis still engaged with information systems whose outputs remained determinate and whose function was meaningfully tied to a specified task. TPB, although developed in social psychology rather than in information systems, shares a closely related implicit assumption about substrate stability: the action-object whose performance is to be predicted is presumed to have known properties, so that beliefs about it can be coherently formed in advance.
Generative, dialogic, probabilistic systems were not part of the empirical world against which any of these theories were calibrated. The substrate has changed in at least four respects that bear directly on the formal apparatus of the inherited theories. First, the same prompt may yield different outputs across runs and across contexts, so that the system’s behaviour is not deterministic in the classical sense; this strains any construct that presumes a stable expectancy about what the system will produce. Second, current LLMs are general-purpose: a single system is used across writing, reasoning, coding, comparing, and decision-supporting tasks, so that “the technology being adopted” is no longer a single well-defined object whose usefulness can be appraised once and for all. Third, the system has no identifiable human author of any particular output, so that the social-norm pathway through which adoption beliefs are partly transmitted in inherited theories - “people I respect think this is a good tool” - points to an object whose authorship is at best diffuse. Fourth, the user’s intent is partly co-constructed during the interaction itself rather than imported into it; this dissolves the temporal separation between “forming an intention to use the tool” and “using the tool”, on which inherited models tacitly rely.
Mogaji et al. (2024) capture part of this in what they call the static nature of TAM: the framework was not built to accommodate evolving, bidirectional engagement with a conversational system. The present analysis extends that diagnosis. The temporal problem is not only that the inherited theories are static while the technology is dynamic; it is that the formal assumptions of the inherited theories - repeatability of outputs, stability of the action object, ex ante evaluability of usefulness, identifiable authorship of system behaviour - were never tested against systems that lack those properties. This is not a refutation of the theories within their original scope. It is a precise statement of why their application to LLM-mediated behaviour is an extrapolation rather than a deployment, and why empirical findings produced by simply re-running TAM- or UTAUT-style instruments against current generative dialogue systems should be read with care.
A.3.3.3 Ontological Layer: Where Cognition Lives
The deeper objection concerns the implicit ontology of the acceptance tradition. In TAM, UTAUT, and TPB, technology is instrumental and cognition is endogenous to the user: the agent forms beliefs about the system, forms an intention to use it, and uses it to execute pre-formed goals. Each of the core constructs in the inherited theories rests on this division of labour. The present sub-section maps the strain construct by construct, because abstract claims about a generalised mismatch are easy to assert and difficult to evaluate; the claim is sharper when the locus of strain in each construct is identified.
In TAM, perceived usefulness is defined as the user’s belief that the system will improve performance on a specifiable task. The construct presumes that the benefit is knowable and stable enough to be evaluated in advance and that performance has a coherent meaning independent of the system. With LLMs, the benefit is co-discovered: the user learns what the system is useful for in the course of using it, and the criterion of “improved performance” often shifts during the interaction as the system reformulates the question, surfaces an option the user had not considered, or alters the user’s sense of what an acceptable answer looks like. Perceived ease of use is defined in terms of the cognitive friction of operation - clicks, learning, time. With LLMs, the central experiential phenomenon is not friction but cognitive co-production: the user types a partial thought and receives a structured continuation that anticipates and reformulates their reasoning. That is not “ease of use” in Davis’s sense; it is a different category of experience that the construct was not built to host.
A second strain on perceived ease of use becomes visible once verification load is brought into view. Heersmink et al. (2024) describe LLM use as combining interactional transparency with internal opacity: the surface of the dialogue is fluent and intelligible to the user, while the conditions under which any given response is correct remain inaccessible. Empirical work on human–AI decision-making (Buçinca, Malaya, & Gajos, 2021; Vasconcelos et al., 2023) demonstrates that users systematically over-rely on fluent AI outputs and that this overreliance is not reduced by attaching explanations to the response. Together these findings carry a measurement consequence for TAM that the present analysis foregrounds. Perceived ease of use as Davis (1989) defined it tracks the operational friction of using a tool - clicks, learning, time - and was calibrated against systems whose outputs were determinate. It does not separate the ease of obtaining a coherent response from the difficulty of verifying its validity. Under LLMs the construct can therefore score high in precisely the conditions Buçinca et al. (2021) and Vasconcelos et al. (2023) identify as those of greatest overreliance risk, and the user has no within-construct signal to distinguish the two cases.
In UTAUT, the same problems apply to performance expectancy and effort expectancy, which are formal generalisations of perceived usefulness and perceived ease of use respectively, with the additional constraint that they presume measurable performance gains on specifiable tasks. Social influence in UTAUT presumes a community of human referents whose use of the system is observed and normatively interpreted by the user. LLM use complicates this on two sides: the system itself participates in dialogue in a way that recruits social cognition (Nass & Moon, 2000; Sundar, 2020; Epley, Waytz, & Cacioppo, 2007), so that the “social” influence operating during a session is partly produced by the system rather than by other people; and the human referent group for LLM use is unstable, partly self-selected, and rapidly evolving, so that the construct’s presumed measurement target is itself moving. Facilitating conditions presume external resources - infrastructure, support - that enable use. With LLMs, the most consequential facilitating conditions are partly internal to the dialogue: prompt construction skill, context-management discipline, and the user’s own willingness to verify outputs. These are not absent from UTAUT in principle, but they are not where the construct expects to find facilitation.
In TPB, attitude towards the behaviour presumes that the behaviour and its attitude object are stable enough to be evaluated coherently; with LLMs, each session may target a different option set generated mid-conversation, so that the “behaviour” being evaluated is itself dialogue-dependent. Subjective norm faces the same instability of referent group noted for UTAUT’s social influence. Perceived behavioural control is the construct on which the strain is sharpest. PBC presumes that the locus of reasoning is internal to the actor: the agent forms an intention based on perceived control over an action whose execution depends on her own resources and capabilities. With LLM-mediated deliberation, the locus of reasoning is partly externalised into a dialogic partner whose contributions shape the goal, the option set, and the criteria of evaluation. The agent is not asking only “can I do this?”, but also, implicitly, “can the joint apparatus of myself and this system do this?”. This is not a niche difficulty: it is precisely the kind of structural problem that Sniehotta, Presseau, and Araújo-Soares (2014) had in mind when they argued that TPB’s sufficiency hypothesis has been falsified. The relevant external influence on behaviour - the LLM - is not absorbed by attitude, subjective norm, or PBC; it acts inside deliberation rather than as a stable upstream input.
The collective ontological problem is therefore not located in any one construct but in the assumption shared across the three theories. TAM, UTAUT, and TPB all treat technology as instrumental and cognition as endogenous: the agent thinks, then the tool acts. LLMs distribute cognition between user and system in a way that none of the three frameworks specifies. This is not a metaphysical objection. It is a measurement objection. When the locus of reasoning is partly outside the user, constructs that measure user-internal beliefs about a stable external object will under-represent the very mechanism that is theoretically active. This is consistent with Bagozzi’s (2007) older argument that TAM lacks an apparatus for the self-regulation that mediates intention and action; the present case sharpens the older argument by specifying that, with LLMs, what is missing is not only self-regulation but co-regulation between user and system within the deliberative episode itself. It is also consistent with Sundar’s (2020) more recent claim that the rise of machine agency requires a new theoretical apparatus for human–AI interaction, since traditional acceptance frameworks treat the machine as a passive object whose properties the user evaluates.
A.3.3.4 Phenomenological Layer: Eight LLM-Mediated Mechanisms Without Integrated Theoretical Homes
The temporal and ontological layers concern, respectively, the substrate against which the inherited theories were calibrated and the assumptions about cognition that those theories carry. The phenomenological layer concerns the empirical phenomena that LLM-mediated activity actually produces and that any model of LLM-assisted evaluative decision-making must be able to host. The claim is not that every underlying psychological process first appeared with LLMs. Several mechanisms have pre-LLM ancestors in automation trust, anthropomorphism, metacognition, cognitive offloading, and decision-support research. The narrower claim is that dialogic, generative, RLHF-tuned, context-retaining systems combine and activate these mechanisms inside the evaluative decision episode in ways that TAM, UTAUT, and TPB do not integrate as a coherent theoretical layer.
Hallucination - the confident generation of plausible but unfounded content - is not error in the deterministic software sense. The system is functioning according to its generative architecture when it produces fluent but unsupported text: it is generating token sequences that are statistically coherent, not necessarily grounded in a verified external state of the world (Ji et al., 2023; Huang et al., 2025; Dang Anh-Hoang, Vu Tran, & Le-Minh Nguyen, 2025). Recent empirical work in high-stakes domains such as law further shows why this is not a trivial reliability problem: hallucinated outputs can be expressed with the same fluency and authority as correct outputs, leaving users without an obvious run-time signal that verification is required (Dahl, Magesh, Suzgun, & Ho, 2024).
Sycophancy - the tendency of dialogue-tuned systems to converge on, flatter, or reinforce the user’s expressed view - is similarly specific to the alignment and preference-optimisation regime of contemporary conversational AI. Sharma et al. (2024) show that RLHF-trained assistants can prefer agreement with the user over truthfulness; Fanous et al. (2025) and Kaur (2025) extend this evidence to systematic sycophancy evaluation and argument-driven stance shifts in multi-turn settings. In inherited adoption models, these phenomena cannot be treated merely as low usefulness or low reliability, because they may increase perceived usefulness precisely by producing fluent, agreeable, and confidence-enhancing output.
Anthropomorphic dialogic interaction has clear pre-LLM ancestry in social-response and anthropomorphism research (Nass & Moon, 2000; Epley, Waytz, & Cacioppo, 2007), but contemporary LLMs intensify the mechanism through sustained, turn-taking natural-language interaction. The user is not only responding to a screen or a menu; the user is interacting with a system that apologises, reasons, explains, adapts tone, and appears to understand the evolving question. Recent work on anthropomorphic cues and mental-state attribution in LLMs shows that such cues can affect trust, perceived competence, and advice acceptance (Cohn et al., 2024; Colombatto et al., 2025).
Stochastic non-determinism - the property that the same or near-identical prompt may yield materially different outputs across runs - further strains the repeatability assumptions implicit in classical acceptance constructs. In software-engineering and applied evaluation settings, repeated LLM queries have been shown to produce non-identical or inconsistent outputs, and even low-temperature settings do not always eliminate non-determinism (Ouyang, Zhang, Harman, & Wang, 2025; Funk et al., 2024). A user may therefore be evaluating not a stable artefact with fixed performance properties, but a probabilistic conversational process whose realised behaviour changes across sessions and turns.
Epistemic opacity in the presence of interactional transparency is another LLM-mediated configuration without a natural home in inherited adoption-intention models. Heersmink et al. (2024) describe LLMs as phenomenologically and interactionally transparent enough to support fluent use, while remaining opaque in their internal generative processes, training-data provenance, and failure modes. Hicks, Humphries, and Slater (2024) sharpen the epistemic problem by arguing that LLM outputs are not truth-governed in the way users may assume, while Cassinadri (2024) places ChatGPT-like systems within a virtue-epistemological problem of cognitive artefact use. This combination matters for TAM because perceived ease of use can be inflated by surface fluency while actual epistemic transparency remains low.
Asymmetric trust calibration then follows: users may update trust, confidence, and reliance faster than the system’s true reliability profile can be assessed. Automation-trust research anticipated parts of this problem (Lee & See, 2004; Parasuraman & Manzey, 2010), but recent AI work makes the calibration issue more acute by showing that confidence, explanation, and metacognitive sensitivity are central to whether users know when to trust AI advice (Lee, Pruitt, Zhou, Du, & Odegaard, 2025).
Context-window adaptation means that the system’s current output is shaped by the sequence of prior turns. The option set shown to the user in turn six is not independent of the assumptions, framings, omissions, or biases introduced in turns one through five. Herlihy, Neville, Schnabel, and Swaminathan (2024) show that LLM-based chatbots can be affected by miscalibrated conversational priors, especially under under-specified requests, while Laban et al. (2026) show that LLMs can become lost in multi-turn conversation by making early assumptions and then over-relying on them rather than recovering cleanly.
Co-construction of the option set is the parallel user-side phenomenon: the user’s alternatives, criteria, and sometimes attitudes are produced inside the dialogue rather than brought to it as fixed upstream inputs. Jakesch et al. (2023) show that co-writing with opinionated language models can shift users’ expressed views and later attitudes, while recent work in AI-mediated persuasion shows that conversational AI can shift political preferences or attitudes under experimental conditions (Argyle et al., 2025; Lin et al., 2025; Hackenburg et al., 2025). In TPB terms, this means that “attitude toward the behaviour” is not always a stable antecedent to the deliberative episode; under LLM mediation, it may be partly generated within that episode.
The theoretical incursion is therefore not the claim that these eight mechanisms have no intellectual ancestry. The incursion is that, under LLM-mediated evaluative interaction, these mechanisms cross the measurement boundaries of the inherited constructs. Hallucination is not merely lower perceived usefulness, because fluency can preserve or increase perceived usefulness while factual grounding collapses. Sycophancy is not merely social influence, because the source of influence is the technological interlocutor itself rather than human peers or social norms.
Anthropomorphic dialogue is not merely hedonic motivation, because it turns the system into a pseudo-social participant in judgement. Stochastic non-determinism is not merely unreliability, because the user is evaluating a probabilistic process rather than a fixed system state. Epistemic opacity with interactional transparency is not merely low ease of use; it is often high ease of use masking low inspectability. Trust calibration is not merely trust as a static antecedent to intention, because reliance is adjusted dynamically inside the interaction. Context-window adaptation is not merely facilitating conditions, because the system’s apparent capability is recursively shaped by prior turns. Co-construction of the option set is not merely attitude or perceived behavioural control, because the system helps produce the alternatives and capability judgements that those constructs treat as already formed.
The two logics therefore converge. The fresh LLM-era literature gives empirical support for the eight mechanisms as real phenomena of generative, dialogic, context-retaining systems. The intellectual-incursion logic explains why their theoretical significance is not exhausted by adding them one by one as moderators, controls, or auxiliary variables. Some mechanisms are LLM-native; others are older mechanisms reconfigured and intensified by LLM affordances. What makes the phenomenological layer necessary is the combined fact that these mechanisms are activated inside the evaluative decision episode and that TAM, UTAUT, and TPB have no integrated construct architecture for modelling how they jointly reshape perceived capacity, calibration, and choice. The phenomenological layer therefore does not claim eight historically originless phenomena; it identifies an LLM-mediated configuration that inherited adoption and intention frameworks can only patch fragment by fragment, but cannot host as an integrated account of evaluative judgement.
A.3.3.5 Construct Proliferation as Confirmatory Evidence
The proliferation pattern is itself diagnostic, and on this point the present argument converges with - and draws empirical support from - the published critique tradition. Bagozzi (2007) used the phrase “stage of chaos” to describe the unprincipled growth of UTAUT-style models. Benbasat and Barki (2007) used the phrase “theoretical chaos and confusion” to describe the same problem in TAM-style extensions. Mogaji et al. (2024) recommended openness to alternative models partly because patch-style extension has not produced integration. Outside the information-systems literature, Shaffer, DeGeest, and Li (2016) formalised the same diagnosis at the methodological level, defining construct proliferation as “the accumulation of ostensibly different but potentially identical constructs representing organizational phenomena” and arguing that the field-level remedy is rigorous discriminant-validity testing of conceptually adjacent constructs rather than further construct addition. The proliferation observed in the AI-behaviour literature has the same signature: the new constructs are not in obvious conflict with TAM, UTAUT, or TPB; they are simply not derivable from them. Each is added as a moderator or as a parallel construct because the parent framework offers no principled location for it.
The pattern continues into 2025. Salam et al. (2025) advance a Revised Artificial Intelligence Device Use Acceptance (RAIDUA) framework that attaches privacy concern, anthropomorphism, and emotional appraisal as new constructs onto a UTAUT-style spine. Xue, Ghazali, and Mahat (2025) propose an Integrated AI-in-Education Acceptance Framework (IAEAF) that adds AI literacy, institutional context, intervention design, and sustained-integration outcomes; the authors themselves motivate their framework as a response to “conceptual redundancy and contextual misalignment” across the proliferating UTAUT2 extensions in their domain - a 2025 echo, in different vocabulary, of Bagozzi’s (2007) “stage of chaos”. Read through Kuhn (1962/2012) and Lakatos (1970), the pattern is familiar. When researchers continue to operate within an existing framework while patching it with successive auxiliary constructs, the framework’s predictions are not refuted; instead, they are increasingly accompanied by qualifications, and ad hoc additions multiply because the underlying frame cannot host the phenomena cleanly. The present analysis does not require the stronger Kuhnian claim that the AI-behaviour literature is undergoing a paradigm shift. The measured claim is sufficient: the inherited frameworks are showing signs of strain under the weight of LLM-specific phenomena; the rate of construct addition is outpacing the rate of theoretical integration; and that pattern has been independently noted by senior commentators in both the 2007 Quo Vadis TAM exchange and the more recent generative-AI-specific critique. A new model in this space, on this view, must do more than predict an additional outcome variable. It must provide a coherent home for the constructs the field has been generating in an ad hoc way. Section A6 returns to this point in the form of a rival-explanation audit; the present sub-section establishes only that proliferation is part of the evidence base for the claim that the inherited frameworks are incomplete, not a separate methodological worry.
A.3.3.6 What the Present Argument Adds Beyond the Published Critiques
§A3.3.1 surveys the published critique tradition; this sub-section identifies what each critic did not specify, so that the position the present model occupies relative to that tradition can be stated precisely. Bagozzi (2007) called for a paradigm shift and proposed a general goal-pursuit decision-making core; that proposal predates the deployment of LLMs and does not specify the cognitive-capacity mechanism that arises when a conversational generative system is present at the moment of deliberation. Benbasat and Barki (2007) called for renewed attention to the system characteristics that actually make systems useful; they did not propose a moderation gate distinguishing empowered use from distorted use. Sniehotta, Presseau, and Araújo-Soares (2014) argued that TPB’s sufficiency hypothesis has been falsified; they did not specify what the relevant unmediated external influences are in AI-mediated deliberation, nor where in the deliberative chain those influences enter. Mogaji, Viglia, Srivastava, and Dwivedi (2024) catalogued five limitations of TAM in the era of generative AI and recommended industry-specific embedding and openness to alternative models; they did not propose a specific construct, mechanism, or extension architecture. Sundar (2020) argued that machine agency requires a new theoretical apparatus and proposed a dual-process framework drawn from the Theory of Interactive Media Effects; that proposal addresses media-effect symbolism and affordance perception more than the formation of decision criteria during high-stakes individual deliberation. Heersmink, de Rooij, Clavel Vazquez, and Colombo (2024) characterised the distinctive epistemological situation of LLM use - interactional transparency with internal opacity - but at the level of philosophy of technology rather than at the level of an empirical model. §A3.3.7 below states the corresponding architectural response.
A.3.3.7 Implication for the Present Model
In one sentence, TAM, UTAUT, and TPB model technology adoption and intention formation; LLM-mediated behaviour at the moment of complex decision additionally requires a model of cognitive partnership during evaluation. The Augmented Behavioural Capacity Model does not seek to replace the inherited frameworks. It accepts that they explain access, uptake, and intention well, treats them as the upstream layer of any complete account, and positions itself as a domain-bounded extension that isolates the LLM-specific cognitive-capacity mechanism that none of the inherited frameworks specifies - and that none of the published critiques of those frameworks has yet specified either. Figure 3 (reported in the manuscript body) visualises the three-layer argument geometrically: the inherited theoretical apparatus is stratified inside the frame; the phenomenological mechanisms unique to LLM-mediated activity sit outside the frame, indicating their lack of a principled home in the existing constructs. Section A5 specifies the architecture by which the proposed extension is organised; Section A6 returns to the construct-proliferation pattern in the form of an audit of rival families that the new construct must withstand.
A symmetric point applies at the dependent-variable end of these theories. TAM and UTAUT terminate in intention to use, actual use, and continued use. TPB terminates in intention and behaviour. None of the inherited frameworks reaches decision quality - the structural complexity, risk posture, and calibration of the choice the user actually makes. In retail-investor decision-making, decision quality in this narrow sense is the outcome that matters: whether LLM assistance changes which strategies the investor is willing to consider, on what evidentiary basis, and with what calibration. The gap the present model addresses is therefore bilateral. On the independent-variable side, no inherited framework specifies the tool-contingent perceived-capacity mechanism that operates inside the deliberative phase under LLM augmentation. On the dependent-variable side, no inherited framework distinguishes adoption-or-intention outcomes from decision-quality outcomes. The Augmented Behavioural Capacity Model is positioned to address both: it inserts a cognitive-capacity layer (Capacity, with PCA as its measurable anchor) and a moderation gate (Calibration) between the inherited adoption-and-intention layer upstream and a proximal decision-quality layer (Choice, bounded to tactic-selection and intended structural complexity) downstream. Read in this light, the model is not another loose extension of TAM, UTAUT, or TPB; it is an attempt to specify where the extension belongs.
Table A5 consolidates the argument across the inherited frameworks named in the heading of this section. Each row identifies what the theory or family explains well within its original scope, the core limitation that emerges under LLM-mediated decision-making, and the corresponding response the ABC Model adopts. The table serves as a one-page orientation map across the theoretical landscape, not as the audit itself. It is a comparative consolidation table, not a meta-analysis and not a systematic review: its function is to map, in one view, what each inherited framework explains within its original scope, where it strains under LLMs, and the ABC response.
Table A5. Consolidation of the §A3.3 argument across the three inherited frameworks.
Table A5. Consolidation of the §A3.3 argument across the three inherited frameworks.
Framework Temporal mismatch Ontological mismatch Phenomenological mismatch ABC implication
TAM Calibrated to stable IS tools User evaluates external tool PU/PEOU cannot house dialogic scaffolding, opacity, hallucination Retain PU/PEOU upstream; add PCA
UTAUT/UTAUT2 Calibrated to determinate task/use systems Performance/effort expectancy assume stable task-tool relation Does not capture internal decision-process transformation or calibration Retain engagement layer; do not equate PE with PCA
TPB Assumes stable behaviour/action object PBC locates control mainly in actor Cannot represent externally scaffolded perceived capacity or co-produced option set Refine PBC into tool-contingent capacity plus calibration
The three-layer analysis establishes the gap that motivates ABC. It does not yet adjudicate every neighbouring construct family that could compete with PCA. That task belongs to the boundary and rival-explanation audit in Section A6. The immediate implication of Section A3 is narrower: inherited acceptance and intention frameworks explain access, uptake, usefulness, effort, social influence, attitude, norm, perceived control, intention, and use, but they do not specify the LLM-specific perceived-capacity mechanism operating inside the evaluative decision episode, nor do they distinguish adoption/intention outcomes from proximal decision-quality outcomes. Section A5 translates this residual zone into the ABC architecture; Section A6 then tests whether adjacent constructs and rival families can absorb it.

A.4. Construct and Element-Status Registry

Figure A1 clarifies how the three-layer critique developed in Section A3 is translated into the element logic of the ABC Model. The mapping is not a one-to-one conversion in which each diagnostic layer mechanically produces one construct. Instead, the three layers play different explanatory roles. The temporal layer establishes the need for a domain-bounded extension by showing that inherited acceptance and intention theories were calibrated to stable technologies rather than dialogic generative systems. The ontological layer motivates the Capacity construct by showing that LLM-assisted decision-making partly relocates reasoning across the user–system interaction rather than leaving cognition wholly inside the individual actor. The phenomenological layer then identifies the concrete mechanism-home and outcome-home problems created by LLM-mediated activity: dialogic scaffolding, metacognitive miscalibration, and altered decision outputs cannot be housed cleanly within the inherited construct space. Figure A1 therefore functions as a bridge between the gap analysis and the construct registry: it shows why ABC requires a domain-bounded model architecture, why Capacity is operationalised through PCA, and why Calibration and Choice are required as distinct model elements rather than as further extensions of adoption, use, or intention.
Figure A1. Translating the Three-Layer Theory-Gap Diagnosis into the ABC Model.
Figure A1. Translating the Three-Layer Theory-Gap Diagnosis into the ABC Model.
Preprints 221466 g0a1
The five elements are heterogeneous in kind: Adoption and Engagement is a boundary stage carrying inherited constructs; Capacity is a single latent construct in the standard reflective sense; Calibration is a battery of diagnostic variables that do not unify into one factor; Choice is a family of outcome variables; Footprint is a domain of realised-behaviour measurement. Process models in behavioural science routinely contain heterogeneous element types, and construct-clarity scholarship requires that such elements be classified before relationships among them are specified, since the type of theoretical object each element is determines what counts as evidence for it (Bacharach, 1989; Edwards & Bagozzi, 2000; MacKenzie, Podsakoff & Podsakoff, 2011; Podsakoff, MacKenzie & Podsakoff, 2016; Suddaby, 2010; Whetten, 1989). Table A6 therefore reports an element-status registry, not a factor model. Its function is to establish what each named element is allowed to mean before §A5 specifies how the elements relate, so that the model is not read as a five-factor reflective structure of the kind familiar from TAM and UTAUT extensions.
The element-status registry should not be read as a sequential mediation model. Capacity and Choice form the direct behavioural-migration spine of ABC: perceived cognitive assistance may alter what the trader considers, selects, structures, or commits to at the decision moment. Calibration is different. It is a diagnostic and moderating apparatus that sits over the Capacity → Choice relationship and conditions how that relationship should be interpreted. In empirical time, calibration may be assessed after a PCA-relevant decision episode; in model logic, however, it is not a stage through which complexity mechanically passes. It is the decision-episode property against which PCA-driven choice is assessed.

A.4.1. Inside the Three Core Stages

Figure A2 displays the three core stages of the ABC Model with their operationalising variables and empirical indicators. The three stages are not three constructs of the same theoretical kind. Capacity is a construct in the standard latent-variable sense; Calibration is a heterogeneous diagnostic apparatus rather than a unified construct; Choice is a set of outcome variables. The figure preserves this asymmetry rather than concealing it, and the three stages are described below in the vocabulary appropriate to each.
Figure A2. The three core stages of the ABC Model, with their operationalising variables and empirical indicators.
Figure A2. The three core stages of the ABC Model, with their operationalising variables and empirical indicators.
Preprints 221466 g0a2
Capacity is operationalised by a single latent construct, Perceived Cognitive Assistance (PCA): the user’s felt expansion of cognitive capability at the moment of a complex trading decision when a large language model is available (Gimmelberg & Ludviga, 2026c). Confirmatory work supports a one-factor PCA model whose items cover five content domains - decision structuring, cognitive-load relief, error checking, scenario comparison, and reflection (Gimmelberg & Ludviga, 2026c). Empirical indicators for the Capacity stage are the validated nine-item PCA scale (PCA-9; Gimmelberg & Ludviga, 2026c), process measures embedded inside a frozen-decision causal-evaluation design, and cognitive-load probes drawn from established workload instruments (Hart & Staveland, 1988; Paas et al., 2003).
Calibration is operationalised by a battery of five diagnostic variables that together distinguish empowered from distorted reasoning under LLM assistance: objective understanding, perceived understanding, overprecision, reliance and deference, and state offloading. Each variable is inherited from a distinct prior literature: perceived understanding from work on the illusion of explanatory depth (Fernbach et al., 2013; Rozenblit & Keil, 2002); overprecision from the calibration taxonomy of Moore and Healy (2008); reliance and deference from research on trust in automation, algorithm aversion, and algorithm appreciation under articulate model output (Buçinca, Malaya & Gajos, 2021; Bussone, Stumpf & O’Sullivan, 2015; Dietvorst, Simmons & Massey, 2015; Lee & See, 2004; Logg, Minson & Moore, 2019; Parasuraman & Riley, 1997); and state offloading from work on cognitive offloading and transactive memory with external systems (Risko & Gilbert, 2016; Sparrow, Liu & Wegner, 2011; Storm & Stone, 2015). These five variables are heterogeneous by design - knowledge states, meta-cognitive judgements, confidence-calibration metrics, behavioural patterns, and cognitive practices - and are deliberately not modelled as indicators of a higher-order Calibration construct. The Calibration stage is therefore not a construct in its own right but a moderating apparatus, formally treated in §A5.3. Empirical indicators of the apparatus are comprehension probes, calibration tasks in which subjective confidence is compared against objective accuracy in the tradition of Lichtenstein, Fischhoff and Phillips (1982), and deference diagnostics that measure how readily users accept articulate model output without independent verification.
Choice refers to the proximal decision output formed at the evaluative moment before execution. It is not realised return, portfolio welfare, or market-level footprint. Nor is decision quality used here to define enhancement or distortion. Decision-quality scoring may later serve as an external criterion for testing whether calibrated enhancement has substantive value, but the classification itself is based on PCA and calibration diagnostics, not on whether the final decision is judged good or bad. In the ABC Model, Choice is an outcome family comprising complexity intentions, strategy evaluation, structural complexity of intended trades, and risk assertion. The shared feature across these outcomes is that they capture what the trader becomes willing to consider, select, structure, or commit to once perceived cognitive capacity changes. Choice is operationalised by four outcome variables : complexity intentions (Complexity intentions capture the investor’s willingness to consider more complex tactics, such as multi-leg or volatility-linked structures. Selected trade-structure complexity captures the objective or rubric-coded sophistication of the specific position the investor selects in a vignette or decision task, such as number of legs, conditionality, hedge structure, volatility exposure, or payoff asymmetry. The distinction is therefore between expansion of the consideration set and the complexity of the chosen intended structure) (willingness to consider multi-leg or volatility-linked structures), strategy evaluation (depth and breadth of strategy comparison), selected trade-structure complexity (the structural sophistication of the position the investor commits to), and risk assertion (the disciplined assignment of risk to a trade thesis). The complexity dimensions are anchored to retail-options behaviour, where the elevation of structural sophistication is the empirically interesting outcome (Bauer, Cosemans & Eichholtz, 2009; Bellofatto, De Winne & D’Hondt, 2018). Choice in this sense is bounded to proximal selection and intended structure; it does not extend to realised execution, which belongs to the Footprint boundary stage and is not advanced as a defended claim of this paper. Indicators are forced-choice vignettes in the tradition of experimental vignette methodology (Aguinis & Bradley, 2014), decision-process quality scored against a fixed rubric, and trade selection at the level of intention rather than execution.

A.4.2. The Structural Asymmetry Across the Three Cs

The asymmetry just described - construct, diagnostic apparatus, outcome family - is not a defect of the architecture. It is a substantive feature of the model, and naming it openly is preferable to dissolving it under a prose convention that the three Cs are constructs of the same kind. Naming it has a further consequence. The asymmetry forces a decision on the formal status of Calibration - moderator, second-order property of PCA, latent class membership, or formative composite - before the propositions implied by the model can be derived cleanly. The choice between formative and reflective specification is not cosmetic: it changes what counts as evidence for the construct and how its indicators are interpreted (Diamantopoulos & Winklhofer, 2001; Jarvis, MacKenzie & Podsakoff, 2003). Section A5.3 makes that decision, and §A5.4 then addresses the within-moderator specification problem created by the directional heterogeneity of the five Calibration variables. The registry does not yet claim that the elements are causally connected; it only establishes what kind of theoretical object each element is. Section A5 therefore moves from classification to relationship architecture.

A.5. ABC Model Architecture and Relationship Logic

A.5.1. The Evaluative Stage as the Locus of the LLM-Specific Mechanism

The ABC Model locates the LLM-specific mechanism inside the deliberative phase of decision-making - the phase at which the investor structures the problem, compares tactics, tests scenarios, and decides what seems cognitively manageable. In TPB terms, this is the phase in which attitudes, subjective norms, and perceived behavioural control are formed before intention crystallises (Ajzen, 1991); in TAM terms, it is the phase in which usefulness and ease judgements operate (Davis, 1989). The inherited frameworks therefore do reach this phase, but they reach it as the site of attitude, usefulness, or control-belief formation rather than as a mechanism specification of that formation under LLM augmentation. ABC’s contribution is not the discovery of a previously unrecognised stage. It is the specification of a tool-contingent perceived-capacity mechanism - PCA - operating inside the deliberative phase that TPB and TAM already touch on, and the claim that this mechanism is theoretically distinguishable from the inherited attitudinal and control constructs that occupy the same temporal territory.
In retail-investor decision-making, this phase is materially consequential. It is where strategy families are appraised, where the structural complexity of a position is judged feasible or not, where ambiguity is converted into a working interpretation, and where confidence is assigned to that interpretation. PCA is the model’s Capacity-layer representation of the user’s felt cognitive manageability inside that phase. The construct is theoretically anchored in the extended-mind and cognitive-scaffolding tradition (Clark & Chalmers, 1998; Hutchins, 1995; Wood, Bruner, & Ross, 1976; Heersmink et al., 2024), but it is not itself a measurement of distributed cognition. It measures the user’s perceived expansion of cognitive capability in the presence of an LLM during a decision episode. The distinction matters because it locates the model’s empirical commitment precisely. ABC does not claim, at the level the construct measures, that cognition is empirically distributed across the user–system boundary. It claims that the user’s perception of being cognitively assisted is a distinct, measurable, domain-relevant construct that inherited frameworks do not specify, and that the distributed-cognition tradition provides the theoretical motivation for why such a perceived-capacity construct should exist as a separate measurement The diagnostic moderation layer is then specified conceptually to distinguish calibrated enhancement from behavioural distortion: high PCA is interpreted as calibrated enhancement when it is accompanied by objective understanding, confidence accuracy, disciplined verification, and retained independent reasoning, but as distortion when it is accompanied by fluency mistaken for evidence quality, overprecision, perceived understanding without objective comprehension, uncritical deference, or unverified cognitive offloading (Alter & Oppenheimer, 2009; Lee & See, 2004; Moore & Healy, 2008; Parasuraman & Manzey, 2010).
Locating the mechanism inside the deliberative phase has two consequences for the model’s structure. First, it distinguishes the ABC Model from acceptance and intention models without contradicting them: ABC does not replace TAM, UTAUT, or TPB, but specifies the LLM-contingent perceived-capacity mechanism that operates within the decision episode those frameworks approach through usefulness, ease, attitude, norm, and perceived-control formation. Second, it identifies the natural site for distortion diagnostics. If the LLM’s effect operates while the user is forming the decision frame, then so do the failure modes - fluency mistaken for evidence quality (Alter & Oppenheimer, 2009), overprecision in confidence intervals (Moore & Healy, 2008), perceived understanding without objective comprehension (Rozenblit & Keil, 2002), deference to articulate output (Parasuraman & Manzey, 2010), and cognitive offloading onto the system (Risko & Gilbert, 2016). Distortion diagnostics therefore belong inside controlled frozen-decision experimental designs and inside longitudinal calibration-tracking studies, rather than as afterthoughts.

A.5.2. Architecture: Three-Stage Core, Two-Stage Boundary

The architecture is composed of five-layer architecture, organised asymmetrically. The three stages corresponding to the model’s three core concepts - Capacity, Calibration, and Choice - constitute the defended theoretical core. The choice to defend an asymmetric core, rather than to defend all five layers as a single integrated structure, follows directly from established principles of theoretical model construction: a model must specify what it claims, what it inherits, and what it extends, and the boundary between these zones must be visible rather than concealed (Bacharach, 1989; Suddaby, 2010; Whetten, 1989). The three core stages are not equally validated. Capacity is empirically anchored through PCA, for which scale-development and confirmatory measurement evidence already exists (Gimmelberg & Ludviga, 2026c). Calibration is formally specified as a diagnostic and moderating apparatus but has not yet been validated as an integrated gateway; its dual-pathway prediction requires designs that combine PCA measurement with objective-comprehension, confidence-accuracy, deference, and offloading diagnostics (Lee & See, 2004; Moore & Healy, 2008; Risko & Gilbert, 2016; Rozenblit & Keil, 2002). Choice is treated as the proximal outcome domain, with partial vignette-based evidence for complexity intentions and broader controlled and longitudinal testing required (Aguinis & Bradley, 2014).
The two stages framing the core - Adoption and Engagement at the upstream end, and Behavioural Footprint at the downstream end - are presented as boundary architecture rather than as directly defended structure. The principle that a model should explicitly mark the surface at which it meets prior literature, and the surface at which it makes empirical commitments it does not itself test, is a standard requirement of theory-building practice (Bacharach, 1989; Wacker, 1998; Whetten, 1989). Adoption inherits from TAM (Davis, 1989) and UTAUT (Venkatesh et al., 2003, 2012), accepted as adequate accounts of pre-decision engagement and not re-tested here. Footprint extends into realised trading behaviour, documented as a downstream extension but not advanced as a defended claim. The architecture is therefore deliberately asymmetric: a three-stage individual-level psychological model at the centre, with an inherited stage upstream and an extension stage downstream. Table A6 maps the five stages onto their theoretical role, main constructs, empirical indicators, and the current empirical status of each layer.
Five stages, rather than three or four, follow from the same logic. A three-stage architecture confined to the defended core would be silent on how the investor arrives at LLM use, and would silently inherit the upstream conditions of access and adoption without naming them; the inherited TAM and UTAUT layer would be carried implicitly rather than as an acknowledged boundary. A four-stage architecture would be inconsistent: adding only Adoption upstream, or only Footprint downstream, would name one boundary while concealing the other, and the model would be unable to state the same boundary discipline at both ends. Five stages is the minimum that lets the inherited adoption tradition appear upstream and the empirical-extension question appear downstream, while keeping both at the boundary rather than inside the defended structure. The architecture therefore commits to two boundary stages it does not defend, in order to make explicit the surface at which the defended core meets the existing literature on one side and the empirical-extension agenda on the other.
Table A6. Architecture of the ABC Model.
Table A6. Architecture of the ABC Model.
Element / layer Status & theoretical role Main constructs / variables Empirical indicators Empirical status & evidence role
1. Adoption and Engagement Boundary stage carrying inherited constructs; the upstream condition of whether and how the investor uses the LLM. Presupposed by ABC, not newly theorised. Perceived usefulness; perceived ease of use; trust; social influence; facilitating conditions; prior use. Use frequency; task breadth; reliance; prompting sophistication. Inherited from TAM (Davis, 1989), UTAUT (Venkatesh et al., 2003, 2012), and TPB (Ajzen, 1991). Enters as context and antecedent covariates; established comparator domain, not re-tested here.
2. Capacity ABC core construct, operationalised through PCA: the user’s felt expansion of cognitive capability at the moment of evaluative decision-making. PCA and its facilitative facets. PCA-9 scale, covering decision structuring, cognitive-load relief, error checking, scenario comparison, and reflection. Current empirical anchor of the model. PCA-9 is content-validated, confirmed as a one-factor scale with adequate reliability and scalar invariance, and shows early incremental predictive evidence (Gimmelberg & Ludviga, 2026c); causal and longitudinal extensions remain future work.
3. Calibration ABC core diagnostic and moderating apparatus, not a latent factor.
Distinguishes calibrated enhancement from inflated capability; high PCA is not automatically warranted, and the apparatus conditions how PCA-driven Choice should be interpreted.
Objective understanding; perceived understanding; overprecision; reliance and deference; state offloading. Comprehension probes; confidence-accuracy calibration tasks; deference diagnostics; offloading measures. Theoretically central but not yet validated as an integrated apparatus. Requires controlled experimental and longitudinal testing combining PCA with the five calibration variables (Lee & See, 2004; Moore & Healy, 2008; Risko & Gilbert, 2016; Rozenblit & Keil, 2002).
4. Choice ABC core outcome family, not a single dependent variable. Captures what the investor becomes willing to consider, select, structure, or commit to once perceived capacity changes, with Calibration conditioning the interpretation. Complexity intentions; strategy evaluation; structural complexity of intended trades; risk assertion. Vignettes; forced choice; decision-quality scoring; intention-level trade-selection tasks. Partially evidenced through vignette-based complexity intentions (Aguinis & Bradley, 2014). Broader controlled and longitudinal tests of the moderated Capacity → Choice path remain future work.
5. Behavioural Footprint Downstream extension, not part of the defended core: whether the individual-level mechanism leaves traces in realised trading behaviour. Multi-leg share; option-trade frequency; concentration; volatility exposure. Trading-log indicators. Extension only; documented separately. Required only if the model claims realised behavioural or market-facing consequences. Not advanced as a defended claim of this paper.
Note. This registry consolidates the former element-status and architecture tables. It is a classification of element types and their current empirical status, not a five-factor reflective model: the elements are heterogeneous in kind - a boundary stage carrying inherited constructs (Layer 1), a single latent construct (Capacity), a heterogeneous diagnostic and moderating apparatus (Calibration, not a unified factor), an outcome family (Choice), and a domain of realised-behaviour measurement (Behavioural Footprint). The empirical-status column reports current evidence, not a research-programme assignment, and should not be read as asserting full layer-by-layer coverage: Layers 2–4 form the defended core but are unequally validated, while Layers 1 and 5 are boundary stages whose validation is not required for ABC to stand as an individual-level psychological account.

A.5.3. The Formal Status of Calibration: Moderating Apparatus

Four formal interpretations of Calibration are theoretically possible, each corresponding to a distinct specification convention in the methodological literature. The first is moderation: the five Calibration variables modify the strength and sign of the path from Capacity to Choice (Aiken & West, 1991; Baron & Kenny, 1986; Hayes, 2018). The second is a second-order property of PCA itself, in which what matters is PCA contingent on comprehension, captured as a product term or as a residualised score. The third is latent class membership, in which users partition into calibrated and miscalibrated subgroups with qualitatively different downstream pathways (Lazarsfeld & Henry, 1968; McCutcheon, 1987). The fourth is a formative composite, in which a calibration index is built additively from the five diagnostics and entered as a measured variable in the path model (Diamantopoulos & Winklhofer, 2001; Jarvis, MacKenzie, & Podsakoff, 2003).
The architecture commits to the moderator interpretation as primary, with latent class preserved as a falsifiable rival for subsequent research. Four reasons motivate this commitment. First, the moderator framing is the most direct translation of the language already used elsewhere in this paper - the claim that high PCA matched by objective understanding, confidence accuracy, verification discipline, and retained independent reasoning is interpreted as calibrated enhancement, whereas high PCA unmatched by those referents is interpreted as distortion, is moderation language, and adopting it openly requires least retrofitting (Mathieu & Taylor, 2006). Second, it nests inside well-understood extension architecture in social-psychological theories of behaviour, where contingency variables routinely moderate paths to action (Baron & Kenny, 1986; MacKinnon, 2008). Third, it is the most statistically tractable form for an early-stage validation programme: interaction terms are testable in modest samples, where mixture models often are not (Aiken & West, 1991; Hayes, 2018). Fourth, the moderator specification does not foreclose the latent-class interpretation. If moderation effects prove sufficiently discontinuous or subgroup-specific in subsequent empirical work, later studies can test whether a latent-class or mixture specification better represents the calibrated and miscalibrated pathways (McCutcheon, 1987). The moderator interpretation is therefore adopted as the primary theoretical form for Claim B, while the latent-class form is sequenced as a downstream alternative rather than ruled out. Within the moderator interpretation, however, a further specification commitment is required because the five Calibration variables are not directionally homogeneous; §A5.4 addresses that commitment and excludes the naive additive specification that the directional heterogeneity would otherwise admit.

A.5.4. Directional Heterogeneity Inside the Moderating Apparatus

The moderator interpretation commits the model to a further specification problem that the asymmetric three-Cs framing of §A4.2 makes visible but does not itself resolve. The five Calibration variables are not directionally homogeneous with respect to the calibration-contingent meaning of PCA. Objective comprehension belongs to the adequacy side of the diagnostic structure: high objective comprehension alongside high PCA is more consistent with calibrated empowerment than with inflated capacity, although by itself it does not establish LLM-caused cognitive expansion. Overprecision belongs to the distortion side: higher confidence relative to accuracy indicates inflated capability that PCA-based choice cannot safely act on (Lichtenstein, Fischhoff, & Phillips, 1982; Moore & Healy, 2008). Three of the five - perceived understanding, reliance/deference, and state offloading - are state- and context-dependent. Each can index productive scaffolded reasoning when accompanied by objective understanding, disciplined verification, and retained independent reasoning, or behavioural distortion when those referents are absent (Lee & See, 2004; Parasuraman & Manzey, 2010; Risko & Gilbert, 2016). Treating the five variables as exchangeable indicators of a single underlying “Calibration” construct - either by combining them additively into a composite or by entering them as parallel components behind a single PCA × Calibration interaction - would mix adequacy and distortion signals. A null result under that specification could reflect genuine absence of moderation, but it could also reflect cancellation between directionally opposed components. This is the measurement-induced indeterminacy that the construct-clarity tradition warns against (Diamantopoulos & Winklhofer, 2001; Jarvis, MacKenzie, & Podsakoff, 2003; MacKenzie, Podsakoff, & Podsakoff, 2011).
Three specification routes remain available within the moderator interpretation, and the architecture takes a position on which is primary, which is alternative, and which is excluded. The first route, miscalibration-gap indexing, treats the operative quantities not as the raw level of any single calibration variable but as pre-specified discrepancies between subjective capacity, confidence, perceived understanding, reliance, or offloading and their matched objective or verification referents. The relevant discrepancies include PCA relative to objective task comprehension or performance; confidence relative to accuracy in the calibration tradition of Lichtenstein et al. (1982) and Moore and Healy (2008); perceived understanding relative to objective understanding in the illusion-of-explanatory-depth tradition (Rozenblit & Keil, 2002); deference relative to independent verification; and offloading relative to retained independent reasoning or an independently defensible choice rationale. These discrepancies should be constructed as sign-aligned standardised gap scores, residualised discrepancy measures, or polynomial response-surface terms rather than assumed to be simple raw-score differences across unlike scales. Once sign-aligned so that higher values indicate greater miscalibration, the gaps can be entered separately, modelled jointly, or combined only after theoretical alignment and measurement checks confirm that aggregation is defensible. This route is sequenced as primary at the architecture level because it respects the directional structure of the underlying calibration diagnostics and produces a moderation specification whose null result is not automatically contaminated by cancellation between adequacy and distortion components. The second route, configural specification, enters the five variables or their matched diagnostic pairs as separate moderators with pre-specified directional expectations and appropriate familywise or false-discovery correction. This preserves heterogeneity at the cost of statistical power and interpretive parsimony, and is retained as an alternative when sample size and instrument granularity permit. The third route, naive additive compositing without discrepancy formation, is explicitly excluded by the architecture because it mixes opposite-direction signals and produces an uninterpretable null condition.
This specification problem is downstream of the matched-domain instrumentation problem identified in §A6.5 and the main article’s §5.2, but it is logically distinct from it. The discrepancy quantities cannot be computed until objective-comprehension, calibration-task, verification-discipline, reliance/deference, and offloading measures are available at the retail-investor decision-episode resolution. The matched-domain battery is therefore a precondition for testing Claim B, but the choice among moderator specifications is a separate architectural commitment that the model makes in advance of the instrumentation work rather than deferring “until the data adjudicate.” The proposition set in §A7 is therefore written to leave both the gap-indexing primary route and the configural alternative operable in subsequent empirical work, while committing the model against the naive-additive specification that would smuggle directional heterogeneity under the appearance of a single moderator.

A.6. Boundary Conditions and Rival Explanations

A.6.1. Methodological Function of the Boundary Step

Section A3 established the explanatory gap: inherited acceptance and intention-based frameworks remain necessary, but they do not specify the LLM-specific mechanism that operates at the evaluative stage of retail-investor decision-making. Section A5 specified the architecture by which the proposed extension is organised: a three-stage core of Capacity, Calibration, and Choice, framed by an inherited Adoption-and-Engagement boundary upstream and a Behavioural Footprint extension downstream. The present section performs a different task. It asks under what conditions the model’s claims are asserted and under what conditions they would be weakened, made redundant, or refuted.
In disciplined model-development practice, this is not an optional or rhetorical step. Boundary specification and rival-explanation analysis are part of theoretical precision rather than supplementary “limitations” prose (Bacharach, 1989; Dubin, 1978; Wacker, 1998; Whetten, 1989). The relevant traditions converge on four obligations that the present section must discharge.
The first obligation comes from the boundary-conditions tradition. Bacharach (1989) requires that a theory specify the values of its variables across which the relationships are asserted, on the principle that a theory which holds for all values of all variables is empirically empty. Whetten (1989) frames the same requirement as the who, where, when question: a theoretical contribution must state who its propositions apply to, in what setting, and during what period. Dubin (1978) treats this as system-state specification: every theoretical system has values of its boundary variables outside which the system as defined no longer obtains. Busse, Kach, and Wagner (2017) develop this into a modern protocol for explicit boundary statement, distinguishing definitional, descriptive, and theory-bounding conditions and providing rules for how each is to be argued. Suddaby (2010) makes the parallel construct-clarity case: scope conditions and semantic boundaries are part of construct meaning, not of construct application.
The second obligation comes from the rival-explanation tradition. Campbell and Fiske (1959) established convergent–discriminant logic as the empirical foundation for distinguishing a focal construct from its neighbours: a construct that cannot be told apart from its rivals on either convergent or discriminant evidence has no theoretical standing. Edwards and Bagozzi (2000) extend this to nomological networks: a construct must be located inside a structure of expected and unexpected relationships, and rival constructs occupy positions within that network rather than outside it. Popper (1959) and Lakatos (1970) supply the falsificationist constraint: a theory must specify the observations under which it would be refuted, and it must distinguish its core claims from the auxiliary protective belt that surrounds them. Auxiliary modifications that absorb every disconfirming observation without affecting the core are exactly what makes a research programme degenerative in Lakatos’s sense.
These traditions produce four functions for the present section. First, the section states the model’s intended domain - who, where, when, and under what conditions the propositions are asserted (Bacharach, 1989; Busse et al., 2017; Whetten, 1989). Second, it protects construct clarity by distinguishing the focal Capacity mechanism from adjacent constructs and measures (Campbell & Fiske, 1959; Edwards & Bagozzi, 2000; Suddaby, 2010). Third, it compares plausible rival explanations and classifies them according to the empirical operation each requires (Bacharach, 1989; Wacker, 1998). Fourth, it identifies the observations that would weaken or refute the model in principle, so that Section A7 can convert those vulnerabilities into formal propositions and a validation agenda (Lakatos, 1970; Popper, 1959; Wacker, 1998).
The construct-proliferation pattern documented in §A3.3.5 enters the present section as premise rather than as content. Section A3 read the rapid expansion of partially overlapping constructs in the AI-behaviour literature as diagnostic evidence of theoretical strain inside the inherited frameworks. The present section accepts that reading and asks the natural follow-on question: now that the inherited frameworks are identified as incomplete and proliferation is identified as evidence of strain, can the proliferating neighbours individually or collectively absorb the ABC Model’s central capacity mechanism? The audit reported in §A6.2 is the operation that answers that question; it does not replay the proliferation argument.
The scope of Section A6 is therefore intermediate. It does not prove that ABC defeats every rival, and it does not run formal empirical tests. It states the boundary conditions, classifies rival explanations, specifies what kind of evidence each rival would require, and identifies the observations that would weaken or refute the model. Section A7 then converts these implications into formal propositions and a staged validation agenda. The contrast between Section A3’s gap argument and the present audit is summarised in Table A7.
Table A7. Methodological function of Section A3 and Section A6.
Table A7. Methodological function of Section A3 and Section A6.
Section Main question Methodological function Output
Section A3 Why are inherited adoption and intention frameworks incomplete for LLM-mediated evaluative reasoning? Gap argument and theoretical motivation Justification for a domain-bounded extension model
Section A6 What could make the ABC Model redundant, mis-scoped, confounded, or empirically misleading? Boundary and rival-explanation audit Boundary conditions, rival classifications, refutation conditions, and implications for the proposition set
A.6.1.1. Pre-Audit Orientation Map: Inherited Frameworks and Neighbouring Explanatory Families
Before reporting the rival-explanation audit, it is useful to separate a conceptual orientation map from the audit itself. Section A3 established the gap left by inherited acceptance and intention theories. The present section asks a different question: whether neighbouring constructs, theories, or diagnostic mechanisms could make ABC redundant, misclassified, or empirically misleading. The table and figure below therefore do not constitute the audit, a meta-analysis, or a final rival register. They provide a pre-audit map of the explanatory neighbourhood that the audit then classifies more formally.
Table A8 summarises the main inherited frameworks and adjacent explanatory families at a high level. Rows 1–3 restate the inherited acceptance/intention frameworks carried forward from Section A3. The remaining rows identify neighbouring families that recur around the ABC mechanism but require different treatments in the audit: some become rival explanations, some become calibration diagnostics, some become boundary conditions, and some become background comparators. The table is therefore an orientation device: it shows why the audit must classify boundary objects by role rather than treat all neighbouring constructs as equivalent psychometric rivals.
Table A8. Pre-audit orientation map: inherited frameworks, adjacent explanatory families, and the ABC response.
Table A8. Pre-audit orientation map: inherited frameworks, adjacent explanatory families, and the ABC response.
Framework or adjacent family What it explains well Core limitation under LLMs ABC response
TAM Usefulness, ease, attitude, adoption intention Treats technology mainly as an evaluated tool; does not capture decision-time cognitive scaffolding Retain PU/PEOU as adoption inputs; add PCA as cognitive-capacity mechanism
UTAUT/UTAUT2 Performance expectancy, effort expectancy, social influence, facilitating conditions, use behaviour Strong adoption/use model, but weak on internal decision-process transformation and calibration Retain as engagement layer; do not treat performance expectancy as equivalent to perceived capacity
TPB Attitude, subjective norm, perceived behavioural control, intention PBC does not distinguish internal ability from externally scaffolded perceived capacity; limited on affect and behaviour quality Refine PBC into tool-contingent capacity plus calibration
Risk/affect extensions Emotion-driven divergence and intention–behaviour gaps Explain affective override but not LLM-specific cognitive scaffolding Use as distortion/context mechanism, not core capacity construct
Human–AI trust/reliance constructs Trust, advice-taking, algorithm aversion/appreciation, over-reliance Explain reliance orientation, but not whether perceived capacity is warranted Treat as calibration diagnostics or rival explanations
Task–technology fit / cognitive fit Fit between task demands and system affordances Explains match between task and tool, but not the user’s felt expansion of cognitive capacity at decision time Treat as rival family and boundary condition
Cognitive offloading / dependence Delegation/substitution of cognitive work Can mimic assistance while reducing independent reasoning Keep outside core PCA; test through state-offloading and deference diagnostics
Note. The adjacent-family rows are broader orientation categories, not the final seven external rival families produced by the audit. The audit in §A6.2 applies a matched-domain admission screen and reclassifies the broader neighbourhood into boundary-object types, external rival families, distortion diagnostics, inherited model alternatives, and model-level rivals.
Table A8 is not audit output; it is pre-audit orientation devices that show the explanatory neighbourhood before the audit classifies that neighbourhood into role-specific boundary objects and rival explanations.
Having located the conceptual neighbourhood, the next subsection reports the actual audit procedure. That procedure expands beyond the orientation map, verifies candidate constructs and instruments, classifies them by threat family, and identifies the seven external rival families that remain on record for boundary treatment and future testing.
A.6.2. Methodological Basis of the Rival-Explanation Audit
The rival-explanation analysis reported in this section follows the same assisted-search, verification, and mapping logic introduced in §A3.1.1, but applies it to a different methodological task. Section A3 used independent LLM-mediated retrieval runs to map the literature supporting the LLM-specificity argument. The present audit uses the same general procedure to test whether the ABC Model’s central capacity mechanism can be absorbed by neighbouring constructs, paradigms, instruments, or model-level alternatives. The purpose is therefore not to conduct a second thematic review, estimate pooled effects, or treat large-language-model outputs as evidence. It is to expose the model’s capacity mechanism to rival explanations that could make it redundant, confounded, distorted, or theoretically underspecified.
The starting point was the existing proximal comparator space around Perceived Cognitive Assistance. Prior validation work tested PCA against four proximal comparator constructs - perceived usefulness, perceived ease of use, trust in the LLM, and trading self-efficacy - and against a baseline of additional controls and covariates, including TPB-related variables, risk tolerance, and LLM usage intensity. The wider audit asked a narrower and more adversarial question: even if PCA survives the tested TAM–TPB neighbourhood, could it still be absorbed by constructs and theories outside that neighbourhood?
Operationally, the audit used an AI-assisted adversarial scoping procedure followed by human verification, consolidation, and methodological classification. A single structured prompt fixed the focal construct definition, the hostile-reviewer task, the rival-construct filter, and eight a priori search domains: self-efficacy and control-belief extensions; personality and individual differences; ability and literacy constructs; information-systems process constructs; human–AI interaction constructs; metacognition and understanding constructs; affect, fluency, and illusion constructs; and distributed-cognition theory. Additional constructs surfaced during execution were retained when relevant but were not treated as additional a priori domains. The prompt-seeded frame contained 41 construct mentions, corresponding to 40 unique labels after one duplicate was collapsed. The same prompt was executed independently across three large-language-model systems in separate sessions. As in §A3.1.1, model outputs were treated as search-surface expansion devices and not as evidentiary authorities; candidate constructs, instruments, source claims, and redundancy arguments were checked against primary records, duplicate and overlapping entries were consolidated, and unsupported or misattributed candidates were removed.
At the source level, the three runs surfaced 94 unique source identities after author–year–title de-duplication, of which 28 appeared in at least two runs and 6 in all three; 66 appeared in one run only. After source and instrument verification, 61 references were retained in the working register and 33 were removed or consolidated on grounds of duplicate identity, author–instrument misattribution, or failure of a construct- and instrument-existence check. At the candidate-register level, the audit produced 55 review-register entries after consolidation, or 54 valid entries after a single unsupported instrument claim was removed.
The verified register was then classified according to the methodological role each candidate plays in the validation logic, because different rival types require different empirical responses. Four threat families were distinguished. Direct substitute threats are constructs operating at the same level of analysis and the same causal role as PCA, where redundancy could be tested by discriminant- or incremental-validity work. Antecedents and confounders are upstream conditions whose appropriate treatment is control, moderation, or longitudinal sequencing. Distortion or contamination mechanisms are not substitutes but interpretation threats: they leave a measured PCA score intact while making it ambiguous, and require diagnostic falsification rather than psychometric contest. Background theories are theoretical lenses or paradigms whose appropriate response is conceptual positioning rather than direct head-to-head testing. The four-family typology is the analytic frame applied to the full register; the seven external rival families are the survivors of the next-stage screen.
For candidates that appeared to be direct substitutes, a matched-domain admission screen was applied. A candidate was treated as eligible for direct latent-variable contest only if it operated at the same level of analysis, the same decision-time horizon, and the same causal role, and only if it had a validated multi-item self-report instrument with a matched item referent. A citation-integrity check was added as a further safeguard after one unsupported instrument claim was identified during verification. This screen produced seven boundary-relevant external rival families that any new construct in this space must reckon with: AI self-efficacy, cognitive absorption, cognitive load and workload, illusion of explanatory depth, AI dependence and over-reliance, task–technology fit, and explanation quality or perceived understanding. These seven are not final empirical comparators. They are external rival families requiring on-record defence and continuing boundary treatment as matched-domain instrumentation becomes available. The risk-mapping logic introduced in §A3.1.1 carries over to the present audit with one additional risk specific to the adversarial scoping task. Hostile-reviewer prompts of the kind used here may elicit adversarial sycophancy: the dialogic counterpart to sycophancy, in which a model overproduces objections matched to the framing of the request rather than objections it would surface unprompted. This risk is bounded by three operations. The eight a priori search domains fix the search frame in advance of the runs and prevent open-ended objection generation from becoming the admission rule. The matched-domain admission screen excludes candidates lacking validated instruments and aligned item referents, removing rivals whose case rests on framing alone. The citation-integrity check operates on the registry itself and removed one unsupported instrument claim during verification. The audit is therefore subject to the same documentary discipline as the literature search: model outputs are search-surface expansion devices, the four-threat-family classification is researcher-applied, and the seven external rival families that survive into the final register are retained on construct- and instrument-existence grounds rather than on the strength of the models’ framing.

A.6.3. Boundary Objects: Not Only Constructs

The boundary analysis in this section does not operate only on constructs. That framing would be too narrow and methodologically misleading. The audit identifies several different kinds of boundary object, and each requires a different empirical response. Treating them as if they were equivalent - as if all could be settled by a single CFA battery - would reproduce exactly the construct-confusion problem the audit was designed to surface. Table A9 distinguishes the seven boundary-object types that enter the analysis.
Table A9. Boundary objects.
Table A9. Boundary objects.
Boundary object Examples Methodological status Proper treatment
Proximal comparator constructs Perceived usefulness; perceived ease of use; trust in the LLM; trading self-efficacy Established measurable constructs close to PCA First-stage construct-boundary battery; CFA/ESEM/HTMT and incremental-validity testing
Additional baseline covariates TPB variables; risk tolerance; LLM usage intensity Controls and covariates in the validation baseline Included to reduce omitted-variable explanations
External rival families AI self-efficacy; cognitive absorption; workload; task–technology fit; perceived understanding Construct families or paradigms outside the tested TAM–TPB neighbourhood Classified by matched-domain fit; direct contest only when an appropriate instrument exists
Distortion diagnostics Processing fluency; overprecision; deference; reliance; state offloading; illusion of understanding Interpretation threats rather than substitute constructs Embedded in subsequent designs as calibration diagnostics
Background theories Distributed cognition; cognitive scaffolding; cognitive offloading; sensemaking Theoretical ancestry or framing lenses Used for positioning; not direct CFA rivals
Inherited model alternatives TAM/UTAUT-only model; TPB/control-belief model Whole alternative explanatory frameworks at the model level Evaluated through whether ABC adds conceptual and predictive value beyond inherited frameworks
Model-level rivals Calibration-only model; TAM+PCA; TPB+PCA; TTF+PCA Alternative explanations of the entire mechanism Compared through competing structural specifications and explanatory parsimony
This distinction resolves what could otherwise read as a “close four versus remote seven” tension. The close four - perceived usefulness, perceived ease of use, trust in the LLM, and trading self-efficacy - are proximal comparator constructs and have already been subjected to discriminant- and incremental-validity testing in prior validation work. The remote seven are not equivalent constructs. They are external rival families, some containing measurable constructs, some operating as paradigms, some functioning as distortion mechanisms, and some serving as broader theoretical lenses. The asymmetry between the two groups is substantive, not rhetorical: the close four are tested-close, whereas the remote seven are mostly blocked-remote - blocked from direct same-level psychometric contest because the matched-domain instruments required for that contest do not yet exist for retail-investor LLM decision episodes. The two groups therefore enter the validation logic through different routes. The close four enter through psychometric testing already underway. The remote seven enter through embedded distortion diagnostics now and through staged future testing as matched-domain instrumentation becomes available.
The methodological rule that follows is role-specific testing. If a rival is a direct substitute, the appropriate test is discriminant and incremental validity: confirmatory factor analysis, exploratory structural equation modelling, heterotrait–monotrait ratios, nested or competing factor models, and prediction beyond rival constructs (Campbell & Fiske, 1959; Clark & Watson, 2019; Henseler, Ringle, & Sarstedt, 2015; MacKenzie, Podsakoff, & Podsakoff, 2011). If a rival is an antecedent or confounder, the appropriate treatment is covariate control, moderation analysis, subgroup stratification, or longitudinal sequencing (Pearl, 2009; Shadish, Cook, & Campbell, 2002). If a rival is a distortion mechanism, the appropriate treatment is diagnostic falsification through objective-comprehension probes, confidence–accuracy calibration, overprecision measures, reliance and deference indicators, and state-offloading measures (Lee & See, 2004; Moore & Healy, 2008; Risko & Gilbert, 2016; Rozenblit & Keil, 2002). If a rival is a background theory, the appropriate response is conceptual positioning rather than direct psychometric contest (Suddaby, 2010; Whetten, 1989). A single empirical operation cannot answer all four threat families, and treating the seven external rivals as if they were a single empirical category would conflate threats whose appropriate responses differ.

A.6.4. Boundary Conditions: Where the Propositions Are Asserted

The main article’s Section 4 stated the model’s scope conditions - domain, population, cultural and regulatory context, model generation, design, and behavioural outcome - as conditions of theoretical applicability. This subsection restates the operative ones in the form the boundary-conditions tradition requires: as the value-ranges across which the propositions are asserted, and outside which they are not claimed to hold (Bacharach, 1989; Busse et al., 2017; Dubin, 1978; Whetten, 1989). The reframing converts a scope description into a falsifiability constraint: a finding that contradicts a proposition outside its stated range is not a refutation, whereas a finding that contradicts it inside the range is.
Five boundaries are asserted. The count is model-determined rather than method-prescribed - the tradition requires that value-ranges be stated but fixes no number - and five is the number of dimensions along which the ABC propositions can be misapplied. Cultural and regulatory context and model-generation era, listed among the main article’s §4 scope conditions, are treated here as generalisability questions rather than boundaries, because they bear on how the model would travel rather than on where its propositions are asserted at this stage. Different cuts are defensible: collapsing mechanism into phenomenon would yield four, and treating cultural and regulatory context and model-generation era as separate boundaries would yield seven.
The phenomenon boundary asserts the propositions for LLM-assisted evaluative reasoning, not technology adoption in general; TAM, UTAUT, and UTAUT2 remain the appropriate frameworks for access, acceptance, usefulness, ease, social influence, and continued use (Davis, 1989; Venkatesh et al., 2003; Venkatesh et al., 2012), which ABC inherits as its upstream Adoption-and-Engagement layer, so findings bearing only on adoption, intention, or use fall outside the range. The population boundary asserts them for self-directed retail investors making cognitively demanding decisions, not professional, institutional, advised, or non-investment users; adjacent populations are candidate generalisation domains for future testing. The task-context boundary asserts them for cognitively demanding activity requiring interpretation, comparison, scenario testing, or structural-complexity assessment, not simple lookup, generic chatbot use, or fully delegated allocation; that low-complexity tasks show weaker effects is itself a proposition (P3) rather than a free parameter. The mechanism boundary asserts them for perceived cognitive capacity operationalised through PCA and moderated by calibration, not objective ability, skill, or performance; because the distinctive predictions concern PCA conditional on calibration, high PCA alone is not a sufficient test condition. The outcome boundary asserts them for proximal behavioural choice - complexity intention, strategy evaluation, structural complexity, and risk assertion - not realised execution, return, or market-facing footprint, which enter only as a downstream extension.
The five boundaries are not arbitrary. Each is an explicit value-range in Bacharach’s (1989) sense and a system-state limit in Dubin’s (1978) sense, marking both where positive evidence supports the model and where negative evidence does not refute it. Each also locates a legitimate generalisation test that future work could attempt through separately validated extension - to professional populations, lower-complexity tasks, alternative mechanism specifications, or realised-outcome measurement - none of which is part of the present claim.

A.6.5. Rival-Explanation Matrix

The rival explanations implied by the audit can now be stated as a falsifiable structure. Table A10 reports, for each rival, the rival claim, the boundary status assigned by the audit, the observation that would weaken or refute the ABC Model under that rival, and the empirical operation appropriate to that rival’s threat family. The “what would weaken or refute” column is the operative one: it converts each rival from a label into a refutation condition that could in principle be observed. The “testing route” column states whether the operation is currently deployable or currently blocked by the absence of matched-domain instrumentation.
Table A10. Rival explanations, refutation conditions, and testing routes. (Sources of the ten rows. The matrix consolidates rivals from three sources rather than from the construct-proliferation audit alone. Two rows - the TAM/UTAUT-only and TPB/control-belief models - are inherited-framework rivals carried forward from §A3 and restated in rival form: they ask whether ABC adds explanatory value beyond the frameworks the model declares as upstream. Seven rows - AI self-efficacy, task–technology fit, cognitive absorption, cognitive load and workload, explanation quality or perceived understanding, illusion of explanatory depth, and dependence/over-reliance - are the seven external rival families produced by the matched-domain screen of the audit reported in §A6.2. One row - the calibration-only model - is an architectural rival produced by the model’s own internal structure, since ABC commits to a combination of PCA and calibration and is therefore exposed to the alternative in which calibration diagnostics carry all the explanatory weight. The audit’s four-threat-family typology (direct substitutes; antecedents and confounders; distortion mechanisms; background theories) is the classification scheme applied inside the boundary-status column, not a separate enumeration of rivals).
Table A10. Rival explanations, refutation conditions, and testing routes. (Sources of the ten rows. The matrix consolidates rivals from three sources rather than from the construct-proliferation audit alone. Two rows - the TAM/UTAUT-only and TPB/control-belief models - are inherited-framework rivals carried forward from §A3 and restated in rival form: they ask whether ABC adds explanatory value beyond the frameworks the model declares as upstream. Seven rows - AI self-efficacy, task–technology fit, cognitive absorption, cognitive load and workload, explanation quality or perceived understanding, illusion of explanatory depth, and dependence/over-reliance - are the seven external rival families produced by the matched-domain screen of the audit reported in §A6.2. One row - the calibration-only model - is an architectural rival produced by the model’s own internal structure, since ABC commits to a combination of PCA and calibration and is therefore exposed to the alternative in which calibration diagnostics carry all the explanatory weight. The audit’s four-threat-family typology (direct substitutes; antecedents and confounders; distortion mechanisms; background theories) is the classification scheme applied inside the boundary-status column, not a separate enumeration of rivals).
Rival explanation What it claims Boundary-status classification What would weaken or refute ABC under this rival Testing route
TAM/UTAUT-only model PCA is repackaged perceived usefulness or performance expectancy under another name. Inherited model rival; proximal comparator space. PCA fails CFA/ESEM/HTMT discriminant tests against perceived usefulness, or adds no incremental prediction of proximal choice once usefulness and performance expectancy are entered. Deployable: discriminant- and incremental-validity testing in the proximal comparator battery; competing factor models.
TPB/control-belief model PCA is trading self-efficacy or perceived behavioural control under another name. Inherited model rival; proximal comparator space. PCA is absorbed by trading self-efficacy or perceived behavioural control and has no residual decision-time role once those are controlled. Deployable: comparator battery including trading self-efficacy, perceived behavioural control, attitude, subjective norms, and risk tolerance.
AI self-efficacy model PCA is confidence in using AI effectively rather than perceived tool-conferred cognitive capacity. External rival family (direct-substitute candidate); possible antecedent. A matched LLM-trading AI self-efficacy measure absorbs PCA variance and removes PCA’s predictive role under matched-domain measurement. Currently blocked: requires a matched-domain AI self-efficacy adaptation with aligned item referent.
Task–technology fit / cognitive fit model PCA is perceived fit between the LLM and a complex trading task. External rival family (direct-substitute candidate). A matched task–technology-fit or cognitive-fit measure explains the same decision-time variance and predicts proximal choice equally well or better. Currently blocked: requires a matched-domain fit scale; competing structural models (TTF-only, TTF+PCA, ABC).
Cognitive absorption model PCA is engagement, flow, immersion, or perceived control during interaction. External rival family (direct-substitute candidate). PCA collapses into absorption or flow-like engagement rather than decision-time reasoning support under matched decision-episode measurement. Currently blocked at decision-episode resolution: requires absorption measurement adapted to LLM-assisted decision episodes.
Cognitive load / workload model PCA is merely reduced workload or lower mental demand. External rival family; possible control and diagnostic. Workload reduction fully explains PCA and its behavioural effects; PCA does not survive once workload is included. Deployable through workload probes embedded in controlled task-complexity designs.
Explanation quality / perceived understanding model PCA is satisfaction with understandable output or perceived explanation quality. External rival family; conceptual-substitute risk. A decision-time perceived-understanding measure absorbs PCA and proximal choice effects. Partly blocked: requires a purpose-built decision-time perceived-understanding scale; partially addressable via comprehension diagnostics in the meantime.
Illusion of explanatory depth / processing-fluency model High PCA reflects fluent explanation and false understanding rather than genuine scaffolding. Distortion diagnostic; contamination mechanism. High PCA predicts confidence or perceived understanding without objective comprehension or decision-quality improvement. Deployable: objective-understanding probes; confidence–accuracy calibration; fluency manipulation; overprecision measures.
Dependence / over-reliance / offloading model PCA is comfortable deference, cognitive substitution, or passive offloading. Distortion diagnostic; negative-pathway rival. High PCA predicts deference, reduced independent checking, state offloading, or reliance escalation rather than active scaffolded reasoning. Deployable: reliance and deference indicators; state-offloading measures; longitudinal tracking of reliance escalation and skill atrophy.
Calibration-only model Downstream choice is explainable by understanding, overprecision, reliance, and offloading diagnostics without a distinct PCA mechanism. Model-level rival. Calibration diagnostics predict proximal choice while PCA adds no residual explanatory role and does not interact with calibration. Deployable: moderated regression or moderated structural-equation modelling comparing PCA+Calibration, Calibration-only, and inherited-framework+Calibration specifications.
Note. “Rival explanation” is used here in the methodological sense established by Campbell and Fiske (1959) and Edwards and Bagozzi (2000) for construct-level rivals, and by Popper (1959) and Lakatos (1970) for theory-level rivals: any alternative account that, if true, would make the focal account redundant, mis-scoped, or refuted. Rivals enter the matrix in three forms. Construct-level rivals (e.g., AI self-efficacy) are individual constructs that could absorb PCA’s variance and are settled by discriminant- and incremental-validity testing. Framework-level rivals (e.g., a TAM/UTAUT-only or TPB-only account) are entire alternative explanatory apparatuses and are settled by competing structural specifications. Architectural rivals (e.g., the calibration-only model) are alternative internal structures of the same explanation and are settled by testing whether the focal mechanism survives when its competitors are entered alongside it. The three forms are reported in a single matrix because all three carry refutation conditions for the ABC Model, but they require different empirical operations as stated in the testing-route column.
The matrix is structured to do the work that Popper (1959) and Lakatos (1970) require of a defensible model. Each row converts an alternative account from a label into a specific observation that, if seen, would weaken or refute the ABC Model. The discriminating predictions are not equally easy to test. Two of the rivals - the TAM/UTAUT-only and TPB/control-belief models - sit inside the proximal comparator space already addressed in prior validation work and are deployable now. Four rivals - workload, illusion of explanatory depth, dependence/over-reliance, and the calibration-only model - are deployable through embedded distortion diagnostics or competing structural specifications in the next round of designs. Four external rival families - AI self-efficacy, task–technology fit, cognitive absorption, and explanation quality / perceived understanding - are blocked from direct same-level psychometric contest because matched-domain instruments do not yet exist for retail-investor LLM decision episodes. Their refutation conditions are stated, but the empirical operations that would test them require instrumentation development.
The asymmetry between deployable and blocked refutation conditions is itself a substantive result of the audit. It is the reason why the empirical priority in subsequent designs is distortion testing rather than head-to-head psychometric contest against every external rival. Stating the blockage explicitly is methodologically required: an audit that presented only the deployable contests and silently omitted the blocked ones would be exactly the kind of selective auxiliary protection that Lakatos (1970) classifies as degenerative. The matrix as written commits the model to refutation conditions it cannot yet test as well as to those it can.

A.6.6. What Would Refute the ABC Model

The model’s defensibility rests on four refutation conditions stated at the level of the architecture rather than at the level of a single rival - what would refute the core of the model, in Lakatos’s (1970) sense, as distinct from refutations of auxiliary or peripheral claims. Stating them at this level keeps the model’s most exposed commitments visible rather than letting them sit inside a longer proposition list where auxiliary modification could absorb them.
The first is construct redundancy. The model is refuted as a contribution if PCA cannot be distinguished, on factor-analytic or predictive grounds, from its proximal comparator space - perceived usefulness, perceived ease of use, trust in the LLM, and trading self-efficacy - under stronger designs (CFA, ESEM, bifactor, HTMT, incremental prediction). Closeness is acceptable and theoretically expected; absorption is not. The boundary claim is residual meaning, not distance.
The second is misclassification of external rivals. The model is weakened by treating AI self-efficacy, cognitive absorption, task–technology fit, perceived understanding, workload, illusion of explanatory depth, and dependence as irrelevant, and weakened in a different way by testing any of them as a direct comparator before a matched-domain instrument exists. It is defensible only if the seven external rivals are classified by threat family and tested through the empirical operation appropriate to that family.
The third is failure of the Calibration overlay. The distinctive claim is not that LLM use raises perceived capacity - that weaker claim could be absorbed by usefulness, fit, self-efficacy, or workload reduction - but that the behavioural meaning of perceived capacity changes across calibration conditions. High PCA is interpreted as calibrated enhancement when matched by objective comprehension, confidence accuracy, disciplined verification, and retained independent reasoning; it is interpreted as behavioural distortion when accompanied by overprecision, uncritical deference, perceived understanding without objective comprehension, or unverified offloading. The model is weakened at the architectural level if PCA predicts the same proximal-choice pattern regardless of calibration diagnostics - that is, if calibration does not moderate the PCA-to-Choice relationship. This condition depends on the §A5.4 specification commitment: naive additive compositing of the calibration components could manufacture a no-moderation result through cancellation of directionally opposing components rather than through genuine absence of moderation, so the architecture excludes that specification. The condition is currently programmatic rather than deployable because, as §A6.5 records, the matched-domain distortion-diagnostic battery it requires does not yet exist as an integrated instrument package. Whether calibrated enhancement later predicts higher ex ante decision quality is a separate criterion-validation question, not a premise of the Claim B classification.
The fourth is failure to survive calibration-only alternatives. If objective understanding, perceived understanding, overprecision, reliance, and offloading diagnostics jointly predict proximal choice while PCA adds no residual explanatory role and does not interact with them, PCA is at most a surface appraisal accompanying calibration rather than the mechanism through which LLM engagement becomes behavioural capacity - and the model’s central mechanism is refuted, even if PCA remains psychometrically distinct from its proximal comparators.
Section A7 converts each condition into a formal proposition with an explicit refutation condition and aligns it with the validation programme that tests it.

A.6.7. Implications for the Falsifiable Proposition Set

The boundary audit produces three structural implications for Section A7.
First, the proposition set must distinguish object types. Construct validity, causal mechanism, calibration moderation, proximal behavioural choice, longitudinal maturation effects, and downstream behavioural footprint are different theoretical objects with different empirical signatures, and the propositions must not be written as if a single empirical operation could test all of them. This is the proposition-level expression of the seven-object distinction in Table A9.
Second, calibration must enter the proposition set as a falsifiable moderator rather than as a conceptual qualifier. The model’s most distinctive claim - that perceived capacity supports empowerment when calibrated and distortion when not - is a moderation claim with a determinate refutation condition: no interaction between PCA and calibration diagnostics in predicting proximal choice. Stating it as anything weaker would forfeit the dual-pathway commitment and reduce the model to a usefulness-with-caveats account.
Third, the validation methods must be role-specific rather than uniform. The four-threat-family typology in §A6.2 implies that direct-substitute threats require discriminant- and incremental-validity tests; antecedent and confounder threats require control, moderation, and longitudinal sequencing; distortion threats require diagnostic falsification through objective-comprehension, confidence–accuracy, reliance, deference, and offloading measures; and model-level rivals require competing structural specifications. The proposition set in Section A7 must therefore align each proposition with the empirical operation appropriate to the threat it engages, rather than nominating a single dataset or design as a universal test.
The maximum claim defensible at the present stage of model development is precise. The ABC Model is a theoretically specified and empirically anchored causal-process architecture. PCA is the validated measurement backbone of the Capacity layer and has survived the first proximal comparator tests against the TAM–TPB neighbourhood. The wider audit identifies seven external rival families that must remain on record, of which most are currently blocked from direct same-level psychometric contest by the absence of matched-domain instruments in retail-investor settings. The model’s most exposed remaining vulnerability is not simple construct redundancy: it is the possibility that perceived cognitive capacity is miscalibrated through fluency, overprecision, deference, over-reliance, or offloading and thereby produces behavioural distortion rather than empowered reasoning. Section A7 turns that vulnerability and the others stated above into formal propositions and a staged validation agenda.

A.7. Falsifiable Propositions

This section operationalises the implications stated in §A6.7 as a specific set of falsifiable propositions for the ABC Model. It does not restate the architecture-level refutation conditions developed in §A6.6, nor the structural implications for the proposition set developed in §A6.7. It executes those implications by stating seven propositions whose refutation conditions are determinate and whose required empirical operations are aligned with the threat families identified in the rival-explanation audit.
Two prefatory points are necessary.
First, the obligation at this stage of model formulation is to advance the model coherently, anchor it empirically in some defensible portion, and specify what would have to be true for the model to fail (Edmondson & McManus, 2007; Lynham, 2002; Van de Ven, 2007). It is not to test all propositions simultaneously. Remaining propositions constitute open commitments handed to the field. This is the standard maturation pattern of a theoretical contribution in the social sciences (Corley & Gioia, 2011; Sutton & Staw, 1995; Whetten, 1989).
Second, the propositions are organised in three classes by their role in sustaining the model. Core propositions (P1–P5) are required to sustain it: if these fail, the model has no reason to be advanced. Developmental propositions (P6) test how the model behaves as use matures, and their failure would narrow its boundary conditions rather than collapse it. Extension propositions (P7) test realised-behaviour reach and are not required to defend the individual-level psychological claim. The classification follows the Lakatosian distinction between hard-core commitments and the protective belt around them (Lakatos, 1970; see also Popper, 1959, on the asymmetry of confirmation and refutation in scientific inference).
The proposition set is also organised around the two central theoretical claims introduced in §A5. P2 is the direct test of Claim A (the behavioural-migration claim: PCA → Choice incremental over inherited frameworks), with P1 as its antecedent precondition (LLM dialogic engagement must actually elevate PCA above static-tool baselines) and P3 as its boundary condition (the path strengthens under cognitively demanding decisions). P5 is the direct test of Claim B (the calibration overlay claim: the dual pathway between empowered and distorted choice). P4 is a discriminant prerequisite for both claims, since neither holds if PCA collapses into perceived usefulness, AI dependence, or passive offloading. P6 refines both claims under increasing user proficiency; P7 extends both into realised behaviour at the Footprint boundary. Claim A’s components are partially anchored - PCA measurement, PCA’s one-step incremental prediction of complexity intentions over inherited proximal comparators, and PCA’s discriminant validity from those comparators - but the end-to-end directed path implied by Claim A has been tested cross-sectionally and not detected, and is therefore not currently anchored as a confirmed mediation chain. Claim B remains theoretically specified, with the architectural refutation condition stated in §A6.6, but is empirically blocked at the operational level because the matched-domain distortion-diagnostic battery required to test P5 does not yet exist as an integrated instrument package in retail-investor research. The proposition set is constructed so that each proposition’s evidentiary status is reported separately rather than aggregated into a single claim of “anchoring.”
Part of the empirical anchoring already exists. PCA has been content-validated against perceived usefulness, perceived ease of use, trust in the LLM, and trading self-efficacy under a naïve-judge sort-and-rate procedure (Colquitt, Sabey, Rodell, & Hill, 2019; MacKenzie, Podsakoff, & Podsakoff, 2011), and confirmatory testing has supported a one-factor PCA model with strong reliability, scalar invariance across trading-experience and LLM-recency strata, and significant incremental prediction of complexity intentions above its proximal comparators. The PCA–perceived-usefulness boundary was close but defensible on HTMT, CFA, and ESEM criteria. P2 and P4 are therefore partially anchored rather than wholly open; the remaining propositions are open commitments whose required empirical operations are stated in Table A11.

A.7.1. The Seven Propositions

Table A11 states each proposition, its refutation condition, and the kind of empirical operation that the proposition’s refutation requires. The third column applies the four-threat-family logic stated in §A6.7: direct-substitute threats require discriminant- and incremental-validity tests (Cronbach & Meehl, 1955; MacKenzie et al., 2011; Shaffer, DeGeest, & Li, 2016); antecedent and confounder threats require control, moderation, and longitudinal sequencing (Aiken & West, 1991; Hayes, 2018; Quintana, 2024); distortion threats require diagnostic falsification through objective-comprehension, confidence–accuracy, reliance/deference, and offloading measures (Lee & See, 2004; Moore & Healy, 2008; Parasuraman & Manzey, 2010; Risko & Gilbert, 2016; Rozenblit & Keil, 2002); and model-level rivals require competing structural specifications (Brown, 2015; Edwards & Bagozzi, 2000; Pearl, 2009).
Table A11. Seven falsifiable propositions: classes, refutation conditions, and required empirical operations.
Table A11. Seven falsifiable propositions: classes, refutation conditions, and required empirical operations.
Proposition / Class Claim role Description and refutation condition Required empirical operation
P1 – Core – Dialogic scaffolding Antecedent precondition for Claim A Interactive LLM support produces higher PCA than static information, market screeners, generic summaries, or non-dialogic tools when investors face cognitively demanding trading decisions. Refuted if LLM support fails to raise PCA above static or non-dialogic tools after controlling for perceived usefulness, trust, and prior skill. Controlled comparison of dialogic vs static or non-dialogic conditions under matched evidence, with covariate adjustment for usefulness, trust, and prior skill (Aguinis & Bradley, 2014; Aiken & West, 1991; Shadish, Cook, & Campbell, 2002).
P2 – Core – Incremental validity Direct test of Claim A PCA predicts proximal behavioural choice - operationalised first as complexity intentions - beyond TAM/UTAUT and TPB variables (perceived usefulness, ease of use, trust, attitude, subjective norm, self-efficacy). Decision-process quality is treated as a secondary corollary. Refuted if PCA adds no residual explanatory value once those comparators are entered jointly. Hierarchical regression or structural equation modelling with TAM/UTAUT/TPB comparators entered jointly; incremental ΔR² and bias-corrected confidence intervals on PCA’s residual prediction (Cronbach & Meehl, 1955; MacKenzie et al., 2011; Shaffer et al., 2016).
P3 – Core – Task-complexity moderation Boundary condition on Claim A Task complexity moderates the PCA-to-Choice path: the incremental effect of PCA on proximal behavioural choice should be stronger in cognitively demanding, ambiguous, multi-step trading decisions than in simple, low-ambiguity decisions. Decision-process quality is a secondary corollary. Refuted if the PCA-to-Choice effect is equal, weaker, or only present in trivial tasks. Within- or between-subjects manipulation of decision complexity; PCA × complexity interaction in predicting proximal choice; conditional-process estimation (Aiken & West, 1991; Hayes, 2018; Pearl, 2009).
P4 – Core – Discriminant validity Discriminant prerequisite shared by Claims A and B PCA remains empirically distinguishable from perceived usefulness, AI dependence, passive cognitive offloading, and over-reliance, while partially relating to perceived usefulness through usefulness judgements. High PCA reflects active scaffolded reasoning, not stopped thinking. Refuted if PCA collapses into perceived usefulness under stronger designs (CFA, ESEM, bifactor) or is fully absorbed by dependence and offloading measures. Confirmatory factor analysis, exploratory structural equation modelling, and bifactor specifications; HTMT and incremental-prediction tests; discriminant-validity contests against dependence and offloading instruments where matched-domain measures exist (Brown, 2015; Edwards & Bagozzi, 2000; MacKenzie et al., 2011; Shaffer et al., 2016).
P5 – Core – Calibration-contingent empowerment / distortion Direct test of Claim B Elevated PCA is interpreted as calibrated enhancement when perceived capacity is matched by objective understanding, confidence accuracy, disciplined verification, and retained independent reasoning, and as behavioural distortion when perceived capacity is paired with overprecision, uncritical deference, perceived understanding without objective comprehension, or unverified offloading. Refuted if PCA predicts the same proximal-choice pattern regardless of pre-specified calibration diagnostics modelled under the §A5.4 miscalibration-gap specification or its configural alternative; in that case Claim B’s calibration-contingent enhancement/distortion distinction is not sustained. This proposition does not claim that calibrated enhancement produces higher decision quality; that is a later criterion-validation test. Embedded distortion-diagnostic battery - objective comprehension, overprecision, perceived understanding, reliance/deference, state offloading, and matched verification or independent-reasoning referents - operationalised under the miscalibration-gap specification committed to in §A5.4 as primary, with the configural specification preserved as an alternative; PCA × miscalibration-gap moderation tests on proximal choice (Buçinca, Malaya, & Gajos, 2021; Lee & See, 2004; Moore & Healy, 2008; Risko & Gilbert, 2016; Rozenblit & Keil, 2002; Schemmer, Kühl, Benz, Bartos, & Satzger, 2023). A later independent decision-quality score may be used as criterion-validity evidence, but not as part of the classification rule.
P6 – Developmental – Proficiency moderation Developmental refinement of Claims A and B LLM proficiency and adoption maturity strengthen the productive PCA pathway and reduce distortion risk. Within-person trajectories should show calibration improving with experience and disciplined verification practice. Refuted if proficiency and maturity do not moderate PCA effects, or if more experienced users are no better calibrated than novices. Longitudinal designs with within-person trajectories; literacy-sensitive specifications tested against undifferentiated usage models; proficiency-stratified moderation tests (Edmondson & McManus, 2007; Quintana, 2024).
P7 – Extension – Behavioural footprint Realised-behaviour extension downstream of both claims Higher LLM engagement and PCA, conditional on calibration, are associated with observable shifts in retail trading behaviour, including multi-leg strategy use, trade frequency, position concentration, and volatility exposure. Refuted if expected behavioural shifts do not follow PCA increases, or if shifts vanish under basic controls for transaction costs, liquidity, market volatility, and execution quality. Trading-record analysis with financial-econometric controls (transaction costs, liquidity, market volatility, execution quality), linking individual-level PCA to realised behavioural shifts (Barber & Odean, 2000, 2001).

A.7.2. Reading the Proposition Set

Three observations follow.
First, P5 - the calibration-contingent enhancement/distortion proposition - is the proposition-level operationalisation of Claim B and consequently the model’s most distinctive theoretical commitment and its hardest empirical claim. Its refutation condition cannot be evaluated with attitudinal or self-report measures alone; it requires the embedded distortion-diagnostic battery noted in Table A11, comprising objective comprehension, overprecision, perceived understanding, reliance/deference, state-offloading indicators, and matched verification or independent-reasoning referents (Buçinca et al., 2021; Lee & See, 2004; Moore & Healy, 2008; Risko & Gilbert, 2016; Rozenblit & Keil, 2002; Schemmer et al., 2023; Si, Goyal, Wu, Zhao, Feng, Daumé, & Boyd-Graber, 2024). The battery is to be operationalised under the miscalibration-gap specification committed to in §A5.4, not as a naive additive composite of directionally heterogeneous components. A null moderation result under the gap specification is eligible to count as evidence against Claim B, subject to instrument quality, statistical power, and design adequacy. A null moderation result under the naive-additive specification is not similarly interpretable, because cancellation between adequacy and distortion components inside the composite cannot be distinguished from genuine absence of moderation. The refutation condition for P5 is therefore conditional on both the matched-domain battery and the §A5.4 specification commitment. Absent that battery and specification, Claim B’s calibration-contingent enhancement/distortion distinction is untested, not supported. A later test linking calibrated enhancement to independently scored ex ante decision quality may provide criterion-validity evidence for the classification, but it is not part of Claim B itself.
Second, the proposition set deliberately omits a free-standing mediation claim of the form “LLM use → PCA → behavioural choice” from the core proposition list. That chain was tested cross-sectionally in the confirmatory PCA work and was not detected: the bias-corrected indirect-effect confidence interval included zero. Two readings of this non-detection are admissible - that the chain is theoretically misspecified, or that it is empirically attenuated in early-stage adoption settings where penetration, capability stability, and prompting literacy are still uneven across participants (MacKinnon, 2008; Mathieu & Taylor, 2006; Quintana, 2024) - and the present evidence does not adjudicate between them. The model’s working position is therefore that the directed path Capacity → Choice is currently supported as a one-step incremental prediction (P2) and as a psychometric distinctiveness claim (the discriminant component of P4), but not as a confirmed end-to-end mediation. P3 and P6 are advanced as independently motivated propositions about task-complexity and proficiency moderation, not as recovery routes for the unconfirmed mediation; their refutation conditions in Table A11 stand or fall on their own evidence and do not absorb the non-detection of the indirect effect. Whether the underlying pathway becomes recoverable under conditions specified by P3 or P6 is an open empirical question that subsequent designs are positioned to answer, not a theoretical inference adopted in advance of those designs.
Third, the set as a whole sustains the Lakatosian discipline introduced in §A6.6 and maps cleanly onto the two central claims of §A5. The hard core distributes as follows: P1 as the antecedent precondition for Claim A’s Capacity layer, P2 and P3 sustaining Claim A through incremental and contextual validity, P4 sustaining the discriminant prerequisite shared by both claims, and P5 sustaining Claim B through the Calibration overlay / diagnostic moderation layer. Refuting any of P2, P3, P4, or P5 would weaken Claim A or Claim B at the architecture level; refuting P1 would remove the precondition on which Claim A rests. The developmental and extension claims (P6, P7) refine and extend the reach of both claims, and their failure would adjust that reach without dismantling either central claim (Lakatos, 1970; Popper, 1959). The proposition set is not contingent on any particular instrument or research line being available now; it requires only that the operations identified in Table A11 be available somewhere in the field’s evidentiary record before the corresponding proposition can be defended or refuted.

References

  1. Agarwal, R., & Karahanna, E. (2000). Time flies when you’re having fun: Cognitive absorption and beliefs about information technology usage. MIS Quarterly, 24(4), 665–694. [CrossRef]
  2. Agnew, J. R., & Szykman, L. R. (2005). Asset allocation and information overload: The influence of information display, asset choice, and investor experience. Journal of Behavioral Finance, 6(2), 57–70. [CrossRef]
  3. Aguinis, H., & Bradley, K. J. (2014). Best practice recommendations for designing and implementing experimental vignette methodology studies. Organizational Research Methods, 17(4), 351–371. [CrossRef]
  4. Aiken, L. S., & West, S. G. (1991). Multiple regression: Testing and interpreting interactions. Sage.
  5. Ajzen, I. (1991). The theory of planned behavior. Organizational Behavior and Human Decision Processes, 50(2), 179–211. [CrossRef]
  6. Ajzen, I. (2015). The theory of planned behaviour is alive and well, and not ready to retire: A commentary on Sniehotta, Presseau, and Araújo-Soares. Health Psychology Review, 9(2), 131–137. [CrossRef]
  7. Almeida Lima, V., Bellei, V., Ballesteros Martins, L. F., & Terlizzi, M. A. (2024). Time flies when you are having fun: Cognitive absorption and beliefs about ChatGPT usage. AIS Transactions on Replication Research, 10, Article 5. [CrossRef]
  8. Alter, A. L., & Oppenheimer, D. M. (2009). Uniting the tribes of fluency to form a metacognitive nation. Personality and Social Psychology Review, 13(3), 219–235. [CrossRef]
  9. Apesteguia, J., Oechssler, J., & Weidenholzer, S. (2020). Copy trading. Management Science, 66(12), 5608–5622. [CrossRef]
  10. Argyle, L. P., Busby, E. C., Gubler, J. R., Lyman, A., Olcott, J., Pond, J., & Wingate, D. (2025). Testing theories of political persuasion using AI. Proceedings of the National Academy of Sciences, 122(18), Article e2412815122. [CrossRef]
  11. Bacharach, S. B. (1989). Organizational theories: Some criteria for evaluation. Academy of Management Review, 14(4), 496–515. [CrossRef]
  12. Bagozzi, R. P. (2007). The legacy of the technology acceptance model and a proposal for a paradigm shift. Journal of the Association for Information Systems, 8(4), 244–254. [CrossRef]
  13. Bahaj, A., Rahimi, H., Chetouani, M., & Ghogho, M. (2025). Gauging overprecision in LLMs: An empirical study [Preprint]. arXiv. [CrossRef]
  14. Bandura, A. (1977). Self-efficacy: Toward a unifying theory of behavioral change. Psychological Review, 84(2), 191–215. [CrossRef]
  15. Barber, B. M., & Odean, T. (2000). Trading is hazardous to your wealth: The common stock investment performance of individual investors. Journal of Finance, 55(2), 773–806. [CrossRef]
  16. Barber, B. M., & Odean, T. (2001). Boys will be boys: Gender, overconfidence, and common stock investment. Quarterly Journal of Economics, 116(1), 261–292. [CrossRef]
  17. Baron, R. M., & Kenny, D. A. (1986). The moderator–mediator variable distinction in social psychological research: Conceptual, strategic, and statistical considerations. Journal of Personality and Social Psychology, 51(6), 1173–1182. [CrossRef]
  18. Bauer, R., Cosemans, M., & Eichholtz, P. (2009). Option trading and individual investor performance. Journal of Banking & Finance, 33(4), 731–746. [CrossRef]
  19. Bellofatto, A., D’Hondt, C., & De Winne, R. (2018). Subjective financial literacy and retail investors’ behavior. Journal of Banking & Finance, 92, 168–181. [CrossRef]
  20. Benbasat, I., & Barki, H. (2007). Quo vadis TAM? Journal of the Association for Information Systems, 8(4), 211–218. [CrossRef]
  21. Bollen, K. A. (1989). Structural equations with latent variables. Wiley.
  22. Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2), 77–101. [CrossRef]
  23. Brenner, L., & Meyll, T. (2020). Robo-advisors: A substitute for human financial advice? Journal of Behavioral and Experimental Finance, 25, Article 100275. [CrossRef]
  24. Brown, T. A. (2015). Confirmatory factor analysis for applied research (2nd ed.). Guilford Press.
  25. Buçinca, Z., Malaya, M. B., & Gajos, K. Z. (2021). To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), Article 188. [CrossRef]
  26. Busse, C., Kach, A. P., & Wagner, S. M. (2017). Boundary conditions: What they are, how to explore them, why we need them, and when to consider them. Organizational Research Methods, 20(4), 574–609. [CrossRef]
  27. Bussone, A., Stumpf, S., & O’Sullivan, D. (2015). The role of explanations on trust and reliance in clinical decision support systems. In 2015 IEEE International Conference on Healthcare Informatics (pp. 160–169). IEEE. [CrossRef]
  28. Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait–multimethod matrix. Psychological Bulletin, 56(2), 81–105. [CrossRef]
  29. Carton, A. M. (2025). The six dimensions of strong theory. Organization Science, 36(4), 1242–1270. [CrossRef]
  30. Cassinadri, G. (2024). ChatGPT and the technology-education tension: Applying contextual virtue epistemology to a cognitive artifact. Philosophy & Technology, 37(1), Article 14. [CrossRef]
  31. Chiambaretto, P., Fernandez, A.-S., & Le Roy, F. (2025). What coopetition is and what it is not: Defining the “hard core” and the “protective belt” of coopetition. Strategic Management Review, 6(1–2), 17–52. [CrossRef]
  32. Clark, A., & Chalmers, D. (1998). The extended mind. Analysis, 58(1), 7–19. [CrossRef]
  33. Clark, L. A., & Watson, D. (2019). Constructing validity: New developments in creating objective measuring instruments. Psychological Assessment, 31(12), 1412–1427. [CrossRef]
  34. Cohn, M., Pushkarna, M., Olanubi, G. O., Moran, J. M., Padgett, D., Mengesha, Z., & Heldreth, C. (2024). Believing anthropomorphism: Examining the role of anthropomorphic cues on trust in large language models. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery. [CrossRef]
  35. Colombatto, C., Birch, J., & Fleming, S. M. (2025). The influence of mental state attributions on trust in large language models. Communications Psychology, 3, Article 84. [CrossRef]
  36. Colquitt, J. A., Sabey, T. B., Rodell, J. B., & Hill, E. T. (2019). Content validation guidelines: Evaluation criteria for definitional correspondence and definitional distinctiveness. Journal of Applied Psychology, 104(10), 1243–1265. [CrossRef]
  37. Compeau, D. R., & Higgins, C. A. (1995). Computer self-efficacy: Development of a measure and initial test. MIS Quarterly, 19(2), 189–211. [CrossRef]
  38. Conner, M. (2015). Extending not retiring the theory of planned behaviour: A commentary on Sniehotta, Presseau, and Araújo-Soares. Health Psychology Review, 9(2), 141–145. [CrossRef]
  39. Corley, K. G., & Gioia, D. A. (2011). Building theory about theory building: What constitutes a theoretical contribution? Academy of Management Review, 36(1), 12–32. [CrossRef]
  40. Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281–302. [CrossRef]
  41. Dahl, M., Magesh, V., Suzgun, M., & Ho, D. E. (2024). Large legal fictions: Profiling legal hallucinations in large language models. Journal of Legal Analysis, 16(1), 64–93. [CrossRef]
  42. Dang Anh-Hoang, D. A., Vu Tran, V. T., & Le-Minh Nguyen, L. N. (2025). Survey and analysis of hallucinations in large language models: Attribution to prompting strategies or model behavior. Frontiers in Artificial Intelligence, 8, Article 1622292. [CrossRef]
  43. Danry, V., Pataranutaporn, P., Groh, M., Epstein, Z., & Maes, P. (2024). Deceptive AI systems that give explanations are more convincing than honest AI systems and can amplify belief in misinformation [Preprint]. arXiv. [CrossRef]
  44. Davis, F. D. (1989). Perceived usefulness, perceived ease of use, and user acceptance of information technology. MIS Quarterly, 13(3), 319–340. [CrossRef]
  45. Davis, F. D., & Granić, A. (2024). The technology acceptance model: 30 years of TAM. Springer. [CrossRef]
  46. Diamantopoulos, A., & Winklhofer, H. M. (2001). Index construction with formative indicators: An alternative to scale development. Journal of Marketing Research, 38(2), 269–277. [CrossRef]
  47. Dietvorst, B. J., Simmons, J. P., & Massey, C. (2015). Algorithm aversion: People erroneously avoid algorithms after seeing them err. Journal of Experimental Psychology: General, 144(1), 114–126. [CrossRef]
  48. DiMaggio, P. J. (1995). Comments on “What theory is not.” Administrative Science Quarterly, 40(3), 391–397. [CrossRef]
  49. Dubin, R. (1978). Theory building (2nd ed.). Free Press.
  50. Dwivedi, Y. K., Rana, N. P., Jeyaraj, A., Clement, M., & Williams, M. D. (2019). Re-examining the unified theory of acceptance and use of technology (UTAUT): Towards a revised theoretical model. Information Systems Frontiers, 21(3), 719–734. [CrossRef]
  51. Dzindolet, M. T., Peterson, S. A., Pomranky, R. A., Pierce, L. G., & Beck, H. P. (2003). The role of trust in automation reliance. International Journal of Human-Computer Studies, 58(6), 697–718. [CrossRef]
  52. Edmondson, A. C., & McManus, S. E. (2007). Methodological fit in management field research. Academy of Management Review, 32(4), 1246–1264. [CrossRef]
  53. Edwards, J. R., & Bagozzi, R. P. (2000). On the nature and direction of relationships between constructs and measures. Psychological Methods, 5(2), 155–174. [CrossRef]
  54. Eigner, E., & Händler, T. (2024). Determinants of LLM-assisted decision-making [Preprint]. arXiv. [CrossRef]
  55. Epley, N., Waytz, A., & Cacioppo, J. T. (2007). On seeing human: A three-factor theory of anthropomorphism. Psychological Review, 114(4), 864–886. [CrossRef]
  56. Fabrigar, L. R., Wegener, D. T., MacCallum, R. C., & Strahan, E. J. (1999). Evaluating the use of exploratory factor analysis in psychological research. Psychological Methods, 4(3), 272–299. [CrossRef]
  57. Fanous, A., Goldberg, J., Agarwal, A. A., Lin, J., Zhou, A., Xu, S., Bikia, V., Daneshjou, R., & Koyejo, S. (2025). SycEval: Evaluating LLM sycophancy. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 8, 893–900.
  58. Fernbach, P. M., Rogers, T., Fox, C. R., & Sloman, S. A. (2013). Political extremism is supported by an illusion of understanding. Psychological Science, 24(6), 939–946. [CrossRef]
  59. Fleming, S. M. (2024). Metacognition and confidence: A review and synthesis. Annual Review of Psychology, 75, 241–268. [CrossRef]
  60. Funk, P. F., Hoch, C. C., Knoedler, S., Knoedler, L., Cotofana, S., Sofo, G., Bashiri Dezfouli, A., Wollenberg, B., Guntinas-Lichius, O., & Alfertshofer, M. (2024). ChatGPT’s response consistency: A study on repeated queries of medical examination questions. European Journal of Investigation in Health, Psychology and Education, 14(3), 657–668. [CrossRef]
  61. Gedikli, F., Jannach, D., & Ge, M. (2014). How should I explain? A comparison of different explanation types for recommender systems. International Journal of Human-Computer Studies, 72(4), 367–382. [CrossRef]
  62. Gimmelberg, D., & Ludviga, I. (2025). Strategic complexity and behavioral distortion: Retail investing under large language model augmentation. International Journal of Financial Studies, 13(4), Article 210. [CrossRef]
  63. Gimmelberg, D., & Ludviga, I. (2026a). Perceived cognitive assistance in LLM-augmented retail trading: Construct definition and content validation. International Journal of Financial Studies, 14(4), Article 83. [CrossRef]
  64. Gimmelberg, D., & Ludviga, I. (2026b). Perceived cognitive assistance and the construct-proliferation challenge: An integrated evaluative judgment that is never finished. Comparator selection, rival scoping, and a falsifiable boundary agenda [Preprint]. SSRN. [CrossRef]
  65. Gimmelberg, D., & Ludviga, I. (2026c). Psychometric validation of the Perceived Cognitive Assistance Scale: Exploratory and confirmatory evidence from LLM-supported decision-making in retail trading [Preprint]. PsyArXiv. [CrossRef]
  66. Glikson, E., & Woolley, A. W. (2020). Human trust in artificial intelligence: Review of empirical research. Academy of Management Annals, 14(2), 627–660. [CrossRef]
  67. Goh, A. Y. H., Hartanto, A., & Majeed, N. M. (2025). Generative artificial intelligence dependency: Scale development, validation, and its motivational, behavioral, and psychological correlates. Computers in Human Behavior Reports, 20, Article 100845. [CrossRef]
  68. Gollwitzer, P. M., & Oettingen, G. (2015). From studying the determinants of action to analysing its regulation: A commentary on Sniehotta, Presseau and Araújo-Soares. Health Psychology Review, 9(2), 146–150. [CrossRef]
  69. Goodhue, D. L. (2007). Comment on Benbasat and Barki’s “Quo Vadis TAM” article. Journal of the Association for Information Systems, 8(4), 219–222. [CrossRef]
  70. Goodhue, D. L., & Thompson, R. L. (1995). Task-technology fit and individual performance. MIS Quarterly, 19(2), 213–236. [CrossRef]
  71. Gregor, S., & Benbasat, I. (1999). Explanations from intelligent systems: Theoretical foundations and implications for practice. MIS Quarterly, 23(4), 497–530. [CrossRef]
  72. Hackenburg, K., Tappin, B. M., Hewitt, L., Saunders, E., Black, S., Lin, H., Fist, C., Margetts, H., Rand, D. G., & Summerfield, C. (2025). The levers of political persuasion with conversational artificial intelligence. Science, 390(6777), Article eaea3884. [CrossRef]
  73. Handler, A., Larsen, K. R., & Hackathorn, R. (2024). Large language models present new questions for decision support. International Journal of Information Management, 79, Article 102811. [CrossRef]
  74. Hart, S. G., & Staveland, L. E. (1988). Development of NASA-TLX (Task Load Index): Results of empirical and theoretical research. In P. A. Hancock & N. Meshkati (Eds.), Human mental workload (pp. 139–183). North-Holland. [CrossRef]
  75. Hayes, A. F. (2018). Introduction to mediation, moderation, and conditional process analysis: A regression-based approach (2nd ed.). Guilford Press.
  76. Heersmink, R., de Rooij, B., Clavel Vázquez, M. J., & Colombo, M. (2024). A phenomenology and epistemology of large language models: Transparency, trust, and trustworthiness. Ethics and Information Technology, 26(3), Article 41. [CrossRef]
  77. Henseler, J., Ringle, C. M., & Sarstedt, M. (2015). A new criterion for assessing discriminant validity in variance-based structural equation modeling. Journal of the Academy of Marketing Science, 43(1), 115–135. [CrossRef]
  78. Herlihy, C., Neville, J., Schnabel, T., & Swaminathan, A. (2024). On overcoming miscalibrated conversational priors in LLM-based chatbots. In Proceedings of the Fortieth Conference on Uncertainty in Artificial Intelligence (Proceedings of Machine Learning Research, Vol. 244, pp. 1599–1620). PMLR.
  79. Hicks, M. T., Humphries, J., & Slater, J. (2024). ChatGPT is bullshit. Ethics and Information Technology, 26(2), Article 38. [CrossRef]
  80. Hollan, J., Hutchins, E., & Kirsh, D. (2000). Distributed cognition: Toward a new foundation for human-computer interaction research. ACM Transactions on Computer-Human Interaction, 7(2), 174–196. [CrossRef]
  81. Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., & Liu, T. (2025). A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, 43(2), Article 42. [CrossRef]
  82. Hutchins, E. (1995). Cognition in the wild. MIT Press.
  83. Huy, L. V., Nguyen, H. T., Vo-Thanh, T., Thinh, N. H. T., & Thi Thu Dung, T. (2024). Generative AI, why, how, and outcomes: A user adoption study. AIS Transactions on Human-Computer Interaction, 16(1), 1–27. [CrossRef]
  84. Jakesch, M., Bhat, A., Buschek, D., Zalmanson, L., & Naaman, M. (2023). Co-writing with opinionated language models affects users’ views. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Article 111, pp. 1–15). Association for Computing Machinery. [CrossRef]
  85. Jarvis, C. B., MacKenzie, S. B., & Podsakoff, P. M. (2003). A critical review of construct indicators and measurement model misspecification in marketing and consumer research. Journal of Consumer Research, 30(2), 199–218. [CrossRef]
  86. Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), Article 248. [CrossRef]
  87. Johns, G. (2006). The essential impact of context on organizational behavior. Academy of Management Review, 31(2), 386–408. [CrossRef]
  88. Kahneman, D., & Tversky, A. (1979). Prospect theory: An analysis of decision under risk. Econometrica, 47(2), 263–291.
  89. Kaur, A. (2025). Echoes of agreement: Argument driven sycophancy in large language models. In Findings of the Association for Computational Linguistics: EMNLP 2025 (pp. 22803–22812). Association for Computational Linguistics. [CrossRef]
  90. Klein, K. J., Dansereau, F., & Hall, R. J. (1994). Levels issues in theory development, data collection, and analysis. Academy of Management Review, 19(2), 195–229. [CrossRef]
  91. Koriat, A. (1997). Monitoring one’s own knowledge during study: A cue-utilization approach to judgments of learning. Journal of Experimental Psychology: General, 126(4), 349–370. [CrossRef]
  92. Kozlowski, S. W. J., & Klein, K. J. (2000). A multilevel approach to theory and research in organizations: Contextual, temporal, and emergent processes. In K. J. Klein & S. W. J. Kozlowski (Eds.), Multilevel theory, research, and methods in organizations: Foundations, extensions, and new directions (pp. 3–90). Jossey-Bass.
  93. Kudina, O., Ballsun-Stanton, B., & Alfano, M. (2025). The use of large language models as scaffolds for proleptic reasoning. Asian Journal of Philosophy, 4(1), Article 24. [CrossRef]
  94. Laban, P., Hayashi, H., Zhou, Y., & Neville, J. (2026). LLMs get lost in multi-turn conversation. In The Fourteenth International Conference on Learning Representations. https://openreview.net/forum?id=VKGTGGcwl6.
  95. Lakatos, I. (1970). Falsification and the methodology of scientific research programmes. In I. Lakatos & A. Musgrave (Eds.), Criticism and the growth of knowledge (pp. 91–196). Cambridge University Press.
  96. Lazarsfeld, P. F., & Henry, N. W. (1968). Latent structure analysis. Houghton Mifflin.
  97. Lee, D., Pruitt, J., Zhou, T., Du, J., & Odegaard, B. (2025). Metacognitive sensitivity: The key to calibrating trust and optimal decision making with AI. PNAS Nexus, 4(5), Article pgaf133. [CrossRef]
  98. Lee, J. D., & See, K. A. (2004). Trust in automation: Designing for appropriate reliance. Human Factors, 46(1), 50–80. [CrossRef]
  99. Lee, Y., Kozar, K. A., & Larsen, K. R. T. (2003). The technology acceptance model: Past, present, and future. Communications of the Association for Information Systems, 12, 752–780. [CrossRef]
  100. Leppink, J., Paas, F., Van der Vleuten, C. P. M., Van Gog, T., & Van Merriënboer, J. J. G. (2013). Development of an instrument for measuring different types of cognitive load. Behavior Research Methods, 45(4), 1058–1072. [CrossRef]
  101. Li, J., Yang, Y., Liao, Q. V., Zhang, J., & Lee, Y.-C. (2025). As confidence aligns: Understanding the effect of AI confidence on human self-confidence in human-AI decision making. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery. [CrossRef]
  102. Lichtenstein, S., Fischhoff, B., & Phillips, L. D. (1982). Calibration of probabilities: The state of the art to 1980. In D. Kahneman, P. Slovic, & A. Tversky (Eds.), Judgment under uncertainty: Heuristics and biases (pp. 306–334). Cambridge University Press.
  103. Liew, T. W., Lim, C. T. M., Khan, M. T. I., & Tan, S.-M. (2025). Banking on voice: AI attributes, technology perceptions, and trust in banking voicebot acceptance. Computers in Human Behavior Reports, 20, Article 100812. [CrossRef]
  104. Lin, H., Czarnek, G., Lewis, B., White, J. P., Berinsky, A. J., Costello, T., Pennycook, G., & Rand, D. G. (2025). Persuading voters using human–artificial intelligence dialogues. Nature, 648(8093), 394–401. [CrossRef]
  105. Logg, J. M., Minson, J. A., & Moore, D. A. (2019). Algorithm appreciation: People prefer algorithmic to human judgment. Organizational Behavior and Human Decision Processes, 151, 90–103. [CrossRef]
  106. Lusardi, A., & Mitchell, O. S. (2014). The economic importance of financial literacy: Theory and evidence. Journal of Economic Literature, 52(1), 5–44. [CrossRef]
  107. Lynham, S. A. (2002). The general method of theory-building research in applied disciplines. Advances in Developing Human Resources, 4(3), 221–241. [CrossRef]
  108. Ma, S., Wang, X., Lei, Y., Shi, C., Yin, M., & Ma, X. (2024). “Are you really sure?” Understanding the effects of human self-confidence calibration in AI-assisted decision making. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery. [CrossRef]
  109. MacKenzie, S. B., Podsakoff, P. M., & Podsakoff, N. P. (2011). Construct measurement and validation procedures in MIS and behavioral research: Integrating new and existing techniques. MIS Quarterly, 35(2), 293–334. [CrossRef]
  110. MacKinnon, D. P. (2008). Introduction to statistical mediation analysis. Lawrence Erlbaum Associates.
  111. Malmqvist, L. (2024). Sycophancy in large language models: Causes and mitigations. arXiv. [CrossRef]
  112. Marchionini, G. (2006). Exploratory search: From finding to understanding. Communications of the ACM, 49(4), 41–46. [CrossRef]
  113. Mathieu, J. E., & Taylor, S. R. (2006). Clarifying conditions and decision points for mediational type inferences in organizational behavior. Journal of Organizational Behavior, 27(8), 1031–1056. [CrossRef]
  114. McCutcheon, A. L. (1987). Latent class analysis. Sage.
  115. Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741–749. [CrossRef]
  116. Mogaji, E., Viglia, G., Srivastava, P., & Dwivedi, Y. K. (2024). Is it the end of the technology acceptance model in the era of generative artificial intelligence? International Journal of Contemporary Hospitality Management, 36(10), 3324–3339. [CrossRef]
  117. Moore, D. A., & Healy, P. J. (2008). The trouble with overconfidence. Psychological Review, 115(2), 502–517. [CrossRef]
  118. Nass, C., & Moon, Y. (2000). Machines and mindlessness: Social responses to computers. Journal of Social Issues, 56(1), 81–103. [CrossRef]
  119. Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., & Lowe, R. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35, 27730–27744.
  120. Ouyang, S., Zhang, J. M., Harman, M., & Wang, M. (2025). An empirical study of the non-determinism of ChatGPT in code generation. ACM Transactions on Software Engineering and Methodology, 34(2), Article 42. [CrossRef]
  121. Paas, F., Tuovinen, J. E., Tabbers, H., & Van Gerven, P. W. M. (2003). Cognitive load measurement as a means to advance cognitive load theory. Educational Psychologist, 38(1), 63–71. [CrossRef]
  122. Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration. Human Factors, 52(3), 381–410. [CrossRef]
  123. Parasuraman, R., & Riley, V. (1997). Humans and automation: Use, misuse, disuse, abuse. Human Factors, 39(2), 230–253. [CrossRef]
  124. Payne, J. W., Bettman, J. R., & Johnson, E. J. (1993). The adaptive decision maker. Cambridge University Press.
  125. Pearl, J. (2009). Causality: Models, reasoning, and inference (2nd ed.). Cambridge University Press.
  126. Podsakoff, P. M., MacKenzie, S. B., & Podsakoff, N. P. (2016). Recommendations for creating better concept definitions in the organizational, behavioral, and social sciences. Organizational Research Methods, 19(2), 159–203. [CrossRef]
  127. Popper, K. R. (1959). The logic of scientific discovery. Hutchinson.
  128. Quintana, R. (2024). Asking-and answering-causal questions using longitudinal data. Quality & Quantity, 58(5), 4679–4701. [CrossRef]
  129. Reynolds, P. D. (1971). A primer in theory construction. Macmillan.
  130. Risko, E. F., & Gilbert, S. J. (2016). Cognitive offloading. Trends in Cognitive Sciences, 20(9), 676–688. [CrossRef]
  131. Rousseau, D. M. (1985). Issues of level in organizational research: Multi-level and cross-level perspectives. Research in Organizational Behavior, 7, 1–37.
  132. Rozenblit, L., & Keil, F. (2002). The misunderstood limits of folk science: An illusion of explanatory depth. Cognitive Science, 26(5), 521–562. [CrossRef]
  133. Rubin, M. (2025). The replication crisis is less of a “crisis” in Lakatos’ philosophy of science than it is in Popper’s. European Journal for Philosophy of Science, 15(1), Article 5. [CrossRef]
  134. Sabbah, J., & Li, F. (2025). When humans and large language models collaborate, problem-finding illuminates. Innovation: Organization & Management. Advance online publication. [CrossRef]
  135. Salam, M., Farooq, M. S., Ikram, A., Shahzad, M., Ali, A., & Jaafar, N. (2025). Revised artificial intelligence device use acceptance (RAIDUA) model: Exploring privacy concerns for socially responsible AI deployment and ethical AI leadership. Journal of Hospitality and Tourism Insights. Advance online publication. [CrossRef]
  136. Sarkar, A. (2024). AI should challenge, not obey. Communications of the ACM, 67(10), 18–21. [CrossRef]
  137. Sarraf, S., Kar, A. K., & Janssen, M. (2024). How do system and user characteristics, along with anthropomorphism, impact cognitive absorption of chatbots-Introducing SUCCAST through a mixed methods study. Decision Support Systems, 178, Article 114132. [CrossRef]
  138. Saunders, B., Sim, J., Kingstone, T., Baker, S., Waterfield, J., Bartlam, B., Burroughs, H., & Jinks, C. (2018). Saturation in qualitative research: Exploring its conceptualization and operationalization. Quality & Quantity, 52(4), 1893–1907. [CrossRef]
  139. Schemmer, M., Kühl, N., Benz, C., Bartos, A., & Satzger, G. (2023). Appropriate reliance on AI advice: Conceptualization and the effect of explanations. In Proceedings of the 28th International Conference on Intelligent User Interfaces (IUI ‘23). Association for Computing Machinery. [CrossRef]
  140. Scherer, R. (2025). Is the Technology Acceptance Model just old wine in new wineskins? Exploring issues for further model development. Journal of University Teaching and Learning Practice, 22(8). [CrossRef]
  141. Schittko, M., Planing, P., & Müller, P. (2026). Stability through human perception: Technology acceptance models’ robustness across various interaction perspectives and comparable technologies. Computers in Human Behavior Reports, 21, Article 100930. [CrossRef]
  142. Schwarzer, R. (2015). Some retirees remain active: A commentary on Sniehotta, Presseau and Araújo-Soares. Health Psychology Review, 9(2), 138–140. [CrossRef]
  143. Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and quasi-experimental designs for generalized causal inference. Houghton Mifflin.
  144. Shaffer, J. A., DeGeest, D., & Li, A. (2016). Tackling the problem of construct proliferation: A guide to assessing the discriminant validity of conceptually related constructs. Organizational Research Methods, 19(1), 80–110. [CrossRef]
  145. Shanahan, M., McDonell, K., & Reynolds, L. (2023). Role play with large language models. Nature, 623(7987), 493–498. [CrossRef]
  146. Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., Cheng, N., Durmus, E., Hatfield-Dodds, Z., Johnston, S. R., Kravec, S., Maxwell, T., McCandlish, S., Ndousse, K., Rausch, O., Schiefer, N., Yan, D., Zhang, M., & Perez, E. (2024). Towards understanding sycophancy in language models. In The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=tvhaxkMKAn.
  147. Si, C., Goyal, N., Wu, T., Zhao, C., Feng, S., Daumé, H., III, & Boyd-Graber, J. (2024). Large language models help humans verify truthfulness-Except when they are convincingly wrong. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1: Long Papers (pp. 1459–1474). Association for Computational Linguistics. [CrossRef]
  148. Simon, H. A. (1955). A behavioral model of rational choice. The Quarterly Journal of Economics, 69(1), 99–118. [CrossRef]
  149. Simon, H. A. (1960). The new science of management decision. Harper & Row.
  150. Singh, A. K., Devkota, S., Lamichhane, B., Dhakal, U., & Dhakal, C. (2023). The confidence-competence gap in large language models: A cognitive study [Preprint]. arXiv. [CrossRef]
  151. Slovic, P. (1987). Perception of risk. Science, 236(4799), 280–285. [CrossRef]
  152. Smart, P., Clowes, R., & Clark, A. (2025). ChatGPT, extended: Large language models and the extended mind. Synthese, 205, Article 242. [CrossRef]
  153. Sniehotta, F. F., Presseau, J., & Araújo-Soares, V. (2014). Time to retire the theory of planned behaviour. Health Psychology Review, 8(1), 1–7. [CrossRef]
  154. Sparrow, B., Liu, J., & Wegner, D. M. (2011). Google effects on memory: Cognitive consequences of having information at our fingertips. Science, 333(6043), 776–778. [CrossRef]
  155. Spatharioti, S. E., Rothschild, D. M., Goldstein, D. G., & Hofman, J. M. (2023). Comparing traditional and LLM-based search for consumer choice: A randomized experiment [Preprint]. arXiv. [CrossRef]
  156. Steyvers, M., Tejeda, H., Kumar, A., Belem, C., Karny, S., Hu, X., Mayer, L. W., & Smyth, P. (2025). What large language models know and what people think they know. Nature Machine Intelligence, 7(2), 221–231. [CrossRef]
  157. Storm, B. C., & Stone, S. M. (2015). Saving-enhanced memory: The benefits of saving on the learning and remembering of new information. Psychological Science, 26(2), 182–188. [CrossRef]
  158. Straub, D. W., & Burton-Jones, A. (2007). Veni, vidi, vici: Breaking the TAM logjam. Journal of the Association for Information Systems, 8(4), 223–229. [CrossRef]
  159. Suddaby, R. (2010). Editor’s comments: Construct clarity in theories of management and organization. Academy of Management Review, 35(3), 346–357. [CrossRef]
  160. Sun, F., Li, N., Wang, K., & Goette, L. (2025). Large language models are overconfident and amplify human bias [Preprint]. arXiv. [CrossRef]
  161. Sundar, S. S. (2020). Rise of machine agency: A framework for studying the psychology of human–AI interaction. Journal of Computer-Mediated Communication, 25(1), 74–88. [CrossRef]
  162. Sutton, R. I., & Staw, B. M. (1995). What theory is not. Administrative Science Quarterly, 40(3), 371–384. [CrossRef]
  163. Tankelevitch, L., Kewenig, V., Simkute, A., Scott, A. E., Sarkar, A., Sellen, A., & Rintel, S. (2024). The metacognitive demands and opportunities of generative AI. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (CHI ‘24). Association for Computing Machinery. [CrossRef]
  164. Todd, P., & Benbasat, I. (1999). Evaluating the impact of DSS, cognitive effort, and incentives on strategy selection. Information Systems Research, 10(4), 356–374. [CrossRef]
  165. Torraco, R. J. (2005). Writing integrative literature reviews: Guidelines and examples. Human Resource Development Review, 4(3), 356–367. [CrossRef]
  166. Tversky, A., & Kahneman, D. (1974). Judgment under uncertainty: Heuristics and biases. Science, 185(4157), 1124–1131. [CrossRef]
  167. Van de Ven, A. H. (2007). Engaged scholarship: A guide for organizational and social research. Oxford University Press.
  168. Vasconcelos, H., Jörke, M., Grunde-McLaughlin, M., Gerstenberg, T., Bernstein, M. S., & Krishna, R. (2023). Explanations can reduce overreliance on AI systems during decision-making. Proceedings of the ACM on Human-Computer Interaction, 7(CSCW1), Article 129. [CrossRef]
  169. Venkatesh, V., Morris, M. G., Davis, G. B., & Davis, F. D. (2003). User acceptance of information technology: Toward a unified view. MIS Quarterly, 27(3), 425–478. [CrossRef]
  170. Venkatesh, V., Thong, J. Y. L., & Xu, X. (2012). Consumer acceptance and use of information technology: Extending the unified theory of acceptance and use of technology. MIS Quarterly, 36(1), 157–178. [CrossRef]
  171. Vessey, I. (1991). Cognitive fit: A theory-based analysis of the graphs versus tables literature. Decision Sciences, 22(2), 219–240. [CrossRef]
  172. von Nordenflycht, A. (2023). Clean up your theory! Invest in theoretical clarity and consistency for higher-impact research. Organization Science, 34(5), 1981–1996. [CrossRef]
  173. Wacker, J. G. (1998). A definition of theory: Research guidelines for different theory-building research methods in operations management. Journal of Operations Management, 16(4), 361–385. [CrossRef]
  174. Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13, Article 14045. [CrossRef]
  175. Wang, B., Rau, P.-L. P., & Yuan, T. (2023). Measuring user competence in using artificial intelligence: Validity and reliability of artificial intelligence literacy scale. Behaviour & Information Technology, 42(9), 1324–1337. [CrossRef]
  176. Wang, K., Lu, Y., & Pan, Z. (2025). Understanding users’ effective use of generative conversational AI from a media naturalness perspective: A hybrid structural equation modeling-artificial neural network (SEM-ANN) approach. Data Science and Management, 8(2), 147–159. [CrossRef]
  177. Wang, Y.-Y., & Chuang, Y.-W. (2024). Artificial intelligence self-efficacy: Scale development and validation. Education and Information Technologies, 29, 4785–4808. [CrossRef]
  178. Webster, J., & Watson, R. T. (2002). Analyzing the past to prepare for the future: Writing a literature review. MIS Quarterly, 26(2), xiii–xxiii.
  179. Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E. H., Le, Q. V., & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35, 24824–24837.
  180. Weick, K. E. (1989). Theory construction as disciplined imagination. Academy of Management Review, 14(4), 516–531. [CrossRef]
  181. Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P.-S., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., Kenton, Z., Brown, S., Hawkins, W., Stepleton, T., Biles, C., Birhane, A., Haas, J., Rimell, L., Hendricks, L. A., … Gabriel, I. (2022). Taxonomy of risks posed by language models. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (pp. 214–229). Association for Computing Machinery. [CrossRef]
  182. Whetten, D. A. (1989). What constitutes a theoretical contribution? Academy of Management Review, 14(4), 490–495. [CrossRef]
  183. Williams, M. D., Rana, N. P., & Dwivedi, Y. K. (2015). The unified theory of acceptance and use of technology (UTAUT): A literature review. Journal of Enterprise Information Management, 28(3), 443–488. [CrossRef]
  184. Wood, D., Bruner, J. S., & Ross, G. (1976). The role of tutoring in problem solving. Journal of Child Psychology and Psychiatry, 17(2), 89–100. [CrossRef]
  185. Xu, Z., Song, T., & Lee, Y.-C. (2025). Confronting verbalized uncertainty: Understanding how LLM’s verbalized uncertainty influences users in AI-assisted decision-making. International Journal of Human–Computer Studies, 197, Article 103455. [CrossRef]
  186. Xue, L., Ghazali, N., & Mahat, J. (2025). A systematic review of UTAUT and UTAUT2 for AI adoption in education. International Journal of Human–Computer Interaction. Advance online publication. [CrossRef]
Figure 1. Nine-step model-development flow used to construct the ABC Model, ordered from problem definition to the closing statement of contribution and evidence status.
Figure 1. Nine-step model-development flow used to construct the ABC Model, ordered from problem definition to the closing statement of contribution and evidence status.
Preprints 221466 g001
Figure 2. Conceptual contrast between ordinary decision support tools and LLMs across six evaluative-stage dimensions. Note. Scores range from 1 (low) to 5 (high) and represent conceptual synthesis values rather than empirical measurements. The figure visualises the comparative profile of ordinary tools and large language models across six evaluative-stage dimensions derived from the 2023–2026 thematic literature review reported in Appendix A. On the metacognitive calibration and distortion risk dimension, a higher score indicates greater exposure to miscalibration and over-reliance, not a more desirable property.
Figure 2. Conceptual contrast between ordinary decision support tools and LLMs across six evaluative-stage dimensions. Note. Scores range from 1 (low) to 5 (high) and represent conceptual synthesis values rather than empirical measurements. The figure visualises the comparative profile of ordinary tools and large language models across six evaluative-stage dimensions derived from the 2023–2026 thematic literature review reported in Appendix A. On the metacognitive calibration and distortion risk dimension, a higher score indicates greater exposure to miscalibration and over-reliance, not a more desirable property.
Preprints 221466 g002
Figure 3. Three-layer insufficiency of inherited models: the temporal (substrate), ontological (locus of cognition), and phenomenological (LLM-mediated mechanism) gaps between inherited adoption and intention theories (TAM, UTAUT, UTAUT2, TPB) and LLM-mediated evaluative reasoning, with the eight phenomenological mechanisms.
Figure 3. Three-layer insufficiency of inherited models: the temporal (substrate), ontological (locus of cognition), and phenomenological (LLM-mediated mechanism) gaps between inherited adoption and intention theories (TAM, UTAUT, UTAUT2, TPB) and LLM-mediated evaluative reasoning, with the eight phenomenological mechanisms.
Preprints 221466 g003
Figure 4. The ABC Model architecture. Note. Solid elements mark the three-stage core (Capacity → Choice, with Calibration as the diagnostic/moderating overlay); dashed elements mark the boundary stages (Adoption and Engagement inherited from TPB/TAM/UTAUT/UTAUT2 upstream; Behavioural Footprint as a downstream extension).
Figure 4. The ABC Model architecture. Note. Solid elements mark the three-stage core (Capacity → Choice, with Calibration as the diagnostic/moderating overlay); dashed elements mark the boundary stages (Adoption and Engagement inherited from TPB/TAM/UTAUT/UTAUT2 upstream; Behavioural Footprint as a downstream extension).
Preprints 221466 g004
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings