A.3.2. Phenomenon and Gap Analysis
Large language models (LLMs) affect the evaluative stage of decision-making differently from ordinary adoption tools because they are not only objects of adoption or external aids. Classical technology-adoption theories explain whether users judge a system as useful, easy to use, socially supported, or controllable (Ajzen, 1991; Davis, 1989; Venkatesh et al., 2003, 2012). Decision-support and fit theories add that technologies can improve task performance when system capabilities and information representations fit the task (Goodhue & Thompson, 1995; Todd & Benbasat, 1999; Vessey, 1991). Explanation research further shows that intelligent systems can justify recommendations and change user evaluations of system output (Gedikli et al., 2014; Gregor & Benbasat, 1999). These literatures remain necessary, but they are incomplete for LLM-assisted judgement because they mostly preserve a separation between the human evaluator and the technological instrument: the tool retrieves, displays, filters, recommends, or automates, while the user remains the main site of interpretation.
LLMs weaken that separation. Their distinctive interface is natural-language, multi-turn, generative, and context-sensitive. The user can externalise a partial thought, receive a structured answer, challenge the answer, request alternatives, ask for counterarguments, and continue the same reasoning thread. Recent decision-support scholarship therefore treats LLMs as a general-purpose decision-support technology rather than another narrow information system, and the naturalness of conversational AI affects users’ cognitive effort and ambiguity (Handler et al., 2024; Wang et al., 2025). Instruction-following and chain-of-thought prompting make the interaction appear reasoning-like, even when the underlying system is not reasoning in a human sense (Ouyang et al., 2022; Wei et al., 2022). Sundar’s account of machine agency, work on social responses to computers, and anthropomorphism research explain why this modality is psychologically consequential: the user increasingly communicates with the machine rather than merely through the machine (Epley et al., 2007; Glikson & Woolley, 2020; Nass & Moon, 2000; Sundar, 2020). Heersmink et al. (2024) sharpen this point for LLMs by arguing that they are sufficiently transparent at the interactional level to support fluent use, yet sufficiently opaque internally to make trust calibration difficult.
The relevant decision point is the evaluative stage, not the adoption stage. Evaluation is where the actor interprets an ambiguous situation, forms a mental representation of what is at stake, compares alternatives, estimates risk, and asks whether action is intelligible and feasible. Bounded rationality and heuristics mean that this stage is always simplified and cue-driven rather than fully optimal (Kahneman & Tversky, 1974; Simon, 1955; Slovic, 1987). It is also self-regulatory: people form beliefs about whether they understand the issue and whether they are capable of acting, and contemporary metacognition research emphasises that confidence judgements are inferential and can diverge from objective performance (Bandura, 1977; Fleming, 2024). In retail investing, the problem is intensified by information overload, uneven and partly subjective financial literacy, attention-driven trading, and recurrent overconfidence (Agnew & Szykman, 2005; Barber & Odean, 2000, 2001; Bellofatto et al., 2018; Lusardi & Mitchell, 2014). Therefore, a technology that changes how the investor builds the decision frame can affect more than adoption intention; it can change the investor’s judgement coordinate itself.
The first pathway is empowerment through cognitive scaffolding. LLMs can structure an analytical path, translate a broad view into candidate tactics, compare alternatives side by side, produce what-if scenarios, and highlight inconsistencies. This is why the strongest theoretical anchor is not ordinary usefulness alone but scaffolding, distributed cognition, extended cognition, and cognitive offloading (Clark & Chalmers, 1998; Hollan et al., 2000; Hutchins, 1995; Risko & Gilbert, 2016; Wood et al., 1976). In this interpretation, the LLM functions as a conversational cognitive artefact: it does not simply provide an answer, but helps organise the user’s reasoning activity. Human-AI decision research supports the importance of this distinction. Cognitive forcing functions can reduce overreliance, explanations can shape reliance, and LLM explanations can increase verification efficiency even when accuracy gains are limited or overreliance appears when the model is wrong (Buçinca et al., 2021; Schemmer et al., 2023; Si et al., 2024). Naturalness of conversational interaction further matters because it changes felt cognitive effort and the experience of ambiguity (Wang et al., 2025).
The second pathway is distortion through metacognitive miscalibration. The same features that make LLMs useful as scaffolds can also inflate perceived understanding, confidence, and control. People often infer competence from fluency, coherence, and processing ease (Alter & Oppenheimer, 2009; Koriat, 1997). They can overestimate what they understand, become overprecise in belief, or mistake a plausible explanation for genuine comprehension (Moore & Healy, 2008; Rozenblit & Keil, 2002). LLMs intensify this risk because they present coherent, human-like, testimony-style output; anthropomorphic cues, mental-state attributions, AI confidence, and verbalised uncertainty can all affect trust, perceived accuracy, or user self-confidence (Cohn et al., 2024; Colombatto et al., 2025; Epley et al., 2007; Li et al., 2025; Xu et al., 2025). Explanations are not a guaranteed cure: in clinical decision support, fuller explanations increased trust but also produced over-reliance (Bussone et al., 2015), and AI-generated explanations can be more persuasive than bare classifications and can amplify belief in misinformation when the explanations are deceptive (Danry et al., 2024). Co-writing with opinionated language models can shift users’ expressed views and even subsequent attitudes (Jakesch et al., 2023). Trust-in-automation research further warns that users can misuse, disuse, or overuse automation when reliance is poorly calibrated (Dzindolet et al., 2003; Glikson & Woolley, 2020; Lee & See, 2004; Parasuraman & Manzey, 2010; Parasuraman & Riley, 1997). Hallucination and sycophancy risks add a specifically generative-AI version of the same problem (Ji et al., 2023; Malmqvist, 2024).
Table A3 reports the resulting LLM specificity matrix. Each row pairs one of the six analytic dimensions specified a priori in §A3.1 with the supporting references mapped to that dimension under the protocol described in Step 6, and the n column gives the count of distinct sources tagged to the dimension.
Table A3.
LLM specificity matrix.
Table A3.
LLM specificity matrix.
| Specificity dimension |
Key finding |
Supporting references (each entry corresponds to one back-matter reference) |
n |
Relevance for PCA and ABC model |
| Dialogic and contextual interaction |
LLMs make the interface conversational and adaptive. The user can ask, challenge, refine, and continue the reasoning thread instead of only retrieving or accepting a fixed output. |
Sundar (2020); Nass and Moon (2000); Heersmink et al. (2024); Ouyang et al. (2022); Wei et al. (2022); Cohn et al. (2024); Colombatto et al. (2025); Xu et al. (2025); Logg et al. (2019); Epley et al. (2007); Glikson and Woolley (2020); Wang et al. (2025) |
12 |
Places the system inside the decision episode as a machine interlocutor rather than as a static adopted tool. |
| Generative and explanatory reasoning |
LLMs generate explanations, comparisons, counterarguments, and stepwise reasoning-like text, but the same generativity can also create plausible but wrong, sycophantic, or persuasively misleading explanations. |
Gregor and Benbasat (1999); Gedikli et al. (2014); Ji et al. (2023); Ouyang et al. (2022); Wei et al. (2022); Si et al. (2024); Malmqvist (2024); Handler et al. (2024); Marchionini (2006) |
9 |
Explains why LLM assistance is not just retrieval: it supplies interpretive structure and justification. |
| Evaluative co-construction and problem framing |
Evaluation is the stage in which the actor interprets ambiguity, compares alternatives, judges risk, and assesses whether action is manageable. LLMs can participate directly in this construction of the decision frame, including by shifting users’ expressed views. |
Simon (1955); Kahneman and Tversky (1974); Slovic (1987); Ajzen (1991); Davis (1989); Goodhue and Thompson (1995); Vessey (1991); Todd and Benbasat (1999); Bandura (1977); Fleming (2024); Jakesch et al. (2023) |
11 |
Bridges technology-adoption logic with behavioural decision theory and metacognitive monitoring. |
| Cognitive scaffolding and offloading |
LLMs can reduce felt cognitive load, externalise partial thoughts, structure multi-step reasoning, and support comparison and scenario simulation in ways consistent with distributed and extended cognition. |
Wood et al. (1976); Hutchins (1995); Clark and Chalmers (1998); Risko and Gilbert (2016); Buçinca et al. (2021); Schemmer et al. (2023); Ma et al. (2024); Li et al. (2025); Hollan et al. (2000); Wang et al. (2025) |
10 |
Provides the theoretical anchor for Perceived Cognitive Assistance as felt expansion of decision-time capability. |
| Metacognitive calibration and distortion risk |
Fluent, confident, human-like explanations can inflate perceived understanding, certainty, and reliance even when the underlying reasoning remains weak or unverified; explanations themselves can amplify acceptance of incorrect or deceptive claims. |
Alter and Oppenheimer (2009); Koriat (1997); Rozenblit and Keil (2002); Moore and Healy (2008); Dzindolet et al. (2003); Lee and See (2004); Parasuraman and Riley (1997); Parasuraman and Manzey (2010); Dietvorst et al. (2015); Logg et al. (2019); Cohn et al. (2024); Colombatto et al. (2025); Li et al. (2025); Xu et al. (2025); Bussone et al. (2015); Danry et al. (2024); Fleming (2024); Glikson and Woolley (2020) |
18 |
Requires calibration diagnostics alongside PCA so empowerment is not confused with distorted confidence. |
| Ecosystem maturity and user familiarity |
Retail investors face overload, uneven literacy (including subjective literacy), attention-driven trading, and overconfidence. Ordinary tools mainly improve access, filtering, automation, or adoption intention; they do not usually co-construct evaluative reasoning. |
Agnew and Szykman (2005); Apesteguia et al. (2020); Barber and Odean (2000); Barber and Odean (2001); Brenner and Meyll (2020); Compeau and Higgins (1995); Davis (1989); Lusardi and Mitchell (2014); Venkatesh et al. (2003); Venkatesh et al. (2012); Bellofatto et al. (2018) |
11 |
Explains why LLM effects are especially important in self-directed trading and investing contexts. |
Table A4 makes the contrast with ordinary adoption tools explicit. The cells record conceptual synthesis scores assigned by the author against a five-point capability rubric: 1 indicates that the dimension is absent from or outside the technology’s design space; 2 that it is present only incidentally or as a side-effect of other features; 3 that it is supported but not a central design property; 4 that it is supported and well-developed; and 5 that it is a defining design property of the technology class. Scoring is single-rater and is presented as a synthesis aid rather than an empirical measurement, but each cell value is corroborated by two independent evidence streams: the 60-reference base anchoring
Table A3, where the per-dimension reference counts give the LLM-side scores their footing in the literature search; and prior published work by the author and collaborator on LLM-versus-baseline cognitive asymmetries in retail-investor decision-making (Gimmelberg & Ludviga, 2025), which provides parallel evidence for the ordinary-tools-versus-LLMs contrast across the six dimensions. The synthesis is reported here to anchor the visual contrast in
Figure 2 (reported in the manuscript body) rather than as an independent empirical contribution.
Table A4.
Conceptual heatmap of ordinary tools versus LLMs across the six specificity dimensions.
Table A4.
Conceptual heatmap of ordinary tools versus LLMs across the six specificity dimensions.
| Dimension |
Ordinary tools |
LLMs |
Interpretation |
| Dialogic and contextual interaction |
1 |
5 |
LLMs sustain multi-turn natural-language exchange and retain context across turns; ordinary tools present fixed interfaces. |
| Generative and explanatory reasoning |
2 |
5 |
LLMs generate explanations, alternatives, and stepwise rationales rather than presenting only fixed outputs. |
| Evaluative co-construction and problem framing |
1 |
5 |
LLMs participate directly in defining the problem, comparing alternatives, and shaping the decision frame. |
| Cognitive scaffolding and offloading |
2 |
5 |
LLMs can structure reasoning steps and reduce felt complexity; this is the design property that anchors Perceived Cognitive Assistance (PCA). |
| Metacognitive calibration and distortion risk |
2 |
5 |
LLMs can inflate perceived understanding, certainty, and reliance through fluency and human-like output; higher score indicates greater exposure to distortion, not a better outcome. |
| Ecosystem maturity and user familiarity |
5 |
3 |
Ordinary tools are the established, well-tested retail-investor toolkit - search, dashboards, automated execution, robo-advice - embedded in the everyday decision workflow and operated with native, low-friction familiarity. LLMs are converging on this territory from two directions: they are absorbing tool-like features (retrieval, structured and visual output, dashboard-style summaries) and are increasingly embedded inside the incumbent platforms themselves. But they are newer to the role and users have not yet built comparable fluency, so the established toolkit retains supremacy on this axis. This is also the most temporally unstable of the six dimensions: the gap is expected to narrow as feature-integration and embedding advance. |
Figure 2 (reported in the manuscript body) visualises the same scores as a radar plot. Plotting the two profiles on shared axes makes the dominant pattern visible at a glance: ordinary adoption tools dominate on the ecosystem-maturity-and-familiarity axis, where they constitute the established retail-investor toolkit; LLMs dominate on the dialogic, generative, evaluative, scaffolding, and calibration-risk axes, which together define the LLM-specific decision-time mechanism the ABC Model is designed to capture.
A.3.3. Limits of Existing Theories: Why TAM, UTAUT, and TPB Are Necessary but Incomplete
The case advanced here is that the established acceptance and intention-based models - Davis’s (1989) Technology Acceptance Model (TAM), Venkatesh, Morris, Davis, and Davis’s (2003) Unified Theory of Acceptance and Use of Technology (UTAUT), Venkatesh, Thong, and Xu’s (2012) UTAUT2, and Ajzen’s (1991) Theory of Planned Behaviour (TPB) - are necessary but incomplete for the analysis of large-language-model-mediated decision-making. The frameworks examined here - TAM, UTAUT, UTAUT2, and TPB - are selected on three convergent criteria. First, they are the dominant inherited frameworks in the AI-acceptance and behavioural-intention literatures, accounting for the bulk of the construct vocabulary that LLM-acceptance studies currently re-use. Second, together they span the adoption → intention → behaviour pipeline that LLM-mediated decision-making appears to disrupt. Third, each has been independently named by the published critique tradition (§A3.3.1) as facing strain under generative AI. The selection is therefore not contestable on grounds of arbitrary target choice: these are the frameworks the field itself has identified as the locus of difficulty.
The case is built in three layers. The temporal layer concerns the empirical substrate against which these theories were calibrated and the substrate on which they are now being applied. The ontological layer concerns the implicit assumptions the theories make about where cognition lives, what counts as a tool, and how intention forms. The phenomenological layer concerns the appearance, in current AI-mediated activity, of mechanisms with no principled location in the inherited construct space. None of the three layers refutes the inherited theories within their original scope. Together they specify why their application to LLM-mediated behaviour at the evaluative stage of decision-making is an extrapolation rather than a deployment, and why a domain-bounded extension is required.
The three layers are organised from least to most demanding. The temporal layer asks whether the inherited theories were ever calibrated against the empirical substrate to which they are now being applied; this is the weakest form of objection and does not refute the theories within their original scope. The ontological layer asks whether the constructs the theories carry can in principle host the new phenomena; this is a stronger objection because it concerns the theories’ implicit assumptions rather than their domain of application. The phenomenological layer asks whether the empirical phenomena that LLM-mediated activity actually produces have a principled location in the inherited construct space; this is the strongest form of objection because it identifies named mechanisms the inherited frameworks cannot host. The three layers are exhaustive in the sense that any disciplined critique of an inherited theory under a new technology must operate at one of these three levels.
The argument is positioned within an established critical tradition rather than offered against an unanimous orthodoxy. A series of published critiques has already documented limitations of each of the three frameworks; some of those critiques explicitly anticipate the type of difficulty the present analysis develops. The first task of this section is therefore to show what the existing critique tradition has already established, what it has not yet specified, and where the present argument adds precision.
A.3.3.1. The Published Critique Tradition: What Is Already on the Record
A first body of critique is internal to the technology-acceptance research programme. The 2007 special issue of the Journal of the Association for Information Systems titled “Quo Vadis TAM” brought together several senior commentators on the trajectory of TAM-based research. Benbasat and Barki (2007) argued that the intense focus on TAM had diverted research attention from other important issues in information-systems use, had produced an illusion of cumulative progress, and had - through repeated, uncoordinated extension - generated a state of theoretical confusion in which it was no longer clear which version of TAM was the canonical one. Bagozzi (2007), in the same issue, offered a more direct theoretical critique: TAM treats the user as a rational evaluator of stable belief inputs; it omits emotional, group, social, and cultural processes; it does not distinguish goal pursuit from action pursuit; and it provides no apparatus for the self-regulation that is known to mediate the gap between intention and action. Bagozzi extended the same critique to UTAUT, observing that its expansion to a model with around forty independent variables for predicting intention and at least eight for predicting behaviour had pushed the acceptance literature towards a “stage of chaos” rather than towards explanatory consolidation. Goodhue (2007), Straub and Burton-Jones (2007), and Lee, Kozar, and Larsen (2003) added complementary points concerning the “more use is better” assumption, the risk of common-method artefacts in TAM-style measurement, and the historical drift of the literature towards an explanatory monoculture. Williams, Rana, and Dwivedi (2015) and Dwivedi, Rana, Jeyaraj, Clement, and Williams (2019) consolidated these concerns in a systematic review and a revised UTAUT proposal that explicitly trimmed and reorganised the model in response to the critique tradition. The position these critiques converge on is not that TAM, UTAUT, or TPB cannot be extended; manifestly they can, and have been, many times. The position is that repeated extension does not on its own produce explanation: it can equally well attach further variables to a frame whose original mechanism no longer carries the phenomenon being studied.
A second body of critique is external to the information-systems literature. In health psychology, Sniehotta, Presseau, and Araújo-Soares (2014) argued that TPB had outlived its usefulness as a complete model of behavioural prediction. The decisive point in their argument is that TPB rests on a sufficiency hypothesis - the claim that all theory-external influences on behaviour are fully mediated by attitudes, subjective norms, and perceived behavioural control - and that this hypothesis has been falsified by evidence that beliefs and contextual factors predict behaviour over and above intentions, and that age, socio-economic status, environmental features, and physical and mental health continue to predict behaviour after TPB constructs are controlled. Ajzen (2015), Conner (2015), Schwarzer (2015), and Gollwitzer and Oettingen (2015) responded to defend, extend, or partially endorse the critique. Ajzen’s defence acknowledged that the theory does not preclude the addition of further predictors and was developed in part by adding perceived behavioural control to the prior theory of reasoned action; that defence is itself consistent with the position that the original three constructs do not exhaust the relevant influences on behaviour, even on a charitable reading of TPB. The relevance of this exchange for the present analysis is that LLMs are precisely the kind of theory-external influence that TPB’s mediation structure does not absorb cleanly: an LLM contributes to the formation of attitudes, the option set on which they are formed, and the perception of control, all simultaneously and inside the deliberative episode rather than as a stable upstream input.
A third body of critique is specific to generative artificial intelligence. Mogaji, Viglia, Srivastava, and Dwivedi (2024) argued that TAM, in its classical and extended forms, faces five concrete limitations in the era of generative AI: an individual-centric perspective that does not capture the inherently relational character of conversational AI use; a limited scope that ties acceptance to a single act of adoption rather than to ongoing co-production with the system; a static nature that does not accommodate evolving and bidirectional engagement over time; cultural applicability constraints that are difficult to handle within a small set of moderators; and a reliance on self-reported measures whose validity is strained when users are uncertain about what they are accepting and what the system is doing on their behalf. The authors recommend embedding TAM within domain context, integrating industry-specific factors, and combining it with alternative methodologies, while remaining open to the possibility that other models will be needed.
The editorial framing of a recent special issue in the Journal of University Teaching and Learning Practice - under the title “Are Technology Acceptance Models still fit for purpose?” - places the same question on a journal-level agenda. The diagnosis has continued past Mogaji et al. (2024) across several venues. Davis and Granić (2024), in a SpringerBriefs retrospective on three decades of TAM, propose a NeuroIS-based path as the future development of the model; the gesture is consequential because Davis was the originator of TAM, and the book reads as a programmatic concession from inside the tradition rather than an external critique. Scherer (2024) presses the case for stronger longitudinal causal-inference designs, a critique to which the acceptance literature’s predominant cross-sectional survey methodology is directly exposed. None of these contributions, however, identifies the cognitive-capacity mechanism that operates inside the deliberative episode itself; the response remains methodological - biometric measurement, longitudinal designs, more careful causal identification - rather than mechanistic.
Taken together, this published critique tradition establishes three points that the present analysis can take as common ground rather than as contested claims. First, the inherited acceptance and intention-based theories have known limitations and have been the object of serious internal and external critique for at least two decades. Second, the limitations are particularly visible when the theories are applied to AI and, more sharply, to generative AI. Third, the standard response in the literature - repeated extension, addition of moderators, and patch-style construct accumulation - has itself been flagged as problematic, both in 2007 (Benbasat & Barki, 2007; Bagozzi, 2007) and in 2024 (Mogaji et al., 2024). What this tradition has not yet specified, however, is the locus of the LLM-specific effect, the cognitive-capacity mechanism that operates at that locus, or the gating logic by which that mechanism produces empowered or distorted behavioural outcomes. That is the gap the present model is designed to occupy. The remainder of this section makes the three-layer argument that defines the gap.
A.3.3.2 Temporal Layer: The Substrate Problem
TAM was developed against a substrate of deterministic, single-purpose enterprise software. Its outputs were authored, in the strict sense that any given screen or report could in principle be traced to a known specification produced by an identifiable human team; its behaviour for any given input was repeatable; and its evaluation was meaningfully ex ante, in the sense that a prospective user could form a belief about “whether this system will help me” before committing to use. UTAUT and UTAUT2 synthesised eight prior adoption models and extended their reach into consumer technology, but the validation samples that anchored the synthesis still engaged with information systems whose outputs remained determinate and whose function was meaningfully tied to a specified task. TPB, although developed in social psychology rather than in information systems, shares a closely related implicit assumption about substrate stability: the action-object whose performance is to be predicted is presumed to have known properties, so that beliefs about it can be coherently formed in advance.
Generative, dialogic, probabilistic systems were not part of the empirical world against which any of these theories were calibrated. The substrate has changed in at least four respects that bear directly on the formal apparatus of the inherited theories. First, the same prompt may yield different outputs across runs and across contexts, so that the system’s behaviour is not deterministic in the classical sense; this strains any construct that presumes a stable expectancy about what the system will produce. Second, current LLMs are general-purpose: a single system is used across writing, reasoning, coding, comparing, and decision-supporting tasks, so that “the technology being adopted” is no longer a single well-defined object whose usefulness can be appraised once and for all. Third, the system has no identifiable human author of any particular output, so that the social-norm pathway through which adoption beliefs are partly transmitted in inherited theories - “people I respect think this is a good tool” - points to an object whose authorship is at best diffuse. Fourth, the user’s intent is partly co-constructed during the interaction itself rather than imported into it; this dissolves the temporal separation between “forming an intention to use the tool” and “using the tool”, on which inherited models tacitly rely.
Mogaji et al. (2024) capture part of this in what they call the static nature of TAM: the framework was not built to accommodate evolving, bidirectional engagement with a conversational system. The present analysis extends that diagnosis. The temporal problem is not only that the inherited theories are static while the technology is dynamic; it is that the formal assumptions of the inherited theories - repeatability of outputs, stability of the action object, ex ante evaluability of usefulness, identifiable authorship of system behaviour - were never tested against systems that lack those properties. This is not a refutation of the theories within their original scope. It is a precise statement of why their application to LLM-mediated behaviour is an extrapolation rather than a deployment, and why empirical findings produced by simply re-running TAM- or UTAUT-style instruments against current generative dialogue systems should be read with care.
A.3.3.3 Ontological Layer: Where Cognition Lives
The deeper objection concerns the implicit ontology of the acceptance tradition. In TAM, UTAUT, and TPB, technology is instrumental and cognition is endogenous to the user: the agent forms beliefs about the system, forms an intention to use it, and uses it to execute pre-formed goals. Each of the core constructs in the inherited theories rests on this division of labour. The present sub-section maps the strain construct by construct, because abstract claims about a generalised mismatch are easy to assert and difficult to evaluate; the claim is sharper when the locus of strain in each construct is identified.
In TAM, perceived usefulness is defined as the user’s belief that the system will improve performance on a specifiable task. The construct presumes that the benefit is knowable and stable enough to be evaluated in advance and that performance has a coherent meaning independent of the system. With LLMs, the benefit is co-discovered: the user learns what the system is useful for in the course of using it, and the criterion of “improved performance” often shifts during the interaction as the system reformulates the question, surfaces an option the user had not considered, or alters the user’s sense of what an acceptable answer looks like. Perceived ease of use is defined in terms of the cognitive friction of operation - clicks, learning, time. With LLMs, the central experiential phenomenon is not friction but cognitive co-production: the user types a partial thought and receives a structured continuation that anticipates and reformulates their reasoning. That is not “ease of use” in Davis’s sense; it is a different category of experience that the construct was not built to host.
A second strain on perceived ease of use becomes visible once verification load is brought into view. Heersmink et al. (2024) describe LLM use as combining interactional transparency with internal opacity: the surface of the dialogue is fluent and intelligible to the user, while the conditions under which any given response is correct remain inaccessible. Empirical work on human–AI decision-making (Buçinca, Malaya, & Gajos, 2021; Vasconcelos et al., 2023) demonstrates that users systematically over-rely on fluent AI outputs and that this overreliance is not reduced by attaching explanations to the response. Together these findings carry a measurement consequence for TAM that the present analysis foregrounds. Perceived ease of use as Davis (1989) defined it tracks the operational friction of using a tool - clicks, learning, time - and was calibrated against systems whose outputs were determinate. It does not separate the ease of obtaining a coherent response from the difficulty of verifying its validity. Under LLMs the construct can therefore score high in precisely the conditions Buçinca et al. (2021) and Vasconcelos et al. (2023) identify as those of greatest overreliance risk, and the user has no within-construct signal to distinguish the two cases.
In UTAUT, the same problems apply to performance expectancy and effort expectancy, which are formal generalisations of perceived usefulness and perceived ease of use respectively, with the additional constraint that they presume measurable performance gains on specifiable tasks. Social influence in UTAUT presumes a community of human referents whose use of the system is observed and normatively interpreted by the user. LLM use complicates this on two sides: the system itself participates in dialogue in a way that recruits social cognition (Nass & Moon, 2000; Sundar, 2020; Epley, Waytz, & Cacioppo, 2007), so that the “social” influence operating during a session is partly produced by the system rather than by other people; and the human referent group for LLM use is unstable, partly self-selected, and rapidly evolving, so that the construct’s presumed measurement target is itself moving. Facilitating conditions presume external resources - infrastructure, support - that enable use. With LLMs, the most consequential facilitating conditions are partly internal to the dialogue: prompt construction skill, context-management discipline, and the user’s own willingness to verify outputs. These are not absent from UTAUT in principle, but they are not where the construct expects to find facilitation.
In TPB, attitude towards the behaviour presumes that the behaviour and its attitude object are stable enough to be evaluated coherently; with LLMs, each session may target a different option set generated mid-conversation, so that the “behaviour” being evaluated is itself dialogue-dependent. Subjective norm faces the same instability of referent group noted for UTAUT’s social influence. Perceived behavioural control is the construct on which the strain is sharpest. PBC presumes that the locus of reasoning is internal to the actor: the agent forms an intention based on perceived control over an action whose execution depends on her own resources and capabilities. With LLM-mediated deliberation, the locus of reasoning is partly externalised into a dialogic partner whose contributions shape the goal, the option set, and the criteria of evaluation. The agent is not asking only “can I do this?”, but also, implicitly, “can the joint apparatus of myself and this system do this?”. This is not a niche difficulty: it is precisely the kind of structural problem that Sniehotta, Presseau, and Araújo-Soares (2014) had in mind when they argued that TPB’s sufficiency hypothesis has been falsified. The relevant external influence on behaviour - the LLM - is not absorbed by attitude, subjective norm, or PBC; it acts inside deliberation rather than as a stable upstream input.
The collective ontological problem is therefore not located in any one construct but in the assumption shared across the three theories. TAM, UTAUT, and TPB all treat technology as instrumental and cognition as endogenous: the agent thinks, then the tool acts. LLMs distribute cognition between user and system in a way that none of the three frameworks specifies. This is not a metaphysical objection. It is a measurement objection. When the locus of reasoning is partly outside the user, constructs that measure user-internal beliefs about a stable external object will under-represent the very mechanism that is theoretically active. This is consistent with Bagozzi’s (2007) older argument that TAM lacks an apparatus for the self-regulation that mediates intention and action; the present case sharpens the older argument by specifying that, with LLMs, what is missing is not only self-regulation but co-regulation between user and system within the deliberative episode itself. It is also consistent with Sundar’s (2020) more recent claim that the rise of machine agency requires a new theoretical apparatus for human–AI interaction, since traditional acceptance frameworks treat the machine as a passive object whose properties the user evaluates.
A.3.3.4 Phenomenological Layer: Eight LLM-Mediated Mechanisms Without Integrated Theoretical Homes
The temporal and ontological layers concern, respectively, the substrate against which the inherited theories were calibrated and the assumptions about cognition that those theories carry. The phenomenological layer concerns the empirical phenomena that LLM-mediated activity actually produces and that any model of LLM-assisted evaluative decision-making must be able to host. The claim is not that every underlying psychological process first appeared with LLMs. Several mechanisms have pre-LLM ancestors in automation trust, anthropomorphism, metacognition, cognitive offloading, and decision-support research. The narrower claim is that dialogic, generative, RLHF-tuned, context-retaining systems combine and activate these mechanisms inside the evaluative decision episode in ways that TAM, UTAUT, and TPB do not integrate as a coherent theoretical layer.
Hallucination - the confident generation of plausible but unfounded content - is not error in the deterministic software sense. The system is functioning according to its generative architecture when it produces fluent but unsupported text: it is generating token sequences that are statistically coherent, not necessarily grounded in a verified external state of the world (Ji et al., 2023; Huang et al., 2025; Dang Anh-Hoang, Vu Tran, & Le-Minh Nguyen, 2025). Recent empirical work in high-stakes domains such as law further shows why this is not a trivial reliability problem: hallucinated outputs can be expressed with the same fluency and authority as correct outputs, leaving users without an obvious run-time signal that verification is required (Dahl, Magesh, Suzgun, & Ho, 2024).
Sycophancy - the tendency of dialogue-tuned systems to converge on, flatter, or reinforce the user’s expressed view - is similarly specific to the alignment and preference-optimisation regime of contemporary conversational AI. Sharma et al. (2024) show that RLHF-trained assistants can prefer agreement with the user over truthfulness; Fanous et al. (2025) and Kaur (2025) extend this evidence to systematic sycophancy evaluation and argument-driven stance shifts in multi-turn settings. In inherited adoption models, these phenomena cannot be treated merely as low usefulness or low reliability, because they may increase perceived usefulness precisely by producing fluent, agreeable, and confidence-enhancing output.
Anthropomorphic dialogic interaction has clear pre-LLM ancestry in social-response and anthropomorphism research (Nass & Moon, 2000; Epley, Waytz, & Cacioppo, 2007), but contemporary LLMs intensify the mechanism through sustained, turn-taking natural-language interaction. The user is not only responding to a screen or a menu; the user is interacting with a system that apologises, reasons, explains, adapts tone, and appears to understand the evolving question. Recent work on anthropomorphic cues and mental-state attribution in LLMs shows that such cues can affect trust, perceived competence, and advice acceptance (Cohn et al., 2024; Colombatto et al., 2025).
Stochastic non-determinism - the property that the same or near-identical prompt may yield materially different outputs across runs - further strains the repeatability assumptions implicit in classical acceptance constructs. In software-engineering and applied evaluation settings, repeated LLM queries have been shown to produce non-identical or inconsistent outputs, and even low-temperature settings do not always eliminate non-determinism (Ouyang, Zhang, Harman, & Wang, 2025; Funk et al., 2024). A user may therefore be evaluating not a stable artefact with fixed performance properties, but a probabilistic conversational process whose realised behaviour changes across sessions and turns.
Epistemic opacity in the presence of interactional transparency is another LLM-mediated configuration without a natural home in inherited adoption-intention models. Heersmink et al. (2024) describe LLMs as phenomenologically and interactionally transparent enough to support fluent use, while remaining opaque in their internal generative processes, training-data provenance, and failure modes. Hicks, Humphries, and Slater (2024) sharpen the epistemic problem by arguing that LLM outputs are not truth-governed in the way users may assume, while Cassinadri (2024) places ChatGPT-like systems within a virtue-epistemological problem of cognitive artefact use. This combination matters for TAM because perceived ease of use can be inflated by surface fluency while actual epistemic transparency remains low.
Asymmetric trust calibration then follows: users may update trust, confidence, and reliance faster than the system’s true reliability profile can be assessed. Automation-trust research anticipated parts of this problem (Lee & See, 2004; Parasuraman & Manzey, 2010), but recent AI work makes the calibration issue more acute by showing that confidence, explanation, and metacognitive sensitivity are central to whether users know when to trust AI advice (Lee, Pruitt, Zhou, Du, & Odegaard, 2025).
Context-window adaptation means that the system’s current output is shaped by the sequence of prior turns. The option set shown to the user in turn six is not independent of the assumptions, framings, omissions, or biases introduced in turns one through five. Herlihy, Neville, Schnabel, and Swaminathan (2024) show that LLM-based chatbots can be affected by miscalibrated conversational priors, especially under under-specified requests, while Laban et al. (2026) show that LLMs can become lost in multi-turn conversation by making early assumptions and then over-relying on them rather than recovering cleanly.
Co-construction of the option set is the parallel user-side phenomenon: the user’s alternatives, criteria, and sometimes attitudes are produced inside the dialogue rather than brought to it as fixed upstream inputs. Jakesch et al. (2023) show that co-writing with opinionated language models can shift users’ expressed views and later attitudes, while recent work in AI-mediated persuasion shows that conversational AI can shift political preferences or attitudes under experimental conditions (Argyle et al., 2025; Lin et al., 2025; Hackenburg et al., 2025). In TPB terms, this means that “attitude toward the behaviour” is not always a stable antecedent to the deliberative episode; under LLM mediation, it may be partly generated within that episode.
The theoretical incursion is therefore not the claim that these eight mechanisms have no intellectual ancestry. The incursion is that, under LLM-mediated evaluative interaction, these mechanisms cross the measurement boundaries of the inherited constructs. Hallucination is not merely lower perceived usefulness, because fluency can preserve or increase perceived usefulness while factual grounding collapses. Sycophancy is not merely social influence, because the source of influence is the technological interlocutor itself rather than human peers or social norms.
Anthropomorphic dialogue is not merely hedonic motivation, because it turns the system into a pseudo-social participant in judgement. Stochastic non-determinism is not merely unreliability, because the user is evaluating a probabilistic process rather than a fixed system state. Epistemic opacity with interactional transparency is not merely low ease of use; it is often high ease of use masking low inspectability. Trust calibration is not merely trust as a static antecedent to intention, because reliance is adjusted dynamically inside the interaction. Context-window adaptation is not merely facilitating conditions, because the system’s apparent capability is recursively shaped by prior turns. Co-construction of the option set is not merely attitude or perceived behavioural control, because the system helps produce the alternatives and capability judgements that those constructs treat as already formed.
The two logics therefore converge. The fresh LLM-era literature gives empirical support for the eight mechanisms as real phenomena of generative, dialogic, context-retaining systems. The intellectual-incursion logic explains why their theoretical significance is not exhausted by adding them one by one as moderators, controls, or auxiliary variables. Some mechanisms are LLM-native; others are older mechanisms reconfigured and intensified by LLM affordances. What makes the phenomenological layer necessary is the combined fact that these mechanisms are activated inside the evaluative decision episode and that TAM, UTAUT, and TPB have no integrated construct architecture for modelling how they jointly reshape perceived capacity, calibration, and choice. The phenomenological layer therefore does not claim eight historically originless phenomena; it identifies an LLM-mediated configuration that inherited adoption and intention frameworks can only patch fragment by fragment, but cannot host as an integrated account of evaluative judgement.
A.3.3.5 Construct Proliferation as Confirmatory Evidence
The proliferation pattern is itself diagnostic, and on this point the present argument converges with - and draws empirical support from - the published critique tradition. Bagozzi (2007) used the phrase “stage of chaos” to describe the unprincipled growth of UTAUT-style models. Benbasat and Barki (2007) used the phrase “theoretical chaos and confusion” to describe the same problem in TAM-style extensions. Mogaji et al. (2024) recommended openness to alternative models partly because patch-style extension has not produced integration. Outside the information-systems literature, Shaffer, DeGeest, and Li (2016) formalised the same diagnosis at the methodological level, defining construct proliferation as “the accumulation of ostensibly different but potentially identical constructs representing organizational phenomena” and arguing that the field-level remedy is rigorous discriminant-validity testing of conceptually adjacent constructs rather than further construct addition. The proliferation observed in the AI-behaviour literature has the same signature: the new constructs are not in obvious conflict with TAM, UTAUT, or TPB; they are simply not derivable from them. Each is added as a moderator or as a parallel construct because the parent framework offers no principled location for it.
The pattern continues into 2025. Salam et al. (2025) advance a Revised Artificial Intelligence Device Use Acceptance (RAIDUA) framework that attaches privacy concern, anthropomorphism, and emotional appraisal as new constructs onto a UTAUT-style spine. Xue, Ghazali, and Mahat (2025) propose an Integrated AI-in-Education Acceptance Framework (IAEAF) that adds AI literacy, institutional context, intervention design, and sustained-integration outcomes; the authors themselves motivate their framework as a response to “conceptual redundancy and contextual misalignment” across the proliferating UTAUT2 extensions in their domain - a 2025 echo, in different vocabulary, of Bagozzi’s (2007) “stage of chaos”. Read through Kuhn (1962/2012) and Lakatos (1970), the pattern is familiar. When researchers continue to operate within an existing framework while patching it with successive auxiliary constructs, the framework’s predictions are not refuted; instead, they are increasingly accompanied by qualifications, and ad hoc additions multiply because the underlying frame cannot host the phenomena cleanly. The present analysis does not require the stronger Kuhnian claim that the AI-behaviour literature is undergoing a paradigm shift. The measured claim is sufficient: the inherited frameworks are showing signs of strain under the weight of LLM-specific phenomena; the rate of construct addition is outpacing the rate of theoretical integration; and that pattern has been independently noted by senior commentators in both the 2007 Quo Vadis TAM exchange and the more recent generative-AI-specific critique. A new model in this space, on this view, must do more than predict an additional outcome variable. It must provide a coherent home for the constructs the field has been generating in an ad hoc way. Section A6 returns to this point in the form of a rival-explanation audit; the present sub-section establishes only that proliferation is part of the evidence base for the claim that the inherited frameworks are incomplete, not a separate methodological worry.
A.3.3.6 What the Present Argument Adds Beyond the Published Critiques
§A3.3.1 surveys the published critique tradition; this sub-section identifies what each critic did not specify, so that the position the present model occupies relative to that tradition can be stated precisely. Bagozzi (2007) called for a paradigm shift and proposed a general goal-pursuit decision-making core; that proposal predates the deployment of LLMs and does not specify the cognitive-capacity mechanism that arises when a conversational generative system is present at the moment of deliberation. Benbasat and Barki (2007) called for renewed attention to the system characteristics that actually make systems useful; they did not propose a moderation gate distinguishing empowered use from distorted use. Sniehotta, Presseau, and Araújo-Soares (2014) argued that TPB’s sufficiency hypothesis has been falsified; they did not specify what the relevant unmediated external influences are in AI-mediated deliberation, nor where in the deliberative chain those influences enter. Mogaji, Viglia, Srivastava, and Dwivedi (2024) catalogued five limitations of TAM in the era of generative AI and recommended industry-specific embedding and openness to alternative models; they did not propose a specific construct, mechanism, or extension architecture. Sundar (2020) argued that machine agency requires a new theoretical apparatus and proposed a dual-process framework drawn from the Theory of Interactive Media Effects; that proposal addresses media-effect symbolism and affordance perception more than the formation of decision criteria during high-stakes individual deliberation. Heersmink, de Rooij, Clavel Vazquez, and Colombo (2024) characterised the distinctive epistemological situation of LLM use - interactional transparency with internal opacity - but at the level of philosophy of technology rather than at the level of an empirical model. §A3.3.7 below states the corresponding architectural response.
A.3.3.7 Implication for the Present Model
In one sentence, TAM, UTAUT, and TPB model technology adoption and intention formation; LLM-mediated behaviour at the moment of complex decision additionally requires a model of cognitive partnership during evaluation. The Augmented Behavioural Capacity Model does not seek to replace the inherited frameworks. It accepts that they explain access, uptake, and intention well, treats them as the upstream layer of any complete account, and positions itself as a domain-bounded extension that isolates the LLM-specific cognitive-capacity mechanism that none of the inherited frameworks specifies - and that none of the published critiques of those frameworks has yet specified either.
Figure 3 (reported in the manuscript body) visualises the three-layer argument geometrically: the inherited theoretical apparatus is stratified inside the frame; the phenomenological mechanisms unique to LLM-mediated activity sit outside the frame, indicating their lack of a principled home in the existing constructs. Section A5 specifies the architecture by which the proposed extension is organised; Section A6 returns to the construct-proliferation pattern in the form of an audit of rival families that the new construct must withstand.
A symmetric point applies at the dependent-variable end of these theories. TAM and UTAUT terminate in intention to use, actual use, and continued use. TPB terminates in intention and behaviour. None of the inherited frameworks reaches decision quality - the structural complexity, risk posture, and calibration of the choice the user actually makes. In retail-investor decision-making, decision quality in this narrow sense is the outcome that matters: whether LLM assistance changes which strategies the investor is willing to consider, on what evidentiary basis, and with what calibration. The gap the present model addresses is therefore bilateral. On the independent-variable side, no inherited framework specifies the tool-contingent perceived-capacity mechanism that operates inside the deliberative phase under LLM augmentation. On the dependent-variable side, no inherited framework distinguishes adoption-or-intention outcomes from decision-quality outcomes. The Augmented Behavioural Capacity Model is positioned to address both: it inserts a cognitive-capacity layer (Capacity, with PCA as its measurable anchor) and a moderation gate (Calibration) between the inherited adoption-and-intention layer upstream and a proximal decision-quality layer (Choice, bounded to tactic-selection and intended structural complexity) downstream. Read in this light, the model is not another loose extension of TAM, UTAUT, or TPB; it is an attempt to specify where the extension belongs.
Table A5 consolidates the argument across the inherited frameworks named in the heading of this section. Each row identifies what the theory or family explains well within its original scope, the core limitation that emerges under LLM-mediated decision-making, and the corresponding response the ABC Model adopts. The table serves as a one-page orientation map across the theoretical landscape, not as the audit itself. It is a comparative consolidation table, not a meta-analysis and not a systematic review: its function is to map, in one view, what each inherited framework explains within its original scope, where it strains under LLMs, and the ABC response.
Table A5.
Consolidation of the §A3.3 argument across the three inherited frameworks.
Table A5.
Consolidation of the §A3.3 argument across the three inherited frameworks.
| Framework |
Temporal mismatch |
Ontological mismatch |
Phenomenological mismatch |
ABC implication |
| TAM |
Calibrated to stable IS tools |
User evaluates external tool |
PU/PEOU cannot house dialogic scaffolding, opacity, hallucination |
Retain PU/PEOU upstream; add PCA |
| UTAUT/UTAUT2 |
Calibrated to determinate task/use systems |
Performance/effort expectancy assume stable task-tool relation |
Does not capture internal decision-process transformation or calibration |
Retain engagement layer; do not equate PE with PCA |
| TPB |
Assumes stable behaviour/action object |
PBC locates control mainly in actor |
Cannot represent externally scaffolded perceived capacity or co-produced option set |
Refine PBC into tool-contingent capacity plus calibration |
The three-layer analysis establishes the gap that motivates ABC. It does not yet adjudicate every neighbouring construct family that could compete with PCA. That task belongs to the boundary and rival-explanation audit in Section A6. The immediate implication of Section A3 is narrower: inherited acceptance and intention frameworks explain access, uptake, usefulness, effort, social influence, attitude, norm, perceived control, intention, and use, but they do not specify the LLM-specific perceived-capacity mechanism operating inside the evaluative decision episode, nor do they distinguish adoption/intention outcomes from proximal decision-quality outcomes. Section A5 translates this residual zone into the ABC architecture; Section A6 then tests whether adjacent constructs and rival families can absorb it.