Preprint
Article

This version is not peer-reviewed.

Mental AI: Post-Deployment Criterion Endogeneity in the Causal Loop of Human Mental Life

Submitted:

24 July 2026

Posted:

29 July 2026

You are already at the latest version

Abstract
Artificial systems are validated against self-reports, expressions, behaviours, clinician judgements, and institutional records. Once an inference about mental life guides a response or decision, those criteria may become descendants of the inference. Agreement can then increase because the model predicted the target, changed the target, altered self-interpretation or expression, selected what became observable, or reshaped the criterion-generation process. I call this evaluation problem post-deployment criterion endogeneity.Mental AI denotes systems in which this problem can arise: systems that use an inference about mental life, operationally incorporate it into a response or decision, and create an inference-indexed, construct-relevant consequence path. This category connects emotion recognition, personality scoring, machine theory of mind, conversational agents, digital mental health, educational analytics, and workplace assessment while retaining important domain and mechanism differences.The framework separates four outcome layers: latent mental state; self-interpretation and report; expressive–behavioural evidence; and social or institutional allocation. It then distinguishes external criteria, whose measurement meaning does not constitutively depend on uptake of a classification, from reflexive criteria partly produced through self-ascription or norm-governed expression. This distinction converts a broad concern about AI influence into an identifiable evaluation problem.Drawing on affective computing, performative prediction, response shift, measurement reactivity, looping effects, AI-mediated selfhood, and bidirectional belief amplification, the article develops ten falsifiable propositions and four decisive experiments. It predicts when ambiguity, persistence, operational authority, and contestability will alter recursive effects. Affective Sovereignty follows as a procedural design implication: preserve refusal, delay, override, and contestation while keeping external evidence available to challenge self-interpretation.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  ;  
Subject: 
Social Sciences  -   Psychology

1. When a Mental Inference Changes Its Own Evidence

An AI system infers anxiety from a person’s language. The inference shapes the system’s follow-up questions, the explanations it offers, and the actions it recommends. At a later assessment, the person’s report aligns more closely with the original prediction. That agreement is compatible with several causal histories. The model may have detected an existing state; the intervention may have changed the state; the label may have reorganised self-interpretation; the interaction may have altered expression; or the deployment may have changed which evidence entered the record. One correlation can therefore contain several scientifically different processes.
Before deployment, validation asks whether an inference corresponds to a criterion generated independently of that inference. After deployment, the inference can become part of the data-generating process. Its output may guide a conversation, intervention, allocation, or institutional decision, and the resulting consequences may later be reused as evidence of accuracy. This creates post-deployment criterion endogeneity: the operational criterion used to evaluate a mental inference is partly produced under conditions altered by deploying that inference. Increased agreement then identifies neither predictive knowledge nor causal influence without an exposure-aware design.
The problem extends across systems ordinarily studied apart. An educational platform infers frustration and changes task difficulty. A conversational agent infers loneliness and adopts increasingly intimate language. A workplace system estimates motivation and changes monitoring or assignment. A clinical model infers risk and changes the order or intensity of care. The common object is not a model architecture or application sector. It is a causal relation in which a computational representation of mental life enters a policy directed back at the represented person or their environment.
Research already establishes the links that make this relation possible. Artificial systems can generate psychologically meaningful representations from language, behaviour, and interaction [1,2,3]. Affective computing supplies mature methods for detecting and modelling emotion [4,5,6,7]. Conversational systems can influence judgement and persuasion [8,9]; human–AI feedback can alter later perceptual, emotional, and social judgements [10]; and persistent agents can become sources of support, attachment, and belief reinforcement [11,12,13,14,15]. The unresolved task is to connect these capacities to the psychology and measurement of what follows after an inference becomes operative.
This article calls that causal class Mental AI:
Definition.Mental AI comprises artificial systems whose inferences about human mental life are operationally taken up in responses or decisions that create an inference-indexed, construct-relevant path capable of changing the inferred attribute, its self-interpretation or expression, or the person’s subsequent social or institutional environment in ways that shape later evidence, action, opportunity, or treatment.
Direction of attribution fixes the scope. Mental AI concerns artificial claims about human minds and the consequences of acting on them; questions about whether artificial systems themselves possess mentality form a separate inquiry [16]. Category membership turns on causal architecture rather than accuracy, marketing language, or observed harm. A false or uncertain inference can acquire the same operative role as a correct one. Recursive consequence is specified ex ante as a pathway; realised change is an empirical outcome. Criterion endogeneity appears at the further stage where downstream data are used to evaluate the earlier inference.
The argument proceeds in six moves. It locates the causal problem across neighbouring traditions; defines three constitutive conditions for Mental AI; separates latent state, self-interpretation, expressive–behavioural evidence, and allocation; formalises criterion endogeneity and its identification requirements; derives ten discriminating propositions; and traces the epistemic and procedural consequences of allowing mental inferences to acquire authority. This order places governance after the data-generating and measurement problem from which it follows.

2. Converging Traditions and the Unresolved Synthesis

2.1. From Representation to Operative Consequence

AI classifications already operate along several independent dimensions. Generative AI names an output capability; agentic AI names a form of goal-directed autonomy; physical AI names embodied sensing and action. Mental AI names a different relation: the target is human mental life, and a representation of that target enters a policy capable of changing subsequent evidence or consequence. A therapeutic robot can therefore be generative, agentic, physical, and Mental AI simultaneously. The categories answer different questions.
Affective computing supplies the principal technical lineage for recognising, modelling, communicating, and responding to emotion [4,5]. Machine theory of mind extends representational targets to beliefs, intentions, and perspectives, with performance that varies across tasks and paradigms [1,17]. Together these fields establish the inference stage. The present framework begins at the next transition: a mental representation becomes a response rule, intervention, or allocation directed toward the represented person.
Social and relational AI establish the importance of interaction, social presence, memory, and sustained relationships. Socioaffective alignment locates evaluation in the social and psychological systems co-produced through those relationships [18]. Psychological competence similarly evaluates how human-facing systems handle cognition, emotion, trust, uncertainty, and decision-making in context [19]. These programmes specify qualities of interaction. Mental AI supplies a causal class that also contains low-relational, high-authority systems such as workplace or clinical scores. Relational persistence and institutional authority can therefore be analysed as separate dimensions rather than collapsed into a single notion of social influence.
Research on AI-mediated selfhood, reflective agency, and epistemic authority supplies the interpretive layer. The algorithmic-self account traces effects on identity, introspection, and emotional expression; recursive-self analysis examines machine representations and de-looping; reflective-agency work identifies design conditions for autonomous meaning-making [20,21,22]. Studies of belief revision and epistemic infrastructure show how computational outputs can acquire authority within a person’s reasoning [23,24]. These insights become experimentally sharper when state, interpretation, expression, and allocation are treated as distinct outcomes.
The bidirectional belief amplification framework provides the closest specified mechanism for a relational mode of Mental AI: human cognitive–emotional biases and chatbot tendencies can mutually reinforce maladaptive beliefs during extended interaction [15]. The Resonant Amplification Framework likewise models how persistent, adaptive dialogue can stabilise correction-resistant interpretations through attachment, co-creation, and linguistic reinforcement [25]. Mental AI places such mechanisms within a wider causal architecture that also includes non-belief targets and non-conversational institutions. Its evaluation question begins when the changed outcome returns as evidence about the earlier inference.
Application fields and normative traditions complete the map. Digital mental-health AI contributes evidence on conversational therapy, triage, symptom monitoring, coaching, and wellness support. Educational, workplace, insurance, marketing, companion, and robotic systems show that the same causal relation crosses sectoral boundaries. Neuroethics and neurorights contribute analyses of mental privacy, identity, agency, and integrity [26,27,28]. Mental inference from text, voice, facial movement, behaviour, and interaction history extends that concern from signal access to the authority acquired by an inference after it is made.

2.2. From Performative Targets to Reflexive Criteria

Performative prediction provides the formal spine: deployment can change the distribution against which a model is subsequently evaluated [29,30]. Causal work on decision support shows how predictions alter observed targets, bias monitoring and retraining, and require comparison with a substantively defined baseline policy [31,32]. Accurate models can participate in harmful self-fulfilling prophecies while retaining strong post-deployment performance [33]. Prediction–allocation theory identifies the underlying epistemic–pragmatic entanglement: once predictions help cause outcomes, accuracy and consequence can no longer be read from the same association without further identification [34].
Psychology and the social sciences supply parallel mechanisms. Self-fulfilling prophecy changes action through expectation [35,36]; rankings and measures reorganise behaviour [37]; interventions can change the meaning of a self-report scale through response shift [38]; measurement itself can be reactive [39,40]; and human classifications can enter the self-understanding and practices of the people classified [41,42]. Proxy-label bias and selective observability further show that the recorded criterion may already be an imperfect or policy-dependent representation of a latent target [43,44].
Mental AI develops this insight in a domain where the target is internally differentiated and partly reflexive. A mental-state inference can affect at least four non-equivalent outcomes: the latent state itself, the person’s interpretation of that state, the person’s outward expression, and the social or institutional allocation that follows. These effects need not move together. A person may alter wording without changing felt affect, adopt a new interpretation without changing physiology, or face changed treatment despite rejecting the label. This differentiation is necessary for both causal inference and ethics.
Table 1. Traditions contributing to the Mental AI synthesis. The rows identify complementary organising questions rather than disciplinary boundaries.
Table 1. Traditions contributing to the Mental AI synthesis. The rows identify complementary organising questions rather than disciplinary boundaries.
Tradition Established contribution Role in the synthesis
Affective computing Computational inference and response concerning affect Supplies the core technical lineage for mental-state inference.
Machine theory of mind Artificial representation of beliefs, intentions, and perspectives Extends the target beyond affect and locates the transition from representation to policy.
Social and relational AI Interaction, social presence, memory, and sustained relationships Supplies mechanisms of persistence while remaining distinct from formal authority.
Psychological competence Context-sensitive support for cognition, emotion, trust, and decision-making Supplies interaction-quality dimensions within the broader causal class.
AI-mediated selfhood and reflective agency Computational mediation of identity, introspection, belief revision, and self-discovery Supplies the interpretive and agency mechanisms linking output to later report.
Digital mental-health AI Assessment and intervention in mental health and well-being Supplies a clinically important application domain and outcome evidence.
Performative prediction Deployment-induced change in targets and data distributions Supplies the formal basis for exposure-aware evaluation.
Neuroethics and neurorights Mental privacy, agency, identity, and integrity Supplies normative analysis of access, authority, and contestability.
A second difference concerns evaluation. In ordinary performative prediction, a changed outcome distribution can be treated as a deployment effect. Prediction–allocation scholarship has already identified the resulting entanglement of epistemic and pragmatic considerations [34]. Mental AI develops the reflexive mental case: the criterion used to assess the original inference may itself be a self-report, behaviour, or institutional record shaped by the system, and producing the criterion can also be an act of self-interpretation or norm-governed expression. Agreement can therefore increase because the model was accurate, because the target changed, or because the evidential criterion moved toward the model’s categories. This article calls the latter evaluation problem post-deployment criterion endogeneity.

2.3. Distinguishing Criterion and Identification Failures

Several established problems can produce misleading validation, but they are not interchangeable. Table 2 separates failures that occur before deployment, through policy-dependent observation, and through intervention-induced change. The distinction is essential because a model can suffer more than one failure at once.

3. Mental AI as a Causal Class

3.1. Three Constitutive Conditions

A system belongs to Mental AI when three conditions are jointly satisfied.

3.1.1. Condition 1: Mental-State Inference

The system generates, adopts, or uses a representation that goes beyond directly observed behaviour to make a claim about a person or group. Relevant targets include affect, mood, intention, belief, desire, motivation, personality, attention, fatigue, stress, vulnerability, mental-health risk, relational meaning, identity, or self-understanding. The claim may be probabilistic, categorical, dimensional, narrative, or embedded in a latent representation. It need not be shown to the user.
Mental inference is not equivalent to collecting self-declared preferences. A music system that stores a user’s explicit genre selection need not infer mental life. If the system infers sadness, loneliness, or personality from listening patterns and uses that inference to alter interaction, the first condition can be satisfied.
Category assignment is construct-functional rather than lexical. A score is not mental merely because it bears a psychological name, and it does not cease to be mental because it is labelled “engagement,” “risk,” or another behavioural or commercial term. Classification depends on what the representation is trained, validated, communicated, or operationally interpreted to stand for. At the episode level, a useful counterfactual test is: holding the observed input fixed, would changing or withholding the mental interpretation change the response or decision? If not, the mental representation has not been operationally taken up in that episode.

3.1.2. Condition 2: Operational Uptake

The inference must affect a response, recommendation, ranking, intervention, restriction, allocation, or decision. An offline research classifier that estimates emotion without affecting anyone is a precursor technology but not a deployed Mental AI episode. Operational uptake occurs when the inference enters a policy, whether the policy is a model’s response strategy, a human decision aid, or an institutional rule.

3.1.3. Condition 3: An Inference-Indexed Recursive Consequence

The response or decision must create or activate a non-trivial, construct-relevant pathway beyond the mere computation or storage of the inference. At least one of the following must be capable of changing: (i) the inferred mental attribute, (ii) the person’s interpretation or report of that attribute, (iii) expression or behaviour later treated as evidence about it, or (iv) the person’s social or institutional opportunity or treatment. The pathway must be indexed to the mental inference. Generic influence is insufficient.
Actual change is not required for category membership. Requiring observed transformation would classify a system only after the consequence occurred. The pathway is specified ex ante as an inference–policy–mediator or operative allocation–outcome chain and tied to the same mental construct or to a theoretically specified downstream consequence. A tutoring system that infers frustration and changes task difficulty can qualify because the intervention can alter frustration, its expression, engagement, and later performance. A hiring system that infers motivation can qualify when that inference changes shortlisting, employment opportunity, or subsequent treatment; an unused internal score does not qualify merely because it exists. A translation system does not qualify merely because translation may incidentally affect mood.
The three conditions are analytically distinct, but they are not assumed to be logically or statistically independent in every deployment. Condition 2 identifies whether the mental inference enters a system or institutional policy. Condition 3 identifies whether that policy has a construct-relevant consequence for the represented person or their environment. In institutional deployments, an operative decision can both instantiate uptake and initiate the consequence pathway. This partial overlap does not collapse the category, because a mental inference may enter an internal policy without acquiring a consequential path back to the represented person. Post-deployment criterion endogeneity is a further condition that arises only when an altered outcome is subsequently reused as evidence for evaluating the inference.

3.2. Non-Cases and Mixed Cases

A generic translation system, a calculator, and a non-personalised meditation recording are not Mental AI merely because they can affect a user’s feelings. A model used to analyse de-identified affective data without any path back to the represented individuals is not a deployed Mental AI system. A system that claims to simulate its own emotion is not Mental AI for that reason alone. The category is about artificial inference directed at human mental life.
Mixed cases are expected. A tutoring system can use an explicit student request to slow down without inferring frustration. The same system becomes Mental AI when it infers frustration, changes task difficulty, and thereby affects engagement or future performance. A clinical risk model may be a Mental AI component even if the final decision is made by a clinician, because operational uptake can occur through human reliance. Research on automation use, misuse, disuse, and overreliance shows why the human decision-maker does not break the causal chain [45,46,47,48,49].

3.3. An Orthogonal Category

Mental AI is not proposed as the next chronological generation after generative or physical AI. It is orthogonal to output capability, autonomy, and embodiment. Figure 1 shows the distinction. The same system can occupy multiple categories. This avoids a common taxonomic error in public discourse, where labels that describe different dimensions are treated as competitors.

4. Mental States, Interpretations, Expressions, and Criteria

4.1. Four Outcome Layers in One Causal Architecture

A central methodological error would be to treat “the person’s mental state” as a single observed variable. Mental life is not directly read from one signal, and self-report, expression, behaviour, physiology, and social treatment are neither identical nor interchangeable. Emotion research has long debated the relation between experience, conceptualisation, expression, and context [50,51,52,53]. Facial movement alone does not provide a universal transparent readout of emotion [53,54]. Introspection and verbal report are informative but fallible, shaped by access, concepts, task demands, and metacognition [55,56,57,58]. Interoceptive accuracy, awareness, and confidence are also distinct [59,60,61].
The framework therefore distinguishes:
  • m t : latent mental state at time t, including affective, cognitive, motivational, or vulnerability-related properties;
  • r t : self-interpretation and report, including labels, narratives, confidence, and explicit self-assessment;
  • x t : observable expression, behaviour, language, physiology, and interaction traces;
  • e t : the social and institutional environment, including opportunities, constraints, surveillance, treatment, and other people’s responses;
  • c t : the contemporaneous decision context available to the system, including the task, user-provided context, deployment rules, and institutional constraints;
  • m ^ t : the system’s inferred representation of mental life;
  • a t : a direct system response to the user;
  • d t : an institutional or third-party decision informed by the inference;
  • u t : exogenous influences not produced by the Mental AI episode.
A minimal dynamic representation is:
m ^ t = f ( x t , r t , e t ) ,
( a t , d t ) = π ( m ^ t , c t ) ,
m t + 1 = g m ( m t , a t , d t , e t , u t ) ,
r t + 1 = g r ( m t + 1 , a t , d t , e t , u t ) ,
e t + 1 = g e ( e t , d t , a t ) ,
x t + 1 = h ( m t + 1 , r t + 1 , e t + 1 , u t + 1 ) .
The absence of m ^ t as a direct cause of r t + 1 is deliberate. A hidden inference cannot alter a person’s report unless it is operationalised through a response, disclosure, decision, or changed environment represented by a t , d t , or e t . The arrows in Equations – identify candidate pathways, not assumed effects. The propositions below test whether those pathways exist, their direction, and the conditions under which they become consequential.
These equations provide a causal bookkeeping schema. They keep distinct variables and outcome layers from being collapsed; identification is supplied at the application level through a measurement model, error structure, functional form, confounder set, observability process, temporal design, and intervention strategy. The causal effect of an inference-guided policy on an outcome Y is conceptually counterfactual:
Δ t + 1 MAI ( Y ) = Y t + 1 π t MAI Y t + 1 π t 0 .
where π t MAI is a policy that uses the mental inference to generate ( a t , d t ) and π t 0 is a substantively specified comparison policy. There is no single domain-general control: no-inference, uncertainty-preserving, human-controlled, and alternative-inference policies answer different causal questions and must not be treated as interchangeable. The estimand is policy-specific rather than an effect of the numerical prediction in isolation. Different outcomes Y must be analysed separately.

4.2. Four Effects

State effect.

The system changes the latent state, such as anxiety, trust, motivation, belief, or loneliness. A supportive intervention may reduce distress; a reinforcing interaction may intensify conviction.

Interpretive effect.

The system changes how the person understands, labels, or narrates an experience. The underlying state may or may not change. An AI label can become a salient hypothesis, a vocabulary, or a default explanation.

Expressive–behavioural effect.

The system changes what the person says, displays, or does in ways that may later be treated as evidence about the inferred construct. This may reflect genuine state change, interpretive change, strategic adaptation, demand characteristics, or avoidance of institutional consequences.

Allocative effect.

The inference changes how institutions or other people treat the person, including task assignment, access, ranking, monitoring, care, employment, or insurance. Allocative effects can occur even when the person rejects the inference and experiences no internal change.
Figure 2. The recursive Mental AI episode. A machine inference becomes causally relevant through direct interaction and institutional allocation. Later reports or behaviours can be both outcomes of the earlier intervention and criteria used to evaluate the model.
Figure 2. The recursive Mental AI episode. A machine inference becomes causally relevant through direct interaction and institutional allocation. Later reports or behaviours can be both outcomes of the earlier intervention and criteria used to evaluate the model.
Preprints 224749 g002
These effects can interact. An allocative decision can subsequently alter a state. An interpretive change can alter expression. A changed expression can be treated as evidence that the original model was accurate. The distinction is therefore analytical, not a claim of independence.

4.3. Post-Deployment Criterion Endogeneity

Suppose a model predicts that a user is anxious. The system repeatedly frames ambiguous experiences as anxiety, recommends avoidance, and asks follow-up questions using the same category. At a later time, the user’s self-report is more consistent with the original prediction. Three explanations are possible: the initial prediction was accurate; the user’s state changed; or the user’s interpretation and reporting shifted toward the system’s frame. These explanations have different scientific and ethical meanings.
The evaluation problem is general. Let T t + 1 denote the target quantity that validation intends to recover and let C t + 1 obs denote the operational criterion that is actually observed. For a latent mental target,
T t + 1 = q * ( m t + 1 ) ,
C t + 1 obs = q ( r t + 1 , x t + 1 , e t + 1 ) .
Because m t + 1 is latent, empirical validation necessarily relies on proxies drawn from report, expression, behaviour, and institutional records. Proxy inadequacy can already compromise validity before deployment [43]; selective labels can determine which outcomes are observed [44]; and an intervention can change the evaluative standard used in self-report, producing response shift [38]. Criterion endogeneity is distinct: if variables entering C t + 1 obs depend on a t or d t , the criterion-generation process itself is post-treatment. Agreement A ( m ^ t , C t + 1 obs ) can then combine predictive validity, target change, selection, response shift, and system-induced criterion movement. Increased agreement is not sufficient evidence of increased accuracy.
This is post-deployment criterion endogeneity: the operational criterion used to evaluate a mental inference is partly generated under conditions altered by deploying that inference. The problem is particularly acute when the criterion is self-report, free text, engagement, adherence, clinician judgement exposed to the score, or an institutional record created after the prediction.
Two ideal types help locate the problem. An external criterion has a measurement rule whose meaning does not depend constitutively on the subject adopting the model’s classification, although the outcome itself may still be changed by deployment. Loan repayment, task completion, and body temperature are examples. A reflexive criterion is partly produced through the person’s interpretation, self-ascription, or norm-governed expression of the construct being measured. Self-reported anxiety, identity narratives, perceived motivation, and free-text accounts of relational meaning are examples. External and reflexive criteria are ideal types rather than an exhaustive binary: clinical ratings, behavioural traces, and performance records can be mixed criteria whose degree of reflexivity depends on self-ascription, social response, and prior system exposure.
General theories of performative prediction can represent deployment-induced distribution change in either class. Mental AI refines performativity for epistemically coupled criteria: with a reflexive criterion, the later measure is simultaneously evidence about the target and an act through which the target is interpreted or expressed. A rise in model–criterion agreement can therefore reflect better prediction, movement of the operational criterion, movement of the underlying state, or several of these at once.
Criterion endogeneity describes an identification property rather than the value of the intervention. Beneficial systems often intend to change outcomes. Sound evaluation reports that causal benefit separately from the predictive accuracy of the inference that initiated it.
Post-deployment data remain informative when evaluation distinguishes at least three questions:
1.
Did the model predict a criterion that was measured independently of its output?
2.
Did the model improve an outcome through a beneficial intervention?
3.
Did the model alter the criterion, expression, or allocation so that subsequent agreement became partly self-produced?
Randomised exposure to interpretations, delayed disclosure, independent criterion assessment, blinded evaluators, alternative-label conditions, and measurement of state, interpretation, expression, and allocation can help distinguish these possibilities.

5. Recursive Leverage Across Functional Modes

5.1. Four Non-Exclusive Functional Modes

Inferential Mental AI.

The primary function is to infer emotion, belief, intention, personality, motivation, attention, fatigue, risk, or another mental attribute. Emotion recognition and zero-shot personality scoring are examples [2,6].

Relational Mental AI.

The system sustains a personalised relationship through memory, adaptation, role continuity, or reciprocal language. Its inferences can accumulate across episodes and become embedded in a shared interaction history. Companion chatbots illustrate this mode [11,12,13,14]. Persistent adaptive dialogue can reinforce particular interpretations through mirroring, elaboration, and relational authority [15,18,25].

Interventional Mental AI.

The system intentionally attempts to alter a state, belief, behaviour, coping strategy, or decision. Digital therapy, coaching, persuasive systems, and adaptive education can operate in this mode [8,9,62,63].

Institutional Mental AI.

The inference enters a formal or consequential allocation by a school, employer, clinician, insurer, platform, welfare agency, or public authority. The European Union’s AI Act recognises the special risk of some emotion-recognition uses in workplaces and educational institutions [64]. Institutional mode is defined by operative standing, not by whether the system communicates directly with the person.
The modes overlap. A mental-health chatbot can be inferential, relational, and interventional. If its risk score is sent to a clinician or insurer, it also becomes institutional. The overlap matters because effects can compound.

5.2. Impact Dimensions

Four dimensions predict the magnitude and durability of recursive effects.

Persistence and personalisation.

Repeated interaction, memory, and adaptation increase the opportunity for cumulative framing and dependence. Persistence is not intrinsically harmful. It can improve continuity and support. It also increases the need to study path dependence and withdrawal effects.

Operational authority.

An optional reflection has lower formal authority than a score that changes access, workload, treatment, or employment. Perceived knowledge or emotional attunement can also give a system substantial informal authority.

Interpretive opacity.

Users may not know which signals were used, what construct was inferred, how uncertainty was represented, or which response was caused by the inference. Opacity weakens correction and can make a probabilistic output appear categorical.

Contestability.

A person may be able to refuse inference, delay its use, correct contextual errors, request an alternative interpretation, override a recommendation, or appeal an allocation. Contestability changes the practical authority of the same model output.
Table 3. Illustrative classification of systems. Entries are conceptual examples, not empirical risk scores.
Table 3. Illustrative classification of systems. Entries are conceptual examples, not empirical risk scores.
System Primary modes Persistence Authority Typical recursive path
One-shot emotion classifier for research Inferential Low Low No deployed path unless output reaches the person or an institution.
Adaptive tutoring system Inferential, interventional, institutional Medium Medium to high Frustration inference changes difficulty, engagement, and future performance records.
Companion chatbot Relational, inferential, interventional High Informal but potentially high Personalised framing affects attachment, belief, self-description, and later dialogue.
Clinical triage model Inferential, institutional Medium High Risk score changes care priority, clinician attention, and subsequent clinical records.
Workplace affect analytics Inferential, institutional Medium to high High Stress or motivation inference changes monitoring, assignment, or evaluation.
Personal reflective journal assistant Inferential, interventional High User-controlled if well designed Suggested labels alter self-interpretation and later journal language.
Social robot in care Physical, relational, inferential, interventional High Context-dependent Inferred affect changes distance, speech, activity, and caregiver response.

6. Convergent Evidence Across Fragmented Literatures

The evidence arrives as a causal chain distributed across literatures: systems generate mental inferences; people and institutions act on them; interaction changes later judgement or expression; and measures, labels, and classifications can be reactive. No component alone establishes the full Mental AI loop. Their conjunction identifies where direct evidence ends and the new experimental programme begins.

6.1. Artificial Systems Can Generate Psychologically Consequential Inferences

Affective computing has produced extensive methods for detecting and modelling affect from facial movement, speech, physiology, text, and multimodal signals [6,7]. Yet construct validity remains a central challenge. Facial movement does not map transparently and universally onto emotion categories [53,54]. The problem is not merely lower accuracy. It is that the construct inferred by a system may differ from lived experience, communicative intention, culturally patterned expression, or situational meaning.
Large language models expand the range of psychological claims that can be generated from ordinary language. Theory-of-mind benchmarks suggest substantial but uneven capacities [1]. Zero-shot personality scoring indicates that brief open-ended text can support predictions related to self-reports, behaviour, and mental-health outcomes [2]. Task-specific studies also show that LLMs can classify psychological text and approach expert reliability in judging empathic communication when constructs, prompts, and benchmarks are carefully validated [65,66]. Natural-language processing can therefore create useful high-inference representations from data originally produced for communication rather than assessment [3]. These results strengthen rather than dissolve the distinction between technical validity and operative authority: a classifier can be reliable for a specified construct and still require separate justification before its output changes a person’s treatment or self-understanding.
Capability and authority are separate variables. A statistically useful representation can warrant a bounded prediction while lacking the evidence, context, or procedural standing required to govern an intervention or allocation.

6.2. Institutional Mental AI Can Change Both Assessment and Assessed Behaviour

Automated hiring provides a direct institutional testbed for Mental AI. Machine-learning systems have been developed to infer Big Five personality traits from verbal, paraverbal, and nonverbal features of asynchronous video interviews. Across several samples, these assessments showed mixed reliability and validity, with substantially stronger results when trained against interviewer reports than against self-reports [67]. This illustrates both the technical feasibility of institutional psychological inference and the dependence of validity on the criterion selected. Institutional evaluation can also change the evidence it receives. In an experiment on automatically evaluated job interviews, applicants who expected algorithmic rather than human evaluation used less deceptive impression management, perceived fewer opportunities to perform, and gave shorter responses [68]. The assessment arrangement therefore altered the expressive–behavioural data available for assessment. Qualitative research with job-seekers who had directly experienced emotion-AI-enabled interviews documented perceived distributive, procedural, and interactional injustices, as well as demands for transparency, privacy, and contestability [69]. The evidence localises the demonstrated effects at the expressive–behavioural and allocative layers: institutional mental inference changes assessed behaviour, perceived obligation, procedural experience, and access to opportunity. Latent-state change and criterion endogeneity require separate designs.

6.3. People Rely on Algorithmic and Conversational Interpretations

Human reliance on automation is neither uniformly excessive nor uniformly deficient. People can over-rely on systems, reject them after observing errors, or prefer algorithmic advice to human judgement depending on context and presentation [45,46,70,71]. Cognitive forcing functions can reduce overreliance by requiring active engagement [48]. In clinical decision support, incorrect AI advice can influence professional judgement, showing that a human in the loop does not eliminate model effects [49].
Mental AI adds two concerns. First, the advice is about the person who receives or is affected by it. Second, perceived psychological insight can carry a distinctive authority. A person may treat an AI interpretation as evidence about an ambiguous internal state, especially when the system appears personalised, confident, or emotionally responsive. Psychological competence research highlights framing, tone, authority, uncertainty handling, and conversational guidance as evaluation dimensions [19].

6.4. Conversational Systems Can Alter Belief, Judgement, and Social Connection

Generative systems can be persuasive, especially when arguments are personalised or delivered in sustained conversation [8,9]. Anthropomorphism and individual differences help explain why some users experience stronger social connection with AI companions [14]. Qualitative and mixed-method studies document perceived support and relationship development with companion chatbots [11,12,13]. A longitudinal social-media quasi-experiment and user interviews report mixed companion outcomes, including greater affective and interpersonal expression alongside increased language associated with loneliness and suicidal ideation [72]. These findings identify relational trajectories as the relevant unit of analysis; designs addressing residual confounding are required for causal attribution.
Controlled and synthesised clinical evidence establishes beneficial pathways alongside risk. A 12-week preregistered trial found improvements in anxiety and well-being from an adaptive conversational intervention, with engagement and benefit moderated by loneliness, social support, and insecure attachment [73]. Reviews and meta-analyses report benefits for some depressive and anxiety outcomes together with heterogeneity, limited effects for some endpoints, and substantial design and evidence-quality constraints [74,75,76]. The resulting account is conditional: vulnerability, engagement, intervention design, comparator, and outcome selection govern both benefit and risk.
Relational style can itself change model behaviour. Training models for greater warmth increased errors and sycophantic affirmation, particularly when users expressed sadness, even when standard capability tests remained relatively stable [77]. Clinical perspectives similarly describe how reassurance and avoidance can be recursively reinforced in anxiety and obsessive–compulsive presentations [78], while mixed-method work suggests that goal-directed, self-regulated use may matter more for perceived support than warmth alone [79]. Socioaffective alignment therefore argues that evaluation should include the social and psychological systems produced through interaction [18]. The Resonant Amplification Framework likewise proposes that persistent adaptive dialogue can stabilise interpretations through attachment, co-creation, and linguistic reinforcement, while recommending delays, source distinctions, and alternative interpretations as circuit breakers [25].
The bidirectional belief amplification framework formalises a particularly important relational mechanism: human cognitive–emotional biases and chatbot tendencies can mutually reinforce maladaptive beliefs across extended interaction [15]. Its evidential domain is high-risk relational interaction; generalisation to ordinary use and non-conversational institutions remains an empirical question. Within that domain, it shows how a system can enter the user’s evidence environment and alter the salience and stability of beliefs.

6.5. Human-AI Feedback Can Alter Later Judgement

Direct experimental evidence establishes one link in the recursive account. Across experiments, biased AI outputs and human interaction altered later perceptual, emotional, and social judgements, while users could underestimate the source of the change [10]. The demonstrated outcome is later judgement. Whether the same process changes self-interpretation, latent state, reflexive report, or exposed validation criteria is the next empirical step.

6.6. Measurement and Classification Can Be Reactive

Questionnaires, rankings, diagnoses, and social expectations can change the behaviour they measure or classify [35,36,37,39,40]. Intervention research further shows that self-report standards and construct understanding can change across measurement occasions, producing response-shift bias [38]. Hacking’s account of interactive kinds and subsequent analysis of psychiatric classification show how human classifications can enter self-understanding and social practice [41,42]. Field evidence from menstrual-cycle tracking apps directly documents prediction-mediated meaning-making in everyday use, including cases in which users regarded the predictions as faulty [80]. Estimating causal magnitude and direction requires complementary experimental designs. Performative prediction and causal-deployment studies provide formal accounts of distribution and target change after a model enters decision-making [29,30,31,32,33].
The framework converts this broad reactivity literature into moderators and differentiated outcomes. Ambiguity, repeated exposure, perceived authority, personalisation, institutional uptake, low contestability, and consequential allocation should increase recursive leverage. State, report, expression, and environment must then be measured separately.

6.7. Human Expression Is a Variable Object, Not a Transparent Criterion

The problem of criterion endogeneity is intensified by the structure of human expression. Self-report is indispensable but not infallible [55,58]. Expression is shaped by social context, conceptual resources, culture, communicative goals, and risk. Large-scale analysis of 351,734 relationship narratives found that narrative elaboration and expressed affect were only weakly coupled, with broad configurations that included elaborated but muted and brief but charged expression [81]. The associated ANAD resource provides derived features for studying such discrepancy without claiming direct access to lived emotion [82].
This human baseline matters for Mental AI. A system trained to expect congruence between narrative complexity and emotional intensity can treat regulated expression as inconsistency or low affect. If users learn which forms of expression systems recognise, they may adapt their language. Improved machine-user agreement could then reflect expressive compression rather than better interpretation.
Adjacent experiments show that AI mediation changes distributions of human expression. Access to generative-AI story ideas increased some measures of individual creativity while reducing the collective diversity of produced stories [83]. Expecting automatic evaluation changed applicants’ response length and impression-management behaviour [68]. These designs establish expressive malleability under AI assistance and evaluation; repeated mental labelling remains the specific manipulation required to test expressive compression.

6.8. Machine Interpretation Has Boundary Conditions

Average benchmark accuracy can conceal failures under semantic ambiguity, normative conflict, or mixed affect. Algorithmic Affective Blunting operationalises dose-dependent degradation of LLM affective interpretation under semantic stress, while the Affective Thermodynamic Relationship examines scaling between normative conflict and response collapse across tested model families [84,85]. Their model sets and operationalisations define the current evidential boundary. Methodologically, they show why Mental AI evaluation requires stress tests and calibrated abstention alongside average performance.

7. Ten Testable Propositions

The framework earns distinct scientific value only through predictions that separate it from generic claims about AI influence. The ten propositions below specify moderators, outcome layers, comparisons, and evidence that would reduce or eliminate the proposed effects. Table 4 summarises the empirical programme.
Proposition 1: Interpretive uptake. When an AI-generated mental-state interpretation is disclosed to a person, subsequent self-interpretation and report will shift toward that interpretation more than in a matched condition where the inference is withheld or presented as one of several uncertain hypotheses. The effect should be distinguished from demand compliance by delayed measures, private reports, behavioural outcomes, and confidence calibration.
Proposition 2: Ambiguity susceptibility. Interpretive uptake will be larger when baseline self-certainty, emotional granularity, or metacognitive confidence is lower. This proposition follows from the idea that external interpretations have greater leverage when endogenous evidence is weak or ambiguous. It could be tested by manipulating signal ambiguity and measuring baseline interoceptive or metacognitive indices [59,61].
Proposition 3: Persistence and personalisation. Holding content quality constant, repeated interaction with a memory-enabled and personalised system will produce greater stability of AI-consistent interpretation and greater resistance to correction than one-shot interaction. A null result across well-powered longitudinal studies would substantially weaken claims about cumulative relational effects.
Proposition 4: Operational authority. The same mental-state inference will exert larger behavioural and allocative effects when presented as an institutionally operative assessment than as optional personal advice. This effect should be separable from perceived accuracy by independently manipulating source authority and evidential quality.
Proposition 5: Contestability buffer. Systems that provide refusal, delay, uncertainty, editable context, alternative interpretations, and appeal will reduce inappropriate uptake and authority transfer without necessarily eliminating useful support. If contestability reliably has no effect on overreliance, correction, or experienced agency, procedural governance claims should be revised.
Proposition 6: Expressive compression. Conceptual work on the algorithmic self and recursive selfhood predicts that repeated algorithmic mediation can homogenise expression and stabilise machine-compatible identities [20,21]. Adjacent experiments show that AI assistance can reduce collective diversity in human-produced stories and that anticipated automated evaluation can change response length and impression-management behaviour [68,83]. Repeated exposure to a restricted set of machine-recognised mental categories should therefore reduce the lexical, semantic, narrative, or mixed-affect diversity of subsequent expression relative to open-language or plural-hypothesis conditions. The claim fails if exposure produces no reproducible contraction after demand effects, topic, exposure dose, and baseline expressive diversity are controlled.
Proposition 7: Post-deployment criterion endogeneity. Agreement between a model’s earlier inference and later self-report or behaviour will be higher when participants have been exposed to the inference-guided response than when criteria are measured independently or under blinded conditions. If exposure does not alter the criterion across domains, the generality of criterion endogeneity should be sharply limited.
Proposition 8: Exploratory mode interaction. Specific combinations of inferential, relational, interventional, and institutional modes will produce non-additive recursive effects. Their direction and magnitude will depend on persistence, exposure dose, perceived and operational authority, and contestability. The proposition is weakened if factorial and longitudinal studies show that mode combinations are fully explained by the independent additive effects of exposure and authority.
Proposition 9: Beneficial scaffolding. Under ambiguous but non-emergency conditions, systems that present calibrated hypotheses, invite correction, distinguish observation from inference, and preserve human choice will produce greater task-relevant benefit and experienced agency than both no-support and authoritative single-label conditions, without increasing inappropriate reliance. The proposition is weakened if these design features provide no reproducible benefit, reduce useful support, or fail to improve calibration and agency.
Proposition 10: Domain asymmetry. Recursive effects and criterion endogeneity will vary with pre-specified properties of the target domain: availability of an independent criterion, inter-rater observability, dependence on self-ascription, semantic contestability, and temporal stability. Effects should be larger where independent criteria are scarce and self-ascription is constitutive of the measure. The proposition is weakened if these properties do not predict effect-size heterogeneity across domains.

8. Epistemic and Design Consequences

8.1. Scaffolding and Delegation

Criterion endogeneity has a direct design implication: evaluation needs conditions under which contrary evidence, contextual correction, and self-revision remain observable. The same technical capability can support those conditions or close them. Scaffolding supplies information, vocabulary, or alternatives while preserving the person’s next act of interpretation. Delegation makes an external inference the privileged default that pre-empts conflict, displaces context, or acquires operative standing beyond its evidential warrant.
Generative AI can become part of an extended or hybrid cognitive system [86]. Incorporation becomes displacement only when the architecture changes who can revise meaning and which interpretation governs action. A nominally human-in-the-loop process can be effectively delegated when time pressure, interface defaults, organisational incentives, or perceived machine authority make disagreement costly. A highly capable system can remain scaffolding when uncertainty is explicit, alternatives are available, and correction changes the operative policy.
First-person standing is procedural rather than infallibilist. It attaches to the person as the continuous bearer of meaning, action, and consequence. People can misunderstand, conceal, or revise aspects of mental life [55,87]; clinicians, peers, and computational systems can disclose patterns unavailable to unaided introspection. Recent work on self-knowledge, quantified-self technologies, epistemic agency, and reflective systems accordingly locates first-person standing in the capacity to interpret, revise, and act [22,23,88,89,90,91,92]. Epistemic injustice supplies the wider account of how credibility and interpretive resources can be distributed unequally [93].

8.2. Affective Sovereignty as a Procedural Design Response

Affective Sovereignty operationalises this procedural standing for emotion AI [94]. Within Mental AI, its core consists of four capacities:
1.
Refusal: the ability to decline inference or its use where practicable;
2.
Delay: the ability to postpone closure or action while context is added;
3.
Override: the ability to correct or supersede a system interpretation in the operative policy;
4.
Contestation: the ability to obtain reasons, challenge consequences, and appeal consequential use.
Interpretive authority should scale with evidential warrant, consequence, and reversibility. Emergency care, safeguarding, and public-risk contexts change the threshold for immediate action; high-stakes institutional use requires correspondingly stronger provenance, uncertainty disclosure, independent review, and appeal than low-stakes reflective support.

8.3. Design Principles

Five design principles follow from the analysis.

Separate observation from inference.

Interfaces should distinguish what was observed from what was inferred and from what action is recommended.

Represent uncertainty and alternatives.

A single fluent label can imply closure. Systems should expose competing interpretations when the evidence supports them.

Preserve criterion independence.

Validation should include outcomes measured without exposure to the model’s interpretation, blinded evaluators, or pre-deployment baselines when possible.

Audit recursive effects longitudinally.

Evaluation should test whether language, self-report, behaviour, and allocation move after exposure, not only whether users report satisfaction.

Make correction operative.

A user’s correction should change future model behaviour or institutional use. Merely collecting disagreement without altering the policy does not provide meaningful contestability.
Figure 3. Conceptual authority-risk space. Persistence and operational authority are distinct. Contestability can reduce effective authority across both dimensions. The figure is a hypothesis-generating map, not a validated scoring instrument.
Figure 3. Conceptual authority-risk space. Persistence and operational authority are distinct. Contestability can reduce effective authority across both dimensions. The figure is a hypothesis-generating map, not a validated scoring instrument.
Preprints 224749 g003

9. Explanatory Commitments and Domain Limits

9.1. A Common Causal Structure Without Domain Equivalence

The category is broad by design and selective by rule. Emotion recognition, companion systems, tutoring, and clinical triage enter the same analysis only when mental inference, operational uptake, and an inference-indexed consequence path occur together. This conjunction identifies a shared causal structure; the four functional modes and the dimensions of persistence, authority, opacity, and contestability preserve domain differences within it. Generic technological influence lies outside the category because the action is not selected by a representation of the person’s mental life.

9.2. Measurement Without Direct Access to Mental State

The latent state m t is a theoretical target rather than a directly observed quantity. The empirical task is therefore to identify which layer changed and how strongly each available measure supports that interpretation. Private report, language, behaviour, physiology, third-party judgement, and institutional records provide complementary evidence with different error and reactivity structures. A change confined to wording supports an expressive interpretation; a change confined to access or treatment supports an allocative interpretation. Claims about latent-state transformation require stronger triangulation. This discipline is the purpose of the four-layer architecture.

9.3. The Explanatory Test for the Category

A scientific category merits retention when it organises observations more economically, reveals distinctions that alter inference, and generates experiments unavailable from a simple list of component fields. Affective computing, social AI, performative prediction, measurement reactivity, and looping-effects research supply the component mechanisms. Mental AI places them into one sequence from construct definition to policy uptake, differentiated consequence, criterion production, and exposure-aware validation. Its incremental value resides in that sequence and in the experimental contrasts it demands.

9.4. Evidential Maturity and Discriminating Commitments

Current evidence is strongest for psychological inference, human reliance, conversational influence, institutional changes in assessed behaviour, measurement reactivity, and AI-induced shifts in later judgement. Long-term state transformation, repeated expressive compression, mode interactions, and cross-domain criterion endogeneity remain open tests. The theory is therefore committed to outcomes that can reduce its scope. It should be narrowed or rejected if cumulative evidence shows that:
1.
disclosure of mental-state inferences does not affect later interpretation, report, expression, behaviour, or allocation beyond transient demand effects;
2.
persistence, personalisation, and authority do not moderate those effects;
3.
independent criteria produce the same validation conclusions as exposed post-deployment criteria;
4.
repeated well-designed studies show that separating state, interpretive, expressive–behavioural, and allocative effects does not improve prediction, causal explanation, construct validity, or governance decisions relative to treating them as a single outcome class;
5.
contestability and uncertainty-preserving design do not change reliance, correction, agency, or outcomes;
6.
the combined sequence offers no explanatory or methodological gain over domain-specific accounts.
These outcomes mark distinct contraction points for the theory. Failure at the criterion-independence comparison would remove the central evaluation claim; failure of the broader moderators would retain only a narrower taxonomy of inference-guided systems.

10. Toward a Science of Mental AI

A research programme on Mental AI should proceed along five coordinated lines.
First, identification and taxonomy should develop reliable coding rules for mental inference, operational uptake, and recursive consequence, grounded in construct definitions, system documentation, and actual policy use.
Second, recursive-effect experiments should randomise exposure to interpretations, authority, uncertainty, alternatives, and contestability. Outcomes should separately measure state, interpretation, expression, behaviour, and allocation.
Third, longitudinal and relational research should examine path dependence, adaptation, memory, withdrawal, correction resistance, and beneficial support over time. Short laboratory interactions cannot establish durable relational effects.
Fourth, institutional studies should examine how mental inferences acquire operative standing in schools, workplaces, clinics, platforms, insurance, and public services. Technical performance must be separated from due process, accountability, and reversibility.
Fifth, evaluation science should develop pre-deployment baselines, independent criteria, exposure-aware validation, and audits for criterion movement. The central comparison is not only model versus human accuracy. It is prediction under conditions where the predictor may help create the later evidence.
The decisive shift is methodological. Before deployment, validity asks whether an inference corresponds to independently generated evidence. After deployment, validity must also ask whether the inference helped produce that evidence. Machine-generated interpretations already enter conversations and institutions through which mental states are expressed, interpreted, acted upon, and sometimes formed. A science that leaves this causal transition unmodelled can mistake influence for knowledge.

Target Article Rationale for Open Peer Commentary

Open peer commentary can resolve four disputes that no single field can adjudicate alone. First, affective computing and cognitive science can test whether the three-condition category captures a genuine causal class across affect, belief, personality, and motivation. Second, psychometrics and philosophy of mind can determine whether external and reflexive criteria mark a principled measurement distinction and how self-interpretation relates to latent state. Third, causal inference and machine learning can specify when target change, response shift, selective observability, and criterion movement are identifiable in deployed systems. Fourth, clinical science, human–computer interaction, organisational psychology, neuroethics, and law can test how benefit, authority, persistence, and contestability vary across relational and institutional settings. The commentary process would therefore do substantive scientific work: refine the construct, expose domain exceptions, compare rival causal models, and determine which propositions survive disciplinary translation.

Author’s Anticipated Commentary Domains

Relevant commentary domains include affective computing; machine theory of mind; cognitive science of self-knowledge and metacognition; human-AI trust and reliance; conversational persuasion; social and relational AI; digital mental health; causal inference and performative prediction; psychometrics and self-report; interoception; AI governance; neuroethics; education; workplace analytics; and philosophy of mind.
Competing interests: The author declares no competing interests.
Funding: This work received no specific external funding.
Ethics statement: Not applicable. This theoretical synthesis reports no new research involving human participants, identifiable human data, animals, or biological materials.
Data and materials: No new empirical data were generated or analysed. The supplementary material contains the structured literature-search record, claim register, comparison matrices, and proposed experimental designs.
Use of generative AI: Generative AI tools were used only for limited language editing. The author independently verified all sources and retained full responsibility for the content.
Status of cited author work: Only published or accepted work by the author is cited as supporting literature. Unpublished and under-review manuscripts are not used as established evidence.

Supplementary Materials

The following supporting information can be downloaded at the website of this paper posted on Preprints.org.

References

  1. Strachan, J.W.A.; Albergo, D.; Borghini, G.; Pansardi, O.; Scaliti, E.; Gupta, S.; Saxena, K.; Rufo, A.; Panzeri, S.; Manzi, G.; et al. Testing Theory of Mind in Large Language Models and Humans. Nat. Hum. Behav. 2024, 8, 1285–1295. [Google Scholar] [CrossRef] [PubMed]
  2. Wright, A.G.C.; Ringwald, W.R.; Vize, C.E.; Eichstaedt, J.C.; Angstadt, M.; Taxali, A.; Sripada, C. Assessing Personality Using Zero-Shot Generative AI Scoring of Brief Open-Ended Text. Nat. Hum. Behav. 2026, 10, 541–555. [Google Scholar] [CrossRef] [PubMed]
  3. Mihalcea, R.; Biester, L.; Boyd, R.L.; Jin, Z.; Perez-Rosas, V.; Wilson, S.; Pennebaker, J.W. How Developments in Natural Language Processing Help Us in Understanding Human Behaviour. Nat. Hum. Behav. 2024, 8, 1877–1889. [Google Scholar] [CrossRef] [PubMed]
  4. Picard, R.W. Affective Computing; MIT Press: Cambridge, MA, 1997. [Google Scholar]
  5. Picard, R.W. Affective Computing: Challenges. Int. J. Hum.-Comput. Stud. 2003, 59, 55–64. [Google Scholar] [CrossRef]
  6. Calvo, R.A.; D’Mello, S. Affect Detection: An Interdisciplinary Review of Models, Methods, and Their Applications. IEEE Trans. Affect. Comput. 2010, 1, 18–37. [Google Scholar] [CrossRef]
  7. Wang, Y.; Song, W.; Tao, W.; Liotta, A.; Yang, D.; Li, X.; Gao, S.; Sun, Y.; Ge, W.; Zhang, W.; et al. A Systematic Review on Affective Computing: Emotion Models, Databases, and Recent Advances. Inf. Fusion 2022, 83–84, 19–52. [Google Scholar] [CrossRef]
  8. Matz, S.C.; Teeny, J.D.; Vaid, S.S.; Peters, H.; Harari, G.M.; Cerf, M. The Potential of Generative AI for Personalized Persuasion at Scale. Sci. Rep. 2024, 14, 4692. [Google Scholar] [CrossRef] [PubMed]
  9. Salvi, F.; Horta Ribeiro, M.; Gallotti, R.; West, R. On the Conversational Persuasiveness of GPT-4. Nat. Hum. Behav. 2025, 9, 1645–1653. [Google Scholar] [CrossRef] [PubMed]
  10. Glickman, M.; Sharot, T. How Human-AI Feedback Loops Alter Human Perceptual, Emotional and Social Judgements. Nat. Hum. Behav. 2025, 9, 345–359. [Google Scholar] [CrossRef] [PubMed]
  11. Ta, V.; Griffith, C.; Boatfield, C.; Wang, X.; Civitello, M.; Bader, H.; DeCero, E.; Loggarakis, A. User Experiences of Social Support From Companion Chatbots in Everyday Contexts: Thematic Analysis. J. Med. Internet Res. 2020, 22, e16235. [Google Scholar] [CrossRef] [PubMed]
  12. Skjuve, M.; Følstad, A.; Fostervold, K.I.; Brandtzaeg, P.B. My Chatbot Companion: A Study of Human-Chatbot Relationships. Int. J. Hum.-Comput. Stud. 2021, 149, 102601. [Google Scholar] [CrossRef]
  13. Pentina, I.; Hancock, T.; Xie, T. Exploring Relationship Development With Social Chatbots: A Mixed-Method Study of Replika. Comput. Hum. Behav. 2023, 140, 107600. [Google Scholar] [CrossRef]
  14. Folk, D.; Heine, S.J.; Dunn, E. Individual Differences in Anthropomorphism Help Explain Social Connection to AI Companions. Sci. Rep. 2025, 15, 36548. [Google Scholar] [CrossRef] [PubMed]
  15. Dohnány, S.; Kurth-Nelson, Z.; Spens, E.; Luettgau, L.; Reid, A.; Gabriel, I.; Summerfield, C.; Shanahan, M.; Nour, M.M. Technological Folie à Deux: Feedback Loops between AI Chatbots and Mental Health. Nat. Ment. Health 2026, 4, 336–345. [Google Scholar] [CrossRef] [PubMed]
  16. Shevlin, H. Three Frameworks for AI Mentality. Front. Psychol. 2026, 17, 1715835. [Google Scholar] [CrossRef] [PubMed]
  17. Binz, M.; Schulz, E. Using Cognitive Psychology to Understand GPT-3. Proc. Natl. Acad. Sci. 2023, 120, e2218523120. [Google Scholar] [CrossRef] [PubMed]
  18. Kirk, H.R.; Gabriel, I.; Summerfield, C.; Vidgen, B.; Hale, S.A. Why Human-AI Relationships Need Socioaffective Alignment. Humanit. Soc. Sci. Commun. 2025, 12, 728. [Google Scholar] [CrossRef]
  19. Economides, M.; Sacher, P.M.; Salzer, S.; Abellar, A.M.; Tsim, F.; Ferrère, A. Psychological Competence as a Missing Dimension in AI Evaluation, 2026. arXiv arXiv:cs. [CrossRef]
  20. Joseph, J. The Algorithmic Self: How AI Is Reshaping Human Identity, Introspection, and Agency. Front. Psychol. 2025, 16, 1645795. [Google Scholar] [CrossRef] [PubMed]
  21. Lungu, B.A. Machines Looping Me: Artificial Intelligence, Recursive Selves and the Ethics of De-Looping. AI Soc.> 2026 Published online. 2025, 41, 1979–1990. [Google Scholar] [CrossRef]
  22. Kim, M.; Wang, W.; Long, J.; Picard, R.; Barczi, N.; Maes, P. Reflective Agency: Ethical and Empirical Framework for AI-Mediated Self-Reflection Systems. Proc. AAAI/ACM Conf. AI Ethics Soc. 2025, 8, 1440–1452. [Google Scholar] [CrossRef]
  23. Coeckelbergh, M. AI and Epistemic Agency: How AI Influences Belief Revision and Its Normative Implications. Soc. Epistemol. 2026, 40, 59–71. [Google Scholar] [CrossRef]
  24. Shin, D. Automating Epistemology: How AI Reconfigures Truth, Authority, and Verification. AI Soc. Published online. 2026, 41, 1553–1559. [Google Scholar] [CrossRef]
  25. Kim, R.S. Interrupting Resonant Amplification: A Mechanistic and Design Framework for Human-AI Interaction. Comput. Hum. Behav. Rep. 2026, 21, 100975. [Google Scholar] [CrossRef]
  26. Yuste, R.; Goering, S.; Agüera y Arcas, B.; et al. Four Ethical Priorities for Neurotechnologies and AI. Nature 2017, 551, 159–163. [Google Scholar] [CrossRef] [PubMed]
  27. Ienca, M.; Andorno, R. Towards New Human Rights in the Age of Neuroscience and Neurotechnology. Life Sci. Soc. Policy 2017, 13, 5. [Google Scholar] [CrossRef] [PubMed]
  28. UNESCO. Recommendation on the Ethics of Neurotechnology. Adopted by the 43rd Session of the General Conference, 11 November 2025, 2025. [Google Scholar]
  29. Perdomo, J.; Zrnic, T.; Mendler-Dünner, C.; Hardt, M. Performative Prediction. In Proceedings of the Proceedings of the 37th International Conference on Machine Learning. PMLR, 2020, Vol. 119, Proceedings of Machine Learning Research. pp. 7599–7609.
  30. Hardt, M.; Mendler-Dünner, C. Performative Prediction: Past and Future. 2023, 2310.16608. [Google Scholar] [CrossRef]
  31. Boeken, P.; Zoeter, O.; Mooij, J.M. Evaluating and Correcting Performative Effects of Decision Support Systems via Causal Domain Shift. In Proceedings of the Proceedings of the Third Conference on Causal Learning and Reasoning. PMLR, 2024, Vol. 236, Proceedings of Machine Learning Research. pp. 551–569.
  32. Feng, J.; Gossmann, A.; Pennello, G.A.; Petrick, N.; Sahiner, B.; Pirracchio, R. Monitoring Machine Learning-Based Risk Prediction Algorithms in the Presence of Performativity. In Proceedings of the Proceedings of the 27th International Conference on Artificial Intelligence and Statistics. PMLR, 2024, Vol. 238, Proceedings of Machine Learning Research. pp. 919–927.
  33. van Amsterdam, W.A.C.; van Geloven, N.; Krijthe, J.H.; Ranganath, R.; Cinà, G. When Accurate Prediction Models Yield Harmful Self-Fulfilling Prophecies. Patterns 2025, 6, 101229. [Google Scholar] [CrossRef] [PubMed]
  34. Zezulka, S.; Genin, K. Prediction, Performativity, and Potential Outcomes: Communicative Rationality in Prediction-Allocation Problems. In Proceedings of the Proceedings of the 2025 ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization, Non-archival position paper. 2025; ACM; p. 299. [Google Scholar] [CrossRef]
  35. Merton, R.K. The Self-Fulfilling Prophecy. Antioch Rev. 1948, 8, 193–210. [Google Scholar] [CrossRef]
  36. Snyder, M.; Tanke, E.D.; Berscheid, E. Social Perception and Interpersonal Behavior: On the Self-Fulfilling Nature of Social Stereotypes. J. Personal. Soc. Psychol. 1977, 35, 656–666. [Google Scholar] [CrossRef]
  37. Espeland, W.N.; Sauder, M. Rankings and Reactivity: How Public Measures Recreate Social Worlds. Am. J. Sociol. 2007, 113, 1–40. [Google Scholar] [CrossRef] [PubMed]
  38. Howard, G.S. Response-Shift Bias: A Problem in Evaluating Interventions with Pre/Post Self-Reports. Eval. Rev. 1980, 4, 93–106. [Google Scholar] [CrossRef]
  39. French, D.P.; Sutton, S. Reactivity of Measurement in Health Psychology: How Much of a Problem Is It? What Can Be Done About It? Br. J. Health Psychol. 2010, 15, 453–468. [Google Scholar] [CrossRef] [PubMed]
  40. Long, P.A.; Huberts, A.S.; Neureiter di Torrero, A.; Otto, L.R.; Rogge, A.A.; Ritschl, V.; Stamm, T.A. The Mere-Measurement Effect of Patient-Reported Outcomes: A Systematic Review and Meta-Analysis. Qual. Life Res. 2025, 34, 1211–1220. [Google Scholar] [CrossRef] [PubMed]
  41. Hacking, I. Rewriting the Soul: Multiple Personality and the Sciences of Memory; Princeton University Press: Princeton, NJ, 1995. [Google Scholar]
  42. Tsou, J.Y. Hacking on the Looping Effects of Psychiatric Classifications: What Is an Interactive and Indifferent Kind? Int. Stud. Philos. Sci. 2007, 21, 329–344. [Google Scholar] [CrossRef]
  43. Guerdan, L.; Coston, A.; Wu, Z.S.; Holstein, K. Ground(less) Truth: A Causal Framework for Proxy Labels in Human-Algorithm Decision-Making. In Proceedings of the Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, 2023; ACM; pp. 688–704. [Google Scholar] [CrossRef]
  44. Wei, D. Decision-Making Under Selective Labels: Optimal Finite-Domain Policies and Beyond. In Proceedings of the Proceedings of the 38th International Conference on Machine Learning. PMLR, 2021, Vol. 139, Proceedings of Machine Learning Research. pp. 11035–11046.
  45. Parasuraman, R.; Riley, V. Humans and Automation: Use, Misuse, Disuse, Abuse. Hum. Factors 1997, 39, 230–253. [Google Scholar] [CrossRef]
  46. Lee, J.D.; See, K.A. Trust in Automation: Designing for Appropriate Reliance. Hum. Factors 2004, 46, 50–80. [Google Scholar] [CrossRef] [PubMed]
  47. Skitka, L.J.; Mosier, K.L.; Burdick, M. Does Automation Bias Decision-Making? Int. J. Hum.-Comput. Stud. 1999, 51, 991–1006. [Google Scholar] [CrossRef]
  48. Buçinca, Z.; Malaya, M.B.; Gajos, K.Z. To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-Assisted Decision-Making. Proc. ACM Hum.-Comput. Interact. 2021, 5, 1–21. [Google Scholar] [CrossRef]
  49. Gaube, S.; Suresh, H.; Raue, M.; Merritt, A.; Berkowitz, S.J.; Lermer, E.; Coughlin, J.F.; Guttag, J.V.; Colak, E.; Ghassemi, M. Do as AI Say: Susceptibility in Deployment of Clinical Decision-Aids. npj Digit. Med. 2021, 4, 31. [Google Scholar] [CrossRef] [PubMed]
  50. Russell, J.A. Core Affect and the Psychological Construction of Emotion. Psychol. Rev. 2003, 110, 145–172. [Google Scholar] [CrossRef] [PubMed]
  51. Lindquist, K.A.; Wager, T.D.; Kober, H.; Bliss-Moreau, E.; Barrett, L.F. The Brain Basis of Emotion: A Meta-Analytic Review. Behav. Brain Sci. 2012, 35, 121–143. [Google Scholar] [CrossRef] [PubMed]
  52. Cowen, A.S.; Keltner, D. Self-Report Captures 27 Distinct Categories of Emotion Bridged by Continuous Gradients. Proc. Natl. Acad. Sci. 2017, 114, E7900–E7909. [Google Scholar] [CrossRef] [PubMed]
  53. Barrett, L.F.; Adolphs, R.; Marsella, S.; Martinez, A.M.; Pollak, S.D. Emotional Expressions Reconsidered: Challenges to Inferring Emotion From Human Facial Movements. Psychol. Sci. Public Interest 2019, 20, 1–68. [Google Scholar] [CrossRef] [PubMed]
  54. Jack, R.E.; Garrod, O.G.B.; Yu, H.; Caldara, R.; Schyns, P.G. Facial Expressions of Emotion Are Not Culturally Universal. Proc. Natl. Acad. Sci. 2012, 109, 7241–7244. [Google Scholar] [CrossRef] [PubMed]
  55. Nisbett, R.E.; Wilson, T.D. Telling More Than We Can Know: Verbal Reports on Mental Processes. Psychol. Rev. 1977, 84, 231–259. [Google Scholar] [CrossRef]
  56. Wilson, T.D.; Schooler, J.W. Thinking Too Much: Introspection Can Reduce the Quality of Preferences and Decisions. J. Personal. Soc. Psychol. 1991, 60, 181–192. [Google Scholar] [CrossRef] [PubMed]
  57. Schooler, J.W. Re-Representing Consciousness: Dissociations Between Experience and Meta-Consciousness. Trends Cogn. Sci. 2002, 6, 339–344. [Google Scholar] [CrossRef] [PubMed]
  58. Corneille, O.; Gawronski, B. Self-Reports Are Better Measurement Instruments Than Implicit Measures. Nat. Rev. Psychol. 2024, 3, 835–846. [Google Scholar] [CrossRef]
  59. Garfinkel, S.N.; Seth, A.K.; Barrett, A.B.; Suzuki, K.; Critchley, H.D. Knowing Your Own Heart: Distinguishing Interoceptive Accuracy From Interoceptive Awareness. Biol. Psychol. 2015, 104, 65–74. [Google Scholar] [CrossRef] [PubMed]
  60. Khalsa, S.S.; Adolphs, R.; Cameron, O.G.; Critchley, H.D.; Davenport, P.W.; Feinstein, J.S.; Feusner, J.D.; Garfinkel, S.N.; Lane, R.D.; Mehling, W.E.; et al. Interoception and Mental Health: A Roadmap. Biol. Psychiatry Cogn. Neurosci. Neuroimaging 2018, 3, 501–513. [Google Scholar] [CrossRef] [PubMed]
  61. Fleming, S.M.; Lau, H.C. How to Measure Metacognition. Front. Hum. Neurosci. 2014, 8, 443. [Google Scholar] [CrossRef] [PubMed]
  62. Fitzpatrick, K.K.; Darcy, A.; Vierhile, M. Delivering Cognitive Behavior Therapy to Young Adults With Symptoms of Depression and Anxiety Using a Fully Automated Conversational Agent (Woebot): A Randomized Controlled Trial. JMIR Ment. Health 2017, 4, e19. [Google Scholar] [CrossRef] [PubMed]
  63. Inkster, B.; Sarda, S.; Subramanian, V. An Empathy-Driven, Conversational Artificial Intelligence Agent (Wysa) for Digital Mental Well-Being: Real-World Data Evaluation Mixed-Methods Study. JMIR mHealth uHealth 2018, 6, e12106. [Google Scholar] [CrossRef] [PubMed]
  64. European Union. Regulation (EU) 2024/1689 Laying Down Harmonised Rules on Artificial Intelligence. Off. J. Eur. Union 2024. [Google Scholar] [CrossRef]
  65. Bunt, H.L.; Goddard, A.; Reader, T.W.; Gillespie, A. Validating the Use of Large Language Models for Psychological Text Classification. Front. Soc. Psychol. 2025, 3, 1460277. [Google Scholar] [CrossRef]
  66. Kumar, A.; Poungpeth, N.; Yang, D.; Farrell, E.; Lambert, B.L.; Groh, M. When Large Language Models Are Reliable for Judging Empathic Communication. Nat. Mach. Intell. 2026, 8, 173–185. [Google Scholar] [CrossRef]
  67. Hickman, L.; Bosch, N.; Ng, V.; Saef, R.; Tay, L.; Woo, S.E. Automated Video Interview Personality Assessments: Reliability, Validity, and Generalizability Investigations. J. Appl. Psychol. 2022, 107, 1323–1351. [Google Scholar] [CrossRef] [PubMed]
  68. Langer, M.; König, C.J.; Hemsing, V. Is Anybody Listening? The Impact of Automatically Evaluated Job Interviews on Impression Management and Applicant Reactions. J. Manag. Psychol. 2020, 35, 271–284. [Google Scholar] [CrossRef]
  69. Pyle, C.; Roemmich, K.; Andalibi, N. U.S. Job-Seekers’ Organizational Justice Perceptions of Emotion AI-Enabled Interviews. Proc. ACM Hum.-Comput. Interact. 2024, 8, 1–42. [Google Scholar] [CrossRef]
  70. Dietvorst, B.J.; Simmons, J.P.; Massey, C. Algorithm Aversion: People Erroneously Avoid Algorithms After Seeing Them Err. J. Exp. Psychol. General. 2015, 144, 114–126. [Google Scholar] [CrossRef] [PubMed]
  71. Logg, J.M.; Minson, J.A.; Moore, D.A. Algorithm Appreciation: People Prefer Algorithmic to Human Judgment. Organ. Behav. Hum. Decis. Process. 2019, 151, 90–103. [Google Scholar] [CrossRef]
  72. Yuan, Y.; Zhang, J.; Aledavood, T.; Zhang, R.; Saha, K. Mental Health Impacts of AI Companions: Triangulating Social Media Quasi-Experiments, User Perspectives, and Relational Lens. In Proceedings of the Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, 2026; ACM; pp. 1–22. Article 382. [Google Scholar] [CrossRef]
  73. Shoshani, A.; Gurfinkel, B.; Kor, A.; Kanarek, O.; Segev, R.; Shafir, O.; Arbel, R.; Ben-Haim, Y. Attachment, Loneliness, and Social Support as Moderators of Conversational AI-Based Mental Health Outcomes. npj Digital Medicine. Online ahead of print. 2026. [CrossRef]
  74. Hang, Y.; Wu, W.; Feng, Y.; Yan, K.; Liu, Y.; Xiao, X.; Qiao, Z. The Effectiveness of CBT-Based NLP-Enabled AI Conversational Agents for Mental Health Intervention: A Systematic Review and Meta-Analysis. npj Digit. Med. Online ahead of print. 2026. [Google Scholar] [CrossRef] [PubMed]
  75. Yip, W.L.T.; Yang, Y.Y.; Wang, Z.L.; Stuckler, D. Efficacy of AI-Delivered Cognitive Behavioral Therapy Interventions for Anxiety and Depressive Symptoms: A Systematic Review. npj Digital Medicine. Online ahead of print. 2026. [CrossRef]
  76. Sohn, J.S.; Ha, B.G.; Park, S.; Kim, J.J.; Lee, E.; Oh, H.; Lee, S.; Kim, E. Systematic Review and Meta Analysis of Chatbots in the Management of Depressive and Anxiety Symptoms. npj Digit. Med. 2026, 9, 377. [Google Scholar] [CrossRef] [PubMed]
  77. Ibrahim, L.; Hafner, F.S.; Rocher, L. Training Language Models to Be Warm Can Reduce Accuracy and Increase Sycophancy. Nature 2026, 652, 1159–1165. [Google Scholar] [CrossRef] [PubMed]
  78. Golden, A.; Aboujaoude, E. A Transdiagnostic Model for How General Purpose AI Chatbots Can Perpetuate OCD and Anxiety Disorders. npj Digit. Med. 2026, 9, 343. [Google Scholar] [CrossRef] [PubMed]
  79. Lan, Y.; Liu, S.; Liu, C.; Chen, H. Trust-Driven Healthy Engagement with Conversational AI for Mental Health Support in Young Adults: A Mixed Methods Study. Sci. Rep. Online ahead of print. 2026. [Google Scholar] [CrossRef] [PubMed]
  80. Zhou, W.; Karaturhan, P.; Weilenmann, A.; Zhu, J. “It Became a Self-Fulfilling Prophecy”: How Lived Experiences Are Entangled with AI Predictions in Menstrual Cycle Tracking Apps. In Proceedings of the Proceedings of the 2026 ACM Designing Interactive Systems Conference, 2026; ACM; pp. 2176–2192. [Google Scholar] [CrossRef]
  81. Kim, R.S. Narrative-Affect Discrepancy as a Regulated Degree of Freedom in 351,734 Relationship Narratives. PLoS ONE 2026, 21, e0348715. [Google Scholar] [CrossRef] [PubMed]
  82. Kim, R.S. ANEST Narrative-Affect Dataset (ANAD v1): A Large-Scale Derived Feature Resource for Quantifying Narrative-Affective Discrepancy. Data Brief. 2026, 66, 112643. [Google Scholar] [CrossRef] [PubMed]
  83. Doshi, A.R.; Hauser, O.P. Generative AI Enhances Individual Creativity but Reduces the Collective Diversity of Novel Content. Sci. Adv. 2024, 10, eadn5290. [Google Scholar] [CrossRef] [PubMed]
  84. Kim, R.S. Algorithmic Affective Blunting Quantifies the Collapse Curve of Interpretative Failure in Large Language Models. In Discover Artificial Intelligence; 2026. [Google Scholar] [CrossRef]
  85. Kim, R.S. The Affective Thermodynamic Relationship: An Empirical Information-Theoretic Scaling Relationship for Normative-Conflict Collapse in Large Language Models. In Communications AI & Computing Accepted and in production; 2026. [Google Scholar]
  86. Clark, A. Extending Minds with Generative AI. Nat. Commun. 2025, 16, 4627. [Google Scholar] [CrossRef] [PubMed]
  87. Moran, R. Authority and Estrangement: An Essay on Self-Knowledge; Princeton University Press: Princeton, NJ, 2001. [Google Scholar]
  88. Doyle, C. Might Technology Undermine First-Person Authority? Erkenntnis Published online. 2026, 91, 81–101. [Google Scholar] [CrossRef]
  89. Doyle, C. Listening to Algorithms: The Case of Self-Knowledge. Eur. J. Philos. 2025, 33, 134–147. [Google Scholar] [CrossRef]
  90. Fan, N. Why AI Is Not an Epistemic Authority Regarding Our Own Minds. Erkenntnis. Online first. 2025. [CrossRef]
  91. Billesbach, G.; Johnson, S.M. Self-Interpretive Freedom and the Politics of Quantified Self Technologies. AI Soc. 2026, 41, 2961–2973. [Google Scholar] [CrossRef]
  92. van Zyl, L.E. The AI-IARA Framework: How to Cultivate Human Agency before Artificial Intelligence Optimizes It a(ny)way. J. Posit. Psychol. Online first. 2026. [Google Scholar] [CrossRef]
  93. Fricker, M. Epistemic Injustice: Power and the Ethics of Knowing; Oxford University Press: Oxford, 2007. [Google Scholar]
  94. Kim, R.S. Formal and Computational Foundations for Implementing Affective Sovereignty in Emotion AI Systems. Discov. Artif. Intell. 2026, 6, 235. [Google Scholar] [CrossRef]
Figure 1. Mental AI is an orthogonal classification dimension. It identifies systems whose inferences about human mental life are operationally coupled to recursively consequential actions.
Figure 1. Mental AI is an orthogonal classification dimension. It identifies systems whose inferences about human mental life are operationally coupled to recursively consequential actions.
Preprints 224749 g001
Table 2. Criterion and identification failures relevant to Mental AI.
Table 2. Criterion and identification failures relevant to Mental AI.
Problem What is compromised Relation to Mental AI
Construct invalidity The inferred construct does not match the phenomenon of interest The system may operationalise a psychologically named but conceptually different attribute.
Proxy-label bias [43] The available label is an imperfect proxy for a latent target Validation can be biased before any recursive effect occurs.
Selective observability [44] Outcomes are observed only under policies that permit them to occur Institutional decisions can determine whose later performance or recovery is measurable.
Response shift [38] An intervention changes the respondent’s evaluative standard or construct understanding A later self-report may not be commensurable with the pre-intervention report.
Performative target change [29,34] Deployment changes the target outcome or its distribution Policy effects must be separated from predictive performance.
Criterion endogeneity The criterion-generation process lies downstream of the inference-guided policy Changed evidence can be mistaken for independent confirmation of the earlier mental inference.
Table 4. Propositions, illustrative tests, outcomes, and falsification conditions.
Table 4. Propositions, illustrative tests, outcomes, and falsification conditions.
Proposition Illustrative comparison Primary outcomes Evidence that would weaken the claim
Interpretive uptake Disclosed label vs withheld label vs plural hypotheses Report shift, confidence, private choice, delayed narrative No reproducible effect beyond demand characteristics
Ambiguity susceptibility High vs low ambiguity; baseline metacognitive confidence Interaction between ambiguity and uptake Equivalent effects across certainty levels and tasks
Persistence Memory-enabled longitudinal system vs matched one-shot system Stability, correction resistance, withdrawal cost No difference after content and exposure are controlled
Operational authority Institutional assessment vs optional advice Behaviour, compliance, allocation, perceived obligation Authority manipulation has no independent effect
Contestability Editable, delayable, appealable output vs fixed output Appropriate reliance, correction, agency, utility No improvement in calibration or agency
Expressive compression Narrow labels vs open language vs plural taxonomy Lexical, semantic, narrative, and mixed-affect diversity No contraction after repeated exposure
Criterion endogeneity Exposed criterion vs blinded independent criterion Prediction-criterion agreement and mediation paths Agreement gains remain identical under independent criteria
Mode interaction Factorial combinations of modes under matched exposure Direction and non-additivity of effects across outcome layers Combinations are fully explained by independent additive effects
Beneficial scaffolding Calibrated contestable hypotheses vs no support vs authoritative label Task benefit, agency, calibration, inappropriate reliance No reproducible advantage or reduced useful support
Domain asymmetry Domains pre-rated on criterion independence, observability, self-ascription, contestability, and stability Property-linked effect-size heterogeneity and criterion movement Pre-specified domain properties do not predict heterogeneity
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.