Preprint
Article

This version is not peer-reviewed.

Personality-Adaptive Conversational AI for Emotional Support: A Simulation Study Integrating Big Five Detection with Zurich Model Regulation

Submitted:

29 May 2026

Posted:

01 June 2026

You are already at the latest version

Abstract
LLM-based conversational agents generate fluent responses but remain limited in adapting their supportive style to individual personality and emotional needs. We address this gap through a theory-grounded personality-adaptive conversational AI pipeline that performs turn-by-turn Big Five trait detection and applies Zurich Model-aligned behavioural regulation, orchestrated with PROMISE (a model-driven framework for state-based LLM orchestration). We generated 20 simulated dialogue sessions (120 assistant turns total) comparing regulated and non-adaptive assistants across two boundary-condition Big Five profiles (all traits set to +1 vs. -1), evaluated using a structured LLM-based evaluator complemented by author-conducted qualitative review. The system achieved near-perfect implementation fidelity (98.3% detection accuracy; 100% regulation adherence). Personality Needs Addressed were met consistently under regulation (100%) but rarely at baseline (8.3%; Cohen's d = 4.42, p < 0.001), reflecting a binary architectural difference rather than a graded dose-response improvement. Emotional tone and relevance did not differ between conditions (d = 0.000 and d = 0.183, ns), confirming selective enhancement of personalisation without degrading generic conversational quality. These findings establish a reproducible detection → regulation → evaluation framework for future empirical validation of personality-adaptive AI-augmented mental health support with human participants.
Keywords: 
;  ;  ;  ;  

1. Introduction

Conversational agents are increasingly used in healthcare to improve health literacy, encourage treatment adherence, support well-being monitoring, and provide emotionally supportive guidance (Dingler et al., 2021; Laranjo et al., 2018). As these systems become more capable and more widely deployed, their effectiveness depends not only on whether they produce relevant information, but also on how they interact with users. In many health contexts, the intended impact emerges through communication, such as whether an agent can build trust, provide reassurance, sustain motivation, and adapt its tone and strategy to the user over time (Dingler et al., 2021; Vasiliu et al., 2021). Designing such interaction qualities remains challenging, especially in sensitive settings where personalisation and consistency are critical (Kocaballi et al., 2022) and where underlying model reliability remains an open concern (Kaddour et al., 2023).
Moreover, health support can rarely be achieved through one-size-fits-all communication. Effective interactions often require behaviours such as empathy, reassurance, encouragement, and persuasion, but these behaviours are not universally appropriate in a fixed form. An empathetic style may strengthen trust and rapport by mirroring effective physician–patient communication (Rapp et al., 2021). However, whether such a style is perceived as supportive may depend on the person and on how the interaction unfolds over time (Han et al., 2022; Juquelier et al., 2025; Seitz, 2024). Similarly, persuasive support in digital health depends on selecting communication strategies that fit the receiver, the goal, and the moment of interaction rather than simply presenting information (Abernethy et al., 2022; Goetz and Schork, 2018; O’Keefe, 2016). If conversational agents are expected to achieve intended outcomes such as literacy, adherence, reassurance, or emotional support, they thus require flexible means of implementing personalised conversational behaviour during the interaction rather than relying solely on static design-time instructions.
Recent advances in large language models (LLMs) have greatly expanded the potential of conversational agents to produce fluent, contextually rich, and human-like dialogue (Vasiliu et al., 2021; Milne-Ives et al., 2020; Wei et al., 2022). This makes them promising building blocks for adaptive health-oriented conversational agents (Ayers et al., 2023; Färber et al., 2023, 2024; Staehelin et al., 2024). However, most LLM-based approaches focus primarily on response generation. Training or fine-tuning domain-specific models remains costly and data-intensive (Kaddour et al., 2023; Strubell et al., 2019; Ding et al., 2023; Hu et al., 2021), while prompt engineering, despite its practical flexibility (Korzynski et al., 2023; White et al., 2023; Fernando et al., 2023; Liu et al., 2023; Hou et al., 2022), often leaves open the central design question of what conversational behaviour should be enacted, for whom, and when. In other words, LLMs can make adaptive interaction plausible, but they do not by themselves provide grounded, controllable, or reproducible personalisation.
Emerging agentic AI and orchestration frameworks offer part of the answer by enabling modular, context-sensitive systems with structured control over goals, prompts, and interaction flow (Wu et al., 2024). PROMISE, for example, supports phase-specific prompts and transition logic through model-driven state-machine orchestration, making it useful for transparent and trust-sensitive applications. Yet structured interaction flow alone is insufficient when the core requirement is not just to progress through phases, but to regulate socially meaningful behaviour in ways that remain responsive to the user. The design gap, therefore, is not merely how to structure multi-turn LLM interactions, but how to integrate grounded personalisation into those structures so that conversational behaviour remains adaptable, theory-informed, and open to evaluation.
This paper addresses that gap through a framework for grounded social behaviour regulation in conversational AI. Rather than treating personalisation as ad hoc prompt variation, we operationalise it as a structured process that combines dynamic personality modelling, theory-grounded behavioural adaptation, and controlled orchestration. In the present instantiation, relatively stable user differences are modelled through turn-by-turn Big Five detection (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism; collectively OCEAN), behavioural adaptation is guided by the Zurich Model of Social Motivation, and interaction control is implemented through PROMISE with a social behaviour regulation extension. Mental health support serves as a particularly demanding demonstration context because emotionally supportive interactions place strong demands on sensitivity, trust, and adaptation, but the broader contribution is a design approach for implementing personalised conversational behaviour in health-oriented conversational agents.
We evaluate this approach in a controlled simulation study comparing regulated and non-adaptive assistants across boundary-condition personality profiles. The study provides a reproducible testbed for examining whether grounded social behaviour regulation can be implemented with high fidelity and whether it selectively improves personality-specific need fulfilment without degrading generic conversational quality. Our results show near-perfect implementation fidelity (98.3% detection accuracy; 100% regulation adherence) and a marked increase in personality-specific need fulfilment under regulation (100% vs. 8.3%; Cohen’s d = 4.42 , p < 0.001 ), while emotional tone and topical relevance remain stable. Taken together, these findings position the work not only as a pipeline for one application setting, but as a concrete answer to the broader design question of how conversational agents can realise grounded personalisation that remains controllable, reproducible, and evaluable.

3. Materials and Methods

This section presents simulation-based experimental methodology testing whether dynamic personality modelling, theoretically grounded regulation, and validated evaluation produce measurable improvements in conversational personalisation. We employ a 2 × 2 factorial design comparing personality-adaptive (regulated) assistants against non-adaptive (baseline) assistants across two extreme personality profiles.

3.1. Overview and Research Objectives

This study addresses three empirical questions:
1.
Detection Accuracy: Can stable personality traits be accurately inferred from conversational cues on a turn-by-turn basis within multi-turn interactions?
2.
Regulation Effectiveness: Does theory-driven personality-adaptive regulation produce measurable improvements in personality-specific need fulfilment beyond generic responses?
3.
Evaluation Validity: How should results be interpreted given reliance on LLM-based evaluation, and what human validation is required to establish evaluation credibility?

3.1.1. Experimental Design

We employed a 2 × 2 factorial design (Montgomery, 2017; Collins et al., 2009) comparing two assistant types (regulated vs. baseline) across two personality profiles (Type A vs. Type B), with 5 dialogue sessions per cell (20 sessions total; 120 dialogue turns). Equivalently, this can be viewed as 10 simulated personas (5 Type A, 5 Type B) each evaluated under two assistant conditions (regulated and baseline). These extreme configurations are intentionally used as boundary-condition personality profiles to verify detection and regulation fidelity under maximal signal conditions.
Regulated assistants implemented dynamic personality detection and trait-aligned behavioural regulation based on the Zurich Model (Quirin et al., 2023; Berlyne, 1960; Hebb, 1949; Bischof, 1985, 1993), updating personality inferences per-turn. Baseline assistants provided emotionally supportive responses without personality-specific adaptation. Each assistant was implemented as a separate GPT-4 instance (OpenAI API, model: gpt-4-0613, accessed June–August 2024) with customised system prompts.
The workflow encompassed four phases: (1) system configuration and personality profile implementation, (2) simulated dialogue generation following standardised protocols, (3) evaluation using a structured comparative framework, and (4) statistical analysis (Figure 6).

3.1.2. Sample Size Considerations

The 20 dialogue sessions (5 per cell; 10 per assistant type when aggregated across profiles) represent distinct GPT-4 instantiations with condition-specific system prompts, not human subjects. This sample size was selected for architectural validation purposes based on: (1) a controlled simulation design with limited stochastic variation, (2) computational cost constraints appropriate for proof-of-concept testing, and (3) focus on demonstrating implementation fidelity rather than estimating population-level effects (Carroll et al., 2007).
Traditional power analysis is inappropriate because: (1) simulations are controlled and not designed to estimate population-level variability, (2) observed ceiling effects (zero variance) violate power analysis assumptions, and (3) the goal is technical validation, not inferential generalisation.
Future human validation will require larger samples (e.g., n 50 –100 per condition), in line with clinical trial standards for statistical power, attrition, and inter-individual variability, to capture real-world heterogeneity absent from this simulation (Faul et al., 2009; Schulz et al., 2010; Chan et al., 2013).

3.1.3. Ethics and Transparency

No human subjects were involved in this study; the user-side prompts/questionnaires were written manually, and the resulting dialogues were generated in a controlled manner using GPT-4-based agents following predefined personality profiles. As the data were synthetic, non-identifiable, and created solely for technical validation, the study fell outside the scope of human-subject research as defined by established ethical frameworks, and formal ethics approval was therefore not required (Association, 2013).
Future studies involving human participants will require full institutional review board approval, comprehensive informed consent addressing AI-based personality profiling and data usage, real-time safety monitoring with clear escalation pathways for psychological distress, and ongoing oversight by qualified mental health professionals (D’Alfonso, 2020; Torous et al., 2020). Ethical design of conversational agents should also address questions of interactional respect and user autonomy (Alberts et al., 2024).

3.2. Personality-Adaptive System Architecture

The personality-adaptive system centres on turn-by-turn Big Five detection and regulation over the user model, integrated in PROMISE (Wu et al., 2024) through state-attached instructions, storage-backed trait values, and finite-state transitions. Its distinguishing mechanisms are: (i) tri-state personality detection, with each OCEAN dimension encoded as { 1 , 0 , + 1 } and an explicit neutral state when evidence is insufficient (Table 1); (ii) cumulative evidence updating, so each user turn refines trait estimates from the full dialogue history rather than from isolated utterances; (iii) confidence-threshold triggering, committing a trait out of neutrality only when detector confidence exceeds τ (specified in Section 3.3); (iv) parallel per-trait detectors, i.e., independent LLM-based probes per dimension scored concurrently; and (v) Zurich Model trait-to-behaviour mapping, translating non-neutral traits into stylistic regulation directives along Security, Arousal, and Affiliation before response generation (Section 3.3.1). These mechanisms define the technical contribution evaluated in our simulation. We operationalize the D–R–E workflow—integrating PROMISE orchestration, turn-by-turn OCEAN detection, and Zurich Model-aligned regulation—within a controlled factorial design and provide systematic statistical outcome analysis.
Within PROMISE, we organise work into detection, regulation, and evaluation functions for structured operation and quality logging. Figure 1 provides the artefact overview, while Figure 2 shows how orchestration and adaptive regulation operate in parallel within the interaction flow. PROMISE supplies model-driven orchestration via hierarchical state-machine modelling. In PROMISE, an interaction is encapsulated as an Agent wrapping a state machine composed of States (with state-specific instruction templates and behaviours) connected by Transitions governed by Decisions (e.g., guards or triggers) and accompanied by Actions (e.g., extraction or summarisation). These elements operate over an interaction storage abstraction implemented as a key–value store (Wu et al., 2024). This structure enables explicit state transitions, systematic composition of instructions from modular parts, and transparent logging of decisions and actions across complex interactions.
We embed the core D–R–E pipeline into PROMISE’s state-orchestrated architecture as follows:
  • Detection: Trait estimation is implemented as a state-driven process that updates a continuous Big Five vector and confidence metrics in the interaction store after each user input. Traits are inferred from linguistic cues in the current interaction context and persist as storage values accessible across states.
  • Regulation: Within response-generation states, prompts are augmented with trait-conditioned instructions that softly modulate the assistant’s response style—such as tone, pacing, and interpersonal framing—without altering semantic content or task objectives. This is achieved by dynamically composing the base state prompt with trait-aligned augmentation prompts at runtime.
  • Evaluation: After each response generation, structured assessment actions execute LLM-based quality checks and log evaluation scores alongside turn identifiers in the interaction store. In the current study, these evaluation outputs are used for offline analysis and do not feed back into detection or regulation during the simulation.
This mapping allows the full pipeline to be inspected through PROMISE’s explicit state/transition structure and logged decisions/actions, improving reproducibility and auditability of adaptation behaviour. The proposed detection–regulation–evaluation logic is conceptually framework-agnostic and could be instantiated in alternative orchestration systems; PROMISE is used here because its declarative state modelling facilitates systematic traceability.
The architecture follows three core principles: (i) Separation of Concerns—independent modules with well-defined interfaces enable parallel development and isolated testing, communicating through standardised data structures (JSON) containing personality vectors, regulation instructions, and evaluation scores; (ii) Scalability—new personality dimensions or regulation strategies can be added without modifying existing components, with trait-to-regulation mapping externalised as configuration files; and (iii) Traceability and Reproducibility—all detection decisions, regulation instructions, and evaluation scores are logged with timestamps and dialogue turn identifiers, and PROMISE’s declarative state modelling ensures all behavioural adaptations are traceable to explicit state transitions.
The system processes messages through two coordinated layers of activity. Figure 2 shows the parallel orchestration and adaptive-regulation flow at the interaction level, while Figure 3 details the OCEAN perception component that converts dialogue evidence into an explicit personality-state signal for downstream adaptation.
PROMISE operationalises complex interactions by composing prompts from modular parts attached to states and transitions (Wu et al., 2024). In typical PROMISE execution, transition decisions (trigger/guard) are evaluated against the accumulated state utterances before generating the next assistant response; if no transition fires, the assistant response is generated from a composed prompt that includes the state prompt and the dialogue history. PROMISE also supports nested state machines (“outer states”) that maintain an aggregated utterance history across inner states and can inject shared role prompts into all contained states, enabling consistent behaviour across multi-stage interactions. Additionally, PROMISE supports placeholder-based instruction templates filled at runtime and actions that extract structured outputs into interaction storage (key–value store), which can be reused by downstream states and system components (Wu et al., 2024). Our integration leverages these mechanisms to (i) persist the current Big Five vector and confidence in storage, (ii) compose trait-aligned regulation instructions into the response-generation prompt, and (iii) log evaluation outputs with turn identifiers for traceability.

3.3. System Implementation

This section describes how the system implements (i) turn-by-turn personality detection and (ii) theory-driven regulation, both orchestrated within PROMISE (Section 3.2).
We represent the user’s Big Five (OCEAN) profile as a five-dimensional vector P = ( O , C , E , A , N ) , with each trait { 1 , 0 , + 1 } (Costa and McCrae, 1992; John et al., 2008; McCrae and John, 1992). In this simulation, the personality vector functions as a predefined reference target for controlled evaluation rather than psychometric ground truth; detection accuracy denotes agreement with the pre-specified simulated profile.
After each user input, the system updates P using cumulative evidence across dialogue history, following a belief-updating logic similar to Bayesian belief updating (Gelman et al., 2013). PROMISE orchestrates this pipeline via hierarchical state machines, parallel trait detectors, and explicit transition logic (Section 3.2). Inference produces trait values (-1/0/+1) and confidence scores, with trait state transitions triggered when confidence exceeds the pre-specified threshold τ = 0.7 (Figure 3). This threshold was chosen as a conservative design trade-off: it reduces premature trait commitments from ambiguous evidence while still allowing adaptation once repeated conversational cues accumulate.
The detection component is implemented as a modular service integrated into PROMISE. It uses GPT-4 (gpt-4-0613) with trait-specific prompts grounded in personality psychology (White et al., 2023; Funder, 1995; Vazire, 2010; OpenAI, 2023; Brown et al., 2020; Reynolds and McDonell, 2021; Liu et al., 2023). The pipeline executes the following stages: (1) message processing, (2) prompt assembly, (3) parallel LLM inference across traits, (4) response parsing and confidence assessment, (5) state update, and (6) downstream transmission to the regulation module.
Average processing time was 2.3 ± 0.8 seconds per turn (measured across 120 turns), with parallel trait detection reducing latency by ~60% versus sequential processing.
Figure 4. Regulation and behaviour component. Personality-state signals are transformed into behaviour cues, conflict-resolved into a single regulation instruction, merged with the base prompt, and logged for traceability before adapted response generation.
Figure 4. Regulation and behaviour component. Personality-state signals are transformed into behaviour cues, conflict-resolved into a single regulation instruction, merged with the base prompt, and logged for traceability before adapted response generation.
Preprints 216002 g004

3.3.1. Theory-Driven Regulation

Regulation translates detected traits into behavioural adaptations grounded in the Zurich Model of Social Motivation (Quirin et al., 2023; Berlyne, 1960; Hebb, 1949; Bischof, 1985, 1993), which links personality to three motivational domains: Security, Arousal, and Affiliation (Section 2.1). Table 2 and Figure 5 present the mapping from Big Five traits to motivational domains and corresponding behavioural adjustment strategies.
After each personality update, PROMISE composes an integrated regulation instruction by concatenating trait-relevant behaviour prompts for all non-neutral traits (Figure 6). When trait recommendations conflict, the system prioritizes emotional safety (Security domain) over stimulation (Arousal domain), consistent with clinical guidance (APA, 2010).
Figure 5. Conceptual mapping from perception to behaviour. Big Five/OCEAN trait signals are interpreted through Zurich Model motivational domains and translated into conversational behaviour effects used to guide response generation.
Figure 5. Conceptual mapping from perception to behaviour. Big Five/OCEAN trait signals are interpreted through Zurich Model motivational domains and translated into conversational behaviour effects used to guide response generation.
Preprints 216002 g005
Figure 6. Proof-of-concept evaluation design. Simulated personality profiles are assigned to baseline and regulated conditions, dialogues are generated under both conditions, and the resulting outputs are prepared for structured comparison and downstream analysis.
Figure 6. Proof-of-concept evaluation design. Simulated personality profiles are assigned to baseline and regulated conditions, dialogues are generated under both conditions, and the resulting outputs are prepared for structured comparison and downstream analysis.
Preprints 216002 g006

3.4. Experimental Materials and Procedure

We implemented two boundary-condition personality profiles representing extremes of human personality variation (McCrae and Costa, 2003; Goldberg, 1993, 1992). Type A (all traits: + 1 , + 1 , + 1 , + 1 , + 1 ) embodied extreme positive trait levels. Type B (all traits: 1 , 1 , 1 , 1 , 1 ) embodied opposite extremes.
These fully polarized profiles were deliberately selected because they produce the strongest possible personality signals, enabling clear-cut validation of detection and regulation mechanisms. While not representative of typical distributions (Costa and McCrae, 1992), such profiles offer superior diagnostic value for validating this implementation before testing with moderate profiles (McCrae and Costa, 1992).
Detection and regulation performance are evaluated against these synthetic reference profiles as ground truth within the simulation environment. Profiles were encoded into GPT-4 system prompts consistently exhibiting target trait levels across turns.
Twenty dialogue sessions (assistant-condition instantiations) were generated across four 2 × 2 cells: Regulated–Type A ( n = 5 ), Regulated–Type B ( n = 5 ), Baseline–Type A ( n = 5 ), Baseline–Type B ( n = 5 ). Regulated assistants received system prompts incorporating dynamic detection modules and trait-aligned regulation instructions, with inferences updating per-turn. Baseline assistants received static supportive response prompts without personality-specific adaptation. Each assistant instantiation used a separate GPT-4 call configuration (model: gpt-4-0613, generation temperature = 0.7, max_tokens = 500). The temperature value was selected to balance conversational naturalness with experimental control: it allows moderate response variation while avoiding the high stochasticity that could obscure condition-specific regulation effects. GPT-4 (gpt-4-0613) was selected because it represented the state of the art in instruction-following at the time of data collection (June–August 2024) and provided the balance of controllability, context length, and response quality required by the PROMISE orchestration framework. The core architectural findings—that explicit regulation instructions produce consistent personality-need fulfillment whereas non-adaptive prompting does not—are expected to generalize to successor models (e.g., GPT-4o, GPT-4-turbo); if anything, stronger base models may yield higher detection accuracy and more nuanced regulation, though independent validation with those models remains necessary.
For each personality type (A and B), we generated 5 regulated and 5 baseline dialogue sessions. Each session consisted of 6 user–assistant exchanges (i.e., 6 assistant responses paired with 6 user messages), resulting in 120 assistant turns total across the full design (2 profiles × 2 assistant types × 5 sessions × 6 assistant turns). All conversations followed a standardised six-exchange structure balancing sufficient interaction depth for valid personality detection with manageable evaluation complexity.
Each session followed a consistent scenario progression across the six exchanges (opening support offer → initial disclosure → elaboration → reframing → coping/next-step guidance → closing), with user messages generated to reflect the target personality profile and assistants responding either with regulated (personality-adaptive) or baseline (non-adaptive) strategies.
User messages were generated by separate GPT-4 instances (model: gpt-4-0613) instantiated with personality profile prompts ensuring consistent trait expression. Conversations were exported into structured Evaluation Matrices (Excel format). Data collection occurred over 4 weeks (June–July 2024).

3.5. Evaluation Procedure

The evaluation measured effectiveness of personality-regulated assistant interactions versus baselines using a structured evaluation framework, with each assistant output scored on detection accuracy, regulation effectiveness, emotional tone appropriateness, relevance, and fulfilment of personality-specific emotional needs. Regulated assistants were evaluated across five criteria: Detection Accuracy, Regulation Effectiveness, Emotional Tone Appropriateness, Relevance & Coherence, and Personality Needs Addressed.
Baseline assistants were assessed on three criteria only (Emotional Tone Appropriateness, Relevance & Coherence, Personality Needs Addressed), as detection and regulation were not implemented. All criteria used a three-point scoring scale (Yes = 2, Not Sure = 1, No = 0). Figure 7 illustrates the structured evaluation framework used to compare both conditions.

3.5.1. LLM-Based Evaluator System

A specialised Evaluator GPT-4-turbo instance was developed using custom-designed prompts for each criterion, positioning the structured evaluator as an application of the emerging “LLM-as-Judge” paradigm for comparative assessment of generated outputs (Zheng et al., 2023; Kim et al., 2024; Liu et al., 2023). In this study, that design is particularly appropriate because the evaluation task involves many assistant turns scored under a fixed rubric across controlled simulated sessions, while the target qualities—including empathy, reassurance, personality-sensitive appropriateness, and regulation quality—are not well captured by simple overlap-based automatic metrics (Liu et al., 2023). The evaluator therefore provided a scalable and repeatable way to apply the same criteria across all outputs at the proof-of-concept stage before resource-intensive human-rating studies. The evaluator assessed responses one row at a time using full dialogue context and the associated personality profile. Using GPT-4-turbo for evaluation reduces exact same-model reuse relative to the GPT-4-0613 dialogue-generation and detection agents, but does not eliminate potential shared-family bias. To mitigate evaluator bias, regulated and baseline assistants were evaluated separately, the evaluator was informed that detection accuracy could improve incrementally across turns, and the temperature was set to 0.3 to reduce stochastic variation. Accordingly, the evaluator was used here as a feasibility-level comparative instrument under controlled simulation conditions rather than as a source of definitive ground-truth judgment (Kim et al., 2024).

3.5.2. Evaluation Validity and Author-Led Qualitative Review

LLM-generated ratings in the present design are informative for structured comparison, but they are not treated as ground truth (Kim et al., 2024). To address “AI-evaluating-AI” concerns, two contributing authors (Samuel Devdas and Jiahua Duojie) reviewed all evaluation matrices and provided qualitative commentary for each rated item. Their assessments largely aligned with the evaluator-generated detections and criterion judgements across the regulated conversations, with one instance marked as uncertain. These author notes provide an internal audit trail supporting the face validity, plausibility, and interpretability of the automated ratings; they do not constitute impartial verification by independent external raters.
As quantitative ratings are LLM-generated and qualitative review was performed by study authors, the findings should be interpreted as feasibility-level, proof-of-concept comparative evidence under controlled simulation conditions, with attendant risk of confirmation bias. This design therefore offers interpretive support rather than quantitative inter-rater reliability or third-party expert certification. Establishing stronger evaluation credibility in future work requires: (i) at least two independent human raters scoring an overlapping subset of turns to quantify inter-rater reliability (e.g., Krippendorff’s α (Krippendorff, 2011)), and (ii) human–AI agreement analysis for automated scoring (e.g., Cohen’s κ (Landis and Koch, 1977; Gisev et al., 2013)).

3.6. Analysis and Reproducibility

The primary outcome was “Personality Needs Addressed,” measured on a three-point scale (Yes = 2, Not Sure = 1, No = 0). Secondary outcomes included Detection Accuracy, Regulation Effectiveness, Emotional Tone Appropriateness, and Relevance & Coherence. Turn-level ratings were collected for each assistant output (10 dialogue sessions per assistant type × 6 assistant turns = 60 outputs per condition) (Little and Rubin, 2019).

3.6.1. Statistical Analysis

The primary unit of analysis for inferential comparisons between regulated and baseline conditions was the conversation ( n = 10 matched pairs: five Type A and five Type B profiles, each with regulated and baseline sessions). For each conversation and metric, we computed the mean rubric score across six assistant turns and tested paired differences using two-tailed paired t-tests (normally distributed metrics) or Wilcoxon signed-rank tests (non-normal metrics; normality assessed via Shapiro–Wilk tests, α = 0.05). Bias-corrected bootstrap confidence intervals (10,000 iterations) were computed on conversation-level mean differences (Efron and Tibshirani, 1994).
For comparability with prior conversational AI studies, we also report turn-level descriptive statistics ( n = 60 assistant turns per condition) and supplementary independent-samples t-tests and Mann–Whitney U tests in Table 3; because turns within a conversation are not independent, these turn-level tests should be interpreted as descriptive supplements and may be anti-conservative. Effect sizes were calculated using Cohen’s d with pooled standard deviation and interpreted using conventional thresholds (small d = 0.2 , medium d = 0.5 , large d = 0.8 ) (Cohen, 1988). For the primary outcome, we also report Cliff’s δ (probability that a randomly chosen regulated score exceeds a randomly chosen baseline score, rescaled to [ 1 , 1 ] ) as a distribution-free complement when conditions show heavy separation on a discrete scale. Statistical significance was set at α = 0.05. All inferential statistics are reported for comparability with prior conversational AI studies and should not be interpreted as evidence of population-level effects (Wasserstein and Lazar, 2016).

3.6.2. Reproducibility and Computational Environment

All statistical analyses were conducted in Python 3.9+ using scipy (v1.9.0) for statistical tests, numpy (v1.23.0) for effect size calculations, and matplotlib/seaborn for visualisations. Experiments used the GPT-4 API (model: gpt-4-0613, OpenAI Python SDK v1.3.5) with temperature = 0.7 for assistant responses, and GPT-4-turbo (model: gpt-4-turbo) with temperature = 0.3 for evaluator responses. All random operations used fixed seed (Python: 42).
Complete system prompts, regulation templates, simulation transcripts, and statistical analysis code are available at GitHub: https://github.com/jahua/zurich-model-chatbot-materials.
Important limitation: GPT-4 API responses exhibit stochastic variation even at temperature = 0 due to probabilistic sampling (Chen et al., 2023). In this study, the fixed random seed (42) and locked model snapshot (gpt-4-0613) are expected to substantially reduce run-to-run variation, though full determinism cannot be guaranteed by design.

4. Results

This section reports (i) dataset distribution, (ii) technical detection/regulation performance, and (iii) comparative effectiveness of regulated versus baseline assistants, followed by qualitative examples illustrating the observed adaptation pattern. As detailed in Section 3.5, these findings should be interpreted as comparative proof-of-concept evidence under a structured rubric-based evaluation procedure.

4.1. Data Quality and Sample Distribution

The evaluation dataset comprised 10 matched conversation pairs across two extreme personality profiles: Type A (OCEAN: +1, +1, +1, +1, +1) and Type B (OCEAN: 1 , 1 , 1 , 1 , 1 ), with 5 conversations per profile. Each conversation contained 6 user–assistant exchanges, yielding 60 rated assistant turns per condition (regulated: n = 60 ; baseline: n = 60 ).
Figure 8 provides a comprehensive overview of sample characteristics, including sample sizes across conditions (Panel A), the distribution of personality profiles (Panel B), and the 2 × 2 experimental design (Panel C). All evaluation metrics achieved 100% data coverage across both conditions, as expected in a fully automated simulation pipeline. Detection and Regulation metrics apply only to the regulated condition.
The balanced 2 × 2 design (5 conversations × 2 profiles × 2 conditions) and near-complete data support valid paired comparisons at the conversation level. Given the simulation’s boundary-condition design and small number of conversations ( n = 10 pairs), formal power analysis was not performed (Section 3.1.2); large effects were expected under maximal signal conditions.

4.2. Detection and Regulation Performance

Regulated assistants achieved near-perfect performance on core technical requirements. Detection accuracy reached 98.3% (59/60 Yes, 1/60 Not Sure) for extreme simulated profiles, with regulation effectiveness achieving 100% (60/60 successful implementations). These results confirm the system reliably executes the designed detection and regulation logic when tested with these profiles. Near-perfect technical performance establishes that the system can identify personality patterns and adjust behaviour accordingly in the vast majority of cases. The critical question becomes whether this technical capability produces actual improvements in conversational quality.

4.2.1. Personality Vector Consistency

Figure 9 and Figure 10 provide diagnostic evidence of detected personality vector patterns across turns. Analysis of 60 detected personality vectors (one per regulated conversation turn) revealed dimension-specific detection fidelity: Extraversion, Agreeableness (Type A), and Neuroticism showed near-perfect alignment with intended profiles (≥87% correct), while Conscientiousness and Agreeableness (Type B) exhibited more frequent Mid (0) detections, reflecting the detector’s use of a neutral state (0) when conversational evidence for extreme trait expression was insufficient. Most conversations showed within-conversation stability; limited within-conversation variation was observed in a small number of sessions (e.g., Extraversion in conversation A-5, Openness in B-5 (where A/B denotes the personality profile type)), consistent with the cumulative-evidence updating mechanism described in Section 3.3.

4.3. Comparative Effectiveness: Regulated vs. Baseline Performance

Having established technical capability, we address the primary research question: whether adaptive regulation produces measurable improvements in conversational quality compared to non-adaptive baselines.

4.3.1. Primary Outcome: Personality Needs Addressed

Table 3 presents turn-level descriptive statistics and supplementary between-group tests; primary inferential comparisons used conversation-level paired tests (Section 3.6.1).
Regulated assistants achieved perfect performance on Personality Needs Addressed ( M = 2.00 , S D = 0.00 ; 100% Yes ratings) compared to baseline’s near-complete failure ( M = 0.20 , S D = 0.58 ; 8.3% Yes ratings), yielding an exceptionally large parametric effect size (Cohen’s d = 4.42 , 95% CI [3.1, 8.1], p < 0.001 ) and an equally extreme non-parametric separation (Cliff’s δ = 0.917). These magnitudes should be interpreted chiefly as a binary capability difference—the regulated stack implements personality-specific regulation on this rubric, whereas the baseline does not—rather than as a graded increment along a continuous intervention scale. Figure 11 visualises this selective enhancement pattern; a fuller interpretation appears in Section 4.3.3.

4.3.2. Secondary Outcomes: Basic Conversational Quality (Ceiling Effects)

Both conditions achieved ceiling performance on Emotional Tone Appropriateness (both M = 2.00 , S D = 0.00 ; 100% Yes) and near-ceiling on Relevance & Coherence (regulated M = 2.00 vs. baseline M = 1.97 ), with negligible effect sizes ( d = 0.000 and d = 0.183 , respectively, both not significant). Emotional Tone Appropriateness was operationalised as whether the assistant’s response conveyed an appropriate affective stance (warmth, empathy, non-judgemental tone) as judged by the LLM evaluator on the Yes/Not Sure/No scale; Relevance & Coherence assessed whether the response was topically relevant and logically consistent with the dialogue history. Both metrics capture baseline conversational quality that a capable LLM should deliver regardless of personality adaptation—the ceiling performance in both conditions therefore indicates that the underlying GPT-4 model already provides high-quality generic support, and that personality adaptation neither helps nor hurts on these dimensions. The near-zero effects reflect ceiling performance in both conditions rather than lack of regulatory benefit, demonstrating that personality adaptation adds targeted value to an already excellent conversational foundation without compromising generic quality. Figure 12 shows the total score distribution across conditions.

4.3.3. Interpretation of the Selective Enhancement Pattern

The exceptionally large effect sizes observed for personality needs (Cohen’s d = 4.42 ; Cliff’s δ = 0.917) represent a binary shift in system capability rather than an incremental improvement in quality. Cliff’s δ near the upper bound indicates near-complete stochastic dominance of regulated over baseline scores on this item but, like d, is expected under a two-condition comparison where one architecture includes the target capability and the other omits it. This magnitude is a direct result of the baseline system’s systematic inability to address personality-specific requirements ( 8.3 % success rate) compared to the regulated assistant’s complete success ( 100 % success rate). In the regulated condition, the absence of variance ( S D = 0.00 ) reflects a ceiling effect where the architecture consistently executed the theory-aligned adaptations across all conversation turns. Conversely, the baseline condition’s failure confirms that standard large language models, while highly competent in general dialogue, do not inherently adapt their behaviour to specific OCEAN profiles without explicit regulatory instructions.
Crucially, this improvement occurred as a “selective enhancement,” meaning that the dramatic gains in personalisation did not come at the cost of fundamental conversational quality. Both assistants maintained ceiling or near-ceiling performance in emotional tone and relevance ( d < 0.2 , non-significant), indicating that the baseline system already provided a high-quality foundation. The adaptive layer thus acts as an additive module that precisely targets individual motivational domains—Security, Arousal, and Affiliation—without interfering with the model’s established linguistic and emotional competence. While these effect sizes were obtained under controlled simulation conditions with extreme profiles, they validate the architectural feasibility of the D–R–E pipeline. Figure 13 and Figure 14 visualise this selective enhancement pattern through paired conversation-level analysis and complete rating distribution composition, respectively.

4.3.4. Statistical Robustness and Implementation Fidelity

The statistical significance of the primary outcome was confirmed at the conversation level using paired t-tests and Wilcoxon signed-rank tests ( n = 10 matched pairs), with identical directional conclusions; bias-corrected bootstrap confidence intervals ( 10 , 000 iterations) further supported robustness. Turn-level independent-samples comparisons in Table 3 yielded the same pattern and are reported for descriptive completeness. These analyses indicate that the difference in addressing personality needs is not a statistical artefact but a systematic outcome of the regulation logic.
Implementation fidelity was similarly high across the 120 analysed turns. The detection module correctly identified the intended extreme profiles in 98.3% of cases (59/60 turns; 1 rated Not Sure), and the regulation module successfully applied the corresponding behavioural adjustments in all turns (60/60; 100% implementation fidelity). This provides a consistent implementation pathway linking theory-driven prompts to the observed conversational adaptations. While real-world deployment with moderate or ambiguous profiles would likely result in smaller effect sizes (estimated at d = 0.5 1.5 ) due to potential detection errors and individual variability, the current results provide an empirical upper bound for the system’s performance under idealised conditions.

4.4. Qualitative Examples Demonstrating Personality Adaptation

Complementing the quantitative results, Figure 15 and Figure 16 illustrate how the regulated and baseline assistants respond to the same user turn under two contrasting personality profiles. Each figure shows one user reply and two assistant responses: the regulated response (left) is conditioned on the detected OCEAN profile and the applied regulation prompt, whereas the baseline response (right) reflects a high-quality but non-personalised conversational strategy. The excerpts are intended to make the “selective enhancement” pattern visible at the dialogue level—i.e., differences emerge primarily in how support is framed and structured to match personality-linked motivational needs, not in basic coherence or politeness.
Across both examples, the qualitative contrast mirrors the quantitative finding that regulation mainly affects personalisation-sensitive aspects of the interaction (framing, pacing, structure, and engagement strategy). Baseline responses can remain highly appropriate on generic quality dimensions, yet still fail the personality-needs criterion because they do not reliably implement a coherent, theory-driven adaptation policy. The figures thus provide an interpretable mechanism for the large improvement in personality-needs ratings observed in Section 4.3: the D–R–E pipeline enables consistent, auditable style regulation grounded in Zurich Model domains (Section 3.3.1), while preserving the baseline model’s already-strong conversational competence.

5. Discussion

This simulation study evaluated a personality-adaptive framework that combines Big Five detection with Zurich Model-driven regulation. Under controlled conditions, the system achieved near-complete implementation fidelity (98.3% detection accuracy, 100% regulation adherence) and exhibited a selective enhancement pattern: near-complete personality need fulfilment under regulation versus minimal baseline success on the same rubric ( d = 4.42 ; Cliff’s δ = 0.917), reflecting a binary architectural contrast rather than a graded clinical treatment effect, while maintaining ceiling or near-ceiling performance on generic conversational quality (emotional tone: d = 0.000 ; relevance: d = 0.183 ). Because quantitative ratings were produced by a structured LLM-as-Judge evaluator and qualitative review was conducted by two contributing authors (not independent third-party raters), these results should be interpreted as feasibility-level, proof-of-concept evidence for the architecture and theory-grounded mapping under controlled simulation conditions, rather than as definitive evidence of human-perceived clinical benefit.
The selective enhancement pattern is more informative than a simple “regulated outperforms baseline” statement for three reasons. First, ceiling effects in emotional tone and coherence indicate that the baseline assistant already provides high-quality generic support; the evaluation is therefore not correcting poor baseline behaviour. Second, the regulated condition adds a distinct capability layer—trait-conditioned response framing—without degrading the baseline’s generic strengths, supporting an AI augmentation model where personality adaptation complements and enhances existing conversational capabilities rather than replacing them. Third, the selectivity of the effect is theoretically coherent: Zurich Model domains (Security, Arousal, Affiliation; Section 3.3.1) are expected to influence personalisation-sensitive dimensions (personality needs) more strongly than general linguistic adequacy.
The magnitude of d = 4.42 should be interpreted as a binary capability difference in this controlled simulation: baseline rarely satisfies the personality-needs criterion (8.3% Yes), whereas regulation yields complete success (100% Yes), producing a very large standardised difference under minimal variance in the regulated condition. In real-world settings, effect sizes are expected to be smaller due to mixed personality profiles, noisier personality inference, longer and less scripted interactions, and individual variability.
These findings are consistent with prior work indicating that personality-aware conversational agents can improve user-aligned interaction outcomes relative to generic approaches (Mairesse and Walker, 2011; Zhou, 2021). In contrast to methods that rely on static profiles or ad hoc heuristics, our framework operationalises a continuous inference-and-regulation loop in which trait estimates are updated from dialogue and mapped to behavioural strategies grounded in the Zurich Model (Section 3.3.1). At the same time, the present study emphasizes evaluation methodology as a key open challenge. LLM-as-Judge scoring enables scalable and structured comparative assessment, but it does not substitute for independent multi-rater human validation. Establishing convergent validity against human judgements and linking conversational metrics to downstream outcomes are necessary steps for translation beyond simulation.

5.1. Research and Design Implications

The results suggest that trait-based regulation provides a complementary capability rather than universally superior conversational performance. Because baseline responses already reach ceiling levels on emotional tone and coherence, the primary room for improvement lies in personalisation—i.e., systematically aligning response style with user-specific motivational needs. This shifts the research frontier from generic “supportiveness” toward controllable adaptation mechanisms that can be layered onto strong base models.
From a theory perspective, the selectivity of the effect provides preliminary support for the Zurich Model framing: personality-adaptive regulation primarily influences domains where individual differences matter (e.g., Security/Arousal/Affiliation-relevant response framing; Section 3.3.1), rather than improving general linguistic competence. The modular D–R–E design also supports targeted iteration: detection and regulation can be improved or replaced independently as long as interfaces and logging remain stable, reducing the risk that personalisation compromises baseline quality.
For practice-oriented systems, the findings suggest the following design principles:
  • Layered deployment: Start from a strong generic baseline and treat personality adaptation as an additive capability; AI augments rather than replaces generic conversational competence.
  • Explicit, updateable personality state: Maintain persistent trait estimates (with uncertainty) updated from conversational evidence.
  • Theory-grounded mapping: Translate traits to behavioural strategies using established personality–motivation frameworks (e.g., the Zurich Model), rather than ad hoc heuristics.
  • Transparent adaptation logic: Log trait estimates, applied regulations, and evaluation outcomes for auditability and debugging.
  • Selective application: Apply adaptation primarily to interactional style and support preferences, not to all content indiscriminately.
For developers of AI-driven digital mental health interventions, these findings suggest prioritizing transparent adaptation logic where personality inferences and regulatory decisions are logged and auditable. As the field of technology-enhanced mental health support matures, ethical safeguards around data privacy, bias mitigation, and human oversight become critical enablers for responsible integration. Future systems should embed mechanisms for uncertainty quantification, allowing interventions to gracefully handle ambiguous personality signals rather than making overconfident inferences. Aligning AI adaptation with patient-centred outcomes (Simon, 2011) remains a key criterion for translating simulation findings into responsible clinical practice. These considerations are especially important given concerns about algorithmic bias, data privacy risks, and the need for equitable access across diverse user populations.

5.2. Strengths, Limitations, and Validation Pathway

The modular D–R–E architecture demonstrated several strengths that support its potential for real-world application. First, architectural modularity enables independent improvement of detection and regulation components without requiring system-wide redesign. The PROMISE-based state machine orchestration provides transparent, auditable adaptation logic, facilitating debugging and compliance with ethical oversight requirements. Second, the scalable evaluation framework combining structured LLM-as-Judge assessment with author-conducted qualitative review offers a pragmatic balance between systematic measurement and interpretive validation at the proof-of-concept stage. Third, the framework is adaptable to alternative personality theories beyond the Big Five and to support domains beyond emotional regulation, suggesting potential for broader application across digital mental health contexts.
From an implementation perspective, these findings demonstrate that AI-augmented personalisation can be layered onto strong baseline conversational models without degrading core quality. The selective enhancement pattern—dramatic gains in personalisation-sensitive dimensions with preserved generic competence—supports incremental deployment strategies where personality adaptation serves as an optional augmentation layer rather than a replacement for established conversational AI capabilities.

5.2.1. Limitations and Barriers

This study has several important limitations. All interactions were simulated between predefined personality profiles and GPT-4-based assistants; no human users participated, and results cannot be generalised to real-world user populations. Profiles were restricted to extreme configurations (all traits high vs. all traits low), not reflecting typical personality distributions. How the framework generalises to moderate or mixed profiles remains an open question: with less extreme trait signals, detection confidence may be lower (more frequent neutral-state assignments), regulation instructions may be less specific, and effect sizes on personality-needs fulfilment are expected to be smaller. Moderate profiles are the priority target for Stage 1 validation (Section 5.2.2). Dialogues were short (six turns per session), unable to capture long-term interaction dynamics. Both system and evaluator rely on the same LLM family, potentially introducing shared biases despite broad alignment between LLM ratings and author review notes in qualitative validation. The work focused on English-language text and Western communication norms; cultural and linguistic generalisability remains unknown. Finally, quantitative ratings were generated by an LLM-based evaluator with author-conducted qualitative review, rather than independent multi-rater human validation with inter-rater reliability analysis; author-led review is not a substitute for blinded or arm’s-length raters.
Beyond technical limitations, several barriers must be addressed for real-world deployment. Privacy considerations are paramount: personality profiling based on conversational data raises questions about user consent, data storage, and potential misuse of inferred trait information. Bias risks include cultural assumptions embedded in Big Five trait models, potential stereotyping based on personality inferences, and unequal performance across demographic groups. Integration challenges involve adapting the system to diverse user needs, ensuring graceful degradation when personality signals are ambiguous, and maintaining performance across varied conversational contexts beyond emotional support scenarios.

5.2.2. Validation Pathway for AI-Augmented Personalisation

Future research should adopt a phased validation approach to systematically address these limitations and barriers.
Stage 1: Extended Proof-of-Concept Validation
  • Test moderate and mixed personality profiles reflecting realistic trait distributions
  • Extend dialogue length to 20+ exchanges to assess long-term adaptation stability
  • Conduct cross-linguistic and cross-cultural validation studies to assess generalisability beyond English and Western norms
Stage 2: Human Subject Validation ( n = 50 –100 per condition)
  • IRB-approved study protocol with comprehensive informed consent addressing AI-based personality profiling and data usage
  • User-reported outcome measures: engagement, satisfaction, perceived support quality, and therapeutic alliance
  • Safety monitoring infrastructure with clear protocols for identifying and responding to user distress
  • Qualitative interviews to understand user experiences with personality-adaptive versus non-adaptive systems
Stage 3: Real-World Deployment Evaluation
  • Multi-site implementation studies across diverse user populations
  • Integration with existing digital mental health platforms to assess workflow compatibility
  • Equity assessments examining performance across demographic groups (age, gender, cultural background, education level)
  • Longitudinal tracking to evaluate sustained engagement and long-term user outcomes
This phased approach prioritizes empirical validation with human participants while maintaining ethical safeguards, moving systematically from controlled validation to ecologically valid deployment contexts.

6. Conclusions

This simulation study demonstrates that integrating Big Five personality detection with Zurich Model-driven regulation in LLM-based conversational agents is technically feasible, achieving near-perfect implementation fidelity (98.3% detection accuracy, 100% regulation adherence) and a 92-percentage-point improvement in personality-specific need fulfilment (Cohen’s d = 4.42 , p < 0.001 ). The critical finding is selective enhancement: dramatic gains on personalisation-sensitive metrics with no degradation in basic conversational quality (emotional tone: d = 0.000 ; relevance: d = 0.183 , ns), confirming that theory-driven adaptation targets individual differences rather than generic interaction quality.
Results represent an upper bound under idealised conditions (boundary-condition profiles, pre-scripted dialogues, synthetic data). Real-world performance will likely be substantially lower (estimated d = 0.5 1.5 ). Human subject validation with moderate personality profiles, longitudinal tracking, user-reported outcomes, and bias audits is essential before real-world deployment. The detection → regulation → evaluation framework serves as an augmentation layer for existing digital mental health interventions, adaptable to alternative personality theories and support domains, providing a methodological template for AI-augmented personalised mental health support grounded in psychological theory.

Supplementary Materials

The following supporting information can be downloaded from the project repository: https://github.com/jahua/zurich-model-chatbot-materials:
  • Supplementary File S1: Complete system prompts for personality detection (5 Big Five trait detectors)
  • Supplementary File S2: Regulation templates for Zurich Model mapping (arousal, security, affiliation)
  • Supplementary File S3: Evaluator GPT system prompt and scoring matrix
  • Supplementary File S4: Complete simulation transcripts (20 conversations, 120 dialogue turns)
  • Supplementary File S5: Statistical analysis code (Python/Jupyter notebooks for effect sizes, paired tests, weighted scoring, and visualisations)
  • Supplementary File S6: Missingness comparison plot (horizontal bars showing <5% missing data across conditions)
  • Supplementary File S7: Personality needs YES-rate by conversation (demonstrating consistent improvement across all 10 conversation pairs)
  • Supplementary File S8: Rating distribution raw counts (YES/NOT SURE/NO composition with value annotations)

Author Contributions

The author contributions follow the CRediT taxonomy. Conceptualization: S. Devdas, G. Lu, A. de Spindler and M. Stieger; methodology: S. Devdas, J. Duojie, G. Lu, A. de Spindler and M. Stieger; software: S. Devdas and A. de Spindler; validation: J. Duojie and S. Devdas; formal analysis: J. Duojie; data curation: J. Duojie and S. Devdas; writing—original draft preparation: J. Duojie; writing—review and editing: M. Stieger, G. Lu and A. de Spindler; visualisation: J. Duojie; project administration: J. Duojie; supervision: G. Lu and M. Stieger. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable. This study involved no human or animal subjects and used synthetic, non-identifiable, AI-generated dialogues. Contributing authors (Samuel Devdas and Jiahua Duojie) provided qualitative evaluation notes on these synthetic materials; no undisclosed third-party raters and no patient data were involved. Future human subject studies will require full institutional review board approval.

Data Availability Statement

Supplementary materials including system prompts, regulation templates, simulation transcripts, and statistical analysis code are available at https://github.com/jahua/zurich-model-chatbot-materials

Acknowledgments

The authors acknowledge the PROMISE framework developers (Wu et al., 2024) whose open architecture informed our system design.

Conflicts of Interest

The authors declare no conflicts of interest. Qualitative review of evaluation matrices was performed by contributing authors (Samuel Devdas and Jiahua Duojie); no external paid evaluator or undisclosed third-party rater contributed to this study.

References

  1. Dingler, T., D. Kwasnicka, J. Wei, E. Gong, and B. Oldenburg. 2021. The use and promise of conversational agents in digital health. Yearbook of Medical Informatics 30: 191–199. [Google Scholar] [CrossRef]
  2. Laranjo, L., A.G. Dunn, H.L. Tong, A.B. Kocaballi, J. Chen, R. Bashir, D. Surian, B. Gallego, F. Magrabi, and A.Y. Lau. 2018. Conversational agents in healthcare: a systematic review. J. of the American Medical Informatics Association 25: 1248–1258. [Google Scholar] [CrossRef]
  3. Vasiliu, L., K. Cortis, R. McDermott, A. Kerr, A. Peters, M. Hesse, J. Hagemeyer, T. Belpaeme, J. McDonald, and R. Villing. 2021. CASIE – Computing affect and social intelligence for healthcare in an ethical and trustworthy manner. Paladyn, J. of Behavioral Robotics 12: 437–453. [Google Scholar] [CrossRef]
  4. Kocaballi, A.B., E. Sezgin, L. Clark, J.M. Carroll, and et al. 2022. Design and Evaluation Challenges of Conversational Agents in Health Care and Well-being: Selective Review Study. J. Med. Internet Res. 24: e38525. [Google Scholar] [CrossRef] [PubMed]
  5. Kaddour, J., J. Harris, M. Mozes, H. Bradley, R. Raileanu, and R. McHardy. 2023. Challenges and Applications of Large Language Models. arXiv arXiv:2307.10169. [Google Scholar] [CrossRef]
  6. Rapp, A., L. Curti, and A. Boldi. 2021. The human side of human-chatbot interaction: A systematic literature review of ten years of research on text-based chatbots. International Journal of Human-Computer Studies 151: 102630. [Google Scholar] [CrossRef]
  7. Han, E., D. Yin, and H. Zhang. 2022. Chatbot Empathy in Customer Service: When It Works and When It Backfires. Proceedings of the SIGHCI 2022 Proceedings, Vol. 1. [Google Scholar]
  8. Juquelier, A., I. Poncin, and S. Hazée. 2025. Empathic chatbots: A double-edged sword in customer experiences. Journal of Business Research 188: 115074. [Google Scholar] [CrossRef]
  9. Seitz, L. 2024. Artificial empathy in healthcare chatbots: Does it feel authentic? Computers in Human Behavior: Artificial Humans 2: 100067. [Google Scholar] [CrossRef]
  10. Abernethy, A., L. Adams, M. Barrett, C. Bechtel, P. Brennan, A. Butte, J. Faulkner, E. Fontaine, S. Friedhoff, J. Halamka, and et al. 2022. The Promise of Digital Health: Then, Now, and the Future. NAM Perspectives.
  11. Goetz, L.H., and N.J. Schork. 2018. Personalized medicine: Motivation, challenges, and progress. Fertility and Sterility 109: 952–963. [Google Scholar] [CrossRef]
  12. O’Keefe, D.J. 2016. Persuasion: Theory and Research. SAGE Publications, Inc. [Google Scholar]
  13. Milne-Ives, M., C. de Cock, E. Lim, M.H. Shehadeh, N. de Pennington, G. Mole, E. Normando, and E. Meinert. 2020. The Effectiveness of Artificial Intelligence Conversational Agents in Health Care: Systematic Review. J. Med. Internet Res. 22: e20346. [Google Scholar] [CrossRef]
  14. Wei, J., X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E.H. Chi, Q.V. Le, and D. Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. Proceedings of the Proceedings of the 36th International Conference on Neural Information Processing Systems, Red Hook, NY, USA; p. NIPS ’22. [Google Scholar]
  15. Ayers, J.W., A. Poliak, M. Dredze, E.C. Leas, Z. Zhu, J.B. Kelley, D.J. Faix, A.M. Goodman, C.A. Longhurst, M. Hogarth, and et al. 2023. Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum. JAMA Internal Medicine 183: 589–596. [Google Scholar] [CrossRef] [PubMed]
  16. Färber, A., A. de Spindler, A. Moser, and G. Schwabe. 2023. Closing the Loop for Patients with Chronic Diseases - from Problems to a Solution Architecture. Proceedings of the The 11th IEEE International Conference on Healthcare Informatics; pp. 1–11. [Google Scholar] [CrossRef]
  17. Färber, A., C. Schwabe, P.H. Stalder, M. Dolata, and G. Schwabe. 2024. Physicians’ and Patients’ Expectations From Digital Agents for Consultations: Interview Study Among Physicians and Patients. JMIR Human Factors 11: e49647. [Google Scholar] [CrossRef]
  18. Staehelin, D., M. Dolata, L. Stöckli, and G. Schwabe. 2024. How Patient-Generated Data Enhance Patient-Provider Communication in Chronic Care: Field Study in Design Science Research. JMIR Medical Informatics 12: e57406. [Google Scholar] [CrossRef] [PubMed]
  19. Strubell, E., A. Ganesh, and A. McCallum. 2019. Edited by A. Korhonen, D. Traum and L. Màrquez. Energy and Policy Considerations for Deep Learning in NLP. In Proceedings of the Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics: pp. 3645–3650. [Google Scholar]
  20. Ding, N., Y. Qin, G. Yang, F. Wei, Z. Yang, Y. Su, S. Hu, Y. Chen, C.M. Chan, W. Chen, and et al. 2023. Parameter-efficient fine-tuning of large-scale pre-trained language models. Nature Machine Intelligence 5: 220–235. [Google Scholar] [CrossRef]
  21. Hu, E.J., Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen. 2021. LoRA: Low-Rank Adaptation of Large Language Models. arXiv arXiv:2106.09685. [Google Scholar]
  22. Korzynski, P., G. Mazurek, P. Krzypkowska, and A. Kurasinski. 2023. Artificial intelligence prompt engineering as a new digital competence: Analysis of generative AI technologies such as ChatGPT. Entrepreneurial Business and Economics Review.
  23. White, J., Q. Fu, S. Hays, M. Sandborn, C. Olea, H. Gilbert, A. Elnashar, J. Spencer-Smith, and D. Schmidt. 2023. A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT. arXiv arXiv:2302.11382. [Google Scholar] [CrossRef]
  24. Fernando, C., D. Banarse, H. Michalewski, and et al. 2023. Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution. arXiv arXiv:2309.16797. [Google Scholar]
  25. Liu, P., W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig. 2023. Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. ACM Computing Surveys 55: 195:1–195:35. [Google Scholar] [CrossRef]
  26. Hou, Y., H. Dong, X. Wang, B. Li, and W. Che. 2022. Edited by N. Calzolari, C.R. Huang, H. Kim, J. Pustejovsky, L. Wanner, K.S. Choi, P.M. Ryu, H.H. Chen, L. Donatelli, H. Ji and et al. MetaPrompting: Learning to Learn Better Prompts. In Proceedings of the Proceedings of the 29th International Conference on Computational Linguistics. International Committee on Computational Linguistics: pp. 3251–3262. [Google Scholar]
  27. Wu, W., J. Heierli, M. Meisterhans, A. Moser, A. Färber, M. Dolata, E. Gavagnin, A. de Spindler, and G. Schwabe. 2024. Edited by S. Islam and A. Sturm. PROMISE: A Framework for Model-Driven Stateful Prompt Orchestration. In Proceedings of the Intelligent Information Systems: CAiSE Forum 2024. Limassol, Cyprus, Springer International Publishing: <i>Lecture Notes in Business Information Processing</i>, June 3–7, Vol. 520. [Google Scholar]
  28. Alisamir, S., and F. Ringeval. 2021. On the Evolution of Speech Representations for Affective Computing: A Brief History and Critical Overview. IEEE Signal Process. Mag. 38: 12–21. [Google Scholar] [CrossRef]
  29. Alsharekh, M. 2022. Facial Emotion Recognition in Verbal Communication Based on Deep Learning. Electronics 11: 2568. [Google Scholar] [CrossRef]
  30. Mairesse, F., and M. Walker. 2011. Controlling User Perceptions of Linguistic Style: Trainable Generation of Personality Traits. Comput. Linguist. 37: 455–488. [Google Scholar] [CrossRef]
  31. Ta, V., C. Griffith, C. Boatfield, X. Wang, M. Civitello, and H. Maffei. 2020. User Experiences of Social Support from Companion Chatbots in Everyday Contexts: Thematic Analysis. J. Med. Internet Res. 22: e16235. [Google Scholar] [CrossRef]
  32. Broadbent, E. 2024. ElliQ, an AI-Driven Social Robot to Alleviate Loneliness: Progress and Lessons Learned. JAR Life 13: 22–28. [Google Scholar] [CrossRef]
  33. Shah, S. 2019. Effectiveness of Digital Technology Interventions to Reduce Loneliness in Adults: A Protocol for a Systematic Review and Meta-Analysis. BMJ Open 9: e029324. [Google Scholar] [CrossRef]
  34. Shah, S. 2021. Evaluation of the Effectiveness of Digital Technology Interventions to Reduce Loneliness in Older Adults: Systematic Review and Meta-Analysis. J. Med. Internet Res. 23: e24712. [Google Scholar] [CrossRef] [PubMed]
  35. Zhou, M. 2021. Designing Effective Interview Chatbots: Automatic Chatbot Profiling and Design Suggestion Generation. Proceedings of the Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’21). [Google Scholar]
  36. Quirin, M., F. Malekzad, D. Paudel, A. Knoll, and M. Mirolli. 2023. Dynamics of Personality: The Zurich Model of Motivation Revived, Extended, and Applied to Personality. J. Pers. 91: 928–946. [Google Scholar] [CrossRef]
  37. Berlyne, D. 1960. Conflict, Arousal, and Curiosity. McGraw-Hill: New York, NY, USA. [Google Scholar]
  38. Hebb, D. 1949. The Organization of Behavior: A Neuropsychological Theory. Wiley: New York, NY, USA. [Google Scholar]
  39. Bischof, N. 1985. Das Rätsel Ödipus: Die biologischen Wurzeln des Urkonfliktes. Piper: Munich, Germany. [Google Scholar]
  40. Bischof, N. 1993. Untersuchungen zur Systemanalyse der sozialen Motivation I: Die Tantalus-Situation. Z. Psychol. 201: 5–43. [Google Scholar]
  41. Bickmore, T., and R. Picard. 2005. Establishing and Maintaining Long-Term Human-Computer Relationships. ACM Trans. Comput.-Hum. Interact. 12: 293–327. [Google Scholar] [CrossRef]
  42. Zheng, Z., L. Liao, Y. Deng, and L. Nie. 2023. Building Emotional Support Chatbots in the Era of LLMs. arXiv arXiv:2308.11584. [Google Scholar] [CrossRef]
  43. Chen, K., X. Kang, X. Lai, and Z. Ni. 2023. Enhancing Emotional Support Capabilities of Large Language Models through Cascaded Neural Networks. Proceedings of the Proceedings of the 2023 4th International Conference on Computer, Big Data and Artificial Intelligence (ICCBD+AI); pp. 318–326. [Google Scholar]
  44. Ouyang, L., J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, and A. Ray. 2022. Training Language Models to Follow Instructions with Human Feedback. Proceedings of the Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS ’22); pp. 27730–27744. [Google Scholar]
  45. Dong, T., F. Liu, X. Wang, Y. Jiang, X. Zhang, and X. Sun. 2024. EmoAda: A Multimodal Emotion Interaction and Psychological Adaptation System. Proceedings of the Proceedings of the Conference on Multimedia Modeling, Cham, Switzerland. [Google Scholar]
  46. Abbasian, M., I. Azimi, A. Rahmani, and R. Jain. 2023. Conversational Health Agents: A Personalized LLM-Powered Agent Framework. arXiv arXiv:2310.02293. [Google Scholar]
  47. Sorino, P. 2024. ARIEL: Brain-Computer Interfaces Meet Large Language Models for Emotional Support Conversation. Proceedings of the Adjunct Proceedings of the 32nd ACM Conference on User Modeling, Adaptation and Personalization (UMAP ’24). [Google Scholar]
  48. Dongre, P. 2024. Physiology-Driven Empathic Large Language Models (EmLLMs) for Mental Health Support. Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (CHI EA ’24). [Google Scholar]
  49. Zhang, H., Y. Chen, M. Wang, and S. Feng. 2024. FEEL: A Framework for Evaluating Emotional Support Capability with Large Language Models. arXiv arXiv:2403.15699. [Google Scholar] [CrossRef]
  50. Zheng, L., W. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, and E. Xing. 2023. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. arXiv arXiv:2306.05685. [Google Scholar]
  51. Kim, S., S. Lee, J. Kim, H. Yoo, Y. Seol, S. Kim, H. Cho, S. Choi, and S. Kim. 2024. Evaluating LLM-as-a-judge with Professional Human Ratings. arXiv arXiv:2402.18139. [Google Scholar]
  52. Montgomery, D. 2017. Design and Analysis of Experiments, 9th ed. ed. Wiley: Hoboken, NJ, USA. [Google Scholar]
  53. Collins, L., J. Dziak, and R. Li. 2009. Design of Experiments with Multiple Independent Variables: A Resource Management Perspective on Complete and Fractional Factorial Designs. Psychol. Methods 14: 202–224. [Google Scholar] [CrossRef] [PubMed]
  54. Carroll, C., M. Patterson, S. Wood, A. Booth, J. Rick, and S. Balain. 2007. A Conceptual Framework for Implementation Fidelity. Implement. Sci. 2: 40. [Google Scholar] [CrossRef] [PubMed]
  55. Faul, F., E. Erdfelder, A. Buchner, and A. Lang. 2009. Statistical Power Analyses Using G*Power 3.1: Tests for Correlation and Regression Analyses. Behav. Res. Methods 41: 1149–1160. [Google Scholar] [CrossRef]
  56. Schulz, K., D. Altman, D. Moher, and C. Group. 2010. CONSORT 2010 Statement: Updated Guidelines for Reporting Parallel Group Randomised Trials. BMJ 340: c332. [Google Scholar] [CrossRef]
  57. Chan, A., J. Tetzlaff, P. Gøtzsche, D. Altman, H. Mann, J. Berlin, K. Dickersin, A. Hróbjartsson, K. Schulz, and W. Parulekar. 2013. SPIRIT 2013 Statement: Defining Standard Protocol Items for Clinical Trials. Ann. Intern. Med. 158: 200–207. [Google Scholar] [CrossRef] [PubMed]
  58. Association, W.M. 2013. World Medical Association Declaration of Helsinki: Ethical Principles for Medical Research Involving Human Subjects. JAMA 310: 2191–2194. [Google Scholar]
  59. D’Alfonso, S. 2020. AI in Mental Health. Curr. Opin. Psychol. 36: 112–117. [Google Scholar] [CrossRef]
  60. Torous, J., K. Myrick, N. Rauseo-Ricupero, and J. Firth. 2020. Digital Mental Health and COVID-19: Using Technology Today to Accelerate the Curve on Access and Quality Tomorrow. JMIR Ment. Health 7: e18848. [Google Scholar] [CrossRef]
  61. Alberts, L., G. Keeling, and A. McCroskery. 2024. Should Agentic Conversational AI Change How We Think about Ethics? Characterising an Interactional Ethics Centred on Respect. arXiv arXiv:2401.09187. [Google Scholar] [CrossRef]
  62. Costa, P., and R. McCrae. 1992. Revised NEO Personality Inventory (NEO-PI-R) and NEO Five-Factor Inventory (NEO-FFI): Professional Manual. Psychological Assessment Resources: Odessa, FL, USA. [Google Scholar]
  63. John, O., L. Naumann, and C. Soto. 2008. Paradigm Shift to the Integrative Big Five Trait Taxonomy: History, Measurement, and Conceptual Issues. Proceedings of the Handbook of Personality: Theory and Research, New York, NY, USA; pp. 114–158. [Google Scholar]
  64. McCrae, R., and O. John. 1992. An Introduction to the Five-Factor Model and Its Applications. J. Pers. 60: 175–215. [Google Scholar] [CrossRef] [PubMed]
  65. Gelman, A., J. Carlin, H. Stern, D. Dunson, A. Vehtari, and D. Rubin. 2013. Bayesian Data Analysis, 3rd ed. ed. CRC Press: Boca Raton, FL, USA. [Google Scholar]
  66. Funder, D. 1995. On the Accuracy of Personality Judgment: A Realistic Approach. Psychol. Rev. 102: 652–670. [Google Scholar] [CrossRef] [PubMed]
  67. Vazire, S. 2010. Who Knows What about a Person? The Self–Other Knowledge Asymmetry (SOKA) Model. J. Pers. Soc. Psychol. 99: 281–303. [Google Scholar] [CrossRef] [PubMed]
  68. OpenAI. 2023. GPT-4 Technical Report, 2303.08774.
  69. Brown, T., B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, and A. Askell. 2020. Language Models are Few-Shot Learners. Proceedings of the Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS ’20); pp. 1877–1901. [Google Scholar]
  70. Reynolds, L., and K. McDonell. 2021. Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm. Proceedings of the Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems; pp. 1–7. [Google Scholar]
  71. Liu, P., W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig. 2023. Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. ACM Comput. Surv. 55: 1–35. [Google Scholar] [CrossRef]
  72. APA. 2010. Practice Guideline for the Treatment of Patients with Major Depressive Disorder, 3rd ed. ed. American Psychiatric Association: Arlington, VA, USA. [Google Scholar]
  73. McCrae, R., and P.T.J. Costa. 2003. Personality in Adulthood: A Five-Factor Theory Perspective, 2nd ed. ed. Guilford Press: New York, NY, USA. [Google Scholar]
  74. Goldberg, L. 1993. The Structure of Phenotypic Personality Traits. Am. Psychol. 48: 26–34. [Google Scholar] [CrossRef]
  75. McCrae, R., and P.T.J. Costa. 1992. Revised NEO Personality Inventory (NEO PI-R) and NEO Five-Factor Inventory (NEO-FFI) Professional Manual. Psychological Assessment Resources: Odessa, FL, USA. [Google Scholar]
  76. Liu, Y., D. Iter, Y. Xu, S. Wang, R. Xu, and C. Zhu. 2023. G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment. arXiv arXiv:2303.16634. [Google Scholar] [CrossRef]
  77. Krippendorff, K. 2011. Computing Krippendorff’s Alpha-Reliability. Departmental Papers (ASC), University of Pennsylvania: Philadelphia, PA, USA. [Google Scholar]
  78. Landis, J., and G. Koch. 1977. The Measurement of Observer Agreement for Categorical Data. Biometrics 33: 159–174. [Google Scholar] [CrossRef]
  79. Gisev, N., J. Bell, and T. Chen. 2013. Interrater Reliability and Agreement in Clinical Research: A Guide to Best Practice. Res. Social Adm. Pharm. 9: 330–337. [Google Scholar] [CrossRef]
  80. Little, R., and D. Rubin. 2019. Statistical Analysis with Missing Data, 3rd ed. ed. Wiley: Hoboken, NJ, USA. [Google Scholar]
  81. Efron, B., and R. Tibshirani. 1994. An Introduction to the Bootstrap. Chapman and Hall/CRC: New York, NY, USA. [Google Scholar]
  82. Cohen, J. 1988. Statistical Power Analysis for the Behavioral Sciences, 2nd ed. ed. Lawrence Erlbaum Associates: Hillsdale, NJ, USA. [Google Scholar]
  83. Wasserstein, R., and N. Lazar. 2016. The ASA’s Statement on p-Values: Context, Process, and Purpose. Am. Stat. 70: 129–133. [Google Scholar] [CrossRef]
  84. Chen, L., M. Zaharia, and J. Zou. 2023. How is ChatGPT’s Behavior Changing over Time. arXiv arXiv:2307.09009. [Google Scholar] [CrossRef]
  85. Simon, G. 2011. Patient-Centered Outcomes in Mental Health Care. JAMA 306: 1141–1142. [Google Scholar]
Figure 1. Three-layer artefact overview. The design contribution is structured into an interaction-flow orchestration layer, a detect–evaluate–adapt regulation layer, and the current theory-grounded instantiation using OCEAN personality signals, Zurich Model motivational domains, and behavioural response effects.
Figure 1. Three-layer artefact overview. The design contribution is structured into an interaction-flow orchestration layer, a detect–evaluate–adapt regulation layer, and the current theory-grounded instantiation using OCEAN personality signals, Zurich Model motivational domains, and behavioural response effects.
Preprints 216002 g001
Figure 2. Parallel orchestration and adaptive-regulation flow. State-machine interaction control and detect–evaluate–adapt personalisation operate in parallel on the same user utterance and merge at prompt assembly before response generation.
Figure 2. Parallel orchestration and adaptive-regulation flow. State-machine interaction control and detect–evaluate–adapt personalisation operate in parallel on the same user utterance and merge at prompt assembly before response generation.
Preprints 216002 g002
Figure 3. OCEAN perception component. User utterances and dialogue context are processed into turn-level Big Five/OCEAN trait signals and accumulated into an explicit personality-state representation for downstream adaptation.
Figure 3. OCEAN perception component. User utterances and dialogue context are processed into turn-level Big Five/OCEAN trait signals and accumulated into an explicit personality-state representation for downstream adaptation.
Preprints 216002 g003
Figure 7. Evaluation framework and scoring pipeline.
Figure 7. Evaluation framework and scoring pipeline.
Preprints 216002 g007
Figure 8. Sample characteristics summary. (A) Distribution of assistant turns analysed; (B) Personality profiles representing experimental arms; (C) 2 × 2 factorial design matrix showing balanced group sizes. All evaluation metrics achieved 100% data coverage across both conditions.
Figure 8. Sample characteristics summary. (A) Distribution of assistant turns analysed; (B) Personality profiles representing experimental arms; (C) 2 × 2 factorial design matrix showing balanced group sizes. All evaluation metrics achieved 100% data coverage across both conditions.
Preprints 216002 g008
Figure 9. Distribution of detected OCEAN personality dimensions across all conversation turns (Regulated condition, N = 60 turns). Grouped bars show frequency of Low ( 1 ), Mid (0), and High ( + 1 ) trait levels for each Big Five dimension; values above bars indicate count and percentage. Detection fidelity is dimension-specific: Extraversion, Agreeableness (Type A turns), and Neuroticism show strong alignment (≥87% intended polarity), while Conscientiousness and Agreeableness (Type B turns) exhibit higher Mid (0) rates, reflecting neutral-state assignments when conversational evidence was insufficient. Overall holistic detection accuracy: 98.3% (59/60 Yes).
Figure 9. Distribution of detected OCEAN personality dimensions across all conversation turns (Regulated condition, N = 60 turns). Grouped bars show frequency of Low ( 1 ), Mid (0), and High ( + 1 ) trait levels for each Big Five dimension; values above bars indicate count and percentage. Detection fidelity is dimension-specific: Extraversion, Agreeableness (Type A turns), and Neuroticism show strong alignment (≥87% intended polarity), while Conscientiousness and Agreeableness (Type B turns) exhibit higher Mid (0) rates, reflecting neutral-state assignments when conversational evidence was insufficient. Overall holistic detection accuracy: 98.3% (59/60 Yes).
Preprints 216002 g009
Figure 10. Personality trait patterns across conversation turns (Regulated condition, N = 60 turns, sorted by Conversation ID). Each column represents one turn; each row represents one OCEAN dimension. Colours indicate trait levels: Blue = High ( + 1 ), Orange = Low ( 1 ), Gray = Mid (0). Vertical lines mark conversation boundaries (each block = 6 turns). A broad colour split is visible between the left (Type A) and right (Type B) halves, with blue dominant in Type A and orange dominant in Type B for most dimensions. Gray (Mid) patterns are more pronounced in Conscientiousness and Agreeableness, reflecting dimensions where neutral-state detection occurs more frequently even under extreme profile conditions.
Figure 10. Personality trait patterns across conversation turns (Regulated condition, N = 60 turns, sorted by Conversation ID). Each column represents one turn; each row represents one OCEAN dimension. Colours indicate trait levels: Blue = High ( + 1 ), Orange = Low ( 1 ), Gray = Mid (0). Vertical lines mark conversation boundaries (each block = 6 turns). A broad colour split is visible between the left (Type A) and right (Type B) halves, with blue dominant in Type A and orange dominant in Type B for most dimensions. Gray (Mid) patterns are more pronounced in Conscientiousness and Agreeableness, reflecting dimensions where neutral-state detection occurs more frequently even under extreme profile conditions.
Preprints 216002 g010
Figure 11. Weighted scores comparison across evaluation metrics (Regulated vs. Baseline). Bars show mean weighted scores (M; 0–2 scale: No = 0, Not Sure = 1, Yes = 2) with standard deviation ( S D ) error bars. Value annotations positioned above error bars. Legend relocated to lower left to avoid overlap. Both conditions achieve ceiling performance on Emotional Tone and Relevance, with dramatic difference only on Personality Needs.
Figure 11. Weighted scores comparison across evaluation metrics (Regulated vs. Baseline). Bars show mean weighted scores (M; 0–2 scale: No = 0, Not Sure = 1, Yes = 2) with standard deviation ( S D ) error bars. Value annotations positioned above error bars. Legend relocated to lower left to avoid overlap. Both conditions achieve ceiling performance on Emotional Tone and Relevance, with dramatic difference only on Personality Needs.
Preprints 216002 g011
Figure 12. Total-score distribution by condition (sum of Emotional Tone + Relevance & Coherence + Personality Needs; maximum = 6). Bars show the percentage of turns at each total-score value, with counts annotated for non-zero cells. The Regulated condition shows complete ceiling performance, with all 60 turns scoring 6/6, whereas Baseline scores are distributed between 2 and 6 ( M d n = 4 , range 2–6). p < 0.001 (independent-samples t-test; supplementary).
Figure 12. Total-score distribution by condition (sum of Emotional Tone + Relevance & Coherence + Personality Needs; maximum = 6). Bars show the percentage of turns at each total-score value, with counts annotated for non-zero cells. The Regulated condition shows complete ceiling performance, with all 60 turns scoring 6/6, whereas Baseline scores are distributed between 2 and 6 ( M d n = 4 , range 2–6). p < 0.001 (independent-samples t-test; supplementary).
Preprints 216002 g012
Figure 13. Selective enhancement demonstrated through paired conversation-level Yes-rate analysis. Lines connect the same conversation under Baseline and Regulated conditions; blue = Type A ( + 1 profiles), amber = Type B ( 1 profiles); bold black line = group M. Mean (M) values annotated beside markers. The dramatic upward shift in Personality Needs (right panel, M = 0.08 1.00 ) contrasts with negligible change in Emotional Tone and Relevance & Coherence (both at ceiling), confirming targeted and selective improvement.
Figure 13. Selective enhancement demonstrated through paired conversation-level Yes-rate analysis. Lines connect the same conversation under Baseline and Regulated conditions; blue = Type A ( + 1 profiles), amber = Type B ( 1 profiles); bold black line = group M. Mean (M) values annotated beside markers. The dramatic upward shift in Personality Needs (right panel, M = 0.08 1.00 ) contrasts with negligible change in Emotional Tone and Relevance & Coherence (both at ceiling), confirming targeted and selective improvement.
Preprints 216002 g013
Figure 14. Response distribution by metric and condition shown as 100% stacked bars. Each bar represents the full response distribution (Yes = teal, Not Sure = gray, No = coral) for Baseline and Regulated. Count and percentage annotated within segments ≥8%. Emotional Tone and Relevance & Coherence show ceiling effects in both conditions (≥98% Yes). Personality Needs reveals the key contrast: Baseline achieves only 8% Yes (5/60) vs. Regulated at 100% Yes (60/60), confirming selective personality-need fulfilment.
Figure 14. Response distribution by metric and condition shown as 100% stacked bars. Each bar represents the full response distribution (Yes = teal, Not Sure = gray, No = coral) for Baseline and Regulated. Count and percentage annotated within segments ≥8%. Emotional Tone and Relevance & Coherence show ceiling effects in both conditions (≥98% Yes). Personality Needs reveals the key contrast: Baseline achieves only 8% Yes (5/60) vs. Regulated at 100% Yes (60/60), confirming selective personality-need fulfilment.
Preprints 216002 g014
Figure 15. Response comparison for a Type B (all traits: 1 ) profile turn (Conversation B-1, turn 2). For the same user input, the regulated response emphasizes validation and a low-pressure stance (supporting Security and reducing Arousal), whereas the baseline response remains broadly supportive but less specifically aligned to the user’s implied motivational state.
Figure 15. Response comparison for a Type B (all traits: 1 ) profile turn (Conversation B-1, turn 2). For the same user input, the regulated response emphasizes validation and a low-pressure stance (supporting Security and reducing Arousal), whereas the baseline response remains broadly supportive but less specifically aligned to the user’s implied motivational state.
Preprints 216002 g015
Figure 16. Response comparison for a Type A (all traits: + 1 ) profile turn (Conversation A-1, turn 3). For the same user input, the regulated response maintains an affirming, growth-oriented framing, while the baseline response shifts toward generic reflective questioning; both are appropriate, but only the regulated condition is explicitly constrained by trait-conditioned regulation logic.
Figure 16. Response comparison for a Type A (all traits: + 1 ) profile turn (Conversation A-1, turn 3). For the same user input, the regulated response maintains an affirming, growth-oriented framing, while the baseline response shifts toward generic reflective questioning; both are appropriate, but only the regulated condition is explicitly constrained by trait-conditioned regulation logic.
Preprints 216002 g016
Table 1. Big Five personality trait operationalisation. Neuroticism is inverse-coded: +1 denotes emotional stability (low neuroticism), and -1 denotes emotional vulnerability (high neuroticism).
Table 1. Big Five personality trait operationalisation. Neuroticism is inverse-coded: +1 denotes emotional stability (low neuroticism), and -1 denotes emotional vulnerability (high neuroticism).
Trait +1 (High) -1 (Low)
Openness Curious, imaginative, open to novelty Prefers routine, resistant to new ideas
Conscientiousness Organised, disciplined, structured Disorganised, impulsive, spontaneous
Extraversion Outgoing, energetic, assertive Reserved, quiet, withdrawn
Agreeableness Cooperative, empathetic, friendly Critical, skeptical, confrontational
Neuroticism (inverse-coded) Calm, emotionally stable, resilient Anxious, emotionally sensitive, insecure
Table 2. Behavioural adjustment strategies mapped by Zurich Model motivational domain. Conscientiousness modulates interaction structure independently of the three core domains.
Table 2. Behavioural adjustment strategies mapped by Zurich Model motivational domain. Conscientiousness modulates interaction structure independently of the three core domains.
Domain Trait +1 (High) -1 (Low)
Security Neuroticism (N) Reassure stability and confidence Offer extra comfort; acknowledge anxieties
Arousal Openness (O) Invite exploration and novelty Focus on familiar topics; reduce novelty
Arousal Extraversion (E) Energetic, sociable tone Calm, low-key style with reflective space
Affiliation Agreeableness (A) Warmth, empathy, collaboration Neutral, matter-of-fact stance
Cross-cutting Conscientiousness (C) Provide organised, structured guidance Flexible, relaxed, spontaneous demeanour
Table 3. Turn-level descriptive comparison of regulated vs. baseline assistants across evaluation metrics (supplementary independent-samples tests). M = mean; S D = standard deviation; n = 60 assistant turns per condition. Metrics scored on 0–2 scale (0 = No, 1 = Not Sure, 2 = Yes). ***  p < 0.001 ; not significant = p > 0.05 . Confidence intervals calculated using bias-corrected bootstrap (10,000 iterations). *Primary outcome. †Cohen’s d is technically undefined when the pooled S D = 0 (both groups at ceiling with S D = 0.00 ); the reported value of 0.000 follows the convention of treating 0/0 = 0 and is included for tabular completeness. Primary inferential analysis used conversation-level paired tests ( n = 10 pairs; Section 3.6.1).
Table 3. Turn-level descriptive comparison of regulated vs. baseline assistants across evaluation metrics (supplementary independent-samples tests). M = mean; S D = standard deviation; n = 60 assistant turns per condition. Metrics scored on 0–2 scale (0 = No, 1 = Not Sure, 2 = Yes). ***  p < 0.001 ; not significant = p > 0.05 . Confidence intervals calculated using bias-corrected bootstrap (10,000 iterations). *Primary outcome. †Cohen’s d is technically undefined when the pooled S D = 0 (both groups at ceiling with S D = 0.00 ); the reported value of 0.000 follows the convention of treating 0/0 = 0 and is included for tabular completeness. Primary inferential analysis used conversation-level paired tests ( n = 10 pairs; Section 3.6.1).
Metric Regulated
M ( S D )
Baseline M
( S D )
Difference Cohen’s d 95% CI p-value
Personality Needs* 2.00 (0.00) 0.20 (0.58) 1.80 4.42 [3.1, 8.1] <0.001***
Emotional Tone 2.00 (0.00) 2.00 (0.00) 0.00 0.000 [-0.2, 0.2] 1.000 not significant
Relevance & Coherence 2.00 (0.00) 1.97 (0.26) 0.03 0.183 [0.0, 0.3] 0.319 not significant
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings