Preprint
Article

This version is not peer-reviewed.

Personality-Adaptive Conversational AI for Emotional Support: A Simulation Study Integrating Big Five Detection with Zurich Model Regulation

Submitted:

20 August 2026

Posted:

21 August 2026

You are already at the latest version

Abstract
LLM-based conversational agents generate fluent responses but remain limited in adapting their supportive style to individual personality and emotional needs. We address this gap through a theory-grounded personality-adaptive conversational AI architecture that performs turn-by-turn Big Five trait detection and applies Zurich Model-aligned behavioural regulation, orchestrated with PROMISE (a model-driven framework for state-based LLM orchestration). We generated 20 simulated dialogue sessions (120 assistant turns) comparing regulated and non-adaptive assistants across two boundary-condition Big Five profiles (all traits +1 vs. −1), scored by a structured LLM-based evaluator with author review retained as an internal audit. Under these conditions the pipeline executed as designed: holistic detection was correct on 59/60 turns (98.3%) and regulation adherence was complete (100%), although dimension-level accuracy was lower (macro mean 83.3%). The turn-level Personality Needs Yes rate rose from 8.3% (5/60) at baseline to 100% (60/60) under regulation, and all 10 conversation pairs favoured regulation; emotional tone was not estimable (both arms at ceiling) and relevance did not differ—a pattern of selective enhancement. Mixed-profile testing reduced dimension-level accuracy to 58.1%, locating the boundary result as an upper bound. A two-rater verify-and-revise audit on an n = 66 overlap supports detection-label convergence (mean linear κ = 0.731) but does not validate the outcome rubrics, which remain LLM-rated. The contribution is a reproducible detection→regulation→evaluation architecture for controlled simulation, implemented on closed cloud-hosted models whose computational footprint and detector interpretability are not characterised here. The Zurich mapping is implemented as a design choice rather than validated as psychological theory; the system is not evaluated as a therapeutic or clinical intervention, and clinical value remains to be demonstrated in human-participant studies.
Keywords: 
;  ;  ;  ;  ;  ;  

1. Introduction

Conversational agents are increasingly used in healthcare to improve health literacy, encourage treatment adherence, support well-being monitoring, and provide emotionally supportive guidance [1,2]. As these systems become more capable and more widely deployed, their effectiveness depends not only on whether they produce relevant information, but also on how they interact with users. In many health contexts, the intended impact emerges through communication, such as whether an agent can build trust, provide reassurance, sustain motivation, and adapt its tone and strategy to the user over time [1,3]. Designing such interaction qualities remains challenging, especially in sensitive settings where personalisation and consistency are critical [4] and where underlying model reliability remains an open concern [5].
Moreover, health support can rarely be achieved through one-size-fits-all communication. Effective interactions often require behaviours such as empathy, reassurance, encouragement, and persuasion, but these behaviours are not universally appropriate in a fixed form. An empathetic style may strengthen trust and rapport by mirroring effective physician–patient communication [6]. However, whether such a style is perceived as supportive may depend on the person and on how the interaction unfolds over time [7,8,9]. Similarly, persuasive support in digital health depends on selecting communication strategies that fit the receiver, the goal, and the moment of interaction rather than simply presenting information [10,11,12]. If conversational agents are expected to achieve intended outcomes such as literacy, adherence, reassurance, or emotional support, they thus require flexible means of implementing personalised conversational behaviour during the interaction rather than relying solely on static design-time instructions.
Recent advances in large language models (LLMs) have greatly expanded the potential of conversational agents to produce fluent, contextually rich, and human-like dialogue [3,13,14]. This makes them promising building blocks for adaptive health-oriented conversational agents [15,16,17,18]. However, most LLM-based approaches focus primarily on response generation. Training or fine-tuning domain-specific models remains costly and data-intensive [5,19,20,21], while prompt engineering, despite its practical flexibility [22,23,24,25,26], leaves the central design question open: what conversational behaviour should be enacted, for whom, and when—providing plausibility without grounded, controllable, or reproducible personalisation.
Emerging agentic AI and orchestration frameworks offer part of the answer by enabling modular, context-sensitive systems with structured control over goals, prompts, and interaction flow [27]. PROMISE, for example, supports phase-specific prompts and transition logic through model-driven state-machine orchestration, making it useful for transparent and trust-sensitive applications. Yet structured interaction flow alone is insufficient when the core requirement is not just to progress through phases, but to regulate socially meaningful behaviour in ways that remain responsive to the user. The challenge is therefore not only architectural—how to structure multi-turn interactions—but substantive: how to embed principled personalisation so that adaptation tracks user differences in an inspectable and evaluable way.
This paper addresses that gap through a framework for grounded social behaviour regulation in conversational AI. Rather than treating personalisation as ad hoc prompt variation, we operationalise it as a structured process that combines dynamic personality modelling, theory-informed behavioural adaptation, and controlled orchestration. In the present instantiation, relatively stable user differences are modelled through turn-by-turn Big Five detection (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism; collectively OCEAN), behavioural adaptation is guided by the Zurich Model of Social Motivation, and interaction control is implemented through PROMISE with a social behaviour regulation extension. This instantiation tests whether that mapping can be executed as designed; it does not validate the Zurich Model as a psychological theory. Mental health support serves as a particularly demanding demonstration context because emotionally supportive interactions place strong demands on sensitivity, trust, and adaptation, but the broader contribution is a design approach for implementing personalised conversational behaviour in health-oriented conversational agents.
We evaluate this approach in a controlled simulation study comparing regulated and non-adaptive assistants across two polarised, boundary-condition personality profiles (all traits + 1 versus 1 ). This boundary-condition simulation is a feasibility test that precedes human-participant validation rather than a substitute for it. The study provides a reproducible testbed for examining whether the regulation can be implemented with high fidelity and whether it selectively raises LLM-rated Personality Needs scores without degrading generic conversational quality. Under extreme profiles, holistic detection was correct on 59 of 60 turns (98.3%) with complete regulation adherence (100%). The turn-level Personality Needs Yes rate, reported as descriptive only, rose from 8.3% (5/60) at baseline to 100% (60/60) under regulation, and all 10 conversation pairs favoured regulation. This complete separation is best read as a binary architectural contrast under maximal signal rather than a portable effect size; mixed-profile testing later reduced dimension-level accuracy to 58.1%, locating the boundary result as an upper bound. Emotional tone was not estimable because both arms were at ceiling, and topical relevance did not differ. Taken together, these findings position the work as an architecture for personalisation that remains controllable, reproducible, and evaluable, not as evidence of clinical effectiveness.

3. Materials and Methods

The evaluation comprised a primary 2 × 2 boundary-condition study (Phase 1A; also called Study A when referring to its scored matrices) and three validation extensions: cross-family rescoring of those outputs (Phase 1B), mixed-profile testing with input replay (Phase 1C), and a two-rater human label audit on a mixed-profile annotation corpus (Phase 2).

3.1. Overview and Research Objectives

This study addresses three empirical questions (aligned with the Results narrative as RQ1–RQ3):
1.
RQ1 (Detection Feasibility): Can the architecture infer turn-by-turn Big Five states with high fidelity under boundary-condition profiles, and how does that fidelity degrade under mixed or subtle profiles?
2.
RQ2 (Regulation Contrast): Does theory-driven personality-adaptive regulation raise LLM-rated Personality Needs scores relative to a non-adaptive baseline without reducing generic conversational quality?
3.
RQ3 (Evaluation Validity): Given LLM-as-judge scoring, what do cross-model rescoring and a two-rater human label audit imply for the credibility and limits of the comparative claims?

3.1.1. Experimental Design

The primary study used a 2 × 2 factorial design [52,53]: regulated versus baseline assistants crossed with Type A versus Type B personality profiles, with 5 sessions per cell (20 sessions; 120 assistant turns). This is equivalent to 10 simulated personas, 5 per profile, evaluated under both assistant conditions. The extreme profiles provide a maximal-signal test of detection and regulation fidelity.
Regulated assistants updated personality inferences after each turn and applied Zurich Model-aligned behavioural guidance [37,38,39,40,41]. Baseline assistants provided emotional support without personality-specific adaptation. Both used separate GPT-4 instances (OpenAI API, gpt-4-0613, accessed June–July 2024) with condition-specific system prompts.
The workflow comprised four steps: (1) system configuration and personality profile implementation, (2) simulated dialogue generation following standardised protocols, (3) evaluation using a structured LLM-based assessor (with author-conducted qualitative review), and (4) statistical analysis.
Regulated and baseline conversations were scenario-matched: the same OCEAN profile and scenario seed were used across conditions, but user utterances were generated independently within each condition. Paired session IDs were constructed as profile–persona indices (e.g., A-1 / B-1 denote the first Type A / Type B persona), with each persona contributing one regulated and one baseline conversation for the conversation-level Wilcoxon analysis. The pairing manifest (pairing_manifest.csv, Supplementary File S4) confirms approximately 0% identical user turns across conditions. Inference is therefore a between-condition comparison within persona, not an input-replay matched-pair design. Both conditions were generated as parallel instantiations within each persona–session (not sequentially, so no condition-order effect arises); the two arms shared the same scenario progression and profile prompt and differed only in the condition-specific system prompt (dynamic detection/regulation modules vs. static supportive baseline). No other prompts, seeds, or topics differed between paired sessions. After the first turn, dialogue histories necessarily diverged because assistant replies differed by condition. In the mixed-profile extension (§Section 4.4), input replay was used: identical user text was presented to both arms, strengthening the paired contrast. The primary study’s independently generated user turns therefore limit strict causal attribution of the headline contrast to the regulation policy; the mixed-profile extension uses input replay to strengthen that paired design.

3.1.2. Sample Size Considerations

The 20 sessions are GPT-4 instantiations, not human subjects. The simulation prioritises architectural fidelity over statistical precision: n = 10 conversation pairs provide a controlled testbed for implementation feasibility rather than a population effect estimate [54]. Computational cost and limited stochastic variation also constrained sample size. Traditional power analysis was not applied because the simulation does not estimate population variability and includes zero-variance ceiling outcomes.
For future human validation (Stage 2 of the validation pathway), we provide a prospective calculation anchored to a defined primary outcome (Personality Needs Addressed, conversation-level). Because simulation effects are upper bounds, we power for substantially smaller real-world effects rather than the observed contrast: detecting a moderate paired effect ( d = 0.5 ; 80% power, α = 0.05 two-sided, Wilcoxon ARE correction) requires ≈35 participants in a within-subject design, and ≈64 per condition in a between-groups design; for a small-to-moderate effect ( d = 0.3 ), ≈92 within-subject or ≈176 per condition between-groups. These estimates refine the general n 50 –100 per-condition guidance from clinical trial standards [55,56,57] and should be finalised using pilot variance estimates.

3.1.3. Ethics and Transparency

No human subjects were involved in this study; the user-side prompts/questionnaires were written manually, and the resulting dialogues were generated in a controlled manner using GPT-4-based agents following predefined personality profiles. As the data were synthetic, non-identifiable, and created solely for technical validation, the study fell outside the scope of human-subject research as defined by established ethical frameworks, and formal ethics approval was therefore not required [58].
Human-participant extensions would require institutional review board approval, comprehensive informed consent addressing AI-based personality profiling and data usage, real-time safety monitoring with clear escalation pathways for psychological distress, and ongoing oversight by qualified mental health professionals [59,60]. Ethical design of conversational agents should also address questions of interactional respect and user autonomy [61].

3.2. Personality-Adaptive System Architecture

The system integrates turn-by-turn Big Five detection and regulation within PROMISE [27]. Five mechanisms define the architecture: (i) tri-state OCEAN values { 1 , 0 , + 1 } , including a neutral state for insufficient evidence (Table 1); (ii) cumulative updating from dialogue history; (iii) commitment only above confidence threshold τ ; (iv) parallel trait-specific detectors; and (v) Zurich Model mapping from non-neutral traits to Security, Arousal, and Affiliation guidance (Section 3.3.1).
PROMISE represents an interaction as an Agent containing States, Transitions, Decisions, and Actions over a key–value store [27]. State-specific prompts and transition logic provide explicit control, while the store persists trait estimates, regulation instructions, and evaluation logs. Figure 1 shows the resulting D–R–E architecture.
The three functions are:
  • Detection: After each user turn, trait-specific detectors evaluate the accumulated dialogue context and update the estimated personality-state vector P ^ = ( O , C , E , A , N ) together with its confidence values in the interaction store. Each trait is represented as 1 , 0, or + 1 ; a directional value is committed only when confidence exceeds the threshold τ = 0.7 , otherwise the trait remains neutral.
  • Regulation: Response-generation states retrieve the current personality state and map non-neutral traits to Zurich Model-aligned behavioural guidance. PROMISE dynamically composes this guidance with the base prompt to adjust tone, pacing, structure, and interpersonal framing while preserving the conversational task objective. Conflicting recommendations are resolved before generation, with Security-related guidance prioritised over stimulation.
  • Evaluation: After each assistant response, structured assessment actions apply the predefined rubrics and store scores with conversation and turn identifiers. These records support traceability and offline analysis; evaluation scores do not modify detection, regulation, or response generation during the simulation.
The D–R–E logic is framework-agnostic; PROMISE provides the declarative state model and logs needed to inspect each adaptation. Independent modules communicate through JSON objects containing personality vectors, regulation instructions, and scores. Trait mappings remain external configuration, and decisions are logged with timestamps and turn identifiers. This design follows three principles: separation of concerns allows detection, regulation, and evaluation to be tested independently; configurability allows additional traits or regulation policies without changing module interfaces; and traceability links each adaptation to an explicit state transition, stored trait estimate, and regulation instruction. Figure 2 details the OCEAN perception component.
During execution, PROMISE evaluates transition guards against accumulated utterances, composes state and role prompts, and writes extracted outputs to the interaction store [27]. Our integration uses these mechanisms to persist Big Five estimates, inject regulation guidance, and log evaluation outcomes.

3.3. System Implementation

This section describes how the system implements (i) turn-by-turn personality detection and (ii) theory-driven regulation, both orchestrated within PROMISE (Section 3.2).
We represent the user’s Big Five (OCEAN) profile as a five-dimensional vector P = ( O , C , E , A , N ) , with each trait { 1 , 0 , + 1 } [62,63,64]. Neuroticism is inverse-coded throughout: + 1 denotes emotional stability and 1 denotes emotional vulnerability (Table 1). In this simulation, the personality vector functions as a predefined reference target for controlled evaluation rather than psychometric ground truth; detection accuracy denotes agreement with the pre-specified simulated profile.
After each user input, the system updates the estimated personality-state vector P ^ using cumulative evidence from the dialogue history, following a belief-updating logic similar to Bayesian belief updating [65]. The estimated state is represented as P ^ = ( O , C , E , A , N ) . PROMISE orchestrates this pipeline via hierarchical state machines, parallel trait detectors, and explicit transition logic (Section 3.2). Inference produces trait values { 1 , 0 , + 1 } and confidence scores, with trait state transitions triggered when confidence exceeds the pre-specified threshold τ = 0.7 (Figure 2). This threshold was chosen as a conservative design trade-off: it reduces premature trait commitments from ambiguous evidence while still allowing adaptation once repeated conversational cues accumulate.
The detection component is implemented as a modular service integrated into PROMISE. It uses GPT-4 (gpt-4-0613) with trait-specific prompts grounded in personality psychology [23,25,66,67,68,69,70]. The pipeline executes the following stages: (1) message processing, (2) prompt assembly, (3) parallel LLM inference across traits, (4) response parsing and confidence assessment, (5) state update, and (6) downstream transmission to the regulation module.
Average processing time was 2.3 ± 0.8 seconds per turn (measured across 120 turns under GPT-4 API hosting). Timing boundaries were not instrumented separately, so this figure is a wall-clock average under hosted API conditions rather than a decomposition into network, queueing, detection, regulation, or generation. Parallel trait detection is expected to reduce wall-clock latency relative to five sequential trait calls, but no sequential baseline was recorded in this study and no percentage reduction is reported. Because the present pipeline relies on a closed-source cloud API, it does not characterise on-device memory footprint or FLOPs. For privacy-sensitive mental-health settings, the same D–R–E / PROMISE state machine could be rehosted on open-weight local models; that would require separate latency, VRAM, and throughput benchmarks and is left to future deployment work. Raw per-turn confidence scores were applied at inference time to trigger state transitions (threshold τ = 0.7 ) but were not persisted in the archived evaluation matrices; the neutral-assignment rate per trait (O: 15.0%, C: 35.0%, E: 6.7%, A: 23.3%, N: 0%; see Table 2) serves as the observable proxy for the proportion of turns in which detector confidence remained below threshold.

3.3.1. Theory-Driven Regulation

Regulation translates detected traits into behavioural adaptations grounded in the Zurich Model of Social Motivation [37,38,39,40,41], which links personality to three motivational domains: Security, Arousal, and Affiliation (Section 2.1). Table 3 and Figure 3 present the mapping from Big Five traits to motivational domains and corresponding behavioural adjustment strategies.
Figure 3 presents the conceptual path from OCEAN signals to motivational domains and response effects. Table 3 makes this path operational by specifying the direction of the prompt-level adjustment for high and low trait values. Neuroticism, Openness, Extraversion, and Agreeableness map to Security, Arousal, and Affiliation, whereas Conscientiousness acts across domains by changing the degree of interaction structure. The mapping is theory-inspired rather than empirically validated: a psychology-trained co-author reviewed the trait-to-domain links for face plausibility, but the study does not include a multi-expert consensus rating or a shuffled-mapping ablation, so links such as Openness→Arousal and Conscientiousness as cross-cutting should be treated as design choices pending formal expert appraisal.
Table 3. Behavioural adjustment strategies mapped by Zurich Model motivational domain. Conscientiousness modulates interaction structure independently of the three core domains.
Table 3. Behavioural adjustment strategies mapped by Zurich Model motivational domain. Conscientiousness modulates interaction structure independently of the three core domains.
Domain Trait +1 (High) -1 (Low)
Security Neuroticism (N) Reassure stability and confidence Offer extra comfort; acknowledge anxieties
Arousal Openness (O) Invite exploration and novelty Focus on familiar topics; reduce novelty
Arousal Extraversion (E) Energetic, sociable tone Calm, low-key style with reflective space
Affiliation Agreeableness (A) Warmth, empathy, collaboration Neutral, matter-of-fact stance
Cross-cutting Conscientiousness (C) Provide organised, structured guidance Flexible, relaxed, spontaneous demeanour
After each personality update, PROMISE composes an integrated regulation instruction by concatenating trait-relevant behaviour prompts for all non-neutral traits (Figure 4). When trait recommendations conflict, the system prioritises emotional safety (Security domain) over stimulation (Arousal domain), consistent with clinical guidance [71].
The resulting instruction is therefore a logged, inspectable transformation of the current personality state rather than an unconstrained rewrite of the response objective. Conflict resolution preserves the Security priority when trait-conditioned cues compete, after which the assembled instruction is combined with the same base task prompt used by the regulated response generator.

3.4. Experimental Materials and Procedure

The primary study used two synthetic boundary profiles: Type A ( + 1 , + 1 , + 1 , + 1 , + 1 ) and Type B ( 1 , 1 , 1 , 1 , 1 ) [62,72,73]. These polarised profiles are uniformly high and uniformly low boundary conditions; high or low standing on a Big Five trait is not uniformly positive or negative. They maximise signal for architectural testing but do not represent typical trait distributions [62]. Detection accuracy therefore measures agreement with a predefined simulation target, not psychometric ground truth.
We generated 5 regulated and 5 baseline sessions for each profile, yielding four cells: Regulated–Type A, Regulated–Type B, Baseline–Type A, and Baseline–Type B ( n = 5 each). Each session contained 6 exchanges, giving 120 assistant turns across the four cells. Separate gpt-4-0613 instances generated personality-conditioned user messages and assistant responses (temperature 0.7; max_tokens=500). The temperature balanced conversational variation against the need to preserve condition-specific control. Sessions followed the same progression: opening support, disclosure, elaboration, reframing, coping guidance, and closing. Scenario and profile prompts were shared across paired conditions, but user turns were generated independently rather than replayed. Conversations were exported to Evaluation Matrices; data collection ran for 4 weeks in June–July 2024. The empirical contribution is the D–R–E orchestration architecture rather than a claim about a single LLM snapshot: later cross-family rescoring with Sonnet 4.6, Gemini 3.1 Flash-Lite, and DeepSeek V4 Flash (Section 4.3 and Section 4.4) preserves the regulated–baseline Personality Needs direction, indicating that the comparative result is not tied to the deprecated gpt-4-0613 generator/evaluator pair alone.

3.5. Validation Extensions

3.5.1. Cross-Model Rescoring

To test same-family evaluator dependence, the 120 Study A outputs per metric were rescored with claude-sonnet-4-6. The blinded prompt included only the user turn, assistant response, and rubric; it excluded profile, regulation directives, and condition labels. Unlike the original GPT-family judge, which scored regulated and baseline matrices separately (Section 3.6.1), the blinded items were presented in alternating condition order (regulated, then baseline, within each turn). The same 120-item file, in that mixed order, was also scored by Gemini 3.1 Flash-Lite and DeepSeek V4 Flash. Cross-family rescoring used temperature 0.0, whereas the original evaluator used temperature 0.3; the contrast is a completed model×temperature check of the Personality Needs direction on the same 120 outputs, not a full prompt×temperature factorial and not a same-model GPT-4-turbo temperature grid. Agreement was assessed with Cohen’s κ , using Gwet AC1 for high-prevalence outcomes (AC1 corrects for prevalence-induced κ deflation when one category dominates), and regulated–baseline contrasts were recomputed with Cliff’s δ .

3.5.2. Mixed-Profile Extension

Generalisation was tested on eight core profiles. The six mixed configurations were Friendly Conventional (FC), Relaxed Creative (RC), Temperamental Uninhibited (TU), Warm Skeptic (WS), Anxious Achiever (AA), and Reserved Organizer (RO). The first three were adapted from the personality configurations identified by Rentfrow et al. [74]; the remaining three were constructed to introduce contrasting mixed-trait patterns. Two single-salient profiles represented subtler signals: Faint Extravert (FE; Extraversion + 1 only) and Anxious Only (AN; Neuroticism 1 only). Two extreme profiles were retained as reference anchors. Unlike Study A, this extension replayed identical user turns across regulated and baseline arms. The core set comprised 24 conversations and was scored by three cross-family judges: Sonnet 4.6, Gemini 3.1 Flash-Lite, and DeepSeek V4 Flash.

3.5.3. Human Label Audit

Two independent non-author raters audited the same Phase 2 OCEAN subset ( n = 66 turns) using a verify-and-revise protocol: one senior research associate with a human–computer interaction background (engineering and architecture; information technology and ambient assisted living) and one senior research associate with a psychology background (business; marketing and economic psychology), both at Lucerne University of Applied Sciences and Arts (staff profiles in Acknowledgments). The n = 66 overlap is the complete-case slice of the first 70 turns issued in the Phase 2 rater packet (four incomplete rows excluded); it is not a separately stratified draw from the full Phase 2 corpus. Raters were blind to condition, generator, intended profile, and each other’s labels, and reviewed pre-filled OCEAN proposals. We report (i) human–machine agreement for each rater as an upper bound under pre-fill anchoring and (ii) human–human agreement on the full overlap as inter-rater reliability. Where the two raters agreed, that value is treated as a consensus label; where they disagreed (19.1% of cells), labels are left unresolved under a consensus-only rule (no machine tie-break) and are not treated as adjudicated ground truth.

3.6. Evaluation Procedure

Regulated outputs were scored on Detection Accuracy, Regulation Effectiveness, Emotional Tone Appropriateness, Relevance & Coherence, and Personality Needs Addressed. Baseline outputs were scored on the three shared outcomes because they contained no detection or regulation module. All criteria used Yes=2, Not Sure=1, and No=0.

3.6.1. LLM-Based Evaluator System

A GPT-4-turbo evaluator used criterion-specific prompts to score each row with the full dialogue context and profile [50,51,75]. A structured LLM-as-Judge procedure was used because the target qualities, including empathy, reassurance, personality-sensitive appropriateness, and regulation fidelity, are not captured well by overlap metrics [75]. Applying a fixed rubric across all outputs provided a scalable proof-of-concept comparison before resource-intensive human rating. Regulated and baseline matrices were evaluated separately; the evaluator was told that detection could update across turns and used temperature 0.3. Using a different GPT-4 snapshot reduced exact model reuse but not shared-family bias, so these ratings are comparative feasibility measures rather than ground truth [51].

3.6.2. Evaluation Validity and Author-Led Qualitative Review

Two contributing authors (Samuel Devdas and Jiahua Duojie) also reviewed all matrices and commented on each item. Their judgements largely aligned with automated ratings, with one uncertain item, but this internal audit is vulnerable to confirmation bias and is not independent validation [51]. Stronger validation requires overlapping ratings from at least two independent raters, inter-rater reliability such as Krippendorff’s α [76], and human–AI agreement such as Cohen’s κ [77,78].

3.7. Analysis and Reproducibility

The primary outcome was “Personality Needs Addressed,” measured on a three-point scale (Yes = 2, Not Sure = 1, No = 0). Secondary outcomes included Detection Accuracy, Regulation Effectiveness, Emotional Tone Appropriateness, and Relevance & Coherence. Turn-level ratings were collected for each assistant output (10 dialogue sessions per assistant type × 6 assistant turns = 60 outputs per condition) [79].

3.7.1. Statistical Analysis

The primary unit of analysis for inferential comparisons between regulated and baseline conditions was the conversation ( n = 10 matched pairs: five Type A and five Type B profiles, each with regulated and baseline sessions). For each conversation and metric, we computed the mean rubric score across six assistant turns and tested paired differences using the exact Wilcoxon signed-rank test (two-tailed). The reported Wilcoxon statistic is the sum of ranks of the positive differences ( T + ); under complete concordance of n = 10 pairs this equals 55. Given the three-point ordinal scale, ceiling effects, and n = 10 , we did not fit a cumulative-logit mixed model; conversation-level exact tests plus explicitly labelled turn-level descriptives are reported instead. Confidence intervals for effect sizes were computed using cluster-bootstrap resampling (10,000 iterations) with conversation as the resampling unit to preserve within-conversation dependence [80].
For comparability with prior conversational AI studies, we also report turn-level descriptive statistics ( n = 60 assistant turns per condition); because turns within a conversation are not independent, these values are descriptive supplements rather than independent inferential observations. Cliff’s δ was computed on the ten conversation-level means per condition, not on the 60 turns: δ = P ( R > B ) P ( R < B ) , estimated from the 10 × 10 pairwise comparisons of regulated versus baseline conversation means. Clustering was handled by reducing each conversation to a single mean before computing δ and by resampling the ten conversation identifiers with replacement (10,000 iterations), recomputing δ on each replicate. When one or both conditions exhibited zero variance (e.g., all values identical), Cohen’s d and related standardised mean differences are undefined and were reported as “not estimable” rather than as d = 0 . Statistical significance was set at α = 0.05 . All inferential statistics are reported as a transparent small-sample summary for this proof-of-concept design and should not be interpreted as evidence of population-level effects [81].

3.7.2. Reproducibility and Computational Environment

All statistical analyses were conducted in Python 3.9+ using scipy (v1.9.0) for statistical tests, numpy (v1.23.0) for effect size calculations, and matplotlib/seaborn for visualisations. Wilcoxon statistics are reported as T + (the sum of ranks of positive differences), not SciPy’s default min ( T + , T ) , which equals 0 under complete concordance. Experiments used the GPT-4 API (model: gpt-4-0613, OpenAI Python SDK v1.3.5) with temperature = 0.7 for assistant responses, and GPT-4-turbo (model: gpt-4-turbo) with temperature = 0.3 for evaluator responses. All random operations used fixed seed (Python: 42).
Complete system prompts, regulation templates, simulation transcripts, statistical analysis code, and analysis-ready tables are archived at Zenodo (https://doi.org/10.5281/zenodo.21995596).
A Python random seed does not make remote API outputs reproducible. OpenAI seed support was not used. The recorded snapshot (gpt-4-0613) and temperature settings are reported for provenance; they do not yield deterministic remote generations [82].

4. Results

The results follow the three research questions in Section 3.1. We first examine detection feasibility under boundary conditions and its degradation under mixed profiles (RQ1). We then compare regulated and baseline responses on personality-specific need fulfilment and generic conversational quality (RQ2). Cross-model rescoring and the two-rater human label audit subsequently assess the robustness and limits of these findings (RQ3). This sequence separates implementation fidelity from comparative outcomes and evaluation validity.
The primary analysis included 10 matched conversation pairs across two extreme personality profiles: Type A (OCEAN: + 1 , + 1 , + 1 , + 1 , + 1 ) and Type B (OCEAN: 1 , 1 , 1 , 1 , 1 ), with 5 conversations per profile. Each conversation contained 6 user–assistant exchanges, yielding 60 rated assistant turns in each condition (regulated: n = 60 ; baseline: n = 60 ). All scored metrics had complete coverage (0% missing ratings); Detection and Regulation applied only to the regulated condition. The balanced design supports conversation-level paired comparisons, although the small sample ( n = 10 pairs) and maximal-signal profiles limit their precision and generalisability.

4.1. Detection and Regulation Performance

Under extreme profiles, holistic detection was correct on 59/60 turns (98.3%; 1 Not Sure, 0 incorrect), and regulation was implemented on all 60 turns.
Dimension-level detection. While holistic detection was high under extreme profiles, dimension-level performance revealed substantial trait-specific variation (Table 2).
Table 4 reports the three-way counts that those accuracy, neutral, and sign-flip rates summarise.
Dimension-level accuracy ranged from 65.0% (Conscientiousness) to 100% (Neuroticism), with a macro mean of 83.3%. Neutral hedging (16.0%) was more common than sign inversion (0.7%), consistent with the conservative threshold τ = 0.7 : uncertain evidence was usually mapped to 0 rather than to the wrong pole. Neuroticism and Extraversion supplied the clearest cues, whereas Conscientiousness and Agreeableness more often failed to cross the threshold. The latter traits may therefore require stronger or more distinctive conversational evidence; the results do not support uniform detectability across OCEAN dimensions. Figure 6 decomposes each trait’s errors into neutral and sign-flip components, and Figure 5 compares trait-level performance across signal tiers.
Research Question 1 (Detection Feasibility): Both pre-specified boundary-condition criteria were met (holistic accuracy > 95 % : 98.3%; dimension-level macro- F 1 > 0.80 : 0.893). This establishes feasibility only for extreme simulated profiles. In the mixed-profile extension, dimension-level accuracy fell to 58.1% and neutral hedging rose to 41.1% (§Section 4.4), so the boundary-condition values are upper bounds rather than population estimates. The ground truth is also simulated: the human-annotated subset (§Section 4.5) provides limited criterion evidence on 66 turns from two independent raters, but pre-filled verify-and-revise labels and unadjudicated rater disagreements cannot yet substitute for validation against a psychometric instrument.
Figure 5. Descriptive per-trait detection accuracy across profile signal tiers. Circles denote extreme profiles (Phase 1A; macro mean 83.3%), squares denote subtle profiles (Phase 1C subset; macro mean 69.4%), and triangles denote mixed profiles (Phase 1C core set; macro mean 58.1%). Traits are categorical and are therefore shown as unconnected groups. Accuracy declined from extreme to mixed profiles for every trait, with Conscientiousness yielding the lowest value in each tier.
Figure 5. Descriptive per-trait detection accuracy across profile signal tiers. Circles denote extreme profiles (Phase 1A; macro mean 83.3%), squares denote subtle profiles (Phase 1C subset; macro mean 69.4%), and triangles denote mixed profiles (Phase 1C core set; macro mean 58.1%). Traits are categorical and are therefore shown as unconnected groups. Accuracy declined from extreme to mixed profiles for every trait, with Conscientiousness yielding the lowest value in each tier.
Preprints 229251 g005
Across signal tiers, the trait comparison reinforces the same boundary: strong performance on maximally separated profiles does not imply equally reliable recovery of individual dimensions when cues overlap or are weak.
Figure 6. Per-trait detection accuracy (Phase 1A, n = 60 turns, extreme profiles). Blue = correct detection; grey = neutral/hedge ( y ^ = 0 ); red = sign error. Macro- F 1 shown per trait. Dashed line = macro mean (83.3%). Results restricted to extreme ± 1 profiles; see Figure 10 for mixed-profile comparison. Ground truth: simulated profiles (LLM-generated), not human-validated psychometric assessments.
Figure 6. Per-trait detection accuracy (Phase 1A, n = 60 turns, extreme profiles). Blue = correct detection; grey = neutral/hedge ( y ^ = 0 ); red = sign error. Macro- F 1 shown per trait. Dashed line = macro mean (83.3%). Results restricted to extreme ± 1 profiles; see Figure 10 for mixed-profile comparison. Ground truth: simulated profiles (LLM-generated), not human-validated psychometric assessments.
Preprints 229251 g006
Personality-vector consistency. Figure 7 shows the 60 regulated-condition personality vectors. Extraversion, Type A Agreeableness, and Neuroticism showed high alignment with the intended profiles; Conscientiousness and Type B Agreeableness had more neutral assignments. Vectors were generally stable within conversations, with limited variation in sessions such as A-5 (Extraversion) and B-5 (Openness).
The vector trajectories therefore support implementation fidelity at the holistic level while locating the residual uncertainty in specific dimensions. Stability within conversations is consistent with cumulative evidence updating, but it is not evidence that every component of the inferred vector is psychometrically correct.

4.2. Regulated vs. Baseline Comparison

Primary outcome: Personality Needs Addressed. Table 5 presents the conversation-level statistical comparison of primary and secondary outcomes under the extreme profiles in Phase 1A.
At the conversation level, all 10 pairs favoured regulation on Personality Needs Addressed (exact Wilcoxon T + = 55.0 , two-sided p = 0.00195 ). The corresponding turn-level Yes rates, reported as descriptive only, were 60 / 60 (100%) versus 5 / 60 (8.3%) at baseline. The corresponding Cliff’s δ = 1.000 indicates complete separation, but the cluster bootstrap is non-estimable under unanimous concordance, so we report the Clopper–Pearson 95% CI [0.692, 1.000] as the honest small-sample bound. This contrast is best interpreted as a binary architectural difference under extreme synthetic profiles rather than a portable dose–response estimate. The primary study used scenario-matched but independently generated user turns, so this headline contrast has weaker causal attribution than the mixed-profile extension, which used input replay and found an attenuated but directionally consistent Personality Needs effect (Section 4.4).
Emotional Tone was not estimable because both arms had zero variance; Cohen’s d is undefined under that condition and is not reported as d = 0 . Cross-model rescoring identifies this as an evaluator ceiling artifact (Section 4.3). Relevance & Coherence remained near ceiling (regulated M = 2.00 , baseline M = 1.97 ; exact Wilcoxon T + = 1.0 , two-sided p = 1.000 ; Cliff’s δ = 0.100 , cluster-bootstrap 95% CI includes 0).
Secondary outcomes: basic conversational quality. Emotional Tone Appropriateness assessed whether a response conveyed an appropriate affective stance, including warmth, empathy, and non-judgemental tone, on the Yes/Not Sure/No rubric. Relevance & Coherence assessed topical relevance and logical consistency with the dialogue history. These secondary outcomes capture generic conversational quality expected from either condition, rather than personality-specific adaptation.
The original evaluator assigned Emotional Tone a common ceiling in both arms ( M = 2.00 , S D = 0.00 ), making an effect non-estimable rather than establishing equivalence. Cross-model rescoring later exposed this as evaluator-specific compression. Relevance & Coherence was also near ceiling, with only thin signal. The secondary outcomes therefore support an absence of observed degradation, but they cannot show that adaptation improves already saturated generic-quality measures. Their near-ceiling values also limit sensitivity to potential trade-offs; evaluator choice affects both absolute levels (Table 6) and the apparent Emotional Tone effect.
Selective enhancement and implementation fidelity. The pattern is selective: the regulated stack changed the personality-sensitive outcome while generic conversational quality remained at or near ceiling. This is consistent with an additive adaptation layer that changes framing, pacing, structure, and engagement strategy without displacing baseline coherence or supportive tone. It does not establish that all relevant dimensions of conversational quality were preserved, because the evaluated secondary rubrics were narrow and ceiling-constrained.
Implementation fidelity was high for the regulated condition: holistic detection was correct on 59/60 turns, and the regulation module applied an adjustment on all 60 turns. However, dimension-level accuracy was lower (macro mean 83.3%; range 65.0–100%). The primary contrast therefore reflects end-to-end operation under idealised profiles, not uniformly accurate recovery of every trait, and holistic success may partly rely on redundant cues across dimensions.
Research Question 2 (Regulation Contrast): Under boundary conditions, the turn-level Personality Needs Yes rate rose from 8.3% ( 5 / 60 ) to 100% ( 60 / 60 ); all 10 conversation pairs favoured regulation (exact Wilcoxon T + = 55.0 , two-sided p = 0.00195 ) without an observed reduction in the two generic-quality outcomes. Complete separation ( δ = 1.000 ; bootstrap non-estimable; Clopper–Pearson bound reported) therefore summarises a binary architectural contrast rather than a graded treatment effect. The Personality Needs direction persisted under Sonnet 4.6 and across three mixed-profile judges, despite weaker detection. The magnitude remains an upper-bound, LLM-rated architectural contrast and is expected to attenuate under realistic profiles and independently human-rated outcomes. The strongest causal design, identical input replay across arms, appears in Section 4.4 and yields attenuated but directionally consistent Personality Needs effects, locating the Phase 1A result as a feasibility upper bound rather than the primary causal claim.

4.3. Cross-Model Re-Evaluation of Existing Outputs

Study A outputs were re-scored with Sonnet 4.6 to test evaluator-family dependence. The procedure was blinded: the evaluator received only the user turn, assistant reply, and outcome rubric, without the personality profile, regulation directives, generator, or condition label. Items were presented in alternating condition order (regulated, then baseline, within each turn). Cross-family rescoring used temperature 0.0 on the same 120 outputs that the original GPT-family evaluator had scored at temperature 0.3, providing a model×temperature check of the Personality Needs direction rather than a full factorial sensitivity grid. For a directional replication claim, item-level agreement need not be strong if both evaluators preserve the regulated–baseline ordering; different severity thresholds are compatible with the same effect direction, so fair κ does not by itself invalidate the contrast. Table 6 compares these ratings with those of the original GPT-family evaluator across n = 120 items per metric.
Sonnet 4.6 preserved the Personality Needs contrast (Cliff’s δ = 0.82 [0.65, 0.99]) but had only fair item-level agreement with the original evaluator ( κ = 0.219 ; [77]). It found an Emotional Tone lift (Cliff’s δ = 0.50 [0.25, 0.86]) absent from the original all-Yes ratings, identifying a ceiling artifact. Relevance & Coherence remained near null (Cliff’s δ = 0.01 ; p = 0.742 ). The replicated quantity is therefore the direction of the regulated–baseline contrast, not interchangeability of individual labels. This distinction provides cross-family internal-validity evidence while leaving external criterion validity unresolved. Figure 8 shows item-level agreement, while Figure 9 compares effect sizes across evaluators and profile tiers.
The condition-specific agreement matrices should consequently be read as diagnostics of rating correspondence, not as the main replication test. Their modest item-level agreement is compatible with preservation of an aggregate effect direction because different label distributions can still yield the same ordering between conditions.
Across evaluators and profile tiers, the forest plot supports a persistent Personality Needs direction but also shows why the Phase 1A point estimate should not be exported as a realistic effect magnitude. Evaluator choice changes both item labels and the apparent secondary-outcome effect.

4.4. Generalization to Mixed Personality Profiles

The mixed-profile inventory comprised six configurations: Friendly Conventional (FC), Relaxed Creative (RC), Temperamental Uninhibited (TU), Warm Skeptic (WS), Anxious Achiever (AA), and Reserved Organizer (RO). FC, RC, and TU were adapted from the configurations reported by Rentfrow et al. [74]; WS, AA, and RO were constructed to add contrasting trait combinations. Two single-salient profiles represented subtler signals: Faint Extravert (FE; Extraversion + 1 only) and Anxious Only (AN; Neuroticism 1 only). Two extreme anchors (Type A and Type B) were retained for reference. Input replay kept user turns identical across regulated and baseline arms, isolating the response-policy difference from variation in user prompts. Table 7 and Table 8 report detection and regulation performance for the core profiles.
Detection weakened when profiles combined less separable cues. The dominant failure mode was neutral hedging rather than confident sign inversion: the detector frequently withheld a directional assignment when evidence did not cross its threshold. This preserves caution but reduces trait coverage and limits downstream claims about profile recovery. The extreme-anchor accuracy of 74.4% ( n = 36 turns) is a Phase 1C reference under input replay, not a re-score of the Phase 1A n = 60 sample (83.3%); the two figures are therefore not a test of detector instability on identical user turns.
For mixed profiles, per-trait accuracy fell to 58.1% and neutral hedging rose to 41.1% (Figure 10). Despite weaker detection, all three judges retained a Personality Needs advantage ( Δ = + 0.79 to + 1.39 ; all p 0.002 ). Relevance & Coherence showed no consistent advantage. Regulation can therefore remain directionally useful when trait recovery is imperfect, but this persistence does not rescue the extreme-profile detection claim. The inventory approximates more realistic combinations than all- + 1 /all- 1 anchors; it is not a population sample, a psychometrically validated profile distribution, or evidence for unrestricted real-world generalisation.
Figure 10. Detection accuracy degradation by signal tier. Blue: mean per-trait accuracy; grey hatched: neutral/hedge rate. The Mixed-tier flip rate (8.1%) is annotated above the bars. Phase 1A extreme macro mean (83.3%) shown for comparison. Arrow: 25.2 percentage-point drop from extreme to mixed tier.
Figure 10. Detection accuracy degradation by signal tier. Blue: mean per-trait accuracy; grey hatched: neutral/hedge rate. The Mixed-tier flip rate (8.1%) is annotated above the bars. Phase 1A extreme macro mean (83.3%) shown for comparison. Arrow: 25.2 percentage-point drop from extreme to mixed tier.
Preprints 229251 g010
The degradation plot makes the RQ1 boundary explicit: extreme-profile accuracy is a maximal-signal benchmark, whereas mixed and subtle tiers better expose abstention under ambiguous cues. Subsequent claims therefore concern persistence of a comparative regulation effect, not reliable identification of realistic users’ full OCEAN vectors.

4.5. Human Validation

Two independent non-author raters (Methods; Acknowledgments) completed a blinded verify-and-revise audit of the same Phase 2 OCEAN subset ( n = 66 turns). Raters were blind to condition, generator, intended profile, and each other’s judgements. Rather than assigning labels de novo, they inspected pre-filled trait labels and retained or revised each proposed cell. Against the machine proposals, the primary rater revised 10.3% of cells and agreed on 89.7% (mean weighted Cohen’s κ = 0.844 ; Table 9); the second rater revised 19.4% and agreed on 80.6% (mean weighted κ = 0.725 ). Because labels were pre-filled, human–machine agreement and κ remain anchoring-inflated upper bounds on detection-label convergence; this audit does not address regulation-outcome validity. On the full overlap, human–human cell agreement was 80.9% (mean linear κ = 0.731 ; unweighted κ = 0.689 ), with Extraversion highest and Conscientiousness lowest. Under a consensus-only adjudication rule, agreeing cells form the consensus label set and the remaining 19.1% of cells are left unresolved (no machine tie-break); they are not treated as ground truth.
Exploratory scoring against the revised primary labels, without detector retuning, favoured zero-shot detection (macro- F 1 = 0.815 ) over the E/A specialist ( F 1 = 0.709 ) and evidence-gated ( F 1 = 0.584 ) arms. Dual-rater reliability is now reportable for the detection-label audit; remaining evaluation boundaries are summarised in Section 5.2.1.
Research Question 3 (Evaluation Validity): Cross-model rescoring supports the direction of the primary contrast, although item-level agreement was only fair, and the two-rater verify-and-revise audit provides convergent evidence for detection labels. These checks strengthen the comparative proof-of-concept claim while leaving adjudicated psychometric ground truth and human-perceived clinical benefit for later stages of the validation pathway.

4.6. Qualitative Examples Demonstrating Personality Adaptation

Figure 11 compares response strategies for Type B (Panel A) and Type A (Panel B). Regulated responses alter framing, pacing, structure, and engagement strategy according to the detected profile, whereas baseline responses remain supportive but non-personalised. The examples make the selective-enhancement mechanism visible: the main contrast lies in how support is organised around personality-linked needs, not in basic fluency, coherence, or politeness.
Across both examples, the qualitative contrast mirrors the quantitative pattern: baseline responses can satisfy generic-quality rubrics yet fail the personality-needs criterion because they do not implement a coherent adaptation policy. The excerpts provide an interpretable account of the observed rating difference, but they are illustrative rather than independent validation and do not establish user preference or clinical benefit.

5. Discussion

The D–R–E architecture executed with high fidelity under boundary conditions (holistic detection 98.3%; regulation adherence 100%) and produced selective enhancement: the turn-level Personality Needs Yes rate rose from 8.3% ( 5 / 60 ) to 100% ( 60 / 60 ), and all 10 conversation pairs favoured regulation, without an observed drop in generic conversational quality. Mixed-profile testing and cross-family rescoring then bounded that result: detection weakens under realistic cues, and the Phase 1A magnitude is an upper bound rather than a portable effect size.
Selective enhancement matters for deployment because it supports an augmentation model: keep a strong generic supportive core, then add trait-conditioned framing only when evidence for adaptation is sufficient. Ceiling effects on Emotional Tone Appropriateness and Relevance & Coherence indicate that regulation did not repair weak baseline behaviour; it changed pacing, structure, and interpersonal stance along Zurich Model domains of Security, Arousal, and Affiliation (Section 3.3.1). In practice, uncertain or conflicting trait signals should therefore trigger non-adaptation or conservative defaults rather than forced personalisation, because over-confident trait inference can stereotype users without improving alliance. Cross-family rescoring with Sonnet 4.6, Gemini 3.1 Flash-Lite, and DeepSeek V4 Flash preserves the Personality Needs direction, which supports treating the contribution as an orchestration architecture rather than a claim about a single 2024 GPT-4 snapshot.
The evidence hierarchy is accordingly clear: boundary-condition testing demonstrates implementation feasibility, mixed-profile testing limits generalisation, and the validation checks constrain claims about evaluator and outcome validity. Quantitative ratings came from a structured LLM-as-Judge procedure, and qualitative review was conducted by two contributing authors rather than independent third-party raters (Section 3.6.2). Cross-model rescoring and the two-rater verify-and-revise audit strengthen the comparative proof-of-concept claim for detection labelling, but they do not validate the regulation-outcome rubrics, adjudicate disagreeing label cells into ground truth, or link scores to downstream user outcomes. The evidence therefore supports architectural feasibility, not human-perceived clinical benefit.
The mixed-profile extension should be read within a narrow scope. It tests whether detection and the Personality Needs contrast survive less separable OCEAN configurations than the Phase 1A anchors, using six Rentfrow-inspired or constructed profiles plus two single-salient subtle cases under identical input replay (Section 4.4). It is not a population sample, a psychometrically validated trait distribution, or a human-outcome study. Within that scope, dimension-level accuracy fell to 58.1% and neutral hedging rose to 41.1%, while three cross-family judges retained a directionally consistent Personality Needs advantage. The practical limitation is therefore incomplete trait recovery under ambiguous cues, which can dilute regulation specificity even when an aggregate personalisation signal remains detectable. Without rebuilding the corpus, the immediate remedies are conservative non-adaptation when confidence stays below τ , explicit reporting of neutral rates alongside accuracy, and treating Phase 1A magnitudes as upper bounds. Longer-term solutions remain those in Section 5.2.2: denser realistic profile sampling, shuffled-mapping ablations, feature-attribution validation, and independent human rating of both detection labels and support outcomes.
These results align with personality-aware dialogue research [30,35]. The framework extends static or heuristic adaptation by updating trait estimates during dialogue and mapping them to Zurich Model guidance (Section 3.3.1). Whether that mapping improves user outcomes remains an empirical question for independent human evaluation.

5.1. Research and Design Implications

Trait-based regulation is a complementary capability, not universally superior conversation. Its target is response style in domains where individual differences matter: Security, Arousal, and Affiliation (Section 3.3.1). The study shows that this mapping can be implemented and detected by the rubric; it does not validate the Zurich Model itself. That would require comparison with competing mappings, a shuffled-mapping ablation, and expert psychologist review of weaker links (notably Openness→Arousal and Conscientiousness as a cross-cutting cue). Those controls were not run in the present revision. They are listed as planned future work (detection-only, regulation-only / oracle-profile, shuffled mapping, and a formal multi-expert mapping review), not as analyses already completed.
The modular D–R–E design supports targeted iteration: detection and regulation can be revised independently while their interfaces and PROMISE logging remain stable. This separation makes it possible to compare alternative trait estimators or motivational mappings without redesigning the full interaction architecture.
For practice-oriented systems, the findings suggest the following design principles:
  • Layered deployment: Add personality adaptation to a strong generic baseline.
  • Updateable personality state: Maintain persistent estimates and uncertainty across turns.
  • Theory-grounded mapping: Translate traits into behaviour through an explicit, testable framework.
  • Transparent logic: Log trait estimates, regulations, and outcomes.
  • Selective application: Adapt interactional style without indiscriminately changing content.
In mental-health applications, uncertainty should trigger neutral or conservative behaviour rather than confident profiling. Translation also requires consent, privacy protection, bias auditing, human oversight, and evaluation against patient-centred outcomes [83].

5.2. Strengths, Limitations, and Validation Pathway

The architecture’s main strengths are modularity and traceability of orchestration decisions. Detection, regulation, and evaluation can be revised independently, while PROMISE logs the state transitions and instructions behind each response. Cross-model rescoring and the human label audit expose evaluation weaknesses rather than relying on a single automated score. These features support reproducible iteration but do not establish clinical effectiveness.

5.2.1. Limitations and Barriers

No human users participated, so the results do not generalise to clinical benefit. The system is not evaluated as a therapeutic or clinical intervention. The simulation did not assess crisis detection, unsafe advice, hallucination, escalation behaviour, dependency risks, or clinically inappropriate reassurance. The primary profiles were extreme and synthetic; mixed-profile testing directly showed lower detection accuracy rather than resolving generalisability. In the mixed tier, neutral hedging reached 41.1% (Table 7), reducing the specificity of downstream regulation even though a judge-detectable Personality Needs advantage persisted (Table 8). Dialogues lasted six turns, leaving long-term stability unknown. In the primary study, paired conditions used scenario-matched but independently generated user turns, which weakens strict causal attribution to the regulation policy despite shared profiles, topics, and progression. The original generator and evaluator shared the GPT family, and cross-family rescoring preserved direction but not strong item-level agreement. PROMISE traceability shows which trait estimate drove each regulation, but the closed GPT-4 detector does not expose which linguistic evidence produced that estimate, and no feature-attribution analysis was performed. Author review was not independent. The human audit used two independent raters on detection labels only (pre-filled verify-and-revise; dual-rater reliability on the n = 66 overlap, mean linear κ = 0.731 ), under a consensus-only rule that leaves the remaining 19.1% of disagreeing cells unresolved. Importantly, that audit does not extend to the regulation-outcome rubrics (Personality Needs Addressed, Emotional Tone Appropriateness, Relevance & Coherence), which remain LLM-rated comparative scores rather than independently human-rated outcomes. Component ablations that would isolate detection, regulation, and mapping specificity were not included in this revision. English-language prompts and Western personality assumptions further limit cultural and linguistic transfer.
Deployment would also require explicit consent for personality inference, secure handling of inferred traits, safeguards against stereotyping, and evidence of equitable performance. Ambiguous signals should produce graceful non-adaptation rather than forced classification. Before human-facing deployment, safety evaluation must include crisis and unsafe-advice testing, hallucination and grounding audits, escalation procedures, dependency-risk assessment, informed consent, privacy protection, clinician oversight, and stopping rules.
If the same pipeline were hosted on open-weight local models, the PROMISE state machine would remain a lightweight control layer; the computational cost would be LLM inference. Each user turn issues five trait-specific detector calls (designed to run in parallel) plus one response-generation call over cumulative dialogue context. Hardware requirements, memory footprint, and per-turn complexity would then depend on the chosen local model, quantisation, batching, context length, and serving stack, and cannot be read from the closed GPT-4 API used here. On-premise or edge hosting for privacy would further constrain VRAM and concurrent-session throughput. We did not measure FLOPs, device memory, or local tokens-per-second; those benchmarks remain planned deployment work rather than results of this study.

5.2.2. Validation Pathway for AI-Augmented Personalisation

The limitations motivate a three-stage validation pathway:
1.
Broader simulation: sample realistic trait distributions, extend dialogues beyond 20 exchanges, compare alternative or shuffled regulation mappings, and test languages and cultures beyond the current setting. Evaluate an auditable, open-weight OCEAN detector using token- and sentence-level SHapley Additive exPlanations (SHAP) [84] or Integrated Gradients [85], interpretable linguistic features, and principal component analysis of contextual embeddings, with the resulting components mapped to each trait by an interpretable regression model. Test attribution stability under paraphrasing and cue-removal perturbations.
2.
Human study: use an IRB-approved protocol and independent raters to annotate both OCEAN labels and the trait-relevant evidence spans in each dialogue, allowing detector attributions to be compared with human-identified evidence. Include validated personality instruments, user-reported support and alliance outcomes, qualitative interviews, and safety monitoring covering crisis detection, unsafe advice, hallucination, escalation behaviour, dependency risk, and clinically inappropriate reassurance. Sample size should follow the prospective calculations in Section 3.1.2 and pilot variance.
3.
Deployment evaluation: test longitudinal outcomes, workflow integration, and equity across sites and demographic groups; for clinical hosting, benchmark open-weight on-premise or edge deployments for latency, memory, and privacy constraints. Extend PROMISE logging to retain each prediction’s influential text spans, attribution strength, confidence, and dialogue context, and augment the existing confidence gate by also withholding adaptation when attributions are unstable across perturbations.
This sequence separates architectural testing from claims about user benefit and real-world effectiveness.

6. Conclusions

The D–R–E architecture combined turn-by-turn Big Five detection with Zurich Model regulation. Under boundary conditions, implementation fidelity was high and the regulated assistant showed selective enhancement: the turn-level Personality Needs Yes rate rose without an observed reduction in Emotional Tone Appropriateness or Relevance & Coherence. Mixed-profile testing reduced dimension-level accuracy to 58.1%, locating the extreme-profile result as an upper bound. A two-rater verify-and-revise audit supported the detection-labelling step on an n = 66 overlap.
These results describe a reproducible architecture for controlled simulation rather than clinical effectiveness. The system is not evaluated as a therapeutic or clinical intervention, and the Zurich mapping is implemented as a design choice rather than validated as psychological theory. Remaining evaluation boundaries are discussed in Section 5.2.1. A powered human study with validated personality measures, independent outcome raters, user-reported outcomes, longitudinal follow-up, and bias audits remains the next empirical step (Section 3.1.2).
Before clinical hosting, two computational studies are still required. Hardware optimisation is needed for on-premise or edge deployment of open-weight models, covering latency, memory, throughput, and privacy constraints that cannot be inferred from the closed GPT-4 API used here. Mathematical XAI should be integrated for detector interpretability, using token- and sentence-level SHapley Additive exPlanations (SHAP) [84] or Integrated Gradients [85], so that trait assignments can be traced to conversational evidence rather than remaining black-box labels.
Study procedures.
OpenAI GPT-4 (gpt-4-0613) generated the synthetic user and assistant turns and performed turn-by-turn Big Five detection. GPT-4-turbo provided the primary structured LLM-as-judge scores. Cross-family rescoring used Claude Sonnet 4.6 on the boundary-condition outputs and Claude Sonnet 4.6, Gemini 3.1 Flash-Lite, and DeepSeek V4 Flash on the mixed-profile set. Model identifiers, sampling settings, prompts, and validation procedures are reported in the Materials and Methods. The authors designed the analyses, executed the statistical code in Python, and checked the synthetic data and reported results using the procedures described in the paper. GenAI did not autonomously produce the reported statistical results or alter the data. Figures are author-constructed schematic and statistical plots, not GenAI-generated research images; Figure 11 displays synthetic dialogue excerpts that are study data.
Research process and manuscript preparation.
OpenAI GPT-5.6 and Anthropic Claude (Opus/Sonnet 4.x), accessed via author writing interfaces during 2025–2026, were used to suggest research questions and support the exploration of research ideas, to edit and debug analysis code, to edit author-written manuscript text, and to check consistency throughout the writing process. AI-generated suggestions served as inputs for consideration rather than as autonomous scientific decisions. The authors critically reviewed all suggestions before incorporating them and verified the manuscript’s factual accuracy, references, scientific claims, and interpretations. The final manuscript reflects the authors’ own understanding, structure, and synthesis.

Supplementary Materials

The following supporting information can be downloaded at the website of this paper posted on Preprints.org. The following supporting information is packaged with this V8.4 deliverable (supplementary/) and mirrored under data/ / scripts/ where noted. The archival copy is at Zenodo (https://doi.org/10.5281/zenodo.21995596): Supplementary File S1: Complete system prompts for personality detection (5 Big Five trait detectors) Supplementary File S2: Regulation templates for Zurich Model mapping (arousal, security, affiliation) Supplementary File S3: Evaluator GPT system prompt and scoring matrix Supplementary File S4: Pairing manifest for the primary study (pairing_manifest.csv; profile–persona pair IDs and identical-user-turn indicator) Supplementary File S5: Complete simulation transcripts (20 conversations, 120 dialogue turns) Supplementary File S6: Statistical analysis code (Python/Jupyter notebooks for effect sizes, paired tests, weighted scoring, and visualisations) Supplementary File S7: Human validation freeze — two-rater audit with complete primary subset ( n = 66 Phase 2 turns; verify-and-revise labels, revision/agreement tables, exploratory detector scores); see supplementary/human_validation_n66/ Supplementary Figure S1: Coverage of scored outcomes (complete coverage; 0% missing ratings on Detection, Regulation, Personality Needs, Emotional Tone, and Relevance & Coherence) Supplementary Figure S2: Personality needs YES-rate by conversation (demonstrating consistent improvement across all 10 conversation pairs) Supplementary Figure S3: Rating distribution raw counts (YES/NOT SURE/NO composition with value annotations)

Author Contributions

The author contributions follow the CRediT taxonomy. Conceptualization: S. Devdas, G. Lu, A. de Spindler and M. Stieger; methodology: S. Devdas, J. Duojie, G. Lu, A. de Spindler and M. Stieger; software: S. Devdas and A. de Spindler; validation (analysis and manuscript checks): J. Duojie and S. Devdas; formal analysis: J. Duojie; data curation: J. Duojie and S. Devdas; writing—original draft preparation: J. Duojie; writing—review and editing: M. Stieger, G. Lu and A. de Spindler; visualisation: J. Duojie; project administration: J. Duojie; supervision: G. Lu and M. Stieger. Independent human label auditing was performed by non-author raters acknowledged below. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable. This study involved no human or animal subjects and used synthetic, non-identifiable, AI-generated dialogues. Contributing authors (Samuel Devdas and Jiahua Duojie) provided qualitative evaluation notes on these synthetic materials, and independent raters contributed a blinded label audit of expressed personality cues in the same synthetic transcripts; no patient data were involved. Future human subject studies will require full institutional review board approval.

Data Availability Statement

Supplementary materials including system prompts, regulation templates, simulation transcripts, statistical analysis code, conversation-pair identifiers, analysis-ready scored matrices, and the dual-rater human-validation freeze are archived at Zenodo (https://doi.org/10.5281/zenodo.21995596).

Acknowledgments

The authors thank Dr. Aliaksei Andrushevich (Senior Research Associate, Engineering & Architecture, iHomeLab; human–computer interaction and ambient assisted living; https://www.hslu.ch/de-ch/hochschule-luzern/ueber-uns/personensuche/profile/?pid=934) and Laura Ebbinghaus (Senior Research Associate, Business, Institute of Communication and Marketing; marketing and economic psychology; https://www.hslu.ch/de-ch/hochschule-luzern/ueber-uns/personensuche/profile/?pid=5015) for their careful contribution as independent raters in the human label audit. The authors also acknowledge the PROMISE framework developers [27] whose open architecture informed our system design.

Conflicts of Interest

The authors declare no conflicts of interest. Qualitative review of evaluation matrices was performed by contributing authors (Samuel Devdas and Jiahua Duojie). Independent rating of expressed personality cues was performed by Dr. Aliaksei Andrushevich and Laura Ebbinghaus, who are acknowledged above and were not authors of this manuscript.

Use of Artificial Intelligence

This statement discloses the use of generative artificial intelligence (GenAI) in the study and during manuscript preparation. GenAI tools are not authors of this paper. The authors critically reviewed and, where necessary, revised the AI-assisted text, code, and suggestions incorporated into this work and take full responsibility for the content of this publication. The evaluation and validation of experimental model outputs, including their limitations, are reported in the Materials and Methods and Discussion.

References

  1. Dingler, T.; Kwasnicka, D.; Wei, J.; Gong, E.; Oldenburg, B. The use and promise of conversational agents in digital health. Yearb. Med. Inform. 2021, 30, 191–199. [Google Scholar] [CrossRef] [PubMed]
  2. Laranjo, L.; Dunn, A.G.; Tong, H.L.; Kocaballi, A.B.; Chen, J.; Bashir, R.; Surian, D.; Gallego, B.; Magrabi, F.; Lau, A.Y. Conversational agents in healthcare: a systematic review. J. Am. Med. Inform. Assoc. 2018, 25, 1248–1258. [Google Scholar] [CrossRef] [PubMed]
  3. Vasiliu, L.; Cortis, K.; McDermott, R.; Kerr, A.; Peters, A.; Hesse, M.; Hagemeyer, J.; Belpaeme, T.; McDonald, J.; Villing, R. CASIE – Computing affect and social intelligence for healthcare in an ethical and trustworthy manner. Paladyn J. Behav. Robot. 2021, 12, 437–453. [Google Scholar] [CrossRef]
  4. Kocaballi, A.B.; Sezgin, E.; Clark, L.; Carroll, J.M.; et al. Design and Evaluation Challenges of Conversational Agents in Health Care and Well-being: Selective Review Study. J. Med. Internet Res. 2022, 24, e38525. [Google Scholar] [CrossRef] [PubMed]
  5. Kaddour, J.; Harris, J.; Mozes, M.; Bradley, H.; Raileanu, R.; McHardy, R. Challenges and Applications of Large Language Models. arXiv 2023, arXiv:2307.10169. [Google Scholar]
  6. Rapp, A.; Curti, L.; Boldi, A. The human side of human-chatbot interaction: A systematic literature review of ten years of research on text-based chatbots. Int. J. Hum.-Comput. Stud. 2021, 151, 102630. [Google Scholar] [CrossRef]
  7. Han, E.; Yin, D.; Zhang, H. Chatbot Empathy in Customer Service: When It Works and When It Backfires. In Proceedings of the SIGHCI 2022 Proceedings, 2022; Vol. 1. [Google Scholar]
  8. Juquelier, A.; Poncin, I.; Hazée, S. Empathic chatbots: A double-edged sword in customer experiences. J. Bus. Res. 2025, 188, 115074. [Google Scholar] [CrossRef]
  9. Seitz, L. Artificial empathy in healthcare chatbots: Does it feel authentic? Comput. Hum. Behav. Artif. Hum. 2024, 2, 100067. [Google Scholar] [CrossRef]
  10. Abernethy, A.; Adams, L.; Barrett, M.; Bechtel, C.; Brennan, P.; Butte, A.; Faulkner, J.; Fontaine, E.; Friedhoff, S.; Halamka, J.; et al. The Promise of Digital Health: Then, Now, and the Future. NAM Perspect. 2022. [Google Scholar] [CrossRef] [PubMed]
  11. Goetz, L.H.; Schork, N.J. Personalized medicine: Motivation, challenges, and progress. Fertil. Steril. 2018, 109, 952–963. [Google Scholar] [CrossRef] [PubMed]
  12. O’Keefe, D.J. Persuasion: Theory and Research; SAGE Publications, Inc., 2016. [Google Scholar]
  13. Milne-Ives, M.; de Cock, C.; Lim, E.; Shehadeh, M.H.; de Pennington, N.; Mole, G.; Normando, E.; Meinert, E. The Effectiveness of Artificial Intelligence Conversational Agents in Health Care: Systematic Review. J. Med. Internet Res. 2020, 22, e20346. [Google Scholar] [CrossRef] [PubMed]
  14. Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E.H.; Le, Q.V.; Zhou, D. Chain-of-thought prompting elicits reasoning in large language models. In Proceedings of the 36th International Conference on Neural Information Processing Systems, Red Hook, NY, USA, 2022; p. NIPS ’22. [Google Scholar]
  15. Ayers, J.W.; Poliak, A.; Dredze, M.; Leas, E.C.; Zhu, Z.; Kelley, J.B.; Faix, D.J.; Goodman, A.M.; Longhurst, C.A.; Hogarth, M.; et al. Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum. JAMA Intern. Med. 2023, 183, 589–596. [Google Scholar] [CrossRef] [PubMed]
  16. Färber, A.; de Spindler, A.; Moser, A.; Schwabe, G. Closing the Loop for Patients with Chronic Diseases - from Problems to a Solution Architecture. In Proceedings of the The 11th IEEE International Conference on Healthcare Informatics, 2023; pp. 1–11. [Google Scholar] [CrossRef]
  17. Färber, A.; Schwabe, C.; Stalder, P.H.; Dolata, M.; Schwabe, G. Physicians’ and Patients’ Expectations From Digital Agents for Consultations: Interview Study Among Physicians and Patients. JMIR Hum. Factors 2024, 11, e49647. [Google Scholar] [CrossRef] [PubMed]
  18. Staehelin, D.; Dolata, M.; Stöckli, L.; Schwabe, G. How Patient-Generated Data Enhance Patient-Provider Communication in Chronic Care: Field Study in Design Science Research. JMIR Med. Inform. 2024, 12, e57406. [Google Scholar] [CrossRef] [PubMed]
  19. Strubell, E.; Ganesh, A.; McCallum, A. Energy and Policy Considerations for Deep Learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics; Korhonen, A., Traum, D., Màrquez, L., Eds.; Association for Computational Linguistics, 2019; pp. 3645–3650. [Google Scholar]
  20. Ding, N.; Qin, Y.; Yang, G.; Wei, F.; Yang, Z.; Su, Y.; Hu, S.; Chen, Y.; Chan, C.M.; Chen, W.; et al. Parameter-efficient fine-tuning of large-scale pre-trained language models. Nat. Mach. Intell. 2023, 5, 220–235. [Google Scholar] [CrossRef]
  21. Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. LoRA: Low-Rank Adaptation of Large Language Models. arXiv 2021, arXiv:2106.09685. [Google Scholar]
  22. Korzynski, P.; Mazurek, G.; Krzypkowska, P.; Kurasinski, A. Artificial intelligence prompt engineering as a new digital competence: Analysis of generative AI technologies such as ChatGPT. Entrep. Bus. Econ. Rev. 2023. [Google Scholar] [CrossRef]
  23. White, J.; Fu, Q.; Hays, S.; Sandborn, M.; Olea, C.; Gilbert, H.; Elnashar, A.; Spencer-Smith, J.; Schmidt, D. A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT. arXiv 2023, arXiv:2302.11382. [Google Scholar]
  24. Fernando, C.; Banarse, D.; Michalewski, H.; et al. Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution. arXiv 2023, arXiv:2309.16797. [Google Scholar]
  25. Liu, P.; Yuan, W.; Fu, J.; Jiang, Z.; Hayashi, H.; Neubig, G. Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. ACM Comput. Surv. 2023, 55, 195:1–195:35. [Google Scholar] [CrossRef]
  26. Hou, Y.; Dong, H.; Wang, X.; Li, B.; Che, W. MetaPrompting: Learning to Learn Better Prompts. In Proceedings of the 29th International Conference on Computational Linguistics; Calzolari, N., Huang, C.R., Kim, H., Pustejovsky, J., Wanner, L., Choi, K.S., Ryu, P.M., Chen, H.H., Donatelli, L., Ji, H., et al., Eds.; International Committee on Computational Linguistics, 2022; pp. 3251–3262. [Google Scholar]
  27. Wu, W.; Heierli, J.; Meisterhans, M.; Moser, A.; Färber, A.; Dolata, M.; Gavagnin, E.; de Spindler, A.; Schwabe, G. PROMISE: A Framework for Model-Driven Stateful Prompt Orchestration. In Proceedings of the Intelligent Information Systems: CAiSE Forum 2024, Limassol, Cyprus, June 3–7, 2024, Proceedings; Islam, S., Sturm, A., Eds.; Springer International Publishing; Lecture Notes in Business Information Processing; 2024; Vol. 520. [Google Scholar]
  28. Alisamir, S.; Ringeval, F. On the Evolution of Speech Representations for Affective Computing: A Brief History and Critical Overview. IEEE Signal Process. Mag. 2021, 38, 12–21. [Google Scholar] [CrossRef]
  29. Alsharekh, M. Facial Emotion Recognition in Verbal Communication Based on Deep Learning. Electronics 2022, 11, 2568. [Google Scholar]
  30. Mairesse, F.; Walker, M. Controlling User Perceptions of Linguistic Style: Trainable Generation of Personality Traits. Comput. Linguist. 2011, 37, 455–488. [Google Scholar] [CrossRef]
  31. Ta, V.; Griffith, C.; Boatfield, C.; Wang, X.; Civitello, M.; Maffei, H. User Experiences of Social Support from Companion Chatbots in Everyday Contexts: Thematic Analysis. J. Med. Internet Res. 2020, 22, e16235. [Google Scholar] [CrossRef] [PubMed]
  32. Broadbent, E. ElliQ, an AI-Driven Social Robot to Alleviate Loneliness: Progress and Lessons Learned. JAR Life 2024, 13, 22–28. [Google Scholar] [CrossRef] [PubMed]
  33. Shah, S. Effectiveness of Digital Technology Interventions to Reduce Loneliness in Adults: A Protocol for a Systematic Review and Meta-Analysis. BMJ Open 2019, 9, e029324. [Google Scholar] [CrossRef] [PubMed]
  34. Shah, S. Evaluation of the Effectiveness of Digital Technology Interventions to Reduce Loneliness in Older Adults: Systematic Review and Meta-Analysis. J. Med. Internet Res. 2021, 23, e24712. [Google Scholar] [CrossRef] [PubMed]
  35. Zhou, M. Designing Effective Interview Chatbots: Automatic Chatbot Profiling and Design Suggestion Generation. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’21), 2021. [Google Scholar]
  36. Dong, T.; Liu, F.; Wang, X.; Jiang, Y.; Zhang, X.; Sun, X. EmoAda: A Multimodal Emotion Interaction and Psychological Adaptation System. In Proceedings of the Conference on Multimedia Modeling, Cham, Switzerland, 2024. [Google Scholar]
  37. Quirin, M.; Malekzad, F.; Paudel, D.; Knoll, A.; Mirolli, M. Dynamics of Personality: The Zurich Model of Motivation Revived, Extended, and Applied to Personality. J. Pers. 2023, 91, 928–946. [Google Scholar] [CrossRef] [PubMed]
  38. Berlyne, D. Conflict, Arousal, and Curiosity; McGraw-Hill: New York, NY, USA, 1960. [Google Scholar]
  39. Hebb, D. The Organization of Behavior: A Neuropsychological Theory; Wiley: New York, NY, USA, 1949. [Google Scholar]
  40. Bischof, N. Das Rätsel Ödipus: Die biologischen Wurzeln des Urkonfliktes; Piper: Munich, Germany, 1985. [Google Scholar]
  41. Bischof, N. Untersuchungen zur Systemanalyse der sozialen Motivation I: Die Tantalus-Situation. Z. Psychol. 1993, 201, 5–43. [Google Scholar]
  42. Bickmore, T.; Picard, R. Establishing and Maintaining Long-Term Human-Computer Relationships. ACM Trans. Comput.-Hum. Interact. 2005, 12, 293–327. [Google Scholar] [CrossRef]
  43. Zheng, Z.; Liao, L.; Deng, Y.; Nie, L. Building Emotional Support Chatbots in the Era of LLMs. arXiv 2023, arXiv:2308.11584. [Google Scholar]
  44. Chen, K.; Kang, X.; Lai, X.; Ni, Z. Enhancing Emotional Support Capabilities of Large Language Models through Cascaded Neural Networks. In Proceedings of the 2023 4th International Conference on Computer, Big Data and Artificial Intelligence (ICCBD+AI), 2023; pp. 318–326. [Google Scholar]
  45. Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A. Training Language Models to Follow Instructions with Human Feedback. In Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS ’22), 2022; pp. 27730–27744. [Google Scholar]
  46. Abbasian, M.; Azimi, I.; Rahmani, A.; Jain, R. Conversational Health Agents: A Personalized LLM-Powered Agent Framework. arXiv 2023, arXiv:2310.02293. [Google Scholar]
  47. Sorino, P. ARIEL: Brain-Computer Interfaces Meet Large Language Models for Emotional Support Conversation. In Proceedings of the 32nd ACM Conference on User Modeling, Adaptation and Personalization (UMAP ’24), 2024. [Google Scholar]
  48. Dongre, P. Physiology-Driven Empathic Large Language Models (EmLLMs) for Mental Health Support. In Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (CHI EA ’24), 2024. [Google Scholar]
  49. Zhang, H.; Chen, Y.; Wang, M.; Feng, S. FEEL: A Framework for Evaluating Emotional Support Capability with Large Language Models. arXiv 2024, arXiv:2403.15699. [Google Scholar]
  50. Zheng, L.; Chiang, W.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. arXiv 2023, arXiv:2306.05685. [Google Scholar]
  51. Kim, S.; Lee, S.; Kim, J.; Yoo, H.; Seol, Y.; Kim, S.; Cho, H.; Choi, S.; Kim, S. Evaluating LLM-as-a-judge with Professional Human Ratings. arXiv 2024, arXiv:2402.18139. [Google Scholar]
  52. Montgomery, D. Design and Analysis of Experiments, 9th ed. ed.; Wiley: Hoboken, NJ, USA, 2017. [Google Scholar]
  53. Collins, L.; Dziak, J.; Li, R. Design of Experiments with Multiple Independent Variables: A Resource Management Perspective on Complete and Fractional Factorial Designs. Psychol. Methods 2009, 14, 202–224. [Google Scholar] [CrossRef] [PubMed]
  54. Carroll, C.; Patterson, M.; Wood, S.; Booth, A.; Rick, J.; Balain, S. A Conceptual Framework for Implementation Fidelity. Sci. 2007, 2, 40. [Google Scholar] [CrossRef] [PubMed]
  55. Faul, F.; Erdfelder, E.; Buchner, A.; Lang, A. Statistical Power Analyses Using G*Power 3.1: Tests for Correlation and Regression Analyses. Behav. Res. Methods 2009, 41, 1149–1160. [Google Scholar] [CrossRef] [PubMed]
  56. Schulz, K.; Altman, D.; Moher, D.; Group, C. CONSORT 2010 Statement: Updated Guidelines for Reporting Parallel Group Randomised Trials. BMJ 2010, 340, c332. [Google Scholar] [CrossRef] [PubMed]
  57. Chan, A.; Tetzlaff, J.; Gøtzsche, P.; Altman, D.; Mann, H.; Berlin, J.; Dickersin, K.; Hróbjartsson, A.; Schulz, K.; Parulekar, W. SPIRIT 2013 Statement: Defining Standard Protocol Items for Clinical Trials. Ann. Intern. Med. 2013, 158, 200–207. [Google Scholar] [CrossRef] [PubMed]
  58. World Medical Association. World Medical Association Declaration of Helsinki: Ethical Principles for Medical Research Involving Human Subjects. JAMA 2013, 310, 2191–2194. [Google Scholar] [PubMed]
  59. D’Alfonso, S. AI in Mental Health. Curr. Opin. Psychol. 2020, 36, 112–117. [Google Scholar] [CrossRef] [PubMed]
  60. Torous, J.; Myrick, K.; Rauseo-Ricupero, N.; Firth, J. Digital Mental Health and COVID-19: Using Technology Today to Accelerate the Curve on Access and Quality Tomorrow. JMIR Ment. Health 2020, 7, e18848. [Google Scholar] [CrossRef] [PubMed]
  61. Alberts, L.; Keeling, G.; McCroskery, A. Should Agentic Conversational AI Change How We Think about Ethics? Characterising an Interactional Ethics Centred on Respect. arXiv 2024, arXiv:2401.09187. [Google Scholar]
  62. Costa, P.; McCrae, R.R. Revised NEO Personality Inventory (NEO-PI-R) and NEO Five-Factor Inventory (NEO-FFI): Professional Manual; Psychological Assessment Resources: Odessa, FL, USA, 1992. [Google Scholar]
  63. John, O.; Naumann, L.; Soto, C. Paradigm Shift to the Integrative Big Five Trait Taxonomy: History, Measurement, and Conceptual Issues. In Proceedings of the Handbook of Personality: Theory and Research; New York, NY, USA, 2008; pp. 114–158. [Google Scholar]
  64. McCrae, R.; John, O. An Introduction to the Five-Factor Model and Its Applications. J. Pers. 1992, 60, 175–215. [Google Scholar] [CrossRef] [PubMed]
  65. Gelman, A.; Carlin, J.; Stern, H.; Dunson, D.; Vehtari, A.; Rubin, D. Bayesian Data Analysis, 3rd ed. ed.; CRC Press: Boca Raton, FL, USA, 2013. [Google Scholar]
  66. Funder, D. On the Accuracy of Personality Judgment: A Realistic Approach. Psychol. Rev. 1995, 102, 652–670. [Google Scholar] [CrossRef] [PubMed]
  67. Vazire, S. Who Knows What about a Person? The Self–Other Knowledge Asymmetry (SOKA) Model. J. Pers. Soc. Psychol. 2010, 99, 281–303. [Google Scholar] [CrossRef] [PubMed]
  68. OpenAI. GPT-4 Tech. Rep. 2303.08774. 2023. [CrossRef]
  69. Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A. Language Models are Few-Shot Learners. In Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS ’20), 2020; pp. 1877–1901. [Google Scholar]
  70. Reynolds, L.; McDonell, K. Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm. In Proceedings of the Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems, 2021; pp. 1–7. [Google Scholar]
  71. APA. Practice Guideline for the Treatment of Patients with Major Depressive Disorder, 3rd ed. ed.; American Psychiatric Association: Arlington, VA, USA, 2010. [Google Scholar]
  72. McCrae, R.R.; Costa, P. Personality in Adulthood: A Five-Factor Theory Perspective, 2nd ed. ed.; Guilford Press: New York, NY, USA, 2003. [Google Scholar]
  73. Goldberg, L. The Structure of Phenotypic Personality Traits. Am. Psychol. 1993, 48, 26–34. [Google Scholar] [CrossRef] [PubMed]
  74. Rentfrow, P.J.; Gosling, S.D.; Jokela, M.; Stillwell, D.J.; Kosinski, M.; Potter, J. Divided We Stand: Three Psychological Regions of the United States and Their Political, Economic, Social, and Health Correlates. J. Personal. Soc. Psychol. 2013, 105, 996–1012. [Google Scholar] [CrossRef] [PubMed]
  75. Liu, Y.; Iter, D.; Xu, Y.; Wang, S.; Xu, R.; Zhu, C. G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment. arXiv 2023, arXiv:2303.16634. [Google Scholar]
  76. Krippendorff, K. Computing Krippendorff’s Alpha-Reliability; Departmental Papers (ASC), University of Pennsylvania: Philadelphia, PA, USA, 2011. [Google Scholar]
  77. Landis, J.; Koch, G. The Measurement of Observer Agreement for Categorical Data. Biometrics 1977, 33, 159–174. [Google Scholar] [CrossRef]
  78. Gisev, N.; Bell, J.; Chen, T. Interrater Reliability and Agreement in Clinical Research: A Guide to Best Practice. Res. Soc. Adm. Pharm. 2013, 9, 330–337. [Google Scholar]
  79. Little, R.; Rubin, D. Statistical Analysis with Missing Data, 3rd ed. ed.; Wiley: Hoboken, NJ, USA, 2019. [Google Scholar]
  80. Efron, B.; Tibshirani, R. An Introduction to the Bootstrap; Chapman and Hall/CRC: New York, NY, USA, 1994. [Google Scholar]
  81. Wasserstein, R.; Lazar, N. The ASA’s Statement on p-Values: Context, Process, and Purpose. Am. Stat. 2016, 70, 129–133. [Google Scholar] [CrossRef]
  82. Chen, L.; Zaharia, M.; Zou, J. How is ChatGPT’s Behavior Changing over Time. arXiv 2023, arXiv:2307.09009. [Google Scholar]
  83. Simon, G. Patient-Centered Outcomes in Mental Health Care. JAMA 2011, 306, 1141–1142. [Google Scholar]
  84. Lundberg, S.M.; Lee, S.I. A Unified Approach to Interpreting Model Predictions. In Advances in Neural Information Processing Systems 30; Curran Associates, Inc., 2017; pp. 4765–4774. [Google Scholar]
  85. Sundararajan, M.; Taly, A.; Yan, Q. Axiomatic Attribution for Deep Networks. In Proceedings of the 34th International Conference on Machine Learning;PMLR; Precup, D., Teh, Y.W., Eds.; 2017; Vol. 70, pp. 3319–3328. [Google Scholar]
Figure 1. Three-layer artefact overview. The design contribution is structured into an interaction-flow orchestration layer, a detect–evaluate–adapt regulation layer, and the current theory-grounded instantiation using OCEAN personality signals, Zurich Model motivational domains, and behavioural response effects.
Figure 1. Three-layer artefact overview. The design contribution is structured into an interaction-flow orchestration layer, a detect–evaluate–adapt regulation layer, and the current theory-grounded instantiation using OCEAN personality signals, Zurich Model motivational domains, and behavioural response effects.
Preprints 229251 g001
Figure 2. OCEAN perception component. User utterances and dialogue context are processed into turn-level Big Five/OCEAN trait signals and accumulated into an explicit personality-state representation for downstream adaptation.
Figure 2. OCEAN perception component. User utterances and dialogue context are processed into turn-level Big Five/OCEAN trait signals and accumulated into an explicit personality-state representation for downstream adaptation.
Preprints 229251 g002
Figure 3. Conceptual mapping from perception to behaviour. Big Five/OCEAN trait signals are interpreted through Zurich Model motivational domains and translated into conversational behaviour effects used to guide response generation.
Figure 3. Conceptual mapping from perception to behaviour. Big Five/OCEAN trait signals are interpreted through Zurich Model motivational domains and translated into conversational behaviour effects used to guide response generation.
Preprints 229251 g003
Figure 4. Regulation and behaviour component. Personality-state signals are transformed into behaviour cues, conflict-resolved into a single regulation instruction, merged with the base prompt, and logged for traceability before adapted response generation.
Figure 4. Regulation and behaviour component. Personality-state signals are transformed into behaviour cues, conflict-resolved into a single regulation instruction, merged with the base prompt, and logged for traceability before adapted response generation.
Preprints 229251 g004
Figure 7. Personality trait patterns across conversation turns (Regulated condition, n = 60 turns, sorted by Conversation ID). Each column represents one turn; each row represents one OCEAN dimension. Colours follow ColorBrewer RdBu (colour-blind safe): blue = High ( + 1 ), orange = Low ( 1 ), grey = Mid (0). Vertical lines mark conversation boundaries (each block = 6 turns). A broad colour split is visible between the left (Type A) and right (Type B) halves, with blue dominant in Type A and orange dominant in Type B for most dimensions. Grey (Mid) patterns are more pronounced in Conscientiousness and Agreeableness, reflecting dimensions where neutral-state detection occurs more frequently even under extreme profile conditions.
Figure 7. Personality trait patterns across conversation turns (Regulated condition, n = 60 turns, sorted by Conversation ID). Each column represents one turn; each row represents one OCEAN dimension. Colours follow ColorBrewer RdBu (colour-blind safe): blue = High ( + 1 ), orange = Low ( 1 ), grey = Mid (0). Vertical lines mark conversation boundaries (each block = 6 turns). A broad colour split is visible between the left (Type A) and right (Type B) halves, with blue dominant in Type A and orange dominant in Type B for most dimensions. Grey (Mid) patterns are more pronounced in Conscientiousness and Agreeableness, reflecting dimensions where neutral-state detection occurs more frequently even under extreme profile conditions.
Preprints 229251 g007
Figure 8. Cross-model agreement matrices for Personality Needs Addressed ratings (original GPT-family evaluator vs. Sonnet 4.6, n = 120 turns). Cells report turn counts separately for the baseline and regulated conditions; outlined diagonal cells indicate exact agreement. Exact agreement was 45.0% (54/120), with Cohen’s κ = 0.219 . In the regulated condition, the original evaluator assigned Yes to all 60 turns, whereas Sonnet 4.6 assigned 16 Yes, 29 Not Sure, and 15 No ratings. The matrices expose systematic differences in rating severity despite preservation of the regulated–baseline effect direction.
Figure 8. Cross-model agreement matrices for Personality Needs Addressed ratings (original GPT-family evaluator vs. Sonnet 4.6, n = 120 turns). Cells report turn counts separately for the baseline and regulated conditions; outlined diagonal cells indicate exact agreement. Exact agreement was 45.0% (54/120), with Cohen’s κ = 0.219 . In the regulated condition, the original evaluator assigned Yes to all 60 turns, whereas Sonnet 4.6 assigned 16 Yes, 29 Not Sure, and 15 No ratings. The matrices expose systematic differences in rating severity despite preservation of the regulated–baseline effect direction.
Preprints 229251 g008
Figure 9. Cliff’s δ with 95% bootstrap CI for Personality Needs (blue) and Emotional Tone (orange) across evaluator families and profile tiers. Open square (†) = ceiling artifact (original GPT-family evaluator / Phase 1A extreme profiles). Phase 1B: Study A outputs re-scored by Sonnet 4.6 ( n = 10 personas); Phase 1C: mixed-profile core set ( n = 24 conversations).
Figure 9. Cliff’s δ with 95% bootstrap CI for Personality Needs (blue) and Emotional Tone (orange) across evaluator families and profile tiers. Open square (†) = ceiling artifact (original GPT-family evaluator / Phase 1A extreme profiles). Phase 1B: Study A outputs re-scored by Sonnet 4.6 ( n = 10 personas); Phase 1C: mixed-profile core set ( n = 24 conversations).
Preprints 229251 g009
Figure 11. Qualitative response comparisons for contrasting personality profiles. (A) Type B (all traits: 1 ; Conversation B-4, turn 1): the regulated reply emphasises validation and a low-pressure stance supporting Security and reducing Arousal, whereas the baseline remains broadly supportive but less specifically aligned to the implied motivational state. (B) Type A (all traits: + 1 ; Conversation A-5, turn 3): the regulated reply maintains an affirming, growth-oriented framing with structured options, whereas the baseline shifts toward generic reflective advice. Both comparisons illustrate how trait-conditioned regulation changes framing while preserving appropriate general support. User turns differ across arms because paired conditions used independently generated inputs (Methods).
Figure 11. Qualitative response comparisons for contrasting personality profiles. (A) Type B (all traits: 1 ; Conversation B-4, turn 1): the regulated reply emphasises validation and a low-pressure stance supporting Security and reducing Arousal, whereas the baseline remains broadly supportive but less specifically aligned to the implied motivational state. (B) Type A (all traits: + 1 ; Conversation A-5, turn 3): the regulated reply maintains an affirming, growth-oriented framing with structured options, whereas the baseline shifts toward generic reflective advice. Both comparisons illustrate how trait-conditioned regulation changes framing while preserving appropriate general support. User turns differ across arms because paired conditions used independently generated inputs (Methods).
Preprints 229251 g011
Table 1. Big Five personality trait operationalisation. Neuroticism is inverse-coded: +1 denotes emotional stability (low neuroticism), and -1 denotes emotional vulnerability (high neuroticism).
Table 1. Big Five personality trait operationalisation. Neuroticism is inverse-coded: +1 denotes emotional stability (low neuroticism), and -1 denotes emotional vulnerability (high neuroticism).
Trait +1 (High) -1 (Low)
Openness Curious, imaginative, open to novelty Prefers routine, resistant to new ideas
Conscientiousness Organised, disciplined, structured Disorganised, impulsive, spontaneous
Extraversion Outgoing, energetic, assertive Reserved, quiet, withdrawn
Agreeableness Cooperative, empathetic, friendly Critical, skeptical, confrontational
Neuroticism (inverse-coded) Calm, emotionally stable, resilient Anxious, emotionally sensitive, insecure
Table 2. Per-trait detection metrics (Phase 1A, n = 60 turns, extreme profiles). Correct = predicted label matches the simulated pole; Neutral = prediction assigned 0; Sign-flip = predicted direction opposite to the simulated pole. P/R are precision and recall for the indicated pole on { 1 , + 1 } . Balanced accuracy equals accuracy because the simulated ground truth contains no true-neutral class. Important limitation: These metrics are computed against machine-generated simulated personality profiles, not validated psychometric ground truth. The human-annotated subset (§Section 4.5; n = 66 turns, two independent raters) provides partial external criterion evidence for the labelling step; these remain provisional benchmarks pending adjudicated psychometric validation.
Table 2. Per-trait detection metrics (Phase 1A, n = 60 turns, extreme profiles). Correct = predicted label matches the simulated pole; Neutral = prediction assigned 0; Sign-flip = predicted direction opposite to the simulated pole. P/R are precision and recall for the indicated pole on { 1 , + 1 } . Balanced accuracy equals accuracy because the simulated ground truth contains no true-neutral class. Important limitation: These metrics are computed against machine-generated simulated personality profiles, not validated psychometric ground truth. The human-annotated subset (§Section 4.5; n = 66 turns, two independent raters) provides partial external criterion evidence for the labelling step; these remain provisional benchmarks pending adjudicated psychometric validation.
Trait Acc. Bal. acc. Neutral Sign-flip P/R ( 1 ) P/R ( + 1 ) Macro- F 1
Openness (O) 85.0% 85.0% 15.0% 0.0% 1.00 / 0.77 1.00 / 0.93 0.917
Conscientiousness (C) 65.0% 65.0% 35.0% 0.0% 1.00 / 0.77 1.00 / 0.53 0.782
Extraversion (E) 93.3% 93.3% 6.7% 0.0% 1.00 / 1.00 1.00 / 0.87 0.964
Agreeableness (A) 73.3% 73.3% 23.3% 3.3% 1.00 / 0.47 0.94 / 1.00 0.802
Neuroticism (N) 100.0% 100.0% 0.0% 0.0% 1.00 / 1.00 1.00 / 1.00 1.000
Macro mean 83.3% 83.3% 16.0% 0.7% 1.00 / 0.80 0.99 / 0.87 0.893
Table 4. Per-trait confusion counts (Phase 1A, n = 60 turns, extreme profiles). Each block is predicted 1 / 0 / + 1 given the simulated ground-truth pole. There is no true-0 class under these profiles. Ground truth is the simulated pole, not a psychometric instrument.
Table 4. Per-trait confusion counts (Phase 1A, n = 60 turns, extreme profiles). Each block is predicted 1 / 0 / + 1 given the simulated ground-truth pole. There is no true-0 class under these profiles. Ground truth is the simulated pole, not a psychometric instrument.
True 1 True + 1
Trait 1 0 + 1 1 0 + 1
Openness (O) 23 7 0 0 2 28
Conscientiousness (C) 23 7 0 0 14 16
Extraversion (E) 30 0 0 0 4 26
Agreeableness (A) 14 14 2 0 0 30
Neuroticism (N) 30 0 0 0 0 30
Table 5. Conversation-level regulation outcomes (Study A, n = 10 persona pairs, extreme profiles). Exact two-sided Wilcoxon signed-rank on conversation means (paired); W is T + , the sum of ranks of positive (regulated > baseline) differences. Cluster-bootstrap 95% CI on Cliff’s δ (10,000 replicates; conversation is the resampling unit). Turn-level Yes rates are descriptive only and are not used for inference. † All 10 pairs favour regulation, so the cluster bootstrap resamples a constant and returns a zero-width interval. We instead report the exact Clopper–Pearson interval for 10/10 concordant pairs, which is the honest small-sample bound; read as an upper-bound contrast only. NE = not estimable (both arms zero-variance; Cohen’s d is undefined). Important limitation: Ratings generated by the original GPT-family evaluator with author qualitative review; not independent human outcome raters. Human validation (§Section 4.5) provides convergent validity evidence on detection labelling ( n = 66 turns, two independent raters; §Section 4.5).
Table 5. Conversation-level regulation outcomes (Study A, n = 10 persona pairs, extreme profiles). Exact two-sided Wilcoxon signed-rank on conversation means (paired); W is T + , the sum of ranks of positive (regulated > baseline) differences. Cluster-bootstrap 95% CI on Cliff’s δ (10,000 replicates; conversation is the resampling unit). Turn-level Yes rates are descriptive only and are not used for inference. † All 10 pairs favour regulation, so the cluster bootstrap resamples a constant and returns a zero-width interval. We instead report the exact Clopper–Pearson interval for 10/10 concordant pairs, which is the honest small-sample bound; read as an upper-bound contrast only. NE = not estimable (both arms zero-variance; Cohen’s d is undefined). Important limitation: Ratings generated by the original GPT-family evaluator with author qualitative review; not independent human outcome raters. Human validation (§Section 4.5) provides convergent validity evidence on detection labelling ( n = 66 turns, two independent raters; §Section 4.5).
Outcome Reg M ( S D ) Base M ( S D ) + / / = T + (exact p) Cliff’s δ [95% CI] Status
Personality Needs 2.000 (0.000) 0.200 (0.450) 10/0/0 55.0 (0.00195) 1.000 [0.692, 1.000]† Estimable — ceiling
Emotional Tone 2.000 (0.000) 2.000 (0.000) 0/0/10 0.0 (1.000) NE Not estimable
Relevance & Coherence 2.000 (0.000) 1.967 (0.105) 1/0/9 1.0 (1.000) 0.100 [0.000, 0.300] Thin signal
Table 6. Cross-model evaluator agreement (Study A; original GPT-family vs. Sonnet 4.6). Cohen’s κ is reported as the primary agreement measure when category prevalence is not extreme. Gwet AC1 is preferred for Emotional Tone and Relevance & Coherence because AC1 corrects for prevalence-induced κ deflation when one category dominates.
Table 6. Cross-model evaluator agreement (Study A; original GPT-family vs. Sonnet 4.6). Cohen’s κ is reported as the primary agreement measure when category prevalence is not extreme. Gwet AC1 is preferred for Emotional Tone and Relevance & Coherence because AC1 corrects for prevalence-induced κ deflation when one category dominates.
Outcome κ Gwet AC1 Direction preserved? Interpretation
Personality Needs 0.219 0.197 Yes (both p = 0.00195 ) Fair; effect replicates
Emotional Tone 0.000 0.348 No Ceiling artifact; AC1 preferred
Relevance & Coh. −0.009 0.719 Yes (near-null both) Null replicates; AC1 preferred
Table 7. Detection accuracy by signal tier, Phase 1C core profiles. Per-trait accuracy is the share of trait-cells matching ground truth; the neutral rate is the share predicted 0; the flip rate is the share of directional ground-truth cells assigned the opposite pole. On mixed and subtle profiles, true-neutral traits exist, so a correct 0 counts in both accuracy and the neutral rate and the three columns need not sum to 100%. Extreme anchors are a reference accuracy only ( n = 36 input-replayed turns); that row’s neutral and flip rates were not a reporting target (—). Primary claims are restricted to core (mixed + subtle) profiles.
Table 7. Detection accuracy by signal tier, Phase 1C core profiles. Per-trait accuracy is the share of trait-cells matching ground truth; the neutral rate is the share predicted 0; the flip rate is the share of directional ground-truth cells assigned the opposite pole. On mixed and subtle profiles, true-neutral traits exist, so a correct 0 counts in both accuracy and the neutral rate and the three columns need not sum to 100%. Extreme anchors are a reference accuracy only ( n = 36 input-replayed turns); that row’s neutral and flip rates were not a reporting target (—). Primary claims are restricted to core (mixed + subtle) profiles.
Tier Conv. Turns Per-trait acc. Neutral rate Flip rate
Mixed (6 profiles) 18 108 58.1% 41.1% 8.1%
Subtle (2 profiles) 6 36 69.4% 52.8% 0.0%
Extreme anchors (ref) 6 36 74.4%
Table 8. Regulation persistence on mixed-profile core conversations ( n = 24 ), three cross-family judges. Mean difference (regulated − baseline) for Personality Needs Addressed (PN) and Emotional Tone Appropriate (ET). Relevance & Coherence: no consistent regulated advantage under any judge.
Table 8. Regulation persistence on mixed-profile core conversations ( n = 24 ), three cross-family judges. Mean difference (regulated − baseline) for Personality Needs Addressed (PN) and Emotional Tone Appropriate (ET). Relevance & Coherence: no consistent regulated advantage under any judge.
Judge PN Δ p ET Δ p
Sonnet 4.6 +0.96 <0.001 +0.67 <0.001
Gemini 3.1 Flash-Lite +1.39 <0.001 +0.63 <0.001
DeepSeek V4 Flash +0.79 0.002 +0.24 0.002
Table 9. Human validation of detection labels ( n = 66 Phase 2 turns, verify-and-revise; two independent raters). Panel A: agreement with pre-filled machine OCEAN labels (upper bound under anchoring). Panel B: human–human agreement on the full overlap. Revision = share of cells changed from the machine proposal.
Table 9. Human validation of detection labels ( n = 66 Phase 2 turns, verify-and-revise; two independent raters). Panel A: agreement with pre-filled machine OCEAN labels (upper bound under anchoring). Panel B: human–human agreement on the full overlap. Revision = share of cells changed from the machine proposal.
A. Human vs. machine (primary rater)
Trait % agree Revision Weighted κ
O 86.4 0.136 0.778
C 92.4 0.076 0.874
E 92.4 0.076 0.916
A 84.8 0.152 0.785
N 92.4 0.076 0.865
Overall 89.7 0.103 0.844
B. Human–human (A vs. B)
Trait % agree Unweighted κ Linear κ
O 75.8 0.608 0.669
C 71.2 0.517 0.577
E 95.5 0.931 0.951
A 83.3 0.732 0.776
N 78.8 0.656 0.683
Overall 80.9 0.689 0.731
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.