Submitted:
29 May 2026
Posted:
01 June 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Related Work
2.1. From Affective Computing to Personality-Aware Digital Support
2.2. LLMs in Mental Health: Capabilities
2.3. Orchestration and Evaluation Challenges
3. Materials and Methods
3.1. Overview and Research Objectives
- 1.
- Detection Accuracy: Can stable personality traits be accurately inferred from conversational cues on a turn-by-turn basis within multi-turn interactions?
- 2.
- Regulation Effectiveness: Does theory-driven personality-adaptive regulation produce measurable improvements in personality-specific need fulfilment beyond generic responses?
- 3.
- Evaluation Validity: How should results be interpreted given reliance on LLM-based evaluation, and what human validation is required to establish evaluation credibility?
3.1.1. Experimental Design
3.1.2. Sample Size Considerations
3.1.3. Ethics and Transparency
3.2. Personality-Adaptive System Architecture
- Detection: Trait estimation is implemented as a state-driven process that updates a continuous Big Five vector and confidence metrics in the interaction store after each user input. Traits are inferred from linguistic cues in the current interaction context and persist as storage values accessible across states.
- Regulation: Within response-generation states, prompts are augmented with trait-conditioned instructions that softly modulate the assistant’s response style—such as tone, pacing, and interpersonal framing—without altering semantic content or task objectives. This is achieved by dynamically composing the base state prompt with trait-aligned augmentation prompts at runtime.
- Evaluation: After each response generation, structured assessment actions execute LLM-based quality checks and log evaluation scores alongside turn identifiers in the interaction store. In the current study, these evaluation outputs are used for offline analysis and do not feed back into detection or regulation during the simulation.
3.3. System Implementation

3.3.1. Theory-Driven Regulation


3.4. Experimental Materials and Procedure
3.5. Evaluation Procedure
3.5.1. LLM-Based Evaluator System
3.5.2. Evaluation Validity and Author-Led Qualitative Review
3.6. Analysis and Reproducibility
3.6.1. Statistical Analysis
3.6.2. Reproducibility and Computational Environment
4. Results
4.1. Data Quality and Sample Distribution
4.2. Detection and Regulation Performance
4.2.1. Personality Vector Consistency
4.3. Comparative Effectiveness: Regulated vs. Baseline Performance
4.3.1. Primary Outcome: Personality Needs Addressed
4.3.2. Secondary Outcomes: Basic Conversational Quality (Ceiling Effects)
4.3.3. Interpretation of the Selective Enhancement Pattern
4.3.4. Statistical Robustness and Implementation Fidelity
4.4. Qualitative Examples Demonstrating Personality Adaptation
5. Discussion
5.1. Research and Design Implications
- Layered deployment: Start from a strong generic baseline and treat personality adaptation as an additive capability; AI augments rather than replaces generic conversational competence.
- Explicit, updateable personality state: Maintain persistent trait estimates (with uncertainty) updated from conversational evidence.
- Theory-grounded mapping: Translate traits to behavioural strategies using established personality–motivation frameworks (e.g., the Zurich Model), rather than ad hoc heuristics.
- Transparent adaptation logic: Log trait estimates, applied regulations, and evaluation outcomes for auditability and debugging.
- Selective application: Apply adaptation primarily to interactional style and support preferences, not to all content indiscriminately.
5.2. Strengths, Limitations, and Validation Pathway
5.2.1. Limitations and Barriers
5.2.2. Validation Pathway for AI-Augmented Personalisation
- Test moderate and mixed personality profiles reflecting realistic trait distributions
- Extend dialogue length to 20+ exchanges to assess long-term adaptation stability
- Conduct cross-linguistic and cross-cultural validation studies to assess generalisability beyond English and Western norms
- IRB-approved study protocol with comprehensive informed consent addressing AI-based personality profiling and data usage
- User-reported outcome measures: engagement, satisfaction, perceived support quality, and therapeutic alliance
- Safety monitoring infrastructure with clear protocols for identifying and responding to user distress
- Qualitative interviews to understand user experiences with personality-adaptive versus non-adaptive systems
- Multi-site implementation studies across diverse user populations
- Integration with existing digital mental health platforms to assess workflow compatibility
- Equity assessments examining performance across demographic groups (age, gender, cultural background, education level)
- Longitudinal tracking to evaluate sustained engagement and long-term user outcomes
6. Conclusions
Supplementary Materials
- Supplementary File S1: Complete system prompts for personality detection (5 Big Five trait detectors)
- Supplementary File S2: Regulation templates for Zurich Model mapping (arousal, security, affiliation)
- Supplementary File S3: Evaluator GPT system prompt and scoring matrix
- Supplementary File S4: Complete simulation transcripts (20 conversations, 120 dialogue turns)
- Supplementary File S5: Statistical analysis code (Python/Jupyter notebooks for effect sizes, paired tests, weighted scoring, and visualisations)
- Supplementary File S6: Missingness comparison plot (horizontal bars showing <5% missing data across conditions)
- Supplementary File S7: Personality needs YES-rate by conversation (demonstrating consistent improvement across all 10 conversation pairs)
- Supplementary File S8: Rating distribution raw counts (YES/NOT SURE/NO composition with value annotations)
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Dingler, T., D. Kwasnicka, J. Wei, E. Gong, and B. Oldenburg. 2021. The use and promise of conversational agents in digital health. Yearbook of Medical Informatics 30: 191–199. [Google Scholar] [CrossRef]
- Laranjo, L., A.G. Dunn, H.L. Tong, A.B. Kocaballi, J. Chen, R. Bashir, D. Surian, B. Gallego, F. Magrabi, and A.Y. Lau. 2018. Conversational agents in healthcare: a systematic review. J. of the American Medical Informatics Association 25: 1248–1258. [Google Scholar] [CrossRef]
- Vasiliu, L., K. Cortis, R. McDermott, A. Kerr, A. Peters, M. Hesse, J. Hagemeyer, T. Belpaeme, J. McDonald, and R. Villing. 2021. CASIE – Computing affect and social intelligence for healthcare in an ethical and trustworthy manner. Paladyn, J. of Behavioral Robotics 12: 437–453. [Google Scholar] [CrossRef]
- Kocaballi, A.B., E. Sezgin, L. Clark, J.M. Carroll, and et al. 2022. Design and Evaluation Challenges of Conversational Agents in Health Care and Well-being: Selective Review Study. J. Med. Internet Res. 24: e38525. [Google Scholar] [CrossRef] [PubMed]
- Kaddour, J., J. Harris, M. Mozes, H. Bradley, R. Raileanu, and R. McHardy. 2023. Challenges and Applications of Large Language Models. arXiv arXiv:2307.10169. [Google Scholar] [CrossRef]
- Rapp, A., L. Curti, and A. Boldi. 2021. The human side of human-chatbot interaction: A systematic literature review of ten years of research on text-based chatbots. International Journal of Human-Computer Studies 151: 102630. [Google Scholar] [CrossRef]
- Han, E., D. Yin, and H. Zhang. 2022. Chatbot Empathy in Customer Service: When It Works and When It Backfires. Proceedings of the SIGHCI 2022 Proceedings, Vol. 1. [Google Scholar]
- Juquelier, A., I. Poncin, and S. Hazée. 2025. Empathic chatbots: A double-edged sword in customer experiences. Journal of Business Research 188: 115074. [Google Scholar] [CrossRef]
- Seitz, L. 2024. Artificial empathy in healthcare chatbots: Does it feel authentic? Computers in Human Behavior: Artificial Humans 2: 100067. [Google Scholar] [CrossRef]
- Abernethy, A., L. Adams, M. Barrett, C. Bechtel, P. Brennan, A. Butte, J. Faulkner, E. Fontaine, S. Friedhoff, J. Halamka, and et al. 2022. The Promise of Digital Health: Then, Now, and the Future. NAM Perspectives.
- Goetz, L.H., and N.J. Schork. 2018. Personalized medicine: Motivation, challenges, and progress. Fertility and Sterility 109: 952–963. [Google Scholar] [CrossRef]
- O’Keefe, D.J. 2016. Persuasion: Theory and Research. SAGE Publications, Inc. [Google Scholar]
- Milne-Ives, M., C. de Cock, E. Lim, M.H. Shehadeh, N. de Pennington, G. Mole, E. Normando, and E. Meinert. 2020. The Effectiveness of Artificial Intelligence Conversational Agents in Health Care: Systematic Review. J. Med. Internet Res. 22: e20346. [Google Scholar] [CrossRef]
- Wei, J., X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E.H. Chi, Q.V. Le, and D. Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. Proceedings of the Proceedings of the 36th International Conference on Neural Information Processing Systems, Red Hook, NY, USA; p. NIPS ’22. [Google Scholar]
- Ayers, J.W., A. Poliak, M. Dredze, E.C. Leas, Z. Zhu, J.B. Kelley, D.J. Faix, A.M. Goodman, C.A. Longhurst, M. Hogarth, and et al. 2023. Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum. JAMA Internal Medicine 183: 589–596. [Google Scholar] [CrossRef] [PubMed]
- Färber, A., A. de Spindler, A. Moser, and G. Schwabe. 2023. Closing the Loop for Patients with Chronic Diseases - from Problems to a Solution Architecture. Proceedings of the The 11th IEEE International Conference on Healthcare Informatics; pp. 1–11. [Google Scholar] [CrossRef]
- Färber, A., C. Schwabe, P.H. Stalder, M. Dolata, and G. Schwabe. 2024. Physicians’ and Patients’ Expectations From Digital Agents for Consultations: Interview Study Among Physicians and Patients. JMIR Human Factors 11: e49647. [Google Scholar] [CrossRef]
- Staehelin, D., M. Dolata, L. Stöckli, and G. Schwabe. 2024. How Patient-Generated Data Enhance Patient-Provider Communication in Chronic Care: Field Study in Design Science Research. JMIR Medical Informatics 12: e57406. [Google Scholar] [CrossRef] [PubMed]
- Strubell, E., A. Ganesh, and A. McCallum. 2019. Edited by A. Korhonen, D. Traum and L. Màrquez. Energy and Policy Considerations for Deep Learning in NLP. In Proceedings of the Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics: pp. 3645–3650. [Google Scholar]
- Ding, N., Y. Qin, G. Yang, F. Wei, Z. Yang, Y. Su, S. Hu, Y. Chen, C.M. Chan, W. Chen, and et al. 2023. Parameter-efficient fine-tuning of large-scale pre-trained language models. Nature Machine Intelligence 5: 220–235. [Google Scholar] [CrossRef]
- Hu, E.J., Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen. 2021. LoRA: Low-Rank Adaptation of Large Language Models. arXiv arXiv:2106.09685. [Google Scholar]
- Korzynski, P., G. Mazurek, P. Krzypkowska, and A. Kurasinski. 2023. Artificial intelligence prompt engineering as a new digital competence: Analysis of generative AI technologies such as ChatGPT. Entrepreneurial Business and Economics Review.
- White, J., Q. Fu, S. Hays, M. Sandborn, C. Olea, H. Gilbert, A. Elnashar, J. Spencer-Smith, and D. Schmidt. 2023. A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT. arXiv arXiv:2302.11382. [Google Scholar] [CrossRef]
- Fernando, C., D. Banarse, H. Michalewski, and et al. 2023. Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution. arXiv arXiv:2309.16797. [Google Scholar]
- Liu, P., W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig. 2023. Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. ACM Computing Surveys 55: 195:1–195:35. [Google Scholar] [CrossRef]
- Hou, Y., H. Dong, X. Wang, B. Li, and W. Che. 2022. Edited by N. Calzolari, C.R. Huang, H. Kim, J. Pustejovsky, L. Wanner, K.S. Choi, P.M. Ryu, H.H. Chen, L. Donatelli, H. Ji and et al. MetaPrompting: Learning to Learn Better Prompts. In Proceedings of the Proceedings of the 29th International Conference on Computational Linguistics. International Committee on Computational Linguistics: pp. 3251–3262. [Google Scholar]
- Wu, W., J. Heierli, M. Meisterhans, A. Moser, A. Färber, M. Dolata, E. Gavagnin, A. de Spindler, and G. Schwabe. 2024. Edited by S. Islam and A. Sturm. PROMISE: A Framework for Model-Driven Stateful Prompt Orchestration. In Proceedings of the Intelligent Information Systems: CAiSE Forum 2024. Limassol, Cyprus, Springer International Publishing: <i>Lecture Notes in Business Information Processing</i>, June 3–7, Vol. 520. [Google Scholar]
- Alisamir, S., and F. Ringeval. 2021. On the Evolution of Speech Representations for Affective Computing: A Brief History and Critical Overview. IEEE Signal Process. Mag. 38: 12–21. [Google Scholar] [CrossRef]
- Alsharekh, M. 2022. Facial Emotion Recognition in Verbal Communication Based on Deep Learning. Electronics 11: 2568. [Google Scholar] [CrossRef]
- Mairesse, F., and M. Walker. 2011. Controlling User Perceptions of Linguistic Style: Trainable Generation of Personality Traits. Comput. Linguist. 37: 455–488. [Google Scholar] [CrossRef]
- Ta, V., C. Griffith, C. Boatfield, X. Wang, M. Civitello, and H. Maffei. 2020. User Experiences of Social Support from Companion Chatbots in Everyday Contexts: Thematic Analysis. J. Med. Internet Res. 22: e16235. [Google Scholar] [CrossRef]
- Broadbent, E. 2024. ElliQ, an AI-Driven Social Robot to Alleviate Loneliness: Progress and Lessons Learned. JAR Life 13: 22–28. [Google Scholar] [CrossRef]
- Shah, S. 2019. Effectiveness of Digital Technology Interventions to Reduce Loneliness in Adults: A Protocol for a Systematic Review and Meta-Analysis. BMJ Open 9: e029324. [Google Scholar] [CrossRef]
- Shah, S. 2021. Evaluation of the Effectiveness of Digital Technology Interventions to Reduce Loneliness in Older Adults: Systematic Review and Meta-Analysis. J. Med. Internet Res. 23: e24712. [Google Scholar] [CrossRef] [PubMed]
- Zhou, M. 2021. Designing Effective Interview Chatbots: Automatic Chatbot Profiling and Design Suggestion Generation. Proceedings of the Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’21). [Google Scholar]
- Quirin, M., F. Malekzad, D. Paudel, A. Knoll, and M. Mirolli. 2023. Dynamics of Personality: The Zurich Model of Motivation Revived, Extended, and Applied to Personality. J. Pers. 91: 928–946. [Google Scholar] [CrossRef]
- Berlyne, D. 1960. Conflict, Arousal, and Curiosity. McGraw-Hill: New York, NY, USA. [Google Scholar]
- Hebb, D. 1949. The Organization of Behavior: A Neuropsychological Theory. Wiley: New York, NY, USA. [Google Scholar]
- Bischof, N. 1985. Das Rätsel Ödipus: Die biologischen Wurzeln des Urkonfliktes. Piper: Munich, Germany. [Google Scholar]
- Bischof, N. 1993. Untersuchungen zur Systemanalyse der sozialen Motivation I: Die Tantalus-Situation. Z. Psychol. 201: 5–43. [Google Scholar]
- Bickmore, T., and R. Picard. 2005. Establishing and Maintaining Long-Term Human-Computer Relationships. ACM Trans. Comput.-Hum. Interact. 12: 293–327. [Google Scholar] [CrossRef]
- Zheng, Z., L. Liao, Y. Deng, and L. Nie. 2023. Building Emotional Support Chatbots in the Era of LLMs. arXiv arXiv:2308.11584. [Google Scholar] [CrossRef]
- Chen, K., X. Kang, X. Lai, and Z. Ni. 2023. Enhancing Emotional Support Capabilities of Large Language Models through Cascaded Neural Networks. Proceedings of the Proceedings of the 2023 4th International Conference on Computer, Big Data and Artificial Intelligence (ICCBD+AI); pp. 318–326. [Google Scholar]
- Ouyang, L., J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, and A. Ray. 2022. Training Language Models to Follow Instructions with Human Feedback. Proceedings of the Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS ’22); pp. 27730–27744. [Google Scholar]
- Dong, T., F. Liu, X. Wang, Y. Jiang, X. Zhang, and X. Sun. 2024. EmoAda: A Multimodal Emotion Interaction and Psychological Adaptation System. Proceedings of the Proceedings of the Conference on Multimedia Modeling, Cham, Switzerland. [Google Scholar]
- Abbasian, M., I. Azimi, A. Rahmani, and R. Jain. 2023. Conversational Health Agents: A Personalized LLM-Powered Agent Framework. arXiv arXiv:2310.02293. [Google Scholar]
- Sorino, P. 2024. ARIEL: Brain-Computer Interfaces Meet Large Language Models for Emotional Support Conversation. Proceedings of the Adjunct Proceedings of the 32nd ACM Conference on User Modeling, Adaptation and Personalization (UMAP ’24). [Google Scholar]
- Dongre, P. 2024. Physiology-Driven Empathic Large Language Models (EmLLMs) for Mental Health Support. Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (CHI EA ’24). [Google Scholar]
- Zhang, H., Y. Chen, M. Wang, and S. Feng. 2024. FEEL: A Framework for Evaluating Emotional Support Capability with Large Language Models. arXiv arXiv:2403.15699. [Google Scholar] [CrossRef]
- Zheng, L., W. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, and E. Xing. 2023. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. arXiv arXiv:2306.05685. [Google Scholar]
- Kim, S., S. Lee, J. Kim, H. Yoo, Y. Seol, S. Kim, H. Cho, S. Choi, and S. Kim. 2024. Evaluating LLM-as-a-judge with Professional Human Ratings. arXiv arXiv:2402.18139. [Google Scholar]
- Montgomery, D. 2017. Design and Analysis of Experiments, 9th ed. ed. Wiley: Hoboken, NJ, USA. [Google Scholar]
- Collins, L., J. Dziak, and R. Li. 2009. Design of Experiments with Multiple Independent Variables: A Resource Management Perspective on Complete and Fractional Factorial Designs. Psychol. Methods 14: 202–224. [Google Scholar] [CrossRef] [PubMed]
- Carroll, C., M. Patterson, S. Wood, A. Booth, J. Rick, and S. Balain. 2007. A Conceptual Framework for Implementation Fidelity. Implement. Sci. 2: 40. [Google Scholar] [CrossRef] [PubMed]
- Faul, F., E. Erdfelder, A. Buchner, and A. Lang. 2009. Statistical Power Analyses Using G*Power 3.1: Tests for Correlation and Regression Analyses. Behav. Res. Methods 41: 1149–1160. [Google Scholar] [CrossRef]
- Schulz, K., D. Altman, D. Moher, and C. Group. 2010. CONSORT 2010 Statement: Updated Guidelines for Reporting Parallel Group Randomised Trials. BMJ 340: c332. [Google Scholar] [CrossRef]
- Chan, A., J. Tetzlaff, P. Gøtzsche, D. Altman, H. Mann, J. Berlin, K. Dickersin, A. Hróbjartsson, K. Schulz, and W. Parulekar. 2013. SPIRIT 2013 Statement: Defining Standard Protocol Items for Clinical Trials. Ann. Intern. Med. 158: 200–207. [Google Scholar] [CrossRef] [PubMed]
- Association, W.M. 2013. World Medical Association Declaration of Helsinki: Ethical Principles for Medical Research Involving Human Subjects. JAMA 310: 2191–2194. [Google Scholar]
- D’Alfonso, S. 2020. AI in Mental Health. Curr. Opin. Psychol. 36: 112–117. [Google Scholar] [CrossRef]
- Torous, J., K. Myrick, N. Rauseo-Ricupero, and J. Firth. 2020. Digital Mental Health and COVID-19: Using Technology Today to Accelerate the Curve on Access and Quality Tomorrow. JMIR Ment. Health 7: e18848. [Google Scholar] [CrossRef]
- Alberts, L., G. Keeling, and A. McCroskery. 2024. Should Agentic Conversational AI Change How We Think about Ethics? Characterising an Interactional Ethics Centred on Respect. arXiv arXiv:2401.09187. [Google Scholar] [CrossRef]
- Costa, P., and R. McCrae. 1992. Revised NEO Personality Inventory (NEO-PI-R) and NEO Five-Factor Inventory (NEO-FFI): Professional Manual. Psychological Assessment Resources: Odessa, FL, USA. [Google Scholar]
- John, O., L. Naumann, and C. Soto. 2008. Paradigm Shift to the Integrative Big Five Trait Taxonomy: History, Measurement, and Conceptual Issues. Proceedings of the Handbook of Personality: Theory and Research, New York, NY, USA; pp. 114–158. [Google Scholar]
- McCrae, R., and O. John. 1992. An Introduction to the Five-Factor Model and Its Applications. J. Pers. 60: 175–215. [Google Scholar] [CrossRef] [PubMed]
- Gelman, A., J. Carlin, H. Stern, D. Dunson, A. Vehtari, and D. Rubin. 2013. Bayesian Data Analysis, 3rd ed. ed. CRC Press: Boca Raton, FL, USA. [Google Scholar]
- Funder, D. 1995. On the Accuracy of Personality Judgment: A Realistic Approach. Psychol. Rev. 102: 652–670. [Google Scholar] [CrossRef] [PubMed]
- Vazire, S. 2010. Who Knows What about a Person? The Self–Other Knowledge Asymmetry (SOKA) Model. J. Pers. Soc. Psychol. 99: 281–303. [Google Scholar] [CrossRef] [PubMed]
- OpenAI. 2023. GPT-4 Technical Report, 2303.08774.
- Brown, T., B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, and A. Askell. 2020. Language Models are Few-Shot Learners. Proceedings of the Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS ’20); pp. 1877–1901. [Google Scholar]
- Reynolds, L., and K. McDonell. 2021. Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm. Proceedings of the Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems; pp. 1–7. [Google Scholar]
- Liu, P., W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig. 2023. Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. ACM Comput. Surv. 55: 1–35. [Google Scholar] [CrossRef]
- APA. 2010. Practice Guideline for the Treatment of Patients with Major Depressive Disorder, 3rd ed. ed. American Psychiatric Association: Arlington, VA, USA. [Google Scholar]
- McCrae, R., and P.T.J. Costa. 2003. Personality in Adulthood: A Five-Factor Theory Perspective, 2nd ed. ed. Guilford Press: New York, NY, USA. [Google Scholar]
- Goldberg, L. 1993. The Structure of Phenotypic Personality Traits. Am. Psychol. 48: 26–34. [Google Scholar] [CrossRef]
- McCrae, R., and P.T.J. Costa. 1992. Revised NEO Personality Inventory (NEO PI-R) and NEO Five-Factor Inventory (NEO-FFI) Professional Manual. Psychological Assessment Resources: Odessa, FL, USA. [Google Scholar]
- Liu, Y., D. Iter, Y. Xu, S. Wang, R. Xu, and C. Zhu. 2023. G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment. arXiv arXiv:2303.16634. [Google Scholar] [CrossRef]
- Krippendorff, K. 2011. Computing Krippendorff’s Alpha-Reliability. Departmental Papers (ASC), University of Pennsylvania: Philadelphia, PA, USA. [Google Scholar]
- Landis, J., and G. Koch. 1977. The Measurement of Observer Agreement for Categorical Data. Biometrics 33: 159–174. [Google Scholar] [CrossRef]
- Gisev, N., J. Bell, and T. Chen. 2013. Interrater Reliability and Agreement in Clinical Research: A Guide to Best Practice. Res. Social Adm. Pharm. 9: 330–337. [Google Scholar] [CrossRef]
- Little, R., and D. Rubin. 2019. Statistical Analysis with Missing Data, 3rd ed. ed. Wiley: Hoboken, NJ, USA. [Google Scholar]
- Efron, B., and R. Tibshirani. 1994. An Introduction to the Bootstrap. Chapman and Hall/CRC: New York, NY, USA. [Google Scholar]
- Cohen, J. 1988. Statistical Power Analysis for the Behavioral Sciences, 2nd ed. ed. Lawrence Erlbaum Associates: Hillsdale, NJ, USA. [Google Scholar]
- Wasserstein, R., and N. Lazar. 2016. The ASA’s Statement on p-Values: Context, Process, and Purpose. Am. Stat. 70: 129–133. [Google Scholar] [CrossRef]
- Chen, L., M. Zaharia, and J. Zou. 2023. How is ChatGPT’s Behavior Changing over Time. arXiv arXiv:2307.09009. [Google Scholar] [CrossRef]
- Simon, G. 2011. Patient-Centered Outcomes in Mental Health Care. JAMA 306: 1141–1142. [Google Scholar]













| Trait | +1 (High) | -1 (Low) |
| Openness | Curious, imaginative, open to novelty | Prefers routine, resistant to new ideas |
| Conscientiousness | Organised, disciplined, structured | Disorganised, impulsive, spontaneous |
| Extraversion | Outgoing, energetic, assertive | Reserved, quiet, withdrawn |
| Agreeableness | Cooperative, empathetic, friendly | Critical, skeptical, confrontational |
| Neuroticism (inverse-coded) | Calm, emotionally stable, resilient | Anxious, emotionally sensitive, insecure |
| Domain | Trait | +1 (High) | -1 (Low) |
| Security | Neuroticism (N) | Reassure stability and confidence | Offer extra comfort; acknowledge anxieties |
| Arousal | Openness (O) | Invite exploration and novelty | Focus on familiar topics; reduce novelty |
| Arousal | Extraversion (E) | Energetic, sociable tone | Calm, low-key style with reflective space |
| Affiliation | Agreeableness (A) | Warmth, empathy, collaboration | Neutral, matter-of-fact stance |
| Cross-cutting | Conscientiousness (C) | Provide organised, structured guidance | Flexible, relaxed, spontaneous demeanour |
| Metric | Regulated M () |
Baseline M () |
Difference | Cohen’s d | 95% CI | p-value |
| Personality Needs* | 2.00 (0.00) | 0.20 (0.58) | 1.80 | 4.42 | [3.1, 8.1] | <0.001*** |
| Emotional Tone | 2.00 (0.00) | 2.00 (0.00) | 0.00 | 0.000 † | [-0.2, 0.2] | 1.000 not significant |
| Relevance & Coherence | 2.00 (0.00) | 1.97 (0.26) | 0.03 | 0.183 | [0.0, 0.3] | 0.319 not significant |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).