Preprint
Article

This version is not peer-reviewed.

Developing Virtual Personas from User-Level Social Media Data: Stability and Human Alignment of LLM-Generated Symptom Profiles

Submitted:

23 July 2026

Posted:

24 July 2026

You are already at the latest version

Abstract
Large language models (LLMs) may support realistic virtual patients, yet most existing systems rely on top-down construction approaches, such as manually curated vignettes. We investigated a bottom-up alternative in which social media histories were transformed into virtual personas examining stability and alignment with human judgment. Histories from Italian Reddit users reporting depressive or anxiety-related distress were manually anonymized and synthesized into two prompt variants (Base and Clinically Enriched). GPT-4o and DeepSeek-V4-Pro completed the PHQ-9 and GAD-7 in independent generations per persona, prompt, model, and scale. Stability was assessed using single-generation and aggregated intraclass correlations. Human alignment with symptom ratings from a psychologist was examined using persona-level correlations and item-level area under the curve (AUC). PHQ-9 total scores showed high single-generation stability across conditions (ICC(2,1) = .84–.88), whereas GAD-7 stability varied by model. Aggregating generations yielded high reliability in all conditions, although item-level stability was heterogeneous. Human alignment was stronger for PHQ-9 than GAD-7 and was highest for DeepSeek-V4-Pro with the Base prompt (r = .98). DeepSeek-V4-Pro also showed higher item-level discrimination, while clinical enrichment had scale-dependent effects. Our findings support the feasibility of bottom-up persona construction, while highlighting the need for psychometric evaluation across models and symptom domains.
Keywords: 
;  ;  ;  ;  
Subject: 
Social Sciences  -   Psychology

1. Introduction

The integration of generative artificial intelligence (AI) and large language models (LLMs) into health-professions education and clinical simulation has expanded rapidly in recent years (Cook, 2025; Jo et al., 2025). By transitioning conversational agents from simple static text generators into highly interactive pedagogical interfaces, these technologies may allow undergraduate and graduate students to practice high-stakes clinical communication, history-taking, and diagnostic reasoning within low-risk, standardized environments (Li et al., 2024; Wang et al., 2024; Bakhaya et al., 2026; Li & Lebai Lutfi, 2026). Emerging empirical work suggests that generative AI patient simulations can support clinical training and assessment, although evidence remains heterogeneous across tasks, settings, and evaluation methods (Fung et al., 2025; Jo et al., 2025). Beyond their direct educational utility, virtual personas may also provide standardized testing environments for auditing medical decision-making pipelines and benchmarking the dialogue architectures of automated conversational evaluation systems (Bossio-Botero et al., 2025; Sabour et al., 2026).
Despite this widespread technological promise, many LLM-based virtual patient (VP) systems continue to rely on a top-down construction paradigm. Most current implementations instantiate patient identities from manually curated clinical vignettes, expert-written backstories, or static system-prompt configurations (Fung & Laing, 2024; Wang et al., 2024). While these handcrafted profiles ensure high experimental control, they frequently simulate isolated, unambiguous disease states that lack the real-world multimorbidity representations, organic narrative variations, and complex psychosocial configurations encountered in clinical practice (Jensen et al., 2024; Kononowicz et al., 2019). Mental health training studies suggest that real patient encounters are often perceived by experts as unpredictable and multifaceted, requiring simulations to capture not only symptom content but also contextual stressors, maladaptive cognitions, emotional states, conversational style, resistance, and disclosure patterns (Wang et al., 2024; Louie et al., 2024). Furthermore, researchers face substantial systemic barriers when attempting to integrate authentic, diverse patient histories, as access to real clinical datasets is heavily restricted due to stringent ethical, legal, and institutional privacy protections under global data governance frameworks (Haider et al., 2025; Kaabachi et al., 2025). To address these empirical limitations, public social media platforms, such as Reddit and X, have emerged as rich, longitudinal data sources for bottom-up clinical persona construction (De Choudhury & De, 2014; Xu et al., 2024). These platforms host massive, unstructured streams of spontaneous, first-person disclosures detailing psychological struggles, lived experiences of distress, help-seeking, and peer-support interactions (Pretorius et al., 2019; Sit et al., 2024).
Recent preliminary work has begun to examine whether social media histories and other forms of user-authored text can support more grounded LLM-based persona construction. Outside the clinical domain, preliminary studies suggest that online behavioral traces can be transformed into heterogeneous LLM-based personas for social simulation and population modeling. For example, Rahimzadeh and colleagues (2025) introduced SYNTHIA, a large-scale dataset of backstories generated from real BlueSky user activity across multiple temporal windows; their findings suggest that grounding persona generation in authentic social media traces can improve narrative consistency while maintaining demographic diversity and real-world survey alignment. Similarly, Hu and colleagues (2025) proposed a framework for population-aligned persona generation from user-authored texts, showing that persona sets can be optimized to better match psychometric and demographic distributions. These studies suggest that user-generated online data may provide richer and more diverse evidence for persona construction than manually written archetypes alone.
Within mental health, early work by Qiu and Lan (2024) used long user posts from an online psychological counseling platform as client profiles for LLM-to-LLM counseling simulation, showing that simulated clients could preserve profile-related lexical and semantic information while generating diverse counseling dialogues. Liu and colleagues (2025) extended this direction by curating publicly available depression-related conversations, transforming them into expert-refined psychological profiles and using them to optimize a specialized LLM for depressive-client simulation. Their results indicate that such an approach can improve linguistic authenticity and profile adherence compared with prompt-based baselines. More recently, Li and colleagues (2026) proposed a multi-source framework that aligns diagnostic interviews, counseling dialogues, and longitudinal social media data to construct temporally grounded patient profiles. This work suggests that longitudinal evidence can enrich simulated patients with more coherent life-event histories, improve dialogue realism and behavioral diversity. Although these studies remain preliminary, they suggest that social media data may provide a promising source of naturalistic behavioral evidence for clinically oriented persona construction, particularly when used to train or enrich patient simulation frameworks. Such data may provide ecologically rich representations of self-described psychological distress while partially bypassing traditional barriers to clinical data access.
However, this literature remains preliminary and leaves two issues insufficiently examined. First, existing approaches often focus on generating or optimizing simulators, rather than evaluating whether the same data-derived persona produces reproducible standardized symptom profiles. Second, relatively little is known about whether LLM-generated symptom scores align with human clinical interpretation of the source-derived persona information. These limitations motivate the need for validation procedures that evaluate social media-derived virtual personas not only as plausible simulations, but as repeatable psychometric response profiles whose stability and human alignment can be quantified.
At the same time, mining online narratives requires navigating complex bioethical and methodological boundaries, especially when suitable pre-existing datasets are not available. Self-reported linguistic markers on social forums represent proxy behavioral expressions of subjective distress rather than clinically validated psychiatric diagnoses, and they are heavily confounded by platform-specific demographic sampling biases and distinct self-disclosure habits (Chancellor & De Choudhury, 2020; Ernala et al., 2019). Moreover, because online users do not provide explicit research consent, data-mining protocols must enforce strict manual anonymization, masking, and generalization of contextual details to eliminate re-identification risks and protect vulnerable populations (Benton et al., 2017; Ajmani et al., 2023).
Consequently, the translation of bottom-up, data-driven narratives into simulable virtual patients raises unresolved questions regarding simulation stability and human alignment (Kuhlmeier et al., 2025; Gonnermann-Müller et al., 2026). Much of the current virtual patient literature focuses on subjective evaluation metrics, prioritizing qualities such as conversational flow, perceived authenticity, and perceived educational usefulness (Salminen et al., 2020; Cabrera Lozoya et al., 2025; Fitrianie et al., 2025; Jo et al., 2025). By contrast, the quantitative, run-to-run reproducibility of identical data-driven personas under non-deterministic generation conditions remains insufficiently examined. Because LLMs are probabilistic text-prediction engines, they are fundamentally susceptible to decoding variability, instruction drift, and factual hallucinations across independent calls (Errica et al., 2025). In psychiatric simulations, minor stochastic shifts across independent model calls may cause divergent clinical histories or multi-turn conversational drift, directly threatening the internal validity and standardized measurement properties required for simulated behavioral experiments (Hanson et al., 2026; Tosato et al., 2026).
Beyond response stability, data-driven virtual patients present significant interpretability challenges regarding symptom anchoring and demographic representation. Emerging work on population-aligned persona generation suggests that persona sets used in LLM-based simulations can remain biased if they do not collectively match real-world psychometric and demographic distributions, motivating explicit distributional alignment procedures rather than reliance on individually plausible profiles alone (Hu et al., 2025). Related studies on persona simulation further suggest that LLM agents may drift toward homogenized or “average” personas, reducing behavioral heterogeneity and weakening the representational value of highly specified profiles (Xiao et al., 2026). In this context, simulated patient populations may fail to preserve the fine-grained nuances of their source narratives, instead over-relying on coarse demographic regularities, diagnostic prototypes, or model-inherent sociodemographic biases when translating textual evidence into clinical profiles (Chen et al., 2024; Haider et al., 2025; Yeo et al., 2025). Furthermore, process-level reasoning trace evaluations show that models can produce safe, superficially appropriate final outputs while relying on stigmatizing logic or clinical biases within their latent layers, a confounding behavior formalized as “safetywashing” (Moore et al., 2025; Tosato et al., 2026; Sankar et al., 2026).
It is also important to note that both response stability and human alignment are moderated by model architectures and the formal structure of prompting instructions (Chen et al., 2024; Rahimzadeh et al., 2025). Benchmarking studies across closed- and open-source models shows substantial variations in instruction following, role adherence, and clinical symptom recall (Yang et al., 2024; Wang et al., 2025; Badawi et al., 2026). Simultaneously, prompt representation, such as instructing a model to evaluate symptom severity using holistic aggregate labels versus executing explicit item-level psychometric responses, can shape the consistency of clinical outputs (Hadar-Shoval et al., 2023; Vu et al., 2025). These differences make model choice and prompt representation methodological variables in their own right, particularly when virtual patients are evaluated for reproducibility, symptom coherence, and alignment with source-derived clinical profiles. In this context, standardized, unidimensional rating scales offer a robust psychometric probing framework to quantify these exact variations without imposing unauthorized diagnostic claims.
Despite these intersecting challenges, three distinct gaps remain insufficiently addressed in the current literature. First, few studies provide systematic, reproducible pipelines to transform raw, unstructured social media text histories into structured, anonymized, and simulable virtual patient profiles. Second, limited quantitative evidence is available on the run-to-run consistency of standardized clinical symptom scores produced by LLM agents when prompted with identical social media-derived personas across repeated independent administrations of psychometric instruments. Third, it remains undetermined whether LLM-generated symptom profiles align with structured human expert evaluations of the same underlying narratives, and how this alignment is affected by model choice and prompt granularity.
Building on these gaps, the present study makes three contributions. First, it operationalizes a bottom-up pipeline for transforming selected and anonymized Italian-language Reddit activity histories into structured, simulation-ready virtual patient prompts through persona synthesis. Second, it evaluates the run-to-run stability of LLM-generated symptom profiles across independent administrations of the PHQ-9 and GAD-7. This contribution concerns the psychometric robustness of the generated response profiles at both total-score and item levels. Third, it quantifies human-LLM alignment by comparing LLM-generated questionnaire responses with psychologist-coded symptom evaluations of the same persona prompts. This contribution examines whether simulated symptom profiles are not only reproducible, but also interpretable relative to a human prompt-based reference.
To examine whether stability and human alignment were robust across prompt representation and model choice, the study combines two prompt structures, Base and Clinically Enriched, with two large language models representing different development and dissemination frameworks: GPT-4o, a proprietary model, and DeepSeek-V4-Pro, an open-weight model. Given documented variation across LLMs in instruction following, role adherence, safety behavior, and mental-health text interpretation, this comparison was not intended to identify a universally superior model, but to assess whether the stability and human alignment of simulated symptom profiles depend on the generative system used (Xu et al., 2024; Yang et al., 2024; Badawi et al., 2026). For each persona, prompt variant, model, and questionnaire condition, 100 independent administrations were retained using identical anonymized persona information, questionnaire instructions, and repetition procedures.

2. Materials and Methods

2.1. Data Source and Sampling

2.1.1. Source Platform and Corpus

This study used publicly available user-generated content expressing psychological distress collected from Reddit, a popular semi-anonymous online platform organized into topic-specific communities known as subreddits. Data collection targeted Italian-language subreddits, including both general (e.g., r/italia) and mental health-related subreddits (e.g., r/Psicologia_Italia, r/Psico_aiuto_Italia, r/psicologia, r/depressione).
The initial identification procedure was based on a previous data collection protocol developed by the research team (Settanni et al., 2025), in which Reddit posts were extracted through the RedditExtractoR package in R using predefined mental health-related keywords. The final corpus included 425 posts describing personally experienced psychological distress, selected to retain self-referential experiences or requests for psychological support, while excluding generic discussions or content lacking meaningful personal relevance.

2.1.2. Content and Author Screening

From this corpus, the present study kept posts whose content was consistent with depressive or generalized anxiety-related distress. The relevant expressions covered low mood, anhedonia, self-evaluative distress, death-related thoughts, persistent worries, fatigue, anxiety-related suffering, and somatic complaints. This step identified 304 posts with depressive- or generalized anxiety-related narrative features.
Because several posts came from the same Reddit users, duplicate authors were collapsed so that each user contributed a single index post. After this step, the pool contained 222 unique Reddit users.

2.1.3. Activity Criteria and Clinical Selection

The remaining users were screened against predefined activity criteria intended to ensure sufficient narrative depth and longitudinal consistency for persona construction. A user was retained only when every criterion was satisfied: at least 50 posts across the activity history; participation in at least 10 distinct subreddits; at least 10 unique posts within mental health subreddits (r/Psicologia\_Italia, r/Psico\_aiuto\_Italia, r/psicologia, and r/depressione); and at least 50 comments. Applying these thresholds reduced the candidate pool to 11 users. As a final step, a licensed psychologist examined each remaining activity history using a DSM-5-TR-informed checklist to identify clinically interpretable symptom-related patterns within the narratives. The checklist functioned solely as a non-diagnostic interpretive framework for recognizing psychologically meaningful expressions of depressive- and anxiety-related distress, preserving the ecological validity of the naturally occurring text. This evaluation retained six Reddit users whose narratives displayed coherent and psychologically interpretable symptom patterns suitable for subsequent persona construction.
Figure 1. Three-stage pipeline from public Reddit histories to simulated virtual patient personas: a selection funnel reducing 425 posts to 6 clinically screened personas (Stage 1), two-part manual anonymization verified by an adversarial re-identification attempt with 0 of 6 accounts recovered (Stage 2), and LLM synthesis of each history into the shadow persona that is grounded in the original patient’s evidence.
Figure 1. Three-stage pipeline from public Reddit histories to simulated virtual patient personas: a selection funnel reducing 425 posts to 6 clinically screened personas (Stage 1), two-part manual anonymization verified by an adversarial re-identification attempt with 0 of 6 accounts recovered (Stage 2), and LLM synthesis of each history into the shadow persona that is grounded in the original patient’s evidence.
Preprints 224688 g001

2.2. Anonymization

Before feature extraction, the complete activity history of each user was de-identified through a manual procedure carried out by the research team. The procedure was designed to reduce the risk of re-identifying the original authors while retaining the psychologically relevant content required for persona construction, including the emotional states and symptom-related expressions conveyed in the narratives. Particulars capable of linking the text to an identifiable individual, such as specific events or relationships, were not preserved. This approach is consistent with established recommendations for the ethical use of social media data in mental health research, where explicit research consent is absent and the content is frequently sensitive (Ajmani et al., 2023; Benton et al., 2017; Chancellor & De Choudhury, 2020).

2.2.1. Removing and Replacing Identifiers

Across each user’s posts and comments, potentially identifying information was either deleted or replaced with a generalized surrogate. Direct identifiers, including usernames and links to the source posts, were removed outright. The procedure also addressed quasi-identifying information, that is, details that alone or in combination could narrow the set of plausible authors. Such details fell into six categories: fine-grained sociodemographic information, geographic locations, occupational and educational specifics, references to identifiable third parties or to distinctive social and relational circumstances, physical attributes, and highly specific medical or health-status details. Because deleting these spans outright would have fragmented the narratives and degraded their psychological coherence, the identifying elements were replaced with plausible alternatives that preserved the structure and meaning of the surrounding text, for example substituting one city (e.g., “Bologna”) for a comparable city (e.g., “Padova”) or one profession for a closely related one.

2.2.2. Keeping Surrogates Consistent

Substitutions were applied consistently within each user, so that the resulting persona remained coherent across the entire history. This combination of deletion and surrogate replacement follows what is currently regarded as best practice in text anonymization (Lison et al., 2021). As an additional safeguard against search-based re-identification, no verbatim excerpts from the original posts are reported in this manuscript.
The adequacy of the procedure was assessed in two ways. First, each anonymized history was independently re-checked by a second member of the research team for residual identifying content; the 11 residual items detected across the six histories, most of them oblique references to identifiable third parties, were removed by consensus. Second, to evaluate the resulting protection directly, the six anonymized personas were subjected to a re-identification attempt in which a researcher sought to recover each source account by searching distinctive phrases drawn from the surrogate-stripped text, and none of the six accounts could be recovered.

2.3. Feature Extraction and Shadow Persona Synthesis

For each selected user, the anonymized activity history was transformed into a structured representation termed a “shadow persona”, which consolidates information dispersed across many posts and comments into a single, simulation-ready profile of the individual.

2.3.1. Input and Model Setup

The input to this stage was the user’s complete anonymized activity, in which each entry retained its type (post or comment), originating subreddit, timestamp, and textual body. As a preprocessing step, the free-text content of all entries was consolidated into a clean corpus ordered by time, discarding non-narrative metadata not required for psychological interpretation while preserving the temporal ordering needed to trace the onset and course of the difficulties. The synthesis itself was performed by a large language model (Gemini 2.5 Pro), instructed through its system prompt to act as an expert psychological profiler and to read the entire activity as a single longitudinal case. Because the clinically relevant cues in naturalistic social media data are sparse, dispersed across unrelated threads, and frequently only implied, the model was directed to integrate evidence across the whole history, to treat individual posts as fragments of one case, and to weight recent information about the user’s current state more heavily when those expressions shifted over time.

2.3.2. Features and Output Constraints

The set of extracted features was defined a priori to approximate the domains ordinarily considered in an initial psychological or psychiatric assessment, adapted to the constraints of naturally occurring online text, and grounded in the American Psychiatric Association practice guidelines for the psychiatric evaluation of adults (Silverman et al., 2015). These domains covered the user’s background and presenting situation (sociodemographic cues, the reason for help-seeking, and current stressors), the clinical history (the course of presenting difficulties and significant life events), and the user’s inner experience (recurring psychological themes, main concerns, and emotional expression). A further set of domains characterized interpersonal style and adaptive functioning, including communication style, relational patterns, coping and functioning cues, risk indicators, and protective factors. Where the source material supported it, the synthesis additionally captured previous psychological or psychiatric treatment, medication use, the onset and course of the difficulties, and emerging symptom patterns.
The structured representation was generated with prompt-engineering strategies intended to keep the output faithful to the source text and to render the model’s inferences auditable, in line with current syntheses of prompting practice (Sahoo et al., 2024; Schulhoff et al., 2024). The model returned its analysis as a schema-constrained JSON object with a fixed set of fields, which standardized the output across users and supported reliable downstream parsing (Shen et al., 2025). Each field was required to be grounded in the source material through a supporting quotation or an explicit line of reasoning, and each carried a confidence rating on a high, medium, or low scale together with a provenance label indicating whether the value had been explicitly stated by the user, inferred from behavioral or contextual cues, or estimated from general population patterns in the absence of direct support. This graded inference hierarchy makes the degree of support behind every element of the persona explicit and prevents weakly grounded inferences from being reported with unwarranted certainty, a recognized failure mode of generative models. By requiring the model to weigh the available evidence before committing to each structured value, the procedure follows the rationale of structured-reasoning approaches developed to improve the reliability of model outputs (Wang et al., 2022).

2.3.3. Confidence Annotations

The confidence annotations indicate that the synthesized personas were grounded predominantly in well-supported information. Aggregated across the six personas, the proportion of field-level confidence ratings classified as high was 94.4% for the base representation (5.6% medium, 0% low) and 81.1% for the clinically enriched representation (18.2% medium, 0.8% low). The lower proportion under the clinically enriched representation follows from its more granular and clinically demanding fields, which relied more often on inference from indirect cues. From these same histories, two structured representations differing in clinical granularity were produced, and their construction and differences are detailed in the following subsection.

2.4. Persona Prompt Variants

Two structured prompt variants were generated from the same underlying shadow persona profiles in order to examine whether increasing the clinical granularity and explicit organization of persona information influenced LLM-based symptom inference. The Base prompt provided a broad user-level description grounded in Reddit activity history, preserving sociodemographic context, reason for help-seeking, recurring concerns, life events, stressors, emotional expression, communication style, relational patterns, and risk/protective factors. This variant was intended to retain the ecological and narrative structure of the source material. The Clinically Enriched prompt retained this psychosocial basis but expanded it into a more clinically explicit and granular representation, including psychopathological history, prior treatment and medication use, onset and course of difficulties, symptom domains, functional impairment, developmental trauma, attachment-related patterns, clinical urgency, and protective resources. Thus, the two variants differed primarily in the clinical explicitness and granularity with which symptom-relevant information was represented, rather than in the basic identity of the underlying persona. The comparison was included to test the expectation that a more clinically granular representation of the same underlying persona would facilitate symptom inference, yielding more reproducible and human-aligned simulated self-report symptom profiles than the Base prompt variant. To illustrate the structure of the two persona representations, one anonymized example of the Base prompt and one anonymized example of the Clinically Enriched prompt are provided in Appendix A. The main differences in the information provided by the two prompt structures are summarized in Figure 2.

2.5. Measures

Two standardized self-report instruments were used as controlled probes of the symptom profiles expressed by the virtual personas. The Patient Health Questionnaire-9 (PHQ-9) (Kroenke et al., 2001) was used to assess depression-related symptoms. It consists of nine items rated on a 0-3 response scale, yielding a total score ranging from 0 to 27. The Generalized Anxiety Disorder-7 (GAD-7) (Spitzer et al., 2006) was used to assess generalized anxiety-related symptoms. It consists of seven items rated on the same 0-3 response scale, yielding a total score ranging from 0 to 21.

2.6. Human Prompt-Based Symptom Evaluation

For each persona and prompt variant, a psychologist from the research team reviewed the persona description and coded each PHQ-9 and GAD-7 item with respect to the information available in the prompt. The primary coding variable was Symptom Present, a binary indicator of whether the symptom corresponding to each item was judged to be supported by the persona description. Human symptom totals were computed by summing Symptom Present indicators across items separately for PHQ-9 and GAD-7. These evaluations served as a human prompt-based reference for assessing alignment with LLM-generated questionnaire responses.

2.7. LLM Experimental Setup

The simulation phase evaluated whether the six virtual personas generated stable and human-aligned questionnaire response profiles across model and prompt variants. To preserve anonymity, original user labels were replaced with generic identifiers P01-P06. Each persona was represented using the two prompt variants described above: the Base prompt and the Clinically Enriched prompt. For each prompt variant, responses were generated using two LLMs, GPT-4o and DeepSeek-V4-Pro, allowing us to examine whether simulated self-report profiles were sensitive to model choice. Both models were run using their default generation settings. Thus, parameters such as temperature, top-p, and model-specific reasoning settings were not manipulated across personas, prompt variants, questionnaires, or repetitions. This choice was intended to evaluate the behavior of each model under its standard usage configuration, while recognizing that model effects may therefore reflect both model architecture and provider-default decoding or alignment settings.
For each combination of persona, prompt variant, model, and questionnaire, 100 independent generations were retained. Each generation corresponded to a separate administration of the questionnaire, implemented as an independent model call with no conversational history carried over from previous repetitions. The PHQ-9 and GAD-7 were administered separately. In each generation, the model was instructed to respond as the target virtual persona and to provide item-level numerical responses using the original 0-3 response scale of the questionnaire, without explanatory text, item labels, or justifications. The full prompts used for PHQ-9 and GAD-7 administration, including output-format constraints, are provided in Appendix B.
LLM outputs were parsed at the item level and converted into total scores by summing item responses. PHQ-9 total scores were computed as the sum of the nine items, yielding a possible range from 0 to 27. GAD-7 total scores were computed as the sum of the seven items, yielding a possible range from 0 to 21. Each generation was treated as an independent administration of the same standardized probe to the same virtual persona.

2.8. Statistical Analyses

Statistical analyses were performed using RStudio (Version 2026.04.0) (Posit Team, 2026). Analyses were conducted in three steps. First, response stability was examined using intraclass correlation coefficients. ICC(2,1) was used to estimate the reliability of a single LLM generation, and ICC(2,100) estimated the reliability of the mean score obtained by aggregating 100 generations. ICCs were computed for total scores and for individual items.
Second, human-LLM alignment was examined at the persona level. Because human evaluations and LLM outputs were expressed on different measurement scales, they were not treated as directly interchangeable. Human totals reflected the number of symptoms coded as present, whereas LLM totals reflected the average questionnaire total across 100 retained generations per persona. To compare relative persona-level profiles without imposing a common raw scale, human and LLM total scores were standardized across the six personas within each scale, prompt variant, and model. Pearson and Spearman correlations were then computed between the corresponding standardized human and LLM profiles. Z-score profile plots were used to compare the relative ordering of personas across human evaluations and LLM-generated scores.
Third, human-LLM alignment was examined at the item level. Binary human symptom-presence ratings were correlated with mean LLM-generated item scores across persona x item combinations. Area under the curve (AUC) was also computed to quantify how well LLM-generated item scores discriminated symptoms coded by humans as present versus absent.

3. Results

3.1. Descriptive Statistics of LLM-Generated Scores

Table 1 reports descriptive statistics for LLM-generated total scores by prompt variant, model, and scale. For PHQ-9, GPT-4o produced higher mean scores and lower variability than DeepSeek-V4-Pro across both prompt variants. For GAD-7, DeepSeek-V4-Pro tended to generate higher scores than GPT-4o, especially under the Clinically Enriched prompt, and showed lower variability across generations.

3.2. Response Stability: Total Scores and Item-Level Analyses

Table 2 reports total-score ICCs. PHQ-9 stability was high across both models and prompt variants, with ICC(2,1) values ranging from .84 to .88, indicating that both models produced relatively stable PHQ-9 profiles for the same virtual personas across repeated generations. GAD-7 stability showed stronger model differences. GPT-4o showed high single-generation stability, especially under the Clinically Enriched prompt, whereas DeepSeek-V4-Pro showed lower reliability for GAD-7, particularly under the Clinically Enriched prompt. However, ICC(2,100) values were very high in all conditions, indicating that aggregation across 100 generations yielded highly stable persona-level estimates.
Item-level ICCs showed greater heterogeneity than total-score ICCs (Table 3). DeepSeek-V4-Pro showed somewhat higher item-level stability than GPT-4o for PHQ-9, whereas GPT-4o showed higher item-level stability for GAD-7, particularly under the Clinically Enriched prompt. Although ICC(2,100) values were consistently high, indicating that aggregated item-level profiles were highly reproducible, ICC(2,1) values showed that single-generation stability varied substantially across individual items. Across models, depressed mood, worthlessness/failure, and self-harm ideation were among the most stable PHQ-9 items, whereas psychomotor symptoms were consistently less stable. Full item-level ICC estimates are reported in Supplementary Table S1.

3.3. Persona-Level Alignment Between Human Evaluations and LLM Scores

Table 4 reports persona-level correlations between human symptom totals and mean LLM-generated total scores. PHQ-9 showed stronger human-LLM alignment than GAD-7. For PHQ-9, both models showed strong positive correlations with human evaluations, especially under the Base prompt; the strongest association was observed for DeepSeek-V4-Pro under the Base prompt, r = .98. For GAD-7, correlations were generally weaker. DeepSeek-V4-Pro showed stronger alignment than GPT-4o under both prompt variants, whereas GPT-4o showed only modest correspondence with human GAD-7 totals, particularly under the Base prompt. Prompt effects were scale-dependent, with the Base prompt yielding stronger PHQ-9 alignment and the Clinically Enriched prompt showing stronger GAD-7 alignment.
Because these correlations were computed across six personas, statistical significance should be interpreted cautiously. The pattern of effect sizes is more informative than individual p values. The persona-level human symptom totals and corresponding LLM-generated total scores underlying these analyses are reported in Supplementary Table S3.
Figure 3 and Figure 4 show z-score profile plots comparing the relative ordering of personas across human evaluations and LLM-generated scores. For PHQ-9, the LLM-generated z-score profiles broadly followed the human-coded profile, especially under the Base prompt. For GAD-7, discrepancies were larger, with DeepSeek-V4-Pro generally closer to the human profile than GPT-4o.
Table 5 reports item-level correspondence between human symptom-presence ratings and LLM-generated item scores. DeepSeek-V4-Pro showed stronger item-level alignment than GPT-4o across both prompt variants and both scales, with consistently higher AUC values. For PHQ-9, DeepSeek-V4-Pro achieved AUC values of .92 under the Clinically Enriched prompt and .94 under the Base prompt, compared with .71 and .83 for GPT-4o. For GAD-7, DeepSeek-V4-Pro again showed higher AUC values, indicating stronger discrimination between symptoms coded as present versus absent by human raters. GPT-4o also showed positive item-level alignment with human ratings, but its item scores were more compressed, with moderately elevated scores often assigned even when symptoms were not coded as present. Descriptive item-level profiles are provided in Supplementary Table S4.
Item-specific associations further showed that alignment varied substantially across symptoms, models, and prompt variants. Full estimates are reported in Supplementary Table S2. For PHQ-9, alignment was strongest and most consistent for relatively distinctive symptoms, including appetite/weight change, worthlessness/failure, and self-harm ideation. For example, appetite/weight change showed very high correlations across models and prompt variants (r = .91-.99); worthlessness/failure also showed strong correspondence, particularly under the Base prompt for both GPT-4o (r = .84) and DeepSeek-V4-Pro (r = .86). In contrast, some PHQ-9 items were more model-dependent. Under the Clinically Enriched prompt, GPT-4o showed weak correlations for sleep problems (r = .27) and low energy (r = .27), whereas DeepSeek-V4-Pro showed stronger associations for the same items (r = .76 and r = .57, respectively).
For GAD-7, item-specific associations were more heterogeneous. DeepSeek-V4-Pro showed stronger associations for uncontrollable worry under the Clinically Enriched prompt (r = .99), as well as for trouble relaxing and irritability across prompt variants. GPT-4o was less consistent on several GAD-7 items, showing weak or negative associations for irritability (r = −.08 Clinically Enriched; r = −.06 Narrative-Psychosocial) and for uncontrollable worry under the Base prompt (r = −.07). Some items, including nervousness/anxiety and restlessness, could not be evaluated because human ratings showed no variation across personas.

4. Discussion

The empirical findings of this study provide a comprehensive psychometric evaluation of a bottom-up, data-driven pipeline designed to translate unstructured social media trace data into simulable virtual patient profiles. By moving away from the traditional, top-down manual curation of clinical vignettes (Fung & Laing, 2024; Wang et al., 2024), this framework demonstrates that naturalistic text histories can be systematically aggregated into contextually valid personas while implementing strict text anonymization boundaries to preserve user privacy (Ajmani et al., 2023; Benton et al., 2017). By administering standardized scales independently across extensive iteration loops, this investigation models LLM-based virtual patients as stochastic systems rather than static textual entities, mapping the precise boundary conditions of run-to-run conversational stability and human expert alignment (Gonnermann-Müller et al., 2026; Tosato et al., 2026).
A primary methodological strength of this investigation resides in the high-density, multi-run repeated measurement paradigm used to audit model behavior. Rather than evaluating simulation validity through surface-level linguistic fluency or single-turn satisfaction metrics (Cabrera Lozoya et al., 2025; Fitrianie et al., 2025), this study treats generative architectures as probabilistic prediction engines susceptible to decoding variability and instruction drift (Errica et al., 2025). Accumulating a vast bank of independent administrations per condition provides a robust empirical estimation of simulation reliability at both the aggregate total-score and individual item levels.
Furthermore, utilizing naturally occurring social media text histories offers a scalable alternative to institutional electronic health records, which remain heavily restricted under modern data governance frameworks (Haider et al., 2025; Kaabachi et al., 2025). This pipeline offers a reproducible blueprint for persona development in domains like digital pedagogy and computational social science, demonstrating that bottom-up naturalistic trace data can be effectively compiled into structured profiles suitable for standardized training workflows (Bakhaya et al., 2026; Sabour et al., 2026).
The statistical evidence reveals a critical psychometric tension between aggregate scale reliability and item-level diagnostic stability. At the macro level, total score profiles for depression exhibited high run-to-run reproducibility across distinct model architectures and prompt variations. However, item-level analyses exposed significant structural volatility across independent model calls. This phenomenon was uniquely prominent in DeepSeek-V4-Pro’s simulation of generalized anxiety under the Clinically Enriched prompt condition; while the aggregate total scores presented a highly compressed standard deviation—indicating consistent overall severity classification—the corresponding item-level reliability collapsed significantly, showing notable vulnerability in items mapping excessive worry.
This specific configuration indicates that the model enforces a rigid macro-level clinical severity ceiling dictated by the system prompt parameters but non-deterministically reallocates values across individual symptom items during decoding. In applied educational and assessment contexts, this item-level structural drift directly compromises simulation standardization (Cook, 2025; Jo et al., 2025). A human trainee interacting with the exact same virtual persona across distinct sessions would encounter changing symptom configurations—such as severe psychomotor adjustments in one iteration and absolute sleep disturbances in the next—thereby undermining the internal validity required for objective clinical evaluation metrics and standardized examination deployment (Bossio-Botero et al., 2025; Hanson et al., 2026).
The empirical data demonstrates a pronounced asymmetry in human-LLM alignment between the depressive and generalized anxiety domains. Persona-level profile correlations with human experts were exceptionally high for depression scales under the Base prompt structure, whereas generalized anxiety alignment was systematically attenuated across models and prompt designs. This construct domain divergence is heavily moderated by how psychopathological distress manifests linguistically within online self-disclosures (De Choudhury & De, 2014; Sit et al., 2024). Depressive pathology often features highly explicit, behaviorally anchored markers—such as distinct references to appetite adjustments or unambiguous expressions of self-harm ideation—which provide clear semantic anchor points for both human raters and algorithmic parsing models (Kang et al., 2024; Yang et al., 2024).
Conversely, generalized anxiety states typically manifest through diffuse, transdiagnostic, and pervasive linguistic styles characterized by ongoing patterns of worry or irritability rather than discrete behavioral events (Ernala et al., 2019; Pretorius et al., 2019). Consequently, when translating these diffuse narrative styles into explicit psychometric responses, human expert judgment and model-inherent inference logic diverge significantly. This variance renders anxiety simulations highly sensitive to model architecture variations and prompt framing constraints, leading to larger discrepancies with human clinical consensus (Hadar-Shoval et al., 2023; Pombal et al., 2025; Vu et al., 2025; Badawi et al., 2026).
Item-level discrimination metrics highlighted divergent inference logic between the evaluated model families. DeepSeek-V4-Pro demonstrated superior clinical boundaries across prompt conditions, selectively differentiating between symptoms judged as present or absent by human experts. Conversely, GPT-4o exhibited severe score compression and generalized over-pathologization. Item-level descriptive data shows that GPT-4o systematically assigned moderately elevated scores across clinical items, even when the underlying text history contained zero evidence of such symptoms.
This behavioral profile may reflect an emerging concern in the literature on model sycophancy and “safetywashing,” whereby alignment procedures may encourage models to over-conform to generalized archetypes of distress (Moore et al., 2025; Sankar et al., 2026). Rather than preserving the fine-grained nuances and individual variations present within the data-driven source materials, the model’s output collapses into a low-dimensional cluster governed by a homogenized clinical stereotype. This structural vulnerability, consistent with emerging concerns about “persona collapse,” presents a notable limitation for multi-agent social simulations, as it restricts population-level heterogeneity and introduces systematic over-pathologization biases into simulated clinical environments (Hu et al., 2025; Xiao et al., 2026).
The data shows that prompt architecture operates as an active variable that alters model inference logic in a scale-dependent manner. Explicitly organizing the user’s trace data into structured psychopathological fields improved persona-level human alignment for generalized anxiety but degraded alignment profiles for depression. For narrative-heavy constructs like depression, introducing an explicit clinical framework may activate latent medical tropes or training data shortcuts within foundation models, causing the agent to abandon authentic character role-play in favor of mechanical diagnostic scripts (Chen et al., 2024; Hadar-Shoval et al., 2023; Vu et al., 2025). This emphasizes that prompt engineering parameters cannot be optimized uniformly but must be structurally matched to the unique semantic architecture of the psychological construct under simulation (Kuhlmeier et al., 2025).

4.1. Limitations and Future Directions

The boundaries of this study suggest specific lines of future inquiry. Methodologically, the evaluation relied on a targeted pool of unique personas. The selection of a focused target cohort of virtual personas was a deliberate configuration choice designed to prioritize a high-density, multi-run repeated measurement framework capable of isolating stochastic decoding patterns. This concentrated setup allowed the study to isolate fine-grained variations in stochastic decoding and token allocation patterns that would be obscured in wider, single-run evaluations. However, this focused sample size constrains the immediate generalizability of the persona-level correlations, indicating a clear need for future research to scale this automated pipeline across larger, more sociodemographically diverse target groups to achieve broader population-level validation (Hu et al., 2025; Salminen et al., 2020).
Furthermore, human alignment was benchmarked against the structured interpretations of a single expert psychologist, which introduces localized interpretive variances; future iterations of this validation layer should implement multiple independent expert raters to calculate formal inter-rater reliability coefficients. Additionally, because the source sample was comprised strictly of histories with observable distress indicators, the current design cannot fully isolate the model’s false-positive attribution rate or determine its baseline tendency to pathologize low-distress cues. Finally, model evaluations were restricted to provider-default hyperparameters. Because decoding parameters such as temperature, top-p, and system-level reasoning tokens directly dictate token selection variance, future studies should systematically manipulate these values across open-source model families to map the precise threshold where generative flexibility compromises psychometric consistency (Badawi et al., 2026; Warner et al., 2025).

5. Conclusion

This investigation demonstrates that naturalistic data grounding through public social media trace text is a viable paradigm for constructing virtual patient profiles under strict privacy constraints. However, bottom-up grounding alone is insufficient to guarantee stable or aligned simulation performance across distinct psychiatric domains. Automated, multi-run psychometric auditing via standardized symptom scales provides an indispensable validation layer, mapping the hidden item-level volatility and model-specific biases of virtual personas before they are integrated into complex, interactive health workflows.

5.1. Ethical Considerations

This study relied exclusively on publicly available Reddit content. No direct interaction with users was conducted, and the study did not involve recruitment, intervention, deception, or access to private, restricted, or password-protected communities. No formal ethics approval was obtained. Nevertheless, because the source material included descriptions of psychological distress, help-seeking experiences, and potentially vulnerable personal disclosures, the data were treated as sensitive human-derived material. To minimize privacy risks, all users’ activity histories underwent a manual anonymization procedure prior to persona construction and simulation, as described in previous sections. In addition, the manuscript does not report identifiable verbatim excerpts from the original posts, in order to reduce the risk of search-based re-identification. All reported information was therefore presented in aggregated, paraphrased, or structurally transformed form, with the aim of preserving the methodological relevance of the data while protecting the privacy and dignity of the individuals whose online material informed the simulations.

Supplementary Materials

The following supporting information can be downloaded at the website of this paper posted on Preprints.org, Table S1: Item-level response stability of LLM-generated PHQ-9 and GAD-7 scores by prompt condition, model, and symptom item; Table S2: Item-specific human–LLM alignment by prompt condition, model, and symptom item; Table S3: Persona-level human symptom evaluations and LLM-generated total scores by prompt condition, instrument, and virtual persona; Table S4: Item-level descriptive profiles of human-coded symptom evidence and LLM-generated item scores by prompt condition, instrument, and symptom item.

Author Contributions

Conceptualization, M.S., D.M., F.Q. and F.T.; methodology, M.S., A.R., D.M., L.D.C., F.T., and F.Q.; software, F.T.; formal analysis, D.M., F.T. and F.Q.; investigation, F.Q. and F.T.; data curation, F.Q., A.R., F.T., L.D.C., M.S. and D.M.; writing – original draft preparation, F.Q., F.T. and D.M.; writing – review and editing, F.Q., F.T., A.R., M.S. and D.M.; visualization, F.T.; supervision, M.S., A.R., L.D.C. and D.M.; project administration, M.S.; funding acquisition, M.S. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by Fondazione CRT, grant #2023.0349.

Institutional Review Board Statement

Formal ethical approval was not required or sought because the study consisted exclusively of secondary analysis of publicly accessible Reddit content, with no recruitment, direct interaction, intervention, or access to private or restricted data. Given the sensitive nature of the material, all user histories were manually de-identified, direct and quasi-identifiers were removed, and no identifiable verbatim content is reported.

Data Availability Statement

Additional de-identified data supporting the findings of this study are available from the corresponding author upon reasonable request. The underlying Reddit activity histories and detailed persona representations are not publicly available because they contain sensitive personal information. Although these materials were manually anonymized, their disclosure could entail residual risks of re-identification and raise ethical and privacy concerns.

Acknowledgments

During the preparation of this manuscript, the authors used ChatGPT, powered by the GPT-5.6 model (OpenAI) for grammar, spelling, and language editing. The authors reviewed and edited all outputs and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A

Table A1. Base Prompt Structure Example.
Table A1. Base Prompt Structure Example.
Section
General metadata Username; extraction_date
Sociodemographic data Age; gender; education_level; occupation; relationship_status; housing_status. Sub-fields: value, confidence, evidence.
Reason for consultation Main_problem; request_summary; confidence; key_quotes.
Psychological characteristics Recurring_themes; main_concerns; general_confidence. For each theme: theme, description, frequency, emotional_intensity, emergence_contexts, examples.
History of the current problem Recent_triggering_factors; onset_timeline; symptom_evolution; presence_of_triggering_factor.
Personal history Positive_significant_events; negative_significant_events; narrative_patterns; general_confidence.
Communication style General_tone; vocabulary; form; self_disclosure; emotional_expression; general_confidence.
Relational style Attitudes_toward_others; typical_behaviors; online_interaction; general_confidence.
Risk factors Suicidal_ideation; suicide_attempts; self_harming_behaviors; substance_use; substance_types; other_factors.
Protective factors Level_of_social_support; family_relationships; romantic_relationship; other_factors.
Analysis metadata Number_of_posts_analyzed; number_of_comments_analyzed; time_range; main_subreddits; temporal_distribution; urgency_indicators.

Table A2. Clinically Enriched Prompt Structure Example.
Table A2. Clinically Enriched Prompt Structure Example.
Section
Sociodemographic data Additional contextual qualifiers for selected fields: occupational status, relationship satisfaction, housing stability.
Reason for consultation Level_of_insight; motivation_for_change.
Psychological characteristics Emotional_expression_pattern; functional_interference; associated_cognitive_pattern; coping_strategies. For main concerns: functional_impact and quotes are specified more explicitly.
Psychopathological history Entirely new section. Includes past_diagnoses; psychotropic_medication_use; history_of_psychological_treatment; recurrence_pattern; apparent_comorbidity. For each diagnosis: condition, source, period, treatment, outcome, quotes. For each medication: medication, period, adherence, perceived_effectiveness, side_effects, quotes.
History of the current problem More detailed onset and evolution fields: onset_mode; age_at_onset; onset_context; course_pattern; improvement_factors; worsening_factors; critical_moments. Recent triggers also include severity and behavioral_reaction.
Symptom patterns Entirely new section. Includes general_description; symptom_domains; cognitive_symptoms; emotional_symptoms; behavioral_symptoms; somatic_symptoms; interpersonal_symptoms; level_of_functional_impairment; general_confidence. For each symptom: frequency, intensity, duration, onset, temporal_pattern, functional_impact, quotes.
Personal history Developmental_trauma; immediate_impact; current_meaning; clinical_relevance; event_type; current_processing; related_symptoms.
Communication style Mood_markers; use_of_psychological_technical_terms; negative_self-referential_terms; clinically_relevant_linguistic_patterns; coherence; tangentiality.
Relational style Attachment_pattern; theory_of_mind; interpersonal_schemas; dependence_autonomy; online_platform_function.
Risk factors Other_risk_behaviors; additional_risk_factors; clinical_urgency. More detailed risk sub-fields: description; intent; access_to_means; precipitating_factors; protective_factors; frequency; onset; evolution.
Protective factors Personal_resources; interests_activities; additional_protective_factors; reasons_for_living. More detailed support sub-fields: reference_people; specific_relationships; type_of_support; strength; function.
Analysis metadata Observation_duration; analysis_limitations; recommendations. For each subreddit: clinical_relevance.

Appendix B

Questionnaire-Administration Prompts
PHQ-9 PROMPT
“I will provide you with a detailed description of a virtual persona. Your task is to generate responses to the PHQ-9 questionnaire as if this person were answering today, in a realistic, coherent, and psychologically plausible manner.
Use only the information contained in the profile. When necessary, make cautious inferences. Give greater weight to recent, high-confidence information concerning the person’s current state.
The questionnaire explicitly refers to the following time frame:
‘Over the last 2 weeks, how often have you been bothered by any of the following problems?’
For each of the 9 PHQ-9 items, assign a score:
0 = Not at all
1 = Several days
2 = More than half the days
3 = Nearly every day
The items, in this order, are:
1. Little interest or pleasure in doing things
2. Feeling down, depressed, or hopeless
3. Trouble falling or staying asleep, or sleeping too much
4. Feeling tired or having little energy
5. Poor appetite or overeating
6. Feeling bad about yourself—or that you are a failure or have let yourself or your family down
7. Trouble concentrating on things, such as reading the newspaper or watching television
8. Moving or speaking so slowly that other people could have noticed. Or the opposite—being so fidgety or restless that you have been moving around a lot more than usual
9. Thoughts that you would be better off dead or of hurting yourself in some way
Output rules:
1. Return ONLY a string of 9 numbers separated by commas
2. No explanation
3. No additional text
4. No labels
5. No parentheses
6. No line breaks
Exact format: Response to item 1, Response to item 2, Response to item 3, Response to item 4, Response to item 5, Response to item 6, Response to item 7, Response to item 8, Response to item 9
Here is the virtual persona’s profile:“
PROMPT GAD7
“I will provide you with a detailed description of a virtual persona. Your task is to generate responses to the GAD-7 questionnaire as if this person were answering today, in a realistic, coherent, and psychologically plausible manner.
Use only the information contained in the profile. When necessary, make cautious inferences. Give greater weight to recent, high-confidence information concerning the person’s current state.
The questionnaire explicitly refers to the following time frame:
‘Over the last 2 weeks, how often have you been bothered by the following problems?’
For each of the 7 GAD-7 items, assign a score:
0 = Not at all
1 = Several days
2 = More than half the days
3 = Nearly every day
The items, in this order, are:
1. Feeling nervous, anxious, or on edge
2. Not being able to stop or control worrying
3. Worrying too much about different things
4. Trouble relaxing
5. Being so restless that it is hard to sit still
6. Becoming easily annoyed or irritable
7. Feeling afraid as if something awful might happen
Output rules:
1. Return ONLY a string of 7 numbers separated by commas
2. No explanation
3. No additional text
4. No labels
5. No parentheses
6. No line breaks
Exact format: Response to item 1, Response to item 2, Response to item 3, Response to item 4, Response to item 5, Response to item 6, Response to item 7
Here is the virtual persona’s profile:“

References

  1. Ajmani, L. H.; Chancellor, S.; Mehta, B.; Fiesler, C.; Zimmer, M.; De Choudhury, M. A systematic review of ethics disclosures in predictive mental health research. 2023 ACM Conference on Fairness, Accountability, and Transparency, Chicago, IL, United States, June 12–15; 2023. [Google Scholar] [CrossRef]
  2. Badawi, A.; Rahimi, E.; Laskar, M. T. R.; Grach, S.; Bertrand, L.; Danok, L.; Dhanesh, P.; Huang, J.; Rudzicz, F.; Dolatabadi, E. When can we trust LLMs in mental health? Large-scale benchmarks for reliable LLM evaluation. 19th Conference of the European Chapter of the Association for Computational Linguistics, Rabat, Morocco, March 24–29; 2026. [Google Scholar] [CrossRef]
  3. Bakhaya, S.; Lehnbom, E. C.; De Carvalho Filho, M. A.; Ma, K. Y.; Svensberg, K. Development and evaluation of AI chatbot tool for written communication training in self-care: Experiences of pharmacy students and faculty. Current Pharmacy Teaching and Learning 18 2026, 102503. [Google Scholar] [CrossRef] [PubMed]
  4. Benton, A.; Coppersmith, G.; Dredze, M. Ethical research protocols for social media health research. First ACL Workshop on Ethics in Natural Language Processing, Valencia, Spain, April 4; 2017. [Google Scholar] [CrossRef]
  5. Bossio-Botero, V.; Yadav, V.; Ouyang, J.; Abbas, A.; Worthington, M. A voice-enabled virtual patient system for interactive training in standardized clinical assessment. arXiv 2025. [Google Scholar] [CrossRef]
  6. Cabrera Lozoya, D.; Conway, M.; De Duro, S. E.; D’Alfonso, S. Leveraging large language models for simulated psychotherapy client interactions: Development and usability study of Client101. JMIR Medical Education 11 2025, e68056. [Google Scholar] [CrossRef] [PubMed]
  7. Chancellor, S.; De Choudhury, M. Methods in predictive techniques for mental health status on social media: A critical review. npj Digital Medicine 3 2020, 43. [Google Scholar] [CrossRef] [PubMed]
  8. Chen, J.; Wang, X.; Xu, R.; Yuan, S.; Zhang, Y.; Shi, W.; Xie, J.; Li, S.; Yang, R.; Zhu, T.; Chen, A.; Li, N.; Chen, L.; Hu, C.; Wu, S.; Ren, S.; Fu, Z.; Xiao, Y. From persona to personalization: A survey on role-playing language agents. Transactions on Machine Learning Research. 2024. Available online: https://openreview.net/forum?id=xrO70E8UIZ.
  9. Cook, D. A. Creating virtual patients using large language models: Scalable, global, and low cost. Medical Teacher 47 2025, 40–42. [Google Scholar] [CrossRef] [PubMed]
  10. De Choudhury, M.; De, S. Mental health discourse on Reddit: Self-disclosure, social support, and anonymity. Proceedings of the International AAAI Conference on Web and Social Media 8 2014, 71–80. [Google Scholar] [CrossRef]
  11. Ernala, S. K.; Birnbaum, M. L.; Candan, K. A.; Rizvi, A. F.; Sterling, W. A.; Kane, J. M.; De Choudhury, M. Methodological gaps in predicting mental health states from social media: Triangulating diagnostic signals. 2019 CHI Conference on Human Factors in Computing Systems, Glasgow, UK, May 4–9; 2019. [Google Scholar] [CrossRef]
  12. Errica, F.; Sanvito, D.; Siracusano, G.; Bifulco, R. What did I do wrong? Quantifying LLMs’ sensitivity and consistency to prompt engineering. 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, Albuquerque, NM, USA, April 29–May 4; 2025. [Google Scholar] [CrossRef]
  13. Fitrianie, S.; Bruijnes, M.; Abdulrahman, A.; Brinkman, W. P. The Artificial Social Agent Questionnaire (ASAQ): Development and evaluation of a validated instrument for capturing human interaction experiences with artificial social agents. International Journal of Human-Computer Studies 199 2025, 103482. [Google Scholar] [CrossRef]
  14. Fung, L.; Laing, R. A proof of concept study on the use of large language models as a client in typed role plays for training therapists. Discover Psychology 4 2024, 201. [Google Scholar] [CrossRef]
  15. Fung, T. C. J.; Chan, S. L.; Lam, C. F. M.; Lam, C. Y.; Cheng, C. C. W.; Lai, M. H.; Ho, C. C. J.; Au, S. L.; Mak, L. Y.; Hu, S.; Phetrasuwan, S.; Granger, J.; Yoon, J. M.; Malik, G.; Moreno, C. C.; Kwok, M. H. P.; Lin, C. C. Effects of generative artificial intelligence (GenAI) patient simulation on perceived clinical competency among global nursing undergraduates: A cross-over randomised controlled trial. BMC Nursing 24 2025, 934. [Google Scholar] [CrossRef] [PubMed]
  16. Gonnermann-Müller, J.; Haase, J.; Leins, N.; Kosch, T.; Pokutta, S. Maintaining stable personas? Examining temporal stability in LLM-based human simulation. 2026 CHI Conference on Human Factors in Computing Systems, Barcelona, Spain, April 13–17; 2026. [Google Scholar] [CrossRef]
  17. Hadar-Shoval, D.; Elyoseph, Z.; Lvovsky, M. The plasticity of ChatGPT’s mentalizing abilities: Personalization for personality structures. Frontiers in Psychiatry 14 2023, 1234397. [Google Scholar] [CrossRef] [PubMed]
  18. Haider, S. A.; Prabha, S.; Gomez-Cabello, C. A.; Borna, S.; Genovese, A.; Trabilsy, M.; Collaco, B. G.; Wood, N. G.; Bagaria, S.; Tao, C.; Forte, A. J. Synthetic patient–physician conversations simulated by large language models: A multi-dimensional evaluation. Sensors 25 2025, 4305. [Google Scholar] [CrossRef] [PubMed]
  19. Hanson, R. W.; Moore, T. M.; Teed, A. R.; Speer, A.; Akbari, M.; Hell, F.; Ahmed, S. S. Large language models for depression assessment: Simulating patients and clinicians in MADRS administration. PsyArXiv 2026. [Google Scholar] [CrossRef]
  20. Hu, Z.; Lian, J.; Xiao, Z.; Xiong, M.; Lei, Y.; Wang, T.; Ding, K.; Xiao, Z.; Yuan, N. J.; Xie, X. Population-aligned persona generation for LLM-based social simulation. arXiv 2025. [Google Scholar] [CrossRef]
  21. Jensen, R. A. A.; Musaeus, P.; Pedersen, K. Virtual patients in undergraduate psychiatry education: A systematic review and synthesis. Advances in Health Sciences Education 29 2024, 329–347. [Google Scholar] [CrossRef] [PubMed]
  22. Jo, Y.-W.; Lee, M.; Yang, H.-J. Large language model-based virtual patient simulations in medical and nursing education: A review. Applied Sciences 15 2025, 11917. [Google Scholar] [CrossRef]
  23. Kaabachi, B.; Despraz, J.; Meurers, T.; Otte, K.; Halilovic, M.; Kulynych, B.; Prasser, F.; Raisaro, J. L. A scoping review of privacy and utility metrics in medical synthetic data. npj Digital Medicine 8 2025, 60. [Google Scholar] [CrossRef] [PubMed]
  24. Kang, M.; Choi, G.; Jeon, H.; An, J. H.; Choi, D.; Han, J. CURE: Context- and uncertainty-aware mental disorder detection. 2024 Conference on Empirical Methods in Natural Language Processing, Miami, FL, USA, November 12–16; 2024. [Google Scholar] [CrossRef]
  25. Kononowicz, A. A.; Woodham, L. A.; Edelbring, S.; Stathakarou, N.; Davies, D.; Saxena, N.; Car, L. T.; Carlstedt-Duke, J.; Car, J.; Zary, N. Virtual patient simulations in health professions education: Systematic review and meta-analysis by the Digital Health Education Collaboration. Journal of Medical Internet Research 21 2019, e14676. [Google Scholar] [CrossRef] [PubMed]
  26. Kroenke, K.; Spitzer, R. L.; Williams, J. B. The PHQ-9: Validity of a brief depression severity measure. Journal of General Internal Medicine 16 2001, 606–613. [Google Scholar] [CrossRef] [PubMed]
  27. Kuhlmeier, F. O.; Hanschmann, L.; Rabe, M.; Luettke, S.; Brakemeier, E.-L.; Maedche, A. Designing an LLM-based behavioral activation chatbot for young people with depression: Insights from an evaluation with artificial users and clinical experts. arXiv 2025. [Google Scholar] [CrossRef]
  28. Li, B.; Jin, B.; Lan, K.; Wang, M.; Wu, M. Synthetic or authentic? Building mental patient simulators from longitudinal evidence. arXiv 2026. [Google Scholar] [CrossRef]
  29. Li, D.; Lebai Lutfi, S. Large language model-based virtual patient systems for history-taking in medical education: Comprehensive systematic review. JMIR Medical Informatics 14 2026, e79039. [Google Scholar] [CrossRef] [PubMed]
  30. Li, Y.; Zeng, C.; Zhong, J.; Zhang, R.; Zhang, M.; Zou, L. Leveraging large language model as simulated patients for clinical education. arXiv 2024. [Google Scholar] [CrossRef]
  31. Lison, P.; Pilán, I.; Sánchez, D.; Batet, M.; Øvrelid, L. Anonymisation models for text data: State of the art, challenges and future directions. 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, Online, August 1–6; 2021. [Google Scholar] [CrossRef]
  32. Liu, S.; Brie, B.; Li, W.; Biester, L.; Lee, A.; Pennebaker, J. W.; Mihalcea, R. Eeyore: Realistic depression simulation via supervised and preference optimization. PsyArXiv 2025. [Google Scholar] [CrossRef]
  33. Louie, R.; Nandi, A.; Fang, W.; Chang, C.; Brunskill, E.; Yang, D. Roleplay-doh: Enabling domain-experts to create LLM-simulated patients via eliciting and adhering to principles. 2024 Conference on Empirical Methods in Natural Language Processing, Miami, FL, USA, November 12–16; 2024. [Google Scholar] [CrossRef]
  34. Moore, J.; Grabb, D.; Agnew, W.; Klyman, K.; Chancellor, S.; Ong, D. C.; Haber, N. Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers. 2025 ACM Conference on Fairness, Accountability, and Transparency, Athens, Greece, June 23–26; 2025. [Google Scholar] [CrossRef]
  35. Pombal, J.; D’Eon, M.; Guerreiro, N. M.; Martins, P. H.; Farinhas, A.; Rei, R. MindEval: Benchmarking language models on multi-turn mental health support. arXiv 2025. [Google Scholar] [CrossRef]
  36. Posit Team. RStudio: Integrated Development Environment for R (version 2026.04.0) [Computer software; Posit Software, PBC, 2026; Available online: https://posit.co/.
  37. Pretorius, C.; Chambers, D.; Coyle, D. Young people’s online help-seeking and mental health difficulties: Systematic narrative review. Journal of Medical Internet Research 21 2019, e13873. [Google Scholar] [CrossRef] [PubMed]
  38. Qiu, H.; Lan, Z. Interactive agents: Simulating counselor–client psychological counseling via role-playing LLM-to-LLM interactions. arXiv 2024. [Google Scholar] [CrossRef]
  39. Rahimzadeh, V.; Monazzah, E. M.; Pilehvar, M. T.; Yaghoobzadeh, Y. SYNTHIA: Scalable grounded persona generation from social media data. arXiv 2025. [Google Scholar] [CrossRef]
  40. Sabour, S.; Ng, T.; Huang, M. PatientHub: A unified framework for patient simulation. arXiv 2026. [Google Scholar] [CrossRef]
  41. Sahoo, P.; Singh, A. K.; Saha, S.; Jain, V.; Mondal, S.; Chadha, A. A systematic survey of prompt engineering in large language models: Techniques and applications. arXiv 2024. [Google Scholar] [CrossRef]
  42. Salminen, J.; Santos, J. M.; Kwak, H.; An, J.; Jung, S.; Jansen, B. J. Persona Perception Scale: Development and exploratory validation of an instrument for evaluating individuals’ perceptions of personas. International Journal of Human-Computer Studies 141 2020, 102437. [Google Scholar] [CrossRef]
  43. Sankar, S.; Nafar, A.; Barman, M.; Heitz, H. K.; Kumar, A.; Tohidi, P.; Li, D.; Hussain, D.; DuBois, R.; Hasheminia, H.; Majzoubi, F. Analyzing LLM reasoning to uncover mental health stigma. arXiv 2026. [Google Scholar] [CrossRef]
  44. Schulhoff, S.; Ilie, M.; Balepur, N.; Kahadze, K.; Liu, A.; Si, C.; Li, Y.; Gupta, A.; Han, H.; Schulhoff, S.; Dulepet, P. S.; Vidyadhara, S.; Ki, D.; Agrawal, S.; Pham, C.; Kroiz, G.; Li, F.; Tao, H.; Srivastava, A.; …; Resnik, P. The Prompt Report: A systematic survey of prompt engineering techniques. arXiv 2024. [Google Scholar] [CrossRef]
  45. Settanni, M.; Quilghini, F.; Toscano, A.; Marengo, D. Assessing the accuracy and consistency of large language models in triaging social media posts for psychological distress. Psychiatry Research 351 2025, 116583. [Google Scholar] [CrossRef] [PubMed]
  46. Shen, Z.; Wang, D. Y.-B.; Mishra, S. S.; Xu, Z.; Teng, Y.; Ding, H. SLOT: Structuring the output of large language models. 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track, Suzhou, China, November 5–9; 2025. [Google Scholar] [CrossRef]
  47. Silverman, J. J.; Galanter, M.; Jackson-Triche, M.; Jacobs, D. G.; Lomax, J. W.; Riba, M. B.; Tong, L. D.; Watkins, K. E.; Fochtmann, L. J.; Rhoads, R. S.; Yager, J. The American Psychiatric Association practice guidelines for the psychiatric evaluation of adults. American Journal of Psychiatry 172 2015, 798–802. [Google Scholar] [CrossRef] [PubMed]
  48. Sit, M.; Elliott, S. A.; Wright, K. S.; Scott, S. D.; Hartling, L. Youth mental health help-seeking information needs and experiences: A thematic analysis of Reddit posts. Youth & Society 56 2024, 24–41. [Google Scholar] [CrossRef]
  49. Spitzer, R. L.; Kroenke, K.; Williams, J. B.; Löwe, B. A brief measure for assessing generalized anxiety disorder: The GAD-7. Archives of Internal Medicine 166 2006, 1092–1097. [Google Scholar] [CrossRef] [PubMed]
  50. Tosato, T.; Helbling, S.; Mantilla-Ramos, Y.-J.; Hegazy, M.; Tosato, A.; Lemay, D. J.; Rish, I.; Dumas, G. Persistent instability in LLM’s personality measurements: Effects of scale, reasoning, and conversation history. Proceedings of the AAAI Conference on Artificial Intelligence 40 2026, 37961–37969. [Google Scholar] [CrossRef]
  51. Vu, D. N. L.; Tan, R.; Moench, L.; Francke, S. J.; Woiwod, D.; Thomas-Odenthal, F.; Stroth, S.; Kircher, T.; Hermann, C.; Dannlowski, U.; Jamalabadi, H.; Ji, S. Roleplaying with structure: Synthetic therapist–client conversation generation from questionnaires. arXiv 2025. [Google Scholar] [CrossRef]
  52. Wang, B.; Sun, Y.; Wang, J.; Yang, H.; Fu, X.; Zhao, Y.; Wei, S.; Wang, S.; Qin, B. CARE-Bench: A benchmark of diverse client simulations guided by expert principles for evaluating LLMs in psychological counseling. arXiv 2025. [Google Scholar] [CrossRef]
  53. Wang, R.; Milani, S.; Chiu, J. C.; Zhi, J.; Eack, S. M.; Labrum, T.; Murphy, S. M.; Jones, N.; Hardy, K.; Shen, H.; Fang, F.; Chen, Z. Z. PATIENT-Ψ: Using large language models to simulate patients for training mental health professionals. 2024 Conference on Empirical Methods in Natural Language Processing, Miami, FL, USA, November 12–16; 2024. [Google Scholar] [CrossRef]
  54. Wang, X.; Wei, J.; Schuurmans, D.; Le, Q.; Chi, E.; Narang, S.; Chowdhery, A.; Zhou, D. Self-consistency improves chain-of-thought reasoning in language models. arXiv 2022. [Google Scholar] [CrossRef]
  55. Warner, A.; LeDue, J.; Cao, Y.; Tham, J.; Murphy, T. H. Synthetic patient and interview transcript creator: An essential tool for LLMs in mental health. Frontiers in Digital Health 7 2025, 1625444. [Google Scholar] [CrossRef] [PubMed]
  56. Xiao, Y.; Zhang, V. J.; Yang, C.; Ma, N.; Xuan, W.; Huang, J.-T. The chameleon’s limit: Investigating persona collapse and homogenization in large language models. arXiv 2026. [Google Scholar] [CrossRef]
  57. Xu, X.; Yao, B.; Dong, Y.; Gabriel, S.; Yu, H.; Hendler, J.; Ghassemi, M.; Dey, A. K.; Wang, D. Mental-LLM: Leveraging large language models for mental health prediction via online text data. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8 2024, 1–32. [Google Scholar] [CrossRef] [PubMed]
  58. Yang, K.; Zhang, T.; Kuang, Z.; Xie, Q.; Huang, J.; Ananiadou, S. MentaLLaMA: Interpretable mental health analysis on social media with large language models. The ACM Web Conference 2024, Singapore, Singapore, May 13–17; 2024. [Google Scholar] [CrossRef]
  59. Yeo, Y. H.; Peng, Y.; Mehra, M.; Samaan, J.; Hakimian, J.; Clark, A.; Suchak, K.; Krut, Z.; Andersson, T.; Persky, S.; Liran, O.; Spiegel, B. Evaluating for evidence of sociodemographic bias in conversational AI for mental health support. Cyberpsychology, Behavior, and Social Networking 28 2025, 44–51. [Google Scholar] [CrossRef] [PubMed]
Figure 2. Comparison of the information included in the Base and Clinically Enriched prompt variants. Both variants were derived from the same shadow persona, but differed in clinical granularity: the Base prompt retained a narrative-psychosocial representation of the user history, whereas the Clinically Enriched prompt added more explicit clinical organization, symptom-domain information, psychopathological history, risk/protective factors, and analysis metadata.
Figure 2. Comparison of the information included in the Base and Clinically Enriched prompt variants. Both variants were derived from the same shadow persona, but differed in clinical granularity: the Base prompt retained a narrative-psychosocial representation of the user history, whereas the Clinically Enriched prompt added more explicit clinical organization, symptom-domain information, psychopathological history, risk/protective factors, and analysis metadata.
Preprints 224688 g002
Figure 3. Z-score profiles for PHQ-9 human and LLM evaluations across anonymized personas for the Base (panel A) and Clinically-enriched (panel B) prompts.
Figure 3. Z-score profiles for PHQ-9 human and LLM evaluations across anonymized personas for the Base (panel A) and Clinically-enriched (panel B) prompts.
Preprints 224688 g003
Figure 4. Z-score profiles for GAD-7 human and LLM evaluations across anonymized personas for the Base (panel A) and clinically-enriched (panel B) prompts.
Figure 4. Z-score profiles for GAD-7 human and LLM evaluations across anonymized personas for the Base (panel A) and clinically-enriched (panel B) prompts.
Preprints 224688 g004
Table 1. Descriptive statistics for LLM-generated total scores.
Table 1. Descriptive statistics for LLM-generated total scores.
Prompt Model Scale n Mean SD Min Max
Clinically Enriched GPT-4o PHQ-9 600 17.21 3.57 10 24
Clinically Enriched DeepSeek V4 Pro PHQ-9 600 15.16 5.67 5 26
Clinically Enriched GPT-4o GAD-7 600 16.29 3.40 11 21
Clinically Enriched DeepSeek V4 Pro GAD-7 600 18.95 1.55 12 21
Base GPT-4o PHQ-9 600 16.75 3.52 8 24
Base DeepSeek V4 Pro PHQ-9 600 15.80 5.74 3 27
Base GPT-4o GAD-7 600 17.27 2.99 11 21
Base DeepSeek V4 Pro GAD-7 600 17.95 2.22 8 21
Table 2. Total-score response stability.
Table 2. Total-score response stability.
Prompt Model Scale ICC(2,1) ICC(2,100)
Clinically Enriched GPT-4o PHQ-9 .887 .999
Clinically Enriched DeepSeek V4 Pro PHQ-9 .880 .999
Clinically Enriched GPT-4o GAD-7 .912 .999
Clinically Enriched DeepSeek V4 Pro GAD-7 .409 .986
Base GPT-4o PHQ-9 .884 .999
Base DeepSeek V4 Pro PHQ-9 .847 .998
Base GPT-4o GAD-7 .777 .997
Base DeepSeek V4 Pro GAD-7 .560 .992
Table 3. Summary of item-level ICC(2,1) estimates.
Table 3. Summary of item-level ICC(2,1) estimates.
Prompt Model Scale Mean ICC Median ICC Range
Clinically Enriched GPT-4o PHQ-9 .612 .609 .227-.870
Clinically Enriched DeepSeek V4 Pro PHQ-9 .712 .753 .351-.816
Clinically Enriched GPT-4o GAD-7 .743 .742 .571-.919
Clinically Enriched DeepSeek V4 Pro GAD-7 .284 .257 .091-.677
Base GPT-4o PHQ-9 .585 .606 .275-.935
Base DeepSeek V4 Pro PHQ-9 .664 .712 .163-.917
Base GPT-4o GAD-7 .596 .611 .389-.721
Base DeepSeek V4 Pro GAD-7 .460 .454 .170-.772
Table 4. Persona-level correlations between human symptom totals and LLM-generated total scores.
Table 4. Persona-level correlations between human symptom totals and LLM-generated total scores.
Prompt Model Scale Pearson r p Spearman rho p
Clinically Enriched GPT-4o PHQ-9 .776 .069 .771 .072
Clinically Enriched DeepSeek V4 Pro PHQ-9 .914 .011 .886 .019
Clinically Enriched GPT-4o GAD-7 .609 .199 .647 .165
Clinically Enriched DeepSeek V4 Pro GAD-7 .824 .044 .853 .031
Base GPT-4o PHQ-9 .941 .005 .928 .008
Base DeepSeek V4 Pro PHQ-9 .989 < .001 .986 < .001
Base GPT-4o GAD-7 .469 .349 .577 .231
Base DeepSeek V4 Pro GAD-7 .786 .064 .698 .123
Table 5. Item-level correspondence between human symptom presence and LLM-generated item scores.
Table 5. Item-level correspondence between human symptom presence and LLM-generated item scores.
Prompt Model Scale Pearson r p Spearman rho p AUC
Clinically Enriched GPT-4o PHQ-9 .399 .003 .364 .007 .710
Clinically Enriched DeepSeek V4 Pro PHQ-9 .728 < .001 .740 < .001 .928
Clinically Enriched GPT-4o GAD-7 .538 < .001 .509 < .001 .806
Clinically Enriched DeepSeek V4 Pro GAD-7 .726 < .001 .785 < .001 .957
Base GPT-4o PHQ-9 .526 < .001 .562 < .001 .833
Base DeepSeek V4 Pro PHQ-9 .741 < .001 .745 < .001 .941
Base GPT-4o GAD-7 .518 < .001 .494 < .001 .793
Base DeepSeek V4 Pro GAD-7 .736 < .001 .667 < .001 .894
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.