Submitted:
04 September 2026
Posted:
07 September 2026
You are already at the latest version
Abstract
Background/Objectives: The burden of Major Depressive Disorder (MDD) calls for assessing the needs and design of innovative add-on interventions. Generative artifi-cial-intelligence (AI) tools are increasingly used in everyday life. Evidence remains limited on how needs assessment and co-participatory design can inform AI chatbot trials for MDD. Methods: This study reports a mixed-methods design. Phase 1 comprised a needs assessment of 511 psychiatric outpatient attendees in Hong Kong. Phase 2 involved a participatory co-design and test-run experiences with 10 lived-experience experts with MDD and 8 multidisciplinary clinical professionals. This process developed the CARES-MDD, a generative AI-chatbot grounded in evidence-based Cogni-tive-Behavioral-Therapy content as an adjunct to standard care. Results: In Phase 1, 66.2% reported using generic AI chatbots for emotional concerns; privacy, anticipated negative evaluation, and concern about burdening others were recurring themes in the qualitative data. 24/7 accessibility, non-judgmental, and localized empathy were the main themes of a desired mental health AI chatbot. In Phase 2, CARES-MDD was co-designed and as-sessed against a 7-domain consensus rubric measuring perceived usefulness and ac-ceptability. Overall, CARES-MDD received favorable expert rubric ratings. Content analysis identified positive and negative experiences which established parameters for future trial methodology. Conclusions: This study identifies needs and yields the co-design of the CARE-MDD purposed as adjunct support for MDD in psychiatric care in Hong Kong, providing the parameters to inform future controlled trials.
Keywords:
generative artificial intelligence
; digital mental health
; young people
; mental health
; depression
; help-seeking
; mental health stigma
; cognitive behavioral therapy
; participatory co-design
1. Introduction
Major Depressive Disorder (MDD) is a leading cause of disability worldwide [1], accounting for substantial disease burden, diminished functional capacity, and elevated suicide risk [2,3]. While empirically validated psychological interventions, most notably Cognitive Behavioral Therapy (CBT),are recommended as first-line treatments, fewer than half of affected individuals receive adequate care [4,5].
Digital mental-health interventions (DHIs) may increase access to information and self-management support; however, sustained engagement, cultural relevance, and integration with routine psychiatric care remain important challenges [6,7]. Critically, the overwhelming majority of mental health DHIs are developed without participatory input from clinical populations, producing tools that feel generic, robotic, and emotionally misaligned for individuals contending with severe depressive symptoms [8,9].
The emergence of Generative Artificial Intelligence(AI) and Large Language Models (LLMs) offers an opportunity to bridge this gap. Unlike rigid, rule-based algorithms, L LLMs can generate responsive language that may be perceived as empathic or validating; however, such outputs are not equivalent to human understanding or clinician-delivered therapy and require empirical safety and efficacy evaluation [10,11,12]. When integrated with evidence-based psychotherapeutic frameworks—specifically CBT, which targets transdiagnostic cognitive and behavioral maintaining mechanisms (e.g., cognitive distortions, avoidance, and rumination; [13,14])—generative conversational agents can deliver flexible, tailored micro-interventions in real time.
However, translating generative AI into a psychiatric adjunct treatment tool requires tuning, localized adaptation, and safety governance .This formative mixed-methods study aimed to identify user-defined requirements, stakeholder-identified failure modes, and governance conditions for the development of Conversational Artificial-intelligence for REducing Symptoms and REmission for Major Depressive Disorder (CARES-MDD), which is a Hong Kong Cantonese AI-supported conversational agent for persons living with depression. The study was designed to inform a pre-specified, staged validation programme for CARES-MDD, including content curation, technical evaluation, safety assurance, and prospective feasibility testing before efficacy trial. Specifically, the study sought to: (1) characterise digital behaviour, generic AI-chatbot use, barriers to emotional disclosure, and desired chatbot features among people attending psychiatric outpatient services; (2) translate these findings, together with feedback from lived-experience experts and clinical professionals, into prioritised AI chatbot design recommendations; and (3) describe stakeholder perceptions of perceived usefulness, acceptability using consensus rubrics and qualitative interviews to collect their positive and negative points when testing CARES-MDD.
2. Materials and Methods
2.1. Study Design
This study employed a mixed-methods, participatory co-design framework, drawing upon established guidelines for DHIs [15,16,17]. Recognizing that the target demographic has important experience and opinions of generative-AI, the investigation was structured into two interrelated phases:
Phase 1 (Needs Assessment): A cross-sectional quantitative and qualitative survey investigating baseline digital habits, help-seeking attitudes, self-stigma, and conversational AI acceptability/ needs among a clinical cohort. Phase 2 (Participatory Co-Design and Test-run Evaluation): A stakeholder co-design/ feedback phase involved adult Lived Experience Experts (LEEs) advisers and multidisciplinary clinical professionals. Phase 1 findings were used to identify candidate design priorities, including privacy and discretion, use pattern, low-burden access and barriers reported. Phase 2 included a workshop of LEEs to integrate the prioritized implementation ideas to co-design the generative AI-chatbot (CARES-MDD), then invited them for test-run CARES-MDD to evaluate its acceptability, perceived usefulness and perceived safety in mixed-methods evaluation, for its potential for progression to a randomized-controlled trial (RCT) design.
2.2. Participants and Recruitment Settings
Phase 1 Cohort. A clinical cohort of 511 outpatient psychiatric attendees was recruited from Child and Adolescent Psychiatric Outpatient Clinics at Queen Mary Hospital Service attending between July 22 and August 21, 2025. Eligible participants possessed a personal smartphone with daily internet access, and demonstrated the ability to read written Chinese and communicate in Cantonese.
Phase 2 Reference Group. The Phase 2 participatory co-design involved the interview of 18 experts, comprising of 10 LEEs, alongside a multidisciplinary expert group of 8 clinical professionals (2 clinical psychologists, three counsellors, 1 psychiatrist, and 2 psychiatric nurses). LEEs had a confirmed clinical diagnosis of Major Depressive Disorder (MDD) based on the Diagnostic and Statistical Manual of Mental Disorders, Fifth Edition (DSM-V) [18], ascertained by case psychiatrists. The LEEs were all aged 18 years or above, restricting the reference group to legal adults was intended to ensure that participants had the developmental maturity to discuss lived experiences of depression, emotional disclosure, and digital mental health support in an interview setting. Exclusion criteria included acute suicidal risk, active psychosis, or concurrent substance use disorders. While the LEEs guided the humanistic and user-experience elements, the clinical experts provided feedback of the content, and appropriateness of the interaction. Furthermore, the expert panel provided opinions on the safety features in practical use.
2.3. Phase 1: Needs Assessment Measures and Data Collection
An electronic questionnaire was developed to capture quantitative metrics and qualitative insights across four primary domains:
Online behavior and AI chatbot use: Their daily digital screen time, generic conversational AI adoption rates, daily online interaction duration, and digital engagement frequency were asked. Depressive symptom severity was assessed using the validated Chinese version of the 9-item Patient Health Questionnaire (PHQ-9) as clinical variable [19].
Help-Seeking and Self-Stigma: Standardized psychometric instruments measured help-seeking behaviors (Help-Seeking Scale, max score 28) and internalized mental health stigma (Self-Stigma Scale, max score 24, where higher scores denote lower stigma), from the Knowledge and Attitudes to Mental Health Scales (KAMHS) [20] which was recently validated in Hong Kong [21].
Barriers to emotional disclosure and help-seeking within existing psychiatric care networks: Participants also completed an open-ended item asking, “What are the reasons that prevent you from sharing emotional difficulties with others?”
Desired AI Chatbot Specifications: To guide the system architecture and feature prioritization, a final open-ended item asked, “What is the most desired feature of an AI chatbot within the existing psychiatric system?”
2.4. Phase 2: Co-Design Workshop, Beta-Runs, and Mixed-Methods Evaluation
To translate the needs assessment findings into actionable AI chatbot development, a participatory co-design workshop was conducted with Lived Experience Experts (LEEs). The engagement protocol was grounded in established Youth Participatory Action Research (YPAR) principles [22] and guidelines for co-designing digital health interventions with young people [23], which used non-medicalized language, culturally relevant humor, and open body language (e.g., maintaining eye-level seating) to actively cultivate trust and rapport. Active listening techniques, such as reflective summarizing and open-ended prompting, were utilized to verify shared understandings, clarify conceptual misinterpretations, and empower LEEs to expand upon their ideas on the prioritization of the AI chatbot theme. This collaborative process directly informed the prioritized implementation strategies for the Conversational Artificial-intelligence for REducing Symptoms and REmission for Major Depressive Disorder (CARES-MDD) intervention (detailed in Appendix A2), bridging the gap between clinical utility and user-centered design.
Following the initial co-design workshops, the expert group was invited for multiple test-run cycles conducted over a development window. The expert reference group attended several supervised, in-person test-runs throughout this period to report any good and bad points about the CARES-MDD. At the conclusion of this iterative cycle, participants completed an evaluative rubric encompassing seven core therapeutic domains, which were generated directly from the prioritized co-design strategies (Appendix A2). Participants rated these domains on a 10-point visual analog scale (1 = completely inadequate, 10 = exceptional), as shown in Table 1.
2.4.1. Description of the CARES-MDD Chatbot
The CARES-MDD application is a text-based, multi-threaded chatbot application developed using a large language model and designed to communicate in natural Hong Kong-style Cantonese. Its conversational content was developed by psychiatrists and clinical psychologists and reviewed by members of the clinical team with reference to evidence-informed, primarily CBT principles. To address clinical safety, CARES-MDD has a keyword-based crisis-signposting module screening for high-risk crisis triggers (e.g., ‘suicidal’, ‘death’, ‘giving up’). Detection of these triggers initiates an automated safety interface,’ temporarily suspends the generation of routine chatbot content, and displaying a static emergency interface with 24/7 direct-dial crisis hotlines. The research team also conducted periodic clinical review of chatbot content and system responses to identify potential safety concerns on a weekly basis. CARES-MDD was designed to support compliance with applicable data-protection requirements. Appropriate measures were implemented to protect participant information and restrict access to study data to authorized members of the research team. It was not intended to diagnose MDD, provide emergency care, replace professional assessment, or autonomously manage suicidal risk. The CARES-MDD system identity, intended use, and technical configuration, together with its safety-screening and crisis-response pathway, are summarised in Appendix Table A1 and Table A3.
2.5. Analysis
We used descriptive statistics and inferential tests to examine closed-ended questions regarding participants’ digital behaviors, help-seeking patterns, and the evaluation rubrics. Quantitative analyses were conducted using Excel and R statistical software (version 4.3.0; R Foundation for Statistical Computing, Vienna, Austria). Between-group comparisons of continuous outcomes Welch’s unequal variances t-tests according to distributional and variance considerations. Continuous variables, including digital engagement durations, psychometric clinical baselines (e.g., PHQ-9) and usability metrics (SUS, WAI-SR), were summarized as means and standard deviations (SD). Categorical variables, such as generic AI adoption rates and interaction frequencies, were reported as frequencies and percentages.
For qualitative data derived from phase 1, short-text responses were coded to categorize non-mutually exclusive recurring themes (e.g., privacy concerns, fear of judgment). To provide cultural context, frequently occurring Chinese terms within these responses were mapped and visualized using word clouds. The qualitative feedback of phase 2 as conducted in an inductive, data-driven iterative process [24]. Due to limitations within the data set (short phrases or sentence fragments), full thematic analysis was not appropriate [25]. Transcripts were read by M.S.Y.Y. and K.Y.H.L. to familiarize themselves with the responses. Next, two coders (M.S.Y.Y. and K.Y.H.L.) undertook the content analysis, generating codes directly from the data rather than from a pre-existing coding framework. Brief descriptive labels (“codes”) were applied to transcript segments, and multiple codes were applied if participants’ narratives presented multiple meanings. Following this, M.S.Y.Y. and K.Y.H.L. met to discuss these coding decisions—specifically conducting inter-rater reliability checks on approximately 30% of the transcripts—and revisions and refinements of the codes were undertaken via negotiated agreement. Following this process, first-order codes (“categories”) were grouped bottom-up into second-order themes based on commonality of meaning. M.S.Y.Y. and K.Y.H.L. then convened to review and refine the final themes.
2.6. Ethical Approvals and Consent Procedures
The assessment received institutional review board approval from the Institutional Review Board of The University of Hong Kong/Hospital Authority Hong Kong West Cluster (UW25-299). Informed consent was obtained from adult participants and their parents (for minors) prior to engagement. Participant confidentiality was secured by decoupling contact information from survey data and assigning de-identified alphanumeric study IDs.
3. Results
3.1. Phase 1 Results: Clinical Needs Assesssment for Barriers and AI Chatbot Usse
The cohort comprised 511 young people (mean age 14.97 ± 1.86 years; 54.0% female, 46.0% male). Analysis of online behavior revealed a significant divergence based on psychological distress levels. Participants recorded a significantly high digital footprint, spending an average of 296.3 minutes per day online. Nearly 2 out of 3 young people (66.2%) are already familiar with and navigating generic AI chatbot, implying the clinical population is already primed for conversational agents. While adoption is high, frequency data highlights a critical gap: 56.4% of active users have only had sporadic, surface-level interactions (fewer than 10 conversations total). This suggests that while young people are curious about AI chatbot, current generic AI chatbot tools are not providing the sustained, meaningful engagement required for psychological intervention for the clinical demographic (Table 2).
Of the 511 participants, 65 (12.7%) provided additional comments in response to the question, “What are the reasons that prevent you from sharing emotional difficulties with others?” Comments were generally brief, comprising a phrase or one to two sentences (total word count = 834). Iterative thematic analysis identified Privacy concerns (n = 22), characterising emotional disclosure as “privacy” or a “personal matter” and expressing concern that their concerns can be gossiped to others; Fear of judgement (n = 16) encompassed anticipated misunderstanding, criticism, stigma, or negative evaluation, as reflected in comments such as “other people’s views and prejudices” and “afraid of being misunderstood”); and Avoidance of burdening others (n = 8) captured concerns that disclosure might negatively affect or inconvenience others, illustrated by “do not want to affect other people’s mood” and “do not want to disturb others”. These are visualised in word cloud in Figure 1.
42 (8.2%) provided additional open-ended comments in response to the question, “What is the most desired feature of an AI chatbot within the existing psychiatric system?” Comments were generally brief, comprising short phrases or single sentences (total word count = 586). Iterative thematic analysis identified four non-mutually exclusive desired features. The most common theme was 24/7 immediate accessibility (), with participants valuing on-demand emotional support between formal clinical appointments (e.g., “always available” and “immediate reply when I cannot sleep”). Privacy and non-judgmental space () reflected the preference to share vulnerable emotions without fear of human evaluation, characterized by comments like “no human judgment” and “completely confidential”. Localized empathy () captured the desire for naturalistic, Hong Kong-style Cantonese interactions that feel culturally resonant rather than robotic (e.g., “sounds like a local friend”). Finally, practical psychological guidance () referred to the need for actionable coping strategies, exemplified by “helps me reframe my negative thoughts”. These are visualised in word cloud in Figure 2.
3.2. Phase 2 Results: Co-participatory Design and Expert Evaluation
Following the Phase 1 needs assessment, 10 LEEs (mean age 22.9 ± 2.9 years; 50% female, 50% male) and 8 clinical experts were invited for the initial workshop for co-design (appendix A2). They were then invited to test run for the acceptability, perceived safety, and feasibility of progression to definitive trial of CARES-MDD. Only the LEEs had the pre- and post-observations adapting the outcome assessors. Quantitative rubric scores across the seven evaluative domains are summarised in Figure 3.
Participants provided generally favourable ratings across all CARES-MDD evaluation domains. Empathy and HK-Cantonese understanding, Accuracy and useful CBT info, and Balance of privacy and clinical governance achieved the highest median scores, reflecting strong acceptability of the chatbot’s culturally resonant communication, perceived usefulness and acceptability, and value clinician oversight. Acceptability of RCT methodology also demonstrated robust feasibility. Conversely, Safety features and adverse digital events alongside Active listening and Socratic questions showed wide interquartile ranges, indicating polarized experiences; some valued these features, while others found automated overrides abrupt and socratic questioning could be demanding during distress. Finally, Supportive pacing and intervention duration yielded the lowest median score, capturing challenges foreseeable if it is positioned as an open-ended, indefinite intervention. The descriptive observations of the CARES-MDD test-run were summaraised in Table A4. Additionally, Table 3 thematically organizes the qualitative feedback underlying these quantitative evaluations.
Among positive experiences, LEEs valued CARES-MDD as a private “tree hole” for emotional disclosure and perceived it as more private than commercial applications, because it was not perceived to use disclosed content for targeted advertising which would pass the information to vendors. The use of natural Hong Kong-style Cantonese was viewed as familiar, emotionally accessible, and appropriate to local expressions. CBT-informed content was generally regarded as clinically appropriate and distinguishable from generic emotional support, particularly when interactions progressed beyond brief exchanges (approximately five conversational turns). To ensure privacy, participants advocated for using pseudonyms to prevent data misattribution. However, they simultaneously supported backend clinician oversight, noting that human supervision of their conversations provided necessary reassurance against potential AI unreliability. Rather than limitations, their sharing of negative experience establish parameters for future trial methodology. Specifically, the cognitive demands of a text-based modality indicate that the intervention is optimally targeted toward a specific demographic: namely, younger, digitally native populations with a higher baseline tolerance for typing. Furthermore, the eventual onset of digital fatigue suggests that CARES-MDD is most viable as a time-limited intervention instead of an open-ended, indefinite intervention. Technologically, the LMM occasionally exhibited inconsistent longitudinal memory, failing to seamlessly recall contextual details from very early sessions, a stark contrast to human psychotherapist. Automated safety overrides were sometimes perceived as abrupt, particularly when triggered by non-suicidal colloquial exclamations such as “死啦” (“oh no”). Some participants expressed residual concern about possible AI hallucinations, despite not reporting such experiences during the study.
4. Discussion
This study reported a mixed-methods study: first, to characterize the online digital behavior, usage of generic AI chatbots for emotional concerns, its barriers, and needs assessment of a desired AI mental health chatbot specifications; second, to co-design and evaluate the iterative test-run of the CARES-MDD among experts living with MDD. By bridging a clinical needs assessment data (Phase 1), we identified (1) barriers of self- stigma and fear of human judgement for sharing emotional concerns in add-on support within existing psychiatric care framework; (2) high technological readiness with generic AI chatbot for emotional concerns ; and (3) the desired features for development of an AI mental health chatbot for this specific clinical demographic. Phase 2 reported an iterative expert co-design process of CARES-MDD involving LEEs and a clinical expert reference group. They prioritized and shared positive experience regarding CARES-MDD’s features of local style communication, CBT-curated RAG instead of generic content, and the balance of privacy and safety monitoring. Rather than limitations, their sharing of negative experience regarding LMM-related risks and intervention pacing highlight key areas requiring reference parameters for future trial methodology.
The Phase 1 data shows that conventional, human-delivered support systems are often hindered by the target demographic’s high self-stigma and fear of judgment. Even among this vulnerable cohort already attending psychiatric outpatient clinics, qualitative themes highlighted concerns about sharing emotions for support regarding anticipated negative evaluation, and a reluctance to burden others [26]. These barriers may reduce willingness to seek help even when clinical need is present, particularly when the clinical population fear that their emotional disclosures will be misunderstood, shared beyond their control, or impose on others.[27,28]. An AI chatbot has the potential as add-on micro-intervention to foreground immediate accessibility without burdening people around, employing a conversational register that approximates support interaction rather than formal assessments.
It was further supported by the findings that most outpatient attendees had high technological readiness, with widespread surface acquaintance with AI chatbots. Users who have interacted briefly with commercial chatbots may hold implicit expectations derived from those platforms (rapid factual responses, task completion, minimal affect) that are poorly matched to the longitudinal, reflective, emotionally containing functions that a therapeutic chatbot must serve [24]. The desired features identified by participants were closely aligned with these barriers. Participants valued immediate availability, a private and non-judgemental conversational space, natural local style communication, and practical psychological guidance. These findings support the need and the essential features of developing an adjunctive AI chatbot (CARES-MDD) that would not shara data to vendors, accessible between appointments, and culturally congruent. The co-design process similarly underscored participants emphasised the importance of conversational tone, appropriate emotional validation, and a style that feels natural rather than formal or clinical.
The Phase 2 evaluation indicated generally favourable perceptions across the seven consensed CARES-MDD domains. Empathy and local style understanding, the accuracy and usefulness of CBT-informed content, and balance of privacy and safety received the highest ratings. Participants valued CARES-MDD as a private “tree hole” for emotional disclosure and perceived it as distinct from commercial applications because it was not seen as using personal content for advertising or commercial targeting. The preference for unique identifiers or pseudonyms further illustrates the importance of maintaining confidence that individual disclosures will not be confused with those of another user. Participants appreciated that CARES-MDD responded in colloquial Cantonese rather than formal written Chinese, and several noted that this linguistic calibration created a sense of relational proximity uncommon in health-technology products [29,30].
Participants did not frame privacy as the absence of clinical oversight. Rather, they supported proportionate human oversight when data governance arrangements were transparent and sensitive personal details were protected. This finding is important for the design of clinically oriented conversational AI: trust may depend on clearly communicating what is monitored, who can access information, what will be documented, and the circumstances under which confidentiality may be overridden because of safety concerns.
CBT-informed content was perceived as more useful than generic reassurance when interactions progressed beyond short exchanges. Participants and experts considered the chatbot more helpful when it could validate distress, clarify the immediate concern, and only subsequently introduce structured reflection or cognitive restructuring. This supports a staged conversational design: initial empathic listening and emotional validation should precede problem-solving or Socratic questioning. The interactions were perceived as more useful after several active exchanges, particularly when a session involved at least five conversational turns. This observation addresses a methodological challenge in DHI research: operationalizing intervention “dosage” outside of traditional psychiatric frameworks. Unlike conventional psychotherapy, which relies on structured 50-minute appointments, unguided digital tools face tremendous variability in user engagement duration. Because meta-analyses indicate that an effective clinical dosage for digital interventions ranges from 30 to 60 minutes per week [31], future RCTs can formally define an “active digital session” as a continuous interaction lasting approximately 10 minutes which estimates to comprise approximately minimum five reciprocal turns.
Several LEEs also reported that abrupt safety override responses, while appropriate from a risk-management standpoint, interrupted the conversational flow at moments when continued empathic engagement may have been more therapeutically productive. Clinical professionals specifically noted that while algorithmic safety gates are non-negotiable from a duty-of-care standpoint, their current implementation in generative AI systems typically operates as a binary interrupt — abruptly redirecting conversation to crisis resources — rather than as a graduated, therapeutically calibrated transition. In human clinical practice, the decision to escalate, contain, or stay present with distress is a moment-to-moment clinical judgement informed by relationship history, non-verbal cues, and dynamic risk appraisal.
A central question explored during the co-design interviews was how clinical users conceptualize the identity of CARES-MDD. Instead, users perceived CARES-MDD prototype as a portable, on-demand conversational support and skills-coaching tool during unassisted hours (e.g., late nights or between outpatient visits). Multiple LEEs affirmed their willingness to recommend CARES-MDD to peers who ‘aren’t ready for traditional therapy yet.’ By allowing distressed individuals to practice cognitive restructuring and emotional articulation in an unobserved, safe environment, the conversational agent serves as an accessible entry point that may lower psychological barriers and build readiness for future formal human psychiatric care.
The integration of generative AI into psychiatric populations raises broader ethical and clinical considerations for the field, particularly regarding unhealthy dependence, hallucination risks, and longitudinal memory constraints [32,33,34]. Their qualitative feedback highlighted a known vulnerability inherent to current LLMs [35]. LLM models unable to dynamically contextualise the psychotherapeutic clinical formulation of the engaging client unlike human therapist, they may be better served as a complementary tool for “augmenting” pre-existing therapist alliance with patients, rather than serving as the therapist alliance in standalone role [36]. In the era of entering into the AI, we should be more cautious in the role of AI chatbot in monitoring its possible safety issues and adverse digital events. Although the LEEs did not experience so, they mentioned the risk of AI hallucinations and psychosis that they overheard. The LLM risks identified in the process are summarized in the subsequent Table 4, aligning with the guideline for an AI chatbot architecture ready for future trials [37].
Participants’ comments regarding long-term engagement should be interpreted as design considerations for future evaluation rather than as evidence of an optimal intervention duration. Participants anticipated digital fatigue if CARES-MDD were available as an open-ended or indefinite intervention. In light of attrition commonly reported in unguided digital mental-health interventions, a time-limited, on-demand format may offer a pragmatic starting point for prospective feasibility testing. An approximately 8-week intervention window could allow participants sufficient opportunity to engage with CBT-informed self-management content while limiting prolonged exposure and the potential for overreliance on the conversational agent [40,41]. To mitigate digital fatigue issue, further RCT trial could integrate the ecological momentary assessment technique which has shown evidence on boosting adherence through self-monitoring and reminder effects [42,43].
Participants generally considered randomisation, including allocation to a psychoeducation-only active digital control, to be fair and acceptable. Future trial design should minimise participant burden, particularly for providing appropriate reimbursement, online modes for assessments; and include follow-up sufficient to examine engagement, acceptability, safety-related events, and longer-term outcomes. While this formative study provided evidence on the architectural feasibility, cultural acceptability, and safety governance boundaries of CARES-MDD, it was not designed to evaluate randomized clinical efficacy or comparative attrition. Therefore, rather than demonstrating clinical superiority, these findings provide the critical structural parameters, such as the minimum 5-turn intervention dosage, the system and clinical measures to counteract the potential LLM risks e.g., pre-consented clinical escalation pathways, required to safely initiate a definitive RCT.
4.1. Strengths, Limitations, and Future Clinical Directions
A strength of this investigation is its mixed-methods methodology, uniting a relatively large clinical sample (N = 511) with more in-depth mixed-method evaluation of LEEs throughout the development lifecycle, and CARES-MDD shows acceptable ecological and cultural validity. However, several limitations must be acknowledged. First, the pre-to-post observations from the uncontrolled beta-testing phase are strictly descriptive; they could not deduce therapeutic efficacy and further clinical trial is needed. Second, it was a convenience sample drawn from a single professional and community network, which limits the generalizability. Nevertheless, given the uncompensated time commitments required for these phases of co-development, leveraging a network with pre-established clinical rapport was essential to maintain experts and accurately capture user-driven needs within existing resource constraints. Third, CARES-MDD is fundamentally without highly experimental technological features. This design choice was driven by the co-design participants, who prioritized clinical reliability and evidence-based safety over technological novelty, which might lead to “blackbox” concerns. While this limits the platform’s technological boundary-pushing, it ensures the intervention maintains the rigorous clinical safety required for vulnerable psychiatric populations.
Future work in developments of generative-AI Chatbots for mental health purpose must prioritize three interrelated agendas to address the universal challenges. Technically, the field should move beyond on aligning global clinical guidelines and standardizing protocols for LLM applications across the disciplines [44,45]. Clinically, the next appropriate evaluative step involves prospective designs, such as full-scale RCTs, to systematically examine the clinical effectiveness. Digital adverse events should be continually assessed in future trials. Conceptually, the pronounced help-seeking barriers and high self-stigma identified in Phase 1 motivate a critical agenda for comparative research: investigating whether initial AI-mediated emotional disclosure can serve as a transitional bridge within a stepped-care pathway, lowering psychological barriers and facilitating subsequent engagement with human-delivered psychiatric care [46]. Addressing these universal challenges will be paramount for integrating generative AI as a safe, regulated, and continuously audited adjunct to global mental health services [12].
5. Conclusions
In conclusion, the needs assessment demonstrates that while the use of generic AI chatbots for emotional support is already prevalent among psychiatric outpatients in Hong Kong. Barriers—namely fear of judgment and burdening others—continue to impede traditional help-seeking behaviors within traditional psychiatric care system. This underscores the necessity that AI chatbots with specified features are needed for individuals living with MDD, even among those already engaged in psychiatric care. By utilizing a needs-driven participatory framework with LEEs, a CBT-grounded generative AI chatbot for MDD, has been co-developed, with mixed-methods findings supporting it is feasible to use end-users and clinical experts. The findings provide preliminary, clinically grounded design parameters and feasibility information for subsequent evaluation. A future trial should compare CARES-MDD plus usual care with an appropriate active digital control plus usual care, incorporate predefined safety and adverse-event monitoring, and evaluate engagement, long-term safety, and clinical outcomes. This mixed—methods study does not establish the efficacy, safety, or therapeutic equivalence of CARES-MDD; these questions require subsequent prospective controlled trials.
Author Contributions
Conceptualization, K.Y.H.L. , W.F.Y. and K.F.C.; methodology, P.H.F.N.; software, K.Y.H.L. and P.K.S.S.; validation, K.Y.H.L., R.W.K.W. and K.F.C.; formal analysis, K.Y.H.L., M.S.Y.Y. , M.C.T.K. , Y.Y.H.C. and P.K.S.S.; investigation, K.Y.H.L., M.S.Y.Y. and P.H.F.N.; resources, K.Y.H.L..; data curation, K.Y.H.L., M.S.Y.Y. and P.K.S.S.; writing—original draft preparation, K.Y.H.L. , W.F.Y., P.H.F.N , M.C.T.K. , Y.Y.H.C. and M.S.Y.Y.; writing—review and editing, K.Y.H.L., M.S.Y.Y., P.K.S.S., R.W.K.W. and K.F.C.; visualization, K.Y.H.L. and R.W.K.W.; supervision, W.F.Y., K.Y.H.L. and K.F.C.; project administration, K.Y.H.L. and M.S.Y.Y.; funding acquisition, K.Y.H.L. All authors have read and agreed to the published version of the manuscript.
Funding
This research was funded by the Seed Fund for PI - Basic Research of the University of Hong Kong, Hong Kong SAR, China, grant number 2502251716.
Institutional Review Board Statement
The study was conducted in accordance with the Declaration of Helsinki, and approved by the Institutional Review Board of The University of Hong Kong/ Hospital Authority Hong Kong West Cluster (no. UW 25-299 and 8/7/2025).
Informed Consent Statement
Informed consent was obtained from all subjects involved in the study.
Data Availability Statement
The data that support the findings of this study are available from the corresponding author upon reasonable request.
Acknowledgments
We thank all the individuals who participated in the study.
Conflicts of Interest
The authors declare no conflicts of interest. The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.
Abbreviations
The following abbreviations are used in this manuscript:
| AI | Artificial intelligence |
| CARES-MDD | Conversational Artificial-intelligence for REducing Symptoms and REmission for Major Depressive Disorder |
| CBT | Cognitive Behavioral Therapy |
| DHIs | Digital Health Interventions |
| DSM-V | Diagnostic and Statistical Manual of Mental Disorders, Fifth Edition |
| GAD-7 | Generalized Anxiety Disorder 7-item scale |
| Gen-AI | Generative Artificial Intelligence |
| HITL | Human-in-the-loop |
| HK | Hong Kong |
| HS | Help-Seeking |
| ISI | Insomnia Severity Index |
| KAMHS | Knowledge and Attitudes to Mental Health Scales |
| LEE | Lived Experience Experts |
| LLM | Large Language Model |
| MDD | Major Depressive Disorder |
| MTR | Mass Transit Railway |
| CBT | Cognitive Behavioral Therapy |
| PHQ-9 | 9-item Patient Health Questionnaire |
| RCT | Randomized Controlled Trial |
| SD | Standard Deviation |
| SS | Self-Stigma |
| YPAR | Youth Participatory Action Research |
Appendix A
Table A1.
CARES-MDD system identity, intended use and technical configuration.
| System name | CARES-MDD (Conversational Artificial-intelligence for REducing Symptoms and REmission for Major Depressive Disorder) |
| Intended users | People attending psychiatric clinic in Hong Kong , currently using smartphone which has Internet access, and speaks Cantonese Chinese |
| Intended use | An adjunctive therapeutic tool delivering Cognitive Behavioral Therapy (CBT) exercises (e.g., cognitive restructuring), empathetic validation, and on-demand emotional support between formal psychiatric outpatient appointments. |
| Clinical boundaries | Not designed for diagnosis, emergency care, medication advice, crisis management, or replacement of standard treatment |
| Language and register | Hong Kong-style Cantonese |
| Underlying model | Requested on demand |
| Deployment setting | API hosted/processed in the server hosted within the universisty, chat content was not transferred to an external model provider |
| Generation approach | Structured system prompting RAG, with fine-tuning ; lexical crisis screening triggers a predefined safety override |
| Sources of the RAG | CBT-informed materials developed/curated by psychiatrists and clinical psychologists and reviewed by the clinical team and expert-curated therapeutic dialogues developed by a multidisciplinary psychiatric team. |
| Memory/context configuration | cross-session information retained during active using. |
| Safety mechanism | Real-time lexical screening for predefined high-risk crisis language; positive triggers suspend routine chatbot content and display a static emergency interface with 24-hour crisis resources |
| Clinical-review arrangement | Authorised clinical team members conducted periodic weekly review of chatbot content and system responses for quality and safety refinement; this was not continuous real-time monitoring or emergency response |
| Participant data used for model training | No. Participant conversation data were not used to train or fine-tune the underlying language model. |
| Material updates during study | No material changes to the model, prompts, safety rules, memory configuration, knowledge resources, or user interface occurred during participant data collection. |
Table A2.
Prioritized Co-Design Implementation Ideas for the CARES-MDD Intervention. Note. Priority rank reflects the order in which implementation ideas were prioritized within each co-design category.
Table A2.
Prioritized Co-Design Implementation Ideas for the CARES-MDD Intervention. Note. Priority rank reflects the order in which implementation ideas were prioritized within each co-design category.
| Evaluation Domain (Co-Design Category) | Priority Rank | Implementation Idea |
|---|---|---|
| Empathy & HK-Cantonese Understanding | 1 | Integrate natural Hong Kong-style Cantonese syntax and culturally appropriate conversational sensibilities to lower disclosure barriers. |
| 2 | Program the agent to explicitly validate negative emotions and distress before attempting to initiate problem-solving frameworks. | |
| 3 | Diversify empathetic phrasing to prevent repetitive, artificial responses during longitudinal interactions. | |
| Accuracy & Useful CBT Info (Therapeutic Fidelity) | 1 | Architect the system using modular knowledge banks (e.g., an eight-knowledge-bank structure) to deliver structured CBT. |
| 2 | Guide users systematically through cognitive restructuring exercises rather than offering superficial advice. | |
| 3 | Implement algorithmic detection of cognitive distortions (e.g., highlighting absolute words such as “always” or “must”) based on user input. | |
| Active Listening & Socratic Questions | 1 | Calibrate the pacing of Socratic questioning to evaluate user readiness, ensuring prompts do not feel demanding during acute distress. |
| 2 | Establish a conversational threshold (e.g., a minimum of five active turns) of supportive listening before transitioning into structured clinical guidance. | |
| 3 | Develop a seamless conversational mechanism allowing users to easily revert to an unstructured “treehole” venting mode when structured CBT is no longer desired. | |
| Supportive Pacing & Intervention Duration | 1 | Bound the testing window to an on-demand use instead of unlimited use to observe any adverse digital events and prevent digital fatigue. |
| 2 | Accommodate severe depressive amotivation and psychomotor slowing by removing immediate session timeouts, allowing asynchronous engagement. | |
| 3 | Optimize the interface to reduce typing burden where clinically appropriate, recognizing the cognitive fatigue associated with major depressive disorder. | |
| Balance of Privacy & Clinical Governance | 1 | Assign distinct pseudonyms or unique login identifiers to prevent data misattribution and ensure anonymity from primary treating teams. |
| 2 | Shield intimate conversational disclosures from the patient’s primary treating psychiatrist while allowing PI to inform case psychiatrists if worsening symptoms. | |
| 3 | Provide transparent, upfront psychoeducation regarding backend clinician oversight to reassure users about AI reliability and data governance. | |
| Safety Features & Adverse Digital Events | 1 | Implement automated crisis contingency messaging and safety overrides when crisis words are detected. |
| 2 | The contingent interface can appear, but can choose to close this interface afterwards, so that can troubleshoot the colloquial venting exclamations. | |
| 3 | Keep clinician oversight to supervise interactions and safeguard against potential large language model hallucinations. | |
| Acceptability of RCT Methodology | 1 | Position and guarantee the platform as a secure, ad-free environment based on uiniversity server to differentiate it from commercial applications and build participant trust. |
| 2 | Give reimbursement for time and allow online assessment (eg via zoom) related to psychometric outcome assessments to minimize participant time burden and during the followup period. | |
| 3 | Allow the other arm also to use active digital control, or else sense unfairness for the control arm |
Table A3.
Safety screening and crisis-response pathway.
| Stage | Implemented process | Trigger/category | User-facing response | Human action | Important limitation |
|---|---|---|---|---|---|
| Initial screening | Lexical-based | Predefined high-risk crisis keywords in Chinese/Cantonese (e.g., ‘suicidal’, ‘death’, ‘giving up’). | None (if negative) | No immediate human action for messages without a trigger. | Not a validated clinical risk assessment tool; highly susceptible to false positives from Cantonese colloquialisms (e.g., ‘死啦’ / ‘Oh no’). |
| Secondary check | Not applicable | Not applicable | Not applicable | Not applicable | No applicable |
| Safety override | Static emergency interface | Immediate upon detection of any initial screening keyword. | Static emergency message advising the user to seek urgent help and displaying 24-hour crisis resources/direct-dial contact options | None in real time | May interrupt conversation abruptly |
| Escalation | Consent is obtained for emergency contact and liaising with their existing professional team in case of safety crisis scalation | Significant risk detected by screening authorised clinical member | Participants are explicitly informed during onboarding and via the consent form that the tool is not an emergency service. | Authorised clinical members conducted periodic weekly review of selected content and system responses for safety and quality refinement; this did not constitute continuous real-time monitoring or emergency response. | The research team needs to manually check user accessed services, contacted a crisis resource, or received a clinical response after viewing the emergency interface. |
| Documentation | Automated event logging of all chat transcripts and triggered overrides. | All user interactions and safety events. | Not applicable | Periodic weekly clinical review of transcripts by the research team. | Trigger-event logs and review records do not by themselves establish whether risk was present, whether a false-positive occurred, whether the user sought help, or whether harm was prevented. |
Table A4.
Descriptive clinical-monitoring observations during CARES-MDD test-run. Note: HDRS-17 = 17-item Hamilton Depression Rating Scale; PHQ-9 = 9-item Patient Health Questionnaire; GAD-7 = 7-item Generalized Anxiety Disorder scale; SD = Standard Deviation. The maximum score for the Working Alliance Inventory is 60, and for Satisfaction with CARES-MDD is 7.
Table A4.
Descriptive clinical-monitoring observations during CARES-MDD test-run. Note: HDRS-17 = 17-item Hamilton Depression Rating Scale; PHQ-9 = 9-item Patient Health Questionnaire; GAD-7 = 7-item Generalized Anxiety Disorder scale; SD = Standard Deviation. The maximum score for the Working Alliance Inventory is 60, and for Satisfaction with CARES-MDD is 7.
| At the beginning of test-runs: Mean (SD) | At the end of the test-runs : Mean (SD) | Clinical Interpretation | |
|---|---|---|---|
| HDRS-17 | 19.33 (6.67) | 14.29 (6.37) | |
| PHQ-9 | 12.22 (4.49) | 8.00 (3.92) | |
| GAD-7 | 10.89 (3.48) | 8.86 (4.85) | |
| Working Alliance Inventory | 43.1 out of 60 (11.9) | Moderate to Strong therapeutic alliance | |
| Satisfaction with CARES-MDD | - | 5.71 (0.76) /7 | High overall satisfaction |
| System Usability Scale | - | 72.14 (6.68) | Above-average usability (>70 threshold) |
Table A5.
This table describes the architecture of CARES-MDD. It is intended to clarify the system’s content sources, module boundaries, and routing approach.
Table A5.
This table describes the architecture of CARES-MDD. It is intended to clarify the system’s content sources, module boundaries, and routing approach.
| Architecture domain | CARES-MDD specification, Content Governance, and Boundaries |
|---|---|
| Base model and deployment | CARES-MDD was implemented as a text-based conversational application using LLM. It is intended an add-on support tool; it was not intended to diagnose MDD, prescribe treatment, provide emergency care, or autonomously manage suicide risk. |
| Knowledge-base design | Clinical content was organised into separate, function-specific knowledge banks. This separation was intended to prevent patient-facing conversations from retrieving therapist-only guidance, detailed suicide-assessment material, medication guidance, or research/governance documents. The knowledge-bank framework comprised CBT foundations (KB01), assessment and formulation (KB02), process-to-intervention rules (KB03), CBT intervention modules (KB04), Traditional Chinese worksheets (KB05), maintenance and relapse prevention (KB06), safety and escalation (KB07), and research/system evidence (KB08). |
| RAG source materials | CBT-informed materials developed/curated by psychiatrists and clinical psychologists and reviewed by the clinical team and expert-curated therapeutic dialogues developed by a multidisciplinary psychiatric team. |
| RAG inclusion criteria | Only short, topic-specific, patient-appropriate knowledge units (“topic cards”) were intended for retrieval. Each unit was to include a defined title, use indications, exclusions, brief patient-facing content, one-question-at-a-time prompts, completion criteria, escalation criteria, source/page metadata, and clinical review status. Detailed suicide-assessment material, diagnostic flowcharts, exposure/ERP hierarchies, trauma-trigger forms, and no-suicide contracts were excluded from unrestricted generation. |
| Assessment and formulation boundary | KB02 was designed to support a structured understanding of the user’s current concern through situation–thought–emotion–behaviour links, goals, triggers, strengths, preferences, functional impact, and maintaining-process hypotheses. |
| Process-routing logic | The architecture proposed a process-to-intervention router (KB03) that identifies one primary maintaining-process hypothesis and selects one approved first-line module. Example mappings included behavioural activation for inactivity/anhedonia; thought monitoring and cognitive restructuring for negative automatic thoughts; emotion labelling and grounding for emotional overload; and structured problem solving for practical difficulties. |
| CBT module boundaries | The CBT-informed modules were intended to provide bounded psychoeducation and optional micro-exercises, including behavioural activation, automatic-thought monitoring, cognitive restructuring, grounding, relaxation, rumination management, problem solving, sleep/routine support, avoidance reduction, and interpersonal support. It was not authorised to provide medication advice or treatment recommendations outside approved content. |
| Traditional Chinese and cultural adaptation | Patient-facing exercises were selected and adapted for Traditional Chinese delivery, with a Hong Kong Cantonese conversational register. The intended design was to guide one brief exercise step at a time rather than deliver a full worksheet. |
| Data governance and privacy | The system was intended to use de-identified participant identifiers and restricted research-team access. Co-design feedback supported clear privacy boundaries, distinct identifiers/pseudonyms, and transparency regarding information that may be reviewed or shared. |
References
- Chan, J.K.N.; Solmi, M.; Lo, H.K.Y.; Chan, M.W.Y.; Choo, L.L.T.; Lai, E.T.H.; Wong, C.S.M.; Correll, C.U.; Chang, W.C. All-Cause and Cause-Specific Mortality in People with Depression: A Large-Scale Systematic Review and Meta-Analysis of Relative Risk and Aggravating or Attenuating Factors, Including Antidepressant Treatment. World Psychiatry 2025, 24, 404–421. [CrossRef]
- Malhi, G.S.; Mann, J.J. Depression. The Lancet 2018, 392, 2299–2312. [CrossRef]
- World Health Organization Depression and Other Common Mental Disorders: Global Health Estimates 2017.
- Rush, A.J.; Trivedi, M.H.; Wisniewski, S.R.; Nierenberg, A.A.; Stewart, J.W.; Warden, D.; Niederehe, G.; Thase, M.E.; Lavori, P.W.; Lebowitz, B.D.; et al. Acute and Longer-Term Outcomes in Depressed Outpatients Requiring One or Several Treatment Steps: A STARD Report. Am. J. Psychiatry 2006, 163, 1905–1917. [CrossRef]
- Thornicroft, G.; Chatterji, S.; Evans-Lacko, S.; Gruber, M.; Sampson, N.; Aguilar-Gaxiola, S.; Al-Hamzawi, A.; Alonso, J.; Andrade, L.; Borges, G.; et al. Undertreatment of People with Major Depressive Disorder in 21 Countries. Br. J. Psychiatry 2017, 210, 119–124. [CrossRef]
- Linardon, J.; Torous, J.; Firth, J.; Cuijpers, P.; Messer, M.; Fuller-Tyszkiewicz, M. Current Evidence on the Efficacy of Mental Health Smartphone Apps for Symptoms of Depression and Anxiety. A Meta-analysis of 176 Randomized Controlled Trials. World Psychiatry 2024, 23, 139–149. [CrossRef]
- Linardon, J.; Fuller-Tyszkiewicz, M.; Firth, J.; Goldberg, S.B.; Anderson, C.; McClure, Z.; Torous, J. Systematic Review and Meta-Analysis of Adverse Events in Clinical Trials of Mental Health Apps. Npj Digit. Med. 2024, 7, 363. [CrossRef]
- Chinsen, A.; Berg, A.; Nielsen, S.; Trewella, K.; Cronin, T.J.; Pace, C.C.; Pang, K.C.; Tollit, M.A. Co-Design Methodologies to Develop Mental Health Interventions with Young People: A Systematic Review. Lancet Child Adolesc. Health 2025, 9, 413–425. [CrossRef]
- Veldmeijer, L.; Terlouw, G.; Os, J.V.; Dijk, O.V.; Veer, J.V. ’t; Boonstra, N. The Involvement of Service Users and People With Lived Experience in Mental Health Care Innovation Through Design: Systematic Review. JMIR Ment. Health 2023, 10, e46590. [CrossRef]
- Heinz, M.V.; Mackin, D.; Trudeau, B.; Bhattacharya, S.; Wang, Y.; Banta, H.A.; Jewett, A.D.; Salzhauer, A.; Griffin, T.; Jacobson, N.C. Evaluating Therabot: A Randomized Control Trial Investigating the Feasibility and Effectiveness of a Generative AI Therapy Chatbot for Depression, Anxiety, and Eating Disorder Symptom Treatment 2024.
- Siddals, S.; Torous, J.; Coxon, A. “It Happened to Be the Perfect Thing”: Experiences of Generative AI Chatbots for Mental Health. Npj Ment. Health Res. 2024, 3, 48. [CrossRef]
- Torous, J.; Linardon, J.; Goldberg, S.B.; Sun, S.; Bell, I.; Nicholas, J.; Hassan, L.; Hua, Y.; Milton, A.; Firth, J. The Evolving Field of Digital Mental Health: Current Evidence and Implementation Issues for Smartphone Apps, Generative Artificial Intelligence, and Virtual Reality. World Psychiatry 2025, 24, 156–174. [CrossRef]
- Hayes, S.C.; Hofmann, S.G. “Third-wave” Cognitive and Behavioral Therapies and the Emergence of a Process-based Approach to Intervention in Psychiatry. World Psychiatry 2021, 20, 363–375. [CrossRef]
- Schefft, C.; Heinitz, C.; Guhn, A.; Brakemeier, E.-L.; Sterzer, P.; Köhler, S. Efficacy and Acceptability of Third-Wave Psychotherapies in the Treatment of Depression: A Network Meta-Analysis of Controlled Trials. Front. Psychiatry 2023, 14. [CrossRef]
- Gonsalves, P.P.; Ansari, S.; Berry, C.; Gonsalves, F.; Iyengar, S.; Kashyap, P.; Mittal, D.; Pal, S.; Razdan, E.; Michelson, D. Co-Designing Digital Mental Health Interventions with Young People: 10 Recommendations from Lessons Learned in Low-and-Middle-Income Countries. Health Educ. J. 2025, 84, 373–384. [CrossRef]
- Peters, S.; Guccione, L.; Francis, J.; Best, S.; Tavender, E.; Curran, J.; Davies, K.; Rowe, S.; Palmer, V.J.; Klaic, M. Evaluation of Research Co-Design in Health: A Systematic Overview of Reviews and Development of a Framework. Implement. Sci. 2024, 19, 63. [CrossRef]
- Warraich, H.J.; Tazbaz, T.; Califf, R.M. FDA Perspective on the Regulation of Artificial Intelligence in Health Care and Biomedicine. JAMA 2025, 333, 241–247. [CrossRef]
- So, E.; Kam, I.; Leung, C.; Chung, D.; Liu, Z.; Fong, S. The Chinese-Bilingual SCID-I/P Project: Stage 1 — Reliability for Mood Disorders and Schizophrenia.
- Kroenke, K.; Spitzer, R.L.; Williams, J.B.W. The PHQ-9 Validity of a Brief Depression Severity Measure. J Gen Intern Med 2001, 16.
- Simkiss, N.J.; Gray, N.S.; Dunne, C.; Snowden, R.J. Development and Psychometric Properties of the Knowledge and Attitudes to Mental Health Scales (KAMHS): A Psychometric Measure of Mental Health Literacy in Children and Adolescents. BMC Pediatr. 2021, 21, 508. [CrossRef]
- Tsoi, J.K.K.; Yuen, S.S.Y.; Kwok, C.; Xu, N.S.; Tsoi, X.L.; Hui, C.C.Y.; Yuen, P.T.Y.; Cheng, A.W.C.; Chu, W.L.; Hao, J.J.X.; et al. Translation and Validation of the Traditional Chinese Version of the Knowledge and Attitudes to Mental Health Scales (KAMHS) in Hong Kong Adolescents. East Asian Arch Psychiatry 2026, In press.
- Cammarota, J.; Fine, M. Revolutionizing Education: Youth Participatory Action Research in Motion; Routledge, 2010; ISBN 978-1-135-91324-3.
- Blease, C.; Tibbs, M.; Balaskas, A.; Liverpool, S.; Hagström, J.; Fitzgerald, A. Coproduction Without Youth? Closing the Participation Gap in Digital Mental Health Research. JMIR Ment. Health 2026, 13, e91739. [CrossRef]
- Braun, V.; Clarke, V. Using Thematic Analysis in Psychology. Qual. Res. Psychol. 2006, 3, 77–101. [CrossRef]
- Marks, D.F.; Yardley, L. Research Methods for Clinical and Health Psychology; SAGE, 2003; ISBN 978-1-4462-3232-3.
- Thornicroft, G. Stigma and Discrimination Limit Access to Mental Health Care. Epidemiol. Psychiatr. Sci. 2008, 17, 14–19. [CrossRef]
- Lee, J.; Lee, D.; Lee, J. Influence of Rapport and Social Presence with an AI Psychotherapy Chatbot on Users’ Self-Disclosure. Int. J. Human–Computer Interact. 2024, 40, 1620–1631. [CrossRef]
- Lee, Y.-C.; Yamashita, N.; Huang, Y.; Fu, W. “I Hear You, I Feel You”: Encouraging Deep Self-Disclosure through a Chatbot. In Proceedings of the Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems; Association for Computing Machinery: New York, NY, USA, April 23 2020; pp. 1–12.
- Grové, C. Co-Developing a Mental Health and Wellbeing Chatbot With and for Young People. Front. Psychiatry 2021, 11, 606041. [CrossRef]
- Sit, H.F.; Ling, R.; Lam, A.I.F.; Chen, W.; Latkin, C.A.; Hall, B.J. The Cultural Adaptation of Step-by-Step: An Intervention to Address Depression Among Chinese Young Adults. Front. Psychiatry 2020, 11. [CrossRef]
- Ma, Q.; Shi, Y.; Zhao, W.; Zhang, H.; Tan, D.; Ji, C.; Liu, L. Effectiveness of Internet-Based Self-Help Interventions for Depression in Adolescents and Young Adults: A Systematic Review and Meta-Analysis. BMC Psychiatry 2024, 24, 604. [CrossRef]
- Elyoseph, Z.; Levkovich, I. Beyond Human Expertise: The Promise and Limitations of ChatGPT in Suicide Risk Assessment. Front. Psychiatry 2023, 14. [CrossRef]
- Keshavan, M.; Torous, J.; Yassin, W. Do Generative AI Chatbots Increase Psychosis Risk? World Psychiatry 2026, 25, 150–151. [CrossRef]
- Meadi, M.R.; Sillekens, T.; Metselaar, S.; Balkom, A. van; Bernstein, J.; Batelaan, N. Exploring the Ethical Challenges of Conversational AI in Mental Health Care: Scoping Review. JMIR Ment. Health 2025, 12, e60432. [CrossRef]
- Tie, J.; Yao, B.; Li, T.; Fang, H.; Ahmed, S.I.; Wang, D.; Zhou, S. “Should I Give Up Now?” Investigating LLM Pitfalls in Software Engineering. ACM Trans. Softw. Eng. Methodol. 2026. [CrossRef]
- Dayan, R.; Uliel, B.; Koplewitz, G. Age against the Machine—Susceptibility of Large Language Models to Cognitive Impairment: Cross Sectional Analysis. BMJ 2024, 387, e081948. [CrossRef]
- Freyer, O.; Wiest, I.C.; Kather, J.N.; Gilbert, S. A Future Role for Health Applications of Large Language Models Depends on Regulators Enforcing Safety Standards. Lancet Digit. Health 2024, 6, e662–e672. [CrossRef]
- Linardon, J.; Cuijpers, P.; Carlbring, P.; Messer, M.; Fuller-Tyszkiewicz, M. The Efficacy of App-Supported Smartphone Interventions for Mental Health Problems: A Meta-Analysis of Randomized Controlled Trials. World Psychiatry 2019, 18, 325–336. [CrossRef]
- Ho, F.Y.Y.; Chung, K.F.; Yeung, W.F.; Ng, T.H.Y.; Cheng, S.K.W. Weekly Brief Phone Support in Self-Help Cognitive Behavioral Therapy for Insomnia Disorder: Relevance to Adherence and Efficacy. Behav. Res. Ther. 2014, 63. [CrossRef]
- Lipschitz, J.M.; Pike, C.K.; Hogan, T.P.; Murphy, S.A.; Burdick, K.E. The Engagement Problem: A Review of Engagement with Digital Mental Health Interventions and Recommendations for a Path Forward. Curr. Treat. Options Psychiatry 2023, 10, 119–135. [CrossRef]
- Li, Y.; Luk, T.T.; Cheung, Y.T.D.; Zhao, S.; Zeng, Y.; Tong, H.S.C.; Lai, V.W.Y.; Wang, M.P. Engagement With a Mobile Chat-Based Intervention for Smoking Cessation: A Secondary Analysis of a Randomized Clinical Trial. JAMA Netw. Open 2024, 7, e2417796. [CrossRef]
- Tucker, I.; Bertotti, M.; Hanafiah, A.; Hossain, S.; Watts, P.; Ahad, M.A.R. Developing a Digital Ecological Momentary Assessment Tool for ‘Real Time’ Evaluation in Implementation Science: Testing through Evaluation of a Novel Digital Social Prescribing Intervention. Front. Public Health 2026, 14. [CrossRef]
- Chirokoff, V.; Tessier, A.; Serre, F.; Dupuy, M.; Auriacombe, M.; Chanraud, S.; Berthoz, S.; Fatseas, M.; Misdrahi, D. Relevance of Ecological Momentary Assessment for Medication Adherence in Clinical Settings: A Precision Psychiatry Approach. Br. J. Clin. Psychol. 2025, 64, 692–701. [CrossRef]
- Knott, M.; Krebs, M.; Kerscher, A. Large Language Models in Healthcare Quality Management: A European Perspective on Process Automation and Compliance. Front. Digit. Health 2026, 8. [CrossRef]
- Ong, J.C.L.; Ning, Y.; Liu, M.; Ma, Y.; Zhao, L.; Singh, K.; Chang, R.T.; Vogel, S.; Lim, J.C.W.; Tan, I.S.K.; et al. Innovating Global Regulatory Frameworks for Generative AI in Medical Devices Is an Urgent Priority. Npj Digit. Med. 2026, 9, 364. [CrossRef]
- Esmaeilzadeh, P.; Mirzaei, T.; Dharanikota, S. Patients’ Perceptions Toward Human–Artificial Intelligence Interaction in Health Care: Experimental Study. J. Med. Internet Res. 2021, 23, e25856. [CrossRef]
Figure 1.
Word cloud of frequently occurring Chinese terms in participants’ qualitative comments regarding barriers to sharing emotional difficulties with others. Larger words represent terms that appeared more frequently in the comments. Prominent terms included “不知道” (“do not know”), “私隱” (“privacy”), “分享” (“sharing”), “感受” (“feelings”), “別人” (“others”), “負面” (“negative”), “壓力” (“pressure”), and “需要” (“need”).
Figure 1.
Word cloud of frequently occurring Chinese terms in participants’ qualitative comments regarding barriers to sharing emotional difficulties with others. Larger words represent terms that appeared more frequently in the comments. Prominent terms included “不知道” (“do not know”), “私隱” (“privacy”), “分享” (“sharing”), “感受” (“feelings”), “別人” (“others”), “負面” (“negative”), “壓力” (“pressure”), and “需要” (“need”).

Figure 2.
Word cloud of frequently occurring Chinese terms in participants’ qualitative comments regarding the most desired features of CARES-MDD. Larger words represent terms that appeared more frequently in the comments. Prominent terms included “隨時” (“anytime”), “安全” (“safe”), “陪伴” (“accompany”), “廣東話” (“Cantonese”), “隱私” (“privacy”), “聆聽” (“listen”), and “即時” (“immediate”).
Figure 2.
Word cloud of frequently occurring Chinese terms in participants’ qualitative comments regarding the most desired features of CARES-MDD. Larger words represent terms that appeared more frequently in the comments. Prominent terms included “隨時” (“anytime”), “安全” (“safe”), “陪伴” (“accompany”), “廣東話” (“Cantonese”), “隱私” (“privacy”), “聆聽” (“listen”), and “即時” (“immediate”).

Figure 3.
Boxplot Distribution of Expert Reference Group Scores Across Seven Evaluative Domains for CARES-MDD (10-Point Visual Analogue Scale; n = 18).
Figure 3.
Boxplot Distribution of Expert Reference Group Scores Across Seven Evaluative Domains for CARES-MDD (10-Point Visual Analogue Scale; n = 18).

Table 1.
AI Evaluative Rubric for CARES-MDD for their acceptability, perception of safety and their opinions on a future trial design.
Table 1.
AI Evaluative Rubric for CARES-MDD for their acceptability, perception of safety and their opinions on a future trial design.
| Evaluation Domain | Assessment Question | Clinical Rationale |
|---|---|---|
| Empathy and HK-Cantonese understanding | Cantonese understanding Did CARES-MDD express empathy and validate negative emotions naturally using locally acceptable language and tone? |
An empathetic, localized understanding of the patients’ motions reduces the barrier to disclosure and builds therapeutic alliance. |
| Accuracy and useful CBT info | Did CARES-MDD provide accurate CBT information and responses rather than providing superficial advice? | Skilled agents must address accurate CBT frameworks (e.g., cognitive restructuring) aids rather than just surface-level symptoms. |
| Active listening and Socratic questions | Did CARES-MDD use active listening followed by appropriately timed Socratic questioning? | Appropriate pacing of validation before Socratic questioning promotes self-exploration without feeling demanding. |
| Supportive pacing and intervention duration | Based on your experience with CARE-MDD’s pacing and technical glitches, what are your recommendations regarding the optimal intervention duration to sustain engagement without causing digital fatigue? | Identifying the technical barriers that hinder interaction, and identify the onset of digital fatigue to determine an acceptable duration of “dosage” for a future definitive RCT while balancing retention. |
| Balance of privacy and clinical goverance | Did CARES-MDD balance the clinical monitoring with clinician oversight with the user’s need for respect of privacy? | Establishing transparent privacy boundaries while maintaining necessary clinical monitoring is essential for ethical governance, ensuring users feel secure enough to disclose without fearing unwarranted data integration. |
| Safety features and adverse digital events |
Did CARES-MDD’s automated safety interface provide a necessary safety net without causing disruptive frustration, and did the interaction induce feelings of being dismissed by a machine? |
Evaluating the trade-off between strict automated risk mitigation and therapeutic flow, while proactively identifying potential iatrogenic harm, safety prerequisite for ethical governance of of a future RCT. |
| Acceptability of RCT methodology | Will persons living with MDD willing to be randomized (including to a psychoeducation control arm) and be okay with proposed evaluation project timeline? | Validates the feasibility of a proposed RCT design, ensuring the recruitment strategy, assessment burden, and control comparators are acceptable to the target clinical demographic. |
Table 2.
Phase 1 Clinical Needs Assessment Findings Note. SD = Standard Deviation; PHQ-9 = Patient Health Questionnaire-9.
Table 2.
Phase 1 Clinical Needs Assessment Findings Note. SD = Standard Deviation; PHQ-9 = Patient Health Questionnaire-9.
| Variable | Result |
|---|---|
| Average Time per day spent online | 296.3 mins/day |
| Barriers to share emotional difficulties within traditional psychiatric system | Privacy concerns, Fear of judgment (negative consequences), Burdening others (causing trouble) |
| Help-Seeking Behavior towards traditional support systems (Overall Score, max 28) | 14.22 ± 5.26 |
| Help-Seeking towards traditional support systems: Depressed vs. Non-depressed | Depressed: 12.37 | Non-depressed: 15.40 (p < 0.001) |
| Self-Stigma (Overall Score, max 24, higher = lower stigma) | 14.05 ± 4.79 |
| Self-Stigma: Depressed vs. Non-depressed | Depressed: 12.10 | Non-depressed: 15.29 (p < 0.001) |
| Prevalence of generic AI chatbot Adoption | 66.2% |
| Average AI chatbot Time Spent (Entire Sample) | 56.9 ± 87.7 mins/day |
| Average AI chatbot Time Spent (Active Users Only) | 85.9 ± 95.5 mins/day |
| Frequency of AI chatbot Interactions (Active Users) | < 10 total conversations: 56.4% 10-50 conversations: 19.8% Several times a week: 13.9% Up to 2 hours daily: 5.1% 2-5 hours daily: 4.9% |
Table 3.
Users’ experiences of good and bad points for the CARES-MDD chatbot.
| Mention level and category | Positive points | Negative points |
|---|---|---|
| Most mentioned | ||
| App-based Modality & Usability | Functions effectively as a private “treehole” where users feel comfortable sharing anything; perceived as much safer than commercial mental health apps due to the absence of targeted advertising based on disclosed content. Its use of natural, Hong Kong-style Cantonese was perceived as familiar and emotionally accessible, and its understanding of local expressions was regarded as appropriate. | Several participants suggested that the intervention might be less suitable for older users or individuals with limited motivation, concentration, or comfort with typing, given the need for regular text input.But text-based format allowed to use CARES-MDD without being overheard and could access it in settings where speaking with another person would be difficult. |
| Therapeutic Features (CBT-informed psychoeducation and self-management prompts) | Experts rated the CBT information as clinically appropriate; the chatbot maintains a structured, supportive tone that differentiates itself from generic emotional venting. The system was perceived as more useful after several active exchanges, particularly when a session involved at least five conversational turns. | Requires at least 5 active interactions to move beyond superficial advice. Socratic questioning can feel demanding if the AI misjudges the user’s readiness, occasionally sounding “fake”. Struggles to switch back to a simple venting mode when structured CBT is no longer desired. |
| Frequently mentioned | ||
| Study Design & Intervention Pacing | The conversational format allowed participants to engage at their own pace and return to the system when support was needed. The extended intervention period provided an opportunity to observe engagement across different stages of use. | Participants identified digital fatigue if unlimited use. Several participants mentioned technical limitations included inconsistent recall of contextual details from earlier sessions, which contrasted with the continuity expected from a human therapist. |
| Moderately mentioned | ||
| Safety features and adverse digital events | Users are comfortable with data syncing to clinical logs provided that distinct pseudonyms (different login names) are used to prevent data mixing. Users value that specific, intimate details are shielded from their primary treating psychiatrists. Some underlying anxiety remains regarding potential AI hallucinations based on external news reports, though not directly experienced. | The automated safety override feels abrupt and can be triggered prematurely by false-positive slang exclamations during normal emotional venting (e.g., ‘死啦’ / ‘Oh no’). Follow-up and systematic and clinical safety monitoring were considered necessary. |
| Less mentioned | ||
| Acceptability of future randomised evaluation | Participants viewed an active-control randomised controlled trial as acceptable and potentially valuable for distinguishing intervention effects from expectancy effects. They considered the proposed outcome measures broadly appropriate and supported longer-term follow-up. | Participants noted the assessments are necessary but spend time, and recommended adequate compensation for participants. Active control as considered as more “fair” to those joining RCT as they would expect to be assigned to aa “helpful” group if participating. |
Table 4.
Stakeholder-Informed Safeguards and Clinical Oversight Considerations for Potential LLM-Related Risks in CARES-MDD.
Table 4.
Stakeholder-Informed Safeguards and Clinical Oversight Considerations for Potential LLM-Related Risks in CARES-MDD.
| Identified LLM Vulnerability | System-level architectural safeguard | Human-in-the-Loop (HITL) & Clinical Oversight |
|---|---|---|
| AI Hallucinations & Reinforcement of Harmful Viewpoints | Retrieval-augmented generation responses are constrained by clinician-reviewed CBT-informed knowledge resources, to mitigate the chance of extrapolating “hallucinatory content” beyond validated, evidence-based frameworks | Samples of transcripts will be reviewed weekly for fidelity to intended content, and potential safety concerns. |
| “Context Rot” (Inconsistent Longitudinal Memory) | The system is explicitly designed for brief, bounded interactions rather than continuous, long-term narrative therapy, months-long narrative therapy, thereby minimizing the model’s reliance on deep historical recall | Users are informed during onboarding that the AI is a present-moment skills coach, managing expectations regarding the AI’s long-term memory capabilities. |
| Iatrogenic Risk of Abrupt Automated Crisis Overrides & Imminent Risk | Supportive crisis-signposting interface: The system triggers a supportive safety interface offering localized crisis hotlines alongside the dialogue, avoiding abrupt conversational termination. | Consent is obtained during onboarding stipulating that, despite using pseudonyms, the research team is authorized to contact the patient’s primary case psychiatrist to advance follow-up if severe risk is detected during periodic reviews. |
| Digital Fatigue & Unhealthy Dependency | The intervention protocol will be structurally bounded to maximum active period to prevent users from reducing the likelihood that chatbot use displaces usual psychiatric care | Evaluates the retention and engagement time of active sessions (defined as >= 10 minutes and >= 5 active conversational turns). Integrates ecological momentary assessment (EMA) techniques which have shown evidence in boosting adherence through self-monitoring and reminder effects. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.