Preprint
Article

This version is not peer-reviewed.

Engagement Phenotypes in a Real-World Tirzepatide-Based Dig-Ital Weight-Loss Program: Unsupervised Clustering of Self-Tracking, Human Coaching, and AI Conversational Support

Submitted:

26 June 2026

Posted:

29 June 2026

You are already at the latest version

Abstract
Background and Objectives: Real-world adherence to medication-supported weight management programs is often low. While digital weight loss services (DWLS) provide multi-modal digital supports to improve engagement and counter attrition, existing literature frequently relies on unidimensional or binary classifications of user engagement. This study used unsupervised machine learning to identify distinct digital engagement phenotypes and evaluated their independent associations with 6-month weight loss outcomes in patients prescribed tirzepatide. Materials and Methods: This retrospective cohort study analyzed deidentified data from 9,470 medication-adherent, complete-case adult patients within an Australian DWLS who initiated tirzepatide between May 20 and December 2, 2025. K-means clustering was performed on four continuous, longitudinal usage metrics: weekly app logins, health coach messaging, automated assistant (JuneBot) messaging, and weight tracking. To evaluate the primary clinical endpoint - 6-month percentage weight loss - unadjusted pairwise comparisons (Tukey HSD) and a fully adjusted ordinary least squares multivariate linear regression model were executed to control for baseline demographic, clinical, and interim behavioral covariates. Results: Four stable engagement phenotypes emerged: non-engaged (n = 1,772), high health-coach engagement (n = 1,141), high JuneBot engagement (n = 1,472), and passive self-trackers (n = 5,085). In the multivariate regression model (Adjusted R2 = 0.3606, p < 0.001), all active phenotypes were strong independent predictors of weight loss success relative to the non-engaged baseline. The adjusted incremental weight loss premiums were similar between the automated conversational cohort (β = 2.28%, SE = 0.24, p < 0.001) and the human health-coaching cohort (β = 2.17%, SE = 0.26, p < 0.001), while passive self-trackers demonstrated a smaller independent premium (β = 1.67%, SE = 0.20, p < 0.001). Conclusions: In a real-world per-protocol cohort, active engagement in a DWLS is associated with significantly improved 6-month weight loss outcomes. Under strict clinical safety guardrails, structured automated conversational support may achieve outcome parity with human coaching, offering a highly scalable mechanism to complement pharmacologic obesity care.
Keywords: 
;  ;  ;  ;  

1. Introduction

Obesity represents one of the most pervasive and severe global health challenges of the 21st century, deeply impacting both public health infrastructure and individual life expectancy. Chronic weight-related conditions - including type 2 diabetes, cardiovascular disease, metabolic dysfunction-associated steatotic liver disease, and various malignancies - contribute significantly to global mortality and place immense financial strain on healthcare systems worldwide [1]. According to recent World Health Organization (WHO) estimates, more than 1 billion individuals globally are living with obesity, a figure that continues to rise across virtually every demographic sector [2]. For decades, clinical interventions relied heavily on behaviorally driven lifestyle modifications; however, long-term weight loss maintenance through willpower and caloric restriction alone has historically yielded low success rates due to complex physiological and metabolic counter-regulatory mechanisms that resist weight reduction [3].
The recent emergence and widespread adoption of glucagon-like peptide-1 receptor agonists (GLP-1 RAs) and dual GIP/GLP-1 receptor agonists, such as tirzepatide, have fundamentally transformed the chronic weight management landscape. These pharmacological therapies provide highly effective, scalable solutions by mimicking incretin hormones to enhance satiety, delay gastric emptying, and regulate central appetite pathways [4], resulting in unprecedented double-digit percentage weight loss in clinical trials [5]. Nevertheless, international health institutions, including the WHO and the National Institute for Health and Care Excellence (NICE), explicitly stress that these medications are not standalone cures. Instead, clinical guidelines dictate that GLP-1 RAs should strictly serve as adjuncts to continuous, comprehensive obesity care, which includes structured behavioral counseling, nutritional restructuring, and physical activity tracking to preserve lean mass and prevent rapid weight regain upon treatment cessation [6].
Despite this clear clinical consensus, delivering continuous, comprehensive obesity care remains exceedingly difficult in traditional face-to-face medical settings. Patients frequently struggle to adhere to intensive lifestyle modification programs due to systemic barriers such as geographic remoteness, high financial costs, limited practitioner availability, and the social stigma often encountered in physical clinic environments. As a result, attrition rates in conventional behavioral interventions remain high, leaving a critical gap in multi-disciplinary support during pharmacological treatment [7].
To overcome these physical limitations, app-mediated Digital Weight Loss Services (DWLS) have expanded rapidly, significantly improving patient access to remote multi-disciplinary teams (MDTs) and structured educational content [6,8]. However, digital health critics argue that simply providing access to a mobile application does not guarantee active patient engagement or continuity of care. Because digital platforms inherently rely on user initiative, large portions of patient cohorts can default to passive usage or disengage entirely, potentially undermining the long-term efficacy of the supported medical therapy [9].
Consequently, healthcare scholars have begun investigating the explicit relationship between digital application engagement patterns and clinical weight loss outcomes [10,11]. While these early publications have consistently found that “more engagement” is associated with better outcomes, they typically operationalise engagement using coarse, often binary or single-metric constructs - for example, classifying patients as “engaged” versus “non-engaged” based on a minimum app-use threshold. This approach is not unique to weight-loss programs: similar binary or unidimensional engagement markers are widely used across digital mental-health, diabetes self-management, and remote cardiac-rehabilitation interventions. Such simplifications obscure the fact that modern digital platforms support multiple, overlapping engagement modalities - including passive biometric self-tracking, asynchronous human coaching, automated conversational agents, and educational content consumption - that may contribute differently to outcomes. A more nuanced understanding of these multi-modal engagement patterns is needed to inform how best to design and allocate scalable digital supports alongside pharmacologic therapy.
To directly address the specific methodological gap presented by the binary classification used in recent literature, this study uses unsupervised clustering to derive naturally occurring digital engagement phenotypes and then examines how these patterns relate to weight-loss outcomes. By shifting from a static binary to a multi-dimensional analysis, we aim to map the naturally occurring, nuanced behavioral trajectories that emerge when patients interact with diverse digital resources. The primary objective of this study is to evaluate the independent effect of distinct digital engagement phenotypes on 6-month percentage weight loss in a large per-protocol cohort of DWLS patients receiving tirzepatide treatment. Through this approach, we seek to determine the extent to which a highly scalable, rule-based AI framework can support patients in a comprehensive medicated DWLS.

2. Materials and Methods

Study Design and Patient Population

This study adopted a retrospective cohort design to evaluate the relationship between distinct digital engagement phenotypes and 6-month weight-loss outcomes among patients utilizing an app-mediated DWLS. The primary objective was to examine the effect of different longitudinal engagement modalities on 6-month percentage weight loss supported by tirzepatide treatment. Investigators followed the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) guidelines throughout the study design and reporting phases [12].
To achieve the study objectives, investigators filtered for patients who adhered to the program over 6 months (Per Protocol). Patients were included in the final analysis if they met the following criteria:
  • They initiated treatment between May 20 and December 2 2025. The former date ensured access to all active automated and human coaching features (Junebot was launched on May 20, 2025), while the latter was exactly 6 months prior to data extraction (2 June, 2026) and enabled all patients the chance to submit follow-up weight data.
  • They received a minimum of 5 Tirzepatide medication orders within a 183-day (6-month) observation window to ensure adequate pharmacological exposure.
  • They provided a verified body weight entry within a strict 6-month clinical window, defined as 173 to 193 days post-program initiation.
All study data were stored in the Juniper central data repository on Google BigQuery, a serverless, cloud-based warehouse. Data were extracted by investigators via the SQL language. The Stanford University Independent Research Board determined that the study did not meet the definition of human subject research as defined in federal regulations 45 CFR 46.102 or 21 CFR 50.3.

Program Overview

The Juniper UK DWLS serves adults with overweight or obesity (BMI ≥25) and operates as an asynchronous, GLP-1 RA-supported clinical program designed for chronic weight management. The journey begins with a comprehensive digital triage in which prospective patients submit a pre-consultation questionnaire covering baseline health markers. These data are reviewed by a pharmacist independent prescriber to assess whether the patient meets the clinical criteria for pharmacological intervention. Patients are eligible for tirzepatide prescription if they have a BMI of ≥30 kg/m2, or a BMI of 27–29.9 kg/m2 in the presence of at least one weight-related comorbidity (National Institute for Health and Care Excellence, 2024), with off-label prescribing considered at the clinical discretion of the prescriber for patients below these thresholds. They also detail absolute contraindications for tirzepatide-supported treatment, which include acute kidney disease, acute pancreatitis, hypoglycaemia, severe gastrointestinal disease, multiple endocrine neoplasia syndrome type 2, a personal or family history of medullary thyroid cancer, and a known sensitivity to tirzepatide or any of the product’s components [13].
Patients who are deemed eligible are then asked to pay a monthly subscription fee. Over the study period, first month fees ranged from 189 to 279 Great Britain Pounds (GBP), and highest fees (for higher tirzepatide doses) ranged from 294 to 339 GBP. Upon completion of the first monthly payment, patients receive access to the Juniper UK application, which includes educational content, a bluetooth-connected weight tracking tool, and unlimited communication with a multidisciplinary (MDT) care team. MDTs consist of a prescribing practitioner, a university-qualified health coach (nutritionist or dietitian), and a medical support officer. Health coaches send patients a fortnightly check-in message unless patients communicate at a higher frequency. Educational content is multimodal, structured around core metabolic pillars including sustainable caloric deficits, macronutrient/protein targets, and physical movement tracking. Adverse events are self-reported via the app using a severity-rating scale, where moderate events trigger ad hoc medical consultations and severe events trigger immediate escalation protocols. All patient communications, tracking logs, and questionnaire responses are stored in Juniper’s central data repository.

JuneBot Functionality

Research In addition to human coaching, the Juniper digital ecosystem features an automated conversational assistant named JuneBot. JuneBot functions under a restricted, rule-based and guideline-driven communication architecture. Rather than executing automated changes or modifying a patient’s active clinical record, it operates purely as an empathetic, non-judgmental information resource and conversational support tool designed to reinforce the program’s baseline programmatic guidelines. A previous study was dedicated to testing the safety of the agent’s communication within the Juniper program [14].
To maintain strict clinical safety boundaries, explicit operational guardrails govern the JuneBot interaction matrix:
  • No Medical Advice: JuneBot is prohibited from delivering personalized clinical assessments, diagnostic evaluations, or prescription adjustments. It only dispenses global, established knowledge bases and pre-approved clinical guidelines.
  • Non-Agentic Framework: The assistant cannot act on behalf of the patient, manipulate medication order frequencies, or execute system-level billing or programmatic operations within the application.
  • No MDT Substitution: JuneBot serves purely as a programmatic augment and is structurally prevented from replacing a practitioner, pharmacist, or university-qualified health coach.
  • Mental Health Guardrails: The system does not manage or process mental health protocols. Conversational inputs that flag psychological distress, severe body dysmorphia, or crises bypass the automated interface entirely and trigger an immediate, mandatory human escalation protocol to the clinical support team.
JuneBot’s informational repository is partitioned into specific behavioral and programmatic domains, including diet, nutrition, and eating behavior; exercise and physical movement; weight loss tracking dynamics; and standardized application/technical support. Regarding conversation frequency, JuneBot operates primarily on an opt-in engagement model, initiating conversational sequences only after a patient logs an initial conversational input. This user-initiated framework is supplemented by a structured longitudinal engagement protocol consisting of a low-frequency, automated weekly “nudge” designed to establish continuous behavioral reinforcement without inducing digital communication fatigue.

Statistical Analysis and Cluster Generation

An unsupervised machine learning approach was implemented using K-means clustering [15] to discover naturally occurring behavioral archetypes without imposing arbitrary categorical thresholds. The clustering algorithm was trained on four continuous behavioral percentage metrics captured longitudinally from program initiation up to the 6-month endpoint:
  • Weekly app engagement percentage: The proportion of weeks a patient actively logged into the mobile application interface.
  • Weekly health coach messaged percentage: The proportion of weeks a patient initiated or replied to conversational messages with a university-qualified human health coach.
  • Weekly JuneBot messaged percentage: The proportion of weeks a patient initiated or replied to conversational messages with JuneBot.
  • Weekly weight track percentage: The proportion of weeks a patient tracked their weight via the Juniper bluetooth scales.

Mathematical Preprocessing and Optimization

Because K-means clustering is highly sensitive to variances in variable scaling, all four continuous engagement metrics were standardized via Z-score transformation (μ=0,σ=1) prior to distance matrix calculation. Missing data points within these critical fields were handled via listwise deletion to ensure complete cases for the clustering algorithm.
To identify a suitable number of clusters (K), the scaled data were evaluated across a search space of 1 ≤ K ≤ 10 using two distinct metrics: the Elbow Method (tracking total within-cluster sum of squares [WSS]) and the Average Silhouette Method [16]. The silhouette width was maximised at K = 2, suggesting a coarse division between low- and high-engagement users, whereas the WSS curve continued to decrease steeply up to approximately K = 4 before flattening. Because K = 2 merged conceptually distinct engagement modes (for example, conversational versus predominantly passive tracking) into a single “engaged” class, we selected candidate solutions in the range K = 3-5 and prioritised those that yielded clinically interpretable phenotypes.
To further assess the robustness of the 4-cluster solution, we computed the gap statistic across K = 1-10 and conducted bootstrap cluster-stability analyses. The gap statistic increased sharply from K = 1 to K = 4-5, with only modest gains beyond this range, and identified K = 5 as the optimal value under the Tibshirani et al. criterion (Supplementary Figure S1, Table S1). Because the gap curve formed a clear plateau between K = 4 and K = 7, and because K = 4 yielded a parsimonious and clinically interpretable separation between non-engaged, conversationally engaged, and passive-tracking phenotypes, we retained the 4-cluster solution for the main analyses. Bootstrap resampling (B = 100) with k-means-based clustering yielded mean Jaccard similarities of 0.99, 0.99, 0.99, and 0.98 for the four clusters, indicating excellent reassignment stability and limited evidence of cluster fragmentation or merging (Supplementary Table S2).
For visualisation purposes only, we then projected the four-dimensional engagement metrics onto the first two principal components, which together explained 81.2% of the variance; this projection is presented in the Results to illustrate the geometric separation of the clusters.

Prototype Profiling and Nominal Variable Generation

Following cluster generation, individual patient cluster assignments were extracted as discrete nominal factors and merged back into the primary unscaled dataset. For clinical interpretation, we assigned descriptive labels based on dominant engagement patterns across the four metrics. One cluster exhibited minimal app use, messaging activity, and weight tracking and was labelled “non-engaged.” A second showed high app use, frequent weight tracking, and frequent messaging with human health coaches (“high health-coach engagement”). A third had similarly high app and tracking activity but predominantly interacted via the automated assistant rather than human coaches (“high JuneBot engagement”). The fourth cluster combined high app use and consistent weight tracking with minimal conversational messaging (“passive self-trackers”). These nominal cluster assignments served as the primary exposure variables for all downstream clinical endpoint evaluations and are described quantitatively in the Results.

Statistical Analysis

The primary clinical endpoint was the observed 6-month percentage weight loss, defined as the percentage change in body weight between program initiation and the closest verified weight measurement within a strict 6-month follow-up window (days 173-193 after initiation). Descriptive statistics for this clinical endpoint were stratified across the four behavioral clusters and summarized as means with standard deviations (SD).
To examine the independent effect of each engagement prototype on 6-month weight loss, a multivariate ordinary least squares linear regression model was constructed. The primary independent variable of interest was the 4-category cluster assignment nominal factor, with the non-engaged group designated as the reference category.
To reduce confounding from baseline and interim differences, the regression model adjusted for various demographic, clinical and behavioural covariates. Demographic variables included age, ethnicity, sex at birth and address remoteness. Clinical markers were baseline BMI, comorbidity count, side effect incidence and GLP-1 RA use in the 6 months prior to program start. Similar to previous studies, behavioural covariates included first month weight track count, weight loss plateau, attainment of a healthy BMI (≤25kg/m2) or a patient’s target weight, program pause count and clinically significant weight loss (≥5%) within the first 3 months. Unstandardized regression coefficients with corresponding 95% confidence intervals (CI) were calculated for each engagement prototype to represent the incremental percentage weight loss achieved relative to the non-engaged baseline.
All statistical pipelines and data visualizations were executed within RStudio (version 2023.06.1). Statistical significance for all downstream analyses was maintained at a two-tailed α=0.05.

3. Results

A total of 39,220 patients initiated tirzepatide treatment within the specified study window. Of these, 17,600 patients met the baseline per-protocol medication adherence criteria fulfilling a minimum of 5 medication orders by 183 days post program initiation. Within this medication-adherent cohort, 11,460 patients successfully provided a verified 6-month body weight entry within the strict day 173–193 clinical follow-up window. A further 1990 patients were omitted from the final analysis matrix due to having changed medication to semaglutide, yielding a final, complete-case study population of 9,470 patients for behavioral cluster generation and outcome modeling.
The baseline demographic and clinical profiles of the 9,470 patients were stratified across the four emergent K-means engagement clusters (Table 1). Pearson’s Chi-squared tests of independence demonstrated that all baseline characteristics varied significantly across the four behavioral phenotypes (p < 0.001). Notably, cluster 2 (high health coach engagement) exhibited the highest concentration of female patients (94%), individuals presenting with a baseline BMI ≥40 kg/m2 (22%), and patients managing a complex comorbidity burden of 3 or more baseline conditions (25%). Conversely, cluster 1 (non-engaged) featured a higher baseline proportion of male patients (27%) and individuals in lower BMI tiers compared to the highly interactive cohorts.

3.1. Engagement Cluster Profiles

The unsupervised K-means clustering algorithm generated distinct behavioral phenotypes based on the four continuous tracking metrics. The average behavioral profiles mapping to each cluster are detailed in Table 2.
  • Cluster 1 (n = 1,772): non-engaged patients, displaying pervasively low interaction across all app modalities, including a baseline weekly app engagement rate of 30.0% and nominal tracking or messaging behaviors.
  • Cluster 2 (n = 1,141): high health coach patients, demonstrating high baseline tracking metrics combined with a high propensity to communicate through the human coaching channel (50.9% of weeks) over automated alternatives (22.4%).
  • Cluster 3 (n = 1,472): high JuneBot group, displaying identical app and tracking frequencies to Cluster 2, but substituting human engagement for dominant interaction with the automated AI interface (46.2% of weeks).
  • Cluster 4 (n = 5,085): passive self-trackers, representing the largest cohort. These users demonstrated high baseline app utilization (86.7%) and consistent biometric weight logging (78.8%) but remained largely silent across human and automated messaging modules.
A post-hoc principal component projection confirmed the mathematical validity of the selected 4-cluster solution, with the first two continuous dimensions explaining 81.2% of the total behavioral variance across the four application engagement metrics. The spatial separation and geometric alignment of the four distinct phenotypes are visually projected via principal component analysis in Figure 1.
Cluster-robustness analyses supported the 4-cluster engagement solution (Table 3). The gap statistic showed a clear improvement in fit up to K ≈ 4–5 with a plateau thereafter, and bootstrap resampling (B = 100) yielded high mean Jaccard similarities for all four clusters (0.99, 0.99, 0.99, 0.98), indicating excellent stability under resampling (Supplementary Figure S1, Tables S1–S2).

Clinical Weight Loss Outcomes

Descriptive evaluation of the primary endpoint revealed distinct clinical trajectories across the engagement groups (Table 3). The unadjusted mean weight loss at 6 months was lowest in the non-engaged baseline cohort (11.7%; ±7.18%). Passive self-trackers achieved an unadjusted mean weight loss of 15.5% (±6.63%). The highest absolute unadjusted weight reductions were observed in both conversational cohorts, with the high health coaching group (16.4%; ±6.38%) and the High JuneBot group (16.4%; ±6.35%) recording identical mean clinical outcomes.

Unadjusted Pairwise Efficacy Comparisons (ANOVA/Tukey HSD)

A one-way ANOVA confirmed highly significant variances in unadjusted 6-month weight loss across the engagement groups (p < 0.001). Post-hoc pairwise comparisons using Tukey’s Honest Significant Difference (HSD) test demonstrated a definitive weight loss advantage for all active engagement phenotypes over the baseline population.
Compared to the non-engaged group, mean weight loss was significantly higher in the passive self-trackers (+3.75%, 95% CI [3.27, 4.22], p < 0.001), the high health coaching group (+4.63%, 95\%, CI [3.98, 5.28], p < 0.001), and the high JuneBot group (+4.63%, 95% CI [4.03, 5.23], p < 0.001). Pairwise testing revealed no statistically significant difference in mean weight loss when comparing the high JuneBot group directly against the high health coaching group ( +0.004%, 95% CI [-0.67, 0.68], p = 0.999). Both conversational phenotypes achieved significantly higher weight loss than passive tracking alone (p < 0.001).

Multivariate Regression Analysis

After controlling for all baseline demographic markers, intake clinical factors, medication dosages, and mid-program clinical tracking variables within the fully adjusted ordinary least squares linear regression model (F = 179, Adjusted R2 = 0.3606, p < 0.001), the engagement prototypes remained strong and highly significant independent predictors of weight loss success.
With the non-engaged group (Cluster 1) assigned as the reference baseline, the adjusted incremental weight loss metrics for each phenotype were:
  • High JuneBot: β = 2.28%, SE = 0.24, t = 9.41, p < 0.001
  • High health coaching: β = 2.17%, SE = 0.26, t = 8.37, p < 0.001
  • Passive self-trackers: β = 1.67%, SE = 0.20, t = 8.24, p < 0.001
Among the adjusted demographic variables, younger age categories were significantly associated with increased weight loss compared to the reference baseline (Age 45–59), with the largest effect observed in patients under 30 (β = 0.65%, SE = 0.19, p < 0.001). Male patients achieved significantly less weight loss than female patients (β = -0.71%, SE = 0.16, p < 0.001). Regarding baseline clinical markers, patients categorized as overweight (β = -2.19%, SE = 0.17, p < 0.001) or presenting with a BMI ≥40 kg/m (β = -0.90%, SE = 0.17, p < 0.001) achieved significantly less weight loss compared to the baseline (BMI 30–34.99) reference group. Baseline comorbidity counts and a history of previous GLP-1 RA use did not significantly impact weight loss outcomes (p > 0.05).
For maximum achieved medication doses, only the 15mg maintenance dose cohort reached statistical significance, demonstrating lower relative weight loss compared to the reference group (β = -2.92%, SE = 1.34, p = 0.029). The reporting of any self-reported side effects was not a statistically significant predictor of 6-month weight loss (β = 0.04%, SE = 0.13, p = 0.738).
All month-1 tracking count categories above the reference baseline (1 submission) were negatively associated with percentage weight loss, with the most pronounced decrease observed in the 11–15 log category (β = -2.19%, SE = 0.36, p < 0.001). Mid-program clinical milestones and outcomes had the largest absolute impact on the model. Achieving clinically significant weight loss (≥ 5\%) within the first 3 months of the program was the single largest positive predictor of final 6-month weight loss (β = 7.48%, SE = 0.17, p < 0.001). Attaining a healthy BMI (25 < kg/m) or a user-defined target weight within 6 months was also a strong positive predictor (β = 3.91\%, SE = 0.14, p < 0.001). Conversely, experiencing a documented weight loss plateau during the program was strongly associated with a reduction in total 6-month weight loss (β = -2.64%, SE = 0.12, p < 0.001). The complete set of unstandardized coefficients, standard errors, t-values, and precise p-values for the regression model are detailed in Supplementary Table S3.

4. Discussion

This retrospective per-protocol analysis of 9,470 tirzepatide-treated patients in an app-mediated digital weight-loss service identified four naturally occurring digital engagement phenotypes and demonstrated that all active engagement profiles - passive self-tracking, high human health-coach engagement, and high JuneBot engagement - were associated with greater 6-month percentage weight loss than a non-engaged phenotype. After adjustment for a broad set of demographic, clinical, medication, and mid-program behavioral covariates, the conversational phenotypes (high health coach and high JuneBot) retained statistically significant and similar incremental effects versus the non-engaged reference group, while passive self-tracking showed a smaller but still significant association. Unadjusted comparisons yielded the same directional pattern, with clinically meaningful absolute differences between the non-engaged cohort and each active engagement group.
Prior studies exploring engagement and outcomes in digital weight-loss interventions have relied predominantly on binary engagement constructs or single-modality measures, limiting insight into how mixed or modality-specific behaviors relate to effectiveness. Johnson et al. [10] and similar work [11] advanced the literature by highlighting an engagement–effectiveness relationship but did not capture the multidimensional engagement patterns our clustering approach revealed. Our results extend this prior work by showing that distinct engagement archetypes, not just engaged versus non-engaged, map to different magnitudes of benefit and that automated conversational engagement can parallel human coaching when implemented within a tightly governed clinical ecosystem.
This study’s findings suggest that, within a real-world medicated DWLS, different modes of digital engagement are associated with differential weight-loss outcomes. The similarity in adjusted effect sizes between high JuneBot engagement and high human health-coach engagement is notable: a rule-based, non-agentic automated conversational tool used at scale produced outcomes statistically indistinguishable from university-qualified human coaching after controlling for baseline and interim confounders. This result aligns with the hypothesis that scalable digital conversational support can augment pharmacologic therapy [17] and may replicate some of the adherence-support and behavioral reinforcement functions traditionally ascribed to human coaching.
However, the observational and retrospective nature of the analysis prohibits causal inference. It is possible that unmeasured patient characteristics, such as intrinsic motivation, digital literacy, socioeconomic factors not captured by remoteness or subscription tier, or prior experience with weight-loss programs, drive both engagement phenotype selection and therapeutic response. Similarly, the per-protocol sampling frame intentionally restricted the analytic cohort to patients who maintained medication orders and provided a 6-month weight, which may concentrate individuals who are more motivated, have fewer adverse effects, or face fewer access barriers. Thus, the associations observed may overestimate effects that would be seen in an intention-to-treat population or in less adherent real-world cohorts.
It is also important to emphasize that a per protocol analysis fails to capture outcomes from the full cohort (i.e., patients who discontinued the program early, delayed medication orders or failed to submit data in the specified timeframe). However, previous studies have already established that 6-month attrition rates in digital chronic care services vary between roughly 40% and 49% [18], and roughly 16% and 40% real-world tirzepatide services [7,8]. The purpose of this study was to examine the relationship between program engagement and effectiveness for the subgroup of patients who adhere to a real-world medicated DWLS over 6 months.

Bridging the Gap Between the Ideal Patient and Real-World Engagement Challenges in Obesity Care

Achieving clinically meaningful weight loss (≥10% of initial body weight) is extremely challenging in overweight and obesity care. Previous studies of unmedicated interventions have revealed that roughly only 20% of individuals successfully achieve and maintain this level of weight loss for one year or longer [19]. While modern obesity medications such as tirzepatide have emerged as a promising solution to the global obesity challenge and are being increasingly prescribed, WHO and NICE stress that they are taken as a supplement to lifestyle interventions, including nutritional counseling, physical activity, and self-management support [6]. From a behavioral medicine perspective, the ideal obesity management patient is one who is intrinsically motivated and able to independently sustain health-promoting behaviors over time [20,21]. However, evidence suggests that such autonomous self-regulation is difficult to achieve, since engagement, self-monitoring and adherence commonly decline during ongoing interventions [22,23]. Our findings provide an interesting perspective on the gap between this theoretical ideal and the challenges in real-world practice. The autonomously engaged subgroup (passive self-trackers) in this study reflects the ideal patient profile envisioned by both behavioral theories and medical doctors, and their outcomes confirm that substantial weight loss (15.5%) can be achieved when engagement is sustained independently. At the same time, the comparable outcomes observed among patients supported by health coaching and AI chatbots (both 16.4%) indicate that other subsets of patients benefit from additional support to maintain engagement. Health coaching and AI-based interventions should therefore not be viewed as alternatives to intrinsic motivation, but rather as mechanisms that support patients struggling with self-engagement while achieving similar weight-loss outcomes.
Collectively, these findings suggest that sustained engagement itself may be more important than the specific source from which it originates, as autonomous engagement, health coaching, and AI-supported engagement were all associated with clinically meaningful six-month weight-loss outcomes. This means that both internal and external sources of engagement can be effective, and that interventions capable of sustaining engagement may help bridge the gap between patients who are self-motivated and those who may struggle to achieve durable weight-loss outcomes independently.

The Nuanced Role of Biometric Self-Monitoring

Interestingly, the multivariate model revealed a paradox regarding weight tracking frequencies. While the unsupervised K-means algorithm clustered highly active patients into phenotypes characterized by robust weekly tracking percentages (Clusters 2, 3, and 4 all averaged >78% weekly tracking metrics), the fully adjusted linear regression demonstrated a negative association between absolute month-1 weight entry counts and 6-month percentage weight loss. For example, logging between 11 and 15 entries was associated with a significant decrease in total weight loss compared to logging a single baseline entry (β = -2.19%, p < 0.001), while all other groups (2-5; 6-10; 16-25 and over 25 entries) lost over one percentage point less than the reference group (baseline weight only).
Previous studies have observed a strong negative trend between the same variable and program retention, which investigators interpreted as possible evidence that month-1 track count is a proxy for patient anxiety. In other words, patients who track excessively after only one dose of tirzepatide are likely anxious about the drug or program’s effectiveness and or have unreasonable expectations about early weight loss. This study’s finding adds possible nuance to the above interpretation.
This negative association must be interpreted conservatively within a per-protocol context. Because every single patient in this analysis was pre-filtered to ensure they had received a minimum of 5 tirzepatide orders in 183 days and submitted a valid 6-month weight entry, these tracking counts represent the frequency of data submission rather than broad program adherence. Clinically, this pattern may reflect a behavioral feedback loop: patients experiencing slow, stalling, or plateauing weight loss trajectories may log their weights significantly more frequently out of anxiety or heightened self-monitoring vigilance. Conversely, individuals experiencing rapid, uninhibited weight loss may feel less compelled to over-index on daily or weekly tracking mechanics, relying instead on stable programmatic habits. This underscores the limitation of viewing tracking data purely as a linear proxy for engagement quality [3]. Nevertheless, this finding, in addition to the consistent finding in previous DWLS studies [8] that excessive first month tracking is the strongest predictor of program attrition, suggests that the early tracking variable needs to be further explored.

Impact of demographic and Clinical Milestones

The multivariate analysis highlighted that mid-program clinical milestones exert the strongest absolute influence on final 6-month weight loss outcomes. Achieving clinically significant weight loss (≥ 5%) within the first 3 months emerged as the single largest positive predictor in the entire model (β = 7.48%, p < 0.001). This suggests that early biological response to tirzepatide is a critical driver of longitudinal success [24], likely serving as a powerful motivational catalyst that reinforces programmatic adherence. Furthermore, attaining a healthy BMI or a user-defined target weight within 6 months maintained a strong independent positive relationship with total percentage weight loss (β = 3.91%, p < 0.001), while experiencing a documented weight loss plateau depressed final outcomes (β = -2.64%, p < 0.001).
Demographically, younger age categories (under 30 and 30–44) and female sex were associated with small but statistically significant weight loss advantages [25], aligning with broader real-world epidemiological observations of incretin therapy cohorts. Interestingly, baseline comorbidity burden and a history of previous GLP-1 RA use did not independently influence final outcomes, suggesting that the combined DWLS and tirzepatide framework remains robust across varying patient profiles.

Public Health Implications

If reproduced in other settings and study designs, the equivalence between structured automated conversational support and human health coaching on short-term weight outcomes could have pragmatic implications for scaling comprehensive obesity care alongside GLP-1 RA access. Automated, rule-based systems can reduce the marginal cost of ongoing behavioral reinforcement [26], increase availability in regions with limited clinician capacity, and provide consistent adherence nudges without replacing essential clinical oversight. Importantly, the AI agent in this study was implemented with explicit non-medical, non-agentic guardrails and mandatory human escalation for flagged mental-health or safety concerns. These safety constraints are essential if automated systems are to be broadly deployed as adjuncts to pharmacotherapy.

Strengths and Limitations

Strengths of this study include a large, clinically relevant sample of 9,470 medication-adherent patients treated with tirzepatide within a real-world digital weight-loss service, the use of longitudinal, multidimensional engagement metrics and an unsupervised clustering approach that avoids arbitrary binary thresholds, and robust preprocessing and reproducibility steps (feature standardization, systematic K selection, multiple random starts with a fixed seed) together with a comprehensive regression adjustment set that reduces confounding from measured covariates. However, it also contained multiple limitations. Firstly, the retrospective, observational design precludes causal inference and leaves room for residual confounding by unmeasured factors such as motivation, socioeconomic status, digital literacy, and other behavioral traits that may determine both engagement and outcomes. Secondly, the per-protocol, complete-case sampling restricts the cohort to medication-adherent patients who furnished a 6-month weight, introducing selection bias and limiting external validity to all initiators or less-adherent populations. Third, handling missing engagement data via listwise deletion may bias results if data are not missing completely at random and alternative imputation approaches were not explored; forth, the primary outcome depends on in-app or Bluetooth-linked weight entries that, despite verification processes, remain subject to measurement variability and potential reporting biases. And finally, JuneBot operates under a highly restricted, rule-based, non-agentic communication framework. Therefore, these outcomes cannot be extrapolated to advanced, generative LLMs or autonomous AI clinical agents, which present entirely different safety, empathy, and operational profiles.

Implications for Future Research

Future work should pursue prospective and randomized designs to test whether routing patients to different support modalities (human coaching, structured automated conversational support, or hybrid models) causally improves outcomes and cost-effectiveness. Qualitative work exploring why patients select and sustain particular engagement patterns, would strengthen external validity and implementation relevance. Studies incorporating richer socioeconomic measures, digital-literacy assessments, and validated measures of motivation and behavioral skills could help disentangle selection effects from intervention effects. Finally, longer-term follow-up is needed to assess maintenance of weight loss, safety, and the durability of engagement phenotypes.

5. Conclusions

In a large per-protocol cohort of tirzepatide-treated patients within a commercial app-mediated DWLS, naturally occurring digital engagement phenotypes were associated with differential 6-month weight-loss outcomes. After extensive covariate adjustment, both high human health-coach engagement and high rule-based automated conversational engagement were independently associated with greater weight loss versus a non-engaged phenotype, and these two conversational modalities demonstrated equivalent adjusted effects. These results are hypothesis-generating and suggest that scalable, safety-constrained automated conversational support could complement pharmacologic obesity treatment. Definitive causal conclusions require prospective, randomized evaluation and broader population sampling.

Supplementary Materials

The following supporting information can be downloaded at the website of this paper posted on Preprints.org, Figure S1: title; Table S1: Gap statistic values and standard errors for K = 1–10 from the clusGap procedure (B = 100 reference datasets); Table S2: Mean Jaccard similarities from bootstrap cluster-stability analysis (B = 100) of the 4-cluster k-means solution. Higher values indicate greater stability under resampling; Table S3: Multivariate regression analysis;.

Author Contributions

Conceptualization, L.T., C.X., L.S. and N.A.; methodology, L.T., C.X., L.S. and J.H.;.; software, L.T., C.X., L.S..;.; validation, L.T., C.X., J.H.;.; formal analysis, L.T., C.X., L.S. and J.H.; investigation, L.T., and J.H.; resources, L.T., C.X., L.S. and N.A.;.; data curation, L.T., J.H., M.T. and L.S.; writing—original draft preparation, L.T., C.X., J.H., L.S., J.A.,., M.T. and N.A.;.; writing—review and editing, L.T., C.X., J.H., L.S., J.A.,., M.T. and N.A.;.;.; visualization, L.T., J.H., L.S.,.; supervision, L.T., M.T.; project administration, L.T., C.X.; funding acquisition. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

The study was conducted in accordance with the Declaration of Helsinki, and approved by the Institutional Review Board of Stanford University Independent Research Board. The Stanford University Independent Research Board determined that the study did not meet the definition of human subject research as defined in federal regulations 45 CFR 46.102 or 21 CFR 50.3.

Data Availability Statement

The data presented in this study are available on request from the corresponding author.

Conflicts of Interest

L.T and L.S are paid a salary at Eucalyptus (Juniper parent company). NA is paid as an advisor at Eucalyptus. C.X is a paid employee at Amigo AI. J.H., M.T., and J.A declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
ANOVA Analysis of Variance
AUD Australian Dollar
BMI Body Mass Index
CI Confidence Interval
DWLS Digital Weight Loss Service
GIP Glucose-Dependent Insulinotropic Polypeptide
GLP-1 RA Glucagon-Like Peptide-1 Receptor Agonist
HSD Honest Significant Difference (referring to Tukey’s post-hoc test)
ITT Intention-To-Treat
LLM Large Language Model
MDT Multidisciplinary Team
NICE National Institute for Health and Care Excellence
NS Not Statistically Significant
OLS Ordinary Least Squares
PCA Principal Component Analysis
PP Per-Protocol
PSM Propensity Score Matching
SD Standard Deviation
SE Standard Error
STROBE Strengthening the Reporting of Observational Studies in Epidemiology
WHO World Health Organization
WSS Within-Cluster Sum of Squares

References

  1. Abdelaal, M.; le Roux, C.W.; Docherty, N.G. Morbidity and mortality associated with obesity. Ann. Transl. Med. 2017, 5, 161. [Google Scholar] [CrossRef] [PubMed]
  2. NCD Risk Factor Collaboration (NCD-RisC). Worldwide trends in underweight and obesity from 1990 to 2022: a pooled analysis of 3,663 population-representative studies with 222 million children, adolescents, and adults. Lancet 2024, 403, 1027–1050. [Google Scholar] [PubMed]
  3. Sumithran, P.; Prendergast, L.A.; Delbridge, E.; Purcell, K.; Shulkes, A.; Kriketos, A.; Proietto, J. Long-term persistence of hormonal adaptations to weight loss. N. Engl. J. Med. 2011, 365, 1597–1604. [Google Scholar] [CrossRef] [PubMed]
  4. Moiz, A.; Filion, K.B.; Tsoukas, M.A.; Yu, O.H.Y.; Peters, T.M.; Eisenberg, M.J. Mechanisms of GLP-1 receptor agonist-induced weight loss: a review of central and peripheral pathways in appetite and energy regulation. Am. J. Med. 2025, 138, 934–940. [Google Scholar] [CrossRef] [PubMed]
  5. Jastreboff, A.M.; Aronne, L.J.; Ahmad, N.N.; Wharton, S.; Connery, L.; Alves, B.; Kiyosue, A.; Zhang, S.; Liu, B.; Bunck, M.C.; SURMOUNT-1 Investigators. Tirzepatide once weekly for the treatment of obesity. N. Engl. J. Med. 2022, 387, 205–216. [Google Scholar] [PubMed]
  6. National Institute for Health and Care Excellence. Tirzepatide for managing overweight and obesity. Technology appraisal guidance TA1026. Available online: https://www.nice.org.uk/guidance/ta1026 (accessed on 2 June 2026).
  7. Hankosky, E.; Chinthammit, C.; Meeks, A.; Huang, A.; Ward, J.; Mojdami, D.; Gibble, T. Real-world use and effectiveness of tirzepatide among individuals without type 2 diabetes: Results from the Optum Market Clarity database. Diabetes Obes. Metab. 2025, 27, 2810–2821. [Google Scholar] [PubMed]
  8. Talay, L.; Hom, J.; Scott, T.; Ahuja, N. Effectiveness and adherence in a tirzepatide-supported digital weight-loss programme in Australia: A real-world observational study. Diabetes Obes. Metab. 2026, 28, 2835–2848. [Google Scholar] [PubMed]
  9. Eysenbach, G. The law of attrition. J. Med. Internet Res. 2005, 7, e11. [Google Scholar] [CrossRef] [PubMed]
  10. Johnson, H.; Huang, D.; Liu, V.; Al Ammouri, M.; Jacobs, C.; El-Osta, A. Impact of digital engagement on weight loss outcomes in obesity management among individuals using GLP-1 and dual GLP-1/GIP receptor agonist therapy: retrospective cohort service evaluation study. J. Med. Internet Res. 2025, 27, e69466. [Google Scholar] [CrossRef] [PubMed]
  11. Lehmann, M.; Jones, L.; Schirmann, F. App engagement as a predictor of weight loss in blended-care interventions: retrospective observational study using large-scale real-world data. J. Med. Internet Res. 2024, 26, e45469. [Google Scholar] [PubMed]
  12. von Elm, E.; Altman, D.G.; Egger, M.; Pocock, S.J.; Gøtzsche, P.C.; Vandenbroucke, J.P.; STROBE Initiative. The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement: guidelines for reporting observational studies. Lancet 2007, 370, 1453–1457. [Google Scholar] [CrossRef] [PubMed]
  13. Eli Lilly and Company. Mounjaro KwikPen — Summary of Product Characteristics. electronic Medicines Compendium (emc). Available online: https://www.medicines.org.uk/emc/product/15481/smpc (accessed on 2 June 2026).
  14. Talay, L.; Lagesen, L.; Yip, A.; Vickers, M.; Ahuja, N. ChatGPT-4o and o1 Preview as Dietary Support Tools in a Real-World Medicated Obesity Program: A Prospective Comparative Analysis. Healthcare 2025, 13, 647. [Google Scholar] [PubMed]
  15. Hartigan, J.A.; Wong, M.A. Algorithm AS 136: A K-means clustering algorithm. J. R. Stat. Soc. Ser. C Appl. Stat. 1979, 28, 100–108. [Google Scholar] [CrossRef]
  16. Rousseeuw, P.J. Silhouettes: a graphical aid to the interpretation and validation of cluster analysis. J. Comput. Appl. Math. 1987, 20, 53–65. [Google Scholar] [CrossRef]
  17. Laranjo, L.; Dunn, A.G.; Tong, H.L.; Bau, A.B.; Gardo, J.; Gandani, C.; Cocos, A. Conversational agents in healthcare: a systematic review. J. Am. Med. Inform. Assoc. 2018, 25, 1248–1258. [Google Scholar] [CrossRef] [PubMed]
  18. Meyerowitz-Katz, G.; Ravi, S.; Arnolda, L.; Feng, X.; Maberly, G.; Astell-Burt, T. Rates of attrition and dropout in app-based interventions for chronic disease: systematic review and meta-analysis. J. Med. Internet Res. 2020, 22, e20283. [Google Scholar] [CrossRef] [PubMed]
  19. Wing, R.R.; Phelan, S. Long-term weight loss maintenance. Am. J. Clin. Nutr. 2005, 82, 222S–225S. [Google Scholar] [CrossRef] [PubMed]
  20. Patrick, H.; Williams, G.C. Self-determination theory: its application to health behavior and complementarity with motivational interviewing. Int. J. Behav. Nutr. Phys. Act. 2012, 9, 18. [Google Scholar] [CrossRef] [PubMed]
  21. Teixeira, P.J.; Silva, M.N.; Mata, J.; Palmeira, L.A.; Markland, D. Motivation, self-determination, and long-term weight control. Int. J. Behav. Nutr. Phys. Act. 2012, 9, 22. [Google Scholar] [PubMed]
  22. Greaves, C.J.; Sheppard, K.E.; Abraham, C.; Hardeman, W.; Roden, M.; Evans, P.H.; Schwarz, P.; IMAGE Study Group. Systematic review of reviews of intervention components associated with increased effectiveness in dietary and physical activity interventions. BMC Public Health 2011, 11, 119. [Google Scholar] [PubMed]
  23. Middleton, K.R.; Anton, S.D.; Perri, M.G. Long-term adherence to health behavior change. Am. J. Lifestyle Med. 2013, 7, 395–404. [Google Scholar] [CrossRef] [PubMed]
  24. Nackers, L.M.; Ross, K.M.; Perri, M.G. The association between rate of initial weight loss and long-term success in obesity treatment: does slow and steady win the race? Int. J. Behav. Med. 2010, 17, 161–167. [Google Scholar] [CrossRef] [PubMed]
  25. Yang, Y.; He, L.; Han, S.; Lin, I.; Wang, M. Sex differences in the efficacy of glucagon-like peptide-1 receptor agonists for weight reduction: a systematic review and meta-analysis. J. Diabetes 2025, 17, e70063. [Google Scholar] [PubMed]
  26. Krishnan, A.; Finkelstein, E.A.; Levine, E.; Foley, P.; Askew, S.; Steinberg, D.; Bennett, G.G. A digital behavioral weight gain prevention intervention in primary care practice: cost and cost-effectiveness analysis. J. Med. Internet Res. 2019, 21, e12201. [Google Scholar] [CrossRef] [PubMed]
Figure 1. Two-dimensional principal component analysis projection of patient engagement phenotypes. Shaded convex hulls visually isolate the geometric boundaries of the four emergent K-means clusters. The horizontal and vertical axes represent the first two dominant principal components extracted from the behavioral metrics matrix.
Figure 1. Two-dimensional principal component analysis projection of patient engagement phenotypes. Shaded convex hulls visually isolate the geometric boundaries of the four emergent K-means clusters. The horizontal and vertical axes represent the first two dominant principal components extracted from the behavioral metrics matrix.
Preprints 220374 g001
Table 1. Baseline patient characteristics stratified by engagement cluster.
Table 1. Baseline patient characteristics stratified by engagement cluster.
Variable Cluster 1
Non-engaged
N=1,772
Cluster 2
High health coach engagement
N=1,141
Cluster 3
High Junebot
engagement
N=1,472
Cluster 4
Passive self-
trackers
N=5,085
p-value
Age category <0.001
45-59 676 (38%) 457 (40%) 603 (41%) 1,984 (39%)
Under 30 210 (12%) 111 (9.6%) 153 (10%) 601 (12%)
30-44 635 (36%) 373 (33%) 418 (28%) 1,810 (35%)
60+ 251 (14%) 201 (17%) 298 (20%) 690 (13%)
Sex at birth <0.001
Female 1,299 (73%) 1,068 (94%) 1,305 (89%) 4,200 (83%)
Male 473 (27%) 73 (6.4%) 167 (11%) 885 (17%)
BMI category (kg/m2) <0.001
<30 501 (28%) 189 (17%) 205 (14%) 947 (19%)
30-34.99 761 (43%) 439 (38%) 570 (39%) 2,036 (40%)
35-39.99 290 (16%) 267 (23%) 346 (24%) 1,115 (22%)
40 and over 220 (12%) 246 (22%) 351 (24%) 987 (19%)
Total comorbidities <0.001
0 986 (56%) 364 (32%) 553 (38%) 2,197 (43%)
1 388 (22%) 271 (24%) 361 (25%) 1,306 (26%)
2 209 (12%) 221 (19%) 255 (17%) 820 (16%)
3+ 189 (11%) 285 (25%) 303 (21%) 762 (15%)
Previous GLP-1 RA use 16 (0.9%) 3 (0.3%) 1 (<0.1%) 15 (0.3%) <0.001
Table 2. Mean feature percentages by emergent engagement cluster.
Table 2. Mean feature percentages by emergent engagement cluster.
Behavioral Metric Cluster 1
Non-engaged
n=1,772
Cluster 2
High health coaching
n=1,141
Cluster 3
High JuneBot
n=1,472
Cluster 4
Passive self-trackers
n=5,085
Weekly App Engagement (%) 30.0% 94.5% 95.0% 86.7%
Weekly Health Coach Messaged (%) 1.97% 50.9% 5.43% 5.25%
Weekly JuneBot Messaged (%) 5.46% 22.4% 46.2% 12.7%
Weekly Weight Track (%) 19.6% 86.2% 87.2% 78.8%
Table 3. Engagement cluster distribution.
Table 3. Engagement cluster distribution.
Cluster Assignment Patient Count (n) Mean Weight Loss (%) Standard Deviation (SD)
Cluster 1: Non-engaged 1,772 11.7% 7.18%
Cluster 4: Passive self-trackers 5,085 15.5% 6.63%
Cluster 2: High health coaching 1,141 16.4% 6.38%
Cluster 3: High JuneBot 1,472 16.4% 6.35%
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings