Submitted:
31 August 2026
Posted:
01 September 2026
You are already at the latest version
Abstract
Background: Symptoms overlap and heterogeneity observed in psychiatric disorders pose a challenge to accurate diagnosis and patient stratification. In low-resource settings such as Nigeria, utilizing available structured clinical data with unsupervised machine learning can reveal latent subtypes, potentially improving diagnosis and treatment. Aims: This study aims to identify meaningful psychiatric subtypes in a Nigerian cohort using clustering of mental state and patient history data. Method: Data from 664 patients with 38 clinical variables were analyzed. Mental state domain included mood, thought content, perception, and insight, while history domain covered psychosocial background, family history, substance use, and past psychiatric episodes. After preprocessing, imputation and one-hot encoding, KMeans clustering was applied independently to the two domains. The optimal cluster number (k = 4) was chosen based on the Elbow Method and silhouette scores for clinical interpretability and model robustness. Cluster stability was assessed with Adjusted Rand Index and Normalized Mutual Information. Chi-square tests evaluated cluster-diagnosis correlation. Results: Mental state clustering demonstrated high stability (ARI = 0.9944; NMI = 0.9912), while history clustering showed moderate agreement (ARI = 0.4490; NMI = 0.5465). Correlation analysis between clusters identified four patient profiles with distinct clinical features: (1) a Social Stressor profile marked by reactive symptoms to acute social triggers without prior history, likely benefiting from psychosocial support; (2) a Psychiatric Chronicity profile of episodic or long-term disorders requiring ongoing management; (3) a Minimal Risk Factor profile representing mostly first-episode cases without clear risk factors, needing careful monitoring; and (4) a Complex Multi-Risk profile involving overlapping personal, familial, and medical risks requiring comprehensive multidisciplinary care. These profiles were significantly associated with diagnosis (p < 0.001). Conclusions:Unsupervised clustering revealed clinically relevant psychiatric subtypes in an underrepresented population, highlighting opportunities for tailored interventions. Further studies should validate these findings and explore their predictive value for treatment outcomes.
Keywords:
clustering
; machine learning
; psychiatric diagnosis
; Nigeria
; subtypes
Introduction
Mental health disorders are a growing public health concern, especially in low- and middle-income countries where resources are scarce. In sub-Saharan Africa, mental health remains crucial but frequently neglected, with challenges including underfunding, a shortage of trained professionals, and pervasive stigma [1]. In Nigeria, approximately 20% of the population is estimated to experience a mental health disorder in their lifetime, yet fewer than 10% of those affected ever receive appropriate care [2]. Also, mental health services tend to be centralized in urban areas [3], leaving rural areas unattended to. As a result, rural patients often rely on traditional or faith healers for care [4], while, mental health systems in many high-income countries (HICs) are predominately anchored in biomedical services delivered by trained professionals, with limited reliance on traditional healing practises5. The shortage of professionals is dire, one recent review notes “only 250 (psychiatrists) for a population of 200 million,” and the African region averages only ~0.9 mental health worker per 100,000 people [3]. As a result, up to 80% of Nigerians with severe mental health needs receive no care [3]. This disparity reflects global inequities, as HICs typically have higher workforce density and more evenly distributed services, enabling earlier intervention and continuity of care6. These factors underscore a large and persistent gap between the burden of mental illness and the availability of adequate mental health care in Nigeria and other similar LMIC contexts.
Assessment Subjectivity
Psychiatric diagnosis still relies heavily on clinician judgment of subjective information. Standard clinical evaluations like Mental State Examination, patient and family history, structured interviews, require the clinician to interpret symptoms, behavior, and report them. However, such assessments can vary widely. For instance, trained psychiatrists show inter-rater variability on basic MSE domains, and much of this variability is attributed to individual interpretation of patient cues [5]. Therefore, diagnostic decisions often depend on a clinician’s experience and cultural perspective. In Nigeria and much of sub-Saharan Africa, cultural beliefs such as attributing mental illness to supernatural causes can further complicate interpretation [6]. Thus, the process is inherently subjective and context-dependent, which can lead to inconsistent diagnoses and delays in treatment.
Data-Driven Approaches in Psychiatry
Integrating data-driven methods into psychiatric care offers a promising way to augment traditional practice. As noted by Insel [7], psychiatry has historically lacked objective measurements, and the emergence of digital tools and analytics could change this. For example, digital phenotyping (using smartphones and sensors) provides “objective and ecological” data on patient behavior, addressing the “lack of objective measurement” that has handicapped psychiatry. Such tool have been implemented in clinical decision-making to complement structured assessments, with evidence suggesting they can mitigate clinician bias and improve diagnostic precision6/7. Moreover, machine learning techniques can uncover complex patterns in clinical data. Unsupervised learning methods like clustering do not rely on existing diagnostic labels; instead they group patients by similarity of symptoms or history. Studies suggest such data-derived subgroups can have clinical utility and clustering approaches have identified patient subtypes that cut across traditional DSM/ICD categories and even predict treatment outcomes better than categorical diagnoses [8]. In one recent study, unsupervised symptom-clustering in a community sample yielded stable patient groups that spanned diagnoses of depression, anxiety, and sleep disorders [9]. These findings imply that algorithmic clusters can reveal meaningful phenotypes overlooked by standard categories.
Aims
This study used a rich clinical dataset from a Nigerian psychiatric facility, comprising 38 structured fields on each patient (demographics, medical/psychiatric history, mental state exam findings, and final diagnosis). We applied unsupervised clustering separately to two domains of this data: current mental state indicators versus psychiatric history variables, to see if natural subgroups emerge. We then tested how these algorithmic clusters relate to the patients’ formal diagnostic labels. By analyzing clusters in each domain and their inter-relationships, we aim to uncover distinct patient profiles that will enhance diagnostic clarity and reveal patterns not captured by existing categories. Ultimately, we seek to inform practical tools such as data-informed triage or early-detection systems that could support clinicians in under-resourced settings. This work aligns with efforts to contextualize mental health innovations for African healthcare by combining local clinical data with advanced analytics which offers a multifaceted approach that could improve patient outcomes beyond what subjective assessment alone can achieve.
Methodology
Data and Preprocessing
We analyzed data from 664 Nigerian psychiatric patients, each with 38 clinical variables such as sex, demographics, history, mental state exam, diagnosis, etc. Only unique subjects were included, no repeated measures or follow-up cases were present, ensuring that all observations are independent.
To prepare the data, we performed several cleaning steps. All string fields were standardized by trimming whitespace and harmonizing synonyms. For example, insight ratings like “G,” “LA,” and “LACKING” were unified as “LACKING.” Mood descriptions like “I FEEL HAPPY,” “I AM HAPPY” were mapped to a set of common categories like HAPPY, SAD, NEUTRAL, WORRIED, RELAXED, UNKNOWN, OTHER, providing a clear and consistent framework for analysis. Occupational titles were grouped manually into broad categories “HEALTHCARE,” “EDUCATION,” “TECHNICAL,” “SERVICE,” etc., to simplify analysis of socioeconomic factors.
After harmonization, we handled missing values separately for categorical and numerical fields. Categorical gaps were filled with the mode and numerical gaps with the median, as these simple imputations preserve the central tendency without assuming a distribution.
Next, we grouped variables into two domains for analysis:
Mental State Attributes: Current clinical assessment findings included mood, affect, thought content, perception, attention, orientation, memory, judgment and insight. These reflect the patient’s present symptoms and mental status.
Historical Attributes: Background factors are the psychosocial history, family psychiatric history, premorbid personality, substance use and previous psychiatric episodes. These capture each patient’s past context and risk factors.
This split follows modern approaches in psychiatric research that combine current symptoms with historic factors. By clustering each domain, we aim to reveal distinct behavioral phenotypes, and then examine how they relate to each other and to diagnosis.
Data Imputation
Amongst the variables, 560 subjects (84.3%) had at least one missing entry and columns with excessive missing values (more that 40% missing) were removed to ensure dataset reliability. After removing these columns, the remaining missing values were imputed, categorical fields were filled with the mode and the numerical field was filled with the median. All 664 subjects had complete data after imputation, ensuring no missing values remained for analysis. Imputation was performed separately for each variable to preserve observed distributions.
Clustering Procedure and Validation
All categorical variables were one-hot encoded, and continuous features were scaled. We then applied K-Means clustering independently on each domain. K-Means was chosen for its speed and simplicity, which are well-suited to moderately large clinical datasets. It also provided a good balance of performance and interpretability. It partitions data into K groups by minimizing within-cluster variance, automatically learning cluster centers.
To select the number of clusters (K), we used two standard techniques. Firstly, the Elbow Method plots the inertia versus K and finds the “elbow” point where adding more clusters yields diminishing returns. Secondly, Silhouette Analysis computes a score (–1 to +1) reflecting how well each point lies within its cluster: scores near +1 mean dense, well-separated clusters, near 0 mean overlapping clusters, and negative means poor assignment. We tested K from 2 up to 10.
Cluster Stability
To quantify stability, for each k we repeated K-means clustering 100 times across distinct random seeds. The consistently low variance in silhouette scores across runs indicated that the solutions were robust to initialization randomness on the full dataset.
In addition, cluster validity was evaluated using the Adjusted Rand Index (ARI) and Normalized Mutual Information (NMI). These indices measure the similarity of cluster assignments across repeated runs, with values closer to 1 indicating near-identical partitions. The high mean ARI and NMI scores, with very small standard deviations, further confirmed that the cluster structures were stable and reproducible.
Dimensionality Reduction and Visualization
For visualization, we reduced the high-dimensional data to two dimensions using UMAP (Uniform Manifold Approximation and Projection). UMAP was selected for its ability to preserve both local and global data structure in complex dataset enabling better identification of subtle, clinically meaningful symptom. Unlike linear methods such as PCA, UMAP captures nonlinear relationships between variables, which is particularly relevant in psychiatric data where symptom profiles often intersect.
While UMAP is stochastic, we mitigated variability by setting a fixed random seed during visualization. The resulting 2D projections enabled interpretation of cluster structure and confirmed that the clusters identified by K-Means were well-separated, supporting the clinical validity of the behavioral subtypes.
The UMAP plots of each domain confirmed that the chosen clusters formed distinct groups rather than diffuse blobs.
Relationship Analysis Using Cross-Tabulation and Chi-Square Tests
To investigate the relationships between the identified clusters and clinical outcomes, cross-tabulations were constructed between the cluster labels and the diagnosis (DIAGN) field. Cross-tabulations provide a way to examine the frequency distribution of variables and identify patterns or associations between them. Additionally, cross-tabulations were generated between the mental state and history clusters to explore intra-patient dimensional associations.
Each cross-tabulation was followed by a Chi-Square Test of Independence to assess the statistical significance of the observed distributions. The Chi-Square Test is a non-parametric test used to determine whether there is a significant association between two categorical variables. The test outputs included the Chi-square statistic, degrees of freedom, and p-values, with a threshold of 0.05 adopted to determine significance.
Ethical Considerations
The authors assert that all procedures contributing to this work comply with the ethical standards of the relevant national and institutional committees on human experimentation and with the Helsinki Declaration of 1975, as revised in 2013. Ethical approval for this study was obtained from the Lagos University Teaching Hospital (LUTH) Health Research Ethics Committee, with NHREC: 19/12/2008a and approval number ADM/DSCST/HREC/APP/4602, dated September 21, 2021. All participants provided written informed consent before participating in the survey. All methods were carried out in accordance with relevant guidelines and regulations.
Results
Sample Characteristics and Demographics
The dataset included 664 psychiatric patients with a moderately balanced gender distribution of 377 females (56.8%) and 287 males (43.2%) and a right-skewed age distribution concentrated in the 25-35 year range, indicating predominance of young adults, reflecting commonly observed demographic patterns in clinical psychiatric population [10]. The clinical insight assessment revealed significant heterogeneity, with the majority of patients demonstrating either poor insight (n=185) or good insight (n=100), suggesting distinct subgroups with varying degrees of illness awareness and treatment engagement capacity.
Missingness and Imputation
The dateset comprised of 38 structured fields, divided into mental state examination variables and historical risk factors. Five fields (TH_POSS, INT_GFK, INT_S_A_D, INT_CAL, INT_PROV) exceeded the 40% missingness threshold and were dropped, leaving 33 fields for analysis. Across these, missingness ranged from 0-37.5%, totaling 2,583 values (11.3% of the dataset). Missing values were imputed and all 664 patients had complete data for analysis.
Cluster Analysis
Mental State Clustering
The elbow method analysis for mental state variables demonstrated optimal clustering between 2-4 clusters, with the most pronounced inflection point occurring at k=3. Silhouette analysis corroborated these findings, showing peak silhouette scores at k=2 (0.172) and k=4 (0.170), with k=3 providing reasonable balance between model complexity and interpretability.
Figure 1.
Mental State Elbow Method & Silhouette Score.

The final 4-cluster solution was selected to capture nuanced clinical presentations while maintaining clinical interpretability and enabling the differentiation of distinct symptom profiles. This methodological approach ensures that the identified clusters represent genuine clinical phenotypes rather than statistical artifacts.
Mental State Cluster Profiles and Clinical Characteristics
Cluster 0: Psychotic Presentation with Insight Impairment (n=327, 49.2%)
This dominant cluster represents patients with kempt appearance, happy mood, reactive affect, logical thought processes, and normal speech patterns, but presenting with persecutory delusions and auditory hallucinations accompanied by poor insight and content assessment. The maintenance of logical thinking processes despite active psychotic symptoms suggests either early-stage psychosis or well-managed chronic conditions with residual positive symptoms. The preserved grooming and reactive affect indicate functional capacity retention, while poor insight represents the primary treatment engagement barrier.
Cluster 1: Psychotic Presentation with Preserved Insight (n=99, 14.9%)
Characterized by kempt appearance, happy mood, but notably dull affect alongside normal thought processes and speech. These patients present with delusions and auditory hallucinations but demonstrate good insight with poor content and good perception. The dull affect suggests possible negative symptom predominance or medication-induced emotional blunting, while preserved insight indicates superior treatment alliance potential and compliance likelihood.
Cluster 2: High-Functioning Psychotic Presentation (n=178, 26.8%)
This cluster encompasses patients maintaining kempt appearance, happy mood, reactive affect, logical thought processes, and normal speech despite persecutory delusions and auditory hallucinations. Critically, they demonstrate consistently good clinical assessments across insight, content, and perception domains. This profile suggests optimal treatment response or mild symptom severity with preserved functional capacity, representing the most favorable psychotic presentation subgroup.
Cluster 3: Mood-Psychotic Comorbid Presentation (n=60, 9.0%)
The smallest cluster exhibits kempt appearance with mood-affect incongruence (happy mood but sad affect), poor thought organization, and delusional content with auditory hallucinations. Poor insight and content assessment with normal judgment characterizes their clinical profile. This presentation pattern suggests possible schizoaffective disorder or psychotic depression, indicating complex treatment requirements and potentially guarded prognosis.
![]() |
Historical Factor Clustering
Historical factor analysis revealed distinct clustering at k=4, with the elbow method showing steady decline in inertia while silhouette scores peaked at k=10 (0.46), suggesting complex underlying historical patterns.
Figure 2.
History Elbow Method & Silhouette Score.

The 4-cluster solution provided clinically interpretable groupings that aligned with established risk factor profiles in psychiatric epidemiology. This dual-clustering approach allows for comprehensive patient characterization across both current presentation and historical trajectory.
Historical Risk Factor Cluster Analysis
Cluster 0: Social Stressor Profile (n=127, 19.1%)
Patients with no psychiatric, medical, or family psychiatric history but positive social history, with normal sexual, forensic, and premorbid functioning. This profile suggests reactive psychiatric presentations triggered by psychosocial stressors, indicating potentially better prognosis with appropriate psychosocial interventions.
Cluster 1: Psychiatric Chronicity Profile (n=182, 27.4%)
The largest historical cluster encompasses patients with previous psychiatric episodes but no medical history, family psychiatric history, or current social stressors. This pattern suggests primary psychiatric disorders with episodic course, indicating need for long-term management strategies and relapse prevention focus.
Cluster 2: Minimal Risk Factor Profile (n=274, 41.3%)
Representing the dominant historical pattern, these patients demonstrate negative histories across all assessed domains. This "clean slate" phenomenon indicates either unidentified risk factors, genetic predisposition, or early-onset primary psychiatric disorders, challenging traditional risk factor models.
Cluster 3: Complex Multi-Risk Profile (n=81, 12.2%)
Patients with positive psychiatric, medical, and family psychiatric histories represent the highest-risk subgroup. This constellation suggests chronic, complex psychiatric illness with genetic loading and medical comorbidities, requiring comprehensive bio-psychosocial treatment approaches.
![]() |
Cluster Validation
To evaluate the reproducibility of the clusters, KMeans clustering ran across 10 random initializations and computed the Adjusted Rand Index (ARI) and Normalized Mutual Information (NMI) between the original and re-generated clusters.
For the Mental State clusters, both ARI (0.9944 ± 0.0015) and NMI (0.9912 ± 0.0019) indicated exceptional stability, confirming the robustness of the identified subtypes.
In contrast, History Clusters yielded moderate agreement (ARI = 0.4490 ± 0.1598, NMI = 0.5465 ± 0.1290), suggesting that historical attributes may not naturally divide into clearly defined clusters. This observation aligns with the multifactorial and overlapping nature of psychiatric risk histories, where unsupervised clustering may capture diffuse groupings rather than discrete categories.
![]() |
UMAP Projection of Cluster Structures
To visualize the spatial distribution of the clusters, Uniform Manifold Approximation and Projection (UMAP) was applied separately to the mental state and historical risk factor. For both dataset, UMAP reduced the high-dimensional features into 2 dimensions (UMAP1 and UMAP2) while preserving the local and global structure. The algorithm was configured with 15 nearest neighbors (n_neighbors = 15), a minimum of distance of 0.1 (min_dist = 0.1), and a fixed random seed (random_state = 42) to ensure reproducibility. Clusters were overlaid on the 2D embedding for visualization.
Figure 5.
UMAP Projection of mental state clusters.

The mental state clusters showed clear separation, supporting their robustness and interpretability. The clusters showed clear spatial separation along the primary UMAP1 axis, suggesting distinct groupings. While Cluster 0 occupied the highest UMAP1 values (indicating potential severity), Cluster 3 demonstrated the most distinct profile in the UMAP2 dimension. This data-driven taxonomy reveals previously uncharacterized heterogeneity in mental states that may correspond to differential treatment responses.
Figure 6.
UMAP Projection of history clusters.

In contrast, the history clusters displayed partial overlap, consistent with the moderate cluster stability observed earlier. This structured heterogeneity may reflect underlying patterns in the historical dataset.
Key Cluster Relationships
The strongest associations reveal three primary patterns:
Figure 7.
Correlation Plot on Clusters.

- Classic psychotic presentations (Mental State Cluster 0) predominantly associate with patients having no significant prior history (History Cluster 2, n=129), suggesting first-episode psychotic disorders;
- Higher functioning psychotic patients (Mental State Cluster 2) similarly cluster with clean historical profiles (History Cluster 2, n=70), indicating better prognosis cases often present without complex comorbidities.
- Mood-psychotic presentations (Mental State Cluster 3) show relatively even distribution across history clusters, suggesting diverse etiological pathways. These patterns indicate that psychotic presentations without significant psychiatric or medical history represent the predominant clinical phenotype, while complex historical profiles are associated with more varied and challenging presentations.
Cluster Overview and Profile
The analysis of the relationship between Mental State Clusters and History Clusters reveals four distinct patient profiles, each with unique characteristics and corresponding clinical implications. These profiles highlight various aspects of psychiatric conditions, ranging from reactive presentations to complex, multi-risk conditions.
Cluster 0: Social Stressor Profile (Likely Acute Stress-Reactive)
- Characteristics: These patients exhibit reactive psychiatric states, triggered primarily by social stressors (e.g., unemployment, relationship breakdowns, social isolation). They show no prior psychiatric or medical history, and have no family history of psychiatric conditions. Their mental states are most commonly associated with situational triggers.
- Clinical Implication: This group is likely to benefit from psychosocial interventions, such as counseling, social support, and environmental stabilization. The prognosis for this group is typically more favorable when the stressor is resolved.
Cluster 1: Psychiatric Chronicity Profile (Likely Long-Term / Episodic Psychiatric Disorders)
- Characteristics: Patients in this group have a personal psychiatric history but no significant family psychiatric history or medical conditions. These individuals often experience episodic psychiatric disorders, with periods of stability followed by relapses.
- Clinical Implication: Long-term management strategies are crucial, with a focus on relapse prevention and ongoing medication adherence. Support systems and psychoeducation are essential for this subgroup.
Cluster 2: Minimal Risk Factor Profile (Likely First-Episode of Acute Cases)
- Characteristics: This cluster represents the majority of patients, displaying no significant historical risk factors. Despite the absence of psychiatric, medical, or familial risk markers, these patients often experience first-time psychiatric episodes, which may point to early-onset conditions or previously undetected vulnerabilities.
- Clinical Implication: These patients require careful monitoring and early intervention, as their risk factors may not be immediately apparent but could emerge with time. Ongoing follow-up is recommended to assess potential emerging risks.
Cluster 3: Complex Multi-Risk Profile (Likely Chronic, Complex, or Severe Psychiatric Conditions)
- Characteristics: This cluster includes patients with multiple intersecting risk factors, such as a personal psychiatric history, family psychiatric history, and medical comorbidities. These patients are at the highest risk for chronic or severe mental health conditions, and their diagnosis are often complicated by both genetic predisposition and biomedical comorbidities.
- Clinical Implication: A comprehensive, multidisciplinary treatment plan is required, addressing both psychiatric and medical needs. These patients may require more intensive care and monitoring due to the complexity of their condition.
Statistical Validation of Cluster-Diagnosis Relationships
All cluster relationships demonstrate exceptionally strong statistical significance. Mental State Clusters show the strongest association with diagnostic outcomes (χ² = 419.52, p < 0.001, df = 144), indicating that symptom presentation patterns are highly predictive of final diagnosis. History Clusters also demonstrate substantial diagnostic relevance (χ² = 278.75, p < 0.001, df = 144), though with lower effect size than mental state patterns. The inter-cluster relationship between mental state and history patterns, while significant (χ² = 45.32, p < 0.001, df = 9), shows moderate strength, confirming that clinical presentation and historical factors represent distinct but related dimensions of psychiatric assessment. These findings validate the clinical utility of both symptom clustering and historical pattern recognition in diagnostic formulation.
Table 1.
Chi-Square Statistic On Cluster Relationships.
| Relationship | Chi-Square Statistic | p-value | Degrees of Freedom | Significance |
| Mental State Cluster vs Diagnosis | 419.52 | < 0.001 | 144 | Highly Significant |
| History Cluster vs Diagnosis | 278.75 | < 0.001 | 144 | Highly Significant |
| Mental State Cluster vs History Cluster | 45.32 | < 0.001 | 9 | Highly Significant |
Discussion
Our unsupervised analysis revealed distinct psychiatric patient subgroups that show the complexity and heterogeneity inherent in psychiatric populations. These clusters emerged from patients’ symptom and history data rather than from pre-assigned diagnostic labels. Pelin et al. [11] identified five diagnostically mixed clusters, ordered by severity, including one dominated by psychotic disorders with the highest symptom burden. This aligns with our finding that some clusters encompassed acute, high-symptom cases often with little prior history that is likely first-episode psychosis), while others had many psychiatric, medical, or social comorbidities consistent with chronic illness trajectories. Our data-driven subgroups cut across traditional diagnoses, reflecting latent patterns of illness. Such clustering approaches have been widely used in psychiatry to decompose inter-individual heterogeneity into more homogeneous subgroups, and can uncover novel patient stratifications. We found a statistically significant association between clusters and formal clinical diagnoses, suggesting these empiric groups have clinical validity.
Implications for Clinical Care
These findings have implications, especially in low-resource settings. Distinguishing a clean-history/high-symptom cluster (potentially first-episode cases) from an established high-risk cluster (multiple comorbidities) could guide targeted triage and treatment. Early intervention is known to improve outcomes in psychosis, and our results suggest that readily available clinical data can predict which patients follow which trajectory. Recent work emphasizes that unsupervised clustering can form a predictive framework for early intervention and personalized treatment. This means that even without expensive tests, clinicians could use routine data to flag patients for more intensive support. For example, those in a high-risk cluster might be fast-tracked to specialized care, while those in an acute-first-episode cluster might receive vigilant follow-up to ensure prompt therapy. These findings can be implicated in mental health tools like SchizoBot [12], a Nigerian chatbot delivering cognitive behavioral therapy for schizophrenia with 93.97% accuracy. While such tools address treatment delivery, these clusters can identify psychiatric subtypes to inform personalized interventions, potentially enhancing the design and cultural relevance of digital mental health tools in diverse populations.
Limitations
A key limitation is the transferability of the cluster classification. Our clusters were derived from one clinical database and depend on the variables recorded. If different patient populations or data collection methods were used, the clustering may differ. Also, unsupervised clustering involves methodological choices such as algorithm and number of clusters which affect outcomes. We tested that our clusters were reasonably robust, but alternative methods or parameter choices might yield different subgroupings.
Finally, while our clusters correlate with diagnoses, we cannot infer causality or progression from cross-sectional data. Longitudinal follow-up would be needed to see how patients move between clusters or respond to specific treatments.
Conclusion
This study demonstrates the usefulness of unsupervised machine learning in uncovering meaningful psychiatric subgroups from structured clinical data in a Nigerian setting. The identified clusters revealed clinically interpretable patterns aligned with different diagnostic and risk profiles. These results suggest that clustering techniques can complement traditional diagnostic processes by offering a finer-grained understanding of patient heterogeneity.
The approach holds promise for improving decision-making in resource-limited environments through early identification of high-risk individuals and tailoring of care pathways. Future work should validate these findings across broader populations and explore integration with longitudinal and outcome data to inform predictive diagnostics and personalized care in mental health.
Data and Code availability
The used and/or analyzed datasets during this study are available from the corresponding author upon reasonable request. The source code used in this study is available at https://github.com/ChibogwuAdaezeEdozie/Beyond-Diagnosis.git. The code repository includes all scripts necessary to reproduce the analyses and results presented in this work.
Author Contributions
Chibogwu Adaeze Edozie investigated the research area, coded, implemented, analyzed results, reviewed and summarized the literature. She also wrote and edited the original draft, managed the research activity planning and execution and the development of ideas according to the research aims. Racheal Rieninwa performed commentary and Revision. Samuel Okodeh collected the data. Isioma Aniukwu was in charge of ethics guidelines review and approval. Ephraim Nwoye performed critical review, commentary, and also managed the research activity planning and execution.
Funding
This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.
Declarations of interest
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
References
- Atewologun FA, Adigun OA, John OO, Hassan HK, Olaleke N, Ahmed MM, Ukoaka BM, Idris NB, Oso TA, Lucero-Prisno DB (2025). A comprehensive review of mental health services across selected countries in sub-Saharan Africa: assessing progress, challenges, and future direction. Discover mental health. 5. 49. 10.1007/s44192-025-00177-7.
- Ozota GO, Sabastine RN, Uduji FC, Okonkwo VC. Nigeria mental health law: Challenges and implications for mental health services. S Afr J Psychiatr. 2024 Apr 19;30:2134. PMID: 38726332; PMCID: PMC11079425. [CrossRef]
- Fadele KP, Igwe SC, Toluwalogo NO, Udokang EI, Ogaya JB, Lucero-Prisno DE 3rd. Mental health challenges in Nigeria: Bridging the gap between demand and resources. Glob Ment Health (Camb). 2024 Feb 16;11:e29. PMCID: PMC10988134. [CrossRef]
- Anjorin O, Hassan Wada Y. Impact of traditional healers in the provision of mental health services in Nigeria. Ann Med Surg (Lond). 2022 Sep 24;82:104755. PMID: 36212734; PMCID: PMC9539773. [CrossRef]
- Letsoalo DL, Ally Y, Tsabedze WF, Mapaling C. Challenging the nexus: Integrating Western psychology and African cultural beliefs in South African mental health care. Psychol Soc. 2024;66(2):45-66.
- Blaabjerg ES, Hemmingsen RPA, Høegh E, Wang AG, Gefke M, Arnfred S. Variability between psychiatrists on domains of the mental status examination. Nord J Psychiatry. 2020 May;74(4):287-292. Epub 2019 Dec 18. PMID: 31852322. [CrossRef]
- Adebayo YO, Oba-Adeoye TA, Ekpo O, Afolabi OE, Adeosun SO, Adeyemi JD, et al. Cross-cultural perspectives on mental health: Understanding variations and promoting cultural competence. World J Adv Res Rev. 2024;23(1):432-9.
- Mouchabac S, Conejero I, Lakhlifi C, Msellek I, Malandain L, Adrien V, et al. Improving clinical decision-making in psychiatry: implementation of digital phenotyping could mitigate the influence of patient’s and practitioner’s individual cognitive biases. Dialogues Clin Neurosci. 2022;23(1):52-61. [CrossRef]
- Ozota GO, Sabastine RN, Uduji FC, Okonkwo VC. Nigeria mental health law: Challenges and implications for mental health services. S Afr J Psychiatr. 2024 Apr 19;30:2134. PMID: 38726332; PMCID: PMC11079425. [CrossRef]
- Insel TR. Digital phenotyping: a global tool for psychiatry. World Psychiatry. 2018 Oct;17(3):276-277. PMID: 30192103; PMCID: PMC6127813. [CrossRef]
- Bzdok D, Meyer-Lindenberg A. Machine Learning for Precision Psychiatry: Opportunities and Challenges. Biol Psychiatry Cogn Neurosci Neuroimaging. 2018 Mar;3(3):223-230. Epub 2017 Dec 6. PMID: 29486863. [CrossRef]
- Hofman A, Lier I, Ikram MA, van Wingerden M, Luik AI. Uncovering psychiatric phenotypes using unsupervised machine learning: A data-driven symptoms approach. Eur Psychiatry. 2023 Feb 21;66(1):e27. PMID: 36804948; PMCID: PMC10044296. [CrossRef]
- Maestre-Miquel C, López-de-Andrés A, Ji Z, de Miguel-Diez J, Brocate A, Sanz-Rojo S, López-Farre A, Carabantes-Alarcon D, Jiménez-García R, Zamorano-León JJ. Gender Differences in the Prevalence of Mental Health, Psychological Distress and Psychotropic Medication Consumption in Spain: A Nationwide Population-Based Study. Int J Environ Res Public Health. 2021 Jun 11;18(12):6350. PMID: 34208274; PMCID: PMC8296165. [CrossRef]
- Pelin, H, Ising, M, Stein, F et al. Identification of transdiagnostic psychiatric disorder subtypes using unsupervised learning. Neuropsychopharmacol. 46, 1895–1905 (2021). [CrossRef]
- Nwoye EO, Adetayo J, Akanmu A, Ezeani I, Omigbodun O, Ogunyemi D, et al. SchizoBot: an artificial neural network chatbot for delivering cognitive behavioral therapy in schizophrenia. JMIR Ment Health. 2023;10:e42566. [CrossRef]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.


