4. Discussion
In this study, we modeled the phenotype of multiple sclerosis as a non-negative mixture of additive symptom modules derived from empirical clinical-note data. The motivation for this approach is grounded in the long-recognized polysymptomatic nature of MS [
6]. MS is defined as an autoimmune demyelinating disease characterized by lesions disseminated in time and space [
26]. Kurtzke’s Expanded Disability Status Scale (EDSS) and Functional Systems Scores (FSS) [
27,
28,
29] codified this clinical heterogeneity by distributing disability across pyramidal, cerebellar, brainstem, sensory, bowel and bladder, visual, and cerebral systems [
29]. Although the FSS provides a clinically intuitive multi-axis framework, it assigned signs and symptoms to functional systems using clinician-defined categories. In contrast, the modules identified here were derived from data-driven co-occurrence patterns among note-derived phenotype features.
This distinction is important because MS phenotype is not stereotyped. Individual patients accumulate different combinations of sensory, motor, cognitive, and autonomic impairments over time, reflecting demyelinating lesions distributed across different CNS locations and disease episodes.
The empirical feature distributions in this cohort supported this view. Most patients had multiple phenotype features documented across their available notes, and the most frequent features included gait impairment, pain, weakness, and sensory symptoms (
Figure 1). Thus, the input data did not consist of isolated single-symptom presentations, but of broad patient-level phenotype profiles accumulated over time. This polysymptomatic structure provided the rationale for modeling MS phenotype as a mixture of latent modules rather than as a set of mutually exclusive clinical categories.
The observed phenotype can be viewed as the additive superposition of system-specific impairments rather than as progression along a single severity axis. NMF provides a natural framework for representing each patient as a non-negative mixture of latent symptom modules. We examined 2-, 3-, 4-, and 5-module solutions and selected the 4-module solution because it provided the clearest clinical interpretation while maintaining reasonable reconstruction error. This choice should be viewed as a balance between model fit and interpretability rather than as evidence for a uniquely optimal number of biological subtypes.
Our four-module solution identified clinically coherent modules corresponding to sensory-visual-pain, ataxic-spastic-falls, cognitive-psychologic-fatigue, and autonomic-bladder-bowel patterns (
Table 2). These empirically derived modules broadly resembled familiar neurological functional domains, including sensory/visual pathways, pyramidal-cerebellar motor systems, distributed cognitive-affective networks, and autonomic/spinal cord pathways with supraspinal modulation.
The most practical summary of this representation is two-dimensional. Each patient can be described by the identity of the dominant module and by the effective number of modules contributing to the phenotype (
Figure 3). The dominant module identifies the principal clinical direction of the phenotype, whereas the effective number of modules quantifies the degree of phenotypic admixture. This distinction is important because two patients may share the same dominant module but differ substantially in whether their phenotype is relatively pure or broadly distributed across multiple modules.
When the four non-negative module weights were normalized to sum to 100%, patients could be visualized within a tetrahedral, or 3-simplex, space. This visualization provides an intuitive geometry for module composition: vertices represent pure module profiles, edges represent two-module mixtures, faces represent three-module mixtures, and interior points represent admixtures of all four modules. The tetrahedron should be interpreted as a convenient visualization of normalized four-module composition rather than as independent evidence for four discrete patient clusters. In this study, most patients occupied admixed regions of the interior phenotype space, although a subset showed relatively pure or strongly dominant patterns near the vertices. These purer phenotypes were most often sensory-visual-pain dominant, whereas relatively pure ataxic-spastic-falls, cognitive-psychologic-fatigue, and autonomic-bladder-bowel phenotypes were less common. Only seven patients occupied true vertices of the tetrahedron, corresponding to pure sensory-visual-pain and pure cognitive-psychologic-fatigue profiles (
Table 3).
The presence of relatively pure vertex-adjacent phenotypes invites comparison with archetypal analysis, a method designed to identify extreme representative profiles on the convex hull of the data. Archetypal approaches have been used to define clinically interpretable extreme disease states from longitudinal clinical data. For example, Trescato et al. [
30] derived archetypal phenotypes in amyotrophic lateral sclerosis and used them to model disease-progression trajectories, illustrating the value of low-dimensional phenotype representations for studying clinical heterogeneity. The present analysis differs in an important respect: whereas archetypal analysis identifies extreme representative disease states, non-negative matrix factorization identifies additive latent phenotype modules that may coexist within the same patient. Thus, the vertices of the tetrahedron should not be interpreted as archetypes discovered by archetypal analysis, but rather as limiting cases of a normalized four-module composition.
The predominance of admixed phenotypes is consistent with clinical experience: as MS evolves, patients often accumulate deficits across multiple functional systems rather than remaining confined to a single domain [
20]. At the same time, the existence of relatively module-pure patients suggests that some individuals may have phenotypes dominated by particular functional systems. Whether these purer patterns reflect biologically meaningful subtypes, stochastic lesion distribution, differences in disease stage, or differences in documentation remains uncertain. If module-dominant groups show distinct MRI lesion topography, regional atrophy, immunologic signatures, fluid biomarkers, genetic risk profiles, or treatment responses, this would support the biological relevance of the module structure. Conversely, if biological markers are similar across module-dominant groups, the modules may primarily reflect the geometry of CNS functional organization under a common pathogenic process.
The principal practical value of this approach is that it converts heterogeneous clinical-note phenotypes into quantitative patient-level module scores. These scores have three components: 1) the identity of the dominant symptom module, 2) the magnitude of the dominant module, and 3) the effective number of modules. These scores may serve as predictors, outcomes, or stratification variables in future models of relapse risk, disability progression, treatment response, MRI lesion distribution, regional atrophy, or biomarker profiles.
Conventional one-dimensional severity scales do not explicitly capture both module dominance and degree of admixture. Even multi-axis clinical instruments such as the Kurtzke Functional Systems Scores provide domain-specific scores but do not directly yield a normalized measure of phenotypic admixture such as the effective number of modules. In this sense, the four-module representation has a regularization-like effect: it trades some feature-level granularity for a more stable and interpretable description of recurring phenotypic patterns.
These findings do not establish distinct MS disease entities. Rather, they provide a quantitative framework for representing MS phenotypic diversity and for generating testable hypotheses about clinically or biologically meaningful subgroups. Future studies should determine whether these NMF-derived module scores are reproducible across cohorts and whether they predict independent measures of disease activity, progression, treatment response, or biological mechanism.
4.1. Relation to Prior MS Phenotype Studies
Prior MS phenotyping studies generally begin with a rectangular patient-by-feature matrix and follow one of two complementary strategies (
Table 4). One strategy uses phenotype features to group patients into a small number of clinically interpretable classes or clusters, making the patient the primary object of analysis. The other examines the structure of the phenotype features themselves, reducing correlated features into a smaller set of latent dimensions, factors, or communities. Thus, while both approaches analyze the same
m ×
n patient-feature matrix, they differ in whether the focus is on categorizing patients or identifying underlying feature structures.
Shahrbanian et al. [
31] used hierarchical clustering to examine relationships among nine MS symptom variables and identified three broad variable clusters: cognitive-emotional symptoms, pain-fatigue-sleep symptoms, and spasticity-balance symptoms. Although clinically intuitive, this analysis produced a relatively coarse symptom grouping rather than a patient-level compositional phenotype model. Gulick [
32] used factor analysis to reduce 22 phenotypic features to 5 factors: skeletal-motor (weakness, ataxia, spasticity), elimination (bowel and bladder), emotional (depression and anxiety), sensory, and brainstem/visual.
Ajdacic-Gross et al. [
33] examined 20 symptom phenotypes in 1942 multiple sclerosis subjects and derived six distinct classes through latent class analysis, which included multiple symptoms (14.1%), gait-balance (13.5%), fatigue-weakness (11.7%), gait-paralysis (23.9%), vision (21.3%), and paresthesia (15.4%). De Nadai et al. [
34] applied latent profile analysis to 11 Multiple Sclerosis Patient-Reported Outcome (MS-PRO) scales, which captured patient-reported impairment across mobility, hand function, vision, fatigue, cognition, bladder/bowel function, sensory symptoms, spasticity, pain, depression, and tremor/coordination domains. They derived 9 distinct subtypes including normal functioning, fatigue, sensory, somatic, somatic plus cognitive, severe disability, poor mobility, moderate disability, and physical symptoms. Howlett-Prieto et al. [
35] applied the Louvain algorithm to detect network communities in a 113 patient by 17 symptom feature array and found five unipartite communities (pain, fatigue, cognitive, sensory, and gait-weakness-hypertonia).
Taken together, these studies support the view that MS phenotype is multidimensional and cannot be adequately represented by a single severity axis. However, they differ from the present study in three important respects. First, several prior studies used patient-reported outcomes or symptom-onset questionnaires, whereas the present analysis used phenotype features extracted from longitudinal clinical notes. Second, clustering, latent class analysis, and latent profile analysis assign patients to discrete groups, whereas NMF represents each patient as a mixture of additive symptom modules. Third, prior cluster or class solutions generally produced mutually exclusive patient categories, while the present approach explicitly quantifies admixture through normalized module weights, module dominance, entropy, and the effective number of modules. Thus, the present findings should be viewed not as a replacement for prior MS phenotyping studies, but as a complementary representation of MS phenotype as a continuous, modular, and compositional space.
4.2. Limitations and Future Directions
This study has several limitations that also define important directions for future work. First, phenotyping was performed by a large language model rather than by manual expert review, although prior work has shown near-human-level performance of the same model on related neurology phenotyping tasks [
24]. Future studies should include additional clinician validation of the automated phenotyping approach, ideally using independent note samples and multiple expert reviewers.
Second, all notes were drawn from a single urban academic safety-net medical center. As a result, the cohort may not be representative of MS populations seen in other clinical settings, particularly with respect to disease severity, disability burden, socioeconomic factors, access to care, and documentation practices. We also did not account for potential differences in documentation style across physicians, clinics, or health systems. The study lacked an independent internal or external validation cohort and should therefore be considered proof-of-concept rather than definitive. Future work should test whether the same four-module structure is reproducible in additional MS cohorts and across institutions.
Third, phenotype features were aggregated across all available notes for each patient to reduce variability in single-note documentation, but we did not explicitly model the number or timing of notes, nor did we examine the temporal evolution of module weights. Longitudinal analyses could determine whether module scores change over time with relapse, remission, progression, or disease-modifying therapy, and whether shifts in module composition provide clinically useful information beyond static phenotype burden.
In addition, phenotype features were scored on a binary basis as absent or present. We did not attempt to grade severity, frequency, laterality, anatomical distribution, or functional impact. As a result, a mild symptom and a severe symptom could contribute equally to the feature matrix. The feature set was intentionally limited to commonly documented MS phenotypes, and some less common manifestations, such as tinnitus, hearing loss, seizures, or other episodic symptoms, were not included. Future studies could extend this framework by incorporating severity grading, anatomical localization, temporal dynamics, and a broader set of MS-related features. As a future enhancement, phenotype features could be mapped to concepts in standard ontologies such as the Human Phenotype Ontology (HPO) or SNOMED CT.
Fourth, we focused on non-negative matrix factorization (NMF) and did not formally compare the four-module solution with alternative dimensionality-reduction, matrix-factorization, or clustering methods, such as principal components analysis, independent components analysis, factor analysis, latent class analysis, latent Dirichlet allocation/topic modeling, archetypal analysis, hierarchical or consensus clustering, Gaussian mixture models, biclustering, or graph/community-detection approaches. NMF was selected because its non-negative, parts-based representation is well suited to modeling additive symptom burden and produces clinically interpretable modules. As with all dimensionality-reduction methods, NMF necessarily involves some loss of feature-level information. In the present analysis, 17 observed phenotype features were reduced to four latent modules, and this compression was reflected in the nonzero relative reconstruction error (
Table 1). Relative reconstruction error is strongly dependent on the structure of the input matrix and should not be compared directly across domains such as image matrices and sparse binary clinical phenotype matrices [
36]. We therefore interpreted the four-module solution as a clinically interpretable approximation of recurring phenotype patterns rather than as a complete reconstruction of the original patient-by-feature matrix. Future studies should determine whether similar phenotype structure is recovered using alternative methods and should compare solutions using stability, reconstruction error, interpretability, and external clinical validity.
We also did not attempt to reconstruct Kurtzke Extended Disability Status Scale (EDSS) or Functional Systems Scores (FSS) from the available notes [
27,
28,
29,
37]. The notes analyzed in this study were routine clinical-care documents rather than standardized research or clinical-trial assessments, and they did not consistently contain the graded severity, functional-system scoring, ambulation distance, or assistance-level information required for valid EDSS or FSS assignment. Future studies linking note-derived module scores to prospectively collected or carefully curated EDSS and FSS measures could determine whether module dominance and phenotypic admixture provide information complementary to established MS disability scales.
Furthermore, phenotype extraction was dependent on the content and quality of clinical documentation. Physicians vary in documentation style, completeness, and attention to neurological detail, and these differences may influence the observed phenotype matrix independently of the patient’s true clinical state. Thus, some variation in phenotype burden may reflect documentation practices rather than biological or clinical differences among patients. Although physician-level documentation quality could potentially be modeled using indirect features such as note length, structure, vocabulary diversity, frequency of negative findings, or evidence of templated text, such an analysis was outside the scope of the present study.
Finally, we did not correlate module scores with independent markers of MS disease activity or progression, such as MRI lesion burden, regional atrophy, relapse rate, disability progression, serum or CSF biomarkers such as neurofilament light chain or GFAP, or genomic, immunologic, or methylomic data. Integrating NMF-derived phenotypic modules with these biological and clinical measures will be necessary to determine whether module-dominant or highly admixed phenotypes correspond to biologically meaningful MS subtypes or clinically useful prognostic groups. Future work should assess whether these modules can predict relapse, disability progression, or treatment response, or whether modules correlate with MRI, biomarker, or genetic data [
25].