Preprint
Article

This version is not peer-reviewed.

Mowat-Wilson Syndrome: Distribution of Pathogenic Variants in ZEB2 Gene Domains and Regions Presents Challenges in Genetic Counseling

Submitted:

10 July 2026

Posted:

13 July 2026

You are already at the latest version

Abstract
Background: Mowat-Wilson syndrome (MWS) is a complex neurodevelopmental and dysmorphic genetic disorder caused by heterozygous loss-of-function variants in the Zinc Finger E-Box Binding Homeobox 2 (ZEB2) gene. Methods & Objective: This comprehensive MWS study examined 301 ClinVar ZEB2 variants with 226 classified as pathogenic and 75 as likely or conflicting pathogenic including mutations, deletions, and insertions in the exonic, intronic, and splice acceptor/donor regions. The study focused on the distribution of these MWS variant sets across the ZEB2 gene, especially within its six functional domains. We used the STRING protein-protein interaction network to infer the molecular and biological functions of the ZEB2 protein. Results & Conclusions: The findings challenge the assumption that incidence of MWS pathogenicity is lower with variants located toward the C-terminal end with intact functional ZEB2 domains and revealed truncating variants near the N-terminus as not significantly different from those near C-terminus. Pathogenicity within functional domains is notably higher. The highest per-amino-acid pathogenicity was found with C-ZF type 6 Zfn, further indicate that maintaining integrity of C-terminal cluster of three ZFNs is equally crucial for this protein to bind bipartite CACCT elements. Noticeable increases are seen in pathogenicity within the SMAD-MH2-binding domain, DNA-binding/homeodomain, CtBP-interacting domain, and C-ZF Zfn types 6-7. Overall, pathogenicity in domain regions was 4.1% higher than in non-domain regions. No significant differences were observed in occurrence of frameshift, nonsense, and missense variants among these two groups of cases. More than 4% of pathogenic cases and over 10% of likely/conflicting pathogenic cases are caused by variants in intronic, splice-acceptor, or splice-donor sites, requiring further studies including clinical investigations Analysis of STRING interactions for ZEB2 identified key molecular and biological functions likely to be affected in this complex neurodevelopmental and dysmorphic genetic disorder and useful for clinical assessment.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  

1. Introduction

1.1. Background and Clinical Findings of Mowat-Wilson Syndrome

Mowat-Wilson syndrome [OMIM # 235730] [MWS] is characterized by numerous mutations and indels in the Zinc Finger E-Box-binding homeobox 2 [ZEB2] gene [HGNC:14881]. Mutations in this gene are also linked to Hirschsprung disease, a key feature of the syndrome, which occurs in 43-57% of the cases (www.omim.org).
Figure 1 shows typical facial frontal and profile views of a female with genetically confirmed Mowat-Wilson syndrome.
As depicted above, MWS is a multisystem neurodevelopmental disorder characterized by distinctive facial features, significant intellectual disability, delayed motor and speech development, and congenital anomalies, including Hirschsprung disease, epilepsy, congenital heart defects, and genitourinary abnormalities. It results from loss-of-function mutations in the ZEB2 gene, which are inherited in an autosomal dominant pattern. The ZEB2 mutation or deletion is typically de novo, with symptoms appearing during pregnancy and in the newborn period. It is a rare autosomal dominant disorder that affects both genders and has a population frequency of about 1 in 60,000 births [https://rarediseases.org/rare-diseases/mowat-wilson-syndrome]. It has been reported in many countries and among different ethnic groups [https://rarediseases.info.nih.gov/diseases/9673/mowat-wilson-syndrome] More than half of people with MWS are born with an intestinal disorder or Hirschsprung disease (HSCR), which involves severe constipation, intestinal blockage, and an enlarged colon due to defects in colonic nerve supply. Chronic constipation also occurs often even without a diagnosis of Hirschsprung disease. HSCR is characterized by issues in the nerve supply of the colon. Some affected individuals may not be diagnosed until childhood or adulthood [https://rarediseases.org; [2,5]]. Among these patients is a rare finding of asplenia, which increases the risk of infections[https://rarediseases.org; https://rarediseases.info.nih.gov/diseases/9673/mowat-wilson-syndrome]. Additionally, characteristic musculoskeletal anomalies include slender, tapering fingers, mild calcaneovalgus, long toes, pes planus, scoliosis, and congenital heart defects [6].
However, the main feature of this syndrome is neurologic and neurodevelopmental involvement [9]. The ZEB2 gene is vital in the embryological development of the neural tube (the early central nervous system) and the neural crest (the early peripheral nervous system). Patients with MWS also exhibit microcephaly, brain malformations, epilepsy, sleep disturbances, and cognitive and behavioral challenges [4,10,11,12,13,14,15,16]. It is believed that ZEB2 plays a key role in nervous system development, as ZEB2 knockout mouse embryos die early due to defects in neural tube and neural crest development. The ZEB2 gene is crucial for nervous system development, and ZEB2 haploinsufficiency leads to severe neurological problems in MWS patients. Specifically, deletions or mutations within ZEB2 cause various cranial and brain abnormalities [7,17,18,19,20,21]. Therefore, managing MWS involves treating Hirschsprung disease and other related conditions that often share clinical features with other genetic syndromes, as outlined below in Table 1.

1.2. Zinc Finger E-Box Binding Homeobox 2 Gene

The ZEB2 gene [Chromosome 2: 144,388,018-144,517,825, reverse strand] encodes a protein belonging to the Zfh1 family of two-handed zinc finger/homeodomain proteins. This protein localizes to the nucleus, where it binds the DNA sequence 5′-CACCT-3′ in various promoters to inhibit transcription by interacting with activated SMADs [45]; https://www.uniprot.org/uniprotkb/O60315/entry. It represses the transcription of E-cadherin and MEOX2 [PubMed: 16061479 & 20516212] and functions as a transcription factor in the Transforming Growth Factor β signaling pathway during early fetal development. The ZEB2 protein is also a master regulator of epithelial-to-mesenchymal transition and mediates trophoblast differentiation [1]. This protein influences the timing and conversion of neuroepithelial cells into radial glial cells during fetal development [2], and it participates in the embryonic development of the neural tube and neural crest. These processes are essential in later stages of nervous system development [3]. Melanocytes, dorsal root ganglia, cranial nerve ganglia, sympathetic ganglionic chains, and the enteric nervous system all derive from the neural crest. Even cells not originating from the neural crest, such as epithelial cells of the colon, kidneys, and skeletal muscles, express ZEB2 RNA [4]; https://www.bgee.org/gene/ENSG00000169554; https://www.Genecards.org. The ZEB2 gene’s pleiotropic expression during development causes various clinical features of MWS.

1.3. ZEB2 Gene Databases and Descriptions

The ClinVar website [https://www.ncbi.nlm.nih.gov/clinvar/] hosts submissions from clinical laboratories, research institutions, and other organizations regarding DNA sequence variants, including interpretations of pathogenicity (e.g., pathogenic, likely pathogenic, and benign) that guide diagnosis, treatment, and genetic counseling. The ClinVar electronic database is vital in medical genetics, helping clinicians, researchers, and patients understand the clinical significance of genetic variations. It also tracks consensus and disagreements to show whether multiple sources agree on a variant’s importance or if there are differing interpretations, which is essential for evaluating reliability. This site includes resources such as OMIM, GeneReviews, and ClinGen, providing broader context for variant analysis. Additionally, each ClinVar variant has a unique RefSeq accession number following a specific format that indicates the sequence type and its version: Nucleotide Accession Numbers: NM_001744.6, NC_003619.1; with a two-letter prefix, a series of digits or alphanumeric characters, and a sequence version number.
Thus, to date, the ClinVar database has documented 216 pathogenic variants of MWS within the protein-coding region of the ZEB2 gene [Supplemental Table 2], along with an additional 10 pathogenic variants located at the intronic/splice acceptor/donor sites of the ZEB2 gene, each with a unique RefSeq accession number [Supplemental Table 3]. Similarly, the ClinVar database contains 67 ZEB2 variants with likely or conflicting MWS pathogenicity [Supplemental Table 4], and 8 more variants with likely or conflicting MWS pathogenicity at the intronic/splice acceptor/donor sites of the ZEB2 gene [Supplemental Table 5]. This subgroup of individuals with likely/conflicting mild features, no malformations, or fewer facial features complicates diagnosis based solely on physical signs. Such individuals may experience mild to moderate intellectual disability rather than severe impairment. Speech abilities tend to be more developed, with some able to produce short sentences by mid-childhood. These individuals generally carry either a missense variant somewhere in the ZEB2 gene or a more typical truncating variant near the end of the ZEB2 gene, resulting in a protein with reduced function rather than complete loss [22] (https://rarediseases.org/rare-diseases/mowat-wilson-syndrome).
It should be noted that all the data generated in this in silico study pertains to the ClinVar documented 216 pathogenic variants [Supplemental Table 2], ClinVar documented 67 ZEB2 variants with likely/conflicting MWS pathogenicity [Supplemental Table 4], and 8 more ClinVar documented variants with likely/conflicting MWS pathogenicity at the intronic/splice acceptor/donor sites of the ZEB2 gene [Supplemental Table 5].

1.4. ZEB2 Protein Structure, Exons and Six Functional Protein Domains

For clarity, ZEB2 exon coordinates, the number of base pairs, and the amino acids in each exon, along with their relative lengths, are displayed in Figure 3 and detailed in Table 6.
The following Figure 2 illustrates the structure of the ZEB2 gene, including its exons and functional protein domains.
The following Figure 3 illustrates the six functional domains of the ZEB2 protein in detail.
As shown in Figure 2 and Figure 3 above, the ZEB2 protein has the following six functional domains [http:/www.Uniprot.org/Uniprot KB]:
  • Functional Domain: NuRD Interacting Motif [NIM] Domain: aa # 14-22: NuRD (nucleosome remodeling and histone deacetylation complex)-interacting motif (NIM) at the N terminus: aa #14-22.
  • Functional Domain: C2H2 Type Zinc Fingers at the N-terminal (NZF): aa # 209-333.
  • Functional Domain: SMAD-MH2 Binding Domain: aa # 437-487.
  • Functional Domain: DNA Binding/Homeodomain-like Domain: aa # 641-708.
  • Functional Domain: C-terminal binding protein (CtBP)-Interacting Domain (CID): aa #757-868.
  • C2H2 Type Zinc Fingers at the C-terminal (CZF): aa # 999-1076.
The ZEB2 protein structure and model are shown in Figure 4 below, which illustrates the AlphaFold 3D structure of the ZEB2 isoform.
Figure 3. ZEB2 protein functional domains: NIM (NuRD Interacting Motif); NZF (N-terminal Zinc Finger clusters with C2H2 types 1-4); SBD (Smad-Binding Domain); HD (Homeodomain-like Domain); CID (CtBP-Interacting Domain); CZF (C-terminal Zinc Finger clusters with C2H2 types 5-7). Modified from Parfenyev et al., 2025 [28].
Figure 3. ZEB2 protein functional domains: NIM (NuRD Interacting Motif); NZF (N-terminal Zinc Finger clusters with C2H2 types 1-4); SBD (Smad-Binding Domain); HD (Homeodomain-like Domain); CID (CtBP-Interacting Domain); CZF (C-terminal Zinc Finger clusters with C2H2 types 5-7). Modified from Parfenyev et al., 2025 [28].
Preprints 222684 g003
Figure 4. AlphaFold 3D structure of ZEB2 Isoform O60315-1, colored by confidence from Very Low to Very High. Data from AlphaFold DB version 2022-11-01, generated with the AlphaFold Monomer v2.0 pipeline [24,25].
Figure 4. AlphaFold 3D structure of ZEB2 Isoform O60315-1, colored by confidence from Very Low to Very High. Data from AlphaFold DB version 2022-11-01, generated with the AlphaFold Monomer v2.0 pipeline [24,25].
Preprints 222684 g004
Figure 5. Classical C2H2 (Cys2His2-type) coordinates with the zinc ion (cyan ball) [26].
Figure 5. Classical C2H2 (Cys2His2-type) coordinates with the zinc ion (cyan ball) [26].
Preprints 222684 g005
Figure 6. ZEB2’s NZF cluster of four C2H2 zinc fingers and CZF cluster of three zinc fingers bind as monomers to bipartite CACCT elements [27].
Figure 6. ZEB2’s NZF cluster of four C2H2 zinc fingers and CZF cluster of three zinc fingers bind as monomers to bipartite CACCT elements [27].
Preprints 222684 g006
Figure 5 below illustrates the classical C2H2 (Cys2His2-type) coordination with the zinc ion.
Exon Coordinates for Each Coding Exon, Total Number of Base Pairs and Amino Acids in Each Exon, and Relative Lengths of Each Exon [http://www.Uniprot.org/Uniprot KB].
Figure 6 below shows the ZEB2-NZF cluster with four C2H2 ZnFs and the CZF cluster with three ZnFs, both binding as monomers to bipartite CACCT elements.

1.5. Types of Mutations Encountered in the ZEB2 Gene

1.5.1. Nonsense Mutation

A nonsense mutation is a point mutation that changes a DNA codon into a premature stop codon, leading to a shortened, incomplete, and often nonfunctional protein. These mutations cause early termination of protein synthesis, and the resulting truncated proteins can cause structural or functional defects in cells, potentially leading to genetic diseases [29,30] [https://biologydictionary.net/nonsense-mutation/].

1.5.2. Frameshift Mutation

This kind of mutation involves inserting or deleting multiple bases that are not multiples of three, which shift the reading frame and often results in a nonfunctional or significantly altered protein. These mutations harm the patient by disrupting protein production and impairing protein function.
Frameshift mutations significantly alter the amino acid sequence, including polar ones, produced after the mutation. Because the sequence changes significantly, the properties of the resulting protein are also altered, greatly affecting its structure and function [31,32,33].
Frameshift mutations alter amino acids, as follows.
  • Shifted reading frame: Frameshift variants occur when nucleotides are inserted or deleted in numbers that are not multiples of three. Since the genetic code is read in triplets called codons, such insertions or deletions alter the reading frame for all subsequent codons.
  • Garbled sequence: The shift in the reading frame results in a completely new, often nonfunctional, sequence of amino acids from the mutation points onward.
  • Altered polarity: Since the codons after the mutation are entirely different, there is no guarantee that a polar amino acid will be replaced by another polar one. Instead, a polar amino acid might be substituted with a nonpolar or charged amino acid, or vice versa. This shift in polarity and other properties is a main reason for the severe functional effects of frameshift mutations.
Impact of frameshift mutations on protein structure and function
Frameshift mutations can cause significant downstream effects on the final protein.
  • Altered protein folding: The sequence of amino acids—containing polar, nonpolar, and charged residues—determines how a protein folds into its specific three-dimensional structure. If the amino acid sequence becomes scrambled, the protein is likely to fold incorrectly, often resulting in a misfolded, nonfunctional form.
  • Premature stop codons: Frameshift mutations often create a new stop codon in the shifted reading frame, causing early termination of protein production and resulting in abnormally short, likely nonfunctional proteins.
  • Altered protein interactions: The surface features of a protein, determined by exposed polar and nonpolar amino acids, are crucial for interactions with other molecules. Changing the surface can cause a frameshift mutation that disrupts proper binding to partners or substrates.
  • Altered enzymatic activity: In an enzyme, the active site is the specific region where a substrate binds. Frameshift mutations can completely render the enzyme unable to perform its catalytic functions [31,32,33].

1.5.3. In-Frame Deletions

An in-frame deletion preserves the reading frame, resulting in a shorter but potentially functional protein. This type of mutation removes DNA bases in multiples of three within a gene’s coding sequence. Such a deletion removes one or more entire codons without altering the reading frame, yielding a protein that may be shorter but still functional or partially functional. Conversely, a deletion not divisible by three causes a frameshift mutation, disrupting the reading frame and usually producing a nonfunctional protein [https://practicalhaemostasis.com/Genetics/genetics_mutational_analysis.html].
Normal Deletion: An in-frame deletion removes a sequence of bases that is a multiple of three (e.g., 3, 6, 9 bases).
  • Preserving the Frame: Because a whole number of codons is removed, the codons following the deletion are still read in their original groups of three.
  • Protein Consequence: The resulting protein will be missing the amino acid(s) specified by the chosen codon(s), but the rest of the amino acid sequence will stay unchanged.
  • Comparison to Frameshift Mutations: An in-frame deletion maintains the genetic reading frame, producing a shortened but possibly functional protein.
The following Figure 7 shows all known 3642 mutations (highlighted and colored nucleotides) within the coding exons as of July 2025.

2. Results

2.1. ClinVar Reported MWS Pathogenic and Likely/Conflicting Pathogenic Amino Acid Variants

Among the 216 MWS pathogenic cases with diverse types of mutations and indels within the coding exons, as depicted in Figure 7 above, 17 involve the same amino acid due to a mutation at any of the three bases within a given codon [Supplemental Table 2]. Therefore, only 199 of the 216 MWS pathogenic cases have unique amino acid variants involved. Out of a total of 226 MWS pathogenic cases, 10 confirmed cases are caused by base pair changes in the intronic/splice acceptor/splice donor sites [Supplemental Table 3], accounting for 4.42% of all reported MWS cases. Eight likely MWS pathogenic cases among the total of 75 were due to base pair changes in the intronic/splice acceptor/splice donor sites [Supplemental Table 3 & 5], representing 10.67% of the reported likely MWS cases. This percentage is more than double that observed among the 226 pathogenic cases, even though the sample size for the likely pathogenic cases is 3 times smaller than that for pathogenic cases. Among the 67 cases with likely or conflicting pathogenicity, 40 are missense variants.
The following Table 9 lists the 7 Missense Variants Among the 216 Cases With MWS Pathogenic Diagnosis and 40 Cases With Missense Variants Among the 67 Cases With Likely or Conflicting Pathogenic Variants.
Among the 67 cases with likely or conflicting pathogenicity within the 9 coding exons, 11 are also reported as pathogenic among the 216 pathogenic cases within the coding exons [Supplemental Table 2 & 4]. Therefore, only 56 amino acids (cases) are exclusively associated with a likely or conflicting pathogenicity determination.
Figure 9 below shows the distribution of MWS pathogenic and likely/conflicting pathogenic amino acid variants along the ZEB2 protein, from the 5′ N-terminal to the 3′ C-terminal: Exons 2-10, which contain 1,214 amino acids and have a mass of 136,447 Da.
  • Red Font: 199 Pathogenic Amino Acid Variants + 17 Repeat Amino Acids with Pathogenic Variants: 199+17=216 cases.
  • Blue Font: 56 exclusively Likely/Conflicting Pathogenic Amino Acid Variants + 11 Amino Acids with Full Pathogenicity as well as with Likely/Conflicting Pathogenicity: 56+11=67 cases.
  • Red Font with Underline: Both Pathogenic & Likely/Conflicting Pathogenic Amino Acid Variant: 10 cases.
  • Frameshift Variant
  • Nonsense Variant
  • Missense Variant
Further analysis of the data reveals the distribution of the 216 MWS pathogenic variants across the functional domains of the ZEB2 protein compared to the non-functional regions [Supplemental Table 2], and the distribution of the 67 likely MWS pathogenic variants within the functional domains versus the non-functional regions [Supplemental Table 4]. Of the 216 MWS pathogenic variants, 83 (38%) are located within the functional domains, while 133 are in the non-domain regions of the ZEB2 protein [Supplemental Table 2, Figure 10 & 12]. Similarly, of the 67 likely MWS pathogenic variants, 22 (33%) occur within functional domains, and 45 occur in non-domain regions of the protein [Supplemental Table 3; Figure 11 & 13].

2.2. In Silico Graphical Comparative Analysis

Our silico graphical comparative analysis examines the presence of MWS pathogenic amino acid variants in 226 case reports with confirmed pathogenicity [Supplemental Table 2& 3], as well as milder or conflicting pathogenic amino acid variants in 75 case reports [Supplemental Table 4 & 5]. This silico study aims to determine whether specific domains or regions are more prone to pathogenic or mildly pathogenic/mildly conflicting amino acid changes, given the scattered distribution of mutations across all nine coding exons of this gene [Figure 2 & 3]. We sought to clarify how complete haploinsufficiency of the ZEB2 protein, caused by pathogenic variants at the N-terminal end, leads to MWS, compared with variants at the C-terminal end, which typically allow near-normal protein production and result in a milder MWS phenotype [6,7,34,35].
A gradual decrease in MWS incidence with increasing protein length may indicate that longer proteins evade nonsense- or missense-mediated RNA decay, resulting in milder phenotypes. However, in this silico study of 216 cases with pathogenic variants within the coding regions and 67 cases with milder, likely, or conflicting pathogenic variants within the same regions [Figure 12 & 13], this idea has been clearly disproven in the absence of a severe MWS phenotype.
The bar diagram in Figure 10 below displays the distribution of 216 MWS pathogenic amino acid variants throughout the entire ZEB2 protein.
In Figure 11 below, a bar diagram displays the distribution of the 67 likely or conflicting MWS pathogenicity across the entire length of the ZEB2 protein.
In Figure 12 below, a bar diagram represents the distribution of MWS pathogenicity across the six functional domains of the ZEB2 protein.
Figure 13 below represents a bar diagram that shows the distribution of 66 likely or conflicting MWS pathogenic cases across the functional domains of the ZEB2 gene.

3. Discussion

3.1. Incidence of MWS Pathogenicity and Potential or Conflicting Pathogenicity Across the ZEB2 Protein and Its Six Functional Domains, Including Each of the Seven C2H2-Type Zinc Fingers (Analyzed Cohorts: n=216 Confirmed Pathogenic; n=67 Likely/Conflicting; Protein Length=1214 Amino Acids)

The bar diagram [Figure 10] clearly shows a consistent level of MWS pathogenicity across the entire length of the ZEB2 protein (aa 1-1200), except at the very tail end (aa 1201-1214). There is also an almost equal distribution of MWS pathogenic cases between the first 600 amino acids from the N-terminal end and the second set of 600 amino acids from the middle to the C-terminal end: 111 cases versus 105 cases among these 216 confirmed MWS pathogenic cases within the ZEB2 gene’s coding exons. Therefore, one could argue that, for MWS pathogenicity, whether due to N-terminal ZEB2 protein deletion or minimal C-terminal protein deletion, there is no significant difference, contradicting the common assumption and earlier claims [23].
There is a consistent trend toward higher per-residue incidence within predefined functional domains compared to non-domain residues across aa1-1200, but this difference does not reach traditional statistical significance in either cohort [Figure 10 & 12] (Domains versus Non-domains, confirmed pathogenic cohort: 83/404=0.2054 cases per residue [20.54% per aa] versus 133/796=0.1671 [16.71% per aa]; Δ=3.83 percentage points; Welch t=1.427, df≈731.2, p=0.154; two-tailed. Likely/conflicting cohort: 27/404=0.0668 cases per residue [6.68% per aa] versus 40/796=0.0503 [5.03% per aa]; Δ=1.65 points; Welch t=1.131, df≈721.2, p=0.258; two-tailed). Except at the NuRD-Interacting Motif [NIM] at aa14-22 (0 cases in both cohorts; 0/9 residues=0.00% per aa), pathogenic variants are found across the remaining annotated domains and outside of domains. Among the N-ZF zinc fingers, per-residue incidence varies, with the highest in C2H2-type 3 (confirmed: 8/23=34.78% per aa; likely/conflicting: 2/23=8.70%) and the lowest in C2H2-type 4 (confirmed: 2/24=8.33%; likely/conflicting: 1/24=4.17%). (N-ZF C2H2-type 1: aa210-234; type 2: aa241-263; type 3: aa282-304; type 4: aa310-333.)
In the confirmed pathogenic group, the per-residue incidence is higher in several functional regions, including the SMAD-MH2 binding domain (aa 437-487: 10 cases; 10/51=19.61% per aa), the DNA-binding/Homeodomain-like region (aa 641-708: 13 cases; 13/68=19.12% per aa), the CtBP-interacting domain (aa 757-859: 11 cases; 11/103=10.68% per aa), and the C-terminal zinc fingers (C-ZF C2H2-type 6: aa 1027-1049: 10 cases; 10/23=43.48% per aa; C-ZF C2H2-type 7: aa 1055-1076: 13 cases; 13/22=59.09% per aa) [Figure 10 & 12].
In the likely/conflicting (milder) cohort, enrichment mainly occurs in the C-terminal half of the protein, especially in the CtBP-interacting domain (aa 757-859: 9 cases; 9/103=8.74% per amino acid) and the C-ZF zinc fingers (C-ZF5: 3/33=9.09% per amino acid; C-ZF6: 4/23=17.39% per amino acid; C-ZF7: 4/22=18.18% per amino acid). In contrast, the SMAD-MH2-binding domain (2/51=3.92% per amino acid) and the DNA-binding/Homeodomain region (1/68=1.47% per amino acid) are at or below the Non-domain baseline [Figure 11 & 13]. Therefore, the highest per-residue incidences in the likely/conflicting cohort are found in C-ZF7 and C-ZF6, followed by C-ZF5 and the CtBP-interacting domain.
Even among the 67 likely or conflicting MWS pathogenicity variants [Figure 11], which can be considered milder cases, the distribution between the first 600 amino acids (aa1–600) and the second 600 amino acids (aa601–1200) is nearly equal (33 versus 34 cases). [A residue-level Welch t-test comparing the mean incidence in aa1–600 versus aa601–1200 shows no difference (Welch t=-0.126, df≈1197.8, p=0.900; two-tailed)], supporting the interpretation that whether a truncating variant occurs toward the N terminus or the C terminus does not correlate with a measurable shift in the overall incidence pattern in this cohort [Figure 11 & 13].
Since the integrity of both zinc finger clusters (C-ZF cluster and Z-ZF cluster) is crucial for ZEB2 to bind its target [Figure 3 & 6], losing either the N-ZF cluster or the C-ZF cluster can lead to the same level of pathogenicity, as observed in this study among the 216 MWS-confirmed cases and the 67 milder (Likely/Conflicting pathogenic) cases in this in silico analysis [27].
Furthermore, the distribution of MWS pathogenic cases across the functional domains NIM, NZF, and SBD at aa14-487—located within the first 600 amino acids—compared to those in the second half of the 600 amino acids, which include the HD, CID, and CZF domains at aa641-1076 (as shown in Figure 12), indicates significantly fewer cases at aa14-487 than at aa641-1076. This suggests that pathogenic variants are more common toward the C-terminal end of the protein, even though the full ZEB2 protein is longer in cases with pathogenic variants at the C-terminal end and thus has more intact functional domains. As shown in Figure 10 and Figure 12, the zinc finger [type 7 at aa1055-1076] in the CZF cluster has double the number of MWS pathogenic cases compared to the NZF-type 1 zinc finger at aa209-234.
The findings above challenge the assumptions and predictions that the incidence of MWS pathogenicity is lower when the pathogenic amino acid variant occurs toward the C-terminal end of the protein [6,7,34,35]. Therefore, the prediction that a pathogenic amino acid variant closer to the C-terminal end would produce a truncated protein more likely to evade RNA-mediated protein decay is not supported by this in situ study. Even if it bypasses such decay, the accumulation of a minimally truncated protein, if rendered non-functional, could lead to haploinsufficiency of the protein [35]. However, all the functional domains of this protein remain intact and unaffected if the amino acid variant occurs beyond aa 1076 at the C-terminal end [Figure 3 & 12].
Furthermore, the occurrence of pathogenic amino acid variants does not decrease toward the C-terminal end of the protein. The observation of six cases of MWS pathogenicity beyond the C-ZF cluster at aa999-1076 and extending to aa1190 [Supplemental Table 2] confirms that MWS can occur even if all the ZEB2 functional domains remain completely intact [Figure 3 and Figure 10 & 12]. This has important implications for prenatal genetic counseling.
Similarly, the 67 reports—likely conflicting—regarding MWS pathogenicity due to amino acid variants in the ZEB2 gene’s coding regions, though based on a relatively small sample size (one-third the size) [Supplemental Table 4; Figure 11 & 13], have been observed throughout the entire ZEB2 protein. These include regions beyond the C-ZF type 7 (aa1055-1076) at aa1118 and aa1179 [Supplemental Table 2 & 4; Figure 10 & 11], which has implications for prenatal genetic counseling. This indicates that an amino acid variant outside the C-ZF functional domain at the C terminus of ZEB2 can also be associated with either MWS pathogenicity or milder forms of the disease.
Among the 216 MWS pathogenic variants, there are no recorded cases between aa1191 and aa1214 (C-terminal), and among the 67 milder likely/conflicting pathogenicity variants, there are no recorded cases between aa1180 and aa1214 (C-terminal). These regions constitute about 2% and 3% of the protein at the very tail end, respectively. Similar to the 216 MWS pathogenic cases [Figure 10], the 67 milder likely/conflicting pathogenicity cases [Figure 11] show peaks in the aa51-100, aa751-800, and aa951-1100 regions. The strongest enrichment peak occurs in aa951-1100 (17 cases across 150 residues; 11.33% per amino acid) compared to the rest of aa1-1200 (4.76% per amino acid; Welch t=2.453, df≈168.6, p=0.0152; two-tailed). The aa51-100 and aa751-800 windows each contain 7 cases over 50 residues (14.00% per amino acid), higher than the rest of aa1-1200 (5.22% per amino acid), but this difference does not reach significance (p<0.05) [Welch t=1.756, df≈50.7, p=0.0850; two-tailed]. Consistent with the caption note about the tail, the extreme C-terminal tail aa1201-1214 has zero cases (0/14 residues), indicating a statistically significant depletion compared to aa1-1200 (Welch t=-8.420, df=1199, p≈1.06×10−16; two-tailed).

3.2. Functional Domains of the ZEB2 Protein: Incidences of MWS Pathogenicity Across Different Domains

3.2.1. NuRD Interacting Motif [NIM] Domain

As shown in Figure 3, the ZEB2 protein has a NuRD (nucleosome remodeling and histone deacetylation complex)-interacting motif (NIM) at the N terminus, amino acids 14-22. The nucleosome remodeling and deacetylation (NuRD) complex regulates chromatin structure. Its activity and the diversity of protein-binding at regulatory regions during cell-state transitions influence RNA polymerase II association at transcription start sites. The production of newly formed transcripts then helps establish lineage-specific transcriptional programs [36].
Except for the NuRD-Interacting Motif [NIM] (aa14-22), which has no cases in either cohort, disease-associated variants are found both within multiple annotated domains and outside of them. Consistent with the pooled domain-versus-non-domain tests above, any overall enrichment within domains is modest and not statistically significant when all domains are combined. ZEB2’s role as a transcriptional repressor involves interactions with various co-effector proteins, with co-repressor binding deemed crucial for this function, e.g., with the NuRD complex [37]. In the literature, only two patients with mild MWS have been documented, both caused by mutations affecting the NuRD-interacting motif. One involves a subtle R22G missense variant that disrupts ZEB2’s interaction with the NuRD complex [38].

3.2.2. C2H2 Type Zinc Fingers 1-4 at the N-Terminal (N-ZF)

Zinc finger proteins participate in various biological processes, such as transcriptional regulation, protein-protein interactions, post-transcriptional control, cell differentiation, epithelial-mesenchymal transition, trophoblast differentiation, and neuronal development and function [1,39,40].
As shown in Figure 3 and Figure 6, four residues (Cys107, Cys112, His125, and His129) coordinate the zinc ion for structural roles and form a ββα motif. This secondary structure supports specific interactions with binding partners, including DNA, RNA, lipids, proteins, and small molecules [26]. The ZEB2 C2H2 Znfs are separated but can engage in paired DNA-binding mechanisms [41]. As shown in Figure 4, the ZEB2 gene’s cluster of four C2H2 Znfs at the N-ZF and a cluster of three Znfs at its C-ZF bind tandemly as monomers to the bipartite CACCT elements [27].
Within the ZEB2 gene cluster of four C2H2 zinc finger (ZnF) domains at N-ZF, this study identifies 20 pathogenic cases [Supplemental Table 2, Figure 12]. These include zinc finger type 1 at amino acids 209-234 (6 cases), type 2 at amino acids 241-263 (4 cases), type 3 at amino acids 282-304 (8 cases), and type 4 at amino acids 310-333 (2 cases). The per-amino-acid pathogenicity rates are 23%, 17%, 35%, and 8%, respectively [Figure 15 & 3]. Although the amino acid lengths are similar (23 vs. 26, 23, and 24), the type 3 zinc finger has more cases and a higher per-amino-acid pathogenicity rate, while type 4 shows the lowest per-amino-acid pathogenicity rate. Among the six cases of type 1 zf, there are two frameshift and four nonsense variants; type 2 zf has three frameshift and one nonsense variant; type 3 zf has six frameshift and two nonsense variants; and type 4 has two frameshift variants [Supplemental Table 2, Figure 12].

3.2.3. SMAD-MH2 Binding Domain [SBD]

As shown in the ZEB2 STRING [String.org] network of interacting proteins [Figure 8 & 18, Table 8 & 10], the most biologically relevant protein-protein interactions of ZEB2 are with SMAD transcription factors. These interactors are key intracellular effectors of the TGFβ superfamily, of which bone morphogenic proteins (BMPs) are a major subfamily [42]. These secreted proteins have diverse functions both in vitro and in vivo and are essential for several developmental processes, including nervous system development [12,42]. Within the confirmed cohort, the SMAD-MH2 binding domain (aa437-487; 10/51=19.61% per aa) does not differ significantly from the rest of aa1-1200 (excluding the domain interval) at the residue level using a Welch t-test (p=0.771; two-tailed).
Among the 83 MWS pathogenic cases in this study, all were confined to the six functional domains of the ZEB2 protein, with 10 (12%) located in the SMAD-MH2 binding domain. This is a notable increase from the baseline of just 2 cases across all domains except the NIM domain [Supplemental Table 2; Figure 12]. Of these 10 cases, 6 resulted from frameshift variants caused by insertions or deletions in the DNA sequence, leading to a shifted reading frame and the production of a different protein from that point onward. Consequently, not only is the SMAD-MH2 binding domain rendered dysfunctional, but the downstream functional domains HD, CID, and CZF-are also lost. This likely causes haploinsufficiency of ZEB2, resulting in the MWS phenotype in these 6 cases. The remaining 4 cases involve nonsense mutations that introduce premature stop codons, producing truncated proteins and the loss of subsequent functional domains:-HD, CID, and CZF.
Within the SMAD-MH2 binding domain, there are 51 amino acids spanning positions 437 to 487, including 437QHLGVGMEAP, 447LLDFPTMNSN, 457LSEVQKVLQI, 467VDNTVSRQKM, 477DCKAEEISKL, and 487K. As shown in Figure 3 and Figure 14, pathogenicity increases by approximately 20% for each amino acid in this domain, slightly exceeding the baseline of 16%. Of the 10 cases involving this domain (Figure 15), 2 involve amino acid variants within a 14-amino-acid stretch inside the larger 51-amino-acid SMAD-binding domain: 437QHLGVGMEAP, 447LLDFPTMNSN, 457LSEVQKVLQI, 467VDNTVSRQKM, 477DCKAEEISKL, and 487K. This segment forms an extended α-helix that may fit into a hydrophobic corridor within the MH2 domain of activated Smads [43]. The 14-residue segment includes four amino acids—two polar Q residues and two nonpolar V residues-forming the tandem repeat (QxVx), which is essential for binding to both TGFβ/Nodal/Activin-Smads and BMP-Smads [Figure 3 and Figure 14; Supplemental Table 2] [43]. Among the 216 cases with MWS pathogenicity in this study, two amino acid variants involve residues at positions 461 (Q: Gln, polar, nonsense variant) and 463 (V: Val, nonpolar, frameshift variant). These are located within the initial QxVx site (Figure 5 and Figure 13), potentially impairing binding to both TGFβ/Nodal/Activin-Smads and BMP-Smads [Figure 18].
The following Figure 14 illustrates the distribution of MWS pathogenic variants within the 51-amino-acid SMAD-binding domain among the 216 MWS pathogenic cases.

3.2.4. DNA Binding/Homeodomain-Like Domain

The DNA-binding/Homeodomain-like region (aa641-708: 13/68=19.12% per aa) does not show significant enrichment compared to the rest of aa1-1200 (excluding the domain interval) (p=0.812; two-tailed). In contrast, C-ZF7 (aa1055-1076: 13/22=59.09% per aa) demonstrates clear enrichment in this residue-level comparison (p=0.014; two-tailed).
The DNA-binding and homeodomain-like region of ZEB2 [Figure 3] is crucial for its function as a transcriptional repressor [Figure 18, Supplemental Table 7 & 10]. This region shares structural similarities with typical homeodomains found in other transcription factors, such as EF1 and Drosophila Zfh-1. However, ZEB2’s homeodomain lacks key conserved residues—specifically, asparagine and arginine in helix 3/4—that are essential for direct DNA binding in canonical homeodomains. Because these residues are missing, ZEB2’s domain may not bind DNA directly; therefore, it is called a “homeodomain-like” sequence rather than a true homeodomain [44,45]. Despite the limitations of its homeodomain-like region, ZEB2 still functions as a transcriptional repressor [44,45]. As shown in Figure 3, the homeodomain-like (HD) segment is flanked by two zinc finger (ZF) clusters, one (NZF) in the N-terminal region and the other (CZF) in the C-terminal region. Its DNA-binding activity mainly derives from these zinc finger clusters, which attach to specific DNA sequences—particularly 5′-CACCT motifs—in target gene promoters [Figure 6] [45]. Therefore, while ZEB2 resembles homeodomain proteins in structure, its DNA-binding and regulatory roles are primarily driven by its zinc fingers rather than the homeodomain itself [45].
The ZEB2 HD segment extends from aa644 to 703 [Figure 3], and this study reports 12 pathogenic MWS cases within this range [Figure 10]. Among the 83 pathogenic cases identified across the six functional domains of the ZEB2 protein in this study, 12 (14%) are located in the DNA Binding / Homeodomain-like Domain, representing a significant increase from the baseline of 2 cases [Supplemental Table 2, Figure 11]. This results in only 18% pathogenicity per amino acid in this domain, due to its larger number of amino acids [Figure 3 & 15], slightly higher than the baseline of 16% pathogenicity [Figure 15]. Of these 12 cases, frameshift and nonsense variants are equally distributed [Supplemental Table 2]. Consequently, not only the HD domain but also the CtBP-Interacting Domain and the C-terminal Zinc Finger clusters, types 6-8, are either lost or rendered nonfunctional.

3.2.5. C-Terminal Binding Protein (CtBP)-Interacting Domain (CID)

The C-terminal binding protein (CtBP)-interacting domain contains four consensus sequences that bind to CtBP1/2 co-repressors, which suppress transcription [23]. CtBP associates ZEB as the promoter to inhibit transcription [46].
The ZEB2-STRING interaction network [Figure 8 & 18, Table 8] reveals direct interactions between ZEB2 and CTBP1 and CTBP2. The CtBP-interacting domain is located at amino acids 757-868 [Figure 5/3 & 10/11; Table 1/2]. Within this 112-amino-acid segment, 13 MWS pathogenic cases were identified in this study [Supplemental Table 2, Figure 11]. Of these, 9 cases involve frameshift variants and 4 involve nonsense variants [Supplemental Table 2]. Due to the frameshift variants, not only is the CID domain rendered dysfunctional, but the three ZF clusters at the C-terminal end of the protein are also lost. This results in the loss of all zinc-finger functions of this protein, since ZEB2’s NZF cluster of four C2H2 Znfs and the CZF cluster of three Znfs tandemly bind as monomers to bipartite CACCT elements, as depicted in Figure 6 [27]. The remaining four cases involve nonsense mutations [Table 1], leading to premature protein termination.
Among the 83 pathogenic cases identified in this study across the six functional domains of the ZEB2 protein, 13 (16%) are found in CID, representing more than a sixfold increase over the baseline of 2 cases [Figure 11]. However, as shown in Figure 16, considering the domain’s enormous size of 112 amino acids, the pathogenicity per amino acid is only 12%.

3.2.6. C2H2 Type Zinc Fingers 5-7 at the C-Terminal (C-ZF)

As shown in Figure 4, the ZEB2 gene’s cluster of three Znfs (5-7) at the C-ZF and the cluster of four C2H2 Znfs (1-4) at the N-ZF bind as monomers to bipartite CACCT elements [27]. In this study, the ZEB2 gene’s cluster of three C2H2 Znfs at the C-ZF includes 28 pathogenic cases [Supplemental Table 2, Figure 11]. These are as follows: zinc finger type 5 at amino acids 999-1021 with 5 cases; type 6 at amino acids 1027-1049 with 10 cases; and type 7 at amino acids 1055-1076 with 13 cases. Their pathogenicity per amino acid is 22%, 43%, and 59%, respectively. Although types 6 and 7 have almost the same number of amino acids as type 5 (23, 22, and 23 amino acids), they show the highest pathogenicity per amino acid at 43% and 59%. Therefore, zinc finger types 6 and 7 demonstrate the highest pathogenicity per amino acid, not only among the three zinc finger clusters in the C-ZF region but also surpassing the four zinc finger clusters in the Z-ZF region [Figure 15]. Among the 5 cases of type 5 zf, there are 2 frameshift and 3 nonsense variants; type 6 zf has 5 frameshift, 3 nonsense, and 2 missense variants; and type 7 zf has 5 frameshift, 6 nonsense, and 2 missense variants [Figure 11 & 16; Supplemental Table 2].

3.2.7. Bar Diagram Showing the Percentage of Pathogenicity Per Amino Acid Across the ZEB2 Functional Domains in the 216 MWS Pathogenic Cases

The following bar diagram seen in Figure15 shows the percentage of MWS pathogenicity for each amino acid across ZEB2’s functional domains.
Figure 15. Bar diagram illustrating the percentage of pathogenicity per amino acid across the ZEB2 functional domains in 216 MWS pathogenic cases [n=216 pathogenic occurrences mapped to ZEB2 protein length L=1214 aa; per-residue incidence: domain % per amino acid (cases/length): NIM=0.00% (0/9), N-ZF1=23.00% (6/26), N-ZF2=17.39% (4/23), N-ZF3=34.78% (8/23), N-ZF4=8.33% (2/24), SMAD-MH2=19.61% (10/51), Homeodomain=17.65% (12/68), CtBP=11.61% (13/112), C-ZF5=21.74% (5/23), C-ZF6=43.48% (10/23), C-ZF7=59.09% (13/22); below-baseline domains include NIM, N-ZF4, and CtBP. [see Figure 3 & 9, and Supplemental Table 2].
Figure 15. Bar diagram illustrating the percentage of pathogenicity per amino acid across the ZEB2 functional domains in 216 MWS pathogenic cases [n=216 pathogenic occurrences mapped to ZEB2 protein length L=1214 aa; per-residue incidence: domain % per amino acid (cases/length): NIM=0.00% (0/9), N-ZF1=23.00% (6/26), N-ZF2=17.39% (4/23), N-ZF3=34.78% (8/23), N-ZF4=8.33% (2/24), SMAD-MH2=19.61% (10/51), Homeodomain=17.65% (12/68), CtBP=11.61% (13/112), C-ZF5=21.74% (5/23), C-ZF6=43.48% (10/23), C-ZF7=59.09% (13/22); below-baseline domains include NIM, N-ZF4, and CtBP. [see Figure 3 & 9, and Supplemental Table 2].
Preprints 222684 g014
The total number of MWS pathogenic cases in the functional domain regions in this study is 83, while in the non-functional regions it is 133. The total number of amino acids in the functional domain regions is 404, compared to 810 in the non-functional regions. Therefore, the combined total of amino acids in the ZEB2 protein is 404 + 810 = 1,214 [functional-domain residues n_in=404; non-domain residues n_out=810]. The percentage of pathogenic cases per amino acid in the functional domains, which consist of 404 amino acids, is 20.54% (=83/404×100=20.54% per aa; mean incidence=83/404=0.2054 cases/residue), whereas in the non-functional domains with 810 amino acids, it is 16.42% (=133/810×100=16.42% per aa; mean incidence=133/810=0.1642 cases/residue) [Figure 14]. Therefore, there is a 4.12% increase in pathogenicity within domain regions compared to non-domain regions (Δ=20.54−16.42=4.12 percentage points; Welch t=1.542, df=722.387, p=0.1236; two-tailed).
Except for the N-ZF C2H2-type 4: aa310-333 and the CtBP Interacting Domain: aa757-868, all other functional domain regions have a higher percentage of pathogenic cases per amino acid than the non-functional domain regions (interdomain regions), which is 16.42% [non-domain baseline=133/810=16.42% per aa; domain % per aa (cases/length): NIM=0.00% (0/9), N-ZF1=23.00% (6/26), N-ZF2=17.39% (4/23), N-ZF3=34.78% (8/23), N-ZF4=8.33% (2/24), SMAD-MH2=19.61% (10/51), Homeodomain=17.65% (12/68), CtBP=11.61% (13/112), C-ZF5=21.74% (5/23), C-ZF6=43.48% (10/23), C-ZF7=59.09% (13/22); domains below baseline include NIM, N-ZF4, and CtBP) [Figure 14/15]. The C-ZF C2H2-type 7: AA: 1055-1076 at the protein’s tail end shows the highest pathogenicity per amino acid: 59% [13/22=59.09% per aa; Welch vs non-domain t=2.715, df=21.345, p=0.0128; two-tailed] [Figure 14/15]. This counters the assumption (C-ZF7: 59.09% vs non-domain 16.42% → Δ=42.67 pp; p=0.0128; C-ZF6: 43.48% → Δ=27.06 pp; p=0.0639) that mutations in the C-terminal region do not cause haploinsufficiency, since most of the protein remains intact, including all functional domains, and even if a C-terminal mutation slightly affects the protein, it results in only milder phenotypes [35].
Pathogenic variants at the tail end of the C-ZF zinc finger cluster, specifically in Zfn type 7, showed that all three Zfn regions within the C-ZF domain are necessary for the functional ZEB2 protein. The ZEB2 gene’s cluster of four C2H2 Zns at N-ZF and the three Zns at its C-ZF bind together as monomers to bipartite CACCT elements to conduct these functions, as shown in Figure 6 [27,41]. Both the N-ZF and C-ZF Zfn clusters must be intact, and even a single disruptive variant at the last Zfn (C2H2 type 7) at the C terminus causes MWS pathogenicity, as shown by this in silico study of 216 MWS pathogenic cases.
The next-highest pathogenicity percentage [43%] appears in the C-ZF C2H2-type 6 (aa1027-1049) zinc finger (zfn) [10/23=43.48% per amino acid; Welch vs. non-domain t=1.949, df=22.464, p=0.0639; two-tailed], which is adjacent to the C-ZF C2H2-type 7 (aa1055-1076). These results indicate that the integrity of the C-terminal cluster of three zinc fingers is crucial for ZEB2 protein function, and that loss-of-function mutations, whether frameshift or nonsense, can lead to haploinsufficiency and pathogenicity in MWS.
The next-highest pathogenicity percentage [35%], caused by frameshift or nonsense variants, occurs at N-ZF C2H2-type 3: aa282-304 (8/23=34.78% per amino acid; Welch vs non-domain t=1.353, df=22.486, p=0.1895; two-tailed), which ranks third among four N-terminal clusters with four zinc finger domains. The resulting truncation of the ZEB2 protein results in the loss of the N-ZF C2H2-type 4 (aa310-333), the SMAD-MH2-binding domain (aa437-487), the DNA-binding/Homeodomain (aa641-708), the CtBP-interacting domain (aa757-868), and the entire C-ZF cluster of three zinc finger domains, thereby eliminating all functional domains of the ZEB2 protein.
As shown in Figure 11 and Figure 13, the distribution of milder [likely/conflicting MWS pathogenicity] cases span the entire length of the protein, peaking at the N-terminal end of the ZEB2 protein at aa 51-150, which should exclude all functional domains except the NIM domain at aa 14-22. A similar peak is observed at aa 751-800, where the CtBP-Interacting Domain is located at aa 751-868 [Figure 13]. The last group of three adjacent peaks occurs at 951-1100, encompassing the C-ZF C2H2 types 6 and 7 at aa 1027-1049 and aa 1055-1076, respectively [Figure 13]. This pattern was seen among the MWS pathogenic cases [Figure 11]. It appears that a mild case with likely or conflicting pathogenicity can occur even when the truncation happens at the N-terminal end of the ZEB2 protein, leaving it devoid of all functional domains except the NIM domain at aa 14-22 [Figure 11]. Furthermore, the milder cases peak at aa 751-800, which includes the CtBP-Interacting Domain at aa 757-868 [Figure 13]. It is not surprising to see milder cases with pathogenic amino acid variants at the C-terminal end of the protein, within the C-ZF C2H2 cluster types 6 and 7 at aa 1027-1076 [Figure 13], and even beyond at aa 1101-1200 [Figure 12 and Figure 13].

3.2.8. Relative Distribution of Pathogenic Variant Types Among the 216 Confirmed MWS Pathogenic Cases and the 67 Mild MWS Cases With Likely or Conflicting Pathogenicity Reports (n=216 Confirmed Pathogenic; n=67 Mild, Likely/Conflicting)

The following bar diagram in Figure 16 illustrates the relative frequencies of MWS pathogenic amino acid variants among the 216 MWS pathogenic cases.
Figure 16. The bar diagram displays the relative frequencies of ZEB2 pathogenic variants among the 216 MWS pathogenic cases (frameshift, 133/216 = 61.57%; nonsense, 75/216 = 34.72%; missense, 7/216 = 3.24%; in-frame deletion, 1/216 = 0.46%) [see Supplemental Table 2 and Figure 3]. A. Occurrence of Frameshift Variants, Nonsense Variants, Missense Variants, and In-frame Deletions among all 216 MWS Pathogenic Cases. B. Occurrence of Frameshift Variants, Nonsense Variants, Missense Variants, and In-frame Deletions within the Functional Domains of the ZEB2 Protein in the 216 MWS Pathogenic Cases. C. Occurrence of Frameshift Variants, Nonsense Variants, Missense Variants, and In-frame Deletions within the Non-Domain Regions of the ZEB2 Protein in 216 MWS Pathogenic Cases.
Figure 16. The bar diagram displays the relative frequencies of ZEB2 pathogenic variants among the 216 MWS pathogenic cases (frameshift, 133/216 = 61.57%; nonsense, 75/216 = 34.72%; missense, 7/216 = 3.24%; in-frame deletion, 1/216 = 0.46%) [see Supplemental Table 2 and Figure 3]. A. Occurrence of Frameshift Variants, Nonsense Variants, Missense Variants, and In-frame Deletions among all 216 MWS Pathogenic Cases. B. Occurrence of Frameshift Variants, Nonsense Variants, Missense Variants, and In-frame Deletions within the Functional Domains of the ZEB2 Protein in the 216 MWS Pathogenic Cases. C. Occurrence of Frameshift Variants, Nonsense Variants, Missense Variants, and In-frame Deletions within the Non-Domain Regions of the ZEB2 Protein in 216 MWS Pathogenic Cases.
Preprints 222684 g015
There are no significant differences in the occurrence of Frameshift, Nonsense, and Missense Variants among the 216 pathogenic cases when grouped as: (A) all 216 MWS pathogenic cases, (B) specifically within ZEB2 protein’s functional domains, and (C) specifically within ZEB2 protein’s non-functional domains [Figure 16] (Domains vs. non-domains, 216 cohort: frameshift t=-1.739, df=167.049, p=0.0839; nonsense t=1.500, df=165.575, p=0.1354; missense t=0.951, df=131.022, p=0.3433; overall 2×4 χ2=4.358, df=3, p=0.2253; two-tailed).
The following bar diagram, Figure 17, shows the relative frequencies of pathogenic variant types among 67 MWS cases with likely or conflicting pathogenicity: Missense 40/67=59.70%; Frameshift 15/67=22.39%; Nonsense 7/67=10.45%; In-frame deletion 1/67=1.49%; Synonymous 4/67=5.97%[see Supplemental Table 4; Figure 3].
Figure 17. The bar diagram shows the relative frequencies of pathogenic variant types among the 67 MWS cases with likely or conflicting pathogenicity: Missense, 40/67 = 59.70%; Frameshift, 15/67 = 22.39%; Nonsense, 7/67 = 10.45%; In-frame deletion, 1/67 = 1.49%; Synonymous, 4/67 = 5.97%. [see Supplemental Table 4; Figure 3]. A. Occurrence of Frameshift Variants, Nonsense Variants, Missense Variants, and In-frame Deletions among All 67 MWS likely/conflicting pathogenic cases. B. Occurrence of Frameshift Variants, Nonsense Variants, Missense Variants, and In-frame Deletions within the Functional Domains of the ZEB2 Protein in the 67 MWS Likely/Conflicting Pathogenic Cases. C. Occurrence of Frameshift Variants, Nonsense Variants, Missense Variants, and In-frame Deletions within the Non-Domain Regions of the ZEB2 Protein among the 67 MWS Cases with Likely or Conflicting Pathogenicity.
Figure 17. The bar diagram shows the relative frequencies of pathogenic variant types among the 67 MWS cases with likely or conflicting pathogenicity: Missense, 40/67 = 59.70%; Frameshift, 15/67 = 22.39%; Nonsense, 7/67 = 10.45%; In-frame deletion, 1/67 = 1.49%; Synonymous, 4/67 = 5.97%. [see Supplemental Table 4; Figure 3]. A. Occurrence of Frameshift Variants, Nonsense Variants, Missense Variants, and In-frame Deletions among All 67 MWS likely/conflicting pathogenic cases. B. Occurrence of Frameshift Variants, Nonsense Variants, Missense Variants, and In-frame Deletions within the Functional Domains of the ZEB2 Protein in the 67 MWS Likely/Conflicting Pathogenic Cases. C. Occurrence of Frameshift Variants, Nonsense Variants, Missense Variants, and In-frame Deletions within the Non-Domain Regions of the ZEB2 Protein among the 67 MWS Cases with Likely or Conflicting Pathogenicity.
Preprints 222684 g016
Similar to the relative frequencies of ZEB2 pathogenic variants among the 216 MWS pathogenic cases, there are no significant differences in the occurrence of Frameshift, Nonsense, and Missense Variants among the 67 pathogenic cases when grouped as: (A) all 67 MWS pathogenic cases, (B) specifically within ZEB2 protein’s functional domains, and (C) within ZEB2 protein’s non-functional domains [Figure 17]. (Domains vs non-domains, 67 cohort: frameshift t=-0.026, df=55.883, p=0.9791; nonsense t=-0.690, df=63.113, p=0.4925; missense t=0.443, df=56.601, p=0.6598; overall 2×4 χ2=1.152, df=3, p=0.7645; two-tailed).
Among the 216 pathogenic cases, frameshift variants are the most common, closely followed by nonsense variants, with missense variants appearing only rarely. The distribution is as follows: frameshift variants make up 133/216 = 61.57%, 45/83 = 54.22%, and 88/133 = 66.17%; nonsense variants account for 75/216 = 34.72%, 34/83 = 40.96%, and 41/133 = 30.83%; and missense variants constitute 7/216 = 3.24%, 4/83 = 4.82%, and 3/133 = 2.26% [Figure 16, Supplemental Table 2].
Among the 216 pathogenic cases, the few missense variants occur within or immediately adjacent to the ZEB2 protein’s most important domain, the C-ZF zinc finger cluster (types 5-7), at the following amino acid positions: 1045 and 1049 (C-ZF zinc finger type 6); 1050 (near C-ZF zinc finger type 6); 1057 and 1071 (C-ZF zinc finger type 7); and 1081 and 1119 (near the C-ZF cluster) [Figure 16, Supplemental Table 2].
In clear contrast, among the 67 likely or conflicting-pathogenicity cases, missense variants are more common than frameshift variants across: (A) all 67 MWS pathogenic cases; (B) specifically within ZEB2 protein’s functional domains; and (C) within ZEB2 protein’s non-functional domains, at 40/67 = 59.70%, 17/27 = 62.96%, and 23/40 = 57.50%, respectively. They are distributed throughout the entire protein from amino acid #9 to #1179 [Figure 17, Supplemental Table 4].
The predominance of missense variants among the mild cases with likely or conflicting pathogenicity [Table 8] is notable when compared to the confirmed-pathogenic cohort and is observed across: (A) all 67 MWS pathogenic cases, (B) specifically within ZEB2 protein’s functional domains, and (C) within ZEB2 protein’s non-functional domains, at 40/67=59.70%, 17/27=62.96%, and 23/40=57.50%, respectively (vs 7/216=3.24% in the confirmed-pathogenic cohort; Welch t=-9.170, df=71.352, p=1.09×10−13; two-tailed; overall 2×4 χ2=128.33, df=3, p=1.24×10−27, excluding synonymous variants). This is followed at a distance by frameshift variants at 15/67 (22.39%), 6/27 (22.22%), and 9/40 (22.50%), respectively [Figure 17].
Among the 67 likely or conflicting pathogenicity cases, nonsense variants account for 7 out of 67 (10.45%), 2 out of 27 (7.41%), and 5 out of 40 (12.50%) across: (A) all 67 MWS pathogenic cases, (B) specifically within ZEB2 protein’s functional domains, and (C) within ZEB2 protein’s non-functional domains, respectively (vs 75 out of 216 = 34.72% in the confirmed-pathogenic cohort; Welch t=4.882, df=171.542, p=2.38×10−6; two-tailed).
The above findings support the earlier prediction [22] that ZEB2 zinc-finger missense mutations lead to hypomorphic alleles and a milder presentation of MWS, based on only six cases of missense mutations with mild phenotypes. This is because a missense mutation changes a single amino acid in the ZEB2 protein via a single nucleotide point mutation likely producing a protein that retains more of its function. In contrast, frameshift mutations alter downstream codons, and nonsense mutations create a stop codon, both of which generally impair the structure and function of the ZEB2 protein. However, if frameshift and nonsense variants occur after the highly conserved C-terminal zinc finger domain, the protein is less likely to cause haploinsufficiency across all functional regions. In these cases, they may result in only mild pathogenicity or even produce a normal phenotype. These predictions suggest that, depending on the mutation’s location and the type of ZEB2 variant, the mutant protein may not be subject to mRNA decay [35].
Since each ZEB2 domain functions independently within the larger ZEB2 protein, a missense mutation in one domain is less likely to disable other domains. Mutations predicted to retain some level of protein function tend to cause milder clinical symptoms [6,47]. For example, two mild cases have been documented with mutations in the NURD Interaction Motif [NIM], which only disrupts the interaction of this mutant allele with the NuRD co-repressor complex, leaving the rest of the ZEB2 protein unaffected [35,37,38]. Frameshift and nonsense mutations produce truncated, smaller proteins that are usually nonfunctional, whereas missense intragenic mutations are often thought to preserve partial or full function, leading to milder symptoms [6], as shown in this study [Table 9; Figure 17]. Zou et al. [47] observed that, in three MWS cases with nonsense mutations and one with a frameshift mutation, severity increased when the mutation was located closer to the amino terminus of the protein, indicating that clinical severity in MWS depends on the position of the truncation within the ZEB2 gene.
Truncating mutations can evade nonsense-mediated decay and be translated, leading to the buildup of the truncated protein. In such cases, the smaller the resulting polypeptide, the more severe the phenotype [47]. Frameshift mutations truncate the functional protein and, depending on their location, can vary in disease severity. Not all ZEB2 alleles found in MWS cause a complete loss of function [6,47].
In this study, frameshift variants are detected throughout the ZEB2 protein (amino acids 8-1190) in 216 MWS pathogenic cases [Supplemental Table 2; Figure 10] (n=133 frameshift variants; 70 out of 133 in aa 1–600 versus 63 out of 133 in aa 601–1214; exact binomial p=0.603 for deviation from a 50:50 split).
Similarly, frameshift variants among the 67 likely or conflicting pathogenicity cases span amino acids 24-1076 [Supplemental Table 4; Figure 12 & 13] (n=15 frameshift variants; 11/15 in aa 1–600 vs 4/15 in aa 601–1214; exact binomial p=0.118 for deviation from a 50:50 split). The frameshift variant closest to the N-terminus among the 216 MWS pathogenic cases is at amino acid 8, which should result in the loss of all functional domains of the ZEB2 protein [Table 1]. Similarly, the frameshift variant closest to the N-terminus, at amino acid 24, should also lead to the loss of all functional domains of the ZEB2 protein, except for the NIM domain at amino acids 14-22 [Supplemental Table 2]. Although there are frameshift variants at the same amino acid positions among these two groups of MWS cases (shared frameshift positions: aa 24 and aa 519), the severity of MWS pathogenicity varies significantly [Supplemental Table 2 & 4]. This difference in pathogenicity for a given amino acid variant between the MWS pathogenic and mild MWS pathogenic groups can be attributed to mosaicism [48,49,50], methylation [51], or possibly the type of mutation involved. These nuances pose a challenge in genetic counseling, highlighting the need for further research and more precise data for clinical use.

3.2.9. MWS Pathogenicity Confirmed Cases and Mild Likely/Conflicting MWS Pathogenic Cases Due to Variants in Intronic or Splice Sites

Further complicating the genetic counseling dilemma, our in-silico study identified 10 additional cases with MWS pathogenicity caused by base-pair variants in intronic splice acceptor and donor sites [Supplemental Table 3]. Therefore, there are 226 cases (216 + 10) with MWS pathogenicity, of which 4.42% are due to these intronic variants. Similarly, this study identified 8 cases of milder MWS pathogenicity [Table 5], in addition to the 67 cases with likely or conflicting pathogenicity caused by base-pair variants in intronic splice acceptor and donor sites. In total, our study involved 75 milder MWS cases [67 + 8], with 10.7% attributable to these intronic variants. Hence, it is essential to consider these specific base-pair variants at intronic splice acceptor and donor sites when assessing MWS pathogenicity.

3.2.10. STRING [String.org] Protein-Protein Interaction Network of the ZEB2 Gene [Figure 8] Validates the Predicted ZEB2-Molecular Functions and Gene Ontology [GO] Biological Processes from PubMed.org, Ensembl.org and Uniprot.org.

See Supplementary Table 10 for a detailed list of ZEB2 STRING [www.String.org] protein-protein interactions that mirror ZEB2 functions and biological processes, as predicted in Supplementary Table 7 based on the data obtained from PubMed.org, Ensembl.org, and UniprotKB-KW.
Figure 18 below shows the significant enrichment of biological processes and functions summarized by STRING, based on the ZEB2 protein-protein interactions data in Supplementary Table 10.
Figure 18. A figurative representation of the significant enrichment of the biological processes summarized through ZEB2 STRING interactions’ strength and signal values listed in Supplementary Table 10. FDR: False Discovery Rate [52].
Figure 18. A figurative representation of the significant enrichment of the biological processes summarized through ZEB2 STRING interactions’ strength and signal values listed in Supplementary Table 10. FDR: False Discovery Rate [52].
Preprints 222684 g017
The predicted ZEB2 molecular functions and Gene Ontology (GO) biological processes from PubMed.org, Ensembl.org, and Uniprot.org (Table 7) were largely confirmed in the String.org protein-protein interaction network for the ZEB2 gene (Table 10, Figure 18). The GO function ‘regulation of cell fate specification’ showed a strong signal of 4.3 (Figure 18) and includes stem cell differentiation and melanocyte cell fate specification. Additionally, as shown in the biological process enrichment figure and table (Figure 18; Table 10), the most prominent process was the negative regulation of transcription by RNA polymerase, with the highest gene count of 20, followed by chromatin organization (15 genes), histone modification (>10 genes), chromatin remodeling (10 genes), and histone deacetylation (<10 genes with a 3.8 signal). Histone deacetylation involves the removal of acetyl groups from histones, reducing DNA access and repressing gene transcription, which is relevant to disease development. This supports the transcriptional inhibitor and repressor functions of the ZEB2 gene. Similarly, the negative regulation of stem cell differentiation (<10 genes; 2.8 signal; FDR: 1.0e-15), the epigenetic regulation of gene expression (7 genes), and the regulation of cell-fate specification (<10 genes, but with maximum signal: 4.3)- all align with the predicted roles of ZEB2 (Figure 18, Table 7 & 10). Furthermore, the regulation of the TGF receptor signaling pathway (<10 genes; signal 2.2) (Figure 18) corresponds with the predicted positive regulation of the TGF beta receptor biological function of ZEB2 (Table 7).

4. Materials and Methods

4.1. In Silico Study of ClinVar Website MWS Database

This in silico research study utilized the ClinVar website MWS database to compile 226 pathogenic ZEB2 gene variants [Supplemental Table 2 & 3] and 75 MWS-likely/conflicting pathogenic ZEB2 variants [Supplemental Table 4 & 5]. The site [https://www.ncbi.nlm.nih.gov/clinvar/ClinVar] is freely accessible and functions as a public archive of reports on human gene variants classified by disease and drug response, with supporting evidence. ClinVar enables access to the relationships asserted between human variants and observed conditions, along with their history. The database reports variants found in patient samples, disease and drug classifications, submitter information, and other supporting data. Variants in submissions are mapped to reference sequences and reported according to the HGVS standard. ClinVar offers data on the website for interactive users, on the FTP site, and via API for those integrating ClinVar into daily workflows and local applications.
Our study thoroughly collected and merged all 301 fully annotated MWS pathogenicity cases, along with those with likely or conflicting classifications from the MWS pathogenic ClinVar database. The data includes serial numbers, affected amino acid positions, accession numbers, reference SNP cluster IDs [rsID], clinical significance and conditions, chromosome locations, mutation names, mutant base pairs, amino acid changes, variant locations, base pair and amino acid alterations, and variant types.
These data were further classified into pathogenic variants in the coding regions of the ZEB2 gene [216 cases: Supplemental Table 2] and pathogenic variants in intronic, splice-acceptor, and splice-donor sites of the gene [10 cases: Supplemental Table 3].
Similarly, data from ClinVar on 75 cases with Likely MWS Pathogenicity or Conflicting MWS Pathogenicity were collected. These were classified as Likely Pathogenic variants in the coding regions of the ZEB2 gene [67 cases: Supplemental Table 4] and as likely pathogenic variants in intronic, splice acceptor, or splice donor sites of the gene [8 cases: Supplemental Table 4 ].
This study mainly analyzes the distribution of over 300 ClinVar-reported MWS pathogenic and likely pathogenic cases to determine their locations within the ZEB2 protein and to examine their impact on clinical outcomes. It explores the pattern of pathogenic variants near the N-terminus compared to the C-terminus of ZEB2, aiming to understand how retaining more intact functional domains influences outcomes compared to complete domain loss. The study also assesses the frequency of pathogenicity per amino acid residue within different functional domains of ZEB2 compared to non-domain regions, and whether certain variant types are more strongly associated with pathogenicity severity. Additionally, it considers variants in the gene’s non-coding regions. Ultimately, it aims to summarize how disruptions in ZEB2’s molecular functions caused by pathogenic variants can lead to diverse abnormalities across multiple organs and systems, especially the nervous system.

4.2. Statistical Analysis

Statistical analysis using various studies was also conducted to examine the ClinVar data, aiming to quantify (i) the distribution of ZEB2 variant incidences along the protein sequence (Figure 9, Figure 10, Figure 11, Figure 12 and Figure 13), (ii) the enrichment of incidences within predefined functional domains compared to non-domain regions (Figure 15), and (iii) the differences in variant-type composition between clinically defined cohorts (Figure 16 and Figure 17). Incidence was defined as a residue-level count: for each amino acid position i, xi represents the number of cases mapping to residue i. For regional summaries (such as 50-aa bins, domain regions, and non-domain regions), the mean incidence was calculated as the arithmetic mean across residues within each region, reported both as cases per residue and as a percentage per amino acid.
Mean incidence: x̄ = (Σ xi) / n
Percent per amino acid: (% per aa) = ( (Σ xi) / n ) × 100 = x̄ × 100
To determine whether the average incidence differs between two regions (for example, domains versus non-domains or a candidate peak bin versus the rest of the protein), two sample Welch’s t-tests were performed. Welch’s t-test was chosen because the regions usually vary in length (n1 ≠ n2) and the incidence variance is not assumed to be equal across regions (s12 ≠ s22), especially in sparse count data where many residues have zero incidence. Two tailed p-values were reported for all t-tests.
Welch t-test: t = (x̄1 − x̄2) / √( s12/n1 + s22/n2 )
Welch–Satterthwaite df: df = ( s12/n1 + s22/n2 )2 / ( (s12/n1)2/(n1−1) + (s22/n2)2/(n2−1) )
Two-tailed p-value: p = 2 × P(Tdf ≥ |t|)
For comparisons of categorical variant-type composition (e.g., frameshift, nonsense, missense, in-frame deletion) between groups (domains vs non-domains; 216 confirmed pathogenic vs 67 mild/likely-conflicting), Pearson’s chi-square tests of independence were used on contingency tables because the outcomes are nominal rather than continuous. Right-tailed p-values from the chi-square distribution were reported.
Chi-square: χ2 = Σ (O − E)2 / E, df = (r−1)(c−1)
All computations were performed in Microsoft Excel (Office 365) using auditable, cell-based formulas. Counts were generated from amino acid position lists using COUNTIF/COUNTIFS and SUMPRODUCT; regional means and percentages were derived from these counts; Welch two-tailed p-values were calculated with T.DIST.2T(|t|, df); and chi-square p-values were obtained with CHISQ.DIST.RT(χ2, df). This method ensures full traceability from input data to the reported statistics across all figures.

4.3. Study of UniProt Protein Knowledgebase

This research study additionally uses the UniProt protein knowledgebase (https://www.uniprot.org/uniprotkb) to define the ZEB2 protein structure and functions [Figure 2 & 3; Table 6]. The UniProt Knowledgebase (UniProtKB) serves as the primary resource for protein functional information, providing accurate, consistent, and detailed annotations. It also includes essential data for each UniProtKB entry—such as the amino acid sequence, protein name or description, taxonomic data, and citation information—and incorporates as much annotation information as possible.

4.4. Study of Ensembl Genome Browser

Additionally, this research uses the Ensembl genome browser [Ensembl genome browser 115] for vertebrate genomes, supporting studies in comparative genomics, evolution, sequence variants, and transcriptional regulation. The Ensembl website [www.Ensembl.org] provides gene annotations, performs multiple sequence alignments, predicts regulatory functions, and compiles disease-related data. Key Ensembl tools include the Variant Effect Predictor (VEP).

4.5. Study of STRING Protein-Protein Interaction Network

Finally, to summarize the molecular functions of ZEB2 and the Gene Ontology (GO) biological processes [Table 7] proposed by Ensembl.org, Uniprot.org, and PubMed.org, this study utilized the STRING protein-protein interaction network [https://string-db.org/]. STRING is a comprehensive resource for exploring protein interactions and relationships across many organisms. It combines data from experimental results, computational predictions, and text mining to provide a detailed overview of protein interactions. Table 7 is located in the Supplementary material, and cited PubMed.org, Ensembl.org, UniProt.org, UniProtKB-KW, as pertinent sources for predicted ZEB2 molecular functions and Gene Ontology [GO] biological processes.
Figure 8, shown below, represents the interactions of the ZEB2 protein with 24 key interactors [listed inTable 8below] from the STRING protein interaction network.
Preprints 222684 i005
NOTES ON INTERACTING GENES:
Figure 8. ZEB2 protein-protein interactions with 24 interactors/25 nodes from the STRING network [Modified from http://www.STRING.org].
Figure 8. ZEB2 protein-protein interactions with 24 interactors/25 nodes from the STRING network [Modified from http://www.STRING.org].
Preprints 222684 g018
Network: Shows current interactions. Neighborhood: Groups of genes frequently found in each other’s genomic neighborhood. Experiments: Co-purification, co-crystallization, Yeast2Hybrid, Genetic Interactions, etc., as imported from primary sources. Co-occurrence: Gene families with similar occurrence patterns across genomes. Databases: Known metabolic pathways, protein complexes, signal transduction pathways, etc., from curated databases. Co-expression: Proteins whose genes show correlated expression across a large number of experiments. Text-mining: Automated, unsupervised text-mining—searching for proteins that are frequently mentioned together. Fusion: Genes that are sometimes fused into single open reading frames. The edges in this figure represent protein-protein associations, intended to be specific and meaningful: proteins jointly contribute to a shared function; this does not necessarily imply physical binding.
Known Interactions are curated from databases and experimentally determined. Predicted Interactions: gene neighborhood, gene fusions, and gene co-occurrence. Others: text mining, co-expression, and protein homology. Network Stats: Number of nodes: 26. Number of edges:137; expected number of edges: 27. Average node degree: 10.5 Average local clustering coefficient: 0.74.
PPI enrichment p-value: < 1.0e-16. Interaction Enrichment: The ZEB2 protein’s STRING network contains more interactions than would be expected for a random set of proteins of the same size and degree distribution drawn from the genome. Such an enrichment indicates that the proteins are at least partially biologically connected as a group with salient functions listed in Table 8 below.
The following Table 8 lists salient functions of 24 ZEB2 protein-protein interactors [25 nodes] observed in the STRING protein interaction network shown in Figure 8 [http://www.STRING.org].
Table 8. Salient functions of each of the 24 ZEB2 protein-protein interactors [25 nodes] in the above STRING protein interaction network [Figure 8] [http://www.STRING.Org]:.
Table 8. Salient functions of each of the 24 ZEB2 protein-protein interactors [25 nodes] in the above STRING protein interaction network [Figure 8] [http://www.STRING.Org]:.
Preprints 222684 i006Preprints 222684 i007Preprints 222684 i008

5. Conclusions

5.1. MWS Pathogenicity, as well as Milder Pathogenicity Occurs Throughout the Entire Length of the ZEB2 Protein

This in silico analysis of 216 MWS-confirmed ClinVar cases and 67 milder ClinVar cases of likely or conflicting pathogenicity found that MWS pathogenicity, as well as milder pathogenicity occurs throughout the entire length of the ZEB2 protein, even beyond the C-ZF type 7 at amino acids 1118 and 1179. This finding has implications for prenatal genetic counseling. The finding that an amino acid variant located even beyond the C-ZF functional domain at the C-terminal end of the ZEB2 protein can confer pathogenicity, ranging from mild to severe, has profound implications for prenatal genetic counseling.

5.2. ZEB2 Base-Pair Variants in Intronic/Splice Acceptor/Splice Donor Sites Also Cause MWS Pathogenicity

Detecting MWS pathogenicity in confirmed cases, as well as in mild cases (with likely or conflicting pathogenicity) due to base-pair variants in intronic/splice acceptor/splice donor sites, makes accurate prenatal genetic counseling even more challenging.

5.3. Incidence of MWS Pathogenicity at ZEB2 Protein Functional Domains Show Significant Correlation

There is a significant correlation between the incidence of MWS pathogenicity and the functional domains of the ZEB2 protein (p=0.0152). Pathogenic variants at the tail end of the C-ZF zinc finger cluster, specifically at Zfn type 7, demonstrate that the integrity of all three Zfn regions within the C-ZF domain is necessary to maintain the ZEB2 protein’s functionality. The ZEB2 gene’s cluster of four C2H2 Znfs at N-ZF and the cluster of three Znfs at its C-ZF bind as monomers to bipartite CACCT elements to perform their functions [27,41]. Therefore, the integrity of the C-terminal cluster of three Zfn is crucial for preserving the overall function of the ZEB2 protein, and its loss of function—either due to a frameshift or nonsense mutation—can lead to haploinsufficiency and MWS pathogenicity as observed in clinical cases.
The distribution of milder [likely/conflicting MWS pathogenicity] cases cover the entire length of ZEB2, with a peak at the N-terminal end (aa 51-150). This region probably makes all functional domains ineffective, except for the NIM domain at aa 14-22. This suggests that a mild case with likely/conflicting pathogenicity can occur even with truncation at the N terminus of the ZEB2 protein, leaving only the NIM domain at aa 14-22. Additionally, milder cases also tend to cluster at aa 751-800, which includes the CtBP-Interacting domain at aa 757-868. It is not surprising to observe milder cases with amino acid variants at the C-terminal end of the protein, including the C-ZF C2H2 cluster types 6 & 7 at aa 1027-1076, and even beyond aa 1101-1200.
Pathogenicity is most prominent in the following functional domains: DNA-binding/Homeodomain, CtBP-interacting domain, and C-ZF C2H2-type 7. Therefore, one could conclude that there is a general increase in the frequency of MWS pathogenicity beyond the baseline across the functional domains and regions of the ZEB2 protein, except in the NZF C2H2-type 4 region: aa 310-333, and that pathogenicity is completely absent at the N-terminal NuRD-Interacting Motif [NIM]: aa 14-22.

5.4. No Significant Differences in the Occurrences of Frameshift, Nonsense, and Missense Variants Among the Pathogenic and the Likely Pathogenic Cases When Grouped

There are no significant differences in the occurrences of Frameshift, Nonsense, and Missense Variants among the 216 pathogenic cases when grouped as: (A) all 216 MWS pathogenic cases, (B) specifically within ZEB2 protein’s functional domains, and (C) within ZEB2 protein’s non-functional domains [Figure 16]. Similarly, there are no significant differences in the occurrences of Frameshift, Nonsense, and Missense Variants among the 67 MWS likely or conflicting pathogenic cases when grouped as: (A) all 67 cases, (B) specifically within ZEB2 protein’s functional domains, and (C) within ZEB2 protein’s non-functional domains [Figure 17].

5.5. Among the Likely or Uncertain Pathogenicity Cases, Missense Variants Significantly Outweigh Frameshift Variants

Among the 216 pathogenic cases, the small number of missense variants occurs exclusively in the ZEB2 protein’s most critical domain, the C-ZF zinc finger cluster (types 5-7). In stark contrast, among the 67 cases with likely or uncertain pathogenicity, missense variants outweigh frameshift variants.
The prevalence of missense variants among mild cases with likely or conflicting pathogenicity is indeed significant: overall p=1.24×10−27. These findings support earlier predictions by Ghoumed et al. [22] that ZEB2 zinc-finger missense mutations lead to hypomorphic alleles and mild Mowat-Wilson syndrome, based on only six cases with mild phenotypes. The reasoning is that a missense mutation changes thereby increasing the likelihood of producing a functionally intact protein. In contrast, frameshift mutations affect subsequent codons, and nonsense mutations create a stop codon, both of which can severely impair the integrity and function of the ZEB2 protein.

5.6. Frameshift Variants Among Pathogenic and Likely Pathogenic Cases Are Randomly Distributed Throughout the Entire Length of ZEB2 Protein

In this study, frameshift variants are randomly distributed throughout the entire length of the ZEB2 protein (amino acids 8-1190) in 216 cases with pathogenic MWS. Similarly, frameshift variants among the 67 cases with likely or conflicting pathogenicity are spread across amino acids 24-1076.

5.7. A Specific Amino Acid Variant Causes MWS Pathogenicity or Milder Pathogenicity

The observed difference in pathogenicity for a specific variant amino acid between the MWS pathogenic and the mild MWS pathogenic groups of ZEB2 protein variant cases [Supplemental Table 2 & 4] could be attributed to mosaicism [48,49,50], methylation [51], or possibly the type of mutation involved. These nuances certainly create a dilemma for genetic counseling and the patient’s medical management.

5.8. STRING ZEB2 Protein-Protein Interaction Network Analysis Confirms ZEB2 Participation in Multiple Systems that Are Affected in Patients with Mowat-Wilson Syndrome

The STRING protein-protein interaction network for ZEB2 gene supports the biological process “regulation of cell fate specification,” which includes stem cell differentiation and melanocyte cell fate specification. With varying signal strengths, the network also confirms the ZEB2 gene’s negative regulation of transcription, chromatin organization, histone modification, chromatin remodeling, histone deacetylation, and positive regulation of TGF beta receptor biological function—each of which participates in multiple systems in patients with Mowat-Wilson syndrome.

Supplementary Materials

The following supporting information can be downloaded at the website of this paper posted on Preprints.org.

Author Contributions

Research design conceptualization: S.K.R. Data curation, all tabulations, formal analysis, introduction, materials and methods, results, interpretations, discussion, conclusions, and references: S.K.R. Figure 1: M.G.B. Figure 2, Figure 3, Figure 4, Figure 5, Figure 6, Figure 7, Figure 8, Figure 9, Figure 10, Figure 11, Figure 12, Figure 13, Figure 14, Figure 15, Figure 16, Figure 17 and Figure 18: S.K.R. Final manuscript drafting: S.K.R. and M.G.B. Statistical analysis: S.K.R. and T.E. Funding acquisition: S.K.R. and M.G.B.

Institutional Review Board Statement

This article is a scholarly report based on a literature review and published clinical and genetic reports on the pathogenicity and likely pathogenicity of MWS in previously deidentified patients.

Data Availability Statement

Because of their enormous size, Tables 2, 3, 4, 5, 7, and 10 are provided as Supplemental files.

Acknowledgments

Authors acknowledge receipt of support from the Mowat-Wilson Syndrome Foundation [MWSF] for S.K.R. We thank Dr. Tayyaba Ejaz [infoways9@gmail.com] for her expert statistical analysis of Figures 10-17 data.

Conflicts of Interest

The authors declare that there are no conflicts of interest. All authors have reviewed and approved the definitive version of the manuscript.

References

  1. DaSilva-Arnold, S.; Kuo, C.; Davra, V.; Remache, Y.; Kim, P.; Fisher, J.; Zamudio, S.; Al-Khan, A.; Birge, R.; Illsley, N. ZEB2, a master regulator of the epithelial-mesenchymal transition, mediates trophoblast differentiation. Mol. Hum. Reprod. 2019, 25(2). [Google Scholar] [CrossRef] [PubMed]
  2. Benito-Kwiecinski, S.; Giandomenico, S.L.; Sutcliffe, M.; Riis, E.S.; Freire-Pritchett, P.; Kelava, I.; Wunderlich, S.; Martin, U.; Wray, G.A.; McDole, K.; Lancaster, M.A. An early cell shape transition drives evolutionary expansion of the human forebrain. Cell 2021, 184(8), 2084–2102, e19. [Google Scholar] [CrossRef] [PubMed]
  3. Carulli, D.; et al. Semaphorins in Adult Nervous System Plasticity and Disease. Front. Synaptic Neurosci. 2021, 13, 672891. [Google Scholar] [CrossRef] [PubMed]
  4. Ricci, E.; Fettab, Anna; Garavellid, Livia; Vignoliu, Aglaia; Canevinia, Mariapaola; Duccio, M.; Cordelli, D.M. Further delineation and long-term evolution of electroclinical phenotype in Mowat Wilson Syndrome. A longitudinal study in 40 individuals. Epilepsy Behav. 2021, Volume 124, 108315. [Google Scholar] [CrossRef]
  5. Bassez, Guillaume; Camand, Olivier J.A; Cacheux, Valère; Kobetz, Alexandra; Dastot-Le Moal, Florence; Marchant, Dominique; Catala, Martin; Abitbol, Marc; Goossens, Michel. Pleiotropic and diverse expression of ZFHX1B gene transcripts during mouse and human development supports the various clinical manifestations of the “Mowat–Wilson” syndrome. Neurobiol. Dis. 2004, Volume 15(Issue 2), Pages 240–250. [Google Scholar] [CrossRef] [PubMed]
  6. Ivanovski, I.; Djuric, O.; Caraffi, S.G.; et al. Phenotype and genotype of 87 patients with Mowat-Wilson syndrome and recommendations for care. Genet Med 2018, 20, 965–75. [Google Scholar] [CrossRef] [PubMed]
  7. Garavelli, L.; Zollino, M.; Mainardi, P.C.; Gurrieri, F.; Rivieri, F.; Soli, F.; Verri, R.; Albertini, E.; Favaron, E.; Zignani, M.; Orteschi, D.; Bianchi, P.; Faravelli, F.; Forzano, F.; Seri, M.; Wischmeijer, A.; Turchetti, D.; Pompilii, E.; Gnoli, M.; Cocchi, G.; Mazzanti, L.; Bergamaschi, R.; De Brasi, D.; Sperandeo, M.P.; Mari, F.; Uliana, V.; Mostardini, R.; Cecconi, M.; Grasso, M.; Sassi, S.; Sebastio, G.; Renieri, A.; Silengo, M.; Bernasconi, S.; Wakamatsu, N.; Neri, G. Mowat-Wilson syndrome: facial phenotype changing with age: study of 19 Italian patients and review of the literature. Am. J. Med. Genet A 2009, 149A(3), 417–26. [Google Scholar] [CrossRef] [PubMed]
  8. St Peter, C.; Hossain, W.A.; Lovell, S.; Rafi, S.K.; Butler, M.G. Mowat-Wilson Syndrome: Case Report and Review of ZEB2 Gene Variant Types, Protein Defects and Molecular Interactions. Int. J. Mol. Sci. 2024, 25(5), 2838. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  9. Cordelli, D.M.; Garavelli, L.; Savasta, S.; Guerra, A.; Pellicciari, A.; Giordano, L.; Bonetti, S.; Cecconi, I.; Wischmeijer, A.; Seri, M.; Rosato, S.; Gelmini, C.; Della Giustina, E.; Ferrari, A.R.; Zanotta, N.; Epifanio, R.; Grioni, D.; Malbora, B.; Mammi, I.; Mari, F.; Buoni, S.; Mostardini, R.; Grosso, S.; Pantaleoni, C.; Doz, M.; Poch-Olivé, M.L.; Rivieri, F.; Sorge, G.; Simonte, G.; Licata, F.; Tarani, L.; Terazzi, E.; Mazzanti, L.; Cerruti Mainardi, P.; Boni, A.; Faravelli, F.; Grasso, M.; Bianchi, P.; Zollino, M.; Franzoni, E. Epilepsy in Mowat–Wilson syndrome: Delineation of the electroclinical phenotype. Am. J. Med. Genet Part A 2013, 161A, 273–284. [Google Scholar] [CrossRef] [PubMed]
  10. Diniz, L.P.; Matias, I.C.P.; Garcia, M.N.; Gomes, F.C.A. Astrocytic control of neural circuit formation: Highlights on TGF-beta signaling. Neurochem. Int. 2014, 78, 18–27. [Google Scholar] [CrossRef] [PubMed]
  11. Hegarty, S.V.; Wyatt, S.L.; Howard, L.; Stappers, E.; Huylebroeck, D.; Sullivan, A.M.; O’Keeffe, G.W. Zeb2 is a negative regulator of midbrain dopaminergic axon growth and target innervation. Sci. Rep. 2017, 7(1), 8568. [Google Scholar] [CrossRef] [PubMed]
  12. Hegarty, S.V.; Sullivan, A.M.; O’Keeffe, G.W. Zeb2: A multifunctional regulator of nervous system development. Prog. Neurobiol. 2015, 132, 81–95. [Google Scholar] [CrossRef] [PubMed]
  13. Rodriguez-Martinez, Griselda, Ivan Velasco, Activin and TGF-β Effects on Brain Development and Neural Stem Cells, CNS & Neurological Disorders—Drug Targets; Volume 11, Issue 7, Year 2012. [CrossRef] [PubMed]
  14. Yam, Patricia T; Charron, Frédéric. Signaling mechanisms of non-conventional axon guidance cues: the Shh, BMP and Wnt morphogens. Curr. Opin. Neurobiol. 2013, Volume 23(Issue 6), 965–973. [Google Scholar] [CrossRef] [PubMed]
  15. Higashi, Y.; Maruhashi, M.; Nelles, L.; Van De Putte, T.; Verschueren, K.; Miyoshi, T.; Yoshimoto, A.; Kondoh, H.; Huylebroeck, D. Generation of the floxed allele of the SIP1 (Smad-Interacting Protein 1) gene for Cre-mediated conditional knockout in the mouse. Genesis 2002, 32, 82–84. [Google Scholar] [CrossRef] [PubMed]
  16. De Putte, T.; Van Maruhashi, M.; Francis, A.; Nelles, L.; Kondoh, H.; Huylebroeck, D.; Higashi, Y. Mice lacking Zfhx1b, the gene that codes for Smad-interacting protein-1, reveals a role for multiple neural crest cell defects in the etiology of hirschsprung disease-mental retardation syndrome. Am. J. Hum. Genet. 2003, 72, 465–470. [Google Scholar]
  17. Adam, M.P.; Schelley, S.; Gallagher, R.; Brady, A.N.; Barr, K.; Blumberg, B.; Shieh, J.T.C.; Graham, J.; Slavotinek, A.; Martin, M.; Keppler-Noreuil, K.; Storm, A.L.; Hudgins, L. Clinical features and management issues in Mowat–Wilson syndrome. Am. J. Med. Genet Part A 2006, 140A, 2730–2741. [Google Scholar] [CrossRef]
  18. Mowat, D.R.; Wilson, M.J.; Goossens, M. Mowat-Wilson syndrome. J. Med. Genet. 2003, 40, 305–310. [Google Scholar] [CrossRef] [PubMed]
  19. Zweier, C.; Albrecht, B.; Mitulla, B.; et al. “Mowat-Wilson” syndrome with and without Hirschsprung disease is a distinct, recognizable multiple congenital anomalies-mental retardation syndrome caused by mutations in the zinc finger homeobox 1B gene. Am. J. Med. Genet 2002, 108, 177–181. [Google Scholar] [CrossRef] [PubMed]
  20. Silengo, M.; et al. Mowat–Wilson syndrome: clinical diagnosis in patients without Hirschsprung disease. Am. J. Med. Genet. Part A 2004, 127A(3), 318–322. [Google Scholar]
  21. Hossain, W.A.; St. Peter, C.; Lovell, S.; Rafi, S.K.; Butler, M.G. ZEB2 Gene Pathogenic Variants Across Protein-Coding Regions and Impact on Clinical Manifestations: A Review. Int. J. Mol. Sci. 2025, 26, 1307. [Google Scholar] [CrossRef] [PubMed]
  22. Ghoumid, J.; Drevillon, L.; Alavi-Naini, S.M.; Bondurand, N.; Rio, M.; Briand-Suleau, A.; Nasser, M.; Goodwin, L.; Raymond, P.; Yanicostas, C.; Goossens, M.; Lyonnet, S.; Mowat, D.; Amiel, J.; Soussi-Yanicostas, N.; Giurgea, I. ZEB2 zinc-finger missense mutations lead to hypomorphic alleles and a mild Mowat-Wilson syndrome. Hum. Mol. Genet. 2013, 22(13), 2652–61. [Google Scholar] [CrossRef] [PubMed]
  23. Birkhoff, J.C.; Huylebroeck, D.; Conidi, A. ZEB2, the Mowat-Wilson Syndrome Transcription Factor: Confirmations, Novel Functions, and Continuing Surprises. Genes 2021, 12, 1037. [Google Scholar] [CrossRef] [PubMed]
  24. Jumper, J.; Evans, R.; Pritzel, A.; et al. Highly accurate protein structure prediction with AlphaFold. Nature 2021, 596, 583–589. [Google Scholar] [CrossRef] [PubMed]
  25. Fleming, J.; Magana, P.; Nair, S.; Tsenkov, M.; Bertoni, D.; Pidruchna, I.; Lima Afonso, M.Q.; Midlik, A.; Paramval, U.; Žídek, A.; Laydon, A.; Kovalevskiy, O.; Pan, J.; Cheng, J.; Avsec, Ž.; Bycroft, C.; Wong, L.H.; Last, M.; Mirdita, M.; Steinegger, M.; Kohli, P.; Váradi, M.; Velankar, S. AlphaFold Protein Structure Database and 3D-Beacons: New Data and Capabilities. J. Mol. Biol. 2025, 437(15), 168967. [Google Scholar] [CrossRef] [PubMed]
  26. Eom, K.S.; Cheong, J.S.; Lee, S.J. Structural Analyses of Zinc Finger Domains for Specific Interactions with DNA. J. Microbiol. Biotechnol. 2016, 26(12), 2019–2029. [Google Scholar] [CrossRef] [PubMed]
  27. Remacle, S.; Kraft, H.; Lerchner, W.; Wuytens, G.; Collart, C.; Verschueren, K.; Smith, J. C.; Huylebroeck, D. SIP1, a novel zinc finger/homeodomain repressor, interacts with Smad proteins and binds to 5′-CACCT sequences in regulatory regions of target genes. Nucleic Acids Res. 1999, 27(23), 4928–4937. [Google Scholar] [CrossRef]
  28. Parfenyev, S.E.; Daks, A.A.; Shuvalov, O.Y.; et al. Dualistic role of ZEB1 and ZEB2 in tumor progression. Biol. Direct 2025, 20, 32. [Google Scholar] [CrossRef] [PubMed]
  29. Clark, David P.; Pazdernik, Nanette J.; McGehee, Michelle R. “Mutations and Repair”. In Molecular Biology; Elsevier, 2019; pp. 832–879. ISBN 9780128132883. [Google Scholar] [CrossRef]
  30. Mort, Matthew; Ivanov, Dobril; Cooper, David N.; Chuzhanova, Nadia A. “A meta-analysis of nonsense mutations causing human genetic disease”. Hum. Mutat. 2008, 29(8), 1037–47. [Google Scholar] [CrossRef] [PubMed]
  31. Hu, J.; Ng, P.C. Predicting the effects of frameshifting indels. Genome Biol. 2012, 13(2), R9. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  32. Losick, Richard; Watson, James D.; Baker, Tania A.; Bell, Stephen; Gann, Alexander; Levine, Michael W. Molecular biology of the gene, 6th ed.; Pearson/Benjamin Cummings: San Francisco, 2008; ISBN 978-0-8053-9592-1. [Google Scholar]
  33. Cox, Michael; Nelson, David R.; Lehninger, Albert L. Lehninger Principles of Biochemistry; W.H. Freeman: San Francisco, 2008; ISBN 978-0-7167-7108-1. [Google Scholar]
  34. Dastot-Le Moal, F.; Wilson, M.; Mowat, D.; et al. ZFHX1B mutations in patients with Mowat-Wilson syndrome. Hum. Mutat. 2007, 28, 313–21. [Google Scholar] [CrossRef] [PubMed]
  35. Meert, L.; Birkhoff, J.C.; Conidi, A.; Poot, R.A.; Huylebroeck, D. Different E-box binding transcription factors, similar neuro-developmental defects: ZEB2 (Mowat-Wilson syndrome) and TCF4 (Pitt-Hopkins syndrome). Rare Dis. Orphan Drugs J. 2022, 1, 8. [Google Scholar] [CrossRef]
  36. Bornelöv, S.; Reynolds, N.; Xenophontos, M.; Gharbi, S.; Johnstone, E.; Floyd, R.; Ralser, M.; Signolet, J.; Loos, R.; Dietmann, S.; Bertone, P.; Hendrich, B. The Nucleosome Remodeling and Deacetylation Complex Modulates Chromatin Structure at Sites of Active Transcription to Fine-Tune Gene Expression. Mol. Cell 2018, 71(1), 56–72.e4. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  37. Verstappen, G.; van Grunsven, L.A.; Michiels, C.; et al. Atypical Mowat-Wilson patient confirms the importance of the novel association between ZFHX1B/SIP1 and NuRD corepressor complex. Hum. Mol. Genet 2008, 17, 1175–83. [Google Scholar] [CrossRef] [PubMed]
  38. Wu, L.; Wang, J.; Conidi, A.; et al. Zeb2 recruits HDAC–NuRD to inhibit Notch and controls Schwann cell differentiation and remyelination. Nat. Neurosci. 2016, 19, 1060–1072. [Google Scholar] [CrossRef] [PubMed]
  39. Liu, X.; Hua, F.; Yang, D.; Lin, Y.; Zhang, L.; Ying, J.; Sheng, H.; Wang, X. Roles of neuroligins in central nervous system development: focus on glial neuroligins and neuron neuroligins. J. Transl. Med. 2022, 20(1), 418. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  40. Epifanova, E.; Babaev, A.; Newman, A.G.; Tarabykin, V. Role of Zeb2/Sip1 in neuronal development. Brain Res. 2019, 1705, 24–31. [Google Scholar] [CrossRef] [PubMed]
  41. Iuchi, S. Three classes of C2H2 zinc finger proteins. Cell Mol. Life Sci. 2001, 58(4), 625–35. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  42. Beth, Bragdon; Moseychuk, Oleksandra; Saldanha, Sven; King, Daniel; Julian, Joanne; Nohe, Anja. Bone Morphogenetic Proteins: A critical review. Cell. Signal. Vol. 2011, Volume 23(Issue 4), 609–620. [Google Scholar] [CrossRef] [PubMed]
  43. Conidi, A.; van den Berghe, V.; Leslie, K.; Stryjewska, A.; Xue, H.; Chen, Y.G.; Seuntjens, E.; Huylebroeck, D. Four Amino Acids within a Tandem QxVx Repeat in a Predicted Extended alpha-Helix of the Smad-Binding Domain of Sip1 Are Necessary for Binding to Activated Smad Proteins. PLoS ONE (print) 2013, 8(10). [Google Scholar] [CrossRef] [PubMed]
  44. Duboule, D. Guidebook to the Homeobox Genes; Oxford University Press: New York, 1994; p. 27±71. [Google Scholar]
  45. Verschueren, K.; Remacle, J.E.; Collart, C.; Kraft, H.; Baker, B.S.; Tylzanowski, P.; Nelles, L.; Wuytens, G.; Su, M.T.; Bodmer, R.; Smith, J.C.; Huylebroeck, D. SIP1, a novel zinc finger/homeodomain repressor, interacts with Smad proteins and binds to 5′-CACCT sequences in candidate target genes. J. Biol. Chem. 1999, 274(29), 20489–98. [Google Scholar] [CrossRef] [PubMed]
  46. Verschueren, Kristin; Remacle, DubouJacques E.; Collart, Clara; Kraft, Harry; Baker, Betty S.; Tylzanowski, f Przemko; Nelles, Luc; Wuytens, Gunther; Su, Ming-Tsan; Bodmer, h,i Rolf; Smith, James C.; Huylebroeck, Danny; Remacle, J.E.; Kraft, H.; Lerchner, W.; Wuytens, G.; Collart, C.; Verschueren, K.; Smith, J.C.; Huylebroeck, D. New mode of DNA binding of multi-zinc finger transcription factors: deltaEF1 family members bind with two hands to two target sites. EMBO J. 1999, 18(18), 5073–84. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  47. Postigo, A.A.; Dean, D.C. ZEB represses transcription through interaction with the corepressor CtBP. Proc. Natl. Acad. Sci. U S A 1999, 96(12), 6683–8. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  48. Zou, D.; Wang, L.; Wen, F.; Xiao, H.; Duan, J.; Zhang, T.; Yin, Z.; Dong, Q.; Guo, J.; Liao, J. Genotype-phenotype analysis in Mowat-Wilson syndrome associated with two novel and two recurrent ZEB2 variants. Exp. Ther. Med. 2020, 20(6), 263. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  49. Geiger, H.; Furuta, Y.; van Wyk, S.; Phillips, J.A., III; Tinker, R.J. The Clinical Spectrum of Mosaic Genetic Disease. Genes 2024, 15, 1240. [Google Scholar] [CrossRef] [PubMed]
  50. Xi, X.; Yang, X. Editorial on Genomic Mosaicism in Human Development and Diseases. Genes 2025, 16, 697. [Google Scholar] [CrossRef] [PubMed]
  51. Yang, X.; et al. ZEB2 is a master switch controlling the tumor-associated macrophage program. Cancer Cell 2025, 43(4), 789–804. [Google Scholar] [CrossRef] [PubMed]
  52. Caraffi, S.G.; van der Laan, L.; Rooney, K.; et al. Identification of the DNA methylation signature of Mowat-Wilson syndrome. Eur. J. Hum. Genet 2024, 32, 619–629. [Google Scholar] [CrossRef] [PubMed]
  53. Benjamini, Y.; Hochberg, Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J. R. Stat. Soc. Ser. B 1995, 57(1), 289–300. [Google Scholar] [CrossRef]
Figure 1. Facial frontal and profile views of a female with Mowat-Wilson syndrome showing medial flare of the eyebrows, depressed and wide nasal root, hypertelorism with downslanting palpebral fissures, ptosis, short prominent nose with a broad nasal tip, short philtrum, full everted lips, downturned corners of mouth, and posteriorly rotated ears with attached earlobes. Additional MWS findings not imaged include brain, cardiac, genitourinary, and kidney anomalies, Hirschsprung disease, and skeletal issues, along with moderate to severe intellectual disabilities, motor and speech delays, and seizures.
Figure 1. Facial frontal and profile views of a female with Mowat-Wilson syndrome showing medial flare of the eyebrows, depressed and wide nasal root, hypertelorism with downslanting palpebral fissures, ptosis, short prominent nose with a broad nasal tip, short philtrum, full everted lips, downturned corners of mouth, and posteriorly rotated ears with attached earlobes. Additional MWS findings not imaged include brain, cardiac, genitourinary, and kidney anomalies, Hirschsprung disease, and skeletal issues, along with moderate to severe intellectual disabilities, motor and speech delays, and seizures.
Preprints 222684 g001
Figure 2. ZEB2 gene: Isoform: O60315-1; Canonical; Genomic location: Chromosome 2:144,389,454-144,517,350; Reverse strand; GRCh38. Total number of exons: 10; Coding exons: 9 (2-10). The total length of the nine coding exons is 3,642 base pairs. Modified from Birkhoff, J.C. et al., 2021 [23].
Figure 2. ZEB2 gene: Isoform: O60315-1; Canonical; Genomic location: Chromosome 2:144,389,454-144,517,350; Reverse strand; GRCh38. Total number of exons: 10; Coding exons: 9 (2-10). The total length of the nine coding exons is 3,642 base pairs. Modified from Birkhoff, J.C. et al., 2021 [23].
Preprints 222684 g002
Figure 7. Spectrum of 3642 mutations in the nine coding exons of the ZEB2 gene: each mutation type is distinctly highlighted and color-coded as follows: 1. Frameshift 2. Inframe deletion3. In frame insertion 4. Missense 5. Protein altering variant. 6. Splice donor 7. Splice region 8. Start lost 9. Stop gained 10.Stop lost 11. Synonymous. Sources: NCBI-ClinVar (https://www.ncbi.nlm.nih.gov/clinvar/), and Gene Cards/MalaCards (https://www.genecards.org; malacards.org/).
Figure 7. Spectrum of 3642 mutations in the nine coding exons of the ZEB2 gene: each mutation type is distinctly highlighted and color-coded as follows: 1. Frameshift 2. Inframe deletion3. In frame insertion 4. Missense 5. Protein altering variant. 6. Splice donor 7. Splice region 8. Start lost 9. Stop gained 10.Stop lost 11. Synonymous. Sources: NCBI-ClinVar (https://www.ncbi.nlm.nih.gov/clinvar/), and Gene Cards/MalaCards (https://www.genecards.org; malacards.org/).
Preprints 222684 g007aPreprints 222684 g007b
Figure 9. Protein sequence of the canonical ZEB2 gene: Distribution of MWS pathogenic and likely or conflicting pathogenic amino acid variants from the 5′ N-terminal to the 3′ C-terminal: Exons 2-8: Number of amino acids: 1,214 [Uniprot # O60315-1].
Figure 9. Protein sequence of the canonical ZEB2 gene: Distribution of MWS pathogenic and likely or conflicting pathogenic amino acid variants from the 5′ N-terminal to the 3′ C-terminal: Exons 2-8: Number of amino acids: 1,214 [Uniprot # O60315-1].
Preprints 222684 g008
Figure 10. The bar diagram shows the reported incidence of 216 MWS pathogenicity along the entire length of the ZEB2 protein (n=216 cases; mean incidence=216/1200=0.180 cases per residue; 50-aa bin counts: mean=9.00, SD=3.58, range=2–17), except at the extreme tail end (aa1201–1214) (n=0/14 cases; mean=0.000; Welch t=-14.721, df=1199.000, p=3.25e-45) [see Supplemental Table 2 & Figure 9].
Figure 10. The bar diagram shows the reported incidence of 216 MWS pathogenicity along the entire length of the ZEB2 protein (n=216 cases; mean incidence=216/1200=0.180 cases per residue; 50-aa bin counts: mean=9.00, SD=3.58, range=2–17), except at the extreme tail end (aa1201–1214) (n=0/14 cases; mean=0.000; Welch t=-14.721, df=1199.000, p=3.25e-45) [see Supplemental Table 2 & Figure 9].
Preprints 222684 g009
Figure 11. The bar diagram shows the reported incidence of 67 Likely/Conflicting Reports of MWS pathogenicity across the ZEB2 protein (aa1-1200). The average incidence is 67/1200 = 0.0558 cases per residue (5.58% per aa). For 24 non-overlapping 50-aa bins (aa1-50 ... aa1151-1200), the bin counts have a mean of 2.79, a standard deviation of 2.08, and range from 1 to 7. The extreme C-terminal tail (aa1201-1214) contains zero cases (0/14 residues), indicating a significant depletion compared to aa1-1200 (Welch t=-8.420, df=1199, p≈1.06×10−16; two-tailed). [see Supplemental Table 4].
Figure 11. The bar diagram shows the reported incidence of 67 Likely/Conflicting Reports of MWS pathogenicity across the ZEB2 protein (aa1-1200). The average incidence is 67/1200 = 0.0558 cases per residue (5.58% per aa). For 24 non-overlapping 50-aa bins (aa1-50 ... aa1151-1200), the bin counts have a mean of 2.79, a standard deviation of 2.08, and range from 1 to 7. The extreme C-terminal tail (aa1201-1214) contains zero cases (0/14 residues), indicating a significant depletion compared to aa1-1200 (Welch t=-8.420, df=1199, p≈1.06×10−16; two-tailed). [see Supplemental Table 4].
Preprints 222684 g010
Figure 12. A bar diagram further illustrates the incidence of MWS pathogenicity across the six functional domains of the ZEB2 protein and within each of the seven C2H2-type zinc finger domains 1-7. There is a correlation between pathogenicity incidence and the functional domains of ZEB2, with a few exceptions.(domains covering aa1–1200: n=86/414 residues, mean=0.208; non-domain mean=0.165; Welch t=1.587, df=756.991, p=0.1130). [see Supplemental Table 2 & Figure 9].
Figure 12. A bar diagram further illustrates the incidence of MWS pathogenicity across the six functional domains of the ZEB2 protein and within each of the seven C2H2-type zinc finger domains 1-7. There is a correlation between pathogenicity incidence and the functional domains of ZEB2, with a few exceptions.(domains covering aa1–1200: n=86/414 residues, mean=0.208; non-domain mean=0.165; Welch t=1.587, df=756.991, p=0.1130). [see Supplemental Table 2 & Figure 9].
Preprints 222684 g011
Figure 13. The bar chart shows the distribution of 67 likely or conflicting MWS pathogenic cases across the functional domains of the ZEB2 gene [See Table 4]. [Within aa1-1200, residues in pooled functional domains have a slightly higher average incidence than residues outside these domains, but this difference is not statistically significant (domains: 27 cases/404 residues = 0.0668 cases per residue [6.68% per aa] vs. non-domains: 40/796 = 0.0503 [5.03% per aa]; Δ=1.65 percentage points; Welch t=1.131, df≈721.2, p=0.258; two-tailed).
Figure 13. The bar chart shows the distribution of 67 likely or conflicting MWS pathogenic cases across the functional domains of the ZEB2 gene [See Table 4]. [Within aa1-1200, residues in pooled functional domains have a slightly higher average incidence than residues outside these domains, but this difference is not statistically significant (domains: 27 cases/404 residues = 0.0668 cases per residue [6.68% per aa] vs. non-domains: 40/796 = 0.0503 [5.03% per aa]; Δ=1.65 percentage points; Welch t=1.131, df≈721.2, p=0.258; two-tailed).
Preprints 222684 g012
Figure 14. Frequency of MWS pathogenic variants in the 51-amino-acid SMAD-binding domain among 216 MWS cases with pathogenic variants [see Supplemental Table 2; Figure 3 and Figure 14].
Figure 14. Frequency of MWS pathogenic variants in the 51-amino-acid SMAD-binding domain among 216 MWS cases with pathogenic variants [see Supplemental Table 2; Figure 3 and Figure 14].
Preprints 222684 g013
Table 1. Shared clinical features of MWS with other genetic syndromes. Source: https://rarediseases.info.nih.gov/diseases/9673/mowat-wilson-syndrome.
Table 1. Shared clinical features of MWS with other genetic syndromes. Source: https://rarediseases.info.nih.gov/diseases/9673/mowat-wilson-syndrome.
Preprints 222684 i001
Table 6. below displays the exon sizes and coordinates for ZEB2, including each exon’s number and length in base pairs and amino acids.
Table 6. below displays the exon sizes and coordinates for ZEB2, including each exon’s number and length in base pairs and amino acids.
Preprints 222684 i002
Table 9. List of 7 Missense Variants Among the 216 Cases With MWS Pathogenic Diagnosis and 40 Cases With Missense Variants Among the 67 Cases With Likely or Conflicting Pathogenic Variants.
Table 9. List of 7 Missense Variants Among the 216 Cases With MWS Pathogenic Diagnosis and 40 Cases With Missense Variants Among the 67 Cases With Likely or Conflicting Pathogenic Variants.
Preprints 222684 i003Preprints 222684 i004
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings