Preprint
Article

This version is not peer-reviewed.

Trio Genome Sequencing Identifies Diagnostic and Candidate Genes in Neurodevelopmental Disorder Unresolved by Prior Testing

Submitted:

16 August 2026

Posted:

17 August 2026

You are already at the latest version

Abstract
Approximately 75% of individuals with neurodevelopmental disorders (NDD) remain without a molecular diagnosis after first-tier genetic testing, and a substantial share of that gap reflects the pace of gene-disease discovery rather than sequencing technology alone (Stefanski et al., 2021). Here we report trio genome sequencing in 36 probands with NDD or epilepsy who remained molecularly unsolved despite prior genetic testing. Likely pathogenic variants were identified in thirteen probands (36%): eight confirmed diagnoses and five high-priority findings pending Sanger validation, including a deep intronic PPP2R5C variant that extends the gene's recognised mutational spectrum. A further 19 candidate genes without an established disease association were prioritised by evidence score, including PRKAR1B, identified independently as a de novo duplication in two unrelated probands but not counted as a diagnosis pending clarification of its disease mechanism. Grouping all 34 findings by functional theme showed genes involved in transcription and chromatin regulation as the largest category. Only 9% of findings were in genome regions inaccessible to exome sequencing, indicating that most diagnostic value reflected thorough, trio-based analysis rather than access to genome-specific sequence. These findings support genome sequencing as a second-tier diagnostic step after negative prior testing, while showing that reassessing gene-disease validity, not sequencing technology, accounted for part of this yield.
Keywords: 
;  ;  ;  ;  ;  

Introduction

Neurodevelopmental disorders, including intellectual disability, developmental delay, autism spectrum disorder, and epilepsy, affect a substantial proportion of children and have highly heterogeneous genetic aetiologies. Exome sequencing (ES) has become a standard diagnostic tool for these conditions, but yield varies widely for each phenotypic group and remains modest overall: a meta-analysis of 103 studies covering over 32,000 individuals reported an overall diagnostic yield of 23.7% across neurodevelopmental disorders, highest for intellectual disability (28.2%) and lowest for autism spectrum disorder (17.1%), with having an intermediate rate overall (24%) that varied substantially by subtype, from 9.3% in epilepsy without intellectual disability to 27.9% in epilepsy with intellectual disability and up to 36.8% in developmental and epileptic encephalopathies (Stefanski et al., 2021). The majority of individuals referred for genetic testing therefore remain without a molecular diagnosis after first-tier testing.
Cases of NDD that remain unresolved after ES or panel testing do not all fail for the same reason. Some carry variants in classes that ES-based methods capture poorly, such as deep intronic changes, structural variants below the resolution of chromosomal microarray, or variants in regions of uneven exome capture. Others carry a variant in a gene not yet recognised as disease-causing at the time of testing. Hundreds of new gene-disease relationships are published each year (Welland et al., 2026), so a variant that was correctly called but filtered out for lack of an established gene-disease association can remain a missed diagnosis until the literature, and the databases built on it, catch up. Systematic reanalysis of existing sequencing data against an updated evidence base has itself been shown to increase diagnostic yield, independent of any change in sequencing technology. A recent large-scale automated reanalysis programme identified new diagnoses in 5% of a previously undiagnosed cohort, with roughly half attributable to new gene- or variant-disease evidence rather than improved analysis alone (Welland et al., 2026). Distinguishing which of these mechanisms is at work in an unresolved cohort matters for how that cohort should be approached: sequencing more of the genome (moving from ES to GS), adopting a trio-based sequencing approach, more thorough reanalysis of existing data, or both.
Here we report trio genome sequencing (GS) analysis in a cohort of 40 probands with neurodevelopmental disorders who remained unresolved after prior genetic testing. We identified pathogenic variants in well-established disease-associated genes, and in recently established diagnostic genes. We prioritise candidate pathogenic variants in genes without an established disease association, using a structured, rule-based, six-domain evidence-scoring system spanning inheritance, variant effect, phenotype match, biology, constraint, and validation. Proband-level outcome classification uses a five-tier framework described in full elsewhere (proband tiering classification, manuscript in preparation); we report only the tier definitions needed to interpret confirmed, pending, and candidate findings here. We further examine how many findings depended specifically on genome-wide coverage rather than systematic reanalysis, and whether the genes identified converge on shared biological themes.

Materials and Methods

Study Design and Cohort

This was a collaborative cohort study across three international sites (Adelaide Australia, Dianalund Denmark, Tokyo Japan), each contributing probands with neurodevelopmental disorders (NDDs) and epilepsy referred for diagnostic genome sequencing (GS); genomic sequencing was conducted commercially at the Queensland University of Technology’s Central Analytic Research Facility and bioinformatic analysis was performed centrally at Adelaide University. Forty unrelated probands with neurodevelopmental phenotypes underwent trio GS. Eligible probands had intellectual disability, developmental delay, autism spectrum disorder, epilepsy, or multiple congenital anomalies, and remained unresolved after prior genetic testing (karyotyping, chromosomal microarray, and/or targeted gene panel). After complete variant analysis, four probands were excluded due to consanguinity, leaving 36 probands in the analytical cohort (19 female, 17 male). Twenty-six of the 36 analytical probands (72%) had documented prior genetic testing before GS enrolment, including trio or single/proband ES (n = 19), 600-gene epilepsy panels, chromosomal microarray, and syndrome-specific testing (methylation analysis for Angelman syndrome, MECP2 sequencing, transferrin isoform analysis, and 7-dehydrocholesterol testing) (Supplementary Table S1). This study was approved by the Human Research Ethics Committee at Adelaide University [HREC protocol 000000329987], with written informed consent obtained from participants or their legal guardians, at the sites of recruitment.

Phenotypic Assessment

Clinical features were mapped to Human Phenotype Ontology (HPO) terms (version 2024-02) by clinician-documented assessment. Phenotype matching for candidate gene assessment (see below) used clinician-documented HPO terms; terms identified retrospectively from gene-disease association literature were tracked separately and reserved for hypothesis generation, not for evidence scoring. Provenance was recorded by set difference: Clinician_HPO comprised all terms from direct clinical assessment; Combined_HPO comprised all terms appearing in variant annotation records; Datamined_HPO was defined as Combined_HPO minus Clinician_HPO.

Genome Sequencing and Variant Calling

Sequencing was performed on an Illumina NovaSeq 6000 using PCR-free libraries with 2 × 150 bp paired-end reads. Mean sequencing depth was 38×, with 95% of bases covered at >20×. Reads were aligned to GRCh38 with novoAlign v4.018 (Novocraft Technologies, 2022). Duplicate marking, base quality score recalibration, and indel realignment were not performed; PCR-free library preparation and the pedigree-aware variant calling workflow below were relied on for base-level accuracy.
Small variants (single-nucleotide variants and indels) were called with GATK HaplotypeCaller v4.1.4.0 (McKenna et al., 2010) in pedigree-aware gVCF mode, followed by joint genotyping across all trios. Hard filtering was used in place of Variant Quality Score Recalibration, which requires training cohorts unavailable for rare NDD studies. HGVS nomenclature was verified with VariantValidator (Freeman et al., 2018) before inclusion in tables; MANE Select or MANE Plus Clinical transcripts were adopted at reporting where available.
Variant prioritisation used slivar v0.2.7 (Pedersen et al., 2021) with the following filters: gnomAD v3.1 (Karczewski et al., 2020) population maximum allele frequency <0.001; HIGH or MODERATE Variant Effect Predictor consequence; minimum depth and genotype-quality thresholds; and inheritance-pattern filters (de novo, autosomal recessive, compound heterozygous, X-linked recessive, X-linked de novo).
Copy-number variants and structural variants were called with ClinSV v1.1 (Minoche et al., 2021), which executes Lumpy and CNVnator to produce candidate calls from split-read, discordant-pair, and depth-of-coverage evidence, then merges and annotates calls with automated quality tranches (High, Pass, Low). Duphold filtering was applied to remove spurious duplication calls. Structural variants were classified as GS-unique if they were smaller than 5 kb; variants located more than 20 bp from an exon boundary, and variants in non-coding RNA genes, were also classified as GS-unique. Structural variants ≥5 kb were classified as non-GS-unique, since they fall within the detection resolution of chromosomal microarray. Prior exome VCF data were unavailable for this cohort, so GS-unique status is an inference rather than an empirical comparison against prior testing.

Variant Classification and Gene-Disease Validity Assessment

Single-nucleotide variants and indels were classified as pathogenic, likely pathogenic, variant of uncertain significance (VUS), likely benign, or benign according to American College of Medical Genetics and Genomics/Association for Molecular Pathology (ACMG/AMP) criteria (Richards et al., 2015), applied using InterVar (Li & Wang, 2017). Copy-number variants and structural variants were classified according to ACMG guidelines for copy-number variant interpretation. Population frequency data were sourced from gnomAD v3.1 (Karczewski et al., 2020). In silico prediction thresholds were CADD PHRED score ≥20 and SpliceAI delta score ≥0.2 for splice-altering predictions.
Gene-disease associations were assessed using the ClinGen Gene-Disease Validity Classification framework (Definitive, Strong, Moderate, Limited, Disputed, Refuted, No Known; ClinGen Consortium, 2025), cross-referenced with OMIM, Orphanet, and Gene2Phenotype. For genes without established ClinGen curation, a literature-based assessment assigned an equivalent validity category based on the number of independent probands reported, functional evidence, and phenotypic specificity.
Phenotype overlap between a proband and a candidate gene-disease association was quantified as recall: the proportion of the gene-disease association’s core HPO features (drawn from GeneReviews or OMIM clinical synopses) present in the proband’s clinician-documented HPO set. Full match (≥70% recall) indicated that all core gene-disease features were present in the proband’s documented phenotype; partial match (30–69%) indicated at least one missing core feature reported in >80% of published cases; limited match (<30%) indicated only non-specific overlap. These thresholds were set by expert consensus among three reviewers.

Proband-Level Classification

Each proband was assigned to one of five tiers (confirmed diagnosis, provisional, candidate with high priority, candidate with low priority, or unresolved) using a previously described framework that integrates ACMG/AMP classification, ClinGen gene-disease validity, phenotype match, independent validation status, and genome-specific detectability (proband tiering classification, manuscript in preparation). In brief, Tier 1 (confirmed diagnosis) required a pathogenic or likely pathogenic variant in a Definitive or Strong ClinGen gene with full phenotype match and completed independent validation; Tier 4 (candidate) comprised variants of uncertain significance, or variants in genes with Moderate, Limited, or No Known ClinGen validity, that did not meet the criteria for Tiers 1–3. Two independent reviewers assigned tiers, with disagreements resolved by consensus with a third senior reviewer.

Gene Evidence Scoring

For Tier 4 candidates in genes without established ClinGen-validated association with NDD or epilepsy, a structured, rule-based, six-domain evidence-scoring system was applied to prioritise candidates for downstream functional validation (Table 1). Each domain (inheritance, variant effect, phenotype match, biology, constraint, and validation) was scored from 0 to 3 against defined criteria and summed to a total evidence score with a range of 0–18; candidate genes with a total score of 7 or more were considered high priority for downstream functional validation. Constraint metrics were not available in gnomAD v3 for a subset of candidate genes, which were scored 0 for that domain and flagged accordingly rather than left blank.

Functional Theme Classification

All 34 reportable genetic findings (confirmed, pending, and candidate) were classified into functional or pathway-based themes by manual curation of the primary literature and gene function annotation (OMIM, UniProt, GeneCards), extending an eight-category scheme (cytoskeletal and axonal biology; ion channels and neurotransmission; neural development and stem cell regulation; signalling and kinase pathways; synaptic and endosomal biology; transcription and chromatin regulation; protein synthesis and translation; emerging or novel) with two additional categories, metabolic and ubiquitin-proteasome system, that were required to accommodate genes not covered by the original scheme. Each finding was assigned to a single best-fitting theme; genes with genuinely dual or ambiguous function (e.g., ciliary genes with cytoskeletal and signalling roles) were assigned by their most literature-supported primary function.

Independent Validation and Segregation Analysis

All Tier 1 and Tier 2–3 variants were validated by Sanger sequencing or an independent method before final tier assignment; Tier 4 variants were validated by Sanger sequencing where primer design was feasible. Structural variants were confirmed by quantitative PCR, multiplex ligation-dependent probe amplification, or chromosomal microarray where appropriate. Validation status was recorded as confirmed (Sanger or independent method concordant with the GS call), pending (validation in progress), false positive (GS call not confirmed by Sanger sequencing), or failed (primer design or PCR amplification unsuccessful). Discrepant Sanger trace reads were reviewed manually before a variant was reclassified as false positive; in one instance, an apparent discrepancy of a candidate indel was resolved as a difference in indel left- versus right-alignment convention between the variant caller and the Sanger trace, rather than true discordance.

Data Sharing

Novel candidate genes will be registered in Matchmaker Exchange (Philippakis et al., 2015) and GeneMatcher (Sobreira et al., 2015) to facilitate identification of additional probands with overlapping phenotypes and variants. Confirmed and candidate variants will be submitted to ClinVar. Gene-disease validity was cross-referenced against ClinGen, OMIM, and PubMed during manuscript preparation.

Statistical Analysis

Descriptive statistics were reported as median (range) for continuous variables and n (%) for categorical variables. Diagnostic yield was calculated as the proportion of probands with a confirmed or high-priority pending diagnostic gene finding (Table 3).

Results

Cohort Characteristics

Forty unrelated probands with neurodevelopmental phenotypes underwent trio genome sequencing analysis. After complete variant analysis, four probands were excluded due to consanguinity, leaving 36 probands (19 female, 17 male) in the analytical cohort (Figure 1, Table 2). Intellectual disability was the most frequent clinical feature, documented in 29 of 36 probands (81%), followed by seizures (29/36, 81%) and developmental delay (19/36, 53%). Autism spectrum disorder or behavioural features were documented in 11 probands (31%) and developmental regression in seven (19%). Twenty-six probands (72%) had documented prior genetic testing before genome sequencing enrolment, including 19 (53%) with prior exome sequencing-based analysis. Probands carried a median of nine HPO terms each (range 1–16).

Diagnostic Yield

Thirteen of 36 probands (36%) received a confirmed or high-priority pending diagnostic gene finding: eight confirmed and five pending Sanger validation (Table 3). Under the companion proband-tiering framework, seven probands (19%) met formal Tier 1 (confirmed diagnosis) criteria and five (14%) met formal Tier 3 (candidate, high priority, unvalidated) criteria, defined as a pathogenic or likely pathogenic variant in a Definitive or Strong ClinGen gene with independent validation pending. An eighth proband, PB39 (PPP2R5C), met the same evidentiary bar for confirmed diagnosis on gene-level curation for this analysis: the variant was pathogenic with full phenotype match and a Sanger-confirmed call, and PPP2R5C gene-disease validity was assessed as Definitive rather than the Limited classification carried in the source ClinGen annotation (see Methods and Discussion). One Tier 3 proband, PB41, carries two independent pending findings (CNKSR2 and CHD3). Two of the eight confirmed genes recurred across unrelated probands: TAF1 (PB02, PB36), both X-linked, syndromic intellectual disability.
Table 3. Confirmed and high-priority pending diagnostic gene findings (n = 13 probands: 8 confirmed, 5 pending Sanger validation).
Table 3. Confirmed and high-priority pending diagnostic gene findings (n = 13 probands: 8 confirmed, 5 pending Sanger validation).
Proband Gene Variant (HGVS c.) Protein change ACMG Inheritance Disease association Status
PB02 TAF1 c.422C>T p.(Pro141Leu) LP X-linked recessive X-linked intellectual disability, syndromic Confirmed
PB04 KAT6B c.5693G>A p.(Arg1898Gln) P De novo Genitopatellar syndrome / SBBYS syndrome Confirmed
PB07 ARID1B c.157207431_157207432delAC* p.? P De novo Coffin-Siris syndrome Confirmed
PB08 GLUL c.3G>C p.(Met1?) LP De novo GLUL deficiency, epileptic encephalopathy Confirmed
PB25 HERC2 c.6555+3G>A p.? P Autosomal recessive HERC2-related intellectual disability Confirmed
PB36 TAF1 c.4753+7_4753+8del p.? P X-linked recessive X-linked intellectual disability, syndromic Confirmed
PB48 MED12 c.6309_6324del p.(Gln2103HisfsTer111) P X-linked recessive Lujan-Fryns / Ohdo syndrome Confirmed
PB39 PPP2R5C c.1608+3140_3141insTTT p.? P De novo Houge-Janssens syndrome 4 Confirmed
PB20 INTS1 c.5970C>G; c.4969C>T p.(Phe1990Leu); p.(Arg1657Cys) LP Compound het. NDD with cataracts (partial) Pending Sanger
PB24 ALG11 c.44G>C; c.-474G>A p.(Arg15Thr); p.? LP Compound het. Congenital disorder of glycosylation Pending Sanger
PB41 CNKSR2 c.2634_2637del p.(Glu879ArgfsTer15) LP X-linked recessive X-linked intellectual disability 46 Pending Sanger
PB41 CHD3 c.277+2T>G p.? LP De novo Snijders Blok-Campeau syndrome Pending (primer redesign)
PB42 PLA2G6 chr22:38130547–38131253 (DEL) p.? (SV) P De novo Neurodegeneration with brain iron accumulation Pending Sanger
PB43 GABRG2 c.316G>A p.(Ala106Thr) LP De novo DEE74 / GEFS+3 / childhood absence epilepsy 2 Pending Sanger
*Genomic coordinate (NC_000006.13); no MANE Select transcript-level HGVS_c was available for this variant. ACMG, American College of Medical Genetics and Genomics classification (P, pathogenic; LP, likely pathogenic). “Pending” status rows carry a pathogenic or likely pathogenic variant in a Definitive or Strong ClinGen gene with partial-to-full phenotype match, pending independent validation; PB41 has two independent pending findings. Full variant-level annotation, including in silico predictions, is provided in Supplementary Table S2.

Confirmed and High-Priority Pending Diagnostic Genes

The eight pathogenic variants identified were: four de novo variants (in KAT6B, ARID1B, GLUL, PPP2R5C), three X-linked variants (two in TAF1, one in MED12), and one autosomal recessive variant in (HERC2, homozygous). All eight had full phenotype match to their associated disease and a Sanger-confirmed or independently validated call.
The five pending findings were in Definitive or Strong ClinGen genes, but await Sanger confirmation (INTS1, ALG11, CNKSR2, PLA2G6, GABRG2) or primer redesign after an initial failed attempt (CHD3). Phenotype match for these five was full in two (ALG11, CHD3) and partial in the remainder.
The PPP2R5C finding in PB39 is a deep intronic insertion (c.1608+3140_3141insTTT). The original description of PPP2R5C-related Houge-Janssens syndrome 4 reported de novo missense or small in-frame variants in all 26 published individuals; a deep intronic mechanism has not previously been reported.

Candidate Genes Without Established Disease Association

Applying the six-domain evidence-scoring system (Table 1) to Tier 4 variants in genes without established ClinGen-validated NDD or epilepsy association, identified 19 candidate genes across 20 proband-gene findings (one gene, PRKAR1B, recurred in two unrelated probands; Table 4). Sixteen of the 19 candidate genes (84%) scored 7 or more and were considered high priority for downstream functional validation. The highest-scoring candidates were RRAGB (PB29, X-linked recessive, 11/18) and TRPC5 (PB46, X-linked recessive, 11/18), followed by RNF123 and TANC1 (10/18 each). Nine of the 19 candidate genes (WASHC2C, LRIG1, HYDIN, TENM4, FAT2, TENT4A, PRKAR1B × 2, CDH4, ADGRG4) lack gene constraint data in gnomAD v3 and were scored 0 for that domain by definition; their evidence scores are likely underestimates of true priority rather than a reflection of weak candidacy.

Variant Classes Uniquely Accessible to Genome Sequencing

Applying the pre-specified GS-unique criteria (Methods) to all 34 reported findings, three (9%) were classified as GS-unique: a deep intronic insertion in PPP2R5C (confirmed, PB39), a structural variant in PLA2G6 (pending, PB42), and a deep intronic variant in RRAGB (candidate, PB29). The remaining 31 findings (91%), including all but one of the confirmed diagnostic genes, fell within the detection resolution of ES or chromosomal microarray on this classification (see Discussion).

Functional Convergence Across Findings

Classifying all 34 reported findings by functional theme (Methods) showed that transcription and chromatin regulation was the largest single category, accounting for seven findings, all among the confirmed and pending diagnostic genes and none among the candidates (TAF1 × 2, KAT6B, ARID1B, MED12, INTS1, CHD3; Figure 2). Signalling and kinase pathways was the largest theme among candidate genes, comprising six findings spread across all three confidence tiers (PPP2R5C, CNKSR2, RRAGB, WNT16, PRKAR1B × 2), and metabolic genes accounted for a further five (GLUL, ALG11, PLA2G6, BHMT, FAM104A/C17orf80). No candidate gene evidence score domain (Table 1) was directly informed by thematic grouping; theme assignment was performed independently of, and after, evidence scoring.

A Recurrent, Brain-Enriched Candidate with an Unresolved Mechanism

PRKAR1B was identified as a candiate gene independently in two unrelated probands (PB44, PB46), both with de novo large duplications (21–25 kb). PRKAR1B is brain-enriched in expression and is established as the gene underlying Marbach-Schaaf neurodevelopmental syndrome. This is an autosomal dominant disorder caused by de novo missense variants and characterised by global developmental delay, autism spectrum disorder, attention-deficit/hyperactivity disorder, and apraxia. No duplication of PRKAR1B has previously been reported in association with this disorder. Both findings are reported as candidates rather than diagnoses (see Discussion).

Phenotype Match

Phenotype match was full in all eight confirmed diagnostic gene findings, and full in two of the five pending findings. Among the 20 candidate-gene findings, phenotype match was full in one (WASHC2C, PB06) and limited in the remaining 19 (95%), consistent with the exploratory, hypothesis-generating status of genes without an established disease association against which to benchmark expected clinical features.

Discussion

Trio genome sequencing analysis in 36 probands with neurodevelopmental disorders unresolved by prior testing yielded a confirmed or high-priority pending diagnostic gene finding in 13 cases (36%), and at least one candidate gene in a further 15. The eight confirmed diagnoses involve established disease genes, including TAF1 in two unrelated probands, and one recently established gene finding, PPP2R5C, that extended the gene’s known mutational spectrum. Nineteen candidate genes without established ClinGen-validated NDD or epilepsy association were prioritised using a structured, six-domain evidence-scoring system, of which 16 met a pre-specified high-priority threshold.
Our confirmed diagnostic yield (8/36, 22%) is comparable to two recent trio-genome sequenced cohorts, also analysed after negative prior genetic testing. López-López et al. (2026) reported a yield of 20.98% among 221 NDD probands with negative targeted sequencing, using solo ES followed by trio pooled-ES. Malmgren et al. (2025) reported a yield of 30% among 321 previously tested, unsolved trios subsequently undergoing GS or ES at a single Swedish centre, compared with 47% among trios with no prior analysis. This gap illustrates how cohort selection (tested-negative versus treatment-naïve) shapes yield comparisons across studies. A third comparable study, Spirito et al. (2026), reported a higher combined yield of 26.4% in an unselected genome-sequencing NDD cohort in Italy. That figure included high-confidence variants of uncertain significance alongside pathogenic and likely pathogenic calls, a broader inclusion criterion than applied here. Together, these comparisons place our yield within, though at the lower end of, published expectations for pre-tested and unresolved NDD cohorts; cohorts without that selection filter tend to report higher rates.
The PPP2R5C finding illustrates a practical consequence of the pace at which gene-disease curation now moves relative to sequencing throughput: the underlying association, Houge-Janssens syndrome 4, was described in the same calendar year this variant was called (Verbinnen et al., 2025), and the ClinGen validity label available for this analysis had not yet been updated to reflect that publication. The variant itself, a deep intronic insertion, was not among the missense or small in-frame changes reported in the original 26-individual description, extending the recognised mutational mechanism for this gene. A similar lag affected RPS23, a Tier 4 candidate whose ClinGen label read “No Known” despite an OMIM-catalogued disease association dating to 2017 (Paolini et al., 2017); the proband’s phenotype extends beyond that condition’s classically described triad, and is reported here as a phenotype-expansion candidate rather than a confirmed diagnosis pending further clinical correlation. Both cases argue for treating externally sourced gene-disease validity labels as provisional and re-checking them against current literature at the point of manuscript preparation, rather than relying on a single upstream annotation pass.
The recurrence of TAF1 pathogenic variant in two unrelated probands, once as a missense substitution and once as a 2-nucleotide deletion approximately 10 bp into the adjacent intron, is consistent with its definitive, well-established role in X-linked syndromic intellectual disability. The second variant was flagged as not GS-unique on the criteria applied here (Methods), meaning it plausibly sat within reach of ES-based capture, and it was typical of the cohort in that respect: only 3 of 34 reported findings (9%) met the pre-specified GS-unique criteria (Results). The diagnostic value of genome sequencing in this cohort therefore lay predominantly in the trio-based, pedigree-aware calling and systematic reanalysis applied after earlier non-diagnostic testing, rather than in reaching variant classes ES-based methods cannot in principle detect; the exceptions, a deep intronic finding in each of PPP2R5C, PLA2G6, and RRAGB, show that genome-wide coverage still mattered for a minority of findings, including the leading candidate gene.
Grouping all 34 findings by functional theme (Figure 2) showed that transcription and chromatin regulation was both the largest theme overall and, notably, the only one populated entirely by confirmed and pending diagnostic genes rather than candidates (TAF1 × 2, KAT6B, ARID1B, MED12, INTS1, CHD3). Chromatin- and transcription-related genes are among the most represented categories of disease-causing genes in neurodevelopmental disorder more broadly (Lintas et al., 2022), and this pattern is consistent with, though not a formal test of, that established prominence. By contrast, signalling and kinase pathways, the largest theme among candidates, spanned all three confidence tiers and included both the strongest formally-scored candidates (RRAGB, TRPC5) and the most biologically compelling one (PRKAR1B), suggesting this theme may be a productive area for continued case-matching and functional follow-up even though no single signalling-pathway gene here has yet reached diagnostic confidence.
The most biologically compelling candidate, PRKAR1B, is also the one whose interpretation is least settled. It was identified independently as a de novo duplication in two unrelated probands, and is the established gene underlying Marbach-Schaaf neurodevelopmental syndrome (Marbach et al., 2021), a disorder caused, in every case reported to date, by de novo missense variants. Its evidence score (9/18, high priority) is itself likely an underestimate, since gnomAD v3 constraint data remain unavailable for the gene; the de novo occurrence of both duplications independently is difficult to attribute to chance, and supports pathogenicity as a general proposition. What it does not resolve is mechanism: missense and copy-number changes are not interchangeable by default, and a duplication could plausibly act through increased gene dosage, through a distinct mechanism unrelated to the missense phenotype, or, despite its de novo origin, not be the cause of the phenotype at all. We have not counted either finding as a diagnosis, and functional evidence, together with additional matched cases through GeneMatcher, will be needed before that question can be resolved.
Several limitations bear on how these findings should be read. Gene-disease validity is a moving target (McGlaughon et al., 2018), as the PPP2R5C and RPS23 findings demonstrate directly; candidate classifications reported here reflect validity as assessed for this analysis and will require periodic reanalysis. The cohort is small (36 analytical probands), which limits the precision of yield estimates, and gnomAD v3 constraint data were unavailable for nine of the 19 candidate genes, so those evidence scores likely underestimate true priority rather than reflect weak candidacy. Finally, no candidate gene reported here has functional evidence beyond in silico prediction and population constraint; the evidence-scoring system is a prioritisation aid, not a substitute for experimental validation. Additional limitations, including HPO term provenance, site-level considerations, and the validation status of individual pending findings, are detailed in Supplementary Note 1.
Taken together, these findings support genome sequencing as a second-tier diagnostic step after negative prior testing in neurodevelopmental disorders, while confirming that a meaningful share of the yield depends on active curation rather than sequencing technology alone. The 16 high-priority candidate genes identified here, including PRKAR1B, will be registered in Matchmaker Exchange and GeneMatcher to support case-matching; functional characterisation of the PRKAR1B duplications, in particular, would clarify whether copy-number gain and missense substitution converge on the same neurodevelopmental phenotype through a shared mechanism.

Supplementary Materials

The following supporting information can be downloaded at website of this paper posted on Preprints.org.

Author Contributions

Conceptualization: A.S.; Data curation: A.S.; Formal analysis: A.S.; Funding acquisition: L.M.D.; Clinical Investigation: R.M., N.S., T.Y.; Genetic Investigation: A.S., M.G.R., L.M.D.; Methodology: A.S.; Project administration: A.S., M.G.R.; Resources: A.S., M.G.R., L.M.D.; Software: A.S.; Supervision: M.G.R., L.M.D.; Validation: Z.S., R.H., S.S.; Visualization: A.S.; Writing—original draft: A.S.; Writing—review and editing: M.G.R., L.M.D.

Institutional Review Board Statement

This study was approved by the Adelaide University Human Research Ethics Committee (HREC protocol 000000329987). Written informed consent was obtained from all participants or their legal guardians.

Data Availability Statement

Original contributions are included in this article and its supplementary information. Additional data are available from the corresponding author on reasonable request.

Acknowledgments

We wish to thank Illumina, Michael Fietz and Nicola Withers for their assistance and support. Declaration of generative AI and AI-assisted technologies in the writing process: During the preparation of this work, the authors used Claude (Anthropic) for grammar checking, reference formatting, and editorial restructuring assistance. The authors reviewed and edited all AI-assisted content and take full responsibility for the content of the published article. AI tools were not used for data analysis, image generation, clinical decision-making, variant classification, or phenotype assignment.

Conflicts of Interest

A.S. is a PhD student at Adelaide University (formerly the University of South Australia) and an employee of Novocraft Technologies, the developer of novoAlign, which was used in this study. The remaining authors declare no conflicts of interest.

References

  1. ClinGen Consortium. (2025). The Clinical Genome Resource (ClinGen): Advancing genomic knowledge through global curation. Genetics in Medicine, 27, Article 101228. [CrossRef]
  2. Freeman, P. J., Hart, R. K., Gretton, L. J., Brookes, A. J., & Dalgleish, R. (2018). VariantValidator: Accurate validation, mapping, and formatting of sequence variation descriptions. Human Mutation, 39(1), 61–68. [CrossRef]
  3. Karczewski, K. J., Francioli, L. C., Tiao, G., et al. (2020). The mutational constraint spectrum quantified from variation in 141,456 humans. Nature, 581, 434–443. [CrossRef]
  4. Li, Q., & Wang, K. (2017). InterVar: Clinical interpretation of genetic variants by the 2015 ACMG-AMP guidelines. American Journal of Human Genetics, 100(2), 267–280. [CrossRef]
  5. Lintas, C., Bottillo, I., Sacco, R., Azzarà, A., Cassano, I., Ciccone, M. P., Grammatico, P., & Gurrieri, F. (2022). Expanding the spectrum of KDM5C neurodevelopmental disorder: A novel de novo stop variant in a young woman and emerging genotype–phenotype correlations. Genes, 13(12), Article 2266. [CrossRef]
  6. López-López, L., Lapeña-Gil, L., Benítez, Y., Serrano, C., Sánchez-Barbero, A. I., Blanco-Kelly, F., López-Grondona, F., Tahsin-Swafiri, S., Lorda-Sánchez, I., Ayuso, C., Mínguez, P., & Almoguera, B. (2026). Accurate and cost-effective workflow integrating trio pooled-WES for novel gene discovery in neurodevelopmental disorders. European Journal of Human Genetics, 34(5), 675–682. [CrossRef]
  7. Malmgren, H., Kvarnung, M., Gustafsson, P., Anderlid, B.-M., Arthur, C., Carlsten, J., De Geer, K., Ehn, E., Grigelioniéné, G., Hammarsjö, A., Helgadottir, H. T., Hellström-Pigg, M., Iwarsson, E., Kuchinskaya, E., Lindelöf, H., Mannila, M., Nilsson, D., Pettersson, M., Rudd, E., Sahlin, E., Tesi, B., Tham, E., Thonberg, H., Westenius, E., Winberg, J., Winerdal, M., Nordenskjöld, M., Johansson-Soller, M., Wirta, V., Nordgren, A., Lindstrand, A., & Lagerstedt-Robinson, K. (2025). Diagnostic yield of 1000 trio analyses with exome and genome sequencing in a clinical setting. Frontiers in Genetics, 16, Article 1580879. [CrossRef]
  8. Marbach, F., Stoyanov, G., Erger, F., Stratakis, C. A., Settas, N., London, E., Rosenfeld, J. A., Torti, E., Haldeman-Englert, C., Sklirou, E., Kessler, E., Ceulemans, S., Nelson, S. F., Martinez-Agosto, J. A., Palmer, C. G. S., Signer, R. H., Andrews, M. V., Grange, D. K., ... Schaaf, C. P. (2021). Variants in PRKAR1B cause a neurodevelopmental disorder with autism spectrum disorder, apraxia, and insensitivity to pain. Genetics in Medicine, 23(8), 1465–1473. [CrossRef]
  9. McGlaughon, J. L., Goldstein, J. L., Thaxton, C., Hemphill, S. E., & Berg, J. S. (2018). The progression of the ClinGen gene clinical validity classification over time. Human Mutation, 39(11), 1494–1504. [CrossRef]
  10. McKenna, A., Hanna, M., Banks, E., Sivachenko, A., Cibulskis, K., Kernytsky, A., Garimella, K., Altshuler, D., Gabriel, S., Daly, M., & DePristo, M. A. (2010). The Genome Analysis Toolkit: A MapReduce framework for analyzing next-generation DNA sequencing data. Genome Research, 20(9), 1297–1303. [CrossRef]
  11. Minoche, A. E., Lundie, B., Peters, G. B., et al. (2021). ClinSV: Clinical grade structural and copy number variant detection from whole genome sequencing data. Genome Medicine, 13, Article 32. [CrossRef]
  12. Novocraft Technologies. (2022). novoAlign (Version 4) [Computer software]. https://www.novocraft.com.
  13. Paolini, N. A., Attwood, M., Sondalle, S. B., Vieira, C. M. d. S., van Adrichem, A. M., di Summa, F. M., O’Donohue, M.-F., Gleizes, P.-E., Rachuri, S., Briggs, J. W., Fischer, R., Ratcliffe, P. J., Wlodarski, M. W., Houtkooper, R. H., von Lindern, M., Kuijpers, T. W., Dinman, J. D., Baserga, S. J., Cockman, M. E., & MacInnes, A. W. (2017). A ribosomopathy reveals decoding defective ribosomes driving human dysmorphism. American Journal of Human Genetics, 100(3), 506–522. [CrossRef]
  14. Pedersen, B. S., Brown, J. M., Dashnow, H., et al. (2021). Effective variant filtering and expected candidate variant yield in studies of rare human disease. npj Genomic Medicine, 6, Article 60. [CrossRef]
  15. Philippakis, A. A., Azzariti, D. R., Beltran, S., Brookes, A. J., Brownstein, C. A., Brudno, M., Brunner, H. G., Buske, O. J., Carey, K., Doll, C., Dumitriu, S., Dyke, S. O. M., den Dunnen, J. T., Firth, H. V., Gibbs, R. A., Girdea, M., Gonzalez, M., Haendel, M. A., Hamosh, A., ... Rehm, H. L. (2015). The Matchmaker Exchange: A platform for rare disease gene discovery. Human Mutation, 36(10), 915–921. [CrossRef]
  16. Richards, S., Aziz, N., Bale, S., Bick, D., Das, S., Gastier-Foster, J., Grody, W. W., Hegde, M., Lyon, E., Spector, E., Voelkerding, K., & Rehm, H. L. (2015). Standards and guidelines for the interpretation of sequence variants: A joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Genetics in Medicine, 17(5), 405–424. [CrossRef]
  17. Sobreira, N., Schiettecatte, F., Valle, D., & Hamosh, A. (2015). GeneMatcher: A matching tool for connecting investigators with an interest in the same gene. Human Mutation, 36(10), 928–930. [CrossRef]
  18. Spirito, G., Trova, S., Treves, G., Farmanli, K., Franzese Canonico, M., Fant, A., Marangoni, S., Furia, F., Charrance, D., Locci, N., Gottardo, S., Groppo, F., Perseghin, V., Toscano, M., Bibbò, E., Donetti Dontin, S., Cargnelutti, C., Colliard, M., Rosina, A., Beoni, A. M., Bérard, C., Coppe, A., Serravalle, P., Landuzzi, F., Musacchia, F., Amoroso, A., Sanges, R., Cavalli, A., Vecchi, M., Obino, L., & Gustincich, S. (2026). New insights into neurodevelopmental disorders by whole genome sequencing of 100 families from Italy. npj Genomic Medicine. [CrossRef]
  19. Stefanski, A., Calle-López, Y., Leu, C., Pérez-Palma, E., Pestana-Knight, E., & Lal, D. (2021). Clinical sequencing yield in epilepsy, autism spectrum disorder, and intellectual disability: A systematic review and meta-analysis. Epilepsia, 62(1), 143–151. [CrossRef]
  20. Verbinnen, I., Douzgou Houge, S., Hsieh, T.-C., Lesmann, H., Kirchhoff, A., Geneviève, D., Brimble, E., Lenaerts, L., Haesen, D., Levy, R. J., Thevenon, J., Faivre, L., ... Janssens, V. (2025). Pathogenic de novo variants in PPP2R5C cause a neurodevelopmental disorder within the Houge-Janssens syndrome spectrum. American Journal of Human Genetics, 112(3), 554–571. [CrossRef]
  21. Welland, M. J., Ahlquist, K. D., De Fazio, P., Austin-Tse, C., Pais, L., Wedd, L., Bryen, S., Rius, R., Franklin, M., Morrison, C., Hall, G., Gauthier, L., Bloemendal, A., Francis, D. I., Mallett, A. J., Mallawaarachchi, A., Lockhart, P. J., Leventer, R., Scheffer, I. E., Howell, K. B., Kassahn, K. S., Scott, H. S., McGaughran, J., Christodoulou, J., Thorburn, D. R., Thompson, B. A., Patel, C. V., Smith, G., O’Donnell-Luria, A., Sadedin, S., Rehm, H. L., Lunke, S., Wander, J., Samocha, K. E., Simons, C., MacArthur, D. G., & Stark, Z. (2026). Automated reanalysis of genomic data for rare disease diagnostics at scale. Nature Medicine, 32(8), 2991–2999. [CrossRef]
Figure 1. Cohort flow diagram. The four outcome groups are mutually exclusive (no proband appears in more than one). “High-priority finding, pending Sanger validation” comprises probands with a pathogenic or likely pathogenic variant in a Definitive or Strong ClinGen gene pending independent validation; one such proband (PB41) carries two independent pending findings, so the group totals 5 probands and 6 findings.
Figure 1. Cohort flow diagram. The four outcome groups are mutually exclusive (no proband appears in more than one). “High-priority finding, pending Sanger validation” comprises probands with a pathogenic or likely pathogenic variant in a Definitive or Strong ClinGen gene pending independent validation; one such proband (PB41) carries two independent pending findings, so the group totals 5 probands and 6 findings.
Preprints 228639 g001
Figure 2. Functional theme distribution across all 34 reported findings, stacked by confidence category. Themes are sorted by total finding count. Theme assignment methodology is described in Methods.
Figure 2. Functional theme distribution across all 34 reported findings, stacked by confidence category. Themes are sorted by total finding count. Theme assignment methodology is described in Methods.
Preprints 228639 g002
Table 1. Six-domain evidence-scoring system for candidate genes without established ClinGen-validated disease association. 
Table 1. Six-domain evidence-scoring system for candidate genes without established ClinGen-validated disease association. 
Domain Data source Score = 3 Score = 2 Score = 1 Score = 0
Inheritance Pedigree-based inheritance calls (slivar) De novo X-linked recessive in a male proband, or compound heterozygous Autosomal recessive Inherited dominant
Variant effect ACMG/AMP evidence codes and predicted consequence Protein-truncating variant, frameshift, or canonical splice disruption Missense with multiple damaging in silico predictions Missense VUS Synonymous or benign
Phenotype match Recall statistic (proband HPO ∩ gene-disease HPO / gene-disease HPO) Full match (≥70% recall) Partial match (30–69% recall) Limited match (<30% recall) Discordant
Biology GTEx-derived tissue expression (Human Protein Atlas, GeneCards); gene function (OMIM, UniProt) Brain-enriched expression + established neurodevelopmental pathway Detectable brain expression + plausible (not disease-established) pathway Plausible biological link without confirmed brain expression, or vice versa Neither
Constraint gnomAD v3 gene constraint metrics (pLI; observed/expected ratio as LOEUF proxy) pLI >0.9 or observed/expected ratio <0.35 Moderate constraint Weak constraint No constraint data available in gnomAD v3
Validation Sanger sequencing trace review Sanger-confirmed with completed parental segregation Sanger-confirmed Pending Failed or false positive
Each domain is scored 0–3 and summed to a total evidence score (range 0–18). A total score of 7 or more was considered high priority for downstream functional validation. ACMG/AMP, American College of Medical Genetics and Genomics/Association for Molecular Pathology; HPO, Human Phenotype Ontology; LOEUF, loss-of-function observed/expected upper bound fraction; pLI, probability of loss-of-function intolerance; VUS, variant of uncertain significance.
Table 2. Clinical characteristics of the analytical cohort (n = 36).
Table 2. Clinical characteristics of the analytical cohort (n = 36).
Characteristic Analytical cohort (n = 36)
Sex
  Female 19 (53%)
  Male 17 (47%)
Developmental delay 19 (53%)
Intellectual disability 29 (81%)
    Severe 14/29 (48%)
    Moderate 3/29 (10%)
    Mild 2/29 (7%)
    Not graded 10/29 (34%)
Seizures (any HPO seizure term) 29 (81%)
Autism spectrum disorder or behavioural features 11 (31%)
Developmental regression 7 (19%)
Dysmorphic features 4 (11%)
Microcephaly 2 (6%)
Hypotonia 2 (6%)
Cerebral palsy 1 (3%)
Visual impairment 1 (3%)
EEG status documented 21/36 (58%)
    Abnormal (incl. multifocal), of those tested 18/20 (90%)*
Neuroimaging status documented 24/36 (67%)
    Abnormal, of those tested 11/22 (50%)*
Documented prior genetic testing 26 (72%)
    Including prior ES 19 (53%)
HPO terms per proband, median (range) 9 (1–16)
Percentages reflect features documented as present in the clinical record; a feature not marked present was not necessarily confirmed absent. *Denominator excludes probands recorded as not tested (1 for EEG, 2 for neuroimaging). Individual-level clinical data are provided in Supplementary Table S1. HPO, Human Phenotype Ontology.
Table 4. Candidate genes without established ClinGen-validated disease association, ranked by evidence score (n = 19 genes, 20 proband-gene findings).
Table 4. Candidate genes without established ClinGen-validated disease association, ranked by evidence score (n = 19 genes, 20 proband-gene findings).
Proband Gene Inheritance Evidence score (/18)
PB29 RRAGB X-linked recessive 11
PB46 TRPC5 X-linked recessive 11
PB01 RNF123 Recessive 10
PB34 TANC1 Recessive 10
PB37 NHSL2 X-linked recessive 9
PB47 FAM104A;C17orf80 Recessive 9
PB30 WNT16 De novo 9
PB01 PCBP4 Recessive 9
PB06 WASHC2C Recessive (biallelic) 9
PB29 LRIG1 Compound heterozygous 9
PB44 PRKAR1B De novo (duplication) 9
PB46 PRKAR1B De novo (duplication) 9
PB45 CDH4 De novo (duplication) 9
PB32 HYDIN Compound heterozygous 8
PB05 TENM4 Compound heterozygous 8
PB14 FAT2 Compound heterozygous 8
PB45 TENT4A De novo 7
PB03 BHMT Recessive 6
PB40 GUCY2F X-linked recessive 6
PB44 ADGRG4 X-linked recessive 5
Evidence score domains and criteria are defined in Table 1. WASHC2C, LRIG1, HYDIN, TENM4, FAT2, TENT4A, PRKAR1B (PB44, PB46), CDH4, and ADGRG4 lack gnomAD v3 constraint data and are scored 0 for that domain; their evidence scores are likely underestimated (see Results). PRKAR1B findings are de novo duplications; the established PRKAR1B-disease mechanism (Marbach-Schaaf neurodevelopmental syndrome) is de novo missense, so the disease-relevance of a duplication mechanism is not established (see Discussion). Full variant-level annotation, including in silico predictions and evidence-score breakdowns, is provided in Supplementary Table S2.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.