Preprint
Article

This version is not peer-reviewed.

Beyond Aneuploidy Screening: Opportunistic Pathogen Detection from Unmapped Reads in Routine Non-Invasive Prenatal Testing

Submitted:

06 September 2026

Posted:

07 September 2026

You are already at the latest version

Abstract
Background: Routine non-invasive prenatal testing (NIPT) generates tens of millions of sequencing reads per sample, the majority of which are aligned to the human genome, while unmapped reads are typically discarded during the analysis. We investigated whether these otherwise unused reads could be repurposed for maternal pathogen screening. Methods: An in-house bioinformatic pipeline was developed to analyze unmapped reads from NIPT sequencing data for pathogen detection. The pipeline was initially applied to a pilot cohort of 2,000 international prenatal cfDNA datasets, followed by an additional 6,595 samples for expanded analysis. Samples identified as pathogen-positive by NextSeq 500 sequencing were subjected to deep resequencing on the NovaSeq 6000 platform. Analytical sensitivity and the minimum detectable fragment threshold were evaluated using commercially available reference materials for CMV and T. gondii, prepared by serial dilution (0–500 copies) and spiked into sheared human genomic DNA mimicking cfDNA prior to library preparation and sequencing. Results: Across approximately 8,600 clinical sequencing datasets, five pathogens associated with maternal and fetal health risks were detected, with a total of 23 pathogen-positive samples identified based on high-confidence read pairs: cytomegalovirus (CMV, n=13), parvovirus B19 (B19V, n=3), varicella-zoster virus (VZV, n=3), herpes simplex virus (HSV, n=2), and Toxoplasma gondii (T. gondii, n=2). Deep resequencing of six representative pathogen-detected samples yielded a 33- to 239-fold increase in pathogen-derived read pairs compared with the corresponding NextSeq data, providing evidence for the reproducibility of the pathogen-associated signals. Mapping-position analysis showed that the recovered pathogen-associated read pairs were evenly distributed across multiple regions of the corresponding pathogen reference genomes rather than being concentrated within a singlesome specific genomic locius. To establish the analytical detection of the pipeline, limit of detection (LOD) analyses were performed using certified reference materials. Importantly, a practical detection criterion of ≥ 2 high-confidence pathogen-associated read pairs per sample was supported by the complete absence of pathogen-derived reads in all negative controls. The LOD was 250 copies for CMV (100% detection, 20/20 replicates) and as low as 1 copy for T. gondii (95% detection, 19/20 replicates). Conclusions: Human-genome-unmapped reads generated during NIPT sequencing, which are typically excluded from downstream NIPT analysis, contain clinically relevant pathogen-derived DNA fragments. By repurposing these reads, this integrated NGS approach simultaneously screens for fetal chromosomal abnormalities (including aneuploidies and microdeletions/microduplications) and multiple congenital pathogens from a single blood draw, without additional laboratory procedures. These findings expand the clinical utility of prenatal cfDNA sequencing beyond genetic fetal aberrations to the detection of microbial infections, highlighting advances in NIPT testing and its potential for clinical translation.
Keywords: 
;  ;  ;  ;  ;  

1. Introduction

Congenital infections are an important contributor to adverse pregnancy and neonatal outcomes, including miscarriage, fetal growth restriction, congenital anomalies, neonatal disease, and long-term neurodevelopmental impairment [1,2]. The TORCH acronym was originally coined by Nahmias et al. in 1971 to collectively describe a group of clinically important pathogens responsible for congenital and perinatal infections: Toxoplasma gondii, Other infections (including parvovirus B19 [B19V], varicella-zoster virus [VZV], and Treponema pallidum), Rubella virus, Cytomegalovirus (CMV), and Herpes simplex virus (HSV) — all of which are associated with neonatal disease and congenital defects at birth [3,4,5,6]. Early identification of maternal infections during pregnancy is therefore essential for appropriate clinical management and prevention of vertical transmission [7]. In current clinical practice, prenatal infectious disease screening relies largely on serological assays that detect pathogen-specific antibodies, particularly against TORCH-associated pathogens [8]. Although serology remains the standard approach for assessing maternal immune status, it has well-recognized limitations in detecting active or asymptomatic infection, and cannot distinguish between primary infection and past exposure without supplementary avidity testing [9]. Furthermore, there is currently no global consensus on standardized prenatal TORCH screening protocols, resulting in significant variability in clinical practice across countries [4]. Consequently, direct detection of pathogen-derived nucleic acids has gained interest as a complementary strategy for identifying active maternal infections during pregnancy [9].
Over the past decade, non-invasive prenatal testing (NIPT) has been widely adopted as a routine screening tool for fetal chromosomal abnormalities [10]. Massively parallel sequencing of maternal plasma cell-free DNA (cfDNA) typically generates tens of millions of sequencing reads per sample, the vast majority of which are aligned to the human reference genome for fetal aneuploidy analysis. A small but consistent fraction of reads, however, fails to map to the human genome and is routinely discarded during standard bioinformatic preprocessing [11].
Accumulating evidence suggests that these human-genome unmappable reads may have clinically useful information for infectious diseases, as circulating microbial and pathogen-derived DNA fragments have been detected at low concentrations in maternal plasma during pregnancy [12,13,14]. Large-scale analysis of NIPT sequencing dataset from over 100,000 pregnant women has demonstrated that microbial cfDNA circulates in peripheral blood during pregnancy, establishing a baseline for the interpretation of pathogen-derived signals within this data type [14]. Several studies have further demonstrated the feasibility of detecting specific viral DNA — most notably CMV — directly within NIPT sequencing data, without the need for dedicated infectious disease assays [15,16]. More broadly, recent advances in metagenomic next-generation sequencing (mNGS) have enabled comprehensive analysis of plasma cell-free DNA, allowing the direct identification of pathogen-derived nucleic acids without prior target selection and detection of a wide range of bacteria, viruses, and parasites [17]. However, most mNGS-based infectious disease studies have required dedicated sequencing experiments specifically designed for pathogen detection. Whether the unmapped reads generated as a by-product of typical NIPT sequencing can be systematically repurposed for concurrent maternal pathogen surveillance — without additional blood sampling, laboratory procedures, or sequencing costs — remains largely unexplored. In this study, we developed an in-house bioinformatic pipeline designed to screen for pathogen-derived DNA fragments within unmapped reads from standard low-depth NIPT sequencing data. The pipeline targets six DNA-based pathogens: cytomegalovirus (CMV), parvovirus B19 (B19V), varicella-zoster virus (VZV), herpes simplex virus (HSV), Treponema pallidum (T. pallidum), and Toxoplasma gondii (T. gondii). Rubella virus, an RNA pathogen, was excluded given that its genome is not detectable by DNA sequencing. We demonstrate that routinely generated unmapped reads from NIPT data can yield clinically meaningful pathogen information beyond fetal aneuploidy screening, supporting the potential integration of maternal infectious disease surveillance into existing prenatal cfDNA testing workflows. Furthermore, given that the prevalence of TORCH-associated pathogens varies substantially across geographic regions, this integrated approach may also provide a basis for population-level maternal infection monitoring.

2. Materials and Methods

2.1. Study Design and NIPT Data Selection

This retrospective study utilized existing sequencing data generated as part of routine non-invasive prenatal testing (NIPT) using samples collected from healthcare institutions across multiple countries and geographic regions between February and September 2025. The sequencing data were generated on an Illumina NextSeq 500 platform (Illumina, San Diego, CA, USA) using 2 × 38-bp paired-end sequencing according to the routine NIPT workflow. To minimize potential geographic bias associated with specific countries or regions, samples were randomly selected from the available dataset.
The study was conducted in two sequential phases. First, 2,000 NIPT samples were selected for a pilot analysis to evaluate the feasibility of detecting pathogen-associated DNA sequences in NIPT data. Following the pilot analysis, data from an additional 6,595 samples were selected using the same sampling strategy for an expanded analysis. In total, sequencing data from 8,595 NIPT samples were included in this study.

2.2. Data Preprocessing

Sequencing reads were processed using Trimmomatic (v0.39) [18] to remove adapter sequences and trim low-quality bases. The preprocessed reads were subsequently aligned to the human reference genome, Genome Reference Consortium Human Build 37 (GRCh37), using BWA-MEM (v0.7.15) [19]. Duplicates were removed using Picard MarkDuplicates (v1.81) [20]. Following alignment to the human reference genome, read pairs for which both mates were unmapped were extracted by SAMtools (v1.8) [21] and subjected to the pathogen DNA screening workflow (Figure 1).

2.3. Pathogen Detection

Read pairs for which both mates were unmapped to the human reference genome were aligned to the reference genomes of the target pathogens using BWA-MEM (v0.7.15). Pathogen reference genome sequences for cytomegalovirus (CMV), parvovirus B19 (B19V), varicella-zoster virus (VZV), herpes simplex virus (HSV), Treponema pallidum (T. pallidum), and Toxoplasma gondii (T. gondii) were obtained from the National Center for Biotechnology Information (NCBI). The corresponding accession numbers are provided in Table 1. Among the reads aligned to each pathogen reference genome, only read pairs with both mates mapped in a properly paired orientation to the corresponding pathogen reference genome were retained. Secondary and supplementary alignments were excluded.
The retained alignments were further filtered according to the following criteria: a mapping quality score (MAPQ) of ≥ 20, an insert size of 38–250 bp, a read length of ≥ 35 bp, and no more than two mismatched bases per read.
To reduce nonspecific or potentially ambiguous alignments, reads mapped on genomic regions annotated as assembly_gap, ncRNA, polyA_site, rRNA, tRNA, tmRNA, or variation in the corresponding GenBank feature annotations were removed using an in-house Python script. These regions were excluded because they were considered more prone to nonspecific or ambiguous alignments. For Toxoplasma gondii, repetitive genomic regions were additionally identified using RepeatMasker (version 4.1.5) [22], and read pairs aligned to these regions were excluded from subsequent analyses.

2.4. BLAST-Based Taxonomic Confirmation

The remaining candidate pathogen-associated read pairs were queried against the National Center for Biotechnology Information (NCBI) core nucleotide database (core_nt) using BLASTN (v2.10.0+) [23]. R1 and R2 reads from each pair were analyzed separately. For each read, the best-supported BLASTN hit was selected primarily based on the highest bit score, followed by the lowest E-value and highest percentage identity in cases of tied alignments. Read pairs were retained only when the best-supported hit for each mate showed a nucleotide sequence identity of ≥ 99%.
The subject accession, organism name, and sequence annotation of the selected BLASTN hits were subsequently reviewed to determine whether the reads matched the corresponding target pathogen or showed evidence of alignment to non-target sequences. Read pairs with best-supported matches to non-target organisms or genomic sequences were excluded from further analysis.
For the pilot analysis, a sample was considered to contain pathogen-associated reads when at least one high-confidence read pair assigned to the same target pathogen was identified. For the expanded analysis, a more stringent detection criterion requiring at least two high-confidence read pairs assigned to the same target pathogen was applied (Figure 2).

2.5. Deep Sequencing-Based Validation

A subset of samples in which pathogen-associated reads were detected during the screening analysis was selected for additional sequence-based validation. Six representative samples were included: two samples with CMV-associated reads and one sample each with reads associated with parvovirus B19 (B19V), herpes simplex virus (HSV), varicella-zoster virus (VZV), and Toxoplasma gondii (T. gondii). Existing NIPT libraries from the selected samples were resequenced on an Illumina NovaSeq 6000 system using an S4 flow cell and a 2 × 38 bp paired-end sequencing configuration. PhiX control library was added at a final proportion of 2% from a 2 nM stock solution. Adapter trimming, quality filtering, removal of human-derived reads, pathogen-reference alignment, and downstream filtering were performed using the same procedures described above.

2.6. Limit of Detection (LOD) Assessment

The analytical limit of detection (LOD) of the pathogen DNA screening workflow was evaluated using commercially available reference materials for CMV and T. gondii (AMPLIRUN DNA CONTROL; Vircell, Granada, Spain). According to the manufacturer, the concentrations of the CMV and T. gondii reference materials were 14,282 ± 327 copies/μL and 12,480 ± 356 copies/μL, respectively.
To simulate pathogen-derived DNA present in a human DNA background, the reference materials were serially diluted and spiked into 10 ng of genomic DNA derived from the human cell line AG09387 (Coriell Cell Repositories, Camden, NJ, USA). Samples containing predefined pathogen DNA inputs of 500, 250, 100, 50, 25, 10, 5, and 1 copy were prepared together with a zero-copy negative control for LOD evaluation.
Initial testing was performed in triplicate across the concentration range to identify the approximate transition between consistent detection and non-detection. Based on the initial results, selected concentrations surrounding the putative detection limit were subsequently evaluated using additional replicates to more precisely estimate the detection rate. For CMV, 250, 100, and 50 copies were evaluated in 20 replicates, whereas other concentrations were evaluated in six replicates following two independent triplicate experiments. For T. gondii, 10, 5, and 1 copy were evaluated in 20 replicates, while higher concentrations were initially evaluated in triplicate. Genomic DNA and reference materials were mixed at high concentrations, subjected to shearing (Covaris M220, Woburn, MA, USA), and subsequently diluted to the target concentrations for downstream processing.
Sequencing libraries were prepared using the same library preparation workflow (NEBNext Ultra II DNA Library Prep Kit for Illumina, New England Biolabs, Ipswich, MA, USA) applied to the routine NIPT samples and sequenced on an Illumina NextSeq 500 or 550Dx platforms using the High Output Reagent Kit v2.5 (2 × 38-bp; paired-end). The resulting sequencing data were analyzed using the same preprocessing, human-read removal, pathogen-reference alignment, alignment filtering, and BLASTN-based taxonomic confirmation procedures described above. A sample was considered positive when at least two high-confidence pathogen-associated read pairs were identified. The observed LOD under the tested conditions was defined as the lowest pathogen DNA input level at which the predefined detection criterion was satisfied in at least 95% of replicate measurements.

3. Results

3.1. Pathogen Detection in the Pilot Cohort

For the pilot analysis, existing NIPT sequencing data from 2,000 samples were analyzed. Human-unmapped read pairs were aligned to the reference genomes of the target pathogens and subjected to the predefined stringent multistep filtering procedure (Figure 2). To minimize nonspecific and potentially ambiguous alignments, only read pairs that satisfied all predefined alignment-quality, fragment-size, read-length, mismatch, genomic-position, repeat-region, and taxonomic-specificity criteria were retained as high-confidence pathogen-associated read pairs.
In the pilot analysis, a sample was considered to contain pathogen-associated reads when at least one high-confidence read pair was identified for a target pathogen. Based on this criterion, pathogen-associated read pairs were identified in 45 of the 2,000 samples (2.3%).
Cytomegalovirus (CMV) was the most frequently detected pathogen, with high-confidence pathogen-associated read pairs identified in 33 samples. This was followed by parvovirus B19 (B19V) and herpes simplex virus (HSV), each detected in four samples, varicella-zoster virus (VZV) in two samples, and Toxoplasma gondii in two samples (Figure 3). No read pairs meeting the predefined detection criteria were identified for T. pallidum.

3.2. Pathogen Detection in Expanded Cohort

The detection of pathogen-associated read pairs in the pilot cohort supported the feasibility of identifying pathogen-associated DNA sequences in NIPT data. To further characterize these signals in a larger cohort, the same pathogen screening pipeline was applied to a total of 8,595 NIPT datasets, including the 2,000 samples analyzed in the pilot phase.
For the expanded analysis, a more stringent detection threshold was applied. Whereas the pilot analysis included samples with at least one high-confidence pathogen-associated read pair, pathogen detection in the expanded cohort required at least two high-confidence read pairs assigned to the same target pathogen.
CMV-associated read pairs were detected most frequently, occurring in 13 samples. Several samples showed relatively high numbers of CMV-associated read pairs, including samples with 96, 31, 13, and 10 read pairs. B19V- and VZV-associated read pairs were each identified in three samples, followed by HSV-associated read pairs in two samples. T. gondii-associated read pairs were also identified in two samples, with three and two read pairs, respectively. Relatively high read-pair counts were observed in two B19V-associated samples, with 40 and 38 read pairs, respectively, and in one HSV-associated sample, with 118 read pairs (Figure 4 and Table 2). No samples met the expanded-cohort detection criterion for T. pallidum. Overall, pathogen-associated read pairs meeting the predefined detection criterion were identified in 23 of 8,595 samples (0.27%).
To assess whether pathogen detection was influenced by differences in sequencing data volume, the distributions of total read counts and human-unmapped read counts were compared between samples with and without detected pathogen-associated read pairs (Figure 5 and Table 3). Both total read counts and human-unmapped read counts showed similar distributions between the two groups, suggesting that the detection of pathogen-associated reads was unlikely to be explained solely by differences in sequencing data volume.

3.3. Sequence-Based Validation of Pathogen-Associated Reads Using Deep Resequencing

Six samples were selected for additional sequence-based validation to represent the pathogen groups identified in the screening analysis. These included two samples with CMV-associated reads and one sample each with read pairs associated with B19V, HSV, VZV, and T. gondii.
Deep sequencing of the six selected samples was performed using the Illumina NovaSeq 6000 platform. In the routine NIPT datasets of these same six samples, a mean of 12,204,689 reads per sample was generated. Deep resequencing generated a mean of 3,102,015,212 reads per sample, corresponding to an approximately 254-fold increase in sequencing output. Similarly, the mean number of reads unmapped to the human reference genome increased from 253,797 in the routine NIPT data to 60,946,654 following deep resequencing, representing an approximately 240-fold increase (Table 4). Despite the substantial increase in sequencing output, the proportion of human-unmapped reads remained broadly comparable between routine NIPT and deep sequencing across the six samples.
Application of the same pathogen screening pipeline to the deep-sequencing data reproduced the pathogen-associated signals observed in the routine NIPT data. In all six samples, the same target pathogen identified in the original screening was detected after deep resequencing. Deep resequencing resulted in a marked increase in the number of high-confidence pathogen-associated read pairs, with increases ranging from approximately 33- to 239-fold relative to the corresponding routine NIPT datasets (Table 5).
Mapping-position analysis showed that the recovered pathogen-associated read pairs were distributed across multiple regions of the corresponding pathogen reference genomes rather than being concentrated within a single genomic locus (Figure S1). This dispersed genomic distribution indicated that the detected signals were supported by reads originating from multiple regions of the corresponding pathogen genomes.
Collectively, the concordance between the initial screening and deep-resequencing results, the increased recovery of high-confidence pathogen-associated read pairs, and their dispersed distribution across the corresponding pathogen genomes provided additional sequence-level support for the pathogen-associated signals identified in the NIPT data.

3.4. Limit of Detection

The analytical sensitivity of the pathogen DNA screening workflow was evaluated using serially diluted reference materials for CMV and T. gondii. Initial testing was performed across a broad range of pathogen DNA inputs to identify the approximate transition between consistent detection and non-detection. Based on the initial results, concentrations surrounding the putative detection limits were subsequently evaluated using additional replicates to more precisely estimate detection rates. A sample was considered positive when at least two high-confidence pathogen-associated read pairs were identified, and the LOD was defined as the lowest nominal pathogen DNA input level at which this detection criterion was met in at least 95% of replicates.
For CMV, all replicates met the predefined detection criterion at 500 copies (6/6) and 250 copies (20/20), with mean high-confidence pathogen-associated read-pair counts of 12.0 and 5.65, respectively. The detection rate decreased to 75% (15/20) at 100 copies and 25% (5/20) at 50 copies. At lower input levels, detection was less consistent, with detection rates of 50% (3/6), 17% (1/6), 0% (0/6), and 17% (1/6) at 25, 10, 5, and 1 copy, respectively. No high-confidence pathogen-associated read pairs were detected in the zero-copy controls. Because 250 copies represented the lowest input level at which the predefined ≥ 95% detection criterion was satisfied, the LOD for CMV was determined to be 250 copies (Figure 6 and Table 6).
In contrast, T. gondii showed substantially higher analytical sensitivity. All replicates met the predefined detection criterion from 500 copies down to 5 copies, including 20 of 20 replicates at both 10 and 5 copies. At an input of 1 copy, 19 of 20 replicates (95%) met the detection criterion, with a mean of 4.20 high-confidence pathogen-associated read pairs. No high-confidence pathogen-associated read pairs were detected in the zero-copy controls. Accordingly, 1 copy, the lowest input level tested, satisfied the predefined LOD criterion for T. gondii. Complete detection (20/20 replicates) was maintained down to 5 copies (Figure 6 and Table 6).
For both pathogens, the number of high-confidence pathogen-associated read pairs generally decreased with decreasing reference-material input. Despite this overall trend, the analytical detection performance differed markedly between CMV and T. gondii, with substantially lower input levels required for reliable detection of T. gondii. These findings indicate pathogen-specific differences in the analytical sensitivity of the sequencing-based screening workflow under the tested experimental conditions.

4. Discussion

The primary novelty of this study is the demonstration that routinely discarded unmapped reads — sequences that do not align to the human genome during standard NIPT bioinformatic processing — contain clinically relevant biological information. Unlike conventional infectious disease testing, this approach requires no additional specimen collection, DNA extraction, library preparation, or sequencing beyond standard NIPT, enabling it to function as an integrated dual-purpose platform that simultaneously screens for both fetal chromosomal abnormalities and pathogen-derived DNA within a single workflow.
Across the entire screening cohort of approximately 8,600 clinical sequencing datasets, CMV was the most frequently detected pathogen, which is broadly consistent with the known global epidemiology of CMV. As the leading cause of congenital infection worldwide [24], CMV has a reported seroprevalence of approximately 45% to over 95% among women of reproductive age, depending on geographic region and socioeconomic factors [25,26]. Despite this high seroprevalence, primary maternal infection during pregnancy still occurs in approximately 1~2% of pregnant women, with incidence varying by country and income level [27,28,29]. While a direct comparison between seropositivity rates and cfDNA detection rates is not straightforward, the presence of CMV cfDNA in plasma reflects active viral replication or reactivation rather than past exposure. Although most maternal CMV infections are asymptomatic, vertical transmission to the fetus may lead to serious long-term sequelae in infected children, including sensorineural hearing loss and neurodevelopmental impairment [30,31]. For this reason, there is an increasing need for routine prenatal screening not only for CMV but also for other pathogens associated with congenital infection risks during pregnancy. However, serological screening programs are not implemented or recommended in most countries. Current clinical practice relies on repeated serological IgG and IgM testing, often performed up to 24 weeks of gestation, placing a considerable burden on both pregnant women and healthcare systems [32,33].
In this context, our findings underscore the potential value of cfDNA-based screening for identifying low-level or asymptomatic maternal infections. By enabling concurrent pathogen screening from a single blood draw already performed for NIPT, this approach eliminates the need for repeated clinic visits and additional laboratory testing, thereby reducing the period of uncertainty and anxiety while awaiting results. It should be noted, however, that this assay is intended solely as a screening tool; confirmatory diagnosis and clinical management should follow the guidance of the treating clinician. Importantly, the current study should be interpreted as a feasibility study demonstrating technical validity rather than a replacement for established diagnostic testing.
Our finding aligns with previous studies demonstrating the feasibility of CMV cfDNA detection within NIPT sequencing data. Faas et al. retrospectively identified CMV cfDNA in maternal plasma from pregnancies and suggested that circulating CMV cfDNA may reflect placental infection rather than strictly maternal infection [34]. Similarly, a proof-of-concept study by Chesnais et al. demonstrated that CMV-specific reads could be detected and quantified using massively parallel shotgun sequencing NIPT data, with a strong correlation between viral load and read counts, supporting the potential of cfDNA as a molecular marker of active infection [11].
Our study extends the clinical utility of NIPT-based pathogen screening beyond CMV by establishing and validating a comprehensive bioinformatic pipeline capable of simultaneously screening multiple congenital pathogens, including Toxoplasma gondii, Parvovirus B19 (B19V), Varicella-Zoster Virus (VZV), Treponema pallidum, Cytomegalovirus (CMV), and Herpes simplex virus (HSV), from a single sequencing dataset.
Independent validation by NovaSeq 6000 resequencing and mapping-position analysis provided orthogonal support for the authenticity of the detected pathogen-derived cfDNA signals. Notably, comparison of total and unmapped read counts across the analyzed samples did not reveal a systematic association between sequencing read metrics and pathogen detection. Thus, the observed pathogen-associated signals were unlikely to be attributable to differences in sequencing depth or unmapped-read abundance. Unlike previous studies focused primarily on CMV detection, our approach demonstrates the feasibility of extending pathogen screening to multiple congenital pathogens using existing low-depth NIPT sequencing data without modifying the underlying NIPT workflow.
The analytical performance of the pipeline was further supported by pathogen-specific limit-of-detection (LOD) test. The LOD for CMV was 250 copies (100% detection rate), while T. gondii was detectable at an input as low as 1 copy (95%, 19/20 replicates) and achieved 100% detection at 5 copies (20/20 replicates), reflecting inherent differences in genome size and genomic complexity, and the biological characteristics of pathogen-derived cfDNA. Although the analytical sensitivity for CMV was lower than that for T. gondii, the analytical sensitivity remained well within a clinically relevant range for a screening assay based on low-depth maternal plasma cfDNA sequencing. Collectively, these findings establish a technically robust and analytically validated framework for pathogen detection from maternal plasma cfDNA and support the feasibility of integrating multi-pathogen screening based on a stepwise bioinformatic filtering pipeline into existing NIPT workflows. Importantly, the practical detection criterion of at least two pathogen-specific paired-end read fragments per sample supported by the complete absence of pathogen-derived reads in all negative control samples, provides a conservative and reproducible threshold for pathogen identification in low-depth NIPT sequencing data. Future prospective studies integrating serological testing, qPCR confirmation, maternal clinical information, and pregnancy outcomes will be required to further establish the clinical validity and utility of this screening approach.

5. Conclusions

In conclusion, our study demonstrates that biologically meaningful unmapped reads generated during routine NIPT sequencing can be repurposed for pathogen detection. By combining a multi-step bioinformatic pipeline with stringent specificity filtering, we established a robust and reproducible analytical framework capable of reliably identifying pathogen-derived cfDNA from low-depth sequencing data. These findings indicate that routine NIPT sequencing data have clinical value beyond fetal chromosomal aneuploidy screening and provide a practical foundation for sequencing-based prenatal pathogen screening without additional laboratory procedures.

Supplementary Materials

The following supporting information can be downloaded at: Preprints.org, Figure S1: Genomic distribution of pathogen-associated reads in deeply resequenced samples.

Author Contributions

Conceptualization, M.-S.L.; methodology, Y.S., J.S., D.P. and B.P.; investigation, Y.S., J.S. and K.S.G.; data curation, Y.K., D.P. and J.K.; formal analysis, C.J.G. and J.P.; visualization, C.J.G., J.P. and J.K.; writing—original draft preparation, J.K., C.J.G. and J.P.; writing—review and editing, J.K. and M.J.K.; supervision, M.J.K. and M.-S.L. All authors have read and agreed to the published version of the manuscript.

Funding

This study was conducted with internal institutional support from Eone-Diagnomics Genome Center Inc. (Incheon, Republic of Korea). No external funding or grant number is applicable.

Institutional Review Board Statement

Not applicable. This study involved the retrospective analysis of de-identified sequencing data generated during routine NIPT testing and did not involve additional sample collection, participant contact, or intervention.

Data Availability Statement

The datasets used and/or analyzed during the current study are available from the corresponding author upon reasonable request.

Acknowledgments

We acknowledge the support of Eone-Diagnomics Genome Center Inc. (EDGC) for this study.

Conflicts of Interest

M.-S.L. is a founder and shareholder of Diagnomics and Eone-Diagnomics Genome Center (EDGC). The other authors are employees of EDGC and have no affiliations with Diagnomics. These affiliations did not influence the study design, data interpretation or conclusions.

Abbreviations

The following abbreviations are used in this manuscript:
NIPT Non-invasive prenatal testing
cfDNA circulating cell-free DNA
CMV Cytomegalovirus
HSV Herpes simplex virus
VZV Varicella-zoster virus
B19V Parvovirus B19
T. gondii Toxoplasma gondii
T. pallidum Treponema pallidum
LOD Limit of detection
MAPQ Mapping quality
BLAST Basic Local Alignment Search Tool
NCBI National Center for Biotechnology Information
GRCh37 Genome Reference Consortium human genome build 37

References

  1. Fortin, O.; Mulkey, S.B. Neurodevelopmental outcomes in congenital and perinatal infections. Curr. Opin. Infect. Dis. 2023, 36, 405–413. [Google Scholar] [CrossRef]
  2. Gordon-Lipkin, E.; Hoon, A.; Pardo, C.A. Prenatal cytomegalovirus, rubella, and Zika virus infections associated with developmental disabilities: Past, present, and future. Dev. Med. Child Neurol. 2021, 63, 135–143. [Google Scholar] [CrossRef]
  3. Nahmias, A.; Walls, K.; Stewart, J.; et al. The ToRCH complex—Perinatal infections associated with toxoplasma and rubella, cytomegalo- and herpes simplex viruses. Pediatr. Res. 1971, 5, 405–406. [Google Scholar] [CrossRef]
  4. David, M.; et al. Fetal and neonatal abnormalities due to congenital syphilis: A literature review. Prenat. Diagn. 2022, 42, 643–655. [Google Scholar] [CrossRef]
  5. Patel, N.; et al. TORCH (toxoplasmosis, other, rubella, cytomegalovirus, herpes simplex virus) infection and the enigma of anomalous fetal development: Pregnancy puzzles. Cureus 2024, 16, e51534. [Google Scholar] [CrossRef]
  6. DeSilva, M.; Munoz, F.M.; McMillan, M.; et al. Congenital anomalies: Case definition and guidelines for data collection, analysis, and presentation of immunization safety data. Vaccine 2016, 34, 6015–6026. [Google Scholar] [CrossRef]
  7. Parums, D.V. A review of emerging viral pathogens and current concerns for vertical transmission of infection. Med. Sci. Monit. 2024, 30, e947335. [Google Scholar] [CrossRef]
  8. Gopalakrishnan, R.; Kandikuppa, R.T. Shining a light on TORCH infections in pregnancy. J. Clin. Infect. Dis. Soc. 2024, 1, 302–308. [Google Scholar] [CrossRef]
  9. Moufarrej, M.N.; et al. Noninvasive prenatal testing using circulating DNA and RNA: Advances, challenges, and possibilities. Annu. Rev. Biomed. Data Sci. 2023, 6, 397–418. [Google Scholar] [CrossRef]
  10. Hui, L.; et al. Position statement from the International Society for Prenatal Diagnosis on the use of non-invasive prenatal testing for the detection of fetal chromosomal conditions in singleton pregnancies. Prenat. Diagn. 2023, 43, 814–828. [Google Scholar] [CrossRef]
  11. Chesnais, V.; Ott, A.; Chaplais, E.; et al. Using massively parallel shotgun sequencing of maternal plasmatic cell-free DNA for cytomegalovirus DNA detection during pregnancy: A proof of concept study. Sci. Rep. 2018, 8, 4321. [Google Scholar] [CrossRef]
  12. Barzon, L.; et al. Next-generation sequencing technologies in diagnostic virology. J. Clin. Virol. 2013, 58, 346–350. [Google Scholar] [CrossRef]
  13. Pan, W.; Ngo, T.T.M.; Camunas-Soler, J.; Song, C.-X.; Kowarsky, M.; et al. Simultaneously monitoring immune response and microbial infections during pregnancy through plasma cfRNA sequencing. Clin. Chem. 2017, 63, 1695–1704. [Google Scholar] [CrossRef]
  14. Tong, X.; et al. Peripheral blood microbiome analysis via noninvasive prenatal testing reveals the complexity of circulating microbial cell-free DNA. Microbiol. Spectr. 2022, 10, e00414-22. [Google Scholar] [CrossRef]
  15. Faas, B.; Astuti, G.; Melchers, W.; et al. Early detection of active human cytomegalovirus (hCMV) infection in pregnant women using data generated for noninvasive fetal aneuploidy testing. eBioMedicine 2024, 100, 104983. [Google Scholar] [CrossRef]
  16. Linthorst, J.; et al. The cell-free DNA virome of 108,349 Dutch pregnant women. Prenat. Diagn. 2023, 43, 448–456. [Google Scholar] [CrossRef]
  17. Linder, K.A.; Miceli, M.H. Impact of metagenomic next-generation sequencing of plasma cell-free DNA testing in the management of patients with suspected infectious diseases. Open Forum Infect. Dis. 2023, 10, ofad385. [Google Scholar] [CrossRef]
  18. Bolger, A.M.; Lohse, M.; Usadel, B. Trimmomatic: A flexible trimmer for Illumina sequence data. Bioinformatics 2014, 30, 2114–2120. [Google Scholar] [CrossRef]
  19. Li, H. Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. arXiv 2013, arXiv:1303.3997. [Google Scholar] [CrossRef]
  20. Broad Institute. Picard Toolkit. 2019. Available online: https://broadinstitute.github.io/picard/.
  21. Li, H.; Handsaker, B.; Wysoker, A.; Fennell, T.; Ruan, J.; Homer, N.; Marth, G.; Abecasis, G.; Durbin, R. 1000 Genome Project Data Processing Subgroup. The Sequence Alignment/Map format and SAMtools. Bioinformatics 2009, 25, 2078–2079. [Google Scholar] [CrossRef]
  22. Smit, A.F.A.; Hubley, R.; Green, P. RepeatMasker Open-4.0. 2013–2015. Available online: http://www.repeatmasker.org (accessed on March 2026).
  23. Camacho, C.; Coulouris, G.; Avagyan, V.; et al. BLAST+: Architecture and applications. BMC Bioinform. 2009, 10, 421. [Google Scholar] [CrossRef]
  24. Mussi-Pinhata, M.M.; et al. Birth prevalence and natural history of congenital cytomegalovirus infection in a highly seroimmune population. Clin. Infect. Dis. 2009, 49, 522–528. [Google Scholar] [CrossRef]
  25. Salomè, S.; Corrado, F.R.; Mazzarelli, L.L.; Maruotti, G.M.; Capasso, L.; Blazquez-Gamero, D.; Raimondi, F. Congenital cytomegalovirus infection: The state of the art and future perspectives. Front. Pediatr. 2023, 11, 1276912. [Google Scholar] [CrossRef]
  26. Cannon, M.J.; Schmid, D.S.; Hyde, T.B. Review of cytomegalovirus seroprevalence and demographic characteristics associated with infection. Rev. Med. Virol. 2010, 20, 202–213. [Google Scholar] [CrossRef]
  27. Colugnati, F.A.; Staras, S.A.; Dollard, S.C.; Cannon, M.J. Incidence of cytomegalovirus infection among the general population and pregnant women in the United States. BMC Infect. Dis. 2007, 7, 71. [Google Scholar] [CrossRef]
  28. Hyde, T.B.; Schmid, D.S.; Cannon, M.J. Cytomegalovirus seroconversion rates and risk factors: Implications for congenital CMV. Rev. Med. Virol. 2010, 20, 311–326. [Google Scholar] [CrossRef]
  29. Ssentongo, P.; Hehnly, C.; Birungi, P.; Roach, M.A.; Spady, J.; Fronterre, C.; et al. Congenital cytomegalovirus infection burden and epidemiologic risk factors in countries with universal screening: A systematic review and meta-analysis. JAMA Netw. Open 2021, 4, e2120736. [Google Scholar] [CrossRef]
  30. Rawlinson, W.D.; Boppana, S.B.; Fowler, K.B.; et al. Congenital cytomegalovirus infection in pregnancy and the neonate: Consensus recommendations for prevention, diagnosis, and therapy. Lancet Infect. Dis. 2017, 17, e177–e188. [Google Scholar] [CrossRef]
  31. Luck, S.E.; Wieringa, J.W.; Blázquez-Gamero, D.; Henneke, P.; Schuster, K.; Butler, K.; et al. Congenital cytomegalovirus: A European expert consensus statement on diagnosis and management. Pediatr. Infect. Dis. J. 2017, 36, 1205–1213. [Google Scholar] [CrossRef]
  32. Faure-Bardon, V.; Fourgeaud, J.; Stirnemann, J.; Leruez-Ville, M.; Ville, Y. Secondary prevention of congenital cytomegalovirus infection with valacyclovir following maternal primary infection in early pregnancy. Ultrasound Obstet. Gynecol. 2021, 58, 576–581. [Google Scholar] [CrossRef]
  33. Leruez-Ville, M.; Ville, Y. Is it time for routine prenatal serological screening for congenital cytomegalovirus? Prenat. Diagn. 2020, 40, 1671–1680. [Google Scholar] [CrossRef]
  34. Faas, B.H.W.; et al. Detection of human cytomegalovirus cell-free DNA in pregnant women with symptomatically infected fetuses: Proof-of-concept study. Ultrasound Obstet. Gynecol. 2025, 65, 470–477. [Google Scholar] [CrossRef]
Figure 1. Overview of the pathogen-associated cfDNA screening workflow using routine NIPT sequencing data.
Figure 1. Overview of the pathogen-associated cfDNA screening workflow using routine NIPT sequencing data.
Preprints 231965 g001
Figure 2. Stringent multistep filtering workflow for the identification of high-confidence pathogen-associated read pairs. *A threshold of ≥1 high-confidence pathogen-associated read pair was used in the pilot analysis, whereas ≥2 read pairs were required for the expanded cohort analysis and LOD assessment.
Figure 2. Stringent multistep filtering workflow for the identification of high-confidence pathogen-associated read pairs. *A threshold of ≥1 high-confidence pathogen-associated read pair was used in the pilot analysis, whereas ≥2 read pairs were required for the expanded cohort analysis and LOD assessment.
Preprints 231965 g002
Figure 3. Pathogen detection in a pilot study. Number of detected cases for each pathogen among 2,000 clinical NGS datasets.
Figure 3. Pathogen detection in a pilot study. Number of detected cases for each pathogen among 2,000 clinical NGS datasets.
Preprints 231965 g003
Figure 4. Detection of pathogen-associated read pairs in 8,595 NIPT samples: (a) Number of samples meeting the detection criterion for each target pathogen. The y-axis indicates the number of detected samples; (b) Distribution of high-confidence pathogen-associated read-pair counts among detected samples. Each point represents one sample. The y-axis represents the log10-transformed number of high-confidence pathogen-associated read pairs per sample. The dashed horizontal line indicates the detection threshold of two read pairs [log10(2) ≈ 0.30].
Figure 4. Detection of pathogen-associated read pairs in 8,595 NIPT samples: (a) Number of samples meeting the detection criterion for each target pathogen. The y-axis indicates the number of detected samples; (b) Distribution of high-confidence pathogen-associated read-pair counts among detected samples. Each point represents one sample. The y-axis represents the log10-transformed number of high-confidence pathogen-associated read pairs per sample. The dashed horizontal line indicates the detection threshold of two read pairs [log10(2) ≈ 0.30].
Preprints 231965 g004
Figure 5. Distribution of sequencing read counts according to pathogen detection status in the pilot and expanded cohorts. (a) Pilot cohort comprising 2,000 NIPT samples. (b) Expanded cohort comprising 8,595 NIPT samples. For each cohort, total sequencing read counts (left) and human-unmapped read counts (right) are shown for samples with and without detected pathogen-associated reads. The y-axis represents the number of sequencing reads per sample. Each point represents an individual sample, and box plots show the median, interquartile range, and whiskers of the read-count distribution within each group. For the human-unmapped read count plots, a discontinuous y-axis was used to accommodate samples with exceptionally high read counts. Samples were classified as detected according to the predefined pathogen detection criteria used for each cohort.
Figure 5. Distribution of sequencing read counts according to pathogen detection status in the pilot and expanded cohorts. (a) Pilot cohort comprising 2,000 NIPT samples. (b) Expanded cohort comprising 8,595 NIPT samples. For each cohort, total sequencing read counts (left) and human-unmapped read counts (right) are shown for samples with and without detected pathogen-associated reads. The y-axis represents the number of sequencing reads per sample. Each point represents an individual sample, and box plots show the median, interquartile range, and whiskers of the read-count distribution within each group. For the human-unmapped read count plots, a discontinuous y-axis was used to accommodate samples with exceptionally high read counts. Samples were classified as detected according to the predefined pathogen detection criteria used for each cohort.
Preprints 231965 g005
Figure 6. Analytical sensitivity of the pathogen DNA screening workflow for CMV and T. gondii. Serial dilutions of CMV and T. gondii reference materials were spiked into a human genomic DNA background and analyzed using the pathogen DNA screening workflow. The x-axis indicates the nominal pathogen DNA input (copies per sample), and the y-axis indicates the number of high-confidence pathogen-associated read pairs detected. Points represent the mean number of high-confidence pathogen-associated read pairs, and error bars indicate the standard deviation across replicate experiments. (a) CMV. (b) T. gondii. A sample was considered positive when at least two high-confidence pathogen-associated read pairs were detected.
Figure 6. Analytical sensitivity of the pathogen DNA screening workflow for CMV and T. gondii. Serial dilutions of CMV and T. gondii reference materials were spiked into a human genomic DNA background and analyzed using the pathogen DNA screening workflow. The x-axis indicates the nominal pathogen DNA input (copies per sample), and the y-axis indicates the number of high-confidence pathogen-associated read pairs detected. Points represent the mean number of high-confidence pathogen-associated read pairs, and error bars indicate the standard deviation across replicate experiments. (a) CMV. (b) T. gondii. A sample was considered positive when at least two high-confidence pathogen-associated read pairs were detected.
Preprints 231965 g006
Table 1. Reference genomes used for alignment of pathogen-associated reads.
Table 1. Reference genomes used for alignment of pathogen-associated reads.
Pathogen Reference Genome Length
CMV
(Cytomegalovirus)
NC_006273.2 (Human herpesvirus 5 strain Merlin, complete genome) 235,646 bp
B19V
(Parvovirus B19)
NC_000883.2 (Human parvovirus B19, complete genome) 5,596 bp
VZV
(Varicella-Zoster Virus)
NC_001348.1 (Human herpesvirus 3, complete genome) 124,884 bp
HSV
(Human herpes virus)
NC_001806.2 (Human herpesvirus 1 strain 17, complete genome) 152,222 bp
T. pallidum
(Treponema pallidum)
NC_021490.2 (Treponema pallidum subsp. pallidum str. Nichols, complete sequence) 1,139,633 bp
T. gondii
(Toxoplasma gondii)
GCF_000006565.2 (Toxoplasma gondii ME49 reference genome TGA4) 65.6 Mb
Table 2. Summary of pathogen detection and high-confidence read-pair counts in the expanded cohort.
Table 2. Summary of pathogen detection and high-confidence read-pair counts in the expanded cohort.
Pathogen Detected samples High-confidence read pairs per sample
CMV 13 2 (n = 7), 3 (n = 2), 10 (n = 1), 13 (n = 1), 31 (n = 1), 96 (n = 1)
B19V 3 2 (n = 1), 38 (n = 1), 40 (n = 1)
VZV 3 2 (n = 2), 10 (n = 1)
HSV 2 2 (n = 1), 118 (n = 1)
T. gondii 2 3 (n = 1), 2 (n = 1)
Total 23
Table 3. Sequencing read characteristics according to pathogen detection status in the pilot and expanded cohorts.
Table 3. Sequencing read characteristics according to pathogen detection status in the pilot and expanded cohorts.
Cohort Detection status Mean total reads Mean human-
unmapped reads
Ratio of human-
unmapped reads (%)
Pilot
cohort
Overall 12,980,727 215,309 1.66
Detected 12,985,178 204,420 1.58
Not detected 12,980,624 215,560 1.66
Expanded
cohort
Overall 13,009,140 237,321 1.83
Detected 12,987,990 219,469 1.70
Not detected 13,009,197 237,369 1.83
*”Detected” indicates samples meeting the predefined pathogen detection criterion: ≥1 high-confidence pathogen-associated read pair in the pilot cohort and ≥2 high-confidence pathogen-associated read pairs in the expanded cohort.
Table 4. Comparison of routine NIPT and deep sequencing metrics for validation samples.
Table 4. Comparison of routine NIPT and deep sequencing metrics for validation samples.
Sample Routine NIPT Deep sequencing
Total reads Human-unmapped reads Unmapped (%) Total reads Human-unmapped reads Unmapped (%)
CMV sample 1 11,826,292 257,155 2.17 2,970,027,024 63,202,637 2.13
CMV sample 2 11,302,592 139,218 1.23 3,346,845,592 38,285,733 1.14
B19V sample 12,007,146 387,557 3.23 2,907,245,236 103,560,635 3.56
HSV sample 13,213,870 226,366 1.71 3,464,427,356 46,546,834 1.34
VZV sample 12,512,740 166,827 1.33 2,974,197,936 38,156,302 1.28
T. gondii sample 12,365,496 345,661 2.8 2,949,348,126 75,927,782 2.57
Mean 12,204,689 253,797 - 3,102,015,212 60,946,654 -
Table 5. Comparison of pathogen-associated read-pair counts between routine NIPT and deep sequencing.
Table 5. Comparison of pathogen-associated read-pair counts between routine NIPT and deep sequencing.
Pathogen Sample Routine NIPT
pathogen-associated read pairs
Deep sequencing
pathogen-associated read pairs
Fold
increase
CMV CMV sample1 96 20,329 211.8×
CMV sample2 10 2,385 238.5×
B19V B19V sample 40 6,258 156.5×
HSV HSV sample 2 65 32.5×
VZV VZV sample 2 209 104.5×
T. gondii T. gondii sample 3 226 75.3×
Table 6. Analytical detection performance of the sequencing-based pathogen screening workflow using CMV and T. gondii spike-in samples.
Table 6. Analytical detection performance of the sequencing-based pathogen screening workflow using CMV and T. gondii spike-in samples.
Spiked
pathogen
Pathogen DNA input (copies) Replicates (n) Mean high-confidence read pairs SD Positive replicates Detection rate (%)
CMV 500 6 12 4.56 6/6 100
250 20 5.65 2.68 20/20 100
100 20 2.4 1.5 15/20 75
50 20 1 1.41 5/20 25
25 6 0.5 0.55 3/6 50
10 6 0.17 0.41 1/6 17
5 6 0 0 0/6 0
1 6 0.17 0.41 1/6 17
0 3 0 0 0/3 0
T. gondii 500 3 3382.67 133.16 3/3 100
250 3 1400 72.69 3/3 100
100 3 483 21.38 3/3 100
50 3 245 22.91 3/3 100
25 3 109 9.85 3/3 100
10 20 52.9 10.43 20/20 100
5 20 25.6 5.1 20/20 100
1 20 4.2 2.48 19/20 95
0 3 0 0 0/3 0
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.