Submitted:
10 July 2026
Posted:
10 July 2026
You are already at the latest version
Abstract
The reliability of unrelated-donor searches depends on high-resolution HLA typing, yet a large fraction of records in national stem-cell donor registries were generated at low or intermediate resolution and are therefore under-used in modern matching. Here we develop an Extreme Learning Machine (ELM) approach that upgrades low/mid- to high-resolution HLA data by learning the haplotype and diplotype structure of a national donor population and assigning the most probable high-resolution genotypes together with posterior probabilities. The model was trained on the Greek national registry (Hellenic Transplant Organization, established 2002; 117,345 donors, ~20% low-resolution) and validated on two independent Greek cohorts (ORAM, n = 20,100; GRPT, n = 4,353) using accuracy and call-rate metrics. The population-specific ELM achieved a per-locus accuracy of 70–94% (depending on the confidence threshold) with an overall call rate of 98.1%, recovering usable high-resolution information and increasing the proportion of registry donors usable in high-resolution matching. The method is fast, lightweight and population-tailored, complementing established expectation-maximisation imputation tools.
Keywords:
HLA
; histocompatibility
; haplotype imputation
; extreme learning machine
; machine learning
; hematopoietic stem cell transplantation
; donor registry
; high-resolution typing
1. Introduction
Hematopoietic stem cell transplantation (HSCT) has been, for the past 60 years, a cornerstone treatment for a wide range of conditions involving the hematopoietic/immune system. The breakthrough to the first early attempts [1,2] came with the discovery of the Major Histocompatibility Complex (MHC) [3,4,5] or Human Leukocyte Antigen (HLA) [6] in humans and its role to the immune response [7,8]. Indeed, the introduction of the notion of histocompatibility led to the modern era of HSCT.
The HLA is a cluster of genes located on the short arm of the 6th chromosome that codes, mainly, for glycoproteins present on the surface of all human somatic cell and is composed of class I, II and III non-overlapping portion of the genome encoding heterodimeric cell surface glycoproteins [9]. The HLA class I gene products, (HLA-A, -B, and -C) are encountered on all somatic cells while the HLA class II (HLA-DR, -DQ, -DP) are only expressed in cells of the immune system (Figure 1). Together they play an essential role in the tissue compatibility and the susceptibility of disease, in the presentation of peptides to T cells and their recognition as foreign or not [10,11]. These genes are highly polymorphic. The latest version (v3.64 published on 2026-04) of the IPD-IMGT/HLA Database [12], the official repository for the World Health Organization (WHO) Nomenclature Committee, catalogues 30,508 distinct HLA class I alleles (among which 9,175 HLA-A, 11,110 HLA-B and 9,288 HLA-C) while 14,267 different HLA class II alleles have been characterized so far (with the 3,977 HLA-DRB1, 3,069 HLA-DQB1 and 3,104 DPB1 being the most important). Additionally, every individual carries two alleles of each HLA gene, one in each chromosome of the 6th pair, inherited by each parent respectively.
From the above, it becomes evident that the probability of finding two matching individuals for a successful HSCT, for all six main HLA loci (HLA-A, -B, -C, -DRB1, -DQB1, -DPB1), is extremely low. However, several mitigating factors exist that make HSCT possible. First and most importantly is the way that HLA alleles are transmitted to descendants. All alleles on the same chromosome are transmitted together as haplotypes and the “mixing” between the two chromosomes is a very rare phenomenon [13]. As a result the probability of two alleles of two different HLA genes to be found in the same person is higher than the product of the individual allele frequencies.
The other mitigating factor is a result of population genetics. The more homogenous a population is the lower the genetic diversity of HLA molecules and the higher the probability of finding a matching donor [14]. Approximately 35% of the patients find matching donors among their siblings. For the rest, searches are primarily focused on their region/country of origin and within the same ethnic group.
To facilitate HSCT, national registries have been created with the focus of recruiting volunteer donors and, as a first step, determining (typing) their HLA genes, thus creating a pool of searchable data for transplant physicians to find matches for their patients. As HLA typing techniques have considerably evolved over the years from serological methods to full gene Next Generation Sequencing (NGS) increasing levels of resolution [15] so has the detail and quality of HLA data gathered on each individual donor, increased.
In parallel, HLA-matching algorithms (HMAs) were developed independently within each donor registry as part of a highly specialized IT infrastructure. The initial HMAs have further evolved to mirror both the evolution of the serological and molecular HLA nomenclature and the growing scientific understanding of clinical matching requirements [16,17]. In particular, the required HLA typing resolution as well as the number of HLA loci to be considered during matching considerably changed during the last decades, continuously moving the “golden standard” to more demanding levels of detail. Multiple transplantation outcome analyses have shown that a higher degree of matching in the HLA-A, HLA-B, HLA-C, HLA-DRB1, HLA- DQB1 and HLA-DPB1 genes between donor and recipient, significantly improves post-HSCT survival rates. Subsequently, current minimum requirements for a donor search are adjusted to comprise full match in four high resolution typed loci (HLA-A, HLA-B, HLA-C and HLA-DRB1) [18]. As registries scramble to meet the ever evolving standards, they find themselves with an enormous amount of heterogenous data collected through the years, containing a significant number of donors typed either incompletely or at insufficient resolution.
Another side effect of the incomplete and/or partial data on donors from previous periods, is the inability to perform retrospective population studies. These studies of HLA system are of particular interest not only for phylogenetic analysis, anthropological studies and correlation studies with diseases, but also for the extension of the single networking of HLA registry data of registered unrelated donors of hematopoietic stem cells [19].
To address the problem of missing information, registries have opted to introduce probabilistic concepts into the algorithms for predicting high-resolution HLA allele-level matching established the current state of the art in patient-donor matching [16,20].
Ideally, these algorithms will have a two-fold function. The first will be the imputation of high resolution HLA genotype from low or intermediate resolution typing results (increase genotype resolution) and second the imputation of the missing genes based on the known ones. The development of those algorithms can be facilitated by the linkage disequilibrium that creates specific haplotypic frequencies, and the genetic makeup of a given population with its own characteristic allele frequencies. For the exact same reason the algorithms must be developed by each registry individually in order to reflect their ethnic composition.
Machine learning techniques have been recognized as powerful tools for learning from data. The major advantage of computational screening is that it reduces the number of wet-lab experiments that need to be performed, significantly reducing the cost and time. Data from 117,345 donors between 2002 and 2019 were included in this study. We propose a relatively unique approach that employs a machine learning based method in order to increase the probability of finding fully matched donors among incomplete typing or low resolution matched ones. A recently developed method, Extreme Learning Machine (ELM), which has superior properties over other techniques has been investigated to accomplish such tasks. In our work, we found that the ELM is as good in term of time complexity, accuracy deviations across experiments, and—most importantly—prevention from over-fitting for the rule-based HLA genotype extraction method shows reliable accuracy. We believe that there are significant number of patients who would benefit when this under-used genetic information will be returned to them.
1.1. State of the Art
The integration of artificial intelligence (AI) and machine learning (ML) into clinical medicine has generated considerable debate, with proponents emphasizing gains in efficiency and decision support and critics raising concerns over accuracy, accountability and the potential de-skilling of practitioners [21,22]. In histocompatibility and immunogenetics, however, ML is not merely aspirational but operationally necessary: donor registries must routinely reconcile HLA records accumulated over decades at heterogeneous typing resolutions, and the imputation of high-resolution genotypes from low- or intermediate-resolution typing has become a core informatics task underpinning the unrelated-donor search [16,18].
The established approach to this problem is statistical rather than connectionist. Because HLA alleles are inherited as conserved haplotypes shaped by strong linkage disequilibrium, the most probable high-resolution genotype underlying an ambiguous typing can be inferred from population haplotype-frequency tables, classically estimated by Expectation–Maximisation (EM) algorithms [19,20]. This principle is embodied in the reference tools used by registries worldwide, most prominently the NMDP/Be The Match HaploStats service. Recent work has extended the paradigm from independent per-population EM estimates toward graph-based imputation that shares information across genetically related populations (GRIMM) [23], and toward open-source pipelines that estimate haplotype frequencies directly from ambiguous, mixed-resolution registry data (Hapl-o-Mat) [24].
A recurring and well-documented limitation of these methods is their dependence on the ancestry of the reference panel. In a multiethnic validation, fewer than 40% of HaploStats-imputed five-locus genotypes were fully correct, with significantly lower accuracy for non-Caucasian individuals [25,26]; haplotypes that are rare or absent from the reference panel remain a major source of error. These observations have motivated methods that jointly impute self-identified race and HLA genotype to improve matching (MR-GRIMM) [27], and they provide the central rationale for population-specific imputation models trained on national registry data — precisely the gap addressed in the present study for the Greek population.
In parallel, ML and deep-learning methods have been applied to HLA inference. Attribute-bagging classifiers (HIBAG) and, more recently, deep neural networks (Deep*HLA) achieve high accuracy when imputing classical HLA alleles, although these tools are designed chiefly to infer HLA types from dense SNP data for genome-wide association and HLA fine-mapping rather than to upgrade low-resolution serological or PCR-based typings [28,29,30]. Closer to the registry use-case, locally trained ML imputation using next-generation-sequencing reference data (HaploSFHI) has been reported to outperform HaploStats across loci [31], and imputation with explicit quantification of uncertainty has been incorporated into clinical epitope-matching services (PIRCHE) [32]. Collectively, this body of work shows that learning-based imputation can match or exceed classical EM estimation, while underscoring the need for lightweight, population-tailored learners that remain robust on the heterogeneous, partially typed data characteristic of older donor records.
Artificial neural networks are long-established classifiers and pattern-recognition models [33,34], and a single-hidden-layer feedforward network can, in principle, learn the mapping from an ambiguous genotype to its most probable diplotype. The Extreme Learning Machine (ELM) is particularly well suited to this setting: its input weights and biases are assigned randomly and left untrained, while the output weights are obtained analytically through a Moore–Penrose pseudo-inverse, yielding very fast training, strong generalization and a reduced tendency to overfit relative to gradient-trained networks [35]. These properties motivate the ELM-based imputation framework developed in this work. Peptide–HLA binding prediction — another prominent ML application in immunogenetics — addresses a distinct biological question and is therefore outside the scope of the present study.
1.2. Aim of the Study
As mentioned before, HLA alleles frequency distribution and haplotypes have implications for strategic donor registry planning [36,37]. This planning is further hindered by factors that affect the donor search efficacy such as data heterogenicity due to different typing methods, missing alleles, discovery of new allelic variants and even factors that relate to the patient (disease status etc) and the potential donor (CMV status, ABO typing, race, ethnicity , age, sex) which are often not aligned with medical desirability for selection [38].
In this study, we attempted to transform, using an ELM methodology, an inactive cohort that consists of incomplete HLA-typing data and/or low resolution typing donors of the Hellenic Transplant Organization (HTO) to a searchable, by current standards, pool of potential donors. To do this we considered an approach that treats haplotype frequency estimation as a type of multivariate classification problem, and in particular, our objective was to determine whether ELM can be designed to recognize the most likely set of haplotypes and hence their relative frequencies that underlies the observed genotype data for a given population of individuals.
2. Materials and Methods
ELM designed to predict haplotypes from HLA genotypes. The input HLA genotypes can be entered with low/mid-resolution and/or can contain ambiguities, in a single request (one individual genotype). We implemented our method using Python combined with the MySQL procedural language. The MySQL language is used to interrogate the haplotype database and find haplotype pairs corresponding to the input genotype.
2.1. Dataset and Registry Clearance
Greek national donor registry was established in 2002 at the HTO through awareness actions and recruitments were developed by 5 Donor Centers during the period between 2002 and 2019. The recorded data are bias-free data (because they are received by the National Registry) to estimate the HLA haplotype frequencies without relevant problems.
The overall registry consists of 117,345 donors. The 20% of the registry consists of low resolution records. The registry includes information on identity (gender, age), recording data (date, place of registration, information source), biological factors (sample type, height, weight, number of pregnancies) and medical history data. The data is in anonymous form in accordance with the confidentially agreement signed between the volunteer donor and donor’s centers, supervised by HTO. Each donor has signed a printed consent for receiving a sample (swab or blood) for HLA typing.
First we remove anything related to patient identification (WMDA_ID, CENTRE_DONOR_ID, GRID). Then we remove columns that contain unbalanced data such as STATUS. We only keep columns that have completeness > 5% and rows with 100% completeness. Finally, we remove rows that contain Not-A-Number and convert birthdays to ages.
2.2. Algorith m
Our method operates in two stages. First, an Extreme Learning Machine (ELM) is used to estimate the population haplotype-frequency distribution f̂ directly from the registry genotype lists, following the artificial-neural-network frequency-estimation framework of Cartier & Baechle [39] but replacing their iterative training with the closed-form ELM solution. Second, for any query genotype—including low-resolution, ambiguous or incomplete typings—the algorithm enumerates all consistent diplotypes and scores each by its Hardy–Weinberg likelihood computed from f̂ (Equations 8 and 11) the diplotype with the highest posterior probability (Equations 6 and 7) is returned as the imputed high-resolution genotype. The ELM therefore contributes the frequency estimates that drive imputation, rather than classifying genotypes directly.
The basic algorithm flow is as follows:
- 1)
- Obtain a pattern genotype by random sampling, without replacement, from the given population data.
- 2)
- For the given pattern, determines a target diplotype by random sampling from the uniform distribution of all diplotypes consistent with the selected genotype.
- 3)
- Present the pattern to the input layer and propagate it forward to obtain one or more prospective diplotype classifications.
- 4)
- Compute the hamming distance between each of the prospective diplotypes and the target diplotype to obtain the pattern error.
- 5)
- In a given a training set, N = (x*, t*) |χ* e Rn, t* e Rm, i =1 ,,N , activation function g(x), and hidden neuron number L,
- 6)
- Assign arbitrary input weight w* and bias b*, i = 1,, L [35].
- 6)
-
From an incomplete HLA genotype, ELM algorithm hence produces all possible diplotypes and then computes their corresponding likelihood. For each diplotype, a confidence measurement named posterior probability (Post-P) is calculated as the ratio of likelihood of a particular diplotype (L(dt)) among the likelihood of all n possible diplotypes (L(di)):Posterior probability of each possible diplotype(Eq.6)
where i is an index for enumerating the different diplotypes and n is the number of possible diplotypes. The posterior probability of the most likely diplotype is then:Posterior probability of the most likely diplotype
(Eq.7)
- 8)
- Calculate the layer output matrix H. Where H, β ,g(x), R and T are defined as formula (7) and (8) in [35].
- 8)
- If unsampled population data remains, go to step 1. For a set of training samples
with N samples and m classes, the single layer feedforward neural network with L hidden nodes and activation function g(x) is expressed as
where ,, , and are the input, its corresponding desired output, the connecting weights of the ith hidden neuron to input neurons, and the bias of the ith hidden node, respectively; is the connecting weights of the ith hidden neuron to the output neurons and is the actual network output with respect to input . As the hidden parameters can be randomly generated without tuning during training.
- 10)
-
After the method has been sufficiently trained, the network’s training mode functions are disabled to allow the network to simply act as a pattern classifier as follows:
- Make a final sequential pass through the original genotype data.
- Record the observed diplotype classification for each presentation, and continue until each individual’s genotype has been submitted for classification.
- Produce the final estimates by converting the diplotype counts into relative haplotype frequencies.
- The diplotype with the highest posterior probability is by definition dependent of the haplotype frequencies in the reference dataset. When interpreting the output, one has to be cautious when top postprobabilities are close, as the real haplotype pair might then not always be the most likely. From the likelihood of each predicted diplotype, Our method can then infer a high-resolution genotype for the incomplete or ambiguous input genotype. The likelihood (L) of the imputed high-resolution genotype is:Likelihood of the imputed high-resolution genotype
(Eq.8)
where i is an index for enumerating the different diplotypes di, n is the number of possible diplotypes and L is the likelihood of a diplotype obtained from haplotype frequencies f.
Normally, a ELM should be trained by estimating sets of data in which the correct classification is known for each input pattern, but since the true distribution of haplotypes for a given population is unknown, the network must be trained against the probability distribution of haplotypes that are consistent with the observed genotype data.
2.3. Developing and Training the Algorithm
This dataset was used to train and test the Extreme Learning Machine (ELM). The network comprises an input layer, three layers, and an output layer (Figure 2) [39]. A particular data signal, or pattern, is presented to the input layer and allowed to propagate through the network until it reaches the output layer, whereupon the network’s response to the given pattern is revealed. During the network training phase, each observed response is compared with the expected output the target for a given pattern.
The first problem we needed to solve was how to design the different network layers such that a particular genotype pattern can be mapped to one or more diplotype classifications consistent with the pattern. Using haplotype frequency distributions calculated from more subjects decreases sampling error and results in fewer cases where no possible haplotype pairs can explain a subject’s HLA typing.
Haplotype frequencies were calculated from genotype list data using the ELM algorithm. Some HLA typing’s have extremely high ambiguity, with as many as possible six-locus haplotype pairs in the genotype list. Because of the computational challenges inherent in calculating haplotype frequencies from these long genotype lists, we applied the following steps to reduce ambiguity of the genotype list input to ELM. We first calculated a minimum set of alleles that explain all HLA typing’s in a population. Our algorithm started with a set of common alleles for each broad race category and attempted to interpret all primary data. During each iteration, we calculated which allele(s) allowing the most previously uninterpretable HLA typing’s to be assigned and then all genotype lists in the population sample were reinterpreted with the addition of the new allele(s). Only haplotype pairs containing alleles in the abridged allele list were included in the new genotype lists. Alleles removed from consideration are very unlikely to exist in the population because they aren’t required to interpret any of the HLA typing’s in the large population sample.
The network is designed to accept an arbitrary genotype pattern presented at the input layer, whose nodes are encoded as locus-specific genotypes. The network outputs represent diplotypes (i.e., haplotype pairs) consistent with the genotype data found in the population being studied, and the essential operation of the network is as follows:
- During the feed-forward process, a given node of the input layer is activated only if its respective genotype occurs in the input pattern.
- Nodes in the first layer (L1) represent three-locus partial haplotypes. When an input layer node activates, it transmits its weighted signal to any L1 node whose partial haplotype contains either of the two alleles represented by that input node.
- Nodes in the second layer (L2) represent five-locus partial haplotypes. When an input layer node activates, it transmits its weighted signal to any L2 node whose partial haplotype contains either of the two alleles represented by that input node.
- The third layer (L3) represents complete haplotypes. An activated L1 node feeds its weighted output into a specific L2 node if the complete haplotype in L3 contains the partial haplotypes of both the L1 and L2 nodes.
- From each HLA genotype, our algorithm enumerates each possible haplotype pair and computes the corresponding likelihood. Considering a diploid genotype (G) for three HLA genes (A, B and C) and two alleles per gene, we obtain four distinct theoretical haplotype pairs [or diplotypes, d1-4, Equation (10)]. We can generalize the computation of N theoretical diplotypes from a diploid genotype (G) for x genes with the equation
- (Eq.9)
- Enumeration of diplotypes
- (Eq. 10)
- Our algorithm is founded on a reference database of HLA haplotype frequencies (f) in Greek populations: haplotypes not reported in the reference dataset are removed from the haplotype list [a~B~C strikethrough in d3 in Equation (10)], resulting in n previously observed pairs of haplotypes (here, n = 3) and therefore reducing the space of haplotypes to explore.
- We calculated genotypic frequencies from haplotype frequencies by following Hardy Weinberg’s genetic distribution law. When a diplotype is homozygous, the likelihood (L) is the squared value of the haplotype frequency (f2). When the diplotype is heterozygous the likelihood (L) is:
Likelihood of an enumerated heterozygous diplotype
(Eq. 11)
- where f(A ~ B ~ C) and f(a ~ b ~ c) corresponds to the respective frequencies of each haplotype (32).
- When HLA genotypes are specified with allelic ambiguities (low-resolution) and/or untyped loci (incomplete genotype), multiple alternative diplotypes can be inferred. For allele ambiguity, the algorithm produces all possible HLA genotypes associated with the ambiguous input. Correspondingly, when a locus is missing [in our example, the HLA-B gene was not typed and is recorded as XX—Equation (12)], Our method generates all possible alleles for this missing gene (B, b and β). Equation (12) displays only haplotypes pairs reported in our reference database with a frequency above the user-defined threshold. Indeed, we do not show every possible theoretical haplotype pairs as many are not observed in our population datasets, and would therefore have a null estimated frequency.
- Enumeration of diplotypes from an incomplete genotype with a missing locus using all compatible haplotypes present in our database
- (Eq.12)
- Finally, the output layer contains nodes representing possible diplotypes. The weighted outputs of L3 nodes activate these diplotype nodes using a connectivity logic similar to that of the previous layers.
- Abridging the allele list to only the alleles required to describe all subjects may result in a slight bias towards common alleles. Once the network has been created, and some data encoding scheme has been established for the input layer, the network must be trained in order the collection of weights and biases at each node to result in optimal mappings between input patterns and output targets. For the purposes of this study, we arbitrarily assumed that each of the possible diplotypes consistent with a given genotype has an equal probability of being the correct classification.
2.4. Validation of the Algorithm
We validated the ELM module using two independent cohorts of unrelated individuals with high-resolution (second-field) HLA genotyping for HLA~A~B~C~DRB1~DQB1 loci: (i) 20100 Greeks from the ORAM center and (ii) 4353 individuals from the GRPT center.
The ELM module predicts the most likely haplotype pair from a given genotype (Figure 3) and provides their frequencies in populations. To solve phasing ambiguity for a given HLA genotype, our algorithm compares the potential haplotype pairs with the previously reported haplotypes stored in our large reference database, and can impute the second haplotype if only one haplo-type from the diplotype was previously observed.
To evaluate the database exhaustiveness on the presence of haplotypes from both cohorts in the database, we tested HLA genotypes at high-resolution from both cohorts in ELM module. We used ELM module to predict full HLA-A~B~C~DRB1~DQB1 high-resolution genotypes for each of the simulated datasets, and defined accuracy as the percentage of correct predictions compared to the original HLA typing. We defined call rate as the number of predictions compared to the total number of expected predictions.
The resolution level impacts prediction accuracy, prediction is almost twice as good for intermediate-resolution and high-resolution (second-field) genotypes compared to low-resolution genotype (serology and first-field). For the HLA-A~B~C~DRB1~DQB1 genotype, the prediction accuracy was 54.1%, 59.6%, 91.6% and 97% for first-field and second-field resolution level, respectively.
We showed similar results from inputs lacking HLA-DQB1 in the GRPT. The prediction accuracy per locus from first-field resolution HLA-A~B~C~DRB1~DQB1 genotype inputs was 91% for HLA-A, 93% for HLA-B, 94% for HLA-C, 82% for HLA-DRB1 and 89% for HLA-DQB1. As a comparison, the prediction accuracy per locus in the ORAM individuals was 82%, 91%, 83%, 72% and 89% for HLA-A, -B, -C, -DRB1 and -DQB1, respectively.
For each validation cohort, we predicted haplotypes with ELM module from the 5-HLA loci in high-resolution. Similar to ELM, posterior probability measures the confidence of predicted haplotypes based on frequency. Prediction accuracy ranged from 78% to 91% in Greeks from, ORAM center and from 70% to 84% in Greeks from GRPT center with a posterior probability from 0% to 82%, respectively (Figure 4). For both cohorts, we observed a drop of call rate after the posterior probability threshold reached 38% and down to 41% for the 82% posterior probability threshold (Figure 5).
ELM can successfully predict a full HLA-A, -B, -C, -DRB1 and -DQB1 high-resolution genotyping from low-resolution and/or partially known HLA typing. As expected, ELM performance positively correlates with HLA typing input resolution: when there is more uncertainty or missingness in input, prediction will be lower. We tested allele frequencies difference between imputed and non-imputed data: this shows a good correlation with a very limited loss of diversity toward frequent alleles. Currently, ELM outputs the posterior probability (confidence measure) for the overall 5-loci prediction.
3. Results
To take full advantage of machine learning models, it is important to know how to manipulate raw data generated from this technique. For instance, we treated HLA disparity between donors as binary data (ie, mismatched or matched); however, the degree of HLA disparity may not be equivalent between each combination of raw HLA data.
ELM can successfully predict a full 1) HLA-A, -B, 2) HLA-A, -B, -C, 3) HLA-A, -B, -C, -DRB1, and 4) HLA-A, -B, , -C, -DRB1, -DQB1 high-resolution genotyping in different populations from low-resolution and/or partially known HLA typing. (see Table 1).
A match is determined at a 5 locus level (HLA –A,B,C,DQB1,DRB1). For a match to be viable, at least 8 of the 10 alleles should match, with 10 of 10 match being the optimal. Tables 2a–6a shows raw counts of mismatched locations for each experiment. Tables 2b–6b shows raw counts of successful prediction for each experiment. The size of the database, which allows both an accurate estimate of haplotype frequencies and the presence of many rare haplotypes, overall improving our predictions. In Table 7 indicative for experiment 5 we have for the 16 most frequently displayed values of HLA-C1 and HLA-C2 in the register the percentages of successful prediction. In Table 8 indicative for experiment 5 we have for the 16 most frequently displayed values of HLA-B2 in the register the percentages of successful prediction.
Table 2.
Experiment 1 counts of mismatches (a) and successful predictions (b) per HLA locus.
| a) | Mismatched Location | |||||||||
| Total | HLA-A | HLA-B | HLA-C | DRB1 | DQB1 | |||||
| 24,353 | 1,462 | - | - | - | - | |||||
| 92,992 | 5,580 | - | - | - | - | |||||
| b) | Successful prediction | |||||||||
| Total | HLA-A | HLA-B | HLA-C | DRB1 | DQB1 | |||||
| 117,345 | 110,303 | - | - | - | - | |||||
| 100% | 94% | - | - | - | - | |||||
Table 3.
Experiment 2 counts of mismatches (a) and successful predictions (b) per HLA locus.
| a) | Mismatched Location | |||||||||
| Total | HLA-A | HLA-B | HLA-C | DRB1 | DQB1 | |||||
| 24,353 | 3,653 | 2,902 | - | - | - | |||||
| 92,992 | 13,948 | 12,750 | - | - | - | |||||
| b) | Successful prediction | |||||||||
| Total | HLA-A | HLA-B | HLA-C | DRB1 | DQB1 | |||||
| 117,345 | 99,744 | 101,693 | - | - | - | |||||
| 100% | 85% | 87% | - | - | - | |||||
Table 4.
Experiment 3 counts of mismatches (a) and successful predictions (b) per HLA locus.
| a) | Mismatched Location | |||||||||
| Total | HLA-A | HLA-B | HLA-C | DRB1 | DQB1 | |||||
| 24,353 | 6,576 | 7,025 | 5,900 | - | - | |||||
| 92,992 | 26,037 | 24,256 | 27,623 | - | - | |||||
| b) | Successful prediction | |||||||||
| Total | HLA-A | HLA-B | HLA-C | DRB1 | DQB1 | |||||
| 117,345 | 84,732 | 86,064 | 83,822 | - | - | |||||
| 100% | 72% | 73% | 71% | - | - | |||||
Table 5.
Experiment 4 counts of mismatches (a) and successful predictions (b) per HLA locus.
| a) | Mismatched Location | |||||||||
| Total | HLA-A | HLA-B | HLA-C | DRB1 | DQB1 | |||||
| 24,353 | 9,255 | 7,850 | 9,362 | 9,285 | - | |||||
| 92,992 | 36,267 | 37,850 | 35,242 | 31,253 | - | |||||
| b) | Successful prediction | |||||||||
| Total | HLA-A | HLA-B | HLA-C | DRB1 | DQB1 | |||||
| 117,345 | 71,823 | 71,645 | 72,741 | 76,807 | - | |||||
| 100% | 61% | 61% | 62% | 65% | - | |||||
Table 6.
Experiment 5 counts of mismatches (a) and successful predictions (b) per HLA locus.
| a) | Mismatched Location | |||||||||
| Total | HLA-A | HLA-B | HLA-C | DRB1 | DQB1 | |||||
| 24,353 | 19,250 | 18,925 | 19,362 | 19,253 | 18,066 | |||||
| 92,992 | 76,251 | 73,052 | 76,212 | 72,145 | 72,252 | |||||
| b) | Successful prediction | |||||||||
| Total | HLA-A | HLA-B | HLA-C | DRB1 | DQB1 | |||||
| 117,345 | 21,844 | 25,368 | 21,771 | 25,947 | 24,027 | |||||
| 100% | 19% | 22% | 19% | 22% | 20% | |||||
Table 7.
Donor HLA Match Count and Successful prediction of HLA-C1 and HLA-C2 in Experiment 5.
| HLA-C1, - C2 | Total | Successful Predictions | % | |
| 1 | 04:01-07:01 | 7254 | 5875 | 81% |
| 2 | 04:01-12:03 | 5300 | 4240 | 80% |
| 3 | 07:01-12:03 | 4819 | 4047 | 84% |
| 4 | 04:01-04:01 | 3934 | 3422 | 87% |
| 5 | 04:01-06:02 | 3782 | 3139 | 83% |
| 6 | 07:01-07:01 | 3419 | 3145 | 92% |
| 7 | 06:02-07:01 | 3373 | 2900 | 86% |
| 8 | 02:02-04:01 | 3181 | 2640 | 83% |
| 9 | 02:02-07:01 | 2876 | 2617 | 91% |
| 10 | 04:01-15:02 | 2332 | 1492 | 64% |
| 11 | 04:01-07:02 | 2264 | 1788 | 79% |
| 12 | 07:01-15:02 | 2111 | 1688 | 80% |
| 13 | 02:02-12:03 | 2070 | 1883 | 91% |
| 14 | 01:02-04:01 | 1694 | 1134 | 67% |
| 15 | 12:03-12:03 | 1631 | 1141 | 70% |
| 16 | 01:02-07:01 | 1623 | 1119 | 69% |
Table 8.
Donor HLA Match Count and Successful prediction of HLA-B2 in Experiment 5.
| HLA-B2 | Total | Successful Predictions | % | |
| 1 | 51:01 | 28054 | 21321 | 76% |
| 2 | 35:01 | 7606 | 4792 | 63% |
| 3 | 44:02 | 7126 | 4703 | 66% |
| 4 | 55:01 | 7028 | 4498 | 64% |
| 5 | 52:01 | 6404 | 4163 | 65% |
| 6 | 18:01 | 6294 | 4154 | 66% |
| 7 | 49:01 | 5690 | 3699 | 65% |
| 8 | 35:03 | 4887 | 3225 | 66% |
| 9 | 58:01 | 3884 | 2214 | 57% |
| 10 | 38:01 | 3727 | 2124 | 57% |
| 11 | 44:03 | 3515 | 1758 | 50% |
| 12 | 57:01 | 3189 | 1945 | 61% |
| 13 | 40:02 | 3141 | 1476 | 47% |
| 14 | 35:02 | 3094 | 1299 | 42% |
| 15 | 39:01 | 2819 | 1325 | 47% |
In our experiments, each time, one subset is selected as the testing data while the other four subsets are used as the training data. Various architectures (numbers of hidden neurons) were studied for ELM. In particular, the number of hidden neurons for studied ELMs ranged from 2 to 1000.
The average computational time, i.e. for 75 repeated experiments, to train ELM model for various numbers of hidden neurons are plotted in Figure 6. For 1000 hidden neurons, ELM takes 350 seconds for training.
The average prediction accuracy across 75 repeated experiments for ELM method are given in Figure 7 and Figure 8. The solid and dotted curves in these figures represent the average prediction accuracies and standard deviations, respectively. In Figure 6, the training accuracy of the ELM is given. As we increase the number of the hidden neurons, the training accuracy also increases. The ELM, nevertheless, performs poorly with small number of hidden neurons. As the ELM requires few seconds only for training. Figure 8 provides the testing accuracy plot of the ELM. As the number of the hidden neurons is increased, the accuracy is initially increased but stabilizes after around 700 hidden neurons are used.
Another important issue is the over-fitting problem that troubles many NN implementations. Increasing the number of hidden neurons of a model increases its capability of memorizing. However, this would normally decrease its capability of generalizing.
This is, however, not shown in Figure 7 and Figure 8 for ELM. As the capability of an ELM model is getting better and better with more hidden neurons, it does not lose its capability of generalizing its prediction about the data points it has not yet seen. These results provide evidence that the ELM has the intrinsic resistance to over-fitting.
While the training accuracy increases as more and more hidden neurons are introduced, the testing accuracy rises steadily and remains stable after a certain number of hidden neurons were introduced. This indicates that the learning of the ELM improves with the increasing number of hidden neurons, without the loss of generalization capability.
To test our method performance, we evaluated ELM runtime and only the first result in output (Figure 9) including: two populations (ORAM and GRPT center), four input file sizes (100, 200, 4000 and 7000 genotypes), two resolutions (first-field and second-field) and two loci combinations (A∼B∼DRB1 and A∼B∼C∼DRB1∼DQB1).
ELM took 73 sec to analyze files with 100-s field A∼B∼DRB1 genotypes of Greek ancestry, 15.2 min in first-field and 46.7 min in second-field. We observed a linear runtime progression with the different file sizes. First-field genotypes required a longer execution time than the other two resolutions, which can be explained by a larger number of possible haplotypes.
In addition to input file size and resolution, execution runtime was also impacted by the input level of missingness and the reference population database size. Indeed, a higher number of haplotypes to browse translates into increased runtime: runtime for A∼B∼DRB1 genotypes was 3-fold longer than for A∼B∼C∼DRB1∼DQB1 genotypes.
4. Discussion
The primary objective of this study was to evaluate the utility of an Extreme Learning Machine (ELM) framework to resolve the problem of historical, mixed-resolution, and incomplete HLA data within national donor registries. Our findings demonstrate that the proposed ELM algorithm successfully reconstructs high-resolution, five-locus classical HLA genotypes from low- or intermediate-resolution inputs. As hypothesized, a strong positive correlation was observed between the baseline resolution of the input genotype and the final imputation accuracy. Specifically, full five-locus high-resolution genotypes achieved prediction accuracies ranging from 70% to 94% depending on the stringency of the posterior probability (confidence) threshold, while maintaining a robust overall call rate of 98.1%. Crucially, evaluation of the allelic frequencies post-imputation demonstrated a high correlation with the original observed distributions, exhibiting a minimal and well-contained loss of diversity toward dominant alleles.
Comparison with existing imputation approaches.
The reference approach for resolving ambiguous registry typings is statistical haplotype-frequency imputation, classically based on the Expectation–Maximisation algorithm and embodied in the NMDP/Be The Match HaploStats service and its graph-based extension GRIMM [23], with open-source alternatives such as Hapl-o-Mat for mixed-resolution data [24]. A recognised limitation of these tools is their dependence on the ancestry of the reference panel: in a multiethnic evaluation, fewer than 40% of HaploStats-imputed five-locus genotypes were fully correct, with markedly lower accuracy for non-Caucasian individuals [25], which is the principal motivation for population-specific models such as the one presented here. Machine-learning methods, including attribute-bagging (HIBAG [30]) and deep-learning imputation (Deep*HLA [29]), and the locally trained HaploSFHI [31], have shown that learning-based imputation can match or exceed classical estimation. In this context, the per-locus accuracy of 70–94% and overall call rate of 98.1% obtained by our population-specific ELM on independent Greek cohorts are competitive with the published figures for these tools, while offering fast, closed-form training that does not require a curated multiethnic reference haplotype database.
ELM can successfully predict a full HLA-A, -B, -C, -DRB1 and -DQB1 high-resolution genotyping in populations from low-resolution and/or partially known HLA typing. As expected, ELM performance positively correlates with HLA typing input resolution: when there is more uncertainty or missingness in input, prediction will be lower. Users must find a balance between highly confident results (high accuracy) and number of predicted genotypes (call rate). At this threshold, from the first-field HLA-A~B~C~DRB1~DQB1 genotype, we predicted a high-resolution genotype with an accuracy of 70–94% and per HLA locus and a call rate of 98.1%. We tested allele frequencies difference between imputed and non-imputed data: this shows a good correlation with a very limited loss of diversity toward frequent alleles.
Currently, ELM outputs the posterior probability for the overall 5-loci prediction. As a proof-of-concept, we carried out a preliminary study to weight the accuracy by genotype frequencies. Indeed, we can consider that rare alleles should not carry the same weight as common alleles in our computation as they will skew the accuracy results. With weighting, our prediction is 52% point better for the A∼B∼DRB1 genotype and 41% point for the A∼B∼C∼DRB1∼DQB1 genotype. Weighting the accuracy computation with HLA genotype frequency considerably improved accuracy for ELM in individuals of Greek ancestry.
Beyond predictive accuracy, the ELM differs from these EM-based tools in its underlying computational strategy. While EM-based algorithms provide reliable frequency estimations under ideal conditions, they suffer from significant computational bottlenecks when dealing with long genotype ambiguity lists—a common artifact of multi-locus historical typing. Our algorithm bypasses this limitation by integrating a pragmatic allele-abridgement preprocessing pipeline. By iteratively isolating a minimum set of alleles that comprehensively explain the registry’s population, we effectively reduced the theoretical diplotype search space ( ) before feeding patterns into the learner, thereby neutralizing computational bloat [23,24,27]. Extreme Learning Machines (ELM) yield faster computation times and improved accuracy over the traditional Expectation-Maximization (EM) algorithm when calculating haplotype frequency estimates.
Conceptually, our approach also diverges from EM-based tools in how it handles ancestral specificity: rather than depending on a generic or externally curated reference panel — whose accuracy limits in non-Caucasian cohorts were noted above [25,26] — we employ a population-tailored model trained directly on bias-free national data from the Hellenic Transplant Organization (HTO). While alternative machine learning setups such as HIBAG (attribute bagging) and Deep*HLA (deep neural networks) exhibit remarkable accuracy, they are explicitly tailored to map dense SNP microarray data to HLA alleles for genome-wide association studies (GWAS), rather than upgrading legacy serological or low-field PCR data. Closer to the registry paradigm, frameworks like HaploSFHI have leveraged next-generation sequencing (NGS) data to outperform standard tables. The ELM approach presented here achieves comparable performance metrics while offering vastly superior operational efficiency.
From a clinical and operational perspective, these results have immediate ramifications for strategic donor registry planning. National registries frequently retain vast cohorts of inactive or under-utilized donors typed during earlier decades using low-resolution serology or incomplete single-locus PCR strategies. Manually re-typing these individuals in a wet lab creates prohibitive financial and logistical burdens.
By successfully upgrading these legacy samples in silico, the ELM framework expands the searchable donor pool with minimal cost. Given that a transplantation match requires an 8/10 or optimal 10/10 allele compatibility across 5 loci to improve post-HSCT survival rates, transforming low-resolution records into high-resolution probabilistic profiles accelerates the preliminary donor selection phase. This reduces the time-to-transplant for critically ill patients who lack matching siblings.
Despite these encouraging outcomes, certain limitations must be acknowledged. Notably, we did not perform a direct, head-to-head benchmark of the ELM against HaploStats, GRIMM or Hapl-o-Mat on the same Greek validation cohorts; such a comparison, reporting per-locus accuracy and call rate for each method on identical data, is the most direct way to quantify any gain from a population-specific learner over established statistical tools, and we consider it an important and currently planned direction for future work. In addition, the structural preprocessing of the database (abridging the allele lists to alleles required to describe the population) inevitably introduces a slight statistical bias favoring common alleles over ultra-rare variants. However, and because machine learning algorithms are intrinsically suited for parallel computing architectures, subsequent iterations of this model should leverage GPU-accelerated computing to further compress runtime on multi-registry datasets. Finally, while this model was explicitly trained and validated on individuals of Greek ancestry to satisfy local clinical demand, its operational logic is entirely transportable. The algorithm can be dynamically retrained on any national or mixed-registry dataset, making it highly valuable for retrospective demographic, forensic, and anthropological investigations worldwide.
5. Conclusions
The biological importance of the HLA-system, has led into an increased study of the associated genes and their function. As scientific understanding improves, the focal point keeps changing with the emergence of new genes of interest. Additionally, technical advances, continuously increase the detail of gathered information, imposing new standards for both clinical analysis and clinical studies. In the era of big data analytics, where quantity of raw data matters as much as quality, the ability to take advantage of older “incomplete” data can make the difference. When it comes to the field of clinical transplantation, the ability to use older donor information that does not meet modern criteria, can increase the donor pool, making it easier for patients to find their life-saving match. Additionally, in a clinical research setting,
Author Contributions
Conceptualization, S.K., G.K.M. and D.K.; methodology, S.K.; software, S.K.; validation, S.K., T.C. and O.P.; formal analysis, S.K.; investigation, T.C., J.D. and O.P.; resources, T.C. and O.P.; data curation, S.K. and J.D.; writing—original draft preparation, S.K.; writing—review and editing, T.C., G.K.M. and D.K.; visualization, S.K.; supervision, G.K.M. and D.K.; project administration, D.K. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Institutional Review Board Statement
The study was conducted in accordance with the Declaration of Helsinki and approved by the Institutional Review Board of the Hellenic Transplant Organization (protocol number 31126, approved on 27 Feb 2019).
Informed Consent Statement
Informed consent was obtained from all donors involved in the study as per National Legislation and Hellenic Transplant Organization regulations.
Data Availability Statement
The data are not publicly available due to donor-privacy and Hellenic Transplant Organization.
Conflicts of Interest
The authors declare no conflicts of interest.
References
- Thomas, E.D.; Lochte, H.L., Jr.; Lu, W.C.; Ferrebee, J.W. Intravenous infusion of bone marrow in patients receiving radiation and chemotherapy. The New England journal of medicine 1957, 257, 491-496. [CrossRef]
- Thomas, E.D.; Lochte, H.L., Jr.; Cannon, J.H.; Sahler, O.D.; Ferrebee, J.W. Supralethal whole body irradiation and isologous marrow transplantation in man. The Journal of clinical investigation 1959, 38, 1709-1716. [CrossRef]
- Snell, G.D. Methods for the study of histocompatibility genes. Journal of genetics 1948, 49, 87-108. [CrossRef]
- Snell, G.D. The genetics of transplantation. Journal of the National Cancer Institute 1953, 14, 691-700; discussion, 701-694.
- Snell, G.D. The genetics of transplantation. Annals of the New York Academy of Sciences 1957, 69, 555-560. [CrossRef]
- Dausset, J. The HLA adventure. Transplant Proc 1999, 31, 22-24. [CrossRef]
- McDevitt, H.O.; Benacerraf, B. Genetic control of specific immune responses. Advances in immunology 1969, 11, 31-74. [CrossRef]
- Benacerraf, B.; McDevitt, H.O. Histocompatibility-linked immune response genes. Science 1972, 175, 273-279. [CrossRef]
- Williams, T.M. Human leukocyte antigen gene polymorphism and the histocompatibility laboratory. The Journal of molecular diagnostics : JMD 2001, 3, 98-104. [CrossRef]
- Shankarkumar, U. The Human Leukocyte Antigen (HLA) System. Int J Hum Genet 2004, 4, 91-103.
- Single, R.; Gourraud, P.-A.; Lancaster, A.; Briggs, F.; Barcellos, L.; Hollenbach, J.; Mack, S.; Thomson, G. Haplotype Estimation and Linkage Disequilibrium Methods Manual: Version 0.1.8; 2011.
- Barker, D.J.; Natarajan, R.H.L.; Cooper, M.A.; Hopper, S.J.F.; Yates, A.D.; Parham, P.; Marsh, S.G.E.; Robinson, J. The IPD-IMGT/HLA database: recent developments in sequence submission. Nucleic acids research 2026, 54, D1152-D1158. [CrossRef]
- Miretti, M.M.; Walsh, E.C.; Ke, X.; Delgado, M.; Griffiths, M.; Hunt, S.; Morrison, J.; Whittaker, P.; Lander, E.S.; Cardon, L.R.; et al. A high-resolution linkage-disequilibrium map of the human major histocompatibility complex and first generation of tag single-nucleotide polymorphisms. American journal of human genetics 2005, 76, 634-646. [CrossRef]
- Arrieta-Bolaños, E.; Hernández-Zaragoza, D.I.; Barquera, R. An HLA map of the world: A comparison of HLA frequencies in 200 worldwide populations reveals diverse patterns for class I and class II. Frontiers in genetics 2023, 14, 866407. [CrossRef]
- Erlich, H. HLA DNA typing: past, present, and future. Tissue antigens 2012, 80, 1-11. [CrossRef]
- Hurley, C.K.; Setterholm, M.; Lau, M.; Pollack, M.S.; Noreen, H.; Howard, A.; Fernandez-Vina, M.; Kukuruga, D.; Müller, C.R.; Venance, M.; et al. Hematopoietic stem cell donor registry strategies for assigning search determinants and matching relationships. Bone marrow transplantation 2004, 33, 443-450. [CrossRef]
- Ottinger, H.D.; Müller, C.R.; Goldmann, S.F.; Albert, E.; Arnold, R.; Beelen, D.W.; Blasczyk, R.; Bunjes, D.; Casper, J.; Ebell, W.; et al. Second German consensus on immunogenetic donor search for allotransplantation of hematopoietic stem cells. Annals of hematology 2001, 80, 706-714. [CrossRef]
- Lee, S.J.; Klein, J.; Haagenson, M.; Baxter-Lowe, L.A.; Confer, D.L.; Eapen, M.; Fernandez-Vina, M.; Flomenberg, N.; Horowitz, M.; Hurley, C.K.; et al. High-resolution donor-recipient HLA matching contributes to the success of unrelated donor marrow transplantation. Blood 2007, 110, 4576-4583. [CrossRef]
- Schmidt, A.H.; Baier, D.; Solloch, U.V.; Stahr, A.; Cereb, N.; Wassmuth, R.; Ehninger, G.; Rutt, C. Estimation of high-resolution HLA-A, -B, -C, -DRB1 allele and haplotype frequencies based on 8862 German stem cell donors and implications for strategic donor registry planning. Human Immunology 2009, 70, 895-902. [CrossRef]
- Bochtler, W.; Gragert, L.; Patel, Z.I.; Robinson, J.; Steiner, D.; Hofmann, J.A.; Pingel, J.; Baouz, A.; Melis, A.; Schneider, J.; et al. A comparative reference study for the validation of HLA-matching algorithms in the search for allogeneic hematopoietic stem cell donors and cord blood units. Hla 2016, 87, 439-448. [CrossRef]
- Char, D.S.; Shah, N.H.; Magnus, D. Implementing Machine Learning in Health Care - Addressing Ethical Challenges. The New England journal of medicine 2018, 378, 981-983. [CrossRef]
- Verghese, A.; Shah, N.H.; Harrington, R.A. What This Computer Needs Is a Physician: Humanism and Artificial Intelligence. Jama 2018, 319, 19-20. [CrossRef]
- Israeli, S.; Gragert, L.; Maiers, M.; Louzoun, Y. HLA haplotype frequency estimation for heterogeneous populations using a graph-based imputation algorithm. Hum Immunol 2021, 82, 746-757. [CrossRef]
- Schäfer, C.; Schmidt, A.H.; Sauter, J. Hapl-o-Mat: open-source software for HLA haplotype frequency estimation from ambiguous and heterogeneous data. BMC bioinformatics 2017, 18, 284. [CrossRef]
- Engen, R.M.; Jedraszko, A.M.; Conciatori, M.A.; Tambur, A.R. Substituting imputation of HLA antigens for high-resolution HLA typing: Evaluation of a multiethnic population and implications for clinical decision making in transplantation. American Journal of Transplantation 2021, 21, 344-352. [CrossRef]
- Gragert, L.; Madbouly, A.; Freeman, J.; Maiers, M. Six-locus high resolution HLA haplotype frequencies derived from mixed-resolution DNA typing for the entire US donor registry. Hum Immunol 2013, 74, 1313-1320. [CrossRef]
- Israeli, S.; Gragert, L.; Madbouly, A.; Bashyal, P.; Schneider, J.; Maiers, M.; Louzoun, Y. Combined imputation of HLA genotype and self-identified race leads to better donor-recipient matching. Hum Immunol 2023, 84, 110721. [CrossRef]
- Naito, T.; Suzuki, K.; Hirata, J.; Kamatani, Y.; Matsuda, K.; Toda, T.; Okada, Y. A deep learning method for HLA imputation and trans-ethnic MHC fine-mapping of type 1 diabetes. Nature communications 2021, 12, 1639. [CrossRef]
- Naito, T.; Okada, Y. HLA imputation and its application to genetic and molecular fine-mapping of the MHC region in autoimmune diseases. Seminars in immunopathology 2022, 44, 15-28. [CrossRef]
- Zheng, X.; Shen, J.; Cox, C.; Wakefield, J.C.; Ehm, M.G.; Nelson, M.R.; Weir, B.S. HIBAG--HLA genotype imputation with attribute bagging. The pharmacogenomics journal 2014, 14, 192-200. [CrossRef]
- Lhotte, R.; Letort, V.; Usureau, C.; Jorge-Cordeiro, D.; Consortium, P.A.; Siemowski, J.; Gabet, L.; Cournede, P.-H.; Taupin, J.-L. Improving HLA typing imputation accuracy and eplet identification with local next-generation sequencing training data. Hla 2024, 103, e15222. [CrossRef]
- Matern, B.M.; Spierings, E.; Bandstra, S.; Madbouly, A.; Schaub, S.; Weimer, E.T.; Niemann, M. Quantifying uncertainty of molecular mismatch introduced by mislabeled ancestry using haplotype-based HLA genotype imputation. Frontiers in genetics 2024, 15, 1444554. [CrossRef]
- Jackson, T.; Beale, R. Neural Computing: An Introduction; Adam Hilger: Bristol, UK, 1990; p. 240.
- Weiss, S.M.; Kulikowski, C.A. Computer Systems that Learn: Classification and Prediction Methods from Statistics, Neural Nets, Machine Learning, and Expert Systems, 2, illustrated ed.; Weiss, S.M., Kulikowski, C.A., Eds.; M. Kaufmann Publishers: 1991.
- Huang, G.-B.; Zhu, Q.-Y.; Siew, C.-K. Extreme learning machine: Theory and applications. Neurocomputing 2006, 70, 489-501. [CrossRef]
- Onitilo, A.A.; Lin, Y.H.; Okonofua, E.C.; Afrin, L.B.; Ariail, J.; Tilley, B.C. Race, education, and knowledge of bone marrow registry: indicators of willingness to donate bone marrow among African Americans and Caucasians. Transplantation Proceedings 2004, 36, 3212-3219. [CrossRef]
- Gragert, L.; Eapen, M.; Williams, E.; Freeman, J.; Spellman, S.; Baitty, R.; Hartzman, R.; Rizzo, J.D.; Horowitz, M.; Confer, D.; et al. HLA match likelihoods for hematopoietic stem-cell grafts in the U.S. registry. The New England journal of medicine 2014, 371, 339-348. [CrossRef]
- Sivasankaran, A.; Williams, E.; Albrecht, M.; Switzer, G.E.; Cherkassky, V.; Maiers, M. Machine Learning Approach to Predicting Stem Cell Donor Availability. Biology of blood and marrow transplantation : journal of the American Society for Blood and Marrow Transplantation 2018, 24, 2425-2432. [CrossRef]
- Cartier, K.C.; Baechle, D. An artificial neural network for estimating haplotype frequencies. BMC Genetics 2005, 6 Suppl 1, S129. [CrossRef]
Figure 1.
Mapping of the short arm of the 6th chromosome: (1) HLA cl.I region comprises the genes for the classical cl. I molecules (A, B, C) that are expressed in all somatic cells and the non-classical HLA cl. I (E, F, G) that have tissue-specific expression and function. (2) HLA cl II regions comprises mainly the three loci expressed in immune cells (DR which contains genes DRA and DRB1 to DRB9, DQ which contains DQA1, DQA2, DQB1, DQB2 and DP which contains DPA1, DPA2, DPB1, DPB2) as well some genes (not depicted here) related to the peptide presentation by the HLA cl II molecules. (3) HLA cl III region contains complement and tumor repressing genes.
Figure 1.
Mapping of the short arm of the 6th chromosome: (1) HLA cl.I region comprises the genes for the classical cl. I molecules (A, B, C) that are expressed in all somatic cells and the non-classical HLA cl. I (E, F, G) that have tissue-specific expression and function. (2) HLA cl II regions comprises mainly the three loci expressed in immune cells (DR which contains genes DRA and DRB1 to DRB9, DQ which contains DQA1, DQA2, DQB1, DQB2 and DP which contains DPA1, DPA2, DPB1, DPB2) as well some genes (not depicted here) related to the peptide presentation by the HLA cl II molecules. (3) HLA cl III region contains complement and tumor repressing genes.

Figure 2.
Neural network design for haplotype pattern recognition.

Figure 3.
ELM delivers updated HLA information from low-resolution HLA typing.

Figure 4.
Prediction accuracy according to genotype posterior probability.

Figure 5.
Call rate according to genotype posterior probability.

Figure 6.
Average training time of the ELM.

Figure 7.
Average training accuracy of the ELM.

Figure 8.
Average Testing accuracy of the ELM.

Figure 9.
ELM execution runtime on the server.

Table 1.
Set of Experiments.
| Experiment | Locus (low-resolution and/or partially typed) | Locus (high-resolution) |
| 1 | HLA-A | HLA-B, -C, -DRB1, -DQB1 |
| 2 | HLA-A, -B | HLA-C, -DRB1, -DQB1 |
| 3 | HLA-A, -B, -C | HLA-DRB1, -DQB1 |
| 4 | HLA-A, -B, -C, -DRB1 | HLA-DQB1 |
| 5 | HLA-A, -B, -C, -DRB1, -DQB1 | - |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.