Preprint
Review

This version is not peer-reviewed.

The Necessity of Batch Effect Assessment and Correction in TCGA RNA-Seq Data

Submitted:

08 September 2026

Posted:

09 September 2026

You are already at the latest version

Abstract
The Cancer Genome Atlas (TCGA) represents a comprehensive and widely utilized resource in cancer research, containing diverse genomic and molecular data from more than 30 human cancer types. Among these datasets, TCGA RNA sequencing (RNA-seq) data are particularly valuable and have become a routine and extensively employed resource for investigating the molecular mechanisms underlying cancer. Numerous studies are published annually based on TCGA RNA-seq data, contributing to the identification of novel diagnostic and prognostic biomarkers and providing insights into the cellular and molecular characteristics of human cancers. However, an important methodological challenge that is often overlooked is the presence of technical variability in TCGA data. TCGA samples are collected from multiple institutions and processed at different times, sequencing centers, and laboratory batches, which can introduce systematic technical variation, commonly referred to as batch effect. These effects may introduce unwanted differences that may affect biological analyses and consequently compromise the validity and reproducibility of downstream analyses. Several approaches have been developed to assess, visualize, and correct batch effect in TCGA RNA-seq data. The appropriate application of these methods can reduce unwanted technical variation while preserving biologically relevant signals, thereby improving the accuracy, robustness, and reliability of subsequent analyses. In this review, we provide an overview of batch effect associated methods in TCGA RNA-seq data, with particular emphasis on their assessment and visualization, as well as the methodological approaches available for their correction.
Keywords: 
;  ;  ;  

1. Introduction

The Cancer Genome Atlas (TCGA) is a widely used and comprehensive resource for accessing genomic and molecular data in cancer research [1]. Numerous studies are published annually utilizing TCGA data to investigate the molecular and clinical characteristics of human cancers. The TCGA database encompasses multiple types of molecular data, including mRNA, miRNA, DNA methylation, and protein profiles, covering more than 30 human cancer types [2,3]. In addition, a wide range of databases, computational packages, and bioinformatics tools have been developed to facilitate the retrieval, processing, integration, and analysis of data generated by TCGA [4,5,6,7]. The use of TCGA data offers several important advantages for cancer research. In addition to providing multiple molecular data types across diverse cancer types, TCGA generally includes relatively large sample cohorts, often comprising more than 100 samples for individual cancer types. Furthermore, comprehensive clinical and pathological information is available for many TCGA samples, including sex, tumor stage, tumor type, and follow-up information. These clinical characteristics enable a broad range of downstream analyses, including survival analysis, clinical correlation studies, and investigations of associations between molecular features and patient outcomes [2].
One of the mostly used data of TCGA project is RNA-seq data. RNA-seq data unveil the transcriptomic changes in cancers and provide novel potential biomarkers and therapeutic options. Using TCGA RNA-seq data, many studies have introduced potential diagnostic and prognostic biomarkers [8,9,10,11]. However, an important yet often overlooked issue is the presence of batch effect, which can introduce unwanted systematic variation into the data, resulting in biased or misleading analytical findings.
Although studies and other sources have documented the presence and potential impact of batch effect in TCGA data [12,13], this critical issue has received relatively limited attention among researchers utilizing TCGA datasets. One possible explanation is the complexity of batch structures within TCGA RNA-seq data, as well as the lack of reliable methodological approaches for detecting, assessing, and correcting batch effect.
In this study, we review the importance of batch effect in TCGA RNA-seq data and present some of the most widely used methods for their assessment and correction. It is important to note that miRNA-seq data are also a type of RNA-seq data; therefore, the considerations discussed here regarding batch effect are equally applicable to miRNA-seq data. Figure 1 visualizes batch effect before and after correction.

2. TCGA RNA-Seq Data

TCGA provides comprehensive genomic data that enable cancer researchers to conduct more precise and in-depth analyses. The project employs a variety of high-throughput technologies to generate large-scale genomic and molecular datasets. These approaches include RNA sequencing (RNA-seq), DNA sequencing (DNA-seq), single-nucleotide polymorphism (SNP)-based platforms, microRNA sequencing (miRNA-seq), reverse-phase protein arrays (RPPA), and array-based DNA methylation profiling [3]. These methods generate various types of molecular data, including gene and exon expression, somatic mutations, copy number variations (CNVs), loss of heterozygosity (LOH), miRNA and protein expression, DNA methylation, and single-nucleotide polymorphisms (SNPs). Among these types of data, RNA-seq are among the most widely used. RNA-seq is a high-throughput and highly precise approach for obtaining transcriptomic information [14]. It enables researchers to identify and quantify both known and novel transcripts, including coding and non-coding RNAs, across a wide range of biological samples [15].

3. Batch Effect in TCGA RNA-Seq Data

Integrating genomic datasets from multiple batches can substantially increase statistical power; however, technical differences between batches may introduce unwanted variation, commonly referred to as batch effect. Batch effects are a major source of technical noise in omics data and can compromise the reliability and reproducibility of downstream analyses [16]. In other words, differences observed between samples may sometimes arise from technical variation rather than from the presence of a disease or a specific biological condition.
Batch effects are a well-recognized challenge in large-scale datasets generated across different sources, institutions, and experimental settings. The Cancer Genome Atlas (TCGA) project involved scientists and managers from the National Cancer Institute (NCI) and the National Human Genome Research Institute (NHGRI), both funded by the U.S. government, and collaborated with numerous institutions across the United States and Europe [3].
In 2015, Kazemian et al. [17] performed differential expression analysis of TCGA RNA-seq data to identify differentially expressed genes between endometrial cancer samples grouped as human papillomavirus 38 positive (HPV38+) and HPV38 negative (HPV38). They found that 61 genes were significantly upregulated or downregulated in the 32 HPV38-positive samples compared with the 136 HPV38-negative samples. However, when 32 samples were randomly selected and compared with the remaining samples, no more than one gene was differentially expressed in more than 98% of the comparisons. They also found that all HPV38-positive samples, although collected from different sources, had been processed on the same sequencing plate (A22K) and in the same batch, while the HPV38-negative samples had been processed in different plates and batches. Therefore, the differences in gene expression may have been caused by a technical issue known as a batch effect rather than by HPV38 itself, meaning that these results do not necessarily show a true association between HPV38 and endometrial cancer. Based on these findings, they recommend that researchers carefully assess batch effect when analyzing RNA-seq data [17].
Each sample in the TCGA RNA-seq dataset is assigned a unique identifier known as the TCGA barcode, which serves as the primary identifier for biospecimen data. The barcode contains information about the project, tissue source site (TSS), participant, sample, vial, portion, analyte, plate, and center. The TSS identifies the site from which the tissue specimen was obtained. The participant identifier represents the individual from whom the specimen was collected, whereas the sample identifier indicates the type of biological sample. The vial identifier indicates the order of the vial within a sequence of samples. The portion identifier specifies the order of a portion within a series of sample portions, typically derived from 100–120 mg of tissue. The analyte identifier indicates the molecular analyte extracted from the sample for analysis, while the plate identifier represents the position of the plate within a sequence of 96-well plates. Finally, the center identifier specifies the sequencing or characterization center responsible for receiving the aliquot for analysis. For example, the TCGA barcode TCGA-02-0001-01B-02D-0181-02 contains the identifiers described in Table 1.
Each of these identifiers can be used to categorize samples into distinct groups and may contribute to significant batch effects. Therefore, systematic assessment and, where necessary, correction of these batch effects is required to ensure the reliability of downstream analyses.

4. Assessment and Visualization of Batch Effect in TCGA RNA-Seq Data

Several tools have been developed to assess batch effect in TCGA RNA-seq data [12,18]. Most of these tools rely on principal component analysis (PCA) to identify potential sources of unwanted variation. By examining the distribution of samples along the major principal components, these approaches can help determine whether the observed differences among TCGA samples are primarily associated with biological factors or may instead be attributable to technical variation between batches. In this section, we review several of these tools.
The “TCGA Batch Effects Viewer” is a web-based tool developed by the MD Anderson Cancer Center to identify, assess, and quantify batch effects in TCGA data. The tool uses methods such as PCA and hierarchical clustering to visualize and evaluate these batch effects. It also provides quantitative measures, such as the Dispersion Separability Criterion (DSC), to assess the degree of variation between batches. When significant batch effects are detected, the tool provides computationally corrected data using appropriate statistical methods.
In a study conducted by Lauss et al. [12] a new R-based package called swamp was introduced. This package can reveal batch effect in TCGA data using heatmap visualizations. Specifically, the swamp R package detects technical biases by applying linear regression to principal components [12]. The package provides a user-friendly approach for assessing batch effect in TCGA RNA-seq data both before and after batch correction. Although swamp offers a simple and effective way to visualize and assess batch effect in TCGA RNA-seq data, it has been implemented in only a limited number of studies [8,11].
PCA-plus [19] is another visualization-based approach designed to facilitate the detection and assessment of batch effect in high-dimensional genomic datasets, including TCGA RNA-seq data. The method builds upon PCA by integrating sample-level information with the principal component structure of the data, thereby enabling the identification of systematic patterns that may be attributable to technical or experimental sources rather than genuine biological variation. In TCGA datasets, PCA-plus can help reveal whether samples cluster according to known or unknown batch-related factors. Hence, PCA-plus provides an intuitive and informative framework for exploratory quality assessment and visualization of batch effect prior to downstream genomic analyses [19].

5. Correction of Batch Effect in TCGA RNA-Seq Data

Following the visualization and assessment of batch effect, significant batch effect should be corrected to minimize their impact on downstream analyses. In this section, we review several methods available for correcting batch effect in RNA-seq data (Table 2). Among these, ComBat-seq is particularly appropriate for TCGA RNA-seq data, whereas other methods may be more suitable under specific experimental or analytical conditions.
In 2020, Zhang et al. [16] indicated that addressing batch effect is essential, particularly in RNA-seq studies, where gene expression data typically consist of skewed and over-dispersed count values rather than continuous data that follow a Gaussian distribution, as assumed by many conventional batch correction approaches. To overcome this limitation, ComBat-seq was developed based on a negative binomial regression framework that accounts for the characteristics of RNA-seq count data while preserving their integer values after batch correction. This feature allows the corrected datasets to remain compatible with widely used differential expression analysis tools that require integer counts. Simulation studies demonstrated that ComBat-seq can improve statistical power and provide better control of false-positive results compared with existing batch adjustment methods, while analysis of real RNA-seq data further confirmed its ability to effectively reduce batch effect and preserve or recover meaningful biological signals [16].
Several studies have used ComBat-seq to remove batch effect from TCGA RNA-seq data [8,23,24,25]. ComBat-seq is implemented in the sva package in R, providing a convenient approach for correcting batch effect in RNA-seq count data [26]. Although ComBat-seq is a widely used and reliable method for batch-effect correction, several other tools and approaches have also been developed for this purpose.
POIBM (POIsson Batch correction through sample Matching) [18] is a batch-effect correction method specifically developed to address technical variation in RNA-seq count data while preserving biologically meaningful differences between samples. Unlike conventional approaches that estimate batch effect primarily from global batch-level distributions, POIBM uses a sample-matching strategy to identify comparable samples across batches and constructs a virtual target sample for each observation based on the most similar samples from other batches. The discrepancy between the observed sample and its corresponding virtual target is then used to estimate and correct the batch effect, with the correction parameters iteratively optimized until convergence. This strategy is particularly advantageous when the biological composition of different batches is unbalanced or when certain biological subgroups are present predominantly in one batch, as it aims to distinguish technical variation from genuine biological heterogeneity and thereby minimize overcorrection. Importantly, POIBM does not require predefined biological replicates or phenotype labels to perform the matching, making it applicable to complex clinical datasets such as TCGA. In the original study, POIBM was evaluated using RNA-seq data from TCGA samples distributed across six batches and demonstrated its ability to reduce technical batch effect while retaining relevant biological signals [18].
Other methods have also been developed for batch-effect correction in RNA-seq data, particularly for use under specific experimental or analytical conditions.
RUVSeq (Remove Unwanted Variation from RNA-seq Data) [20] is a normalization and unwanted-variation correction framework developed specifically for RNA-seq count data. The method models observed read counts as a function of known biological covariates and latent factors representing unwanted technical variation, such as batch effect, library preparation, or other nuisance sources. RUVSeq provides three complementary approaches for estimating these unwanted factors: RUVg, which uses negative-control genes such as housekeeping or spike-in genes; RUVs, which uses technical replicate or negative-control samples with constant biological conditions; and RUVr, which estimates unwanted variation from residuals obtained from an initial generalized linear model. The estimated unwanted factors are subsequently incorporated as covariates into a generalized linear model for differential expression analysis, allowing the effects of technical variation to be controlled while preserving biological variation of interest. [20].
NPM (Nearest-Pair Matching) [21] is a non-parametric batch-effect correction method designed to identify and correct latent technical variation in omics data without requiring explicit batch information. The method takes a normalized and log-transformed gene expression matrix together with phenotype labels and identifies nearest-neighbor pairs across different biological conditions based on sample similarity, using Pearson correlation or Euclidean distance. By matching each sample with the most similar samples from the opposite phenotype group, NPM generates a fully paired dataset in which systematic differences among the resulting pairs, referred to as “pairing effects,” are interpreted as batch-related variation. These pairing effects are subsequently removed and the corrected data are reconstructed to their original dimensions by averaging values from duplicated samples. In this way, NPM aims to reduce latent batch effect while preserving biologically relevant differences between phenotypic groups, making it particularly useful for clinical omics datasets in which batch information may be unavailable or incompletely characterized [21].
MultiBaC is an R package to remove batch effect in multi-omics data such as RNA-seq, miRNA-seq and DNA methylation data obtained from TCGA [22]. MultiBaC contains a diversity of graphical outputs to facilitate model validation and assess the effectiveness of batch-effect correction [22].
Given the availability of these batch-effect correction methods, TCGA RNA-seq data users are strongly encouraged to consider and address this important issue before conducting downstream analyses. However, existing methods still have several limitations. For example, in TCGA data, the presence of multiple types of batch factors means that several batch effects may be statistically significant within the same dataset. These multiple batch effects may need to be addressed simultaneously, requiring appropriate statistical methods and specialized tools. In fact, correction methods may not adequately accommodate multiple batch factors in a single analytical step, potentially requiring sequential correction procedures that may introduce additional bias or remove biologically relevant variation. Therefore, there is a need for improved methods that can effectively address complex and multiple batch effects. Appropriate statistical approaches are essential for accurately identifying and correcting significant batch effects in TCGA data while preserving genuine biological variation and ensuring reliable and reproducible results.

6. Conclusion

It is well established that unwanted noise and unmodeled sources of variation such as batch effect can dramatically reduce the accuracy of statistical inference in genomic experiments. This issue is particularly relevant to TCGA RNA-seq data which represent a widely utilized resource for genomic research in cancer biology. By assessing and correcting batch effect in TCGA RNA-seq data, researchers can distinguish potentially confounding technical variation from biologically meaningful structure, ultimately leading to more robust, reliable, and precise results. In this review, we provided an overview of several widely used approaches for the detection, visualization, assessment, and correction of batch effect in TCGA RNA-seq datasets. Although a range of methods is currently available for identifying and mitigating batch-related variation, the development of more robust, accurate, and flexible methodologies is still required.

Funding

This work did not receive any specific grant from funding agencies in the public, commercial or profit sectors.

Data Availability Statement

No data were generated or used in this study. All datasets and software packages mentioned in this study are publicly available online.

Conflicts of Interest

The author declares no conflict of interest.

References

  1. Chang, K.; Creighton, C.J.; Davis, C.; Donehower, L.; Drummond, J.; Wheeler, D.; et al. The Cancer Genome Atlas Pan-Cancer analysis project. Nat. Genet. 2013, 45(10), 1113–20. [Google Scholar] [CrossRef] [PubMed]
  2. Liu, J.; Lichtenberg, T.; Hoadley, K.A.; Poisson, L.M.; Lazar, A.J.; Cherniack, A.D.; et al. An Integrated TCGA Pan-Cancer Clinical Data Resource to Drive High-Quality Survival Outcome Analytics. Cell 2018, 173(2), 400–16.e11. [Google Scholar] [CrossRef] [PubMed]
  3. Tomczak, K.; Czerwińska, P.; Wiznerowicz, M. The Cancer Genome Atlas (TCGA): an immeasurable source of knowledge. Contemp. Oncol. (Pozn) 2015, 19(1a), A68–77. [Google Scholar] [PubMed]
  4. Chandrashekar, D.S.; Karthikeyan, S.K.; Korla, P.K.; Patel, H.; Shovon, A.R.; Athar, M.; et al. UALCAN: An update to the integrated cancer data analysis platform. Neoplasia 2022, 25, 18–27. [Google Scholar] [CrossRef] [PubMed]
  5. Colaprico, A.; Silva, T.C.; Olsen, C.; Garofano, L.; Cava, C.; Garolini, D.; et al. TCGAbiolinks: an R/Bioconductor package for integrative analysis of TCGA data. Nucleic Acids Res. 2016, 44(8), e71. [Google Scholar] [CrossRef] [PubMed]
  6. de Bruijn, I.; Kundra, R.; Mastrogiacomo, B.; Tran, T.N.; Sikina, L.; Mazor, T.; et al. Analysis and Visualization of Longitudinal Genomic and Clinical Data from the AACR Project GENIE Biopharma Collaborative in cBioPortal. Cancer Res. 2023, 83(23), 3861–7. [Google Scholar] [CrossRef] [PubMed]
  7. Tang, Z.; Kang, B.; Li, C.; Chen, T.; Zhang, Z. GEPIA2: an enhanced web server for large-scale expression profiling and interactive analysis. Nucleic Acids Res. 2019, 47(W1), W556–w60. [Google Scholar] [CrossRef] [PubMed]
  8. Donyavi, M.H.; Salehi-Mazandarani, S.; Nikpour, P. Comprehensive competitive endogenous RNA network analysis reveals EZH2-related axes and prognostic biomarkers in hepatocellular carcinoma. Iran. J. Basic Med. Sci. 2022, 25(3), 286–94. [Google Scholar] [CrossRef] [PubMed]
  9. Jafarnia, P.; Mahdevar, M.; Peymani, M. LENG8-AS1: A Prognostic Biomarker in Colorectal Cancer-Differential Expression and Clinical Implications. Indian J. Clin. Biochem. 2026, 41(1), 105–12. [Google Scholar] [CrossRef] [PubMed]
  10. Liu, M.; Luo, K.; Zhao, H.; Li, Z.; Cai, Y.; Zeng, L.; et al. BANF1 as a potential prognostic biomarker associated with tumor-intrinsic programs and a complex immune landscape in lung adenocarcinoma. Discov. Oncol. 2026, 17(1), 814. [Google Scholar] [CrossRef] [PubMed]
  11. Salehi-Mazandarani, S.; Nikpour, P. Analysis of a Four-Component Competing Endogenous RNA Network Reveals Potential Biomarkers in Gastric Cancer: An Integrated Systems Biology and Experimental Investigation. Adv. BioMed Res. 2023, 12, 238. [Google Scholar] [CrossRef] [PubMed]
  12. Lauss, M.; Visne, I.; Kriegner, A.; Ringnér, M.; Jönsson, G.; Höglund, M. Monitoring of technical variation in quantitative high-throughput datasets. Cancer Inform. 2013, 12, 193–201. [Google Scholar] [CrossRef] [PubMed]
  13. Rasnic, R.; Brandes, N.; Zuk, O.; Linial, M. Substantial batch effects in TCGA exome sequences undermine pan-cancer analysis of germline variants. BMC Cancer 2019, 19(1), 783. [Google Scholar] [CrossRef] [PubMed]
  14. Wang, Z.; Gerstein, M.; Snyder, M. RNA-Seq: a revolutionary tool for transcriptomics. Nat. Rev. Genet. 2009, 10(1), 57–63. [Google Scholar] [CrossRef] [PubMed]
  15. Miller, D.F.; Yan, P.S.; Buechlein, A.; Rodriguez, B.A.; Yilmaz, A.S.; Goel, S.; et al. A new method for stranded whole transcriptome RNA-seq. Methods 2013, 63(2), 126–34. [Google Scholar] [CrossRef] [PubMed]
  16. Zhang, Y.; Parmigiani, G.; Johnson, W.E. ComBat-seq: batch effect adjustment for RNA-seq count data. NAR Genom. Bioinform. 2020, 2(3). [Google Scholar] [CrossRef] [PubMed]
  17. Kazemian, M.; Ren, M.; Lin, J.-X.; Liao, W.; Spolski, R.; Leonard Warren, J. Possible Human Papillomavirus 38 Contamination of Endometrial Cancer RNA Sequencing Samples in The Cancer Genome Atlas Database. J. Virol. 2015, 89(17), 8967–73. [Google Scholar] [CrossRef] [PubMed]
  18. Holmström, S.; Hautaniemi, S.; Häkkinen, A. POIBM: batch correction of heterogeneous RNA-seq datasets through latent sample matching. Bioinformatics 2022, 38(9), 2474–80. [Google Scholar] [CrossRef] [PubMed]
  19. Zhang, N.; Casasent, T.D.; Casasent, A.K.; Kumar, S.V.; Wakefield, C.; Broom, B.M.; et al. PCA-Plus: Enhanced principal component analysis with illustrative applications to batch effects and their quantitation. bioRxiv 2024. [Google Scholar] [CrossRef] [PubMed]
  20. Risso, D.; Ngai, J.; Speed, T.P.; Dudoit, S. Normalization of RNA-seq data using factor analysis of control genes or samples. Nat. Biotechnol. 2014, 32(9), 896–902. [Google Scholar] [CrossRef] [PubMed]
  21. Zito, A.; Martinelli, A.; Masiero, M.; Akhmedov, M.; Kwee, I. NPM: latent batch effects correction of omics data by nearest-pair matching. Bioinformatics 2025, 41(3). [Google Scholar] [CrossRef] [PubMed]
  22. Ugidos, M.; Nueda, M.J.; Prats-Montalbán, J.M.; Ferrer, A.; Conesa, A.; Tarazona, S. MultiBaC: an R package to remove batch effects in multi-omic experiments. Bioinformatics 2022, 38(9), 2657–8. [Google Scholar] [CrossRef] [PubMed]
  23. Govarthan, P.K.; Agastinose Ronickom, J.F.; Swaminathan, R. Identification of Cervical Cancer Biomarkers Using Gene Co-Expression Networks and Machine Learning Methods. Stud. Health Technol. Inform. 2026, 336, 398–402. [Google Scholar] [CrossRef] [PubMed]
  24. Jansen, R.J.; Munro, S.A.; Antwi, S.O.; Rabe, K.G.; Sicotte, H. Bulk RNA-seq deconvolution heterogeneity across paired pancreatic cancer human samples. Front Genet. 2025, 16, 1662924. [Google Scholar] [CrossRef] [PubMed]
  25. Wroblewski, T.H.; Karabacak, M.; Seah, C.; Yong, R.L.; Margetis, K. Radiomic Consensus Clustering in Glioblastoma and Association with Gene Expression Profiles. Cancers 2024, 16(24). [Google Scholar] [CrossRef] [PubMed]
  26. Leek, J.T.; Johnson, W.E.; Parker, H.S.; Jaffe, A.E.; Storey, J.D. The sva package for removing batch effects and other unwanted variation in high-throughput experiments. Bioinformatics 2012, 28(6), 882–3. [Google Scholar] [CrossRef] [PubMed]
Figure 1. Batch effect visualization before and after batch effect correction. In the left part of the figure, samples are distinguished based on the batches instead on their types of biological condition including normal samples and samples from cancer patients. In the right part, batch effect is corrected and samples are distinguished based on their biological condition.
Figure 1. Batch effect visualization before and after batch effect correction. In the left part of the figure, samples are distinguished based on the batches instead on their types of biological condition including normal samples and samples from cancer patients. In the right part, batch effect is corrected and samples are distinguished based on their biological condition.
Preprints 232354 g001
Table 1. Identifiers of the TCGA barcode TCGA-02-0001-01B-02D-0181-02.
Table 1. Identifiers of the TCGA barcode TCGA-02-0001-01B-02D-0181-02.
TCGA 02 0001 01 B 02 D 0181 02
Project TSS Participant Sample Vial Portion Analyte Plate Center
Table 2. Methods for correction of batch effect in RNA-seq data.
Table 2. Methods for correction of batch effect in RNA-seq data.
Method Identified batch Unknown batch1 Point Ref.
ComBat-seq * Useful method for TCGA RNA-seq data [16]
RUVSeq * Useful to identify and correct latent batch effects [20]
POIBM * Based on a sample-matching strategy [18]
NPM * Useful to identify and correct latent batch effects [21]
MultiBaC * Batch correction in multi omics data integration [22]
1 Unknown or hidden batch effect refer to technical or experimental variations that are not explicitly recorded or accounted for in the experimental design, but can introduce systematic differences between samples.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.