Submitted:
03 January 2023
Posted:
04 January 2023
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Materials and Methods
2.1. Data collection and preprocessing
2.2. Feature encoding schemes
2.2.1. One-Hot (OH) encoding
2.2.2. ZSCALE encoding
2.2.3. Word-embedding (WE) encoding
2.2.4. Enhanced Amino Acid Composition (EAAC) encoding
2.2.5. Enhanced Grouped Amino Acids Content (EGAAC) encoding
2.3. The architecture of deep-learning classifiers
- (1)
- Input layer. Each sequence is converted into a feature vector with One-Hot encoding.
- (2)
- The convolution layer. It contains two convolution sublayers followed by two sequentially connected blocks. each block includes a convolution sublayer and a max pooling sublayer. There are 128 convolution kernels with the sizes of 1 and 3 for the first and second convolution sublayers, respectively. A dropout layer with a rate of 0.7 follows each convolution kernel to prevent potential overfitting. In these two blocks, there were 128 convolution kernels with a size of 9 and 10 for these two convolution sublayers of two blocks, respectively; the parameters pool_size of the max-pooling sublayer was set as 2; the dropout rate was set to 0.5. The rectified linear unit (ReLU) is considered the activation function.
- (3)
- Fully connected layer. It contains a dense sublayer with 128 neurons without flattening and a global average pooling sublayer to calculate and output an average value.
- (4)
- Output layer: This layer contains a single neuron, activated by a sigmoid function, to output the probability score (within the range from 0 to 1), indicating the likelihood of the crosstalk. If the probability score of an input sequence is greater than a specified threshold, the central serine in the sequence is predicted as a crosstalk site.
2.4. Performance evaluation
3. Results and discussion
3.1. Construction and functional investigation of the pSADPr datasets
3.2. Construction and evaluation of CNN-based classifiers
3.3. Construction and evaluation of stacking ensemble learning classifiers
3.4. Comparison of CNN-based models and stacking ensemble models
3.5. Construction of the online EdeepSADPr predictor
4. Conclusion
Supplementary Materials
Author Contributions
Funding
Conflicts of Interest
References
- Zolnierowicz, S. and M. Bollen, Protein phosphorylation and protein phosphatases. De Panne, Belgium, September 19-24, 1999. EMBO J, 2000. 19(4): p. 483-8. [CrossRef]
- Nowak, K., et al., Engineering Af1521 improves ADP-ribose binding and identification of ADP-ribosylated proteins. Nat Commun, 2020. 11(1): p. 5199. [CrossRef]
- Brustel, J., et al., Linking DNA repair and cell cycle progression through serine ADP-ribosylation of histones. Nat Commun, 2022. 13(1): p. 185. [CrossRef]
- Larsen, S.C., et al., Systems-wide Analysis of Serine ADP-Ribosylation Reveals Widespread Occurrence and Site-Specific Overlap with Phosphorylation. Cell Rep, 2018. 24(9): p. 2493-2505 e4. [CrossRef]
- Peng, M., et al., Identification of enriched PTM crosstalk motifs from large-scale experimental data sets. J Proteome Res, 2014. 13(1): p. 249-59. [CrossRef]
- Venne, A.S., L. Kollipara, and R.P. Zahedi, The next level of complexity: crosstalk of posttranslational modifications. Proteomics, 2014. 14(4-5): p. 513-24. [CrossRef]
- Luo, F., et al., DeepPhos: prediction of protein phosphorylation sites with deep learning. Bioinformatics, 2019. 35(16): p. 2766-2773. [CrossRef]
- Sha, Y., et al., DeepSADPr: A Hybrid-learning Architecture for Serine ADP-ribosylation site prediction. Methods, 2021. [CrossRef]
- Buch-Larsen, S.C., et al., Mapping Physiological ADP-Ribosylation Using Activated Ion Electron Transfer Dissociation. Cell Reports, 2020. 32(12). [CrossRef]
- Hendriks, I.A., S.C. Larsen, and M.L. Nielsen, An Advanced Strategy for Comprehensive Profiling of ADP- ribosylation Sites Using Mass Spectrometry- based Proteomics. Molecular & Cellular Proteomics, 2019. 18(5): p. 1010-1026. [CrossRef]
- Hornbeck, P.V., et al., PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse. Nucleic Acids Res, 2012. 40(Database issue): p. D261-70. [CrossRef]
- Sha, Y., et al., DeepSADPr: A hybrid-learning architecture for serine ADP-ribosylation site prediction. Methods, 2022. 203: p. 575-583. [CrossRef]
- Huang, Y., et al., CD-HIT Suite: a web server for clustering and comparing biological sequences. Bioinformatics, 2010. 26(5): p. 680-2. [CrossRef]
- Li, W. and A. Godzik, Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences. Bioinformatics, 2006. 22(13): p. 1658-9. [CrossRef]
- Wang, D., et al., MusiteDeep: a deep-learning based webserver for protein post-translational modification site prediction and visualization. Nucleic Acids Res, 2020. 48(W1): p. W140-W146. [CrossRef]
- Chen, Z., et al., iFeature: a Python package and web server for features extraction and selection from protein and peptide sequences. Bioinformatics, 2018. 34(14): p. 2499-2502. [CrossRef]
- Zhang, L., et al., DeepKhib: A Deep-Learning Framework for Lysine 2-Hydroxyisobutyrylation Sites Prediction. Front Cell Dev Biol, 2020. 8: p. 580217. [CrossRef]
- Chen, Y.Z., et al., SUMOhydro: a novel method for the prediction of sumoylation sites based on hydrophobic properties. PLoS One, 2012. 7(6): p. e39195. [CrossRef]
- Ge, L. Improving text classification with word embedding. in IEEE International Conference on Big Data. 2018.
- Lyu, X.R., et al., DeepCSO: A Deep-Learning Network Approach to Predicting Cysteine S-Sulphenylation Sites. Frontiers in Cell and Developmental Biology, 2020. 8. [CrossRef]
- Wei, X.L., et al., DeepKcrot: A Deep-Learning Architecture for General and Species-Specific Lysine Crotonylation Site Prediction. Ieee Access, 2021. 9: p. 49504-49513. [CrossRef]
- Vacic, V., L.M. Iakoucheva, and P. Radivojac, Two Sample Logo: a graphical representation of the differences between two sets of sequence alignments. Bioinformatics, 2006. 22(12): p. 1536-7. [CrossRef]
- Wang, C., et al., GPS 5.0: An Update on the Prediction of Kinase-specific Phosphorylation Sites in Proteins. Genomics Proteomics Bioinformatics, 2020. 18(1): p. 72-80. [CrossRef]
- Mishra, A., P. Pokhrel, and M.T. Hoque, StackDPPred: a stacking based prediction of DNA-binding protein from sequence. Bioinformatics, 2019. 35(3): p. 433-441. [CrossRef]
- Basith, S., G. Lee, and B. Manavalan, STALLION: a stacking-based ensemble learning framework for prokaryotic lysine acetylation site prediction. Brief Bioinform, 2022. 23(1). [CrossRef]
- Zhang, L., et al., SBP-SITA: A sequence-based prediction tool for S-itaconation. bioRxiv, 2021.
- Xiong, Y., et al., PredT4SE-Stack: Prediction of Bacterial Type IV Secreted Effectors From Protein Sequences Using a Stacked Ensemble Method. Front Microbiol, 2018. 9: p. 2571. [CrossRef]








| Classifier | SN | SP | ACC | MCC | AUC |
| Ten-fold Cross-validation | |||||
| CNNOH | 0.599±0.031 | 0.694±0.001 | 0.649±0.016 | 0.294±0.031 | 0.712±0.020 |
| CNNZSCALE | 0.598±0.059 | 0.694±0.001 | 0.649±0.025 | 0.293±0.058 | 0.705±0.030 |
| CNNWE | 0.591±0.089 | 0.694±0.001 | 0.644±0.044 | 0.285±0.088 | 0.696±0.043 |
| CNNEAAC | 0.523±0.040 | 0.694±0.001 | 0.611±0.021 | 0.219±0.040 | 0.659±0.016 |
| CNNEGAAC | 0.488±0.034 | 0.694±0.001 | 0.595±0.018 | 0.185±0.034 | 0.621±0.029 |
| Independent test | |||||
| CNNOH | 0.608±0.034 | 0.694±0.000 | 0.653±0.016 | 0.303±0.033 | 0.700±0.010 |
| CNNZSCALE | 0.583±0.037 | 0.694±0.000 | 0.641±0.018 | 0.278±0.036 | 0.692±0.017 |
| CNNWE | 0.557±0.058 | 0.694±0.000 | 0.628±0.028 | 0.253±0.057 | 0.682±0.022 |
| CNNEAAC | 0.500±0.016 | 0.694±0.000 | 0.601±0.008 | 0.197±0.016 | 0.637±0.008 |
| CNNEGAAC | 0.488±0.044 | 0.694±0.000 | 0.595±0.021 | 0.185±0.043 | 0.621±0.016 |
| Classifier | SN | SP | ACC | MCC | AUC |
| Cross-validation | |||||
| CNNO+Z+W | 0.618±0.029 | 0.694±0.001 | 0.657±0.014 | 0.313±0.029 | 0.719±0.021 |
| CNNO+Z+W+E | 0.621±0.030 | 0.694±0.001 | 0.658±0.015 | 0.315±0.030 | 0.719±0.019 |
| CNNO+Z+W+E+EG | 0.617±0.039 | 0.694±0.001 | 0.657±0.019 | 0.311±0.039 | 0.718±0.022 |
| Independent test | |||||
| CNNO+Z+W | 0.578±0.009 | 0.694±0.000 | 0.638±0.004 | 0.274±0.009 | 0.704±0.003 |
| CNNO+Z+W+E | 0.584±0.012 | 0.694±0.000 | 0.641±0.006 | 0.279±0.012 | 0.703±0.002 |
| CNNO+Z+W+E+EG | 0.597±0.022 | 0.694±0.000 | 0.647±0.011 | 0.292±0.021 | 0.703±0.002 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).