Submitted:
02 June 2023
Posted:
05 June 2023
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Related Work
3. Audio Datasets
3.1. XRey
3.2. Hi-Fi TTS
3.3. Tux
3.4. Data Analysis
3.4.1. Phonetic frequency
3.4.2. SNR
3.4.3. Uttering speed
4. Voice Cloning System
5. Experimental framework
5.1. Postprocessed Datasets
- Only audios of durations between 1 and 10 seconds were used in order to reduce variability and increase batch size during training.
5.2. Quality measurement
5.2.1. MOS estimators
5.2.2. Alignment metrics
| Algorithm 1 Character alignment algorithm using the attention matrix of Tacotron-2 |
|
6. Evaluation results and discussion
6.1. Evaluation of XRey
6.2. Evaluation of HQ speakers trained on 3 h
6.3. Evaluation of HQ speakers trained on the whole corpora
7. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| CNN | Convolutional Neural Network |
| GAN | Generative Adversarial Network |
| GMM | Gaussian Mixture Model |
| HMM | Hidden Markov Model |
| HQ | High Quality, refers to speakers Tux and Hi-Fi TTS |
| MFCC | Mel Frequency Cepstral Coefficient |
| MOS | Mean Opinion Score |
| MFA | Montreal Forced Aligner |
| NISQA | Non-Intrusive Speech Quality Assessment |
| SNR | Signal-to-Noise Ratio |
| TTS | Text-to-Speech |
| VAD | Voice Activity Detection |
References
- González-Docasal, A.; Álvarez, A.; Arzelus, H. Exploring the limits of neural voice cloning: A case study on two well-known personalities. In Proceedings of the Proc. IberSPEECH 2022, 2022; pp. 11–15. [Google Scholar] [CrossRef]
- Dale, R. The voice synthesis business: 2022 update. Natural Language Engineering 2022, 28, 401–408. [Google Scholar] [CrossRef]
- van den Oord, A.; Dieleman, S.; Zen, H.; Simonyan, K.; Vinyals, O.; Graves, A.; Kalchbrenner, N.; Senior, A.; Kavukcuoglu, K. WaveNet: A Generative Model for Raw Audio. In Proceedings of the Proc. 9th ISCA Workshop on Speech Synthesis Workshop (SSW 9), 2016; p. 125.
- Lo, C.C.; Fu, S.W.; Huang, W.C.; Wang, X.; Yamagishi, J.; Tsao, Y.; Wang, H.M. MOSNet: Deep Learning-Based Objective Assessment for Voice Conversion. In Proceedings of the Proc. Interspeech 2019, 2019; pp. 1541–1545. [Google Scholar] [CrossRef]
- Cooper, E.; Huang, W.C.; Toda, T.; Yamagishi, J. Generalization Ability of MOS Prediction Networks, 2022. arXiv:2110.02635 [eess]. [CrossRef]
- Cooper, E.; Yamagishi, J. How do Voices from Past Speech Synthesis Challenges Compare Today? In Proceedings of the Proc. 11th ISCA Speech Synthesis Workshop (SSW 11), 2021; pp. 183–188. [CrossRef]
- Shen, J.; Pang, R.; Weiss, R.J.; Schuster, M.; Jaitly, N.; Yang, Z.; Chen, Z.; Zhang, Y.; Wang, Y.; Skerrv-Ryan, R.; et al. Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions. In Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018; pp. 4779–4783. [CrossRef]
- Kong, J.; Kim, J.; Bae, J. HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis. In Proceedings of the Advances in Neural Information Processing Systems. Curran Associates, Inc., 2020, Vol. 33; pp. 17022–17033.
- Mittag, G.; Naderi, B.; Chehadi, A.; Möller, S. NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets. In Proceedings of the Interspeech 2021. ISCA, 2021; pp. 2127–2131. [CrossRef]
- Veerasamy, N.; Pieterse, H. Rising Above Misinformation and Deepfakes. In Proceedings of the International Conference on Cyber Warfare and Security, 2022, Vol. 17; pp. 340–348.
- Pataranutaporn, P.; Danry, V.; Leong, J.; Punpongsanon, P.; Novy, D.; Maes, P.; Sra, M. AI-generated characters for supporting personalized learning and well-being. Nature Machine Intelligence 2021, 3, 1013–1022. [Google Scholar] [CrossRef]
- Salvador Dalí Museum. Dalí Lives (via Artificial Intelligence). https://thedali.org/press-room/dali-lives-museum-brings-artists-back-to-life-with-ai/. Accessed: 2023-05-16.
- Aholab, University of the Basque Country. AhoMyTTS. https://aholab.ehu.eus/ahomytts/. Accessed: 2023-05-16.
- Doshi, R.; Chen, Y.; Jiang, L.; Zhang, X.; Biadsy, F.; Ramabhadran, B.; Chu, F.; Rosenberg, A.; Moreno, P.J. Extending Parrotron: An end-to-end, speech conversion and speech recognition model for atypical speech. In Proceedings of the ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE; 2021; pp. 6988–6992. [Google Scholar] [CrossRef]
- Google Research. Project Euphonia. https://sites.research.google/euphonia/about/. Accessed: 2023-05-16.
- Jia, Y.; Cattiau, J. Recreating Natural Voices for People with Speech Impairments. https://ai.googleblog.com/2021/08/recreating-natural-voices-for-people.html/, 2021. Accessed: 2023-05-16.
- The Story Lab. XRey. https://open.spotify.com/show/43tAQjl2IVMzGoX3TcmQyL/. Accessed: 2023-05-16.
- Chou, J.c.; Lee, H.Y. One-Shot Voice Conversion by Separating Speaker and Content Representations with Instance Normalization. In Proceedings of the Interspeech 2019. ISCA, 2019; pp. 664–668. [CrossRef]
- Qian, K.; Zhang, Y.; Chang, S.; Yang, X.; Hasegawa-Johnson, M. AutoVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss. In Proceedings of the Proceedings of the 36th International Conference on Machine Learning. PMLR, 2019; pp. 5210–5219, ISSN 2640-3498.
- Qian, K.; Zhang, Y.; Chang, S.; Xiong, J.; Gan, C.; Cox, D.; Hasegawa-Johnson, M. Global Prosody Style Transfer Without Text Transcriptions. In Proceedings of the Proceedings of the 38th International Conference on Machine Learning; Meila, M.; Zhang, T., Eds. PMLR, 2021, Vol. 139, Proceedings of Machine Learning Research, pp. 8650–8660.
- Kaneko, T.; Kameoka, H. CycleGAN-VC: Non-parallel Voice Conversion Using Cycle-Consistent Adversarial Networks. In Proceedings of the 2018 26th European Signal Processing Conference (EUSIPCO); IEEE: Rome, 2018; pp. 2100–2104. [Google Scholar] [CrossRef]
- Zhou, K.; Sisman, B.; Li, H. Transforming Spectrum and Prosody for Emotional Voice Conversion with Non-Parallel Training Data. In Proceedings of the Proc. Odyssey 2020 The Speaker and Language Recognition Workshop, 2020; pp. 230–237. [CrossRef]
- Zhou, K.; Sisman, B.; Li, H. Vaw-Gan For Disentanglement And Recomposition Of Emotional Elements In Speech. In Proceedings of the 2021 IEEE Spoken Language Technology Workshop (SLT), 2021; pp. 415–422. [CrossRef]
- van Niekerk, B.; Carbonneau, M.A.; Zaïdi, J.; Baas, M.; Seuté, H.; Kamper, H. A Comparison of Discrete and Soft Speech Units for Improved Voice Conversion. In Proceedings of the ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022; pp. 6562–190. [CrossRef]
- Hsu, W.N.; Bolte, B.; Tsai, Y.H.H.; Lakhotia, K.; Salakhutdinov, R.; Mohamed, A. HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units. IEEE/ACM Transactions on Audio, Speech and Language Processing 2021, 29, 3451–3460. [Google Scholar] [CrossRef]
- Ren, Y.; Hu, C.; Tan, X.; Qin, T.; Zhao, S.; Zhao, Z.; Liu, T.Y. FastSpeech 2: Fast and High-Quality End-to-End Text to Speech, 2022, [arXiv:eess.AS/2006.04558]. [CrossRef]
- Kim, J.; Kim, S.; Kong, J.; Yoon, S. Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search. In Proceedings of the Advances in Neural Information Processing Systems; Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., Lin, H., Eds.; Curran Associates, Inc., 2020; Vol. 33, pp. 8067–8077. [Google Scholar]
- Casanova, E.; Shulby, C.; Gölge, E.; Müller, N.M.; de Oliveira, F.S.; Candido Jr., A.; da Silva Soares, A.; Aluisio, S.M.; Ponti, M.A. SC-GlowTTS: An Efficient Zero-Shot Multi-Speaker Text-To-Speech Model. In Proceedings of the Proc. Interspeech 2021, 2021; pp. 3645–3649. [Google Scholar] [CrossRef]
- Mehta, S.; Kirkland, A.; Lameris, H.; Beskow, J.; Éva Székely.; Henter, G.E. OverFlow: Putting flows on top of neural transducers for better TTS, 2022, [arXiv:eess.AS/2211.06892]. [CrossRef]
- Kumar, K.; Kumar, R.; de Boissiere, T.; Gestin, L.; Teoh, W.Z.; Sotelo, J.; de Brébisson, A.; Bengio, Y.; Courville, A.C. MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis. In Proceedings of the Advances in Neural Information Processing Systems; Wallach, H., Larochelle, H., Beygelzimer, A., Alché-Buc, F.d., Fox, E., Garnett, R., Eds.; Curran Associates, Inc., 2019; Vol. 32. [Google Scholar]
- gil Lee, S.; Ping, W.; Ginsburg, B.; Catanzaro, B.; Yoon, S. BigVGAN: A Universal Neural Vocoder with Large-Scale Training, 2023, [arXiv:cs.SD/2206.04658]. [CrossRef]
- Bak, T.; Lee, J.; Bae, H.; Yang, J.; Bae, J.S.; Joo, Y.S. Avocodo: Generative Adversarial Network for Artifact-free Vocoder, 2023, [arXiv:eess.AS/2206.13404]. [CrossRef]
- Kim, J.; Kong, J.; Son, J. Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech, 2021, [arXiv:cs.SD/2106.06103]. [CrossRef]
- Casanova, E.; Weber, J.; Shulby, C.; Junior, A.C.; Gölge, E.; Ponti, M.A. YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for everyone, 2023, [arXiv:cs.SD/2112.02418]. [CrossRef]
- Wang, C.; Chen, S.; Wu, Y.; Zhang, Z.; Zhou, L.; Liu, S.; Chen, Z.; Liu, Y.; Wang, H.; Li, J.; et al. Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers, 2023. arXiv:2301.02111 [cs, eess]. [CrossRef]
- Bakhturina, E.; Lavrukhin, V.; Ginsburg, B.; Zhang, Y. Hi-Fi Multi-Speaker English TTS Dataset, 2021. [CrossRef]
- Zen, H.; Dang, V.; Clark, R.; Zhang, Y.; Weiss, R.J.; Jia, Y.; Chen, Z.; Wu, Y. LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech. In Proceedings of the Interspeech 2019. ISCA, 2019; pp. 1526–1530. [CrossRef]
- Torres, H.M.; Gurlekian, J.A.; Evin, D.A.; Cossio Mercado, C.G. Emilia: a speech corpus for Argentine Spanish text to speech synthesis. Language Resources and Evaluation 2019, 53, 419–447. [Google Scholar] [CrossRef]
- Gabdrakhmanov, L.; Garaev, R.; Razinkov, E. RUSLAN: Russian Spoken Language Corpus for Speech Synthesis. In Proceedings of the Speech and Computer; Salah, A.A.; Karpov, A.; Potapova, R., Eds.; Springer International Publishing: Cham, 2019. Lecture Notes in Computer Science. pp. 113–121. [Google Scholar] [CrossRef]
- Srivastava, N.; Mukhopadhyay, R.; K R, P.; Jawahar, C.V. IndicSpeech: Text-to-Speech Corpus for Indian Languages. In Proceedings of the Proceedings of the Twelfth Language Resources and Evaluation Conference; European Language Resources Association: Marseille, France, 2020; pp. 6417–6422. [Google Scholar]
- Ahmad, A.; Selim, M.R.; Iqbal, M.Z.; Rahman, M.S. SUST TTS Corpus: A phonetically-balanced corpus for Bangla text-to-speech synthesis. Acoustical Science and Technology 2021, 42, 326–332. [Google Scholar] [CrossRef]
- Casanova, E.; Junior, A.C.; Shulby, C.; Oliveira, F.S.d.; Teixeira, J.P.; Ponti, M.A.; Aluísio, S. TTS-Portuguese Corpus: a corpus for speech synthesis in Brazilian Portuguese. Language Resources and Evaluation 2022, 56, 1043–1055. [Google Scholar] [CrossRef]
- Panayotov, V.; Chen, G.; Povey, D.; Khudanpur, S. Librispeech: An ASR corpus based on public domain audio books. In Proceedings of the 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2015; pp. 5206–5210. [CrossRef]
- Zandie, R.; Mahoor, M.H.; Madsen, J.; Emamian, E.S. RyanSpeech: A Corpus for Conversational Text-to-Speech Synthesis. In Proceedings of the Interspeech 2021. ISCA, 2021; pp. 2751–2755. [CrossRef]
- Torcoli, M.; Kastner, T.; Herre, J. Objective Measures of Perceptual Audio Quality Reviewed: An Evaluation of Their Application Domain Dependence. IEEE/ACM Transactions on Audio, Speech, and Language Processing 2021, 29, 1530–1541. [Google Scholar] [CrossRef]
- Mittag, G.; Möller, S. Deep Learning Based Assessment of Synthetic Speech Naturalness. In Proceedings of the Interspeech 2020. ISCA, 2020; pp. 1748–1752. [CrossRef]
- Huang, W.C.; Cooper, E.; Tsao, Y.; Wang, H.M.; Toda, T.; Yamagishi, J. The VoiceMOS Challenge 2022. In Proceedings of the Interspeech 2022. ISCA, 2022; pp. 4536–4540. [CrossRef]
- McAuliffe, M.; Sonderegger, M. Spanish MFA acoustic model v2.0.0a. Technical report, https://mfa-models.readthedocs.io/acoustic/Spanish/SpanishMFAacousticmodelv200a.html, 2022.
- McAuliffe, M.; Sonderegger, M. English MFA acoustic model v2.0.0a. Technical report, https://mfa-models.readthedocs.io/acoustic/English/EnglishMFAacousticmodelv200a.html, 2022.
- Ito, K.; Johnson, L. The LJ Speech Dataset. https://keithito.com/LJ-Speech-Dataset/, 2017.
- Kingma, D.P.; Ba, J. Adam: A Method for Stochastic Optimization. In Proceedings of the ICLR (Poster); Bengio, Y.; LeCun, Y., Eds., 2015.
- Rethage, D.; Pons, J.; Serra, X. A Wavenet for Speech Denoising. In Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018; pp. 5069–5073. [CrossRef]
- Yamagishi, J. English multi-speaker corpus for CSTR voice cloning toolkit. https://datashare.ed.ac.uk/handle/10283/3443/, 2012.
| 1 | |
| 2 | |
| 3 | |
| 4 | |
| 5 | |
| 6 | |
| 7 | |
| 8 |









| Speaker | Min. | Max | Mean | Median | Stdev |
|---|---|---|---|---|---|
| XRey | 0.72 | 32.81 | 21.15 | 21.27 | 3.84 |
| Tux valid | -20.34 | 97.90 | 36.48 | 37.36 | 8.75 |
| Hi-Fi TTS 92 | -17.97 | 73.86 | 40.57 | 39.48 | 9.62 |
| Speaker | Min. | Max | Mean | Median | Stdev |
|---|---|---|---|---|---|
| XRey | 11.68 | 20.11 | 15.46 | 15.42 | 1.09 |
| Tux valid | 6.25 | 22.22 | 12.20 | 12.23 | 1.49 |
| Hi-Fi TTS 92 | 7.46 | 26.32 | 15.03 | 15.14 | 1.46 |
| Speaker | All | 1High SNR | 2Utt. Speed | 1,2SNR & speed | ||||
| Files | Hours | Files | Hours | Files | Hours | Files | Hours | |
| XRey | 1 075 | 3:13 | 588 | 1:49 | 792 | 2:25 | 493 | 1:33 |
| 3 | 2 398 | 2:46 | 1 451 | 1:35 | 1 978 | 2:08 | 1 249 | 1:23 |
| Tux valid | 52 398 | 53:46 | 44 549 | 47:21 | 38 395 | 44:16 | 33 889 | 40:33 |
| 3 | 46 846 | 45:08 | 40 111 | 39:43 | 35 503 | 36:46 | 31 345 | 33:31 |
| 3 h partition | 3 092 | 3:00 | 2 649 | 2:39 | 2 326 | 2:27 | 2 061 | 2:15 |
| Hi-Fi TTS 92 | 35 296 | 27:18 | 31 634 | 25:02 | 25 975 | 21:40 | 23 996 | 20:16 |
| 3 | 33 589 | 26:34 | 30 374 | 24:27 | 25 838 | 21:19 | 23 889 | 19:58 |
| 3 h partition | 3 131 | 3:00 | 2 835 | 2:45 | 2 486 | 2:31 | 2 301 | 2:21 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).