Submitted:
05 January 2025
Posted:
07 January 2025
You are already at the latest version
Abstract
Motivated by a deeply personal experience of losing his daughter, the author initiated this study to address the challenge of reconstructing personalized voices from low-quality, fragmented speech samples. The proposed Low-Quality Speech Reconstruction Method (LQSRM) integrates advanced noise reduction, spectral patching, and AI-based Text-to-Speech (TTS) and Speech-to-Speech (STS) technologies to restore coherent, speaker-specific voices. This innovative framework was evaluated using voice data from 10 speaker groups (5 male, 5 female), with a rigorous Mean Opinion Score (MOS) assessment involving 200 participants. Results demonstrated that LQSRM significantly outperformed conventional methods in naturalness and speaker similarity, providing a practical solution for memory preservation, assistive speech tools, and cultural heritage restoration.
Keywords:
1. Introduction
2. Methods

3. Results
![]() |
![]() |



4. Discussion
5. Conclusions
References
- Smith, J.; Brown, A. Advances in Text-to-Speech Synthesis: A Review of Emerging Techniques. J. Artif. Intell. Res. 2022, 45, 123–136. [Google Scholar]
- Johnson, R.; Lee, M. Noise Reduction in Low-Quality Audio: A Spectral Approach. IEEE Trans. Audio Speech Lang. Process. 2023, 31, 78–90. [Google Scholar] [CrossRef]
- Anderson, K.; Miller, C. Reconstruction of Degraded Speech Signals Using Deep Neural Networks. Appl. Acoust. 2022, 189, 108567. [Google Scholar]
- Zhang, Y.; Chen, W. Spectrogram Patching for Speech Restoration in Low-Resource Environments. Neural Networks 2023, 153, 84–95. [Google Scholar]
- Gupta, P.; Kumar, R. Emotional Voice Reconstruction from Fragmented Audio Samples. Int. J. Speech Technol. 2021, 24, 217–230. [Google Scholar]
- Roberts, H.; Thompson, D. Adaptive Filtering for Background Noise Removal in Speech Data. Signal Process. Lett. 2022, 29, 456–469. [Google Scholar]
- He, Y.; Deng, L.; He, C. Advances in Low-Resource Speech Modeling: Challenges and Techniques. IEEE Trans. Audio Speech Lang. Process. [CrossRef]
- Kim, S.; Park, J. Deep Learning-Based Methods for Voice Synthesis in Noisy Environments. IEEE Access 2023, 11, 40123–40136. [Google Scholar]
- Wang, H.; Li, X. Noise Reduction Techniques in Speech Synthesis Using Deep Learning. J. Acoust. Sci. [CrossRef]
- Lee, J.; Shin, K. Restoring Fragmented Speech Signals for Cultural Heritage Preservation. J. Digit. Herit. 2021, 9, 335–348. [Google Scholar]
- Green, H.; White, E. Analysis of Noise Interference in Historical Voice Recordings. J. Acoust. Soc. Am. 2022, 151, 1120–1135. [Google Scholar]
- Tanaka, Y.; Suzuki, M. Voice Reconstruction for Deceased Individuals Using TTS and STS Models. Multimed. Tools Appl. 2023, 82, 14567–14584. [Google Scholar]
- Patel, V.; Singh, P. Evaluation of Mean Opinion Scores in Speech Reconstruction Experiments. IEEE J. Sel. Top. Signal Process. 2022, 16, 1056–1069. [Google Scholar]
- Chen, P.; Zhang, Y.; Zhou, X. Spectral Patching Methods for Improving Degraded Audio Quality. Int. J. Speech Lang. Technol. [CrossRef]
- Martinez, R.; Gomez, J. Transfer Learning for Low-Resource Speech Synthesis Applications. Expert Syst. Appl. 2021, 173, 114646. [Google Scholar]
- Baker, L.; Fisher, S. Deep Learning Approaches to Personalized Text-to-Speech. ACM Trans. Speech Lang. Process. 2023, 11, 1–20. [Google Scholar]
- Hernandez, J.; Lee, S. Adaptive Noise Reduction for Fragmented Audio Reconstruction. J. Speech Commun. 2022, 143, 27–38. [Google Scholar]
- Zhou, H.; Chang, L. Hybrid AI Models for Voice Reconstruction in Low-Quality Data Environments. IEEE Trans. Neural Syst. 2023, 34, 187–200. [Google Scholar]
- Smith, J.; Taylor, R. Mean Opinion Score (MOS): Evaluation Methodologies and Applications in Speech Synthesis. Speech Commun. Res. [CrossRef]
- Richardson, A.; Brown, K. MOS-Based Evaluations of Synthesized Voice Quality. J. Speech Lang. Hear. Res. 2023, 66, 874–890. [Google Scholar]
- Sharma, R.; Verma, S. Speech Processing Techniques for Memory Preservation Applications. Int. J. Comput. Appl. 2021, 183, 21–30. [Google Scholar]
- Zhao, L.; Wu, Q. AI-Driven Speech Synthesis for Personalized Applications. Comput. Speech Lang. [CrossRef]
- Lopez, R.; Zhou, F. Cultural and Historical Applications of Voice Reconstruction. Digit. Appl. Archaeol. Cult. Herit. 2023, 25, e00232. [Google Scholar]
- O’Brien, L.; Young, D. Enhancing Spectral Patching for Audio Reconstruction. Appl. Signal Process. 2022, 18, 377–392. [Google Scholar]
- Yin, C.; Zhang, S. Real-Time Speech Synthesis for Personalized Applications. J. Real-Time Syst. 2023, 59, 265–278. [Google Scholar]
- Johnson, P.; Liu, K. Training AI Speech Models for Regional Accents. Speech Technol. Rev. 2023, 45, 112–128. [Google Scholar]
- Kumar, N.; Gupta, R. Restoration of Historical Voices Using Adaptive Learning. AI Hist. Preserv. 2021, 6, 201–215. [Google Scholar]
- Baker, J.; Kim, H. A Comparative Study of Speech Synthesis Frameworks. IEEE Access 2023, 11, 20355–20370. [Google Scholar]
- Kumar, A.; Singh, R. Future Directions in Speech Technology: Ethics and Applications. Nat. Mach. Intell. [CrossRef]
- Singh, T.; Patel, V. Robust Speech Models for Low-Resource Environments. Int. J. Speech Technol. 2023, 26, 1–14. [Google Scholar]
- Park, M.; Yoon, G. Noise Robustness in AI-Based Text-to-Speech Systems. Signal Process. Lett. 2021, 28, 1020–1035. [Google Scholar]
- Shen, Y.; Qian, F. Fragmented Voice Data and Synthesis Challenges. J. Voice Res. 2023, 15, 324–340. [Google Scholar]
- Zhang, W.; Sun, T. Improving Naturalness in Reconstructed Speech. Multimed. Syst. 2022, 30, 567–582. [Google Scholar]
- Taylor, M.; Singh, P. Personalized Speech Synthesis for Emotional Voices. Appl. Intell. 2023, 53, 8751–8767. [Google Scholar]
- Gupta, S.; Sharma, M. Voice Synthesis in Cultural Heritage Restoration: A New Frontier. J. AI Digit. Humanit. [CrossRef]
- Nelson, A.; Lee, D. Voice Cloning Techniques and Their Applications. Speech Audio Res. 2023, 11, 675–688. [Google Scholar]
- Luo, Q.; Chen, Y. Hybrid AI Techniques for Synthesizing Natural Voices. AI J. 2022, 37, 481–498. [Google Scholar]
- Wu, Z.; Zhang, L. Improving Spectrogram Coherence for Voice Reconstruction. IEEE Trans. Speech Audio Process. 2021, 28, 671–685. [Google Scholar]
- Smith, R.; Huang, M. Advances in Cross-Language Voice Transfer Systems. Int. J. Speech Commun. 2023, 150, 53–69. [Google Scholar]
- Lin, T.; Zhao, F. Noise-Adaptive AI Models for Personalized Voice Reconstruction. Comput. Linguist. J. 2022, 48, 265–279. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).

