Submitted:
30 December 2024
Posted:
30 December 2024
You are already at the latest version
Abstract
Keywords:
1. Introduction
1.1. What Is a Colloquial Malayalam Speech Converter API?
1.2. Technological Components Used in the Colloquial Malayalam Speech Converter API
- Speech Recognition (ASR): This is the core component of the system, converting spoken Malayalam dialects into text. The API leverages Wav2Vec 2.0 and SpeechBrain models to accurately capture the various nuances in Malayalam dialects, ensuring the correct conversion of colloquial speech to standard Malayalam.
- Translation Engine: After the speech-to-text conversion, a translation engine (e.g., GoogleTrans API) is used to translate the standardized Malayalam text into English. This translation facilitates cross-linguistic communication and makes the output useful for non-Malayalam speakers.
- Text-to-Speech (TTS): Once the translation is completed, the Google Text-to-Speech (gTTS) engine is employed to convert the translated English text back into speech. This ensures that the output is available in both text and auditory forms, providing accessibility for users who prefer spoken output.
- Natural Language Processing (NLP): NLP models help enhance the system’s understanding of different colloquial expressions and dialect variations. This component ensures that the API can accurately interpret a variety of regional Malayalam phrases and dialectal speech patterns.
1.3. Linguistic Diversity in Malayalam
1.4. Relevance in Kerala and India
1.5. Global and Regional Impact
2. Literature Review
3. Conclusions
References
- V. Brydinskyi, D. Sabodashko, Y. Khoma, M. Podpora, A. Konovalov, and V. Khoma, "Enhancing Automatic Speech Recognition With Personalized Models: Improving Accuracy Through Individualized Fine-Tuning," IEEE Access, vol. 12, pp. 116649-116656, 2024. [CrossRef]
- H. Inaguma, K. Duh, T. Kawahara, and S. Watanabe, "Multilingual End-to-End Speech Translation," arXiv, 2019. [CrossRef]
- R. S. A. Podila, G. S. S. Kommula, R. K., S. Vekkot, and D. Gupta, "Telugu Dialect Speech Dataset Creation and Recognition Using Deep Learning Techniques," 2022 IEEE 19th India Council International Conference (INDICON), Kochi, India, 2022, pp. 1-6. [CrossRef]
- Yerramreddy D. R. , J. Marasani, P. S. V. Gowtham, G. Harshit, and Anjali, "Speech Recognition Paradigms: A Comparative Evaluation of SpeechBrain, Whisper and Wav2Vec2 Models," 2024 IEEE 9th International Conference for Convergence in Technology (I2CT), Pune, India, 2024, pp. 1-6. [CrossRef]
- Penkova B, M. Mitreska, K. Ristov, K. Mishev, and M. Simjanoska, "Learning Translation Model to Translate Croatian Dialects to Modern Croatian Language," 2023 46th MIPRO ICT and Electronics Convention (MIPRO), Opatija, Croatia, 2023, pp. 1083-1088. [CrossRef]
- Kaneb and M. Kodad, "Literature Review: NLP Techniques for Arabic Dialect Recognition," 2024 International Conference on Circuit, Systems and Communication (ICCSC), Fes, Morocco, 2024, pp. 1-5. [CrossRef]
- Unni V. , N. Joshi, and P. Jyothi, "Coupled Training of Sequence-to-Sequence Models for Accented Speech Recognition," ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain, 2020, pp. 8254-8258. [CrossRef]
- X. Shi, et al., "The Accented English Speech Recognition Challenge 2020: Open Datasets, Tracks, Baselines, Results and Methods," ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada, 2021, pp. 6918-6922. [CrossRef]
- Chen L. W. and Rudnicky A. , "Exploring Wav2Vec 2.0 Fine-Tuning for Improved Speech Emotion Recognition," arXiv, 2021. [CrossRef]
- Gao Q., H. Wu, Y. Sun, and Y. Duan, "An End-to-End Speech Accent Recognition Method Based on Hybrid CTC/Attention Transformer ASR," ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada, 2021, pp. 7253-7257. [CrossRef]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).