Submitted:
25 September 2025
Posted:
26 September 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
- Speech Denoising Module: We introduce a noise suppression component before the audio encoding stage to purify the input signal, reducing the impact of background noise and enhancing the usability of audio features.
- Residual-based Audio-Visual Encoder: We propose a novel encoder architecture that leverages residual connections to improve cross-modal feature fusion and alignment, thereby enhancing modeling accuracy and visual output quality.
- Extensive Experimental Validation: We conduct comprehensive quantitative and qualitative experiments across multiple datasets to evaluate the proposed method in terms of lip-sync accuracy, facial detail restoration, and visual realism. The results consistently demonstrate superior synthesis quality and robustness compared to existing approaches.
2. Related Works
2.1. Audio-Driven Talking Head Synthesis
2.2. Speech Denoising Techniques
3. RAE-NeRF Architecture Design
3.1. DualPathZipformer Blocks
3.1.1. Down-UPSampleStacks
3.1.2. ZipformerBlock
3.1.3. Bypass
3.2. Residual-Based Audio-Visual Encoder
3.2.1. Network Architecture
- First FCResBlock: Input dimension = 512, Output dimension = 256;
- Second FCResBlock: Input dimension = 256, Output dimension = 128;
- Third FCResBlock: Input dimension = 128, Output dimension = .
3.2.2. Definition of FCResBlock

3.2.3. Down-UPSampleStacks
3.2.4. ZipformerBlock
3.2.5. Bypass
3.3. Residual-Based Audio-Visual Encoder
3.3.1. Network Architecture
- First FCResBlock: Input dimension = 512, Output dimension = 256;
- Second FCResBlock: Input dimension = 256, Output dimension = 128;
- Third FCResBlock: Input dimension = 128, Output dimension = .
3.3.2. Definition of FCResBlock

3.3.3. Overall Module Workflow
3.4. Tri-Plane Hash Representation
4. Experiments
4.1. Experimental Settings
4.2. Denoising Performance Analysis
4.3. Quantitative Evaluation
4.4. Qualitative Evaluation

4.5. Audio-video Encoder
5. Conclusion
References
- Thies, J.; Elgharib, M.; Tewari, A.; Theobalt, C.; Nießner, M. Neural voice puppetry: Audio-driven facial reenactment. In Proceedings of the Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XVI 16. Springer, 2020, pp. 716–731.
- Peng, Z.; Luo, Y.; Shi, Y.; Xu, H.; Zhu, X.; Liu, H.; He, J.; Fan, Z. Selftalk: A self-supervised commutative training diagram to comprehend 3d talking faces. In Proceedings of the Proceedings of the 31st ACM International Conference on Multimedia, 2023, pp. 5292–5301.
- Kim, H.; Garrido, P.; Tewari, A.; Xu, W.; Thies, J.; Niessner, M.; Pérez, P.; Richardt, C.; Zollhöfer, M.; Theobalt, C. Deep video portraits. ACM transactions on graphics (TOG) 2018, 37, 1–14. [CrossRef]
- Chen, L.; Maddox, R.K.; Duan, Z.; Xu, C. Hierarchical cross-modal talking face generation with dynamic pixel-wise loss. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 7832–7841.
- Prajwal, K.; Mukhopadhyay, R.; Namboodiri, V.P.; Jawahar, C. A lip sync expert is all you need for speech to lip generation in the wild. In Proceedings of the Proceedings of the 28th ACM international conference on multimedia, 2020, pp. 484–492.
- Zhou, Y.; Han, X.; Shechtman, E.; Echevarria, J.; Kalogerakis, E.; Li, D. Makelttalk: speaker-aware talking-head animation. ACM Transactions On Graphics (TOG) 2020, 39, 1–15. [CrossRef]
- Zhou, H.; Sun, Y.; Wu, W.; Loy, C.C.; Wang, X.; Liu, Z. Pose-controllable talking face generation by implicitly modularized audio-visual representation. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 4176–4186.
- Lu, Y.; Chai, J.; Cao, X. Live speech portraits: real-time photorealistic talking-head animation. ACM Transactions on Graphics (ToG) 2021, 40, 1–17. [CrossRef]
- Zhang, C.; Zhao, Y.; Huang, Y.; Zeng, M.; Ni, S.; Budagavi, M.; Guo, X. Facial: Synthesizing dynamic talking face with implicit attribute learning. In Proceedings of the Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 3867–3876.
- Guan, J.; Zhang, Z.; Zhou, H.; Hu, T.; Wang, K.; He, D.; Feng, H.; Liu, J.; Ding, E.; Liu, Z.; et al. Stylesync: High-fidelity generalized and personalized lip sync in style-based generator. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 1505–1515.
- Wang, J.; Qian, X.; Zhang, M.; Tan, R.T.; Li, H. Seeing what you said: Talking face generation guided by a lip reading expert. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 14653–14662.
- Zhang, Z.; Hu, Z.; Deng, W.; Fan, C.; Lv, T.; Ding, Y. Dinet: Deformation inpainting network for realistic face visually dubbing on high resolution video. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence, 2023, Vol. 37, pp. 3543–3551.
- Zhong, W.; Fang, C.; Cai, Y.; Wei, P.; Zhao, G.; Lin, L.; Li, G. Identity-preserving talking face generation with landmark and appearance priors. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 9729–9738.
- Peng, Z.; Hu, W.; Shi, Y.; Zhu, X.; Zhang, X.; Zhao, H.; He, J.; Liu, H.; Fan, Z. Synctalk: The devil is in the synchronization for talking head synthesis. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 666–676.
- Li, J.; Zhang, J.; Bai, X.; Zhou, J.; Gu, L. Efficient region-aware neural radiance fields for high-fidelity talking portrait synthesis. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 7568–7578.
- Tang, J.; Wang, K.; Zhou, H.; Chen, X.; He, D.; Hu, T.; Liu, J.; Zeng, G.; Wang, J. Real-time neural radiance talking portrait synthesis via audio-spatial decomposition. arXiv preprint arXiv:2211.12368 2022.
- Müller, T.; Evans, A.; Schied, C.; Keller, A. Instant neural graphics primitives with a multiresolution hash encoding. ACM transactions on graphics (TOG) 2022, 41, 1–15. [CrossRef]
- Chen, L.; Li, Z.; Maddox, R.K.; Duan, Z.; Xu, C. Lip movements generation at a glance. In Proceedings of the Proceedings of the European conference on computer vision (ECCV), 2018, pp. 520–535.
- Das, D.; Biswas, S.; Sinha, S.; Bhowmick, B. Speech-driven facial animation using cascaded gans for learning of motion and texture. In Proceedings of the Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXX 16. Springer, 2020, pp. 408–424.
- KR, P.; Mukhopadhyay, R.; Philip, J.; Jha, A.; Namboodiri, V.; Jawahar, C. Towards automatic face-to-face translation. In Proceedings of the Proceedings of the 27th ACM international conference on multimedia, 2019, pp. 1428–1436.
- Meshry, M.; Suri, S.; Davis, L.S.; Shrivastava, A. Learned spatial representations for few-shot talking-head synthesis. In Proceedings of the Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 13829–13838.
- Song, L.; Wu, W.; Qian, C.; He, R.; Loy, C.C. Everybody’s talkin’: Let me talk as you want. IEEE Transactions on Information Forensics and Security 2022, 17, 585–598. [CrossRef]
- Vougioukas, K.; Petridis, S.; Pantic, M. Realistic speech-driven facial animation with gans. International Journal of Computer Vision 2020, 128, 1398–1413. [CrossRef]
- Zhou, H.; Liu, Y.; Liu, Z.; Luo, P.; Wang, X. Talking face generation by adversarially disentangled audio-visual representation. In Proceedings of the Proceedings of the AAAI conference on artificial intelligence, 2019, Vol. 33, pp. 9299–9306.
- Sun, Y.; Zhou, H.; Wang, K.; Wu, Q.; Hong, Z.; Liu, J.; Ding, E.; Wang, J.; Liu, Z.; Hideki, K. Masked lip-sync prediction by audio-visual contextual exploitation in transformers. In Proceedings of the SIGGRAPH Asia 2022 Conference Papers, 2022, pp. 1–9.
- Wang, S.; Li, L.; Ding, Y.; Fan, C.; Yu, X. Audio2head: Audio-driven one-shot talking-head generation with natural head motion. arXiv preprint arXiv:2107.09293 2021.
- Mildenhall, B.; Srinivasan, P.P.; Tancik, M.; Barron, J.T.; Ramamoorthi, R.; Ng, R. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM 2021, 65, 99–106. [CrossRef]
- Wang, X.; Wang, C.; Liu, B.; Zhou, X.; Zhang, L.; Zheng, J.; Bai, X. Multi-view stereo in the deep learning era: A comprehensive review. Displays 2021, 70, 102102. [CrossRef]
- Zhang, P.; Zhou, L.; Bai, X.; Wang, C.; Zhou, J.; Zhang, L.; Zheng, J. Learning multi-view visual correspondences with self-supervision. Displays 2022, 72, 102160. [CrossRef]
- Wang, C.; Wang, X.; Zhang, J.; Zhang, L.; Bai, X.; Ning, X.; Zhou, J.; Hancock, E. Uncertainty estimation for stereo matching based on evidential deep learning. pattern recognition 2022, 124, 108498. [CrossRef]
- Zhang, J.; Wang, X.; Bai, X.; Wang, C.; Huang, L.; Chen, Y.; Gu, L.; Zhou, J.; Harada, T.; Hancock, E.R. Revisiting domain generalized stereo matching networks from a feature consistency perspective. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 13001–13011.
- Zhang, Y.; Chen, Y.; Bai, X.; Yu, S.; Yu, K.; Li, Z.; Yang, K. Adaptive unimodal cost volume filtering for deep stereo matching. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence, 2020, Vol. 34, pp. 12926–12934.
- Liu, X.; Xu, Y.; Wu, Q.; Zhou, H.; Wu, W.; Zhou, B. Semantic-aware implicit neural audio-driven video portrait generation. In Proceedings of the European conference on computer vision. Springer, 2022, pp. 106–125.
- Ye, Z.; Jiang, Z.; Ren, Y.; Liu, J.; He, J.; Zhao, Z. Geneface: Generalized and high-fidelity audio-driven 3d talking face synthesis. arXiv preprint arXiv:2301.13430 2023.
- Chatziagapi, A.; Athar, S.; Jain, A.; Rohith, M.; Bhat, V.; Samaras, D. LipNeRF: What is the right feature space to lip-sync a NeRF? In Proceedings of the 2023 IEEE 17th International Conference on Automatic Face and Gesture Recognition (FG). IEEE, 2023, pp. 1–8.
- Kim, E.; Seo, H. SE-Conformer: Time-Domain Speech Enhancement Using Conformer. In Proceedings of the Interspeech, 2021, pp. 2736–2740.
- Kong, Z.; Ping, W.; Dantrey, A.; Catanzaro, B. Speech denoising in the waveform domain with self-attention. In Proceedings of the ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, pp. 7867–7871.
- Zhao, S.; Ma, B.; Watcharasupat, K.N.; Gan, W.S. FRCRN: Boosting feature representation using frequency recurrence for monaural speech enhancement. In Proceedings of the ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, pp. 9281–9285.
- Valentini-Botinhao, C.; Yamagishi, J. Speech enhancement of noisy and reverberant speech for text-to-speech. IEEE/ACM Transactions on Audio, Speech, and Language Processing 2018, 26, 1420–1433. [CrossRef]
- Yu, G.; Li, A.; Zheng, C.; Guo, Y.; Wang, Y.; Wang, H. Dual-branch attention-in-attention transformer for single-channel speech enhancement. In Proceedings of the ICASSP 2022-2022 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2022, pp. 7847–7851.
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Advances in neural information processing systems 2017, 30.
- Luo, Y.; Chen, Z.; Yoshioka, T. Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech separation. In Proceedings of the ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020, pp. 46–50.
- Subakan, C.; Ravanelli, M.; Cornell, S.; Bronzi, M.; Zhong, J. Attention is all you need in speech separation. In Proceedings of the ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021, pp. 21–25.
- Chen, H.; Yu, J.; Weng, C. Complexity scaling for speech denoising. In Proceedings of the ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, pp. 12276–12280.
- Burchi, M.; Vielzeuf, V. Efficient conformer: Progressive downsampling and grouped attention for automatic speech recognition. In Proceedings of the 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2021, pp. 8–15.
- Kim, S.; Gholami, A.; Shaw, A.; Lee, N.; Mangalam, K.; Malik, J.; Mahoney, M.W.; Keutzer, K. Squeezeformer: An efficient transformer for automatic speech recognition. Advances in Neural Information Processing Systems 2022, 35, 9361–9373.
- Yao, Z.; Guo, L.; Yang, X.; Kang, W.; Kuang, F.; Yang, Y.; Jin, Z.; Lin, L.; Povey, D. Zipformer: A faster and better encoder for automatic speech recognition. arXiv preprint arXiv:2310.11230 2023.
- Wang, H.; Tian, B. ZipEnhancer: Dual-Path Down-Up Sampling-based Zipformer for Monaural Speech Enhancement. arXiv preprint arXiv:2501.05183 2025.
- Guo, Y.; Chen, K.; Liang, S.; Liu, Y.J.; Bao, H.; Zhang, J. Ad-nerf: Audio driven neural radiance fields for talking head synthesis. In Proceedings of the Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 5784–5794.
- Shen, S.; Li, W.; Zhu, Z.; Duan, Y.; Zhou, J.; Lu, J. Learning dynamic facial radiance fields for few-shot talking head synthesis. In Proceedings of the European conference on computer vision. Springer, 2022, pp. 666–682.
- Loshchilov, I. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101 2017.
- Roux, J.L.; Wisdom, S.; Erdogan, H.; Hershey, J.R. SDR - Half-baked or Well Done? In Proceedings of the ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019, pp. 626–630. [CrossRef]
- Kubichek, R. Mel-cepstral distance measure for objective speech quality assessment. In Proceedings of the Proceedings of IEEE Pacific Rim Conference on Communications Computers and Signal Processing, 1993, Vol. 1, pp. 125–128 vol.1. [CrossRef]
- Zhang, R.; Isola, P.; Efros, A.A.; Shechtman, E.; Wang, O. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 586–595.
- Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; Hochreiter, S. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems 2017, 30.
- Zhang, W.; Liu, Y.; Dong, C.; Qiao, Y. Ranksrgan: Generative adversarial networks with ranker for image super-resolution. In Proceedings of the Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 3096–3105.
- Mittal, A.; Soundararajan, R.; Bovik, A.C. Making a "completely blind" image quality analyzer. IEEE Signal processing letters 2012, 20, 209–212. [CrossRef]
- Mittal, A.; Moorthy, A.K.; Bovik, A.C. No-reference image quality assessment in the spatial domain. IEEE Transactions on image processing 2012, 21, 4695–4708. [CrossRef]






| Methods | PSNR ↑ | SSIM ↑ | LPIPS ↓ | FID ↓ |
|---|---|---|---|---|
| Wav2Lip | 34.849174 | 0.987339 | 0.016931 | 5.581930 |
| ER-NeRF | 36.577897 | 0.997066 | 0.008733 | 5.736807 |
| SyncTalk | 42.492073 | 0.999254 | 0.003496 | 1.548124 |
| RAE-NeRF | 42.642164 | 0.999274 | 0.003332 | 1.401045 |
| Methods | LMD ↓ | BRISQUE ↓ | NIQE ↓ |
|---|---|---|---|
| Wav2Lip | 2.0289 | 52.9446 | 7.4609 |
| ER-NeRF | 1.9174 | 52.9446 | 7.4267 |
| SyncTalk | 1.9961 | 49.9054 | 7.3830 |
| RAE-NeRF | 1.9090 | 49.8398 | 7.3829 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).