Submitted:
14 August 2026
Posted:
18 August 2026
You are already at the latest version
Abstract
Talking face generation (TFG), the synthesis of photorealistic speaking video from a portrait and an audio signal, has moved through five architectural generations since 2016: Long Short-Term Memory (LSTM) lip-sync systems, Generative Adversarial Networks (GANs), Neural Radiance Fields (NeRFs), Diffusion models, and now Diffusion Transformers (DiT) and Gaussian Splatting. With real-time photorealism, TFG is entering healthcare, education, identity management, and public communication. This systematic literature review synthesises 92 primary studies (January 2016 to December 2025), selected from 18,343 records across six repositories under Kitchenham’s SLR guidelines and PRISMA 2020, to answer six research questions on architecture evolution, dataset bias, evaluation metrics, generative paradigm shifts, ethical safeguards, and socio-technical deployment. DiT and Gaussian Splatting together account for roughly 36% of the most recent studies retrieved and GAN usage has fallen below 3%. Training datasets are demographically skewed, and the dominant metrics (PSNR, SSIM, SyncNet) do not measure what deployment requires. Most seriously, only one of the 92 systems embeds any technical safeguard (a non-compliant post-hoc watermark), even though the EU AI Act (Regulation 2024/1689), the US TAKE IT DOWN Act (2025), and Coalition for Content Provenance and Authenticity (C2PA) v2.0 entered enforcement during the review period. None integrates C2PA-compliant watermarking, consent, or provenance mechanisms. We propose three deployment-fitness metrics: Perceived Trustworthiness Score (PTS), Cross-Cultural Authenticity Index (CCAI), and Deepfake Detectability Rate (DDR); and the Responsible Talking Face Generation in Socio-Technical Systems (RTFG-STS) Framework, grounded in AI4People principles and socio-technical systems theory, with five operational mechanisms aligned to EU AI Act Article 14.
Keywords:
talking face generation
; socio-technical systems
; deepfake governance
; systematic literature review
; responsible AI
; human-AI interaction
1. Introduction
1.1. Background and Motivation
Modern deep learning can now synthesize a photorealistic speaking face from a single portrait photograph and an audio clip. Talking face generation (TFG), illustrated in Figure 1, is the automated creation of temporally coherent video in which a face speaks in synchrony with a given audio stream. The problem sits at the intersection of computer vision, audio signal processing, and generative modelling, and in one decade it has grown from a narrow scientific question into a technology with commercial weight and social consequences.
The technical lineage is short but dense. Bregler et al. [1] showed in the 1990s that vocal-tract articulation is tightly coupled with visible facial motion, making data-driven synthesis plausible, and the statistical models of the 2000s formalized the mapping without approaching perceptual realism. Deep learning changed the trajectory. The Synthesizing Obama system of Suwajanakorn et al. [2] demonstrated that broadcast-quality lip synchronization could be learned directly from audio-visual corpora, without hand-crafted features. Generative Adversarial Networks (GANs) [3] then became the generative backbone of choice, and Wav2Lip [4] achieved lip-sync fidelity good enough for in-the-wild use. Large audio-visual corpora such as VoxCeleb [5], VoxCeleb2, and HDTF [6] supplied the training signal to push generalization beyond single-speaker settings.
Three further transitions followed from around 2021. Transformer architectures such as FaceFormer [7] brought long-range audio-visual context modelling; Neural Radiance Fields (NeRFs) [8] were adapted to dynamic portraits in AD-NeRF [9], giving 3D-consistent talking heads; and Diffusion Models [10] then displaced GANs, with DiffTalk [11] and DreamTalk [12] delivering probabilistic, high-fidelity synthesis with better training stability and diversity. By 2024, Diffusion Transformers (DiT) and Gaussian Splatting had begun to remove the long-standing latency barrier: VASA-1 [13] produces 512×512 video at 45 fps in real time, and OmniHuman-1 [14] extends the DiT paradigm to full-body digital human animation.
1.2 Applications and Open Problems
TFG has become socio-technical in the sense of Trist and Bamforth [15] and later socio-technical systems (STS) theory [16]: its deployment consequences are now inseparable from the human and institutional contexts it operates in. In healthcare, TFG-based conversational agents are one response to a projected global shortage of 11.1 million health workers by 2030 [17]; AI physician avatars can deliver patient education with perceived empathy comparable to human clinicians [18], and real-time systems such as MuseTalk [19] and VASA-1 [13] make live telehealth interaction feasible, with emotionally adaptive models such as EmotiveTalk [20] pointing toward mental-health applications. In education, embodied agents with realistic faces improve engagement and retention over disembodied voices or text [21], support virtual reality (VR) anatomy and clinical-communication training [22], and, with dyadic systems such as INFP [23], open up conversational skills training. In media and personalized communication, TFG enables multilingual dubbing, lip-readable synthetic media for hearing-impaired audiences, and consumer applications of latent-space face synthesis [13,24,25]; public institutions can use the same capability for low-cost multilingual civic communication.
1.3. The Regulatory Turn
The potential harms of TFG have triggered a wave of regulation. The EU AI Act (Regulation 2024/1689), adopted in 2024 and fully enforceable from August 2026, classifies deepfakes and synthetic facial media as a transparency risk under Article 50: generated content must be labelled, users must know when they are interacting with an AI-generated face, and Article 50(2) with Recital 134 mandates watermarking, metadata identifiers, and provenance tooling so synthetic media stays machine-detectable [26]. In the United States, the TAKE IT DOWN Act (Public Law 119-12, May 2025) criminalizes distribution of non-consensual AI-generated intimate imagery, with prison terms of up to two years (three where the subject is a minor) and a 48-hour platform takedown obligation [27]. C2PA v2.0 provides cryptographically signed provenance manifests already adopted by major platforms, and China’s synthetic media labelling rules took effect on 1 September 2025 [28].
The timing matters as much as the content. VASA-1 [13] and OmniHuman-1 [14] were published and deployed before any of these instruments entered force. That gap between capability and governance is a familiar pattern in socio-technical systems [16], and closing it requires the kind of systems-level account of TFG that this review sets out to provide.
1.4. Gap, Objectives, and Research Questions
Surveys of TFG and adjacent fields have appeared steadily since 2022 [29,30,31,32,33,34], and Section 2 analyses them in detail: none addresses the regulatory environment, dataset equity, or the socio-technical systems TFG is deployed into, and no existing review is simultaneously rigorous, current to 2024-2025, and comprehensive across technical, ethical, and socio-technical dimensions.
This review addresses that gap through six research questions:
RQ1: What deep learning architectures have dominated TFG from 2016 to 2025, and what trajectory characterizes their evolution?
RQ2: How do biases in the datasets used to train and evaluate TFG models affect generalization, and what are the equity implications for deployment?
RQ3: What metrics are currently used to evaluate TFG systems, and how adequately do they capture the requirements of socio-technical deployment?
RQ4: How have Diffusion Transformers, Gaussian Splatting, and related modern architectures advanced the quality, efficiency, and controllability of TFG relative to GANs and classical NeRFs?
RQ5: What ethical safeguards (detection, consent, watermarking, provenance) are integrated into state-of-the-art TFG systems, and how does this align with emerging regulatory requirements?
RQ6: How does TFG embed within, and generate risks and opportunities for, the socio-technical systems in which it is deployed, including healthcare, education, identity management, and democratic governance?
1.5. Contributions and Structure
This review contributes: (i) an architectural taxonomy extended to Gaussian Splatting and DiT, mapping the TFG design space from 2016 to 2025; (ii) the first systematic assessment of demographic representation across the most widely used TFG benchmark datasets; (iii) a critical analysis of the evaluation metric ecosystem and three proposed deployment-fitness metrics, the Perceived Trustworthiness Score (PTS), Cross-Cultural Authenticity Index (CCAI), and Deepfake Detectability Rate (DDR); (iv) documentation of the safeguard gap, showing that no reviewed system natively integrates watermarking, consent, or provenance, mapped against the EU AI Act, the TAKE IT DOWN Act, and C2PA v2.0; (v) a socio-technical analysis of four deployment domains; (vi) the Responsible Talking Face Generation in Socio-Technical Systems (RTFG-STS) Framework, grounded in Floridi et al.’s AI4People principles [35] and operationalized through five mechanisms aligned to EU AI Act Article 14; and (vii) a formalized three-stage process pipeline for TFG.
2. Related Surveys and Reviews
Secondary literature on TFG and its neighbors falls into three groups: deepfake generation and detection reviews, TFG-specific surveys, and reviews of adjacent domains. We assess each against four dimensions: methodological rigor, temporal coverage, thematic scope, and socio-technical perspective. Figure 2 summarizes the comparison and shows that no prior review satisfies all four.
2.1. Deepfake Generation and Detection
Rana et al. [29] surveyed deepfake detection (2018-2020) with a declared methodology adapted from Kitchenham and Charters [36]; generation, datasets, and ethics are excluded by design, and the window predates the transformer, diffusion, and Gaussian Splatting transitions. Heidari et al. [33] extended this line to 2021 with an artefact-based detection taxonomy, closing before diffusion models arrived, whose outputs defeat detectors calibrated to GAN artefacts. Romero Moreno [37] analyses deepfakes through human rights law and evaluates the EU AI Act; it is a legal analysis rather than a technical survey, but its treatment of consent and provenance as rights-enabling mechanisms directly informs this review’s RQ5 and the framework in Section 5.
2.2. Talking Face and Talking Head Surveys
Zhen et al. [30] surveyed talking-head generation (2017-2022) with a 2D/3D, pipeline/end-to-end classification but declare no methodology: no search string, databases, criteria, or quality assessment, and the survey ends before DiffTalk [11], VASA-1 [13], and OmniHuman-1 [14] existed. Toshpulatov et al. [31] add technical depth and a Convolutional Neural Network (CNN)/GAN/NeRF taxonomy, likewise without a protocol, and a frontier placed at NeRFs is now a full architectural generation out of date. Bigioi and Corcoran [32] frame the field through multilingual dubbing, also without a declared method. Rakesh et al. [34] is the only prior survey to reach diffusion and Gaussian Splatting, but it too follows no systematic method and is entirely technical: nothing on ethics, governance, or deployment. Nisar et al. [38] take an application-oriented view, among the first to treat TFG as socially consequential, though the coverage is illustrative rather than systematic.
2.3. Adjacent Domains
Three adjacent literatures matter here. Audio-visual speech learning supplied both the field’s dominant training corpus (VoxCeleb [5], used in over 25 of the 92 included studies) and its dominant synchrony metric, SyncNet [39], built for correspondence detection in broadcast television rather than naturalness assessment, a design intent central to the RQ3 analysis. 3D face reconstruction supplied the parametric priors that intermediate-representation systems such as SadTalker [40] and GeneFace [41] depend on; such priors average over their training demographics and can encode bias geometrically, which is relevant to RQ2. The embodied conversational agent (ECA) literature supplies the socio-technical evidence base: Provoost et al. [21] found ECA-administered clinical interviews can match human-administered assessment in specific contexts, and Bickmore and Cassell [42] showed that trust in agents depends on human-like appearance and conversational naturalness. These findings motivate RQ6 and the proposed RTFG-STS Framework.
2.4. Comparative Analysis and Gaps
Four gaps follow from Figure 2. First, methodological rigor: only the two detection reviews [29,33] declare a systematic method, and every survey of TFG generation itself [30,31,32,34] is non-reproducible, which biases coverage toward the most-cited work. Second, temporal obsolescence: no methodologically rigorous review reaches the diffusion or Gaussian Splatting eras at all. Third, a governance blind-spot: no TFG survey addresses the EU AI Act [26], the TAKE IT DOWN Act [27], C2PA v2.0, or China’s labelling rules [28], despite these creating concrete design obligations within the review period. Fourth, no socio-technical analysis: every existing review treats TFG as a benchmark optimization problem. The literature explains how to build a talking face system; it says almost nothing about what happens when one enters a clinical workflow, a curriculum, a KYC (Know Your Customer) pipeline, or a political information environment. This review is designed to close all four gaps at once.
3. Methodology
3.1. Review Framework and Rationale
The review follows the SLR methodology of Kitchenham and Charters [36] in three phases (planning, conducting, reporting), with study identification and reporting per PRISMA 2020 [43]. A scoping review would be inappropriate because RQ3 and RQ5 need quality-filtered synthesis to support defensible conclusions. A meta-analysis is impossible because TFG studies report incompatible metric subsets of FID, SyncNet, PSNR, SSIM, LPIPS, and FVD on incompatible test sets; making statistical pooling meaningless in this context. The Kitchenham protocol accommodates qualitative synthesis of heterogeneous studies while keeping selection transparent and auditable. Figure 3 presents the PRISMA 2020 flow diagram.
3.2. Scope
The review covers deep learning approaches to TFG across the full pipeline from feature extraction to rendering: audio-driven single-image animation, audio-plus-video reenactment, 3D-aware synthesis (NeRF and Gaussian Splatting), diffusion and DiT architectures, and detection or safeguard work directly tied to TFG assessment. Coverage runs from January 2016, when deep recurrent architectures first reached talking face synthesis, to December 2025, fixed at the fourth search iteration in March 2026. Pre-2016 methods (HMMs, DBNs, AAMs) fall outside the deep learning pipeline the research questions examine. Because RQ5 and RQ6 require it, the disciplinary scope extends past computer vision into applied ethics, AI governance, and STS analysis.
3.3. Search Strategy
Keywords were identified in three stages: a scoping search that yielded 47 highly cited papers and 31 candidate terms in three semantic clusters; forward snowballing through the 20 most-cited candidates, adding five terms; and a 2025 revision adding frontier terminology. The final Boolean string, applied to titles, abstracts, and keywords, was:
(“talking face” OR “talking head” OR “audio-driven face” OR “speech-driven face” OR “audio-visual synthesis” OR “portrait animation” OR “lip sync” OR “lip synchronization” OR “video reenactment” OR “deepfake” OR “facial animation deep learning” OR “neural talking” OR “video dubbing” OR “face reenactment” OR “Gaussian splatting avatar” OR “diffusion portrait” OR “facial diffusion transformer” OR “audio-driven digital human”)
The string was executed in tiers: a core tier (“talking face,” “talking head,” “lip sync”) in every iteration; an extended tier (deepfake, portrait animation, reenactment, dubbing) from the second iteration; and a frontier tier only in the fourth, so that terminology which did not exist before 2025 could not distort coverage of older work. Six repositories were searched for disciplinary relevance and complementary coverage: IEEE Xplore, ScienceDirect, the ACM Digital Library, SpringerLink, Google Scholar with Web of Science, and arXiv, where landmark systems such as VASA-1 appeared months before formal publication. Four iterations ran between March 2023 and March 2026; a single frozen search would have missed the diffusion and Gaussian Splatting literature entirely. Table 1 reports retrieval counts.
3.4 Selection Criteria and Procedure
Table 2 states the inclusion and exclusion criteria, which operationalize the scope above. All of I1-I6 must hold for inclusion and any one of E1-E6 excludes.
Selection ran in three stages. Deduplication (DOI match, then title-plus-first-author matching at a 0.95 Levenshtein threshold) reduced 18,343 records to 8,741, and coarse title screening against I3-I5 left 448 for abstract review. Two reviewers screened all 448 abstracts independently (Cohen’s κ = 0.84), excluding 118 (off-topic DL application n = 64; not primary research n = 31; temporal n = 14; language n = 9); eight full texts proved unretrievable, and the remaining 322 were assessed against all criteria and the quality instrument below. Of these, 230 were excluded (opaque methodology n = 78; topic mismatch n = 91; tutorial or chapter n = 31; other n = 30), leaving 92 studies in the synthesis corpus.
3.5. Quality Assessment
Each of the 322 full texts was scored on five criteria adapted from Kitchenham and Charters [36]: problem clarity (QA1), methodological transparency (QA2), evaluation rigor (QA3), limitation acknowledgement (QA4), and contribution significance (QA5). Each criterion scores 1.0, 0.5, or 0.0; the inclusion threshold of 3.0 of 5.0 excludes fundamental deficiencies while admitting honestly reported early-stage work. Scores of the 322 full texts ranged from 0.5 to 5.0; 92 studies (28.6%) met the threshold, with a mean of 4.27 (SD = 0.39) among those included. QA1 was the criterion most often fully satisfied (93.5%), QA4 the most often only partial (41.3% scored 0.5), and QA2 was the most discriminating: studies scoring 0.0 on methodological transparency never exceeded 2.0 overall, making opacity the main driver of exclusion.
3.6. Extraction and Synthesis
A structured extraction form, piloted on 10 studies and refined once, captured bibliographic metadata, technical characteristics, datasets and metrics with quantitative results, five ethical and governance sub-fields (detection, watermarking/provenance, consent, risk discussion, regulatory reference), and deployment characteristics. The lead author extracted all 92 studies and a co-author independently verified a stratified 25% subsample (23 studies), with disagreements resolved against the source text and, in four cases, by email to the original authors. Given the heterogeneity of study types and outcomes, we adopted qualitative narrative synthesis following Cochrane guidance [44]: thematic synthesis within each RQ, followed by a cross-RQ analysis of patterns no single question exposes; quantitative results are tabulated where comparable but never statistically pooled. Where fewer than three independent studies support a finding, this is flagged in Section 4. The full bibliographic record of all 92 included studies, with extracted characteristics and quality scores, is provided as supplementary material (Appendix A).
3.7. Corpus Characteristics
The corpus comprises 56 conference papers (60.8%), 24 journal articles (26.1%), and 12 preprints (13.0%), consistent with a field whose flagship venues are conferences: CVPR leads with 22, followed by ECCV (9), ACM Multimedia (8), ICCV (7), and others. Annual output grows from 2 included studies in 2016-2017 to a roughly 20 per year across 2023-2025.
3.8. Threats to Validity
Four threats deserve note. The English-language restriction and database selection likely underrepresent Chinese-language venues, a real limitation given how much TFG research originates from Chinese institutions. Publication bias favors positive results; including arXiv preprints softens but does not remove it. The frontier described here is accurate to March 2026 and may be partially superseded at reading time; the multi-iteration search was designed to limit this decay. Finally, quality scoring involves judgement, particularly on QA3 and QA5; the κ = 0.84 agreement and explicit scoring descriptors constrain but do not eliminate subjectivity. RQ6 rests on a thinner evidence base than RQ1-RQ4, since primary studies rarely report deployment outcomes; the analysis there leans on the ECA literature and regulatory documents.
4. Data Synthesis and Analysis
4.1. Corpus Overview
The 92 studies span 2016 to 2025 and nine architecture families, drawing on 19 distinct named training and evaluation corpora (plus custom single-identity datasets) and 17 distinct evaluation metrics. GAN-based work is the largest single family (22 studies, 23.9%), a legacy of its 2018-2022 dominance, but the three post-GAN families together, Diffusion Models (16, 17.4%), Diffusion Transformers (7, 7.6%), and Gaussian Splatting (8, 8.7%), account for 33.7% of the corpus and overtake it. Two numbers foreshadow the RQ5 findings: only 8 studies (8.7%) contain any ethical content, and only 1 (1.1%) embeds a safeguard mechanism. Fewer than half (44, 47.8%) state a deployment domain at all.
4.2. The TFG Process Pipeline
Across all nine families, virtually every reviewed system decomposes into three sequential modules: feature extraction, a mapping network, and a transformation network (Figure 4). This shared scaffold is what makes cross-era comparison possible in RQ1, RQ3, and RQ4.
Feature extraction encodes the source image and audio sequence into latent representations:
where is an image encoder (a face-recognition backbone, or a VAE encoder in diffusion systems) and an audio encoder (a mel-spectrogram CNN or a pre-trained speech encoder). The mapping network learns the cross-modal correspondence between audio features and facial motion:
This is the most architecturally variable component in the corpus, realised as recurrent lip-displacement regression (GAN era), audio-visual cross-attention (transformers), audio-conditioned deformation fields (NeRF), or conditioning vectors injected into the denoising network (diffusion); may cover lips, 3DMM expression coefficients, pose, gaze, and blinks. The transformation network then synthesises each frame from the source image and motion parameters:
The form of is what distinguishes the architectural eras: an adversarially trained generator with combined adversarial, lip-synchrony (SyncNet [39]), identity, and perceptual losses in GAN systems; iterative denoising conditioned on identity and motion in diffusion systems [10]; and, in Gaussian Splatting systems, differentiable rasterisation of learned 3D Gaussians whose position, opacity, covariance, and colour parameters are jointly optimised from the audio-conditioned motion prior.
4.3. RQ1: Architecture Evolution, 2016-2025
Figure 5 shows the distribution of included studies by primary architecture family and year. Five phases are identifiable, each anchored by a specific breakthrough.
Phase I, recurrent foundations (2016-2018). LSTM and bidirectional RNN models mapped mel-spectrogram frames to lip displacements: Fan et al. [45] set the deep recurrent baseline, and Suwajanakorn et al. [2] reached broadcast quality with 17 hours of single-speaker footage, at the cost of strict speaker dependence.
Phase II, GAN ascendancy (2019-2022). Adversarial training lifted the quality ceiling: Zhou et al. [46] disentangled audio and visual representations for speaker-agnostic synthesis, Wav2Lip [4] achieved the first in-the-wild lip sync by training against a pre-trained SyncNet discriminator, and MakeItTalk [47] and PC-AVS [48] added speaker-aware animation and pose control.
Phase III, transformers and NeRFs (2021-2023). FaceFormer [7] replaced recurrence with self-attention for long-range temporal coherence, while AD-NeRF [9] conditioned a radiance field on audio, giving 3D-consistent talking portraits at the price of hours of per-identity optimization and inference far below real time.
Phase IV, diffusion (2023-2024). DiffTalk [11] and DreamTalk [12] showed that audio- and identity-conditioned diffusion beats GANs on diversity and artefacts without mode collapse; SadTalker [40] added pose-controllable synthesis without per-identity tuning; and VASA-1 [13] closed the phase with a disentangled face latent space and flow matching, at 512×512 and 45 fps in real time.
Phase V, DiT and Gaussian Splatting (2024-2025). Two paradigms now co-define the frontier: DiTs replace the U-Net with a transformer over latent tokens and scale quality with compute [49], with OmniHuman-1 [14] dissolving the boundary between TFG and full-body digital humans, while Gaussian Splatting systems (GSTalker [50]; GaussianEmoTalker [51]; VASA-3D [52]) replace volume rendering with fast rasterization and remove the latency problem outright. Table 3 summarizes the full taxonomy.
The trajectory is a series of paradigm shifts, not a smooth curve, and each shift has an identifiable cause. GANs gave way to diffusion because mode collapse produced systematic artefacts in high-motion regions (teeth and hair most visibly). NeRFs gave way to Gaussian Splatting because volume rendering could not meet real-time requirements. U-Net diffusion is giving way to DiT because transformers scale better with parameter count. Phase V is distinguished by simultaneity: DiT-scale models produce the highest quality but demand server-class hardware, Gaussian Splatting runs in real time at quality sufficient for most deployments, and a hybrid, DiT-generated motion rendered through Gaussian Splatting, is the most promising route to closing that gap. Of the 41 corpus studies dated 2024 or later, 15 (36.6%) are DiT or Gaussian Splatting systems, while only one is GAN-based (2.4%).
4.4. RQ2: Dataset Biases and Generalization
Figure 6 shows dataset usage across the corpus: VoxCeleb2 dominates (60 studies, 65.2%), followed by HDTF (36, 39.1%) and MEAD (18, 19.6%); because most systems train and evaluate on several corpora, counts sum to more than the corpus size. A further 12 studies (13.0%) rely wholly or partly on custom single-identity corpora, and the 3D vertex-animation corpora VOCASET and BIWI (3 studies each) serve the FaceFormer line of work. Figure 6(b) characterizes the most used corpora.
The most consistent finding is under-documentation. Of the 19 named corpora identified, only CREMA-D reports ethnic diversity statistics in its own paper, and CREMA-D appears in exactly one included study. VoxCeleb2, which underpins nearly two-thirds of the corpus, reports subject counts and languages but nothing systematic about ethnicity, geography, or age. Studies running cross-demographic evaluations on it ([S16], [S21], [S31]) consistently note that South Asian, Middle-Eastern, African, and Latin American faces are severely underrepresented, several estimating these groups at under 8% of subjects combined. HDTF is built from 362 news and political broadcast subjects, predominantly white and East Asian male speakers of American or British English.
The skew is measurable in performance. Multiple studies report degraded lip-sync confidence when VoxCeleb2-trained models are evaluated on South Asian, African, and Middle-Eastern subjects; exact figures vary, but the direction is consistent and matches the face-recognition fairness literature [64]. The practical meaning is stark: a telehealth avatar trained on the dominant benchmark will articulate less accurately for the faces and languages of Sub-Saharan Africa, South Asia, and the Middle East, precisely where clinician shortages make AI-mediated care most valuable.
Two further gaps compound this. Affective coverage is thin: only MEAD (18 studies) and CREMA-D (1 study) carry emotion labels, and neither exceeds 91 subjects, so the emotion disentanglement reported by systems such as EmotiveTalk [20] and EAT [65] rests on very small affective corpora. Linguistic coverage is thinner still: 14 of the 19 named corpora are English-only, the VoxCeleb pair and CelebV-HQ are nominally multilingual but Western-dominant in practice, and ViCo (2 studies) is the sole substantively non-English corpus. Models trained on English phonemes produce implausible lip shapes for tonal languages (Mandarin, Thai, Yoruba) and for complex consonant clusters (Arabic, Polish).
The problem is not dataset size but diversity and documentation. Fixing it requires demographic audits as a publication prerequisite, open release of phonemically and ethnically diverse corpora, and demographic-stratified evaluation of the kind standard in face-recognition fairness work [64]; EU AI Act Article 10, requiring training data representative of the served population, now supplies regulatory pressure in the same direction.
4.5. RQ3: Evaluation Metrics and Deployment Fitness
Figure 7 shows metric frequency across the corpus. Distribution- and synchrony-level measures lead: FID appears in 60 studies (65.2%) and SyncNet-derived synchrony scores in 51 (55.4%), with the signal-fidelity pair SSIM (43, 46.7%) and PSNR (34, 37.0%) close behind and human MOS evaluation in a third of studies (33, 35.9%). Figure 7(b) assesses each metric family against three deployment-fitness criteria.
Three problems stand out. First, the dominance of PSNR and SSIM is hard to defend. Both measure pixel correspondence to a specific reference, which conflates quality with reference proximity: a slightly blurred but naturally animated face can score below a sharp but temporally incoherent one. The community has known PSNR is a poor perceptual proxy since at least 2012 [66]; its persistence in 2024-2025 papers is evaluation inertia, and it matters: a video that scores well on PSNR while looking uncanny to a patient is not fit for use as a virtual clinician.
Second, SyncNet [39] was trained on English broadcast correspondences, and no included study validates it on non-English phonemic inventories. It also measures correspondence, not naturalness: minimal lip movement scores well on neutral phonemes while missing the jaw openings, bilabial closures, and fricative patterns that make speech look real. Models optimized hard against SyncNet drift toward stiff, over constrained lip motion that is metrically accurate and perceptually unconvincing.
Third, the metrics that matter most for deployment remain rare or shallow. CSIM, the only standardized identity-preservation measure, appears in 15.2% of studies, and FVD, which captures the temporal coherence that per-frame metrics miss, in 13.0%. Human MOS evaluation is more widespread (35.9%) but typically shallow: panels of fewer than 30 raters are the norm, below accepted sample sizes for reliable MOS work, and no study recruits raters from its intended deployment population. The one metric whose adoption is clearly rising, FPS (15.2%), tracks the field’s real-time turn rather than any perceptual or governance property. Human perception remains the fitness measure that deployment actually turns on.
We therefore propose three metric dimensions as a research agenda. The Perceived Trustworthiness Score (PTS) is a human evaluation protocol in which raters drawn from the target deployment population (patients for healthcare avatars, students for educational agents) score credibility, warmth, and professional appropriateness on a multi-item Likert scale; it supplements rather than replaces MOS. The Cross-Cultural Authenticity Index (CCAI) extends SyncNet-style evaluation across at least five typologically distinct language families (for example Indo-European, Sino-Tibetan, Afro-Asiatic, Niger-Congo, Dravidian), exposing the cross-lingual limits of English-trained models and creating an incentive for diverse training data. The Deepfake Detectability Rate (DDR) is the proportion of generated frames correctly classified as synthetic by an ensemble of state-of-the-art detectors; it bears directly on EU AI Act Article 50, and systems below 80% DDR should carry mandatory C2PA watermarking as a compensating control. All three are proposed with provisional parameters and need validation through large, demographically diverse human studies before serving as standards.
4.6. RQ4: DiT and Gaussian Splatting Versus GANs and NeRFs
Figure 8 compares four architecture eras on visual quality, lip synchrony, and real-time capability; Table 4 reports benchmark data for representative systems.
On quality, the FID trajectory is unambiguous: GAN and early hybrid systems report FIDs in the 41-48 range (2021-2023), early diffusion at 35-47 in 2023, VASA-1 at 22.1, OmniHuman-1 at 18.4, a 55% relative reduction from the GAN baseline in four years, a pattern matching image generation more broadly, where DiT architectures keep improving predictably with scale.
On latency, 2024-2025 resolved the long-standing trade-off in three separate ways: flow matching (VASA-1 [13]) cut inference-time function evaluations without losing diversity; Gaussian rasterization (GSTalker [50]; VASA-3D [52]) replaced volume rendering for an order-of-magnitude speedup; and autoregressive streaming (Teller [60]) predicts motion tokens at sub-frame latency for continuous live output. Three architecturally unrelated solutions arriving at once suggest the barrier is genuinely down, not narrowly circumvented.
Controllability advanced in parallel and is invisible to FID and SyncNet. Wav2Lip [4] controlled lips only; PC-AVS [48] added pose; SadTalker [40] added stylistic motion via 3DMM coefficients; EmotiveTalk [20] decouples audio into content and emotion for independent affective control; and VASA-1 [13] disentangles pose, expression, gaze, blink, and speaking style into orthogonal latent dimensions with fine-grained post-hoc editing. For deployment this is not a luxury: a tutor that cannot adapt its emotional register teaches worse, and a clinician avatar with one fixed expression reads as uncanny.
4.7. RQ5: Ethical Safeguards and Regulatory Alignment
Figure 9 presents the starkest finding in the corpus. Of 92 included studies, only 8 (8.7%) contain any ethical content at all, and exactly 1 (1.1%) integrates a technical safeguard: a frequency-domain watermark added as a post-processing step, non-C2PA-compliant and untested against compression. No study implements a consent protocol. No study cites the EU AI Act, the TAKE IT DOWN Act, C2PA, GDPR Article 9, or any other governance instrument as a design constraint. Table 5 details the inventory.
Regulatory instruments are now in force and the gap is comprehensive. EU AI Act Article 50(2) requires machine-detectable marking by August 2026 (no reviewed system embeds a compliant manifest); Recital 134 requires marks robust to compression and conversion (no study tests robustness); Article 10 requires representative training data (no study audits its dataset demographically); and Article 14 requires meaningful human oversight (no system includes an oversight or revocation mechanism). The TAKE IT DOWN Act’s consent conditions go unaddressed, and C2PA is not referenced in a single included paper.
We do not read this as individual negligence: publication incentives reward benchmark quality, ethics review does not yet require safeguard implementation, and liability-creating regulation arrived only at the end of the review period. The result is a compliance cliff: a large body of deployable state-of-the-art systems facing EU AI Act obligations from August 2026 while satisfying none of them. Section 5 offers an operational path across.
4.8. RQ6: TFG in Socio-Technical Systems
Figure 10 shows the coded deployment domains. Most studies (48, 52.2%) specify no domain, treating TFG as a context-free technical problem; among the rest, entertainment leads (18, 19.6%), then education (9), healthcare (7), identity management (5), governance (3), and accessibility (2). The analysis below covers the four domains most consequential for this review, drawing on the 24 domain-specific studies plus the adjacent literature from Section 2.
Healthcare. DeTore et al. [67] found TFG-powered avatars could administer validated clinical instruments (PHQ-9, GAD-7) with response quality comparable to human-administered assessment; Haider et al. [18] found avatar-delivered patient education matched human instruction on comprehension while cutting clinician time by 34%. The risk profile has three parts: patients must know they are talking to an AI (an Article 50 obligation no healthcare-focused study reports implementing); the RQ2 performance gaps translate into reduced comprehension of clinical information, with patient safety implications; and liability for incorrect AI-delivered medical advice is unassigned in every jurisdiction we could identify.
Education. Chheang et al. [22] showed higher knowledge retention with TFG-powered VR anatomy assistants than with text; EmotiveTalk [20] sketches emotionally adaptive tutoring at scale, though the cross-lingual limits from RQ2 constrain language learning uses. The risks are distinctive: interactive tutors that read student facial affect are collecting biometric data governed by GDPR Article 9 and FERPA, which no included study addresses; the authority effect of face-to-face communication [42] may lead students to under-scrutinize AI-delivered content; and convergence on a few commercial avatar products could homogenize pedagogy.
Identity management. This is where the threat is already quantifiable. Venkatesh et al. [68] show that presentation attack detection (PAD) systems generalize poorly to attack types unseen in training. PAD calibrated on GAN artefacts (spectral inconsistencies, flicker, blink anomalies) can fail outright against the smooth, physiologically coherent output of flow-matching systems such as VASA-1 [13] or Gaussian Splatting systems such as GSTalker [50], which produce none of those signatures. The iProov threat intelligence report [69] documents sharply rising real-world PAD bypasses through 2023-2024, coinciding with diffusion TFG availability, and neither VASA-1 nor GSTalker has been formally evaluated against current liveness detection in any peer-reviewed study we identified. The institutional consequences are already on record: the USD 25 million Arup fraud [70], and Pindrop’s finding that over a third of a 300-profile applicant sample was wholly fabricated. A convincing deepfake video interview now costs a consumer GPU and a public model checkpoint. In systems terms, this erodes the trust infrastructure of digital identity itself [16]; no single countermeasure suffices, which is why Section 5 argues for coordinated governance across standards, law, platforms, and institutional practice.
Governance and democratic information. Political deepfakes of real politicians were documented in at least seven national election contexts between 2022 and 2024, and photorealistic talking faces raise the credibility and virality of false information, with lower-income and lower-digital-literacy populations disproportionately exposed. Detection is caught in an arms race: as generation improves (RQ4), detectors trained on older outputs decay. Mandatory provenance (EU AI Act Article 50; C2PA v2.0) is a structural exit from that race, binding content cryptographically to its generating model at creation so verification no longer depends on out-detecting generation quality. Table 6 summarizes the four domains.
4.9. Cross-RQ Patterns
Four patterns emerge only at the intersection of the research questions.
Optimization-deployment misalignment. The field optimizes PSNR, SSIM, and SyncNet; the deployment contexts that matter require identity preservation, cross-cultural authenticity, trustworthiness, and detectability, which the field does not measure. This reflects incentive structure, not oversight, and it will persist until venues and funders require PTS, CCAI, and DDR alongside FID.
Democratization and concentration at once. Gaussian Splatting and lightweight diffusion push high-fidelity TFG onto consumer hardware, while DiT-scale models (OmniHuman-1 [14]; EMO2 [71]) concentrate frontier capability in organizations with large compute budgets. The result is a two-tier field in which the hardest-to-detect systems are least accessible for beneficial use, and the most accessible systems are both the most useful and the easiest to misuse.
The absent safeguard as structural failure. That 91.3% of studies contain no ethical content and 98.9% no safeguard reflects the field’s incentive design. Ethics review at NeurIPS (2021) and CVPR (2023) has not measurably shifted the pattern, suggesting domain-specific requirements (watermarking, consent, DDR reporting) need to become submission conditions.
Governance lag as predictable risk. Each capability jump (2019 GAN photorealism, 2023 diffusion fidelity, 2024 real-time Gaussian Splatting) was followed by roughly 18-24 months of misuse before any governance response. The EU AI Act’s technology-neutral design may weather the next transition better than capability-specific rules; whether enforcement is agile enough is open.
5. The RTFG-STS Framework
5.1. Theoretical Foundations
The framework’s ethical core is the AI4People report of Floridi et al. [35], which joined the four bioethical principles of Beauchamp and Childress [72] (beneficence, non-maleficence, autonomy, justice) with a fifth, AI-specific principle, explicability. AI4People underpins the European Commission’s AI ethics guidelines and the philosophy of the EU AI Act [26], and it specifies what an ethical system must do, not only what it must avoid; on its own, though, it addresses the AI system in isolation. Socio-technical systems theory [15,16] supplies the missing half: a virtual clinician must be judged by its behavior inside a clinical workflow, its effect on patient trust, and its interaction with consent regulation, not by FID alone. Bijker’s concept of socio-technical stabilization [16] explains why the governance lag in Section 4.9 is costly: governance built after a technology hardens into institutional practice faces one that is already difficult to reshape. The RTFG-STS Framework therefore applies AI4People’s principles at the level of the deployed socio-technical system and adds a sixth principle, systemic embeddedness, which AI4People lacks and the STS analysis demands.
5.2. Six Ethical Principles of the RTFG-STS Framework
Table 7 states the six principles, their interpretation for TFG, and the operational mechanisms (Section 5.3) that implement each.
5.3. Five Operational Mechanisms
M1: C2PA-compliant watermarking and provenance (P2, P5). Every deployed TFG system must embed a cryptographically signed provenance manifest per C2PA v2.0, satisfying EU AI Act Article 50(2) and Recital 134:
The manifest encodes the provenance claim, the generating model’s identity hash, a timestamp, hashes of the inputs, and the issuer’s certificate, signed with a key on the C2PA Trust List. Recital 134 requires robustness to H.264/H.265 compression, rescaling, and format conversion, a requirement the corpus’s single watermarking study does not meet; DCT-domain embedding with error-correction coding is current best practice pending a C2PA-TFG robustness standard, and systems should report the proportion of frames whose manifest survives compression, treating results below 80% as a barrier to high-risk deployment.
M2: Risk-tiered consent (P2, P3). Consent obligations scale with deployment risk, in three tiers aligned to the EU AI Act’s classification (Table 8). At Tier 3, consent is stored as a Verifiable Credential:
where the DIDs identify the consent subject and system operator, consentScope bounds permitted uses, and the revocation endpoint lets the subject withdraw consent, obliging removal of synthesized content within 48 hours, the TAKE IT DOWN Act timeframe [27].
M3: Socio-Technical Impact Assessment (P1, P4, P6). Before deployment beyond controlled research, the system undergoes a STIA, modelled on the EU AI Act’s Fundamental Rights Impact Assessment (Article 27) but extended to the full deployment context, and maintained as a living document across major revisions. Its six modules assess the target socio-technical system and the TFG system’s role in it; demographic performance equity (stratified evaluation plus CCAI); effects on the behavior, trust, and wellbeing of interacting humans (via PTS); the institutional processes and authority structures the system will intersect; the feedback loops likely to emerge over time; and regulatory compliance against the EU AI Act, the TAKE IT DOWN Act, GDPR, C2PA, and jurisdiction-specific rules, with a remediation timeline for gaps.
M4: Deepfake detectability audit (P2, P4, P5). DDR is defined as
where is the i-th generated frame and an ensemble of at least three state-of-the-art detectors, computed after H.264 compression at 2,000 kbps to approximate platform distribution. Three thresholds apply: at DDR ≥ 80%, content is reliably machine-detectable and a C2PA manifest suffices; between 60% and 80%, C2PA plus interaction-point disclosure is mandatory and Tier 3 deployment requires a regulatory waiver; below 60%, Tier 2-3 deployment is prohibited pending improvement, with research use permitted under manifest. The 80% floor follows NIST SP 800-76-2 guidance for detection rates in government identity contexts [73]; the 60% boundary matches the empirical detection floors reported for GAN-era detectors facing novel synthesis modalities [33,68]. Both are provisional pending a standardized TFG-specific DDR benchmark suite.
M5: Human oversight layer (P3, P5, P6). For Tier 3 deployments, a human must stand between generated output and consequential action, per EU AI Act Article 14. Concretely: a qualified clinician reviews clinical content an avatar delivers, with the AI nature disclosed to the patient; a trained reviewer verifies identity through a non-video channel before any decision rests on a video interview, with C2PA verification mandatory; a certified forensic reviewer attests provenance for TFG content in legal proceedings; and TFG video of public figures carries Article 50-compliant disclosure under the editorial responsibility of a publisher who accepts liability.
5.4. Application Across Deployment Domains
Table 9 translates the framework into concrete obligations for the four domains analysed in RQ6.
5.5. Framework Limitations
The framework is a normative proposal derived from the synthesis findings, AI4People, and STS theory; it has not been validated in live deployments, and pilot studies in healthcare and education are needed to test whether the STIA and consent tiers are operationally feasible. It is calibrated to the regulatory picture of December 2025 and should be re-checked against regulation annually. DDR depends on the currency of its detection ensemble: an ensemble calibrated to GAN artefacts says nothing valid about Gaussian Splatting output, so it must be refreshed annually, ideally against a public benchmark maintained by standards bodies. Finally, the framework specifies what responsible deployment requires but cannot enforce it; enforcement will come from venue ethics policies, funder conditions, and supervisory authorities, and advocacy toward venue-level adoption at CVPR, NeurIPS, and ACM MM is the most direct route to uptake.
6. Conclusions
6.1. Principal Findings
Across 92 primary studies from 2016 to 2025, the picture is of a field whose technical progress has consistently outrun its governance and evaluation. Architecturally (RQ1), TFG crossed five phases in nine years; the 2024-2025 frontier is contested between DiT and Gaussian Splatting, which together account for roughly 36% of the most recent studies and have jointly resolved the historic latency-quality trade-off. The training data beneath this achievement (RQ2) skews toward East Asian and Western European faces, with no dataset paper in the corpus reporting a demographic audit, so geographic and linguistic bias is written into the model weights serving the very populations the data underrepresents. Evaluation (RQ3) still runs on PSNR, SSIM, and SyncNet, none of which measures identity preservation, cross-cultural naturalness, trustworthiness, or detectability. Modern architectures (RQ4) delivered a 55% FID reduction over the GAN baseline and three independent routes to real time, yet this capability exists without safeguards (RQ5): zero systems with compliant watermarking, zero consent protocols, zero references to any regulatory instrument as a design constraint. The consequences (RQ6) are no longer hypothetical: USD 25 million lost to a single deepfake video call, systematic liveness bypass, and measurable erosion of trust in digital identity.
Read through STS theory, these findings share one cause: the field’s technological frame, its goals, metrics, and rewards, is built entirely around benchmark optimization, with no structural place for governance, equity, or deployment fitness. Widening that frame is beyond any individual researcher; it needs coordinated action from venues, funders, regulators, and the institutions now deploying TFG.
6.2. Implications
For researchers, the most actionable change is ethics-by-design reporting: DDR alongside FID and SyncNet, disclosed demographic audits of training data, and a stated deployment context with a fitness assessment against it. The cost is modest next to building a frontier system. For practitioners, the message is blunter: no currently available state-of-the-art system meets the obligations EU AI Act Article 50 imposes from August 2026, so deploying VASA-1, MuseTalk, or OmniHuman-1 in a patient-facing, student-facing, or KYC context today without C2PA watermarking, Tier 3 consent, and human oversight is a compliance liability with a dated deadline. For policymakers, three interventions exceed any current instrument: public funding for an open, demographically diverse, multilingual TFG dataset program; extension of ISO/IEC SC37 presentation attack standards to TFG synthesis attacks; and platform liability for hosting synthetic media without C2PA verification, which does not depend on winning the detection arms race.
6.3. Limitations
Four limitations apply. The English-language restriction underrepresents Chinese-language venues despite Chinese institutions being among the field’s most active contributors. The regulatory analysis is dated to December 2025 and may be sharpened by the EU AI Office’s forthcoming Code of Practice. The RQ6 analysis rests on 24 domain-specific studies plus adjacent literature, so its claims about political communication and education are propositions to test rather than settled findings.
6.4. Closing Statement
TFG can now synthesize photorealistic, real-time, emotionally controllable speech video from one image and an audio clip. Whether we can build it is settled; whether we can govern it is not. The EU AI Act, the TAKE IT DOWN Act, and C2PA v2.0 are the regulatory answers, and the RTFG-STS Framework is ours: embed consent, provenance, equity, and deployment fitness as design prerequisites rather than afterthoughts. Whether the next nine years produces systems that serve the rural clinic and the multilingual classroom without also arming the fraudster and the misinformed is not a technical question. It is a socio-technical one, and it deserves a socio-technical answer.
Author Contributions
Conceptualization, R. Salahudeen; methodology, R. Salahudeen and A. Garba; software, R. Salahudeen and A. Garba; data curation, R. Salahudeen; writing—original draft preparation, R. Salahudeen; writing—review and editing, R. Salahudeen; visualization, R. Salahudeen; supervision, M. Fonkam, N.R. Vajjhala and S.B. Junaidu. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Acknowledgments
“During the preparation of this manuscript/study, the author(s) used MS Word Gmail and Google Drive for the purposes of editing, correspondence, sharing and versioning. The authors have reviewed and edited the output and take full responsibility for the content of this publication.
Conflicts of Interest
“The authors declare no conflicts of interest.”.
Appendix A
Studies are organised by primary architecture family and chronological order within each family. QA scores are out of 5.0; minimum inclusion threshold = 3.0. Arch. = Architecture; Diff. = Diffusion; Transf. = Transformer; GS = Gaussian Splatting; DiT = Diffusion Transformer. Citations marked † are arXiv preprints confirmed as active at time of search.
Table A1.
Complete List of Included Primary Studies (N = 92).
| No. | Authors (Year) | Abbreviated Title | Venue |
Arch. Family |
Training Dataset(s) |
Key Metrics Reported |
QA Score |
| LSTM / RNN (n = 5) | |||||||
| S1 | Fan et al. (2016) [45] | A Deep Bidirectional LSTM Approach for Video-Realistic Talking Head | Multimed. Tools Appl. | LSTM/RNN | Custom (single-spk) | PSNR, SSIM, MOS | 3.5 |
| S2 | Suwajanakorn et al. (2017) [2] | Synthesizing Obama: Learning Lip Sync from Audio | ACM TOG (SIGGRAPH) | LSTM/RNN | Obama Weekly Address | Lip Sync Error, MOS | 4.8 |
| S3 | Song et al. (2018) [53] | Talking Face Generation by Conditional Recurrent Adversarial Network | IJCAI 2018 | LSTM/RNN | Grid, LRW | PSNR, SSIM, LMD | 3.8 |
| S4 | Vougioukas et al. (2020) [74] | Realistic Speech-Driven Facial Animation with GANs | IJCV 2020 | LSTM/RNN | GRID, VoxCeleb | FID, PSNR, SyncNet | 4.0 |
| S5 | Chen et al. (2019) [75] | Hierarchical Cross-Modal Talking Face Video Generation | ACM MM 2019 | LSTM/RNN | LRW, VoxCeleb | PSNR, SSIM, LMD | 3.5 |
| CNN-based (n = 7) | |||||||
| S6 | Zhou et al. (2018) [54] | VisemeNet: Audio-Driven Animator-Centric Speech Animation | ACM TOG (SIGGRAPH) | CNN | CMU-MOSI, LRW | MOS, CPBD | 4.0 |
| S7 | Siarohin et al. (2019) [76] | First Order Motion Model for Image Animation | NeurIPS 2019 | CNN | VoxCeleb, BAIR | FID, SSIM, AKD | 4.5 |
| S8 | Jamaludin et al. (2019) [77] | You Said That?: Synthesising Talking Faces from Audio | IJCV 2019 | CNN | BBC News, LRS2 | Sync Error, PSNR | 3.8 |
| S9 | Wiles et al. (2018) [78] | X2Face: A Network for Controlling Face Generation | ECCV 2018 | CNN | VoxCeleb, iHead | SSIM, FID, MOS | 3.8 |
| S10 | Wang et al. (2021) [79] | One-Shot Free-View Neural Talking-Head Synthesis | CVPR 2021 | CNN | VoxCeleb2, HDTF | FID, CSIM, AED | 4.2 |
| S11 | Cheng et al. (2022) [80] | VideoReTalking: Audio-Based Lip Sync for Talking Head Editing | SIGGRAPH Asia 2022 | CNN | HDTF, VoxCeleb2 | SyncNet, LMD, FID | 4.0 |
| S12 | Das et al. (2020) [81] | Speech-Driven Facial Animation Using Cascaded GANs | ECCV 2020 | CNN | VoxCeleb, GRID | PSNR, SSIM, SyncNet | 3.5 |
| GAN-based (n = 22) | |||||||
| S13 | Zhou et al. (2019) [46] | Talking Face Generation by Adversarially Disentangled A-V Representation | AAAI 2019 | GAN | VoxCeleb, LRW | LMD, PSNR, MOS | 4.5 |
| S14 | Prajwal et al. (2020) [4] | Wav2Lip: A Lip Sync Expert Is All You Need for Speech to Lip Generation | ACM MM 2020 | GAN | LRS2, VoxCeleb2 | SyncNet, LSE-D, LSE-C | 5.0 |
| S15 | Zhou et al. (2020) [47] | MakeItTalk: Speaker-Aware Talking-Head Animation | ACM TOG 2020 | GAN | VoxCeleb, GRID | LMD, FID, MOS | 4.5 |
| S16 | Salahudeen et al. (2024) [24] | Photo-Realistic Talking Face Generation under Latent Space Manipulation | IEEE TCE 2024 | GAN | VoxCeleb2, MEAD | FID, SyncNet, PSNR, SSIM | 4.2 |
| S17 | Salahudeen & Siu (2023) [25] | Activate Your Face in Virtual Meeting Platform | ICCE-Taiwan 2023 | GAN | VoxCeleb2 | SyncNet, SSIM | 3.2 |
| S18 | Zhou et al. (2021) [48] | Pose-Controllable Talking Face Generation (PC-AVS) | CVPR 2021 | GAN | VoxCeleb2 | FID, SyncNet, CSIM | 4.8 |
| S19 | Ji et al. (2021) [82] | Audio-Driven Emotional Video Portraits | CVPR 2021 | GAN | MEAD, VoxCeleb2 | FID, PSNR, MOS | 4.2 |
| S20 | Wang et al. (2021) [83] | Audio2Head: Audio-Driven One-Shot Talking-Head Generation | IJCAI 2021 | GAN | VoxCeleb, HDTF | FID, SyncNet, SSIM | 4.0 |
| S21 | Hong et al. (2022) [84] | Depth-Aware Generative Adversarial Network for Talking Head | CVPR 2022 | GAN | VoxCeleb2, HDTF | FID, CPBD, LMD | 4.2 |
| S22 | Yin et al. (2022) [85] | StyleHEAT: One-Shot High-Resolution Editable Talking Face via StyleGAN | ECCV 2022 | GAN | VoxCeleb2, HDTF | FID, PSNR, SSIM, LPIPS | 4.5 |
| S23 | Liang et al. (2022) [86] | Expressive Talking Head Generation with Granular Audio-Visual Control | CVPR 2022 | GAN | MEAD, VoxCeleb2 | FID, MOS, AUE | 4.0 |
| S24 | Song et al. (2022) [87] | Talking Face Generation with Multilingual TTS | CVPR 2022 | GAN | HDTF, LRW | SyncNet, FID, LMD | 4.0 |
| S25 | Lu et al. (2021) [88] | Live Speech Portraits: Real-Time Photorealistic Talking-Head Animation | ACM TOG 2021 | GAN | Custom | PSNR, SSIM, MOS | 4.5 |
| S26 | Zhu et al. (2021) [89] | Arbitrary Talking Face Generation via Attentional A-V Coherence Learning | IJCAI 2021 | GAN | VoxCeleb, MEAD | FID, SyncNet, SSIM | 3.8 |
| S27 | Wen et al. (2020) [90] | Photorealistic Audio-Driven Video Portraits | IEEE TVCG 2020 | GAN | Custom (Obama) | PSNR, SSIM, MOS | 4.0 |
| S28 | Zhang et al. (2021) [6] | Flow-Guided One-Shot Talking Face Generation | ACM MM 2021 | GAN | VoxCeleb2, HDTF | FID, SyncNet, LPIPS | 4.0 |
| S29 | Hwang et al. (2023) [91] | DisCoHead: Disentangled Control of Head Pose and Facial Expressions | ICASSP 2023 | GAN | VoxCeleb2, HDTF | FID, LMD, SyncNet | 3.8 |
| S30 | Ji et al. (2022) [92] | EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion Model | ACM SIGGRAPH 2022 | GAN | MEAD, VoxCeleb2 | FID, PSNR, SSIM, AUE, MOS | 4.5 |
| S31 | Yi et al. (2022) [93] | Predicting personalized head movement from short video and speech signal | IEEE TMM 2022 | GAN | VoxCeleb2, MEAD | PSNR, SSIM, LMD, MOS | 4.0 |
| S32 | Wang et al. (2022) [94] | One-Shot Talking Face Generation from Single-Speaker A-V Correlation | AAAI 2022 | GAN | VoxCeleb2 | FID, SyncNet, CSIM | 3.8 |
| S33 | Peng et al. (2023) [95] | SelfTalk: Self-Supervised Training for 3D Talking Faces | ACM MM 2023 | GAN | MEAD, VoxCeleb2 | FID, AUE, MOS | 4.0 |
| S34 | Chen et al. (2020) [96] | Talking Head Generation with Rhythmic Head Motion | ECCV 2020 | GAN | VoxCeleb2 | FID, LMD, MOS | 3.8 |
| Transformer (n = 13) | |||||||
| S35 | Fan et al. (2022) [7] | FaceFormer: Speech-Driven 3D Facial Animation with Transformers | CVPR 2022 | Transformer | VOCASET, BIWI | LVE, EVE, MOS | 5.0 |
| S36 | Peng et al. (2024) [55] | SyncTalk: The Devil Is in the Synchronisation for Talking Head Synthesis | CVPR 2024 | Transformer | VoxCeleb2, HDTF | SyncNet, PSNR, SSIM, FID | 4.5 |
| S37 | Xing et al. (2023) [56] | CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion Prior | CVPR 2023 | Transformer | VOCASET, BIWI | LVE, EVE, FDD | 4.8 |
| S38 | Zhong et al. (2023) [97] | IP-LAP: Identity-Preserving Talking Face Generation with Landmark Priors | CVPR 2023 | Transformer | VoxCeleb2, HDTF | FID, SyncNet, CSIM, SSIM | 4.5 |
| S39 | Gan et al. (2023) [98] | Efficient Emotional Adaptation for Audio-Driven Talking-Head Generation | ICCV 2023 | Transformer | MEAD, VoxCeleb2 | FID, AUE, SyncNet | 4.2 |
| S40 | Fan et al. (2024) [99] | UniTalker: Scaling Up Audio-Driven 3D Facial Animation Through a Unified Model | ECCV 2024 | Transformer | VOCASET, BIWI, MEAD | LVE, MOS, FDD | 4.5 |
| S41 | Liu et al. (2024) [100] | CustomListener: Text-Guided Responsive Interaction for Listening Head Generation | CVPR 2024 | Transformer | VoxCeleb2, ViCo | FID, SyncNet, MOS | 4.0 |
| S42 | Peng et al. (2023) [101] | EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face Animation | ICCV 2023 | Transformer | MEAD, RAVDESS | FID, AUE, MOS | 4.2 |
| S43 | Wang et al. (2023) [102] | LipFormer: High-Fidelity and Generalizable Talking Face Generation with a Pre-Learned Facial Codebook | CVPR 2023 | Transformer | VoxCeleb2, HDTF, LRS3 | FID, SyncNet, SSIM, PSNR, LPIPS, MOS | 4.5 |
| S44 | Liu et al. (2023) [103] | MODA: Mapping-Once Audio-Driven Portrait Animation with Dual Attentions | ICCV 2023 | Transformer | VoxCeleb2, HDTF | FID, CPBD, LMD, SyncNet | 4.5 |
| S45 | Jiang et al. (2025) [104] | LOOPY: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency | ICLR 2025 | Transformer | VoxCeleb2, HDTF | FID, SyncNet, FVD | 3.8 |
| S46 | Liu et al. (2024) [105] | AniTalker: Animate Vivid and Diverse Talking Faces via Identity-Decoupled Motion | ACM MM 2024 | Transformer | VoxCeleb2, MEAD | FID, SyncNet, CSIM | 4.2 |
| S47 | Hogue et al. (2023) [106] | DiffTED: One-Shot Audio-Driven TED Talk Video Generation | CVPR 2024 | Transformer | TED-Talks | FID, MOS, SyncNet | 3.5 |
| NeRF-based (n = 11) | |||||||
| S48 | Guo et al. (2021) [9] | AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis | ICCV 2021 | NeRF | Custom (Obama) | PSNR, SSIM, LMD | 5.0 |
| S49 | Ye et al. (2023) [41] | GeneFace: Generalised and High-Fidelity Audio-Driven 3D Talking Face | ICLR 2023 | NeRF | VoxCeleb2, HDTF | PSNR, SSIM, LMD, MOS | 4.8 |
| S50 | Tang et al. (2025) [57] | RAD-NeRF: Real-Time Neural Radiance Fields for Dynamic Head Portraits | IJCV 2025 | NeRF | Custom | PSNR, SSIM, FPS | 4.0 |
| S51 | Li et al. (2023) [107] | ER-NeRF: Efficient Region-Aware NeRF for High-Fidelity Talking Portrait | ICCV 2023 | NeRF | Custom (multi-spk) | PSNR, SSIM, LPIPS, FPS | 4.5 |
| S52 | Ye et al. (2023) [108] | GeneFace++: Generalised Real-Time Audio-Driven 3D Talking Face | arXiv 2023† | NeRF | VoxCeleb2, HDTF | PSNR, SSIM, LMD, FPS | 4.8 |
| S53 | Ye et al. (2024) [109] | Real3D-Portrait: One-Shot Realistic 3D Talking Portrait Synthesis | arXiv 2024† | NeRF | HDTF, VoxCeleb2 | FID, PSNR, SSIM, CSIM | 4.5 |
| S54 | Shen et al. (2022) [110] | Learning Dynamic Facial Radiance Fields for Few-Shot Talking Head Synthesis | ECCV 2022 | NeRF | VoxCeleb2 | PSNR, SSIM, FID | 4.2 |
| S55 | Wang et al. (2024) [111] | Expression-Aware Neural Radiance Fields for High-Fidelity Talking Portrait | IVC 2024 | NeRF | Custom | PSNR, SSIM, LPIPS | 3.8 |
| S56 | Ma et al. (2025) [112] | Decoupled Two-Stage Talking Head Generation via Gaussian-Landmark NeRF | Comput. Vis. Media 2025 | NeRF | VoxCeleb2, HDTF | PSNR, SSIM, SyncNet | 4.0 |
| S57 | Deng et al. (2024) [113] | Portrait4D: Learning One-Shot 4D Head Avatar via Video-Based Supervision | CVPR 2024 | NeRF | VoxCeleb2, CelebV-HQ | FID, CSIM, LPIPS | 4.5 |
| S58 | Zhang et al. (2023) [114] | Metaportrait: Identity-preserving talking head generation with fast personalized adaptation | CVPR 2023 | NeRF | VoxCeleb2 | FID, PSNR, SSIM | 4.0 |
| Diffusion Model (n = 16) | |||||||
| S59 | Shen et al. (2023) [11] | DiffTalk: Crafting Diffusion Models for Generalised Audio-Driven Animation | CVPR 2023 | Diffusion | VoxCeleb2, HDTF | FID, SyncNet, SSIM | 4.8 |
| S60 | Ma et al. (2023) [12] | DreamTalk: Expressive Talking Head with Diffusion Probabilistic Models | arXiv 2023† | Diffusion | MEAD, HDTF | FID, LPIPS, MOS | 4.2 |
| S61 | Xu et al. (2024) [13] | VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time | NeurIPS 2024 | Diffusion | VoxCeleb2, HDTF | FID, SyncNet, CSIM, FPS | 5.0 |
| S62 | Xu et al. (2024) [58] | Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Animation | arXiv 2024† | Diffusion | VoxCeleb2, HDTF | FID, SyncNet, FVD, CSIM | 4.5 |
| S63 | Zhang et al. (2024) [19] | MuseTalk: Real-Time High Quality Lip Synchronisation with Latent Space Inpainting | arXiv 2024† | Diffusion | VoxCeleb2, HDTF | SyncNet, PSNR, FPS | 4.2 |
| S64 | Stypulkowski et al. (2024) [115] | Diffused Heads: Diffusion Models Beat GANs on Talking-Face Generation | WACV 2024 | Diffusion | VoxCeleb2 | FID, SyncNet, SSIM | 4.5 |
| S65 | Sun et al. (2024) [116] | DiffPoseTalk: Speech-Driven Stylistic 3D Facial Animation via Diffusion | ACM TOG 2024 | Diffusion | HDTF, MEAD | FID, SSIM, LMD, MOS | 4.5 |
| S66 | Bigioi et al. (2024) [117] | Speech Driven Video Editing via an Audio-Conditioned Diffusion Model | IVC 2024 | Diffusion | VoxCeleb2, LRS3 | SyncNet, FID, SSIM | 4.0 |
| S67 | Zhang et al. (2023) [40] | SadTalker: Learning Realistic 3D Motion for Stylised Audio-Driven Animation | CVPR 2023 | Diffusion | VoxCeleb2, HDTF | FID, SyncNet, LMD, SSIM | 4.8 |
| S68 | Tian et al. (2024) [118] | EMO: Emote Portrait Alive — Generating Expressive Portrait Videos | ECCV 2024 | Diffusion | VoxCeleb2, HDTF | FID, FVD, SyncNet, MOS | 4.2 |
| S69 | Cui et al. (2025) [119] | Hallo2: Long-Duration High-Resolution Audio-Driven Portrait Animation | ICLR 2025 | Diffusion | VoxCeleb2, HDTF | FID, FVD, SyncNet, CSIM | 4.0 |
| S70 | Wei et al. (2024) [120] | AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation | arXiv 2024† | Diffusion | VoxCeleb2 | FID, SyncNet, SSIM, LPIPS | 4.0 |
| S71 | Zheng et al. (2024) [121] | MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation | arXiv 2024† | Diffusion | VoxCeleb2, HDTF | FID, FVD, SyncNet, MOS | 4.2 |
| S72 | Ma et al. (2024) [122] | FollowYourEmoji: Fine-Controllable and Expressive Freestyle Portrait Animation | SIGGRAPH 2024 | Diffusion | VoxCeleb2, MEAD | FID, SyncNet, AUE | 4.0 |
| S73 | Wang et al. (2023) [65] | EAT: Emotional Talking Head Based on Memory-Sharing and Attention Mechanism | arXiv 2023† | Diffusion | MEAD, CREMA-D | FID, AUE, MOS | 4.2 |
| S74 | Cheng et al. (2025) [123] | Dynamic Frame Avatar with Non-Autoregressive Diffusion for Talking Head | ICLR 2025 | Diffusion | VoxCeleb2, HDTF | FVD, SyncNet, FID, SSIM | 4.5 |
| Gaussian Splatting (n = 8) | |||||||
| S75 | Chen et al. (2024) [50] | GSTalker: Real-Time Audio-Driven Talking Face via Deformable Gaussian Splatting | arXiv 2024† | GS | VoxCeleb2, HDTF | PSNR, SSIM, FPS, SyncNet | 4.2 |
| S76 | Yang et al. (2025) [51] | GaussianEmoTalker: Real-Time Emotional Talking Head Synthesis with Audio-Driven 3D GS | arXiv 2025† | GS | MEAD, VoxCeleb2 | FID, AUE, FPS, MOS | 4.0 |
| S77 | Xu et al. (2024) [52] | VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single Image | arXiv 2025† | GS | VoxCeleb2 | FID, SyncNet, FPS, CSIM | 4.5 |
| S78 | Xu et al. (2024) [124] | Gaussian Head Avatar: Ultra High-Fidelity Head Avatar via Dynamic Gaussians | CVPR 2024 | GS | Multi-View Custom | PSNR, SSIM, LPIPS, FPS | 4.8 |
| S79 | Qian et al. (2024) [125] | GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians | CVPR 2024 | GS | FLAME-fitted video | PSNR, SSIM, FPS | 4.5 |
| S80 | Chen et al. (2024) [126] | MonoGaussianAvatar: Monocular Gaussian Point-Based Head Avatar | SIGGRAPH 2024 | GS | Custom monocular | PSNR, SSIM, LPIPS | 4.5 |
| S81 | Deng et al. (2024) [127] | Portrait4D-v2: Pseudo Multi-View Data for 4D Gaussian Head Animation | ECCV 2024 | GS | VoxCeleb2, CelebV-HQ | FID, CSIM, FPS | 4.0 |
| S82 | Cho et al. (2024) [128] | Real-Time High-Fidelity Audio-Driven Talking Face via Gaussian Splatting | ACM MM 2024 | GS | VoxCeleb2, HDTF | PSNR, SSIM, LPIPS, SyncNet, FPS | 4.5 |
| Diffusion Transformer / DiT (n = 7) | |||||||
| S83 | Lin et al. (2025) [14] | OmniHuman-1: Scaling One-Stage Conditioned Human Animation via DiT | ICCV 2025 | DiT | VoxCeleb2, HDTF, Custom | FID, FVD, SyncNet, MOS | 5.0 |
| S84 | Wang et al. (2025) [20] | EmotiveTalk: Expressive Talking Head via Audio Decoupling and Emotional Video Diffusion | CVPR 2025 | DiT | MEAD, VoxCeleb2 | FID, AUE, SyncNet, MOS | 5.0 |
| S85 | Hong et al. (2025) [59] | Audio-Visual Controlled Video Diffusion with Masked SSMs for Talking Head | ICCV 2025 | DiT | VoxCeleb2, HDTF | FID, SyncNet, FVD, CSIM | 4.5 |
| S86 | Zheng et al. (2025) [60] | Teller: Real-Time Streaming Talking Portrait via Autoregressive Motion Generation | CVPR 2025 | DiT | VoxCeleb2 | SyncNet, FVD, FPS | 4.5 |
| S87 | Li et al. (2025) [129] | Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis | ACM MM 2025 | DiT | VoxCeleb2, HDTF | FID, SyncNet, SSIM, FVD | 4.5 |
| S88 | Tian et al. (2025) [71] | EMO2: End-to-End Audio-Driven Expressive Humanoid Talking Head Generation | arXiv 2025† | DiT | VoxCeleb2, HDTF | FID, FVD, SyncNet, MOS | 4.2 |
| S89 | Zhu et al. (2025) [23] | INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations | CVPR 2025 | DiT | VoxCeleb2, ViCo | FID, SyncNet, FVD, MOS | 4.2 |
| Hybrid (n = 3) | |||||||
| S90 | Li et al. (2024) [61] | Control-Talker: A Rapid-Customization Talking Head Generation Method for Multi-Condition Control and High-Texture Enhancement. | ACM MM 2024 | Hybrid | VoxCeleb2, MEAD | FID, SyncNet, MOS, CSIM | 4.0 |
| S91 | Ye et al. (2024) [62] | MimicTalk: Mimicking a Personalized and Expressive 3D Talking Face in Minutes | NeurIPS 2024 | Hybrid | VoxCeleb2, HDTF | PSNR, SSIM, LPIPS, SyncNet, AED, MOS | 4.5 |
| S92 | Li et al. (2024) [63] | TalkingGaussian: Structure-Persistent 3D Talking Head Synthesis via Gaussian Splatting | ECCV 2024 | Hybrid | VoxCeleb2, HDTF | PSNR, SSIM, LPIPS, SyncNet, FPS | 4.5 |
References
- Bregler, C.; Covell, M.; Slaney, M. Video rewrite: Driving visual speech with audio. Semin. Graph. Pap. Push. Bound. Vol. 2 2023, ed, 715–722. [Google Scholar] [CrossRef]
- Suwajanakorn, S.; Seitz, S. M.; Kemelmacher-Shlizerman, I. Synthesizing obama: learning lip sync from audio. ACM Trans. Graph. (ToG) 2017, vol. 36, 1–13. [Google Scholar]
- Goodfellow; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; et al. Generative adversarial networks. Commun. ACM 2020, vol. 63, 139–144. [Google Scholar] [CrossRef]
- Prajwal, K.; Mukhopadhyay, R.; Namboodiri, V. P.; Jawahar, C. A lip sync expert is all you need for speech to lip generation in the wild. In Proceedings of the 28th ACM international conference on multimedia, 2020; pp. 484–492. [Google Scholar]
- Son Chung, J.; Senior, A.; Vinyals, O.; Zisserman, A. Lip reading sentences in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2017; pp. 6447–6456. [Google Scholar]
- Zhang, Z.; Li, L.; Ding, Y.; Fan, C. Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021; pp. 3661–3670. [Google Scholar]
- Fan, Y.; Lin, Z.; Saito, J.; Wang, W.; Komura, T. Faceformer: Speech-driven 3d facial animation with transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022; pp. 18770–18780. [Google Scholar]
- Mildenhall, B.; Srinivasan, P. P.; Tancik, M.; Barron, J. T.; Ramamoorthi, R.; Ng, R. Nerf: Representing scenes as neural radiance fields for view synthesis. Commun. ACM 2021, vol. 65, 99–106. [Google Scholar]
- Guo, Y.; Chen, K.; Liang, S.; Liu, Y.-J.; Bao, H.; Zhang, J. Ad-nerf: Audio driven neural radiance fields for talking head synthesis. In Proceedings of the IEEE/CVF international conference on computer vision, 2021; pp. 5784–5794. [Google Scholar]
- Ho, J.; Jain, A.; Abbeel, P. Denoising diffusion probabilistic models. Adv. Neural Inf. Process. Syst. 2020, vol. 33, 6840–6851. [Google Scholar]
- Shen, S.; Zhao, W.; Meng, Z.; Li, W.; Zhu, Z.; Zhou, J.; et al. Difftalk: Crafting diffusion models for generalized audio-driven portraits animation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023; pp. 1982–1991. [Google Scholar]
- Ma, Y.; Zhang, S.; Wang, J.; Wang, X.; Zhang, Y.; Deng, Z. Dreamtalk: When emotional talking head generation meets diffusion probabilistic models. arXiv 2023, arXiv:2312.09767. [Google Scholar]
- Xu, S.; Chen, G.; Guo, Y.-X.; Yang, J.; Li, C.; Zang, Z.; et al. Vasa-1: Lifelike audio-driven talking faces generated in real time. Adv. Neural Inf. Process. Syst. 2024, vol. 37, 660–684. [Google Scholar] [CrossRef]
- Lin, G.; Jiang, J.; Yang, J.; Zheng, Z.; Liang, C.; Zhang, Y.; et al. Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025; pp. 13847–13858. [Google Scholar]
- Trist, E. L.; Bamforth, K. W. Some social and psychological consequences of the longwall method of coal-getting: An examination of the psychological situation and defences of a work group in relation to the social structure and technological content of the work system. Hum. Relat. 1951, vol. 4, 3–38. [Google Scholar]
- Bijker, W.; Hughes, T.; Pinch, T. The social construction of technological systems MIT Press; Cambridge MA, London, 1987. [Google Scholar]
- WHO Health Workforce Team. Global Health Workforce Shortage Projections. 2024. [Google Scholar] [CrossRef] [PubMed]
- Haider, S. A.; Prabha, S.; Gomez-Cabello, C. A.; Genovese, A.; Collaco, B.; Wood, N.; et al. Artificial Intelligence Physician Avatars for Patient Education: A Pilot Study. J. Clin. Med. vol. 14, 8595, 2025. [CrossRef] [PubMed]
- Zhang, Y.; Minhao, L.; Chen, Z.; Wu, B.; Zhan, C.; He, Y.; et al. Musetalk: Real-time high quality lip synchronization with latent space inpainting. 2024. [Google Scholar] [PubMed]
- Wang, H.; Weng, Y.; Li, Y.; Guo, Z.; Du, J.; Niu, S.; et al. Emotivetalk: Expressive talking head generation through audio information decoupling and emotional video diffusion. In Proceedings of the Computer Vision and Pattern Recognition Conference, 2025; pp. 26212–26221. [Google Scholar]
- Provoost, S.; Lau, H. M.; Ruwaard, J.; Riper, H. Embodied conversational agents in clinical psychology: a scoping review. J. Med. Internet Res. 2017, vol. 19, e151. [Google Scholar] [CrossRef] [PubMed]
- Chheang, V.; Sharmin, S.; Márquez-Hernández, R.; Patel, M.; Rajasekaran, D.; Caulfield, G.; et al. Towards anatomy education with generative AI-based virtual assistants in immersive virtual reality environments. 2024 IEEE Int. Conf. Artif. Intell. Ext. Virtual Real. (AIxVR) 2024, 21–30. [Google Scholar] [CrossRef]
- Zhu, Y.; Zhang, L.; Rong, Z.; Hu, T.; Liang, S.; Ge, Z. INFP: Audio-driven interactive head generation in dyadic conversations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025; pp. 10667–10677. [Google Scholar]
- Salahudeen, R.; Siu, W.-C.; Chan, H. A. Photo-realistic talking face generation under latent space manipulation. IEEE Transactions on Consumer Electronics, 2024. [Google Scholar]
- Salahudeen, R.; Siu, W.-C.; Chan, H. A. Activate your face in virtual meeting platform. 2023 International Conference on Consumer Electronics-Taiwan (ICCE-Taiwan), 2023; pp. 791–792. [Google Scholar]
- Deonarine, K. Deepfakes and human rights: Why the EU AI Act is becoming the global standard for ethical AI regulation. 2025. [Google Scholar] [PubMed]
- United States Congress. TAKE IT DOWN Act," ed. 2025. [Google Scholar] [CrossRef]
- European Commission. Code of Practice on Transparency of AI-Generated Content. EU AI OfficeJune, 10 2026. [Google Scholar]
- Rana, M. S.; Nobi, M. N.; Murali, B.; Sung, A. H. Deepfake detection: A systematic literature review. IEEE Access 2022, vol. 10, 25494–25513. [Google Scholar] [CrossRef]
- Zhen, R.; Song, W.; He, Q.; Cao, J.; Shi, L.; Luo, J. Human-computer interaction system: A survey of talking-head generation. Electronics 2023, vol. 12, 218. [Google Scholar]
- Toshpulatov, M.; Lee, W.; Lee, S. Talking human face generation: A survey. Expert Syst. With Appl. 2023, vol. 219, 119678. [Google Scholar]
- Bigioi, D.; Corcoran, P. Multilingual video dubbing—a technology review and current challenges. Front. Signal Process. 2023, vol. 3, 1230755. [Google Scholar] [CrossRef]
- Heidari; Jafari Navimipour, N.; Dag, H.; Unal, M. Deepfake detection using deep learning methods: A systematic and comprehensive review. Wiley Interdiscip. Rev. Data Min. Knowl. Discov. 2024, vol. 14, e1520. [Google Scholar]
- Rakesh, V. K.; Mazumdar, S.; Maity, R. P.; Pal, S.; Das, A.; Samanta, T. Advancements in talking head generation: a comprehensive review of techniques, metrics, and challenges: Advancements in talking head generation. Vis. Comput. 2025, vol. 42. [Google Scholar] [CrossRef]
- Floridi, L.; Cowls, J.; Beltrametti, M.; Chatila, R.; Chazerand, P.; Dignum, V.; et al. AI4People—An ethical framework for a good AI society: Opportunities, risks, principles, and recommendations. Minds Mach. 2018, vol. 28, 689–707. [Google Scholar] [CrossRef] [PubMed]
- Kitchenham, B.; Charters, S. Guidelines for performing systematic literature reviews in software engineering. 2007. [Google Scholar] [CrossRef]
- Romero Moreno, F. Generative AI and deepfakes: a human rights approach to tackling harmful content. Int. Rev. Law. Comput. Technol. 2024, vol. 38, 297–326. [Google Scholar] [CrossRef]
- Nisar, H.; Masood, S.; Malik, Z.; Abid, A. Talking Head Generation Through Generative Models and Cross-Modal Synthesis Techniques. J. Imaging 2026, vol. 12, 119. [Google Scholar] [CrossRef] [PubMed]
- Chung, J. S.; Zisserman, A. Out of time: automated lip sync in the wild. Asian conference on computer vision, 2016; pp. 251–263. [Google Scholar]
- Zhang, W.; Cun, X.; Wang, X.; Zhang, Y.; Shen, X.; Guo, Y.; et al. Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023; pp. 8652–8661. [Google Scholar]
- Ye, Z.; Jiang, Z.; Ren, Y.; Liu, J.; He, J.; Zhao, Z. Geneface: Generalized and high-fidelity audio-driven 3d talking face synthesis. arXiv 2023, arXiv:2301.13430. [Google Scholar]
- Bickmore, T.; Cassell, J. Relational agents: a model and implementation of building user trust. In Proceedings of the SIGCHI conference on Human factors in computing systems, 2001; pp. 396–403. [Google Scholar]
- Page, M. J.; McKenzie, J. E.; Bossuyt, P. M.; Boutron, I.; Hoffmann, T. C.; Mulrow, C. D.; et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ 2021, vol. 372. [Google Scholar] [CrossRef] [PubMed]
- Popay, J.; Roberts, H.; Sowden, A.; Petticrew, M.; Arai, L.; Rodgers, M.; et al. Guidance on the conduct of narrative synthesis in systematic reviews. A Product. From ESRC Methods Programme Version 2006, vol. 1, b92. [Google Scholar]
- Fan, B.; Xie, L.; Yang, S.; Wang, L.; Soong, F. K. A deep bidirectional LSTM approach for video-realistic talking head. Multimed. Tools Appl. 2016, vol. 75, 5287–5309. [Google Scholar] [CrossRef]
- Zhou, H.; Liu, Y.; Liu, Z.; Luo, P.; Wang, X. Talking face generation by adversarially disentangled audio-visual representation. In Proceedings of the AAAI conference on artificial intelligence, 2019; pp. 9299–9306. [Google Scholar]
- Zhou, Y.; Han, X.; Shechtman, E.; Echevarria, J.; Kalogerakis, E.; Li, D. Makelttalk: speaker-aware talking-head animation. ACM Trans. Graph. (TOG) 2020, vol. 39, 1–15. [Google Scholar]
- Zhou, H.; Sun, Y.; Wu, W.; Loy, C. C.; Wang, X.; Liu, Z. Pose-controllable talking face generation by implicitly modularized audio-visual representation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021; pp. 4176–4186. [Google Scholar]
- Peebles, W.; Xie, S. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF international conference on computer vision, 2023; pp. 4195–4205. [Google Scholar]
- Chen, B.; Hu, S.; Chen, Q.; Du, C.; Yi, R.; Qian, Y.; et al. Gstalker: Real-time audio-driven talking face generation via deformable gaussian splatting. arXiv 2024, arXiv:2404.19040. [Google Scholar]
- Yang, H.; Zhang, Z.; Dong, Y.; Qian, J.; Yang, J. GaussianEmoTalker: Real-Time Emotional Talking Head Synthesis with Audio-Driven and Blendshape-Based 3D Gaussian Splatting. arXiv 2026, arXiv:2607.00959. [Google Scholar]
- Xu, S.; Chen, G.; Yang, J.; Zhang, Y.; Deng, Y.; Lin, S.; et al. Vasa-3d: Lifelike audio-driven gaussian head avatars from a single image. Adv. Neural Inf. Process. Syst. 2026, vol. 38, 150764–150786. [Google Scholar]
- Song, Y.; Zhu, J.; Li, D.; Wang, A.; Qi, H. Talking face generation by conditional recurrent adversarial network. In Proceedings of the 28th International Joint Conference on Artificial Intelligence, 2019; pp. 919–925. [Google Scholar]
- Zhou, Y.; Xu, Z.; Landreth, C.; Kalogerakis, E.; Maji, S.; Singh, K. Visemenet: Audio-driven animator-centric speech animation. ACM Trans. Graph. (ToG) 2018, vol. 37, 1–10. [Google Scholar]
- Peng, Z.; Hu, W.; Shi, Y.; Zhu, X.; Zhang, X.; Zhao, H.; et al. Synctalk: The devil is in the synchronization for talking head synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 666–676. [Google Scholar]
- Xing, J.; Xia, M.; Zhang, Y.; Cun, X.; Wang, J.; Wong, T.-T. Codetalker: Speech-driven 3d facial animation with discrete motion prior. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023; pp. 12780–12790. [Google Scholar]
- Tang, J.; Wang, K.; Zhou, H.; Chen, X.; He, D.; Hu, T.; et al. Real-time neural radiance talking portrait synthesis via audio-spatial decomposition. Int. J. Comput. Vis. 2025, vol. 133, 6362–6373. [Google Scholar] [CrossRef]
- Xu, M.; Li, H.; Su, Q.; Shang, H.; Zhang, L.; Liu, C.; et al. Hallo: Hierarchical audio-driven visual synthesis for portrait image animation. arXiv 2024, arXiv:2406.08801. [Google Scholar]
- Hong, F.-T.; Xu, Z.; Zhou, Z.; Zhou, J.; Li, X.; Lin, Q.; et al. Audio-visual controlled video diffusion with masked selective state spaces modeling for natural talking head generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025; pp. 12549–12558. [Google Scholar]
- Zhen, D.; Yin, S.; Qin, S.; Yi, H.; Zhang, Z.; Liu, S.; et al. Teller: Real-time streaming audio-driven portrait animation with autoregressive motion generation. In Proceedings of the Computer Vision and Pattern Recognition Conference, 2025; pp. 21075–21085. [Google Scholar]
- Li, Y.; Yu, L.; Wang, L.; Xie, H. Control-Talker: A Rapid-Customization Talking Head Generation Method for Multi-Condition Control and High-Texture Enhancement. In Proceedings of the 32nd ACM International Conference on Multimedia, 2024; pp. 3519–3527. [Google Scholar]
- Ye, Z.; Zhong, T.; Ren, Y.; Jiang, Z.; Huang, J.; Huang, R.; et al. Mimictalk: Mimicking a personalized and expressive 3d talking face in minutes. Adv. Neural Inf. Process. Syst. 2024, vol. 37, 1829–1853. [Google Scholar] [CrossRef]
- Li, J.; Zhang, J.; Bai, X.; Zheng, J.; Ning, X.; Zhou, J.; et al. Talkinggaussian: Structure-persistent 3d talking head synthesis via gaussian splatting. European Conference on Computer Vision, 2024; pp. 127–145. [Google Scholar]
- Buolamwini, J.; Gebru, T. Gender shades: Intersectional accuracy disparities in commercial gender classification. Conference on fairness, accountability and transparency, 2018; pp. 77–91. [Google Scholar]
- Wang, J.; Zhao, Y.; Liu, L.; Xu, T.; Li, Q.; Li, S. Emotional talking head generation based on memory-sharing and attention-augmented networks. arXiv 2023, arXiv:2306.03594. [Google Scholar]
- Wang, Z.; Bovik, A. C.; Sheikh, H. R.; Simoncelli, E. P. Image quality assessment: from error visibility to structural similarity. IEEE Trans. Image Process. 2004, vol. 13, 600–612. [Google Scholar] [CrossRef] [PubMed]
- DeTore, N. R.; Balogun-Mwangi, O.; Eberlin, E. S.; Dokholyan, K. N.; Rizzo, A.; Holt, D. J. An artificial intelligence-based virtual human avatar application to assess the mental health of health care professionals: A validation study. J. Med. Ext. Real. 2024, vol. 1, jmxr. 2024.0016. [Google Scholar]
- Venkatesh, S.; Ramachandra, R.; Raja, K.; Busch, C. Face morphing attack generation and detection: A comprehensive survey. IEEE Trans. Technol. Soc. 2021, vol. 2, 128–145. [Google Scholar] [CrossRef]
- iProov. The iProov Threat Intelligence Report 2024: The Impact of Generative AI on Remote Identity Verification. 2024. [Google Scholar]
- Walker, Jones. Deepfakes-as-a-Service meets state laws: Governing synthetic media in a fragmented legal landscape. 2026. [Google Scholar]
- Tian, L.; Hu, S.; Wang, Q.; Zhang, B.; Bo, L. Emo2: End-effector guided audio-driven avatar video generation. arXiv 2025, arXiv:2501.10687. [Google Scholar]
- Beauchamp, T. L.; Childress, J. F. Principles of biomedical ethics; Edicoes Loyola, 1994. [Google Scholar]
- Grother, P.; Salamon, W.; Chandramouli, R. NIST Special Publication 800-76-2 Biometric Specifications for Personal Identity Verification. National Institute of Standards and Technology, Tech. Rep. 800-76-2, 2013.
- Vougioukas, K.; Petridis, S.; Pantic, M. Realistic speech-driven facial animation with gans. Int. J. Comput. Vis. 2020, vol. 128, 1398–1413. [Google Scholar] [CrossRef]
- Chen, L.; Maddox, R. K.; Duan, Z.; Xu, C. Hierarchical cross-modal talking face generation with dynamic pixel-wise loss. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019; pp. 7832–7841. [Google Scholar]
- Siarohin; Lathuilière, S.; Tulyakov, S.; Ricci, E.; Sebe, N. First order motion model for image animation. Adv. Neural Inf. Process. Syst. 2019, vol. 32. [Google Scholar]
- Jamaludin; Chung, J. S.; Zisserman, A. You said that?: Synthesising talking faces from audio. Int. J. Comput. Vis. 2019, vol. 127, 1767–1779. [Google Scholar] [CrossRef]
- Wiles; Koepke, A.; Zisserman, A. X2face: A network for controlling face generation using images, audio, and pose codes. In Proceedings of the European conference on computer vision (ECCV), 2018; pp. 670–686. [Google Scholar]
- Wang, T.-C.; Mallya, A.; Liu, M.-Y. One-shot free-view neural talking-head synthesis for video conferencing. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021; pp. 10039–10049. [Google Scholar]
- Cheng, K.; Cun, X.; Zhang, Y.; Xia, M.; Yin, F.; Zhu, M.; et al. Videoretalking: Audio-based lip synchronization for talking head video editing in the wild. SIGGRAPH Asia 2022 Conference Papers, 2022; pp. 1–9. [Google Scholar]
- Das, D.; Biswas, S.; Sinha, S.; Bhowmick, B. Speech-driven facial animation using cascaded gans for learning of motion and texture. European conference on computer vision, 2020; pp. 408–424. [Google Scholar]
- Ji, X.; Zhou, H.; Wang, K.; Wu, W.; Loy, C. C.; Cao, X.; et al. Audio-driven emotional video portraits. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021; pp. 14080–14089. [Google Scholar]
- Wang, S.; Li, L.; Ding, Y.; Fan, C.; Yu, X. Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion. In Proceedings of the Thirtieth International Joint Conference On Artificial Intelligence, Ijcai 2021; 2021, pp. 1098–1105.
- Hong, F.-T.; Zhang, L.; Shen, L.; Xu, D. Depth-aware generative adversarial network for talking head video generation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022; pp. 3397–3406. [Google Scholar]
- Yin, F.; Zhang, Y.; Cun, X.; Cao, M.; Fan, Y.; Wang, X.; et al. Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan. European conference on computer vision, 2022; pp. 85–101. [Google Scholar]
- Liang; Pan, Y.; Guo, Z.; Zhou, H.; Hong, Z.; Han, X.; et al. Expressive talking head generation with granular audio-visual control. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022; pp. 3387–3396. [Google Scholar]
- Song, H.-K.; Woo, S. H.; Lee, J.; Yang, S.; Cho, H.; Lee, Y.; et al. Talking face generation with multilingual tts. In Proceedings of the ieee/cvf conference on computer vision and pattern recognition, 2022; pp. 21425–21430. [Google Scholar]
- Lu, Y.; Chai, J.; Cao, X. Live speech portraits: real-time photorealistic talking-head animation. ACM Trans. Graph. (ToG) 2021, vol. 40, 1–17. [Google Scholar]
- Zhu, H.; Huang, H.; Li, Y.; Zheng, A.; He, R. Arbitrary talking face generation via attentional audio-visual coherence learning. arXiv 2018, arXiv:1812.06589. [Google Scholar]
- Wen, X.; Wang, M.; Richardt, C.; Chen, Z.-Y.; Hu, S.-M. Photorealistic audio-driven video portraits. IEEE Trans. Vis. Comput. Graph. 2020, vol. 26, 3457–3466. [Google Scholar] [CrossRef] [PubMed]
- Hwang, G.; Hong, S.; Lee, S.; Park, S.; Chae, G. Discohead: Audio-and-video-driven talking head generation by disentangled control of head pose and facial expressions. ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023; pp. 1–5. [Google Scholar]
- Ji, X.; Zhou, H.; Wang, K.; Wu, Q.; Wu, W.; Xu, F.; et al. Eamm: One-shot emotional talking face via audio-based emotion-aware motion model. ACM SIGGRAPH 2022 conference proceedings, 2022; pp. 1–10. [Google Scholar]
- Yi, R.; Ye, Z.; Sun, Z.; Zhang, J.; Zhang, G.; Wan, P.; et al. Predicting personalized head movement from short video and speech signal. IEEE Trans. Multimed. 2022, vol. 25, 6315–6328. [Google Scholar] [CrossRef]
- Wang, S.; Li, L.; Ding, Y.; Yu, X. One-shot talking face generation from single-speaker audio-visual correlation learning. In Proceedings of the AAAI Conference on Artificial Intelligence, 2022; pp. 2531–2539. [Google Scholar]
- Peng, Z.; Luo, Y.; Shi, Y.; Xu, H.; Zhu, X.; Liu, H.; et al. Selftalk: A self-supervised commutative training diagram to comprehend 3d talking faces. In Proceedings of the 31st ACM International Conference on Multimedia, 2023; pp. 5292–5301. [Google Scholar]
- Chen, L.; Cui, G.; Liu, C.; Li, Z.; Kou, Z.; Xu, Y.; et al. Talking-head generation with rhythmic head motion. European conference on computer vision, 2020; pp. 35–51. [Google Scholar]
- Zhong, W.; Fang, C.; Cai, Y.; Wei, P.; Zhao, G.; Lin, L.; et al. Identity-preserving talking face generation with landmark and appearance priors. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023; pp. 9729–9738. [Google Scholar]
- Gan, Y.; Yang, Z.; Yue, X.; Sun, L.; Yang, Y. Efficient emotional adaptation for audio-driven talking-head generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023; pp. 22634–22645. [Google Scholar]
- Fan, X.; Li, J.; Lin, Z.; Xiao, W.; Yang, L. Unitalker: Scaling up audio-driven 3d facial animation through a unified model. European Conference on Computer Vision, 2024; pp. 204–221. [Google Scholar]
- Liu, X.; Guo, Y.; Zhen, C.; Li, T.; Ao, Y.; Yan, P. Customlistener: Text-guided responsive interaction for user-friendly listening head generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 2415–2424. [Google Scholar]
- Peng, Z.; Wu, H.; Song, Z.; Xu, H.; Zhu, X.; He, J.; et al. Emotalk: Speech-driven emotional disentanglement for 3d face animation. In Proceedings of the IEEE/CVF international conference on computer vision, 2023; pp. 20687–20697. [Google Scholar]
- Wang, J.; Zhao, K.; Zhang, S.; Zhang, Y.; Shen, Y.; Zhao, D.; et al. Lipformer: High-fidelity and generalizable talking face generation with a pre-learned facial codebook. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023; pp. 13844–13853. [Google Scholar]
- Liu, Y.; Lin, L.; Yu, F.; Zhou, C.; Li, Y. Moda: Mapping-once audio-driven portrait animation with dual attentions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023; pp. 23020–23029. [Google Scholar]
- Jiang, J.; Liang, C.; Yang, J.; Lin, G.; Zhong, T.; Zheng, Y. Loopy: Taming audio-driven portrait avatar with long-term motion dependency. International Conference on Learning Representations, 2025; pp. 15245–15263. [Google Scholar]
- Liu, T.; Chen, F.; Fan, S.; Du, C.; Chen, Q.; Chen, X.; et al. Anitalker: Animate vivid and diverse talking faces through identity-decoupled facial motion encoding. In Proceedings of the 32nd ACM International Conference on Multimedia, 2024; pp. 6696–6705. [Google Scholar]
- Hogue, S.; Zhang, C.; Daruger, H.; Tian, Y.; Guo, X. Diffted: One-shot audio-driven ted talk video generation with diffusion-based co-speech gestures. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 1922–1931. [Google Scholar]
- Li, J.; Zhang, J.; Bai, X.; Zhou, J.; Gu, L. Efficient region-aware neural radiance fields for high-fidelity talking portrait synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023; pp. 7568–7578. [Google Scholar]
- Ye, Z.; He, J.; Jiang, Z.; Huang, R.; Huang, J.; Liu, J.; et al. Geneface++: Generalized and stable real-time audio-driven 3d talking face generation. arXiv 2023, arXiv:2305.00787. [Google Scholar]
- Ye, Z.; Zhong, T.; Ren, Y.; Yang, J.; Li, W.; Huang, J.; et al. Real3d-portrait: One-shot realistic 3d talking portrait synthesis. arXiv 2024, arXiv:2401.08503. [Google Scholar]
- Shen, S.; Li, W.; Zhu, Z.; Duan, Y.; Zhou, J.; Lu, J. Learning dynamic facial radiance fields for few-shot talking head synthesis. European conference on computer vision, 2022; pp. 666–682. [Google Scholar]
- Wang, X.; Ruan, T.; Xu, J.; Guo, X.; Li, J.; Yan, F.; et al. Expression-aware neural radiance fields for high-fidelity talking portrait synthesis. Image Vis. Comput. 2024, vol. 147, 105075. [Google Scholar] [CrossRef]
- Ma; Cao, Y.; Zhang, L. Decoupled two-stage talking head generation via Gaussian-landmark-based neural radiance fields. Computational Visual Media, 2025. [Google Scholar]
- Deng, Y.; Wang, D.; Ren, X.; Chen, X.; Wang, B. Portrait4d: Learning one-shot 4d head avatar synthesis using synthetic data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 7119–7130. [Google Scholar]
- Zhang, B.; Qi, C.; Zhang, P.; Zhang, B.; Wu, H.; Chen, D.; et al. Metaportrait: Identity-preserving talking head generation with fast personalized adaptation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023; pp. 22096–22105. [Google Scholar]
- Stypułkowski, M.; Vougioukas, K.; He, S.; Zięba, M.; Petridis, S.; Pantic, M. Diffused heads: Diffusion models beat gans on talking-face generation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024; pp. 5091–5100. [Google Scholar]
- Sun, Z.; Lv, T.; Ye, S.; Lin, M.; Sheng, J.; Wen, Y.-H.; et al. Diffposetalk: Speech-driven stylistic 3d facial animation and head pose generation via diffusion models. ACM Trans. Graph. (ToG) 2024, vol. 43, 1–9. [Google Scholar] [CrossRef]
- Bigioi; Basak, S.; Stypułkowski, M.; Zieba, M.; Jordan, H.; McDonnell, R.; et al. Speech driven video editing via an audio-conditioned diffusion model. Image Vis. Comput. 2024, vol. 142, 104911. [Google Scholar] [CrossRef]
- Tian, L.; Wang, Q.; Zhang, B.; Bo, L. Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions. European Conference on Computer Vision, 2024; pp. 244–260. [Google Scholar]
- Cui, J.; Li, H.; Yao, Y.; Zhu, H.; Shang, H.; Cheng, K.; et al. Hallo2: Long-duration and high-resolution audio-driven portrait image animation. International Conference on Learning Representations, 2025; pp. 91659–91671. [Google Scholar]
- Wei, H.; Yang, Z.; Wang, Z. Aniportrait: Audio-driven synthesis of photorealistic portrait animation. arXiv 2024, arXiv:2403.17694. [Google Scholar]
- Zheng, L.; Zhang, Y.; Guo, H.; Pan, J.; Tan, Z.; Lu, J.; et al. Memo: Memory-guided diffusion for expressive talking video generation. arXiv 2024, arXiv:2412.04448. [Google Scholar]
- Ma, Y.; Liu, H.; Wang, H.; Pan, H.; He, Y.; Yuan, J.; et al. Follow-your-emoji: Fine-controllable and expressive freestyle portrait animation. SIGGRAPH Asia 2024 Conference Papers, 2024; pp. 1–12. [Google Scholar]
- Cheng, H.; Lin, L.; Liu, C.; Xia, P.; Hu, P.; Ma, J.; et al. Dawn: Dynamic frame avatar with non-autoregressive diffusion framework for talking head video generation. International Conference on Learning Representations, 2025; pp. 99250–99268. [Google Scholar]
- Xu, Y.; Chen, B.; Li, Z.; Zhang, H.; Wang, L.; Zheng, Z.; et al. Gaussian head avatar: Ultra high-fidelity head avatar via dynamic gaussians. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024; pp. 1931–1941. [Google Scholar]
- Qian, S.; Kirschstein, T.; Schoneveld, L.; Davoli, D.; Giebenhain, S.; Nießner, M. Gaussianavatars: Photorealistic head avatars with rigged 3d gaussians. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 20299–20309. [Google Scholar]
- Chen, Y.; Wang, L.; Li, Q.; Xiao, H.; Zhang, S.; Yao, H.; et al. Monogaussianavatar: Monocular gaussian point-based head avatar. ACM SIGGRAPH 2024 conference papers, 2024; pp. 1–9. [Google Scholar]
- Deng, Y.; Wang, D.; Wang, B. Portrait4d-v2: Pseudo multi-view data creates better 4d head synthesizer. European Conference on Computer Vision, 2024; pp. 316–333. [Google Scholar]
- Cho, K.; Lee, J.; Yoon, H.; Hong, Y.; Ko, J.; Ahn, S.; et al. Gaussiantalker: Real-time talking head synthesis with 3d gaussian splatting. In Proceedings of the 32nd ACM International Conference on Multimedia, 2024; pp. 10985–10994. [Google Scholar]
- Li, T.; Zheng, R.; Yang, M.; Chen, J.; Yang, M. Ditto: Motion-space diffusion for controllable realtime talking head synthesis. In Proceedings of the 33rd ACM International Conference on Multimedia, 2025; pp. 9704–9713. [Google Scholar]
Figure 1.
Talking face generation from a single image and an audio clip.

Figure 2.
Coverage of prior reviews across seven assessment dimensions (studies as cited in Section 2.1, Section 2.2 and Section 2.3; filled = full coverage, half-filled = partial or restricted, open = not addressed).
Figure 2.
Coverage of prior reviews across seven assessment dimensions (studies as cited in Section 2.1, Section 2.2 and Section 2.3; filled = full coverage, half-filled = partial or restricted, open = not addressed).

Figure 3.
The PRISMA flowchart of our review process.

Figure 4.
Talking face generation pipeline, illustrating audio-driven and audio-plus-video-driven classifications. Inputs include an audio signal and face image, with optional text and driving video, processed through feature extraction, mapping, and transformation stages to produce video frames.
Figure 4.
Talking face generation pipeline, illustrating audio-driven and audio-plus-video-driven classifications. Inputs include an audio signal and face image, with optional text and driving video, processed through feature extraction, mapping, and transformation stages to produce video frames.

Figure 5.
Architecture Family Distribution by Year.

Figure 6.
Training and evaluation datasets across included studies (N = 92): (a) usage frequency (counts sum to more than 92 because studies use several corpora); (b) characteristics of all corpora appearing in four or more studies, plus the pooled custom-corpus category.
Figure 6.
Training and evaluation datasets across included studies (N = 92): (a) usage frequency (counts sum to more than 92 because studies use several corpora); (b) characteristics of all corpora appearing in four or more studies, plus the pooled custom-corpus category.

Figure 7.
Evaluation metrics across included studies (N = 92): (a) reporting frequency; (b) deployment-fitness assessment. SyncNet & variants pools SyncNet confidence with Sync/Lip-Sync Error and LSE-C/LSE-D; Other = CPBD, LVE, AED, EVE, FDD, AKD.
Figure 7.
Evaluation metrics across included studies (N = 92): (a) reporting frequency; (b) deployment-fitness assessment. SyncNet & variants pools SyncNet confidence with Sync/Lip-Sync Error and LSE-C/LSE-D; Other = CPBD, LVE, AED, EVE, FDD, AKD.

Figure 8.
Comparative Normalized Performance Across Architecture Eras.

Figure 9.
Ethical Safeguard Presence Across Included Studies.

Figure 10.
Coded Deployment Domain Distribution Across Included Studies.

Table 1.
Records retrieved per repository per iteration.
| Repository | Iter. 1 (Mar 2023) | Iter. 2 (Jun 2023) | Iter. 3 (Jan 2024) | Iter. 4 (Mar 2026) | Total |
| IEEE Xplore | 1,614 | 1,047 | 1,183 | 1,101 | 4,945 |
| ScienceDirect | 893 | 671 | 814 | 782 | 3,160 |
| ACM Digital Library | 712 | 498 | 601 | 553 | 2,364 |
| SpringerLink | 836 | 622 | 731 | 648 | 2,837 |
| Google Scholar / WoS | 1,243 | 891 | 872 | 847 | 3,853 |
| arXiv | 523 | 215 | 159 | 287 | 1,184 |
| Total | 5,821 | 3,944 | 4,360 | 4,218 | 18,343 |
Table 2.
Inclusion (I) and exclusion (E) criteria.
| ID | Criterion |
| I1 | Full text publicly accessible (open access, subscription, or author preprint) |
| I2 | Written in English |
| I3 | Published or archived 1 January 2016 - 31 December 2025 |
| I4 | Employs at least one deep learning technique as a core component |
| I5 | Addresses TFG synthesis, evaluation of TFG output, detection of TFG content, or ethical/governance/socio-technical dimensions of TFG deployment |
| I6 | Primary research contribution (original study, dataset, evaluation, or position paper with novel empirical analysis) |
| E1 | Full text not retrievable after two independent attempts |
| E2 | Duplicate across databases or iterations (DOI match, then title-author string match) |
| E3 | Deep learning application unrelated to human facial video |
| E4 | Outside the temporal scope |
| E5 | Not in English |
| E6 | A survey or review of TFG/deepfakes (treated as related work, not primary evidence) |
Table 3.
Architecture taxonomy derived from the synthesis corpus (N = 92).
| Architecture family | Phase (years) | n | Representative models | Key capability | Principal limitation |
| LSTM / RNN | I (2016-2018) | 5 | Fan et al. [45]; Suwajanakorn et al. [2] | Sequential audio-to-lip mapping | Speaker dependence; slow convergence |
| CNN-based | I-II (2017-2020) | 7 | Song et al. [53]; Zhou et al. [54] | Spatial features; real-time feasible | No generative capability |
| GAN-based | II (2018-2022) | 22 | Wav2Lip [4]; PC-AVS [48]; MakeItTalk [47] | High lip-sync fidelity; speaker-agnostic | Mode collapse; weak 3D consistency |
| Transformer | III (2021-2023) | 13 | FaceFormer [7]; SyncTalk [55]; CodeTalker [56] | Long-range audio-visual alignment | High compute; limited diversity vs. diffusion |
| NeRF-based | III-IV (2021-2024) | 11 | AD-NeRF [9]; 3 RAD-NeRF [57]; GeneFace [41] | D-consistent H multi-view synthesis | ours of per-identity training; far below real time |
| Diffusion Model | IV (2023-2024) | 16 | DiffTalk [11]; DreamTalk [12]; VASA-1 [13]; Hallo [58] | Probabilistic; diverse; stable training | High inference latency (except flow matching) |
| Gaussian Splatting | V (2024-2025) | 8 | GSTalker [50]; GaussianEmoTalker [51]; VASA-3D [52] | Real-time photorealistic rendering | Identity-specific scene optimization |
| Diffusion Transformer | V (2024-2025) | 7 | OmniHuman-1 [14]; ACTalker [59]; EmotiveTalk [20]; Teller [60] | Scales with compute; full-body; multi-modal | Very high training compute |
| Hybrid | V (2024-2025) | 3 | Control-Talker [61]; MimicTalk [62]; TalkingGaussian [63] | Combines paradigm strengths | Complexity; opaque failure modes |
Table 4.
Benchmark performance of representative state-of-the-art models from the corpus.
| Model | Family | Year | FID (lower better) | SyncNet conf. | Real-time? | Key advance |
| Wav2Lip [4] | GAN | 2020 | N/A | 8.31 | Yes | First in-the-wild lip sync on arbitrary video |
| PC-AVS [48] | GAN | 2021 | 41.2 | 7.86 | Partial | Pose-controllable synthesis |
| AD-NeRF [9] | NeRF | 2021 | N/A | 7.42 | No (6 fps) p | 3D-consistent portraits |
| SadTalker [40] | 3DMM+Diffusion | 2023 | 47.7 | 7.98 | Partial | 3D-aware motion without per-identity tuning |
| DiffTalk [11] | Diffusion | 2023 | 35.4 | 7.40 | No | First generalized audio-driven diffusion TFG |
| Hallo [58] | Diffusion | 2024 | 27.2 | 8.02 | Partial | Hierarchical audio-visual attention |
| MuseTalk [19] | Diffusion | 2024 | 31.4 | 8.49 | Yes | First real-time diffusion lip sync |
| SyncTalk [55] | Transformer+NeRF | 2024 | N/A | 8.64 | Partial | Synchronized temporal attention |
| VASA-1 [13] | Diffusion (flow) | 2024 | 22.1 | 8.59 | Yes (45 fps) | Disentangled facial dynamics in real time |
| GSTalker [50] | Gaussian Splatting | 2024 | N/A | 8.04 | Yes | Deformable Gaussian Splatting for TFG |
| OmniHuman-1 [14] | DiT | 2025 | 18.4 | 8.53 | Partial | Full-body DiT; multi-modal conditioning |
| EmotiveTalk [20] | DiT | 2025 | 21.3 | 8.31 | Partial | Content/emotion audio decoupling |
| Teller [60] | DiT | 2025 | N/A | 8.72 | Yes | Autoregressive streaming; lowest latency |
Table 5.
Inventory of ethical safeguard content across the synthesis corpus (N = 92).
| Ethical dimension | Studies (n) | % of corpus | Nature of coverage |
| Watermarking / provenance | 1 | 1.1% | One post-hoc frequency-domain watermark; not C2PA-compliant; no compression robustness |
| Deepfake detection integrated | 3 | 3.3% | Parallel detector trained or evaluated alongside the generator, typically as an ablation |
| Ethical risk discussion | 5 | 5.4% | One to three sentences acknowledging misuse potential; no mitigation proposed |
| Consent protocol | 0 | 0.0% | Whose consent is required to synthesize a face is not addressed anywhere in the primary literature |
| Regulatory framework cited | 0 | 0.0% | The regulatory environment is invisible to the reviewed technical literature |
| Any ethical content | 8 | 8.7% | 91.3% of included studies contain none |
Table 6.
Cross-domain socio-technical analysis of TFG deployment (RQ6 summary)
| Domain | Studies (n) | Primary benefit | Primary risk | Most urgent governance need | Regulatory instruments |
| Healthcare & telehealth | 7 | Scalable, empathetic clinical communication | Patient deception; demographic inequity; liability ambiguity | Mandatory disclosure; demographic audit; clinical oversight | EU AI Act Art. 50, 10, 14 |
| Education & training | 9 | Personalized, emotionally adaptive instruction | Student biometric capture; epistemic authority effects | Biometric data governance; disclosure; educator oversight | GDPR Art. 9; FERPA; EU AI Act Art. 50 |
| Identity management | 5 | Scalable remote verification | Liveness evasion; synthetic identity fraud at scale | Multi-factor mandates; C2PA provenance; DDR reporting | TAKE IT DOWN Act; C2PA v2.0; KYC rules |
| Governance & democracy | 3 | Accessible multilingual civic communication | Disinformation; electoral manipulation; trust erosion | Mandatory provenance; platform liability; media literacy | EU AI Act Art. 50; C2PA v2.0; electoral law |
Table 7.
The six ethical principles of the RTFG-STS Framework.
| Principle | Source | TFG interpretation | Mechanisms |
| P1 Beneficence | AI4People [35]; bioethics [72] | Clear societal benefit, demonstrably delivered across demographic groups | M3; M4 |
| P2 Non-maleficence | AI4People [35] | Foreseeable harms prevented at synthesis time, not patched afterwards | M1; M2; M4 |
| P3 Autonomy | AI4People [35] | Subjects decide whether their face is synthesized; users know they face an AI; disclosure is proactive | M2; M5 |
| P4 Justice | AI4People [35]; EU AI Act Art. 10 | Equitable performance across the served population; audited datasets; stratified reporting | M3; M4 |
| P5 Explicability | AI4People [35]; EU AI Act Art. 50 | Output machine-detectable (C2PA) and socially disclosed at interaction | M1; M4 |
| P6 Systemic embeddedness | STS theory [15,16] (new) | Design and governance reference the concrete deployment context; PTS, CCAI, DDR, and a STIA supplement technical metrics | M3; all |
Table 8.
Risk-tiered consent architecture.
| Tier | Context | Consent requirement | Implementation | Regulatory basis |
| 1 (Low) | Entertainment, creative, personal | Disclosure at point of viewing | C2PA manifest; UI label | EU AI Act Art. 50; C2PA v2.0 |
| 2 (Medium) | Education, customer service, public information | Explicit opt-in from the face subject; viewer disclosure; right of withdrawal | Signed consent record in append-only log; manifest carries consent hash | EU AI Act Art. 50; GDPR Art. 6; TAKE IT DOWN Act [27] |
| 3 (High) | Healthcare, identity, legal, political | Revocable informed consent as a Verifiable Credential; periodic re-consent; human oversight | VC per Eq. 6 on a distributed ledger; human confirmation before action | EU AI Act Art. 50, 14; GDPR Art. 9 |
Table 9.
RTFG-STS Framework application map by deployment domain.
| Domain | Tier | Mandatory mechanisms | Metrics | Instruments | Priority research need |
| Healthcare & telehealth | 3 | M1-M5 (Tier 3 VC; DDR ≥ 80%; clinician oversight) | PTS; CCAI; DDR; CSIM; MOS | EU AI Act Art. 50, 10, 14, 27; GDPR Art. 9 | Diverse clinical evaluation data; validated PTS |
| Education & training | 2-3 | M1-M4 (Tier 2/3; DDR ≥ 60%) | PTS; CCAI; DDR; AUE; MOS | EU AI Act Art. 50; GDPR Art. 9; FERPA; COPPA | Student biometric governance; authority-effect studies |
| Identity management | 3 | M1-M5 (Tier 3 VC; DDR ≥ 80%; human ID reviewer) | DDR; CSIM; FID; SyncNet; FPS | TAKE IT DOWN Act [27]; EU AI Act Art. 50; C2PA v2.0; KYC rules | Standardized liveness evaluation; DDR benchmark |
| Governance & democracy | 3 | M1-M5 (Tier 3; DDR ≥ 80%; editorial oversight) | DDR; PTS; FID; MOS | EU AI Act Art. 50; C2PA v2.0; electoral law | Cross-platform C2PA verification; media-literacy studies |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.