Preprint
Review

This version is not peer-reviewed.

Talking Face Generation in Socio-Technical Systems: A Systematic Review of Deep Learning Architectures, Deployment Contexts, and Governance Frameworks

Submitted:

14 August 2026

Posted:

18 August 2026

You are already at the latest version

Abstract
Talking face generation (TFG), the synthesis of photorealistic speaking video from a portrait and an audio signal, has moved through five architectural generations since 2016: Long Short-Term Memory (LSTM) lip-sync systems, Generative Adversarial Networks (GANs), Neural Radiance Fields (NeRFs), Diffusion models, and now Diffusion Transformers (DiT) and Gaussian Splatting. With real-time photorealism, TFG is entering healthcare, education, identity management, and public communication. This systematic literature review synthesises 92 primary studies (January 2016 to December 2025), selected from 18,343 records across six repositories under Kitchenham’s SLR guidelines and PRISMA 2020, to answer six research questions on architecture evolution, dataset bias, evaluation metrics, generative paradigm shifts, ethical safeguards, and socio-technical deployment. DiT and Gaussian Splatting together account for roughly 36% of the most recent studies retrieved and GAN usage has fallen below 3%. Training datasets are demographically skewed, and the dominant metrics (PSNR, SSIM, SyncNet) do not measure what deployment requires. Most seriously, only one of the 92 systems embeds any technical safeguard (a non-compliant post-hoc watermark), even though the EU AI Act (Regulation 2024/1689), the US TAKE IT DOWN Act (2025), and Coalition for Content Provenance and Authenticity (C2PA) v2.0 entered enforcement during the review period. None integrates C2PA-compliant watermarking, consent, or provenance mechanisms. We propose three deployment-fitness metrics: Perceived Trustworthiness Score (PTS), Cross-Cultural Authenticity Index (CCAI), and Deepfake Detectability Rate (DDR); and the Responsible Talking Face Generation in Socio-Technical Systems (RTFG-STS) Framework, grounded in AI4People principles and socio-technical systems theory, with five operational mechanisms aligned to EU AI Act Article 14.
Keywords: 
;  ;  ;  ;  ;  

1. Introduction

1.1. Background and Motivation

Modern deep learning can now synthesize a photorealistic speaking face from a single portrait photograph and an audio clip. Talking face generation (TFG), illustrated in Figure 1, is the automated creation of temporally coherent video in which a face speaks in synchrony with a given audio stream. The problem sits at the intersection of computer vision, audio signal processing, and generative modelling, and in one decade it has grown from a narrow scientific question into a technology with commercial weight and social consequences.
The technical lineage is short but dense. Bregler et al. [1] showed in the 1990s that vocal-tract articulation is tightly coupled with visible facial motion, making data-driven synthesis plausible, and the statistical models of the 2000s formalized the mapping without approaching perceptual realism. Deep learning changed the trajectory. The Synthesizing Obama system of Suwajanakorn et al. [2] demonstrated that broadcast-quality lip synchronization could be learned directly from audio-visual corpora, without hand-crafted features. Generative Adversarial Networks (GANs) [3] then became the generative backbone of choice, and Wav2Lip [4] achieved lip-sync fidelity good enough for in-the-wild use. Large audio-visual corpora such as VoxCeleb [5], VoxCeleb2, and HDTF [6] supplied the training signal to push generalization beyond single-speaker settings.
Three further transitions followed from around 2021. Transformer architectures such as FaceFormer [7] brought long-range audio-visual context modelling; Neural Radiance Fields (NeRFs) [8] were adapted to dynamic portraits in AD-NeRF [9], giving 3D-consistent talking heads; and Diffusion Models [10] then displaced GANs, with DiffTalk [11] and DreamTalk [12] delivering probabilistic, high-fidelity synthesis with better training stability and diversity. By 2024, Diffusion Transformers (DiT) and Gaussian Splatting had begun to remove the long-standing latency barrier: VASA-1 [13] produces 512×512 video at 45 fps in real time, and OmniHuman-1 [14] extends the DiT paradigm to full-body digital human animation.
1.2 Applications and Open Problems
TFG has become socio-technical in the sense of Trist and Bamforth [15] and later socio-technical systems (STS) theory [16]: its deployment consequences are now inseparable from the human and institutional contexts it operates in. In healthcare, TFG-based conversational agents are one response to a projected global shortage of 11.1 million health workers by 2030 [17]; AI physician avatars can deliver patient education with perceived empathy comparable to human clinicians [18], and real-time systems such as MuseTalk [19] and VASA-1 [13] make live telehealth interaction feasible, with emotionally adaptive models such as EmotiveTalk [20] pointing toward mental-health applications. In education, embodied agents with realistic faces improve engagement and retention over disembodied voices or text [21], support virtual reality (VR) anatomy and clinical-communication training [22], and, with dyadic systems such as INFP [23], open up conversational skills training. In media and personalized communication, TFG enables multilingual dubbing, lip-readable synthetic media for hearing-impaired audiences, and consumer applications of latent-space face synthesis [13,24,25]; public institutions can use the same capability for low-cost multilingual civic communication.

1.3. The Regulatory Turn

The potential harms of TFG have triggered a wave of regulation. The EU AI Act (Regulation 2024/1689), adopted in 2024 and fully enforceable from August 2026, classifies deepfakes and synthetic facial media as a transparency risk under Article 50: generated content must be labelled, users must know when they are interacting with an AI-generated face, and Article 50(2) with Recital 134 mandates watermarking, metadata identifiers, and provenance tooling so synthetic media stays machine-detectable [26]. In the United States, the TAKE IT DOWN Act (Public Law 119-12, May 2025) criminalizes distribution of non-consensual AI-generated intimate imagery, with prison terms of up to two years (three where the subject is a minor) and a 48-hour platform takedown obligation [27]. C2PA v2.0 provides cryptographically signed provenance manifests already adopted by major platforms, and China’s synthetic media labelling rules took effect on 1 September 2025 [28].
The timing matters as much as the content. VASA-1 [13] and OmniHuman-1 [14] were published and deployed before any of these instruments entered force. That gap between capability and governance is a familiar pattern in socio-technical systems [16], and closing it requires the kind of systems-level account of TFG that this review sets out to provide.

1.4. Gap, Objectives, and Research Questions

Surveys of TFG and adjacent fields have appeared steadily since 2022 [29,30,31,32,33,34], and Section 2 analyses them in detail: none addresses the regulatory environment, dataset equity, or the socio-technical systems TFG is deployed into, and no existing review is simultaneously rigorous, current to 2024-2025, and comprehensive across technical, ethical, and socio-technical dimensions.
This review addresses that gap through six research questions:
RQ1: What deep learning architectures have dominated TFG from 2016 to 2025, and what trajectory characterizes their evolution?
RQ2: How do biases in the datasets used to train and evaluate TFG models affect generalization, and what are the equity implications for deployment?
RQ3: What metrics are currently used to evaluate TFG systems, and how adequately do they capture the requirements of socio-technical deployment?
RQ4: How have Diffusion Transformers, Gaussian Splatting, and related modern architectures advanced the quality, efficiency, and controllability of TFG relative to GANs and classical NeRFs?
RQ5: What ethical safeguards (detection, consent, watermarking, provenance) are integrated into state-of-the-art TFG systems, and how does this align with emerging regulatory requirements?
RQ6: How does TFG embed within, and generate risks and opportunities for, the socio-technical systems in which it is deployed, including healthcare, education, identity management, and democratic governance?

1.5. Contributions and Structure

This review contributes: (i) an architectural taxonomy extended to Gaussian Splatting and DiT, mapping the TFG design space from 2016 to 2025; (ii) the first systematic assessment of demographic representation across the most widely used TFG benchmark datasets; (iii) a critical analysis of the evaluation metric ecosystem and three proposed deployment-fitness metrics, the Perceived Trustworthiness Score (PTS), Cross-Cultural Authenticity Index (CCAI), and Deepfake Detectability Rate (DDR); (iv) documentation of the safeguard gap, showing that no reviewed system natively integrates watermarking, consent, or provenance, mapped against the EU AI Act, the TAKE IT DOWN Act, and C2PA v2.0; (v) a socio-technical analysis of four deployment domains; (vi) the Responsible Talking Face Generation in Socio-Technical Systems (RTFG-STS) Framework, grounded in Floridi et al.’s AI4People principles [35] and operationalized through five mechanisms aligned to EU AI Act Article 14; and (vii) a formalized three-stage process pipeline for TFG.
Section 2, Section 3, Section 4, Section 5 and Section 6 present, in turn, prior surveys, methodology, the synthesis by research question, the RTFG-STS Framework, and conclusions.

3. Methodology

3.1. Review Framework and Rationale

The review follows the SLR methodology of Kitchenham and Charters [36] in three phases (planning, conducting, reporting), with study identification and reporting per PRISMA 2020 [43]. A scoping review would be inappropriate because RQ3 and RQ5 need quality-filtered synthesis to support defensible conclusions. A meta-analysis is impossible because TFG studies report incompatible metric subsets of FID, SyncNet, PSNR, SSIM, LPIPS, and FVD on incompatible test sets; making statistical pooling meaningless in this context. The Kitchenham protocol accommodates qualitative synthesis of heterogeneous studies while keeping selection transparent and auditable. Figure 3 presents the PRISMA 2020 flow diagram.

3.2. Scope

The review covers deep learning approaches to TFG across the full pipeline from feature extraction to rendering: audio-driven single-image animation, audio-plus-video reenactment, 3D-aware synthesis (NeRF and Gaussian Splatting), diffusion and DiT architectures, and detection or safeguard work directly tied to TFG assessment. Coverage runs from January 2016, when deep recurrent architectures first reached talking face synthesis, to December 2025, fixed at the fourth search iteration in March 2026. Pre-2016 methods (HMMs, DBNs, AAMs) fall outside the deep learning pipeline the research questions examine. Because RQ5 and RQ6 require it, the disciplinary scope extends past computer vision into applied ethics, AI governance, and STS analysis.

3.3. Search Strategy

Keywords were identified in three stages: a scoping search that yielded 47 highly cited papers and 31 candidate terms in three semantic clusters; forward snowballing through the 20 most-cited candidates, adding five terms; and a 2025 revision adding frontier terminology. The final Boolean string, applied to titles, abstracts, and keywords, was:
(“talking face” OR “talking head” OR “audio-driven face” OR “speech-driven face” OR “audio-visual synthesis” OR “portrait animation” OR “lip sync” OR “lip synchronization” OR “video reenactment” OR “deepfake” OR “facial animation deep learning” OR “neural talking” OR “video dubbing” OR “face reenactment” OR “Gaussian splatting avatar” OR “diffusion portrait” OR “facial diffusion transformer” OR “audio-driven digital human”)
The string was executed in tiers: a core tier (“talking face,” “talking head,” “lip sync”) in every iteration; an extended tier (deepfake, portrait animation, reenactment, dubbing) from the second iteration; and a frontier tier only in the fourth, so that terminology which did not exist before 2025 could not distort coverage of older work. Six repositories were searched for disciplinary relevance and complementary coverage: IEEE Xplore, ScienceDirect, the ACM Digital Library, SpringerLink, Google Scholar with Web of Science, and arXiv, where landmark systems such as VASA-1 appeared months before formal publication. Four iterations ran between March 2023 and March 2026; a single frozen search would have missed the diffusion and Gaussian Splatting literature entirely. Table 1 reports retrieval counts.
3.4 Selection Criteria and Procedure
Table 2 states the inclusion and exclusion criteria, which operationalize the scope above. All of I1-I6 must hold for inclusion and any one of E1-E6 excludes.
Selection ran in three stages. Deduplication (DOI match, then title-plus-first-author matching at a 0.95 Levenshtein threshold) reduced 18,343 records to 8,741, and coarse title screening against I3-I5 left 448 for abstract review. Two reviewers screened all 448 abstracts independently (Cohen’s κ = 0.84), excluding 118 (off-topic DL application n = 64; not primary research n = 31; temporal n = 14; language n = 9); eight full texts proved unretrievable, and the remaining 322 were assessed against all criteria and the quality instrument below. Of these, 230 were excluded (opaque methodology n = 78; topic mismatch n = 91; tutorial or chapter n = 31; other n = 30), leaving 92 studies in the synthesis corpus.

3.5. Quality Assessment

Each of the 322 full texts was scored on five criteria adapted from Kitchenham and Charters [36]: problem clarity (QA1), methodological transparency (QA2), evaluation rigor (QA3), limitation acknowledgement (QA4), and contribution significance (QA5). Each criterion scores 1.0, 0.5, or 0.0; the inclusion threshold of 3.0 of 5.0 excludes fundamental deficiencies while admitting honestly reported early-stage work. Scores of the 322 full texts ranged from 0.5 to 5.0; 92 studies (28.6%) met the threshold, with a mean of 4.27 (SD = 0.39) among those included. QA1 was the criterion most often fully satisfied (93.5%), QA4 the most often only partial (41.3% scored 0.5), and QA2 was the most discriminating: studies scoring 0.0 on methodological transparency never exceeded 2.0 overall, making opacity the main driver of exclusion.

3.6. Extraction and Synthesis

A structured extraction form, piloted on 10 studies and refined once, captured bibliographic metadata, technical characteristics, datasets and metrics with quantitative results, five ethical and governance sub-fields (detection, watermarking/provenance, consent, risk discussion, regulatory reference), and deployment characteristics. The lead author extracted all 92 studies and a co-author independently verified a stratified 25% subsample (23 studies), with disagreements resolved against the source text and, in four cases, by email to the original authors. Given the heterogeneity of study types and outcomes, we adopted qualitative narrative synthesis following Cochrane guidance [44]: thematic synthesis within each RQ, followed by a cross-RQ analysis of patterns no single question exposes; quantitative results are tabulated where comparable but never statistically pooled. Where fewer than three independent studies support a finding, this is flagged in Section 4. The full bibliographic record of all 92 included studies, with extracted characteristics and quality scores, is provided as supplementary material (Appendix A).

3.7. Corpus Characteristics

The corpus comprises 56 conference papers (60.8%), 24 journal articles (26.1%), and 12 preprints (13.0%), consistent with a field whose flagship venues are conferences: CVPR leads with 22, followed by ECCV (9), ACM Multimedia (8), ICCV (7), and others. Annual output grows from 2 included studies in 2016-2017 to a roughly 20 per year across 2023-2025.

3.8. Threats to Validity

Four threats deserve note. The English-language restriction and database selection likely underrepresent Chinese-language venues, a real limitation given how much TFG research originates from Chinese institutions. Publication bias favors positive results; including arXiv preprints softens but does not remove it. The frontier described here is accurate to March 2026 and may be partially superseded at reading time; the multi-iteration search was designed to limit this decay. Finally, quality scoring involves judgement, particularly on QA3 and QA5; the κ = 0.84 agreement and explicit scoring descriptors constrain but do not eliminate subjectivity. RQ6 rests on a thinner evidence base than RQ1-RQ4, since primary studies rarely report deployment outcomes; the analysis there leans on the ECA literature and regulatory documents.

4. Data Synthesis and Analysis

4.1. Corpus Overview

The 92 studies span 2016 to 2025 and nine architecture families, drawing on 19 distinct named training and evaluation corpora (plus custom single-identity datasets) and 17 distinct evaluation metrics. GAN-based work is the largest single family (22 studies, 23.9%), a legacy of its 2018-2022 dominance, but the three post-GAN families together, Diffusion Models (16, 17.4%), Diffusion Transformers (7, 7.6%), and Gaussian Splatting (8, 8.7%), account for 33.7% of the corpus and overtake it. Two numbers foreshadow the RQ5 findings: only 8 studies (8.7%) contain any ethical content, and only 1 (1.1%) embeds a safeguard mechanism. Fewer than half (44, 47.8%) state a deployment domain at all.

4.2. The TFG Process Pipeline

Across all nine families, virtually every reviewed system decomposes into three sequential modules: feature extraction, a mapping network, and a transformation network (Figure 4). This shared scaffold is what makes cross-era comparison possible in RQ1, RQ3, and RQ4.
Feature extraction encodes the source image x I and audio sequence x A into latent representations:
F I = E I x I ,
F A = E A x A ,
where E I is an image encoder (a face-recognition backbone, or a VAE encoder in diffusion systems) and E A an audio encoder (a mel-spectrogram CNN or a pre-trained speech encoder). The mapping network learns the cross-modal correspondence between audio features and facial motion:
M t = f m a p F t A , F I ,
This is the most architecturally variable component in the corpus, realised as recurrent lip-displacement regression (GAN era), audio-visual cross-attention (transformers), audio-conditioned deformation fields (NeRF), or conditioning vectors injected into the denoising network (diffusion); M t may cover lips, 3DMM expression coefficients, pose, gaze, and blinks. The transformation network then synthesises each frame from the source image and motion parameters:
x ^ t = G x I , M t ,
The form of G is what distinguishes the architectural eras: an adversarially trained generator with combined adversarial, lip-synchrony (SyncNet [39]), identity, and perceptual losses in GAN systems; iterative denoising conditioned on identity and motion in diffusion systems [10]; and, in Gaussian Splatting systems, differentiable rasterisation of learned 3D Gaussians whose position, opacity, covariance, and colour parameters are jointly optimised from the audio-conditioned motion prior.

4.3. RQ1: Architecture Evolution, 2016-2025

Figure 5 shows the distribution of included studies by primary architecture family and year. Five phases are identifiable, each anchored by a specific breakthrough.
Phase I, recurrent foundations (2016-2018). LSTM and bidirectional RNN models mapped mel-spectrogram frames to lip displacements: Fan et al. [45] set the deep recurrent baseline, and Suwajanakorn et al. [2] reached broadcast quality with 17 hours of single-speaker footage, at the cost of strict speaker dependence.
Phase II, GAN ascendancy (2019-2022). Adversarial training lifted the quality ceiling: Zhou et al. [46] disentangled audio and visual representations for speaker-agnostic synthesis, Wav2Lip [4] achieved the first in-the-wild lip sync by training against a pre-trained SyncNet discriminator, and MakeItTalk [47] and PC-AVS [48] added speaker-aware animation and pose control.
Phase III, transformers and NeRFs (2021-2023). FaceFormer [7] replaced recurrence with self-attention for long-range temporal coherence, while AD-NeRF [9] conditioned a radiance field on audio, giving 3D-consistent talking portraits at the price of hours of per-identity optimization and inference far below real time.
Phase IV, diffusion (2023-2024). DiffTalk [11] and DreamTalk [12] showed that audio- and identity-conditioned diffusion beats GANs on diversity and artefacts without mode collapse; SadTalker [40] added pose-controllable synthesis without per-identity tuning; and VASA-1 [13] closed the phase with a disentangled face latent space and flow matching, at 512×512 and 45 fps in real time.
Phase V, DiT and Gaussian Splatting (2024-2025). Two paradigms now co-define the frontier: DiTs replace the U-Net with a transformer over latent tokens and scale quality with compute [49], with OmniHuman-1 [14] dissolving the boundary between TFG and full-body digital humans, while Gaussian Splatting systems (GSTalker [50]; GaussianEmoTalker [51]; VASA-3D [52]) replace volume rendering with fast rasterization and remove the latency problem outright. Table 3 summarizes the full taxonomy.
The trajectory is a series of paradigm shifts, not a smooth curve, and each shift has an identifiable cause. GANs gave way to diffusion because mode collapse produced systematic artefacts in high-motion regions (teeth and hair most visibly). NeRFs gave way to Gaussian Splatting because volume rendering could not meet real-time requirements. U-Net diffusion is giving way to DiT because transformers scale better with parameter count. Phase V is distinguished by simultaneity: DiT-scale models produce the highest quality but demand server-class hardware, Gaussian Splatting runs in real time at quality sufficient for most deployments, and a hybrid, DiT-generated motion rendered through Gaussian Splatting, is the most promising route to closing that gap. Of the 41 corpus studies dated 2024 or later, 15 (36.6%) are DiT or Gaussian Splatting systems, while only one is GAN-based (2.4%).

4.4. RQ2: Dataset Biases and Generalization

Figure 6 shows dataset usage across the corpus: VoxCeleb2 dominates (60 studies, 65.2%), followed by HDTF (36, 39.1%) and MEAD (18, 19.6%); because most systems train and evaluate on several corpora, counts sum to more than the corpus size. A further 12 studies (13.0%) rely wholly or partly on custom single-identity corpora, and the 3D vertex-animation corpora VOCASET and BIWI (3 studies each) serve the FaceFormer line of work. Figure 6(b) characterizes the most used corpora.
The most consistent finding is under-documentation. Of the 19 named corpora identified, only CREMA-D reports ethnic diversity statistics in its own paper, and CREMA-D appears in exactly one included study. VoxCeleb2, which underpins nearly two-thirds of the corpus, reports subject counts and languages but nothing systematic about ethnicity, geography, or age. Studies running cross-demographic evaluations on it ([S16], [S21], [S31]) consistently note that South Asian, Middle-Eastern, African, and Latin American faces are severely underrepresented, several estimating these groups at under 8% of subjects combined. HDTF is built from 362 news and political broadcast subjects, predominantly white and East Asian male speakers of American or British English.
The skew is measurable in performance. Multiple studies report degraded lip-sync confidence when VoxCeleb2-trained models are evaluated on South Asian, African, and Middle-Eastern subjects; exact figures vary, but the direction is consistent and matches the face-recognition fairness literature [64]. The practical meaning is stark: a telehealth avatar trained on the dominant benchmark will articulate less accurately for the faces and languages of Sub-Saharan Africa, South Asia, and the Middle East, precisely where clinician shortages make AI-mediated care most valuable.
Two further gaps compound this. Affective coverage is thin: only MEAD (18 studies) and CREMA-D (1 study) carry emotion labels, and neither exceeds 91 subjects, so the emotion disentanglement reported by systems such as EmotiveTalk [20] and EAT [65] rests on very small affective corpora. Linguistic coverage is thinner still: 14 of the 19 named corpora are English-only, the VoxCeleb pair and CelebV-HQ are nominally multilingual but Western-dominant in practice, and ViCo (2 studies) is the sole substantively non-English corpus. Models trained on English phonemes produce implausible lip shapes for tonal languages (Mandarin, Thai, Yoruba) and for complex consonant clusters (Arabic, Polish).
The problem is not dataset size but diversity and documentation. Fixing it requires demographic audits as a publication prerequisite, open release of phonemically and ethnically diverse corpora, and demographic-stratified evaluation of the kind standard in face-recognition fairness work [64]; EU AI Act Article 10, requiring training data representative of the served population, now supplies regulatory pressure in the same direction.

4.5. RQ3: Evaluation Metrics and Deployment Fitness

Figure 7 shows metric frequency across the corpus. Distribution- and synchrony-level measures lead: FID appears in 60 studies (65.2%) and SyncNet-derived synchrony scores in 51 (55.4%), with the signal-fidelity pair SSIM (43, 46.7%) and PSNR (34, 37.0%) close behind and human MOS evaluation in a third of studies (33, 35.9%). Figure 7(b) assesses each metric family against three deployment-fitness criteria.
Three problems stand out. First, the dominance of PSNR and SSIM is hard to defend. Both measure pixel correspondence to a specific reference, which conflates quality with reference proximity: a slightly blurred but naturally animated face can score below a sharp but temporally incoherent one. The community has known PSNR is a poor perceptual proxy since at least 2012 [66]; its persistence in 2024-2025 papers is evaluation inertia, and it matters: a video that scores well on PSNR while looking uncanny to a patient is not fit for use as a virtual clinician.
Second, SyncNet [39] was trained on English broadcast correspondences, and no included study validates it on non-English phonemic inventories. It also measures correspondence, not naturalness: minimal lip movement scores well on neutral phonemes while missing the jaw openings, bilabial closures, and fricative patterns that make speech look real. Models optimized hard against SyncNet drift toward stiff, over constrained lip motion that is metrically accurate and perceptually unconvincing.
Third, the metrics that matter most for deployment remain rare or shallow. CSIM, the only standardized identity-preservation measure, appears in 15.2% of studies, and FVD, which captures the temporal coherence that per-frame metrics miss, in 13.0%. Human MOS evaluation is more widespread (35.9%) but typically shallow: panels of fewer than 30 raters are the norm, below accepted sample sizes for reliable MOS work, and no study recruits raters from its intended deployment population. The one metric whose adoption is clearly rising, FPS (15.2%), tracks the field’s real-time turn rather than any perceptual or governance property. Human perception remains the fitness measure that deployment actually turns on.
We therefore propose three metric dimensions as a research agenda. The Perceived Trustworthiness Score (PTS) is a human evaluation protocol in which raters drawn from the target deployment population (patients for healthcare avatars, students for educational agents) score credibility, warmth, and professional appropriateness on a multi-item Likert scale; it supplements rather than replaces MOS. The Cross-Cultural Authenticity Index (CCAI) extends SyncNet-style evaluation across at least five typologically distinct language families (for example Indo-European, Sino-Tibetan, Afro-Asiatic, Niger-Congo, Dravidian), exposing the cross-lingual limits of English-trained models and creating an incentive for diverse training data. The Deepfake Detectability Rate (DDR) is the proportion of generated frames correctly classified as synthetic by an ensemble of state-of-the-art detectors; it bears directly on EU AI Act Article 50, and systems below 80% DDR should carry mandatory C2PA watermarking as a compensating control. All three are proposed with provisional parameters and need validation through large, demographically diverse human studies before serving as standards.

4.6. RQ4: DiT and Gaussian Splatting Versus GANs and NeRFs

Figure 8 compares four architecture eras on visual quality, lip synchrony, and real-time capability; Table 4 reports benchmark data for representative systems.
On quality, the FID trajectory is unambiguous: GAN and early hybrid systems report FIDs in the 41-48 range (2021-2023), early diffusion at 35-47 in 2023, VASA-1 at 22.1, OmniHuman-1 at 18.4, a 55% relative reduction from the GAN baseline in four years, a pattern matching image generation more broadly, where DiT architectures keep improving predictably with scale.
On latency, 2024-2025 resolved the long-standing trade-off in three separate ways: flow matching (VASA-1 [13]) cut inference-time function evaluations without losing diversity; Gaussian rasterization (GSTalker [50]; VASA-3D [52]) replaced volume rendering for an order-of-magnitude speedup; and autoregressive streaming (Teller [60]) predicts motion tokens at sub-frame latency for continuous live output. Three architecturally unrelated solutions arriving at once suggest the barrier is genuinely down, not narrowly circumvented.
Controllability advanced in parallel and is invisible to FID and SyncNet. Wav2Lip [4] controlled lips only; PC-AVS [48] added pose; SadTalker [40] added stylistic motion via 3DMM coefficients; EmotiveTalk [20] decouples audio into content and emotion for independent affective control; and VASA-1 [13] disentangles pose, expression, gaze, blink, and speaking style into orthogonal latent dimensions with fine-grained post-hoc editing. For deployment this is not a luxury: a tutor that cannot adapt its emotional register teaches worse, and a clinician avatar with one fixed expression reads as uncanny.

4.7. RQ5: Ethical Safeguards and Regulatory Alignment

Figure 9 presents the starkest finding in the corpus. Of 92 included studies, only 8 (8.7%) contain any ethical content at all, and exactly 1 (1.1%) integrates a technical safeguard: a frequency-domain watermark added as a post-processing step, non-C2PA-compliant and untested against compression. No study implements a consent protocol. No study cites the EU AI Act, the TAKE IT DOWN Act, C2PA, GDPR Article 9, or any other governance instrument as a design constraint. Table 5 details the inventory.
Regulatory instruments are now in force and the gap is comprehensive. EU AI Act Article 50(2) requires machine-detectable marking by August 2026 (no reviewed system embeds a compliant manifest); Recital 134 requires marks robust to compression and conversion (no study tests robustness); Article 10 requires representative training data (no study audits its dataset demographically); and Article 14 requires meaningful human oversight (no system includes an oversight or revocation mechanism). The TAKE IT DOWN Act’s consent conditions go unaddressed, and C2PA is not referenced in a single included paper.
We do not read this as individual negligence: publication incentives reward benchmark quality, ethics review does not yet require safeguard implementation, and liability-creating regulation arrived only at the end of the review period. The result is a compliance cliff: a large body of deployable state-of-the-art systems facing EU AI Act obligations from August 2026 while satisfying none of them. Section 5 offers an operational path across.

4.8. RQ6: TFG in Socio-Technical Systems

Figure 10 shows the coded deployment domains. Most studies (48, 52.2%) specify no domain, treating TFG as a context-free technical problem; among the rest, entertainment leads (18, 19.6%), then education (9), healthcare (7), identity management (5), governance (3), and accessibility (2). The analysis below covers the four domains most consequential for this review, drawing on the 24 domain-specific studies plus the adjacent literature from Section 2.
Healthcare. DeTore et al. [67] found TFG-powered avatars could administer validated clinical instruments (PHQ-9, GAD-7) with response quality comparable to human-administered assessment; Haider et al. [18] found avatar-delivered patient education matched human instruction on comprehension while cutting clinician time by 34%. The risk profile has three parts: patients must know they are talking to an AI (an Article 50 obligation no healthcare-focused study reports implementing); the RQ2 performance gaps translate into reduced comprehension of clinical information, with patient safety implications; and liability for incorrect AI-delivered medical advice is unassigned in every jurisdiction we could identify.
Education. Chheang et al. [22] showed higher knowledge retention with TFG-powered VR anatomy assistants than with text; EmotiveTalk [20] sketches emotionally adaptive tutoring at scale, though the cross-lingual limits from RQ2 constrain language learning uses. The risks are distinctive: interactive tutors that read student facial affect are collecting biometric data governed by GDPR Article 9 and FERPA, which no included study addresses; the authority effect of face-to-face communication [42] may lead students to under-scrutinize AI-delivered content; and convergence on a few commercial avatar products could homogenize pedagogy.
Identity management. This is where the threat is already quantifiable. Venkatesh et al. [68] show that presentation attack detection (PAD) systems generalize poorly to attack types unseen in training. PAD calibrated on GAN artefacts (spectral inconsistencies, flicker, blink anomalies) can fail outright against the smooth, physiologically coherent output of flow-matching systems such as VASA-1 [13] or Gaussian Splatting systems such as GSTalker [50], which produce none of those signatures. The iProov threat intelligence report [69] documents sharply rising real-world PAD bypasses through 2023-2024, coinciding with diffusion TFG availability, and neither VASA-1 nor GSTalker has been formally evaluated against current liveness detection in any peer-reviewed study we identified. The institutional consequences are already on record: the USD 25 million Arup fraud [70], and Pindrop’s finding that over a third of a 300-profile applicant sample was wholly fabricated. A convincing deepfake video interview now costs a consumer GPU and a public model checkpoint. In systems terms, this erodes the trust infrastructure of digital identity itself [16]; no single countermeasure suffices, which is why Section 5 argues for coordinated governance across standards, law, platforms, and institutional practice.
Governance and democratic information. Political deepfakes of real politicians were documented in at least seven national election contexts between 2022 and 2024, and photorealistic talking faces raise the credibility and virality of false information, with lower-income and lower-digital-literacy populations disproportionately exposed. Detection is caught in an arms race: as generation improves (RQ4), detectors trained on older outputs decay. Mandatory provenance (EU AI Act Article 50; C2PA v2.0) is a structural exit from that race, binding content cryptographically to its generating model at creation so verification no longer depends on out-detecting generation quality. Table 6 summarizes the four domains.

4.9. Cross-RQ Patterns

Four patterns emerge only at the intersection of the research questions.
Optimization-deployment misalignment. The field optimizes PSNR, SSIM, and SyncNet; the deployment contexts that matter require identity preservation, cross-cultural authenticity, trustworthiness, and detectability, which the field does not measure. This reflects incentive structure, not oversight, and it will persist until venues and funders require PTS, CCAI, and DDR alongside FID.
Democratization and concentration at once. Gaussian Splatting and lightweight diffusion push high-fidelity TFG onto consumer hardware, while DiT-scale models (OmniHuman-1 [14]; EMO2 [71]) concentrate frontier capability in organizations with large compute budgets. The result is a two-tier field in which the hardest-to-detect systems are least accessible for beneficial use, and the most accessible systems are both the most useful and the easiest to misuse.
The absent safeguard as structural failure. That 91.3% of studies contain no ethical content and 98.9% no safeguard reflects the field’s incentive design. Ethics review at NeurIPS (2021) and CVPR (2023) has not measurably shifted the pattern, suggesting domain-specific requirements (watermarking, consent, DDR reporting) need to become submission conditions.
Governance lag as predictable risk. Each capability jump (2019 GAN photorealism, 2023 diffusion fidelity, 2024 real-time Gaussian Splatting) was followed by roughly 18-24 months of misuse before any governance response. The EU AI Act’s technology-neutral design may weather the next transition better than capability-specific rules; whether enforcement is agile enough is open.

5. The RTFG-STS Framework

5.1. Theoretical Foundations

The framework’s ethical core is the AI4People report of Floridi et al. [35], which joined the four bioethical principles of Beauchamp and Childress [72] (beneficence, non-maleficence, autonomy, justice) with a fifth, AI-specific principle, explicability. AI4People underpins the European Commission’s AI ethics guidelines and the philosophy of the EU AI Act [26], and it specifies what an ethical system must do, not only what it must avoid; on its own, though, it addresses the AI system in isolation. Socio-technical systems theory [15,16] supplies the missing half: a virtual clinician must be judged by its behavior inside a clinical workflow, its effect on patient trust, and its interaction with consent regulation, not by FID alone. Bijker’s concept of socio-technical stabilization [16] explains why the governance lag in Section 4.9 is costly: governance built after a technology hardens into institutional practice faces one that is already difficult to reshape. The RTFG-STS Framework therefore applies AI4People’s principles at the level of the deployed socio-technical system and adds a sixth principle, systemic embeddedness, which AI4People lacks and the STS analysis demands.

5.2. Six Ethical Principles of the RTFG-STS Framework

Table 7 states the six principles, their interpretation for TFG, and the operational mechanisms (Section 5.3) that implement each.

5.3. Five Operational Mechanisms

M1: C2PA-compliant watermarking and provenance (P2, P5). Every deployed TFG system must embed a cryptographically signed provenance manifest per C2PA v2.0, satisfying EU AI Act Article 50(2) and Recital 134:
C 2 P A I t = S i g n M a n i f e s t I t , m e t a d a t a , m o d e l h a s h , K i s s u e r ,
The manifest encodes the provenance claim, the generating model’s identity hash, a timestamp, hashes of the inputs, and the issuer’s certificate, signed with a key on the C2PA Trust List. Recital 134 requires robustness to H.264/H.265 compression, rescaling, and format conversion, a requirement the corpus’s single watermarking study does not meet; DCT-domain embedding with error-correction coding is current best practice pending a C2PA-TFG robustness standard, and systems should report the proportion of frames whose manifest survives compression, treating results below 80% as a barrier to high-risk deployment.
M2: Risk-tiered consent (P2, P3). Consent obligations scale with deployment risk, in three tiers aligned to the EU AI Act’s classification (Table 8). At Tier 3, consent is stored as a Verifiable Credential:
V C I = S i g n D I D s u b j e c t D I D i s s u e r c o n s e n t S c o p e t i m e s t a m p r e v o c a t i o n E n d p o i n t , K s u b j e c t ,
where the DIDs identify the consent subject and system operator, consentScope bounds permitted uses, and the revocation endpoint lets the subject withdraw consent, obliging removal of synthesized content within 48 hours, the TAKE IT DOWN Act timeframe [27].
M3: Socio-Technical Impact Assessment (P1, P4, P6). Before deployment beyond controlled research, the system undergoes a STIA, modelled on the EU AI Act’s Fundamental Rights Impact Assessment (Article 27) but extended to the full deployment context, and maintained as a living document across major revisions. Its six modules assess the target socio-technical system and the TFG system’s role in it; demographic performance equity (stratified evaluation plus CCAI); effects on the behavior, trust, and wellbeing of interacting humans (via PTS); the institutional processes and authority structures the system will intersect; the feedback loops likely to emerge over time; and regulatory compliance against the EU AI Act, the TAKE IT DOWN Act, GDPR, C2PA, and jurisdiction-specific rules, with a remediation timeline for gaps.
M4: Deepfake detectability audit (P2, P4, P5). DDR is defined as
D D R = 1 N i = 1 N 1 D e n s e m b l e f i = " AI - generated " ,
where f i is the i-th generated frame and D e n s e m b l e an ensemble of at least three state-of-the-art detectors, computed after H.264 compression at 2,000 kbps to approximate platform distribution. Three thresholds apply: at DDR ≥ 80%, content is reliably machine-detectable and a C2PA manifest suffices; between 60% and 80%, C2PA plus interaction-point disclosure is mandatory and Tier 3 deployment requires a regulatory waiver; below 60%, Tier 2-3 deployment is prohibited pending improvement, with research use permitted under manifest. The 80% floor follows NIST SP 800-76-2 guidance for detection rates in government identity contexts [73]; the 60% boundary matches the empirical detection floors reported for GAN-era detectors facing novel synthesis modalities [33,68]. Both are provisional pending a standardized TFG-specific DDR benchmark suite.
M5: Human oversight layer (P3, P5, P6). For Tier 3 deployments, a human must stand between generated output and consequential action, per EU AI Act Article 14. Concretely: a qualified clinician reviews clinical content an avatar delivers, with the AI nature disclosed to the patient; a trained reviewer verifies identity through a non-video channel before any decision rests on a video interview, with C2PA verification mandatory; a certified forensic reviewer attests provenance for TFG content in legal proceedings; and TFG video of public figures carries Article 50-compliant disclosure under the editorial responsibility of a publisher who accepts liability.

5.4. Application Across Deployment Domains

Table 9 translates the framework into concrete obligations for the four domains analysed in RQ6.

5.5. Framework Limitations

The framework is a normative proposal derived from the synthesis findings, AI4People, and STS theory; it has not been validated in live deployments, and pilot studies in healthcare and education are needed to test whether the STIA and consent tiers are operationally feasible. It is calibrated to the regulatory picture of December 2025 and should be re-checked against regulation annually. DDR depends on the currency of its detection ensemble: an ensemble calibrated to GAN artefacts says nothing valid about Gaussian Splatting output, so it must be refreshed annually, ideally against a public benchmark maintained by standards bodies. Finally, the framework specifies what responsible deployment requires but cannot enforce it; enforcement will come from venue ethics policies, funder conditions, and supervisory authorities, and advocacy toward venue-level adoption at CVPR, NeurIPS, and ACM MM is the most direct route to uptake.

6. Conclusions

6.1. Principal Findings

Across 92 primary studies from 2016 to 2025, the picture is of a field whose technical progress has consistently outrun its governance and evaluation. Architecturally (RQ1), TFG crossed five phases in nine years; the 2024-2025 frontier is contested between DiT and Gaussian Splatting, which together account for roughly 36% of the most recent studies and have jointly resolved the historic latency-quality trade-off. The training data beneath this achievement (RQ2) skews toward East Asian and Western European faces, with no dataset paper in the corpus reporting a demographic audit, so geographic and linguistic bias is written into the model weights serving the very populations the data underrepresents. Evaluation (RQ3) still runs on PSNR, SSIM, and SyncNet, none of which measures identity preservation, cross-cultural naturalness, trustworthiness, or detectability. Modern architectures (RQ4) delivered a 55% FID reduction over the GAN baseline and three independent routes to real time, yet this capability exists without safeguards (RQ5): zero systems with compliant watermarking, zero consent protocols, zero references to any regulatory instrument as a design constraint. The consequences (RQ6) are no longer hypothetical: USD 25 million lost to a single deepfake video call, systematic liveness bypass, and measurable erosion of trust in digital identity.
Read through STS theory, these findings share one cause: the field’s technological frame, its goals, metrics, and rewards, is built entirely around benchmark optimization, with no structural place for governance, equity, or deployment fitness. Widening that frame is beyond any individual researcher; it needs coordinated action from venues, funders, regulators, and the institutions now deploying TFG.

6.2. Implications

For researchers, the most actionable change is ethics-by-design reporting: DDR alongside FID and SyncNet, disclosed demographic audits of training data, and a stated deployment context with a fitness assessment against it. The cost is modest next to building a frontier system. For practitioners, the message is blunter: no currently available state-of-the-art system meets the obligations EU AI Act Article 50 imposes from August 2026, so deploying VASA-1, MuseTalk, or OmniHuman-1 in a patient-facing, student-facing, or KYC context today without C2PA watermarking, Tier 3 consent, and human oversight is a compliance liability with a dated deadline. For policymakers, three interventions exceed any current instrument: public funding for an open, demographically diverse, multilingual TFG dataset program; extension of ISO/IEC SC37 presentation attack standards to TFG synthesis attacks; and platform liability for hosting synthetic media without C2PA verification, which does not depend on winning the detection arms race.

6.3. Limitations

Four limitations apply. The English-language restriction underrepresents Chinese-language venues despite Chinese institutions being among the field’s most active contributors. The regulatory analysis is dated to December 2025 and may be sharpened by the EU AI Office’s forthcoming Code of Practice. The RQ6 analysis rests on 24 domain-specific studies plus adjacent literature, so its claims about political communication and education are propositions to test rather than settled findings.

6.4. Closing Statement

TFG can now synthesize photorealistic, real-time, emotionally controllable speech video from one image and an audio clip. Whether we can build it is settled; whether we can govern it is not. The EU AI Act, the TAKE IT DOWN Act, and C2PA v2.0 are the regulatory answers, and the RTFG-STS Framework is ours: embed consent, provenance, equity, and deployment fitness as design prerequisites rather than afterthoughts. Whether the next nine years produces systems that serve the rural clinic and the multilingual classroom without also arming the fraudster and the misinformed is not a technical question. It is a socio-technical one, and it deserves a socio-technical answer.

Author Contributions

Conceptualization, R. Salahudeen; methodology, R. Salahudeen and A. Garba; software, R. Salahudeen and A. Garba; data curation, R. Salahudeen; writing—original draft preparation, R. Salahudeen; writing—review and editing, R. Salahudeen; visualization, R. Salahudeen; supervision, M. Fonkam, N.R. Vajjhala and S.B. Junaidu. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Acknowledgments

“During the preparation of this manuscript/study, the author(s) used MS Word Gmail and Google Drive for the purposes of editing, correspondence, sharing and versioning. The authors have reviewed and edited the output and take full responsibility for the content of this publication.

Conflicts of Interest

“The authors declare no conflicts of interest.”.

Appendix A

Studies are organised by primary architecture family and chronological order within each family. QA scores are out of 5.0; minimum inclusion threshold = 3.0. Arch. = Architecture; Diff. = Diffusion; Transf. = Transformer; GS = Gaussian Splatting; DiT = Diffusion Transformer. Citations marked † are arXiv preprints confirmed as active at time of search.
Table A1. Complete List of Included Primary Studies (N = 92).
Table A1. Complete List of Included Primary Studies (N = 92).
No. Authors (Year) Abbreviated Title Venue Arch.
Family
Training
Dataset(s)
Key Metrics
Reported
QA
Score
LSTM / RNN (n = 5)
S1 Fan et al. (2016) [45] A Deep Bidirectional LSTM Approach for Video-Realistic Talking Head Multimed. Tools Appl. LSTM/RNN Custom (single-spk) PSNR, SSIM, MOS 3.5
S2 Suwajanakorn et al. (2017) [2] Synthesizing Obama: Learning Lip Sync from Audio ACM TOG (SIGGRAPH) LSTM/RNN Obama Weekly Address Lip Sync Error, MOS 4.8
S3 Song et al. (2018) [53] Talking Face Generation by Conditional Recurrent Adversarial Network IJCAI 2018 LSTM/RNN Grid, LRW PSNR, SSIM, LMD 3.8
S4 Vougioukas et al. (2020) [74] Realistic Speech-Driven Facial Animation with GANs IJCV 2020 LSTM/RNN GRID, VoxCeleb FID, PSNR, SyncNet 4.0
S5 Chen et al. (2019) [75] Hierarchical Cross-Modal Talking Face Video Generation ACM MM 2019 LSTM/RNN LRW, VoxCeleb PSNR, SSIM, LMD 3.5
CNN-based (n = 7)
S6 Zhou et al. (2018) [54] VisemeNet: Audio-Driven Animator-Centric Speech Animation ACM TOG (SIGGRAPH) CNN CMU-MOSI, LRW MOS, CPBD 4.0
S7 Siarohin et al. (2019) [76] First Order Motion Model for Image Animation NeurIPS 2019 CNN VoxCeleb, BAIR FID, SSIM, AKD 4.5
S8 Jamaludin et al. (2019) [77] You Said That?: Synthesising Talking Faces from Audio IJCV 2019 CNN BBC News, LRS2 Sync Error, PSNR 3.8
S9 Wiles et al. (2018) [78] X2Face: A Network for Controlling Face Generation ECCV 2018 CNN VoxCeleb, iHead SSIM, FID, MOS 3.8
S10 Wang et al. (2021) [79] One-Shot Free-View Neural Talking-Head Synthesis CVPR 2021 CNN VoxCeleb2, HDTF FID, CSIM, AED 4.2
S11 Cheng et al. (2022) [80] VideoReTalking: Audio-Based Lip Sync for Talking Head Editing SIGGRAPH Asia 2022 CNN HDTF, VoxCeleb2 SyncNet, LMD, FID 4.0
S12 Das et al. (2020) [81] Speech-Driven Facial Animation Using Cascaded GANs ECCV 2020 CNN VoxCeleb, GRID PSNR, SSIM, SyncNet 3.5
GAN-based (n = 22)
S13 Zhou et al. (2019) [46] Talking Face Generation by Adversarially Disentangled A-V Representation AAAI 2019 GAN VoxCeleb, LRW LMD, PSNR, MOS 4.5
S14 Prajwal et al. (2020) [4] Wav2Lip: A Lip Sync Expert Is All You Need for Speech to Lip Generation ACM MM 2020 GAN LRS2, VoxCeleb2 SyncNet, LSE-D, LSE-C 5.0
S15 Zhou et al. (2020) [47] MakeItTalk: Speaker-Aware Talking-Head Animation ACM TOG 2020 GAN VoxCeleb, GRID LMD, FID, MOS 4.5
S16 Salahudeen et al. (2024) [24] Photo-Realistic Talking Face Generation under Latent Space Manipulation IEEE TCE 2024 GAN VoxCeleb2, MEAD FID, SyncNet, PSNR, SSIM 4.2
S17 Salahudeen & Siu (2023) [25] Activate Your Face in Virtual Meeting Platform ICCE-Taiwan 2023 GAN VoxCeleb2 SyncNet, SSIM 3.2
S18 Zhou et al. (2021) [48] Pose-Controllable Talking Face Generation (PC-AVS) CVPR 2021 GAN VoxCeleb2 FID, SyncNet, CSIM 4.8
S19 Ji et al. (2021) [82] Audio-Driven Emotional Video Portraits CVPR 2021 GAN MEAD, VoxCeleb2 FID, PSNR, MOS 4.2
S20 Wang et al. (2021) [83] Audio2Head: Audio-Driven One-Shot Talking-Head Generation IJCAI 2021 GAN VoxCeleb, HDTF FID, SyncNet, SSIM 4.0
S21 Hong et al. (2022) [84] Depth-Aware Generative Adversarial Network for Talking Head CVPR 2022 GAN VoxCeleb2, HDTF FID, CPBD, LMD 4.2
S22 Yin et al. (2022) [85] StyleHEAT: One-Shot High-Resolution Editable Talking Face via StyleGAN ECCV 2022 GAN VoxCeleb2, HDTF FID, PSNR, SSIM, LPIPS 4.5
S23 Liang et al. (2022) [86] Expressive Talking Head Generation with Granular Audio-Visual Control CVPR 2022 GAN MEAD, VoxCeleb2 FID, MOS, AUE 4.0
S24 Song et al. (2022) [87] Talking Face Generation with Multilingual TTS CVPR 2022 GAN HDTF, LRW SyncNet, FID, LMD 4.0
S25 Lu et al. (2021) [88] Live Speech Portraits: Real-Time Photorealistic Talking-Head Animation ACM TOG 2021 GAN Custom PSNR, SSIM, MOS 4.5
S26 Zhu et al. (2021) [89] Arbitrary Talking Face Generation via Attentional A-V Coherence Learning IJCAI 2021 GAN VoxCeleb, MEAD FID, SyncNet, SSIM 3.8
S27 Wen et al. (2020) [90] Photorealistic Audio-Driven Video Portraits IEEE TVCG 2020 GAN Custom (Obama) PSNR, SSIM, MOS 4.0
S28 Zhang et al. (2021) [6] Flow-Guided One-Shot Talking Face Generation ACM MM 2021 GAN VoxCeleb2, HDTF FID, SyncNet, LPIPS 4.0
S29 Hwang et al. (2023) [91] DisCoHead: Disentangled Control of Head Pose and Facial Expressions ICASSP 2023 GAN VoxCeleb2, HDTF FID, LMD, SyncNet 3.8
S30 Ji et al. (2022) [92] EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion Model ACM SIGGRAPH 2022 GAN MEAD, VoxCeleb2 FID, PSNR, SSIM, AUE, MOS 4.5
S31 Yi et al. (2022) [93] Predicting personalized head movement from short video and speech signal IEEE TMM 2022 GAN VoxCeleb2, MEAD PSNR, SSIM, LMD, MOS 4.0
S32 Wang et al. (2022) [94] One-Shot Talking Face Generation from Single-Speaker A-V Correlation AAAI 2022 GAN VoxCeleb2 FID, SyncNet, CSIM 3.8
S33 Peng et al. (2023) [95] SelfTalk: Self-Supervised Training for 3D Talking Faces ACM MM 2023 GAN MEAD, VoxCeleb2 FID, AUE, MOS 4.0
S34 Chen et al. (2020) [96] Talking Head Generation with Rhythmic Head Motion ECCV 2020 GAN VoxCeleb2 FID, LMD, MOS 3.8
Transformer (n = 13)
S35 Fan et al. (2022) [7] FaceFormer: Speech-Driven 3D Facial Animation with Transformers CVPR 2022 Transformer VOCASET, BIWI LVE, EVE, MOS 5.0
S36 Peng et al. (2024) [55] SyncTalk: The Devil Is in the Synchronisation for Talking Head Synthesis CVPR 2024 Transformer VoxCeleb2, HDTF SyncNet, PSNR, SSIM, FID 4.5
S37 Xing et al. (2023) [56] CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion Prior CVPR 2023 Transformer VOCASET, BIWI LVE, EVE, FDD 4.8
S38 Zhong et al. (2023) [97] IP-LAP: Identity-Preserving Talking Face Generation with Landmark Priors CVPR 2023 Transformer VoxCeleb2, HDTF FID, SyncNet, CSIM, SSIM 4.5
S39 Gan et al. (2023) [98] Efficient Emotional Adaptation for Audio-Driven Talking-Head Generation ICCV 2023 Transformer MEAD, VoxCeleb2 FID, AUE, SyncNet 4.2
S40 Fan et al. (2024) [99] UniTalker: Scaling Up Audio-Driven 3D Facial Animation Through a Unified Model ECCV 2024 Transformer VOCASET, BIWI, MEAD LVE, MOS, FDD 4.5
S41 Liu et al. (2024) [100] CustomListener: Text-Guided Responsive Interaction for Listening Head Generation CVPR 2024 Transformer VoxCeleb2, ViCo FID, SyncNet, MOS 4.0
S42 Peng et al. (2023) [101] EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face Animation ICCV 2023 Transformer MEAD, RAVDESS FID, AUE, MOS 4.2
S43 Wang et al. (2023) [102] LipFormer: High-Fidelity and Generalizable Talking Face Generation with a Pre-Learned Facial Codebook CVPR 2023 Transformer VoxCeleb2, HDTF, LRS3 FID, SyncNet, SSIM, PSNR, LPIPS, MOS 4.5
S44 Liu et al. (2023) [103] MODA: Mapping-Once Audio-Driven Portrait Animation with Dual Attentions ICCV 2023 Transformer VoxCeleb2, HDTF FID, CPBD, LMD, SyncNet 4.5
S45 Jiang et al. (2025) [104] LOOPY: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency ICLR 2025 Transformer VoxCeleb2, HDTF FID, SyncNet, FVD 3.8
S46 Liu et al. (2024) [105] AniTalker: Animate Vivid and Diverse Talking Faces via Identity-Decoupled Motion ACM MM 2024 Transformer VoxCeleb2, MEAD FID, SyncNet, CSIM 4.2
S47 Hogue et al. (2023) [106] DiffTED: One-Shot Audio-Driven TED Talk Video Generation CVPR 2024 Transformer TED-Talks FID, MOS, SyncNet 3.5
NeRF-based (n = 11)
S48 Guo et al. (2021) [9] AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis ICCV 2021 NeRF Custom (Obama) PSNR, SSIM, LMD 5.0
S49 Ye et al. (2023) [41] GeneFace: Generalised and High-Fidelity Audio-Driven 3D Talking Face ICLR 2023 NeRF VoxCeleb2, HDTF PSNR, SSIM, LMD, MOS 4.8
S50 Tang et al. (2025) [57] RAD-NeRF: Real-Time Neural Radiance Fields for Dynamic Head Portraits IJCV 2025 NeRF Custom PSNR, SSIM, FPS 4.0
S51 Li et al. (2023) [107] ER-NeRF: Efficient Region-Aware NeRF for High-Fidelity Talking Portrait ICCV 2023 NeRF Custom (multi-spk) PSNR, SSIM, LPIPS, FPS 4.5
S52 Ye et al. (2023) [108] GeneFace++: Generalised Real-Time Audio-Driven 3D Talking Face arXiv 2023† NeRF VoxCeleb2, HDTF PSNR, SSIM, LMD, FPS 4.8
S53 Ye et al. (2024) [109] Real3D-Portrait: One-Shot Realistic 3D Talking Portrait Synthesis arXiv 2024† NeRF HDTF, VoxCeleb2 FID, PSNR, SSIM, CSIM 4.5
S54 Shen et al. (2022) [110] Learning Dynamic Facial Radiance Fields for Few-Shot Talking Head Synthesis ECCV 2022 NeRF VoxCeleb2 PSNR, SSIM, FID 4.2
S55 Wang et al. (2024) [111] Expression-Aware Neural Radiance Fields for High-Fidelity Talking Portrait IVC 2024 NeRF Custom PSNR, SSIM, LPIPS 3.8
S56 Ma et al. (2025) [112] Decoupled Two-Stage Talking Head Generation via Gaussian-Landmark NeRF Comput. Vis. Media 2025 NeRF VoxCeleb2, HDTF PSNR, SSIM, SyncNet 4.0
S57 Deng et al. (2024) [113] Portrait4D: Learning One-Shot 4D Head Avatar via Video-Based Supervision CVPR 2024 NeRF VoxCeleb2, CelebV-HQ FID, CSIM, LPIPS 4.5
S58 Zhang et al. (2023) [114] Metaportrait: Identity-preserving talking head generation with fast personalized adaptation CVPR 2023 NeRF VoxCeleb2 FID, PSNR, SSIM 4.0
Diffusion Model (n = 16)
S59 Shen et al. (2023) [11] DiffTalk: Crafting Diffusion Models for Generalised Audio-Driven Animation CVPR 2023 Diffusion VoxCeleb2, HDTF FID, SyncNet, SSIM 4.8
S60 Ma et al. (2023) [12] DreamTalk: Expressive Talking Head with Diffusion Probabilistic Models arXiv 2023† Diffusion MEAD, HDTF FID, LPIPS, MOS 4.2
S61 Xu et al. (2024) [13] VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time NeurIPS 2024 Diffusion VoxCeleb2, HDTF FID, SyncNet, CSIM, FPS 5.0
S62 Xu et al. (2024) [58] Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Animation arXiv 2024† Diffusion VoxCeleb2, HDTF FID, SyncNet, FVD, CSIM 4.5
S63 Zhang et al. (2024) [19] MuseTalk: Real-Time High Quality Lip Synchronisation with Latent Space Inpainting arXiv 2024† Diffusion VoxCeleb2, HDTF SyncNet, PSNR, FPS 4.2
S64 Stypulkowski et al. (2024) [115] Diffused Heads: Diffusion Models Beat GANs on Talking-Face Generation WACV 2024 Diffusion VoxCeleb2 FID, SyncNet, SSIM 4.5
S65 Sun et al. (2024) [116] DiffPoseTalk: Speech-Driven Stylistic 3D Facial Animation via Diffusion ACM TOG 2024 Diffusion HDTF, MEAD FID, SSIM, LMD, MOS 4.5
S66 Bigioi et al. (2024) [117] Speech Driven Video Editing via an Audio-Conditioned Diffusion Model IVC 2024 Diffusion VoxCeleb2, LRS3 SyncNet, FID, SSIM 4.0
S67 Zhang et al. (2023) [40] SadTalker: Learning Realistic 3D Motion for Stylised Audio-Driven Animation CVPR 2023 Diffusion VoxCeleb2, HDTF FID, SyncNet, LMD, SSIM 4.8
S68 Tian et al. (2024) [118] EMO: Emote Portrait Alive — Generating Expressive Portrait Videos ECCV 2024 Diffusion VoxCeleb2, HDTF FID, FVD, SyncNet, MOS 4.2
S69 Cui et al. (2025) [119] Hallo2: Long-Duration High-Resolution Audio-Driven Portrait Animation ICLR 2025 Diffusion VoxCeleb2, HDTF FID, FVD, SyncNet, CSIM 4.0
S70 Wei et al. (2024) [120] AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation arXiv 2024† Diffusion VoxCeleb2 FID, SyncNet, SSIM, LPIPS 4.0
S71 Zheng et al. (2024) [121] MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation arXiv 2024† Diffusion VoxCeleb2, HDTF FID, FVD, SyncNet, MOS 4.2
S72 Ma et al. (2024) [122] FollowYourEmoji: Fine-Controllable and Expressive Freestyle Portrait Animation SIGGRAPH 2024 Diffusion VoxCeleb2, MEAD FID, SyncNet, AUE 4.0
S73 Wang et al. (2023) [65] EAT: Emotional Talking Head Based on Memory-Sharing and Attention Mechanism arXiv 2023† Diffusion MEAD, CREMA-D FID, AUE, MOS 4.2
S74 Cheng et al. (2025) [123] Dynamic Frame Avatar with Non-Autoregressive Diffusion for Talking Head ICLR 2025 Diffusion VoxCeleb2, HDTF FVD, SyncNet, FID, SSIM 4.5
Gaussian Splatting (n = 8)
S75 Chen et al. (2024) [50] GSTalker: Real-Time Audio-Driven Talking Face via Deformable Gaussian Splatting arXiv 2024† GS VoxCeleb2, HDTF PSNR, SSIM, FPS, SyncNet 4.2
S76 Yang et al. (2025) [51] GaussianEmoTalker: Real-Time Emotional Talking Head Synthesis with Audio-Driven 3D GS arXiv 2025† GS MEAD, VoxCeleb2 FID, AUE, FPS, MOS 4.0
S77 Xu et al. (2024) [52] VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single Image arXiv 2025† GS VoxCeleb2 FID, SyncNet, FPS, CSIM 4.5
S78 Xu et al. (2024) [124] Gaussian Head Avatar: Ultra High-Fidelity Head Avatar via Dynamic Gaussians CVPR 2024 GS Multi-View Custom PSNR, SSIM, LPIPS, FPS 4.8
S79 Qian et al. (2024) [125] GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians CVPR 2024 GS FLAME-fitted video PSNR, SSIM, FPS 4.5
S80 Chen et al. (2024) [126] MonoGaussianAvatar: Monocular Gaussian Point-Based Head Avatar SIGGRAPH 2024 GS Custom monocular PSNR, SSIM, LPIPS 4.5
S81 Deng et al. (2024) [127] Portrait4D-v2: Pseudo Multi-View Data for 4D Gaussian Head Animation ECCV 2024 GS VoxCeleb2, CelebV-HQ FID, CSIM, FPS 4.0
S82 Cho et al. (2024) [128] Real-Time High-Fidelity Audio-Driven Talking Face via Gaussian Splatting ACM MM 2024 GS VoxCeleb2, HDTF PSNR, SSIM, LPIPS, SyncNet, FPS 4.5
Diffusion Transformer / DiT (n = 7)
S83 Lin et al. (2025) [14] OmniHuman-1: Scaling One-Stage Conditioned Human Animation via DiT ICCV 2025 DiT VoxCeleb2, HDTF, Custom FID, FVD, SyncNet, MOS 5.0
S84 Wang et al. (2025) [20] EmotiveTalk: Expressive Talking Head via Audio Decoupling and Emotional Video Diffusion CVPR 2025 DiT MEAD, VoxCeleb2 FID, AUE, SyncNet, MOS 5.0
S85 Hong et al. (2025) [59] Audio-Visual Controlled Video Diffusion with Masked SSMs for Talking Head ICCV 2025 DiT VoxCeleb2, HDTF FID, SyncNet, FVD, CSIM 4.5
S86 Zheng et al. (2025) [60] Teller: Real-Time Streaming Talking Portrait via Autoregressive Motion Generation CVPR 2025 DiT VoxCeleb2 SyncNet, FVD, FPS 4.5
S87 Li et al. (2025) [129] Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis ACM MM 2025 DiT VoxCeleb2, HDTF FID, SyncNet, SSIM, FVD 4.5
S88 Tian et al. (2025) [71] EMO2: End-to-End Audio-Driven Expressive Humanoid Talking Head Generation arXiv 2025† DiT VoxCeleb2, HDTF FID, FVD, SyncNet, MOS 4.2
S89 Zhu et al. (2025) [23] INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations CVPR 2025 DiT VoxCeleb2, ViCo FID, SyncNet, FVD, MOS 4.2
Hybrid (n = 3)
S90 Li et al. (2024) [61] Control-Talker: A Rapid-Customization Talking Head Generation Method for Multi-Condition Control and High-Texture Enhancement. ACM MM 2024 Hybrid VoxCeleb2, MEAD FID, SyncNet, MOS, CSIM 4.0
S91 Ye et al. (2024) [62] MimicTalk: Mimicking a Personalized and Expressive 3D Talking Face in Minutes NeurIPS 2024 Hybrid VoxCeleb2, HDTF PSNR, SSIM, LPIPS, SyncNet, AED, MOS 4.5
S92 Li et al. (2024) [63] TalkingGaussian: Structure-Persistent 3D Talking Head Synthesis via Gaussian Splatting ECCV 2024 Hybrid VoxCeleb2, HDTF PSNR, SSIM, LPIPS, SyncNet, FPS 4.5

References

  1. Bregler, C.; Covell, M.; Slaney, M. Video rewrite: Driving visual speech with audio. Semin. Graph. Pap. Push. Bound. Vol. 2 2023, ed, 715–722. [Google Scholar] [CrossRef]
  2. Suwajanakorn, S.; Seitz, S. M.; Kemelmacher-Shlizerman, I. Synthesizing obama: learning lip sync from audio. ACM Trans. Graph. (ToG) 2017, vol. 36, 1–13. [Google Scholar]
  3. Goodfellow; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; et al. Generative adversarial networks. Commun. ACM 2020, vol. 63, 139–144. [Google Scholar] [CrossRef]
  4. Prajwal, K.; Mukhopadhyay, R.; Namboodiri, V. P.; Jawahar, C. A lip sync expert is all you need for speech to lip generation in the wild. In Proceedings of the 28th ACM international conference on multimedia, 2020; pp. 484–492. [Google Scholar]
  5. Son Chung, J.; Senior, A.; Vinyals, O.; Zisserman, A. Lip reading sentences in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2017; pp. 6447–6456. [Google Scholar]
  6. Zhang, Z.; Li, L.; Ding, Y.; Fan, C. Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021; pp. 3661–3670. [Google Scholar]
  7. Fan, Y.; Lin, Z.; Saito, J.; Wang, W.; Komura, T. Faceformer: Speech-driven 3d facial animation with transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022; pp. 18770–18780. [Google Scholar]
  8. Mildenhall, B.; Srinivasan, P. P.; Tancik, M.; Barron, J. T.; Ramamoorthi, R.; Ng, R. Nerf: Representing scenes as neural radiance fields for view synthesis. Commun. ACM 2021, vol. 65, 99–106. [Google Scholar]
  9. Guo, Y.; Chen, K.; Liang, S.; Liu, Y.-J.; Bao, H.; Zhang, J. Ad-nerf: Audio driven neural radiance fields for talking head synthesis. In Proceedings of the IEEE/CVF international conference on computer vision, 2021; pp. 5784–5794. [Google Scholar]
  10. Ho, J.; Jain, A.; Abbeel, P. Denoising diffusion probabilistic models. Adv. Neural Inf. Process. Syst. 2020, vol. 33, 6840–6851. [Google Scholar]
  11. Shen, S.; Zhao, W.; Meng, Z.; Li, W.; Zhu, Z.; Zhou, J.; et al. Difftalk: Crafting diffusion models for generalized audio-driven portraits animation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023; pp. 1982–1991. [Google Scholar]
  12. Ma, Y.; Zhang, S.; Wang, J.; Wang, X.; Zhang, Y.; Deng, Z. Dreamtalk: When emotional talking head generation meets diffusion probabilistic models. arXiv 2023, arXiv:2312.09767. [Google Scholar]
  13. Xu, S.; Chen, G.; Guo, Y.-X.; Yang, J.; Li, C.; Zang, Z.; et al. Vasa-1: Lifelike audio-driven talking faces generated in real time. Adv. Neural Inf. Process. Syst. 2024, vol. 37, 660–684. [Google Scholar] [CrossRef]
  14. Lin, G.; Jiang, J.; Yang, J.; Zheng, Z.; Liang, C.; Zhang, Y.; et al. Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025; pp. 13847–13858. [Google Scholar]
  15. Trist, E. L.; Bamforth, K. W. Some social and psychological consequences of the longwall method of coal-getting: An examination of the psychological situation and defences of a work group in relation to the social structure and technological content of the work system. Hum. Relat. 1951, vol. 4, 3–38. [Google Scholar]
  16. Bijker, W.; Hughes, T.; Pinch, T. The social construction of technological systems MIT Press; Cambridge MA, London, 1987. [Google Scholar]
  17. WHO Health Workforce Team. Global Health Workforce Shortage Projections. 2024. [Google Scholar] [CrossRef] [PubMed]
  18. Haider, S. A.; Prabha, S.; Gomez-Cabello, C. A.; Genovese, A.; Collaco, B.; Wood, N.; et al. Artificial Intelligence Physician Avatars for Patient Education: A Pilot Study. J. Clin. Med. vol. 14, 8595, 2025. [CrossRef] [PubMed]
  19. Zhang, Y.; Minhao, L.; Chen, Z.; Wu, B.; Zhan, C.; He, Y.; et al. Musetalk: Real-time high quality lip synchronization with latent space inpainting. 2024. [Google Scholar] [PubMed]
  20. Wang, H.; Weng, Y.; Li, Y.; Guo, Z.; Du, J.; Niu, S.; et al. Emotivetalk: Expressive talking head generation through audio information decoupling and emotional video diffusion. In Proceedings of the Computer Vision and Pattern Recognition Conference, 2025; pp. 26212–26221. [Google Scholar]
  21. Provoost, S.; Lau, H. M.; Ruwaard, J.; Riper, H. Embodied conversational agents in clinical psychology: a scoping review. J. Med. Internet Res. 2017, vol. 19, e151. [Google Scholar] [CrossRef] [PubMed]
  22. Chheang, V.; Sharmin, S.; Márquez-Hernández, R.; Patel, M.; Rajasekaran, D.; Caulfield, G.; et al. Towards anatomy education with generative AI-based virtual assistants in immersive virtual reality environments. 2024 IEEE Int. Conf. Artif. Intell. Ext. Virtual Real. (AIxVR) 2024, 21–30. [Google Scholar] [CrossRef]
  23. Zhu, Y.; Zhang, L.; Rong, Z.; Hu, T.; Liang, S.; Ge, Z. INFP: Audio-driven interactive head generation in dyadic conversations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025; pp. 10667–10677. [Google Scholar]
  24. Salahudeen, R.; Siu, W.-C.; Chan, H. A. Photo-realistic talking face generation under latent space manipulation. IEEE Transactions on Consumer Electronics, 2024. [Google Scholar]
  25. Salahudeen, R.; Siu, W.-C.; Chan, H. A. Activate your face in virtual meeting platform. 2023 International Conference on Consumer Electronics-Taiwan (ICCE-Taiwan), 2023; pp. 791–792. [Google Scholar]
  26. Deonarine, K. Deepfakes and human rights: Why the EU AI Act is becoming the global standard for ethical AI regulation. 2025. [Google Scholar] [PubMed]
  27. United States Congress. TAKE IT DOWN Act," ed. 2025. [Google Scholar] [CrossRef]
  28. European Commission. Code of Practice on Transparency of AI-Generated Content. EU AI OfficeJune, 10 2026. [Google Scholar]
  29. Rana, M. S.; Nobi, M. N.; Murali, B.; Sung, A. H. Deepfake detection: A systematic literature review. IEEE Access 2022, vol. 10, 25494–25513. [Google Scholar] [CrossRef]
  30. Zhen, R.; Song, W.; He, Q.; Cao, J.; Shi, L.; Luo, J. Human-computer interaction system: A survey of talking-head generation. Electronics 2023, vol. 12, 218. [Google Scholar]
  31. Toshpulatov, M.; Lee, W.; Lee, S. Talking human face generation: A survey. Expert Syst. With Appl. 2023, vol. 219, 119678. [Google Scholar]
  32. Bigioi, D.; Corcoran, P. Multilingual video dubbing—a technology review and current challenges. Front. Signal Process. 2023, vol. 3, 1230755. [Google Scholar] [CrossRef]
  33. Heidari; Jafari Navimipour, N.; Dag, H.; Unal, M. Deepfake detection using deep learning methods: A systematic and comprehensive review. Wiley Interdiscip. Rev. Data Min. Knowl. Discov. 2024, vol. 14, e1520. [Google Scholar]
  34. Rakesh, V. K.; Mazumdar, S.; Maity, R. P.; Pal, S.; Das, A.; Samanta, T. Advancements in talking head generation: a comprehensive review of techniques, metrics, and challenges: Advancements in talking head generation. Vis. Comput. 2025, vol. 42. [Google Scholar] [CrossRef]
  35. Floridi, L.; Cowls, J.; Beltrametti, M.; Chatila, R.; Chazerand, P.; Dignum, V.; et al. AI4People—An ethical framework for a good AI society: Opportunities, risks, principles, and recommendations. Minds Mach. 2018, vol. 28, 689–707. [Google Scholar] [CrossRef] [PubMed]
  36. Kitchenham, B.; Charters, S. Guidelines for performing systematic literature reviews in software engineering. 2007. [Google Scholar] [CrossRef]
  37. Romero Moreno, F. Generative AI and deepfakes: a human rights approach to tackling harmful content. Int. Rev. Law. Comput. Technol. 2024, vol. 38, 297–326. [Google Scholar] [CrossRef]
  38. Nisar, H.; Masood, S.; Malik, Z.; Abid, A. Talking Head Generation Through Generative Models and Cross-Modal Synthesis Techniques. J. Imaging 2026, vol. 12, 119. [Google Scholar] [CrossRef] [PubMed]
  39. Chung, J. S.; Zisserman, A. Out of time: automated lip sync in the wild. Asian conference on computer vision, 2016; pp. 251–263. [Google Scholar]
  40. Zhang, W.; Cun, X.; Wang, X.; Zhang, Y.; Shen, X.; Guo, Y.; et al. Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023; pp. 8652–8661. [Google Scholar]
  41. Ye, Z.; Jiang, Z.; Ren, Y.; Liu, J.; He, J.; Zhao, Z. Geneface: Generalized and high-fidelity audio-driven 3d talking face synthesis. arXiv 2023, arXiv:2301.13430. [Google Scholar]
  42. Bickmore, T.; Cassell, J. Relational agents: a model and implementation of building user trust. In Proceedings of the SIGCHI conference on Human factors in computing systems, 2001; pp. 396–403. [Google Scholar]
  43. Page, M. J.; McKenzie, J. E.; Bossuyt, P. M.; Boutron, I.; Hoffmann, T. C.; Mulrow, C. D.; et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ 2021, vol. 372. [Google Scholar] [CrossRef] [PubMed]
  44. Popay, J.; Roberts, H.; Sowden, A.; Petticrew, M.; Arai, L.; Rodgers, M.; et al. Guidance on the conduct of narrative synthesis in systematic reviews. A Product. From ESRC Methods Programme Version 2006, vol. 1, b92. [Google Scholar]
  45. Fan, B.; Xie, L.; Yang, S.; Wang, L.; Soong, F. K. A deep bidirectional LSTM approach for video-realistic talking head. Multimed. Tools Appl. 2016, vol. 75, 5287–5309. [Google Scholar] [CrossRef]
  46. Zhou, H.; Liu, Y.; Liu, Z.; Luo, P.; Wang, X. Talking face generation by adversarially disentangled audio-visual representation. In Proceedings of the AAAI conference on artificial intelligence, 2019; pp. 9299–9306. [Google Scholar]
  47. Zhou, Y.; Han, X.; Shechtman, E.; Echevarria, J.; Kalogerakis, E.; Li, D. Makelttalk: speaker-aware talking-head animation. ACM Trans. Graph. (TOG) 2020, vol. 39, 1–15. [Google Scholar]
  48. Zhou, H.; Sun, Y.; Wu, W.; Loy, C. C.; Wang, X.; Liu, Z. Pose-controllable talking face generation by implicitly modularized audio-visual representation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021; pp. 4176–4186. [Google Scholar]
  49. Peebles, W.; Xie, S. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF international conference on computer vision, 2023; pp. 4195–4205. [Google Scholar]
  50. Chen, B.; Hu, S.; Chen, Q.; Du, C.; Yi, R.; Qian, Y.; et al. Gstalker: Real-time audio-driven talking face generation via deformable gaussian splatting. arXiv 2024, arXiv:2404.19040. [Google Scholar]
  51. Yang, H.; Zhang, Z.; Dong, Y.; Qian, J.; Yang, J. GaussianEmoTalker: Real-Time Emotional Talking Head Synthesis with Audio-Driven and Blendshape-Based 3D Gaussian Splatting. arXiv 2026, arXiv:2607.00959. [Google Scholar]
  52. Xu, S.; Chen, G.; Yang, J.; Zhang, Y.; Deng, Y.; Lin, S.; et al. Vasa-3d: Lifelike audio-driven gaussian head avatars from a single image. Adv. Neural Inf. Process. Syst. 2026, vol. 38, 150764–150786. [Google Scholar]
  53. Song, Y.; Zhu, J.; Li, D.; Wang, A.; Qi, H. Talking face generation by conditional recurrent adversarial network. In Proceedings of the 28th International Joint Conference on Artificial Intelligence, 2019; pp. 919–925. [Google Scholar]
  54. Zhou, Y.; Xu, Z.; Landreth, C.; Kalogerakis, E.; Maji, S.; Singh, K. Visemenet: Audio-driven animator-centric speech animation. ACM Trans. Graph. (ToG) 2018, vol. 37, 1–10. [Google Scholar]
  55. Peng, Z.; Hu, W.; Shi, Y.; Zhu, X.; Zhang, X.; Zhao, H.; et al. Synctalk: The devil is in the synchronization for talking head synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 666–676. [Google Scholar]
  56. Xing, J.; Xia, M.; Zhang, Y.; Cun, X.; Wang, J.; Wong, T.-T. Codetalker: Speech-driven 3d facial animation with discrete motion prior. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023; pp. 12780–12790. [Google Scholar]
  57. Tang, J.; Wang, K.; Zhou, H.; Chen, X.; He, D.; Hu, T.; et al. Real-time neural radiance talking portrait synthesis via audio-spatial decomposition. Int. J. Comput. Vis. 2025, vol. 133, 6362–6373. [Google Scholar] [CrossRef]
  58. Xu, M.; Li, H.; Su, Q.; Shang, H.; Zhang, L.; Liu, C.; et al. Hallo: Hierarchical audio-driven visual synthesis for portrait image animation. arXiv 2024, arXiv:2406.08801. [Google Scholar]
  59. Hong, F.-T.; Xu, Z.; Zhou, Z.; Zhou, J.; Li, X.; Lin, Q.; et al. Audio-visual controlled video diffusion with masked selective state spaces modeling for natural talking head generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025; pp. 12549–12558. [Google Scholar]
  60. Zhen, D.; Yin, S.; Qin, S.; Yi, H.; Zhang, Z.; Liu, S.; et al. Teller: Real-time streaming audio-driven portrait animation with autoregressive motion generation. In Proceedings of the Computer Vision and Pattern Recognition Conference, 2025; pp. 21075–21085. [Google Scholar]
  61. Li, Y.; Yu, L.; Wang, L.; Xie, H. Control-Talker: A Rapid-Customization Talking Head Generation Method for Multi-Condition Control and High-Texture Enhancement. In Proceedings of the 32nd ACM International Conference on Multimedia, 2024; pp. 3519–3527. [Google Scholar]
  62. Ye, Z.; Zhong, T.; Ren, Y.; Jiang, Z.; Huang, J.; Huang, R.; et al. Mimictalk: Mimicking a personalized and expressive 3d talking face in minutes. Adv. Neural Inf. Process. Syst. 2024, vol. 37, 1829–1853. [Google Scholar] [CrossRef]
  63. Li, J.; Zhang, J.; Bai, X.; Zheng, J.; Ning, X.; Zhou, J.; et al. Talkinggaussian: Structure-persistent 3d talking head synthesis via gaussian splatting. European Conference on Computer Vision, 2024; pp. 127–145. [Google Scholar]
  64. Buolamwini, J.; Gebru, T. Gender shades: Intersectional accuracy disparities in commercial gender classification. Conference on fairness, accountability and transparency, 2018; pp. 77–91. [Google Scholar]
  65. Wang, J.; Zhao, Y.; Liu, L.; Xu, T.; Li, Q.; Li, S. Emotional talking head generation based on memory-sharing and attention-augmented networks. arXiv 2023, arXiv:2306.03594. [Google Scholar]
  66. Wang, Z.; Bovik, A. C.; Sheikh, H. R.; Simoncelli, E. P. Image quality assessment: from error visibility to structural similarity. IEEE Trans. Image Process. 2004, vol. 13, 600–612. [Google Scholar] [CrossRef] [PubMed]
  67. DeTore, N. R.; Balogun-Mwangi, O.; Eberlin, E. S.; Dokholyan, K. N.; Rizzo, A.; Holt, D. J. An artificial intelligence-based virtual human avatar application to assess the mental health of health care professionals: A validation study. J. Med. Ext. Real. 2024, vol. 1, jmxr. 2024.0016. [Google Scholar]
  68. Venkatesh, S.; Ramachandra, R.; Raja, K.; Busch, C. Face morphing attack generation and detection: A comprehensive survey. IEEE Trans. Technol. Soc. 2021, vol. 2, 128–145. [Google Scholar] [CrossRef]
  69. iProov. The iProov Threat Intelligence Report 2024: The Impact of Generative AI on Remote Identity Verification. 2024. [Google Scholar]
  70. Walker, Jones. Deepfakes-as-a-Service meets state laws: Governing synthetic media in a fragmented legal landscape. 2026. [Google Scholar]
  71. Tian, L.; Hu, S.; Wang, Q.; Zhang, B.; Bo, L. Emo2: End-effector guided audio-driven avatar video generation. arXiv 2025, arXiv:2501.10687. [Google Scholar]
  72. Beauchamp, T. L.; Childress, J. F. Principles of biomedical ethics; Edicoes Loyola, 1994. [Google Scholar]
  73. Grother, P.; Salamon, W.; Chandramouli, R. NIST Special Publication 800-76-2 Biometric Specifications for Personal Identity Verification. National Institute of Standards and Technology, Tech. Rep. 800-76-2, 2013.
  74. Vougioukas, K.; Petridis, S.; Pantic, M. Realistic speech-driven facial animation with gans. Int. J. Comput. Vis. 2020, vol. 128, 1398–1413. [Google Scholar] [CrossRef]
  75. Chen, L.; Maddox, R. K.; Duan, Z.; Xu, C. Hierarchical cross-modal talking face generation with dynamic pixel-wise loss. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019; pp. 7832–7841. [Google Scholar]
  76. Siarohin; Lathuilière, S.; Tulyakov, S.; Ricci, E.; Sebe, N. First order motion model for image animation. Adv. Neural Inf. Process. Syst. 2019, vol. 32. [Google Scholar]
  77. Jamaludin; Chung, J. S.; Zisserman, A. You said that?: Synthesising talking faces from audio. Int. J. Comput. Vis. 2019, vol. 127, 1767–1779. [Google Scholar] [CrossRef]
  78. Wiles; Koepke, A.; Zisserman, A. X2face: A network for controlling face generation using images, audio, and pose codes. In Proceedings of the European conference on computer vision (ECCV), 2018; pp. 670–686. [Google Scholar]
  79. Wang, T.-C.; Mallya, A.; Liu, M.-Y. One-shot free-view neural talking-head synthesis for video conferencing. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021; pp. 10039–10049. [Google Scholar]
  80. Cheng, K.; Cun, X.; Zhang, Y.; Xia, M.; Yin, F.; Zhu, M.; et al. Videoretalking: Audio-based lip synchronization for talking head video editing in the wild. SIGGRAPH Asia 2022 Conference Papers, 2022; pp. 1–9. [Google Scholar]
  81. Das, D.; Biswas, S.; Sinha, S.; Bhowmick, B. Speech-driven facial animation using cascaded gans for learning of motion and texture. European conference on computer vision, 2020; pp. 408–424. [Google Scholar]
  82. Ji, X.; Zhou, H.; Wang, K.; Wu, W.; Loy, C. C.; Cao, X.; et al. Audio-driven emotional video portraits. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021; pp. 14080–14089. [Google Scholar]
  83. Wang, S.; Li, L.; Ding, Y.; Fan, C.; Yu, X. Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion. In Proceedings of the Thirtieth International Joint Conference On Artificial Intelligence, Ijcai 2021; 2021, pp. 1098–1105.
  84. Hong, F.-T.; Zhang, L.; Shen, L.; Xu, D. Depth-aware generative adversarial network for talking head video generation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022; pp. 3397–3406. [Google Scholar]
  85. Yin, F.; Zhang, Y.; Cun, X.; Cao, M.; Fan, Y.; Wang, X.; et al. Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan. European conference on computer vision, 2022; pp. 85–101. [Google Scholar]
  86. Liang; Pan, Y.; Guo, Z.; Zhou, H.; Hong, Z.; Han, X.; et al. Expressive talking head generation with granular audio-visual control. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022; pp. 3387–3396. [Google Scholar]
  87. Song, H.-K.; Woo, S. H.; Lee, J.; Yang, S.; Cho, H.; Lee, Y.; et al. Talking face generation with multilingual tts. In Proceedings of the ieee/cvf conference on computer vision and pattern recognition, 2022; pp. 21425–21430. [Google Scholar]
  88. Lu, Y.; Chai, J.; Cao, X. Live speech portraits: real-time photorealistic talking-head animation. ACM Trans. Graph. (ToG) 2021, vol. 40, 1–17. [Google Scholar]
  89. Zhu, H.; Huang, H.; Li, Y.; Zheng, A.; He, R. Arbitrary talking face generation via attentional audio-visual coherence learning. arXiv 2018, arXiv:1812.06589. [Google Scholar]
  90. Wen, X.; Wang, M.; Richardt, C.; Chen, Z.-Y.; Hu, S.-M. Photorealistic audio-driven video portraits. IEEE Trans. Vis. Comput. Graph. 2020, vol. 26, 3457–3466. [Google Scholar] [CrossRef] [PubMed]
  91. Hwang, G.; Hong, S.; Lee, S.; Park, S.; Chae, G. Discohead: Audio-and-video-driven talking head generation by disentangled control of head pose and facial expressions. ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023; pp. 1–5. [Google Scholar]
  92. Ji, X.; Zhou, H.; Wang, K.; Wu, Q.; Wu, W.; Xu, F.; et al. Eamm: One-shot emotional talking face via audio-based emotion-aware motion model. ACM SIGGRAPH 2022 conference proceedings, 2022; pp. 1–10. [Google Scholar]
  93. Yi, R.; Ye, Z.; Sun, Z.; Zhang, J.; Zhang, G.; Wan, P.; et al. Predicting personalized head movement from short video and speech signal. IEEE Trans. Multimed. 2022, vol. 25, 6315–6328. [Google Scholar] [CrossRef]
  94. Wang, S.; Li, L.; Ding, Y.; Yu, X. One-shot talking face generation from single-speaker audio-visual correlation learning. In Proceedings of the AAAI Conference on Artificial Intelligence, 2022; pp. 2531–2539. [Google Scholar]
  95. Peng, Z.; Luo, Y.; Shi, Y.; Xu, H.; Zhu, X.; Liu, H.; et al. Selftalk: A self-supervised commutative training diagram to comprehend 3d talking faces. In Proceedings of the 31st ACM International Conference on Multimedia, 2023; pp. 5292–5301. [Google Scholar]
  96. Chen, L.; Cui, G.; Liu, C.; Li, Z.; Kou, Z.; Xu, Y.; et al. Talking-head generation with rhythmic head motion. European conference on computer vision, 2020; pp. 35–51. [Google Scholar]
  97. Zhong, W.; Fang, C.; Cai, Y.; Wei, P.; Zhao, G.; Lin, L.; et al. Identity-preserving talking face generation with landmark and appearance priors. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023; pp. 9729–9738. [Google Scholar]
  98. Gan, Y.; Yang, Z.; Yue, X.; Sun, L.; Yang, Y. Efficient emotional adaptation for audio-driven talking-head generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023; pp. 22634–22645. [Google Scholar]
  99. Fan, X.; Li, J.; Lin, Z.; Xiao, W.; Yang, L. Unitalker: Scaling up audio-driven 3d facial animation through a unified model. European Conference on Computer Vision, 2024; pp. 204–221. [Google Scholar]
  100. Liu, X.; Guo, Y.; Zhen, C.; Li, T.; Ao, Y.; Yan, P. Customlistener: Text-guided responsive interaction for user-friendly listening head generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 2415–2424. [Google Scholar]
  101. Peng, Z.; Wu, H.; Song, Z.; Xu, H.; Zhu, X.; He, J.; et al. Emotalk: Speech-driven emotional disentanglement for 3d face animation. In Proceedings of the IEEE/CVF international conference on computer vision, 2023; pp. 20687–20697. [Google Scholar]
  102. Wang, J.; Zhao, K.; Zhang, S.; Zhang, Y.; Shen, Y.; Zhao, D.; et al. Lipformer: High-fidelity and generalizable talking face generation with a pre-learned facial codebook. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023; pp. 13844–13853. [Google Scholar]
  103. Liu, Y.; Lin, L.; Yu, F.; Zhou, C.; Li, Y. Moda: Mapping-once audio-driven portrait animation with dual attentions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023; pp. 23020–23029. [Google Scholar]
  104. Jiang, J.; Liang, C.; Yang, J.; Lin, G.; Zhong, T.; Zheng, Y. Loopy: Taming audio-driven portrait avatar with long-term motion dependency. International Conference on Learning Representations, 2025; pp. 15245–15263. [Google Scholar]
  105. Liu, T.; Chen, F.; Fan, S.; Du, C.; Chen, Q.; Chen, X.; et al. Anitalker: Animate vivid and diverse talking faces through identity-decoupled facial motion encoding. In Proceedings of the 32nd ACM International Conference on Multimedia, 2024; pp. 6696–6705. [Google Scholar]
  106. Hogue, S.; Zhang, C.; Daruger, H.; Tian, Y.; Guo, X. Diffted: One-shot audio-driven ted talk video generation with diffusion-based co-speech gestures. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 1922–1931. [Google Scholar]
  107. Li, J.; Zhang, J.; Bai, X.; Zhou, J.; Gu, L. Efficient region-aware neural radiance fields for high-fidelity talking portrait synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023; pp. 7568–7578. [Google Scholar]
  108. Ye, Z.; He, J.; Jiang, Z.; Huang, R.; Huang, J.; Liu, J.; et al. Geneface++: Generalized and stable real-time audio-driven 3d talking face generation. arXiv 2023, arXiv:2305.00787. [Google Scholar]
  109. Ye, Z.; Zhong, T.; Ren, Y.; Yang, J.; Li, W.; Huang, J.; et al. Real3d-portrait: One-shot realistic 3d talking portrait synthesis. arXiv 2024, arXiv:2401.08503. [Google Scholar]
  110. Shen, S.; Li, W.; Zhu, Z.; Duan, Y.; Zhou, J.; Lu, J. Learning dynamic facial radiance fields for few-shot talking head synthesis. European conference on computer vision, 2022; pp. 666–682. [Google Scholar]
  111. Wang, X.; Ruan, T.; Xu, J.; Guo, X.; Li, J.; Yan, F.; et al. Expression-aware neural radiance fields for high-fidelity talking portrait synthesis. Image Vis. Comput. 2024, vol. 147, 105075. [Google Scholar] [CrossRef]
  112. Ma; Cao, Y.; Zhang, L. Decoupled two-stage talking head generation via Gaussian-landmark-based neural radiance fields. Computational Visual Media, 2025. [Google Scholar]
  113. Deng, Y.; Wang, D.; Ren, X.; Chen, X.; Wang, B. Portrait4d: Learning one-shot 4d head avatar synthesis using synthetic data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 7119–7130. [Google Scholar]
  114. Zhang, B.; Qi, C.; Zhang, P.; Zhang, B.; Wu, H.; Chen, D.; et al. Metaportrait: Identity-preserving talking head generation with fast personalized adaptation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023; pp. 22096–22105. [Google Scholar]
  115. Stypułkowski, M.; Vougioukas, K.; He, S.; Zięba, M.; Petridis, S.; Pantic, M. Diffused heads: Diffusion models beat gans on talking-face generation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024; pp. 5091–5100. [Google Scholar]
  116. Sun, Z.; Lv, T.; Ye, S.; Lin, M.; Sheng, J.; Wen, Y.-H.; et al. Diffposetalk: Speech-driven stylistic 3d facial animation and head pose generation via diffusion models. ACM Trans. Graph. (ToG) 2024, vol. 43, 1–9. [Google Scholar] [CrossRef]
  117. Bigioi; Basak, S.; Stypułkowski, M.; Zieba, M.; Jordan, H.; McDonnell, R.; et al. Speech driven video editing via an audio-conditioned diffusion model. Image Vis. Comput. 2024, vol. 142, 104911. [Google Scholar] [CrossRef]
  118. Tian, L.; Wang, Q.; Zhang, B.; Bo, L. Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions. European Conference on Computer Vision, 2024; pp. 244–260. [Google Scholar]
  119. Cui, J.; Li, H.; Yao, Y.; Zhu, H.; Shang, H.; Cheng, K.; et al. Hallo2: Long-duration and high-resolution audio-driven portrait image animation. International Conference on Learning Representations, 2025; pp. 91659–91671. [Google Scholar]
  120. Wei, H.; Yang, Z.; Wang, Z. Aniportrait: Audio-driven synthesis of photorealistic portrait animation. arXiv 2024, arXiv:2403.17694. [Google Scholar]
  121. Zheng, L.; Zhang, Y.; Guo, H.; Pan, J.; Tan, Z.; Lu, J.; et al. Memo: Memory-guided diffusion for expressive talking video generation. arXiv 2024, arXiv:2412.04448. [Google Scholar]
  122. Ma, Y.; Liu, H.; Wang, H.; Pan, H.; He, Y.; Yuan, J.; et al. Follow-your-emoji: Fine-controllable and expressive freestyle portrait animation. SIGGRAPH Asia 2024 Conference Papers, 2024; pp. 1–12. [Google Scholar]
  123. Cheng, H.; Lin, L.; Liu, C.; Xia, P.; Hu, P.; Ma, J.; et al. Dawn: Dynamic frame avatar with non-autoregressive diffusion framework for talking head video generation. International Conference on Learning Representations, 2025; pp. 99250–99268. [Google Scholar]
  124. Xu, Y.; Chen, B.; Li, Z.; Zhang, H.; Wang, L.; Zheng, Z.; et al. Gaussian head avatar: Ultra high-fidelity head avatar via dynamic gaussians. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024; pp. 1931–1941. [Google Scholar]
  125. Qian, S.; Kirschstein, T.; Schoneveld, L.; Davoli, D.; Giebenhain, S.; Nießner, M. Gaussianavatars: Photorealistic head avatars with rigged 3d gaussians. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 20299–20309. [Google Scholar]
  126. Chen, Y.; Wang, L.; Li, Q.; Xiao, H.; Zhang, S.; Yao, H.; et al. Monogaussianavatar: Monocular gaussian point-based head avatar. ACM SIGGRAPH 2024 conference papers, 2024; pp. 1–9. [Google Scholar]
  127. Deng, Y.; Wang, D.; Wang, B. Portrait4d-v2: Pseudo multi-view data creates better 4d head synthesizer. European Conference on Computer Vision, 2024; pp. 316–333. [Google Scholar]
  128. Cho, K.; Lee, J.; Yoon, H.; Hong, Y.; Ko, J.; Ahn, S.; et al. Gaussiantalker: Real-time talking head synthesis with 3d gaussian splatting. In Proceedings of the 32nd ACM International Conference on Multimedia, 2024; pp. 10985–10994. [Google Scholar]
  129. Li, T.; Zheng, R.; Yang, M.; Chen, J.; Yang, M. Ditto: Motion-space diffusion for controllable realtime talking head synthesis. In Proceedings of the 33rd ACM International Conference on Multimedia, 2025; pp. 9704–9713. [Google Scholar]
Figure 1. Talking face generation from a single image and an audio clip.
Figure 1. Talking face generation from a single image and an audio clip.
Preprints 228369 g001
Figure 2. Coverage of prior reviews across seven assessment dimensions (studies as cited in Section 2.1, Section 2.2 and Section 2.3; filled = full coverage, half-filled = partial or restricted, open = not addressed).
Figure 2. Coverage of prior reviews across seven assessment dimensions (studies as cited in Section 2.1, Section 2.2 and Section 2.3; filled = full coverage, half-filled = partial or restricted, open = not addressed).
Preprints 228369 g002
Figure 3. The PRISMA flowchart of our review process.
Figure 3. The PRISMA flowchart of our review process.
Preprints 228369 g003
Figure 4. Talking face generation pipeline, illustrating audio-driven and audio-plus-video-driven classifications. Inputs include an audio signal and face image, with optional text and driving video, processed through feature extraction, mapping, and transformation stages to produce video frames.
Figure 4. Talking face generation pipeline, illustrating audio-driven and audio-plus-video-driven classifications. Inputs include an audio signal and face image, with optional text and driving video, processed through feature extraction, mapping, and transformation stages to produce video frames.
Preprints 228369 g004
Figure 5. Architecture Family Distribution by Year.
Figure 5. Architecture Family Distribution by Year.
Preprints 228369 g005
Figure 6. Training and evaluation datasets across included studies (N = 92): (a) usage frequency (counts sum to more than 92 because studies use several corpora); (b) characteristics of all corpora appearing in four or more studies, plus the pooled custom-corpus category.
Figure 6. Training and evaluation datasets across included studies (N = 92): (a) usage frequency (counts sum to more than 92 because studies use several corpora); (b) characteristics of all corpora appearing in four or more studies, plus the pooled custom-corpus category.
Preprints 228369 g006
Figure 7. Evaluation metrics across included studies (N = 92): (a) reporting frequency; (b) deployment-fitness assessment. SyncNet & variants pools SyncNet confidence with Sync/Lip-Sync Error and LSE-C/LSE-D; Other = CPBD, LVE, AED, EVE, FDD, AKD.
Figure 7. Evaluation metrics across included studies (N = 92): (a) reporting frequency; (b) deployment-fitness assessment. SyncNet & variants pools SyncNet confidence with Sync/Lip-Sync Error and LSE-C/LSE-D; Other = CPBD, LVE, AED, EVE, FDD, AKD.
Preprints 228369 g007
Figure 8. Comparative Normalized Performance Across Architecture Eras.
Figure 8. Comparative Normalized Performance Across Architecture Eras.
Preprints 228369 g008
Figure 9. Ethical Safeguard Presence Across Included Studies.
Figure 9. Ethical Safeguard Presence Across Included Studies.
Preprints 228369 g009
Figure 10. Coded Deployment Domain Distribution Across Included Studies.
Figure 10. Coded Deployment Domain Distribution Across Included Studies.
Preprints 228369 g010
Table 1. Records retrieved per repository per iteration.
Table 1. Records retrieved per repository per iteration.
Repository Iter. 1 (Mar 2023) Iter. 2 (Jun 2023) Iter. 3 (Jan 2024) Iter. 4 (Mar 2026) Total
IEEE Xplore 1,614 1,047 1,183 1,101 4,945
ScienceDirect 893 671 814 782 3,160
ACM Digital Library 712 498 601 553 2,364
SpringerLink 836 622 731 648 2,837
Google Scholar / WoS 1,243 891 872 847 3,853
arXiv 523 215 159 287 1,184
Total 5,821 3,944 4,360 4,218 18,343
Table 2. Inclusion (I) and exclusion (E) criteria.
Table 2. Inclusion (I) and exclusion (E) criteria.
ID Criterion
I1 Full text publicly accessible (open access, subscription, or author preprint)
I2 Written in English
I3 Published or archived 1 January 2016 - 31 December 2025
I4 Employs at least one deep learning technique as a core component
I5 Addresses TFG synthesis, evaluation of TFG output, detection of TFG content, or ethical/governance/socio-technical dimensions of TFG deployment
I6 Primary research contribution (original study, dataset, evaluation, or position paper with novel empirical analysis)
E1 Full text not retrievable after two independent attempts
E2 Duplicate across databases or iterations (DOI match, then title-author string match)
E3 Deep learning application unrelated to human facial video
E4 Outside the temporal scope
E5 Not in English
E6 A survey or review of TFG/deepfakes (treated as related work, not primary evidence)
Table 3. Architecture taxonomy derived from the synthesis corpus (N = 92).
Table 3. Architecture taxonomy derived from the synthesis corpus (N = 92).
Architecture family Phase (years) n Representative models Key capability Principal limitation
LSTM / RNN I (2016-2018) 5 Fan et al. [45]; Suwajanakorn et al. [2] Sequential audio-to-lip mapping Speaker dependence; slow convergence
CNN-based I-II (2017-2020) 7 Song et al. [53]; Zhou et al. [54] Spatial features; real-time feasible No generative capability
GAN-based II (2018-2022) 22 Wav2Lip [4]; PC-AVS [48]; MakeItTalk [47] High lip-sync fidelity; speaker-agnostic Mode collapse; weak 3D consistency
Transformer III (2021-2023) 13 FaceFormer [7]; SyncTalk [55]; CodeTalker [56] Long-range audio-visual alignment High compute; limited diversity vs. diffusion
NeRF-based III-IV (2021-2024) 11 AD-NeRF [9]; 3 RAD-NeRF [57]; GeneFace [41] D-consistent H multi-view synthesis ours of per-identity training; far below real time
Diffusion Model IV (2023-2024) 16 DiffTalk [11]; DreamTalk [12]; VASA-1 [13]; Hallo [58] Probabilistic; diverse; stable training High inference latency (except flow matching)
Gaussian Splatting V (2024-2025) 8 GSTalker [50]; GaussianEmoTalker [51]; VASA-3D [52] Real-time photorealistic rendering Identity-specific scene optimization
Diffusion Transformer V (2024-2025) 7 OmniHuman-1 [14]; ACTalker [59]; EmotiveTalk [20]; Teller [60] Scales with compute; full-body; multi-modal Very high training compute
Hybrid V (2024-2025) 3 Control-Talker [61]; MimicTalk [62]; TalkingGaussian [63] Combines paradigm strengths Complexity; opaque failure modes
Table 4. Benchmark performance of representative state-of-the-art models from the corpus.
Table 4. Benchmark performance of representative state-of-the-art models from the corpus.
Model Family Year FID (lower better) SyncNet conf. Real-time? Key advance
Wav2Lip [4] GAN 2020 N/A 8.31 Yes First in-the-wild lip sync on arbitrary video
PC-AVS [48] GAN 2021 41.2 7.86 Partial Pose-controllable synthesis
AD-NeRF [9] NeRF 2021 N/A 7.42 No (6 fps) p 3D-consistent portraits
SadTalker [40] 3DMM+Diffusion 2023 47.7 7.98 Partial 3D-aware motion without per-identity tuning
DiffTalk [11] Diffusion 2023 35.4 7.40 No First generalized audio-driven diffusion TFG
Hallo [58] Diffusion 2024 27.2 8.02 Partial Hierarchical audio-visual attention
MuseTalk [19] Diffusion 2024 31.4 8.49 Yes First real-time diffusion lip sync
SyncTalk [55] Transformer+NeRF 2024 N/A 8.64 Partial Synchronized temporal attention
VASA-1 [13] Diffusion (flow) 2024 22.1 8.59 Yes (45 fps) Disentangled facial dynamics in real time
GSTalker [50] Gaussian Splatting 2024 N/A 8.04 Yes Deformable Gaussian Splatting for TFG
OmniHuman-1 [14] DiT 2025 18.4 8.53 Partial Full-body DiT; multi-modal conditioning
EmotiveTalk [20] DiT 2025 21.3 8.31 Partial Content/emotion audio decoupling
Teller [60] DiT 2025 N/A 8.72 Yes Autoregressive streaming; lowest latency
Table 5. Inventory of ethical safeguard content across the synthesis corpus (N = 92).
Table 5. Inventory of ethical safeguard content across the synthesis corpus (N = 92).
Ethical dimension Studies (n) % of corpus Nature of coverage
Watermarking / provenance 1 1.1% One post-hoc frequency-domain watermark; not C2PA-compliant; no compression robustness
Deepfake detection integrated 3 3.3% Parallel detector trained or evaluated alongside the generator, typically as an ablation
Ethical risk discussion 5 5.4% One to three sentences acknowledging misuse potential; no mitigation proposed
Consent protocol 0 0.0% Whose consent is required to synthesize a face is not addressed anywhere in the primary literature
Regulatory framework cited 0 0.0% The regulatory environment is invisible to the reviewed technical literature
Any ethical content 8 8.7% 91.3% of included studies contain none
Table 6. Cross-domain socio-technical analysis of TFG deployment (RQ6 summary)
Table 6. Cross-domain socio-technical analysis of TFG deployment (RQ6 summary)
Domain Studies (n) Primary benefit Primary risk Most urgent governance need Regulatory instruments
Healthcare & telehealth 7 Scalable, empathetic clinical communication Patient deception; demographic inequity; liability ambiguity Mandatory disclosure; demographic audit; clinical oversight EU AI Act Art. 50, 10, 14
Education & training 9 Personalized, emotionally adaptive instruction Student biometric capture; epistemic authority effects Biometric data governance; disclosure; educator oversight GDPR Art. 9; FERPA; EU AI Act Art. 50
Identity management 5 Scalable remote verification Liveness evasion; synthetic identity fraud at scale Multi-factor mandates; C2PA provenance; DDR reporting TAKE IT DOWN Act; C2PA v2.0; KYC rules
Governance & democracy 3 Accessible multilingual civic communication Disinformation; electoral manipulation; trust erosion Mandatory provenance; platform liability; media literacy EU AI Act Art. 50; C2PA v2.0; electoral law
Table 7. The six ethical principles of the RTFG-STS Framework.
Table 7. The six ethical principles of the RTFG-STS Framework.
Principle Source TFG interpretation Mechanisms
P1 Beneficence AI4People [35]; bioethics [72] Clear societal benefit, demonstrably delivered across demographic groups M3; M4
P2 Non-maleficence AI4People [35] Foreseeable harms prevented at synthesis time, not patched afterwards M1; M2; M4
P3 Autonomy AI4People [35] Subjects decide whether their face is synthesized; users know they face an AI; disclosure is proactive M2; M5
P4 Justice AI4People [35]; EU AI Act Art. 10 Equitable performance across the served population; audited datasets; stratified reporting M3; M4
P5 Explicability AI4People [35]; EU AI Act Art. 50 Output machine-detectable (C2PA) and socially disclosed at interaction M1; M4
P6 Systemic embeddedness STS theory [15,16] (new) Design and governance reference the concrete deployment context; PTS, CCAI, DDR, and a STIA supplement technical metrics M3; all
Table 8. Risk-tiered consent architecture.
Table 8. Risk-tiered consent architecture.
Tier Context Consent requirement Implementation Regulatory basis
1 (Low) Entertainment, creative, personal Disclosure at point of viewing C2PA manifest; UI label EU AI Act Art. 50; C2PA v2.0
2 (Medium) Education, customer service, public information Explicit opt-in from the face subject; viewer disclosure; right of withdrawal Signed consent record in append-only log; manifest carries consent hash EU AI Act Art. 50; GDPR Art. 6; TAKE IT DOWN Act [27]
3 (High) Healthcare, identity, legal, political Revocable informed consent as a Verifiable Credential; periodic re-consent; human oversight VC per Eq. 6 on a distributed ledger; human confirmation before action EU AI Act Art. 50, 14; GDPR Art. 9
Table 9. RTFG-STS Framework application map by deployment domain.
Table 9. RTFG-STS Framework application map by deployment domain.
Domain Tier Mandatory mechanisms Metrics Instruments Priority research need
Healthcare & telehealth 3 M1-M5 (Tier 3 VC; DDR ≥ 80%; clinician oversight) PTS; CCAI; DDR; CSIM; MOS EU AI Act Art. 50, 10, 14, 27; GDPR Art. 9 Diverse clinical evaluation data; validated PTS
Education & training 2-3 M1-M4 (Tier 2/3; DDR ≥ 60%) PTS; CCAI; DDR; AUE; MOS EU AI Act Art. 50; GDPR Art. 9; FERPA; COPPA Student biometric governance; authority-effect studies
Identity management 3 M1-M5 (Tier 3 VC; DDR ≥ 80%; human ID reviewer) DDR; CSIM; FID; SyncNet; FPS TAKE IT DOWN Act [27]; EU AI Act Art. 50; C2PA v2.0; KYC rules Standardized liveness evaluation; DDR benchmark
Governance & democracy 3 M1-M5 (Tier 3; DDR ≥ 80%; editorial oversight) DDR; PTS; FID; MOS EU AI Act Art. 50; C2PA v2.0; electoral law Cross-platform C2PA verification; media-literacy studies
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.