Submitted:
25 August 2026
Posted:
26 August 2026
You are already at the latest version
Abstract
Deepfake detection is usually measured on curated benchmarks, yet synthetic video often reaches audiences through social-media platforms that compress, resize, re-encode, crop, and repost it before downstream detection. The practical question is therefore not whether a detector separates real from fake under laboratory conditions, but whether the evidence it depends on still exists at the moment a decision must be made. This survey reviews Ncorpus works on generative audio-video deepfakes across the full lifecycle of creation, platform distribution, detection, provenance, and remediation. We organize the review around five forensic assumptions that detectors implicitly rely on: that a manipulated region leaves a boundary, that a generator leaves a stable fingerprint, that test data resemble training data, that forensic signals survive processing, and that low-level clues are sufficient. Using these assumptions as a common lens, we present an evidence-based taxonomy of generation methods, review signal-driven, learning-based, reasoning-based, agentic, and adversarially robust detection, and examine provenance mechanisms including content credentials and watermarking alongside emerging disclosure and takedown regulation. Two findings recur. First, foundation-model generation can weaken several assumptions at once, since fully synthetic and jointly generated audio-video content may remove compositing boundaries and reduce exploitable cross-modal inconsistencies. Second, the literature is concentrated on creation and detection, with distribution, provenance, and remediation together accounting for PctTail of the corpus. We close with evaluation protocols that report which assumptions a benchmark actually exercises. Project Page: https://vectorinstitute.github.io/deepfakes-survey-2026/
Keywords:
deepfakes
; video generation
; media provenance
; foundation models
; diffusion
; forensics
; agentic AI
1. Introduction
Over the past decade, synthetic media generation has progressed through successive waves of generative modeling, from variational autoencoders and generative adversarial networks (GANs) to diffusion models and, more recently, large-scale video foundation models, making fabricated content increasingly realistic and convincing [1]. At the same time, synthetic content is increasingly distributed through social media platforms, where recommendation algorithms, compression, moderation, and user interactions shape its visibility and spread. DataReportal [2] and Statista [3] estimate approximately 5.24 billion active social media user identities worldwide at the start of 2025, with major platforms such as Facebook, YouTube, Instagram, TikTok, and Telegram operating at billion-user scale. Once uploaded, a synthetic video may be transformed through platform-specific compression, algorithmic promotion, editing, and cross-platform reposting in ways that controlled laboratory evaluations do not fully capture.
A deepfake is, therefore, not simply a video to classify; it passes through a lifecycle of creation, upload, platform processing, recommendation, viewing, reporting, review, and eventual removal, labelling, or continued spread. Media transformations within this lifecycle, such as compression, re-encoding, cropping, and reposting, can weaken or alter the forensic traces on which detectors rely, including sensor noise, physiological inconsistencies, and generative artifacts [4,5]. Despite these real-world dynamics, deepfake detection research has largely evolved independently of how deepfakes are distributed online. The challenge therefore extends beyond high detection accuracy in controlled settings to whether forensic evidence remains reliable throughout real-world distribution.
Figure 1.
Structure of this survey. Each leaf lists representative works and, where forensic (A1–A5) applies.
Figure 1.
Structure of this survey. Each leaf lists representative works and, where forensic (A1–A5) applies.

The fragility of passive forensic cues under real-world distribution has also motivated proactive approaches based on watermarking and content provenance. For example, Anthropic has announced embedded text watermarks and C2PA-based signed provenance metadata for Claude-generated content1. Motivated by these challenges, this survey bridges the gap between laboratory-focused deepfake detection with real-world deployment. Rather than treating detection as an isolated classification problem, we examine the deepfake ecosystem across its full lifecycle.
Scope This survey examines deepfakes from generation to real-world social media deployment. We first review the deepfake production pipeline and generative foundations, then examine detection methods and the social media lifecycle, including creation, distribution, detection, provenance, and remediation. Throughout, we use five forensic assumptions to explain how generation and deployment conditions affect detector reliability.
Search Methodology We conducted a systematic literature search spanning January 2017 to August 2026 across IEEE Xplore, the ACM Digital Library, SpringerLink, ScienceDirect, OpenReview, PMLR, the AAAI Digital Library, and arXiv. Queries combined {deepfake, face forgery, synthetic media, AI-generated video} with {detection, generation, robustness, watermarking, provenance, evaluation, social media, in the wild}, supplemented by forward and backward citation tracking. We also included foundational pre-2017 methods, official regulatory and standards documents, and widely used open-source implementations.
Candidates were screened for relevance to deepfake generation, forensic detection, platform distribution, provenance, remediation, or evaluation, and duplicates were removed. Non-peer-reviewed works were retained only if they met a minimum citation threshold or introduced a post-2024 foundation-model generator or evaluation lacking a peer-reviewed account. The final corpus comprises 212 works, of which 74% are from peer-reviewed venues. The median publication year is 2024 and 52% appeared in 2024 or later. Figure 2 summarizes the screening process.
Within the corpus, we prioritized works that (i) introduce novel methods with empirical evaluation, (ii) address post-2022 generators, (iii) provide cross-dataset, cross-generator, or real-world evaluation, or (iv) contribute widely adopted datasets, protocols, or implementations. Figure 3 summarizes the corpus, which is concentrated on detection and generation, with comparatively few works addressing distribution, remediation, and provenance, revealing an important imbalance in the existing literature.
Relation to Prior SurveysTable 1 compares this survey with representative prior reviews. Existing surveys primarily focus on generation and detection taxonomies [112,113,114,116,117,118], reliability [115], proactive defense [1], modality-specific analysis [119], or model architectures [120]. Although some cover temporal and audio-visual evidence, few connect deepfake generation and forensic detection with the platform conditions under which synthetic media is processed, propagated, verified, and remediated. Agentic detection also remains largely absent. Our survey addresses this gap by integrating generative foundations and forensic detection with a social-media lifecycle perspective, including platform processing, provenance, agentic methods, and real-world audio-video deepfakes.
Contributions Our contributions are: (1) We introduce the A1–A5 framework to explain when and why detectors fail. (2) We develop a generation taxonomy linking manipulation types to the traces they leave. (3) We review detection methods spanning signal-based, learning-based, reasoning-based, agentic, and adversarial approaches. (4) We formalize the social-media lifecycle as transformations of forensic evidence and quantify a literature imbalance: creation and detection account for 75% of the corpus, versus 12% for distribution, provenance, and remediation. (5) We connect provenance and real-world evaluation to A1–A5, mapping benchmarks to these assumptions and outlining deployment requirements. To the best of our knowledge, this is the first survey to unify foundation-model-era generation, forensic detection, and the social-media lifecycle.
2. Background, Generative Modeling and Forensics
This section establishes the conceptual and technical foundations for this survey.
2.1. Deepfake Video Attack Pipeline
We formalize the deepfake video attack as a five-stage pipeline, as illustrated in Figure 4.
Data Acquisition. Attackers may gather target images or videos from social media, public appearances, and other web sources, making social platforms both a data source and a distribution channel. Traditional face-swap systems [15] typically require hundreds or thousands of identity-specific frames and per-target training. Recent methods reduce these requirements through training-free identity conditioning from a single reference image [43,44] (targeted synthesis), while text-to-video foundation models can generate realistic synthetic individuals without target-specific data [121] (untargeted synthesis).
Preprocessing. The collected media may undergo face detection, alignment, landmark extraction, and segmentation to produce normalized face crops suitable for identity-specific generation [122]. Cropping and alignment resample the source and overwrite capture-level regularities. The crop boundary can also determine where a compositing seam appears when the generated face is returned to the frame. In audio-driven pipelines, speech is converted into mel-spectrograms or learned embeddings that guide facial motion [123], limiting the visual signal to information preserved by that representation.
Generation. The synthesis stage uses generative models to produce the manipulated face, image, or video. Traditional pipelines relied mainly on autoencoders [15] and GANs, whereas modern systems increasingly use diffusion and flow-matching models [11,124], autoregressive transformers [12], and 3D-aware neural rendering [14]. This shift toward full-frame foundation-model generation reduces the localized artifacts and stable fingerprints exploited by earlier forensic detectors.
Post-Processing. The generated output may be refined through blending, alpha compositing, colour harmonization, resolution, temporal smoothing, and codec re-encoding [15]. In such deepfakes, these operations improve consistency between the manipulated region and the surrounding frame, whereas fully synthetic videos may undergo appearance and temporal enhancement. Adversarial perturbations may also target the features on which a detector relies [125].
Dissemination. The final video is distributed through online channels such as social media. Platform-specific processing, including transcoding, resolution reduction, cropping, overlays, and re-encoding, can alter high-frequency statistics, remove metadata, and weaken links to provenance manifests [6]. These operations can suppress generation traces while introducing platform artifacts. Repeated downloading and reuploading across platforms compound this effect, widening the gap between benchmark and real-world detection conditions.
2.2. Generative Modelling Foundations for Video Synthesis
We review the generative paradigms behind deepfake video in this section.
Autoencoders and Variational Autoencoders (VAE) Autoencoders map inputs to a lower-dimensional latent representation and reconstruct them through a decoder; VAEs instead regularize the encoded distribution toward a prior, trading reconstruction accuracy for a smoother, sampleable latent space. Early face-swapping systems such as DeepFaceLab [15] and FaceSwap [126] paired a shared encoder with identity-specific decoders, blending a reconstructed face into an authentic frame. VAEs persist as the encoders and decoders of modern diffusion models.
Generative Adversarial Networks (GANs) Based Video Synthesis GANs, which learn through competition between a generator and a discriminator, dominated face synthesis and manipulation before the shift toward diffusion. StyleGAN [9] improved controllable synthesis, FaceShifter [127] and SimSwap [16] enabled identity transfer, and FOMM [19] and StyleGAN-V [128] extended it to motion and time. Their convolutional upsampling leaves periodic frequency traces, which much of the early detection literature was built to read.
Diffusion-Based Video Generation Diffusion models generate video by reversing a noising process, usually in the latent space of a pretrained autoencoder [11]. Because the decoder is shared across outputs of a model family, its reconstruction bias becomes a forensic trace. Video diffusion extends image backbones with temporal modelling [129,130]; more recent diffusion transformers operate on space-time tokens, while flow-matching approaches learn continuous transformations between noise and data. Diffusion also supports localized editing and inpainting, which can leave boundaries absent from fully generated video.
Autoregressive (AR) and Transformer-Based Video Models AR models factorize the joint distribution of video tokens as a product of conditionals, so the discrete tokenizer sits between the model and the output and leaves quantization traces of its own. CogVideo [12] and VideoPoet [13] extend Transformer-based generation to video, editing, and multimodal synthesis, while hybrid systems increasingly combine autoregressive and diffusion components [131,132].
3D-Aware Neural Rendering: NeRFs and Gaussian Splatting 3D-aware approaches explicitly model geometry, appearance, and motion. Neural Radiance Fields represent scenes as continuous volumetric functions, whereas 3D Gaussian Splatting uses rasterized Gaussian primitives for efficient rendering [14]. Each leaves rendering artifacts: floaters and view-dependent inconsistency for NeRFs, and splat-edge visibility and popping for Gaussian Splatting. These representations suit controllable talking-head synthesis, while newer hybrid approaches combine audio or motion conditioning with diffusion.
2.3. Forensic Signals in Deepfake Videos
Digital media forensics2 predates deepfakes by nearly two decades [4]. Early studies showed that authentic media contain regularities introduced during capture and compression, while manipulation can disturb these patterns. Sensor noise, resampling artifacts, and compression statistics can reveal camera origin and processing history, while video forensics extends these cues to temporal edits such as frame duplication or deletion [4,133]. The same principle was later applied to synthetic media, where GANs were found to leave model-related fingerprints [55].
Framework for Forensic Assumptions Deepfake detectors rely on assumptions about the evidence available, such as the presence of an editing boundary, a recurring generator fingerprint, or similarity between training and test data. We group these expectations into five assumptions (A1–A5), which form the analytical spine of this survey.
A1: A manipulated region leaves a boundary. Many manipulations replace only part of the frame, such as a swapped face or an inpainted region, while leaving the rest unchanged. Compositing the generated content into the original frame may leave differences in texture, resolution, colour, or compression, and Face X-Ray detects such blending traces [53]. The assumption, therefore, holds only when manipulated content is composited against an authentic region and fails for fully synthetic video, where every pixel is generated.
A2: A generator leaves a stable fingerprint. A detector may assume that videos produced by the same model share recurring statistical patterns. Upsampling in GAN generators, for example, can leave frequency-domain traces that support detection or attribution [55]. However, these patterns may change across models, versions, fine-tuning methods, and post-processing. A detector trained on one generator’s fingerprint may therefore fail on another.
A3: Test videos resemble the training data. Detectors usually perform better when the videos seen during testing are similar to those used during training. In real-world deployment, however, they may encounter new generators, manipulation methods, identities, scenes, or recording conditions. This difference between training and deployment limits cross-generator and cross-dataset performance [51].
A4: Forensic signals survive processing. A detector may assume that relevant forensic detail survives processing. Benchmark videos are often original or only moderately compressed [47], whereas social-media videos may be resized, filtered, compressed, or re-encoded several times. These operations can weaken useful traces and introduce unrelated platform artifacts [6].
A5: Low-level clues are enough. Many detectors rely mainly on pixels, textures, frequencies, or frame-level motion. They may not examine whether the scene makes sense, whether the audio matches the speaker, or whether objects behave consistently. As generated videos become more realistic, semantic and multimodal reasoning may be needed to identify contradictions that low-level detectors miss [80].
In the generation table (Table 2) the symbols report whether an assumption holds for a method’s outputs. In the detection table (Table 3), they indicate whether a method depends on an assumption or mitigates its failure. In the dataset table (Table 4), they identify the assumptions a benchmark can stress-test. Fewer filled circles in the generation table therefore indicates a harder detection problem, not a more capable generator.
3. Taxonomy of Generative Deepfake Videos
This section organizes deepfakes by manipulation type and the forensic evidence available to detectors.3
3.1. Identity-Centric Deepfakes
Identity-centric deepfakes alter the appearance or behaviour of a person. Many early methods modified a localized facial region while preserving the surrounding frame. Recent systems can generate the full head or portrait, reducing the amount of authentic facial content available for comparison.
Face Swapping. Face swapping transfers a source identity onto a target video while preserving the target pose, expression, and scene context. Methods in this category span successive architectural generations: autoencoder-based systems such as DeepFaceLab [15], GAN-based methods such as FaceShifter [127] and SimSwap [16], and diffusion-based methods such as DiffFace [150], DiffSwap [151] and REFace [17]. Identity transfer through implicit 3D representations has also been explored [152]. Because the fake face (driver) is placed into a real frame (source), the manipulated region may be localizable through compositing seams(A1), while the generator family determines what else is recoverable (A2).
Facial Reenactment. Facial reenactment transfers driver expressions, gaze, or head pose while preserving the source identity. Face2Face [18] uses explicit facial reconstruction, FOMM [19] uses learned keypoints and motion representations, and LivePortrait [153] supports efficient portrait animation. Unlike face swapping, reenactment generally preserves much of the target appearance while altering facial dynamics. Detectors must therefore examine motion, expression, and identity dynamics rather than relying only on static appearance. The frame is authentic and nothing is pasted in, so A1 weakens and the evidence is motion rather than a seam.
Talking-Head Synthesis. Talking-head synthesis generates a speaking portrait from a reference image (source) and driver audio. Existing approaches include mesh-based systems such as SadTalker [154], neural-rendering methods such as ER-NeRF [155] and GaussianTalker [144], and diffusion-based methods such as EMO [20], Hallo 3 [147], and VASA-1 [21]. These methods may generate most or all facial motion from a single static image, so there may be no localized face-swap boundary. When the full portrait is generated, A1 does not apply and only temporal and audio-visual consistency remain.
Lip-Sync Manipulation. Lip-sync manipulation modifies mouth motion in a target video to match driver audio while preserving most of the original frame. Wav2Lip [22] is a widely used example, while Diff2Lip [156], and VideoReTalking [157] use diffusion or multi-stage processing to improve visual quality. Because edits may be confined to the mouth region, global video statistics can be dominated by authentic content, while compression further weakens local traces. Mouth-region analysis and audio-visual synchronization are therefore especially important. A1 holds only around the mouth, while A4 becomes particularly fragile because compression erases small-region traces first.
3.2. Body-Centric Manipulations
Body-centric methods modify human pose, gesture, movement, or appearance. They may preserve the original face, replace the full person, or synthesize the entire body within an existing or generated scene.
Pose and Gesture Transfer. Pose and gesture transfer animates a source person using motion extracted from a driver video or pose sequence. Everybody Dance Now [23] demonstrated pose-driven motion transfer, while MagicAnimate [24] and Animate Anyone [25] use diffusion-based generation to improve appearance consistency and temporal quality. These methods may leave inconsistencies around limbs, hands, clothing, occlusions, or interactions with the background. A1 holds only when the generated person is composited into an authentic scene.
Full-Body Puppetry and Attribute Editing. Full-body puppetry extends pose transfer by controlling a person movement, gesture, and appearance. DisCo [26] disentangles pose, background, and human appearance for controllable dance generation, Champ [27] conditions on a parametric 3D body model for improved shape and pose fidelity, and UniAnimate [28] extends generation to longer sequences. Attribute editing may instead modify clothing, body appearance, or other visual characteristics while preserving identity and motion. As the edit expands from a local seam to full-frame synthesis, A1 progressively weakens and may disappear entirely.
3.3. Audio-Centric Manipulations
Audio-centric deepfakes alter speech content, speaker identity, or vocal characteristics. Unlike visual manipulation, audio has no spatial boundary. Detection instead relies on acoustic, temporal, linguistic, and speaker-related evidence.
Speech Synthesis and Voice Cloning. Speech synthesis, commonly referred to as text-to-speech (TTS), generates an utterance from text, while voice cloning produces speech that imitates a target speaker. Neural codec language models such as VALL-E [29] synthesize speech from a few seconds of reference audio, YourTTS [30] supports zero-shot multilingual cloning, and diffusion and flow-matching systems such as NaturalSpeech 2 [31] and VoiceBox [32] improve prosody and speaker similarity. A2 is load-bearing here, since detection assumes that each model family leaves a stable acoustic fingerprint.
Voice Conversion and Speech Editing. Voice conversion changes the perceived speaker while preserving the linguistic content of an utterance, using disentangled content and speaker representations [158] or self-supervised speech features. Speech-editing methods instead modify selected words or temporal segments while retaining the surrounding authentic audio [35]. Partial editing creates the audio analogue of video inpainting, and pairing manipulated audio with authentic video defeats visual-only detectors. This violates A5 because a detector relying on only one modality may miss manipulation confined to the other.
3.4. Scene-Centric Deepfakes
Scene-centric methods generate or modify substantial portions of a video scene. Unlike localized identity manipulation, these methods may contain little or no camera-captured reference content. The detection task consequently shifts from locating an edited region to determining whether the full sequence was generated or substantially altered.
Text-to-Video (T2V) Generation. T2V models generate complete video sequences from natural-language prompts. Recent T2V systems such as Open Sora [36], CogVideoX [38], HunyuanVideo [159], and Wan 2.1 [39] generate coherent multi-second to minute-scale sequences. Because the full scene may be synthetic, no authentic-synthetic compositing boundary exists, invalidating A1. Unseen generators may also fall outside the training distribution, challenging A3.
Image-to-Video (I2V) Animation. I2V methods animate a static image using text, motion, or other conditioning signals. Stable Video Diffusion [40] generates video from an image, while AnimateDiff [41] introduces reusable motion modules. Although the conditioning image may provide an identity or appearance reference, the generated frames remain synthetic; forensic evidence therefore depends partly on whether the source image itself is authentic. A1 depends on the conditioning image: an authentic source gives a reference boundary, whereas a generated one does not.
Video-to-Video (V2V) Editing. V2V editing modifies existing footage through inpainting, instruction-guided editing, object replacement, or broader scene transformation. InstructVid2Vid [140] and Wan 2.1-VACE [39] support controllable video editing, while ProPainter [42] performs temporally consistent inpainting and object removal. The spatial extent of these edits can vary from a small object to most of the scene. Localized edits may leave boundaries, while global transformations such as changing weather, lighting, or visual style may have diffuse boundaries that are difficult to localize. A1 holds for local edits and fails for global transformations, where boundaries are diffuse.
3.5. Hybrid and Multimodal Pipelines
Modern deepfakes may combine several manipulation categories and model families within one workflow. These pipelines blur the boundaries between identity, body, audio, and scene manipulation and can modify the forensic traces introduced at earlier stages.
Joint Audio-Visual Manipulation. Joint audio-visual manipulation alters both the visible speaker and the accompanying voice. For example, a cloned voice may be combined with talking-head synthesis or lip-sync editing. Such content can remain consistent across both modalities, reducing the effectiveness of detectors that assume one channel is authentic. Detection must therefore examine not only synchronization but also whether the voice, facial behaviour, identity, and semantic content are mutually consistent. A5 fails here because cross-modal consistency alone may no longer distinguish authentic from manipulated content.
Multi-Stage Generation and Editing. A multi-stage pipeline may first generate a scene, inject a target identity using methods such as InstantID [43] or IP-Adapter [44], or align mouth motion using Wav2Lip [22] , add a cloned voice, and then apply video editing and post-processing. Each stage may weaken, replace, or add forensic signals. The final output therefore cannot always be attributed to a single generator or manipulation type. A detector trained on isolated face swaps, lip-sync edits, or text-to-video outputs may not generalize to their combination. A2 fails, because no single fingerprint survives a chain of generators.
3.6. Cross-Cutting Temporal and Partial Manipulations
Any manipulation category may affect only part of a video duration or spatial extent. Segment splicing, frame insertion or deletion, and short manipulated windows can produce sequences that remain predominantly authentic. Datasets such as LAV-DF [45] and AV-Deepfake1M [46] include temporally localized manipulations, often altering a short word, phrase, or corresponding visual segment. In this setting, a video-level label provides limited information; the task is to identify the manipulated interval or region. Temporal and partial manipulations therefore cuts across identity-, body-, audio-, scene-, and multimodal deepfakes. The forensic question therefore shifts from whether the video is manipulated to whether the manipulated interval or region can be localized.
3.7. Forensics Mapping
Across the taxonomy, A1 weakens as manipulation expands from localized edits to full-frame generation. A2 remains relevant but is challenged by chained generators that obscure fingerprints. A5 is weakened when jointly generated audio and video since cross-modal and semantic inconsistencies may be reduced. A3 and A4 are not properties of the manipulation at all. They are set by what a detector was trained on and by what the media has been through since, so they can fail for any row in Table Section 3.7, which is why the lifecycle in §5 is part of the same problem.
4. Detection, Localization, and Attribution of Manipulated and Fully Synthetic Videos
This section reviews methods for detecting, localizing, and attributing manipulated and fully synthetic videos.
4.1. Problem Formulation
Deepfake detection comprises several related tasks that differ in output granularity and operational purpose. Let denote an audio-video sample, where and when audio is unavailable. We next formalize the main forensic tasks.
Binary Authenticity Verification. The most common formulation treats detection as binary classification. Given , a detector produces , where a higher value indicates stronger evidence of manipulation. The binary decision is , where is a decision threshold selected for the intended operating conditions. This formulation underlies widely used benchmarks such as FaceForensics++ [47], Celeb-DF [48], and DFDC [49]. Binary verification provides a simple decision but does not identify the manipulation type, affected region, or likely generation process.
Fine-Grained Deepfake Classification. Fine-grained classification identifies the type of synthetic alteration present in a sample. Benchmarks range from manipulation-specific labels in FaceForensics++ to broader taxonomies such as ForgeryNet [50] and DF40 [51], while FakeAVCeleb [52] distinguishes visual, audio, and combined audio-video manipulations. For K deepfake categories and one authentic category, single-label classification produces , with . Multi-label classification instead predicts . Naming a manipulation assumes that categories remain distinguishable, which a pipeline combining face swapping, lip-sync editing, and voice cloning no longer guarantees (A2).
Spatial and Temporal Localization. Localization methods identify where and when synthetic modifications occur. For a video with T frames of height H and width W, spatial localization produces , where each entry represents the probability that a pixel is manipulated. Temporal localization produces , where each entry represents the probability that a frame belongs to a manipulated segment. Face X-Ray predicts blending-boundary masks [53], ForgeryNet provides pixel-level annotations [50], and UMMAFormer localizes manipulated temporal segments using audio and visual streams [54]. For fully synthetic video, spatial localization may reduce to a full-frame mask because no authentic region remains (A1).
Generator Attribution. Generator attribution seeks to identify the model, model family, or generation pipeline responsible for a synthetic sample. For G known generators and one unknown-source category, attribution produces , with . The output assigns the sample to a known generator or an unknown source. GAN-generated media can contain model-specific fingerprints [55], but these traces may change across model variants, fine-tuning, and multi-stage pipelines (A2). Attribution systems should therefore support an “unknown” outcome. Unlike attribution, provenance relies on externally recorded information or credentials rather than media-internal evidence (Section 5.3).
Operating Points, Prevalence, and Calibration. The threshold is often treated as an implementation detail, while evaluation emphasizes threshold-independent measures such as AUC. In deployment, however, performance at the selected operating point is critical. Let denote synthetic-media prevalence, the true-positive rate, and the false-positive rate. Precision is
For example, at , , and , precision is only about . Under low prevalence, even a small FPR can create substantially more false alarms than correct detections, a deployment burden not captured by AUC alone. Detection scores should therefore be calibrated so that confidence reflects observed correctness.
4.2. Signal-Driven Forensic Detection
Signal-driven methods target specific statistical, spatial, temporal, physiological, or compression properties of media. Some use handcrafted measurements, while others train neural networks to extract forensic evidence.
Spatial and Identity-Consistency Cues Spatial methods examine blending boundaries, texture discontinuities, facial geometry, and identity consistency. Face X-Ray estimates boundary maps produced by face compositing [53]. ID-Reveal compares facial identity representations with the surrounding spatio-temporal context using ArcFace embeddings [56,176]. Such boundary-based evidence may be absent in fully synthetic frames (A1).
Frequency-Domain and Reconstruction-Based Detection Frequency-based detectors exploit spectral patterns introduced by generation pipelines, including artifacts associated with convolutional upsampling [177]. However, diffusion and alias-free architectures can weaken these signatures [57]. Reconstruction-based methods instead measure differences after projecting media through a generative model’s latent representation; AEROBLADE [58], for example, uses reconstruction error from a latent diffusion autoencoder. Both approaches depend on generator-related traces that may change across architectures or processing pipelines (A2).
Temporal, Physiological, and Audio-Visual Cues Video detectors can use information unavailable in isolated frames, including physiological signals, facial motion, optical flow, and frame-to-frame consistency. Physiological cues such as blinking, gaze, and remote photoplethysmography were effective against earlier generators that modelled individual frames better than long-term facial behaviour. Their reliability decreases as temporal modelling improves and is affected by motion, illumination, resolution, and compression. AltFreezing [66] learns complementary spatial and temporal forgery features, while audio-visual approaches examine speech–mouth correspondence, as in LipForensics and self-supervised correspondence methods [178]. Such mismatch cues weaken when audio and video are generated jointly and synchronized by construction (A5).
Acoustic and Speech-Forensic Cues Audio deepfakes include text-to-speech synthesis, voice conversion, and localized speech editing. Detection uses spectral, temporal, phase, vocoder, neural-codec, or learned waveform representations, with representative systems including RawNet2 [60] and AASIST [61]. Performance can decline across unseen generators and after compression, resampling, noise, or platform transcoding; consequently, audio detection often depends on stable generator traces that survive subsequent processing (A2, A4).
Compression and Codec Forensics Codec analysis examines quantization, prediction, and compression-history inconsistencies. Double-compression traces may indicate that a video was edited and re-encoded. However, generated and authentic videos often use the same codecs, and social-media platforms routinely transcode uploaded content. Platform processing can therefore remove generation traces while introducing new compression patterns unrelated to manipulation. Codec traces primarily reflect processing history, which platform transcoding may overwrite before detection (A4).
4.3. Learning-Based Detection Architectures
Learning-based detectors infer discriminative representations directly from data. Their architectures have progressed from frame-level convolutional networks to temporal transformers, adapted foundation models, and multimodal systems.
CNN-Based Frame-Level Detection Early deepfake detectors adapted CNNs such as Xception, ResNet, and EfficientNet for frame-level classification, often achieving strong within-dataset performance but weaker cross-dataset transfer. Later approaches improve generalization through augmentation and localized forensic learning; for example, Self-Blended Images generates pseudo-forgeries without relying on a fixed generator [62]. Nevertheless, performance remains sensitive to differences between training and deployment distributions (A3).
Temporal and Transformer-Based Architectures Transformer-based detectors model long-range spatial and temporal relationships across video frames. Representative approaches include TALL [64], which compactly represents temporal information, and AltFreezing [66], which alternates spatial and temporal optimization. Although these architectures integrate broader evidence than frame-level CNNs, they do not eliminate distribution shift to unseen forgery types or fully synthetic video (A3).
Foundation-Model-Adapted Detectors Recent detectors adapt large pre-trained vision encoders to improve transfer across manipulation methods. UnivFD [67], for example, uses frozen CLIP representations for cross-generator detection, while newer approaches adapt pre-trained features through lightweight forensic modules or selective fine-tuning [179]. Face-specific pre-training provides a complementary direction [69,70]. However, broader pre-training can reduce rather than eliminate sensitivity to unseen deployment distributions (A3).
Multimodal Detection Frameworks Multimodal systems combine evidence from video, audio, text, metadata, or provenance. Audio-visual detectors measure temporal alignment, phonetic consistency, or correspondence between vocal and facial characteristics. SyncNet-based approaches estimate alignment between speech and mouth motion [180], while LipForensics and self-supervised audio-visual methods learn deviations from natural speech dynamics [178,181]. Joint audio-video generation presents a more difficult case, since synchronization may be internally consistent even when the whole event is synthetic. Multimodal detection must therefore extend beyond temporal alignment toward semantic, provenance, and source-based verification. Cross-modal mismatch cues are most informative when the modalities are generated or manipulated independently; their value may decrease when a single model jointly generates consistent audio and video (A5)
Detectors for Fully Synthetic Video T2V and I2V generation removes the compositing structure assumed by many facial-manipulation detectors because entire frames or clips may be synthetic. UNITE therefore classifies full-frame content using domain-agnostic vision features [71], while DeMamba applies a bidirectional state-space module over frozen encoder features and is evaluated across contemporary video generators [72]. Song et al. [182] instead learn multimodal representations targeting diffusion-generated video. Without compositing seams, these methods rely on statistical and temporal evidence that must remain stable across generators (A2), transfer beyond the training distribution (A3), and survive processing (A4).
4.4. Generalization and Adaptation
A major obstacle to deployment is performance loss under distribution shift. High scores on a benchmark do not imply comparable performance on contemporary social-media content. For example, Deepfake-Eval-2024 contains approximately 44 hours of video collected from 88 web sources and reports an average 50% AUC reduction for evaluated open-source video detectors relative to their original benchmarks [73].
Generalization strategies operate at three levels. Data-level methods increase diversity through multiple generators, manipulation types, domains, and realistic augmentations [53,62,74]. Model-level methods seek representations that separate transferable forensic evidence from identity, scene, or generator-specific characteristics [75,76,77]. Learning-level methods use domain generalization, meta-learning, continual learning, or test-time adaptation to respond to distribution shift [78,79]. These strategies broaden the conditions represented during training but cannot guarantee coverage of future generators (A3).
4.5. Reasoning-Based and Agentic Detection
Most detectors produce a score through a fixed inference procedure. Vision-language and agent-based systems, instead, formulate forensic analysis as evidence collection, interpretation, and decision-making. Their main potential lies in combining heterogeneous evidence and producing explanations rather than replacing specialized signal detectors.
Vision-Language and Knowledge-Grounded Detection Vision-language models extend forensic analysis from low-level artifacts toward semantic and contextual reasoning. GPT-4V can recognize some visible manipulation cues but remains less accurate than specialized detectors [80], while forensic systems such as FakeShield [81], Veritas [82], and Omni-Fake-R1 [83] combine authenticity assessment with localization or explanation. Such models can identify physical and contextual inconsistencies but may miss subtle statistical artifacts or provide plausible yet unfaithful explanations. They are therefore better treated as complementary evidence rather than definitive forensic judges (A5).
Knowledge-Grounded Forensic Reasoning Forensic adaptation and knowledge grounding can improve the relevance of VLM reasoning. Recent systems combine authenticity assessment with localization and explanation, including FakeShield [81], Veritas [82], and Omni-Fake-R1 [83]. Such approaches connect low-level perception with semantic interpretation and can identify physical or contextual inconsistencies. However, these cues may weaken as generation becomes more coherent, motivating their combination with signal-level and provenance evidence (A5).
Agent-Based Forensic Frameworks Agent-based systems coordinate several models, tools, or reasoning roles. AIFo combines forensic tools, evidence-gathering agents, reasoning components, and structured debate for AI-generated image detection [84]. Agent4FaceForgery uses LLM-based agents to simulate forgery creation and social interaction, producing more realistic multimodal training data [85]. It is therefore better understood as an agent-based data-generation and evaluation framework than as a direct multi-agent detector. Agentic systems require multiple inference calls and may be too expensive for platform-scale screening. Coordinating several tools aggregates their evidence and their assumptions, and the weakest component still sets the floor.
4.6. Adversarial Robustness and Evasion
In this section, we consider deliberate evasion, where an adversary intentionally modifies or suppresses the signals used for detection. Conventional white-box, black-box, and transferable attacks have been extensively studied in seminal literature. Here, we focus instead on two emerging forms of evasion that extend beyond conventional per-sample perturbations.
Learned anti-forensic evasion. Recent methods learn transformations specifically designed to suppress forensic traces in synthetic media. Deep-dithering models, for example, can reduce generative artifacts through a learned image-to-image transformation [86], while GANFR explicitly separates and suppresses GAN fingerprint features in spatial and frequency domains [87]. Such attacks directly challenge A2 by suppressing generator-specific fingerprints and may also challenge A3 when the resulting media falls outside the detector’s training distribution. However, current evidence is concentrated largely on generated images and GAN-based synthesis; the effectiveness of such approaches against modern audio-video and diffusion-based generators remains less established.
Attacks on reasoning-based detectors. The reasoning-based systems discussed in §4.5 introduce an additional attack surface because visual content can also carry natural-language instructions. Recent multimodal prompt-injection attacks show that text embedded within images can redirect the behaviour of vision-language models [88]. In a deepfake-detection setting, overlays, misleading authenticity claims, or other semantic distractors could therefore interfere with how a reasoning-based detector interprets forensic evidence. Agentic systems that retrieve external evidence may further enlarge this attack surface. Robustness evaluation should consequently assess not only prediction accuracy but also whether the model’s reasoning remains consistent with the forensic evidence on which its decision should depend.
Both forms of evasion may interact with platform processing. Transcoding and compression can weaken an adversarial signal, but they may also remove the forensic traces required by the detector. Attack success should therefore be evaluated after platform-representative processing [6], rather than only on the adversarial media produced directly by the attack. Evasion can deliberately suppress generator fingerprints (A2) or shift inputs outside the detector’s training distribution (A3).
4.7. Forensics Mapping
The A1–A5 framework summarizes the conditions underlying current detection methods. Boundary-based approaches depend on compositing evidence (A1); frequency, reconstruction, acoustic, and attribution methods depend on stable generator traces (A2); supervised detectors remain vulnerable to distribution shift (A3); signal-level evidence may be weakened by platform processing (A4); and increasingly coherent generation reduces low-level and cross-modal inconsistencies (A5). No single detector is therefore reliable across all conditions. Table 3 maps representative methods to these assumptions.
5. Deepfakes Across the Social Media Lifecycle
The pipeline in §2.1 covers deepfake creation; here, we examine how distribution, provenance, and remediation affect forensic evidence after content enters social media. These stages account for only 12% of the reviewed corpus, compared with 75% for creation and detection (Figure 3(a)), revealing a substantial gap in research on how forensic evidence survives deployment.
5.1. Distribution I: Platform Upload and Forensic Transformation
Platform processing, including resizing, compression, resampling, and re-encoding, can weaken manipulation traces in both the visual and audio streams. Compression weakens the compositing boundaries associated with A1, while resampling and re-encoding alter the spectral fingerprints associated with A2. These operations also introduce artifacts unrelated to synthesis, which complicates the distinction between generation and distribution traces [6]. Repeated downloading, reposting, cropping, overlays, and screen recording compound these effects. The forensic signal available at detection may therefore differ substantially from that present at generation, challenging A4.
5.2. Distribution II: Algorithmic Propagation and Cross-Platform Reposting
After upload, a deepfake becomes part of a time-varying propagation cascade shaped by sharing, recommendation systems, and cross-platform reposting. We represent the cascade observed up to time t as a directed graph , where nodes represent observed posts and directed edges represent reposting, remixing, or cross-platform transfer. Each repost may undergo transformations such as transcoding, cropping, overlays, or re-encoding. The complete graph is rarely observable because platforms do not expose full provenance or engagement histories. Relationships between posts must therefore be inferred from perceptual matching, audio fingerprints, semantic similarity, or available provenance evidence. Propagation is also shaped by platform ranking, account popularity, network structure, and prior engagement [89]. Propagation-aware evaluation should consider not only whether a deepfake is detected but also how far it spreads. Let denote the earliest observed upload and the detection time. Useful measures include
where denotes the platform hosting post , so that the three quantities represent time to detection, posts observed before detection, and platforms reached, respectively.
5.3. Provenance, Watermarking, and Content Authenticity
Passive detection infers manipulation from the observed media. Provenance instead records evidence at creation or editing time, reducing dependence on forensic traces that may later be weakened by platform processing. We consider signed metadata, in-generation watermarking, and post-hoc watermarking, including their robustness limits and relationship to A1–A5.
Signed Metadata and Content Credentials. The Coalition for Content Provenance and Authenticity (C2PA) specifies cryptographically signed manifests describing an asset, its ingredients, and the actions applied to it [91]. Verification establishes the integrity and signer of the manifest and whether the bound asset remains consistent with the signed record. Two limitations remain. First, provenance is only as trustworthy as its signer. Second, metadata may be stripped or invalidated during re-encoding. C2PA addresses the latter through durable credentials that use perceptual fingerprints or embedded watermarks to recover stripped manifests, thereby inheriting the robustness limits of those signals.
Watermarking During and After Generation. In-generation watermarking modifies the generator or sampling process so that outputs contain a recoverable signal. Representative methods include Stable Signature, which embeds a signature through the decoder [92]; Tree-Ring, which places a structured pattern in the initial noise [93]; and Gaussian Shading, which encodes information in the latent noise [94]. Model-level fingerprinting instead modifies the training process so that the resulting generator reproduces a predefined mark [200]. These approaches can support detection or attribution, but verification may require model-specific keys, decoders, or inversion procedures. They also depend on generator participation: a watermark designed for one model family may not transfer to another.
Post-hoc watermarking embeds a signal into finished media and can therefore be applied independently of the generator. For video, VideoSeal [95] are designed to remain detectable under transformations such as compression, frame loss, cropping, and temporal editing. For audio, AudioSeal [96] provides localized watermark detection, allowing short synthetic segments within otherwise authentic speech to be identified, while WavMark [97] embeds recoverable payloads in short audio windows. Across modalities, these methods trade payload capacity and imperceptibility against robustness to platform processing and resynthesis.
Removal, Forgery, and Evaluation Under Attack. Watermarks face two broad adversarial risks. Removal attacks aim to erase the mark while preserving perceptual quality. Regeneration through diffusion models can remove many imperceptible image watermarks, while optimization-based and universal attacks reveal similar weaknesses [98,99]. Forgery attacks create the opposite failure: an adversary estimates or copies a watermark pattern and applies it to authentic media, potentially creating false attribution [100]. This is particularly concerning when the presence of a watermark is treated as evidence of synthetic generation. Robustness claims should therefore be interpreted relative to a clearly stated attack model. WAVES evaluates image watermarks under distortion, regeneration, and adversarial attacks and reports substantial degradation compared with distortion-only evaluation [101]. Similar stress testing has been extended to audio [201], while evaluation of video watermarks under realistic platform transcoding chains remains comparatively limited.
Relation to the Forensic Assumptions. Signed provenance metadata sits largely outside the assumptions underlying passive forensic detection because its main failure modes are stripping and signer trust. Watermarking, however, remains subject to similar signal degradation. An embedded watermark must survive compression, resizing, cropping, and re-encoding, corresponding directly to A4. Recovery may also depend on a specific decoder, latent space, or generator family, creating a dependency similar to A2. Watermarks and passive forensic traces may therefore be degraded by many of the same platform transformations, particularly compression, resizing, cropping, and re-encoding (A4).
5.4. Platform Governance and Remediation
Deepfake governance is becoming more explicit in regulation. Articles 50(2) and 50(4) of the EU AI Act require synthetic content to be machine-readably marked and detectable, and deepfake content to be disclosed [102]. China similarly requires labeling of deep-synthesis content that may cause public confusion [103], and election-specific restrictions have emerged, as in South Korea’s regulation of AI-generated content in election campaigns [202]. Other measures target harmful applications rather than synthetic content in general. The U.S. TAKE IT DOWN Act addresses non-consensual intimate deepfakes and requires covered platforms to remove reported content within a specified period [104]. Australia criminalizes non-consensual distribution of technologically created or altered sexual material [105], while the UK and Canada have introduced comparable measures [106,107]. Governance is thus moving beyond detection toward disclosure, platform responsibility, and harm-specific remediation.
Marking obligations of this kind bind identifiable providers. Open-weight generators, forked checkpoints, and locally run pipelines may fall outside provider-based marking obligations and can be used to produce targeted non-consensual or political content. The reachable coverage of provenance is therefore bounded by the compliant share of the generator population rather than by the technical robustness of any watermarking scheme, leaving passive detection as the only available evidence for the remainder.
Human review often forms part of downstream platform moderation, yet it is rarely modelled in deepfake-detection research. Human identification accuracy can approach chance under some conditions and varies with content type, familiarity with the depicted person, and presentation conditions [108]. Human review therefore does not automatically resolve uncertainty left by upstream automated systems.
Lifecycle Summary.Figure 5 summarizes how forensic evidence changes across the lifecycle. Generation, platform processing, and reposting can progressively weaken media-internal signals, while provenance, contextual evidence, and human review provide complementary evidence. The lifecycle perspective therefore shifts deepfake analysis from a static-file problem to an evolving-evidence problem: the central question is not only whether content is fake, but what evidence remains when a decision must be made.
6. Datasets and Evaluation
This section reviews the evolution of deepfake benchmarks, evaluation tasks and metrics, and deployment-oriented protocols for stress-testing forensic assumptions.
6.1. Benchmark Evolution
Deepfake benchmarks have evolved alongside generation technology. Early datasets such as FaceForensics++ [47] and Celeb-DF [48] focused primarily on facial manipulation under relatively controlled conditions. Later benchmarks introduced greater manipulation diversity, additional generators, perturbations, localization annotations, and in-the-wild media, enabling evaluation beyond within-dataset binary classification [50].
The benchmark landscape has also expanded beyond visual face manipulation. Audio datasets cover synthetic speech, voice conversion, codec effects, and unseen spoofing attacks, while audio–visual datasets include audio-only, video-only, and joint manipulations, increasingly with temporal localization. Recent benchmarks such as AV-Deepfake1M [203] expand large-scale audio–visual evaluation, while Deepfake-Eval-2024 [73] evaluates contemporary media collected from online sources, exposing substantial gaps between benchmark and real-world performance.
6.2. Evaluation Tasks and Metrics
Evaluation metrics should match the forensic task. Binary detection commonly uses AUC, average precision (AP), accuracy, or equal error rate (EER), particularly for audio spoofing. Fine-grained manipulation classification and generator attribution benefit from class-balanced measures such as macro-F1 when class frequencies are unequal [52]. Spatial localization typically uses intersection over union (IoU) or Dice, whereas temporal localization commonly uses temporal IoU and mean average precision over manipulated segments [45,203].
Metric choice alone, however, does not determine whether an evaluation is realistic. High within-dataset performance may reflect similarity between training and test distributions rather than robustness to unseen generators or deployment conditions [204]. Results are also difficult to compare across studies when training sets, compression levels, sampling procedures, operating thresholds, or test protocols differ. These conditions should therefore be reported alongside headline metrics.
6.3. Recommended Evaluation Protocol and Forensic Assumptions
Deployment-oriented evaluation should extend beyond a single random train–test split. At minimum, evaluation should include an in-domain baseline and cross-dataset or cross-generator testing. Where relevant, robustness tests should introduce platform-representative transformations such as compression, resizing, re-encoding, noise, and cropping. Multimodal systems should additionally be evaluated separately on audio-only, video-only, and joint manipulations to identify modality-specific shortcuts.
The A1–A5 column in Table 4 indicates which forensic assumptions each benchmark can help stress-test. Controlled compositing datasets probe boundary evidence (A1); multi-generator benchmarks test dependence on generator-specific traces and train–test alignment (A2–A3); transformed and in-the-wild media test signal survival under processing (A4); and multimodal or fully synthetic benchmarks test whether low-level and cross-modal evidence remains sufficient as generation becomes more coherent (A5). Robust evaluation therefore requires complementary datasets and protocols rather than reliance on a single benchmark score.
Table 5.
Representative datasets for deepfake detection and evaluation. Rows are selected to capture major shifts in manipulation diversity, modality, localisation, robustness, cross-generator generalisation, and foundation-model-era content. Mod.: I = image, A = audio, V = video, AV = audio–visual. The final column identifies forensic assumptions that each dataset can help stress-test: A1 = boundary availability; A2 = stable generator fingerprints; A3 = train–test similarity; A4 = survival through processing; and A5 = sufficiency of low-level cues.
Table 5.
Representative datasets for deepfake detection and evaluation. Rows are selected to capture major shifts in manipulation diversity, modality, localisation, robustness, cross-generator generalisation, and foundation-model-era content. Mod.: I = image, A = audio, V = video, AV = audio–visual. The final column identifies forensic assumptions that each dataset can help stress-test: A1 = boundary availability; A2 = stable generator fingerprints; A3 = train–test similarity; A4 = survival through processing; and A5 = sufficiency of low-level cues.
![]() |
7. Findings
Applying the A1–A5 coding scheme uniformly across the reviewed generation methods, detectors, platform transformations, and benchmarks yields six findings.
F1. Research emphasizes A3 more than A4. Of the 50 detectors coded in Table 3, 36 explicitly mitigate distribution shift (A3), whereas 39 depend on forensic signals surviving processing (A4) and only 10 mitigate its failure. Five of these ten are reasoning or agentic systems, whose multiple inference or tool calls may limit their suitability for platform-scale screening. Thus, although A4 is widely assumed by existing methods, comparatively few address its failure.
F2. Boundary evidence is becoming less central. In Table 3, 41 of 50 detectors neither depend on A1 nor explicitly mitigate its failure. Table 2 also shows residual boundary evidence declining from High for autoencoder- and GAN-based face manipulation to Low for diffusion, autoregressive, and flow-based full-frame generation. A1 therefore remains most relevant to localized manipulations such as compositing, lip-sync editing, body transfer rather than fully synthetic generation.
F3. Generalization is often evaluated through a narrow cross-dataset shift. Eighteen of the 50 detectors report FaceForensics++ → Celeb-DF v2 as a principal cross-dataset evaluation. Both benchmarks primarily represent pre-foundation-model facial manipulation, so this protocol measures transfer between related face-manipulation distributions rather than open-world generalization to contemporary generators. A3 is therefore the most frequently mitigated assumption, while its evaluation remains comparatively narrow.
F4. Foundation-model content remains under-evaluated. Of the 50 detectors, 34 report no evaluation on diffusion- or foundation-model-generated media, two provide partial evaluation, and 14 explicitly evaluate such content. This represents a gap in the evaluation record rather than evidence that earlier detectors necessarily fail. Existing methods should therefore be re-evaluated on contemporary generation paradigms before performance gaps are attributed solely to architectural limitations.
F5. Lifecycle effects are more often described than quantified. Of the nine entries in Table 4, three report quantitatively measured effects and one has partial evidence. Repost depth and screen recording lack reported degradation curves, and no work in the reviewed corpus reports time to detection, posts observed before detection, or platforms reached. This mirrors the broader corpus imbalance: only 12% of the 212 reviewed works address distribution, provenance, and remediation, compared with 75% covering creation and detection.
F6. Benchmarks emphasize generator and distribution diversity over signal survival. Across the 22 benchmarks in Table 4, generator diversity (A2) is exercised by 21 and train–test dissimilarity (A3) by 20, whereas signal survival under processing (A4) is exercised by 12 and boundary availability (A1) by 13; only three exercise all five assumptions. Current benchmark coverage therefore favors generator and distribution diversity, while substantially fewer benchmarks test whether forensic evidence survives realistic processing. This gap is consistent with the large performance degradation reported on contemporary circulated media [73].
Two consequences follow. First, provenance cannot fully compensate for this evaluation gap because its coverage depends on adoption by generators and platforms; open-weight, forked, or locally operated pipelines may remain outside participating provenance ecosystems regardless of watermark robustness. Second, the findings motivate greater emphasis not only on improving classifiers, but on measuring which forensic signals survive real distribution pipelines and reporting detector performance under those conditions.
8. Discussion
The findings in §7 describe where forensic evidence is lost and where the literature has looked for it. This section states what follows for how detectors are built, evaluated, deployed, and then sets out the questions the survey leaves open.
8.1. Impact of This Study
For method design. The A1–A5 framework replaces the question of whether a detector generalizes with the question of which evidence it requires. A method that depends on a compositing boundary, a stable fingerprint, or an uncompressed signal can state that dependency and be evaluated against the conditions that remove it, rather than against a single cross-dataset score that conflates all five. Combining detectors that share the same failed assumption does not, by itself, resolve that failure; ensembles and agentic pipelines should therefore report whether their components rely on independent evidence.
For evaluation. Benchmark results become comparable only when the conditions that produced them are reported. We recommend that a benchmark declare which of A1–A5 it exercises, and that a detector report its operating threshold, calibration, and performance after platform-representative processing alongside the headline metric. Under the prevalence conditions, a threshold-independent score does not determine deployed precision, so the reporting requirement is not a formality.
For platforms and governance. Treating a deepfake as an evolving evidence trail rather than a file to classify identifies where intervention is still possible. Evidence that generation never produced cannot be recovered downstream, evidence that platform processing removes cannot be recovered by a better classifier, and evidence that provenance records is available only for content whose producer chose to record it. Detection, provenance, disclosure, and remediation are therefore stages of one evidence chain, and a policy that assumes any one of them is sufficient inherits the failures of the others.
Figure 6 sketches what this looks like operationally. The addition is not a new architecture but a record of which evidence a prediction used and whether that evidence was reliable, which is the reporting requirement F1 and F5 imply.
8.2. Open Questions
Q1. How far does a fingerprint transfer across a model lineage (A2)? Attribution and frequency-domain detection assume a stable trace, yet fine-tuning, distillation, multi-stage chaining (§Section 3.5), and anti-forensic transformations each remove it, and both are evaluated almost only on base checkpoints. Attribution should be measured across checkpoint families and generation chains, and allowed to answer unknown. Q2. What evidence remains when nothing is composited (A1, A5)? Fully generated video leaves no authentic region for comparison, and jointly generated audio and video are synchronized by construction, so both cues fail together. The open task is to establish which of the remaining signals, such as decoder bias, temporal dynamics, or physical plausibility, transfer across generator families rather than within one. Q3. What should open-set detection look like (A3)? Generators are released faster than forensic datasets are rebuilt, so unseen architectures are the default deployment case rather than the exception. Generator-agnostic representations, authentic-media modelling, continual adaptation, and open-set recognition all need protocols in which the test-time generator, manipulation type, and domain are absent from training. Q4. How should weak and heterogeneous evidence be combined (A4)? When media-level signals are degraded, a decision must draw on provenance, context, calibrated uncertainty, and human review, but combining sources also combines their assumptions and the weakest one sets the floor. Deployment evaluation should therefore report calibration, abstention, and time to detection rather than accuracy alone. Q5. Can evaluation keep pace with generation? Static benchmarks age as generators are released, and Table 4 shows that existing ones exercise different subsets of A1–A5. A refreshed framework should admit new generators only after detector training and increasingly draw on circulated rather than curated media. Maintaining such evaluation requires continual data collection, annotation, and benchmark updating as generation technology evolves.
8.3. Limitations of This Review
Four scope decisions constrain this review. First, the search covers selected databases and English-language publications, which may under-represent work in other languages or regional venues. Second, the citation threshold applied to non-peer-reviewed work may exclude recent preprints that have not yet accumulated citations, biasing against newer methods apart from the 2025-2026 exceptions described above. Third, the review is video-centric; image-only and audio-only forensic methods appear only where they inform video analysis. Fourth, our treatment of distribution covers publicly observable platforms, so synthetic media circulating through private messaging and closed groups remains outside both this review and the available evidence base.
9. Conclusion
Deepfake detection depends on what evidence survives from generation to distribution. Our A1–A5 framework shows that foundation models weaken traditional forensic traces at creation, while social-media processing further removes them before detection. Modern generators may leave no clear boundary or stable fingerprint (A1, A2), differ from training distributions (A3), and generate synchronized audio and video that weaken cross-modal cues (A5). Distribution then adds compression, resizing, cropping, and reposting effects that further suppress forensic evidence (A4). Yet distribution, provenance, and remediation account for only 12% of the corpus. No single defense is sufficient: passive detectors face evasion and distribution shift, watermarks can be removed, provenance can be stripped or absent, and attribution becomes difficult in multi-stage pipelines. Detectors that perform near ceiling on curated benchmarks can lose roughly half their AUC on circulated media. Future benchmarks should therefore evolve with new generators and clearly state which A1–A5 assumptions they test.
Funding
Resources used in preparing this research were provided, in part, by the Province of Ontario, the Government of Canada through CIFAR, and companies sponsoring the Vector Institute. This research was funded by the EU’s Horizon Europe project AIXPERT (ID 101214389)
References
- Nguyen-Le, H.H.; Tran, V.T.; Nguyen, T.; Le-Khac, N.A. A survey on proactive deepfake defense: Disruption and watermarking. ACM Comput. Surv. 2025, 58, 1–37. [Google Scholar] [CrossRef]
- DataReportal. Global Social Media Statistics. 2025. Available online: https://datareportal.com/social-media-users (accessed on 2026-04-17).
- Statista. Most Popular Social Networks Worldwide as of February 2025, by Number of Monthly Active Users. 2025. Available online: https://www.statista.com/statistics/272014/global-social-networks-ranked-by-number-of-users/ (accessed on 2026-04-17).
- Farid, H. Image forgery detection. IEEE Signal Process. Mag. 2009, 26, 16–25. [Google Scholar] [CrossRef]
- Chernyshev, M.; Baig, Z.; Syed, N.; Doss, R.; Shore, M. Large language models in digital forensics: capabilities, challenges and future directions. Forensic Sci. Int. Digit. Investig. 2026, 56, 302043. [Google Scholar] [CrossRef]
- Montibeller, A.; Shullani, D.; Baracchi, D.; Piva, A.; Boato, G. Bridging the Gap: A Framework for Real-World Video Deepfake Detection via Social Network Compression Emulation. In Proceedings of the Proceedings of the 1st on Deepfake Forensics Workshop: Detection, Attribution, Recognition, and Adversarial Challenges in the Era of AI-Generated Media, 2025; pp. 29–36. [Google Scholar]
- Kingma, D.P.; Welling, M. Auto-Encoding Variational Bayes. In Proceedings of the International Conference on Learning Representations (ICLR), 2014. [Google Scholar]
- Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative Adversarial Nets. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2014; pp. 2672–2680. [Google Scholar]
- Karras, T.; Laine, S.; Aila, T. A style-based generator architecture for generative adversarial networks. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019; pp. 4401–4410. [Google Scholar]
- Ho, J.; Jain, A.; Abbeel, P. Denoising Diffusion Probabilistic Models. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2020. [Google Scholar]
- Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; Ommer, B. High-Resolution Image Synthesis with Latent Diffusion Models. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022; pp. 10684–10695. [Google Scholar]
- Hong, W.; Ding, M.; Zheng, W.; Liu, X.; Tang, J. CogVideo: Large-Scale Pretraining for Text-to-Video Generation via Transformers. In Proceedings of the International Conference on Learning Representations (ICLR), 2023. [Google Scholar]
- Kondratyuk, D.; Yu, L.; Gu, X.; Lezama, J.; Huang, J.; Hornung, R.; Adam, H.; Akbari, H.; Alon, Y.; Biber, V.; et al. VideoPoet: A Large Language Model for Zero-Shot Video Generation. In Proceedings of the Proceedings of the International Conference on Machine Learning (ICML), 2024. [Google Scholar]
- Kerbl, B.; Kopanas, G.; Leimkühler, T.; Drettakis, G.; et al. 3d gaussian splatting for real-time radiance field rendering. ACM Trans. Graph. 2023, 42, 139–1. [Google Scholar] [CrossRef]
- Perov, I.; Gao, D.; Chervoniy, N.; Liu, K.; Marangonda, S.; Umé, C.; Facenheim, C.S.; RP, L.; Jiang, J.; Zhang, S.; et al. DeepFaceLab: Integrated, flexible and extensible face-swapping framework. arXiv 2020, arXiv:2005.05535. [Google Scholar]
- Chen, R.; Chen, X.; Ni, B.; Ge, Y. SimSwap: An Efficient Framework for High Fidelity Face Swapping. In Proceedings of the Proceedings of the 28th ACM International Conference on Multimedia (MM ’20), 2020; pp. 2003–2011. [Google Scholar]
- Baliah, S.; Lin, Q.; Liao, S.; Liang, X.; Khan, M.H. REFace: Realistic and Efficient Face Swapping: A Unified Approach with Diffusion Models. In Proceedings of the Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) arXiv, 2025. [Google Scholar]
- Thies, J.; Zollhöfer, M.; Stamminger, M.; Theobalt, C.; Nießner, M. Face2Face: Real-Time Face Capture and Reenactment of RGB Videos. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2016; pp. 2387–2395. [Google Scholar]
- Siarohin, A.; Lathuilière, S.; Tulyakov, S.; Ricci, E.; Sebe, N. First order motion model for image animation. 2019, Vol. 32. [Google Scholar]
- Tian, L.; Wang, Q.; Zhang, B.; Bo, L. Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions 2024, 244–260.
- Xu, S.; Chen, G.; Guo, Y.X.; Yang, J.; Li, C.; Zang, Z.; Zhang, Y.; Tong, X.; Guo, B. Vasa-1: Lifelike audio-driven talking faces generated in real time. Adv. Neural Inf. Process. Syst. 2024, 37, 660–684. [Google Scholar] [CrossRef]
- Prajwal, K.; Mukhopadhyay, R.; Namboodiri, V.P.; Jawahar, C. A lip sync expert is all you need for speech to lip generation in the wild. In Proceedings of the Proceedings of the 28th ACM international conference on multimedia, 2020; pp. 484–492. [Google Scholar]
- Chan, C.; Ginosar, S.; Zhou, T.; Efros, A. Everybody dance now. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE, 2019; pp. 5932–5941. [Google Scholar]
- Xu, Z.; Zhang, J.; Liew, J.H.; Yan, H.; Liu, J.W.; Zhang, C.; Feng, J.; Shou, M.Z. Magicanimate: Temporally consistent human image animation using diffusion model. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE, 2024; pp. 1481–1490. [Google Scholar]
- Hu, L.; Gao, X.; Zhang, P.; Sun, K.; Zhang, B.; Bo, L. Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. [Google Scholar]
- Wang, T.; Li, L.; Lin, K.; Zhai, Y.; Lin, C.C.; Yang, Z.; Zhang, H.; Liu, Z.; Wang, L. Disco: Disentangled control for realistic human dance generation. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE, 2024; pp. 9326–9336. [Google Scholar]
- Zhu, S.; Chen, J.L.; Dai, Z.; Dong, Z.; Xu, Y.; Cao, X.; Yao, Y.; Zhu, H.; Zhu, S. Champ: Controllable and consistent human image animation with 3d parametric guidance. In Proceedings of the European Conference on Computer Vision, 2024; Springer; pp. 145–162. [Google Scholar]
- Wang, X.; Zhang, S.; Gao, C.; Wang, J.; Zhou, X.; Zhang, Y.; Yan, L.; Sang, N. Unianimate: Taming unified video diffusion models for consistent human image animation. Sci. China Inf. Sci. 2025, 68, 200103. [Google Scholar] [CrossRef]
- Wang, C.; Chen, S.; Wu, Y.; Zhang, Z.; Zhou, L.; Liu, S.; Chen, Z.; Liu, Y.; Wang, H.; Li, J.; et al. Neural codec language models are zero-shot text to speech synthesizers. arXiv 2023, arXiv:2301.02111. [Google Scholar]
- Casanova, E.; Weber, J.; Shulby, C.D.; Junior, A.C.; Gölge, E.; Ponti, M.A. Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone. In Proceedings of the International conference on machine learning. PMLR, 2022; pp. 2709–2720. [Google Scholar]
- Shen, K.; Ju, Z.; Tan, X.; Liu, E.; Leng, Y.; He, L.; Qin, T.; Bian, J.; et al. Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers. Proc. Int. Conf. Learn. Represent. 2024, Vol. 2024, 698–722. [Google Scholar]
- Le, M.; Vyas, A.; Shi, B.; Karrer, B.; Sari, L.; Moritz, R.; Williamson, M.; Manohar, V.; Adi, Y.; Mahadeokar, J.; et al. Voicebox: Text-guided multilingual universal speech generation at scale. Adv. Neural Inf. Process. Syst. 2023, 36, 14005–14034. [Google Scholar] [CrossRef]
- Qian, K.; Zhang, Y.; Chang, S.; Yang, X.; Hasegawa-Johnson, M. Autovc: Zero-shot voice style transfer with only autoencoder loss. In Proceedings of the International Conference on Machine Learning. PMLR, 2019; pp. 5210–5219. [Google Scholar]
- Li, J.; Tu, W.; Xiao, L. Freevc: Towards high-quality text-free one-shot voice conversion. In Proceedings of the ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE, 2023; pp. 1–5. [Google Scholar]
- Wang, X.; Thakker, M.; Chen, Z.; Kanda, N.; Eskimez, S.E.; Chen, S.; Tang, M.; Liu, S.; Li, J.; Yoshioka, T. Speechx: Neural codec language model as a versatile speech transformer. IEEE/ACM Trans. Audio Speech Lang. Process. 2024, 32, 3355–3364. [Google Scholar] [CrossRef]
- Tech, H.P.C.-A.I. Open-Sora: Democratizing Efficient Video Production for All. 2024. Available online: https://github.com/hpcaitech/Open-Sora.
- Google DeepMind. Veo: High-Fidelity Video Generation Model. 2024. Available online: https://deepmind.google/technologies/veo.
- Yang, Z.; Teng, J.; Zheng, W.; Ding, M.; Huang, S.; Xu, J.; Yang, Y.; Hong, W.; Zhang, X.; Feng, G.; et al. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv 2024, arXiv:2408.06072. [Google Scholar]
- Wan, T.; Wang, A.; Ai, B.; Wen, B.; Mao, C.; Xie, C.W.; Chen, D.; Yu, F.; Zhao, H.; Yang, J.; et al. Wan: Open and advanced large-scale video generative models. arXiv 2025, arXiv:2503.20314. [Google Scholar]
- Blattmann, A.; Dockhorn, T.; Kulal, S.; Mendelevitch, D.; Kilian, M.; Lorenz, D.; Levi, Y.; English, Z.; Voleti, V.; Letts, A.; et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv 2023, arXiv:2311.15127. [Google Scholar]
- Guo, Y.; Yang, C.; Rao, A.; Liang, Z.; Wang, Y.; Qiao, Y.; Agrawala, M.; Lin, D.; Dai, B. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. 2023. [Google Scholar] [CrossRef]
- Zhou, S.; Li, C.; Chan, K.C.; Loy, C.C. Propainter: Improving propagation and transformer for video inpainting. In Proceedings of the Proceedings of the IEEE/CVF international conference on computer vision, 2023; pp. 10477–10486. [Google Scholar]
- Wang, Q.; Bai, X.; Wang, H.; Qin, Z.; Chen, A.; Li, H.; Tang, X.; Hu, Y. Instantid: Zero-shot identity-preserving generation in seconds. arXiv 2024, arXiv:2401.07519. [Google Scholar]
- Ye, H.; Zhang, J.; Liu, S.; Han, X.; Yang, W. Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models. arXiv 2023, arXiv:2308.06721. [Google Scholar]
- Cai, Z.; Stefanov, K.; Dhall, A.; Hayat, M. Do you really mean that? content driven audio-visual deepfake dataset and multimodal method for temporal forgery localization. In Proceedings of the 2022 International Conference on Digital Image Computing: Techniques and Applications (DICTA); IEEE, 2022; pp. 1–10. [Google Scholar]
- Cai, Z.; Ghosh, S.; Adatia, A.P.; Hayat, M.; Dhall, A.; Gedeon, T.; Stefanov, K. AV-Deepfake1M: A large-scale LLM-driven audio-visual deepfake dataset. In Proceedings of the Proceedings of the 32nd ACM international conference on multimedia, 2024; pp. 7414–7423. [Google Scholar]
- Rossler, A.; Cozzolino, D.; Verdoliva, L.; Riess, C.; Thies, J.; Nießner, M. Faceforensics++: Learning to detect manipulated facial images. In Proceedings of the Proceedings of the IEEE/CVF international conference on computer vision, 2019; pp. 1–11. [Google Scholar]
- Li, Y.; Yang, X.; Sun, P.; Qi, H.; Lyu, S. Celeb-df: A large-scale challenging dataset for deepfake forensics. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE, 2020; pp. 3204–3213. [Google Scholar]
- Dolhansky, B.; Bitton, J.; Pflaum, B.; Lu, J.; Howes, R.; Wang, M.; Ferrer, C.C. The deepfake detection challenge (dfdc) dataset. 2020. [Google Scholar] [CrossRef]
- He, Y.; Gan, B.; Chen, S.; Zhou, Y.; Yin, G.; Song, L.; Sheng, L.; Shao, J.; Liu, Z. Forgerynet: A versatile benchmark for comprehensive forgery analysis. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021; pp. 4360–4369. [Google Scholar]
- Yan, Z.; Yao, T.; Chen, S.; Zhao, Y.; Fu, X.; Zhu, J.; Luo, D.; Wang, C.; Ding, S.; Wu, Y.; et al. Df40: Toward next-generation deepfake detection. Adv. Neural Inf. Process. Syst. 2024, 37, 29387–29434. [Google Scholar] [CrossRef]
- Khalid, H.; Tariq, S.; Kim, M.; Woo, S.S. FakeAVCeleb: A novel audio-video multimodal deepfake dataset. arXiv 2021, arXiv:2108.05080. [Google Scholar]
- Li, L.; Bao, J.; Zhang, T.; Yang, H.; Chen, D.; Wen, F.; Guo, B. Face x-ray for more general face forgery detection. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020; pp. 5001–5010. [Google Scholar]
- Zhang, R.; Wang, H.; Du, M.; Liu, H.; Zhou, Y.; Zeng, Q. Ummaformer: A universal multimodal-adaptive transformer framework for temporal forgery localization. In Proceedings of the Proceedings of the 31st ACM international conference on multimedia, 2023; pp. 8749–8759. [Google Scholar]
- Yu, N.; Davis, L.S.; Fritz, M. Attributing Fake Images to GANs: Learning and Analyzing GAN Fingerprints. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019; pp. 7556–7566. [Google Scholar]
- Cozzolino, D.; Rössler, A.; Thies, J.; Nießner, M.; Verdoliva, L. Id-reveal: Identity-aware deepfake video detection. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE, 2021; pp. 15088–15097. [Google Scholar]
- Corvi, R.; Cozzolino, D.; Zingarini, G.; Poggi, G.; Nagano, K.; Verdoliva, L. On the Detection of Synthetic Images Generated by Diffusion Models. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023; pp. 1–5. [Google Scholar]
- Ricker, J.; Lukovnikov, D.; Fischer, A. Aeroblade: Training-free detection of latent diffusion images using autoencoder reconstruction error. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE, 2024; pp. 9130–9140. [Google Scholar]
- Liu, W.; She, T.; Liu, J.; Li, B.; Yao, D.; Liang, Z.; Wang, R. Lips are lying: Spotting the temporal inconsistency between audio and visual in lip-syncing deepfakes. Adv. Neural Inf. Process. Syst. 2024, 37, 91131–91155. [Google Scholar] [CrossRef]
- Tak, H.; Patino, J.; Todisco, M.; Nautsch, A.; Evans, N.; Larcher, A. End-to-end anti-spoofing with rawnet2. In Proceedings of the ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE, 2021; pp. 6369–6373. [Google Scholar]
- Jung, J.w.; Heo, H.S.; Tak, H.; Shim, H.j.; Chung, J.S.; Lee, B.J.; Yu, H.J.; Evans, N. Aasist: Audio anti-spoofing using integrated spectro-temporal graph attention networks. In Proceedings of the ICASSP 2022-2022 IEEE international conference on acoustics, speech and signal processing (ICASSP); IEEE, 2022; pp. 6367–6371. [Google Scholar]
- Shiohara, K.; Yamasaki, T. Detecting deepfakes with self-blended images. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022; pp. 18720–18729. [Google Scholar]
- Nguyen, D.; Mejri, N.; Singh, I.P.; Kuleshova, P.; Astrid, M.; Kacem, A.; Ghorbel, E.; Aouada, D. Laa-net: Localized artifact attention network for quality-agnostic and generalizable deepfake detection. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 17395–17405. [Google Scholar]
- Xu, Y.; Liang, J.; Jia, G.; Yang, Z.; Zhang, Y.; He, R. Tall: Thumbnail layout for deepfake video detection. In Proceedings of the Proceedings of the IEEE/CVF international conference on computer vision, 2023; pp. 22658–22668. [Google Scholar]
- Liu, H.; Tan, Z.; Tan, C.; Wei, Y.; Wang, J.; Zhao, Y. Forgery-aware adaptive transformer for generalizable synthetic image detection. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 10770–10780. [Google Scholar]
- Wang, Z.; Bao, J.; Zhou, W.; Wang, W.; Li, H. Altfreezing for more general video face forgery detection. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023; pp. 4129–4138. [Google Scholar]
- Ojha, U.; Li, Y.; Lee, Y.J. Towards Universal Fake Image Detectors that Generalize Across Generative Models. In Proceedings of the CVPR, 2023. [Google Scholar]
- Cui, X.; Li, Y.; Luo, A.; Zhou, J.; Dong, J. Forensics adapter: Adapting clip for generalizable face forgery detection. In Proceedings of the Proceedings of the Computer Vision and Pattern Recognition Conference, 2025; pp. 19207–19217. [Google Scholar]
- Wang, G.; Lin, F.; Wu, T.; Liu, Z.; Ba, Z.; Ren, K. Fsfm: A generalizable face security foundation model via self-supervised facial representation learning. In Proceedings of the 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE, 2025; pp. 24364–24376. [Google Scholar]
- Cai, Z.; Ghosh, S.; Stefanov, K.; Dhall, A.; Cai, J.; Rezatofighi, H.; Haffari, R.; Hayat, M. Marlin: Masked autoencoder for facial video representation learning. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE, 2023; pp. 1493–1504. [Google Scholar]
- Kundu, R.; Xiong, H.; Mohanty, V.; Balachandran, A.; Roy-Chowdhury, A.K. Towards a universal synthetic video detector: From face or background manipulations to fully ai-generated content. In Proceedings of the 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE, 2025; pp. 28050–28060. [Google Scholar]
- Chen, H.; Hong, Y.; Huang, Z.; Xu, Z.; Gu, Z.; Li, Y.; Lan, J.; Zhu, H.; Zhang, J.; Wang, W.; et al. Demamba: Ai-generated video detection on million-scale genvideo benchmark. Sci. China Inf. Sci. 2026, 69, 162103. [Google Scholar] [CrossRef]
- Chandra, N.A.; Lee, H.; Murtfeldt, R.; Qiu, L.; Karmakar, A.; Tanumihardja, E.; Farhat, K.; Caffee, B.; Lee, C.; Choi, J.; et al. Deepfake-eval-2024: A multi-modal in-the-wild benchmark of deepfakes circulated in 2024. arXiv 2025, arXiv:2503.02857. [Google Scholar]
- Wang, A.; Islam, M.; Xu, M.; Ren, H. Curriculum-based augmented fourier domain adaptation for robust medical image segmentation. IEEE Trans. Autom. Sci. Eng. 2023, 21, 4340–4352. [Google Scholar] [CrossRef]
- Yan, Z.; Zhang, Y.; Fan, Y.; Wu, B. Ucf: Uncovering common features for generalizable deepfake detection. In Proceedings of the Proceedings of the IEEE/CVF international conference on computer vision, 2023; pp. 22412–22423. [Google Scholar]
- Cao, J.; Ma, C.; Yao, T.; Chen, S.; Ding, S.; Yang, X. End-to-end reconstruction-classification learning for face forgery detection. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE, 2022; pp. 4103–4112. [Google Scholar]
- Kong, C.; Luo, A.; Bao, P.; Yu, Y.; Li, H.; Zheng, Z.; Wang, S.; Kot, A.C. Moe-ffd: Mixture of experts for generalized and parameter-efficient face forgery detection. IEEE Transactions on Dependable and Secure Computing, 2025. [Google Scholar]
- Tran, V.N.; Kwon, S.G.; Lee, S.H.; Le, H.S.; Kwon, K.R. Generalization of forgery detection with meta deepfake detection model. IEEE Access 2022, 11, 535–546. [Google Scholar] [CrossRef]
- Kim, M.; Tariq, S.; Woo, S.S. Cored: Generalizing fake media detection with continual representation using distillation. In Proceedings of the Proceedings of the 29th ACM International Conference on Multimedia, 2021; pp. 337–346. [Google Scholar]
- Jia, S.; Lyu, R.; Zhao, K.; Chen, Y.; Yan, Z.; Ju, Y.; Hu, C.; Li, X.; Wu, B.; Lyu, S. Can ChatGPT Detect DeepFakes? A Study of Using Multimodal Large Language Models for Media Forensics. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2024; pp. 4324–4333. [Google Scholar]
- Xu, Z.; Zhang, X.; Li, R.; Tang, Z.; Huang, Q.; Zhang, J. Fakeshield: Explainable image forgery detection and localization via multi-modal large language models. Proc. Int. Conf. Learn. Represent. 2025, Vol. 2025, 31186–31216. [Google Scholar]
- Tan, H.; Tan, Z.; Shi, S.; Liu, A.; Song, C.; Zhu, H.; Wang, W.; Wan, J.; Lei, Z.; et al. Veritas: Generalizable deepfake detection via pattern-aware reasoning 2026. 2026, 66420–66477. [Google Scholar]
- Li, T.; Huang, Z.; Wen, H.; He, Y.; Li, X.; Zhu, B.; Duan, W.; Chen, C.; Fu, Z.; Dong, Y.; et al. OMNI-fake: benchmarking unified multimodal social media deepfake detection. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2026; pp. 30299–30311. [Google Scholar]
- Liang, M.; Qu, Y.; Jiang, Y.; Backes, M.; Zhang, Y. From Evidence to Verdict: An Agent-Based Forensic Framework for AI-Generated Image Detection. arXiv 2025, arXiv:2511.00181. [Google Scholar]
- Lai, Y.; Yu, Z.; Wang, J.; Shen, L.; Xu, Y.; Cao, X. Agent4FaceForgery: Multi-Agent LLM Framework for Realistic Face Forgery Detection. arXiv 2025, arXiv:2509.12546. [Google Scholar]
- Xie, H.; Ni, J.; Zhang, J.; Zhang, W.; Huang, J. Evading generated-image detectors: A deep dithering approach. Signal Process. 2022, 197, 108558. [Google Scholar] [CrossRef]
- Lu, Y.; Liu, J.; Zhang, R. GANFR: GAN fingerprint removal network for image anti-forensics. Knowl.-Based Syst. 2025, 114134. [Google Scholar] [CrossRef]
- Nagaraja, N.; Zhang, L.; Wang, Z.; Zhang, B.; Patil, P. Image-based prompt injection: Hijacking multimodal llms through visually embedded adversarial instructions. In Proceedings of the 2025 3rd International Conference on Foundation and Large Language Models (FLLM); IEEE, 2025; pp. 916–922. [Google Scholar]
- Cheng, L.; Guo, R.; Shu, K.; Liu, H. Causal understanding of fake news dissemination on social media. In Proceedings of the Proceedings of the 27th ACM SIGKDD conference on knowledge discovery & data mining, 2021; pp. 148–157. [Google Scholar]
- Jeong, S.; Kim, J.; Sundar, S.S.; Han, J. Multimodal Spatiotemporal Forecasting of Deepfake Propagation on Social Media. Proc. Proc. ACM Web Conf. 2026, 2026, 9656–9664. [Google Scholar]
- Coalition for Content Provenance and Authenticity. Content Credentials: C2PA Technical Specification, Version 2.3. 2026. Available online: https://spec.c2pa.org/specifications/specifications/2.3/specs/C2PA_Specification.html (accessed on 2026-08-09).
- Fernandez, P.; Couairon, G.; Jégou, H.; Douze, M.; Furon, T. The stable signature: Rooting watermarks in latent diffusion models. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE, 2023; pp. 22409–22420. [Google Scholar]
- Wen, Y.; Kirchenbauer, J.; Geiping, J.; Goldstein, T. Tree-rings watermarks: Invisible fingerprints for diffusion images. Adv. Neural Inf. Process. Syst. 2023, 36, 58047–58063. [Google Scholar] [CrossRef]
- Yang, Z.; Zeng, K.; Chen, K.; Fang, H.; Zhang, W.; Yu, N. Gaussian shading: Provable performance-lossless image watermarking for diffusion models. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE, 2024; pp. 12162–12171. [Google Scholar]
- Fernandez, P.; Elsahar, H.; Yalniz, I.Z.; Mourachko, A. Video seal: Open and efficient video watermarking. arXiv 2024, arXiv:2412.09492. [Google Scholar]
- Roman, R.S.; Fernandez, P.; Défossez, A.; Furon, T.; Tran, T.; Elsahar, H. Proactive detection of voice cloning with localized watermarking. arXiv 2024, arXiv:2401.17264. [Google Scholar]
- Chen, G.; Wu, Y.; Liu, S.; Liu, T.; Du, X.; Wei, F. Wavmark: Watermarking for audio generation. arXiv 2023, arXiv:2308.12770. [Google Scholar]
- Zhao, X.; Zhang, K.; Su, Z.; Vasan, S.; Grishchenko, I.; Kruegel, C.; Vigna, G.; Wang, Y.X.; Li, L. Invisible image watermarks are provably removable using generative ai. Adv. Neural Inf. Process. Syst. 2024, 37, 8643–8672. [Google Scholar] [CrossRef]
- Jiang, Z.; Zhang, J.; Gong, N.Z. Evading watermark based detection of ai-generated content. In Proceedings of the Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security, 2023; pp. 1168–1181. [Google Scholar]
- Saberi, M.; Sadasivan, V.S.; Rezaei, K.; Kumar, A.; Chegini, A.; Wang, W.; Feizi, S. Robustness of ai-image detectors: Fundamental limits and practical attacks. Proc. Int. Conf. Learn. Represent. 2024, Vol. 2024, 27500–27526. [Google Scholar]
- An, B.; Ding, M.; Rabbani, T.; Agrawal, A.; Xu, Y.; Deng, C.; Zhu, S.; Mohamed, A.; Wen, Y.; Goldstein, T.; et al. Waves: Benchmarking the robustness of image watermarks. arXiv 2024, arXiv:2401.08573. [Google Scholar]
- European Parliament and Council of the European Union. Regulation (EU) 2024/1689 Laying Down Harmonised Rules on Artificial Intelligence (Artificial Intelligence Act), 2024. Off. J. Eur. Union 50(4).
- Cyberspace Administration of China and Ministry of Industry and Information Technology and Ministry of Public Security. Provisions on the Administration of Deep Synthesis Internet Information Services, 2022. Order No. 12. 2023. [Google Scholar]
- United States Congress. Tools to Address Known Exploitation by Immobilizing Technological Deepfakes on Websites and Networks Act (TAKE IT DOWN Act), 2025. Public Law 119-12, 139 Stat. 55.
- Parliament of Australia. Criminal Code Amendment (Deepfake Sexual Material) Act 2024, 2024. Act No. 78 of 2024. 2 September 2024.
- Parliament of the United Kingdom. Data (Use and Access) Act 2025, 2025. 2025 c. 18, Section 138: creating, or requesting the creation of, a purported intimate image of an adult.
- Department of Justice Canada. Protecting Victims Act: Legislation to Protect Victims and Keep Kids Safe from Predators, 2026. Backgrounder on Bill C-16, including amendments concerning non-consensual sexual deepfakes. received Royal Assent, June 18, 2026. [Google Scholar]
- Groh, M.; Epstein, Z.; Firestone, C.; Picard, R. Deepfake detection by human crowds, machines, and machine-informed crowds. Proc. Natl. Acad. Sci. 2022, 119, e2110013119. [Google Scholar] [CrossRef]
- Wang, X.; Delgado, H.; Tak, H.; Jung, J.w.; Shim, H.j.; Todisco, M.; Kukanov, I.; Liu, X.; Sahidullah, M.; Kinnunen, T.; et al. ASVspoof 5: Crowdsourced speech data, deepfakes, and adversarial attacks at scale. arXiv 2024, arXiv:2408.08739. [Google Scholar]
- Frank, J.; Schönherr, L. Wavefake: A data set to facilitate audio deepfake detection. arXiv 2021, arXiv:2111.02813. [Google Scholar]
- Ma, L.; Xue, Z.; Wang, Y.; Yan, Z.; Xu, J.; Jiang, X.; Yu, H.; Liao, Y.; Bi, Z. Your One-Stop Solution for AI-Generated Video Detection. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2026; pp. 4458–4470. [Google Scholar]
- Tolosana, R.; Vera-Rodriguez, R.; Fierrez, J.; Morales, A.; Ortega-Garcia, J. Deepfakes and beyond: A survey of face manipulation and fake detection. Inf. Fusion 2020, 64, 131–148. [Google Scholar] [CrossRef]
- Mirsky, Y.; Lee, W. The creation and detection of deepfakes: A survey. ACM Comput. Surv. (CSUR) 2021, 54, 1–41. [Google Scholar] [CrossRef]
- Malik, A.; Kuribayashi, M.; Abdullahi, S.M.; Khan, A.N. DeepFake detection for human face images and videos: A survey. Ieee Access 2022, 10, 18757–18775. [Google Scholar] [CrossRef]
- Wang, T.; Liao, X.; Chow, K.P.; Lin, X.; Wang, Y. Deepfake detection: A comprehensive survey from the reliability perspective. ACM Comput. Surv. 2024, 57, 1–35. [Google Scholar] [CrossRef]
- Croitoru, F.A.; Hiji, A.; Hondru, V.; Ristea, N.C.; Irofti, P.; Popescu, M.; Rusu, C.; Ionescu, R.; Khan, F.; Shah, M. Deepfake media generation and detection in the generative ai era: A survey and outlook. ACM Computing Surveys, 2024. [Google Scholar]
- Xie, S.; Qiao, T.; Li, S.; Zhang, X.; Zhou, J.; Feng, G. DeepFake detection in the AIGC era: A survey, benchmarks, and future perspectives. Inf. Fusion 2025, 103740. [Google Scholar] [CrossRef]
- Pei, G.; Zhang, J.; Hu, M.; Zhang, Z.; Wang, C.; Wu, Y.; Zhai, G.; Yang, J.; Tao, D. Deepfake generation and detection: A benchmark and survey. ACM Comput. Surv. 2026, 58, 1–41. [Google Scholar] [CrossRef]
- Hashmi, A.; Shahzad, S.A.; Lin, C.W.; Tsao, Y.; Wang, H.M. Understanding Audiovisual Deepfake Detection: Techniques, Challenges, Human Factors, and Perceptual Insights. IEEE Comput. Intell. Mag. 2026, 21, 38–54. [Google Scholar] [CrossRef]
- Raza, A.; Basit, A.; Amin, A.; Arfeen, Z.A.; Masud, M.I.; Fayyaz, U.; Jumani, T.A. A comprehensive review of deepfake detection techniques: From traditional machine learning to advanced deep learning architectures. AI 2026, 7, 68. [Google Scholar] [CrossRef]
- Google. Generate videos with Veo 3.1 in Gemini API. 2026. Available online: https://ai.google.dev/gemini-api/docs/video (accessed on 2026-07-30).
- Zhang, K.; Zhang, Z.; Li, Z.; Qiao, Y. Joint face detection and alignment using multitask cascaded convolutional networks. IEEE Signal Process. Lett. 2016, 23, 1499–1503. [Google Scholar] [CrossRef]
- Baevski, A.; Zhou, Y.; Mohamed, A.; Auli, M. wav2vec 2.0: A framework for self-supervised learning of speech representations. Adv. Neural Inf. Process. Syst. 2020, 33, 12449–12460. [Google Scholar]
- Esser, P.; Kulal, S.; Blattmann, A.; Entezari, R.; Müller, J.; Saini, H.; Levi, Y.; Lorenz, D.; Sauer, A.; Boesel, F.; et al. Scaling rectified flow transformers for high-resolution image synthesis. In Proceedings of the Forty-first international conference on machine learning, 2024. [Google Scholar]
- Carlini, N.; Farid, H. Evading Deepfake-Image Detectors with White- and Black-Box Attacks. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2020; pp. 2804–2813. [Google Scholar]
- Deepfakes Contributors. Faceswap: Deepfakes Software for All. 2017. Available online: https://github.com/deepfakes/faceswap (accessed on 2026-03-09).
- Li, L.; Bao, J.; Yang, H.; Chen, D.; Wen, F. Advancing high fidelity identity swapping for forgery detection. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020; pp. 5074–5083. [Google Scholar]
- Skorokhodov, I.; Tulyakov, S.; Elhoseiny, M. StyleGAN-V: A Continuous Video Generator with the Price, Image Quality and Perks of StyleGAN2. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022; pp. 3626–3636. [Google Scholar]
- Blattmann, A.; Rombach, R.; Ling, H.; Dockhorn, T.; Kim, S.W.; Fidler, S.; Kreis, K. Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023; pp. 22563–22575. [Google Scholar]
- Ho, J.; Salimans, T.; Gritsenko, A.; Chan, W.; Norouzi, M.; Fleet, D.J. Video Diffusion Models. Proc. Adv. Neural Inf. Process. Syst. (NeurIPS) 2022, Vol. 35, 8633–8646. [Google Scholar] [CrossRef]
- Li, T.; Tian, Y.; Li, H.; Deng, M.; He, K. Autoregressive Image Generation without Vector Quantization. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2024; Vol. 37. [Google Scholar]
- Polyak, A.; Zohar, A.; Brown, A.; et al. Movie Gen: A Cast of Media Foundation Models. arXiv 2024, arXiv:2410.13720. [Google Scholar]
- Milani, S.; Fontani, M.; Bestagini, P.; Barni, M.; Piva, A.; Tagliasacchi, M.; Tubaro, S. An overview on video forensics. APSIPA Trans. Signal Inf. Process. 2012, 1, e2. [Google Scholar] [CrossRef]
- Wang, Y.; Chen, X.; Zhu, J.; Chu, W.; Tai, Y.; Wang, C.; Li, J.; Wu, Y.; Huang, F.; Ji, R. Hififace: 3d shape and semantic prior guided high fidelity face swapping. arXiv 2021, arXiv:2106.09965. [Google Scholar]
- InsightFace. inswapper: Face Swapping Model. 2023. Available online: https://github.com/deepinsight/insightface (accessed on 2026-08-01).
- Nirkin, Y.; Keller, Y.; Hassner, T. FSGAN: Subject Agnostic Face Swapping and Reenactment. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019; pp. 7184–7193. [Google Scholar]
- Rosberg, F.; Aksoy, E.E.; Alonso-Fernandez, F.; Englund, C. FaceDancer: Pose- and Occlusion-Aware High Fidelity Face Swapping. In Proceedings of the Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2023; pp. 3454–3463. [Google Scholar]
- Hong, F.T.; Zhang, L.; Shen, L.; Xu, D. Depth-aware generative adversarial network for talking head video generation. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022; pp. 3397–3406. [Google Scholar]
- Tulyakov, S.; Liu, M.Y.; Yang, X.; Kautz, J. MoCoGAN: Decomposing Motion and Content for Video Generation. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2018; pp. 1526–1535. [Google Scholar]
- Qin, B.; Li, J.; Tang, S.; Chua, T.S.; Zhuang, Y. Instructvid2vid: Controllable video editing with natural language instructions. In Proceedings of the 2024 IEEE International Conference on Multimedia and Expo (ICME); IEEE, 2024; pp. 1–6. [Google Scholar]
- Jiang, Z.; Han, Z.; Mao, C.; Zhang, J.; Pan, Y.; Liu, Y. Vace: All-in-one video creation and editing. In Proceedings of the 2025 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE, 2025; pp. 17191–17202. [Google Scholar]
- Guo, Y.; Chen, K.; Liang, S.; Liu, Y.J.; Bao, H.; Zhang, J. Ad-nerf: Audio driven neural radiance fields for talking head synthesis. In Proceedings of the Proceedings of the IEEE/CVF international conference on computer vision, 2021; pp. 5784–5794. [Google Scholar]
- Ye, Z.; Jiang, Z.; Ren, Y.; Liu, J.; He, J.; Zhao, Z. Geneface: Generalized and high-fidelity audio-driven 3d talking face synthesis. arXiv 2023, arXiv:2301.13430. [Google Scholar]
- Cho, K.; Lee, J.; Yoon, H.; Hong, Y.; Ko, J.; Ahn, S.; Kim, S. Gaussiantalker: Real-time talking head synthesis with 3d gaussian splatting. In Proceedings of the Proceedings of the 32nd ACM International Conference on Multimedia, 2024; pp. 10985–10994. [Google Scholar]
- Bar-Tal, O.; Chefer, H.; Tov, O.; Herrmann, C.; Paiss, R.; Zada, S.; Ephrat, A.; Hur, J.; Liu, G.; Raj, A.; et al. Lumiere: A space-time diffusion model for video generation. In Proceedings of the SIGGRAPH Asia 2024 Conference Papers, 2024; pp. 1–11. [Google Scholar]
- Girdhar, R.; Singh, M.; Brown, A.; Duval, Q.; Azadi, S.; Rambhatla, S.S.; Shah, A.; Yin, X.; Parikh, D.; Misra, I. Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning. arXiv 2024, arXiv:2311.10709. [Google Scholar]
- Cui, J.; Li, H.; Zhan, Y.; Shang, H.; Cheng, K.; Ma, Y.; Mu, S.; Zhou, H.; Wang, J.; Zhu, S. Hallo3: Highly dynamic and realistic portrait image animation with video diffusion transformer. In Proceedings of the 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE, 2025; pp. 21086–21095. [Google Scholar]
- Chen, Z.; Cao, J.; Chen, Z.; Li, Y.; Ma, C. Echomimic: Lifelike audio-driven portrait animations through editable landmark conditions. Proc. Proc. AAAI Conf. Artif. Intell. 2025, Vol. 39, 2403–2410. [Google Scholar] [CrossRef]
- Labs, B.F. FLUX. 2024. Available online: https://github.com/black-forest-labs/flux.
- Kim, K.; Kim, Y.; Cho, S.; Seo, J.; Nam, J.; Lee, K.; Kim, S.; Lee, K. Diffface: Diffusion-based face swapping with facial guidance. Pattern Recognit. 2025, 163, 111451. [Google Scholar] [CrossRef]
- Zhao, W.; Rao, Y.; Shi, W.; Liu, Z.; Zhou, J.; Lu, J. DiffSwap: High-Fidelity and Controllable Face Swapping via 3D-Aware Masked Diffusion. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023; pp. 8568–8577. [Google Scholar]
- Pijarowski, M.; Wala, J.; Tabor, J.; Spurek, P. ImplicitDeepfake: Plausible Face-Swapping through Implicit Deepfake Generation using NeRF and Gaussian Splatting. arXiv 2024, arXiv:2402.06390. [Google Scholar]
- Guo, J.; Zhang, D.; Liu, X.; Zhong, Z.; Zhang, Y.; Wan, P.; Zhang, D. Liveportrait: Efficient portrait animation with stitching and retargeting control. arXiv 2024, arXiv:2407.03168. [Google Scholar]
- Zhang, W.; Cun, X.; Wang, X.; Zhang, Y.; Shen, X.; Guo, Y.; Shan, Y.; Wang, F. Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE, 2023; pp. 8652–8661. [Google Scholar]
- Li, J.; Zhang, J.; Bai, X.; Zhou, J.; Gu, L. Efficient Region-Aware Neural Radiance Fields for High-Fidelity Talking Portrait Synthesis. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2023; pp. 7568–7578. [Google Scholar]
- Mukhopadhyay, R.; Prajwal, K.R.; Namboodiri, V.P.; Jawahar, C.V. Diff2Lip: Audio Conditioned Diffusion Models for Lip-Synchronization. arXiv 2024, arXiv:2308.09716. [Google Scholar]
- Cheng, K.; Cun, X.; Zhang, Y.; Xia, M.; Yin, F.; Zhu, M.; Wang, X.; Wang, J.; Wang, N. Videoretalking: Audio-based lip synchronization for talking head video editing in the wild. In Proceedings of the SIGGRAPH Asia 2022 Conference Papers, 2022; pp. 1–9. [Google Scholar]
- Polyak, A.; Adi, Y.; Copet, J.; Kharitonov, E.; Lakhotia, K.; Hsu, W.N.; Mohamed, A.; Dupoux, E. Speech resynthesis from discrete disentangled self-supervised representations. arXiv 2021, arXiv:2104.00355. [Google Scholar]
- Kong, W.; Tian, Q.; Zhang, Z.; Min, R.; Dai, Z.; Zhou, J.; Xiong, J.; Li, X.; Wu, B.; Zhang, J.; et al. Hunyuanvideo: A systematic framework for large video generative models. arXiv 2024, arXiv:2412.03603. [Google Scholar]
- Ye, F.; Hua, M.; Zhang, P.; Li, X.; Sun, Q.; Zhao, S.; He, Q.; Wu, X. DreamID: High-Fidelity and Fast Diffusion-based Face Swapping via Triplet ID Group Learning. In Proceedings of the ACM SIGGRAPH Asia 2025 Conference Papers, 2025. [Google Scholar] [CrossRef]
- Xu, C.; He, K.; Zhu, J.; Ge, Y.; Li, W.; Wang, C. HiFiVFS: High Fidelity Video Face Swapping. arXiv 2024, arXiv:2411.18293. [Google Scholar]
- Wang, R.; Chen, Y.; Xu, S.; He, T.; Zhu, W.; Song, D.; Chen, N.; Tang, X.; Hu, Y. DynamicFace: High-Quality and Consistent Face Swapping for Image and Video using Composable 3D Facial Priors. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2025; pp. 13438–13447. [Google Scholar]
- Zhao, J.; Zhang, H. Thin-Plate Spline Motion Model for Image Animation. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022; pp. 3657–3666. [Google Scholar]
- Thies, J.; Elgharib, M.; Tewari, A.; Theobalt, C.; Nießner, M. Neural Voice Puppetry: Audio-Driven Facial Reenactment. In Proceedings of the Proceedings of the European Conference on Computer Vision (ECCV), 2020; pp. 716–731. [Google Scholar]
- Bounareli, S.; Tzelepis, C.; Argyriou, V.; Patras, I.; Tzimiropoulos, G. Diffusionact: Controllable diffusion autoencoder for one-shot face reenactment. In Proceedings of the 2025 IEEE 19th International Conference on Automatic Face and Gesture Recognition (FG); IEEE, 2025; pp. 1–11. [Google Scholar]
- Drobyshev, N.; Casademunt, A.B.; Vougioukas, K.; Landgraf, Z.; Petridis, S.; Pantic, M. Emoportraits: Emotion-enhanced multimodal one-shot head avatars. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE, 2024; pp. 8498–8507. [Google Scholar]
- Wang, S.; Li, L.; Ding, Y.; Fan, C.; Yu, X. Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion. In Proceedings of the Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI), 2021. [Google Scholar]
- Zhou, Y.; Han, X.; Shechtman, E.; Echevarria, J.; Kalogerakis, E.; Li, D. MakeItTalk: Speaker-Aware Talking-Head Animation. ACM Transactions on Graphics (Proc. SIGGRAPH Asia), 2020; 39. [Google Scholar]
- Wei, H.; Yang, Z.; Wang, Z. AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animations. arXiv 2024, arXiv:cs. [Google Scholar]
- Wang, C.; Tian, K.; Zhang, J.; Guan, Y.; Luo, F.; Shen, F.; Jiang, Z.; Gu, Q.; Han, X.; Yang, W. V-express: Conditional dropout for progressive training of portrait video generation. arXiv 2024, arXiv:2406.02511. [Google Scholar]
- Jia, Y.; Zhang, Y.; Weiss, R.J.; Wang, Q.; Shen, J.; Ren, F.; Chen, Z.; Nguyen, P.; Pang, R.; Moreno, I.L.; et al. Transfer Learning from Speaker Verification to Multispeaker Text-to-Speech Synthesis. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2018. [Google Scholar]
- Team, K.; Chen, J.; Ci, Y.; Du, X.; Feng, Z.; Gai, K.; Guo, S.; Han, F.; He, J.; He, K.; et al. Kling-Omni Technical Report. arXiv 2025, arXiv:2512.16776. [Google Scholar]
- Runway. Gen-3 Alpha. 2024. Available online: https://runwayml.com/research/gen-3-alpha (accessed on 2026-03-18).
- HaCohen, Y.; Chiprut, N.; Brazowski, B.; Shalem, D.; Moshe, D.; Richardson, E.; Levin, E.; Shiran, G.; Zabari, N.; Gordon, O.; et al. Ltx-video: Realtime video latent diffusion. arXiv 2024, arXiv:2501.00103. [Google Scholar]
- Pika Labs. Pika 2.0. 2024. Available online: https://pika.art/ (accessed on 2026-03-18).
- Deng, J.; Guo, J.; Xue, N.; Zafeiriou, S. ArcFace: Additive Angular Margin Loss for Deep Face Recognition. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019; pp. 4690–4699. [Google Scholar]
- Durall, R.; Keuper, M.; Keuper, J. Watch your up-convolution: Cnn based generative deep neural networks are failing to reproduce spectral distributions. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE, 2020; pp. 7887–7896. [Google Scholar]
- Feng, C.; Chen, Z.; Owens, A. Self-supervised video forensics by audio-visual anomaly detection. In Proceedings of the proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023; pp. 10491–10503. [Google Scholar]
- Yermakov, A.; Cech, J.; Matas, J.; Fritz, M. Deepfake detection that generalizes across benchmarks. In Proceedings of the Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2026; pp. 773–783. [Google Scholar]
- Chung, J.S.; Zisserman, A. Out of time: automated lip sync in the wild. In Proceedings of the Asian conference on computer vision, 2016; Springer; pp. 251–263. [Google Scholar]
- Haliassos, A.; Vougioukas, K.; Petridis, S.; Pantic, M. Lips don’t lie: A generalisable and robust approach to face forgery detection. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021; pp. 5039–5049. [Google Scholar]
- Song, X.; Guo, X.; Zhang, J.; Li, Q.; Bai, L.; Liu, X.; Zhai, G.; Liu, X. On learning multi-modal forgery representation for diffusion generated video detection. Adv. Neural Inf. Process. Syst. 2024, 37, 122054–122077. [Google Scholar] [CrossRef]
- Birla, L.; Saikia, T.; Gupta, P. AVENUE: A novel deepfake detection method based on temporal convolutional network and rPPG information. ACM Trans. Intell. Syst. Technol. 2025, 16, 1–16. [Google Scholar] [CrossRef]
- Yasser, B.; Hani, J.; El-Gayar, S.; Amgad, O.; Ahmed, N.; Ebied, H.M.; Amr, H.; Salah, M. Deepfake detection using efficientnet and xceptionnet. In Proceedings of the 2023 Eleventh International Conference on Intelligent Computing and Information Systems (ICICIS); IEEE, 2023; pp. 598–603. [Google Scholar]
- Nguyen, H.H.; Yamagishi, J.; Echizen, I. Use of a capsule network to detect fake images and videos. arXiv 2019, arXiv:1910.12467. [Google Scholar]
- Qian, Y.; Yin, G.; Sheng, L.; Chen, Z.; Shao, J. Thinking in frequency: Face forgery detection by mining frequency-aware clues. In Proceedings of the European conference on computer vision, 2020; Springer; pp. 86–103. [Google Scholar]
- Zhao, H.; Zhou, W.; Chen, D.; Wei, T.; Zhang, W.; Yu, N. Multi-attentional deepfake detection. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021; pp. 2185–2194. [Google Scholar]
- Liu, H.; Li, X.; Zhou, W.; Chen, Y.; He, Y.; Xue, H.; Zhang, W.; Yu, N. Spatial-phase shallow learning: rethinking face forgery detection in frequency domain. In Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE, 2021; pp. 772–781. [Google Scholar]
- Zhao, T.; Xu, X.; Xu, M.; Ding, H.; Xiong, Y.; Xia, W. Learning self-consistency for deepfake detection. In Proceedings of the Proceedings of the IEEE/CVF international conference on computer vision, 2021; pp. 15023–15033. [Google Scholar]
- Larue, N.; Vu, N.S.; Struc, V.; Peer, P.; Christophides, V. Seeable: Soft discrepancies and bounded contrastive learning for exposing deepfakes. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE, 2023; pp. 20954–20964. [Google Scholar]
- Lin, Y.; Song, W.; Li, B.; Li, Y.; Ni, J.; Chen, H.; Li, Q. Fake it till you make it: Curricular dynamic forgery augmentations towards general deepfake detection. In Proceedings of the European conference on computer vision, 2024; Springer; pp. 104–122. [Google Scholar]
- Zhuang, W.; Chu, Q.; Tan, Z.; Liu, Q.; Yuan, H.; Miao, C.; Luo, Z.; Yu, N. UIA-ViT: Unsupervised inconsistency-aware method based on vision transformer for face forgery detection. In Proceedings of the European conference on computer vision, 2022; Springer; pp. 391–407. [Google Scholar]
- Shao, R.; Wu, T.; Nie, L.; Liu, Z. Deepfake-adapter: Dual-level adapter for deepfake detection. Int. J. Comput. Vis. 2025, 133, 3613–3628. [Google Scholar] [CrossRef]
- Khan, S.A.; Dang-Nguyen, D.T. Clipping the deception: Adapting vision-language models for universal deepfake detection. In Proceedings of the Proceedings of the 2024 International Conference on Multimedia Retrieval, 2024; pp. 1006–1015. [Google Scholar]
- Yan, Z.; Luo, Y.; Lyu, S.; Liu, Q.; Wu, B. Transcending forgery specificity with latent space augmentation for generalizable deepfake detection. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 8984–8994. [Google Scholar]
- Koutlis, C.; Papadopoulos, S. DiMoDif: Discourse modality-information differentiation for audio-visual deepfake detection and localization. arXiv 2024, arXiv:2411.10193. [Google Scholar]
- Ma, L.; Yan, Z.; Guo, Q.; Liao, Y.; Yu, H.; Zhou, P. Detecting ai-generated video via frame consistency. In Proceedings of the 2025 IEEE International Conference on Multimedia and Expo (ICME); IEEE, 2025; pp. 1–6. [Google Scholar]
- Zhang, Y.; Colman, B.; Guo, X.; Shahriyari, A.; Bharaj, G. Common sense reasoning for deepfake detection. In Proceedings of the European conference on computer vision, 2024; Springer; pp. 399–415. [Google Scholar]
- Yu, P.; Fei, J.; Gao, H.; Feng, X.; Xia, Z.; Chang, C.H. Unlocking the capabilities of large vision-language models for generalizable and explainable deepfake detection. arXiv 2025, arXiv:2503.14853. [Google Scholar]
- Yu, N.; Skripniuk, V.; Abdelnabi, S.; Fritz, M. Artificial fingerprinting for generative models: Rooting deepfake attribution in training data. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE, 2021; pp. 14428–14437. [Google Scholar]
- Liu, H.; Guo, M.; Jiang, Z.; Wang, L.; Gong, N.Z. Audiomarkbench: Benchmarking robustness of audio watermarking. Adv. Neural Inf. Process. Syst. 2024, 37, 52241–52265. [Google Scholar] [CrossRef]
- National Election Commission of the Republic of Korea. 90 Days Until National Assembly Elections: Election Campaigns Using Deepfake Restricted, 2024. Discussing Article 82-8 of the Public Official Election Act concerning election campaigns using AI-based deepfake videos.
- Cai, Z.; Kuckreja, K.; Ghosh, S.; Chuchra, A.; Khan, M.H.; Tariq, U.; Gedeon, T.; Dhall, A. Av-deepfake1m++: A large-scale audio-visual deepfake benchmark with real-world perturbations. In Proceedings of the Proceedings of the 33rd ACM International Conference on Multimedia, 2025; pp. 13686–13691. [Google Scholar]
- Zi, B.; Chang, M.; Chen, J.; Ma, X.; Jiang, Y.G. Wilddeepfake: A challenging real-world dataset for deepfake detection. In Proceedings of the Proceedings of the 28th ACM international conference on multimedia, 2020; pp. 2382–2390. [Google Scholar]
- Xiong, X.; Patel, P.; Fan, Q.; Wadhwa, A.; Selvam, S.; Guo, X.; Qi, L.; Liu, X.; Sengupta, R. Talkingheadbench: A multi-modal benchmark & analysis of talking-head deepfake detection. In Proceedings of the 2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV); IEEE, 2026; pp. 4139–4149. [Google Scholar]
- Wang, X.; Yamagishi, J.; Todisco, M.; Delgado, H.; Nautsch, A.; Evans, N.; Sahidullah, M.; Vestman, V.; Kinnunen, T.; Lee, K.A.; et al. ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech. Comput. Speech Lang. 2020, 64, 101114. [Google Scholar] [CrossRef]
- Liu, X.; Wang, X.; Sahidullah, M.; Patino, J.; Delgado, H.; Kinnunen, T.; Todisco, M.; Yamagishi, J.; Evans, N.; Nautsch, A.; et al. Asvspoof 2021: Towards spoofed and deepfake speech detection in the wild. arXiv 2022, arXiv:2210.02437. [Google Scholar]
- Xie, Y.; Lu, Y.; Fu, R.; Wen, Z.; Wang, Z.; Tao, J.; Qi, X.; Wang, X.; Liu, Y.; Cheng, H.; et al. The codecfake dataset and countermeasures for the universally detection of deepfake audio. IEEE Trans. Audio Speech Lang. Process. 2025, 33, 386–400. [Google Scholar] [CrossRef]
- Croitoru, F.A.; Hondru, V.; Popescu, M.; Ionescu, R.T.; Khan, F.S.; Shah, M. Mavos-dd: Multilingual audio-video open-set deepfake detection benchmark. arXiv 2025, arXiv:2505.11109. [Google Scholar]
- Barrington, S.; Bohacek, M.; Farid, H. The deepspeak dataset. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2026; pp. 1893–1902. [Google Scholar]
- Liu, J.; Wang, J.; Hou, S.; Ren, M.; Wu, H.; Ma, L.; Pei, R.; He, Z. Beyond face swapping: A diffusion-based digital human benchmark for multimodal deepfake detection. In Proceedings of the ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE, 2026; pp. 14087–14091. [Google Scholar]
- Xia, S.; Li, P.; Liu, X.; Zhang, D.; Guo, X.; Li, Z. AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMs. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2026; pp. 35416–35426. [Google Scholar]
| 1 |
Anthropic, “How Claude marks AI-generated content,” Claude Support, 2026. |
| 2 | Here, a forensic signal is any regularity that provides evidence about a video’s origin, processing, or manipulation. |
| 3 | Here, the source is the output identity, while the driver provides motion, expression, or speech. |
Figure 2.
Literature screening and selection process.

Figure 3.
Overview of the reviewed literature corpus. (a) Creation and detection account for 75% of the corpus, compared with 13% for distribution, provenance, and remediation combined. (b) Publication sources by venue type; 73% of the corpus is peer-reviewed. (c) Publications per year by primary generation/detection role; the 2026 bar is partial (January–August). (d) Cumulative generation and detection coverage over time. Stage and role labels follow the section in which each work is discussed.
Figure 3.
Overview of the reviewed literature corpus. (a) Creation and detection account for 75% of the corpus, compared with 13% for distribution, provenance, and remediation combined. (b) Publication sources by venue type; 73% of the corpus is peer-reviewed. (c) Publications per year by primary generation/detection role; the 2026 bar is partial (January–August). (d) Cumulative generation and detection coverage over time. Stage and role labels follow the section in which each work is discussed.

Figure 4.
Deepfake video pipeline.

Figure 5.
Deepfake lifecycle and forensic evidence degradation. Generation, platform processing, and distribution can weaken evidence associated with assumptions A1–A5, while provenance, contextual information, and human review provide complementary signals for verification and remediation.
Figure 5.
Deepfake lifecycle and forensic evidence degradation. Generation, platform processing, and distribution can weaken evidence associated with assumptions A1–A5, while provenance, contextual information, and human review provide complementary signals for verification and remediation.

Figure 6.
Proposed evidence-aware detection pipeline.

Table 1.
Comparison of this survey with prior surveys across major deepfake research dimensions. Coverage reflects the primary emphasis of each survey.
Table 1.
Comparison of this survey with prior surveys across major deepfake research dimensions. Coverage reflects the primary emphasis of each survey.
![]() |
Table 2.
Comparison of generative paradigms, methods, manipulation scope, and forensics. A1–A5 indicate detector-side assumptions: ● generally holds, ◐ partially holds, and generally does not. Fewer filled circles indicate a harder detection problem.
Table 2.
Comparison of generative paradigms, methods, manipulation scope, and forensics. A1–A5 indicate detector-side assumptions: ● generally holds, ◐ partially holds, and generally does not. Fewer filled circles indicate a harder detection problem.
![]() |
A1: a manipulated region leaves a detectable boundary. A2: the generator leaves a sufficiently stable statistical fingerprint. A3: test data resemble the detector’s training distribution. A4: forensic signals survive compression and other processing. A5: media-internal low-level clues are sufficient for detection. Evidence (Evid.): usable forensic evidence remaining for a detector, from High (multiple independent traces) to Low (few or none).
Table 3.
Representative generative deepfake methods, their characteristic forensic traces, and whether assumptions A1–A5 typically hold when detecting outputs from each method. The symbols describe the applicability of each assumption to detection, not a capability used by the generation method: ● means the assumption generally holds, ◐ means it holds partially or conditionally, and ○ means it generally does not hold. Para. (paradigm): AE = autoencoder; G = GAN; D = diffusion model; AR = autoregressive model; 3D = 3D-aware neural rendering; Flow = flow matching; DiT = diffusion transformer; Hyb = hybrid.
Table 3.
Representative generative deepfake methods, their characteristic forensic traces, and whether assumptions A1–A5 typically hold when detecting outputs from each method. The symbols describe the applicability of each assumption to detection, not a capability used by the generation method: ● means the assumption generally holds, ◐ means it holds partially or conditionally, and ○ means it generally does not hold. Para. (paradigm): AE = autoencoder; G = GAN; D = diffusion model; AR = autoregressive model; 3D = 3D-aware neural rendering; Flow = flow matching; DiT = diffusion transformer; Hyb = hybrid.
![]() |
![]() |
![]() |
Table 4.
Representative deepfake detection methods organised by methodological family and their relationship to assumptions A1–A5. Results are not directly comparable across differing evaluation protocols.
Table 4.
Representative deepfake detection methods organised by methodological family and their relationship to assumptions A1–A5. Results are not directly comparable across differing evaluation protocols.
![]() |
![]() |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.







