Preprint
Article

This version is not peer-reviewed.

Invisible Image Watermark Attacks: A Survey

Submitted:

19 September 2026

Posted:

20 September 2026

You are already at the latest version

Abstract
Invisible image watermarking supports copyright protection, provenance tracing, and AI-generated content governance. However, deep encoder-decoder and generative watermarks, public image editors, and latent-space reconstruction have transformed watermark attacks into a rapidly evolving landscape beyond conventional distortion-centered taxonomies and evaluation protocols. To the best of our knowledge, this is the first survey dedicated to attacks on invisible image watermarks, covering foundational attacks and representative research from 2018 to 2026 across classical and diffusion-based image watermarks. We organize more than sixty target watermarks into five classes (W1--W5) and introduce a unified multilabel taxonomy of seven attack families: conventional signal-processing, geometric and synchronization, regeneration and purification, adversarial perturbation, semantic editing, latent-space inversion, and forgery and provenance attacks. Our threat model separates attack objectives, attacker knowledge and access capabilities, and fidelity constraints, while treating data conditions as an orthogonal dimension. We provide grading references for attack effectiveness and image fidelity and consider forensic detectability as an additional evaluation dimension. Three findings emerge: robustness does not imply security, apparently stronger watermark classes expose new attack surfaces, and forgery may be more practical than removal. Sequential attack chains remain an open evaluation concern. We conclude with open problems in unified evaluation, forensic detectability, and cryptographic attribution.
Keywords: 
;  ;  ;  ;  

1. Introduction

1.1. Background, Motivation, and Security Implications

With the development of Internet technologies, digital content has become increasingly easy to acquire, store, and distribute. At the same time, copyright protection, provenance tracing, and content governance have become more pressing. Digital watermarking [1,2] is an important mechanism for protecting the copyright of digital images. Invisible image watermarking embeds human-imperceptible identification information into protected images, thereby supporting copyright ownership confirmation and subsequent verification. With the rapid spread of generative artificial intelligence and diffusion models, its role has expanded from traditional copyright protection to AI-generated content (AIGC) governance, provenance tracing, image ownership verification, and model- or platform-level attribution.
Digital watermarking research comprises two closely related perspectives: watermarking schemes and watermark attacks. From the defender’s perspective, watermarking schemes seek to improve imperceptibility and robustness so that embedded evidence remains visually unobtrusive and detectable after common image processing. From the attacker’s perspective, watermark attacks deliberately cause detection or decoding failure, or fabricate watermark evidence and false provenance. Robustness to benign or anticipated distortions therefore does not by itself establish security against adaptive removal, evasion, or forgery. These two perspectives constrain and stimulate each other, driving the development of digital watermarking through an attack-defense arms race.
The generative era has substantially expanded this attack surface. Deep encoder-decoder watermarks [21,22,23,24,25,107] and diffusion or generative watermarks [104,105,106] introduce new dependencies on neural features, VAE representations, sampling trajectories, and latent-space structures. At the same time, public image editors and generative reconstruction tools allow attackers to rewrite images while preserving visible semantics. Diffusion regeneration can wash out hidden residual signals [42,44], VAE encoding and decoding can filter imperceptible evidence, latent-space and semantic editing can bypass pixel-level defenses [65,70,71], and spectral, averaging, or proxy-model attacks can remove or forge watermark evidence with limited visual and semantic degradation [34,37,87]. Provenance systems can also be targeted through substitution or forgery rather than simple watermark removal [87,88].
The consequences extend beyond image-quality degradation. Successful attacks can cause detection failure, destroy recoverable payloads, create false provenance, invalidate copyright evidence, trigger ownership disputes, and misattribute content to another user, model, or platform [42,87,88]. Conventional robustness taxonomies and evaluation protocols centered on common distortions are insufficient for this broader threat surface. Therefore, watermark robustness is not equivalent to watermark security [43,47]. This motivates a systematic reassessment of the security boundary of invisible image watermarking from an attack-centric perspective. Such an assessment also supports attack-driven defense: exposing weaknesses through adversarial analysis provides a principled basis for improving watermark design and restoring balance between attacks and defenses [8].

1.2. Research Landscape and Evolution of Technical Paths

Early taxonomies describe watermark attacks as robustness, presentation, interpretation, legal, geometric, cryptographic, protocol, or unauthorized-manipulation problems [3,4,5,6,7,8]. These classifications primarily reflect an era in which attacks were treated as signal degradation, desynchronization, or unauthorized modification and defenses focused on stronger embedding, geometric invariance, and synchronization correction.
The technical path subsequently shifted from fixed distortions to learned removal and adversarial detector evasion, and then to regeneration, semantic editing, latent-space inversion, and provenance forgery. Deep encoder–decoder watermarks became representative attack targets after 2018 [22,107,119]; diffusion and latent-space watermarks later introduced dependencies on public generative priors and sampling trajectories [104,105,106]. Recent attacks exploit spectral concentration, generative reconstruction, public VAEs, detector boundaries, and transferable watermark evidence under increasingly limited access [34,35,42,44,57,70,71,72,73,87,88]. This evolution motivates the four-level and seven-category taxonomy developed in Section 5. A detailed chronological account is provided in Appendix G.

1.3. Comparison with Existing Surveys

Existing surveys either continue to use taxonomies centered on common distortions [26,27], focus on watermark design, defense-side methods, or systematized knowledge [102,108], or discuss robustness benchmarks themselves [99]. Few take watermark attacks as the main thread or systematically cover both invisible image watermarks and diffusion or AIGC image watermarks, and they generally lack a unified characterization, cross-comparison, and threat model for emerging attacks. Table 1 compares this survey with representative prior surveys along five dimensions: focus and modality, whether attacks are the main theme, whether invisible image watermarks are covered, whether diffusion or AIGC image watermarks are covered, and whether an attack taxonomy or threat model is provided. In the table, ✓ denotes full coverage, ○ denotes partial coverage, and a dash denotes little or no coverage.
Table 1. Comparison Between This Survey and Existing Surveys
Table 1. Comparison Between This Survey and Existing Surveys
Survey Year Focus / Modality Attack-Centered Invisible Image Watermarks Diffusion / AIGC Image Watermarks Attack Taxonomy / Threat Model
Wan et al. [109] 2022 Traditional and learning-based robust image watermarking, including geometric invariance, HDR, DIBR, SCI, and point-cloud watermarking ○ ✓ — ○
Wang et al. [110] 2023 Data hiding based on deep learning, unifying watermarking and steganography, mainly covering image, audio, and video ○ ○ — ○
Luo et al. [111] 2024 Deep image watermarking design, including CNN, INN, and encoder-decoder robustness frameworks ○ ✓ ○ ○
Nguyen-Le et al. [108] 2025 Proactive deepfake defense: disruption and watermarking across visual/audio modalities ○ ○ ○ ✓
Ye et al. [112] 2026 LLM text/code/multimodal watermarking and fingerprinting, including training-, logits-, and token-sampling-based methods ○ — — ○
Cao et al. [113] 2026 AIGC image watermarking design, robustness, and security, with emphasis on diffusion-model watermarks ○ ✓ ✓ ✓
This survey 2026 Attacks on invisible image and diffusion watermarks, covering removal, forgery, evasion, and evaluation ✓ ✓ ✓ ✓
As Table 1 shows, existing surveys either center on the design of robust watermarking schemes [109,110,111], focus on active defense against deepfakes and multimodal scenarios [108], or address watermarking for LLMs and AI-generated content [112,113]. They generally do not take attacks as the main thread. To the best of our knowledge, this paper is the first systematic survey dedicated to invisible image watermark attacks. It covers both traditional invisible image watermarks and watermarks for diffusion-generated or other AI-generated images, and provides a unified attack taxonomy and threat model.

1.4. Main Research Questions

This paper focuses on attacks against invisible image watermarks. Its scope includes attacks on spatial-domain and transform-domain watermarks, deep encoder-decoder watermarks, generative or diffusion watermarks, latent-space attacks, provenance and attribution attacks, and attacks that preserve visual quality and semantic fidelity.
This survey addresses three questions about invisible image watermark attacks. First, how can the increasingly complex attack landscape be organized into seven attack families under four abstraction levels: signal, model, semantic, and attribution? Second, what are the essential differences among attack paradigms in terms of attacker knowledge and victim-system access capabilities, orthogonal data conditions, fidelity constraints, and attack objectives, including removal, evasion, and forgery? Third, how vulnerable are different watermark classes to different attack families, and how should they be objectively evaluated under a multidimensional framework that jointly considers objective-specific attack effectiveness, image fidelity, and forensic detectability [99,100]?

1.5. Main Contributions

This paper provides a systematic survey of attacks against invisible image watermarks. Its main contributions are as follows.
The first contribution is a unified hierarchical taxonomy and threat model. We propose a multilabel hierarchical taxonomy for invisible image watermark attacks, covering seven attack families: conventional signal-processing attacks, geometric and synchronization attacks, regeneration and purification, adversarial perturbation, semantic editing, latent-space inversion, and forgery and provenance attacks. We also establish a unified threat model that clearly distinguishes attack objectives such as removal, evasion, and forgery, as well as their inherent trade-offs [43].
The second contribution is a systematic review of foundational classical baselines together with representative modern attacks published from 2018 to 2026. We cover classical signal-processing, geometric, adversarial, and copy-attack baselines [3,4,5,6,7,8,26,33,77,85], as well as representative attacks in the generative era, including regeneration [42,44], adversarial optimization [57,58,59], latent-space inversion [70,71,72,73,74], semantic editing [65,66], and provenance forgery [87,88,89,93]. We further provide a watermark classification used throughout the paper, denoted as W1–W5, a mapping between attack effectiveness and image-fidelity grades, and a summary table by watermark class for representative works.
The third contribution is an analysis of research gaps and open problems. From the perspectives of unified benchmarks [99,101], forensic detectability [100], and systematized knowledge [102], we summarize open problems in watermark robustness evaluation in the generative era. These include the lack of unified benchmarks, the fact that forgery can be easier than removal in some scenarios, no-box and query-free attacks, sequential attack chains, and anti-forensics. We also distill three core observations: robustness is not equivalent to security [43,47], apparently stronger watermark classes expose new attack surfaces [71,73,87], and forgery can be more practical than removal [87,88]. Sequential attack chains are treated as an open evaluation concern rather than as an established attack result.

1.6. Organization

The rest of this paper is organized as follows. Section 2 presents the watermark foundations and the W1–W5 classification, with detailed class profiles in Appendix A. Section 3 describes the threat model. Section 4 introduces the evaluation principles, while Appendix B provides the complete grading framework and Appendix F details the benchmark studies. Section 5 presents the attack taxonomy. Section 6 analyzes the mechanisms, representative works, limitations, and family-level defense implications of the seven attack families; Appendix C provides detailed quantitative results and method-specific defense considerations, and Appendix D presents the cross-category vulnerability analysis. Section 7 discusses open challenges and future directions, and Section 8 concludes the paper. Appendix E summarizes the unified threat model, attack taxonomy, operational conditions, and scope boundaries, while Appendix G provides the detailed historical evolution.

2. Background and Watermark Classification

2.1. Fundamentals of Digital Watermarking and the Attack-Defense Arms Race

Functionally, a complete invisible image watermarking system usually consists of two components: a watermark embedder and a watermark extractor or detector. Where applicable, a secret key controls the embedder, which writes the watermark signal into the original image and aims to produce a watermarked image that is perceptually indistinguishable from the original to human observers. The extractor or detector then recovers the payload from an image that may have undergone processing or attack, or decides whether the watermark is present. According to the output form, watermarks are divided into zero-bit watermarks and multi-bit watermarks. A zero-bit watermark supports only a present-or-absent decision and usually outputs a detection statistic or p-value, whereas a multi-bit watermark recovers a payload and is commonly evaluated by bit error rate (BER) or bit accuracy. A watermarking method is evaluated along three mutually constrained dimensions: imperceptibility, which requires the visual and semantic differences between the watermarked and original images to be imperceptible and is commonly measured by peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and learned perceptual image patch similarity (LPIPS); robustness, which requires the watermark to remain correctly extractable or detectable after conventional image processing and deliberate attacks; and capacity, which measures the amount of information that can be embedded. As discussed in Section 1, watermarking methods on the defender side and watermark attacks on the attacker side co-evolve through an attack-defense arms race. By attack objective, attacks are divided into removal, evasion, and forgery: removal disables or erases watermark evidence so that the signal can no longer be reliably extracted or detected; evasion induces a false negative from the detector without necessarily erasing the underlying watermark signal; and forgery plants another party’s watermark on unauthorized images or impersonates the claimed provenance. These objectives form the attack-objective dimension of the unified threat model presented in Section 3.

2.2. Overview of Invisible Image Watermark Classes (W1–W5)

Table 2 organizes the invisible image watermarks studied in the attack literature into five classes (W1–W5).
Table 2. Overview of the five categories of invisible image watermarks (W1–W5)
Table 2. Overview of the five categories of invisible image watermarks (W1–W5)
Category Representative Methods Main Embedding Domain Semantics Typical Vulnerabilities (attacks it fears most)
W1 — Classical frequency-/spatial-/statistical-domain watermarks DwtDctSvd, LSB, spread-spectrum (Cox), quaternion Fourier transform, moment-domain (PHT/QPHFM), SVD, QIM, RRWID Pixel/spatial domain, transform domain (DWT/DCT/SVD), moment domain Mostly content-agnostic and nonsemantic; some use content-adaptive placement or strength allocation Removed outright by regeneration/denoising (low-perturbation embedding); geometric desynchronization; spread-spectrum key estimation; multi-image averaging and copy-forgery.
W2 — Deep multi-bit / steganographic watermarks HiDDeN, StegaStamp, RivaGAN, MBRS, CIN, RoSteALS, RAWatermark, SSL Watermarking, TrustMark, SepMark, WAM, UDH, InvisMark, VINE Pixel domain / feature domain / autoencoder latent Mostly content-adaptive and nonsemantic; content-agnostic exceptions include RoSteALS and RAWatermark Decoder-query or white-box adversarial perturbation; transfer attacks; optimization-based adaptive attacks; regeneration; forgery.
W3 — In-generation, fine-tuned, or trigger-based generative watermarks Stable Signature, LaWa, AquaLoRA, PTW, Yu, trigger-based (Safe-SD) Generator / VAE-decoder / LoRA parameters / text trigger Content-adaptive, nonsemantic Decoder-rooted members: VAE replacement, regeneration, or purification with latent estimation; transfer attacks; spectral-magnitude erasure; Counterfeit Extractor forgery.
W4 — Diffusion latent-space / zero-bit / semantic watermarks Tree-Ring, RingID, Gaussian Shading, PRC, WIND, ROBIN, ZoDiac, GaussMarker, SEAL, SFW, DiffuseTrace, WatermarkDM Initial noise / DDIM latent / sampling trajectory / pseudorandom code Mostly content-agnostic but generation-bound; semantic or content-aware variants also exist (mostly zero-bit) Latent-space inversion with public priors; trajectory deflection; multi-image averaging and subtraction; boundary leakage; geometric-phase attacks.
W5 — Geometry-synchronized, physically robust, or cryptographic-attribution watermarks PrintCamera, SSyncOA, GeoWM, Tree-Ring-SynTag, GauShad-SynTag, Robust-Wide, Tree-Ring-CoSDA, GS-CoSDA, MetaSeal Pixel/spatial sync templates; diffusion latent + sync tags; semantic-cryptographic binding Mixed: content-agnostic, content-adaptive, or content-bound, depending on the underlying scheme Potential (largely untested) risks: heavy regeneration, deformation outside the training distribution, novel view synthesis / micro geometric desynchronization, and white-box tampering or denial of service disruption.
The classification considers three dimensions: embedding domain, relation to the generation process, and semanticity. The first four classes primarily describe underlying embedding and generation mechanisms, whereas W5 is a cross-cutting, defense-oriented class. A hardened W5 variant may therefore retain an underlying W1–W4 mechanism while also being discussed under W5. This survey includes only invisible image watermarks and excludes visible watermarks; audio, text, and video watermarks; and pure model intellectual property watermarks. In the semanticity column, content-agnostic means that the payload or detection statistic is unrelated to image semantics; content-adaptive means that the perturbation allocation depends on image or model features but does not encode semantics; and semantic watermarks bind watermark evidence to image semantics. Relation to the generation process is reported separately in the class description and embedding-domain columns.
Throughout the remainder of this survey, W1–W5 refer to the five watermark classes summarized in Table 2. Detailed profiles of these classes are provided in Appendix A.

3. Threat Model

To characterize the attacks discussed in Section 6 in a unified way, this section builds a threat model for invisible image watermark attacks along three core dimensions: attack objective, attacker knowledge and access capability, and fidelity constraint. These dimensions specify what the attacker aims to achieve, what victim-system capability boundary the attacker operates under, and what range of image distortion is acceptable. Data and tool conditions are recorded as orthogonal descriptors, while forensic stealth is treated as an optional additional requirement rather than as image fidelity.

3.1. Formal Setup

Let x denote the original image, m the payload, and k the secret key. The watermark embedder E produces a watermarked image x w by embedding m into x under the key k:
x w = E ( x , m , k ) .
For a zero-bit watermark, m can be treated as an empty payload or a fixed identifier. The detector or extractor D then outputs either a zero-bit detection result or the recovered payload under the key k:
D ( x w , k ) = z , for zero - bit detection , m ^ , for payload extraction ,
where z may be a binary decision, detection score, distance, or statistical p-value, and m ^ denotes the extracted payload.
Let D A denote the attacker’s data condition, which may contain a watermarked input, a clean target or cover image, one or more reference images, paired or unpaired samples, or known messages. Let R A denote independent tools and resources, such as a public surrogate model, that do not themselves provide access to the victim system. An attack may produce an attacked image x ′ , a counterfeit verifier or extractor D ˜ , or both:
( x ′ , D ˜ ) = A atk ( D A ; K , Q , R A ) ,
where A atk denotes the attack algorithm, K denotes the attacker’s knowledge of the victim system, and Q denotes victim-system query or internal access. For an image-only attack, D ˜ = ⌀ ; for a Counterfeit Extractor attack, the image may remain unchanged while D ˜ carries the forged verification behavior.
Because peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and learned perceptual image patch similarity (LPIPS) have different metric directions, this survey uses a distortion function Δ to represent perceptual distortion. The fidelity constraint is written as
Δ ( x ′ , x ref ) ≤ ϵ .
Here, x ref is the content-bearing image that should be preserved, usually x w for removal or evasion and the clean target or cover image for forgery. The constraint applies when the attack produces an image. The function Δ can be instantiated by LPIPS, 1 − SSIM , pixel distance, or another distortion measure, and it can also be equivalently expressed as a lower bound on PSNR or SSIM.
Therefore, an attack can be characterized by the following tuple:
τ = ( O , K , Q , ϵ ; D A , R A , S ) ,
where O denotes the attack objective, K and Q denote the attacker’s knowledge and victim-system access capability, and ϵ specifies the fidelity budget. The descriptors after the semicolon record the orthogonal data condition D A , independent tools or resources R A , and the optional forensic-stealth requirement S.

3.2. Attack Objectives: Removal, Evasion, and Forgery

By their primary effect on detection, extraction, or attribution, attacks have three objectives: removal, evasion, and forgery. Removal and evasion both cause verification failure, whereas forgery causes false acceptance, false attribution, or ownership ambiguity. Under the diffusion-purification assumptions of Saberi et al. [43], detector-level evasion and spoofing errors exhibit a source-specific trade-off; this result should not be generalized to every removal and forgery mechanism.
Removal and evasion reduce the true positive rate, so that a watermarked image is judged as unwatermarked, or drive the extracted payload toward random guessing, reflected by a falling TPR or a bit error rate (BER) approaching 0.5. Removal emphasizes disabling or erasing the watermark signal in the image, whereas evasion emphasizes inducing a false negative from the detector, and in practice the two are often achieved together. Representative mechanisms include regeneration and purification [42,44], adversarial perturbation [57,58], latent-space inversion [70,71], semantic editing [65,66], and geometric desynchronization [77,78,79].
Forgery increases the false positive rate or produces false attribution. It can appear as spoofing, which plants target watermark evidence in an arbitrary clean image so that the image is falsely judged as watermarked or is misattributed to another party [88,89,91,93]. It can also appear as Counterfeit Extractor forgery, in which the attacker trains a Counterfeit Extractor that outputs an attacker-specified watermark from an arbitrary image [94]. A third form is ambiguity or invertibility forgery, which derives a seemingly valid watermark or original image from another party’s watermarked image and thereby creates dual ownership claims that weaken the uniqueness of ownership verification [85,86,92].
The relation between removal and forgery depends on the watermark mechanism and threat model. Saberi et al. [43] establish a lower bound on the sum of evasion and spoofing errors after diffusion purification for low-perturbation watermarking methods, rather than a universal monotonic relation between embedding strength and forgeability. Empirically, generative watermarks can remain difficult to remove under some settings while still admitting low-access forgery [87,88]. This is one reason why watermark robustness is not equivalent to watermark security. For inversion-based semantic watermarks, for example, no-box forgery is achievable from a single reference image and an unrelated surrogate diffusion model [87].

3.3. Attacker Knowledge and Access

By the attacker’s knowledge of and access to the victim watermarking system, the threat model defines four capability profiles. For reporting, an attack is assigned the strongest applicable victim-system profile. Black-box access adds victim-system queries to the no-box profile; gray-box access adds partial victim-system knowledge and may coexist with black-box queries; and white-box access exposes the relevant victim internals. This cumulative reporting rule preserves the four labels while keeping query access, partial knowledge, and independent public tools conceptually distinct.
In a no-box setting, the attacker cannot access the victim detector, extractor, API, model parameters, or secret key and cannot query the target system. This setting uses the weakest knowledge and access assumptions and is closest to open deployment, yet it still supports many strong attacks, including diffusion-based regeneration [42], controllable regeneration from clean noise [44], Deep Image Prior [46], frequency-aware denoising [35], dual-domain natural projection [41], semantic-guided partial regeneration [65], proxy-model variants that remove or forge latent-space watermarks from a single image [70,87], feature-leakage attacks that enable single-image evasion or forgery against learning-based post-hoc watermarks [90], and transferable adversarial attacks that train surrogate models offline without querying the target detector [59].
In a black-box setting, the attacker can query the victim detector or API with input images and receive detection decisions, extracted bits, or generated outputs, but cannot access the victim’s internal parameters or key. Representative attacks include query-based adversarial evasion [57]. The query budget is the key constraint at this level, and some methods seek to reduce the number of required queries. A method that makes no target-system queries is classified as no-box or gray-box according to its remaining victim-system knowledge.
In a gray-box setting, the attacker holds partial knowledge of the victim system, such as the watermarking algorithm type, a public VAE or diffusion backbone known to be part of or closely matched to the victim pipeline, or a reproducible surrogate key, but does not know the complete internal parameters or secret key. Gray-box attacks may also query a victim API; under the cumulative rule, the partial victim knowledge determines the reported profile. An independent public model used only as an external attack tool does not by itself constitute gray-box victim access. Typical examples include attacks on Tree-Ring that query a hosted generator and use a public VAE matched to the victim’s latent representation as a surrogate inversion model [71], and adaptive adversarial removal that queries a victim generator while reproducing a differentiable surrogate key of the target scheme without the true key [58].
In a white-box setting, the attacker has full internal access to the victim components relevant to the attack, such as the encoder, detector, decoder, or generator parameters and gradients. Secret-key access is included only when the evaluated attack exposes or requires it. This is the strongest assumption and is mainly used for security stress testing or adaptive attack evaluation. Representative methods include generator-level removal through pivotal tuning [60] and white-box perturbations against the decoder [57].
Beyond these profiles, paired, unpaired, single-image, or multi-image data; the availability of in-distribution samples or known messages; independent public models; and training requirements are orthogonal data and tool conditions. They should be specified separately for each concrete attack and do not change the victim-access label unless they reveal victim-system knowledge, queries, or internals.
Appendix E consolidates these definitions with the orthogonal data and tool profile, the optional forensic-stealth requirement, and the operational conditions of the seven attack families.

4. Evaluation Metrics

Evaluation of watermark attacks usually involves two independent dimensions: attack effectiveness, which measures whether a watermark is removed, evaded, or forged, and image fidelity, which measures whether the attacked image still preserves acceptable visual quality and semantic consistency. Attack effectiveness and image fidelity are evaluated independently; a successful removal, evasion, or forgery may still incur unacceptable fidelity loss. This section reviews attack-effectiveness metrics and image-fidelity metrics, while Appendix B provides the detailed grading references, scheme-specific decision thresholds, and comparative metric analysis.

4.1. Attack-Effectiveness Metrics

Attack-effectiveness metrics can be divided according to the output form of the watermark. For multi-bit payloads, the main metrics include bit error rate (BER) and bit accuracy (Bit-Acc or BA), for which movement toward 0.5 indicates that extraction is approaching random guessing; values far below or above 0.5 can retain systematically inverted information rather than indicate further destruction. Removal rate (RR) is defined as R R = 1 − 2 | B E R − 0.5 | , so a value approaching 1 indicates near-complete randomization [38]. The general criterion is that the decoded result is no better than random guessing, which means that BA or BER approaches 0.5 [37,44]. For zero-bit or detection-based watermarks, the main metrics include the area under the ROC curve (AUC/AUROC), where movement toward 0.5 indicates loss of separability; a value below 0.5 can indicate inverted score ordering and may still evade a fixed detector, but it does not establish that recoverable separability has disappeared [37,43,71]. Other metrics include TPR at x% FPR (TPR@x%FPR), where a lower value favors the attacker and a value approaching 0 indicates successful removal or evasion [42,44,99,101]; statistical p-values for schemes using the corresponding no-watermark null hypothesis, where threshold crossing can indicate removal or forgery under the source protocol and the significance threshold is paper-specific [65,70]; and the Inverse-Distance score of Tree-Ring, with a threshold of 0.0141 [34]. At the attack level, attack success rate (ASR), evasion rate, and detection rate can further be used for aggregation. Removal or evasion attacks focus on reducing the detection rate from 1.00 to a value close to 0, whereas forgery attacks focus on the proportion of clean images that are falsely judged as watermarked [44,57,73]. Key-recovery attacks are usually measured by normalized correlation (NC), where a higher value indicates stronger key recovery [39]. Detection-based metrics depend on FPR calibration: the same BA value or test statistic may correspond to different success decisions under different FPR settings. Since different watermarking schemes adopt non-identical success criteria, Appendix B summarizes the decision lines of representative schemes in Table B2.

4.2. Image Fidelity and Quality Metrics

Image-fidelity evaluation should combine pixel, structural, and perceptual evidence. PSNR, SSIM, and LPIPS are the principal image-level metrics, while FID is used for relative distributional comparison under a fixed protocol [35,42,46,71]. Geometric attacks additionally require aligned metrics when coordinate displacement would otherwise dominate pixel error [78]. Distributional, semantic, no-reference, and protocol-specific scores do not have universal absolute quality thresholds and should be interpreted by direction and by change from a clean or no-attack baseline. Appendix B provides the complete metric inventory, grading references, and scheme-specific decision lines.
Detailed grading thresholds, scheme-specific decision lines, and comparative metric analysis are provided in Appendix B. Unified benchmarks operationalize these principles under shared attack suites and calibrated operating points. WAVES and WIBE improve cross-method comparison, Erasing the Invisible evaluates unknown and partially known targets, and Forensic Stealth shows that successful removal and high image fidelity do not imply that attacked outputs are forensically indistinguishable [99,100,101,103]. Detailed benchmark profiles and results are provided in Appendix F.

5. Attack Taxonomy

This section organizes the seven attack families reviewed in Section 6 under four abstraction levels: signal, model, semantic, and attribution. The resulting four-level and seven-category taxonomy connects attack mechanisms with target watermark classes and attack objectives, while the historical evolution of these technical paths is reviewed in Section 1.2. The taxonomy also corresponds to the watermark classes W1 to W5 in Section 2 and the threat model in Section 3.

5.1. A Four-Level and Seven-Category Attack Taxonomy

If the watermarking system is viewed as a processing pipeline composed of embedding, transmission, generation or editing, detection, and attribution, attacks can be grouped into four levels according to their main point of intervention. As the abstraction level rises, the attack target extends from low-level image signals to generative models, semantic content, and ownership attribution.
Signal-level attacks act directly on spatial-domain signals, transform-domain signals, or geometric synchronization relationships. They include conventional signal-processing attacks, discussed in Section 6.1, which weaken or erase watermark information through denoising, filtering, compression, quantization, frequency manipulation, spread-spectrum key estimation, statistical estimation, or learning-based denoising. They also include geometric and synchronization attacks, discussed in Section 6.2, which disrupt the spatial alignment required for reliable detection or decoding through rotation, scaling, cropping, random bending, micro-geometric displacement, or novel-view synthesis.
Model-level attacks rely on generative models, detection models, or latent-space modeling mechanisms. They include regeneration and purification attacks, discussed in Section 6.3, which reconstruct the image through diffusion purification, VAE reconstruction, Deep Image Prior, or controllable regeneration from clean noise; adversarial perturbation attacks, discussed in Section 6.4, which optimize against a target or surrogate detector under white-box, query-based, or transfer settings; and latent-space inversion attacks, discussed in Section 6.6, which map images into a diffusion latent space and manipulate inverted latent variables, noise, or sampling trajectories, exploit detection boundary leakage, or fine-tune watermark-bearing decoder parameters to achieve removal, evasion, or forgery.
Semantic-level attacks correspond mainly to semantic editing attacks, discussed in Section 6.5. These methods modify or reorganize image content while preserving visible semantic consistency, using instruction-driven editing, local regeneration, inpainting, or trajectory deflection to weaken or remove watermark evidence.
Attribution-level attacks correspond to forgery and provenance attacks, discussed in Section 6.7. Rather than treating image-signal corruption as their only objective, these attacks target ownership claims, provenance tracing, and attribution credibility through watermark transfer, spoofing, Counterfeit Extractor attacks, invertibility or ambiguity constructions, and removal that enables plagiarism or competing provenance claims.
Table 3 summarizes the correspondence among the four abstraction levels, seven attack families, evaluated watermark classes, attack objectives, and representative studies.
Table 3. Overview of the Four-Level × Seven-Category Attack Taxonomy
Table 3. Overview of the Four-Level × Seven-Category Attack Taxonomy
Level Attack Category (Detailed Section) Evaluated Target Watermark Classes Attack Objective Representative References
Signal-level Conventional signal-processing attacks (Section 6.1) W1–W4, mainly W1/W2 Removal / forgery [5,34,35,36,37,38,39]
Signal-level Geometric and synchronization attacks (Section 6.2) W1–W4 Removal / evasion [26,33,77,78,79,83,84]
Model-level Regeneration and purification attacks (Section 6.3) W1–W4 Removal [40,41,42,43,44,45,46,47,48,49,51,52,53,54,55]
Model-level Adversarial perturbation attacks (Section 6.4) W1–W4 Removal / evasion [31,57,58,59,60,61]
Model-level Latent-space inversion attacks (Section 6.6) W1–W4, mainly W3/W4 Removal / evasion / forgery [70,71,72,73,74]
Semantic-level Semantic editing attacks (Section 6.5) W1–W4 Removal [65,66,67,68]
Attribution-level Forgery and provenance attacks (Section 6.7) W1–W4 Forgery / evasion / removal [85,86,87,88,89,90,91,92,93,94,98]

6. Attack Methods: Mechanisms, Evaluation, and Defense Implications

This section expands on the seven mechanisms introduced in Section 5. Each category analyzes the attack mechanism and exploitable weakness, summarizes representative works and attack settings, and provides a survey-level synthesis of strengths, limitations, target watermark classes, and threat-model significance. Family-level defense implications remain in the main text, while detailed quantitative results, fidelity measurements, and method-specific defense considerations are provided in Appendix C. Throughout this section, the target watermark classes are annotated under a unified taxonomy, including W1 classical frequency-/spatial-/statistical-domain watermarks, W2 deep multi-bit / steganographic watermarks, W3 in-generation, fine-tuned, or trigger-based generative watermarks, W4 diffusion latent-space, zero-bit, or semantic watermarks, and W5 geometry-synchronized, physically robust, or cryptographic-attribution watermarks. Attack effectiveness and image fidelity are judged according to the metric grading scheme in Section 4.

6.1. Conventional Signal-Processing Attacks

Conventional signal-processing attacks emerged early and generally have a low implementation barrier. Rather than requiring full knowledge of the watermarking algorithm, they usually treat a watermarked image as the superposition of a host image and an estimable residual. Statistical estimation, spectral erasure, frequency-aware denoising, content-agnostic multi-image averaging, key estimation, and learning-based denoising then weaken or filter watermark evidence in the pixel domain, frequency domain, or statistical feature space. Geometric and synchronization attacks instead disrupt spatial alignment and are discussed separately in Section 6.2.

6.1.1. Statistical Estimation and Perceptual Remodulation Removal Attacks

These attacks estimate the cover image with a maximum a posteriori or minimum mean square error estimator and remodulate the residual with the opposite sign. A noise visibility function places the residual in edge and texture regions, exposing a matched-filter decoder to non-Gaussian noise. Under the power-spectrum condition, cover-image estimation error is maximized when the watermark spectrum follows the cover spectrum. This construction exploits the statistical estimability of an additive, content-agnostic watermark from a single-image residual without the key.
Voloshynovskiy et al. [5] (Optimal Estimation, SP 2001) present this attack within a second-generation watermarking benchmark. The denoising and remodulation components are discussed here, while copy and desynchronization are assigned to Section 6.7 and Section 6.26.2, respectively. The attack uses weighted PSNR and Watson Total Perceptual Error as visibility constraints.
This no-box historical baseline shows that perceptually constrained cover estimation can expose reusable W1 watermark residuals without detector access; its scheme-specific BER and fidelity results are reported in Appendix C.1.

6.1.2. Global Spectral-Magnitude Erasure Attacks

Robust watermark detectors may rely on characteristic energy patterns in the global spectrum, especially in Fourier magnitude. Spectral erasure attacks optimize separate perturbations for high- and low-frequency magnitudes, combine them with cropping, and constrain the result with a perceptual budget. High-frequency components emphasize fine textures, whereas low-frequency components carry more structural and semantic information. The vulnerability arises when reliable detection depends on global spectral-magnitude patterns, including patterns in semantically relevant frequency bands.
Kassis and Hengartner [34] (UnMarker, S&P 2025) require no detector queries or auxiliary training data and have no access to model internals or the secret key; the attack is therefore no-box in this survey. UnMarker optimizes high- and low-frequency magnitudes in two stages and combines the resulting perturbations with cropping under a perceptual constraint. This route targets both non-semantic watermark signals and schemes whose evidence remains tied to semantically relevant spectral bands.
UnMarker shows that global spectral structure can expose a shared attack surface across W2–W4, although its optimization cost, visible artifacts, and reduced applicability to regionalized or object-bound evidence limit practicality. Detailed target-specific results and runtime are reported in Appendix C.1.

6.1.3. Frequency-Aware Denoising and Corrupt-and-Restore Attacks

The energy of post-hoc and fine-tuned watermarks is often concentrated in high-frequency edge and texture regions. MarkSweep first amplifies this signal with edge-aware Gaussian perturbations and then suppresses it with learnable band decomposition and frequency-aware fusion denoising. Box-Blur instead follows a corrupt-and-restore route that combines box blurring with deblurring. Neither image-only route requires decoder access. Both exploit watermark evidence that remains concentrated in separable high-frequency components; when the generator is available, a separate overwriting variant can also retrain it to replace the original watermark.
Cao et al. [35] (MarkSweep, ICASSP 2026) propose a no-box, sub-second two-stage attack trained only with clean and AI-generated images. Their analysis uses the data processing inequality and the Fano bound to motivate information suppression, while the experiments show reduced bit accuracy at roughly one thousandth of the runtime of UnMarker. Wu et al. [36] (Box-Blur / When There Is No Decoder, ICICS 2025) present three strategies: adding noise to edge regions, applying box blur with k = 9 followed by FFTformer deblurring, and fine-tuning a Stable Diffusion generator with a surrogate decoder to overwrite the original watermark.
The image-only variants provide efficient no-box removal of W2 and W3 post-hoc watermarks but are weaker against in-generation or distribution-preserving W4 schemes; the generator-overwriting variant instead requires white-box access. Detailed effectiveness, fidelity, and runtime results are reported in Appendix C.1.

6.1.4. Content-Agnostic Multi-Image Averaging Attacks

Under the same key, content-agnostic watermarks often superimpose an approximately fixed pattern across different images. Averaging a batch of same-key watermarked images can therefore estimate this pattern. Subtracting the estimate can remove the watermark, while adding it to clean images can achieve forgery. The exploitable weakness is the low per-image variance of the embedded pattern under a shared key.
Yang et al. [37] (Averaging, NeurIPS 2024) present a no-box attack whose experiments reveal a clear contrast between content-agnostic and content-adaptive schemes. Detection or extraction can approach random performance for content-agnostic schemes, whereas content-adaptive schemes generally resist pattern estimation more effectively. The same estimated pattern supports removal by subtraction and forgery by addition.
This no-box route supports both removal and forgery across W1–W4 but generally requires hundreds to thousands of same-key images and is less reliable for content-adaptive or strongly image-bound embedding. Detailed AUC, detection, and fidelity results are reported in Appendix C.1.

6.1.5. Secret-Carrier and Equivalent-Key Estimation Attacks in the Known-Message Attack Setting

These attacks estimate a secret carrier or equivalent key under a known-message attack (KMA) setting, which specifies available watermark-message observations rather than model access. Two-Stage SS casts key recovery as binary classification, optimizes a sigmoid-relaxed decoder, and removes the watermark from discrete wavelet transform sub-bands with a quaternion U-shaped network. EK Estimation samples unit vectors, retains those that decode the known message within a prescribed error range, and averages them to recover the carrier direction. Both methods exploit the linear estimability of additive spread-spectrum embedding and information leakage from carrier reuse.
You and Zhou [38] (Two-Stage SS QCNN, TMM 2024) first estimate the key and then use a quaternion U-shaped convolutional network with a removal-rate loss over wavelet sub-bands, prioritizing image fidelity during removal. You et al. [39] (EK Estimation, TMM 2023) use Monte Carlo averaging over equivalent keys to blindly recover the carriers of additive and improved spread-spectrum schemes, outperforming FastICA and maximum likelihood estimation.
Because both studies assume knowledge of the named target spread-spectrum embedding family, they are reported as gray-box under the access taxonomy in Section 3; the paired watermarked-signal/message observations are an orthogonal KMA data condition. Two-Stage SS demonstrates partial removal, whereas EK Estimation shows that repeated known-message observations can expose an equivalent carrier. Both remain limited by strong data requirements and applicability to classical linear embedding; the quantitative recovery and fidelity evidence is reported in Appendix C.1.
Defense implications. Overall, conventional signal-processing attacks exploit statistically estimable residuals, predictable spectral concentration, cross-image watermark consistency, and reusable secret carriers. Defenses should therefore combine content-adaptive or per-image randomized embedding, cryptographically seeded key diversification, multi-band redundancy, and stronger coupling between watermark evidence and image content or distribution-preserving latent structures. Detailed quantitative results, fidelity measurements, and method-specific defense considerations are provided in Appendix C.1.

6.2. Geometric and Synchronization Attacks

Geometric and synchronization attacks target the coordinate assumptions used before watermark detection or decoding. Global affine transformations, local random bending, micro-geometric displacement, latent viewpoint changes, and object relocation can misalign correlation peaks, phase patterns, or marked regions while largely preserving image content. Their impact differs across watermark families because spatial redundancy, binding during generation, and explicit synchronization provide different forms of protection. The discussion groups these attacks into five mechanisms and draws on classical attacks [33,77], modern studies of attack surfaces and limits [78,79,83,84], geometric robustness evaluations in GeoWM, SynTag, and PrintCamera [80,81,82], and the historical survey by Licks and Jordan [26].

6.2.1. Classical Local Random-Bending Desynchronization

Correlation and matched-filter detectors assume that embedded signals remain aligned with the image coordinate system. Classical attacks violate this assumption through nonlinear resampling, random bending, jitter, mosaics, or smooth displacement fields. These operations shift correlation peaks or fragment marked regions, causing detection failure without physically erasing the payload.
Petitcolas et al. [33] introduce StirMark, jitter, and mosaic attacks against first-generation systems. Licks and Jordan [26] later organize geometric attacks within a common synchronization perspective. Barni et al. [77] extend local bending through local permutation with cancelation and duplication (LPCD) and the Markov field desynchronization attack (MF-DA). These geometric variants use a no-box setting, require no attack training, and operate on one watermarked image without keys, detector access, or queries.
These attacks expose a broad W1 synchronization weakness but do not necessarily destroy the payload; registration or spatial redundancy may restore extraction, and early-system results do not transfer uniformly to modern watermark families. The source-specific detection result is reported in Appendix C.2.

6.2.2. Micro-Geometric Phase-Perturbation Desynchronization

Latent-space and phase-based watermarks are not inherently invariant to geometric displacement. Micro-geometric perturbation introduces small spatial shifts that disrupt the phase coherence of Fourier rings or latent patterns without reconstructing the image. The exploitable weakness is the detector’s dependence on precise phase and coordinate correspondence.
Kong et al. [78] propose MarkCleaner, which uses Frequency Band Masking (FBM), Spatial Random Masking (SRM), a visual encoder, and a 2D Gaussian splatting decoder to render a slightly displaced output. The model is trained with geometrically perturbed targets on 5,000 MS-COCO images. At inference, it operates in a no-box setting on one suspicious image, without paired clean and watermarked examples, detector access, watermark knowledge, or detection queries.
MarkCleaner weakens a broad W1–W4 target set, but its deliberate micro-displacement means that the output is alignable rather than pixel-exact. Its evaluation is also limited in generator and resolution coverage; detailed effectiveness and throughput results are reported in Appendix C.2.

6.2.3. Latent-Space Viewpoint Transformation

Diffusion latent-space, zero-bit, and semantic watermarks often lack viewpoint invariance. A viewpoint transformation attack partially inverts the image, changes its spatial coordinates in latent space, and regenerates a coherent alternative view. This process preserves scene semantics while breaking the pixel and latent alignment on which watermark detection depends.
Shamshad et al. [79] propose RAVEN, which combines partial DDIM inversion, latent viewpoint modulation, view-guided correspondence attention, and CIELAB color and contrast transfer. It uses a frozen public diffusion model and requires no attack training, watermark knowledge, detector access, or queries. The no-box attack processes one watermarked image at a time and does not require the original prompt or multiview data.
RAVEN is strongest against W4 semantic watermarks and also weakens several W1–W3 and multi-bit W4 schemes, although some payloads remain recoverable. Viewpoint strength creates a tradeoff between suppression and structural fidelity; detailed TPR, bit-accuracy, and runtime results are reported in Appendix C.2.

6.2.4. Object-Relocation Crop-Paste Desynchronization

A globally embedded watermark may lose synchronization when a protected object is cropped, translated, rotated, or scaled and then pasted into another image. The attack changes the position and pose of the marked region, making evidence tied to the global coordinate system difficult to locate and decode. The underlying weakness is the absence of object-level synchronization.
Zhao et al. [84] study this crop-paste setting while introducing SSyncOA, a defense that segments the marked object, normalizes its geometry, and embeds the payload within the aligned object region. The attack uses a no-box setting and requires no training; it needs a watermarked source image and a destination image but no key, detector access, or queries. SSyncOA is trained separately as a defensive watermarking system.
The setting captures realistic object relocation and compositing, but the evidence is limited to single-object operations and does not establish robustness under multiple interacting objects or full scene reconstruction. The comparative bit-accuracy results are reported in Appendix C.2.

6.2.5. Coding-Theoretic Limits of Geometric Desynchronization

Some cryptographic watermarks encode pseudorandom codewords in latent symbol signs and recover them through error correction. Geometric operations such as cropping and resizing can corrupt these signs before decoding. Once the corruption rate crosses the code’s admissible region, robust verification and soundness cannot both be maintained, which yields a coding-level limit rather than a specific attack network.
Francati et al. [83] prove that the critical corruption rate is α * = 1 − 1 / q for a q-ary alphabet and α * = 1 / 2 for binary codewords. They also show that crop and resize can drive a PRC image watermark toward this boundary. The analysis uses a no-box setting with one watermarked image and requires neither attack training, detector queries, nor key access.
The result establishes a sharp boundary for W4 PRC watermarks: binary decoding fails as the pre-decoding symbol error rate approaches 50%. Its limitation is equally important. The theorem does not provide a general constructive attack with standardized fidelity control, so it should be interpreted as an impossibility boundary rather than a measured tradeoff between image fidelity and attack effectiveness.
Defense implications. Overall, geometric and synchronization attacks exploit dependence on global registration, phase alignment, and evidence tied to fixed locations. Defenses should combine explicit registration or geometric normalization, local or object-level redundancy, transformation-aware training, and coding margins that remain valid under coordinate drift.
Detailed quantitative results, aligned fidelity measurements, and method-specific defense considerations are provided in Appendix C.2.

6.3. Regeneration and Purification Attacks

Many regeneration and purification attacks first corrupt a watermarked image and then reconstruct it with a generative prior. Core methods inject bounded noise into the image or its latent representation before applying a VAE, diffusion model, or learned reconstruction network, so that reconstruction attenuates the watermark signal. The broader category also includes quality-preserving random walks, re-watermarking, and learned networks that do not follow a strict two-stage pipeline. These methods exploit reconstruction instability, repeated content rewriting, localized watermark energy, or interference between watermark signals. The central trade-off concerns the perturbation budget: low-perturbation post-hoc watermarks are often removed by ordinary purification, whereas high-perturbation or diffusion-native watermarks may require regeneration from clean noise, targeted reconstruction, or overwriting. This section reviews 15 attack studies in six categories. TrustMark [50] and RRWID [56] are included as defensive schemes and robust targets.

6.3.1. Bounded-Noise Regeneration and Diffusion Purification

Post-hoc watermarks are often embedded as low-amplitude perturbations. Bounded Gaussian noise can mask these perturbations in the image or latent representation, after which a VAE or diffusion model projects the corrupted input back toward the natural-image manifold. The reconstruction preserves dominant visual content but may not preserve the weaker watermark signal. This mechanism exploits the tension between watermark invisibility, which favors a small embedding budget, and resistance to quality-preserving reconstruction.
Zhao et al. [42] (WatermarkAttacker, NeurIPS 2024) formalize this no-box setting and prove that watermarks within a finite ℓ 2 budget are removable under their assumptions. Saberi et al. [43] (Diffusion Purification, ICLR 2024) analyze the limits of detection-based watermarks and evaluate no-box diffusion purification; their separate substitute-model and spoofing routes use stronger access assumptions and are not treated as purification mechanisms here. Shamshad et al. [53] (First-Place Solution, NeurIPS 2024 challenge) combine adaptive VAE fine-tuning, test-time VAE optimization, and color and contrast restoration. The competition’s official named-target track is treated as gray-box in this survey because the watermark family is known, while its black-box track permits only external observations and limited leaderboard queries.
Bounded-noise regeneration is reliable against low-perturbation W1 and W2 watermarks, but stronger W2 signals and diffusion-native W4 schemes expose its limits. Increasing the corruption strength can improve removal at a higher fidelity cost; detailed removal and detection results are reported in Appendix C.3.

6.3.2. Controllable Regeneration from Clean Noise

Controllable regeneration does not denoise the watermarked image directly. It starts from clean Gaussian noise and reconstructs the image under semantic and structural constraints, thereby replacing both pixel-level and latent watermark evidence while preserving recognizable content. The exploitable weakness is that a watermark attached to the original image or sampling trajectory is unlikely to survive when the generation process is restarted from an independent noise sample.
Liu et al. [44] (CtrlRegen, ICLR 2025) use a DINOv2 semantic adapter and a Canny-edge ControlNet to guide diffusion from clean noise. The attack is no-box with respect to the watermarking system and requires only the watermarked image and public reconstruction components; it does not query the target detector or use its parameters.
CtrlRegen extends regeneration to W1–W4, including Tree-Ring, but complete redrawing weakens pixel-level correspondence even when semantic and no-reference quality remain high. Detailed TPR and bit-accuracy results are reported in Appendix C.3.

6.3.3. Region- and Band-Targeted Reconstruction Attacks

Watermark energy may concentrate in low-saliency regions or particular frequency bands rather than being distributed uniformly. Region-targeted methods inject noise according to a saliency mask before reverse diffusion, whereas band-targeted methods exploit the spectral bias of an untrained convolutional network and stop reconstruction before high-frequency details are fitted. Both routes exploit localization in the watermark signal and seek to preserve portions of the image that contribute most to perceived content.
Alam et al. [45] (SADRE, WWW 2025) apply region-adaptive latent noise and reverse diffusion under a saliency mask. Liang et al. [46] (DIP, TMLR 2025) optimize an untrained deep image prior on a single watermarked image and use early stopping to suppress high-frequency evidence. Both are no-box attacks with no paired-data requirement; SADRE uses a public diffusion prior, while DIP performs per-image optimization without a pretrained remover.
Targeted reconstruction can reduce distortion, but effectiveness depends on the watermark’s spatial and spectral placement. SADRE provides weakening without establishing detector evasion, whereas DIP is strongest on high-frequency W1 and W2 schemes and weaker on other embedding domains. Detailed results are reported in Appendix C.3.

6.3.4. Quality-Preserving Random Walks and Impossibility

A quality-preserving random walk repeatedly rewrites content while consulting a quality oracle and a perturbation oracle. In the image setting, diffusion inpainting supplies the rewriting operation. The process is scheme-agnostic and key-free, and it can gradually wash out watermark evidence while each individual step remains within the permitted quality region. The resulting impossibility statement is conditional on these oracle and random-walk assumptions rather than a universal claim about all image watermarks.
Zhang et al. [47] (WiTS, ICML 2024) establish this theoretical result and instantiate it for image watermarks with a model weaker than the attacked model. The attack is no-box with respect to the watermarking system; access to the quality and perturbation oracles is a separate operational assumption and does not imply detector or key access.
WiTS identifies a conditional security boundary rather than a uniformly decisive empirical attack. Its image results support gradual removal under the stated oracle assumptions, but their practical force depends on the availability and reliability of those oracles; the scheme-specific p-values are reported in Appendix C.3.

6.3.5. Re-Watermarking Attacks

Re-watermarking places a second watermark over an already watermarked image. The new signal may interfere with extraction of the original payload or overwrite evidence used by a zero-bit detector. For Stable Signature [W3] and Tree-Ring [W4], the attacker can use ZoDiac [W4] as the overwriting scheme. This route exploits watermark superposability and the absence of a mechanism that distinguishes the registered watermark from a later embedding operation.
Bulychev et al. [48] (Re-Watermarking, arXiv 2026) first classify the victim scheme and then re-embed a watermark from the same method or a compatible family. The end-to-end pipeline is no-box with respect to the victim detector and key. At inference, it requires one watermarked image and access to public or attacker-controlled embedding methods; its scheme classifier is trained offline on labeled watermarked images from the candidate schemes.
Re-watermarking affects W2–W4 without reconstructing the image from scratch, but its effectiveness is not uniform across RoSteALS, ZoDiac, WAM, and Tree-Ring. The evidence demonstrates cross-watermark interference rather than guaranteed overwriting; detailed bit-accuracy and quality-budget results are reported in Appendix C.3.

6.3.6. Learning-Based Denoising and Reconstruction Attack Networks

Learning-based attacks train CNNs, GANs, or diffusion networks to map watermarked images toward clean-image distributions. Supervised denoisers treat the watermark as a residual, unpaired methods learn distributional translation without one-to-one clean targets, bidirectional diffusion models learn both watermark residuals and restoration noise, and dual-domain methods combine frequency reconstruction with semantic refinement. These routes exploit the separability of moment-domain, frequency-domain, and deep post-hoc watermark signals from the visual content that the network is trained to preserve.
Wang et al. [40] (HAI-WAN, Signal Processing 2025) use supervised data generated with a known target scheme, corresponding to gray-box training followed by no-box inference. Yan et al. [49] (UQP-WR, TDSC 2025) train a modified DDPM-UNet generator and discriminator on unpaired watermarked and public clean images; collecting watermarked outputs through a generation API is a black-box training condition, while inference requires no detector queries. Wang et al. [52] (BDMWA, TCE 2026) use bidirectional diffusion and transfer from trained W1/W2 schemes to unseen targets. Meshram and Chandrasekaran [41] (D2RA/DAWN, ICML 2026) require only one watermarked image at attack time and combine an 8 × 8 DCT U-Net with image-to-image diffusion and optional color or style correction. HIWANet [54], RD-IWAN [55], and Concealed Attack [51] use supervised, target-specific training data but perform no-box inference without a key or decoder.
Learned reconstruction is strongest on W1 and post-hoc W2 watermarks, while transfer to diffusion-native W4 schemes remains mixed and can incur substantial distortion. Supervised variants often favor fidelity over complete removal; detailed BER, attack-success, and fidelity results are reported in Appendix C.3.
Defense implications. Regeneration and purification attacks exploit low-amplitude embedding, localized spectral energy, reconstruction-unstable evidence, watermark superposability, and mappings that learned restorers can separate from image content. Defenses should combine binding to image content or the generation process, redundancy across spatial and frequency bands, per-image randomized embedding, robustness training against generative reconstruction, and cryptographic provenance with regeneration forensics. These measures can raise the attack cost, but they do not remove the theoretical limits associated with repeated quality-preserving rewriting.
Detailed quantitative results, fidelity measurements, and method-specific defense considerations are provided in Appendix C.3.

6.4. Adversarial Perturbation Attacks

Adversarial perturbation attacks formulate watermark removal or evasion as optimization against the decision boundary of a detector, decoder, or surrogate. Depending on the available access, an attacker differentiates through the target, queries its output, reproduces a surrogate key, or transfers perturbations from locally trained models. The central weakness is that watermark verification often exposes a differentiable, queryable, or reproducible decision rule. Because the optimization targets decoded bits or detector scores rather than reconstructing the entire image, successful attacks can preserve high perceptual fidelity. The in-scope studies reviewed here form five mechanism classes and cover W1–W4 watermarks; related model-IP work is treated only as a scope boundary in Appendix E.

6.4.1. Substitute Models and Transferable Adversarial Examples

When an attacker cannot differentiate through the target detector, a substitute detector or an ensemble of surrogate watermarking models can approximate its decision geometry. Gradients computed on these local models produce perturbations that transfer to the unseen target. This route exploits shared watermark features, including high-frequency evidence and common decoder representations, rather than access to the true secret key.
Quiring and Rieck [31] (AdvML, EUSIPCO 2018) train a compact substitute detector on high-frequency coefficients from watermarked images generated with the same key and from unwatermarked images, apply optimization in the style of Carlini and Wagner, and use the target detector only for confirmatory queries. This is a black-box setting, and the same-key watermarked images are a separate data condition. Hu et al. [59] (Transfer Attack, ICLR 2025) remove access to the target system entirely: in a no-box setting, they train dozens of diverse surrogate watermarking models, up to 100, on independently collected data and jointly optimize inverse-decoding targets for transfer to unseen decoders.
Substitute-based attacks show that robustness to conventional distortions or bounded certification does not ensure resistance to transferred adversarial examples. Their strongest evidence concerns W2 and Stable Signature [W3] with explicit bit-decoding interfaces, while constructing representative surrogates remains costly. Detailed evasion results are reported in Appendix C.4.

6.4.2. Bit-Flipping Evasion with White-Box Gradients and Black-Box Queries

Multi-bit watermark detectors often determine watermark presence from the agreement between decoded and reference bits. A direct attack that flips every bit may be caught by a double-tail detector, so bit-flipping evasion instead drives bit accuracy toward 0.5, where the decoded payload resembles a random string. White-box variants optimize through decoder gradients, whereas black-box variants search the decision boundary through detector queries. Both exploit the sensitivity of a bit-decoding interface to small, targeted perturbations.
Jiang et al. [57] (WEvade, CCS 2023) instantiate both settings. WEvade-W-II has white-box access to the decoder but not the true watermark, and optimizes toward a randomly chosen target payload. WEvade-B-Q has only black-box access to the detector’s binary decision and adapts HopSkipJump with JPEG-based initialization and early stopping. In either setting, the data condition is a single watermarked image; the distinction is the decoder or query access used during per-image optimization.
WEvade is highly effective on the tested W1 and W2 multi-bit watermarks, but it requires a differentiable or queryable bit-decoding interface and therefore does not directly cover zero-bit or semantic verification. Query-based use also incurs per-image cost; detailed evasion and query results are reported in Appendix C.4.

6.4.3. Adaptive Adversarial Optimization with Surrogate Keys

If the key-generation and verification structure can be reproduced locally, an attacker can construct a differentiable surrogate key and optimize either bounded pixel noise or a learned compression model. The perturbation or compressor is trained against the surrogate and then transferred to the unknown target key. This mechanism exploits low effective key diversity and reproducible verification structure rather than direct access to the true key.
Lukas et al. [58] (Adaptive Attack, ICLR 2024) apply adversarial noising to Tree-Ring and tune a Stable Diffusion autoencoder for adversarial compression against WatermarkDM, DWT, DWT-SVD, and RivaGAN. Under this survey’s threat model, the setting is gray-box because the attacker knows the watermarking algorithm and uses an open surrogate generator or a reproducible surrogate key, while lacking the target key and verification access. Non-differentiable DWT variants require a locally trained ResNet-50 extractor; this training requirement is a data and computation condition, not a stronger target-access level.
Adaptive optimization can outperform fixed post-processing under the reported protocol, with adversarial compression more consistent than adversarial noising. Its reach depends on public algorithm knowledge, a representative surrogate generator, and reproducible key structure; detailed detection results are reported in Appendix C.4.

6.4.4. Feature-Aware Residual Adversarial Watermark Removal

Feature-aware residual attacks train an image-to-image network to estimate the spatial residual associated with a watermark and subtract it from the watermarked image. Branch convolutions capture features at different receptive fields, while attention modules emphasize channels that predict the residual. This approach exploits the separability and spatial regularity of post-hoc watermark evidence, but it deliberately balances bit corruption against image reconstruction quality.
Chen et al. [61] (FAADW, IEEE MultiMedia 2025) combine branch convolution, deep feature extraction, and feature attention in a conditional GAN. The model is trained primarily against StegaStamp and evaluated for transfer to DwtDctSvd, HiDDeN, and SSL Watermarking. Attack-time inference is no-box because only the watermarked image is supplied and no target detector is queried. Its supervised preparation uses paired original and StegaStamp-watermarked images produced with a pretrained encoder; this paired-data requirement is orthogonal to the no-box inference label.
FAADW transfers across W1 and W2 while retaining moderate fidelity, but the evidence supports payload disruption rather than complete erasure and does not include a calibrated detector-evasion threshold. Transfer to zero-bit, semantic, or diffusion-native watermarks remains unestablished; detailed BER and fidelity results are reported in Appendix C.4.

6.4.5. Reverse Pivotal Tuning for Generator-Parameter Watermarks

For a watermark embedded in generator parameters, an attacker can reverse the embedding process by fine-tuning the generator on a small clean-image set. This changes the protected model rather than post-processing individual outputs. The exploitable weakness is that parameter-space evidence introduced through pivotal tuning may be overwritten by later optimization while the generator retains its image-synthesis capability.
Lukas and Kerschbaum [60] evaluate Reverse Pivotal Tuning against their PTW watermark for GAN generators (USENIX Security 2023). Effective reverse pivotal tuning assumes white-box access to generator parameters and uses 50–200 real images as an orthogonal data condition. The paper also evaluates black-box output distortions, but these baselines do not approach the removal performance of parameter fine-tuning.
Reverse pivotal tuning exposes the reversibility of W3 generator-parameter watermarks under strong white-box access, while the failure of ordinary output distortions marks an access-dependent boundary. It is therefore a stress test of model-level persistence rather than a low-access attack on standalone images; detailed capacity results are reported in Appendix C.4.
Attacks on box-free model watermarks target model intellectual property rather than standalone watermarked images and therefore fall outside the W1–W5 scope; their relation to the present taxonomy is summarized in Appendix E.
Defense implications. Adversarial perturbation attacks exploit exposed decoder gradients, queryable decisions, transferable feature geometry, reproducible surrogate keys, watermark evidence that learned restorers can separate from image content, and reversible generator parameters. Defenses should therefore combine authenticated and rate-limited verification, per-image cryptographic keys, higher key diversity, detector randomization or adversarial training, stronger coupling between watermark evidence and image content, and integrity protection for watermarked generators. These measures should be evaluated against both direct optimization and transfer from independently trained surrogates.
Detailed quantitative results, fidelity measurements, and method-specific defense considerations are provided in Appendix C.4.

6.5. Semantic Editing Attacks

Semantic editing attacks weaken or remove watermarks by rewriting image content or its generative trajectory while preserving the image’s principal semantics. Unlike conventional perturbations, they use instruction-driven editing, semantic masks, partial regeneration, or stochastic diffusion resampling to replace the pixel and latent correlations that carry watermark evidence. The three mechanisms reviewed here are instruction-driven global or local editing, content-preserving background inpainting, and stochastic hidden-trajectory deflection. SemanticRegen and SHIFT are direct attack studies, whereas Robust-Wide and VINE / W-Bench provide defense and benchmark evidence for this attack surface. AWD-AGP [69] protects visible watermarks and is outside the scope of this survey.

6.5.1. Instruction-Driven Image Editing

Instruction-driven editing applies a general-purpose editor, such as InstructPix2Pix, UltraEdit, MagicBrush, or a ControlNet inpainting variant, to a watermarked image. A text instruction guides global restyling or local replacement while preserving the visible subject and intended meaning. The edit nevertheless rewrites the pixel distribution and disrupts fixed spatial or frequency residuals. This mechanism exploits weak coupling between post-hoc watermark evidence and the semantic content that the editor is asked to retain.
The attack effect is characterized mainly through benchmark and defense studies rather than a dedicated attack algorithm. Lu et al. [68] (VINE / W-Bench, ICLR 2025) evaluate 11 invisible image watermarking methods under global and local instruction-driven editing as well as image regeneration. Hu et al. [67] (Robust-Wide, ECCV 2024) treat instruction-driven editing as a semantic distortion when designing an edit-robust watermark. From the attacker’s perspective, these operations are no-box: they require one watermarked image and a public editor, but no target detector, key, or query access.
Instruction-driven editing is operationally realistic because it coincides with ordinary image-editing workflows and attacks W1 and W2 without a specialized removal pipeline. Global edits may alter intended content, local edits are often weaker, and fidelity depends on the instruction and editor. W-Bench indicates that spatial redundancy helps but does not by itself bind watermark evidence to image meaning; detailed TPR results are reported in Appendix C.5.

6.5.2. Content-Preserving Background Inpainting and Partial Regeneration

Some post-hoc, generative, and zero-bit watermarks distribute evidence across background regions that can be regenerated without changing the salient subject. A content-preserving attack first describes the image with a vision-language model, identifies prominent objects with language-guided segmentation, and applies diffusion inpainting to the complementary region. This differs from the global reconstruction attacks in Section 6.3 because the semantic mask determines which content is preserved and which region is rewritten. The weakness is insufficient binding between watermark evidence and the salient content that establishes the image’s identity.
Tallam et al. [65] (SemanticRegen, arXiv 2025) implement this route with BLIP2 captioning, LangSAM foreground masks, language-model prompt rewriting, and Stable Diffusion inpainting. The attack is no-box and training-free with respect to the target watermark system: it uses one watermarked image and public semantic and generative tools, without the watermark key, detector access, or detector queries. Its evaluation covers W1–W4 targets, including DwtDct, StegaStamp, Stable Signature, and Tree-Ring.
SemanticRegen shows that selective rewriting can remove or weaken W1–W4 evidence while preserving the retained foreground, although StegaStamp does not reach detector evasion under this survey’s grading references. Its limitations include dependence on captioning and segmentation, inpainting artifacts, and fidelity measures that emphasize retained regions; detailed p-value, bit-accuracy, and masked-SSIM results are reported in Appendix C.5.

6.5.3. Stochastic Hidden-Trajectory Deflection

Diffusion watermarks often assume that deterministic inversion recovers a trajectory or initial-noise structure consistent with the embedded evidence. Stochastic hidden-trajectory deflection applies partial forward diffusion and then injects fresh noise during ancestral reverse resampling, forcing the image onto a different latent path. The reconstructed image can remain semantically similar while its recovered noise is statistically decoupled from the watermark trajectory. The exploited weakness is therefore trajectory consistency rather than a particular decoder or key.
Bao et al. [66] (SHIFT, arXiv 2026) use a public latent diffusion model to encode the watermarked image, add partial forward noise, and perform stochastic reverse resampling. Under this survey’s threat model, SHIFT is no-box because it requires no knowledge of the target watermark algorithm, key, prompt, verifier, or parameters and makes no target-system queries. The public diffusion model and the single input image are tool and data conditions; the attack requires no retraining, fine-tuning, or per-watermark adaptation.
SHIFT provides broad, training-free transfer across W4 trajectory- and noise-based designs, but stronger deflection can alter fine details and its CLIP and FID evidence supports only protocol-relative comparison. It is placed here because it performs semantics-preserving stochastic rewriting rather than target-latent inversion or detector-boundary optimization; detailed attack-success and fidelity results are reported in Appendix C.5.
Defense implications. Semantic editing attacks exploit weak binding between watermark evidence and salient image content, regenerable background regions, and stable diffusion trajectories. Defenses should distribute evidence across semantically important regions, bind verification to image content, train against global and local instruction-driven edits as well as stochastic resampling, and combine watermark verification with editing and regeneration forensics. Evaluation should report both detection at a strict false positive rate and semantic fidelity so that robustness is not claimed merely because meaningful editing is suppressed.
Detailed quantitative results, fidelity measurements, and method-specific defense considerations are provided in Appendix C.5.

6.6. Latent-Space Inversion Attacks

Latent-space inversion attacks recover an approximate diffusion latent, initial-noise representation, or denoised latent and then manipulate that representation or its associated decoder parameters. Their entry points include single-image optimization, surrogate detection in a recovered VAE space, optical-flow warping, latent-symbol boundary manipulation, and decoder fine-tuning. These mechanisms primarily affect W3 and W4 watermarks whose evidence is tied to latent noise, inversion trajectories, codeword signs, or watermark-bearing decoder weights, although NFPA also evaluates W1 and W2 targets. The literature reveals two recurring weaknesses: inversion is non-unique, and latent evidence may remain separable or mutable even when it is robust to pixel-level distortions. This section reviews five attack papers in five subcategories. CoSDA [75] and ZoDiac [76] provide defense-side evidence and robust comparison targets.

6.6.1. Single-Image Latent-Space Inversion for Removal and Forgery

For latent-space and initial-noise watermarks, the mapping from an image to its initial noise is not one-to-one. An attacker can therefore optimize a small RGB perturbation while using VAE encoding and DDIM inversion to move the recovered representation across a watermark region. Moving a watermarked image out of that region supports removal, whereas moving a clean cover image toward the region inferred from a reference supports forgery. The attack exploits both inversion non-uniqueness and the geometric separability of watermark-bearing latents.
Jain et al. [70] (Single-Image Removal / Single-Image Forgery, arXiv 2025) implement this route without the diffusion U-Net, watermark key, or target-detector queries. Forgery uses one watermarked reference carrying the target pattern together with a clean cover image; removal uses the watermarked image itself. Under this survey’s taxonomy, using a public proxy VAE is no-box, while direct access to the victim VAE raises the setting to gray-box. The reference and cover images remain data conditions rather than access labels.
The principal finding is an asymmetry between low-access forgery and removal: forgery transfers broadly across the tested W4 schemes, whereas removal is more target dependent and may require greater fidelity loss. This exposes weak ownership binding but depends on a usable victim or proxy VAE; detailed success and fidelity results are reported in Appendix C.6.

6.6.2. Public-VAE Surrogate-Detector Attacks

When a target pipeline uses a public or reproducible VAE, an attacker can map watermarked and comparison images into an approximate latent space, train a surrogate classifier there, and optimize an image perturbation against the surrogate with projected gradient descent. The perturbation changes the VAE-recovered representation and can transfer to the undisclosed detector. This route exploits latent separability and architectural reuse rather than direct access to the watermark key or decision threshold.
Lin and Juarez [71] (CrackBark, USENIX 2025) train this surrogate on Tree-Ring outputs collected from the target generator and either non-watermarked generator outputs or public images. The attack then transfers a PGD perturbation from the recovered-latent classifier to the real Tree-Ring detector. It is gray-box under this survey’s taxonomy because the attacker uses the same or a similar public VAE, but it requires no watermark key, detector threshold, detector internals, or detector queries. The collected outputs and public images are separate training-data conditions.
CrackBark shows that public VAE priors can expose a transferable Tree-Ring attack surface without detector queries. Its main limitation is architectural dependence, so it does not establish a model-independent route for every latent watermark; detailed ROC-AUC and perturbation results are reported in Appendix C.6.

6.6.3. Semantic-Preserving Removal via Latent-Space Optical-Flow Shifts

Latent-space and initial-noise watermarks may lose correspondence after a small optical-flow deformation of their inverted noise. A next-frame prediction attack first applies DDIM inversion, warps the recovered latent noise along bounded horizontal, vertical, or combined flow directions, and denoises a two-frame latent sequence with frame attention. The generated second frame is used as the attacked image. This route exploits sensitivity to latent viewpoint shifts while using temporal consistency as an image-preservation prior.
Qiu et al. [72] (NFPA, NeurIPS 2025) instantiate the attack with a public text-to-image diffusion model, DDIM inversion, bounded flow search, and an empty prompt because the original prompt is unknown. Under this survey’s taxonomy, NFPA is no-box because it requires no knowledge of the target watermark system and receives no target-detector feedback or target-system queries. Its separate tool and data conditions are one watermarked image, a public generation pipeline, and no additional training data. Although the mechanism borrows a two-frame generative prior, the evaluated targets are invisible image watermarks and the output used for evaluation is a single attacked image.
NFPA transfers across W1–W4 without watermark-specific adaptation and is strongest against Stable Signature and several latent-noise designs, while some multi-bit schemes retain residual detectability. Its cost, inversion error, public-model dependence, and simple flow patterns limit generality; detailed TPR and fidelity results are reported in Appendix C.6.

6.6.4. Detection-Boundary Leakage on Latent-Symbol Watermarks

For latent-symbol watermarks such as PRC, watermarked starting points can be computationally indistinguishable from unwatermarked Gaussian samples while the detector boundary retains exploitable geometry. A boundary-leakage attack sorts latent coordinates by absolute value and flips the signs nearest the code boundary, increasing the codeword error rate at lower distortion than isotropic noise. Because sign flipping preserves the marginal Gaussian form, sample indistinguishability alone does not guarantee resistance to removal.
Lee et al. [73] (Boundary Leakage, CCS 2025) study a known-scheme setting in which the key remains hidden, the attacker observes same-key watermarked images, and the starting latent is recovered exactly or approximated through DDIM inversion. This is gray-box access in the present taxonomy because the PRC detection structure is known, but the attack does not require direct detector queries. The paper analyzes exact inversion and two practical inversion settings and also proves a matching boundary-hiding result based on a secret latent transformation.
Boundary Leakage establishes that cryptographic undetectability and removal resistance are different properties. Its advantage depends on recoverable boundary geometry and accurate inversion and collapses when the boundary is hidden. The analysis is deepest for PRC rather than all W4 schemes; detailed distortion and bit-flip results are reported in Appendix C.6.

6.6.5. Model-Level Removal by Decoder Parameter Fine-Tuning

Stable Signature roots a fixed watermark in fine-tuned VAE decoder parameters. A model-level attacker can estimate denoised latents for non-watermarked natural images and then fine-tune the watermarked decoder to reconstruct those images. The modified decoder subsequently generates outputs that evade watermark verification without a separate per-image removal step. The exploitable weakness is that watermark persistence depends on mutable, accessible decoder weights.
Hu et al. [74] (StableSig-Unstable, arXiv 2024) assume white-box access to the watermarked denoising layers and decoder and use 4,000 non-watermarked images from each attacking dataset. The E-aware route also uses the encoder and diffusion process to estimate latents directly. The E-agnostic route omits those two components, estimates each latent through successive pixel and Watson-VGG optimization, and then fine-tunes the decoder. E-agnostic therefore reduces the required component access but remains white-box with respect to the watermark-bearing decoder.
A single decoder modification affects future outputs and is therefore more scalable than repeated per-image removal, but it requires extensive component access and substantial one-time computation. The attack is relevant to open-source or leaked-weight deployments rather than API-only generators; detailed evasion and runtime results are reported in Appendix C.6.
Defense implications. Latent-space inversion attacks exploit non-unique inversion, reusable public VAEs, separable recovered latents, unstable coordinate correspondence, exposed code boundaries, and mutable decoder parameters. Defenses should combine content- or message-bound watermark evidence with per-image keys, diversify or conceal latent transformations and detector geometry, test robustness across mismatched inversion pipelines, align latent drift when inversion is unavoidable, and protect watermark-bearing model weights with authentication and integrity checks. Verification should also compare evidence across independent latent or image representations so that one recoverable coordinate system does not determine ownership.
Detailed quantitative results, fidelity measurements, and method-specific defense considerations are provided in Appendix C.6.

6.7. Forgery and Provenance Attacks

Forgery and provenance attacks act primarily at the attribution layer, but their immediate objectives are not limited to forgery and spoofing. Detector evasion and watermark removal also belong here when they enable plagiarism, competing ownership claims, or provenance ambiguity. Across these objectives, two observations recur. First, robustness can increase forgeability when persistent watermark features become stable enough to estimate or transfer. Second, successful extraction alone does not establish ownership unless the recovered evidence is cryptographically bound to the image, claimant, and key. We organize 11 in-scope attack studies into six subcategories and discuss WIND [96] and MetaSeal [97] as defense studies. Related attacks on model IP watermarks fall outside the survey scope and are discussed only as boundary cases in Appendix E.

6.7.1. Classical Invertibility and Copy Attacks

Watermarks without cryptographic binding may permit two classical forms of attribution attack. An invertibility attack constructs a counterfeit original and counterfeit watermark that also explain a released watermarked image, whereas a copy attack estimates watermark evidence and transfers it to another cover image with perceptual weighting. Both routes exploit additive or otherwise separable evidence that is insufficiently bound to the content and the genuine key.
Craver et al. [86] (Invertibility Attack, IEEE JSAC 1998) construct counterfeit ownership explanations from either one watermarked image or two watermarked images. Kutter et al. [85] (Copy Attack, SPIE 2000) use predictive denoising to estimate a watermark from a source watermarked image and paste it into a target cover image with noise visibility weighting. Both attacks are no-box under the present taxonomy: they require image data but no detector, model, key, or query access.
The classical studies established that robustness against removal does not prevent an adversary from producing rival ownership evidence. Their historical importance lies in separating detectability from trustworthy attribution. Their practical reach is narrower against modern content-bound or cryptographically authenticated watermarks, and the original studies do not provide image-fidelity measurements that can be graded under the framework of this survey.

6.7.2. Surrogate-Model Forgery of Semantic and Latent-Space Watermarks

Semantic and latent-space watermarks may leave transferable structure in inverted diffusion noise. A surrogate-model attacker can recover that structure with a proxy diffusion model, imprint it into the latent representation of a cover image, reprompt from the recovered noise, or plant it through a regenerative restoration backbone. The attack succeeds when watermark-bearing latent features survive changes in the inversion, denoising, or restoration architecture.
Müller et al. [87] (Black-Box Forgery, CVPR 2025) recover watermark structure with an unrelated proxy diffusion model and use Imprinting or Reprompting to forge Gaussian Shading and Tree-Ring. Zhu et al. [89] (PnP / OptFree-Forgery, arXiv 2025) combine proxy inversion with StableSR, DiffBIR, or SeeSR to plant the estimated latent into an arbitrary cover without per-image optimization. Both attacks are no-box in this survey because they do not query or inspect the victim watermark system. The required watermarked reference images and public proxy or restoration models are data and tool conditions, not victim-system access.
The results undermine the assumption that hiding the victim generator is sufficient when W4 evidence transfers through public latent priors. Architecture compatibility is the main limitation: transfer is strong for compatible latent and denoising structures but weakens across substantially different backbones. Detailed detection and attribution results are reported in Appendix C.7.

6.7.3. Low-Access Forgery of Post-Hoc Watermarks

Post-hoc watermarks can expose learnable residual distributions or image-independent features. An attacker may learn the common watermark distribution from a collection of outputs or isolate transferable evidence from one watermarked image, then implant that evidence into a cover image or manipulate it for evasion. The exploitable weakness is weak dependence between the embedded evidence and the individual image content.
Dong et al. [88] (WMCopier, NeurIPS 2025) train an unconditional diffusion model on multiple images produced by the same watermarking scheme and manipulate DDIM trajectories to implant the learned distribution. Souček et al. [93] (One-Shot Forging, NeurIPS 2025) use an offline preference model to extract transferable evidence from one watermarked image. Ba et al. [90] (Robust-Leak, arXiv 2025) use channel-aware feature extraction from one watermarked image for evasion and target-dependent forgery on the evaluated W2 targets. All three are no-box; their differing collections, single-image references, and offline models are data or training conditions rather than victim-system access.
Low-access forgery is practical across W1–W3, but input-dependent schemes remain harder targets and strict low-FPR calibration reduces apparent acceptance. This indicates that stronger content dependence can limit the transferability of leaked residual evidence; detailed forged bit-accuracy and acceptance results are reported in Appendix C.7.

6.7.4. Learning-Based Forgery and Counterfeit Extractor

Learning-based forgery follows two distinct routes. The first approximates the embedding map and generates images that imitate the target watermark. The second fabricates the verifier itself, producing an attacker-chosen message for victim outputs without changing those images. Both exploit ownership tests that accept a recovered message without authenticating the extractor, key, or relation between the evidence and the image.
Wang et al. [91] (WatermarkFaker, ICME 2021) train a U-Net and PatchGAN on paired original and watermarked samples to synthesize forged images. This is no-box access because the paired samples are a data condition and the target system is neither queried nor inspected during the attack. Luo et al. [94] (Counterfeit Extractor, SPL 2025) use black-box access: the attacker queries a victim watermarked SDM and a clean SDM, then trains a ResNet-18 extractor to return a forged message on victim outputs and random-like predictions on clean outputs.
The learned embedding route is stronger on spatial W1 schemes than in the DCT domain, whereas the forged-verifier route exposes a protocol failure rather than an image transformation. Because the images are unchanged, image fidelity cannot reveal Counterfeit Extractor; detailed target-message results are reported in Appendix C.7.

6.7.5. Ambiguity Attacks on Trigger-Based Watermarks

Trigger-based diffusion watermarks associate ownership with a secret prompt that reproduces a designated watermark image. An ambiguity attack instead optimizes a forged prompt in token-embedding space until the same watermarked model reproduces evidence accepted as the owner’s. The weakness is non-unique verification: any prompt that reaches the acceptance condition may support a competing claim.
Yuan et al. [92] (Ambiguity Attack, Signal Processing 2024) optimize forged prompts through a frozen text encoder and UNet so that the watermarked Stable Diffusion model regenerates the owner’s watermark image. The attack is white-box in the present taxonomy because optimization backpropagates through these victim-model components; the observed watermark image and prompt initialization remain separate data conditions.
A trigger is not an ownership credential unless its uniqueness is enforced. The attack applies only to trigger-based or backdoor-style W3 schemes and does not establish forgery of post-hoc or latent-noise watermarks. Its reliance on victim-model gradients also makes it less applicable to closed services than the no-box attacks above.

6.7.6. Removal and Plagiarism in Attention Space

Attention-space attacks optimize inverse-latent anchors and cross-attention shims to generate content that remains semantically close to a protected image while weakening its watermark. Their primary objective is removal, but the resulting replicas can support plagiarism and provenance ambiguity. The exploitable weakness is that watermark evidence can be separated from content through a generative attention trajectory.
Zou et al. [98] (Neural Plagiarism, ICCV 2025) combine anchor search, shim optimization, and cross-attention perturbation against W1–W4 targets. The attack is no-box with respect to the victim watermark system because it neither queries nor inspects that system. Gradient access is required only for an independent attack diffusion model used to optimize the anchors and shims, and therefore does not constitute white-box access to the victim.
The attack strongly suppresses several W1–W3 targets while Tree-Ring [W4] remains detectable, ruling out a universal-removal claim and pointing instead to coexisting-watermark ambiguity for latent-pattern schemes. Visible fidelity costs in some settings make forensic analysis an important complement; detailed TPR results are reported in Appendix C.7.
Defense implications. Forgery and provenance attacks exploit weak content binding, stable residual leakage, learnable embedding maps, static or reusable keys, non-unique triggers, and verification protocols that do not authenticate the extractor or claimant. Defenses should combine cryptographic binding between the image, message, key, and identity with per-image keys, authenticated verifiers, controlled query interfaces, and dedicated forgery and plagiarism forensics. Robustness tests should also include transfer, replay, counterfeit-extractor, and competing-claim scenarios rather than removal alone.
Detailed quantitative results, fidelity measurements, and method-specific defense considerations are provided in Appendix C.7.

7. Open Challenges and Future Directions

Taken together, the attack taxonomy in Section 6 and the benchmark evidence detailed in Appendix F reveal two related trends. Many recent attacks reduce the victim-system knowledge or query access required, although their computational and data costs remain method dependent. At the same time, the long-standing objectives of removal, evasion, and forgery are being pursued against broader target surfaces, including provenance and attribution systems, while mechanisms now include bit flipping, desynchronization, and overwriting. Evaluation and defense, however, have not fully kept pace with these shifts. We group the open problems into four areas.

7.1. Evaluation Comparability and Theoretical Limits

Existing results are difficult to compare across studies. As summarized in Section 4 and detailed in Appendix F, WAVES [99] and Erasing the Invisible [101] use normalized aggregate quality-degradation scores, whereas WIBE [103] reports standard fidelity and robustness metrics within its own pipeline; these scores are not directly transferable, while earlier studies use bit error rate (BER), TPR@1%FPR, or area under the ROC curve (AUROC). Attack coverage is also uneven. Distortion, regeneration, and adversarial attacks are repeatedly measured, but semantic editing (§6.5) remains sparse outside W-Bench [68], latent-space inversion (§6.6) lacks broad standardized coverage, and forensic detectability [100] lacks a shared protocol. Theoretical results are similarly fragmented: regeneration has provable removability results [42], strong watermarks face impossibility results under quality and perturbation oracle/random-walk assumptions [47], and PRC-style coding layers have a critical corruption rate of α * = 1 / 2 under the binary independent-symbol model [83]. A unified leaderboard should jointly report objective-specific attack effectiveness and image fidelity at explicit low-FPR operating points, include forensic detectability, and state access and data budgets so that removal, evasion, and forgery remain comparable.

7.2. Expansion of the Threat Surface

Forgery remains underestimated relative to removal. Forgery is orthogonal to removal and can be more damaging for forensic reasoning (§6.7), yet it has received less systematic attention. Müller et al. [87], Dong et al. [88] (WMCopier), and Luo et al. [94] show that attackers can implant watermark evidence into arbitrary images from a single reference image, from a collected watermarked-image set, or from outputs obtained through black-box model queries; they can also manufacture ownership evidence without modifying the image. In their stronger reported configurations, forgery metrics, including bit accuracy and detection or attribution success, reach approximately 85% to 99%, although weaker settings and targets produce substantially lower results. Forgery and removal show a structural asymmetry. For distribution-preserving Gaussian Shading [W4], forgery is nearly perfect, whereas removal reaches only 34%–59% [70] (§6.6). The causes and limits of this asymmetry remain insufficiently characterized. Several attacks operate with limited victim-system access. No-box and detector-query-free attacks, such as MarkSweep [35], WMCopier [88], and Transfer Attack [59], show that watermarks can be removed or forged without interacting with the detector, which limits the effectiveness of defenses based only on rate limiting or query screening.
Multi-watermark settings remain underexplored. In real deployments, multiple watermarks may be stacked; an attacker may selectively remove or forge one watermark while preserving the others, thereby creating more covert misattribution. Practical threats may also involve sequential attack chains, such as geometric desynchronization followed by regeneration and adversarial perturbation. WAVES reports low Tree-Ring [W4] detection performance under separate regeneration and adversarial variants [99], which motivates the evaluation of composed attacks but does not establish their effectiveness. Most studies still evaluate single-step attacks.

7.3. Physical and Viewpoint Chains

This survey focuses on static invisible image watermarks, so future work remains within image-domain geometric, viewpoint, and physical attack chains. Existing image attacks have already begun to use a 2D Gaussian splatting decoder to disrupt latent-space phase coherence, as in MarkCleaner [78] (§6.2). The physical chain is another major gap. Print-camera watermarks [W5] [82] must survive printing and camera recapture, which introduce partly non-differentiable or difficult-to-model degradations. Their attacks and defenses cannot be directly inferred from digital-domain results, and benchmarks for physically realizable image watermark removal and forgery are still lacking.

7.4. Compliance, Legal Evidence, and Anti-Forensics

Watermarks are ultimately used for copyright determination, provenance tracing, and attribution, so attacks can have compliance and legal consequences. Forgery and invertibility attacks, such as the Copy Attack [85] and the attack by Craver et al. [86] in §6.7, can create ownership disputes and make arbitration difficult for the first embedder. This calls for ownership-resolution protocols designed for legal evidence chains, where timestamps, key escrow, and unforgeable content binding are incorporated into the evidentiary process. Cryptographic watermarking is an important direction. SoK [102] identifies pseudorandom code watermarks as a promising route, and MetaSeal [W5] [97] shows that content-dependent cryptographic binding can make forgery fail verification. However, Lee et al. [73] (Boundary Leakage, CCS 2025) show that PRC [W4] can still be attacked, which indicates that undetectability is not equivalent to removal resistance. Cryptographic schemes therefore still require further hardening under coding-theoretic limits [83]. Anti-forensics is an additional evaluation layer. Forensic Stealth [100] shows that the six evaluated remover-specific forensic detectors identify attacked outputs at a 1% false positive rate with true positive rates of 99.24%–99.97%. The next step for attackers may be forensic stealth, while defenders should make forensic detectability a routine evaluation metric and design objective.

8. Conclusion

A review of foundational attacks and modern developments from 2018 to 2026 shows that the core change is not merely stronger watermark removal, but the expansion of attack mechanisms, target surfaces, and low-access settings. Many recent attacks require less victim-system knowledge or query access, although their computational and data costs remain method dependent. Across these settings, the formal objectives remain removal, evasion, and forgery, while modern mechanisms include bit flipping, desynchronization, and overwriting, and provenance and attribution systems have become prominent targets. This leads to three observations.
First, robustness is not equivalent to security. Conventional robustness evaluations mainly measure resistance to benign distortions and do not by themselves establish security against adaptive attacks or forgery. Section 6.4 and 6.7 show that some robust and imperceptible watermarks may still leak features that can be extracted, transferred, or exploited for forgery. Robustness and security therefore remain in tension.
Second, apparently stronger watermarks expose new attack surfaces. Diffusion latent-space, zero-bit, or semantic watermarks [W4] are often regarded as stronger defenses. However, 2025 and 2026 works challenge them through latent-space inversion, such as NFPA [72] and CrackBark [71] in §6.6, detection-boundary leakage, such as Boundary Leakage [73] against PRC [W4], and cross-architecture forgery, such as Müller et al. [87]. Cryptographic-attribution watermarks [W5] are not yet directly compromised at the content level under acceptable-fidelity, in-scope attack evaluations surveyed here, but they face boundary risks such as key leakage, strong regeneration, and denial of service. These results show that private priors and undetectability are not equivalent to security.
Third, forgery can be more practical and more dangerous than removal in several scenarios. Forgery is orthogonal to removal. An attacker can create misattribution under single-image, black-box, or even image-preserving conditions (§6.7). The consequence is that innocent content may be assigned to a specific source, which is harder to clarify after the fact than simple watermark erasure.
Accordingly, watermark evaluation should retain the two core dimensions defined in Section 4: objective-specific attack effectiveness and image fidelity. Attack effectiveness should be reported separately for removal, evasion, and forgery, while forensic detectability should be included as an additional evaluation dimension. Evaluation should also be paired with verification protocols designed for legal evidence chains. Watermarking remains necessary for copyright protection and AI-generated content provenance tracing, but watermarking alone is insufficient against the attack taxonomy reviewed above. A more viable direction is to combine distribution-preserving or in-generation embedding, cryptographic provenance tracing and attribution binding, forensic detection, and evaluation against sequential attack chains, so that verifiable security boundaries can be stated under explicit threat models.
  • Acknowledgments: This work was supported in part by the Taishan Scholar Program under Grant tsqnz20250747; in part by the National Natural Science Foundation of China under Grants 62502250, 62406051, 62302249, 62541206, and 62272255; and in part by the Young Talent Lifting Project for Science and Technology in Shandong under Grant SDAST2025QTB030.

A. Detailed Profiles of Watermark Classes

The following profiles provide detailed descriptions of W1–W5 together with their distinct advantages, attack surfaces, and defense implications.

A.1. W1: Classical Frequency-/Spatial-/Statistical-Domain Watermarks

This class is grounded in classical signal processing. It embeds payload bits or detection statistics in the least significant bit (LSB), the DCT/DWT/SVD transform domains, the spread-spectrum domain, the quaternion Fourier transform domain, the moment domain, or statistically selected pixel sets, and it usually does not rely on deep generative models. Representative methods include the DwtDct invisible watermark used in Stable Diffusion implementations [57], the DWT-DCT-SVD variant DwtDctSvd [114], LSB spatial-domain watermarking, the spread-spectrum watermark of Cox et al. [115], quaternion Fourier transform watermarking [116], SVD-based watermarking [117], quantization index modulation (QIM), a quantization-based informed-embedding method related to dirty-paper coding [118], and the moment-domain robust reversible watermark RRWID [56]. This class also includes commercial systems such as Digimarc and several commercial targets reported in the StirMark paper [33]. These methods are nonsemantic: some use content-adaptive placement or embedding strength, whereas others are content-agnostic under the definition in Section 2. Their main weaknesses are as follows. Low-strength embedded watermarks are removed by regeneration, blurring, or denoising. Geometric desynchronization, such as random bending and StirMark-style distortion, markedly reduces detection reliability. For spread-spectrum watermarks, key estimation under known-message or multi-observation conditions lets an attacker read or forge the mark. Multi-image averaging together with copy and forgery attacks further weakens ownership verification.
Extraction assumptions vary across W1: many listed schemes support blind extraction, whereas classical Cox-style spread-spectrum detection requires the original or a reference image and successful registration [115]. Their practical advantages are that they require no training, have low deployment cost, remain relatively interpretable, and benefit from mature engineering implementations. Because W1 signals can be targeted through signal processing, geometric desynchronization, adversarial optimization, averaging, and forgery, this class has one of the broadest attack surfaces in the taxonomy. Defenses should combine multi-band redundancy and error-correcting codes with synchronization templates or Fourier–Mellin and log-polar invariant-domain embedding. Key diversification and content-adaptive, non-repeating templates can further reduce equivalent-key estimation, averaging, and copy-attack risks.

A.2. W2: Deep Multi-Bit / Steganographic Watermarks

This class is built around end-to-end deep neural encoder-decoder architectures, generative adversarial networks, autoencoders, or neural feature alignment. It embeds multi-bit payloads in the pixel domain, the feature domain, or an autoencoder latent space, and it is among the most widely used classes of invisible image watermarking methods. Representative methods include the foundational HiDDeN [22] and the high-perturbation StegaStamp [107], the attention-based GAN watermark RivaGAN [119], used as an image/frame-level baseline in attack studies, the JPEG-resistant MBRS [120], CIN [121], which combines invertible and non-invertible mechanisms, the flow-based invertible FIN [122], the screen-shooting-resistant PIMoG [123], the autoencoder-latent steganographic method RoSteALS [124], RAWatermark as evaluated in the averaging study [37], the self-supervised feature watermark SSL Watermarking [125], the provenance-oriented SepMark [126], the arbitrary-resolution TrustMark [50], the high-fidelity InvisMark [127], the localizable WAM [128], the general-purpose deep hiding method UDH [129], the attention-guided ARWGAN [130], the tamper-localizing dual watermark EditGuard [131], and VINE [68], which uses generative priors to improve robustness against editing. Many of these methods allocate perturbation adaptively according to image or model features but do not encode semantics, so they are content-adaptive and nonsemantic. Content-agnostic exceptions, including RoSteALS and RAWatermark as categorized in the averaging evaluation [37], retain more reusable cross-image structure. Their typical weaknesses are as follows. An attacker can query the decoder or apply adversarial perturbation under white-box access, as in WEvade [57]; train a surrogate model for transfer attacks that break even smoothed variants with certified robustness [59]; remove the watermark adaptively through optimization [58]; or remove it through regeneration or purification. Content-agnostic members are vulnerable to multi-image averaging, while stable learnable residual features can also enable one-shot single-image forgery [93].
The main advantages of W2 are high capacity, stronger robustness to common post-processing, and end-to-end optimization of fidelity and robustness; some schemes also support tamper localization or deepfake detection. Its attack surface additionally includes semantic editing. High-perturbation schemes such as StegaStamp are more resistant to ordinary purification than lower-perturbation schemes [42], but controllable regeneration and semantic editing can still weaken them [44,68]. Defenses should combine spatial redundancy with low- and mid-frequency allocation, editing- and regeneration-aware training, decoder-query screening, and per-image non-reproducible keys.

A.3. W3: In-Generation, Fine-Tuned, or Trigger-Based Generative Watermarks

This category modifies the generator, the VAE decoder, LoRA or model parameters, or introduces text-trigger conditions, so that the generative model injects the watermark during image generation. Representative methods include Stable Signature [105], which embeds a 48-bit watermark by fine-tuning the LDM decoder; LaWa [132], which embeds the watermark in the latent space during generation; AquaLoRA [133], a watermark LoRA for customized Stable Diffusion models; PTW [60], based on pivotal tuning of GAN generators; the generator-level artificial fingerprints of Yu et al. [134]; and backdoor-style trigger watermarks such as Safe-SD [135], which produce a watermarked image under a specific prompt. These methods are content-adaptive through the generation process or model parameters, but the watermark is not explicitly bound to image semantics. Their vulnerabilities depend on the underlying architecture. For decoder-rooted latent diffusion members such as Stable Signature, controllable diffusion regeneration can weaken or remove the watermark [44], while model purification combined with latent-space estimation can recover an unwatermarked decoder [74]. Transfer attacks against Stable Signature [59], spectral-magnitude erasure against evaluated W3 targets such as PTW and Stable Signature (UnMarker [34]), and frequency-aware denoising (MarkSweep [35]) also weaken detection. Counterfeit Extractor forges attribution for Stable Signature and AquaLoRA, together with the W4-adjacent DiffuseTrace [94].
DiffuseTrace and WatermarkDM can also be discussed along the generation pipeline, although their latent-space detection properties place them closer to W4. Because some W3 watermarks are inserted during generation and require no post-processing, they can be more resistant to conventional signal-level attacks than post-hoc schemes, although this advantage depends on the specific embedding mechanism. Their additional attack surface includes decoder fine-tuning and overwriting; under black-box query access, Counterfeit Extractor can fabricate messages for Stable Signature, AquaLoRA, and DiffuseTrace with about 99% accuracy without modifying the image [94]. Defenses should protect decoder parameters, bind evidence to image content, use adversarial and editing-aware fine-tuning, introduce cryptographic attribution, and keep trigger conditions secret and difficult to reverse-optimize.

A.4. W4: Diffusion Latent-Space, Zero-Bit, or Semantic Watermarks

This category binds detection cues to the diffusion sampling noise, the DDIM or VAE latent space, the sampling trajectory, a pseudorandom code, or a semantic structure, and relies mostly on zero-bit detection or latent-code verification. Representative methods include Tree-Ring [104], which embeds concentric Fourier rings in the initial noise, and its multi-key extension RingID [136]; Gaussian Shading [106], which is distribution-preserving and provably performance-lossless; PRC [137], which writes a cryptographic pseudorandom code into the sign bits of the latent space; WIND [96], a two-stage noise watermark; ROBIN [138], which uses adversarially optimized embedding; ZoDiac [76], optimized over VAE latents and DDIM sampling; GaussMarker [139], a dual-domain watermark over high-frequency components and noise trajectories; the semantic-aware SEAL [140]; the frequency-domain semantic watermark SFW [141]; DiffuseTrace [142], a multi-bit latent diffusion watermark; and WatermarkDM [143], an in-generation watermark for diffusion models. These methods are tightly coupled with the generation process, and for most of them, detection depends on latent-space patterns, noise structure, or codeword matching rather than image semantics, while semantic watermarks additionally bind to semantic structure. Their vulnerabilities concentrate in the latent space and the sampling trajectory. Proxy inversion with a public VAE substantially lowers detection, as CrackBark drops the ROC-AUC of Tree-Ring from 0.993 to 0.153 [71]. Single-image inversion broadly enables forgery, while removal remains target dependent and is ineffective for RingID and WIND in the reported settings [70]. Trajectory deflection with SHIFT [66] supports removal; multi-image averaging and subtraction support removal and, for some schemes, forgery [37]; detection-boundary leakage corrupts or removes PRC codewords [73]; and geometric-phase perturbations cause removal or evasion [78,79]. For PRC and related sign-coded latent watermarks, decoding fails as the latent sign-error rate approaches 50% [83].
Content-agnostic W4 schemes can retain low embedding distortion, but content agnosticism does not itself confer attack resistance and may expose reusable cross-image structure to averaging [37]. Distribution-preserving designs such as Gaussian Shading can reduce fidelity cost [106], while resistance to conventional signal-level and ordinary regeneration attacks depends on the specific embedding and detector. Defenses should combine drift alignment and compensation sampling, geometric synchronization tags, private VAEs, per-image keys, forgery-resistant content binding, and stronger error correction within coding-theoretic limits.

A.5. W5: Geometry-Synchronized, Physically Robust, or Cryptographic-Attribution Watermarks

W5 is a cross-cutting, defense-oriented class that targets failure modes such as geometric desynchronization, print-camera distortion, local editing, and attribution forgery. Its hardened variants may retain an underlying W1–W4 embedding mechanism while adding synchronization templates, geometric normalization, physical noise modeling, content binding, or cryptographic verification. Representative methods include PrintCamera [82], a print-camera robust watermark; SSyncOA [84], which achieves object-aligned self-synchronization; GeoWM [80], which improves geometric robustness with a Swin Transformer and deformable convolution; the Tree-Ring-SynTag and GauShad-SynTag variants [81], which add geometry-sensitive synchronization tags to W4 inversion-based watermarks; Robust-Wide [67], which targets instruction-driven editing; the Tree-Ring-CoSDA and GS-CoSDA variants [75], which strengthen inversion robustness; and MetaSeal [97], a cryptographic-attribution watermark that binds content through semantic features and an ECDSA signature. Their semanticity therefore depends on the underlying scheme: W5 includes content-agnostic, content-adaptive, and content-bound designs. The W5 schemes designed for geometric or physical robustness are markedly more resistant to their corresponding distortions, including rotation, shearing, affine transformation, and print-camera transmission. A demonstrated residual weakness is strong purification: at the strongest reported SynTag setting, diffusion purification reduces TPR to zero, but PSNR also falls to 15.44 dB, so this result does not establish a practical acceptable-fidelity removal compromise. Other risks remain largely potential and untested for this class: local or nonrigid deformation outside the training distribution, novel view synthesis and micro geometric desynchronization for designs that depend on synchronization, and white-box tampering or denial of service disruption. As discussed in Appendix D, W5 watermarks do not yet show direct content-level compromises under acceptable-fidelity, in-scope attack evaluations covered by this survey. Under the replay, mixup, and substitute-INN attacks evaluated in MetaSeal, its cryptographic content binding resists forgery under the stated cryptographic assumptions; the residual threat shifts toward denial of service through transformations such as heavy cropping and severe regeneration [97].
At the family level, defenses should combine Fourier–Mellin or log-polar invariant-domain embedding, spatial-transformer correction, geometric and physical simulation during training, and content-dependent private-key binding, as in MetaSeal, while modeling the print-camera channel end to end.

B. Detailed Evaluation Framework and Grading References

This appendix consolidates the complete grading references for attack effectiveness and image fidelity. Metrics with stable absolute ranges receive explicit grading bands, whereas distributional, semantic, no-reference, and protocol-specific metrics are interpreted by direction and relative change against clean or no-attack baselines. Geometric attacks should additionally report registered AlignPSNR and AlignSSIM where applicable.

B.1. Grading Table for Attack Effectiveness

Table 4 provides a four-level grading scheme for attack effectiveness. For multi-bit watermarks, the grading mainly uses BER and its complementary metric BA. For zero-bit watermarks, the grading mainly uses AUC/AUROC, TPR@x%FPR, and the p-value. Unlike separability and detection-rate metrics, p-values are interpreted by threshold crossing rather than by a graded magnitude above the significance threshold. Since the threshold for being judged as watermarked varies with the source protocol and its FPR calibration, Table B2 lists representative scheme-specific decision thresholds together with the reported FPR setting. An entry marked “FPR not reported” denotes a scheme threshold for which the cited source does not provide an explicit FPR calibration. For multi-bit watermarks, BA or BER approaching 0.5 is usually considered successful randomization. For zero-bit watermarks, AUC/AUROC approaching 0.5 indicates loss of separability, while a detection statistic below the detection threshold or a significant decrease in the detection rate under a fixed FPR usually indicates successful removal or evasion. Values on the opposite side of 0.5 can indicate systematic inversion rather than additional information destruction and should be interpreted under the source detector protocol.
Table B1. Attack-Effectiveness Levels for Multi-Bit and Zero-Bit Watermarks
Table B1. Attack-Effectiveness Levels for Multi-Bit and Zero-Bit Watermarks
Metric Weak Attack (Poor) Effective Attack Good Attack Strong Attack Notes and Sources
BER (toward 0.5) < 0.25 or > 0.75 0.25–0.35 or 0.65–0.75 0.35–0.40 or 0.60–0.65 0.40–0.60 Distance from 0.5 measures payload randomization; values near 0 or 1 retain decodable or inverted structure [38,40]
BA (toward 0.5; = 1 − BER ) > 0.75 or < 0.25 0.65–0.75 or 0.25–0.35 0.60–0.65 or 0.35–0.40 0.40–0.60 Bit accuracy; values near 0 may indicate systematic inversion rather than destruction [37,44]
AUC / AUROC (toward 0.5) > 0.90 or < 0.10 0.80–0.90 or 0.10–0.20 0.60–0.80 or 0.20–0.40 0.40–0.60 Loss of separability is measured by distance from 0.5; values below 0.5 indicate inverted ranking, although they may evade a fixed detector [43,71]
TPR@x%FPR (↓) > 0.80 0.40-0.80 0.10-0.40 → 0 Detection rate at a low false-positive operating point; WAVES / Erasing use 0.1% FPR [99,101]
p-value under a source-specific hypothesis test Use the paper-specific null hypothesis and significance threshold; p-values are not graded by magnitude across studies Threshold crossing can indicate removal or forgery only under the source detector protocol [65,70]
Removal Rate (RR) (↑) < 0.25 0.25–0.50 0.50–0.75 > 0.75 Derived from R R = 1 − 2 | B E R − 0.5 | ; values approaching 1 indicate complete payload randomization [38]
Objective-specific ASR / Forgery Success Rate (↑) < 0.40 0.40–0.70 0.70–0.90 > 0.90 Fraction of samples for which the stated removal, evasion, or forgery objective succeeds; the objective and acceptance rule must be reported [44,57,73,87,88,89,91,93]
NC for equivalent-key estimation (↑) < 0.60 0.60–0.80 0.80–0.90 > 0.90 Normalized correlation between the true and estimated carriers in the key-recovery setting; this direction should not be generalized to detector NCC or forgery realism [39]
Table B2. Scheme-Specific Decision Thresholds and Reported FPR Settings
Table B2. Scheme-Specific Decision Thresholds and Reported FPR Settings
Scheme and protocol Decision Threshold FPR / Protocol Source
StegaStamp (WatermarkAttacker) BA 0.614 (59/96 bits) 1% FPR [42]
StegaStamp (UnMarker) BA 0.63 Scheme threshold; FPR not reported [34]
SSL BA 0.719 (23/32 bits) 1% FPR [42]
HiDDeN BA 0.73; double-tail τ* ≈ 0.83 Scheme threshold / 10−4 FPR, respectively [35,57]
Stable Signature BA 0.69; τ = 0.77 Scheme threshold / below 10−5 FPR, respectively [35,74]
Yu (GAN generation) BA 0.61–0.63 Scheme threshold; FPR not reported [34,35]
PTW (GAN generation) BA 0.70 Scheme threshold; FPR not reported [34,35]
UDH τ* = 0.613 one-sided / 0.621 two-sided 10−4 FPR [57]
Generic k-bit τ = 0.9 / 0.83 / 0.73 for 20 / 30 / 64-bit messages ≤ 10−4 FPR [59]
Gaussian Shading (Müller et al.) BA 0.70703 for detection and attribution (k = 256) 10−6 FPR [87]
Gaussian Shading (PnP) Detection BA 0.5976525; attribution BA 0.6484375 10−3 FPR; 1,000-user attribution [89]
PRC Sign-flip rate approaching ~50% breaks BP decoding Coding-theoretic criterion; FPR not applicable [73,83]

B.2. Image Quality Classification Table

Table B3 presents four empirical reference bands for image quality. For metrics whose absolute ranges are relatively stable within the reviewed protocols, the table reports four survey-level categories: poor, fair, moderate, and good. These bands are comparative aids rather than universal thresholds across datasets or evaluation pipelines. For distribution-level, semantic-level, no-reference, or protocol-specific metrics, the table reports the metric direction and representative observed ranges only; these metrics should be compared against clean or no-attack baselines rather than read as fixed thresholds from Poor to Good.
Table B3. Image Quality Levels After Attack
Table B3. Image Quality Levels After Attack
Metric Poor Fair Moderate Good Sources and Anchors
PSNR (dB) (↑) 0–20 20–25 25–30 30+; clean watermarked images ≈ 37–39 [35] Crop 14.27 / Noise 26.22 / Blur 31.60; [42] 39.38
SSIM (↑) < 0.50 0.50–0.80 0.80–0.90 0.90–1.00; clean ≈ 0.98 [35] 0.46 / 0.53 / 0.92; [42] 0.983
MS-SSIM (↑) < 0.50 0.50–0.80 0.80–0.95 0.95–1.0 [87] high-fidelity forgery/removal examples report about 0.98–0.997
LPIPS (↓) > 0.30 0.20–0.30 0.10–0.20 < 0.10; near-imperceptible < 0.05 [35] Crop 0.43 → JPEG 0.02 / UnMarker 0.01
Watson-TPE (↓) > 30; rejected 7.6–30 4.1–7.6 ≤ 4.1; invisible [5] GT = 4.1 / LT1 = 7.6 / LT2 = 30
90th-percentile Q, 0–255 (↓) > 15 10–15 5–10 < 5 [46] DwtDctSvd 4.66 / RoSteALS 12.8 / StegaStamp 15.9
FID (↓) — — — Relative only; lower is closer to the clean/no-attack distribution; clean watermarked images ≈ 5–10 in [42] [42,71,78]
SIFID (↓) — — — Relative only; lower is closer at the single-image level; clean baseline ≈ 0.02–0.05 in [37] [37]
CLIP-FID (↓) — — — Relative only; lower is closer in CLIP-feature space; observed ≈ 4–11 in [44] [44]
CLIP Score (↑) — — — Relative only; higher or near-unchanged indicates better prompt-image alignment [72,78]
CLIP-I (↑) — — — Relative only; higher or near-unchanged indicates better image-to-image semantic similarity Directional reference; no fixed cross-paper bands
A-FINE (↓) — — — Relative only; lower indicates less degradation; observed ≈ 29–67 [35]
Q-Align (↑) — — — Relative/no-reference only; higher indicates better perceived quality [44]
LIQE / PickScore (↑) — — — Direction only; higher is better LIQE [44]; PickScore [41]
Wasserstein / NQD (↓) — — — Direction only; no absolute grading Wasserstein [45]; NQD [48,99]
Note. Only PSNR, SSIM, MS-SSIM, LPIPS, Watson-TPE, and 90th-percentile Q are used here as fixed grading references from Poor to Good. FID, SIFID, CLIP-FID, CLIP-Score, CLIP-I, A-FINE, Q-Align, LIQE, PickScore, Wasserstein, and NQD are relative, distributional, semantic, no-reference, or protocol-specific metrics and are reported by direction and by change against clean or no-attack baselines.

B.3. Comparative Analysis of Evaluation Metrics

Although evaluation metrics vary widely, they ultimately serve two basic judgments: whether the attack is effective and whether image fidelity remains acceptable. Attack effectiveness reduces to whether the watermark can still be reliably detected or extracted, whereas image fidelity focuses on whether the difference between the attacked image and the unattacked baseline remains within an acceptable range. Attack effectiveness should be judged against a scheme-specific threshold or calibrated operating point, while fidelity evaluation generally requires an unattacked image or an unattacked watermarked image as the reference. For example, multi-bit watermarks usually use decoding no better than random guessing as the failure line, which means that BA or BER approaches 0.5 [37,44]. Zero-bit watermarks usually use the loss of detector separability as the failure line, which means that AUC/AUROC approaches 0.5 or TPR drops significantly; for detectors using a statistical hypothesis test, the p-value must cross the paper-specific significance threshold under the source null hypothesis [42,43]. On the image-quality side, the unattacked watermarked image is commonly used as the anchor. For instance, clean DwtDctSvd reports a PSNR of 39.38, an SSIM of 0.983, and an FID range of 5.28 to 9.62 [42].
Differences among metrics mainly appear at four levels. First, attack-effectiveness metrics differ with the output form of the watermark. Multi-bit watermarks mainly use BER and BA [35], whereas zero-bit or semantic watermarks more often use AUC/AUROC [37,43,71], TPR@x%FPR [42,44,99,101], the p-value [65,70], and Inverse-Distance [34]. Second, image-fidelity metrics differ with the type of distortion. PSNR mainly characterizes pixel error, LPIPS reflects deep perceptual distance, FID measures distributional realism, and Watson-TPE and wPSNR focus more on human-visible distortion. Third, the same attack may lead to different conclusions under different metrics. For example, geometric translation can substantially reduce PSNR while leaving CLIP-based semantic alignment almost unchanged, so MarkCleaner introduces optical-flow-calibrated AlignPSNR and AlignSSIM to reduce the interference of geometric displacement in quality evaluation [78]. Fourth, metrics differ in their suitability for grading. PSNR, SSIM, LPIPS, Watson-TPE, and Q anchor ranges from Poor to Good within fixed intervals, whereas FID, SIFID, CLIP-FID, CLIP-Score, A-FINE, Q-Align, LIQE, PickScore, Wasserstein, and NQD are better suited to relative or directional comparison and should not be used as strict absolute quality thresholds. In addition, success thresholds themselves vary with FPR calibration. For example, StegaStamp uses a decision line of BA 0.614 at 1% FPR [42], UnMarker uses 63% [34], SSL uses 0.719 [42], and Gaussian Shading uses 0.707 at 1e-6 FPR [87]. The same BA value may therefore correspond to different success judgments across schemes and FPR settings.
These differences arise mainly from three factors. First, watermarking mechanisms differ. Multi-bit watermarks contain a directly decodable payload and are therefore naturally suited to BER or BA measurement, whereas diffusion zero-bit watermarks usually have no payload that can be compared bit by bit and rely more on statistical tests or detection scores. Second, the definition of image quality evolves with technical stages. During the conventional signal-processing period from 1998 to 2001, evaluation mainly concerns whether distortion is visible to the human eye; second-generation benchmarks therefore introduce wPSNR and Watson-TPE and show that images with the same PSNR may differ substantially in perceptual visibility as measured by TPE [5]. During the generative period from 2023 to 2026, evaluation pays more attention to whether images remain realistic and natural at the perceptual, semantic, and distributional levels, so LPIPS, FID/SIFID, and CLIP-FID are used more widely. Third, evaluation goals extend beyond the attack objectives of removal, evasion, and forgery to include forensic stealth as an additional evaluation dimension. The evaluation system therefore incorporates scheme-specific p-value decisions for detectors using the relevant null hypothesis, such as Tree-Ring-family tests, where forgery and removal cross opposite sides of the source protocol’s significance threshold [70]; attribution thresholds such as the Gaussian Shading attribution line of 0.648 at 1e-3 [89]; and the AUROC of forensic detectors [100].
The evolution of the metric system also tracks how the field has developed. Fidelity metrics move from signal-visibility measures such as wPSNR and Watson-TPE to deep perceptual and distributional realism measures such as LPIPS, FID, and CLIP-FID, which corresponds to the shift of attacks from noise addition and denoising to regeneration. A regeneration attack may have a relatively low PSNR but still bring FID close to the unattacked baseline, so distribution-level metrics are needed to reveal its actual cost. Success judgment also shifts from a single similarity threshold, such as the StirMark similarity score of 3.74 falling below the threshold of 21.08 [33], to low-false-positive evaluation calibrated by FPR, such as TPR@0.1%FPR [99,101] and scheme-specific p-values [65,70], which reflects increasing attention to deployment credibility. Evaluation frameworks further expand from simple removal tests to multi-objective criteria covering removal, evasion, forgery, and forensic indistinguishability [100], indicating a longer threat chain. Evaluation protocols also move from paper-specific settings toward unified benchmarks such as WAVES, Erasing the Invisible, and WIBE [99,101,103], improving cross-method comparability. Some thresholds carry theoretical meaning only for particular constructions: for PRC and related sign-coded latent watermarks, a sign-error rate approaching 1/2 is associated with information-theoretic undecodability [73,83]. Separately, impossibility results for strong watermarks show that, under quality and perturbation oracle/random walk assumptions, diffusion image watermarks can be removed [47].
A set of empirical results from MarkSweep illustrates the necessity of evaluating attack effectiveness and image fidelity separately. Cropping reduces PSNR to 14.27 dB and SSIM to 0.46 while increasing LPIPS to 0.43, but it does not cross the scheme-specific removal thresholds for any of the four evaluated watermarks. Severe degradation therefore does not by itself make an attack effective. JPEG preserves high image fidelity, with LPIPS as low as 0.02, but its effectiveness is target-dependent: it often fails to remove the watermark, although it reduces HiDDeN BA to 0.65, below that scheme’s 0.73 threshold. UnMarker crosses the reported removal thresholds for all four targets, yet its fidelity evidence is mixed: LPIPS ranges from 0.01 to 0.21, whereas PSNR ranges from 15.84 to 23.54 dB and SSIM from 0.42 to 0.62 [34,35]. These results show that neither visible degradation nor one favorable fidelity metric is sufficient to establish a practical attack. Only by placing the attack-effectiveness grading in Table 4 and the image-quality grading in Table B3 side by side can one judge whether an attack satisfies both effectiveness and usability requirements.

C. Detailed Quantitative Results and Method-Specific Defense Considerations

C.1. Conventional Signal-Processing Attacks

Table C1 summarizes the seven papers in this section row by row according to attack method × target watermark family. Each row corresponds to one target watermark family. If the same attack spans multiple watermark families, it is split into multiple rows so that the evaluation data correspond one-to-one to the target watermark category. Attack effectiveness and image fidelity are judged according to the grading framework introduced in Section 4 and detailed in Appendix B. The method-specific evidence following the table reports detailed quantitative results, fidelity measurements, and defense considerations.
Table C1. Summary of Conventional Signal-Processing Attacks
Table C1. Summary of Conventional Signal-Processing Attacks
Attack method (author [ref.]) Target watermark (family + algorithm) Attack effectiveness Image fidelity Threat model
Voloshynovskiy et al. [5] (Optimal Estimation) [W1] Schemes A, B, C, Pereira DCT, Digimarc Effective Good No-box
Kassis and Hengartner [34] (UnMarker) [W2] HiDDeN, StegaStamp Effective Poor to good across metrics No-box
Kassis and Hengartner [34] (UnMarker) [W3] Yu1, Yu2, PTW, Stable Signature Effective Poor to good across metrics No-box
Kassis and Hengartner [34] (UnMarker) [W4] Tree-Ring Borderline effective Poor to good across metrics No-box
Cao et al. [35] (MarkSweep) [W2] HiDDeN Effective Moderate No-box
Cao et al. [35] (MarkSweep) [W3] Stable Signature, Yu, PTW Mixed Moderate No-box
Wu et al. [36] (Box-Blur, image-only blur/deblur) [W2] SD+HiDDeN Effective Relative distributional degradation only No-box
Wu et al. [36] (generator overwriting) [W2] SD+HiDDeN Effective Relative distributional quality only White-box
Yang et al. [37] (Averaging) [W1] DwtDctSvd, DwtDct Mixed (DwtDctSvd: success; DwtDct: failure) Good to moderate No-box
Yang et al. [37] (Averaging) [W2] RoSteALS, RAWatermark, RivaGAN, SSL, HiDDeN Mixed (RoSteALS, RAWatermark: success; RivaGAN, SSL, HiDDeN: failure) Moderate No-box
Yang et al. [37] (Averaging) [W3] Stable Signature Failed Not applicable No-box
Yang et al. [37] (Averaging) [W4] Tree-Ring, Gaussian Shading Effective (Tree-Ring, Gaussian Shading: success) Poor No-box
You and Zhou [38] (Two-Stage SS) [W1] Additive SS, ISS, CW-ISS, RNW, TSW Weak to effective Good Gray-box
You et al. [39] (EK Estimation) [W1] Additive SS, ISS, CW-ISS, NW, TSW Effective for key recovery Not reported for removal/forgery Gray-box

C.1.1. Method-Specific Quantitative Evidence and Defense Considerations

The following method-specific results and defense considerations are drawn from the corresponding studies [5,34,35,36,37,38,39].
Optimal Estimation and perceptual remodulation. In the attack-specific remodulation experiment, the detector fails on schemes A, B, and C [W1], BER reaches 60%–80%, and attacked-image PSNR is 35.52, 35.85, and 38.51 dB, respectively. The separate survival totals of 0.79, 1.75, and 3.52 are aggregate scores over six benchmark attack categories and are therefore not used to grade Optimal Estimation alone; they show only that scheme C, based on nonadaptive discrete cosine transform (DCT), is the most resistant overall. Image quality is also constrained by weighted PSNR and Watson-TPE; when noise visibility exceeds the rejection threshold, the image is judged unacceptable. The watermark power spectrum should be shaped in relation to the host power spectrum so that cover estimation becomes less reliable. Content-adaptive and key-dependent embedding can further weaken single-image statistical estimation, while information-theoretic embedding, including improved spread spectrum (ISS) or Quantization Index Modulation (QIM), can reduce residual estimability.
UnMarker. With 10% cropping under an LPIPS budget, the optimized high- and low-frequency perturbations reduce the detection rate below 50%. For HiDDeN [W2], bit accuracy decreases from 99.16% to 58.0% and the detection rate drops to 0. For StegaStamp [W2], bit accuracy decreases from 99.95% to 61.49% and the detection rate drops to 43%. For Yu1, Yu2, PTW, and Stable Signature [W3], bit accuracy decreases to 59.08%, 53.03%, 62.75%, and 52.82%, respectively. For Tree-Ring [W4], inverse distance decreases from 0.022 to 0.014, close to the source threshold of 0.0141, and the detection rate decreases from 100% to 40%. Across these targets, LPIPS is 0.05 to 0.15, while reported Fréchet Inception Distance (FID) spans 11.9 to 60.73. The MarkSweep comparison reports PSNR of approximately 15.8 to 23.5 dB and SSIM of 0.42 to 0.62 [35]. These metrics indicate that the fidelity assessment is protocol dependent; FID is treated only as a relative distributional measure. Defenses should avoid concentrating watermark energy in predictable frequency bands, use multi-band redundancy, and bind evidence more strongly to image content, a distribution-preserving latent space, or initial noise. Adversarial training or purification alone may be insufficient because UnMarker removes spectral evidence from the carrier rather than merely suppressing a transient perturbation.
MarkSweep and Box-Blur. MarkSweep reduces HiDDeN [W2] bit accuracy to 51.32%, with PSNR 27.85 dB, SSIM 0.82, and LPIPS 0.23. For Stable Signature and Yu [W3], bit accuracy is 66.83% and 59.24%, respectively, whereas PTW remains at 72.19%. The corresponding outputs have PSNR of approximately 28.69 to 29.75 dB, SSIM of 0.86 to 0.88, and LPIPS of 0.17 to 0.20. The no-box Box-Blur image attack reduces the bit accuracy of model-specific SD+HiDDeN [W2] to 0.379 before deblurring and 0.491 after deblurring, with FID 99.61 and 78.33. The separate white-box generator-overwriting variant reaches 0.6479 in the strongest reported 32-bit setting, with FID 41.81. Neither variant reports PSNR or SSIM, so FID supports only variant-specific relative distributional comparison rather than an absolute image-quality grade. Defenses can distribute evidence across low, middle, and high frequency bands and couple the payload to distribution-preserving latent structures. Generator overwriting additionally motivates generator-weight protection and integrity checks.
Content-agnostic averaging. For Tree-Ring [W4], AUC decreases from 1.000 to 0.241, while PSNR, SSIM, and LPIPS are 15.47 dB, 0.548, and 0.425. The AUC below 0.5 indicates inverted score ordering under the fixed detector rather than complete loss of recoverable separability. For DwtDctSvd [W1], bit accuracy decreases from 1.000 to 0.317 under the in-distribution setting, with PSNR 38.67 dB, SSIM 0.977, and LPIPS 0.013. DwtDct [W1] remains at 0.998. For RoSteALS [W2], bit accuracy decreases from 0.994 to 0.244, with PSNR 24.77 dB, SSIM 0.838, and LPIPS 0.059. RAWatermark [W2] is weakened from AUC 0.714 to 0.502, whereas RivaGAN, SSL Watermarking, and HiDDeN [W2] retain bit accuracies of 0.967, 0.917, and 0.961. Stable Signature [W3] remains at 0.998. Gaussian Shading [W4] decreases to 0.462, although its low PSNR of 10.12 dB mainly reflects the watermark itself rather than additional averaging degradation. Content-adaptive or per-image randomized embedding, cryptographically seeded key diversification, and multi-key mixing directly address the shared-pattern assumption. In the reported three-key mitigation ablation, multi-key mixing increases Tree-Ring AUC from below 0.2 to above 0.7.
Secret-carrier and equivalent-key estimation. Two-Stage SS reaches removal rates of 0.20 to 0.38 on Additive SS, ISS, CW-ISS, RNW, and TSW [W1], with PSNR 37.10 to 38.59 dB and SSIM 0.987 to 0.989. EK Estimation achieves carrier normalized correlations of 87.84% to 95.78%. For most schemes, the bit error rate obtained with the estimated carrier is close to that obtained with the true key, while NW retains a larger gap. The EK experiments use an embedding PSNR of approximately 40 dB, but this is an experimental condition rather than the fidelity of a removal or forgery output. EK Estimation evaluates hundreds to 2,600 observations, while Two-Stage SS uses BOWS2 owner images and BOSSBase public images, with 1,000 training and 1,000 test images. Defenses should avoid reuse of a single secret carrier, diversify keys through cryptographic per-image seeding, and limit exposure of watermarked-signal and message observations. Nonlinear embedding or information-theoretic security constraints can further reduce linear carrier estimability.

C.2. Geometric and Synchronization Attacks

Table C2 summarizes the desynchronization attacks and coding limits in this section by attack method and target watermark family. Defense and survey papers are used as countermeasures or quantitative references for vulnerability and are not listed as separate attack rows. The PRC row combines the empirical crop-and-resize attack with the coding-theoretic limit that explains its failure boundary. Attack effectiveness and image fidelity are judged according to the grading framework introduced in Section 4 and detailed in Appendix B. The supporting measurements and method-specific defense considerations follow the table. For geometric attacks, raw PSNR may be low because of global image displacement, so aligned metrics after registration must also be considered.
Table C2. Summary of Geometric and Synchronization Attacks
Table C2. Summary of Geometric and Synchronization Attacks
Attack method (authors [ref.]) Target watermark (family + algorithm) Attack effectiveness Image fidelity Threat model
Petitcolas et al. [33] (StirMark and jitter) [W1] NEC/Cox, SysCoP Effective Not graded (qualitative only) No-box
Barni et al. [77] (LPCD/MF-DA) [W1] SS-DFT, SS-DWT, dirty-paper trellis, orthogonal dirty paper, Cox SS-DCT Effective Not graded (qualitative only) No-box
Kong et al. [78] (MarkCleaner) [W1] DwtDct Effective Poor raw / moderate aligned No-box
Kong et al. [78] (MarkCleaner) [W2] SSL Watermarking, StegaStamp, WOFA, VINE Effective Poor raw / moderate aligned No-box
Kong et al. [78] (MarkCleaner) [W3] Stable Signature Effective Poor raw / moderate aligned No-box
Kong et al. [78] (MarkCleaner) [W4] Tree-Ring, RingID, Gaussian Shading, T2SMark, HSTR, HSQR Effective Poor raw / moderate aligned No-box
Shamshad et al. [79] (RAVEN) [W1] DwtDct, DwtDctSvd Effective Relative semantic/distributional quality reported globally No-box
Shamshad et al. [79] (RAVEN) [W2] RivaGAN, TrustMark, StegaStamp, VINE Mixed (RivaGAN, TrustMark: effective; StegaStamp, VINE: weakened only) Relative semantic/distributional quality reported globally No-box
Shamshad et al. [79] (RAVEN) [W3] Stable Signature Effective Relative semantic/distributional quality reported globally No-box
Shamshad et al. [79] (RAVEN) [W4] Tree-Ring, ZoDiac, HSTR, RingID, HSQR, ROBIN, Gaussian Shading Effective Relative semantic/distributional quality reported globally No-box
Zhao et al. [84] (SSyncOA crop-paste) [W2] ARWGAN, RoSteALS Mixed (RoSteALS: effective; ARWGAN: partial) Not reported after attack No-box
Francati et al. [83] (PRC crop-and-resize / coding limit) [W4] PRC Effective Not graded (qualitative only) No-box

C.2.1. Method-Specific Quantitative Evidence and Defense Considerations

Cross-family geometric sensitivity. The evaluation in SynTag [81] reports TPRs under geometric distortion of 0.004 for DwtDctSvd [W1] and 0.016 for Gaussian Shading [W4]. RivaGAN [W2], MBRS [W2], Stable Signature [W3], and LaWa [W3] retain TPRs of 0.788, 0.744, 0.760, and 0.764, respectively. GeoWM [80] likewise reports that under a strong affine setting with scaling parameter 50 and rotation 180 ∘ , bit accuracy falls to 69.59% for SSL Watermarking [W2] but remains 85.19% for MBRS [W2]. These results show that spatial redundancy and binding during generation change the response to geometric distortion, although neither property alone guarantees synchronization robustness. GeoWM introduces geometric distortion layers during training, SynTag adds an explicit synchronization tag to inversion-based generative watermarks, and PrintCamera [82] combines perspective simulation with spatial transformation before decoding.
Classical local desynchronization. StirMark [33] reduces the correlation similarity of the NEC/Cox DCT spread-spectrum watermark [W1] to 3.74, below its decision threshold of 21.08, while SysCoP reports that no valid watermark can be found after jitter. The paper describes the image changes qualitatively as visually slight and does not report PSNR or SSIM. It also presents a separate white-box implementation attack on Digimarc/PictureMarc in which debugger access permits the copyright identifier to be overwritten; this result should not be conflated with the no-box geometric variants. In the experiments of Barni et al. [77], detections fall from 6 to 0 for SS-DFT and orthogonal dirty paper, and from 6 to 3 for dirty-paper trellis. LPCD does not defeat SS-DWT, which remains at 6 detections, whereas MF-DA reduces it to 0 at σ = 3 . The paper evaluates fidelity qualitatively through smooth displacement fields. The countermeasures reviewed by Licks and Jordan [26] include invariant domains, synchronization templates, autocorrelation-based self-synchronization, feature-point embedding, and registration before decoding.
Micro-geometric phase perturbation. MarkCleaner [78] evaluates 12 W1–W4 schemes. The TPR is 0.000 for Stable Signature and 0.001 for Tree-Ring, while the reported bit accuracies approach random guessing. Residual TPRs are higher for HSTR and HSQR at 0.051 and 0.022. Raw PSNR and SSIM are 16.52 dB and 0.444, whereas AlignPSNR and AlignSSIM after registration reach 25.30 dB and 0.820. The difference confirms that geometric displacement, rather than content replacement, accounts for much of the raw pixel error. A detector-side response is to estimate and correct the displacement before decoding; combining registration with invariant-domain or spatially redundant evidence may further reduce sensitivity to phase shifts.
Latent viewpoint transformation. For Tree-Ring, ZoDiac, HSTR, RingID, HSQR, and ROBIN [W4], RAVEN [79] reports TPR@1%FPR values of 0.020, 0.067, 0.025, 0.018, 0.015, and 0.012. Bit accuracy for DwtDct, DwtDctSvd, RivaGAN, TrustMark, Stable Signature, and Gaussian Shading falls within 0.472–0.540, whereas StegaStamp and VINE remain at 0.575 and 0.588 and are therefore weakened rather than fully randomized. The reported global FID of 40.18 and CLIP text alignment of 0.328 indicate relative distributional and semantic quality only within the paper’s comparison protocol; they do not define an absolute fidelity grade. Runtime is about six seconds per image. RAVEN is sensitive to viewpoint strength, and stronger transformations can introduce structural or color changes. Method-specific defenses include viewpoint-aware training and synchronization tags that permit alignment before detection or inversion.
Object relocation. In the crop-paste comparison reported by Zhao et al. [84], bit accuracy is 51.00% for RoSteALS and 63.50% for ARWGAN [W2], whereas SSyncOA128 [W5] retains 98.70%. Under the Section 4 criterion that bit accuracy should approach 0.5 for effective payload randomization, the result is effective for RoSteALS but establishes only partial weakening for ARWGAN. The separate combined-noise model reports 98.08% bit accuracy, PSNR 43.93 dB, and SSIM 0.9956; these values characterize the combined-noise setting rather than the crop-paste experiment. The method uses segmentation-guided object localization and geometric normalization so that the watermark region follows the object through translation, rotation, and scaling. These results apply to the studied single-object setting and do not establish performance for multiple interacting objects or complete scene reconstruction.
PRC crop-and-resize attack and coding limit. Francati et al. [83] report that cropping 15 pixels from each side of a 512 × 512 image and resizing it back causes detection to fail in more than 1,000 trials, while the pre-decoding sign-error rate approaches 50%. Fidelity is assessed only through side-by-side visual evidence, without PSNR, SSIM, LPIPS, or FID, and is therefore not graded. The binary coding boundary is α * = 1 / 2 , while the general q-ary bound is α * = 1 − 1 / q . The theorem assumes independent symbol corruption, whereas crop-and-resize is a structured image transformation, so the empirical result illustrates the same numerical regime rather than directly instantiating the theorem’s tampering channel. Additional redundancy or stronger error correction may improve tolerance below the boundary but cannot remove the upper bound; coding protection should therefore be paired with geometric synchronization.

C.3. Regeneration and Purification Attacks

Table C3 summarizes the 15 attack papers in this section by attack method and target watermark lineage, with one row corresponding to one target lineage. TrustMark [50] and RRWID [56] are cited in the main text as robust targets and defensive countermeasures, but are not listed as attack rows. Text and video targets are excluded under the scope of invisible image watermarks. Attack effectiveness and image fidelity are judged according to the grading framework introduced in Section 4 and detailed in Appendix B; the supporting metric values and method-specific defense considerations are reported after the table.
Table C3. Summary of Regeneration and Purification Attacks
Table C3. Summary of Regeneration and Purification Attacks
Attack Method (Author [Ref.]) Target Watermark (Lineage + Algorithm) Attack Effectiveness Image Fidelity Threat Model
Zhao et al. [42] (WatermarkAttacker) [W1] DwtDctSvd Effective Good No-box
Zhao et al. [42] (WatermarkAttacker) [W2] RivaGAN, SSL Watermarking, StegaStamp Mixed (RivaGAN, SSL Watermarking: success; StegaStamp: success only under heavy noise) Good to poor No-box
Zhao et al. [42] (WatermarkAttacker) [W4] Tree-Ring Failed Not applicable No-box
Saberi et al. [43] (Diffusion Purification) [W1] DwtDct, DwtDctSvd Effective Fair to moderate No-box
Saberi et al. [43] (Diffusion Purification) [W2] RivaGAN, MBRS; StegaStamp Mixed Fair to moderate No-box
Saberi et al. [43] (Diffusion Purification) [W4] WatermarkDM; Tree-Ring Mixed Fair to moderate No-box
Liu et al. [44] (CtrlRegen) [W1] DwtDctSvd Effective Poor No-box
Liu et al. [44] (CtrlRegen) [W2] RivaGAN, SSL Watermarking, StegaStamp Effective Poor No-box
Liu et al. [44] (CtrlRegen) [W3] Stable Signature Effective Poor No-box
Liu et al. [44] (CtrlRegen) [W4] Tree-Ring Effective Poor No-box
Alam et al. [45] (SADRE) [W1] DwtDct, DwtDctSvd Not established (payload corruption only) Good No-box
Alam et al. [45] (SADRE) [W2] RivaGAN, StegaStamp, EditGuard Not established (payload corruption only) Good No-box
Alam et al. [45] (SADRE) [W4] Tree-Ring Not established Good No-box
Liang et al. [46] (DIP) [W1] DwtDctSvd Effective Good No-box
Liang et al. [46] (DIP) [W2] RivaGAN, SSL Watermarking; RoSteALS, StegaStamp Mixed Good to poor No-box
Liang et al. [46] (DIP) [W4] Tree-Ring Failed Not applicable No-box
Zhang et al. [47] (WiTS) [W1] DwtDct Effective Protocol-specific relative quality No-box
Zhang et al. [47] (WiTS) [W3] Stable Signature Effective under paper-specific threshold Protocol-specific relative quality No-box
Bulychev et al. [48] (Re-Watermarking) [W2] StegaStamp, Pixel Seal, WAM, RoSteALS Mixed (StegaStamp, Pixel Seal: success; WAM, RoSteALS: partial) Protocol-specific quality budget No-box
Bulychev et al. [48] (Re-Watermarking) [W3] Stable Signature Effective Protocol-specific quality budget No-box
Bulychev et al. [48] (Re-Watermarking) [W4] Tree-Ring, ZoDiac Mixed Protocol-specific quality budget No-box
Wang et al. [40] (HAI-WAN) [W1] QPHFMs, QDFT, LSB Weak to effective Good Gray-box training / No-box inference
Yan et al. [49] (UQP-WR) [W1] DwtDct, DwtDctSvd, Digimarc Effective Good Black-box training / No-box inference
Yan et al. [49] (UQP-WR) [W2] RivaGAN, HiDDeN, StegaStamp Mixed Good to moderate Black-box training / No-box inference
Wang et al. [52] (BDMWA, bidirectional diffusion) [W1] LSB, DCT, PHFMs Partial to effective Good No-box
Wang et al. [52] (BDMWA, bidirectional diffusion) [W2] HiDDeN, MBRS Mixed (HiDDeN: success; MBRS: partial) Good No-box
Wang et al. [52] (BDMWA, bidirectional diffusion) [W4] GaussMarker Mixed Good No-box
Meshram and Chandrasekaran [41] (D2RA/DAWN) [W1] DwtDct, DwtDctSvd Effective Poor to moderate No-box
Meshram and Chandrasekaran [41] (D2RA/DAWN) [W2] RivaGAN, TrustMark, SSL Watermarking, InvisMark, WAM Effective Poor to moderate No-box
Meshram and Chandrasekaran [41] (D2RA/DAWN) [W4] ZoDiac, Tree-Ring, PRC, Gaussian Shading, SFW Mixed (ZoDiac, Tree-Ring, PRC: success; Gaussian Shading, SFW: failure) Poor No-box
Wang et al. [54] (HIWANet) [W1] QPHFMs, LSB, QFT Weak Good No-box
Wang et al. [55] (RD-IWAN) [W1] QPHFMs, LSB, QFT, SVD Weak Good No-box
Li and Wang [51] (Concealed Attack) [W1] QEM frequency-moment watermark Weak Good to moderate No-box
Shamshad et al. [53] (First-Place Solution) [W2] Modified StegaStamp Effective Protocol-specific quality score Gray-box named-target track; Black-box separate hidden-target track
Shamshad et al. [53] (First-Place Solution) [W4] Tree-Ring Effective Protocol-specific quality score Gray-box named-target track; Black-box separate hidden-target track

C.3.1. Method-Specific Quantitative Evidence and Defense Considerations

Bounded-noise regeneration and diffusion purification. WatermarkAttacker [42] achieves removal rates above 99% for DwtDctSvd [W1] and SSL Watermarking [W2], and 98% for RivaGAN [W2], while keeping PSNR above 30 dB. StegaStamp [W2] requires substantially stronger diffusion noise, with PSNR around 17 dB and SSIM around 0.38, and Tree-Ring [W4] remains detectable with TPR between 0.994 and 1.000. At t = 0.2 , the diffusion-purification route of Saberi et al. [43] reduces AUROC to 0.542–0.644 for RivaGAN [W2], DwtDct and DwtDctSvd [W1], WatermarkDM [W4], and MBRS [W2], with PSNR around 26 dB and SSIM around 0.72. StegaStamp and Tree-Ring retain AUROC values of 0.966 and 0.976. The First-Place Solution [53] reduces the detection score in the black-box track to 0.043, corresponding to about 95.7% removal, with a challenge quality-degradation score of 0.136. Increasing the embedding budget and spatial redundancy can improve resistance but directly competes with imperceptibility. TrustMark [50] resists Zhao-style regeneration at the tested setting, and RRWID [56] resists VAE removal and sanitization but remains vulnerable to diffusion-class purification. Distribution-preserving or generation-related evidence, combined with cryptographic provenance and regeneration forensics, provides a broader defense than reliance on post-hoc robustness alone.
Controllable regeneration from clean noise. CtrlRegen [44] reduces TPR@1%FPR to 0.00–0.06 and drives bit accuracy toward 0.5 for DwtDctSvd [W1], RivaGAN, SSL Watermarking, StegaStamp [W2], and Stable Signature [W3]. Tree-Ring [W4] falls to 0.12. Pixel-level PSNR is approximately 19 dB because the content is redrawn, although the paper reports strong no-reference image-quality scores. Increasing only the watermark perturbation budget is unlikely to address clean-noise regeneration; content-bound or semantic evidence should instead be combined with regeneration forensics.
Region- and band-targeted reconstruction. SADRE [45] reduces bit recovery accuracy to 0.40–0.48 for DwtDct and DwtDctSvd [W1], RivaGAN, StegaStamp, and EditGuard [W2], while reporting PSNR of 32–35 dB and SSIM of 0.85–0.95. These results indicate weakening, but the paper does not provide detector-evasion measurements. Its Tree-Ring [W4] result requires caution because bit recovery accuracy is not the native metric of a zero-bit detector. DIP [46] exceeds 90% evasion on high-frequency targets such as DwtDctSvd [W1], RivaGAN, and SSL Watermarking [W2], with typical successful-evasion quality of 35–37 dB PSNR, SSIM of 0.92–0.97, and 90%-quantile pixel differences of 6–8 at strict thresholds. It performs poorly on low- and mid-frequency RoSteALS and StegaStamp [W2] and does not reliably evade Tree-Ring [W4]. Multi-band redundancy and low- or middle-frequency content coupling reduce the risk that one localized reconstruction route removes all watermark evidence.
Quality-preserving random walks. Under the paper-specific tests of WiTS [47], the average p-value of Stable Signature [W3] rises to 0.059, with average CLIP-Score changing from 33.91 to 33.40. DwtDct [W1] reaches an average p-value of 0.206, with CLIP-Score changing from 35.64 to 35.51; a representative image reaches a p-value of 0.235 while its CLIP-Score changes from 34.82 to 33.60. Values above the paper-specific significance threshold indicate that detection is no longer significant, but larger p-values should not be interpreted as proportionally stronger removal. The conditional impossibility result motivates cryptographic provenance and attack forensics because no single watermark can be assumed information-theoretically irremovable under repeated oracle-guided rewriting.
Re-watermarking. Under its Normalized Quality Degradation budget, Re-Watermarking [48] produces absolute bit-accuracy drops of 25 to 48.5 percentage points across the five in-scope multi-bit image schemes for which bit accuracy is reported. Pixel Seal [W2] drops by 48.5 points, Stable Signature [W3] by 45 points, StegaStamp [W2] by 41 points, WAM [W2] by 29 points, and RoSteALS [W2] by 25 points. Its automatic routing classifier is trained offline on 500 labeled images per class. ZoDiac overwriting drives Stable Signature TPR toward zero; for Tree-Ring [W4], the main tables report TPR values of 0.280 and 0.370, while a figure-level operating point approaches zero. Defenses should use key-bound location randomization, authenticate the registered embedding event, and detect abnormal secondary watermark insertion.
HAI-WAN and UQP-WR. HAI-WAN [40] reports average BER values of 0.1517 on QPHFMs, 0.2453 on QDFT, and 0.3022 on LSB [W1]. On Lena, BER is 0.152 at a noise standard deviation of 25, still far from random guessing, while average PSNR is 32.15 dB. The method uses a known target scheme to synthesize supervised training data. UQP-WR [49] raises BER to 40%–44% for DwtDct and DwtDctSvd [W1], RivaGAN and StegaStamp [W2], with PSNR around 37–40 dB for DwtDct, DwtDctSvd, RivaGAN, and Digimarc, 32.01 dB for HiDDeN, and 29.16 dB for StegaStamp; SSIM is approximately 0.90–0.98. HiDDeN is the hardest target, with BER of 23.84%, and Digimarc reaches a 93% attack success rate. UQP-WR removes the need for paired samples, although watermarked outputs are still collected through a black-box generation interface.
BDMWA and D2RA/DAWN. BDMWA [52] raises BER on LSB, DCT, and PHFMs [W1] and HiDDeN [W2], with PSNR of 42.64–48.90 dB on CelebA. Removal remains partial for several schemes, including PHFMs at approximately 0.27–0.28 BER. Transfer to unseen MBRS [W2] and GaussMarker [W4] is weaker than same-scheme testing. D2RA/DAWN [41] reaches attack success rates of 99.8% and 98.4% on DwtDct and DwtDctSvd [W1], and 85.4%–100% on RivaGAN, TrustMark, SSL Watermarking, InvisMark, and WAM [W2]. The corresponding fidelity values are approximately 28.2 dB PSNR, 0.54–0.57 SSIM, and 0.47–0.48 LPIPS. On ZoDiac, Tree-Ring, and PRC [W4], success rates are 92.2%, 70.2%, and 67.8%; Tree-Ring fidelity falls to 14.56 dB PSNR, 0.46 SSIM, and 0.64 LPIPS. Gaussian Shading and SFW are not effectively attacked in the reported setting.
HIWANet, RD-IWAN, and Concealed Attack. HIWANet [54] reports BER of approximately 0.10–0.25 on QPHFMs [W1], with PSNR around 31–32 dB. RD-IWAN [55] reports a default QPHFMs BER of 0.0774, illustrating the residual evidence left by restoration toward the original image. Concealed Attack [51] weakens the QEM frequency-moment watermark [W1], but its BER remains modest; its outputs have PSNR of approximately 27–29 dB and SSIM around 0.99. These supervised methods favor fidelity over complete removal. Defenses against the broader learned-reconstruction family should distribute evidence across low and middle frequency bands, couple it to image content or generation semantics, and evaluate transfer against both target-specific and universal removers.

C.4. Adversarial Perturbation Attacks

Table C4 summarizes the in-scope adversarial attacks on W1–W4 watermark schemes reviewed in this section, with one row for each attack method and target watermark lineage. Box-free model watermarks [62,63,64] are model IP watermarks outside the survey scope and are discussed only as boundary cases in Appendix E; accordingly, they are not included in this table. Tree-Ring and related schemes that appear only in Hu’s side robustness study but are not actually attacked are also excluded. Attack effectiveness and image fidelity are judged according to the grading framework introduced in Section 4 and detailed in Appendix B; the supporting values and method-specific defense considerations follow the table.
Table C4. Summary of Adversarial Perturbation Attacks
Table C4. Summary of Adversarial Perturbation Attacks
Attack Method (Author [Ref.]) Target Watermark (Lineage + Algorithm) Attack Effectiveness Image Fidelity Threat Model
Quiring and Rieck [31] (AdvML) [W1] Broken Arrows Effective Good Black-box
Jiang et al. [57] (WEvade) [W1] DwtDct Effective Relative perturbation evidence only White-box / Black-box
Jiang et al. [57] (WEvade) [W2] HiDDeN, UDH Effective Relative perturbation evidence only White-box / Black-box
Lukas et al. [58] (Adaptive Attack) [W1] DWT, DWT-SVD Effective Good Gray-box
Lukas et al. [58] (Adaptive Attack) [W2] RivaGAN Effective Good Gray-box
Lukas et al. [58] (Adaptive Attack) [W4] Tree-Ring, WatermarkDM Effective Good Gray-box
Hu et al. [59] (Transfer Attack) [W1] DwtDct Effective Good No-box
Hu et al. [59] (Transfer Attack) [W2] HiDDeN, StegaStamp, Smoothed HiDDeN, Smoothed StegaStamp Effective Good No-box
Hu et al. [59] (Transfer Attack) [W3] Stable Signature Effective Good No-box
Chen et al. [61] (FAADW) [W1] DwtDctSvd Effective payload disruption Moderate No-box
Chen et al. [61] (FAADW) [W2] StegaStamp, HiDDeN, SSL Watermarking Effective to good payload disruption Moderate No-box
Lukas and Kerschbaum [60] (Reverse Pivotal Tuning against PTW) [W3] PTW Mixed (white-box: success; black-box: failure) Relative distributional change only White-box / Black-box

C.4.1. Method-Specific Quantitative Evidence and Defense Considerations

Substitute and transfer attacks. Quiring and Rieck [31] report removal on 100% of the evaluated Broken Arrows [W1] images, with PSNR values of 35.6 and 38.8 dB in the reported settings. Hu et al. [59] obtain evasion rates of 1.0 on Stable Signature [W3] and Smoothed HiDDeN [W2] with 100 surrogate models, and approximately 0.78 on HiDDeN and StegaStamp [W2]. Reported SSIM remains above 0.92, although the normalized ℓ ∞ perturbation can be nontrivial and the per-image optimization cost is high. The transfer attack does not attack Tree-Ring, RingID, or Gaussian Shading; those schemes occur only in a side robustness study. Defenses should reduce reusable decoder structure, diversify keyed verification across images, and evaluate transfer outside any certified perturbation radius rather than treating bounded certification as general attack resistance.
WEvade. For 30-bit HiDDeN [W2], whose reported detector threshold is approximately 0.83, WEvade drives bit accuracy toward 0.5 and reaches an evasion rate near 1. The white-box ℓ ∞ perturbation is approximately 0.008, while WEvade-B-Q reaches the same evasion objective with a perturbation of approximately 0.032 under a 30-query budget. UDH [W2], evaluated with 256-bit payloads and thresholds of 0.613/0.621, and the 136-bit DwtDct watermark [W1] used with Stable Diffusion are also evaded at rates near 1. These perturbation magnitudes provide relative fidelity evidence but do not support a fixed grade from Poor to Good under the Section 4 references; the source-specific detector thresholds should likewise not be generalized across decoders. Defenses should protect decoder access, rate-limit and screen detector queries, randomize the decision interface, and include double-tail and adversarially trained verification in evaluation.
Adaptive Attack. Lukas et al. [58] report that the main adaptive preparation can be completed in less than one GPU hour. Adversarial noising reduces Tree-Ring [W4] detection to 5.2% at ϵ = 2 / 255 with LPIPS 0.032. The optimized compression attack reduces detection to 2.0% for WatermarkDM [W4], 0.8% for DWT [W1], 1.9% for DWT-SVD [W1], and 6.3% for RivaGAN [W2], with LPIPS values of 0.056, 0.077, 0.075, and 0.073, respectively. FID and CLIP-Score do not differ significantly from the watermarked baselines under the reported protocol and are interpreted only as relative distributional and semantic evidence. Pure noising is less reliable, leaving WatermarkDM near 30% and RivaGAN near 67% detection in the corresponding tests; non-differentiable DWT variants also require a trained ResNet-50 surrogate extractor. Defenses should increase key entropy, prevent local reproduction of verifier structure, and stress-test each deployment against adaptive compressors trained once and then reused rather than only fixed codecs.
FAADW. On its primary StegaStamp [W2] target, FAADW [61] reports BER 0.32, PSNR 27.31–27.33 dB, and SSIM 0.862. Transfer experiments report BER/PSNR/SSIM of 0.31/26.75 dB/0.887 on DwtDctSvd [W1], 0.36/26.21 dB/0.894 on HiDDeN [W2], and 0.30/27.59 dB/0.912 on SSL Watermarking [W2]. These BER values fall within the Effective and Good payload-disruption bands of Appendix B but remain below the 0.5 random-guessing reference, and the paper does not report a calibrated detector-evasion threshold. The corresponding image fidelity is Moderate under the Section 4 references. Defenses should make residual evidence less spatially separable, distribute it across complementary features, and train against learned residual cancellation as well as generic noise.
Reverse pivotal tuning. With 200 real images, the white-box attack studied with PTW [60] reduces watermark capacity on StyleGAN2, StyleGAN-XL, and StyleGAN3 from 43.05/48.79/40.33 to 4.91/4.52/4.59. FID changes from 5.4/2.67/6.61 to 5.47/3.52/6.65, respectively, and is interpreted only relative to each model’s baseline. By contrast, the strongest evaluated black-box distortion reduces StyleGAN2 capacity only to 32.86 while changing FID to 11.51. The resulting Mixed judgment in Table C4 therefore separates effective white-box parameter tuning from ineffective black-box output distortion. Defenses should protect generator integrity, monitor unauthorized fine-tuning, and bind ownership evidence to authenticated model states rather than relying on a parameter watermark alone.

C.5. Semantic Editing Attacks

Table C5 summarizes the semantic editing attacks in this subsection by attack method and target watermark lineage. Robust-Wide and VINE are defense studies and are therefore not listed as independent attack rows; W-Bench supplies the attack-side evidence for the instruction-driven editing rows. For instruction-driven editing, W-Bench reports semantic preservation and clean baseline quality rather than an attacked-image fidelity grade directly comparable with Section 4 and Appendix B. The supporting values and method-specific defense considerations follow the table.
Table C5. Summary of Semantic Editing Attacks
Table C5. Summary of Semantic Editing Attacks
Attack Method (Author [Ref.]) Target Watermark (Lineage + Algorithm) Attack Effectiveness Image / Semantic Fidelity Threat Model
Instruction-driven editing, UltraEdit-Global, evaluated in W-Bench [68] [W1] DwtDct, DwtDctSvd Effective Semantic preservation only No-box
Instruction-driven editing, UltraEdit-Global, evaluated in W-Bench [68] [W2] RivaGAN, MBRS, SSL Watermarking, CIN, EditGuard; PIMoG, TrustMark, StegaStamp, SepMark Mixed Semantic preservation only No-box
Tallam et al. [65] (SemanticRegen) [W1] DwtDct Effective Good No-box
Tallam et al. [65] (SemanticRegen) [W2] StegaStamp Partial Good No-box
Tallam et al. [65] (SemanticRegen) [W3] Stable Signature Effective Good No-box
Tallam et al. [65] (SemanticRegen) [W4] Tree-Ring Effective Good No-box
Bao et al. [66] (SHIFT) [W4] SEAL, SFW, WIND, PRC, RingID, ROBIN, Tree-Ring, GaussMarker, Gaussian Shading Effective Relative semantic/distributional quality reported globally No-box

C.5.1. Method-Specific Quantitative Evidence and Defense Considerations

Instruction-driven editing. Under UltraEdit Global in W-Bench [68], TPR@0.1%FPR falls to 0.06%, 1.16%, 4.02%, 4.14%, 7.50%, 10.58%, and 17.00% for DwtDct [W1], EditGuard [W2], DwtDctSvd [W1], RivaGAN [W2], MBRS [W2], SSL Watermarking [W2], and CIN [W2], respectively. PIMoG, TrustMark, StegaStamp, and SepMark [W2] retain 40.14%, 43.48%, 51.24%, and 51.84%, producing the Mixed W2 judgment in Table 5. W-Bench evaluates semantic preservation with task-specific CLIP measures, so these results do not provide a single attacked-image fidelity grade under Appendix B. VINE-Robust retains 86.86% TPR in the same UltraEdit Global setting and approaches 100% TPR in several other image-editing settings. Robust-Wide [67] trains through Partial Instruction-driven Denoising Sampling Guidance (PIDSG), a differentiable approximation of semantic editing; removing PIDSG raises edited-image BER from 2.6579% to 50.1558%. VINE adapts SDXL-Turbo as a conditional generative watermark encoder and fine-tunes VINE-Robust against InstructPix2Pix through a straight-through estimator. Defenses should combine training against both global and local editors with spatial redundancy and explicit content binding, because redundancy alone does not prevent the editor from rewriting every watermark-bearing region.
SemanticRegen. Tallam et al. [65] report a Tree-Ring [W4] mean p-value of 0.10, above the paper-specific 0.05 significance threshold; the detector therefore no longer rejects the no-watermark hypothesis. Masked SSIM is 0.95 and masked PSNR is 31.71 dB. Stable Signature [W3] and DwtDct [W1] reach bit accuracies of 0.49 and 0.51. StegaStamp [W2] reaches 0.70, which falls below SemanticRegen’s own 0.75 criterion but remains above the FPR-calibrated StegaStamp reference of 0.614 in Table B2; this survey therefore grades the result as Partial because the payload is weakened without detector evasion. Masked SSIM is approximately 0.94 for StegaStamp, Stable Signature, and DwtDct, but this metric emphasizes retained regions and should not be interpreted as full-image preservation. Defenses should place authenticated evidence in salient and background regions, test selective inpainting explicitly, and detect inconsistencies introduced by captioning, segmentation, or diffusion-based background replacement.
SHIFT. Bao et al. [66] report attack success rates of 98% on SEAL, 100% on SFW, 98% on WIND, 97% on PRC, 95% on RingID, 98% on ROBIN, 98% on Tree-Ring, 99% on GaussMarker, and 97% on Gaussian Shading [W4], for an average of 97.8%. The reported global CLIP-Score is 32.19 and FID is 73.47, compared with FID 106.831 for the black-box baseline and 96.431 for the removal baseline. These CLIP and FID values support only relative semantic and distributional comparison under the source protocol, not an absolute image-quality grade. Defenses should test stochastic reverse samplers rather than only deterministic reconstruction, verify trajectory consistency across independent inversions, and combine forensics for diffusion-mediated edits with watermark evidence that is not tied to a single recoverable noise path.

C.6. Latent-Space Inversion Attacks

Table C6 summarizes the five attack papers in this section by attack method and target watermark lineage. CoSDA [75] and ZoDiac [76] are defense studies and are not listed as attack rows. Single-Image forgery and removal are separated because their effectiveness and fidelity differ substantially. The complete metric values, source-specific judgments, and method-specific defense considerations follow the table.
Table C6. Summary of Latent-Space Inversion Attacks
Table C6. Summary of Latent-Space Inversion Attacks
Attack Method (Author [Ref.]) Target Watermark (Lineage + Algorithm) Attack Effectiveness Image Fidelity Threat Model
Jain et al. [70] (Single-Image, forgery) [W4] RingID, WIND, Gaussian Shading, Tree-Ring Effective Moderate No-box / Gray-box
Jain et al. [70] (Single-Image, removal) [W4] Tree-Ring, Gaussian Shading, RingID/WIND Mixed (Tree-Ring: success; Gaussian Shading: partial; RingID/WIND: failure) Poor to good No-box / Gray-box
Lin and Juarez [71] (CrackBark) [W4] Tree-Ring Effective Good Gray-box
Qiu et al. [72] (NFPA) [W1] DwtDct Effective Relative semantic/distributional quality only No-box
Qiu et al. [72] (NFPA) [W2] StegaStamp, SSL Watermarking, RivaGAN Effective (residual TPR highest for RivaGAN) Relative semantic/distributional quality only No-box
Qiu et al. [72] (NFPA) [W3] Stable Signature Effective Relative semantic/distributional quality only No-box
Qiu et al. [72] (NFPA) [W4] Gaussian Shading, RingID, Tree-Ring Effective Relative semantic/distributional quality only No-box
Lee et al. [73] (Boundary Leakage) [W4] PRC Effective Moderate Gray-box
Hu et al. [74] (StableSig-Unstable) [W3] Stable Signature Effective Moderate White-box

C.6.1. Method-Specific Quantitative Evidence and Defense Considerations

Single-Image inversion. Under SDv2.0 with λ = 2 × 10 4 , Jain et al. [70] report forgery success of 100% for RingID, WIND, and Gaussian Shading and 90.72% for Tree-Ring. Representative forged-image fidelity is approximately 30–31 dB PSNR and 0.80–0.82 SSIM. Removal is less uniform: Tree-Ring reaches 97.80% success but falls to 19.71 dB PSNR and 0.77 SSIM, whereas Gaussian Shading reaches 34.73% at the representative setting and 59.13% at λ = 1 × 10 4 . RingID and WIND removal is ineffective in the reported experiments. Defenses should bind the latent watermark to image content or an authenticated message and use per-image keys so that a region inferred from one reference cannot be reused as a generic ownership target.
CrackBark. For Tree-Ring on Stable Diffusion v2.1, the VAE-recovered surrogate trained with watermarked and non-watermarked outputs reduces ROC-AUC from 0.993 to 0.108 and TPR@1%FPR from 0.968 to 0.023; using public comparison images yields ROC-AUC 0.153 and TPR@1%FPR 0.039 [71]. Because both AUROC values fall below 0.5, they indicate reversed score ranking under the fixed detector and should not be interpreted as greater information destruction than an AUROC of 0.5. The corresponding LPIPS values are 0.021 and 0.012, while the FID values are 2.684 and 0.694 under the source protocol. A 16-channel VAE or a VAE-free Guided Diffusion pipeline weakens transfer, with the latter retaining ROC-AUC near 0.787. Defenses should reduce reuse of a fixed public latent representation, diversify the detector across independent features, and evaluate surrogate transfer under both matched and mismatched VAEs.
NFPA. Under NFPA-xy, TPR@1%FPR is 0.00 for Stable Signature and Gaussian Shading, 0.01 for DwtDct, 0.02 for StegaStamp and RingID, 0.07 for Tree-Ring, 0.09 for SSL Watermarking, and 0.13 for RivaGAN, producing an average of 0.04 across eight targets [72]. Average FID is 69.48 compared with 66.57 without attack, and the reported CLIP-Score remains approximately 0.32; both are protocol-level semantic and distributional evidence rather than per-image absolute grades. Defenses should test latent warping and viewpoint shifts explicitly, align inversion drift across geometric changes, and supplement watermark verification with semantic-edit forensics.
Boundary Leakage. At ℓ 2 budgets of 4, 6, and 8, the stealthy sign-flip attack changes 0.078, 0.102, and 0.123 of PRC bits, compared with 0.009, 0.015, and 0.019 under white noise [73]. Flipping about 5% of the bits requires distortion near 3.0 for the stealthy attack and 45.0 for white noise. Under the practical null-prompt inversion setting at budget 12, the stealthy attack reaches ASR 0.43 with SSIM 0.83, whereas white noise reaches ASR 0.14 with SSIM 0.73. The paper’s secret orthonormal latent transformation hides the boundary direction and provably reduces the attacker’s advantage to the white-noise level under its stated assumptions. Implementations should also assess transformation cost and resistance to richer query oracles.
StableSig-Unstable. Across ImageNet, MS-COCO, and Conceptual Captions, the E-aware and E-agnostic routes obtain evasion rates of approximately 0.94–1.00 and bit accuracies below 0.66 [74]. In the reported utility table, E-aware fine-tuning takes 14.197 minutes and yields FID 18.15, PSNR 29.50 dB, and SSIM 0.86; E-agnostic fine-tuning takes 8777.885 minutes and yields FID 25.68, PSNR 29.40 dB, and SSIM 0.86. FID is interpreted relative to the source baseline. Defenses should authenticate decoder weights, detect unauthorized fine-tuning, restrict access to watermark-bearing components, and avoid treating open decoder parameters as a durable ownership root.
Defense-side evidence. CoSDA [75] combines compensation sampling with latent drift alignment. It raises the reported distorted-image TPR of Tree-Ring from 0.642/0.639 to 1.000/1.000 and that of Gaussian Shading from 0.960/0.967 to 1.000/1.000, while keeping FID and CLIP close to the watermark-free Stable Diffusion baseline. ZoDiac [76] instead optimizes and blends a latent Fourier watermark into an existing image; on the reported MS-COCO regeneration test, its detection rate changes from 0.998 to 0.988 while several post-hoc baselines collapse. These results support drift compensation and content-conditioned latent optimization as defenses, but neither removes the need to test proxy VAEs, boundary leakage, and model-parameter tampering.

C.7. Forgery and Provenance Attacks

Table C7 summarizes the 11 in-scope attack studies in this section by attack method and target watermark lineage. WIND and MetaSeal provide defense-side evidence and are not listed as attack rows. Model IP watermark attacks are excluded and discussed separately in Appendix E. The complete metric values, source-specific judgments, and method-specific defense considerations follow the table.
Table C7. Summary of Forgery and Provenance Attacks
Table C7. Summary of Forgery and Provenance Attacks
Attack Method (Author [Ref.]) Target Watermark (Lineage + Algorithm) Attack Effectiveness Image Fidelity Threat Model
Craver et al. [86] (Invertibility) [W1] Cox SS, Pitas statistical watermark Effective forgery Not reported No-box
Kutter et al. [85] (Copy Attack) [W1] Spatial spread-spectrum Software A/B Effective forgery Not reported No-box
Müller et al. [87] (Black-Box Forgery) [W4] Gaussian Shading, Tree-Ring Mixed across victim architectures Fair to moderate No-box
Zhu et al. [89] (PnP / OptFree-Forgery) [W4] Gaussian Shading, Tree-Ring Effective forgery Good to moderate No-box
Dong et al. [88] (WMCopier) [W1] DwtDct Effective forgery Good No-box
Dong et al. [88] (WMCopier) [W2] HiDDeN, RivaGAN Effective forgery Good No-box
Dong et al. [88] (WMCopier) [W3] Stable Signature Effective forgery Good No-box
Souček et al. [93] (One-Shot Forging) [W2] CIN, MBRS, TrustMark Mixed (CIN, MBRS: effective; TrustMark: limited at strict FPR) Good No-box
Ba et al. [90] (Robust-Leak) [W1] DwtDct, DwtDctSvd Effective evasion Good No-box
Ba et al. [90] (Robust-Leak) [W2] StegaStamp, RivaGAN, PIMoG, HiDDeN, CIN Effective evasion; mixed forgery (RivaGAN: weak) Good No-box
Wang et al. [91] (WatermarkFaker) [W1] LSB, LSB-M, LSB-MR, 8×8 DCT Mixed Good No-box
Yuan et al. [92] (Ambiguity Attack) [W3] SDM trigger watermarks, QR/Lena/Girl Effective forgery Not graded White-box
Luo et al. [94] (Counterfeit Extractor) [W3] Stable Signature, AquaLoRA Effective forgery Image unchanged Black-box
Luo et al. [94] (Counterfeit Extractor) [W4] DiffuseTrace Effective forgery Image unchanged Black-box
Zou et al. [98] (Neural Plagiarism, removal) [W1] DwtDctSvd Effective removal Fair to moderate No-box
Zou et al. [98] (Neural Plagiarism, removal) [W2] RivaGAN Effective removal Fair to moderate No-box
Zou et al. [98] (Neural Plagiarism, removal) [W3] Stable Signature Effective removal Fair to moderate No-box
Zou et al. [98] (Neural Plagiarism, removal) [W4] Tree-Ring Weak Not applicable No-box

C.7.1. Method-Specific Quantitative Evidence and Defense Considerations

Classical invertibility and copy attacks, and defense-side evidence. Craver et al. [86] demonstrate that counterfeit originals can yield ownership claims comparable with genuine claims for Cox spread-spectrum and Pitas statistical watermarks [W1], but the evidence is primarily conceptual and does not include image-fidelity measurements. Kutter et al. [85] make copied watermark evidence detectable in spatial spread-spectrum Software A and Software B [W1], but likewise report no PSNR or SSIM. Defenses should use non-invertible, content-dependent embedding and cryptographic ownership binding. On the defense side, WIND [96] reports average transformation-attack detection accuracies of 0.404 for Tree-Ring, 0.926 for RingID with 32 keys, and 0.887 for WINDfast with 128-bit identifiers. These protocol-specific values show that attribution robustness varies sharply even within W4. MetaSeal [97] instead authenticates a content-bound signature through cryptographic verification, illustrating why a recovered but unauthenticated message should not by itself establish ownership.
Surrogate-model forgery. Müller et al. [87] report detection and attribution success of 1.00 for Gaussian Shading [W4] in the representative Imprint-Forgery setting, with PSNR about 22.13 dB and LPIPS 0.160. The fixed metrics support fidelity ranging from Fair to Moderate. Tree-Ring detection success reaches 1.00 on SD2.1 but falls to 0.12–0.23 on FLUX, so the attack result is mixed across victim architectures rather than uniformly effective. In its strongest settings, PnP [89] reaches decoding and attribution success of 1.00 for Gaussian Shading, with bit accuracy of 0.79–0.94, and also forges Tree-Ring. Its evaluation includes PSNR, SSIM, LPIPS, DISTS, and no-reference quality metrics, which support a range from Good to Moderate rather than a universal fidelity claim. Defenses should bind latent evidence to authenticated content or identity and test forgery across unrelated proxy, restoration, and target architectures.
Low-access post-hoc forgery. WMCopier [88] reports forged bit accuracy and acceptance of 99.34% and 95.90% for HiDDeN [W2], 98.04% and 94.60% for Stable Signature [W3], 95.74% and 90.90% for RivaGAN [W2], and 89.19% and 60.20% for DwtDct [W1]. The acceptance threshold is calibrated to a clean-image FPR of 10 − 6 , and forged-image PSNR lies between 31 and 34 dB. With one reference image, One-Shot Forging [93] reaches forged bit accuracy of 1.00 for CIN, 0.83 for MBRS, and 0.61 for TrustMark [W2], with PSNR 31.3 dB in the reported setting; the TrustMark ROC shows limited acceptance at strict low-FPR operating points, so the cross-target result is mixed. Robust-Leak [90] obtains evasion rates of 0.96/0.95 for DwtDct/DwtDctSvd [W1], 1.00 for StegaStamp and RivaGAN, 0.87 for PIMoG, 0.78 for CIN, and 0.77 for HiDDeN [W2]. Its forgery success is 1.00 on PIMoG, StegaStamp, and CIN, 0.91 on HiDDeN, and 0.18 on RivaGAN, yielding effective evasion but target-dependent forgery. Fidelity remains Good overall, with SSIM 0.81–0.90 and PSNR 33–37 dB. Defenses should reduce reusable residual structure, use per-image keys and stronger content dependence, and train forgery detectors on transferred as well as directly embedded evidence.
Learned embedding maps and counterfeit extractors. For WatermarkFaker [91], image-level fidelity is distinct from similarity between the extracted forged watermark and the target watermark. The LSB and LSB-M [W1] forged images have PSNR/SSIM of 30.963/0.941 and 44.555/0.987, while the corresponding watermark-level values are 36.805/0.986 and 42.112/0.999. For the 8 × 8 DCT watermark, the image-level PSNR and SSIM are 32.668 and 0.920, but the watermark-level values fall to 9.772 and 0.545, which supports the Mixed effectiveness judgment despite Good image fidelity. Counterfeit Extractor [94] reaches above 99% target-message accuracy on victim outputs for Stable Signature and AquaLoRA [W3] and DiffuseTrace [W4], while remaining near random, approximately 49%, on clean-model outputs. Because it changes the verifier rather than the image, image fidelity is unchanged. Defenses should authenticate both the verifier and recovered signature, bind them to the genuine key and claimant, and monitor or limit query patterns that support counterfeit-extractor training.
Trigger ambiguity. For QR, Lena, and Girl trigger watermarks [W3], the forged prompts of Yuan et al. [92] all exceed the paper-specific cosine-similarity threshold of δ = 0.90 . The paper also reports PSNR, SSIM, and CLIP image cosine similarity, and the forged QR output remains scannable, but these task-specific values do not provide a general image-fidelity grade. Defenses should authenticate trigger ownership, test uniqueness against gradient-based prompt search, and avoid verification rules under which multiple prompts can support the same ownership claim.
Attention-space removal and plagiarism. Neural Plagiarism [98] reduces DwtDctSvd [W1] bit accuracy to 0.52, detection accuracy to 0.01, and TPR@1%FPR to 0, with PSNR 25.27 dB, SSIM 0.73, and FID 41.69. For RivaGAN [W2], bit accuracy is 0.56 and TPR@1%FPR is 0, with PSNR 25.29 dB and SSIM 0.73. Stable Signature [W3] also reaches TPR@1%FPR of 0, with PSNR 24.08 dB and SSIM 0.80, whereas Tree-Ring [W4] remains detectable at 0.98 in the removal setting and is affected mainly through coexisting-watermark ambiguity. The W1–W3 fixed metrics therefore support fidelity ranging from Fair to Moderate; FID is interpreted only relative to the source protocol. Defenses should couple watermark evidence to content across attention and latent representations and combine watermark verification with plagiarism, tamper, and competing-claim forensics.

12. Cross-Category Attack–Watermark Vulnerability Analysis

The preceding seven subsections discuss the attack mechanisms category by category. To help defenders examine which attacks affect each watermark class and to what extent, this section consolidates representative papers from the seven attack categories into a summary table by watermark class, shown in Table D1. Each row corresponds to an attack paper, and each column corresponds to one watermark class from W1 to W5 as defined in Section 2. Each cell reports the attack effect and image fidelity grade for the corresponding watermark class, while uncovered cases are marked as not covered. The cells follow the format attack effect / image fidelity. The first term uses attack-outcome labels including Effective, Weak, Weak to effective, Partial, Effective to good, Mixed, Failed, Not established, Forgery, Removal, and Evasion. Here Partial denotes partial weakening that does not reach a full effective-attack judgment under the Section 4 criteria, whereas Mixed denotes different outcomes across targets, victim architectures, or access settings. The second term uses image-fidelity labels from Section 4, including Poor, Fair, Moderate, and Good. Entries marked relative, protocol-specific, or semantic-only report evidence that does not support a directly comparable absolute image-fidelity grade. Methods with missing quality reports, qualitative-only evidence, or only conceptual arguments are marked as not reported, qualitative, or conceptual. Cross-modal targets, model IP targets, and defense-only studies are excluded according to the scope defined earlier. Some methods have multiple objectives. Averaging and Robust-Leak directly demonstrate forgery in addition to removal or evasion. EK Estimation demonstrates equivalent-key recovery that could in principle enable subsequent removal or forgery, but does not evaluate the fidelity of those outputs; Two-Stage SS evaluates removal rather than forgery. Single-Image Removal / Single-Image Forgery supports both removal and forgery, although removal remains difficult for distribution-preserving watermarks. Table D1 reports representative results, and the detailed values are given in Tables C1–C7.
Table D1 reveals three main patterns. First, classical frequency-/spatial-/statistical-domain watermarks (W1) and deep multi-bit / steganographic watermarks (W2) are covered by almost all attack mechanisms and therefore have the broadest attack surface. Second, in-generation, fine-tuned, or trigger-based generative watermarks (W3) mainly expose vulnerabilities under regeneration, transferable adversarial attacks, latent-space inversion, and forgery attacks. Third, diffusion latent-space, zero-bit, or semantic watermarks (W4) can resist many signal-level attacks, but are primarily affected by latent-space inversion, semantic editing, and micro-geometric phase perturbation. They also show an asymmetry in which forgery can be easier than removal. Geometry-synchronized, physically robust, or cryptographic-attribution watermarks (W5) do not yet show direct content-level compromises under acceptable-fidelity, in-scope attack evaluations and mainly appear as defenses or comparison baselines. The SynTag purification stress test is not entered as an attack row because it is reported within a defense study and reaches TPR 0 only when PSNR falls to 15.44 dB, outside the acceptable-fidelity condition used for this conclusion.
Table D1. Attacks × Watermark Categories Matrix
Table D1. Attacks × Watermark Categories Matrix
Attack Method (Author [Ref.], Short Name) Category W1 W2 W3 W4 W5
Voloshynovskiy et al. [5] (Optimal Estimation) Conventional signal-processing Effective / Good — — — —
Kassis and Hengartner [34] (UnMarker) Conventional signal-processing — Effective / Poor to good across metrics Effective / Poor to good across metrics Borderline / Poor to good across metrics —
Cao et al. [35] (MarkSweep) Conventional signal-processing — Effective / Moderate Mixed / Moderate — —
Wu et al. [36] (Box-Blur) Conventional signal-processing — Effective / relative-only — — —
Yang et al. [37] (Averaging) Conventional signal-processing Mixed / Good to moderate Mixed / Moderate Failed Effective / Poor —
You and Zhou [38] (Two-Stage SS) Conventional signal-processing Weak to effective / Good — — — —
You et al. [39] (EK Estimation) Conventional signal-processing Effective key recovery / Not reported for removal or forgery — — — —
Petitcolas et al. [33] (StirMark) Geometric Effective / Not graded — — — —
Barni et al. [77] (LPCD/MF-DA) Geometric Effective / Not graded — — — —
Kong et al. [78] (MarkCleaner) Geometric Effective / Poor raw, moderate aligned Effective / Poor raw, moderate aligned Effective / Poor raw, moderate aligned Effective / Poor raw, moderate aligned —
Shamshad et al. [79] (RAVEN) Geometric Effective / Relative Mixed / Relative Effective / Relative Effective / Relative —
Zhao et al. [84] (SSyncOA crop-paste) Geometric — Mixed / Not reported — — —
Francati et al. [83] (PRC Crop-and-Resize / Coding Bound) Geometric — — — Effective / Not graded —
Zhao et al. [42] (WatermarkAttacker) Regeneration Effective / Good Mixed / Good to poor — Failed / Not applicable —
Saberi et al. [43] (Diffusion Purification) Regeneration Effective / Fair to moderate Mixed / Fair to moderate — Mixed / Fair to moderate —
Liu et al. [44] (CtrlRegen) Regeneration Effective / Poor Effective / Poor Effective / Poor Effective / Poor —
Alam et al. [45] (SADRE) Regeneration Not established / Good Not established / Good — Not established / Good —
Liang et al. [46] (DIP) Regeneration Effective / Good Mixed / Good to poor — Failed / Not applicable —
Zhang et al. [47] (WiTS) Regeneration Effective / protocol-specific relative quality — Effective / protocol-specific relative quality — —
Bulychev et al. [48] (Re-Watermarking) Regeneration — Mixed / protocol-specific Effective / protocol-specific Mixed / protocol-specific —
Wang et al. [40] (HAI-WAN) Regeneration Weak to effective / Good — — — —
Yan et al. [49] (UQP-WR) Regeneration Effective / Good Mixed / Good to moderate — — —
Wang et al. [52] (BDMWA) Regeneration Partial to effective / Good Mixed / Good — Mixed / Good —
Meshram and Chandrasekaran [41] (D2RA) Regeneration Effective / Poor to moderate Effective / Poor to moderate — Mixed / Poor —
Wang et al. [54] (HIWANet) Regeneration Weak / Good — — — —
Wang et al. [55] (RD-IWAN) Regeneration Weak / Good — — — —
Li and Wang [51] (Concealed) Regeneration Weak / Good to moderate — — — —
Shamshad et al. [53] (First-Place) Regeneration — Effective / protocol-specific — Effective / protocol-specific —
Quiring and Rieck [31] (AdvML) Adversarial Effective / Good — — — —
Jiang et al. [57] (WEvade) Adversarial Effective / Relative perturbation evidence only Effective / Relative perturbation evidence only — — —
Lukas et al. [58] (Adaptive Attack) Adversarial Effective / Good Effective / Good — Effective / Good —
Hu et al. [59] (Transfer Attack) Adversarial Effective / Good Effective / Good Effective / Good — —
Chen et al. [61] (FAADW) Adversarial Effective payload disruption / Moderate Effective to good payload disruption / Moderate — — —
Lukas and Kerschbaum [60] (Reverse Pivotal Tuning against PTW) Adversarial — — Mixed / Relative distributional change only — —
Instruction-driven editing, evaluated in W-Bench [68] Semantic Effective / Semantic preservation only Mixed / Semantic preservation only — — —
Tallam et al. [65] (SemanticRegen) Semantic Effective / Good Partial / Good Effective / Good Effective / Good —
Bao et al. [66] (SHIFT) Semantic — — — Effective / Relative semantic/distributional quality only —
Jain et al. [70] (Single-Image) Latent-space — — — Forgery / Moderate; Removal / Poor to good —
Lin and Juarez [71] (CrackBark) Latent-space — — — Effective / Good —
Qiu et al. [72] (NFPA) Latent-space Effective / Relative semantic/distributional quality only Effective / Relative semantic/distributional quality only Effective / Relative semantic/distributional quality only Effective / Relative semantic/distributional quality only —
Lee et al. [73] (Boundary Leakage) Latent-space — — — Effective / Moderate —
Hu et al. [74] (StableSig-Unstable) Latent-space — — Effective / Moderate — —
Craver et al. [86] (Invertibility) Forgery and provenance Forgery / Not reported — — — —
Kutter et al. [85] (Copy Attack) Forgery and provenance Forgery / Not reported — — — —
Müller et al. [87] (Black-Box Forgery) Forgery and provenance — — — Mixed forgery / Fair to moderate —
Zhu et al. [89] (PnP) Forgery and provenance — — — Forgery / Good to moderate —
Dong et al. [88] (WMCopier) Forgery and provenance Forgery / Good Forgery / Good Forgery / Good — —
Souček et al. [93] (One-Shot Forging) Forgery and provenance — Mixed forgery / Good — — —
Ba et al. [90] (Robust-Leak) Forgery and provenance Effective evasion / Good Effective evasion, mixed forgery / Good — — —
Wang et al. [91] (WatermarkFaker) Forgery and provenance Mixed / Good — — — —
Yuan et al. [92] (Ambiguity Attack) Forgery and provenance — — Forgery / Not graded — —
Luo et al. [94] (Counterfeit Extractor) Forgery and provenance — — Forgery / Image unchanged Forgery / Image unchanged —
Zou et al. [98] (Neural Plagiarism) Forgery and provenance Effective removal / Fair to moderate Effective removal / Fair to moderate Effective removal / Fair to moderate Weak removal / Not applicable —

E. Unified Attack Taxonomy and Operational Conditions

Table E1 summarizes the attack entry points, evaluated target watermark classes, threat models, and orthogonal data and tool conditions of the seven attack families corresponding to §6.1–§6.7. The threat model definitions follow §3 and use the strongest applicable victim-system profile. No-box means offline attacks without victim detector access or victim-system queries; black-box means query access to a victim detector or API; gray-box means partial victim-system knowledge and may coexist with victim queries; and white-box means full internal access to the relevant victim parameters, gradients, keys, or software components. Independent public models used only as attack tools are recorded under data and tool conditions and do not by themselves change the victim-access category.
Table E1. Mechanisms, Threat Models, and Operational Conditions of Seven Attack Families
Table E1. Mechanisms, Threat Models, and Operational Conditions of Seven Attack Families
Attack Family Attack Entry Point Evaluated Target Watermark Classes Threat Model Data and Tool Conditions
6.1 Conventional Signal-Processing Statistical residual estimation, spectral erasure, frequency-aware denoising, cross-image pattern averaging, and carrier or equivalent-key estimation weaken, recover, or reuse signal-level watermark evidence. W1–W4, mainly W1 and W2 No-box / Gray-box / White-box Single image for most signal attacks; clean and AI-generated images for offline MarkSweep training; a pretrained deblurrer for image-only Box-Blur; an independent surrogate decoder and fine-tuning data for generator overwriting; multi-image same-key batches for averaging; KMA additionally requires paired watermarked signals and their corresponding known messages
6.2 Geometric and Synchronization Attacks RST, affine, perspective, local bending, latent-space phase perturbations, object relocation through crop-paste operations, and crop-and-resize codeword corruption disrupt geometric alignment or pre-decoding symbols. W1–W4 No-box One watermarked image for most attacks; geometrically perturbed MS-COCO training data for MarkCleaner; a public frozen diffusion model for RAVEN; SSyncOA additionally requires a destination cover image
6.3 Regeneration and Purification Noise, rewriting, overwriting, or learned restoration disrupt watermark evidence before or during reconstruction with natural-image priors. W1–W4 No-box / Black-box / Gray-box Single watermarked image for direct attacks; independent public VAE or diffusion tools for regeneration; public or attacker-controlled embedders plus labeled scheme-classification data for automatic re-watermarking; paired, unpaired, target-specific, or large-scale training data for learned variants
6.4 Adversarial Perturbation Gradient-, query-, or transfer-based optimization targets detectors, decoders, surrogate keys, or generator parameters. W1–W4 image-watermark targets; model IP watermarks are outside scope and discussed in Appendix E No-box / Black-box / Gray-box / White-box Single watermarked images for per-image optimization; same-key samples for substitute training; independent surrogate data or paired training data for learned transfer; a public surrogate generator for adaptive optimization; 50–200 real images for reverse pivotal tuning
6.5 Semantic Editing Instruction editing, semantically masked inpainting, and stochastic trajectory deflection rewrite pixel or latent correlations while preserving principal image content. W1–W4 No-box Single watermarked image plus a public instruction editor, semantic segmentation and inpainting pipeline, or diffusion model; no attack training is required in the reviewed attacks
6.6 Latent-Space Inversion Latent inversion enables optimization, warping, surrogate detection, and boundary manipulation; parameter attacks fine-tune watermark-bearing weights. W1–W4, chiefly W3 and W4 No-box / Gray-box / White-box Single-Image: reference and cover or removal images plus an independent proxy VAE for no-box variants; CrackBark: target outputs and public or non-watermarked samples; NFPA: one image and a public diffusion model; Boundary Leakage: same-key samples; StableSig-Unstable: non-watermarked images
6.7 Forgery and Provenance Attacks Invertibility constructions, watermark copying or transfer, learned embedding imitation, counterfeit verification, trigger ambiguity, and attention-space removal compromise attribution or create provenance ambiguity. W1–W4 No-box / Black-box / White-box One or two reference images, same-scheme collections, paired data, prompt initialization, an independent attack diffusion model, or victim and clean SDM output collections for Counterfeit Extractor
In this table, knowledge of a named target scheme or access to a victim-matched component is represented by Gray-box rather than by a data-condition label. Direct access to watermark-bearing decoder, denoiser, or generator parameters and gradients is represented by White-box. For §6.7, queries to victim and clean SDMs account for Black-box access in Counterfeit Extractor, while backpropagation through victim text-encoder and UNet components accounts for White-box access in the Ambiguity Attack. Gradients used only within an independent attack model do not change the victim-access profile.
In this table, knowledge of a named target scheme or access to a victim-matched component is represented by Gray-box rather than by a data-condition label. Direct access to watermark-bearing decoder, denoiser, or generator parameters and gradients is represented by White-box. For §6.7, queries to victim and clean SDMs account for Black-box access in Counterfeit Extractor, while backpropagation through victim text-encoder and UNet components accounts for White-box access in the Ambiguity Attack. Gradients used only within an independent attack model do not change the victim-access profile.

E.1. Unified Threat-Model Summary

Table E2 summarizes the three core threat-model dimensions, the orthogonal data and tool profile, and the optional forensic-stealth requirement adopted in this survey.
Table E2. Threat Models for Attacks on Invisible Image Watermarks
Table E2. Threat Models for Attacks on Invisible Image Watermarks
Dimension Setting Meaning Representative Attacks / References
Attack objective Removal / evasion Disable or erase watermark evidence, or induce a false negative; reduce TPR for detection-based schemes and drive BA or BER toward the binary random-guessing level of 0.5 for multi-bit schemes [42,44,57,65,71,77]
Attack objective Forgery / spoofing Increase false positives or fabricate false attribution / dual ownership [85,87,88,89,92,94]
Knowledge / access No-box No victim detector, extractor, API, model parameters, secret key, or target-system query access [35,41,42,44,46,59,70,87]
Knowledge / access Black-box Query access to the victim detector/API, but no access to victim parameters [57]
Knowledge / access Gray-box Partial victim-system knowledge, possibly combined with victim queries, but no complete relevant internals or secret key [58,71]
Knowledge / access White-box Full internal access to the relevant victim components; key access only when assumed by the attack [57,60]
Orthogonal condition Data and tools Single-image, multi-image, paired, unpaired, known-message, or in-distribution data; independent public models; and offline training requirements [38,39,59,70,71]
Fidelity constraint Perceptual fidelity Constrains per-image distortion using PSNR, SSIM, LPIPS, or related measures; FID is used only for relative distributional comparison under a fixed protocol [42,50,99]
Additional requirement Forensic stealth Requires attacked outputs to avoid reliable forensic discrimination under the stated forensic protocol [100]

E.2. Scope Boundary: Model IP Watermarks

Box-free model watermarking protects the intellectual property of an image-to-image model by inserting ownership evidence into the model’s outputs, whereas this survey studies attacks on standalone watermarked images in the W1–W5 lineages. The distinction concerns the protected object, not an additional threat-model level. An et al. examine this separate setting through Decoder Gradient Shields (DGS) [62], black-box output-interface removal and overwriting [63], and query-based reverse engineering (QBRE) [64]. DGS is a defense that reorients or suppresses gradients exposed by a watermark decoder and reports a defense success rate of 1.00 under its evaluated attacks. The two attack studies instead use model or API queries to learn a remover, an overwriter, or a surrogate of the hidden embedding transformation.
High-Freq Overwriting [95] belongs to the same model IP boundary. It queries a released watermarked image-processing model, overwrites high-frequency output evidence with a deep steganographic network, and uses the resulting input-output pairs to train a surrogate model. This is black-box access to the protected model, with query outputs serving as a separate training-data condition. Because the protected object is the model and the attack objective is model extraction rather than compromise of a standalone W1–W5 image watermark, the method is excluded from Section 6.7, Table C7, and the cross-category matrix.
The source-specific results further show why these works should not be merged with attacks on post-hoc image watermarks. The EGG remover in Box-Free Blackbox [63] reports removal success 1.00 with PSNR up to 41.81 dB and MS-SSIM 0.9992, while QBRE [64] reports removal success 1.00 with PSNR 32.05–34.69 dB and MS-SSIM 0.988–0.992 across its principal settings. By comparison, WEvade-B-Q, which was designed for image-watermark detector evasion, reaches only 0.14–0.53 removal success when reused as a baseline against these model IP watermarks. These results depend on the model-output interface and have no direct analogue for a standalone watermarked image, so they are excluded from the W1–W5 attack tables and cross-category matrix.

F. Detailed Image-Watermark Evaluation Benchmark Profiles

Evaluation benchmarks usually do not propose a single attack method. Instead, they use standardized attack suites to assess the robustness of different watermark families and place otherwise scattered results under a unified, reproducible, and cross-comparable protocol. Their value lies in improving comparability across studies and in empirically corroborating the complementary vulnerability map developed in Section 6 and summarized in Appendices C and D. This appendix covers six benchmark works, organized into four groups: standardized robustness benchmarks (WAVES and WIBE); competition-style stress tests (Erasing the Invisible and its winning solution); forensic detectability as a new evaluation dimension (Forensic Stealth); and systematized knowledge and security definitions (SoK).

F.1. Standardized Robustness Benchmarks

Standardized robustness benchmarks evaluate multiple watermarking schemes with fixed attack suites and unified scoring protocols. An et al. [99] (WAVES, ICML 2024) establish a three-way taxonomy of distortions, regeneration, and adversarial attacks, covering 26 attack variants. WAVES jointly scores watermark robustness by the true positive rate at a fixed 0.1% false positive rate and image degradation by a WAVES quality score aggregated from eight metrics. It also introduces two adversarial attack families, adversarial embedding and surrogate-detector attacks. Yakushev et al. [103] (WIBE, ASE 2025) provide a broader benchmarking framework in terms of supported methods, evaluation metrics, cluster-parallel execution, and visualization. WIBE covers 17 watermarking methods, 23 attacks, and 10 metrics, and demonstrates the framework on StegaStamp, SSL Watermarking, TrustMark, and WAM under seven attacks: JPEG compression, rotation, Gaussian blur, Gaussian noise, center cropping, BM3D, and deep image prior. Together, WAVES and WIBE provide a common basis for comparison and calibrate detection at the low-false-alarm operating point of 0.1% FPR.

F.2. Competition-Style Stress Tests

Competition-style stress tests evaluate removal attacks on unknown and partially known watermarks through public challenge settings. Ding et al. [101] (Erasing the Invisible, NeurIPS 2024 competition) define two tracks: a black-box track in which the watermarking method is hidden and a beige-box track in which the watermarking methodology label is provided. Here, beige-box is the organizers’ official track name, which corresponds approximately to gray-box access in the four-level taxonomy adopted by this survey. The competition covers five invisible image watermark families, namely Stable Signature, Tree-Ring, StegaStamp, DWT-DCT, and RivaGAN, and uses a two-dimensional score based on detection rate and image fidelity. As a methodological proposal, this work defines the tracks and scoring protocol but does not report final leaderboard results. The winning solution is given by Shamshad et al. [53] (First-Place Solution, NeurIPS 2024 challenge). It combines adaptive VAE fine-tuning, test-time VAE optimization, and CIELAB color transfer for modified StegaStamp, uses spatial translation for a Tree-Ring variant, and divides the black-box images into four artifact clusters for cluster-specific removal. This solution reduces the detection scores to 0.043 and 0.037 on the black-box and beige-box tracks, respectively, which corresponds to more than 95.7% removal, with protocol-specific quality-degradation scores of 0.136 and 0.153. These results indicate that VAE-based regeneration and artifact-specific restoration pipelines can attack both known and unknown invisible image watermarks when adapted to the observed watermark family.

F.3. Forensic Detectability: A Third Evaluation Dimension

Forensic detectability extends evaluation beyond removal effectiveness and image fidelity. Goonatilake and Ateniese [100] (Forensic Stealth, arXiv 2026) show that even when a remover suppresses the watermark and preserves high image fidelity, its output can still be identified as attacked at a 1% false positive rate with a true positive rate above 99% for the evaluated removers. Using clean images as the reference, this benchmark evaluates six state-of-the-art removers, namely UnMarker, WatermarkAttacker, CtrlRegen+, NFPA, Boundary Leakage, and WiTS. The reason is that regeneration and optimization processes leave spectral and statistical residuals. The implication is that watermark removal is not equivalent to forensic stealth. This observation raises the bar for attackers and establishes forensic detectability as an evaluation axis that should be reported alongside removal success and fidelity.

F.4. Systematized Knowledge and Security Definitions

Systematized knowledge provides the conceptual basis for empirical benchmarks. Zhao et al. [102] (SoK, IEEE S&P 2025) provide definitions and security notions that are relevant to image watermarks for AI-generated content. They classify watermarking schemes, threat models, including removal and forgery, and security properties, including image fidelity, low false positive rate, robustness, unforgeability, message capacity, and efficiency. They also identify cryptographically undetectable pseudorandom-code watermarks as a promising but still incomplete direction. This work does not conduct attack evaluation itself. Rather, it provides threat definitions and a map of known vulnerabilities for the field.

F.5. Vulnerability Map and Limitations Revealed by Benchmarks

These benchmarks expose scheme-specific weaknesses under unified protocols and support the conclusions of the detailed comparisons in Appendices C and D. WAVES shows that regeneration drives the true positive rate of Stable Signature [W3] to 0.001, and that diffusion rinsing and surrogate-detector adversarial attacks reduce Tree-Ring [W4] to roughly 0.44–0.50. The surrogate-detector variants achieve this with low quality degradation, with a WAVES quality score around 0.14, whereas multi-round rinsing incurs higher degradation, with a WAVES quality score around 0.47. StegaStamp [W2] is the most robust among the main WAVES targets, retaining 0.9–1.0 after many attacks, although rotation, resized cropping, geometric combinations, and heavy blur reduce its average detection performance to 0.36–0.54. The blur setting gives 0.41 with a WAVES quality score of 1.198, indicating a large quality-degradation cost. DwtDct [W1] appears fragile in the WAVES appendix plots, and SSL Watermarking [W2] is fragile in the WIBE demonstration under several common distortions. WIBE also quantifies imperceptibility differences across schemes, with PSNR ranging from 29.13 dB for StegaStamp to 44.18 dB for TrustMark. Taken together with the summary tables, these results show that vulnerability is mechanism-dependent: regeneration strongly affects both low-perturbation post-processing watermarks and Stable Signature, whereas schemes that better resist signal-level attacks may remain exposed to latent-space inversion or semantic editing.
The benchmarks themselves also have limitations. First, benchmarks are not attack methods, and their coverage remains uneven. WAVES mainly studies three targets, Tree-Ring, Stable Signature, and StegaStamp, while DwtDct and MBRS appear only in the appendix; the competition proposal and SoK intentionally do not provide quantitative attack results. Second, metric heterogeneity hinders cross-paper comparison. Existing studies use bit accuracy, TPR@1%FPR, or AUROC. WAVES introduces a normalized aggregate quality-degradation score, which Erasing the Invisible adopts in its competition protocol, whereas WIBE reports standard fidelity and robustness metrics within its own benchmarking pipeline; these scores and metrics cannot be directly transferred across protocols. Third, benchmark suites lag behind newer attack vectors. Existing benchmarks focus on distortion, regeneration, and adversarial attacks, while semantic editing and latent-space inversion attacks that can compromise robust latent-space watermarks are not yet sufficiently covered.
Table F1. Overview of Image-Watermark Evaluation Benchmarks
Table F1. Overview of Image-Watermark Evaluation Benchmarks
Benchmark / Work (Author [Ref.]) Type Attack Families and Scale Covered Watermarks Main Metrics Key Findings
An et al. [99] (WAVES) Robustness benchmark Distortion, regeneration, and adversarial attacks; 26 variants Tree-Ring, Stable Signature, StegaStamp as primary targets TPR@0.1%FPR versus WAVES quality score Regeneration nearly suppresses Stable Signature detection; specific surrogate-detector and gray-box embedding attacks weaken Tree-Ring; StegaStamp is the most robust among the primary targets
Yakushev et al. [103] (WIBE) Benchmark framework 7 demonstration attacks; library contains 23 attacks StegaStamp, SSL, TrustMark, WAM in the demo; 17 watermarking methods in the library 10 metrics, including TPR@0.1%FPR Broader watermark-method support and execution features; PSNR ranges from 29.13–44.18 dB
Ding et al. [101] (Erasing the Invisible competition) Competition methodology Distortion, regeneration, and adversarial attacks; black-box and beige-box tracks Stable Signature, Tree-Ring, StegaStamp, DWT-DCT, RivaGAN Detection rate × image quality as a two-dimensional score Defines the protocol but does not itself provide a leaderboard
Shamshad et al. [53] (Winning Solution) Competition-winning method VAE fine-tuning + test-time optimization + CIELAB restoration Modified StegaStamp, Tree-Ring variant, and unknown black-box clusters Detection score, quality score, and total score Black-box 0.043 / 0.136 / 0.143; beige-box 0.037 / 0.153 / 0.157
Goonatilake and Ateniese [100] (Forensic Stealth) Forensic evaluation Forensic detection on outputs of six removers Removers: UnMarker, WatermarkAttacker, CtrlRegen+, NFPA, Boundary Leakage, WiTS TPR@1%FPR for detectable residual traces Removal is not the same as leaving no trace; the six remover-specific forensic detectors achieve TPRs of 99.24%–99.97% at 1% FPR
Zhao et al. [102] (SoK) Systematization of knowledge No attack evaluation Definitions and security notions relevant to image watermarks Threat models and security properties Provides threat definitions and a vulnerability map

G. Detailed Historical Evolution of Invisible Image Watermark Attacks

Watermark attacks aim either to make embedded watermark information unextractable or undetectable, or to fabricate watermark evidence and false provenance. Since watermark attacks were first studied, several taxonomies have been proposed. Existing classifications include robustness attacks, presentation attacks, interpretation attacks, and legal attacks [3]; simple attacks, synchronization attacks, ambiguity attacks, and removal attacks [4]; cryptographic attacks, geometric attacks, protocol attacks, and removal attacks [5]; unintentional and intentional attacks [6]; and unauthorized embedding, unauthorized detection, and unauthorized removal [7]. Building on these classifications, Cao et al. [8] propose a more systematic taxonomy from the perspectives of non-technical and technical attacks, and place the above attack types under these two categories. Among them, robustness attacks are the most widely studied.
To resist robustness attacks, defenders have proposed many strategies and produced a large body of work. For signal-processing attacks, researchers design watermarking schemes in the spatial domain [9,10], transform domain [11,12], and feature space [13,14]. For geometric attacks, they develop strategies based on geometric invariants [15,16], synchronization correction [17,18], and local feature regions. Other watermarking schemes are designed to resist gradient-descent attacks, sensitivity attacks, and perturbation attacks [19,20]. This early line of work treats attacks mainly as signal degradation, desynchronization, or unauthorized manipulation and accordingly strengthens watermark robustness at the embedding and detection stages.
Between 2018 and 2022, deep learning changed both watermark design and the corresponding attack surface. Before this period, Haribabu et al. [21] introduced a watermarking algorithm based on an autoencoder in 2015; this direction was followed by HiDDeN [22], the end-to-end differentiable framework of Ahmadi et al. [23], the generative adversarial design of Hao et al. [24], and the resolution-independent method of Lee et al. [25]. Post-hoc deep watermarks such as HiDDeN, StegaStamp [107], and RivaGAN [119] subsequently became representative attack targets; RivaGAN is also used as a baseline for images or individual video frames in attack studies. Although some of these systems retain low bit error rates or nearly lossless payload recovery after conventional distortions, such robustness does not establish security against adaptive attacks. Conventional distortion attacks may also damage visual quality, limiting their applicability in settings with stringent data accuracy requirements, such as military, medical, and remote sensing applications.
During the same period, attacks based on neural networks began to move beyond conventional signal processing. In 2020, Nam et al. [28] proposed the Watermarking Attack Network (WAN), based on residual dense blocks, to disable several mainstream watermarking schemes while preserving image quality. Sharma et al. [29] presented an attack based on an autoencoder alongside a robust hybrid watermarking technique, and Geng et al. [30] proposed a real-time convolutional removal attack that disrupts extraction without prior knowledge. Quiring and Rieck [31] used adversarial learning and a surrogate watermark detector to attack unknown watermarks under black-box access. In 2021, Hatoum et al. [32] applied a fully convolutional denoising network to spread-spectrum and other watermarks, framing the attack as quality-enhancing denoising. These studies mark the transition from fixed distortions toward learned removal and adaptive detector evasion.
From 2023 to 2024, diffusion models, latent-space generative modeling, and public image editors shifted the field toward model-level and generative attacks. Tree-Ring [104] and Stable Signature [105] embedded watermark evidence into diffusion latent spaces or generative models, while Gaussian Shading [106] further developed this generative watermarking direction. On the attack side, Zhao et al. [42] introduced a regeneration-based removal method, first released on arXiv in 2023 and later published at NeurIPS 2024, and WEvade [57] studied adversarial detector evasion. In 2024, WAVES [99] provided a unified benchmark across attack categories and evaluation metrics; Saberi et al. [43] analyzed fundamental limits and the trade-off between evasion and spoofing; WiTS [47] examined the security limitations of strong watermarks under quality and perturbation oracle or random walk assumptions; and no-box averaging attacks against Tree-Ring began to appear [37]. This period established regeneration, adversarial evasion, unified benchmarking, and theoretical removability as central research directions.
In 2025 and 2026, the attack landscape expanded toward universal spectral removal, controllable regeneration, latent-space inversion, semantic editing, and attribution forgery. UnMarker [34] introduced a general spectral attack against multiple post-processing watermarks and some diffusion latent-space or zero-bit watermarks under no-box, data-free, and query-free conditions, while MarkSweep [35] performed frequency-aware denoising and CtrlRegen [44] used controllable regeneration from clean noise. Single-image inversion [70], CrackBark [71], NFPA [72], and Boundary Leakage [73] respectively exploited latent reconstruction, public VAE priors, next-frame prediction, and PRC detection boundary leakage for removal or codeword flipping. Müller et al. [87] studied one-reference semantic watermark forgery and removal using a proxy diffusion model and DDIM inversion, while attacks targeting provenance [88], transfer attacks [59], and one-shot forging [93] further lowered the barrier to misattribution and forgery. D2RA/DAWN [41] and SHIFT [66] demonstrated continued pressure from generative and semantic editing attacks, and MetaSeal [97] introduced cryptographic attribution to address forgery and provenance risks. Overall, watermark attacks have evolved from signal degradation and desynchronization to model reconstruction, semantic manipulation, latent-space inversion, and attribution-level attacks, forming an attack landscape that conventional distortion-centered taxonomies and evaluation systems no longer fully cover.

References

  1. Z. Ma, W. Zhang, H. Fang, X. Dong, L. Geng, and N. Yu. 2021. Local Geometric Distortions Resilient Watermarking Scheme Based on Symmetry. IEEE TCSVT 31, 12 (2021), 4826-4839. [CrossRef]
  2. Z. Yue, Z. Li, Y. Yang, et al. 2020. Research on a histogram 2Bin multi-ary image digital watermarking algorithm. Acta Electronica Sinica 48, 3 (2020), 531-537.
  3. S. Craver, B.-L. Yeo, and M. Yeung. 1998. Technical trials and legal tribulations. Comm. ACM 41, 7 (1998), 45-54.
  4. F. Hartung, J. Su, and B. Girod. 1999. Spread spectrum watermarking: malicious attacks and counterattacks. In Proc. SPIE (1999), 147-158.
  5. S. Voloshynovskiy, S. Pereira, V. Iquise, and T. Pun. 2001. Attack modelling: towards a second generation watermarking benchmark. Signal Processing 81, 6 (2001), 1177-1214.
  6. B. Vassaux, et al. 2002. Survey on attacks in image and video watermarking. In Proc. SPIE (2002), 169-179.
  7. I. J. Cox, M. L. Miller, J. A. Bloom, J. Fridrich, and T. Kalker. 2008. Digital Watermarking and Steganography (2nd ed.). Morgan Kaufmann, Burlington, MA, USA.
  8. Y. Cao, K. Wang, D. Wang, and J. Wang. 2007. A new classification of digital image watermarking attack methods. Application Research of Computers 24, 4 (2007), 144-146.
  9. G. Hua, Y. Xiang, and L. Zhang. 2020. Informed Histogram-Based Watermarking. IEEE SPL 27 (2020), 236-240.
  10. T. Zong, et al. 2015. Robust histogram shape-based method for image watermarking. IEEE TCSVT 25, 5 (2015), 717-729.
  11. Y. Shen, et al. 2021. A DWT-SVD based adaptive color multi-watermarking scheme. ESWA 168 (2021), 114414.
  12. X. Liu, et al. 2017. Fractional Krawtchouk Transform With an Application to Image Watermarking. IEEE TSP 65, 7 (2017), 1894-1908.
  13. C. Wang, X. Wang, C. Zhang, et al. 2018. Stereo image zero-watermarking algorithm based on trinion polar harmonic-Fourier moments and chaotic mapping. Scientia Sinica Informationis 48, 1 (2018), 79-99.
  14. M. Yamni, et al. 2021. Image watermarking using separable fractional moments of Charlier-Meixner. J. Franklin Inst. 358, 4 (2021), 2535-2560.
  15. R. Hu and S. Xiang. 2021. Cover-Lossless Robust Image Watermarking Against Geometric Deformations. IEEE TIP 30 (2021), 318-331.
  16. B. Ma, et al. 2020. Robust image watermarking using invariant accurate polar harmonic Fourier moments and chaotic mapping. Signal Processing 172 (2020), 107544.
  17. C. Wang, et al. 2017. Geometric correction based color image watermarking. Signal Processing 134 (2017), 197-208.
  18. X. Wang, et al. 2020. Synchronization correction-based robust digital image watermarking. Pattern Anal. Appl. 23, 2 (2020), 933-951.
  19. H. R. Shahdoosti and M. Salehi. 2018. Transform-based watermarking algorithm maintaining perceptual transparency. IET Image Process. 12, 5 (2018), 751-759.
  20. X. Zhang and S. Wang. 2007. Watermarking Scheme Capable of Resisting Sensitivity Attack. IEEE SPL 14, 2 (2007), 125-128.
  21. K. Haribabu, et al. 2015. A robust digital image watermarking technique using auto encoder based CNNs. In IEEE WCI (2015).
  22. J. Zhu, R. Kaplan, J. Johnson, and L. Fei-Fei. 2018. HiDDeN: Hiding Data With Deep Networks. In ECCV (2018), 682-697.
  23. M. Ahmadi, et al. 2020. ReDMark: Framework for residual diffusion watermarking based on deep networks. ESWA 146 (2020), 113157.
  24. K. Hao, G. Feng, and X. Zhang. 2020. Robust image watermarking based on generative adversarial network. China Communications 17, 11 (2020), 131-140.
  25. J.-E. Lee, Y.-H. Seo, and D.-W. Kim. 2020. CNN-Based Digital Image Watermarking Adaptive to the Resolution of Image and Watermark. Applied Sciences 10, 19 (2020), 6854.
  26. V. Licks and R. Jordan. 2005. Geometric attacks on image watermarking systems. IEEE MultiMedia 12, 3 (2005), 68-78.
  27. M. Tanha, et al. 2012. An overview of attacks against digital watermarking and their respective countermeasures. In CyberSec (2012), 265-270.
  28. S.-H. Nam, et al. 2020. WAN: Watermarking Attack Network. arXiv preprint arXiv:2008.06255 (2020). https://arxiv.org/abs/2008.06255.
  29. S. S. Sharma and V. Chandrasekaran. 2020. A robust hybrid digital watermarking technique against a powerful CNN-based adversarial attack. MTA 79, 43 (2020), 32769-32790.
  30. L. Geng, et al. 2020. Real-time attacks on robust watermarking tools in the wild by CNN. J. Real-Time Image Process. 17, 3 (2020), 631-641.
  31. E. Quiring and K. Rieck. 2018. Adversarial Machine Learning Against Digital Watermarking. In EUSIPCO (2018), 519-523.
  32. M. W. Hatoum, et al. 2021. Using Deep learning for image watermarking attack. Signal Process.: Image Commun. 90 (2021), 116019.
  33. F. A. P. Petitcolas, R. J. Anderson, and M. G. Kuhn. 1998. Attacks on Copyright Marking Systems (StirMark). In Information Hiding (IH’98) (1998). LNCS 1525, 218-238.
  34. A. Kassis and U. Hengartner. 2025. UnMarker: A Universal Attack on Defensive Image Watermarking. In IEEE S&P (2025). https://arxiv.org/abs/2405.08363.
  35. J. Cao, et al. 2026. MarkSweep: A No-Box Removal Attack on AI-Generated Image Watermarking via Noise Intensification and Frequency-Aware Denoising. In ICASSP (2026). https://arxiv.org/abs/2602.15364.
  36. X. Wu, et al. 2025. When There Is No Decoder: Removing Watermarks from Stable Diffusion Models in a No-box Setting. In ICICS (2025). https://arxiv.org/abs/2507.03646.
  37. P. Yang, H. Ci, Y. Song, and M. Z. Shou. 2024. Steganalysis on Digital Watermarking: Is Your Defense Truly Impervious? (Can Simple Averaging Defeat Modern Watermarks?). In NeurIPS (2024). https://arxiv.org/abs/2406.09026.
  38. J. You and Y. Zhou. 2024. Two-Stage Watermark Removal Framework for Spread Spectrum Watermarking. IEEE TMM 26 (2024), 7687-7699.
  39. J. You, et al. 2023. Estimating the Secret Key of Spread Spectrum Watermarking Based on Equivalent Keys. IEEE TMM 25 (2023), 2459-2473.
  40. C. Wang, et al. 2025. Highly applicable and imperceptible watermark attack network (HAI-WAN). Signal Processing 230 (2025), 109840.
  41. P. S. Meshram and V. Chandrasekaran. 2026. Low-Compute Watermark Removal via Dual-Domain Natural Projection (D2RA/DAWN). In ICML (2026). https://arxiv.org/abs/2510.07538.
  42. X. Zhao, et al. 2024. Invisible Image Watermarks Are Provably Removable Using Generative AI (WatermarkAttacker). In NeurIPS (2024). https://arxiv.org/abs/2306.01953.
  43. M. Saberi, et al. 2024. Robustness of AI-Image Detectors: Fundamental Limits and Practical Attacks. In ICLR (2024). https://arxiv.org/abs/2310.00076.
  44. Y. Liu, et al. 2025. Image Watermarks are Removable Using Controllable Regeneration from Clean Noise (CtrlRegen). In ICLR (2025). https://arxiv.org/abs/2410.05470.
  45. I. Alam, et al. 2025. Saliency-Aware Diffusion Reconstruction for Effective Invisible Watermark Removal (SADRE). In WWW Companion (2025). [CrossRef]
  46. H. Liang, T. Li, and J. Sun. 2025. A Baseline Method for Removing Invisible Image Watermarks using Deep Image Prior (DIP). In TMLR (2025). https://arxiv.org/abs/2502.13998.
  47. H. Zhang, et al. 2024. Watermarks in the Sand: Impossibility of Strong Watermarking for Generative Models (WiTS). In ICML (2024). https://arxiv.org/abs/2311.04378.
  48. M. Bulychev, et al. 2026. Watermarks Attack Watermarks: Re-Watermarking as a Generic Removal Strategy. arXiv preprint arXiv:2605.16796 (2026). https://arxiv.org/abs/2605.16796.
  49. F. Yan, et al. 2025. Universal and Quality-Preserving Watermark Removal based on Unpaired Learning (UQP-WR). IEEE Transactions on Dependable and Secure Computing (2025). [CrossRef]
  50. T. Bui, S. Agarwal, and J. Collomosse. 2025. TrustMark: Robust Watermarking and Watermark Removal for Arbitrary Resolution Images. In ICCV (2025). https://arxiv.org/abs/2311.18297.
  51. Q. Li and X. Wang. 2021. Concealed Attack for Robust Watermarking Based on Generative Model and Perceptual Loss. IEEE Transactions on Circuits and Systems for Video Technology (2021). [CrossRef]
  52. C. Wang, et al. 2026. Enhanced Watermark Breaker: A Watermarking Attack Method Based on Bidirectional Diffusion Model (BDMWA). IEEE TCE 72, 2 (2026), 3100-3111.
  53. F. Shamshad, et al. 2025. First-Place Solution to NeurIPS 2024 Invisible Watermark Removal Challenge. arXiv preprint arXiv:2508.21072 (2025). https://arxiv.org/abs/2508.21072.
  54. C. Wang, et al. 2024. HIWANet: A high imperceptibility watermarking attack network. EAAI 133 (2024), 108039.
  55. C. Wang, et al. 2022. RD-IWAN: Residual Dense Based Imperceptible Watermark Attack Network. IEEE TCSVT 32, 11 (2022), 7460-7472.
  56. B. Guo, et al. 2025. Robust Reversible Watermarking With Invisible Distortion Against VAE Watermark Removal (RRWID). IEEE TIP 34 (2025), 6386-6401.
  57. Z. Jiang, J. Zhang, and N. Z. Gong. 2023. Evading Watermark based Detection of AI-Generated Content (WEvade). In ACM CCS (2023). https://arxiv.org/abs/2305.03807.
  58. N. Lukas, et al. 2024. Leveraging Optimization for Adaptive Attacks on Image Watermarks. In ICLR (2024). https://arxiv.org/abs/2309.16952.
  59. Y. Hu, Z. Jiang, M. Guo, and N. Z. Gong. 2025. A Transfer Attack to Image Watermarks. In ICLR (2025). https://arxiv.org/abs/2403.15365.
  60. N. Lukas and F. Kerschbaum. 2023. PTW: Pivotal Tuning Watermarking for Pre-Trained Image Generators. In USENIX Security (2023). https://arxiv.org/abs/2304.07361.
  61. M. Chen, D. Guo, and H. Yao. 2025. Feature-Attention-Mechanism-Based Attack for Deep Robust Watermarking (FAADW). IEEE MultiMedia 32, 1 (2025), 83-91.
  62. H. An, G. Hua, W. Du, H. Cao, Y. Tao, G. Xu, S. Rahardja, and Y. Fang. 2026. Decoder Gradient Shields: A Family of Provable and High-Fidelity Methods Against Gradient-Based Box-Free Watermark Removal (DGS). IEEE Transactions on Dependable and Secure Computing 23, 3 (2026), 5015-5028. [CrossRef]
  63. H. An, G. Hua, Z. Lin, and Y. Fang. 2026. Box-Free Model Watermarks Are Prone to Black-Box Removal Attacks. IEEE Transactions on Pattern Analysis and Machine Intelligence (2026). https://arxiv.org/abs/2405.09863.
  64. H. An, et al. 2026. Removing Box-Free Watermarks for Image-to-Image Models via Query-Based Reverse Engineering (QBRE). In AAAI (2026). https://arxiv.org/abs/2507.18034.
  65. K. Tallam, et al. 2025. Removing Watermarks with Partial Regeneration using Semantic Information (SemanticRegen). arXiv preprint arXiv:2505.08234 (2025). https://arxiv.org/abs/2505.08234.
  66. R. Bao, et al. 2026. SHIFT: Stochastic Hidden-Trajectory Deflection for Removing Diffusion-based Watermarks. arXiv preprint arXiv:2603.29742 (2026). https://arxiv.org/abs/2603.29742.
  67. R. Hu, et al. 2024. Robust-Wide: Robust Watermarking against Instruction-driven Image Editing. In ECCV (2024). https://arxiv.org/abs/2402.12688.
  68. S. Lu, et al. 2025. Robust Watermarking Using Generative Priors Against Image Editing (VINE / W-Bench). In ICLR (2025). https://arxiv.org/abs/2410.18775.
  69. M. Lyu, Y. Huang, and A. W.-K. Kong. 2023. Adversarial Attack for Robust Watermark Protection Against Inpainting-based and Blind Watermark Removers (AWD-AGP). In ACM MM (2023), 8396-8405.
  70. A. Jain, et al. 2025. Forging and Removing Latent-Noise Diffusion Watermarks Using a Single Image. arXiv preprint arXiv:2504.20111 (2025). https://arxiv.org/abs/2504.20111.
  71. J. Lin and M. Juarez. 2025. A Crack in the Bark: Leveraging Public Knowledge to Remove Tree-Ring Watermarks (CrackBark). In USENIX Security (2025). https://arxiv.org/abs/2506.10502.
  72. H. Qiu, et al. 2025. The Future Unmarked: Watermark Removal in AI-Generated Images via Next-Frame Prediction (NFPA). In NeurIPS (2025).
  73. D. Z. Lee, H. Fang, H. Wang, and E.-C. Chang. 2025. Removal Attack and Defense on AI-generated Content Latent-based Watermarking (Boundary Leakage). In ACM CCS (2025), 2174-2188. https://arxiv.org/abs/2509.11745.
  74. Y. Hu, et al. 2024. Stable Signature is Unstable: Removing Image Watermark from Diffusion Models. arXiv preprint arXiv:2405.07145 (2024). https://arxiv.org/abs/2405.07145.
  75. H. Fang, et al. 2025. CoSDA: Enhancing the Robustness of Inversion-Based Generative Image Watermarking Framework. In AAAI (2025), 2888-2896.
  76. L. Zhang, X. Liu, A. V. Martin, C. X. Bearfield, Y. Brun, and H. Guan. 2024. Attack-Resilient Image Watermarking Using Stable Diffusion (ZoDiac). In NeurIPS (2024). [CrossRef]
  77. M. Barni, G. D’Angelo, and N. Merhav. 2007. Expanding the Class of Watermark De-synchronization Attacks. In ACM MM&Sec (2007), 195-204.
  78. X. Kong, et al. 2026. MarkCleaner: High-Fidelity Watermark Removal via Imperceptible Micro-Geometric Perturbation. arXiv preprint arXiv:2602.01513 (2026). https://arxiv.org/abs/2602.01513.
  79. F. Shamshad, N. Lukas, K. Nandakumar, et al. 2026. RAVEN: Erasing Invisible Watermarks via Novel View Synthesis. arXiv preprint arXiv:2601.08832 (2026). https://arxiv.org/abs/2601.08832.
  80. L. Ma, et al. 2024. A Geometric Distortion Immunized Deep Watermarking Framework with Robustness Generalizability (GeoWM). In ECCV (2024).
  81. H. Fang, et al. 2025. SynTag: Enhancing the Geometric Robustness of Inversion-based Generative Image Watermarking. In ICCV (2025), 15416-15425.
  82. C. Qin, et al. 2024. Print-Camera Resistant Image Watermarking With Deep Noise Simulation and Constrained Learning. IEEE TMM 26 (2024), 2164-2177.
  83. D. Francati, Y. N. Goonatilake, et al. 2026. The Coding Limits of Robust Watermarking for Generative Models. arXiv preprint arXiv:2509.10577 (2026). https://arxiv.org/abs/2509.10577.
  84. C. Zhao, H. Ling, S. Xie, et al. 2024. SSyncOA: Self-synchronizing Object-aligned Watermarking to Resist Cropping-paste Attacks. In ICME (2024). https://arxiv.org/abs/2405.03458.
  85. M. Kutter, S. Voloshynovskiy, and A. Herrigel. 2000. The Watermark Copy Attack. In Proc. SPIE 3971: Security and Watermarking of Multimedia Content II (2000).
  86. S. Craver, et al. 1998. Resolving Rightful Ownerships with Invisible Watermarking Techniques: Limitations, Attacks, and Implications. IEEE JSAC 16, 4 (1998), 573-586.
  87. A. Müller, D. Lukovnikov, J. Thietke, et al. 2025. Black-Box Forgery Attacks on Semantic Watermarks for Diffusion Models. In CVPR (2025), 20937-20946. https://arxiv.org/abs/2412.03283.
  88. Z. Dong, et al. 2025. WMCopier: Forging Invisible Image Watermarks on Arbitrary Images. In NeurIPS (2025). https://arxiv.org/abs/2503.22330.
  89. C. Zhu, et al. 2025. Optimization-Free Universal Watermark Forgery with Regenerative Diffusion Models (PnP). arXiv preprint arXiv:2506.06018 (2025). https://arxiv.org/abs/2506.06018.
  90. Z. Ba, et al. 2025. Robust Watermarks Leak: Channel-Aware Feature Extraction Enables Adversarial Watermark Manipulation (Robust-Leak / DAPAO). arXiv preprint arXiv:2502.06418 (2025). https://arxiv.org/abs/2502.06418.
  91. R. Wang, C. Lin, Q. Zhao, and F. Zhu. 2021. Watermark Faker: Towards Forgery of Digital Image Watermarking. In ICME (2021). https://arxiv.org/abs/2103.12489.
  92. Z. Yuan, L. Li, Z. Wang, and X. Zhang. 2024. Ambiguity Attack Against Text-to-Image Diffusion Model Watermarking. Signal Processing 221 (2024), 109509.
  93. T. Souček, S.-A. Rebuffi, P. Fernandez, et al. 2025. Transferable Black-Box One-Shot Forging of Watermarks via Image Preference Models. In NeurIPS (2025). https://arxiv.org/abs/2510.20468.
  94. H. Luo, L. Li, and X. Zhang. 2025. A Watermark Forgery Attack Against Stable Diffusion Model Watermarking. IEEE Signal Processing Letters 32 (2025), 3580-3584.
  95. H. Chen, T. Zhu, C. Liu, et al. 2024. High-frequency Matters: An Overwriting Attack and defense for Image-processing Neural Network Watermarking. IEEE TSC 17, 4 (2024), 1565-1579. https://arxiv.org/abs/2302.08637.
  96. K. Arabi, B. Feuer, R. T. Witter, C. Hegde, and N. Cohen. 2025. Hidden in the Noise: Two-Stage Robust Watermarking for Images (WIND). In ICLR (2025). https://arxiv.org/abs/2412.04653.
  97. T. Zhou, R. Ding, G. Liu, et al. 2026. MetaSeal: Defending Against Image Attribution Forgery Through Content-Dependent Cryptographic Watermarks. In TMLR (2026).
  98. Z. Zou, B. Gong, and L. Wang. 2025. Attention to Neural Plagiarism: Diffusion Models Can Plagiarize Your Copyrighted Images! In ICCV (2025), 19546-19556.
  99. B. An, M. Ding, T. Rabbani, et al. 2024. WAVES: Benchmarking the Robustness of Image Watermarks. In ICML (2024). PMLR 235, 1456-1492. https://arxiv.org/abs/2401.08573.
  100. Y. N. Goonatilake and G. Ateniese. 2026. Removing the Watermark Is Not Enough: Forensic Stealth in Generative-AI Watermark Removal. arXiv preprint arXiv:2605.09203 (2026). https://arxiv.org/abs/2605.09203.
  101. M. Ding, T. Rabbani, B. An, et al. 2024. NeurIPS 2024 Competition: Erasing the Invisible: A Stress-Test Challenge for Image Watermarks. In NeurIPS Competition (2024).
  102. X. Zhao, S. Gunn, M. Christ, et al. 2025. SoK: Watermarking for AI-Generated Content. In IEEE S&P (2025). https://arxiv.org/abs/2411.18479.
  103. A. Yakushev, et al. 2025. WIBE: Watermarks for generated Images - Benchmarking & Evaluation. In ASE (2025), 4033-4036. [CrossRef]
  104. J. Wen, J. Kirchenbauer, J. Geiping, and T. Goldstein. 2023. Tree-Ring Watermarks: Fingerprints for Diffusion Images that are Invisible and Robust. In NeurIPS (2023).
  105. P. Fernandez, G. Couairon, H. Jégou, M. Douze, and T. Furon. 2023. The Stable Signature: Rooting Watermarks in Latent Diffusion Models. In ICCV (2023).
  106. Z. Yang, K. Zeng, K. Chen, et al. 2024. Gaussian Shading: Provable Performance-Lossless Image Watermarking for Diffusion Models. In CVPR (2024).
  107. M. Tancik, B. Mildenhall, and R. Ng. 2020. StegaStamp: Invisible Hyperlinks in Physical Photographs. In CVPR (2020).
  108. H.-H. Nguyen-Le, V.-T. Tran, T. Nguyen, and N.-A. Le-Khac. 2025. A Survey on Proactive Deepfake Defense: Disruption and Watermarking. ACM Computing Surveys 58, 5 (2025).
  109. W. Wan, J. Wang, Y. Zhang, J. Li, H. Yu, and J. Sun. 2022. A comprehensive survey on robust image watermarking. Neurocomputing 488 (2022), 226-247. [CrossRef]
  110. Z. Wang, O. Byrnes, H. Wang, R. Sun, C. Ma, H. Chen, Q. Wu, and M. Xue. 2023. Data Hiding With Deep Learning: A Survey Unifying Digital Watermarking and Steganography. IEEE Trans. Computational Social Systems 10, 6 (2023), 2985-2999. [CrossRef]
  111. Y. Luo, X. Tan, and Z. Cai. 2024. Robust Deep Image Watermarking: A Survey. Computers, Materials & Continua 81, 1 (2024), 133-160. [CrossRef]
  112. P. Ye, H. Ren, Z. Li, A. Yan, H. Yan, S. Wang, and J. Li. 2026. Securing Large Language Models: A Survey of Watermarking and Fingerprinting Techniques. ACM Computing Surveys 58, 7 (2026), Article 187. [CrossRef]
  113. J. Cao, Q. Li, Z. Zhang, J. Ni, and R. Lu. 2026. Secure and Robust Watermarking for AI-generated Images: A Comprehensive Survey. arXiv preprint arXiv:2510.02384 (2026). https://arxiv.org/abs/2510.02384.
  114. X.-B. Kang, F. Zhao, G.-F. Lin, and Y.-J. Chen. 2018. A novel hybrid of DCT and SVD in DWT domain for robust and invisible blind image watermarking with optimal embedding strength. Multimedia Tools and Applications 77, 11 (2018), 13197-13224. [CrossRef]
  115. I. J. Cox, J. Kilian, F. T. Leighton, and T. Shamoon. 1997. Secure Spread Spectrum Watermarking for Multimedia. IEEE TIP 6, 12 (1997), 1673-1687. [CrossRef]
  116. X.-Y. Wang, C.-P. Wang, H.-Y. Yang, and P.-P. Niu. 2013. A robust blind color image watermarking in quaternion Fourier transform domain. Journal of Systems and Software 86, 2 (2013), 255-277. [CrossRef]
  117. N. M. Makbol and B. E. Khoo. 2013. Robust blind image watermarking scheme based on redundant discrete wavelet transform and singular value decomposition. AEU - International Journal of Electronics and Communications 67, 2 (2013), 102-112. [CrossRef]
  118. B. Chen and G. W. Wornell. 2001. Quantization Index Modulation: A Class of Provably Good Methods for Digital Watermarking and Information Embedding. IEEE Trans. Information Theory 47, 4 (2001), 1423-1443.
  119. K. A. Zhang, L. Xu, A. Cuesta-Infante, and K. Veeramachaneni. 2019. Robust Invisible Video Watermarking with Attention (RivaGAN). arXiv preprint arXiv:1909.01285 (2019). https://arxiv.org/abs/1909.01285.
  120. Z. Jia, H. Fang, and W. Zhang. 2021. MBRS: Enhancing Robustness of DNN-based Watermarking by Mini-Batch of Real and Simulated JPEG Compression. In ACM MM (2021). [CrossRef]
  121. R. Ma, M. Guo, Y. Hou, F. Yang, Y. Li, H. Jia, and X. Xie. 2022. Towards Blind Watermarking: Combining Invertible and Non-invertible Mechanisms (CIN). In ACM MM (2022). [CrossRef]
  122. H. Fang, Y. Qiu, K. Chen, J. Zhang, W. Zhang, and E.-C. Chang. 2023. Flow-Based Robust Watermarking with Invertible Noise Layer for Black-Box Distortions (FIN). In AAAI (2023). [CrossRef]
  123. H. Fang, Z. Jia, Z. Ma, E.-C. Chang, and W. Zhang. 2022. PIMoG: An Effective Screen-shooting Noise-Layer Simulation for Deep-Learning-Based Watermarking Network. In ACM MM (2022). [CrossRef]
  124. T. Bui, S. Agarwal, N. Yu, and J. Collomosse. 2023. RoSteALS: Robust Steganography using Autoencoder Latent Space. In CVPRW (2023). [CrossRef]
  125. P. Fernandez, A. Sablayrolles, T. Furon, H. Jégou, and M. Douze. 2022. Watermarking Images in Self-Supervised Latent Spaces. In ICASSP (2022), 3054-3058. https://arxiv.org/abs/2112.09581.
  126. X. Wu, X. Liao, and B. Ou. 2023. SepMark: Deep Separable Watermarking for Unified Source Tracing and Deepfake Detection. In ACM MM (2023). [CrossRef]
  127. R. Xu, et al. 2025. InvisMark: Invisible and Robust Watermarking for AI-generated Image Provenance. In WACV (2025). [CrossRef]
  128. T. Sander, P. Fernandez, A. Durmus, T. Furon, and M. Douze. 2024. Watermark Anything with Localized Messages (WAM). arXiv preprint arXiv:2411.07231 (2024). https://arxiv.org/abs/2411.07231.
  129. X. Zhang, R. Li, J. Yu, Y. Xu, W. Li, and J. Zhang. 2024. EditGuard: Versatile Image Watermarking for Tamper Localization and Copyright Protection. In NeurIPS (2020).
  130. J. Huang, T. Luo, L. Li, G. Yang, H. Xu, and C.-C. Chang. 2023. ARWGAN: Attention-Guided Robust Image Watermarking Model Based on GAN. IEEE Transactions on Instrumentation and Measurement (2023). [CrossRef]
  131. X. Zhang, R. Li, J. Yu, Y. Xu, W. Li, and J. Zhang. 2024. EditGuard: Versatile Image Watermarking for Tamper Localization and Copyright Protection. In CVPR (2024). [CrossRef]
  132. A. Rezaei, M. Akbari, S. R. Alvar, A. Fatemi, and Y. Zhang. 2024. LaWa: Using Latent Space for In-Generation Image Watermarking. In ECCV (2024). [CrossRef]
  133. W. Feng, W. Zhou, J. He, J. Zhang, T. Wei, G. Li, T. Zhang, W. Zhang, and N. Yu. 2024. AquaLoRA: Toward White-box Protection for Customized Stable Diffusion Models via Watermark LoRA. In ICML (2024). PMLR 235, 13423-13444.
  134. N. Yu, V. Skripniuk, S. Abdelnabi, and M. Fritz. 2021. Artificial Fingerprinting for Generative Models: Rooting Deepfake Attribution in Training Data. In ICCV (2021). [CrossRef]
  135. Z. Ma, G. Jia, B. Qi, and B. Zhou. 2024. Safe-SD: Safe and Traceable Stable Diffusion with Text Prompt Trigger for Invisible Generative Watermarking. In ACM MM (2024). [CrossRef]
  136. H. Ci, P. Yang, Y. Song, and M. Z. Shou. 2024. RingID: Rethinking Tree-Ring Watermarking for Enhanced Multi-key Identification. In ECCV (2024). [CrossRef]
  137. S. Gunn, X. Zhao, and D. Song. 2025. An Undetectable Watermark for Generative Image Models (PRC). In ICLR (2025). https://arxiv.org/abs/2410.07369.
  138. H. Huang, Y. Wu, and Q. Wang. 2024. ROBIN: Robust and Invisible Watermarks for Diffusion Models with Adversarial Optimization. In NeurIPS (2024). [CrossRef]
  139. K. Li, Z. Huang, X. Hou, and C. Hong. 2025. GaussMarker: Robust Dual-Domain Watermark for Diffusion Models. In ICML (2025). PMLR 267, 34688-34701.
  140. K. Arabi, R. T. Witter, C. Hegde, and N. Cohen. 2025. SEAL: Semantic Aware Image Watermarking. In ICCV (2025). https://arxiv.org/abs/2503.12172.
  141. S. J. Lee and N. I. Cho. 2025. Semantic Watermarking Reinvented: Enhancing Robustness and Generation Quality with Fourier Integrity (SFW; including HSTR/HSQR). In ICCV (2025). https://arxiv.org/abs/2509.07647.
  142. L. Lei, K. Gai, J. Yu, and L. Zhu. 2024. DiffuseTrace: A Transparent and Flexible Watermarking Scheme for Latent Diffusion Model. arXiv preprint arXiv:2405.02696 (2024). https://arxiv.org/abs/2405.02696.
  143. Y. Zhao, T. Pang, C. Du, X. Yang, N.-M. Cheung, and M. Lin. 2023. A Recipe for Watermarking Diffusion Models (WatermarkDM). arXiv preprint arXiv:2303.10137 (2023). https://arxiv.org/abs/2303.10137.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.