Digital watermarking is the technique of embedding identifying or authenticating information directly into digital content such as images, audio, video, text, or model outputs, ideally so the mark is imperceptible yet recoverable. A watermark may be robust, surviving compression and editing, or fragile, breaking on tampering to signal alteration. It is increasingly used to mark AI-generated media for provenance and to support copyright protection and content authentication.

Semantic Classification

Content

Compositional Relationships (Components)

SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:hasPart ai:WatermarkEncoder))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:hasPart ai:WatermarkDetector))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:hasPart ai:WatermarkPayload))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:hasPart ai:RobustnessProperty))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:hasPart ai:ImperceptibilityProperty))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:hasPart ai:WatermarkCapacity))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:hasPart ai:CoverMedium))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:hasPart ai:CryptographicKey))

Dependency Relationships

SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:requires ai:CryptographicKey))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:requires ai:StatisticalSignalModel))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:requires ai:CoverMedium))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:requires ai:ErrorCorrectingCode))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:dependsOn ai:SignalProcessing))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:dependsOn ai:InformationTheory))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:dependsOn ai:Cryptography))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:dependsOn ai:DeepLearning))

Capability Relationships

SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:enables ai:ContentAuthentication))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:enables ai:CopyrightProtection))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:enables ai:AIOriginDeclaration))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:enables ai:LeakTracing))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:enables ai:ProvenanceVerification))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:enables ai:TamperDetection))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:supports ai:DeepfakeDetection))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:supports ai:RegulatoryCompliance))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:supports ai:DataProvenance))

Implementation Relationships

SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:implements ai:SteganographicEmbedding))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:implements ai:FrequencyDomainEncoding))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:implements ai:LatentSpacePerturbation))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:implements ai:TokenProbabilityBiasing))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:implements ai:SpreadSpectrumSignalling))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:uses ai:Steganography))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:uses ai:ErrorCorrectingCode))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:uses ai:ZeroKnowledgeProof))

Reduction Relationships

SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:reducesTo ai:InformationHiding))
SubClassOf(ai:DigitalWatermarking
  ObjectSomeValuesFrom(ai:reducesTo ai:ContentProvenance))
SubClassOf(ai:RobustWatermark
  ObjectSomeValuesFrom(ai:reducesTo ai:DigitalWatermarking))
SubClassOf(ai:FragileWatermark
  ObjectSomeValuesFrom(ai:reducesTo ai:DigitalWatermarking))
SubClassOf(ai:TextWatermark
  ObjectSomeValuesFrom(ai:reducesTo ai:DigitalWatermarking))
SubClassOf(ai:LatentDiffusionWatermark
  ObjectSomeValuesFrom(ai:reducesTo ai:DigitalWatermarking))
SubClassOf(ai:ModelWatermark
  ObjectSomeValuesFrom(ai:reducesTo ai:DigitalWatermarking))

About

  • Digital watermarking traces its conceptual lineage to centuries-old techniques in analogue media: paper manufacturers embedded translucent designs (watermarks) into paper sheets during manufacture by introducing a thinner area in the paper that allows light to pass through differently, visible when the paper is held up to light. Banknote printers used invisible inks (fluorescent under UV), microprint text (only legible under magnification), and security threads embedded into the paper substrate. These analogue precursors demonstrate that the fundamental problem — embedding identifying information invisibly into a physical or digital substrate, detectable by authorised parties but imperceptible to the general population — predates the digital era by centuries and has deep roots in security printing, intelligence tradecraft, and media authentication.
  • In the digital era, formal technical frameworks emerged in the 1990s primarily driven by the music and film industries’ urgent need to protect commercial audio and video assets from digital piracy. The proliferation of CD and DVD ripping tools created a crisis of unauthorised reproduction, and content owners demanded a technical mechanism to embed ownership and distribution-channel information that would survive copying and re-encoding. Bender, Gruhl, Morimoto, and Lu at the MIT Media Lab formalised the spread-spectrum approach in 1996: a watermark signal is spread across many frequency components of an image or audio sample by pseudo-random modulation, echoing spread-spectrum radio communications where a signal is spread across a wide bandwidth so it appears as noise to any observer not possessing the spreading code. Cox, Kilian, Leighton, and Shamoon at NEC Research and Bellcore simultaneously published the foundational paper “Secure Spread Spectrum Watermarking for Multimedia” in IEEE Transactions on Image Processing (1997), establishing the theoretical underpinning and demonstrating that embedding a mark in the perceptually significant components of the discrete cosine transform (DCT) or wavelet transform domain maximises robustness against compression and editing while remaining within the human just-noticeable difference (JND) threshold for perceived image quality.
  • The DCT-domain spread-spectrum approach became the basis for commercial watermarking products throughout the broadcast and film industries, including Digimarc’s image watermarking platform and Civolution’s (later Kantar Media’s) broadcast watermarking systems. These products embedded session-specific or distribution-channel identifiers in broadcast video signals, enabling post-hoc tracing of leaks to specific affiliates or distribution points. The Digital Rights Management industry adopted watermarking as a complement to encryption: whereas encryption prevents unauthorised access to the content, a watermark survives decryption and persists in the plaintext media, providing a forensic trail even after authorised decryption and subsequent unauthorised redistribution.
  • The landscape was fundamentally transformed in the late 2010s and early 2020s by the emergence of deep neural network-based encoder–decoder watermarking frameworks. The HiDDeN framework (Zhu, Kaplan, Johnson, Fei-Fei, NeurIPS 2018) introduced end-to-end differentiable training of a watermark encoder and detector pair, co-optimising imperceptibility (measured by L2 or LPIPS loss against the unwatermarked image) and robustness jointly under a differentiable noise layer that approximates common image distortions including JPEG compression, cropping, brightness adjustment, and Gaussian noise. This learnt approach significantly outperformed hand-crafted frequency-domain methods, particularly against adversarial removal attacks that exploit knowledge of the specific frequency coefficients used. The StegaStamp framework (Tancik, Mildenhall, Ng, CVPR 2020) extended DNN watermarking to photographic print-and-scan recovery, enabling QR-code-equivalent payload recovery from photographs of printed pages, demonstrating robustness to extreme geometric and photometric distortions.
  • The generative AI era opened an entirely new embedding paradigm: rather than inserting a watermark post-hoc into a finished content artefact, the watermark can be baked into the generative process itself, making it fundamentally more difficult to remove than any post-hoc technique applied after generation. Google DeepMind’s SynthID system, initially released in 2023 for Imagen-generated images and expanded throughout 2024 to cover text (Gemini), audio (Lyria, NotebookLM), and video, exemplifies this approach. For image generation, SynthID perturbs the initial noise latent of the diffusion process so that the resulting pixel distribution, while visually identical to unwatermarked outputs, carries a detectable statistical signature in certain frequency patterns. For text generation, SynthID-Text biases the token sampling logit scores at inference time, creating a statistical signature in the distribution of generated tokens detectable by a matched classifier without access to the model weights. By late 2024, SynthID had been applied to over ten billion images and video frames across Google’s AI services, and a consumer-facing SynthID detector was released in 2025 enabling pixel-level watermark localisation to highlight suspected watermarked regions within an image.
  • For Large Language Model text outputs more broadly, token-level watermarking was formalised by Kirchenbauer, Geiping, Wen, Katz, Miers, and Goldstein at the University of Maryland in a 2023 ICML paper. Their green-list/red-list scheme partitions the vocabulary into green and red token sets conditioned on the preceding token context using a shared secret key, then biases the model’s sampling to preferentially select green tokens. A detector lacking the model weights but possessing the key can verify the watermark by computing the proportion of green tokens in the text: genuine watermarked text exhibits a statistically significant green-list excess compared to unwatermarked text. This scheme requires no model retraining and adds negligible inference latency, but is vulnerable to paraphrase attacks by a model that rearranges sentences while preserving semantics. Tree-Ring Watermarking (Wen et al., NeurIPS 2023) addressed the image domain by embedding the watermark pattern in the Fourier frequency space of the initial latent noise, achieving robustness to regeneration attacks — the hardest known removal attack, in which an adversary passes the watermarked image through a separate diffusion model to “scrub” the mark — without any additional training of the embedding or detection networks.

Technical Approaches / Architecture

  • Spatial-domain embedding: The simplest and historically earliest approach modifies pixel values directly in the spatial (pixel) domain, most commonly by substituting the least-significant bit (LSB) of each pixel’s colour channel with a watermark bit, exploiting the human visual system’s inability to distinguish pixel values differing by 1 out of 255. LSB substitution offers very high payload capacity (one bit per channel per pixel) and zero perceptual distortion at 1-bit depth, but is entirely fragile to JPEG re-compression, which rounds coefficient values and destroys LSB information. More sophisticated spatial-domain approaches modify pixel values by a small additive or multiplicative amount derived from a pseudo-random mask, achieving greater robustness than pure LSB but still falling short of frequency-domain methods.
  • Frequency-domain embedding: Rather than modifying pixel values directly, frequency-domain methods transform the image into a frequency representation (Discrete Cosine Transform, Discrete Wavelet Transform, or Discrete Fourier Transform) and embed the watermark into the transform coefficients, then invert the transform. Embedding in mid-frequency DCT coefficients provides an ideal balance: these coefficients contribute to visible textures but not to the very high frequencies that JPEG compression discards, achieving robustness to JPEG while remaining imperceptible. The Cox et al. spread-spectrum method adds a Gaussian-distributed watermark vector to the N largest DCT coefficients (after the DC coefficient), achieving proven robustness to arbitrary linear operations and collusion attacks under certain statistical conditions.
  • Spread-spectrum watermarking: Pioneered by Cox et al. (1997), spread-spectrum watermarking distributes the watermark energy across all frequency components by modulating a pseudo-random carrier sequence with the watermark payload, analogous to direct-sequence spread-spectrum (DSSS) radio transmission. The correlator detector applies the matched filter (cross-correlation with the known carrier sequence) to recover the payload from any received version of the watermarked content. Spread-spectrum provides high robustness against noise addition, filtering, and compression at the cost of low payload capacity (typically 1–64 bits per image) and requires careful parameter tuning to remain perceptually invisible in images with strong uniform textures.
  • Deep neural network (DNN) encoder–decoder: End-to-end trained encoder–decoder architectures represent the state-of-the-art for robust image watermarking. The encoder is a Convolutional Neural Network that maps (cover image, payload bitstring) → watermarked image, constrained by a perceptual quality loss. The decoder is a CNN that maps any (possibly distorted) version of the watermarked image → recovered payload bits. Training incorporates a differentiable noise layer simulating JPEG compression, cropping, brightness/contrast adjustment, blurring, and additive noise, forcing the encoder–decoder pair to learn representations resilient to these distortions. The joint optimisation of imperceptibility and robustness under adversarial noise conditions yields substantial improvements over hand-crafted methods. HiDDeN (2018) was the foundational framework; subsequent work including MBRS (2021), RivaGAN (2019), and DWTDCT (2020) extended the approach with improved architectures.
  • Latent diffusion embedding (Tree-Ring, SynthID-Image, Stable Signature): A new paradigm specific to Diffusion Model-based image generators. Rather than post-hoc embedding into a completed image, the watermark is incorporated into the initial noise latent from which the image is generated, so the watermark propagates through the entire denoising process and is inherently woven into the image’s statistical structure. Tree-Ring Watermarking (Wen et al., 2023) embeds a circular watermark pattern in the Fourier spectrum of the initial Gaussian noise; the pattern is preserved through DDIM inversion and persists through image edits, cropping, and even regeneration by a separate diffusion model. The Stable Signature (Fernandez et al., ICCV 2023) fine-tunes the latent diffusion model’s decoder so all generated images carry a specific multi-bit signature. SynthID-Image operates analogously at the level of the diffusion model’s noise prediction process.
  • Token probability biasing for text watermarking: In LLM-based text watermarking, the watermark is embedded at inference time by shifting the probability distribution over the next-token vocabulary before sampling. The Kirchenbauer et al. (2023) scheme defines a green-list subset of the vocabulary for each context window (derived from the previous token and a shared secret key) and adds a small positive logit bias (typically δ=2.0) to all green-list tokens before softmax sampling. The resulting text contains a statistical excess of green tokens detectable by a one-tailed z-test on the green-token count. SynthID-Text employs a more sophisticated multi-bit tournament sampling scheme that embeds a structured binary message rather than a single detection bit. Both approaches share the property that the watermark introduces negligible distortion to text quality (measured by perplexity) when δ is small and the vocabulary is large.
  • Robust/fragile dual-mode and semi-fragile designs: Robust watermarks are optimised to persist; fragile watermarks are designed to break on any modification, serving tamper-evidence applications where the absence or corruption of the watermark proves manipulation. Semi-fragile watermarks occupy a designed intermediate: they survive JPEG compression at quality 75+ (considered benign normalisation) but break under any content-modifying operation such as copy-paste, object insertion, or face swap. Medical imaging watermarking (DICOM) frequently uses semi-fragile schemes to distinguish authorised PACS processing from malicious pixel-level manipulation.
  • Zero-knowledge proof (ZKP) watermark verification: A frontier technique (zkDL++, Bagad et al., 2025) that allows a content creator to prove to a verifier that a specific watermark is present in a piece of content, without revealing the watermark key or the embedding function. ZKP watermark verification decouples detection from key disclosure, enabling verifiable provenance claims in adversarial or privacy-sensitive settings, and represents the cryptographic frontier of Content Provenance authentication for AI-generated content.

Major Families and Variants

  • Image Watermarking: The most mature subdomain, spanning DCT-domain spread-spectrum (classical), DNN encoder-decoder (HiDDeN, StegaStamp, MBRS), latent diffusion (SynthID-Image, Tree-Ring, RingID, Stable Signature, ShapeMark), and frequency-domain DNN hybrids (DWT-DCT-SVD fusion). Image watermarks are evaluated on: imperceptibility (PSNR, SSIM, LPIPS against unwatermarked), robustness (bit accuracy or detection rate after JPEG, crop, rotate, blur, resize, colour jitter, and regeneration attack), and capacity (bits per image or per pixel).
  • Audio Watermarking: Commercial systems (Civolution/Kantar Media, Digimarc) embed multi-bit session identifiers in the audio signal using psychoacoustic models that identify frequency components masked by the audio and embed the watermark below the masking threshold. DNN-based audio watermarking systems (AudioSeal, WavMark, 2023–2024) extend the encoder-decoder framework to the temporal audio domain, learning embeddings robust to MP3/AAC compression, pitch shifting, time stretching, and additive background noise. SynthID Audio extends the approach to AI-generated audio from Lyria and NotebookLM.
  • Video Watermarking: Video watermarking applies per-frame image watermarking with additional temporal coherence constraints ensuring the watermark pattern is consistent across frames to prevent detection by frame-differencing. SynthID Video extends image and audio methods across all frames of AI-generated video. Broadcast watermarking systems (A/V sync marking, Nielsen audio codes) are highly mature, with robustness to tape copying, transcoding, and even brief audio dropout.
  • Text Watermarking: The newest major subdomain, activated by the rise of Large Language Model deployment. Paradigms include: token probability biasing (Kirchenbauer 2023 green-list); semantic watermarking (embedding information in paraphrase-invariant syntactic or semantic structures); multi-bit robust text watermarking (Xu et al., 2025, combining learnt features with error-correcting codes); and autoregressive image watermarking through lexical biasing (2026), which applies LLM-style token biasing to image autoregressive models.
  • Model Watermarking: A distinct subdomain embedding ownership signals into the weights or activation patterns of trained Neural Network models rather than into content they produce. Model watermarks enable verification of model IP: an owner can prove to a third party that a suspicious model is a copy of their own by querying a secret set of trigger inputs and verifying the expected output pattern. This is relevant for protecting proprietary generative AI model weights.
  • Fragile / Tamper-Evidence Watermarking: Watermarks designed for integrity verification rather than identification. Any pixel-level modification — splicing, object insertion, face swap — destroys the watermark, providing a tamper alarm. Used in medical imaging (DICOM integrity), legal document authentication, and identity credential verification. Block-based fragile schemes localise tampering to specific image regions.

Use Cases

  • AI-generated content labelling (regulatory): The EU AI Act Article 50, enforceable 2 August 2026, requires that generative AI outputs be marked in a machine-readable format detectable as artificially generated. Watermarking provides a content-bound label that persists even if file metadata is stripped, resaved, or converted, unlike pure metadata-based approaches such as XMP tags. The European Commission’s Code of Practice on AI Content Transparency (final version expected May-June 2026) explicitly references both C2PA metadata credentials and embedded watermarks as compliant mechanisms. Platforms including Google (SynthID), Meta, and Microsoft have announced watermarking integration across their generative AI services in anticipation of this enforcement date.
  • Copyright and ownership tracing in media production: Major film studios (Warner, Universal, Disney) embed distributor-specific watermarks in pre-release screener copies sent to critics and award voters, enabling post-hoc leak attribution. Record labels and streaming services (Spotify, Apple Music) use audio watermarking to trace session-specific identifiers embedded in high-resolution audio files distributed to radio stations and licensees. AI model providers (Google, Stability AI, Midjourney) are implementing watermarks in generated outputs to assert model provenance and support copyright claims over AI-assisted creative works.
  • Social media platform content moderation: Meta, Google, and LinkedIn have committed to integrating watermark detection into their content moderation pipelines to identify and label AI-generated images and video in feeds. LinkedIn uses C2PA content credentials (itself compatible with SynthID watermarks) to display provenance labels on AI-generated profile images. The Content Authenticity Initiative (CAI), which administers C2PA, counts Adobe, Microsoft, Sony, Canon, and Nikon among its members, with camera manufacturers beginning to embed C2PA content credentials at the point of capture.
  • Medical imaging integrity and chain-of-custody: PACS (Picture Archiving and Communication System) integrity in hospital settings is a regulatory requirement in many jurisdictions. Semi-fragile or fragile watermarks embedded in DICOM-format medical images (radiographs, MRI, CT, pathology slides) provide a pixel-level tamper alarm: any modification to the image breaks the watermark, flagging potential manipulation in a forensic or legal context. UK NHS digital pathology deployments at multiple trusts are evaluating DICOM watermarking for chain-of-custody compliance.
  • Deepfake detection complementarity: Watermark detectors operated at media ingestion points (social media upload APIs, news agency wire submission systems) can flag content whose expected watermark is absent (potentially manipulated and stripped) or that contains a forged watermark (inconsistent with any known issuing system), providing a complementary signal to discriminative Deepfake Detection models. The absence of an expected SynthID watermark in an image claimed to come from Google’s Gemini, for example, is evidence of manipulation.
  • Broadcast and streaming digital rights management: Commercial broadcast watermarking systems (Nielsen audio watermarks, Civolution/Kantar NexGuard video watermarks) carry session-specific or affiliate-specific identifiers in broadcast video and audio streams, enabling real-time and post-hoc tracking of unauthorised redistribution on illegal streaming platforms. Kantar Media’s Arbitron audio watermarking is embedded in virtually all US broadcast audio content for audience measurement purposes. These systems are extremely mature, having operated at billion-impressions scale for decades.
  • Scientific data integrity: Scientific image watermarking is emerging as a tool for combating research fraud (image manipulation in biomedical publications), complementing existing perceptual hash-based duplicate detection. Fragile watermarks embedded at the point of microscopy capture can demonstrate that a published image has not been selectively cropped, brightness-adjusted, or composited, supporting reproducibility and integrity claims.

Theoretical Foundations

  • The information-theoretic foundation of watermarking was established by Claude Shannon’s noisy channel coding theory (1948) and the seminal “writing on dirty paper” capacity result by Max Costa (1983) and the “writing on dirty paper” capacity result by Costa (1983). Costa proved that if a transmitter knows the interference (the cover signal into which the watermark is embedded) non-causally, the channel capacity is the same as if the interference were absent — implying that the watermark encoder can in principle pre-cancel the cover signal’s effect and use the full channel capacity for the payload. This theoretical result underpins the achievability of robust, high-capacity watermarking at the Shannon limit, though practical systems remain well below this bound due to computational constraints and imperfect knowledge of the decoder’s operating conditions.
  • The detection-theoretic framework models watermark detection as a binary hypothesis test: H0 (no watermark present) versus H1 (watermark present). The false positive rate (FPR) is the probability of detecting a watermark in unwatermarked content; the true positive rate (TPR) or detection power is the probability of correctly detecting a present watermark. The receiver operating characteristic (ROC) curve traces the trade-off between TPR and FPR as the detection threshold varies; for a fixed FPR (e.g., 1%), the maximum achievable TPR is the Neyman-Pearson optimal test. In Kirchenbauer’s text watermarking scheme, the z-score test is asymptotically equivalent to the Neyman-Pearson optimal test for the binomial green-token detection problem, with performance improving as text length increases.
  • The information-hiding game between the watermark embedder and the attacker is formalised as a security game in which the adversary is computationally bounded. The security of a keyed watermarking scheme is defined analogously to semantic security of encryption: no polynomial-time attacker can distinguish a watermarked from an unwatermarked carrier with non-negligible advantage, assuming only the distribution of watermarked content and not the secret key. This security definition motivates key management requirements: keys must be long enough (256 bits minimum) to resist brute-force and collision attacks, and must be rotated to limit the amount of content exposed under a single key.
  • The “no free lunch” theorem for watermarking (Piet et al., 2024) demonstrates that quality (text naturalness measured by perplexity), detection rate (minimum text length for reliable detection), and tamper resistance (robustness to paraphrase attacks) are inherently in tension: any watermarking scheme that improves on one metric does so at the cost of at least one other. This result parallels the fundamental impossibility in steganography between undetectability and message rate (Cachin 1998), confirming that the trade-off structure of watermarking is a fundamental feature of information hiding rather than an artefact of specific schemes. Specifically, increasing the watermark strength (logit bias δ in text watermarking) improves detection rate but degrades text quality; increasing the vocabulary partition ratio improves detection sensitivity but reduces capacity per token; and any scheme robust to paraphrase attacks must embed information in paraphrase-invariant semantic features, which by definition carry less textual information per unit length.
  • The connection between watermarking and error-correcting codes is fundamental: the payload embedded by a watermark is best understood as a codeword in an error-correcting code (ECC), where the “channel noise” is the combination of the cover signal, any post-processing distortions, and adversarial removal attacks. Robust watermarking schemes increasingly incorporate BCH codes, Reed-Solomon codes, LDPC codes, or polar codes to provide payload redundancy: even if some watermark bits are corrupted by distortions or attacks, the ECC can recover the full payload from the surviving bits. SynthID-Text employs tournament sampling with error-correcting structure; Tree-Ring Watermarking embeds a structured Fourier pattern that functions as a low-rate error-correcting code across the spatial frequency components of the image.

Attack Taxonomy and Adversarial Robustness

  • The watermarking security model distinguishes three classes of adversary: a passive observer who attempts only to detect whether a watermark is present (steganalysis / detection attack); a semi-active attacker who attempts to destroy or degrade the watermark payload below the detection threshold (removal attack); and an active attacker who attempts to forge a valid watermark or transplant a watermark from one item of content to another (forgery / transplantation attack). The security of a watermarking system must be evaluated against all three threat classes.
  • Removal attacks — additive noise: Adding Gaussian or uniform noise to the watermarked content attempts to mask the watermark signal. Spread-spectrum and DNN-based watermarks are generally robust to this because the watermark energy is distributed across many frequency components or learnt to be resilient to noise during training. The watermark capacity can be estimated from signal-to-noise considerations: a watermark embedded at a given signal strength (measured in peak signal-to-noise ratio, PSNR) can survive noise up to a threshold determined by the detector’s statistical sensitivity.
  • Removal attacks — geometric transformations: Rotation, scaling, translation, cropping, and perspective warping are classical attacks because they break the spatial alignment assumed by correlation-based detectors. Template-based watermarks (embedding a known reference pattern whose detection enables alignment recovery) provide robustness to geometric attacks. DNN-based detectors trained on geometrically augmented data learn spatial invariance directly.
  • Removal attacks — compression: JPEG quantisation in the DCT domain discards high-frequency information, which is precisely where many early spatial-domain watermarks were embedded. JPEG-robust watermarks use mid-frequency DCT coefficients (quantisation indices 2–15) that survive standard JPEG quality settings 75–95. Neural compression codecs (VVC, BPG, CompressAI) are harder to predict; DNN watermarks trained on differentiable JPEG approximations generalise poorly to neural codecs, motivating training on diverse compression channels.
  • Removal attacks — content-aware image editing: Automatic object removal, face swapping, inpainting, and style transfer in commercial photo editors (Adobe Photoshop generative fill, DALL-E 2 inpainting) destroy watermarks in the edited regions while leaving them in unedited background areas. Semi-fragile watermarks in the edited regions provide localised tamper evidence; robust watermarks may partially survive in the remaining unedited content.
  • Removal attacks — regeneration (diffusion-based scrubbing): The hardest known image watermark removal attack involves running the watermarked image through a separate Diffusion Model (using DDIM inversion to find the approximate noise latent, then regenerating from a nearby latent). This can remove most watermark patterns while preserving visual content because diffusion regeneration replaces the statistical structure of the pixel distribution without fundamentally altering perceptual content. Tree-Ring and SynthID-Image, which embed watermarks in the fundamental noise latent from which the image is generated, are substantially more robust to regeneration attacks than post-hoc DNN-embedded watermarks.
  • Forgery attacks: An attacker with knowledge of the watermarking system but not the secret key attempts to construct a valid watermark from scratch or to copy a legitimate watermark from one piece of content to another (transplantation). The security against forgery depends on the secrecy of the embedding key: without the key, a statistical test should detect invalid watermarks. Key management and rotation are therefore critical operational requirements for production watermarking systems.
  • Detection attacks (steganalysis): Even without recovering the payload, an attacker may wish to detect whether a given piece of content is watermarked — for example, to selectively scrub content before redistribution. Statistical steganalysis uses the Rich Model (a feature vector computed from high-order statistics of pixel residuals) or trained CNN classifiers to distinguish watermarked from unwatermarked images. DNN-based watermarking methods trained to minimise statistical detectability (security-aware training) are more resistant to steganalysis than classic methods.
  • The robustness–imperceptibility–capacity trade-off (fundamental limits): Information theory places fundamental limits on the trade-off between watermark capacity (bits per pixel or bits per symbol), imperceptibility (measured by PSNR or distortion measure), and robustness (survival rate under a given channel distortion level). At fixed imperceptibility, increasing capacity reduces robustness, and vice versa. The channel capacity for information hiding was analysed by Costa (1983) as an instance of writing on dirty paper — the fundamental capacity limit is the same as the channel capacity without the interference term, suggesting that clever encoder design can approach the theoretical maximum. In practice, DNN-based methods are much closer to this theoretical limit than classical methods.

Academic Context

  • The field emerged formally in the mid-1990s at the MIT Media Lab (Bender, Gruhl, Morimoto, Lu, 1996) and at NEC Research (Cox et al., 1997). The information-theoretic foundations were established by Cachin (1998), who analysed the capacity of information hiding channels and the fundamental limits of steganographic detectability. The first international conference dedicated to the field was the Information Hiding Workshop (1996, Cambridge, UK), which remains a leading venue alongside the IEEE International Workshop on Information Forensics and Security (WIFS) and the ACM Workshop on Multimedia and Security.
  • The textbook by Cox, Miller, Bloom, Fridrich, and Kalker (2007, Morgan Kaufmann) consolidates the classical theory. The Petitcolas, Anderson, and Kuhn survey (1999, Proceedings of the IEEE) remains the canonical taxonomy, distinguishing watermarking from steganography and fingerprinting. Steganalysis (the detection of watermarks or steganographic content) has developed in parallel as an adversarial discipline, with the RS (Regular-Singular) steganalysis technique (Fridrich et al., 2001) and the SRM (Spatial Rich Model) feature set (Fridrich and Kodovsky, 2012) as major landmarks.
  • Deep learning-era watermarking was pioneered by HiDDeN (Zhu et al., NeurIPS 2018). The StegaStamp work (Tancik et al., CVPR 2020) extended the approach to print-scan recovery. Generative model-specific watermarking accelerated after 2022, with Tree-Ring Watermarks (Wen et al., NeurIPS 2023), Stable Signature (Fernandez et al., ICCV 2023), and SynthID-Image (Google DeepMind, Nature 2024 / arXiv:2510.09263) as landmark results. The comprehensive SoK: Watermarking for AI-Generated Content survey (Piet et al., arXiv:2411.18479, 2024) provides the most complete current systematisation of the field across modalities, attack vectors, and evaluation metrics.
  • LLM text watermarking was formalised by Kirchenbauer et al. (ICML 2023). The No Free Lunch in LLM Watermarking trade-off analysis (Piet et al., 2024) demonstrates that quality, size, and tamper resistance are inherently competing objectives with no universal optimum. Cryptographic watermark verification via ZKP (Bagad et al., zkDL++, 2025) represents the frontier of verifiable provenance without key disclosure, combining techniques from Zero Knowledge Proof cryptography with learnt watermarking embeddings.

Benchmark Datasets and Evaluation

  • BOSS (Break Our Steganographic System) corpus: 10,000 greyscale images at 512×512 pixels, the standard benchmark for spatial-domain steganography and steganalysis evaluation. The BOSS corpus was used in the BOWS2 (Break Our Watermarking System 2) international competition to evaluate the security of spatial-domain embedding against statistical steganalysis using SRM features. It remains the reference set for comparing classical LSB-domain watermarking robustness against steganalysis detection.
  • Kodak Photo CD: 24 losslessly compressed photographic images at 768×512 or 512×768, selected to cover a range of scene types (indoor, outdoor, portrait, landscape). PSNR, SSIM, and LPIPS (Learned Perceptual Image Patch Similarity) watermarking imperceptibility scores are routinely reported on Kodak as the de-facto standard for classical watermarking quality evaluation, enabling direct comparison across DWT-DCT, HiDDeN, Tree-Ring, and SynthID-Image methods.
  • CLIC (Challenge on Learned Image Compression) datasets: Organised by CVPR/ICCV compression workshops, CLIC professional and mobile image sets are used to evaluate robustness of learned watermarks against neural-network-based compression (BPG, VVC, and learned codecs such as CompressAI), which constitutes the hardest known form of benign distortion for image watermarks because learned codecs can fundamentally alter statistical structure in ways that destroy frequency-domain watermarks.
  • WatermarkBench (2024): A unified multi-method benchmark for image watermarking evaluating imperceptibility (PSNR, SSIM, LPIPS), robustness across 16 distortion types (JPEG, crop, rotate, blur, colour jitter, noise, brightness, contrast, hue, saturation, and regeneration), and bit accuracy (BER), enabling direct comparison across HiDDeN, DWT-DCT, StegaStamp, Tree-Ring, and SynthID-Image. Establishes standardised evaluation protocol reducing the irreproducibility problem in watermarking evaluation caused by varied distortion parameters across papers.
  • MARKMYWORDS benchmark (2024): Comprehensive benchmark for LLM text output watermarking methods, focusing on three primary metrics: quality (measured as perplexity increase relative to unwatermarked text), size (minimum text length in tokens for reliable detection at 1% false positive rate), and tamper resistance (bit accuracy after paraphrase attack using T5-paraphraser, word substitution, and deletion). Tests Kirchenbauer green-list, SynthID-Text tournament sampling, MPAC, and Exponential Minimum sampling schemes across GPT-2, LLaMA-2-7B, and Mistral-7B base models.
  • AVI (AI-generated Video Identification) benchmarks (2025–2026): Emerging standardised benchmark suite for video watermark robustness, evaluating detection rate and payload recovery after transcoding (H.264, H.265, AV1), resolution change, frame rate conversion (24 to 30 fps), cropping, temporal clipping, and screen-capture-and-reencoding. Evaluates SynthID Video and frame-level image watermark variants across AI-generated video from Sora, Gen-2, and Lumiere.
  • Evaluation metrics summary:
    • Imperceptibility: PSNR (Peak Signal-to-Noise Ratio, dB — higher is better, typically 35–45 dB target), SSIM (Structural Similarity, 0–1, target >0.99), LPIPS (Learned Perceptual Image Patch Similarity, lower is better, target <0.01)
    • Robustness: Bit Accuracy (fraction of payload bits correctly recovered after distortion, target >0.9), True Positive Rate at 1% FPR (binary watermark detection), Area Under ROC Curve (AUC)
    • Capacity: Payload bits per image (typical range: 1 bit for binary detection, 48–256 bits for multi-bit identification)
    • Text quality: Perplexity increase (ΔPPl — lower is better, target <5%), BLEU score retention, BERTScore retention after watermarking

Current Landscape (2026)

  • Regulatory pressure as the primary growth driver: The EU AI Act Article 50, enforceable from 2 August 2026, is the single most significant regulatory event for digital watermarking in the technology industry’s history. It mandates machine-readable marking of AI-generated images, video, audio, and text in a format detectable by automated tools, with implementation guidance provided by the European Commission’s Code of Practice on AI Content Transparency (first draft 17 December 2025, finalised May-June 2026). Non-compliant AI providers face fines of up to 1.5% of global annual turnover, creating powerful economic incentives for rapid deployment of certified watermarking systems across all generative AI platforms operating in the EU market.
  • SynthID at billion-scale deployment: Google DeepMind’s SynthID had watermarked over ten billion images and video frames across Imagen, Gemini, and NotebookLM by the end of 2024, marking a watershed moment in which production-scale AI watermarking was demonstrated to be operationally feasible. The 2025 SynthID detector consumer experience enables users to upload images for watermark assessment with highlighted suspected-watermarked regions, bringing watermark detection out of the laboratory and into everyday use. Google’s commitment to open-sourcing SynthID-Text tooling for the broader LLM community lowers barriers to adoption.
  • C2PA plus embedded watermark convergence: Industry practice is converging on a two-layer architecture: C2PA cryptographic content credentials provide chain-of-custody provenance metadata (who created the content, with which tool, when, and with what edits), while an embedded watermark provides a content-bound mark surviving credential stripping. Magiclight.AI (2026) is among the platforms building end-to-end C2PA-plus-watermark pipelines explicitly targeting the EU AI Act enforcement deadline. The C2PA Technical Specification v2.0 (2024) includes normative references to embedded watermarks as complementary mechanism.
  • Adversarial robustness escalation: The most challenging attack vector remains the regeneration attack (passing a watermarked image through a separate diffusion model), which can strip many watermarks while preserving image quality. ShapeMark (arXiv:2603.09454, 2026) introduces a diversity-preserving watermarking scheme for diffusion models designed to resist regeneration attacks. The MarkDiffusion open-source toolkit (arXiv:2509.10569, 2025) lowers barriers to experimenting with and comparing generation-time watermarking approaches.
  • Provenance-watermark desynchronisation as an emerging risk: The arXiv:2603.02378 (2026) paper “Authenticated Contradictions from Desynchronized Provenance and Watermarking” identifies a real operational risk: when a content item’s C2PA metadata provenance and embedded watermark give contradictory signals (e.g., metadata says human-created, watermark says AI-generated), it can create “authenticated contradictions” that undermine trust rather than establishing it. This motivates strict operational protocols for co-embedding watermarks and C2PA credentials at the point of generation.
  • Text watermark vulnerability: The 2026 paper “Breaking Semantic-Aware Watermarks via LLM-Guided Coherence-Preserving Semantic Injection” (arXiv:2602.21593) demonstrates a practical attack on semantic text watermarking schemes, highlighting that the adversarial arms race in text watermarking remains active. This parallels the adversarial examples problem in image classification, with both communities converging on certified robustness approaches that provide provable guarantees rather than empirical robustness under known attacks.
  • Adoption in non-AI content: The technology is also seeing adoption for non-AI content provenance: camera manufacturers Canon, Nikon, and Sony are embedding C2PA content credentials (with optional pixel-level watermarks) at the point of camera capture, enabling photographers to establish authentic photographic origin in an era of ubiquitous AI image generation. This broadens the watermarking ecosystem beyond AI-generated content to encompass verified authentic content.

UK Context

  • Academic centres: The Alan Turing Institute (London) and the Centre for Data Ethics and Innovation (CDEI, now part of DSIT) have published policy analyses of AI provenance, watermarking mandates, and the EU AI Act’s implications for UK-based AI providers post-Brexit. Queen Mary University of London’s Centre for Digital Music has active research in audio watermarking and music provenance. BBC Research and Development’s Media Integrity and Provenance team has longstanding expertise in broadcast watermarking for rights management of iPlayer content and has contributed to C2PA Working Group standards. University College London’s Intelligent Systems Lab has published on adversarial robustness in DNN-based watermarking. Edinburgh’s School of Informatics has research interests in generative model provenance and content authentication.
  • Regulatory alignment: The UK has taken a principles-based, innovation-first approach to AI regulation under the AI White Paper (2023) and subsequent DSIT frameworks, diverging from the EU’s mandatory watermarking requirements under Article 50. As of mid-2026, the UK has not enacted equivalent mandatory AI content labelling legislation, though government consultations in 2025 explicitly considered this. Ofcom’s 2025/26 strategic approach to AI regulation notes that regulated sectors may voluntarily adopt C2PA content credentials (LinkedIn, for example, uses them to label AI-generated images) but imposes no requirement. The UK AI Safety Institute evaluates watermarking technologies as part of its AI evaluation toolkit.
  • Industry deployment: BBC R&D has developed and deployed broadcast watermarking in collaboration with commercial vendors for iPlayer content rights management, with watermarks enabling detection of authorised versus unauthorised redistribution of BBC content on illegal streaming platforms. ITV Studios and Channel 4 use Kantar Media NexGuard watermarking for pre-release screener tracking. ARM Holdings (Cambridge) develops inference silicon — Cortex-A and Ethos NPU families — increasingly used in on-device watermark embedding pipelines for mobile AI applications. The BBC, ITV, and Channel 4 are additionally involved in the GLARE (Global Alliance for Responsible Media Enhanced) watermarking interoperability initiative.
  • Northern England context: The University of Manchester’s Department of Computer Science has research groups working at the intersection of media forensics, Deep Learning, and content authentication. The Manchester-based UKRI Centre for Doctoral Training in AI for Media (AI4Media CDT) works on synthetic media detection and watermarking as part of the UK’s AI for Creative Industries research programme. The University of Leeds School of Computing has research in digital provenance and PACS security relevant to NHS digital pathology. Newcastle University’s Digital Humanities group researches AI content attribution. Leeds Teaching Hospitals NHS Trust and Sheffield Teaching Hospitals NHS Foundation Trust have active clinical informatics projects on DICOM image integrity and tamper-evident watermarking for digital pathology chain-of-custody.

Key Terminology

  • Watermark encoder: The embedding module that maps (cover content, watermark payload) → watermarked content, implementing the chosen embedding paradigm (DCT, DNN, latent diffusion, token biasing).
  • Watermark detector: The verification module that maps watermarked (or potentially unwatermarked) content → detection decision (watermark present/absent) and optional payload recovery. May be a statistical test (Kirchenbauer z-test), a trained DNN decoder, or a correlation-based matcher.
  • Cover medium: The host content into which the watermark is embedded — image, audio file, video sequence, text document, or AI model activation.
  • Payload / message: The information embedded by the watermark, ranging from a single detection bit (robust vs. fragile) to a 256-bit identifier encoding origin system, timestamp, and user ID.
  • Imperceptibility (IQA): Image Quality Assessment of the watermarked output relative to the unwatermarked original, measured by PSNR, SSIM, or LPIPS. The goal is that the watermarked and unwatermarked versions are perceptually indistinguishable to human observers.
  • Robustness: The ability of the watermark to survive common image processing operations (JPEG compression, resizing, rotation, brightness/contrast adjustment) and adversarial removal attacks (regeneration, diffusion-based removal, neural erasure).
  • Regeneration attack: The hardest known watermark removal attack: running a watermarked image through a separate generative model (typically a diffusion model) to produce a new image with similar visual content but without the watermark. Tree-Ring and SynthID-Image are the most robust current methods against this attack.
  • Fragile watermark: A watermark whose integrity breaks on any content modification, serving as a tamper alarm rather than a persistent identifier.
  • Semi-fragile watermark: Survives benign operations (JPEG compression, minor brightness adjustment) but breaks on content-modifying operations (face swap, object insertion, geometric distortion beyond a threshold).
  • Steganalysis: The adversarial discipline of detecting the presence of hidden information in a carrier medium, regardless of whether the payload can be recovered. Statistical steganalysis applies feature-based classifiers (SRM, CNN-based) to identify statistical anomalies introduced by embedding.

Future Directions (2026–2030)

  • On-device watermark embedding at inference time: The AI inference stack is migrating rapidly from cloud data centres to on-device deployment on ARM Cortex and Apple Neural Engine hardware. Watermark embedding at model inference time (SynthID-style) must be implemented with negligible latency overhead on NPU hardware, motivating research in hardware-efficient watermarking architectures that exploit fixed-point arithmetic and sparse computation patterns native to neural processing units. Target latency budget is under 1 ms additional inference overhead on Ethos-U65 class NPUs.
  • Universal watermark standards (ISO/IEC WG4 and JPEG AI): ISO/IEC JTC1/SC29/WG4 (JPEG) and JPEG AI are developing standardised watermark payload formats, detection APIs, and interoperability requirements that allow cross-vendor watermark verification: a detector from vendor B should be able to verify a watermark embedded by vendor A’s system. This is a critical requirement for multi-platform provenance pipelines where content passes through multiple platforms before reaching the final consumer. The JPEG Trust initiative (JPEG 9 WG4) explicitly addresses this cross-vendor detection interoperability requirement.
  • Verifiable provenance via ZKP integration: Zero-knowledge proof-based watermark verification (zkDL++) will mature from research prototype to production-scale deployment, enabling content creators to publish verifiable provenance proofs without disclosing proprietary watermarking keys. ZKP-based proofs will be integrated into blockchain-anchored provenance ledgers for archival media, scientific data, and legal documentation. The integration of ZKP watermark proofs with W3C Verifiable Credentials standards is expected to enable interoperable provenance claims across platform boundaries.
  • Multi-modal coherent watermarking: As AI systems generate coherent multimedia packages (video + audio + captions + structured metadata), watermarks must be co-embedded across modalities with joint detectability and cross-modal consistency checks. Stripping a watermark from the audio track while leaving it in the video creates detectable cross-modal inconsistency; adversarial systems are expected to target such asymmetries, motivating joint multi-modal embedding schemes. Research in 2025 on temporally coherent video-audio watermarks is a precursor to fully joint multi-modal approaches.
  • Adversarial hardening via game-theoretic co-training: The arms race between watermarking and removal/forgery attacks is expected to converge on adversarial co-training frameworks analogous to Generative Adversarial Networks, where a watermark encoder-detector pair is explicitly hardened against a simultaneously trained watermark remover network, yielding encoder-decoder designs certified robust against the strongest known removal attacks. This adversarial watermarking training framework is sometimes called “watermark GAN” or W-GAN in the literature, distinct from Wasserstein GANs.
  • Quantum-safe watermark key cryptography: The cryptographic keys underpinning watermark payload authentication will require migration to post-quantum key exchange and signature schemes (NIST PQC standards: CRYSTALS-Kyber for key encapsulation, CRYSTALS-Dilithium for digital signatures, SPHINCS+ as a stateless hash-based alternative, all finalised by NIST in 2024) to ensure long-term archival provenance claims remain cryptographically secure against quantum-capable adversaries targeting historical media archives with retroactive decryption attacks.
  • Regulatory harmonisation and mutual recognition: Post-2026 regulatory dialogue between the EU (AI Act Article 50), UK (AI Safety Institute frameworks and potential Digital Information legislation), and US (NIST AI Risk Management Framework, White House Executive Order 14110 AI content labelling provisions) is expected to produce mutually recognised watermark standards and assessment frameworks enabling a single watermarking implementation to satisfy multiple regulatory regimes simultaneously, reducing compliance fragmentation for multinational AI providers operating across all three jurisdictions.
  • Foundation model watermarking: As the AI industry consolidates around a small number of large foundation models fine-tuned for specific applications, watermarking research is shifting from content watermarking (marking outputs) to model watermarking (marking weights). Model-level watermarks that propagate to all fine-tuned derivatives would enable the foundation model provider to trace all downstream outputs to their model, creating a comprehensive model-lineage provenance system compatible with content watermarking at the output level.

Research and Literature

    1. Cox, I.J., Kilian, J., Leighton, T., Shamoon, T. (1997). Secure spread spectrum watermarking for multimedia. IEEE Transactions on Image Processing, 6(12), 1673–1687. Foundational spread-spectrum framework establishing DCT-domain embedding theory.
    1. Bender, W., Gruhl, D., Morimoto, N., Lu, A. (1996). Techniques for data hiding. IBM Systems Journal, 35(3–4), 313–336. First formal treatment of digital data hiding techniques; established the field’s terminology.
    1. Cox, I.J., Miller, M., Bloom, J., Fridrich, J., Kalker, T. (2007). Digital Watermarking and Steganography, 2nd edition. Morgan Kaufmann. Definitive textbook covering classical and DNN approaches.
    1. Petitcolas, F.A.P., Anderson, R.J., Kuhn, M.G. (1999). Information hiding — a survey. Proceedings of the IEEE, 87(7), 1062–1078. Canonical taxonomy distinguishing watermarking, steganography, and fingerprinting.
    1. Cachin, C. (1998). An information-theoretic model for steganography. Proc. 2nd Workshop Information Hiding, LNCS 1525, 306–318. Capacity-theoretic foundation for information hiding.
    1. Zhu, J., Kaplan, R., Johnson, J., Fei-Fei, L. (2018). HiDDeN: Hiding Data With Deep Networks. ECCV 2018. Foundational DNN encoder-decoder watermarking framework.
    1. Tancik, M., Mildenhall, B., Ng, R. (2020). StegaStamp: Invisible Hyperlinks in Physical Photographs. CVPR 2020. Extension of DNN watermarking to photographic print-and-scan recovery.
    1. Wen, Y., Kirchenbauer, J., Geiping, J., Goldstein, T. (2023). Tree-Ring Watermarks: Fingerprints for Diffusion Images that are Invisible and Robust. NeurIPS 2023. arXiv:2305.20030. Fourier-domain latent diffusion watermarking robust to regeneration attacks.
    1. Kirchenbauer, J., Geiping, J., Wen, Y., Katz, J., Miers, I., Goldstein, T. (2023). A Watermark for Large Language Models. ICML 2023. Foundational green-list/red-list text watermarking for LLMs.
    1. Piet, J. et al. (2024). SoK: Watermarking for AI-Generated Content. arXiv:2411.18479. Comprehensive systematisation-of-knowledge survey across all modalities and attack types.
    1. Google DeepMind (2024). SynthID-Image: Image watermarking at internet scale. arXiv:2510.09263. Production-scale diffusion watermarking covering 10+ billion images.
    1. Ci, H. et al. (2024). RingID: Rethinking Tree-Ring Watermarking for Enhanced Multi-Key Identification. Extension of Tree-Ring to multi-bit multi-key regime.
    1. Xu, Y. et al. (2025). Watermarking Language Models for Many Adaptive Users. arXiv:2405.11109. Robust multi-bit text watermark resilient to adaptive paraphrase attacks.
    1. Piet, J. et al. (2024). No Free Lunch in LLM Watermarking: Trade-offs in Watermarking Design Choices. arXiv:2402.16187. Quality-size-tamper resistance trade-off analysis for text watermarking.
    1. Bagad, P. et al. (2025). zkDL++: Zero-knowledge cryptographic watermark verification. Zero-knowledge proof-based provenance without key disclosure.
    1. Anonymous (2025). MarkDiffusion: An Open-Source Toolkit for Generative Watermarking of Latent Diffusion Models. arXiv:2509.10569. Open-source diffusion watermarking framework and benchmark.
    1. Anonymous (2026). ShapeMark: Robust and Diversity-Preserving Watermarking for Diffusion Models. arXiv:2603.09454. Regeneration-attack-resistant diversity-preserving watermarking.
    1. Anonymous (2026). Authenticated Contradictions from Desynchronized Provenance and Watermarking. arXiv:2603.02378. Operational risk analysis of C2PA-watermark desynchronisation.
    1. Anonymous (2026). Autoregressive Images Watermarking through Lexical Biasing: An Approach Resistant to Regeneration Attack. arXiv:2506.01011. Token-biasing approach applied to autoregressive image generation.
    1. Anonymous (2026). Breaking Semantic-Aware Watermarks via LLM-Guided Coherence-Preserving Semantic Injection. arXiv:2602.21593. Practical attack on semantic LLM text watermarks.
    1. Anonymous (2025). Guidance Watermarking for Diffusion Models. arXiv:2509.22126. Classifier-guidance-based watermark embedding during diffusion.
    1. European Commission (2025). Code of Practice on the Transparency of AI-Generated Content. First draft 17 December 2025; finalised May-June 2026. Sets Article 50 implementation framework.
    1. C2PA Coalition for Content Provenance and Authenticity (2024). C2PA Technical Specification v2.0. Standard for content credentials compatible with embedded watermarks.
    1. Mareen, H. et al. (2023). A Unified Frequency Domain-Based Watermarking Method for Deep Neural Networks. IEEE Transactions on Information Forensics and Security. Unification of frequency-domain DNN watermarking.
    1. Fernandez, P. et al. (2023). The Stable Signature: Rooting Watermarks in Latent Diffusion Models. ICCV 2023. Meta AI latent-decoder fine-tuning watermarking.
    1. Zhao, X. et al. (2023). Invisible Image Watermarks Are Provably Removable Using Spectral Methods. NeurIPS 2023. Vulnerability analysis motivating regeneration-robust designs.
    1. Lukas, N. et al. (2023). Leveraging Optimization for Adaptive Attacks on Image Watermarks. ICLR 2024. Adaptive attack framework evaluating worst-case watermark robustness.
    1. MDPI Mathematics (2025). Digital Watermarking Technology for AI-Generated Images: A Survey. doi:10.3390/math13040651. Recent comprehensive survey of AIGC-specific watermarking methods.

Provenance