Term
AI Watermarking
A method AI providers use to mark generated content so machines can recognise it. In text this is an invisible statistical pattern in word choice; in files it is signed provenance metadata. The goal is evidence that a model was involved in producing the content.
AI watermarking — explained in more detail
An AI watermark is a mark a provider embeds into its model’s output so that it can later be verified by machine whether a piece of content came from that AI. Unlike a visible notice (created with AI), the mark is invisible to readers and survives ordinary copy and paste.
Two families of techniques exist in practice, and they differ fundamentally:
1. Statistical text watermarking. At every word a language model has several near-equivalent continuations to choose from. When watermarking, that choice is not made neutrally but nudged according to a secret key. A single word reveals nothing — but across a sufficiently long text those choices form a statistical pattern that a detector holding the key can recognise. The best-known published approach of this kind is SynthID-Text by Google DeepMind (2024); Anthropic’s implementation is related to it.
2. Signed provenance metadata. Here the content itself is untouched. Instead, a cryptographically signed record is attached to the generated file, recording which model produced it. The industry standard is C2PA / Content Credentials. The upside: unambiguously verifiable. The downside: metadata is lost on copying, conversion or screenshotting.
Anthropic uses both. Text marking applies to models released on or after 2 August 2026 — at launch that means Claude Fable 5.1 and Claude Mythos 5.1. Older models such as Claude Opus 5 or Sonnet 5 are not marked retroactively.
A watermark is evidence, not proof
Statistical text watermarks produce a probability, not a yes/no answer. They need a minimum length to register, lose signal under rewriting and translation, and cannot rule out that a human edited the text. Conversely, a missing watermark does not mean a text is human-written — it may simply come from a model that does not mark its output.
Example / practical relevance
An editorial team drafts a specialist article with Claude Fable 5.1 and publishes it after human revision. The raw draft carried a watermark; after heavy rewriting the signal is weakened or gone, depending on how deep the edits went. A second team has the same model write a code library — the code itself is largely unmarked, while the docstrings and documentation are ordinary prose and can carry the pattern.
Not everyone can check for the watermark today: detection runs through a provider interface that is initially available as a private preview for selected organisations. For website owners the practical takeaway is that you cannot verify third-party texts yourself, but should assume platforms and auditors will be able to.
Distinguishing it from similar terms
AI detectors (tools that estimate an AI probability) guess from stylistic features and have high error rates — a watermark instead checks an embedded signal against a known key. Visible labelling under the EU AI Act is a legal disclosure duty towards humans, whereas a watermark is a machine-readable mark; Article 50 of the AI Act requires both in different situations. Digital signatures prove a file is unaltered but say nothing about AI involvement.
Discover more
AI Workflows by Keyword: How We Make Recurring Routines Enforceable
A typed keyword triggers a fixed AI routine — and every single step must be committed before the next one appears. Why that's the actual trick.
GlossaryAttention Mechanism
Computational method in transformer models that weights which parts of the input text are most relevant for predicting the next token — instead of treating every part as equally important.
EncyclopediaLLM Hallucinations — Causes and Remedies
Why LLMs confidently invent falsehoods, what types of hallucinations exist, and which remedies actually help — RAG, source enforcement, verification.