Watermarks in AI Text: What Actually Gets Marked
Watermarks in AI Text: What Actually Gets Marked
Something shifted in the summer of 2026 that almost nobody noticed in day-to-day work with AI: new language models no longer hand over their text neutrally, but with a built-in provenance mark. Anthropic marks every model released on or after 2 August 2026 — starting with Claude Fable 5.1 and Claude Mythos 5.1.
That sounds like a footnote for lawyers. It isn’t. Anyone producing content at scale should understand what is actually being marked, how reliable it is and — above all — what it does not do. Because two things are currently being overestimated: the control providers gain from it, and the risk publishers run because of it.
Short version
- There are two mechanisms: a statistical watermark in the text and signed provenance metadata on files.
- Anthropic uses both, for models released on or after 2 August 2026 (Fable 5.1, Mythos 5.1). Older models such as Opus 5 are not marked retroactively.
- The text watermark needs length to register and loses signal under heavy rewriting.
- Code is largely unmarked — comments, docstrings and documentation, being ordinary prose, can carry the pattern.
- Almost nobody can check it right now: detection is a private preview for selected organisations.
- It is not a direct Google ranking factor. It does matter for compliance and trust.
How a text watermark works technically
A language model produces text word by word. At each position it has not one correct continuation but a probability distribution over many — and several are often practically equivalent. “The result was surprising” and “The result was unexpected” are the same thing to a reader.
That freedom is exactly what the watermark exploits. Instead of sampling neutrally, the method nudges the choice slightly in a particular direction according to a secret key. A single word gives nothing away — the nudge sits inside the noise. But across hundreds of words a statistical pattern emerges that a detector holding the same key can find. The best-known published approach of this kind is SynthID-Text, presented by Google DeepMind in 2024; Anthropic’s method is related to it.
Three properties follow directly, and they matter more in practice than the technique itself:
It needs length. A two-liner contains too few decisions for a pattern to be distinguishable from chance. The longer the text, the more confident the verdict.
It is probabilistic. The detector does not say “made by AI” but “this pattern is so unlikely that model involvement is plausible”. That is evidence, not proof.
It does not survive heavy edits. Substantially rewriting, translating or radically shortening a text destroys part of the marked word choice. Light editing usually leaves the watermark intact; a genuine rewrite does not.
The second mechanism: signed provenance metadata
The second route works entirely differently and is often confused with the first. Here the content stays untouched. Instead, a cryptographically signed record is attached to the generated file, recording who produced it and when. The industry standard is C2PA, or Content Credentials, backed by a broad alliance of media, camera and software companies.
The practical difference is sharp:
| | Text watermark | Signed metadata (C2PA) | |---|---|---| | Lives in | the words themselves | an attachment on the file | | Survives copy-paste | yes | no | | Survives screenshots | yes (as text) | no | | Survives rewriting | partially | yes (unchanged) | | Strength of claim | probability | unambiguous, cryptographic | | Removable | hard but possible | trivial (delete metadata) |
Each method has its own hole — which is precisely why they complement each other.
Where watermarks fail
The most interesting gap concerns the very discipline these models are most used for: code is largely unmarked. The reason is straightforward. Prose offers many equivalent phrasings whose selection can be nudged. Code does not — a variable is named what it must be named, and syntax cannot be “phrased differently” without changing the program. A watermark would need room that simply isn’t there.
The side effect: comments, docstrings and accompanying documentation are ordinary prose — and do carry the pattern. An AI-generated README can be marked while the function beneath it is not.
Other holes worth knowing:
- Short outputs — subject lines, meta descriptions, product titles, social snippets: too short for a reliable signal.
- Translation — a text run through another system carries that system’s word choice, not the marked one.
- Model mixing — drafting with a marking model and finalising with a non-marking one leaves no usable signal.
- Older models — anything before the cut-off, including strong models like Claude Opus 5, does not mark.
Who can actually detect the watermark?
This is where public excitement parts ways with reality. A statistical watermark can only be checked with the matching key — and the key sits with the provider. Anthropic offers verification through a detection interface that initially runs as a private preview for selected organisations, in particular bodies with corresponding verification duties. There is no complete public overview of which model marks to what extent.
For you as a website owner that means two concrete things. First: you cannot check third-party texts yourself — anyone selling you a tool that “detects Claude watermarks” is selling you a classic AI detector with the well-known error rates. Second: you should assume platforms, clients and auditors will gain this capability in the medium term.
How the other providers position themselves
Anthropic is not alone, but the landscape is patchier than the headlines suggest. The decisive split runs between text and media: images and video are marked by almost everyone, text by almost nobody.
| Provider | Text watermark | Media / metadata | |---|---|---| | Anthropic (Claude) | yes, for models from 2 Aug 2026 | signed provenance on files | | Google (Gemini) | yes, SynthID-Text | SynthID for image, audio, video | | OpenAI (GPT, Sora) | no publicly deployed method | C2PA Content Credentials in image/video | | Meta (Llama, Meta AI) | no | visible AI labels + image metadata | | Microsoft | no | C2PA in its own image products | | Mistral, xAI and others | no public commitment known | inconsistent to none |
As of September 2026. This moves fast — check the current state with the provider before relying on it.
That Google and Anthropic lead on text is no accident: both either did the underlying research or adapted it. OpenAI demonstrably developed a text method but has not rolled it out — the fear of losing users to non-marking competitors is a real counter-argument in this market.
The legal background: EU AI Act, Article 50
The 2 August 2026 cut-off is no coincidence. Article 50 of the AI Act requires providers of AI systems to mark synthetic content in a machine-readable format so it is detectable as artificially generated. That is exactly what a watermark is. Providers are not acting out of idealism; they are meeting an obligation.
Important for you: the regulation distinguishes between providers of AI systems and deployers, meaning those who use them. Machine-readable marking is a provider duty — you need do nothing. Deployer duties apply in other situations, such as deepfakes or text on matters of public interest, where disclosure to humans is required.
Not legal advice
This overview is carefully researched but is not legal advice. Whether and how Article 50 applies to your specific publication depends on the individual case — with meaningful volumes of AI content, a legal review is worth it.
What this means for SEO and content production
The obvious worry: “Will Google now find out my text came from an AI and penalise me?” The answer is a clear no — but for several reasons at once, and each is worth stating.
Google does not hold the key. Anthropic’s watermark is only verifiable with Anthropic’s detector. Google cannot simply read other providers’ watermarks.
Google does not penalise AI content per se anyway. The Search Central position has been unchanged for years: quality and usefulness are assessed, not the production method. What gets penalised is volume without value — and that needs no watermark; you can see it in the text.
Something does change nonetheless. Proof becomes technically possible, and that shifts what clients, platforms and auditors expect. Claiming today that you deliver exclusively hand-written content while the draft is marked carries a risk that did not exist before 2026.
In practice:
- Don’t hide anything. Deliberately stripping a watermark is high effort for a poor return — and exactly the thing that looks bad later.
- Equally, don’t over-dramatise. A marked draft that has been edited, fact-checked and enriched with your own experience is a legitimate work product. That is precisely the point of a clean AI content workflow.
- Review your contracts and claims. If you promise clients “100% human-written”, that is now a checkable assertion. Describe honestly what your process does.
- Choose models deliberately. Not to evade marking, but because you should know which of your pipelines produce marked output and which do not.
- Don’t forget documentation. In code projects it is the READMEs and docstrings that carry the signal — not the code.
Conclusion
Watermarks in AI text are neither the loss of control they are debated as, nor a formality. They are solid but incomplete evidence, running quietly in new Anthropic models since 2 August 2026, with verification open to only a few for now. For rankings, nothing changes. For what you can honestly claim about your own production process, quite a lot does.
The sensible approach is the same as before, minus the excuses: AI for speed and structure, humans for truth, voice and accountability — and an honest description of that to the outside world.