Anthropic Starts Watermarking: New Claude Models Mark Their Text
Anthropic now marks the output of its new models. Every Claude model released on or after 2 August 2026 weaves an invisible watermark into generated text and attaches signed provenance metadata to generated files. First up are Claude Fable 5.1 and Claude Mythos 5.1, both released on 1 September 2026. Older models such as Claude Opus 5 or Sonnet 5 are not marked retroactively.
What Anthropic shipped
- Cut-off 2 August 2026 — models from that date mark their output at launch.
- Affected so far: Claude Fable 5.1 and Claude Mythos 5.1.
- Two mechanisms: a statistical watermark in the text and signed provenance metadata on files.
- The text method is related to SynthID-Text (Google DeepMind, 2024): word choice is nudged slightly according to a secret key.
- Code stays largely unmarked — comments, docstrings and documentation are prose and do carry the pattern.
- The detection API initially runs as a private preview for selected organisations.
What applied before
Until the summer of 2026 there was practically no reliable provenance check for AI text. Images and video had been marked by several providers for a while — Google with SynthID, OpenAI and Microsoft via C2PA metadata in their image products. Text was the gap: anyone wanting to know whether an article came from a model had to rely on statistical AI detectors that guess from stylistic features and are known for high error rates.
The reason for the gap was economic rather than technical. Text watermarking methods had existed for years; Google DeepMind published SynthID-Text in 2024. Almost nobody rolled them out — the fear of losing users to non-marking competitors was a real counter-argument in a market of interchangeable models.
What applies now
1. The text itself carries the mark. At every position a language model has several equivalent continuations to choose from. When watermarking, that choice is nudged according to a secret key. A single word reveals nothing; across a sufficiently long text a statistical pattern emerges that a detector holding the key can find. The method is invisible and survives copy-paste.
2. Files additionally get signed metadata. The second mechanism leaves the content untouched and instead attaches a cryptographically signed record to the generated file. That is unambiguously verifiable but lost on copying, conversion or screenshotting — the two methods cover each other’s gaps.
3. Almost nobody can verify it yet. A statistical watermark is only checkable with the matching key, and the key sits with the provider. Anthropic offers a detection interface that initially runs as a private preview for selected organisations, particularly those with corresponding verification duties. There is no complete public overview of which model marks to what extent.
4. Code is the big gap. Prose offers many equivalent phrasings whose selection can be nudged. Code does not — syntax cannot be phrased differently without changing the program. The very discipline Fable is marketed hardest for therefore stays largely unmarked. The accompanying documentation, by contrast, is ordinary prose and does carry the pattern.
The regulatory background
The cut-off date is no coincidence. Article 50 of the EU AI Act requires providers of AI systems to mark synthetic content in a machine-readable format. That is exactly what a watermark does. Anthropic is meeting an obligation here, not imposing a voluntary restraint — and 2 August 2026 coincides with the date those transparency obligations became applicable.
Important context: machine-readable marking is a provider duty. Anyone merely using AI systems is a deployer and has different obligations — such as disclosure to humans for deepfakes or for text on matters of public interest.
How the others position themselves
Anthropic is in small company on text. The dividing line runs between media and text: almost everyone marks images and video, almost nobody marks text.
| Provider | Text watermark | Media / metadata | |---|---|---| | Anthropic (Claude) | yes, for models from 2 Aug 2026 | signed provenance on files | | Google (Gemini) | yes, SynthID-Text | SynthID for image, audio, video | | OpenAI (GPT, Sora) | no publicly deployed method | C2PA Content Credentials | | Meta (Llama, Meta AI) | no | visible AI labels + image metadata | | Microsoft | no | C2PA in its own image products | | Mistral, xAI and others | no public commitment known | inconsistent to none |
As of September 2026.
Assessment
The obvious worry — “Google will now spot my AI text and penalise me” — does not hold, for two independent reasons. Google does not have the key: Anthropic’s watermark is only verifiable with Anthropic’s detector. And Google assesses quality and usefulness anyway, not the production method; what gets penalised is volume without value, and that is visible without any watermark.
What does change is the evidentiary situation. Until now, “we wrote this ourselves” was a claim practically nobody could check. It now becomes technically checkable — for a few today, and probably for clients, platforms and auditors in the medium term. Agencies and editorial teams that blanket-promise “100% human-written” should revisit their wording before someone else does.
At the same time the mark is weaker than the headlines suggest. It needs length, it yields a probability rather than proof, it loses signal under heavy rewriting, and it barely registers on short outputs such as meta descriptions or subject lines. A marked draft that has been edited and enriched with your own experience remains a perfectly legitimate work product — how that runs cleanly is covered in the reference piece Watermarks in AI Text.