Claude watermarks all chats

The truth about Claude watermarks

Kyle explains Claude's hidden watermarks during the livestream.

Full video here (20k views and rising! It’s a hot one!)

Claude is putting an invisible watermark inside the text it generates.

And the internet is flipping out.

With…l’ll be honest…some reason!

People are getting very upset about this and (sorry) letting their emotions get in the way of, you know, the actual truth of the matter.

I’m here to give you the skinny. What’s ACTUALLY happening. This is already important! We don’t need the additional made up stuff.

The facts: Anthropic says new Claude models launched in the EU on or after 2 August support marking from day one. Older models are being updated. The detector and technical documentation are still coming.

We’ve been living with water marked models for 2 weeks or so already.

So this is a real change. It just needs the caveats left attached so you do not sound like a dumbass when talking about it….

One mark travels in the words. One travels with the file.

Anthropic is adding two different marks.

Claude uses an embedded text watermark and signed C2PA metadata for supported files.

The first is an invisible watermark embedded directly into generated text. Anthropic says it moves with the words when you copy and paste them and may survive some editing.

The second is signed C2PA provenance metadata on supported SVG, PNG and JPG files. That travels with the file, but normal exports, resaving and screenshots can strip metadata.

The text version is the interesting one. It is applied at model level, so the same supported model can mark output from Claude, the API, Claude Code, Cowork and supported cloud partners.

Nobody outside Anthropic knows the exact method yet. Although that doesn’t stop people from spouting off!

It’s is highly likely (probable, lol) that it’s a statistical watermark. Basically the statistical distribution of tokens are shifted slightly in ways that make the text identifiable as Claude generated.

This is not as crude as just inserting certain words. It’s MUCH more subtle - small nudges and tweaks across tokens that will be largely imperceptible to human readers.

 Alex Cui has a very good technical explanation of how statistical text watermarks can reweight the next word a model chooses. Highly recommend giving it a read.

Google has publicly documented something similar for SynthID Text. It does not mean Claude uses Google's method.

If they are though this is important: no-one has reliably broken SynthID since 2024. So people saying “lol I’ll just remove it” are talking nonsense.

The EU change is going global

You may be thinking “well I’m not in the UK this doesn’t make any difference to me”.

In fact many people have been telling me that on Youtube.

Wrong.

Sorry.

On Monday I broke down the four actual Article 50 transparency rules. This is the product consequence. Worth a re-cap to get the foundations of what is happening.

This is a GLOBAL rollout. Claude are adapting ALL of their models GLOBALLY.

Anthropic could make one Claude for the EU and another for everybody else. Or it could add the mark once and ship the same models worldwide.

It chose worldwide.

The EU rule changed Claude worldwide because Anthropic is shipping one product behaviour.

We have seen this before. The EU required USB-C, and Apple eventually replaced Lightning rather than manufacturing a separate European iPhone.

Money talks here. Not falling in line means i) losing access to the EU market and ii) fines.

Far easier just to play ball.

It’s a bit like having a dinner party with a whole bunch of people with different dietary restrictions and needing to make ONE dish for all. How do you deal with this? You make a dish that the fussiest person will eat. Everyone else will just put up with it.

In this case the EU is the fussy eater.

Do remember that this is also not just Anthropic. They didn’t just wake up one day and decide to annoy their users. Google has used SynthID for years. OpenAI also use watermarks in images and audio and says it plans to extend provenance signals to text (probably before December when the fines really kick in). Meta, Microsoft and Mistral have also signed the provider section of the EU code.

All of the Western labs are falling in line. Except xAI. Because, well, Elon Musk.

Claude is simply the current punching bag because Anthropic explained its rollout first. They stuck their head above the parapet and promptly had their head blown off. Can’t catch a break.

A mark does not mean Claude wrote it

Super important here because it’s going to ruin some people’s lives. Employers, schools and clients are going to misuse this.

A detected Claude mark signals processing, not original authorship.

Know this: you can write an entire document yourself, ask Claude to proofread it and BOOM its watermarked. The same can happen when you translate your own article, summarise your meeting notes or convert a file. Watermarked.

Claude touched it. Claude did not necessarily author it. Doesn’t matter - it’s flag with the watermark.

Anthropic's own wording is that a detected mark means the content may have been processed by Claude. It cannot tell you who had the idea, how much a human wrote, whether the facts are correct or whether anybody cheated.

This is a problem.

ANY touch will create a positive result when tested. Which schools, universities, employers etc. are going to assume means all AI.

Instead treat the mark like a fingerprint showing contact with the tool. It is not an authorship certificate.

No mark does not mean human either

The opposite conclusion is just as dodgy. False negatives.

No detected Claude mark does not prove that a document was written by a human.

A detector may find nothing because the text came from an older Claude model, the passage is too short, somebody heavily rewrote it, the file metadata disappeared or another AI made it. Perfectly innocuous reasons.

OR

The people who deliberately want to cheat have the strongest reason to strip the mark.

The honest person who used Claude for a grammar pass may get flagged. The dishonest person who knows how to work around the detector may sail through.

They now have an even stronger incentive to cheat and strip the watermark (if possible). Because they now live in a world where everything else flags as AI. Their marginal gain from cheating the system just rose. Cheaters prosper.

All this said we do NOT know if it’ll be possible to strip the watermark. Anthropic has not published the model coverage, minimum passage length, confidence thresholds, false-positive rates, false-negative rates or even the detector launch date yet.

So anybody claiming they can give you a definitive Claude verdict today is talking bullshit. I’m not saying we can or we can’t. We just have no information.

Keep your own evidence

So…creators. And business owners. Do not rely on a watermark to explain how you made something. If you think you’ll need to one day prove you made something start tracking provenance now.

A practical provenance workflow for Claude-assisted work.

Record the model and product you used. Keep your original notes and source files. Save the human edits and approval trail. Test what happens when files pass through your normal export workflow.

A shared AI vault gives you somewhere to keep that history. This is also why Anthropic paid millions for clean human data and then destroyed millions of physical books: knowing where content came from is becoming extremely valuable.

To the Task,

Kyle