Key takeaways
- Anthropic is adding invisible watermarks to Claude-generated text, but the mark only shows that Claude processed the text, not that Claude wrote it.
- Anthropic acknowledges that substantial rewriting and paraphrasing can make the watermark undetectable.
- That means people trying to remove the signal can often do so, while ordinary users may unknowingly carry it after using Claude for editing or proofreading.
- WriteHuman’s position is simple: people should control their final writing, and invisible provenance signals should not be treated as proof of cheating or authorship.
This week, Anthropic announced they will start stamping an invisible watermark into everything Claude writes. It's part of complying with the EU's AI Act, and on paper it sounds reasonable: mark AI text, so people can tell what's AI-generated. I run a company that removes exactly this kind of signal, so I've spent more time than most with the technology behind it. The short version is that this watermark won't do what most people think it does, and the people it does affect are mostly the ones who least deserve it.
How the watermark actually works
From what Anthropic has described, and from what one of the leads at Claude confirmed on X, the method is close to Google's SynthID. It's the same basic idea a group of researchers at the University of Maryland introduced back in 2023. As the model writes, it quietly nudges its word choices toward a secret "green list" of tokens. No single word looks unusual, but across a few hundred words the pattern becomes statistically detectable if you hold the key. SynthID does the same thing by tweaking the probabilities during generation. Anthropic bakes it in token by token as the text streams out.
It's smart in theory, but it's also fragile.
Why paraphrasing breaks it
The watermark lives in the exact sequence of words the model picked. Change enough of those words and the signal falls apart. This isn't a bug anyone can patch. There's a 2024 paper out of Harvard and Boston University, "Watermarks in the Sand," that proves it: under normal assumptions, you can't build a text watermark that survives a determined rewrite without also wrecking the quality of the text. Case in point: Gemini watermarks their text. Who do you know that uses Gemini to help them write?
The empirical results back up the math. A single paraphrasing pass drops detection on these methods sharply. Free tools already strip Google's SynthID. We've been doing this at WriteHuman for years, and none of the recent watermarking work changes the underlying picture. If someone wants the mark gone, it's gone.
Here's what bothers me
If the watermark is trivial to remove, who's left carrying it?
The person who didn't know it was there.
Think of a student who writes English as a second language. She wrote her essay herself, then asked Claude to fix her grammar and tighten a few sentences. She pastes the cleaned-up text into her document and hands it in. She has no idea anything was embedded, no reason to run it through a paraphraser, no intent to deceive anyone. Her paper now carries the same mark as an essay that was generated wholesale by someone who never wrote a word of it.
The watermark can't tell those two people apart. Anthropic has said as much: the mark shows that Claude touched the text, not that Claude wrote it. Proofreading and ghost-writing look identical to it.
That's backwards. The tool lets everyone acting in bad faith walk right through, and quietly tags the people who were only trying to write more clearly.
We already know how this story goes
This isn't hypothetical, because we've watched AI detectors make the same mistake for over two years. A Stanford team tested seven detectors on essays written by non-native English speakers and found they falsely flagged more than 61% of them as AI. On writing by native speakers, the same tools were nearly perfect. The reason: careful, plain, correct writing reads as "too predictable," and predictability is what these systems treat as AI-generated.
It has done real damage. Students have been pulled into misconduct hearings over false flags. Enough universities decided these tools couldn't be trusted that many switched them off. And plenty of our customers now run their own, completely original writing through humanizers for one reason only: so they won't be accused. Not to cheat. To defend themselves.
Watermarking fixes none of this. If anything it adds to the pile: one more invisible signal that gets read as guilt no matter what actually happened.
What we actually believe
WriteHuman has a stake in this. We help people control whether their writing reads as AI generated, so discount what I'm saying however you think is fair.
But the principle we care about doesn't depend on our business. People should own their words. Someone who used a tool to help them communicate is not a cheater, and a system that can't tell the difference shouldn't get to brand them one. Transparency that only catches the honest isn't transparency. It's theater, and the bill for it lands on the people with the least standing to argue back.
Anthropic's watermark will be removed by anyone who cares to remove it. Everyone else will carry a mark they never knew was there, judged by tools that were already getting this wrong. That's the part of this story worth paying attention to.



