New Sept 3: Enhanced model updated for GPTZero and Pangram
AI Detection

Claude's Watermark Punishes the Wrong People

Anthropic's new text watermark is easy to remove and easy to miss. The writers most likely to get flagged are the ones who did nothing wrong.

Ivan JacksonIvan Jackson4 min read
Claude's Watermark Punishes the Wrong People

Key takeaways

  • Anthropic is adding invisible watermarks to Claude-generated text, but the mark only shows that Claude processed the text, not that Claude wrote it.
  • Anthropic acknowledges that substantial rewriting and paraphrasing can make the watermark undetectable.
  • That means people trying to remove the signal can often do so, while ordinary users may unknowingly carry it after using Claude for editing or proofreading.
  • WriteHuman’s position is simple: people should control their final writing, and invisible provenance signals should not be treated as proof of cheating or authorship.

This week, Anthropic announced they will start stamping an invisible watermark into everything Claude writes. It's part of complying with the EU's AI Act, and on paper it sounds reasonable: mark AI text, so people can tell what's AI-generated. I run a company that removes exactly this kind of signal, so I've spent more time than most with the technology behind it. The short version is that this watermark won't do what most people think it does, and the people it does affect are mostly the ones who least deserve it.

How the watermark actually works

From what Anthropic has described, and from what one of the leads at Claude confirmed on X, the method is close to Google's SynthID. It's the same basic idea a group of researchers at the University of Maryland introduced back in 2023. As the model writes, it quietly nudges its word choices toward a secret "green list" of tokens. No single word looks unusual, but across a few hundred words the pattern becomes statistically detectable if you hold the key. SynthID does the same thing by tweaking the probabilities during generation. Anthropic bakes it in token by token as the text streams out.

It's smart in theory, but it's also fragile.

Why paraphrasing breaks it

The watermark lives in the exact sequence of words the model picked. Change enough of those words and the signal falls apart. This isn't a bug anyone can patch. There's a 2024 paper out of Harvard and Boston University, "Watermarks in the Sand," that proves it: under normal assumptions, you can't build a text watermark that survives a determined rewrite without also wrecking the quality of the text. Case in point: Gemini watermarks their text. Who do you know that uses Gemini to help them write?

The empirical results back up the math. A single paraphrasing pass drops detection on these methods sharply. Free tools already strip Google's SynthID. We've been doing this at WriteHuman for years, and none of the recent watermarking work changes the underlying picture. If someone wants the mark gone, it's gone.

Here's what bothers me

If the watermark is trivial to remove, who's left carrying it?

The person who didn't know it was there.

Think of a student who writes English as a second language. She wrote her essay herself, then asked Claude to fix her grammar and tighten a few sentences. She pastes the cleaned-up text into her document and hands it in. She has no idea anything was embedded, no reason to run it through a paraphraser, no intent to deceive anyone. Her paper now carries the same mark as an essay that was generated wholesale by someone who never wrote a word of it.

The watermark can't tell those two people apart. Anthropic has said as much: the mark shows that Claude touched the text, not that Claude wrote it. Proofreading and ghost-writing look identical to it.

That's backwards. The tool lets everyone acting in bad faith walk right through, and quietly tags the people who were only trying to write more clearly.

We already know how this story goes

This isn't hypothetical, because we've watched AI detectors make the same mistake for over two years. A Stanford team tested seven detectors on essays written by non-native English speakers and found they falsely flagged more than 61% of them as AI. On writing by native speakers, the same tools were nearly perfect. The reason: careful, plain, correct writing reads as "too predictable," and predictability is what these systems treat as AI-generated.

It has done real damage. Students have been pulled into misconduct hearings over false flags. Enough universities decided these tools couldn't be trusted that many switched them off. And plenty of our customers now run their own, completely original writing through humanizers for one reason only: so they won't be accused. Not to cheat. To defend themselves.

Watermarking fixes none of this. If anything it adds to the pile: one more invisible signal that gets read as guilt no matter what actually happened.

What we actually believe

WriteHuman has a stake in this. We help people control whether their writing reads as AI generated, so discount what I'm saying however you think is fair.

But the principle we care about doesn't depend on our business. People should own their words. Someone who used a tool to help them communicate is not a cheater, and a system that can't tell the difference shouldn't get to brand them one. Transparency that only catches the honest isn't transparency. It's theater, and the bill for it lands on the people with the least standing to argue back.

Anthropic's watermark will be removed by anyone who cares to remove it. Everyone else will carry a mark they never knew was there, judged by tools that were already getting this wrong. That's the part of this story worth paying attention to.

Share
Trusted by 5M+ writers

Make AI text sound truly human.

Drop in any AI-generated draft and get back writing that reads like you wrote it yourself.

  • Works with GPTZero, Originality, Turnitin, and Copyleaks
  • Keeps your voice, tone, and meaning intact
  • One click. No prompts to memorize.
Try it free

Free to start. No card required.

Editor’s pick

Pangram AI Detector Review 2026 blog cover with the orange Pangram logo on a dark background.

Pangram AI Detector Review 2026: The Tool Substack Uses

An honest review of Pangram, the AI detector powering Substack's reader-facing scans: accuracy claims, outside validation, pricing, blind spots, and the pre-publish self-check writers should adopt.

Popular this month

  1. 01ChatGPT Zero (GPTZero): Free AI Check and What It Flags
  2. 02WriteHuman Review 2026: Does It Work? Scores and Pricing
  3. 03Rewritify.com Review (2026): Should You Switch, and When to Skip It
  4. 04Grammarly AI Humanizer Tested (2026): Does It Work?
  5. 05The Best Free AI Humanizer Tools That Actually Work in 2026

Follow on Google

Make WriteHuman a preferred sourceSee our latest posts higher in Google Top Stories

Latest

  1. yesterday

    UnAIMyText Review (2026): What You Need to Know

  2. 2d ago

    Smodin Review 2026: How Good Is Its AI Humanizer?

  3. 1w ago

    14 AI Humanizers Tested: September 2026 AI Humanizer Rankings

  4. 2w ago

    HIX AI Review (2026): Does the Humanizer Actually Work?

  5. 2w ago

    SemiHuman AI Review: A Budget Humanizer in a Crowded Field

Browse by topic

Get the WriteHuman app

iOS & Android

Humanize, detect, and rewrite from anywhere. 5 free humanizations when you install.

See the app

Ready to humanize your writing?

Try WriteHuman free and make your AI-generated text sound naturally human.

Try WriteHuman Free