Key takeaways
- WriteHuman finished #1 of 14 tools in the September 2026 AI humanizer rankings for the second month in a row, with a composite score of 78.29.
- The top pass rate in the field was 90.1%, the only result above 90% across all five detectors.
- Ten of the 14 tools were penalized for meaning drift or length inflation, changing text or padding it out to slip past detectors.
- Walter Writes had the fourth best pass rate on the board at 83.9% but finished tenth after a 12-point penalty.
- Newcomer SupWriter debuted at #5 as the field grew from 12 tools to 14.
- Every result was produced on entry-level plans, not premium tiers.
HumanizerBench ran its September cycle over the first two days of the month, and the results are in. WriteHuman finished first for the second month running, with a composite score of 78.29 and a 90.1 percent detector bypass rate. Both numbers led the board. Both are also better than the numbers that won us August.
If you are new to this series, HumanizerBench is an independent benchmark that puts AI humanizers through the same test every month. Each tool rewrites the same 33 prompts, and every output is checked against five major AI detectors. GPTZero, Originality.ai, Copyleaks, Winston AI and ZeroGPT all weigh in. The scoring rewards writing that passes detection while still saying what it was supposed to say, and every prompt, raw output, detector verdict and scoring script is published in a public GitHub repo. Anyone can rerun the math, which is the whole reason we treat this as the benchmark worth writing about.
August was our first month on top. September makes it two in a row, and the more interesting story is what happened to our score along the way.
How the September board finished
The field grew to 14 tools this cycle. Here is the top five.
Rank | Tool | Composite score | Detector bypass rate |
|---|---|---|---|
1 | WriteHuman | 78.29 | 90.1% |
2 | Stealth Writer | 75.39 | 86.4% |
3 | HIX Bypass | 73.97 | 86.3% |
4 | AI Humanize io | 72.94 | 81.5% |
5 | SupWriter | 72.80 | 80.3% |
There was real movement underneath us. Stealth Writer climbed a spot to take second. HIX Bypass moved up one to third. SupWriter is brand new to the benchmark and debuted at five, which is a strong first showing in a field this crowded.
We say some version of this every month because it stays true every month. The competition is good and it keeps improving. That is exactly what makes holding the top spot mean something.
Three cycles, three score climbs
The number we watch most closely is not the rank. Rank depends on what everyone else does. The score is about us.
In July we scored 73.07. In August we scored 76.69. This month we scored 78.29.

That is a climb of more than five points across two months, on the same benchmark, against the same five detectors, with a scoring script anyone can read. The test did not get friendlier. Our output got harder to flag, and it did that without giving anything back on meaning or readability, which the composite score also measures.
We ship model improvements between cycles, and this chart is where that work shows up. A leaderboard position can move for all sorts of reasons. A score that climbs three cycles straight really only has one explanation.
Crossing the 90 percent line
Last month our bypass rate was 89.1 percent. This month it is 90.1 percent, and it is the only rate above 90 in the field.
Here is what that means in practice. When you run a piece of writing through WriteHuman, nine times out of ten it comes back clearing the detectors people actually use. Not one hand-picked detector on a good day. The five that show up in real classrooms, real editorial reviews and real content pipelines.
The gap behind us is not small either. The next best rate on the board is 86.4 percent. At the other end of the table, one tool posted a bypass rate of exactly zero. The spread between the best and worst tools in this category is enormous, which is worth remembering the next time a landing page promises that something is simply "undetectable."

Ten tools took penalties. We took none.
This is the September story that matters most, and it needs a little setup.
HumanizerBench does not only score whether your text slips past detectors. It also subtracts points when a tool games the test, mainly in two ways. The first is meaning drift, where the rewritten text no longer says what the original said. The second is length inflation, where a tool pads the writing with extra words because longer, waffling text is easier to sneak past a detector.
This cycle, ten of the fourteen tools were penalized for one or both. That includes the second place finisher, Stealth Writer, which took a meaning drift penalty. Undetectable AI was docked six points for inflating length. Walter Writes is the cleanest illustration of why any of this matters. Its bypass rate was 83.9 percent, fourth best on the entire board, but a twelve point penalty for drifting meaning and padding length dropped it all the way to tenth place.
WriteHuman took no penalty at all. Nothing for drift, nothing for inflation. Only four of the fourteen tools finished with a clean score, and we were the highest ranked of them.
We think the penalty column is the most underrated part of this benchmark. Beating a detector by garbling a sentence is not humanizing. It is just a different way of ruining your writing. The entire point of this product is that your text comes back sounding like you, saying what you meant, and passing. The penalty column is where the benchmark checks that. Ours is empty.
Around the rest of the board
A few results further down are worth a note.
Undetectable AI finished sixth. Last month it fell hard after expanding texts by around 40 percent, and while it clawed back some ground this cycle, it was penalized for length inflation again. Same habit, same cost.
Grammarly finished twelfth with a 41.1 percent bypass rate. In fairness, Grammarly preserves meaning well, and that makes sense given what it is. It is a grammar and clarity product that added a humanizer feature, not a tool built around beating detection. Its score mostly tells you that a bolt-on rewriter and a purpose-built humanizer are two different things.
Penlify finished last with a bypass rate of 0.0 percent. Every one of its outputs was flagged by the detectors. There is not much analysis to add to a number like that.
Same test, same entry-level plans
One methodology detail we want to repeat, because it changes how you should read every number in this post. All of September's results came from entry-level plans. No enterprise tiers, no special access, no upgraded model unlocks that a normal customer would never see.
The 78.29 and the 90.1 percent are what the base WriteHuman plan produced. The benchmark tests the product people actually buy, which is exactly how it should be.
The full September standings
Rank | Tool | Score | Bypass rate | Penalty |
|---|---|---|---|---|
1 | WriteHuman | 78.29 | 90.1% | None |
2 | Stealth Writer | 75.39 | 86.4% | −1.0 |
3 | HIX Bypass | 73.97 | 86.3% | None |
4 | AI Humanize io | 72.94 | 81.5% | −1.0 |
5 | SupWriter | 72.80 | 80.3% | None |
6 | Undetectable.ai | 71.75 | 82.1% | −6.0 |
7 | Humanize AI Pro | 70.07 | 70.8% | None |
8 | Phrasly | 67.58 | 82.1% | −6.0 |
9 | StealthGPT | 65.10 | 79.6% | −5.0 |
10 | Walter Writes | 61.70 | 83.9% | −12.0 |
11 | Super Humanizer | 60.97 | 64.4% | −5.0 |
12 | Grammarly | 60.35 | 41.1% | −3.0 |
13 | StealthBypass | 39.38 | 78.6% | −23.0 |
14 | Penlify | 36.57 | 0.0% | −5.0 |
For anyone who wants the formula, the composite score weighs detector bypass at 42 percent, meaning preservation at 32 percent, readability at 16 percent and consistency at 10 percent, with penalties subtracted after that. The full methodology, every prompt and every detector verdict are public at humanizerbench.com.
Two months is a streak. We want a longer one.
The October cycle will run in a few weeks, the field will get stronger, and we will publish the results here either way. That is the deal with an open benchmark. You do not get to only show up when you win.
In the meantime, the best way to evaluate any of this is the same as it has always been. Take something you wrote with AI, run it through WriteHuman, and put the output in front of whichever detector you trust. Then read it out loud and see if it still sounds like you. That second test is the one most tools on this list are failing.




