New Sept 16: Enhanced model updated for GPTZero and Pangram
AI Detection

Why AI Writing Is Full of Em Dashes (and How to Fix It)

Why language models overuse em dashes, how the mark became the internet's favorite AI tell, and a three-pass method for fixing a draft.

Ivan JacksonIvan Jackson6 min read
AI robot hand editing a document filled with em dashes on a glowing purple background.

Key takeaways

  • A March 2026 preprint measured GPT-4.1 producing about nine em dashes per 1,000 words even under instructions to avoid them, while Meta's Llama models produce almost none. The habit is a training choice, not a law of nature.
  • The strongest explanation is corpus skew: digitized print books from eras when em dash usage peaked, compounded by fine-tuning on markdown-heavy text.
  • Through 2025 the "ChatGPT hyphen" became mainstream shorthand for AI text, covered by the Washington Post and Rolling Stone, until OpenAI publicly promised restraint that November.
  • One em dash proves nothing. Density plus co-occurring tells is what reads as AI, and fixing a draft means changing rhythm, not just swapping characters.

In November 2025, the CEO of OpenAI logged on to X to celebrate a punctuation fix. Sam Altman announced that ChatGPT would finally honor a custom instruction to stop using em dashes, calling it a "small-but-happy win." Sit with that for a second. A frontier AI lab shipped punctuation obedience as a feature, because one horizontal line had become the most recognized fingerprint of machine writing on the internet.

The short answer to the em dash AI mystery is that language models absorbed the habit from their training data, where older print books and heavily edited longform prose use the mark far more than everyday writing does, and modern fine-tuning locked it in so deeply that some models keep producing em dashes even when explicitly told not to. The punctuation was never the problem. The density was.

Why does ChatGPT use so many em dashes?

Nobody outside the labs can say with certainty, because training data and fine-tuning recipes are secret. But the evidence that has accumulated since 2024 points at two compounding causes, and rules out a few popular theories along the way.

The training corpus did it

The most careful public analysis comes from engineer Sean Goedecke, who worked through the leading hypotheses in October 2025. His key observation: GPT-3.5, released in late 2022, did not overuse em dashes. The habit shows up in later models, which means something changed in how they were built. His best candidate is training data. Between GPT-3.5 and the GPT-4o era, labs hungry for high-quality text added huge amounts of scanned print books, and published English from the mid-1800s used em dashes at roughly 0.35 percent of all words, the historical peak. Modern general English sits closer to 0.25 percent, about 2.5 per 1,000 words. Feed a model a century of dash-happy print and it learns that good prose leans on the mark.

Goedecke also tested a theory that circulated widely: that RLHF contractors in Kenya and Nigeria nudged models toward regional English conventions. The numbers say no. His corpus check found Nigerian English uses em dashes at a small fraction of the general English rate, so if anything the raters should have trained the habit out.

Fine-tuning made it permanent

A March 2026 arXiv preprint, The Last Fingerprint: How Markdown Training Shapes LLM Prose, adds the second half of the story. The author frames chatbot em dashes as markdown leaking into prose. Models are fine-tuned on mountains of structured, markdown-formatted text, and when you ask them to drop the formatting, everything vanishes except the dashes. The paper measured GPT-4.1 at about nine em dashes per 1,000 words even under explicit suppression instructions, three to four times the human baseline, while Llama models produced essentially none. A comparison of base and instruction-tuned models showed the latent tendency exists before RLHF ever runs, which means preference tuning amplified an inherited habit rather than inventing one.

The popular folk theories fare worse. Em dashes do not meaningfully save tokens, since most could be commas. And the idea that the mark keeps a model's options open mid-sentence fails to explain why GPT-3.5 avoided it. The boring answer wins. The models write like their data, and their data was unusually fond of the dash.

How the em dash became a tell

Readers noticed before journalists did. Moderators of large subreddits started flagging polished posts with a suspicious dash density in 2024, and freelance editors began quietly stripping the mark from client work so it would not raise eyebrows. By April 2025 the argument reached the mainstream when the Washington Post covered the fight between people who saw the em dash as an AI giveaway and writers who refused to surrender it. Rolling Stone chronicled the rise of the phrase "ChatGPT hyphen" as internet shorthand. By August 2025, The Ringer was publishing a full defense of the mark, pointing out that writers from Emily Dickinson to Nietzsche wielded it expressively long before transformers existed.

The accusation stung because it hit real writers hardest. People who had used em dashes their whole careers suddenly found LinkedIn commenters diagnosing their prose as machine output. Some kept the mark on principle. Plenty of professionals, marketers, and agencies simply purged it, deciding that being right about punctuation was not worth losing a client's trust. When Altman announced the custom-instructions fix in November 2025, it was clear that enough paying users wanted the dashes gone that punctuation control became a roadmap item.

Are em dashes a sign of AI writing?

On their own, no. Humans invented the mark and some of the best human stylists use it constantly. What reads as AI is the pattern. Dashes at metronomic density, often a matched pair per paragraph setting off an aside, in prose that also carries the other machine habits. A writer who drops one expressive dash into a punchy paragraph looks nothing like a model averaging nine per 1,000 words across an entire document.

It is also worth separating human suspicion from machine detection. AI detectors do not check punctuation with a ruler. They model how predictable the text is token by token, so an em dash purge alone will not meaningfully move a score. If you want to know how a draft actually reads to a classifier, run it through a free AI detector and treat the result as a probability estimate, not a verdict. The em dash matters more with human readers, who now pattern-match on it instantly, fairly or not.

The tells that replaced the em dash

Once a tell becomes famous, it stops working as a tell, and the em dash is now famous. Attention has moved to the cluster of habits that survive a punctuation pass:

  • Vocabulary. Words like "delve," "tapestry," "landscape," and "leverage" spike in model output. We keep a running list in our breakdown of the most common AI words.

  • Contrast scaffolding. Sentences built as "not just X but Y," and the reflexive rule of three, where every list has exactly three parallel items.

  • Uniform rhythm. Paragraphs of nearly identical length, sentences of nearly identical shape, and a summarizing final paragraph that restates everything above it.

  • Formatting residue. Bold-led list items, headers for 200-word answers, and hedges like "it's important to note" padding every claim.

The full field guide lives in our piece on AI tells in 2026, which tracks how the giveaway list has shifted as models patch their most notorious habits.

How do I stop AI from using em dashes?

Work in three passes, from upstream to down.

First, instruct the model. Since OpenAI's November 2025 change, a custom instruction like "never use em dashes" actually holds in ChatGPT. Put it in permanent custom instructions rather than the chat itself, because style requests made mid-conversation decay as the context grows. Other models comply unevenly, which is exactly what the 2026 preprint measured, so verify rather than assume.

Second, do a manual pass. Search the draft for the character and make a deliberate choice at each hit. An aside becomes parentheses or paired commas. A dramatic reveal becomes a colon. A hard pivot becomes a period and a new sentence. Resist the temptation to swap every dash for a comma, which trades a famous tell for comma splices and a droning rhythm. This pass usually improves the writing anyway, since each dash forces you to decide what the sentence is actually doing.

Third, fix the rhythm, because punctuation was never the whole problem. A dash-free draft that still marches in identical sentence shapes will still read as generated. This is where a rewrite at the sentence level beats find-and-replace. An AI humanizer like WriteHuman restructures phrasing, varies cadence, and swaps the stock vocabulary instead of just deleting characters, and you can try it on the homepage without an account. The free plan covers three humanizations a month at 250 words each, and paid plans start at $20 per month. Compare the before and after through the AI detector and you can see how much of the machine texture a real rewrite removes versus a punctuation purge.

Frequently asked questions

Sources (6)
  1. 1.
  2. 2.
  3. 3.
  4. 4.
  5. 5.
  6. 6.
Share
Trusted by 5M+ writers

Make AI text sound truly human.

Drop in any AI-generated draft and get back writing that reads like you wrote it yourself.

  • Works with GPTZero, Originality, Turnitin, and Copyleaks
  • Keeps your voice, tone, and meaning intact
  • One click. No prompts to memorize.
Try it free

Free to start. No card required.

Editor’s pick

Pangram AI Detector Review 2026 blog cover with the orange Pangram logo on a dark background.

Pangram AI Detector Review 2026: The Tool Substack Uses

An honest review of Pangram, the AI detector powering Substack's reader-facing scans: accuracy claims, outside validation, pricing, blind spots, and the pre-publish self-check writers should adopt.

Popular this month

  1. 01ChatGPT Zero (GPTZero): Free AI Check and What It Flags
  2. 02Grammarly AI Humanizer Tested (2026): Does It Work?
  3. 03Does ChatGPT Watermark Text? What OpenAI Shipped in 2026
  4. 04Rewritify.com Review (2026): Should You Switch, and When to Skip It

Follow on Google

Make WriteHuman a preferred sourceSee our latest posts higher in Google Top Stories

Latest

  1. 4d ago

    What Is AI Slop? The Word Defining Content in 2026

  2. 6d ago

    Google's August 2026 Spam Update and AI-Assisted Content

  3. 1w ago

    LinkedIn's 2026 AI Slop Crackdown: Why Posts Lose Reach

  4. 1w ago

    Substack Readers Can Now Scan Your Newsletter for AI

  5. 1w ago

    SynthID Is Becoming the Default AI Watermark: What to Know

Browse by topic

Get the WriteHuman app

iOS & Android

Humanize, detect, and rewrite from anywhere. No account needed to start.

See the app
Chrome extension

Catch it before you hit send.

Grammar and humanizing right inside Gmail, LinkedIn, Notion and every other text box.

Get the extension

Ready to humanize your writing?

Try WriteHuman free and make your AI-generated text sound naturally human.

Try WriteHuman Free