Three ink strokes painted on cream paper and stacked from shortest to longest, with the longest one, the em dash, in bright blue
Last updated on

Why does AI use so many em dashes?


AI uses so many em dashes because the writing it learned from is full of them.

That’s the short answer, and it’s less satisfying than the theories going around. Settled: ask ChatGPT, Claude, or Gemini for a few paragraphs and you’ll get long dashes holding the clauses together, often more than once a screen. Not settled: the exact reason. No AI company has published one.

The essentials:

  • A chatbot writes by predicting the next likely piece of text, and the writing it was trained on — books, journalism, essays — uses the long dash constantly.
  • It’s a bad test for spotting AI. We counted: this site’s articles use em dashes at almost exactly the rate Moby-Dick does.
  • Since November 2025, ChatGPT will mostly drop them if you ask in your custom instructions.

What is an em dash?

It’s the long one: —. The name comes from typesetting, where it was the width of a capital M.

English has three horizontal marks and they do different jobs. The hyphen (-) glues words together, as in well-known. The en dash (–) covers a span, as in 2020–2026. The em dash (—) breaks a sentence open, either to shove in an aside or to land a final thought with a bit of drama.

Nothing about it is new or unusual. Emily Dickinson used dashes so freely that her early editors quietly regularized them into ordinary commas and periods, and it took until Thomas H. Johnson’s 1955 edition to put them back.

Why does AI use so many em dashes?

Because of what it read, and because it’s copying the parts of what it read that were graded highest.

A chatbot doesn’t decide what to say and then write it down. It predicts text a chunk at a time, picking each piece based on what usually follows in the material it was trained on. That’s the same machinery behind why chatbots sometimes state things that aren’t true: the model is producing what looks likely, not what it knows.

During training, not all text counts equally. Mustafa Ocal, an AI researcher at Florida International University who works on detecting AI-written text, describes it as a grading system: engineers score the sources, so a celebrated novelist’s work might be graded a 10 out of 10 while a student blog post gets a 3, and the model learns to imitate the tens. Published prose is exactly where the em dash lives, so the models absorbed the punctuation habits along with everything else. If you want the longer version of how that works, we’ve written about what a large language model actually is.

That much most researchers agree on, though no lab has published its actual weighting. Beyond it, people are guessing.

The most careful attempt we’ve read comes from software engineer Sean Goedecke, who went looking for the cause and ended up somewhere specific: AI companies started scanning old print books for high-quality training material, and English used dashes more heavily in the 1800s than it does now. Learn to write from books published in 1890 and you’ll punctuate a bit like 1890. He’s careful to call it speculation, and he’s right to. He also tested a competing theory about the English dialects of the contractors who help fine-tune these models, measured it against a corpus, and threw it out when the numbers didn’t support it.

Which leaves an odd situation. One of the most recognizable features of AI writing, and the people who study it are still trading hypotheses.

Is an em dash proof that something was written by AI?

No. It’s weak evidence at best, and weak evidence is not what people are using it for.

Here’s the problem with using it as a test. A signal only tells you something if it’s common in one group and rare in the other. A cough in flu season isn’t meaningless, it’s just nowhere near a diagnosis. That’s the em dash. Nobody has published a reliable count of how often AI writing actually uses it, so anyone quoting you a threshold is bluffing — and the human side of the comparison is high enough to swallow almost any accusation.

We wanted a number rather than an argument, so we counted. We took plain-text editions from Project Gutenberg, the free online archive of books whose copyright has expired, stripped the archive’s own header and footer, and counted the em dash character. Then we ran the same count over the 37 articles already published on this site.

TextEm dashesPer 1,000 words
The Great Gatsby (1925)4178.7
Moby-Dick (1851)1,7288.1
This site, 37 articles (counted Sept 2026)4417.8
Great Expectations (1861)1,1516.2

Melville used 1,728 em dashes in one novel. Nobody has ever suggested he had help.

Three honest caveats, and the second one cuts against us.

These are counts of particular editions, and transcriptions differ. We ran Pride and Prejudice too, and that edition contains almost no em dash characters at all, because whoever typed it up used plain hyphens. Publishing that as “Austen didn’t use em dashes” would have been nonsense, so we left it out.

These are also old books, which is exactly the corpus the leading theory says the models learned from. So they’re not an independent human control. What they show is narrower and still useful: the mark has always been ordinary in published prose, and a writer using it heavily is doing something writers have done for two centuries.

And our own articles aren’t a human control either. We use AI tools for research and first drafts, and a person edits and fact-checks everything before it goes up. Read our 7.8 as a line that sits on both sides of the divide at once, which is the whole problem with the test.

What happens when people get accused over an em dash?

It gets used anyway, and it misfires on the people least equipped to argue back.

In May 2026, Nike posted a line celebrating the tennis player Jannik Sinner: “This isn’t just history — it’s his story in the making.” Social media decided the dash gave it away and piled on. Maybe a model wrote it and maybe a copywriter did; the crowd had no way to know, and convicted on punctuation.

That’s an embarrassment for a sportswear company. It’s heavier for everyone else. On ResearchGate, academics are asking each other whether em dashes have become a false positive in peer review, with reviewers questioning papers over punctuation that has always been standard in scholarly writing. And plenty of writers have simply stopped using a mark they like, on the reasoning that being right is no defense against being accused. If you’re sending work somewhere it’ll be screened, like a resume, that pressure is real even though the test is junk.

The accusation is the damage. There’s no appeal process for a hunch.

How do you stop ChatGPT from using em dashes?

Ask in your custom instructions rather than in the message itself, and check the result.

This used to be genuinely difficult. For years, users told ChatGPT to stop, watched it agree, and watched it do the same thing in the next sentence. On November 14, 2025, OpenAI’s Sam Altman posted that it had been fixed: “Small-but-happy win: If you tell ChatGPT not to use em-dashes in your custom instructions, it finally does what it’s supposed to do!” The change landed with GPT-5.1, released two days earlier, which was pitched partly on following instructions more closely.

It works better than what came before. It isn’t absolute — users still report one slipping through now and then — so if the output is going anywhere that matters, read it before you send it. Custom instructions live in your ChatGPT settings and apply to every new chat, which is why they hold better than a request buried in one conversation. Our guide to writing better prompts covers the rest of that panel.

Other chatbots have no equivalent announcement, so the fallback is the old one: ask, then find-and-replace what’s left.

Worth asking yourself first whether you want to. Stripping the dashes doesn’t make AI text yours, and it doesn’t make your own writing safer. It mostly makes both flatter.

So what does give AI writing away?

Not punctuation, and not the detector tools either.

The habits that hold up are about substance rather than style: check whether a specific claim is actually true, because that’s where AI writing fails. Our guide to spotting AI-generated content works through that properly, including why the detector tools don’t work.

You may have noticed this article uses em dashes. It would have been a strange piece to write without them.

Punctuation isn’t the only habit these models picked up from us. They’re just as predictable when you ask for a number: why does AI always pick 17?

More plain-English explanations of how these tools work in our AI explained section.

About the picture: the three ink strokes at the top were generated with ChatGPT for this article. In a piece about who wrote what, it would be odd not to say so.

Frequently asked questions

Why does ChatGPT use so many em dashes?

Because the writing it learned from is full of them. A chatbot predicts the next piece of text based on patterns in what it read, and the material that AI companies treat as high-quality writing — published books, magazine journalism, essays — uses the long dash heavily. Nobody at the AI companies has published a full explanation, so the specifics beyond that are still guesswork.

Is an em dash proof that something was written by AI?

No. A test only works when the thing you're checking for is common in one group and rare in the other, and nobody has published a reliable count of how often AI writing uses the em dash, so the test has never actually been measured. What we can measure is the human side. We counted the em dashes in Project Gutenberg's plain-text edition of Moby-Dick and found 1,728 of them, about 8.1 per thousand words, and the 37 other articles on this site averaged 7.8 per thousand when we counted in September 2026. A mark that ordinary in published prose can't carry an accusation on its own.

How do I stop ChatGPT from using em dashes?

Put it in your custom instructions rather than in each individual message. On November 14, 2025, OpenAI's Sam Altman posted that if you tell ChatGPT not to use em dashes in custom instructions, it 'finally does what it's supposed to do.' It isn't perfect and users still report the occasional one slipping through, so check the output before you send it anywhere that matters.

Should I stop using em dashes in my own writing?

Only if you want to. Plenty of writers have dropped a punctuation mark they like because they're afraid of being accused, which is a real cost of a test that doesn't work. If you're submitting something that will be screened — a job application, a school assignment, a journal paper — knowing that some readers treat the long dash as a red flag is useful context for your own decision.

What actually gives away AI writing?

No single mark or word does, and the detector tools are unreliable. The more useful habit is to judge the substance instead of the style: pick a specific claim and check whether it's true, and look for firsthand detail a model couldn't have. AI writing tends to go wrong in checkable ways, like invented citations and wrong dates, rather than in punctuation.