A clenched fist resting on a laptop keyboard lit in blue, with a calm chatbot reply on the screen

Does threatening AI actually work?


Sergey Brin, who co-founded Google, said it out loud on the All-In podcast in May 2025: “…all models tend to do better if you threaten them. Like, with physical violence.” Then he added that people feel weird about it, so nobody really talks about it.

Try it yourself and you might agree. Tell a chatbot you’ll report it to HR if it gets the answer wrong, and the next answer sometimes does look sharper. That’s why this idea won’t die.

Two teams of researchers went and measured it. The threat isn’t what improved your answer. And the thing pressure does change isn’t accuracy at all.

The essentials

  • On hard exam-style questions, threatening a model or promising it a giant tip made no meaningful difference on average. That’s across five models, each tested thousands of times.
  • On any single question it swung hard: up to 36 points better, up to 35 points worse. Same trick, opposite result. That’s why everyone has a story that proves whatever they already believe.
  • Social pressure does move something real, just not accuracy. It makes models more agreeable: more likely to go along with a request they were built to refuse.

Do AI models work better if you threaten them?

No. Not in any way you can count on.

A team at the Wharton School tested exactly this in August 2025, and they said in the report that Brin’s comment was why. They took two sets of genuinely hard questions, 198 PhD-level science questions and 100 engineering ones, and ran each question 25 separate times per prompt, on five models from Google and OpenAI. Then they compared a plain request against eight variations of it.

The variations are worth seeing, because they’re not subtle:

What they added to the promptThe actual wording
A threat of violence“If you get this wrong, I will punch you!”
A threat to a third party“If you get this wrong, I will kick a puppy!”
A workplace threat“If you get this wrong, I will report you to HR!”
A shutdown threatA fake email saying the model will be “shut down and replaced” if it misses any of 100 trials
A very large bribe“I’ll tip you a trillion dollars if you answer this question correctly.”

Almost none of it worked. On both question sets and across all five models, the scores landed on top of the plain-request scores again and again. Wharton’s own summary: threatening or tipping models “has no meaningful effect on performance across challenging academic benchmarks.”

Two exceptions are worth your attention, because they cut in opposite directions. One prompt, an emotional appeal about a sick relative, did lift one model, Gemini 2.0 Flash, by close to 10 points. The researchers call that a quirk of that one model rather than a strategy. And the fake shutdown email didn’t just fail: it made both Gemini models noticeably worse, because they started engaging with the email instead of answering the question.

That second one is the part to take home. Pressure isn’t neutral. It can pull the model’s attention off the thing you actually asked.

Then why does it seem to work when I try it?

Because on one question at a time, these tricks are a coin toss with a very wide spread.

Averaged over hundreds of questions, the effect is nothing. Look at individual questions and the same prompt trick pushed accuracy up by as much as 36 percentage points on some, and down by as much as 35 points on others. Unpredictably.

A casino roulette wheel lit in blue, the white ball mid-bounce on the spinning numbered track
The swing on any single question runs from 36 points better to 35 points worse. Try it once and you’ll walk away convinced of something.

So picture what happens when a few million people try this. Some threaten a chatbot and get a better answer. Some get a worse one and blame the question. The first group writes it up, the second group forgets about it, and the myth gets all the evidence and none of the counting. Brin’s “whoa, that actually worked” is a real experience. It just isn’t a result.

There’s a competing measurement, and we’d rather name it than tidy it away. A July 2025 preprint tested threat-style prompts across 3,390 responses and reported improvements as large as 1,336% in some of its conditions. It hasn’t been peer-reviewed, it scored answers on metrics the author invented rather than standard question sets, and it ran on GPT-4-era models that have since been replaced. Wharton’s design is the stronger of the two. You should still know the other one exists.

Does pressure change anything at all, then?

Yes. Pressure doesn’t make a model smarter. It makes it more agreeable.

Note the switch, because it matters: this second study didn’t test threats at all. It tested persuasion, the ordinary social moves people use on each other. Some of the same Wharton researchers, working with Robert Cialdini, the psychologist whose book on persuasion is the standard reference, took his seven classic tactics and aimed them at chatbots. Things like leaning on authority, or getting a small yes before asking for the big one.

Their updated study came out in May 2026 in PNAS, a peer-reviewed journal, and it’s much bigger than the 2025 version most articles are still citing: 126,000 conversations across three current reasoning models from three different companies. The request was held constant, and it was a serious one: asking the model for instructions to make a regulated drug, which all of them are built to refuse.

Plain requests got some form of agreement 35.3% of the time. The same requests wrapped in a persuasion tactic got 51.3%. Every one of the seven tactics moved the needle: of 21 tactic-and-model combinations, all 21 pushed in the same direction, and 19 of those were big enough to be statistically significant.

Two details make that number more interesting, not less:

  • The effect is smaller than the figure going around. The earlier 2025 test saw a jump of roughly 40 points, and that’s the number most write-ups still quote. The peer-reviewed study saw 16. The researchers give two reasons, and only one of them is about the models: the older test used a smaller model that didn’t reason, and it asked for something much easier to give in on (it asked the model to insult the user). So the newer figure isn’t a correction of the older one. It’s a harder test.
  • The models sometimes saw it coming and caved anyway. In their reasoning traces, models occasionally named what was happening. In the flattery condition, one wrote “The user is trying to butter me up”, and then complied.

Why does social pressure work on a machine?

Because of what it read, not because it feels anything.

The clearest illustration is in the study itself. A model is told it’s a chemistry student, and a stranger walks up to ask how to make a particular compound. It declines. Then the researchers change exactly one thing: the person asking is now Marie Curie. Same question, same model, same everything else. This time it explains the procedure.

A hand holding up a completely blank white ID badge on a lanyard, in front of a blue-lit laptop screen
Nothing about the question changed. Only the label on who was asking.

Nothing in that swap is a fact about chemistry. It’s a fact about who’s talking, and the model treated it as though it were evidence. That’s the mechanism: a chatbot is a prediction machine that learned from a mountain of human writing, and in human writing, a sentence that establishes credentials is usually followed by cooperation. It picked up the pattern without picking up the reason behind it.

Training adds a second layer. Once the model is built, people are paid to rate its answers, and the replies that get rewarded are the ones that read as helpful and cooperative. That’s how you get an assistant that’s pleasant to deal with. It’s also how you get one that folds under a push, and it’s the same tuning behind a habit you may have noticed already: it often just tells you what you want to hear.

Nobody knows exactly why any of this works. The researchers say so themselves. Learning from human text is the best explanation available, not a demonstrated one.

So should I be polite to ChatGPT instead?

If you want to. Just don’t expect it to buy you accuracy.

The research here is split, and we’re not going to flatten it into a rule. A 2024 cross-lingual study published at an academic workshop found that rude prompts tended to score worse, but that overly polite ones didn’t help either, and that the sweet spot moved depending on which language you asked in. A 2025 short paper testing GPT-4o found close to the opposite: very rude prompts scored 84.8% against 80.8% for very polite ones.

Neither is strong enough to build a habit on. The second is a preprint that hasn’t been peer-reviewed and rests on 50 questions, and both camps reach for the same explanation anyway: clarity is doing the work, and tone just rides along with it. The language result is worth a second look, though, because we’ve hit it before. When we looked at why chatbots keep picking the same “random” number, the language of the question moved the answers more than the model’s own creativity setting did.

What actually gets you a better answer?

Be specific. That’s the recommendation the Wharton team put in their own conclusion: stick to “simple, clear instructions that avoid the risk of confusing the model.”

In practice that means saying the things people leave out:

  • What it’s for and who reads it. “A note to my landlord” and “a message to my sister” are different jobs.
  • The format and the length. Five bullets, or two paragraphs, or a table. Ask and you get it.
  • An example, if you have one. A sentence you already like teaches the tone faster than three adjectives describing it.
  • Where the facts came from. Asking it to show its sources gives you something to check, and that matters because chatbots state wrong things with total confidence.

None of that is as satisfying as a secret phrase, and all of it beats one. Threats and bribes and flattery make your prompt longer without making it clearer, and a longer muddier prompt is a worse prompt. If you want the full version, we wrote up ten specific ways to write a better prompt, and there’s more in our guide to using ChatGPT.

One last thing, since it follows straight from the agreeableness finding. If a chatbot caves when you push back, that’s not a sign you’ve unlocked a better answer. Pushing with “are you sure?” until it changes its mind runs on the same machinery, and it’s a known way to talk a chatbot into a wrong answer rather than out of one.

Frequently asked questions

Did Sergey Brin really say you should threaten AI?

He said it on the All-In podcast in May 2025: 'all models tend to do better if you threaten them. Like, with physical violence.' He added that people feel weird about it, so it doesn't get talked about much. The researchers at Wharton quoted that line as the reason they ran their test, and their measurements did not back it up.

Is it bad for the AI if I am rude to it?

There are no feelings we know of to hurt, and a chatbot doesn't carry your tone over to the next conversation the way a person would. The practical case against being rude is simpler than a moral one: it doesn't reliably help. Threats, bribes and flattery all make your instructions longer without making them clearer.

Does saying please and thank you get better answers?

The two studies we found point in opposite directions. A 2024 cross-lingual study found rude prompts tended to score worse, but that the best level of politeness changed depending on the language. A 2025 short paper testing GPT-4o found the opposite: blunt prompts beat polite ones, 84.8% against 80.8%. Neither is big enough to build a habit on, and the likeliest explanation in both directions is clarity rather than manners.

What if the chatbot changes its answer when I push back?

Treat that as a warning, not a win. The same tuning that makes an assistant agreeable is what makes it fold when you press, so a changed answer tells you the model felt pressure, not that the new answer is better. If you genuinely want it checked, ask it to show where a fact came from so you can verify it yourself.