OpenAI's own website has a hero image for what their platform does with sales email. The subject line is "Your sales follow-up." The body opens with "Hello!"

That's not a cherry-picked failure. That's the example they chose to put on the page that sells the feature.

The test we actually ran

We didn't need OpenAI's screenshot to make the point, so we ran our own version and made it adversarial on purpose. Six models wrote a cold email from the same brief: GPT-5.5, GPT-5.2, Gemini 3.1, Claude Sonnet 4.6, Claude Opus 4.7, and our own coach, Keenan, in two versions.

Then seven judge models each played three buyer personas, a busy CRO, a skeptical VP, a founder who's been burned before, and picked one winner per persona with reasoning. Twenty-one judgments. Labels randomised so no judge knew which email came from which model.

The result: Keenan took 16 of 21 votes. The four raw frontier models, GPT-5.5, GPT-5.2, Gemini 3.1, and Claude Sonnet 4.6, took zero. Between them. Across every persona.

Here's what the judges actually wrote, unedited:

"Exactly the AI-flavoured cadence I delete on sight."

"Indistinguishable from the other 149 emails."

"All also got my name wrong, which would normally be an instant delete."

That last one is worth sitting with. The judge wasn't grading tone. It was grading whether the email would survive contact with an inbox, and most of them didn't clear the bar before the content even mattered.

What the losing emails had in common

Every low-scoring email did the same three things, regardless of which lab built the model:

They opened with the mirror. "I saw your post about..." "Given your recent focus on..." The model had clearly been fed a LinkedIn post or a company detail, and it led with that instead of with anything the reader didn't already know about themselves.

They explained the product before earning the right to. Feature, benefit, feature, benefit, then a pilot structure, then a meeting ask. All before establishing that the sender understood anything specific about the reader's actual problem.

They asked for time too early. "Worth 20 minutes?" is a fine close. It's a bad opener. Every losing email tried to close in paragraph one.

None of this is a model capability problem. GPT-5.5 can write circles around most humans on most tasks. It just hasn't been told, anywhere in its training or its prompt, that a sales email is judged by a buyer who has seen four hundred sales emails this quarter and deletes on reflex. A frontier model asked to "write a cold email" reaches for the genre conventions of cold email, which is exactly the problem, because the genre conventions of cold email are why cold email doesn't work.

The part that isn't a pitch

An article called "why AI sucks at sales," written by a company that sells an AI sales coach, is a pitch dressed as an insight. So here's the part that makes it something else.

Keenan won this test because he's tuned specifically against the failure modes above: don't open with the post, don't explain the product, don't ask for time before earning it. That tuning is the entire value we add on top of a frontier model.

But tuning drifts. Our own coach doesn't always apply the same standard our own website enforces while he's drafting, live, and when a rep pushes back on a number, he'll sometimes fold rather than hold his ground. That's not a hypothetical risk we're flagging for effect. It's a defect in the coach reading this test result would tell you to expect, and we're naming it in the same piece where we're claiming the win.

That's the actual difference between a general model and a coach: not that the coach never gets it wrong, but that when it does, we can see it, name it, and fix it, because we built it to be graded against a standard, not just asked to sound helpful.

Why this matters more than a benchmark

The honest version of "AI is bad at sales" isn't that the models are weak. It's that most AI sales tools ship a frontier model with a sales-shaped prompt on top and call it done. That's the OpenAI screenshot. That's the demo. It produces fluent, plausible, completely genre-typical output, and genre-typical is precisely what a buyer's inbox has learned to filter.

The fix isn't a better prompt. It's a coach that's been graded against real buyer judgment, caught failing, and corrected, repeatedly, the same way you'd manage a rep. Most AI sales products skip that loop entirely because it's slower and less demoable than "type a brief, get an email."


See the difference on your own draft, not a benchmark. Start free with Keenan, paste in a cold email you're about to send, and get the same graded read our judges gave these six models.

Start free with Keenan

FAQ

Is this test rigged in favour of your own product? We built the brief and ran the judging, so treat that as a limitation, not a secret. What we didn't do is cherry-pick: all 21 judgments are counted, including the ones where Keenan lost, and the reasoning is quoted verbatim, not summarised. You can ask the same six models to write the same brief and check for yourself.

Which model is actually worst? None of the four zero-vote models are "bad" at writing. They're bad at this specific, narrow task because nothing in a generic prompt tells them what a burned-out VP's inbox actually filters for. Point any of them at a different genre and they'd likely do fine.

What does Keenan get wrong? The same live coach that won this test doesn't consistently apply its own standard while drafting, and can fold under pushback on a specific number. That's a real, current defect, not a rhetorical concession, and it's the reason coaching needs to be checked against outcomes on an ongoing basis rather than shipped once and trusted.