Why does AI recommend us when we ask, but not when a buyer asks?

By Greg Rosner
Founder of PitchKitchen · Author of StoryCraft for Disruptors
· 7 min read

TL;DR
AI engines name you on the general category question and drop you when a buyer adds one ordinary detail about themselves. Clovion AI, measuring 69,120 multi-turn conversations, found re-asking the same question keeps about 90% of the recommended list while adding "for a small team" keeps 28%. A separate arXiv preprint by Dmitrij Zatuchin found monthly recommendation share almost frozen across seven months at a Spearman of 0.994. Both hold: category ownership is stable, your match to a specific buyer is not. The detail that flips the list is a segment question, so the fix is naming who you're for in plain words.
Adding four words to a question removes 62% of the companies AI just recommended.
The four words are "for a small team." Ask an answer engine who the best vendors in your category are and you get a list. Add one ordinary fact about your own company, and most of that list disappears.
Here's the direct answer. AI names you on the general question because it can match you to a category. It drops you the moment a buyer describes their situation, because your positioning never said who you're for, and the model has nothing to match the buyer against.
Your place in the first answer tells you almost nothing about whether you survive the second one.
Why does the answer change when a buyer adds one detail?
Clovion AI, an Oslo research firm, ran 69,120 multi-turn conversations across Claude, ChatGPT and Gemini in 36 B2B software and fintech categories. Re-asking the identical question kept about 90% of the recommended list intact. Adding "for a small team" kept 28%. On large-enterprise framings the churn reached roughly 72%.
Two cautions before you spend anything on that number. Clovion sells answer-engine optimization, so this is a vendor publishing research about the problem it sells into. And it doesn't disclose when the data was collected. Greg Jarboe wrote it up in Search Engine Journal in July 2026, which puts the reading almost three months old by now.
Read it as directional, not as physics.
What makes it worth reading at all is that a completely unrelated party measured the same instability from the other side.
Is our AI visibility score even stable?
Dmitrij Zatuchin's preprint on arXiv, revised on September 23, 2026, put 250 brand-free category queries to GPT-5.2, Gemini 3 Flash and Perplexity sonar-pro, five runs each, across 50 brands in five industries. He ran the whole thing in February and again in September.
Across those seven months, recommendation share barely budged. The cross-date correlation came in at a Spearman of 0.994.
Both readings are true at the same time, and holding both is the whole point. Across months, the leaderboard is close to frozen. Inside a single conversation, it churns. Category ownership is stable. Your match to one specific buyer is what moves.
A founder reading a monthly visibility dashboard is looking at one frame of a moving picture. The tool isn't lying. It's answering a narrower question than the one you think you asked, which is the same trap we've written about in how to read AI visibility metrics without falling for inflated numbers.
Zatuchin is careful about what his work can carry, and you should be too. It's a preprint, not peer-reviewed, and he states plainly that the statistics describe system output and identify no causal mechanism. Nothing here proves that rewriting your positioning makes a model recommend you. It shows what the output does.
Which detail actually does the damage?
Look at which detail collapses the list in the Clovion data. It isn't a feature question. It isn't a price question. It's a segment question.
For a small team. For a large enterprise.
That's the first question of the Three Questions Test getting measured in public. Who is this for. A company that never answers it in plain words can be named in the abstract and can't be matched to anybody real.
This is what Solution-Centric Marketing produces at the end of the pipeline. You describe the product beautifully, you never describe the buyer, and an engine asked about a specific buyer has no reason to reach for you.
The buyer's follow-up question is the qualifying round. Most companies never find out they lost it.
What happened when we measured our own qualified prompts?
Our own numbers cut both ways, and that's why they're worth showing. In a 30-day window ending June 29, 2026, across a tracked set of 70 buyer prompts, PitchKitchen scored 72% on "best messaging consultant for mid-market SaaS $10-50M," 78% on "best B2B messaging firms for fintech," and 57% on "best B2B messaging consultants for growth-stage." Over the same stretch, on the wide generic ChatGPT field, we sat between 7% and 10%.
That's the Clovion pattern running backwards. We're weaker on the broad question and stronger on the qualified one.
There's no mystery in it. Every page on our site names the band we serve, $5M to $75M in revenue, and names the situation we serve it in. When a buyer adds their size, the engine has something to match.
Now the honest part. Those readings came from a measurement tool we've since retired, and they measure a surface where the engine can fetch our pages, so they reflect our own pages by construction. Thirty-nine of those 70 prompts still scored zero. We survive qualification. We don't win everywhere, and the difference matters.
Is there any room left, or does somebody already own every category?
Zatuchin found that 7.6% of category queries returned no brand consistently owning the answer. He also notes, provisionally, that most of those gaps probably reflect which brands he happened to sample rather than a genuinely empty room.
Take the number with his caveat attached. Roughly one query in thirteen has no settled owner, and some of that is an artifact of the sample.
It's still the most useful thing in either study. The vacancies sit in the specific questions, not the broad ones, because the broad ones were decided years ago by whoever the web talked about most. That's the mechanism behind why AI recommends your competitors and not you.
Nobody has to beat the incumbent on the general question. The general question was never the one that closes.
What do we actually do this quarter?
Three moves, in order, and you can run all of them without hiring anybody.
- 1Run your prompt set twice, not once. Take the 15 to 20 buyer questions you already track, ask each one plain, then ask it again with your buyer's real situation attached: their size, their industry, their constraint. The gap between the two lists is your actual problem, and almost nobody measures it. Our method for building that prompt set is in how do you measure whether AI engines are recommending your B2B company.
- 2Read your homepage for the segment sentence. Not the value proposition. One sentence that names who you're for, specifically enough that a stranger could disqualify themselves. If it isn't there, the engine is guessing on your behalf every time a buyer adds a detail.
- 3Fix the sentence before you touch the tooling. No amount of schema markup gives a model something to match when your positioning never named a buyer.
The free Brand Signal Score runs the homepage half of that in a few minutes. It's a 19-criteria read on narrative clarity, trust, AI-readiness and conversion, and it'll tell you whether a stranger, or a model, can work out who you're for.
The rest of it is a conversation your team has to have, and it's uncomfortable, because naming who you're for means naming who you're not for.
Our buyers have been describing themselves to an engine long before they ever describe themselves to us. Some version of the shortlist gets built without anybody visiting a website.
You can't be in the room for that. You can only decide, in advance, whether the sentence they meet is specific enough to survive them.
Questions People Ask
FAQ
Why does ChatGPT recommend our company sometimes and not other times?
Because you're re-decided on every turn of a conversation, not ranked once. Clovion AI measured 69,120 multi-turn conversations and found that re-asking the identical question keeps about 90% of the recommended list, while adding one buyer detail such as "for a small team" keeps only 28%. The engine isn't being inconsistent. It's matching a more specific request, and companies that never state who they serve have nothing to be matched against.
Does a monthly AI visibility score tell us whether AI recommends us?
It tells you whether you own the category question, which is stable. An arXiv preprint by Dmitrij Zatuchin, revised September 2026, re-ran 250 category queries seven months apart and found recommendation share almost unchanged at a Spearman correlation of 0.994. What a single-turn score can't see is whether you survive a buyer's follow-up. Run your prompt set twice, once plain and once with a real buyer situation attached, and compare the two lists.
Which kind of buyer detail makes AI drop a company from its recommendation?
A segment detail, not a feature or price detail. In the Clovion data the collapse came from ordinary self-descriptions such as "for a small team" or "for a large enterprise," and the large-enterprise framing churned roughly 72% of the list. That's the first question of the Three Questions Test measured in public: if your positioning never names who you're for, the model has no basis for keeping you once the buyer says who they are.
Is every category already owned by an incumbent in AI answers?
No. Zatuchin found 7.6% of category queries returned no brand consistently owning the answer, though he notes provisionally that most of those gaps may reflect which brands his study sampled rather than genuinely empty rooms. Take the figure with that caveat. The broader point holds: the vacancies sit in specific, qualified questions rather than the broad ones, which were largely settled by whoever the web already talked about most.
Can we fix this with schema markup or an AEO tool?
Not on its own. Markup helps an engine parse a page that already says something matchable. It can't supply a buyer definition that your positioning never made. Fix the segment sentence first, then make it machine-readable. The free Brand Signal Score reads your homepage against 19 criteria covering narrative clarity, trust, AI-readiness and conversion, which tells you whether a stranger or a model can work out who you serve.