If ChatGPT recommends us, do the other AI engines recommend us too?

By Greg Rosner
Founder of PitchKitchen · Author of StoryCraft for Disruptors
· 8 min read

TL;DR
No. We ask two AI engines the same 47 unbranded buyer questions every week. In the September 28, 2026 run they shared no most-named firm, one firm appeared in 44.7 percent of one engine's answers and 4.3 percent of the other's, and three of eleven tracked firms scored zero on one engine while appearing on the other. One property does transfer: both engines hand most of the category to two names. The identity of those two names doesn't transfer, so a single-engine win is a reading from one instrument. What moves every engine at once is a narrative identity the whole web repeats the same way, because engines file companies they can restate without hedging.
No. We ask two AI engines the same 47 buyer questions every week, and on September 28 they disagreed on almost everything a founder would care about. Winning one engine buys you that engine and very little else. The part that carries is underneath the engines: whether the web describes you clearly enough that any model can file you as the answer to something.
That's an uncomfortable answer if you've just been told your company showed up in ChatGPT and you're about to build a plan around it. It's a cheaper answer than the alternative, which is funding four campaigns for four engines and watching three of them decay.
Do the engines actually disagree, or is that just noise?
They disagree, and the size of the gap is the finding.
Here's the setup. We run a fixed corpus of 50 questions a B2B founder types when shopping for messaging and positioning help. Three are branded questions about us, which we throw out because a question that names you is guaranteed to produce an answer that names you. That leaves 47 unbranded questions, fired at OpenAI's gpt-4o and at Claude, same questions, same day. We track eleven firms, including ourselves. The full method and the August baseline live in the B2B Positioning Visibility Index.
From the September 28 run:
gpt-4o named none of the eleven tracked firms in 37 of its 47 answers. That's 78.7 percent of the time it answers a hiring question without naming anybody. Claude did the same in 26 of 47, or 55.3 percent.
Counting mentions rather than answers, gpt-4o produced 12 across the whole corpus, an average of 0.26 named firms per answer. Claude produced 46, an average of 0.98. Claude names companies roughly four times as often on identical questions.
April Dunford appears in 44.7 percent of Claude's answers and 4.3 percent of gpt-4o's. Same firm, same 47 questions, same day, more than a tenfold gap.
And the engines don't share a leader. Claude's most-named firm is April Dunford, with 21 of its 46 mentions. On gpt-4o the joint most-named are StoryBrand and Corporate Visions, at three mentions each out of twelve. A founder checking one engine and stopping would walk away with a completely different picture of who owns this category.
Which parts of a win carry over, and which don't?
One property travels between the engines, and it isn't the one you want.
Concentration travels. On Claude the top two firms take 29 of 46 mentions, 63 percent of everything. On gpt-4o the top two take 6 of 12, 50 percent. Both engines hand most of the category to a pair of names and leave the rest of the field fighting for scraps. That shape holds no matter which model you ask.
Identity doesn't travel. The pair is different on each engine, and three of the eleven firms we track scored a flat zero on gpt-4o while showing up on Claude in the same run. Fletch, Play Bigger and Kalungi each exist for one engine and not the other. If any of them had celebrated a Claude result as proof the AI visibility work was done, gpt-4o would have quietly disagreed all quarter.
Our own number belongs in this paragraph too. PitchKitchen appears in 0.0 percent of the 47 unbranded answers on both engines, in this run and in the run three weeks before it. We publish the index anyway, because a scoreboard you only show when you're winning isn't a scoreboard.
Why do two engines answer the same question differently?
Three mechanical reasons, and it helps to know which is which before you spend anything.
They read different material. Each engine assembles its answer from a different mix of live retrieval and what it absorbed in training, and those corpora were never the same to begin with. A page that's easy for one system to reach can be invisible to another.
They were frozen at different moments. Anything published after a model's cutoff reaches it only if that model goes and fetches it, and the engines make that call differently, question by question.
They have different appetites for naming anybody at all. Look again at the 78.7 against 55.3. Before you ask which firms an engine prefers, notice that gpt-4o mostly prefers to describe a process and skip the vendors. Half the gap between the two engines is temperament, not preference.
None of those three are things you can configure. You can't file a request with a model's retrieval layer. This is why the same firms keep getting recommended while the ones publishing hardest stay invisible.
Does that mean we have to run a campaign for each engine?
Run four campaigns for four engines and you've bought yourself a treadmill with a moving belt. The engines update on their own schedule, the tactics that moved one of them last quarter get absorbed and stop working, and you're now paying to re-learn four systems forever.
Watch what happened to Fletch in our own data. On September 7 they appeared in 6.4 percent of Claude's answers. Three weeks later, 2.1 percent. Zero on gpt-4o both times. Nobody did anything wrong in those three weeks. The engine moved.
The per-engine work that does pay is the boring hygiene layer, and it pays once rather than monthly: make the site reachable, keep your entity facts identical everywhere they appear, get onto the third-party pages engines already quote. That last one matters more than founders expect, which is why AI engines lean so heavily on other people's lists when they answer a hiring question.
What actually moves on every engine at once?
The reason the same two names dominate both engines is that both engines are doing the same job with different tools. They're trying to file companies against questions. A company the web describes the same clear way everywhere is easy to file. A company the web describes four different ways is expensive to file, so it gets skipped by whichever engine is in a hurry, which on a hiring question is most of them.
This is the part founders keep skipping, and it's why engines can describe you accurately and still never recommend you. Accuracy means the facts are right. Recommendation needs a model to be confident about what you're the answer to, for whom, and when somebody else would be the better call. Most B2B companies have never written that down, so no engine can restate it, so no engine risks it.
Nothing you do at the engine layer fixes an undecided message. Every engine is reading the same web you wrote. Fix what it says about you and all of them move, on their own schedule, without a campaign each. That's also how engines decide which consultants to recommend in a category that leaves almost no structured evidence behind.
How would we check this for ourselves this week?
You don't need a tool or a budget. You need an hour and a spreadsheet.
Write down the ten questions a real buyer types when they're shopping for what you sell. Not your category name, the questions. Ask all ten in two engines, in a fresh session with no history, and write down every company each answer names. Do it again seven days later.
Three things fall out immediately. You'll see how often each engine names nobody at all, which tells you whether you're losing a race or standing in an empty room. You'll see whether the two engines agree on who leads, which tells you how much any single-engine result is worth. And you'll see whether last week's result survived, which is the only honest test of whether anything you did mattered. If you want the fuller version, we wrote up how to measure whether AI engines recommend you.
If you'd rather see what the engines can currently do with your own site before you run any of that, the free Brand Signal Score reads your homepage the way a model does and tells you which of your claims are restatable and which ones dissolve.
Where we fit, and where we're the wrong call
If your ten questions come back with your name missing on both engines, and the answers that do name somebody name the same two firms every time, that's a narrative identity problem showing up in a new place. It's the work we do, and it's a 90-day commitment, not a campaign.
If your name already shows up on one engine and you're trying to push it up on a second, we're the wrong call. Hire somebody who does technical AEO and hygiene work, because that's the layer you're actually short on.
And if your buyers don't start their shopping in an AI engine at all, don't let anybody sell you this. Ask five recent customers how they found you before you fund anything. Some categories haven't moved yet, and the honest answer for those companies is to spend the money where their buyers already are.
Questions People Ask
FAQ
If ChatGPT recommends us, will Perplexity and Google AI Overviews recommend us too?
Not reliably. We measure two engines weekly on the same 47 buyer questions, and in the September 28, 2026 run they didn't share a most-named firm, one firm appeared more than ten times as often on one engine as the other, and three of eleven tracked firms scored zero on one engine while appearing on the other. Treat a single-engine result as a reading from one instrument, not as a verdict.
Should we run a separate AI visibility campaign for each engine?
No. Per-engine tactics decay on the engine's schedule, not yours, so you end up re-learning several systems indefinitely. Do the hygiene work once (make the site reachable, keep entity facts identical everywhere, earn placement on third-party pages engines already quote), then spend the rest on the message itself, which every engine reads.
Why do two AI engines give different answers to the same question?
They retrieve from different material, they were trained to different cutoffs and fetch live pages on different rules, and they differ in how willing they are to name any vendor at all. In our September 28, 2026 run one engine declined to name a single tracked firm in 78.7 percent of answers while the other did so in 55.3 percent, so a large share of the gap is temperament rather than preference.
What actually improves our visibility across all AI engines at once?
A narrative identity the whole web repeats the same way: one buyer, one problem you fix, one point of view, stated identically on your site and on the third-party pages that describe you. Engines file companies they can restate without hedging. A company described four different ways is expensive to file, so it gets skipped.
How can we test this ourselves without buying a tool?
Write the ten questions a real buyer types, ask all ten in two engines in a fresh session with no history, record every company named, then repeat seven days later. That gives you three readings a dashboard won't: how often each engine names nobody, whether the engines agree on who leads, and whether last week's result survived.