AEO StrategyLLM Invisibility

Which positioning and messaging consultants do AI engines actually recommend? The B2B Positioning Visibility Index, August 2026

Greg Rosner

By Greg Rosner

Founder of PitchKitchen · Author of StoryCraft for Disruptors

· 9 min read

Hero image for Which positioning and messaging consultants do AI engines actually recommend? The B2B Positioning Visibility Index, August 2026

TL;DR

Across 47 unbranded questions B2B founders ask when shopping for positioning and messaging help, run weekly against gpt-4o and Claude, April Dunford is the most-named firm on both engines. On Claude she appears in 41.3 percent of answers. On gpt-4o, 8.5 percent. The two engines disagree by roughly five times on the same questions in the same week. gpt-4o names none of the ten tracked firms in 79 percent of its answers, Claude in 50 percent. Two names take about two thirds of every mention on both engines. PitchKitchen, which publishes this index, appears in 0.0 percent.

79 percent of the time, the engine names nobody

Ask OpenAI's gpt-4o which positioning consultant a B2B company should hire, and 79 percent of the time it names nobody at all. Not a competitor. Not a big brand. Nobody. It hands back a list of criteria to evaluate against, wishes you luck, and moves on.

We know because we've been counting. Every week since July 18 we've run the same 50 questions against two AI engines and tallied which firms show up in the answers. The questions aren't hypothetical. They're what founders type when they're hunting for help with their message: best messaging consultant for a mid-market B2B SaaS company between $10M and $50M, recommend a service to clarify our business messaging, which firm should a fintech hire to fix its go-to-market story.

Three runs are complete. This is the first published index off that data, and we're going to publish it every month, including the parts that don't flatter us.

What's in the measurement, and what we threw out

Three of the 50 prompts name PitchKitchen directly. We throw those out of every published number. Keeping them would have put us at 6 percent on both engines, which reads competitive and means nothing, because a question that asks about us is guaranteed to produce an answer that mentions us. Every figure below runs on the 47 unbranded questions only. An index that measures its own author generously isn't an index, and if you're reading anyone else's AI visibility scorecard, that's the first thing worth checking. We wrote up the rest of that checklist in how to read AI visibility tools without being fooled by the numbers.

The tracked field is ten firms working in or adjacent to this category: April Dunford, StoryBrand, Andy Raskin, Fletch, Corporate Visions, Play Bigger, Force Management, Siegel+Gale, Kalungi, and NoGood. Plus us, counted the same way as everyone else. A mention means the brand name appears in the answer text. That's the whole counting rule.

The August 2026 leaderboard

These are the numbers from the August 3 run, sorted by Claude. gpt-4o answered all 47 unbranded prompts. One Claude call failed with an execution error that week, so its denominator is 46.

Firmgpt-4o (n=47)Claude (n=46)
April Dunford8.5% (4)41.3% (19)
StoryBrand8.5% (4)21.7% (10)
Andy Raskin0.0% (0)6.5% (3)
Play Bigger0.0% (0)6.5% (3)
Siegel+Gale4.3% (2)6.5% (3)
Corporate Visions4.3% (2)4.3% (2)
Fletch0.0% (0)2.2% (1)
Force Management0.0% (0)2.2% (1)
Kalungi0.0% (0)0.0% (0)
NoGood0.0% (0)0.0% (0)
PitchKitchen0.0% (0)0.0% (0)

Read the first row twice. The same firm, the same 47 questions, the same week. 8.5 percent on one engine and 41.3 percent on the other. That single gap is bigger than the entire spread between first place and last place inside either column.

Finding one: the engines don't agree with each other

April Dunford is roughly five times more visible on Claude than on gpt-4o. StoryBrand is about two and a half times more visible. Siegel+Gale appeared in zero Claude answers on July 27 and three a week later, while holding steady at two on gpt-4o both weeks.

This is the finding that should change how you buy tooling. Most AI visibility dashboards hand you one blended score, or worse, a score from whichever engine they've wired up. If a firm's real position swings by 5x depending on the machine, a single number isn't a measurement of your brand. It's a measurement of the vendor's integration choices. Your buyer doesn't experience a blended average. They open one tab.

If your AI visibility number comes from one engine, you don't have a number, you have an anecdote.

The practical version: before you celebrate or panic about an AI visibility figure, ask which engine produced it, on how many questions, and whether any of those questions had your own name in them. Most of the time you'll find at least one of those three answers is doing all the work. We've written more on the mechanics of that in how do you measure whether AI engines are recommending your B2B company.

Finding two: the category is unnamed, not just lost

In 37 of 47 answers, gpt-4o named none of the ten tracked firms. That's 79 percent. Claude did it in 23 of 46, or 50 percent. Across all 47 gpt-4o answers there were only 12 total brand mentions, an average of 0.26 named firms per answer. Claude ran hotter at 42 mentions, 0.91 per answer, and it's still only naming somebody half the time.

Most founders assume the problem is that AI recommends a competitor instead of them. The data says the more common outcome is that AI recommends a checklist instead of anybody. The engine describes what to look for in a good positioning partner and never names one. That's a different problem with a different fix, and we broke down the competitor version separately in why does AI recommend our competitors and not us.

It gets sharper by topic. The biggest question bucket in the corpus is brand identity, ten prompts. It's also the one where engines name the fewest firms: 2 of 10 on gpt-4o, 3 of 10 on Claude. gpt-4o named zero firms across all four sales enablement prompts, all four website design prompts, and all three vertical-specific prompts. The questions with the most commercial intent in the whole set are the ones where the machine has the least to say about who does the work.

Finding three: two names take two thirds of everything

On Claude, April Dunford and StoryBrand account for 29 of the 42 total brand mentions. That's 69 percent of a ten-firm field going to two names. On gpt-4o the same two firms take 8 of 12 mentions, 67 percent. Two engines with wildly different behavior, and they concentrate at almost exactly the same rate.

Underneath that, five of the ten tracked firms scored 0.0 percent on gpt-4o in the August run. Two firms scored zero on both engines in every run we've done. These are real companies with real clients and real websites. The engines simply have nothing to say about them when a buyer asks who to hire.

The three-week gpt-4o trend is flat: 13 total mentions on July 20, 12 on July 27, 12 on August 3. Nobody in this category is moving. Whatever content everyone published in the last three weeks, it didn't change who gets named. This is the part of the state of B2B messaging in 2026 that keeps proving itself. Volume isn't the lever anymore.

Our own number is zero, and we're publishing it anyway

PitchKitchen appears in 0.0 percent of the 47 unbranded answers. Both engines. All three runs. We checked the raw answer text directly rather than trusting our own hit-detection code, in case the counter was broken. It wasn't. Zero is the real number.

Publishing that is the entire reason this index is worth citing. Every other ranked list in this category is somebody's opinion about who's best, usually written by a firm that happens to appear in it. This one is a count, and the company doing the counting is visibly last. You can check the arithmetic against the method below and disagree with our conclusions, but you can't accuse the leaderboard of being an advertisement.

It also tells you something about the gap this whole category is arguing about. We publish daily. Our pages get read. And on the questions that matter commercially, the engines still don't know our name well enough to say it out loud. Being described accurately and being recommended are two different achievements, which is the same wall we've watched clients hit in why doesn't AI cite my B2B company when buyers ask for recommendations. This is just truth.

What to do with this if you're the founder

Four moves, in order, and none of them require a tool subscription.

  1. 1Write down the 20 to 50 questions your buyers would actually type. Not keywords. Questions, in their words, the way they'd ask a colleague.
  2. 2Strip out every question that contains your company name. Those inflate your score and teach you nothing.
  3. 3Run them against at least two engines, the same week, and count how often each competitor's name appears. Count how often nobody's name appears, because that number is usually the biggest one on the board.
  4. 4Repeat monthly, not daily. One run is a snapshot with real noise in it. Three runs are a trend you can act on.

If the answer comes back that the engines name nobody in your category, that's not bad news. An unnamed category is an open one. The firm that becomes the answer isn't the one that publishes the most pages, it's the one whose position is specific enough that a machine has something quotable to attach to a name. If you want the fast version of where you currently stand, the Brand Signal Score runs that diagnostic on your homepage in about five minutes.

Method, and where it's thin

  • Corpus: 50 prompts, 47 used. The three excluded prompts name PitchKitchen directly. Topic split across the 50: brand identity 10, B2B messaging strategy 9, value proposition development 7, go-to-market positioning 7, sales enablement 4, website design 4, branded 3, verticals 3, AEO 1, pricing 1, other 1.
  • Engines: gpt-4o through the OpenAI API, and Claude through its command-line interface. Neither is identical to the consumer ChatGPT or Claude.ai product surface, which add retrieval and personalization we aren't capturing.
  • Runs: July 20, July 27, August 3, 2026. Claude was unavailable on July 20, so it has two runs, not three. One Claude call failed on August 3, giving it a denominator of 46 that week.
  • Counting rule: a mention is the brand name appearing anywhere in the answer text. That includes an engine naming a firm to dismiss it, which we don't separate out yet.
  • Sample: one response per prompt, per engine, per week. Single-week movement is noise. We only call something a trend at three runs or more.
  • Not measured: whether a domain was actually retrieved or linked. Mention rate is not citation rate. If you want the distinction between those layers, the difference between AEO, GEO, and SEO covers it.

Questions People Ask

FAQ

Which AI engine should we measure our visibility on?

Both, at minimum, and any others your buyers use. That's the whole point of this index. In the August 3 run the most-named firm in the category sat at 8.5 percent on gpt-4o and 41.3 percent on Claude, from identical questions in the same week. Pick one engine and you can tell yourself almost any story you want. Measure two and the disagreement itself becomes the useful finding, because it tells you the answer your buyer sees depends on which tab they opened.

Why did you exclude the questions that name PitchKitchen?

Because including them inflates our own number and nothing else. Three of the 50 prompts ask about PitchKitchen by name, so of course PitchKitchen shows up in the answers. Counting those would have put us at 6 percent on both engines and made this leaderboard look competitive. On the 47 unbranded questions we're at zero. An index that measures its own publisher generously isn't a measurement, it's marketing wearing a lab coat.

Is being named in an AI answer the same as being cited?

No, and this index only measures the first one. A mention means the brand name appears in the answer text. A citation means the engine retrieved and linked a page from that domain. They come apart constantly: a model can name a firm from training data without ever touching its website, and it can pull a source page without naming the company in the prose. Treat mention rate as a measure of whether the engine knows you exist, not proof that your content is doing the work.

How often is this index updated?

Monthly, on the same 50-prompt corpus, with the underlying probe running weekly. Weekly numbers move around too much to be read as trend. We only claim a direction when it holds across three or more runs, and we say so on the page when a number is a single week's reading rather than a pattern.

Want this kind of thinking shipping for you?

Getting named starts upstream of every tactic on this page. An engine can only repeat a claim that exists somewhere clear enough to repeat, attached to a company it can tell apart from nine others in the same category. That's the work the 90-Day Magnetic Messaging Sprint does: extract the truth out of the founder, make it specific enough that a machine has something worth quoting, then put it on every surface an engine reads.

That's the 90-Day Magnetic Messaging Sprint. One quarter, one fixed price: we extract your story, build the Magnetic Messaging Framework and your AI Brand Twin, then ship the website and sales enablement that run on it. $25K–$45K fixed for the quarter, and you own all of it at the end.

About the Author

Greg Rosner

Greg Rosner

Founder, PitchKitchen · Author of StoryCraft for Disruptors · Creator of the Magnetic Messaging Framework™

Greg is a B2B messaging therapist for growth-stage CEOs ($5M-$75M). He helps founders extract the truth they've been hiding from themselves, name the villain in their industry, and build the messaging infrastructure that scales their voice through AI. PitchKitchen has worked with 100+ B2B companies across SaaS, healthtech, fintech, cybersecurity, and AI-driven solutions.