How do we prove our AI is real and not just marketing?

By Greg Rosner
Founder of PitchKitchen · Author of StoryCraft for Disruptors
· 7 min read

TL;DR
You can't prove your AI is real by describing it better. Across 302 scored B2B homepages, not one landed at the bottom of our AI-Parmesan Index and 181 earned full marks for specific, mechanistic AI claims, yet 78 of those 181 are still in the Invisible band. Specificity is table stakes. What separates the pages buyers believe is a claim with a cost attached: a refusal line naming what your AI won't do, who it isn't for, or where a human stays in the loop, plus an owner for the claim (a named method, original research, a documented framework). Companies with full AI specificity and no authority signals average 15.1 of 38. The ones with both average 25.1.
You can't prove your AI is real by describing it better. Every serious company describes it well now, which is exactly why description stopped counting as evidence. A buyer believes an AI claim when they can check it without you in the room, and the fastest way to make a claim checkable is to write the sentence a company faking it couldn't afford to write.
That sentence has a shape. It names a limit. It says what your AI won't do, who it isn't for, or what a human still has to handle. We call it the refusal line, and it does more work on a homepage than three paragraphs of architecture ever will.
Here's the uncomfortable part. When we scored 302 B2B homepages this year, not one landed at the bottom of our AI-Parmesan Index. Zero. The buzzword-only homepage everybody worries about being has mostly gone extinct, and the companies who cleaned theirs up are still invisible. Specific AI language is table stakes now.
Why doesn't more technical detail make our AI claim believable?
Because detail and proof aren't the same currency, and buyers learned the difference the hard way.
A founder who hears "your AI claim isn't landing" almost always reaches for the same fix: add the model, add the pipeline, add the retrieval step, add a diagram. The claim gets longer. Credibility stays flat, because everything added is still something the buyer has to take on faith. A longer claim is still a claim.
Economists have a name for the thing that actually moves a skeptical audience. They call it a costly signal, an idea running back through Michael Spence's work on job-market signalling and Amotz Zahavi's handicap principle in biology: a signal carries information only when it would be expensive or painful to fake. A degree signals something because it's hard to get. A warranty signals something because honoring it costs money. "We use a proprietary AI engine" signals nothing, because it's free to type.
Your buyer isn't running that theory consciously. They're running the field version of it, which is a screen, because they've been burned by a demo that worked in the room and collapsed in month four. We wrote the buyer's side of that screen in how to vet an agency's AI claims before you believe the demo. You're the one being screened now.
What separates a claim a buyer checks from one they skip?
Cost, and the cost that matters here is the buyer's. Every claim on your page carries a check cost: the time a buyer would have to spend, without your help, to find out whether it's true. When that cost is high, the claim gets skipped no matter how true it is. Skipped claims contribute nothing, and a page made entirely of skipped claims is why founders tell us their site undersells a genuinely good product.
A refusal line collapses check cost to zero, because it does the checking for them. Nobody publishes a limit they haven't hit. When you write "this doesn't work on unstructured PDFs yet, so we start customers on the API feed," the buyer doesn't verify it. They believe it instantly, and then they believe the sentence next to it.
A refusal line carries a cost, and in a market where language is free that's the only kind of claim that carries information. That's the same logic behind why buyers stopped believing marketing claims, applied to the one category where the hype ran hottest.
Why did this get harder in 2026?
Three things stacked up.
- 1The floor rose. Two years ago a specific AI sentence was a differentiator. Now it's the baseline, and our own scoring data says the same thing. Sounding substantive is no longer a signal of being substantive.
- 2The claim went legal. AI-washing moved from a marketing embarrassment to a litigation category, with securities cases turning on whether the AI did what the page said it did. Buyers at regulated companies now route your homepage through legal before procurement.
- 3The buyer got a research assistant. Founders aren't the only ones using AI. Your prospect asks ChatGPT what you do before they ever load your site, and an engine that can't find a checkable, attributable claim will summarize you as one of several similar vendors. That's how a real capability turns into a category description.
We wrote up that second shift in AI-Parmesan just became a securities problem. Put the three together and your AI claim is being graded by a more skeptical human, a more literal lawyer, and a machine that only repeats what it can extract. All three reward the same thing, which is a claim with an owner and a boundary.
What do we see when we score 302 B2B homepages on their AI claims?
Our 2026 messaging index scores 302 B2B homepages against the 19 signals in the Brand Signal Score. One of those signals is the AI-Parmesan Index, which asks whether a page's AI claims are specific and mechanistic or sprinkled on top of nothing. The field averages 19.18 out of 38. Here's the cut we ran for this piece.
| What we measured | Result |
|---|---|
| Homepages scoring at the bottom of the AI-Parmesan Index | 0 of 302 |
| Homepages scoring full marks on it (AI claims specific and mechanistic) | 181 of 302 |
| Of those 181, how many still land in the Invisible band | 78 |
| Those 181 companies' average total score | 19.78 of 38 (field average 19.18) |
| Full marks on AI specificity but zero authority signals | 50 companies, averaging 15.10 |
| Full marks on AI specificity and full authority signals | 28 companies, averaging 25.07 |
| Homepages that name an alternative or a trade-off at all | 24 of 302 |
Read the first two rows together and the strategy problem is obvious. Nobody is failing at being specific about their AI, and being specific moves the total score by roughly half a point. Sixty percent of the field earns full marks on the signal founders worry about most, and 78 of them are still functionally invisible.
The rows underneath are where the movement lives. Same AI-claim quality, ten-point spread, and the variable is whether the page gives the claim an owner: a named method, original research, a framework, a body of work that explains why this team gets to say this. The 278 companies who never name an alternative are the same story from the other side. They've made their AI claim unfalsifiable, and unfalsifiable reads as unchecked.
These are associations across one scored field, not proof of cause. We're not claiming a refusal line raises your score by ten points. We're saying the pages that carry an owner and a boundary cluster at the top, and the pages that carry neither cluster at the bottom, and that pattern held across every cut we ran.
How do we write a refusal line without shrinking the product?
The fear is real and it's usually backwards. Founders think naming a limit hands the deal to a competitor. In practice it disqualifies the buyer who was going to churn in month five and accelerates the one who was already a fit, which is the same mechanic behind why deals die in no decision when nobody gives the buyer a reason to feel certain.
Four places to look for yours:
- The input boundary. What data does your AI need that most prospects don't have clean yet? Say it, and say what you do in the meantime.
- The human boundary. Where does a person still review, approve, or override the output? A named human-in-the-loop step is the single most credible sentence a health, finance, or security buyer can read.
- The scope boundary. What adjacent job does your AI deliberately not do? Naming the thing you left to someone else proves you made a choice instead of a claim.
- The confidence boundary. Where is the model good, and where is it merely useful? A published accuracy range with the conditions attached beats an unqualified superlative every time.
Then give the claim an owner. That's the second half of the data above, and most founders skip it because it feels like bragging. It's attribution. A named method, a piece of original research, a documented framework, a defined term you use consistently: these turn a floating claim into somebody's claim. Documenting them in one place is the whole point of the Magnetic Messaging Framework, and the mechanics of making engines trust them are in how to build entity authority AI engines trust.
“A claim nobody can disprove is a claim nobody has to believe.”
... Greg Rosner, PitchKitchen
What does this look like on a real page?
A healthtech company we scored leads with the category sentence everyone leads with, then does something almost nobody does. It names the enemy directly, calls out operative reports written from memory hours after a procedure, and quantifies what that costs with sourced numbers. Its AI claim isn't "AI-powered documentation." It's a named, trademarked method that converts one specific input into one specific output, with the accuracy of the status quo printed on the page next to it.
That last part is the refusal line doing its job. Publishing the baseline you're beating is a limit: it tells a buyer exactly where the ceiling is and what you're claiming to move. It scored in our top band without a single sentence about model architecture.
Compare that to the far more common version. "Our AI platform surfaces intelligent insights across your workflow." Specific enough to clear the buzzword screen, checkable by nobody, ownable by anybody. It's the shape we named in AI-Parmesan, the B2B marketing plague nobody is naming, and cleaning up the vocabulary without adding a cost signal just produces a fluent version of the same problem. Sharpening your point of view alone has the same limit, which we cover in how to talk about AI without sounding like everybody else. A point of view gets you read. Belief needs a cost signal sitting on top of it.
What should we change this week?
- 1Print your homepage and mark every AI claim with a C or an F. C means a buyer could check it in ten minutes without talking to you. F means they'd have to take your word for it. Count the ratio. Most founders find zero Cs.
- 2Write one refusal line. Pick the input, human, scope, or confidence boundary you're most confident about and put it in plain language above the fold or in the first product section.
- 3Give your strongest AI claim an owner. Attach it to a named method, a number you generated, or a documented framework. Unattributed claims get summarized. Attributed claims get quoted.
- 4Name one alternative honestly, including doing nothing. You'll be one of 24 companies out of 302 who do, which is a differentiator that costs you nothing but nerve.
If you want an outside read before you rewrite anything, run your homepage through the Brand Signal Score. It scores the same 19 signals we used for the data above, including the AI-Parmesan Index, so you'll see exactly where your AI claim sits against the field instead of guessing. And if the harder question is how your whole team should be using AI without producing more of this, that's what we work through in the AI Workforce Clinic. The first one's free.
The proving problem is really a cost problem. Right now every sentence about your AI is free to write, which means it's free to ignore. Put a price on one of them and watch what happens to the rest.
Questions People Ask
FAQ
How do we prove our AI is real to a skeptical buyer?
Publish a claim they can check without you. The fastest version is a refusal line: one sentence naming what your AI won't do, who it isn't for, or where a human still reviews the output. Nobody publishes a limit they haven't hit, so a boundary reads as evidence in a way that more technical description never does. Then attach your strongest claim to a named method, a number you generated, or a documented framework so it belongs to somebody.
Doesn't adding technical detail make an AI claim more credible?
Not on its own. Detail is still something the buyer has to take on faith, and it's free to produce. Across 302 B2B homepages we scored in 2026, 181 earned full marks for specific, mechanistic AI claims and their average total score was 19.78 out of 38 against a field average of 19.18. Sixty percent of the field is already specific, which means specificity has stopped distinguishing anyone.
Won't naming what our AI can't do cost us deals?
It costs you the wrong ones. A published limit disqualifies buyers who would have churned once they hit it, and it accelerates the ones who already fit, because it removes the doubt they were carrying silently. Only 24 of the 302 homepages we scored acknowledge any alternative or trade-off, so honesty here is close to a free differentiator.
What is AI-washing, and how is it different from bad AI messaging?
AI-washing is marketing a capability your product doesn't meaningfully have, which since 2024 has become a securities and litigation issue rather than only a credibility one. Bad AI messaging is describing a real capability in language nobody can check. Both produce the same buyer response, but only one of them can end up in front of a regulator.
How do AI search engines treat unverifiable AI claims?
They flatten them. An engine summarizing your company repeats what it can extract and attribute, so a floating claim like 'AI-powered platform' gets folded into a generic category description alongside your competitors. A claim tied to a named method, a specific input and output, or original research is quotable, and quotable claims are the ones that survive into an answer.
