AI-ParmesanAI Brand Twin

How to vet an agency's AI claims before you believe the demo

Greg Rosner

By Greg Rosner

Founder of PitchKitchen · Author of StoryCraft for Disruptors

· 8 min read

Hero image for How to vet an agency's AI claims before you believe the demo

TL;DR

To vet an agency's AI claims, stop grading the demo and start measuring the priming gap: the distance between the context a human hand-loaded into the model for your 45-minute call and the context that model will have on a random Tuesday in month four. Ask five questions. Who loaded the context, and how long did it take? What happens on a topic nobody prepped? Where does our narrative live when the engagement ends, and do we own it? Does the model get sharper as it learns us? Can we see one raw first draft nobody edited? Proprietary AI that can't survive those five questions is AI-Parmesan with a login screen.

Why does every agency suddenly have proprietary AI?

Every agency deck in B2B has an AI slide now. Proprietary engine, trained models, a platform with a name and a lowercase logo. Two years ago that slide was a differentiator. Today it's table stakes, which means it has stopped carrying information. When every vendor in the bake-off claims the same capability, the claim tells a buyer nothing.

What still carries information is the demo. An agency shares a screen, types a prompt about your company, and something genuinely good comes out. Your positioning, in your category, with your actual competitor named. It lands. You believe it, because you just watched it happen in front of you.

Here's the part nobody shows you. The reason that output was good has almost nothing to do with the model and almost everything to do with what a human being loaded into it in the hour before your call. Someone read your site. Someone pasted your deck and your last two case studies. Someone wrote three paragraphs of instructions about your buyer and your category. Then they typed a prompt in front of you, and you graded the model.

You didn't watch a capability. You watched a rehearsal. And the distance between those two things is where most AI-agency engagements quietly go to die.

What is the priming gap?

The priming gap is the distance between the context a human hand-loaded into a model for your demo and the context that same model will have on a random Tuesday in month four. A small gap means the AI claim is real and systematized. A large gap means you're buying one talented person doing prompt engineering at agency rates, and you'll feel it the week they rotate off your account.

Every good AI output sits downstream of context. The model supplies fluency. The context supplies the company. When an agency preps a demo, they close the gap by hand: a senior strategist spends ninety minutes becoming a temporary expert on you and pours all of it into the prompt window. The output is sharp because a human made it sharp, one time, for a room they were trying to win.

Then the contract starts. Now it's a coordinator running the same tool across eleven accounts with no ninety minutes to spare on any of them. The model is identical. The context is gone. The work comes back as the industry average with your logo on it, which is the exact thing you hired them to escape. That's AI-Parmesan: The B2B Marketing Plague Nobody Is Naming, sprinkled at the vendor level instead of yours.

Why is this worse in 2026 than it was two years ago?

AI brought the cost of deliverables to zero. That's the macro shift everybody's discussed. The buying consequence gets discussed far less: it also brought the cost of a convincing demo to zero. A vendor can produce a beautiful, specific, on-brand artifact about your company with twenty minutes of prep. Two years ago that artifact was evidence of capability, because producing it took real skill and real hours. Now it's evidence of twenty minutes.

Which means the way founders have always evaluated agencies has quietly stopped working. You looked at the portfolio, the writing samples, the spec work they did on your homepage to win the deal. All of it is cheap to generate now. Polish is free. Specificity is nearly free. The signal you used to read is gone, and most buyers haven't replaced it with anything.

The claim got louder at the same time. "Proprietary AI" has become the new "full-service": a statement that sounds like a differentiator and functions as wallpaper. When a founder can't separate two vendors on the AI claim, the decision falls back to price or chemistry, and both of those pick the wrong firm about as often as a coin flip. We wrote about the pre-AI version of this failure in Why does every marketing agency we hire fail?. The AI layer didn't create that problem. It just made it faster and harder to see.

The same inflation shows up in the numbers vendors show you, which is a related trap worth knowing about. How can we read AI visibility metrics without falling for inflated numbers? covers the measurement side of the same instinct: grade the instrument, not just the reading.

What five questions should you ask before you believe the demo?

Run these in the room, out loud, while the demo is still on screen. Each one measures the priming gap rather than the output quality. A vendor with real capability answers all five without flinching. A vendor selling AI-Parmesan will reach for the roadmap on at least three of them.

  1. 1Who loaded the context for this demo, and how long did it take them? You want a name and a number. "Our strategist spent about ninety minutes on your site and your deck" is an honest answer and a useful one, because now you know exactly what that output costs to reproduce. "It just pulls from the web" is either untrue or a warning, and either way you should keep pulling the thread.
  2. 2Run it live on something nobody prepped. Hand them a topic off-script: an objection your team heard last week, a segment they've never seen, a competitor they didn't research. Watch the first raw output before anyone edits it. Real capability degrades gracefully here. A rehearsal falls off a cliff, and everyone in the room can see the moment it happens.
  3. 3Where does our narrative actually live, and do we own it? Ask to see the artifact. If the answer is a set of prompts inside their platform, you're renting your own story and you start from zero the day you leave. If it's a documented framework you hold and can hand to any vendor, model, or new hire, you're buying an asset that outlives the engagement.
  4. 4Does the model get sharper as it learns us, or does it stay the same? A system that improves has a mechanism you can point at: a document that gets updated, feedback that goes somewhere specific, a version number. A system that stays the same is a wrapper with good sales support. Ask what changed between month one and month six on their last account, and listen for whether they can name it.
  5. 5Show us one first draft you didn't edit. This is the hardest ask and the most revealing. Every agency shows finished work. Almost none will show what came out of the machine before a human rescued it. The distance between those two documents is the actual size of the AI claim, and a vendor who declines has answered the question anyway.

Notice what's missing from that list. Nothing about which model they use. Model choice is Model Theater, comparison-shopping engines as though the engine is the variable. Every serious vendor reaches the same frontier models you do, on the same day you do. The variable is what gets fed in. How does AI training on brand messaging actually work? walks the mechanism if you want the longer version.

What do we see across the agencies founders bring us?

Founders forward us these decks constantly, usually with some version of "does this hold up?" Across the ones we've read, the AI claims sort into four shapes, and only one of them reliably survives month three.

What the vendor saysWhat it usually meansWhat to ask next
We built a proprietary AI trained on our client workA general model with a system prompt assembled from other companies' patterns. It averages their book of business, which is the opposite of the differentiation you're paying for.Whose narrative is in the training? If it's everyone's, it's nobody's.
Our AI learns your brand voiceUsually real, and usually shallow. Tone gets captured well. Strategy doesn't. You get copy that sounds like you and argues nothing.Show us where the position lives, not just the tone.
We use AI to speed up execution, the strategy stays humanThe most honest claim in the category, and frequently the strongest vendor in the room.Then who makes the strategic calls, and what document do they write them into?
Our platform delivers AI-powered insightsA dashboard. Sometimes a good one. Almost never a messaging capability.What decision did this change for a client last quarter?

The pattern underneath all four rows is the same. Every one of these vendors has the same models you have. The difference between them sits entirely upstream: whether anyone has done the work of writing down who your company is, who it's for, what it's against, and where it's taking the buyer. That documented narrative is what we build as a Magnetic Messaging Framework (MMF), a strategic narrative system built around four anchors: category design, villain framing, an old-way / new-way contrast, and a promised-land outcome. Greg Rosner, founder of PitchKitchen and author of Story Craft for Disruptors, developed it across more than 300 founder engagements. Without something like it in the room, an agency's AI has nothing specific to work from and defaults to the average of the internet.

This is just truth. The tool was never the moat. The context is.

The demo isn't the only thing founders over-index on when they pick a firm. Portfolio Hypnosis: why enterprise software companies keep hiring the wrong messaging agency covers the other half of this decision, the one where the work looks beautiful and still can't survive the room you're not in.

How does this play out in practice?

Here's a composite drawn from engagements we've run, with the identifying details changed.

A cybersecurity company around $18M in revenue ran a three-vendor bake-off. All three claimed proprietary AI. Vendor two gave the best demo by a distance: it produced a category-level positioning statement live in the room that made the CEO sit forward in his chair. They signed vendor two the following week, and nobody in the room thought it was a close call.

Month one was excellent. Month two was good. By month four the CEO was rewriting most of what came back, and the pieces had started sounding like every other vendor in his category again. Nothing had broken. The senior strategist who prepped that demo had rotated onto a new account, and everything she knew about the company left with her, because it had never been written down anywhere except her prompt window.

When we ran the five questions retroactively, vendor two failed three of them: no named context owner, no artifact the client owned, no mechanism for the model to get sharper over time. Vendor three, the one that lost on demo quality, would have passed all five. Their demo looked worse precisely because they hadn't hand-primed it. They'd shown the founder what their system actually produces cold, and it cost them the deal.

The fix wasn't a different agency. It was writing the narrative down first and handing the same document to the same vendor. Output quality recovered inside a month, with the same team and the same models.

What does this mean for you?

If you're evaluating an agency on its AI claim right now, the demo you're about to watch has already been primed. That's not dishonest, it's normal sales prep, and you'd do the same thing. The mistake is grading the output instead of measuring the gap sitting behind it.

The deeper move is to stop making the vendor's context problem your context problem. Once your narrative is documented and you own it, the AI question mostly answers itself. Any competent agency, running any competent model, produces work that sounds like you, because you handed them the thing that makes it sound like you. That's the whole idea behind an AI Brand Twin, PitchKitchen's trained AI voice model built on the foundation of a completed Magnetic Messaging Framework. It works because your documented narrative is what's inside it. The engine underneath is the same one everybody has.

  1. 1Before your next vendor call, write the five questions on a card and ask them in the room. Note which ones get answered with a roadmap instead of a fact. Three roadmap answers is a no.
  2. 2Ask every finalist for one unedited first draft, on a topic you pick that morning. Whoever says yes is showing you their real capability. Whoever declines has told you what you needed to know.
  3. 3Write your own narrative down before you hire anyone, and make sure you own the file. It's the only version of this asset that survives an agency change, a model change, and the strategist who rotates off your account.

We build Magnetic Messaging Frameworks for founder-led B2B companies in the $5M-$75M range, and most of the time we're fixing a broken marketing message and an underperforming website for a CEO whose sales stalled because the message wasn't doing the work. No agency's AI will invent your position for you. It'll repeat whatever's already on your site, faster and in more places. Before you write a check to anyone, find out where your own message stands with the free Brand Signal Score.

Questions People Ask

FAQ

How do you vet an agency's AI claims before you sign?

Measure the priming gap instead of grading the demo. Ask who loaded the context for the demo and how long it took, run the tool live on a topic nobody prepped, ask where your narrative lives and whether you own it, ask what mechanism makes the model sharper over time, and ask to see one unedited first draft. A real capability answers all five. A thin one reaches for the roadmap.

Why did the agency's AI demo look so much better than the actual work?

Because a senior person spent an hour becoming a temporary expert on your company and poured that into the prompt window before your call. That hour is the whole difference. Once the contract starts, a coordinator runs the same tool across eleven accounts with no hour to spare, and the output drifts back to the industry average with your logo on it.

What does 'proprietary AI' actually mean when an agency says it?

Usually a general model with a system prompt and some workflow tooling around it. That can be genuinely useful. The question worth asking is whose narrative is inside it. If the answer is that it learns from all their client work, it's averaging somebody else's book of business, which is the opposite of the differentiation you're buying.

Should we care which AI model an agency uses?

Barely. Model choice is Model Theater, comparison-shopping engines as though the engine is the variable. Every serious vendor has access to the same frontier models you do. The variable is context: whether anyone has documented who your company is, who it's for, what it's against, and where it's taking the buyer. Fluency comes from the model. Specificity comes from you.

What should we own at the end of an engagement with an AI-enabled agency?

A documented narrative you can hand to any vendor, model, or employee, in a format that isn't locked inside their platform. If the only place your positioning exists is a set of prompts in their tool, you're renting your own story, and you start from zero the day you leave. Portability is the test.

Want this kind of thinking shipping for you?

You're about to pay an agency to solve a context problem that belongs to you. The 90-Day Magnetic Messaging Sprint hands you the thing every vendor's AI is missing: your narrative, extracted from you and your team, documented as a Magnetic Messaging Framework you own and can hand to any agency, any model, any new hire. Before you sign anything, find out where your own message stands with the free Brand Signal Score at pitchkitchen.com/brand-signal-score.

That's the 90-Day Magnetic Messaging Sprint. One quarter, one fixed price: we extract your story, build the Magnetic Messaging Framework and your AI Brand Twin, then ship the website and sales enablement that run on it. $25K–$45K fixed for the quarter, and you own all of it at the end.

About the Author

Greg Rosner

Greg Rosner

Founder, PitchKitchen · Author of StoryCraft for Disruptors · Creator of the Magnetic Messaging Framework™

Greg is a B2B messaging therapist for growth-stage CEOs ($5M-$75M). He helps founders extract the truth they've been hiding from themselves, name the villain in their industry, and build the messaging infrastructure that scales their voice through AI. PitchKitchen has worked with 100+ B2B companies across SaaS, healthtech, fintech, cybersecurity, and AI-driven solutions.