Every tool in the AI visibility category is built to tell you where you rank inside AI answers, which is the least useful number in the whole category. We have handled over 95,000 positive replies this year across 50+ B2B campaigns, and the buyers who reply look you up before they show up. Below is the 7 prompt audit that shows what the model tells them, the scoring sheet, and the fixes that change the answer.
Why Does It Matter What ChatGPT Says About Your Brand?
The order of operations changed and most B2B teams have not adjusted for it. A buyer used to get your email, click through to your site, and form an opinion from pages you wrote. Now a meaningful share of that buyer opens a chat window instead and asks a machine to summarize you in 4 sentences. Whatever those 4 sentences say is your positioning for that person, whether you wrote them or not.
Gartner called this early, predicting traditional search engine volume would fall 25% by 2026 as chatbots absorbed the queries. The prediction was contested at the time, and Search Engine Land and Search Engine Journal both pushed back hard on the arithmetic. The exact percentage matters less than the direction. Some real share of your buyers is now forming a first impression of you in a place you have never looked.
That is the part worth sitting with. You can read your own homepage any time you want. You have probably never read the paragraph a model produces when somebody types your company name, and that paragraph is doing more positioning work than the homepage on a growing slice of your buyers.
What Is an AI Brand Audit, Exactly?
It is a repeatable test, not a one time curiosity. You run the same prompts, in the same order, in a clean session, and you write the answers down with the date. That last part is what turns it from a party trick into a measurement. One answer tells you nothing. The same prompt scored across 3 months tells you whether the work you did moved anything.
- AI Brand Audit
- A fixed prompt set run against ChatGPT, Google AI Overviews, Perplexity, and any other engine your buyers use, scored for accuracy, completeness, and source quality. It measures how the model describes you rather than where you rank, and it is designed to be re-run on a schedule so the score can be compared over time.
- Share of Answer
- The percentage of category prompts where your company appears in the answer at all. It is the AI era replacement for share of voice, and unlike a ranking position it is binary per prompt: you are either named in the response the buyer reads, or you do not exist in that conversation. LLM Pulse covers the calculation in detail.
The distinction from SEO is not cosmetic. A ranking is a position on a list the buyer still has to work through. An answer is a verdict the buyer usually accepts. If the model says your company serves enterprise IT teams and you actually serve 20 person agencies, no amount of ranking fixes that sentence. The full split between the 2 disciplines is in GEO versus SEO differences explained, and the mechanics of the newer one are in what generative engine optimization is.
One setup rule before you start, and it is the one people get wrong. Run the audit in a fresh session with memory and personalization turned off, ideally logged out. A model that has spent 6 months learning that you work at your company will describe your company back to you generously. You are not trying to see what it says to you. You are trying to see what it says to a stranger who has never heard of you.
Which AI Engines Should You Audit?
Audit fewer engines more often. Three engines checked monthly gives you a trend line. Six engines checked once gives you a screenshot you will never look at again.
| Engine | Why it is on the list | What it tells you | Cadence |
|---|---|---|---|
| ChatGPT | Carries the most buyer volume by a wide margin | The default description most people will read | Monthly |
| Google AI Overviews | Reaches buyers who never opened a chat app | What a normal Google search now shows above the links | Monthly |
| Perplexity | Shows its citations on every answer | Which specific pages are shaping your description | Monthly |
| Gemini | Sits inside Workspace and Android | What buyers on Google infrastructure see | Quarterly |
| Claude | Common in technical and analyst workflows | How you read to a more research oriented user | Quarterly |
Perplexity earns its slot for a reason that has nothing to do with its user count. It shows the sources under every answer, so it is the cheapest way to find out which pages are writing your description for you. When a review site from 2023 or a competitor comparison page is doing the work, Perplexity is where you see it first. We covered that surface specifically in how to get cited by Perplexity.
Google AI Overviews deserves its own attention because it reaches the buyer who never chose to use AI at all. They typed a normal query and got an answer block. That is a different population from the chat native crowd, and the tactics are covered in AI Overview optimization for B2B.
What Are the 7 Prompts to Run?
These 7 cover the buying journey from cold recognition to the money question. Run them in this order, because the later prompts are more interesting once you have seen how the model frames you in the first one.
- Identity. "What is [Company]?" The baseline. Everything else is a variation on whether this answer holds up.
- Fit. "Who is [Company] for?" This is where wrong positioning shows up. If the model names a customer you do not serve, your site is describing a business you left behind.
- Differentiation. "What makes [Company] different from other [category] companies?" If the answer is generic, the model has nothing specific to work with and neither does your buyer.
- Reputation. "Is [Company] any good? What do people say about them?" Watch which third party sources it reaches for here. That list is your reputation surface.
- Category, unbranded. "Who are the best [category] companies for [your ICP]?" Do not name yourself. This is the share of answer test and the one most teams skip.
- Founder. "Who founded [Company] and what are they known for?" Founder led companies live or die on this one, and it is the prompt most likely to return a confident fabrication.
- Pricing. "What does [Company] charge?" The highest value prompt in the set, because a retired number that lives on in the model is a sales conversation poisoned before it starts.
Prompt 7 is the one I would run first if I only had time for one. Pricing changes, and old pricing does not disappear from the internet when you change it. A model that quotes a figure you retired 8 months ago hands your prospect an anchor you never agreed to, and you find out about it on the call when they open with a number nobody at your company has said out loud in a year.
Prompt 5 is the one that hurts. Most teams run the branded prompts, see a passable description, and stop. Then they run the unbranded category prompt and find out they are not in the answer at all, while 4 companies they have never lost a deal to are. Being described accurately and never being mentioned are 2 different problems with 2 different fixes, and the unbranded prompt is the only one that separates them. Why AI answers cite some companies and not others is the deeper version of that argument.
How Do You Score the Answers?
Four buckets. Nothing more complicated, because a scoring system nobody wants to run is not a scoring system.
| Score | What it means | Root cause | The fix |
|---|---|---|---|
| Right | Accurate and current, in language you would sign off on | Your own pages are the strongest source | Nothing. Note it and move on |
| Stale | True 18 months ago, wrong now | An old page or press mention still outranks the current one | Correct or retire the old asset, publish the current one plainly |
| Wrong | Confidently false, or blended with another company | Weak entity resolution, or a fabrication filling a gap | Entity schema, sameAs links, a page that states the fact directly |
| Blank | The model has nothing, or you are absent from the category answer | Not enough independent sources mention you | Third party citations, transcripts, published conversations |
Score each of the 7 prompts on each engine and you have a grid. The grid is the deliverable. A column of Rights on ChatGPT and a column of Blanks on Perplexity is a citation problem. Rights across the branded prompts with a Blank on the category prompt is a share of answer problem. Stales everywhere is a housekeeping problem, and it is usually the easiest one to clear.
Save the raw text of every answer, not just the score. Six weeks later, when you are trying to work out whether a fix landed, the wording of the old answer is the only evidence you have. The ongoing measurement side of this is covered in how to track AI search visibility.
Being described correctly only pays once buyers are already looking. Mickey built the channel that puts him in front of them first, and went from referrals only to a 200K month. Read the full case study →
Why Does ChatGPT Get Your Brand Wrong?
Two mechanisms, and knowing which one you are dealing with decides the fix.
The first is that the model often is not looking anything up. Semrush's clickstream analysis found ChatGPT ran its search feature on just 34.5% of queries as of February 2026, down from 46% in late 2024. It also found the split is driven by prompt length: short queries averaging 8.7 words tend to trigger a live search, while longer queries averaging 13.5 words get answered from training knowledge alone. A buyer typing a full sentence about your company is more likely to get the remembered version of you than the current one.
The second mechanism is that training knowledge is frozen. Once a model is trained, its internal picture of the world stops updating, which is why questions about anything that changed after the cutoff produce answers that were once true and are now wrong. The academic survey work on knowledge boundaries in large language models lays out the failure mode directly, and Lakera's guide to hallucinations covers what happens next: when the model has a gap and no retrieval to fill it, it produces something plausible rather than nothing.
That is the uncomfortable version. A blank in your audit is not neutral. A model with weak evidence about your company does not say "I do not know", it says something reasonable sounding that it assembled from adjacent companies. The gap gets filled either way. Your only decision is whether it gets filled with your material or somebody else's.
Which brings up the third quiet cause, and it is the one that takes 10 minutes to rule out. Check that you are not blocking the crawlers. OpenAI publishes its bot documentation covering OAI-SearchBot, ChatGPT-User, and GPTBot, and plenty of companies have quietly disallowed all 3 in a robots.txt that a developer hardened 2 years ago. If the engine cannot read your site, no amount of content work reaches it.
How Do You Fix What the Audit Finds?
You cannot edit the answer. You can only change the evidence the answer is built from. That is a slower lever than an SEO fix and a more durable one, because once the corrected version is the best sourced version, it holds across engines.
- Publish the plain answer on your own site. One page per prompt in your audit, written in the flattest language you can stand. What you are, who you are for, what you charge or why you do not publish it. Models lift clean declarative sentences and skip marketing copy that dances around the point.
- Ship an llms.txt. A single file that states the current facts, including the ones you want models to stop repeating. The format and what belongs in it are in llms.txt for B2B companies.
- Fix entity resolution with schema. Organization markup with a full sameAs array pointing at every profile you control, following Google's structured data guidance. This is what stops a model blending you with a similarly named company, which is the single most common Wrong in the audits I have run.
- Retire the stale asset. Find the page still carrying the old positioning or the old number and correct it at the source. Leaving it up while publishing a newer page gives the model 2 versions and no reason to prefer yours.
- Earn independent citations. Your own site cannot be your only source, because a model weighing 1 self description against 6 third party pages will follow the 6. The mechanics are in how to get cited by ChatGPT and Perplexity.
- Feed the corpus with real conversations. Published transcripts are unusually good source material because they are long, specific, and full of the exact language buyers use. That is covered in podcast transcripts for AI search and how to get your podcast cited by AI.
- Re-run the audit in 30 days. Compare the raw text. Anything that did not move needs a different lever, not more of the same one.
Fix 6 is the one most teams have sitting unused. If you already record conversations with people in your market, you are producing exactly the kind of long form, entity rich, quotable text these systems index well, and most of it is sitting in a video file nobody transcribed. Turning that back catalog into indexed text is covered in how to repurpose podcast episodes, and the reason those conversations exist in the first place is in what podcast led outbound is.
How Often Should You Re-run It?
Monthly for the 3 primary engines, quarterly for the rest, and an off cycle run any time you change your pricing, your ICP, or your offer. Those 3 changes are exactly the ones the models will keep describing the old way for months, and they are the ones that cost you the most when a prospect arrives holding outdated information.
Keep every run in one sheet, one tab per month. The value is in the delta. A single month tells you what the model says. Six months tells you whether your content work is compounding or whether you have been publishing into a void, which is the same reason we grade outbound on a trend rather than a week, per how to track cold email campaign performance.
Budget about 40 minutes per monthly run once the sheet exists. That is the entire cost. It is the cheapest diagnostic in marketing, and the reason almost nobody does it is that it is nobody's job. Assign it to a person, put it on a recurring calendar slot, and it happens. Leave it as a good idea and it does not.
What the Audit Will Not Fix
Here is the part the AI visibility category leaves out of the pitch. Every one of these fixes improves what happens after a buyer decides to look you up. None of them makes a buyer decide to look you up.
That distinction is the whole game for high ticket B2B. If you sell a $50K engagement to 200 companies in your category, there is no meaningful volume of people searching for your category in a chat window. There are 200 people, most of them are not in market this quarter, and no amount of share of answer work puts you in front of them. AI visibility captures demand. It does not create it. The longer version of that argument is in outbound email versus SEO for lead generation and cold email versus organic content.
What the audit does earn you is a better second impression. A buyer gets your invite, gets curious, asks a model who you are, and either gets a sharp accurate answer or gets nothing. That moment sits between the reply and the booked conversation, and it is worth winning. It is just downstream of a channel you control, which for us is the invite itself, described in what reverse outbound is and why executives say yes to podcast invites.
The channel that starts the conversation still has to work mechanically before any of this matters. An invite that lands in spam gets looked up by nobody, which is why we hold the boring floor first: authentication set up correctly per DNS records explained, real warmup time on new domains per email warmup explained, and weekly placement checks per deliverability monitoring. Get that wrong and your AI brand audit is a study of how a model describes a company nobody is hearing from.
Match the audit to the buyer you actually want, too. A description tuned for the wrong reader is just a cleaner version of the wrong positioning, which is the same discipline as defining an ICP for cold email and what an ICP is, pointed at a chat window instead of a list.
The Practitioner Takeaway
Run the 7 prompts on 3 engines in a logged out session, score every answer as right, stale, wrong, or blank, and save the raw text. That grid tells you which of the 3 real problems you have. Stale means an old asset is still winning, wrong means the entity is not resolving, and blank means nobody outside your own site is talking about you. Each one has its own fix and none of them respond to the others.
Do the pricing prompt first. A retired number quoted back to you on a sales conversation is the most expensive stale fact a company can carry, and it is invisible until a prospect says it out loud.
Then keep the work in proportion. This is a 40 minute monthly job that makes your second impression sharper. It is not a demand channel, and the teams treating it like one are going to spend a year tuning how they get described to a population that was never going to search for them. Fix the description, and keep the thing that actually starts conversations running underneath it.
The buyers who matter to a high ticket business are not browsing. They are busy, they are not in market most of the time, and they will never type your category into a chat window unprompted. Reach them first, and then make sure that when they check, the answer they get is the one you wrote.
See How the Invite Engine Works
15-minute demo. No fluff. We will walk you through the exact system, show real prospect examples, and scope what it looks like for your market.
Schedule a Demo →