Most AI visibility advice tells you to buy a tool and watch a score. That score is the least useful number in the whole exercise. We run outbound for 50+ B2B companies and have handled over 95,000 positive replies this year, so we see what buyers do before they respond, and a growing share of them look you up inside an AI answer first. Below is what to measure, what to ignore, and the monthly routine.

What does tracking AI search visibility actually mean?

AI search visibility tracking is running a fixed set of buyer questions through ChatGPT, Claude, Perplexity, Gemini, and Google AI Overviews on a schedule, then logging 4 things each run: whether you were named, how you were described, which sources got cited, and who else appeared. Presence rate across repeated runs is the metric, not a position.

The mental model people bring to this is search rankings, and it breaks immediately. There is no list of 10 positions. There is one answer, and either your company is inside it or it is not. So there is nothing to rank, which means the only honest unit of measurement is frequency: out of 20 runs of the same question, how many named you.

The second thing that breaks is the idea that a single check tells you anything. Generation is probabilistic. Ask the same assistant the same question 5 times and you will get 5 different answers, sometimes with a different shortlist of companies. A study published in the Journal of General Internal Medicine queried 6 commercial models 5 times each on identical prompts and documented real variation in what came back. Anyone who screenshots one answer and calls it a baseline is measuring a coin flip.

So the method is boring on purpose. Fixed questions, repeated runs, logged answers, compared month over month. The value is in the log, not in any single reading.

AI search visibility
How often, and how accurately, an AI assistant names your company when someone asks a question your buyers actually ask. Distinct from AI referral traffic, which only counts the small share of those mentions that carried a clickable link and got clicked.
Presence rate
The share of runs of a given prompt where your company appears by name. If a category question names you in 6 of 20 runs, presence rate is 30% for that prompt. It is the closest thing to a real KPI in this space because it survives model variance.

If you have not read the piece on why AI answers cite some companies and not others, read it alongside this one. That article covers what decides the shortlist. This one covers how to watch it move.

Why does your analytics show none of this?

Because analytics can only record a session that started with a click, and most AI mentions never produce one.

When an assistant recommends 4 companies in a comparison, describes what your business does, or tells a buyer your pricing model, none of that is a link. It is an impression on a surface you have no instrumentation on. The buyer forms an opinion, closes the tab, and later shows up in your inbox already decided, or never shows up at all. Nothing in your reporting stack registers that anything happened.

Even the mentions that do carry links convert to clicks at a fraction of the old rate. Pew Research Center tracked the browsing behavior of 900 United States adults and found users clicked a search result 8% of the time when an AI summary was present, against 15% when it was not. Roughly half the click behavior, on the same query set.

Ahrefs measured the same effect from the publisher side. Their 300,000 keyword study found the presence of an AI Overview correlated with a 34.5% lower clickthrough rate for the top ranking page, and their December 2025 update put the figure at 58%. The direction has been consistent every time anyone has run the numbers.

Then there is the crawl side, which is the part most teams have never looked at. Cloudflare publishes crawl to refer ratios showing how many pages an AI platform fetches for every visitor it sends back. Google sits around 14 pages crawled per referral. In the July 2025 data, OpenAI was at roughly 1,091 to 1 and Anthropic around 38,000 to 1. Your content is being read at enormous volume and cited at a tiny fraction of that, and none of the reading shows up anywhere in GA4.

8%
of Google sessions with an AI summary produced a result click, against 15% without one, per Pew Research
58%
clickthrough drop for position 1 pages where an AI Overview appears, in the December 2025 Ahrefs update
1,091:1
pages crawled per referral sent by OpenAI in July 2025, against roughly 14:1 for Google

Google has started closing part of the gap. In June 2026 it introduced generative AI performance reporting in Search Console, covering impressions, pages, countries, devices, and dates for URLs that appeared in AI features. That is a real improvement and it is still partial. AI Mode is folded into the overall performance report rather than broken out, and the generative AI reporting does not carry click data, so it tells you that you showed up without telling you what showing up earned. Search Engine Land has tracked the rollout, which has been incremental by region.

The practical read: the platform reporting will always lag, and it only ever covers Google. Everything ChatGPT, Claude, and Perplexity say about you is dark unless you go and look yourself.

What are the 4 numbers worth tracking?

Four, and they answer different questions. Most teams collect one of them and wonder why the reporting feels hollow.

Get outbound insights, weekly
Tactics, benchmarks, and playbooks from 50+ B2B outbound campaigns. No spam, unsubscribe anytime.
You are in. Check your inbox.
  1. Presence rate. The share of runs where you were named, per prompt and per assistant. This is the headline number. Track it per prompt rather than as a single average, because the average hides the only interesting pattern, which is that you probably appear on brand questions and vanish on category questions.
  2. Description accuracy. Whether the category, the mechanism, and the pricing model the assistant states about you are correct. This matters more than presence for anyone already getting named, because being described as the generic version of your category is worse than being skipped. Score it simply: correct, vague, or wrong.
  3. Citation set. Which domains the assistant pulled to support the answer. This is the most actionable of the four and almost nobody logs it. If the same 5 domains keep appearing for your category questions, you now know exactly where the next 5 mentions need to live.
  4. Referral quality. What the small volume of AI referral traffic does after it lands. Small volume, high intent. Semrush found in its AI search traffic study that the average visit from an AI source was worth roughly 4.4 times the average visit from traditional organic search on conversion. Judge this channel on what it does, not on how much of it there is.

Notice what is missing. There is no composite score, no share of voice percentage, and no rank. Those are all derivatives that compress the 4 real signals into one number that cannot be acted on. When a composite score drops, you learn nothing about what to fix. When description accuracy on Perplexity goes from correct to wrong, you know exactly what to go do.

A number you cannot act on is not a metric. It is a mood.

How do you build a prompt panel that means anything?

The panel is the whole method, and the mistake is writing prompts a marketer would type instead of prompts a buyer would type.

Fifteen to 25 prompts is the right size. Fewer and model variance drowns the signal. More and nobody keeps up the routine past month two. Split them across 5 categories.

Run each prompt 3 to 5 times, in a fresh session, with memory and personalization turned off. This part is not optional. A logged-in assistant that has watched you research your own company for a year will name you every time, and the number you produce will be a flattering lie.

Run across at least 4 surfaces: ChatGPT, Claude, Perplexity, and Google AI Overviews, with Gemini added if your buyers live in Google Workspace. They disagree more than people expect. It is normal to have your category right on one and your mechanism attributed to a competitor on another.

Log to a spreadsheet with one row per run: date, assistant, prompt, named yes or no, description verdict, competitors named, domains cited. That is 7 columns, and the whole panel takes about 30 minutes once you have done it twice. Never overwrite last month's tab. The entire value of this exercise is the comparison over time.

How often should you run it, and what does movement look like?

Monthly for the panel, weekly for referral traffic. Anything more frequent on the panel is measuring noise and calling it a trend.

Expect the first 2 months to look flat, because they will be. The signals that move presence rate are off-site mentions and third party descriptions, and those accumulate over quarters. Ahrefs analyzed 75,000 brands and found brand web mentions correlated with AI Overview visibility at 0.664, against 0.218 for backlinks. Mentions have to be earned, published, indexed, and then seen often enough to shift a model's confidence. That is not a 30 day loop.

What real movement looks like: a category prompt that named you in 3 of 20 runs in January naming you in 8 of 20 by June. A description that read "an AI outreach tool" in month one reading "a done-for-you outbound service that books recorded conversations with senior buyers" in month five. A citation set that used to be 5 competitor blogs now including 2 pages of yours plus a podcast transcript.

What fake movement looks like: your composite score in a tool jumping 12 points in a week. Models get updated, sampling windows shift, and a tool changes its prompt set without telling you. Never report a week over week change from a third party score to anyone who will make a decision on it.

One more discipline that pays for itself. Log which assistant got which detail wrong, not just that something was wrong. The fix is almost always a specific sentence on a specific page, or a specific missing third party description, and knowing which surface failed points straight at it.

Being described correctly only pays once buyers have a reason to look you up. Mickey Hardy went from referrals only to a 200K month once the invites started going out. Read the full case study →

Which tracking method should you use?

There are 5 real options and they answer different questions. Most teams need 3 of them and pay for the one that answers the least.

Method What it answers What it misses Effort How often
Manual prompt panel Presence rate, description accuracy, citation set, competitor set, across every assistant Nothing structural. It is just labor, and it does not scale past roughly 25 prompts 30 minutes a month once the sheet exists Monthly
Dedicated AI visibility platform The same 4 signals, sampled far more often and stored for you The composite score is not measurable from outside a model. Prompt sets change without notice Setup plus a monthly subscription Monthly review of the raw answers, not the score
Search Console generative AI report Whether your URLs appeared in Google AI features, by page, country, device, and date, per Google's June 2026 announcement Google only. No click data on the generative AI reporting, and AI Mode is not broken out Free, already in your account Monthly
Server log and bot analysis Whether GPTBot, ClaudeBot, PerplexityBot, and Google-Extended reach your deep pages or only the homepage Crawling is not citing. High crawl volume with no mentions is common, per the Cloudflare crawl to click data An afternoon to set up, using the published crawler user agents Quarterly
GA4 referral segment What the small volume of AI referral traffic does after it lands, and what it converts at Only counts clicked links. Blind to the roughly 4 in 5 mentions that carry no link at all One segment to build Weekly

The honest recommendation for a B2B company under 50 people: run the manual panel, build the GA4 segment, glance at Search Console, and skip the platform until the panel is genuinely too big to run by hand. The tools are not bad. They are sampling engines sold as scoreboards, and the sampling is the part you are actually buying.

What should you ignore?

Four things soak up attention and return nothing.

There is a fifth, softer trap. Do not confuse being cited with being described accurately. A Stanford audit of generative search engines, Evaluating Verifiability in Generative Search Engines, found that only 51.5% of generated sentences were fully supported by the citations attached to them. Your name appearing next to a source does not mean the sentence about you is right, which is exactly why description accuracy is a separate column in the log.

How do you move the number once you can see it?

Measurement is the cheap half. The work that moves presence rate happens on domains you do not own, and it is slower and less technical than any GEO checklist implies.

Start with the citation set you have been logging, because it is a map. If Reddit threads, YouTube descriptions, and 3 trade blogs keep supporting the answers in your category, that is where the next mention needs to be. You are not competing for a slot on your own website. You are competing for a sentence in somebody else's.

Recorded interviews are the highest ratio version of that work. One conversation produces a transcript page, a video with a title and description, an episode page on the host's site, clips, and usually a post from each side. That is 5 or 6 independent artifacts describing your company, in your own words, on domains you do not control, out of a single hour. The mechanics are in podcast transcripts for AI search and how to get your podcast cited by AI, and the repurposing pass that multiplies it is in repurposing episodes.

Running the show yourself flips the direction. Instead of pitching to be a guest and waiting on somebody else's calendar, you invite the people you want to be associated with, and every recording produces the same artifact stack with your brand attached to all of it. That mechanism is podcast led outbound, the operating version is a podcast acquisition system, and the commercial case, which has nothing to do with AI, is in podcast lead generation and why executives say yes to podcast invites.

None of that runs without the boring layer underneath it. Invitations only reach senior buyers if the sending setup clears filters, which is email deliverability, SPF, DKIM, and DMARC, warming a new domain, domain reputation, and ongoing deliverability monitoring with regular inbox placement tests. Who you point it at is the ideal customer profile question, worked through for outbound in defining an ICP for cold email and building the guest list.

On our side that engine carries one commitment: 30 recorded conversations with your ideal buyers in 90 days, or your money back. Editing is included, the recordings are yours, and the invites go out over email only. The AI visibility is a byproduct of running it, not the reason to run it.

The other half of moving the number is publishing things a model has a reason to cite. Original data is the strongest version, because there is no way to be the source of a number you did not measure. Our cold email reply rate benchmarks, podcast lead generation benchmarks, and state of AI outbound exist for exactly that reason. Structure helps too, once you are already on the shortlist: self contained blocks, an answer capsule in the first 120 words, and clean schema, all covered in getting cited by ChatGPT and Perplexity, getting cited by Perplexity, and AI Overviews for B2B. The difference between the two jobs is laid out in GEO versus SEO and LLM citation versus SEO traffic.

Where this lands

Tracking AI visibility is a log, not a dashboard. Fixed prompts, repeated runs, 4 columns, compared monthly. That is the entire method, and it beats every score you can buy because you can act on every cell of it.

The reason to do it at all is not the reporting. It is that a buyer who asks an assistant about your category is making a shortlist decision before you know they exist, and right now most companies have no idea what is being said in that moment. Thirty minutes a month turns that from invisible into a spreadsheet, and a spreadsheet you can work on.

Set the panel this week, record the first month as a zero point, and resist reading anything into it until month three. The signals underneath are slow by design, because a model that changed its mind about a company every week would be useless to the person asking. Meanwhile keep the pipeline that does not depend on any of this running, which is the discipline in tracking campaign performance and the metrics that predict podcast revenue.

The companies that show up in AI answers are the ones other people describe. The log just tells you whether that is working yet.

See How the Invite Engine Works

15 minute demo. No fluff. We will walk you through the exact system, show real prospect examples, and scope what it looks like for your market.

Schedule a Demo