Most advice on AI visibility tells you to restructure your pages so a model can read them, and page structure is not what decides who gets named. We run outbound for 50+ B2B companies, so we watch how senior buyers vet a vendor before they reply, and that check now starts inside an AI answer. Below is what actually decides the shortlist, the data behind it, and the work that puts you on it.

Why does AI name some companies and skip others?

A model picks the companies before it picks the pages. It assembles a shortlist of brands it is confident exist and belong in the category, then finds pages to support that shortlist. Confidence comes from how often and how consistently the open web describes you, so a company nobody else writes about rarely makes the list.

That order is the whole thing, and almost every GEO checklist gets it backwards. The checklists assume the model reads the web, finds the best page, and cites it. What actually happens is closer to how a person answers the same question. Somebody asks you to name the good outbound agencies in your city. You produce 3 or 4 names off the top of your head, then you go look them up to make sure you are not saying something wrong. The lookup confirms the shortlist. It does not build it.

If your company was not in the recall set, no amount of page polish gets you into the answer, because the model never went looking for you. This is why two companies with identical websites can have completely different AI visibility. One of them gets talked about somewhere other than its own domain. The other does not.

AI citation
A named reference to a company, page, or source inside a generated answer from an assistant such as ChatGPT, Claude, Perplexity, or Google AI Overviews. Distinct from a ranking, because there is no list of 10 positions. There is one answer, and you are either in it or you are invisible.
Entity confidence
How certain a model is that a company exists, what category it belongs in, and what it does. Built from repeated, consistent descriptions across many independent sources. Low entity confidence is why an assistant hedges, describes you generically, or swaps your mechanism for a competitor's.

The practical version: your website controls how you get described. It does not control whether you get mentioned. Those are two different jobs and most companies are only doing the first one. The mechanics of the split are covered in GEO vs SEO, and the reason citation and traffic behave differently is in LLM citation vs SEO traffic.

What signals actually decide the shortlist?

There is real data on this now, and it points hard in one direction.

Ahrefs analyzed 75,000 brands against AI Overview visibility and ran correlations across every signal they could measure. Brand web mentions came out on top at 0.664. Brand anchors landed at 0.527. Brand search volume at 0.392. Backlinks, the signal an entire industry has been built around, came in at 0.218.

Read the top 3 again. Web mentions, anchor text describing you, and how many people search your name. All 3 happen on properties you do not own. The highest scoring on-page factor in that study does not crack the top of the list, and the gap is not small. Mentions beat backlinks by roughly 3 times.

The obvious caveat applies. Correlation is not causation, and a brand that gets written about in trade press is usually a brand that has customers, a category position, and something worth writing about. Mentions are a proxy for being real. They are not a lever you pull in isolation, which is exactly why buying a pile of mentions from a syndication service does very little.

Signal Where it lives Strength How fast you can move it
Brand web mentions Other people's domains Strongest measured correlation, 0.664 in the Ahrefs 75K study Months. Every mention has to be earned and then indexed
Brand anchors Other people's domains 0.527 Months, and mostly a byproduct of the mentions above
Brand search volume Search engines 0.392 Quarters. This is downstream of everything else you do
Backlinks Other people's domains 0.218, roughly a third of plain mentions Months, and the effort is better spent on the description than the link
Schema markup Your own pages Not a shortlist factor. Decides how accurately you get described One afternoon, per schema.org
Answer capsules and chunking Your own pages Not a shortlist factor. Decides which page gets pulled One pass through your library
llms.txt Your own domain Effectively zero on its own 2 hours, and worth it for other reasons, per the llms.txt breakdown

The table splits cleanly into two halves. The top half decides whether you appear. The bottom half decides how you appear once you do. Both matter, and companies spend almost all their effort on the half that only matters second.

Why do off-site mentions beat anything you publish yourself?

Because a model has no way to verify a claim you make about yourself, and every way to verify a claim other people make about you.

Get outbound insights, weekly
Tactics, benchmarks, and playbooks from 50+ B2B outbound campaigns. No spam, unsubscribe anytime.
You are in. Check your inbox.

Every company on the internet says it is the leading provider of something. That sentence carries no information, because it appears on 40,000 homepages and the model has seen all of them. A sentence in somebody else's article that says your company runs outbound for B2B agencies and does it by inviting buyers onto a podcast carries a lot of information, because the writer had no reason to say it unless it was true.

This is the same instinct a buyer has. Nobody believes a homepage. Everybody believes a description from a third party who was not paid to write it. Models inherited that structure from the training data, not from a rule somebody wrote.

The second reason is repetition across independent sources. One mention is an anecdote. Forty mentions across 25 domains, all describing the company the same way, is a fact the model can state without hedging. That is what entity confidence is, and it is why consistency matters more than volume. Fifteen sources describing you 15 different ways produces a model that describes you vaguely, which is functionally the same as not being described at all.

You cannot argue your way into an AI answer from your own domain. You can only be corroborated into one.

Third, a mention on a live, frequently crawled domain gets refreshed. Your about page might get fetched every few weeks. A discussion thread or a news post on a high traffic domain gets fetched constantly, which means anything said about you there stays current in the retrieval layer instead of aging out.

Which sources do AI answers pull from most?

The answer surprises people who expected company websites and trade publications.

A Semrush analysis of over 150,000 AI citations found Reddit accounting for 40.1% of references, Wikipedia at 26.3%, and YouTube at 23.5%. Search Engine Land reported the same broad pattern across ChatGPT, Google AI Mode, Gemini, Perplexity, and AI Overviews, with Reddit first and YouTube, LinkedIn, Wikipedia, and Forbes filling out the top 5.

A Q1 2026 audit synthesizing 9 independent datasets put Wikipedia at 13.15% and Reddit at 11.97% of United States ChatGPT citations, with major newspapers absent from the top 20 entirely. Visual Capitalist has a readable version of the same ranking across models.

The exact percentages move around, sometimes violently, and anybody quoting one number as permanent is selling something. What holds across every dataset is the shape: platforms where people talk about companies outrank the companies themselves. Discussion, video, encyclopedic reference, and professional profiles. Not marketing sites.

0.664
correlation between brand web mentions and AI Overview visibility across 75,000 brands
0.218
correlation for backlinks in the same dataset, roughly a third as strong
51.5%
of generated sentences fully supported by their citations in a Stanford audit

There is a strategic read buried in that list. You are not competing for a slot on your own website. You are competing for a mention inside a Reddit thread, a YouTube description, a LinkedIn post, or somebody else's article. That is a completely different job than publishing more pages, and it is the job almost nobody in B2B is staffed for.

Does anything on your own pages still matter?

Yes, and the distinction is worth being precise about, because the answer is not no.

On-page work decides which of your pages gets pulled once the model has already decided your company belongs in the answer. That is a tiebreaker, and tiebreakers matter when the alternative is a competitor's page being quoted in a paragraph that is nominally about you.

Three things carry most of the weight there.

  1. Self contained blocks. A model lifts a chunk, not a page. A block of 50 to 150 words that answers one question completely, with no setup and no dependency on the paragraph above it, is the unit that gets extracted. A page that answers 6 questions half way loses to a page that answers 1 question fully.
  2. An answer capsule near the top. Put the plain language summary in the first 120 words. Analyses of citation behavior consistently find that a large share of quoted content comes from the top third of a page, and burying the summary under 800 words of context wastes the most extractable block you have.
  3. Schema markup on every page. FAQPage, Article, Organization, and Person schemas hand a model labeled facts instead of prose it has to infer from. Google's structured data documentation is the reference, and none of it requires guessing.

What does not carry weight, despite the hype: an llms.txt file on its own, keyword density, word count for its own sake, and any of the AI visibility scores sold by tools that cannot see inside a model. The full argument on the file is in how to write an llms.txt for a B2B company, and the practical citation tactics are in getting cited by ChatGPT and Perplexity and getting cited by Perplexity specifically. For the Google surface, AI Overview optimization for B2B covers the same ground.

Being described correctly only pays once buyers have a reason to look you up at all. Mickey Hardy went from referrals only to a 200K month once the invites started going out. Read the full case study →

Why does AI get your company wrong even when it names you?

Getting cited and getting described accurately are separate problems, and the second one is worse than most companies realize.

A Stanford audit of generative search engines, published as Evaluating Verifiability in Generative Search Engines, found that on average only 51.5% of generated sentences were fully supported by the citations attached to them, and only 74.5% of citations actually supported the sentence they were attached to. Roughly half of what an assistant says with a source next to it is not fully backed by that source.

For a company, that shows up in 3 recognizable ways.

All 3 come from the same root cause. Low entity confidence makes a model interpolate, and interpolation always lands on the average of the category. The fix is not more content. It is the same sentence, about the same mechanism, repeated in enough independent places that there is nothing left to interpolate.

How do you get mentioned in places you do not own?

This is the actual work, and it is slower and less technical than a GEO checklist implies.

Being interviewed is the highest leverage version, and the reason is mechanical rather than clever. One recorded conversation produces a transcript page, a video with a description, a set of clips, an episode page on the host's site, and usually a post from the guest and a post from the host. That is 5 or 6 independent artifacts describing your company, in your own words, on domains you do not control, from a single hour. Nothing else in B2B marketing has that ratio. The argument in full is in podcast transcripts for AI search and how to get your podcast cited by AI.

It also explains why YouTube keeps showing up near the top of citation studies. Video is not cited because models watch video. It is cited because video ships with a title, a description, and a transcript, which is 3 pieces of indexable text per upload. Spoken content is invisible until it is text on a page, so publish the transcript every time. Repurposing episodes is the same idea applied across surfaces.

Running the show yourself flips the direction entirely. Instead of pitching to be a guest and waiting, you invite the people you want to be associated with, and each recording produces the same artifact stack with your brand attached to every one of them. That mechanism is covered in podcast led outbound and the operating version in what a podcast acquisition system is. The commercial reason to do it, which has nothing to do with AI, is in podcast lead generation.

Two more sources worth naming. Publish original numbers from your own book, because models cite the source of a statistic and there is no way to be the source of a number you did not measure. Our cold email reply rate benchmarks and podcast lead generation benchmarks exist for that reason. And answer real questions where real people ask them, in your own name, without a link. That is the slowest item on the list and the closest thing to a moat, because it cannot be handed to a content team.

What to skip: press release syndication that puts the same paragraph on 200 low quality domains, paid listicles that rank nowhere, and any service selling AI mentions as a product. Repetition without independence does not build entity confidence. It builds a pattern that looks like exactly what it is.

How do you check whether any of this is working?

Two checks. One takes 20 minutes a month and the other takes a log file.

The monthly check is to ask the assistants directly. Open ChatGPT, Claude, Perplexity, and Gemini in clean sessions with no memory, and ask each one 4 questions: what does this company do, who is it for, how is it priced, and who are the alternatives. Keep the answers in a doc so you can see drift. Run all 4 assistants rather than the one you personally use, because they disagree more than people expect. One will have your category right and your mechanism wrong. Another will confidently attach a number to you that came from a competitor.

Log which assistant got which detail wrong, because the fix is almost always a specific page or a specific missing sentence rather than a general effort. Knowing which one saves you the afternoon.

The deeper check is server logs. Filter for GPTBot, ClaudeBot, PerplexityBot, and Google-Extended and look at what they fetch and how often. OpenAI documents its crawler behavior at platform.openai.com, and Cloudflare Radar is a decent public read on how AI crawler traffic is moving across the web generally. What you are looking for is whether the crawlers reach your deep pages at all, or only ever touch the homepage.

One thing not to do: buy a visibility score from a tool and treat the number as truth. No external tool can see inside a model. Every score is sampled prompts run on a schedule, which is a fine directional signal and a terrible reporting metric. Sample the prompts your buyers would actually type and read the answers yourself.

Where this actually lands

The companies that get cited are the ones other people describe. Everything else is a tiebreaker, and the tiebreakers only run once you have already made the list.

So the priority order is simple, even though the work is not. Get described accurately somewhere other than your own domain, repeatedly, in consistent language, by sources that had no obligation to describe you at all. Then make sure your own pages are clean enough that when the model comes to confirm the shortlist it already built, it finds a self contained block, a schema block, and a summary in the first 120 words instead of a hero image and a video autoplay.

The uncomfortable part is that this is not a marketing tactic with a 30 day window. Web mentions accumulate over quarters. Entity confidence is a slow variable by design, because a model that changed its mind about a company every week would be useless. Anybody promising to move it in a sprint is describing something else.

The other uncomfortable part: being described accurately by an assistant only pays once buyers have a reason to look you up. That reason comes from a pipeline, not from a schema block. Getting into the consideration set in the first place runs through why executives say yes to podcast invites, building the right guest list, and a sending setup that actually clears the filters, which is covered across email deliverability, SPF, DKIM, and DMARC, warming a new domain, and domain reputation. Who you point it at is the ideal customer profile question, worked through for outbound in defining an ICP for cold email. What happens after the recording is in turning guests into clients, and the wider numbers are in the state of AI outbound.

An assistant can only name a company it has heard about from somebody other than that company. Build the reasons for other people to talk. The citations are the receipt, not the strategy.

See How the Invite Engine Works

15 minute demo. No fluff. We will walk you through the exact system, show real prospect examples, and scope what it looks like for your market.

Schedule a Demo