Every GEO checklist starts with a page. Add the schema, tighten the intro, wait to get cited. We publish more than 260 articles on outbound, and the ones that get named in AI answers are almost never the ones we tuned hardest, they are the ones sitting inside a finished cluster. Below, why retrieval rewards coverage over polish, and the build order for a cluster a model treats as the source.

What Is Topical Authority in AI Search?

Topical authority in AI search is how consistently a model finds your domain across every angle of a subject, not how well one page ranks for one query. Assistants split a question into several sub-queries and retrieve separately for each one. The domain that keeps surfacing gets treated as the source and gets cited.

The phrase is borrowed from SEO, and the borrowing causes most of the confusion. In classic search, topical authority was a soft ranking input. One page still competed against one page, and covering a subject broadly was a way to help that page win.

In AI search, coverage is the mechanism itself. A model answering a question does not run one query and read the top result. It decomposes the question, runs several retrievals, and assembles an answer from whatever came back. Your page is not competing for a slot. Your domain is competing to be present across a spread of sub-queries you never see.

Topical authority
The degree to which a retrieval system finds one domain relevant across the full question space of a subject. It is measured by how often that domain appears among the sources pulled for related sub-queries, and it is subject specific. A site can hold deep authority on one topic and none at all on the topic next to it.
Query fan-out
The step where an assistant rewrites one user question into several narrower searches, retrieves documents for each, and merges the results before writing. It is why a question you never targeted can surface your page, and why a page that answers only the headline question often gets left out of the merge.

Two things follow, and both are uncomfortable for anyone who has spent a decade on page level work. Polish on a single page has a low ceiling, because that page can only satisfy the sub-queries it happens to cover. And a thin page inside a complete cluster often outperforms a heavily worked page standing alone, because it is the only document that answers its specific sub-query cleanly.

Why One Strong Page No Longer Wins

The mechanics here are not a theory about what models prefer. They are a description of how retrieval augmented generation works when you watch it run.

Ask an assistant which agency to hire for B2B podcast lead generation and it will not search that string. It will search for podcast lead generation agencies, then for what podcast lead generation costs, then for whether it works, then for alternatives, then for complaints. Five retrievals, five different sets of documents, one answer written from the union.

A company with one flagship page shows up in one of those five. A company with a page for each shows up in all five, and the model reads that repetition as authority rather than coincidence. Semrush analysis of 126 million AI search prompts found that sites with a clear topical focus are cited roughly 2.4 times more often than generalists, which is the same effect measured from the outside.

This is also why the training question is a distraction. The bots doing this work are not the training crawlers. OpenAI documents OAI-SearchBot as a separate agent from GPTBot, and Anthropic documents Claude-SearchBot the same way. Retrieval happens live, against whatever is published today. Your cluster does not have to wait for a model to be retrained to start getting cited, which is the single most encouraging fact in this entire subject.

The volume of sourcing matters too. ChatGPT cites an average of 15 sources per response, where Gemini cites around 3. The wider the source list, the more room there is for a focused domain to appear more than once inside a single answer, which is the strongest signal a model can send that it considers you the reference.

There is a second reason to care, and it is the one that changes the budget conversation. Most AI reading never produces a click. Pew Research tracked 68,879 real Google searches and found users clicked a link on 8% of pages carrying an AI summary against 15% of pages without one, and clicked a link inside the summary itself just 1% of the time. The citation is the impression. If you are still grading this work on sessions, you are measuring the 1% and ignoring the rest, which we unpack in LLM citation vs SEO traffic.

The same gap shows up from the infrastructure side. Cloudflare now publishes a crawl to refer ratio on Radar, tracking how many pages each AI platform fetches for every visitor it sends back. The ratios are lopsided and they are getting worse. Reading that as theft is one option. Reading it as a channel where the impression is the product, and the click was never the point, is the more useful one for a company that sells a service rather than pageviews.

None of that is a reason to stop caring about search. Gartner predicted in 2024 that search engine volume would fall 25% by 2026, and plenty of practitioners have argued the number was too aggressive. The prediction being early does not change the direction. Buyers are asking assistants questions they used to type into a search bar, and the answer they get names two or three companies rather than ten blue links. The shortlist got shorter, so being on it matters more.

How Is Topical Authority Different From Domain Authority?

Domain authority is a third party score that travels with your whole site. Topical authority is earned per subject and does not transfer. A publisher with enormous link equity can lose an AI citation on a narrow B2B question to a site nobody has heard of, because the small site answered the sub-query and the large one did not.

Get outbound insights, weekly
Tactics, benchmarks, and playbooks from 50+ B2B outbound campaigns. No spam, unsubscribe anytime.
You are in. Check your inbox.

Holding both models in your head at once is the hard part, so here is the split written out.

Dimension Domain authority (classic SEO) Topical authority (AI search)
Unit that competes One page against one page for one query One domain against one domain across a spread of sub-queries
Primary signal Links pointing at the domain Coverage depth and consistency across the subject
Does it transfer between topics Yes, it is site wide No, it is earned per subject
What a thin page does Dilutes the site and risks a quality problem Helps, if it is the only clean answer to its sub-query
How you win Make the strongest page on the query Leave no question in the space unanswered
How you check it Rank tracking on a keyword list Prompt tracking on a buyer question set
What success looks like Position 1 and a click Named in the answer, click optional
Time to move Weeks to months 90 to 180 days, because indexing and embedding both lag

The row that trips people up is the fourth. A thin post is a liability in classic SEO and an asset in retrieval, provided it is genuinely the cleanest answer to one specific question. That is not permission to publish filler. It is permission to publish a 900 word page that answers exactly one thing instead of holding it back for a 4,000 word guide that never gets written. The broader split is covered in GEO vs SEO.

How Many Pages Does a Cluster Actually Need?

Everyone wants a number. There is not one, and chasing a number is how content teams end up with 40 posts that all answer the same question in slightly different words.

The right unit is the question, not the post. Write down every question a real buyer asks between first hearing about your category and signing, then check which ones you have a dedicated page for. Most B2B topics resolve into 15 to 40 distinct questions once they are on paper, which is why finished clusters tend to land in that range. The number is an output, not a target.

Four question types keep showing up, and a cluster missing any one of them has a visible hole:

  1. Definition questions. What is this, what is it not, what is the difference between this and the adjacent thing. These are the most citable pages you will ever publish, because they are what a model reaches for when a user asks a category question cold.
  2. Evaluation questions. Does it work, is it worth it, what does it cost, what goes wrong. Buyers ask assistants these before they ask a vendor, and a company that refuses to answer them publicly is simply absent from that part of the conversation.
  3. Comparison questions. This versus that, this versus doing nothing, this versus hiring someone. Tables get pulled into answers at a high rate because they are already structured, which we cover in why comparison tables get cited by AI.
  4. Execution questions. How do I do it, in what order, what do I need first. These convert worst and cite best, because they are where the specifics live.

Write one page per question. Do not let a page cover two, because a page covering two answers neither in a way a retriever can lift cleanly. The mechanics of tying them together are in internal linking for LLM retrieval.

What Does a Finished Cluster Look Like?

Ours on AI search visibility runs about 20 pages, and it is easier to show than describe. The definition layer covers how to rank in AI search and what GEO changed. The execution layer covers FAQ schema, entity consistency, llms.txt, robots.txt for AI crawlers and an LLM info page. The measurement layer covers auditing how ChatGPT describes your brand, tracking AI search visibility and refresh cadence. The negative space is covered too, in what content AI assistants skip.

Each page answers one question. None of them is the definitive 5,000 word guide to AI search, and that is the point. When an assistant fans a question about AI visibility into six retrievals, we are eligible for six of them instead of one.

Three properties separate a cluster from a pile of posts on the same theme:

1 page
per buyer question, with no overlap between pages
2.4x
more citations for focused sites than generalists, per Semrush
15
sources cited in an average ChatGPT response

The first property is coverage, which is above. The second is consistency. Every page has to describe the same company in the same words, because a model assembling an answer from five of your pages is also deciding what kind of company you are. Describe yourself three different ways across three pages and the model picks one, usually the least flattering. This is the entity problem, and it is the most common own goal we see.

The third is structure. A page that leads with a short standalone answer, defines its terms, and puts comparisons in a table is easy to lift. A page that buries the answer under 600 words of context is not. The Princeton and Georgia Tech GEO study tested nine content strategies across a 10,000 query benchmark and found that adding relevant statistics and quotations from credible sources produced the biggest visibility gains, with the best methods beating the baseline by roughly 41% on one of the two headline metrics. Structure and sourcing are not garnish. They are the two levers with published evidence behind them.

Mickey Hardy built his authority the same way, through recorded conversations rather than a content calendar, and went from referrals only to a $200K month. Read the full case study →

Where the Citations Actually Come From

Here is the part most GEO advice skips. A 20 page cluster is 20 pages of writing, and a 40 page cluster is a full time job for a quarter. The strategy is sound and almost nobody executes it, because the bottleneck was never the plan. It was the source material.

Writing 30 posts from your own head produces 30 posts that read like your own head. They are fine, they cover the questions, and they contain nothing a competitor could not also write. Models notice. Coverage of the retrieval shift keeps landing on the same conclusion: undifferentiated summary content is exactly what a model can now generate itself, so it has no reason to cite yours.

What a model cannot generate is a specific operator saying a specific thing in their own words. Named people, real numbers, direct quotes, and the messy detail of how something actually went. That is the material the GEO research found lifts citations, and it is the material a blog post written alone in a doc never contains.

The cheapest way to produce it at volume is to record conversations with people who already have the expertise. One 45 minute interview yields a transcript, an episode page, several question level pages and a stock of quotable primary material with a real name attached. That is a week of a writer's output from an hour of somebody's time, which is the argument in podcast transcripts as AI search fuel and how to get your podcast cited by AI.

It also compounds in a way writing does not. Every guest is a named entity with their own footprint, and an episode page connects your domain to theirs in a way a model can follow. The search side effects are laid out in what a podcast does for your search footprint, and the markup that makes those pages legible is in schema markup for podcast pages.

How Do You Build a Cluster Without a Content Team?

The order matters more than the effort. Most teams start by writing and stall around post 9. Start by producing source material and the writing becomes editing, which is a job a much smaller team can sustain.

  1. Write the question list first. 20 to 40 buyer questions, pulled from sales calls and from what buyers actually type. Not keywords. Questions. If you have not defined who is asking, start at how to define your ICP, because a question list for the wrong buyer builds authority nobody is searching for.
  2. Book the conversations. Invite the operators who live inside those questions onto a recorded interview. The invitation converts far better than a pitch does, which is the whole premise of invite versus pitch, and the mechanics are in how to invite guests to your B2B podcast.
  3. Publish the episode as a page, not a player. Full transcript, real headings, a summary at the top. An embedded audio player with a 40 word description is invisible to retrieval.
  4. Cut the transcript into question pages. Each question on the list gets one page, built around what the guest actually said, with the quote and the name attached. This is where repurposing stops being a social exercise and starts being an authority one.
  5. Wire the cluster together. Every page links to the others that answer adjacent questions, in descriptive anchor text. Keep the entity description identical everywhere, and publish the machine readable layer: schema on every page, an llm-info page, an llms.txt written per this guide, and a robots.txt that actually lets the retrieval crawlers in.
  6. Then measure, and only then tune. Run the question list against the assistants monthly and fix the questions where you are absent.

Step 2 is the one that decides whether the rest happens. Everything after it is production work with a clear finish line. Getting 30 credible operators to sit down with you is the part that needs a system, and it is the part we run for clients: the invites go out by email, the recordings land on the client's own calendar, and every episode is edited and published. The guarantee is 30 recorded conversations with your ideal buyers in 90 days, or your money back. The client owns the show and owns every recording.

If the invitation side is new to you, does podcast lead generation actually work and the first 30 days of a podcast acquisition system cover what the first quarter looks like. If the sending infrastructure is new, start with email warmup, because invites that land in spam produce neither guests nor pages.

How Long Does It Take, and What Do You Measure?

Plan on 90 to 180 days before the answers move in a way you can point at. Two lags stack on top of each other. Crawling and indexing take days to weeks, and re-embedding a corpus into a retrieval index takes longer still, which is why a single page fix usually shows nothing for 4 to 8 weeks.

That lag is the reason most companies quit at week 5, and the reason the ones who do not are still holding the position two years later. It is a slow moat, which makes it a real one.

Measure it with a fixed prompt set, not with analytics. Take the same 20 to 40 questions, run them across ChatGPT, Claude, Gemini, Perplexity and Google AI Mode on a schedule, and record four states per question: named and linked, named only, linked only, or absent. Share of answers is the number that matters. The method is written up in how to track your AI search visibility.

Watch the linked only state closely. Semrush found that 62% of AI citations are what they call ghost citations, where an engine uses a domain as a source but never names the brand in the text. Those are worth having and they are worth knowing about, because a dashboard counting only brand mentions will report failure on work that is landing. The fix is usually entity level rather than content level, and it starts with entity consistency and auditing how the assistants describe you.

One more measurement note. Do not grade this on traffic. Given that only 1% of users click a link inside an AI summary, a channel that is working can look flat in analytics for months. Judge it on whether you are in the answer, which is the argument in LLM citation vs SEO traffic.

The Call We Would Make

If you have 10 posts scattered across 6 topics, do not write an eleventh. Pick the one topic you want to be known for, write the question list, and finish it. A complete cluster on one subject beats partial coverage of six, every time, because retrieval is deciding per subject and partial coverage loses every one of them.

If you have already picked the topic and the writing is where it stalls, the problem is supply, not strategy. Go get the source material by talking to people who know things, publish what they said, and let the cluster assemble itself out of the transcripts. That is the same reason we run guest conversations as an acquisition channel rather than a content one. The pages are the byproduct. The conversations were always the asset.

What nobody should do is keep tuning one page and waiting. That page is competing for a slot that no longer exists, in a system that stopped asking one question at a time. The work moved up a level, from the page to the subject, and the companies that notice first get a two year head start while everyone else is still editing a title tag.

See How the Invite Engine Works

15 minute demo. No fluff. We will walk you through the exact system, show real prospect examples, and scope what it looks like for your market.

Book A Call