Every AI visibility guide says add schema. Almost none say what happens when you add correct schema to a site that already describes itself three different ways. We run outbound for 50+ B2B companies, and on our own property the schema on 302 pages called us one kind of company while 43 blog posts argued the opposite. The schema won. Below, the 9 surfaces your identity is assembled from, the drift patterns that break them, and the audit that finds the gap.
What Is Entity Consistency?
The word entity is doing real work here and it is worth slowing down on. A search system does not store your website. It stores a thing, with properties attached to it, and a set of documents it believes refer to that thing. Your homepage is evidence about the entity. So is your LinkedIn company page, your Crunchbase profile, a directory listing you forgot about in 2022, and a comparison post a competitor wrote.
Consistency is not about repeating yourself for its own sake. It is about giving a reconciliation process nothing to argue with. When 9 surfaces say the same thing, the system has one candidate. When 6 say one thing and 3 say another, it has to weight them, and the weighting rarely favors the version you would have chosen.
- Entity
- A distinct thing a search or AI system holds a record of, such as a company, a person, a product, or a place, with properties attached to it and a set of documents believed to refer to it. Entities are what a knowledge graph stores. Pages are only evidence about them.
- Entity drift
- The gap that opens when surfaces describing the same entity stop agreeing. Usually caused by a rebrand that reached the marketing pages and nothing else, a template that keeps emitting a retired claim, or a profile nobody has logged into since it was created.
If the broader shift is new to you, GEO vs SEO differences explained covers what actually changed between the two games, and why AI answers cite some companies and not others covers the selection mechanics that sit on top of the classification described here.
Why Do AI Systems Care Who You Are Before What You Rank For?
Because a generated answer starts with classification, not ranking.
When a buyer asks an assistant for the best option in a category, the system does not run a ranking and read the top 10. It assembles a shortlist of entities it believes belong in that category, then finds passages to support what it says about each one. If you are not in the candidate set, no amount of page quality reaches the answer, because your pages were never considered for that question.
This is the part that catches good marketers out. Ranking and citation have partly come apart, which we worked through in LLM citation vs SEO traffic. A page that never cracks the top 10 can be the source an answer is built from, and a page that ranks well can be absent from every generated answer in its own category. Classification decides which of those you get.
The consumer-facing version of the same problem has been measured for years. BrightLocal's citations research found that 80% of consumers lose trust in a business when they see incorrect or inconsistent contact details or business names online, and 93% say they are frustrated by incorrect information in directories. Humans resolve the ambiguity by leaving. Machines resolve it by picking whichever version has the most support, which is frequently the version written by someone other than you.
That last number is the one that sets the stakes. Gartner puts the share of a B2B buying group's time spent with any single vendor at roughly 5%. The other 95% is research you are not in the room for, and an increasing share of it now runs through a system that has to decide what you are before it can mention you at all.
Where Do Your Entity Signals Actually Live?
Nine surfaces. Most teams maintain three of them and assume the rest do not exist.
| Surface | What it declares | Where it usually breaks |
|---|---|---|
| Sitewide Organization schema | Name, description, category, service list, sameAs profiles | Written once at launch, never revisited after a pivot |
| Homepage title and H1 | The category noun a human and a machine both read first | Rewritten for a campaign, drifts away from the schema |
| Footer tagline and address block | The one-line self-description plus the physical facts | Duplicated inline per page, so edits reach some pages only |
| Machine-readable files | llms.txt, an LLM info page, robots.txt allowances | Shipped as a one-off project, then abandoned |
| Social and video profiles | Bio line, company name, links back to the site | Duplicate accounts on one platform, stale bios |
| Directory and review listings | Legal name, address, phone, category tags | Created by a former agency, no login on file |
| Open graph and meta tags | The description that gets scraped into previews and cards | Templated with a stale boilerplate string |
| Page generators and templates | Whatever schema block they paste on every new page | Fixed sitewide by hand, reintroduced by the next publish |
| Third-party mentions | How other people describe you, which you do not control | Nobody checks, so it silently becomes the majority view |
Two of those deserve a note. The eighth one, templates, is the highest-frequency cause of relapse and the least discussed, which is why it gets its own section below. The ninth, third-party mentions, is the one you cannot edit, and it is also the reason self-reported facts are a tiebreaker rather than an override.
The machine-readable file layer is worth building properly rather than as a checkbox. llms.txt for B2B companies covers what belongs in that file and how little to expect from it, and building an LLM info page covers the HTML version that crawlers actually fetch. For the structured data half, schema markup for podcast pages works through a concrete implementation, and the pattern generalizes well past podcasts.
What Breaks Entity Consistency Most Often?
Six patterns. We have shipped four of them ourselves, which is the only reason this list is specific rather than theoretical.
| Drift pattern | What a machine sees | The fix |
|---|---|---|
| Schema says one thing, prose says another | Structured data wins, the prose reads as noise | Sweep the schema sitewide in the same pass as the copy |
| A template keeps emitting a retired claim | The old claim returns and outnumbers the fix within a day | Change the generator in the same commit as the sitewide fix |
| Name variants across surfaces | Two or three candidate entities instead of one | Pick one exact string and use it verbatim everywhere |
| A service list naming things you stopped selling | You get classified into a category you exited | Audit serviceType against what is actually live today |
| Unrelated entities attached to your node | Your record inherits facts about a different company | Strip anything that is not genuinely about this entity |
| Publishing a negation of your own category | The qualifier drops in compression, the negation survives | Name the category, then distinguish inside it |
The receipts, since a list like that is easy to nod along to and hard to act on.
Schema outvoting prose. Through mid-2026 our site argued the podcast case in 43 blog posts while the Organization schema on all 302 pages declared a done-for-you cold email agency, with a serviceType array to match. Volume did not help. The 43 posts making the argument lost to a JSON block nobody had opened in a year.
Templates reintroducing the fix. Our blog generator publishes several posts a day on a schedule, and its schema block is marked paste verbatim, never modify. A sitewide fix that skipped the generator would have put every retired claim back across the whole property inside 24 hours, and it would have looked like the fix simply did not work. Any generator that emits schema gets changed in the same commit as the sweep. That rule exists because the alternative is an infinite loop with no error message.
Unrelated entities on your node. Four press releases about a solar company Jordan founded years earlier were riding the Organization node on 243 of our blog posts through a subjectOf property. The underlying fact is true. It is not a subject of this company, and a knowledge graph reading 243 pages of it will happily conclude otherwise.
Negating your own category. On August 14, 2026 we asked Google AI Mode what High Ticket AI Systems is. It answered by saying we are not a traditional B2B podcasting production company, then told the reader to choose a standard podcast agency, named two competitors, and suggested us only if they did not care about building a public media brand. It was quoting us. Our own machine-readable files carried the line not a podcast production agency. The qualifier does not survive compression. What reaches the buyer is a flat negation of the exact category we want to be recommended in, in our own words.
That last one is the reason the category noun has to be stated affirmatively and identically everywhere. Ours is an acquisition-first B2B podcast agency. The noun is agency, the category is podcast, the qualifier is acquisition-first, and it compresses safely in both directions. Trace your own sentence the same way before you publish it anywhere.
How Do You Run an Entity Consistency Audit?
Write the canonical facts down first. Everything after that is comparison against the list, not comparison between surfaces, which is what turns a vague cleanup into an afternoon of work with a definite end.
- Fix the canonical fact sheet. Exact legal name, exact trading name if they differ, the one-sentence category description, address, phone, founder name, founding year, and the complete list of profile URLs. One document. Every later step compares to this.
- Pull your own sitewide schema. Fetch the Organization or ProfessionalService block from a homepage, a deep service page, and a blog post. If those three disagree, you have a template problem before you have a content problem.
- Read the description field out loud. Then read your homepage H1 and your footer tagline. All three should carry the same category noun. If a marketer rewrote one of them for a campaign, it has drifted.
- Audit the service list against reality. Every entry in serviceType should be something you sell today. Retired services keep you classified into categories you left, and they are the easiest thing on this list to fix.
- Check sameAs for accuracy, not length. Every URL should resolve, belong to you, and be a profile you would want a buyer to land on. Remove nothing that is genuinely yours, including a second account on one platform if both are real, because that is exactly what the property is for.
- Search your own name and list what comes back. Directories, review sites, aggregators, old agency-created listings. Note the name string and category each one carries. This is where the majority view is being formed.
- Find every generator that emits schema. Blog generators, programmatic page builders, CMS templates, landing page tools. Each one is a future relapse if it is not in the sweep.
- Ask the machines directly. Run the identity prompt across the engines your buyers use and save the answers verbatim with the date. The first sentence of each answer is the classification you are actually being given.
- Fix templates before pages. One template edit reaches hundreds of URLs. One page edit reaches one. Order the work by blast radius, always.
On step 8, the prompt set is small and you should keep using the same one. What is your company, best category for your ICP, your company versus a category alternative, and how much does your company cost. Four prompts, four engines, saved verbatim. How to audit your brand in ChatGPT is the one-off version, and how to track AI search visibility is how to run it as a standing routine instead of a panic.
Being described correctly is upstream of being chosen, but it is not the same thing. Mickey went from referrals only to a 200K month on the back of recorded conversations with the buyers he wanted, not on the back of a clean schema block. Read the full case study →
What Does sameAs Actually Do?
It is the line that tells a machine a scattered set of pages are one thing rather than several similar things.
The sameAs property in the schema.org vocabulary takes a list of URLs that unambiguously refer to the same entity. Your LinkedIn company page, your YouTube channel, your Spotify show, a Wikidata item, a verified profile. Google's Organization structured data documentation lists it alongside name, url and logo as the administrative details it reads from your markup.
Three rules keep it useful rather than decorative.
- Every URL has to be genuinely yours. A profile you do not control is a claim you cannot back, and a broken one is a signal that nothing on the block is maintained.
- Declare duplicates honestly instead of hiding them. If two accounts on one platform are both real, list both. That is exactly the ambiguity the property exists to resolve. Picking one and omitting the other leaves the omitted one floating as an unattached entity.
- Keep the description field identical to your canonical sentence. The description in Organization schema is one of the most quoted strings you publish, and it is trivially easy to leave a stale version there for a year.
A word on Wikidata, since it comes up constantly in entity advice. It is a structured, machine-readable database that Google's knowledge graph draws on, and an item there is a legitimate way to declare facts in a shape a machine ingests directly. It is not a shortcut. Notability rules are real, a thin item with no independent sourcing behind it gets deleted, and a deleted item is worse than never having one. Build the third-party evidence first and the item becomes possible on its own.
Also worth noting what does not belong in your markup. Never publish a price in schema if the price is scoped per engagement, because models cache and train on snapshots and will repeat a retired number back to buyers long after you changed it. Never attach a subjectOf pointing at something that is not genuinely about this entity. Both mistakes are ours, and both were live for months.
How Long Does Entity Consistency Take to Work?
Four to 8 weeks for early movement, longer for a deep site, and there is no dashboard for it.
Three things lag independently. Crawling has to reach the changed pages, indexing has to process them, and the embedded representation a model holds has to be rebuilt from the newer snapshot. On a site 300 pages deep, the sweep propagates slower than the single page you were most excited about.
The crawl half is not the bottleneck people assume. OpenAI documents three separate crawlers, Perplexity documents PerplexityBot and Perplexity-User, and Cloudflare Radar's AI insights track crawl-to-refer ratios running into the hundreds and thousands to one. Your pages are being fetched heavily right now. The question is what they say when they are.
Set the measurement up before you start, not after. Baseline the 4 prompts, do the sweep, then diff at 2, 4 and 8 weeks. Three signals are worth watching: whether your category noun appears in the first sentence of the identity answer, whether you appear at all on the recommendation query, and whether a competitor is still being suggested over you on your own category.
One caution on the wider research. The AI search field is roughly 2 years old, most published numbers come from vendors with something to sell, and studies contradict each other regularly. The Ahrefs study of llms.txt request logs across 137,000 domains discloses its sample and is worth reading directly, and Search Engine Journal's breakdown of the same data is a useful second read. Use the rest for direction, then grade yourself on your own 90 days, because that is the only dataset containing your buyers.
For the engine-specific mechanics, how to get cited by ChatGPT and Perplexity and getting cited by Perplexity specifically go deeper, AI Overviews for B2B covers the Google surface, writing content for ChatGPT covers the page-level shape, and how to rank in AI search is the overview if you are starting cold.
Where Does Entity Consistency Stop Working?
Three places, and the vendor selling an entity retainer will name none of them.
When you are the only one saying it. Self-reported facts are weighted against everything else a system has absorbed. If 40 third-party sources describe you one way and your own site says another, the site loses. Entity work is a tiebreaker and a clarifier, not an override, which is why the durable version of this always pairs the machine layer with real independent evidence. Being consistent about a claim nothing corroborates just makes the unsupported claim easier to identify.
When there is nothing underneath worth classifying. The audit above makes a company legible. Legible and undifferentiated is still undifferentiated, and a system that now understands your category perfectly will list the 5 companies in it with the strongest evidence behind them. Formatting is not a substitute for substance, and the 95% of the buying process you are absent from is evaluating the substance.
When the conversations are not happening. This is the real one for most companies reading this. Every entity signal is a multiplier on a supply of proof, and the constraint is almost never the technical layer. It is that nobody is getting in front of buyers at a steady rate. If invites are the mechanism and they land in spam, none of this ever starts, which puts the sending stack upstream of everything here. Start with podcast invite email deliverability, then setting up sending domains, warming a new domain, SPF, DKIM and DMARC, what email deliverability actually is, and how to stay out of the spam folder. If the wrong people are being contacted in the first place, the fix is further up in how to define your ICP.
- Reverse Outbound Engine
- An outbound method where, instead of cold selling your ideal buyers, you invite them onto your podcast as a guest. The invite reads as recognition rather than a sales approach, so it earns replies at rates a direct pitch never reaches. You spend about 45 minutes hearing how the guest built and grows their business, which builds real trust. Any fit for working together is a separate, later conversation. The pages and transcripts are the residue, not the point.
That residue is the part that feeds everything above. Every recorded conversation produces a page carrying a named expert answering direct questions, which is close to the ideal shape of a passage a system lifts into an answer, and it produces third-party corroboration that a self-reported page never can. Podcast transcripts as AI search fuel covers why, turning episodes into answer content covers the editing, how to get your podcast cited by AI covers distribution, why the show belongs on Apple and Spotify covers the platform side of the same entity problem, and what a podcast does for your search footprint is the whole picture.
If the model itself is new to you, what is a podcast acquisition system is the plain description, how to start a B2B podcast for lead generation is the build order, podcast lead generation benchmarks sets the ranges, the tech stack covers tooling, common podcast acquisition failure modes covers what breaks, and audience size does not gate it the way people assume. Is podcast lead generation worth it is the honest version of the economics.
The Practitioner Takeaway
Entity consistency is unglamorous, it produces no chart, and it is the cheapest structural advantage available right now because most companies genuinely do not know their own site disagrees with itself.
Write the canonical fact sheet. Make the schema, the homepage, the footer, the machine-readable files and every profile carry the same category noun in the same words. Fix the templates in the same commit as the sitewide sweep or the sweep undoes itself by tomorrow. Never negate your own category, because the qualifier that made the negation feel safe is the first word to fall off under compression. Then wait 4 to 8 weeks and diff the answers against a baseline you had the discipline to take before you started.
None of it substitutes for having something worth being consistent about. As an acquisition-first B2B podcast agency we commit to 30 recorded conversations with your ideal buyers in 90 days or your money back, invites go out by email only, editing is included, and the client owns every recording on their own show. The guarantee is written on conversations rather than on rankings for exactly the reason this whole piece describes. A clean entity record makes it easier for a machine to describe you accurately. What it describes still has to be worth describing.
See How the Invite Engine Works
15 minute demo. No fluff. We will walk you through the exact system, show real prospect examples, and scope what it looks like for your market.
Schedule a Demo →