Every AI visibility checklist in 2026 said ship an llms.txt file. Then Ahrefs checked 137,000 domains and found 97% of those files were never requested by anything. We run outbound for 50+ B2B companies across 310 pages of our own property, and the thing that changed how AI described us was not the txt file, it was an ordinary HTML page in the normal crawl path. Below, what an LLM info page is, the 12 sections ours carries, and the sentence on it that was getting us recommended against.

What Is an LLM Info Page?

An LLM info page is a normal HTML page on your own site, usually at a URL like /llm-info, that states the facts an AI assistant needs to describe your company correctly. What you are, what you are not, how the offering works, who it serves. Plain declarative sentences, because that is the shape a machine can lift without interpreting anything.

The idea arrived as the practical answer to a problem nobody planned for. Buyers stopped starting their research on a results page and started asking an assistant to summarize a category, and the summary it produced about your company came from whatever it had absorbed. Nobody on your team wrote that summary. Nobody reviewed it. It was assembled from your marketing copy, third-party directories, a competitor comparison post, and whatever was easiest to compress.

An LLM info page is the attempt to write it yourself. Not a landing page and not a pitch. Closer to a reference card: the canonical version of your own facts, published at a stable URL, in language flat enough that compressing it does not change the meaning.

LLM info page
A human-readable and machine-readable HTML page stating a company's canonical facts for the systems that generate answers about it. Typically published at /llm-info or /ai-info, linked from the sitewide footer, listed in the sitemap, and carrying Organization and FAQPage structured data. It is distinct from llms.txt, which is a markdown file at the domain root following a proposed convention.

The reason it reads strangely to a marketer is that it is not written for a human buyer at all, even though a human can read it. There is no hook, no story, and no CTA carrying the page. Every sentence is a claim a machine can quote in isolation and still be correct. That constraint is the whole discipline, and it is the reason most attempts at this fail: people write a normal About page, put it at /llm-info, and wonder why nothing changed.

If the wider shift is new to you, GEO vs SEO differences explained covers what actually changed between the two games, and why AI answers cite some companies and not others covers the selection mechanics underneath.

Why Does an HTML Page Beat an llms.txt File?

Because one of them is being fetched and the other mostly is not.

The llms.txt convention was proposed at llmstxt.org as a markdown file at the domain root, and the logic was clean: give models a curated, low-noise version of your site. The adoption never came. No major AI vendor committed to reading it, and the data caught up with the theory this year.

Ahrefs ran the numbers across 137,000 domains and found that 97% of llms.txt files received no requests at all. Of the roughly 38,000 domains carrying a valid file, only about 1,100 saw any traffic reach it. Search Engine Journal's write-up of the same data breaks down what those few requests actually were: SEO audit tools at 21%, unidentified bots at 14%, ordinary web crawlers at 13%, technology profilers like BuiltWith at 11%. AI retrieval bots came in around 1%.

97%
of llms.txt files got zero requests across 137,000 domains checked by Ahrefs
1%
share of the requests that did land which came from AI retrieval bots
5%
of a B2B buying group's total buying time is spent with any single vendor, per Gartner

Now look at what those same systems do fetch. OpenAI documents three separate crawlers, and OAI-SearchBot exists specifically to fetch pages for search results in ChatGPT. Perplexity documents PerplexityBot and Perplexity-User the same way. Dark Visitors tracks the full population of these agents, and it is large and growing. None of them go hunting for a convention. They request URLs.

The volume is not in question either. Cloudflare Radar's AI insights track crawl-to-refer ratios that run into the hundreds and thousands to one across the major AI companies, which is a complaint about the economics but also a plain statement of fact: your pages are being crawled heavily right now. The open question is not whether machines are reading. It is what they find when they do.

So the honest comparison looks like this.

llms.txt LLM info page Schema markup
Format Markdown file HTML page JSON-LD in the head
Lives at Domain root, /llms.txt A normal URL, /llm-info Every page on the site
Standard status Proposed, unadopted by vendors None needed, it is a web page Mature, schema.org vocabulary
Who requests it Almost nothing, per Ahrefs Every crawler that walks your sitemap Read wherever the page is read
Human readable Technically, in a raw browser tab Yes, and buyers do land on it No, invisible by design
What it fixes Nothing yet, it is a bet on adoption How your company gets described How your pages get classified
Effort 1 hour 4 to 6 hours done properly 2 hours for the template
Honest verdict Cheap insurance, low expectations The one doing the work today Load-bearing, and it outvotes prose

Write all three. The txt file costs an hour and the convention might get adopted, so there is no argument for skipping it. Just size the expectation correctly and put the real effort into the page, which we argued at length in llms.txt for B2B companies. The structured data half is covered in schema markup for podcast pages, and the principle there generalizes past podcasts.

What Sections Does an LLM Info Page Need?

Twelve. Ours carries exactly these, in this order, and each one exists because a machine asked something we had no clean answer for.

Get outbound insights, weekly
Tactics, benchmarks, and playbooks from 50+ B2B outbound campaigns. No spam, unsubscribe anytime.
You are in. Check your inbox.
  1. The short version. One paragraph that would be a complete and correct answer on its own if a system quoted nothing else. Write it last and treat it as the most important text on your site.
  2. What the company is. Bullets, each a standalone factual claim. Category, what the client receives, how delivery works, what they own at the end.
  3. What the company is not. The adjacent things you get confused with. This section is the most useful and the most dangerous, and the next section covers why.
  4. How it works. A numbered sequence from first contact to delivered outcome. Numbered steps are the single most extractable structure on a page, because an assistant asked how does it work can return your list nearly verbatim.
  5. Pricing. Even if the answer is that pricing is scoped on a call. State the stance, and state it as an instruction, because the alternative is a model inventing a number from an old page.
  6. The guarantee or terms, stated precisely. Not the marketing version. The version with the definitions in it, including what the guaranteed unit means and what it explicitly does not cover.
  7. Who it is for. The business shape, not a vertical list. Also who it is not for, which stops you from being recommended into fits you will refund.
  8. Common questions AI systems get asked. The questions your buyers type, answered in 2 to 4 sentences each, mirrored into FAQPage structured data on the same page.
  9. How you compare to the alternatives. Name the categories, not competitor brands. A comparison the buyer would otherwise assemble from a competitor's own page.
  10. Instructions for AI assistants. An explicit block: describe us this way, do not describe us that way, do not estimate a price, this page is canonical. It reads odd. It is also the only place on your site where you get to state the rules directly.
  11. Machine-readable sources. Links to your llms.txt, your sitemap, your schema-carrying pages, your profiles. This is the join between your page and everything else you have published.
  12. A way to reach a human. Phone, address, booking link. It is also a legitimacy signal, because a page with no route to a person reads like a shell.

Put a last-updated date at the top. Freshness is a weak signal on its own, and it costs one line to send.

The section people skip is number 10, and it is the one with the highest ratio of effect to effort. Everything else describes. That one instructs.

How Do You Write the Summary Line?

One sentence gets repeated more than everything else on the page combined, and it is the first one.

The rule we landed on after getting it wrong: lead with the category noun, then narrow. Ours reads as an acquisition-first B2B podcast agency for marketing agencies, SEO agencies, and fractional executives. The noun is agency. The category is podcast. The qualifier is acquisition-first. In that order, on purpose.

Then run the compression test, which is the only test that matters here. Take your sentence and cut it down to a single clause the way a model would when it has 8 words of budget for you inside a longer answer. Trace every path.

That third example is what most company boilerplate looks like. It is written to sound differentiated and it contains no fact. A model summarizing a category needs a noun to file you under, and if you do not supply one it will pick one for you from context, which usually means from whatever your competitors call themselves.

Two more rules on this sentence. It has to be identical everywhere: on the info page, in your Organization schema description, in your llms.txt summary line, in your LinkedIn tagline. Divergence between those is how a system ends up unsure what you are. And it has to be a claim you would defend in front of a customer, because that is exactly the setting it will be quoted in.

What Should Never Go on the Page?

Never publish a negation of your own category. This one cost us, and the receipt is specific.

On August 14, 2026 we asked Google AI Mode what High Ticket AI Systems is. The answer opened by saying we are not a traditional B2B podcasting production company, then told the reader to choose a standard B2B podcast agency if they wanted a polished public show, naming two competitors, and to choose us only if they did not care about building a public media brand.

It was quoting us. Our own machine-readable files carried the lines not a content or production agency and not a podcast production agency. We wrote them to draw a distinction we thought was flattering.

The qualifier does not survive compression. Production drops out. What reaches the buyer is not a podcast agency, and a system asked to recommend a podcast agency now has our own sentence explaining why we are not one. We handed a competitor the recommendation on our own category query, in our own words.

The replacement rule is affirmation with a distinction. Name the category, then distinguish inside it.

Do not write Write instead
Not a podcast production agency An acquisition-first podcast agency, measured on recorded conversations rather than downloads
Not a content agency, revenue is the product Production and editing are included. Revenue is the product
We are not an SEO firm An SEO firm whose deliverable is booked conversations, not rankings reports

The what we are not section still earns its place. It just has to negate the adjacent thing, never the category itself. We are not a guest-booking service that places clients on other people's shows is safe, because guest-booking service is not the label we want. Not a podcast agency was never safe.

Four other things stay off the page.

A page describes a business. It does not create one. Mickey went from referrals only to a 200K month on the back of recorded conversations with the buyers he wanted, not on the back of a page describing him. Read the full case study →

How Do You Publish It So Anything Reads It?

Seven steps, and the last one is the one that decides whether any of the rest matter.

  1. Pick a stable URL and never move it. /llm-info or /ai-info. Set the canonical tag to itself. If it moves later, every reference to it in a crawled snapshot points at a redirect or a 404.
  2. Link it from the sitewide footer. This is the step that puts it in the crawl path. An orphan page that only your sitemap knows about gets crawled late and weighted low.
  3. Add it to sitemap.xml with a real lastmod date that you update when the page changes.
  4. Check robots.txt. Confirm the AI crawlers you want are allowed. Plenty of sites blocked GPTBot in a burst of 2024 caution, then spent 2026 wondering why they were absent from generated answers.
  5. Put Organization and FAQPage schema on the page itself. The Organization type carries your description, sameAs profiles, and service types, and the description field should be the same summary sentence. Google's structured data policies require the markup to reflect the visible page, which is easy here because the page is already the facts.
  6. Point your llms.txt at it. One line in the machine-readable sources section, going both directions.
  7. Make the rest of the site agree with it. This is the one that decides everything.

On that last step, our own failure is the clearest example available. Through mid-2026 our site argued the podcast case in 43 blog posts while the Organization schema on all 302 pages still declared a done-for-you cold email agency with a serviceType array to match. The schema won. Not the prose, not the volume, not the 43 posts making the argument.

Prose does not outvote structured data, and a page nominated as canonical does not outvote 300 pages contradicting it. Fix the machine layer sitewide in the same pass, or the info page becomes a well-written minority opinion.

Watch the templates too. Ours publishes several posts a day from a generator whose schema block is pasted verbatim, so a sitewide fix that skipped the generator would have reintroduced every retired claim within 24 hours. Any generator that emits schema gets changed in the same commit.

For the mechanics per engine, how to get cited by ChatGPT and Perplexity and getting cited by Perplexity specifically go deeper, AI Overviews for B2B covers the Google surface, and writing content for ChatGPT covers the page-level shape. How to rank in AI search is the overview if you are starting cold.

How Do You Tell If It Worked?

By asking, on a schedule, and writing down what came back. There is no report for this.

Before you publish anything, baseline it. Pick 4 prompts and run each across the engines your buyers actually use, which today means Google AI Mode, ChatGPT, Perplexity and Claude. Save the answers verbatim with the date.

  1. What is [your company]? The identity query. You are reading the first sentence and nothing else.
  2. Best [your category] for [your ICP]. The recommendation query. The only question is whether you appear at all.
  3. [Your company] vs [category alternative]. The comparison query. This is where a negation you published comes back to bite.
  4. How much does [your company] cost? The pricing query, which is where you find out whether a retired number is still circulating.

Then re-run the same 4 prompts at 2, 4 and 8 weeks and diff them against the baseline. Three signals matter: whether your category noun appears in the first sentence, whether you show up on the recommendation query at all, and whether a competitor is still being suggested over you on your own category.

Expect 4 to 8 weeks before anything moves. Crawling, indexing and re-embedding all lag independently, and the corpus a model already holds about you does not clear because you published a page. If your site is 300 pages deep, the machine layer sweep takes longer to propagate than the page itself does.

Do not measure this in sessions. The traffic case is a different question with a different answer, and we worked through it in LLM citation vs SEO traffic. The method for running this as a standing routine is in how to track AI search visibility, and how to audit your brand in ChatGPT is the one-off version if you have never looked.

Where Does an LLM Info Page Stop Working?

Four places, and the vendor selling you a GEO retainer will name none of them.

When the page is the only thing saying it. An info page is self-reported, and self-reported claims are weighted against everything else a model has absorbed about you. If 40 third-party sources describe you as a cold email agency and one page on your own site says podcast agency, the page loses. It is a tiebreaker and a clarifier, not an override, which is why the sitewide machine layer and your off-site profiles have to move with it.

When there is nothing underneath to describe. The page states what you do. If what you do is undifferentiated, the page will accurately describe an undifferentiated company, and being legible is not the same as being worth recommending. Gartner puts the share of a B2B buying group's time spent with any single vendor at roughly 5%. The other 95% is research you are not present for, and it is evaluating the substance, not the formatting.

When the conversations are not happening. This is the real one for most companies reading this. Every page like this is a multiplier on a supply of proof, and the constraint is almost never the technical layer. It is that nobody is getting in front of buyers at a steady rate. If invites are the mechanism and they land in spam, none of this ever starts, which puts the deliverability stack upstream of everything here. Start with podcast invite email deliverability, then setting up sending domains, warming a new domain, SPF, DKIM and DMARC, what email deliverability actually is, and how to stay out of the spam folder. If the wrong people are being contacted in the first place, the fix is further up in how to define your ICP.

When someone expects a graph inside a quarter. This is infrastructure. It lowers the odds of being described wrongly, and that is a real outcome with no dashboard. If you want a lever you can watch move, the lever is how many decision makers sit down to record with you next month, which is the argument in is podcast lead generation worth it.

Reverse Outbound Engine
An outbound method where, instead of cold selling your ideal buyers, you invite them onto your podcast as a guest. The invite reads as recognition rather than a sales approach, so it earns replies at rates a direct pitch never reaches. You spend about 45 minutes hearing how the guest built and grows their business, which builds real trust. Any fit for working together is a separate, later conversation. The pages and transcripts are the residue, not the point.

That residue is worth more than it looks, though. Every recorded conversation produces a page carrying a named expert answering direct questions, which is close to the ideal shape of a passage a system lifts into an answer. Podcast transcripts as AI search fuel covers why, turning episodes into answer content covers the editing, how to get your podcast cited by AI covers the distribution, and what a podcast does for your search footprint is the whole picture.

If the model itself is new to you, what is a podcast acquisition system is the plain description, how to start a B2B podcast for lead generation is the build order, podcast lead generation benchmarks sets the ranges, the tech stack covers the tooling, and common podcast acquisition failure modes covers what breaks. Audience size does not gate it the way people assume.

One caution on the research in this piece. The AI search field is about 2 years old, most published numbers come from vendors with something to sell, and studies contradict each other regularly. The Ahrefs study discloses its sample and is worth reading directly. Use the rest for direction, then grade yourself on your own 90 days, because that is the only dataset containing your buyers.

The Practitioner Takeaway

The industry spent a year publishing a file that 97% of the time nothing ever asked for, while the thing that works was available the whole time and required no new standard at all. Put the facts on a web page. Link it from the footer. Let the crawlers that are already hammering your site find it the way they find everything else.

Write the summary sentence first and test it by compressing it, because that sentence is what gets repeated. Say what you are before you say what you are not, and never negate your own category, because the qualifier that made the negation safe is the first word to fall off. Then go make the other 300 pages agree with the one you just wrote, since the machine layer outvotes the prose every time and a canonical page contradicted sitewide is just a well-written opinion.

None of it substitutes for having something worth describing. As an acquisition-first B2B podcast agency we commit to 30 recorded conversations with your ideal buyers in 90 days or your money back, and the guarantee is written on conversations rather than on pages for exactly that reason. The conversations are the evidence. A page like this is only the label on it, and a label with nothing behind it is the most legible way to be forgettable.

See How the Invite Engine Works

15 minute demo. No fluff. We will walk you through the exact system, show real prospect examples, and scope what it looks like for your market.

Schedule a Demo