Most llms.txt guides sell the file as an AI visibility cheat code, and it is not one. We run outbound for 50+ B2B companies, so we watch how senior buyers check a vendor before they reply, and that check now starts inside an AI answer. Below is what the file actually does, the structure we publish, and what decides whether an assistant describes your company correctly.
What is an llms.txt file?
The format was proposed by Jeremy Howard at Answer.AI and is documented at llmstxt.org. The idea is simple. A model reading your site has a limited context window, your pages are wrapped in navigation and scripts, and the parts that describe your business are scattered across a homepage, an about page, and a pricing page that all say slightly different things. An llms.txt collapses that into one file a model can read in a few thousand tokens.
The structure is a heading with your company name, a blockquote summary, a few sections of plain prose, and then lists of links with one line of context each. No HTML, no scripts, no navigation. Just the facts, in the order you want them read.
- llms.txt
- A markdown file served at yourdomain.com/llms.txt that summarizes what a company does and lists the pages an AI system should read to describe it accurately. It is a proposed convention rather than a ratified web standard, and no major model provider has committed to crawling it on a schedule.
- Generative Engine Optimization (GEO)
- The practice of getting a company named and cited inside AI answers rather than ranked inside a list of blue links. The mechanics differ from classic search, which is covered in GEO vs SEO. llms.txt is one small tactic inside GEO, not the whole discipline.
Does llms.txt actually work in 2026?
Mostly no, and the honest answer is worth more to you than the optimistic one.
Ahrefs analyzed 137,000 sites and found that 97% of llms.txt files received zero requests over the study month. Of the 3% that got any traffic at all, 96% of the requests came from bots, and most of those bots were SEO audit tools and tech profilers rather than AI assistants. Named AI tools accounted for a small slice, led by GPTBot and Claude-Code, and Claude-Code is a coding agent, not a consumer search assistant. Search Engine Journal covered the same dataset and landed in the same place.
Google has been saying a version of this for over a year. John Mueller has described llms.txt as not something done for search, and called it a temporary crutch that mostly saves tokens for AI coding tools. Nothing in Google Search Central treats it as a signal. OpenAI documents its crawlers at platform.openai.com and describes robots.txt behavior, not llms.txt behavior.
So the file is not a ranking lever and it is not a citation lever. That still leaves a narrow, real use, which is the next section. First, here is where it sits against the files that do carry weight.
| File | What it does | Who honors it | Worth your time? |
|---|---|---|---|
| robots.txt | Controls which crawlers may fetch which paths | Google, Bing, OpenAI, Anthropic, and most well behaved bots, per robotstxt.org | Yes. This is the one that actually gates access |
| sitemap.xml | Lists every canonical URL and when it changed | Every major search crawler | Yes. Non negotiable |
| Schema markup | Labels entities and answers in machine readable JSON-LD, per schema.org | Search engines and, in practice, the models trained on their output | Yes. This is the highest leverage GEO work on the page itself |
| llms.txt | Curated summary and a routing list of your important pages | Nobody on a schedule. Some agents when handed the URL directly | Only after the three above are done |
Then why write one at all?
Three reasons, and only one of them is about crawlers.
It forces you to write down what you actually are. This is the real return, and it has nothing to do with AI. Writing an llms.txt means deciding, in one paragraph, what your company sells, who it is for, and who it is not for. Most B2B companies cannot do that cleanly. Their homepage says one thing, their sales deck says another, and their proposal says a third. The moment you try to compress all of it into 40 words, every contradiction surfaces. We rewrote our own positioning twice while writing ours, and both rewrites made the sales conversation shorter.
Agents fetch it when a human hands them your URL. The Ahrefs number measures unprompted crawling. It does not measure what happens when a buyer pastes your domain into ChatGPT or Claude and asks what you do. In that flow, the agent fetches live, and a clean file at a predictable path is cheap to read. That is a real moment, and for a high-ticket B2B company it happens more often than a random crawl ever will.
It costs an afternoon and expires nothing. There is no downside case. It does not compete with your sitemap, it does not touch your rankings, and if a provider does start honoring it in 2027, you are already there. Treat it as insurance with a small premium, not as a growth channel.
What it is not: a substitute for having pages worth citing. If your site has nothing specific on it, a summary file just summarizes nothing faster.
What goes in a B2B llms.txt, section by section?
Here is the structure we publish, in order. Every section earns its place by answering a question a buyer or an assistant actually asks.
- H1 with the exact company name. Use the legal, full name, spelled the way you want it repeated. If a shorter form is acceptable, say so explicitly in the file. Assistants copy what you give them, including your mistakes.
- A blockquote summary, 40 to 70 words. This is the single most important block in the file. It is what gets lifted verbatim. Write it as a standalone paragraph that makes sense to somebody who has never heard of you, and put the category, the buyer, and the mechanism in it.
- What the company is. Bullets. Concrete deliverables, not adjectives. What is included, what the client receives, what they own at the end.
- What the company is not. The most underused section in every llms.txt on the internet. Models fill gaps with the nearest category, so if you do not say what you are not, you get lumped in with whatever is adjacent. We explicitly say we are not a podcast production agency and not a guest booking service, because both are the wrong nearest neighbor.
- Pricing, or an explicit instruction about pricing. If you publish a number, put it here. If you do not, say so and say it plainly, so a model does not infer one from an old page or a competitor. Ours says the number is scoped on a call and that any figure attached to the company elsewhere is out of date.
- The promise you make and how it is defined. Whatever you commit to, define the unit precisely. Vague commitments get restated by assistants in a stronger form than you wrote, which is a problem you find out about on a sales call.
- How it works, numbered. A model can lift a numbered process cleanly. Prose gets paraphrased and loses steps.
- Who it is for. The fit test, in the shape of a business rather than a headcount band. This is the same discipline as writing an ideal customer profile, and the same work pays off when you go to define an ICP for cold email.
- Core pages, with one line of context each. Not a sitemap dump. 8 to 15 links, each with a sentence explaining what a reader gets there. The context line matters more than the URL.
- Topic clusters. Group your deep content by subject so an assistant answering a narrow question can find the narrow page. Ours groups the deliverability library separately from the podcast acquisition library, because those answer different questions.
Two formatting rules. Keep the whole file under roughly 5,000 words so it fits comfortably in a working context window, and write in plain declarative sentences. Markdown headings, bullets, and links only. No tables, no HTML, no cleverness.
You will also see llms-full.txt referenced, which is a second file holding the full text of your key pages rather than links to them. Skip it unless you have a documentation site. For a B2B services company it duplicates content that already exists at a real URL, it goes stale the moment a page changes, and it gives an assistant a second version of your facts to disagree with the first. One file, kept current, beats two files drifting apart.
Serve it as plain text at a fixed path with no redirect chain. The whole value depends on an agent finding it at the address it expects, and a 301 hop or a login wall in front of it is the same as not publishing it. Add it to your deploy so it ships with the site rather than living as a one time upload somebody forgets.
How do you write the summary line an assistant will repeat?
The blockquote at the top of the file is the block that gets quoted. Treat it like a headline you are going to read out loud, because that is functionally what happens.
Four things belong in it. The category you are in, stated in the words a buyer would use. The buyer you serve. The mechanism, meaning the specific thing you do that is different. And one disqualifier, so the model has a boundary.
A weak version reads like this: A leading provider of innovative B2B growth solutions helping companies scale. Every word in that sentence survives being deleted, which means none of them are doing work. An assistant handed that sentence will describe you as a generic agency, because that is the only information present.
A strong version names the mechanism. Ours says we run cold email that invites a client's ideal buyers onto the client's own podcast, runs the recording, and converts guests into sales conversations afterward, and that the opening ask is an invitation rather than a pitch. That last clause is the differentiator, and it is the clause that comes back when somebody asks an assistant how we are different from an SDR agency. The mechanism is the memorable part, which is the same reason it carries the weight in podcast led outbound and in a podcast acquisition system.
If your summary line could be pasted onto a competitor's site without anybody noticing, it is not a summary. It is filler, and a model will treat it that way.
Getting described correctly only pays once something is driving buyers to look you up in the first place. Mickey Hardy went from referrals only to a 200K month once the invites started going out. Read the full case study →
What actually gets a B2B company cited by AI?
This is where the effort belongs. An llms.txt is a 2 hour job. The list below is the actual work, roughly in order of return.
- Pages that answer one question completely. A model lifts a self contained block. A page that answers 6 questions half way gets skipped for a page that answers 1 question fully. This is the mechanic behind getting cited by ChatGPT and Perplexity.
- Structured data on every page. FAQPage, Article, Organization, and Person schemas give a model labeled facts instead of prose it has to infer from. This is the highest leverage on-page work available and it is entirely under your control.
- An answer capsule near the top of every page. 40 to 60 words, plain language, no setup. It is the block that gets extracted verbatim, and burying it below 800 words of throat clearing wastes it.
- Entity consistency across the web. Same company name, same founder name, same location, everywhere. Mismatches split the entity and the model hedges. LLM citation behaves differently from SEO traffic here, because citation is about confidence in an entity rather than authority on a keyword.
- Original data nobody else has. Models cite the source of a number. If you publish benchmarks from your own book, you become that source. Our cold email reply rate benchmarks and podcast lead generation benchmarks exist for exactly this reason.
- Transcripts of anything you record. Spoken content is invisible until it is text on a page. The full argument is in podcast transcripts for AI search and how to get your podcast cited by AI.
- Third party mentions. Being described by somebody else, on their domain, carries weight your own site cannot generate. Cloudflare Radar is a decent public read on how much AI crawler traffic is moving across the web generally, and it is not evenly distributed. It concentrates on sites that get referenced.
Notice what is missing from that list. None of it is a file at the root of your domain. The file is a router. These are the destinations, and a router that points at nothing routes nobody.
How do you check whether any of this is working?
Two checks, one cheap and one that takes a system.
The cheap one is to ask the assistants directly. Open ChatGPT, Claude, Perplexity, and Gemini in a clean session with no memory, and ask each one what your company does, who it is for, how it is priced, and how it compares to the obvious alternative. Do it monthly and keep the answers in a doc. You are watching for three failures: the wrong category, a price you never published, and a competitor's mechanism attributed to you. All three are fixable, and all three are invisible until you look. Getting cited by Perplexity covers the citation side of the same audit.
Run the same questions across all 4 assistants rather than just the one you use, because they disagree more than people expect. One will have your category right and your mechanism wrong. Another will confidently attach a number to you that came from a competitor's pricing page. Write down which assistant got which detail wrong, because the fix is usually a specific page rather than a general effort, and knowing which page saves the afternoon.
The harder check is server logs. Filter for GPTBot, ClaudeBot, PerplexityBot, and Google-Extended, and see what they fetch and how often. If your llms.txt shows zero hits after 90 days, you are in the 97%, and that is the expected result rather than a failure to debug. What matters more in that log is whether the crawlers are reaching your deep pages at all.
The mistakes that make an llms.txt worse than nothing
A bad file is not neutral. It is a confident, machine readable version of your worst copy, sitting at a predictable path.
- Dumping your sitemap into it. 400 links with no context is a sitemap with extra steps. Curate to the 10 or 15 pages that carry the argument.
- Marketing language in the summary. Adjectives get discarded. Mechanisms get repeated. Write the mechanism.
- Leaving stale facts in it. A file at a fixed path becomes the canonical version of you. Ours drifted for weeks after a role changed internally, and the file kept describing a process that had already moved. Put it on the same review cadence as your homepage.
- Contradicting your own pages. If the file says one thing and the pricing page says another, you have taught a model to hedge. Pick one and fix the other.
- Treating it as the GEO plan. The file is the smallest item on the list. If it is the only thing you have shipped, you have optimized the signpost and skipped the building.
- Publishing it and never testing it. Paste your own URL into an assistant and ask it to describe you. If the answer is wrong, the file is wrong.
Where this actually lands
Write the file. It takes an afternoon, it forces you to say what you are in plain language, and it costs nothing to keep. Just do it with the right expectation, which is that almost nobody will crawl it on their own, and the return you get is a cleaner set of facts about your company rather than a spike in AI citations.
The visibility work that pays is upstream of the file and always has been. Pages that answer one question completely, schema on everything, an answer capsule near the top, original numbers from your own book, transcripts of what you record, and a consistent entity across the web. That is the list. The llms.txt is the index card taped to the front of it.
And the part nobody wants to hear: being described accurately by an AI only matters once buyers have a reason to look you up. Getting into that consideration set is a different job, run through podcast lead generation, a warmed sending setup that clears the filters, and invites that senior buyers actually open. The infrastructure side of that is covered across email deliverability, SPF, DKIM, and DMARC, DNS records for deliverability, domain reputation, setting up sending domains, warming a new domain, warmup mechanics, staying out of the spam folder, deliverability monitoring, multi domain sending, and secondary domains. The conversation side runs through why executives say yes to podcast invites, building a guest list, and what counts as a positive reply.
An assistant can only describe a company it has a reason to know about. The file makes the description accurate. The pipeline is what makes anybody ask.
See How the Invite Engine Works
15 minute demo. No fluff. We will walk you through the exact system, show real prospect examples, and scope what it looks like for your market.
Schedule a Demo →