Most advice about getting cited by AI starts and ends with schema markup. We have shipped more than 250 articles on this site testing what engines actually pull, and the largest study on the question found that 1,885 pages adding schema saw citations barely move. Below, the block that does get lifted, the exact shape it takes, and how to write one for every heading on the page.
What Is an Answer Capsule?
The block you just read is one. Nothing in it depends on the sentence before it. No word in it points at something further up the page. If you deleted this entire article and left only that box, a reader would still walk away knowing what an answer capsule is and roughly how long it should be.
That property is the whole point, and it is rarer than it sounds. Open almost any B2B article and try to lift one paragraph out cleanly. Most of them open with a transition, refer to a concept introduced two sections earlier, and put the actual answer in the fourth sentence after the writer has finished warming up. A human reading top to bottom never notices. A machine picking one passage to quote cannot use any of it.
- Answer capsule
- A short, self-contained block of prose, usually 40 to 60 words, positioned immediately beneath a heading and written to answer that heading in isolation. Analysis by Averi of content that AI engines cite found an identifiable capsule of this kind present in close to three quarters of cited material. The format is sometimes called an answer block or answer-first structure, and the underlying idea is the same in all three: the answer leads, the context follows.
The naming varies by who is writing about it. What does not vary is the structural test. Cover the page. Read the block. Does it still say something true and complete? If yes, an engine can use it. If no, it will retrieve your page, find nothing worth quoting, and cite somebody else while your content sits in the candidate set doing nothing.
Why Do AI Engines Lift Some Blocks and Skip Others?
It helps to see the three steps as three separate decisions, because they fail in different ways. DeepSmith frames the split cleanly: chunking is a storage decision made at index time, retrieval is a recall decision made when the question arrives, and passage extraction is a selection decision about which sentence inside the surfaced chunk gets quoted.
Most content teams are working on the middle one. They think about keywords and topic coverage, which is retrieval. Almost nobody works on the third one, which is the only step that produces a citation with your name attached.
Averi's breakdown of chunking and embeddings describes the mechanics: the system breaks your content into fragments, converts each fragment into a vector, and pulls back only the fragments whose vectors sit close to the question. Your page is never evaluated as a page. It is evaluated as a bag of fragments, and the fragment is the unit that wins or loses.
The attrition through those steps is brutal. Work published by AuthorityTech on how answer engines choose sources puts it at roughly 15 percent of retrieved pages actually getting cited, and only about 5 percent of retrieved content reaching the person who asked the question.
Read those three numbers together and the strategy writes itself. Getting retrieved is table stakes and a lot of decent content clears it. Getting selected out of the retrieved set is the bottleneck, and the thing that most reliably separates the cited from the merely retrieved is whether a quotable block exists at all. We covered the wider version of this in why AI answers cite some companies and not others, and the mechanics of the engines themselves in how to rank in AI search.
Why Did Schema Markup Barely Move Citations?
This is the finding that reorganized how we prioritize GEO work, so it is worth stating precisely rather than as a slogan. The Ahrefs research team identified pages that added JSON-LD between August 2025 and March 2026, matched them against control pages with similar prior citation levels, and measured the 30 days before against the 30 days after. As Search Engine Roundtable reported, AI Overviews showed a small decline for the schema group rather than a lift.
The caveat in the study matters as much as the headline, and most of the coverage dropped it. The pages tested were already receiving 100 or more AI Overview citations before any schema was added. These were pages the engines already understood. For a page that has never been cited, structured data may well help it get crawled, parsed, and correctly attributed to your company. That is a real job and it is why we still ship every one of these articles with seven schema blocks in the head.
What the study does kill is the idea that markup is the lever. If you add schema to a page with no extractable passage in it, you have described a page that still has nothing worth quoting. Schema is the label on the jar. The capsule is what is in the jar.
The right way to hold both facts at once is a division of labor. Use structured data for identity and understanding, covered in schema markup for podcast pages, entity consistency for AI search, and llms.txt for B2B companies. Use capsules for the sentence that gets quoted. Teams that do one and skip the other lose, and skipping the capsule is by far the more common failure.
How Do You Write an Answer Capsule?
Seven rules cover almost every capsule we write. They look obvious written down and they are broken constantly, including by writers who know them.
- Answer in sentence one. Not a definition of the topic, not a note about why the topic matters, not a promise that you are about to answer. The answer. If your first sentence could be deleted without losing information, delete it and start with the second.
- Repeat the subject by name. An engine lifting the block has no access to your heading. "It usually takes 6 to 8 weeks" is useless out of context. "A new sending domain usually needs 6 to 8 weeks of warmup" survives extraction.
- Carry one qualifier. A number, a range, a timeframe, or a condition. Unqualified claims read as marketing and get passed over for a source that committed to something specific.
- Stay inside 40 to 60 words. Under 40 and you usually dropped the qualifier. Over 80 and you are covering two ideas, which forces the engine to cut, and a cut quote is a quote you no longer control.
- Use plain sentence structure. No nested clauses, no colons doing structural work, no lists inside the capsule. The block gets rendered as prose inside somebody else's answer.
- Do not sell inside it. A capsule that names your product is a capsule an engine skips as promotional. Earn the citation with the answer, and let the attribution link do the selling.
- Test it cold. Copy the block into a blank document, read it, and ask whether a stranger could act on it. This is the only check that matters and it takes 10 seconds.
Rule 2 is the one that breaks most often. Writers naturally use pronouns to avoid repetition, because repetition reads clumsy to a human moving down the page in order. In a capsule, that instinct is exactly backwards. Repetition is what makes the block portable.
The related trap is the wind up. Phrases like "when it comes to", "in today's market", and "there are a few things to consider here" spend the first sentence of a capsule saying nothing. Under normal reading they are invisible. Under extraction they are fatal, because the sentence an engine most often quotes is the first one. Animalz's rundown of answer engine citation techniques lands in the same place from a different angle: the structural pattern beats the prose quality.
What Does a Weak Capsule Look Like Next to a Strong One?
Here is the same answer written both ways, for the question "how long does it take to warm up a cold sending domain".
| Element | Weak version | Strong version |
|---|---|---|
| First sentence | "Domain warmup is something a lot of teams ask about." | "A new sending domain needs 3 to 6 weeks of warmup before it carries production volume." |
| Subject naming | "It depends on how you are set up." | "Warmup speed depends on authentication, sending volume, and early reply rate." |
| Qualifier | None. "It varies quite a bit." | "Start at roughly 5 sends a day and step up gradually." |
| Length | 22 words, missing the answer | 52 words, complete |
| Survives extraction | No. Every sentence needs the page. | Yes. Reads correctly standing alone. |
| Reads well for humans | Yes, in sequence | Yes, in sequence and out of it |
Notice that the weak version is not badly written. It is friendly, it flows, and inside a full article nobody would flag it in an edit. It fails on one axis only, and it is the axis that decides whether a machine can use it.
The strong version also does something the weak one cannot: it gives the engine a fact to attribute. Our own working numbers on this sit in how to warm up a new email domain and email warmup explained, and the authentication half is in what SPF, DKIM, and DMARC actually do. Each of those pages carries its own capsule, which is why they get pulled into answers about deliverability rather than sitting in the retrieved set.
Citations are a means, not the scoreboard. Mickey went from referrals only to a 200K month by putting his buyers in the room rather than waiting for them to find him. Read the full case study →
How Many Capsules Should One Page Have?
The reason follows from chunking. If the engine never sees your page as a page, then a page level summary is the wrong unit of work. Every H2 is a separate entry point, every entry point answers a different question, and a section without a capsule is a section that can be retrieved and then discarded during selection.
Placement inside the section is not flexible. The capsule goes directly under the heading with nothing between them. Put an image, a transition sentence, or a "before we get into this" paragraph in that gap and you have separated the question from the answer at exactly the point where the chunk boundary is most likely to fall.
The strongest capsule on the page belongs under the first H2, inside the opening 120 words. Engines and human readers both weight the top of a document, and the first capsule is the one most likely to be pulled for the broad version of your topic. Everything below it competes for the long tail. MADX's guide to structuring content for answer engines arrives at a similar hierarchy from the agency side.
One nuance worth naming: the paragraph directly under a capsule has to earn its place. If it repeats what the capsule just said in longer words, the page reads padded and the reader learns to skip the prose. Use the paragraph after a capsule for the detail, the evidence, or the exception. That is the pattern we follow on every article, and it is covered further in how to structure content for ChatGPT and GEO vs SEO differences explained.
How Do You Tell Whether Capsules Are Working?
Rank tracking cannot answer this question. A page can hold position 3 in organic results and never appear in a single generated answer, and the reverse happens too. You need a separate measurement loop, run on a schedule, against a fixed question set.
The part teams skip is logging the quoted sentence itself. Recording a yes or no on whether you were cited tells you the score. Recording which 45 words got lifted tells you why, and it turns the next revision into an edit rather than a guess. Over a few months that log becomes the most useful content asset you own, because it is a direct readout of what the engines consider quotable in your category.
The traffic side is worth watching too, for a reason that is easy to miss. Volume from AI search is still small next to organic for most B2B companies, but the visitors behave differently. Semrush's study of AI search and SEO traffic found visitors arriving from AI search converting at several times the rate of traditional organic visitors, and its clickstream analysis of ChatGPT referrals put outbound referral growth from ChatGPT at 206 percent across 2025. Small and rising, with better intent per visit.
Our full measurement setup is in how to track AI search visibility, the audit version is in how to audit your brand in ChatGPT, and the reason we do not judge this work on sessions is in LLM citation vs SEO traffic.
What Happens When You Apply This to a Podcast Library?
This is where the topic stops being theoretical for us. Every client show produces recordings full of specific, first hand answers from operators who actually run the thing being discussed. That is exactly the material engines want to quote, and in raw transcript form almost none of it is usable.
The fix is mechanical. Pull the 5 or 6 real questions the episode answered, write each one as an H2, write a capsule under each that says what the guest said in a form that stands alone, and link out to the full episode underneath. The episode page stops being a wall of text and becomes 6 retrievable sections. We wrote the longer version of this in turning episodes into answer content and podcast transcripts for AI search.
There is a compounding effect worth being honest about, because it is slower than most GEO promises. A capsule does not get you cited this week. It makes the page eligible for a selection step it was previously failing, and eligibility only pays once retrieval catches up. Pixis's explainer on how AI engines read the internet is a good primer on why crawl and index lag makes 4 to 8 weeks a realistic first read.
What makes this work for us is that the podcast is not the acquisition channel. The invite is. We invite a client's ideal buyers onto their own show by email, the recording is the conversation, and our commitment is 30 recorded conversations with those buyers in 90 days or their money back. Editing is included, the client owns the show and every recording, and invites go out by email only. The search footprint the library builds afterward is a second return on work already done, which is the honest way to frame it. More on the mechanics in what a podcast does for your search footprint and what reverse outbound is.
The Part Nobody Wants to Hear
Answer capsules are boring work. There is no tool to buy, no integration to wire, and no dashboard that lights up when you finish. You reread your own headings, decide what each one actually promises, and write 50 honest words underneath. On a 40 page site that is a week of unglamorous editing, which is precisely why the technique stays available while everyone argues about markup.
The rest of the field is currently spending its GEO budget on the layer a controlled study just showed does not move selection, on pages that contain no quotable sentence. That gap is the opportunity, and it will not stay open forever. Once capsules become the default in a category, they stop being an edge and go back to being table stakes, the same way meta descriptions did.
Start with the pages you already want to be known for. Write one capsule per heading, lead with the answer, name the subject, carry a number, and read each block cold before you ship it. Then log which sentences the engines quote back at you over the next 8 weeks, because that log is the only feedback loop in this discipline that tells you the truth.
See How the Invite Engine Works
15 minute demo. No fluff. We will walk you through the exact system, show real prospect examples, and scope what it looks like for your market.
Schedule a Demo →