How to get cited by ChatGPT and AI Overviews

AI engines split pages into chunks, embed them, retrieve the most relevant few, and synthesise an answer. To be cited, make each section a self-contained answer: lead with a 40–60 word direct response, use question-shaped headings, and put facts in tables with real sources.
7 min read · Updated

How AI engines actually read your page

This is the part worth internalising, because everything else follows from it.

An AI engine does not read your page the way a person does. It splits the page into chunks, converts each chunk into a vector, retrieves the handful most relevant to the question, and synthesises an answer from those chunks.

You are optimising the chunk, not the page. A brilliant article whose individual sections do not stand alone will lose to a plainer one whose sections do.

Lead with the answer

Every page should answer its own title question in the first 40–60 words. Self-contained, no preamble, containing the primary entity.

The test: if that paragraph were lifted out and shown to someone with no other context, would it read as a complete answer? If not, rewrite it.

This single habit drives more citation than anything else on this list. It is also the one most often skipped, because writers are trained to build up to a conclusion.

Use question-shaped headings

Phrase H2s as the question a user actually types. "How do AI engines read your page?" beats "AI Engine Architecture".

Then make each section independently answer its own heading. This is where most pages fail.

Make chunks independent

Because sections are retrieved in isolation, pronouns that reach back to earlier sections destroy retrievability.

  • ❌ "It works by splitting the content into pieces."
  • ✅ "An AI engine works by splitting the page into chunks."

Restate the subject. It reads slightly repetitive to a human reading top to bottom, and it is the difference between being quoted and being skipped.

Write definitional sentences

X is a Y that Z. Copularly explicit, no hedging.

"Topical authority is a search engine's assessment that your site comprehensively covers a subject" is quotable. "When we talk about topical authority, there are several things to consider" is not.

Put facts in tables

Structured blocks extract far more reliably than the same facts in prose.

Format Extraction reliability
Table High
Bulleted list High
Prose paragraph Moderate
Narrative with facts scattered Low

If a section contains three or more comparable facts, it should probably be a table.

Cite real sources

Named entities, specific dates, exact numbers, each traceable to a real source.

Vague authority language — "studies show", "experts agree" — is worthless to a retrieval system, which cannot verify it and has no reason to prefer your unsourced claim over anyone else's.

Never fabricate a statistic or a source. Beyond the obvious integrity problem, engines increasingly cross-check, and a page with an invented citation is worse than a page with no citation.

Keep it fresh

Visible dateModified, current-year data, and updates when facts change. AI engines weight recency heavily, more so than classic search does for most queries.

Stay consistent across your site

If two of your pages state different figures for the same fact, both lose citation confidence. Check facts across the corpus, not just within a page.

Publish the machine surfaces

Emit the files that make your content easy to consume:

  • llms.txt — a curated markdown map of the site
  • Markdown mirrors of each page, so engines read clean text instead of scraping HTML
  • FAQPage and DefinedTerm JSON-LD where the content genuinely warrants it
  • Clean sitemaps and feeds

Be honest about status: llms.txt and ai.txt are emerging conventions rather than ratified standards, and not all engines honour them. They cost almost nothing and carry no downside, which is the whole argument for shipping them.

Which content types get cited most?

Glossary and definitional pages, by a wide margin. A clean definition is the easiest possible thing for a model to quote.

Comparison tables and original research follow. Long narrative essays are cited least, because there is rarely a self-contained chunk worth lifting.

That ranking is a good reason to build your glossary early — see the content types guide for how.

Frequently asked questions

Is AEO different from SEO?

It is a subset with different emphasis. The structural work overlaps heavily, but AEO weights chunk independence, answer-first writing and factual density far more than traditional SEO does.

Does llms.txt actually work?

It is an emerging convention, not a ratified standard, and it is not universally honoured. It costs almost nothing to publish and the downside is zero, so it is worth doing — but do not expect it to carry your strategy.

Keep reading

Or let Omnitopical do all of this.

Everything in this guide, run automatically on one domain — mapped, written, interlinked, and published with 25 machine files. $99/mo.

Map my site's topical authority