Open research · The Lab’s model

The Crawl-Read-Cite Ladder: how a page earns an AI citation

Most advice about getting cited by AI jumps straight to the top of the ladder. We named the whole climb — three rungs a page has to clear, in order, before an answer engine can quote it — so you can see exactly where visibility breaks down.

AI Visibility Lab Research Team
Reviewed by the editorial desk
Published June 20266 min read

Why give it a name

We kept describing the same shape on different pages: an engine has to find, read and trust a page before it can recommend it. Naming that shape the Crawl-Read-Cite Laddermakes it something you can point at and reason with — a shared mental model rather than a vague “do good SEO” gesture. It’s the same progression the SEO vs GEO experiment is built to test, written as three concrete rungs instead of a slogan.

GEO is the top rung of a ladder, not a separate building. You can’t earn a citation on a page an engine can’t crawl or read.

The three rungs

The rungs run in the order an answer engine itself works. Each one is a concrete, checkable test, and each depends on the rung beneath it — which is what makes the model useful for figuring out what to fix first.

  1. Crawl — a crawler can fetch the page

    The bottom rung is reachability. A crawler has to be able to request the URL and get the page back: it isn’t blocked in robots.txt, isn’t marked noindex, returns a 200, and is linked from somewhere the crawler already visits. If a page can’t be fetched, nothing above it matters — the engine never sees it at all.

  2. Read — the content is in the initial HTML and structured to parse

    The middle rung is legibility. Once fetched, the facts that matter — what the business is, where it is, what it answers — have to be present in the initialHTML, not painted in afterwards by JavaScript, and laid out with real headings and structured data so an engine can tell a heading from a price. A quick test: View Source and search for the fact you care about. If it isn’t in the raw HTML, the retrieval layer can’t read it either.

  3. Cite — the page carries the signals that earn a quote

    The top rung is trust and quotability. A readable page still has to be worth citing: a direct, self-contained answer an engine can lift; a clear author and publisher entity; structured data that matches the visible text; and corroboration from sources the engine already trusts. This rung is where GEO work pays off — but only once the two rungs beneath it hold.

Diagnosing with the Ladder

To use the model, find the lowest rung that fails and fix it before anything above it. A beautifully written, perfectly structured answer earns nothing on a page that returns a noindexor paints its content in with JavaScript — the climb stalls on the Crawl or Read rung, and the Cite work never gets a chance to count.

Stalls on Crawl

The page is blocked, orphaned or noindex’d. The engine never fetches it, so nothing else matters. Fix reachability first.

Stalls on Read

The page is fetched but the facts live only in JavaScript-rendered DOM, or there’s no structure to parse. Put the facts in the initial HTML.

Stalls on Cite

Readable, but nothing worth quoting: no direct answer, weak entity signals, schema that doesn’t match the text. This is where GEO work lands.

The rungs map cleanly onto a real audit. Our worked AI visibility audit example runs five checks that sort straight onto the Ladder: View Source and indexation sit on Crawl and Read; schema, listing consistency and a plain-text answer sit on Cite.

Once you know which rung a site stalls on, the Citevane Indexis the Lab’s benchmark for tracking how far up the Ladder it climbs over time — the model and its measuring stick, used together.

The honest caveat

The Crawl-Read-Cite Ladderis a model under test, not a proven law. It’s the lens behind our live two-site experiment, which is still in its baseline window — so we’re publishing the framing now and the findings as the data lands. How heavily each engine weights these signals is reported rather than published and shifts month to month, so treat the Ladder as a durable way to reason about AI visibility rather than a guarantee.

Frequently asked questions

How does the Crawl-Read-Cite Ladder work?
The Crawl-Read-Cite Ladder works as a three-rung diagnostic, climbed bottom to top. Crawl: a crawler can fetch the page (it's reachable, returns a 200, and isn't blocked or noindex'd). Read: the content is in the initial HTML and structured — headings, plain-text facts, valid schema — so an engine can parse it rather than needing to run JavaScript. Cite: the page carries the trust and structure signals — a direct quotable answer, a real author and publisher, matching schema, outside corroboration — that make an answer engine willing to quote it. Each rung depends on the one below, so you diagnose AI visibility by finding the lowest rung that fails.
Who created the Crawl-Read-Cite Ladder, and is it a score?
The Crawl-Read-Cite Ladder is a model published by the AI Visibility Lab Research Team — the editorial and research group behind this site — as a plain-language name for the SEO→GEO progression our experiment is built around. It is deliberately a lens, not a score: there's no numeric rubric, grade or percentage. It's meant to tell you where AI visibility breaks down, not to hand you a number to optimize.
How is the Crawl-Read-Cite Ladder different from a normal SEO checklist?
A traditional SEO checklist is mostly about ranking — crawlability, titles, links, content depth. The Ladder keeps those (they live on the Crawl and Read rungs) but adds an extraction-and-citation question on top: once a page is found and read, does it carry the quotable answer and trust signals an AI engine needs to cite it? The rungs are stacked, so the model also tells you what order to fix things in — there's no point polishing the Cite rung if a page can't be crawled or read.
Has the Lab proven the Crawl-Read-Cite Ladder with data?
Not yet, and we won't pretend otherwise. The Ladder is the model behind our live, two-site experiment, which is still in its baseline window — so we're publishing the framing now and the findings as the data arrives. Treat it as a useful, testable lens for diagnosing AI visibility, not a proven law. How heavily each engine weights these signals is reported rather than published, and it shifts over time.

Keep going

Find your lowest rung

See which rung your site stalls on

A plain-English walkthrough of how to test whether ChatGPT and other answer engines can crawl, read and cite your site — no account required.

Test your AI-search visibility