How AI search engines choose what to cite
An AI answer engine cites a page when it can reach it, lift a standalone answer from it, judge that answer trustworthy, and find the claim corroborated somewhere else. Failing any one of those four makes the other three irrelevant.
Get an Agent Trust Gap BriefFour gates, and they are sequential
This is why sites that rank well often go uncited. Ranking optimises gate three. If gate one is closed — a CDN quietly blocking an AI crawler your robots.txt permits — nothing downstream matters.
1. Access
Can the crawler fetch the page? robots.txt, CDN rules, WAF behaviour, and whether content requires JavaScript to appear. Most surprises live here and almost nobody verifies by request.
2. Extractability
Is there a clean answer near the top that survives being lifted out of context? Content that only makes sense after three paragraphs of preamble gets summarised, not quoted.
3. Trust
Named sources, real dates, stated methodology, visible authorship. Unsourced assertion is the most common reason a page is read and not cited.
4. Corroboration
Does the claim appear anywhere you do not control? Single-source claims are the hardest for an engine to surface confidently.
What actually moves citation
- A definition-first opening. Answer the head question in the first 40–60 words, completely, in a form that stands alone. This single change moves extraction more than any other on-page edit.
- Named primary sources inline. Not "studies show" — the instrument, the section, the date, the link.
- Real verification dates. A visible last-verified date that is true. A stale date is worse than none.
- Structured data that matches what renders. FAQPage schema where FAQs actually appear. Schema describing content that is not on the page is a trust signal in the wrong direction.
- Consistent entity description. The same description of who you are, everywhere. Ambiguity suppresses confident citation.
- Something published off your own domain. An open repository, a schema under a permissive licence, a talk, a profile. This is the most neglected and the highest leverage.
What is measurable and what is not
Referral parameters, grounding queries and observed citations give partial visibility. There is no complete citation analytics for AI answers, and anyone claiming otherwise is overstating. Say which parts of your measurement are observed and which are inferred — the honesty is also a trust signal. More in GEO and AEO services.
Common questions
Why does my site rank well but never get cited?
Usually extractability or corroboration. Ranking rewards relevance and authority; citation additionally requires a liftable standalone answer and a claim that appears somewhere other than your own domain.
Does llms.txt improve AI citation?
Some AI crawlers read it; Google ignores it. It is cheap and low-risk to publish. It is not a ranking mechanism and treating it as one misreads what it does.
What is the single highest-leverage change?
A definition-first opening on your most commercially important pages — a complete standalone answer in the first 40–60 words.
How important is off-site corroboration?
Very, and it is the most neglected. A claim that exists only on your own site gives an answer engine no independent reason to repeat it.
Can AI citations be measured?
Only partially. Referral parameters, grounding queries and observed citations help. Complete attribution is not currently available and should not be promised.
Last verified 2026-08-07. Sources are named inline. Not legal advice.
By Julian Joseph, Founder, ApexClaw. Written from direct work operating autonomous systems under governance. Reviewed against the claims policy: sourced, first-party, or labelled.