AEO Audit Checklist (2026): 32 Checks in 6 Stages
A practical AEO audit checklist for 2026 - 32 concrete checks across crawler access, snippet eligibility, content structure, structured data, entity signals, and measurement. With the evidence for what actually works.
Quick answer
An AEO audit answers one question in six stages: can AI answer engines reach your content, are they allowed to quote it, can they extract it, do they trust it, do they cite it, and can you tell? Run the stages in that order, because a failure at stage one makes every later optimization worthless. The checklist below is 32 concrete checks you can work through in an afternoon, and it flags the two popular tactics whose published evidence does not support the effort they usually get.
This is the technical companion to our free 30-minute AI visibility audit, which scores whether you appear. This one diagnoses why.
Stage 1: Crawler access (7 checks)
Nothing downstream matters if the bots cannot fetch you. Separate the two crawler categories, because they carry different business decisions.
Retrieval bots put you into live answers. Blocking them removes you from AI search results:
OAI-SearchBot/1.4- surfaces sites in ChatGPT searchClaude-SearchBot- Claude search quality and retrievalPerplexityBot/1.0- Perplexity indexing, explicitly not used for foundation model trainingGooglebot- powers Google Search and, through it, AI Overviews and AI ModeApplebot- Spotlight, Siri, Safari
Training crawlers are a separate call: GPTBot/1.4, ClaudeBot, CCBot, Bytespider, meta-externalagent/1.1, Amazonbot, Google-Extended.
The checks:
- Fetch your live robots.txt and confirm each retrieval bot above is allowed.
- Confirm no blanket
User-agent: *Disallow: /rule is shadowing a later allow. - Check for WAF, bot-management, or CDN rules blocking these agents above robots.txt. This is the most common silent failure, and your robots.txt will look perfect while it happens.
- Verify in server logs that the retrieval bots are actually fetching, not just permitted.
- Confirm your training-crawler policy is a deliberate decision, documented, rather than a default someone inherited.
- Note which fetchers ignore robots.txt anyway. Perplexity’s docs state
Perplexity-Usergenerally ignores robots.txt; Meta saysmeta-externalfetcherandfacebookexternalhitmay bypass it; Google’s user-triggered fetchers includingGoogle-Agentgenerally ignore it too. - Check crawl volume against referrals. Cloudflare reported 2026 crawl-to-referral ratios ranging from 118:1 to nearly 50,000:1 across AI companies, against roughly 5:1 for traditional search. Heavy crawling with no referrals is a signal, not a bug.
Stage 2: Snippet eligibility (5 checks)
This stage catches the failure nobody looks for. Google’s documentation states that to appear in AI Overviews and AI Mode a page needs to be indexed and eligible to be shown with a snippet. Old SEO decisions routinely break that.
- Grep templates and page head for
nosnippet. Any page carrying it is disqualified from the surfaces that quote you. - Check
data-nosnippetattributes on the content blocks that answer your key questions. - Check
max-snippetvalues. A tight character limit set years ago for SERP reasons throttles what an engine can quote. - Confirm no
noindexon pages you expect to be cited, including paginated and filtered variants. - Confirm nobody has blocked
Google-Extendedbelieving it controls AI Overviews. Google states it governs Gemini training and grounding only and does not affect Search inclusion or rankings. Blocking it costs you Gemini and buys you nothing on AI Overviews.
Stage 3: Content extractability (6 checks)
Answer engines quote passages, not pages.
- Does the direct answer to the page’s core question appear in the first 100 words, in plain declarative language?
- Are headings phrased as questions a buyer would actually ask?
- Is comparison data in real HTML tables rather than images or CSS-only layouts?
- Does primary content render without JavaScript? Test with JS disabled.
- Are claims attributed with specifics - named sources, dates, figures - rather than asserted?
- Is each page’s content self-contained enough to be quoted without the surrounding site for context?
One caution from Google’s own AI optimization guidance: it warns against artificially chunking content to game extraction. Structure for readers who skim, not for a parser you are imagining.
Stage 4: Structured data, honestly (5 checks)
This is where the industry oversells, so here is the evidence before the checklist.
Google’s documentation says plainly that there are no additional requirements to appear in AI Overviews or AI Mode and no special schema.org structured data that you need to add. Ahrefs tested it properly: a quasi-experiment on 1,885 pages that added JSON-LD against roughly 4,000 controls, measured 30 days either side, found AI Overviews citations down 4.6%, with AI Mode and ChatGPT changes statistically indistinguishable from zero. Their conclusion was that adding schema produced no major uplift on any platform.
So implement schema for what it demonstrably does, which is rich results and commerce surfaces:
- Organization schema, complete and consistent with your other listings.
- Product or SoftwareApplication schema on commercial pages.
- Article or BlogPosting with accurate author and date fields, since freshness signals do correlate with citation.
- Validate everything with the Rich Results Test and the Schema.org validator.
- Note that Google removed FAQ rich result documentation in June 2026 and FAQ rich results are no longer shown. Keep FAQ content for readers, but stop counting it as a rich-result play.
Our deeper guide to schema markup for AI citations covers implementation detail for the types that still earn their place.
Stage 5: Entity and source signals (5 checks)
This is where the actual leverage is, and it is the least automatable stage.
- Run your 20 priority prompts and record which sources each engine cites, not just whether you appear.
- Check your presence on each of those cited sources. Review sites, Reddit threads, industry directories, comparison pages.
- Verify Wikipedia and Wikidata representation if your category has it.
- Check description consistency across every third-party listing. Conflicting descriptions weaken entity confidence.
- Identify the two or three sources appearing most often in your category’s answers and treat closing that gap as the quarter’s priority.
If an engine consistently cites a competitor’s comparison page or a Reddit thread you are not in, that is your diagnosis. No dashboard produces it and no schema fixes it.
Stage 6: Measurement (4 checks)
- Turn on AI crawler observability. Cloudflare AI Crawl Control is included on all Cloudflare plans at no extra cost and shows which AI bots reach you.
- Enable AI Performance reporting in Bing Webmaster Tools, which entered public preview in February 2026.
- Set a weekly prompt baseline across priority queries. Our AI visibility tools pricing comparison covers what that costs, including the free options.
- Track crawl volume and citation share as separate metrics. They fail for different reasons and a single blended score hides both.
On llms.txt
Publish one if you want, but sequence it last. Google’s documentation states that creating machine-readable AI text files will neither harm nor help visibility, because Google Search ignores them. An Ahrefs study across 137,210 domains found 97% of llms.txt files received zero traffic, and that no AI bot ever requested a file that did not exist. A separate server-log analysis of roughly 268,000 agent requests found only 37 of about 770 llms.txt fetches came from named AI assistants. What did work in that study was HTTP content negotiation, with one coding agent taking Markdown 76% of the time when offered it.
It costs an hour and it is harmless. It is not a strategy. Our llms.txt implementation guide has the format if you want it.
Why the order matters
Most AEO work fails because it starts at stage four. Teams ship schema, publish llms.txt, and wait, while a bot-management rule quietly blocks OAI-SearchBot and a max-snippet directive from 2023 caps what anyone can quote. Stages one and two take an afternoon and invalidate or validate everything else.
The context worth holding: zero-click behaviour is now the norm, with SparkToro and Similarweb measuring 68% of US Google searches ending without a click in early 2026, and Pew finding users clicked a result on 8% of visits with an AI summary against 15% without. Being cited is increasingly the outcome, not the click. That is the argument for auditing citation conditions rather than rankings.
If you would rather have this run for you across the major engines, with competitor benchmarking and a prioritized fix list, that is exactly our GEO Readiness Audit. For the remediation itself, GEO Technical Implementation does stages one through four. And for the full discipline around it, start with our complete guide to generative engine optimization.
Frequently Asked Questions
What is an AEO audit?
An AEO audit is a systematic check of whether AI answer engines can reach your content, are allowed to quote it, can extract it cleanly, and actually cite it. It runs in six stages: crawler access, snippet eligibility, content structure, structured data, entity and source signals, and measurement. Unlike a prompt-tracking dashboard, which reports whether you are visible, an audit tells you which specific technical or content condition is blocking you.
Which AI crawlers should I allow in robots.txt?
At minimum the retrieval bots, because blocking them removes you from live AI answers: OAI-SearchBot (ChatGPT search), Claude-SearchBot, PerplexityBot, and Googlebot. Training crawlers - GPTBot, ClaudeBot, CCBot, Bytespider, meta-externalagent - are a separate business decision you can make independently. Note that several user-initiated fetchers including Perplexity-User, meta-externalfetcher, and Google's user-triggered fetchers generally ignore robots.txt entirely.
Does blocking Google-Extended remove me from AI Overviews?
No, and this is one of the most common misunderstandings in AEO. Google's documentation is explicit that Google-Extended controls Gemini model training and Gemini grounding only, and that it does not affect inclusion in Google Search or act as a ranking signal. AI Overviews and AI Mode are powered by Googlebot. To control appearance there you use snippet directives such as nosnippet, data-nosnippet, and max-snippet, or noindex.
Is llms.txt worth implementing in 2026?
The honest answer is that the evidence is weak. Google's documentation states directly that Google Search ignores these files, and an Ahrefs study of 137,210 domains found 97% of llms.txt files received zero traffic, with no AI bot ever requesting a file that did not exist. A separate server-log study found only 37 of roughly 770 llms.txt fetches came from named AI assistants. It is cheap and harmless, so publish one if you like, but do not sequence it ahead of crawler access and content structure.
Does schema markup increase AI citations?
Less than the industry claims. Google's own documentation says there is no special structured data required to appear in AI Overviews or AI Mode. Ahrefs ran a quasi-experiment on 1,885 pages that added JSON-LD against 4,000 controls and found AI Overviews citations fell 4.6% while AI Mode and ChatGPT changes were statistically indistinguishable from zero. Implement schema for rich results and commerce surfaces where it demonstrably works, not as a citation lever.
Complementary NomadX Services
Get Recommended by AI.
Book a free 30-minute GEO strategy call. We check what ChatGPT, Perplexity, and Gemini say about your product right now - and show you how to improve it.
Talk to an Expert