AI search stopped being a pilot. By May 2026, Similarweb measured citations in US ChatGPT answers at roughly 6.8%, up from 1.6% a year earlier, while ChatGPT’s share of generative AI web traffic fell from ~76% to ~53% as Gemini and Claude grew. G2’s March 2026 survey of 1,076 B2B software buyers found 51% now start research in an AI chatbot and 85% think more highly of a vendor an AI chatbot recommends. Google, meanwhile, states plainly that AI Overviews and AI Mode have no special requirements — no llms.txt, no AI markup, no special schema. Indexability, crawlability, and people-first content still decide what gets cited. Run this checklist before you spend another dollar on “GEO strategy” work.
1. Crawl & index
Google surfaces supporting links in AI Overviews and AI Mode from pages that are indexed and snippet-eligible — nothing more (and nothing less) than the Search technical requirements. AI Mode’s query fan-out pulls links from a wider, more diverse set of pages than classic search, so every crawlable page in your topic area is a potential citation, not just your money pages.
- Primary content is available in HTML without client-side rendering traps (inspect rendered HTML via URL Inspection)
- Canonical URL is correct and the page is indexed
- Page is in the XML sitemap and linked internally from a hub or relevant article
- Priority pages are snippet-eligible: no
nosnippet,max-snippet:0, or accidentalnoindex - Page experience passes Core Web Vitals — cited pages still need to load for the user who clicks through
2. AI crawler access
Search bots and training bots are different products with different tokens: Googlebot, OAI-SearchBot, and PerplexityBot determine whether your site appears in AI answers, while GPTBot, ClaudeBot, Google-Extended, and Bytespider are about model training and grounding. Cloudflare’s network data shows Bytespider (ByteDance) accessed 40.4% of protected sites and GPTBot 35.5% — but only ~3% of the top million sites block AI bots at all, so most blocks you do implement are still table stakes. OpenAI and Perplexity both say robots.txt changes take about 24 hours to propagate, and Perplexity’s user-triggered fetcher generally ignores robots.txt.
- robots.txt allows
Googlebot,OAI-SearchBot, andPerplexityBotif you want AI answer visibility - Training bots (
GPTBot,ClaudeBot,Google-Extended,Bytespider,CCBot) have an intentional, documented policy — allow or disallow per crawler, not a blanket wildcard - Server logs confirm AI crawlers can reach your content — no WAF or geo-rule blocks on the published IP ranges (Perplexity and OpenAI publish JSON IP lists; Anthropic publishes its bot list)
- No aggressive rate limiting that also starves search crawlers; if you must throttle, use
Crawl-delayonly where supported - After robots.txt edits, wait ~24 hours and verify the fetch worked before changing anything else
- An
llms.txtfile exists only if other systems (docs platforms, agents) will use it — Google ignores it for Search, per its June 2026 clarification
3. Entities & authorship
AI systems attribute answers to sources they can identify. Consistent naming, real author pages, and clear organizational identity make your content easier to attribute than anonymous commodity text. Google’s Preferred Sources feature — now live in AI Overviews and AI Mode — lets readers badge your domain as preferred, but it only works at domain level and only if readers choose it, so you have to ask.
- Company, product, and service names are consistent across pages and mentions
- Author bylines exist with a bio that demonstrates expertise on the topic
- Organization page and contact details are consistent and up to date
- Site is checked in the source preferences tool and readers are invited to add it as preferred (domain-level only, per Google’s docs)
- Each priority page has a unique point of view or first-hand material — Google explicitly calls commodity content (“7 Tips for First-Time Homebuyers”) weak for AI features
4. Answer architecture
AI responses are grounded in retrieval: the system pulls whole pages, then cites the supporting sentence. Pages that state the answer early, define terms where they’re used, and structure comparisons plainly are easier to quote accurately. Queries are also getting longer and more conversational — Similarweb shows users trusting AI to interpret intent — so answer full questions rather than engineering keyword variants. Covering related subtopics helps with fan-out queries, but building hundreds of thin variants for them is scaled-content spam.
- Question-led pages state the direct answer within the first ~100 words
- Key terms are defined where they first appear
- H2/H3 structure mirrors how buyers phrase the question, not just the keyword
- Comparison or table sections exist for “vs” and evaluation queries
- Genuinely distinct subtopics have their own sections or pages (fan-out coverage) without thin duplicate content
- The page answers the question once, clearly — no keyword padding or AI-pleasing rewrites
5. Citation-worthy evidence
Citation rates are low but climbing: ~6.8% of US ChatGPT answers carried citations by May 2026, and categories built on checkable facts (travel ~23%, automotive ~20%) get cited three to five times more often than advice-heavy ones like professional services (under 4%). The practical implication: in B2B, where citation rates are at the bottom, verifiable claims, dated statistics, and named sources are what separate citable pages from restatements. ChatGPT’s May 2026 homepage-link update also pushed ~6 in 10 referred visits to homepages, so brand recognition compounds even when a page isn’t quoted.
- Every statistic carries a source link and a publication date
- Original data or benchmarks include a methodology note (sample, period, tool)
- Fast-moving claims show a visible updated date
- Definitions and processes attribute named frameworks or authors
- No fabricated metrics — numbers match the cited source exactly
- The key answer is written as a short, standalone, quotable statement
6. Structured data & schema
Google is explicit: there is no special schema for AI Overviews or AI Mode, and structured data is not a ranking lever in AI features. Worse, the FAQ rich result — which had already been limited in 2023 — was fully removed from Google Search results starting May 7, 2026, and Google retired its documentation the following month. Keep markup only where a real rich result still exists, and make sure it matches visible content, because spam policies now explicitly apply to generative AI responses.
-
ArticleandOrganizationJSON-LD is valid and matches visible content - FAQPage markup has been removed or is maintained only for non-Google consumers — it no longer earns rich results
- No
Review/aggregateRatingmarkup where reviews aren’t visible on the page - JSON-LD passes the Rich Results Test
- Schema effort is budgeted only where a live rich result exists (breadcrumb, product, article, review snippet)
7. Trust, freshness & verification
Google’s helpful-content standards and spam policies apply to AI features — inauthentic mentions, scaled content, and undeclared incentives are the fastest way to get filtered out of both classic results and AI answers. On measurement: Search Console’s Generative AI performance report shows impressions from AI Overviews and AI Mode, and Similarweb tracks a 43%+ AI Overview appearance rate in US search. Remember that when an AI Overview appears, CTR falls — a Seer Interactive study cited by Similarweb measured 0.64% organic CTR with an AI Overview present versus 3.97% without — so declining CTR with stable rankings is demand absorbed by answers, not necessarily lost demand.
- Affiliate, sponsor, and incentive relationships are disclosed on-page
- Claims that age carry dates; stale content is updated or removed on a schedule
- Spam-policy review: no scaled content, no inauthentic mentions, no hidden incentives
- Search Console Generative AI performance report is reviewed monthly for AI Overviews and AI Mode impressions
- AI Overviews presence is tracked at keyword level separately from classic rank; CTR drops are diagnosed against it
- Brand mentions inside AI answers are sampled quarterly across ChatGPT, Gemini, Perplexity, and Claude
How to run this checklist
- Owner: content/SEO lead. Assign one person per failing checkbox; unresolved items roll into the next cycle.
- Cadence: full pass quarterly, plus a quick pass within two weeks of any major Google core update or AI feature change.
- Evidence: keep the Search Console Generative AI report export, robots.txt, and server-log crawler sample with each pass so trends are comparable.
- Follow-up: use the AI visibility audit sprint to turn failures into an execution plan, and topic cluster build to fix thin or orphaned coverage.
Citations
Sources & references
- AI features and your website | Google Search CentralGoogle Search Central
- Optimizing your website for generative AI features on Google SearchGoogle Search Central
- Perplexity CrawlersPerplexity



