◈
Is your site in the AI training data?
Frontier models learn from a handful of massive datasets — and Common Crawl is the biggest of them all. If your site isn't indexed, LLMs have no baseline knowledge of your brand, making AI Overview citations nearly impossible. This audit checks training corpus presence, knowledge graph grounding, and whether all 6 major training crawlers can actually reach your pages.
◈Common Crawl IndexThe primary training source for GPT-4, Gemini, Claude, and Llama models
⬡Wikipedia & WikidataHighest-weighted corpora — models over-index on structured knowledge bases
⊞6 Live Crawler TestsGPTBot, Google-Extended, ClaudeBot, PerplexityBot, CCBot, and Common Crawl tested live
Querying Common Crawl index…
Checking recent crawls for your domain. This can take 15–25 seconds.
① Common Crawl
② Internet Archive
③ Knowledge Graph
④ Crawler Tests
Common Crawl Capture History
What This Means for AI Citations
Common CrawlPrimary LLM Training Source
Common Crawl underpins the training data for GPT-4, Gemini, Claude, LLaMA, and nearly every frontier model. Pages indexed here form the baseline "knowledge" these models have about your brand — before any retrieval augmentation.
Internet Archive (Wayback Machine)
Indexed Pages Sample
These are the exact URLs from your domain that appear in Common Crawl and may be part of AI training datasets.
Training Crawler Access Tests
robots.txt permission is only half the story — a WAF rule or CDN bot-blocking policy can silently 403 a crawler your robots.txt welcomes. Each crawler below was tested live with its real user-agent string.
Why This Matters
Wikipedia and Wikidata are the highest-weighted corpora in LLM training — models reference them far more than average web pages. A Wikipedia article and a Wikidata entity ID create a strong "anchor" that other content about your brand can connect to, dramatically improving recall accuracy and citation consistency across all AI platforms.
Recommendations — Ordered by Impact