Common Crawl checker: Is your page in Common Crawl's record?
This free Common Crawl checker shows whether your page is in Common Crawl's record, kept since 2008 with about one new crawl a month. Type a page or a whole site to see which crawls saved it, open the copy it kept, or see why the bot left without it.
Every mark is one crawl of a few billion pages.A few early crawls, then about one a month.
What Common Crawl is
Common Crawl is a nonprofit that has archived the public web since 2008 and gives the data away free. Its bot is called CCBot.
Many AI training datasets start from these crawls, so being in the record means your page is in the raw material. Whether a model learns from it depends on what its builders keep.
A missing page usually has a plain reason: a firewall turned the bot away, robots.txt told it to stay out, the address redirects or errors, or the bot simply hasn't reached it yet.
Read more on Common Crawl's own site, or in our guide to what Common Crawl is.
Three sites, checked against the record
We checked three sites against the record on October 6, 2026, and got three different answers. Each line runs from July 1 to September 30, 2026.
- A crawl running
- Saved a copy
- Blocked
- Blocked by robots.txt
-
getseedsrighthere.com was saved in September.In July the site's firewall answered the bot with HTTP 403, and the saved response is Cloudflare's block page.
-
nytimes.com keeps the bot out with robots.txt.As of the September 2026 crawl, its robots.txt has a section for CCBot that says Disallow: /, which is why the site is missing.
-
eliteaeo.ai isn't in yet, and that's fine.The site went public Sep 25, 2026, after the September crawl had finished. It's new, not blocked.
These answers are from October 6, 2026. Run a check to see what Common Crawl has today.
Captures
What the results mean
Each result means something plain, and these are the questions people ask most after a check. For a real check read crawl by crawl, see how the Common Crawl checker reads each result.
What is Common Crawl?
Common Crawl is a nonprofit that has archived the public web since 2008 and gives the data away for free. It runs a new crawl about once a month, and each one holds a few billion pages. Many of the datasets used to train AI language models start from these crawls, so being in them is one way your pages reach what models learn.
Why isn't my page in it?
Your page may not be in Common Crawl yet because CCBot only finds pages through links and sitemaps: links from pages it already knows, and the sitemaps listed in robots.txt. A new site, or a page with few links pointing to it, can take several crawls to show up. A robots.txt rule that blocks CCBot keeps it out for as long as the rule stands, and so do the AI bot blockers some hosts and CDNs offer.
What does a redirect or error result mean?
A redirect or error result means CCBot reached the address but didn't save a page there. A 301 or 302 sent it to another URL, so check that one too; the result has a button that does it for you. A 403 or 429 usually means a firewall or rate limit turned the bot away, which is worth raising with your host. On Cloudflare, start with AI Crawl Control.
Why is the saved copy nearly empty?
A saved copy is nearly empty when the page builds its text with JavaScript, because CCBot saves the HTML your server sends and doesn't run scripts. That near-empty shell is what datasets built from Common Crawl get.
Does being in Common Crawl mean an AI model learned my page?
Being in Common Crawl doesn't mean an AI model learned your page. Being in the archive is the first step: AI labs filter Common Crawl hard before training, dropping duplicates, thin pages and spam, and a model trained on a new crawl takes months to ship. Live AI search tools such as ChatGPT search and Perplexity read pages through their own crawlers and search indexes, so Common Crawl shapes what models know, not what they find live.
Does the checker store what I check?
The checker doesn't store what you check. The check runs in your browser, and the lookups go from your browser straight to Common Crawl. Like any website, the server that hosts this page keeps logs of the page addresses people visit, and a share link carries the address you checked, so opening one is logged like any other visit.