How the Common Crawl checker works

The Common Crawl checker runs in your browser: it asks Common Crawl's public URL index about your address, crawl by crawl, and tells you what the bot found and why.

Reading Common Crawl's URL index

The check starts by reading Common Crawl's URL index, and your browser does that directly.

Common Crawl runs a new crawl about once a month, and each crawl has an index: every address its bot fetched, when, the HTTP status it got, and where the saved copy sits in the archive.

The check asks the index for the newest 3, 6 or 12 crawls, one at a time with a short pause between. "This page" looks up one exact address; http, https, www and no www count as the same page. "Whole site" covers every address on that host, and "Site and subdomains" adds addresses like blog.example.com too.

For some results it also pulls a saved record or two from the archive, as the sections below explain.

What each result means

Most results mean one of five things, and the check names each one in plain words.

Below, a real check of the getseedsrighthere.com home page from October 6, 2026, shows three of them. Any other answer reads "Reached it, got HTTP" and the status, such as 404 for a missing page or 500 for a server error, and Common Crawl kept no copy.

Example: the getseedsrighthere.com home page in the newest 3 crawls, checked October 6, 2026. Each crawl ran for two weeks; each column is one day.
  • Saved a copy
  • Blocked
  • Redirect
  • Not in this crawl
  • Side scale: hour of the day, UTC

July 2026CC-MAIN-2026-30Jul 10 to 23

Blocked: got HTTP 403

August 2026CC-MAIN-2026-34Aug 7 to 20

Not in this crawl

September 2026CC-MAIN-2026-39Sep 4 to 17

Saved 3 copies

Saved a copy

When Common Crawl saved a copy, it stored the page itself in that crawl. For this home page the September 2026 row reads "Saved 3 copies", and the "Open the newest copy" button shows exactly what was kept.

Not in this crawl

Not in this crawl means that crawl's index has no record of the address. Each crawl reaches part of the web, so a quiet month is normal. New sites wait their turn: eliteaeo.ai went public on Sep 25, 2026, after the September crawl ended. It's new, not blocked.

Got a 301 redirect

A row reads "Got a 301 redirect" when the bot reached the address and was sent on to another one, and the "Check the redirect target" button follows it for you. In September the http:// version of this home page sent the bot on to the https:// version, which it saved.

Blocked: got HTTP 403

A row reads "Blocked: got HTTP 403" when the bot asked for the page and the server refused it, and 401, 429 or a similar refusal reads the same way. This home page answered CCBot with HTTP 403 on Jul 12, 2026, at 08:49 UTC, and the check named who sent it: "Blocked by Cloudflare."

Blocked by robots.txt

Blocked by robots.txt means the site's robots.txt tells CCBot to stay out, and Common Crawl obeys it, so the check says "Your robots.txt blocks CCBot." nytimes.com has Disallow: / in a section for CCBot, which is why its pages are missing.

How the blocker is named

The check names the blocker from the response Common Crawl saved, because the archive keeps what its bot was shown, error pages included.

When a crawl shows a refusal, the check downloads that saved response and reads its headers and text. Security services leave marks: Cloudflare's block page says "Sorry, you have been blocked", and Wordfence, Sucuri, Imunify360, SiteGround, ModSecurity, Akamai, Imperva, DataDome, HUMAN, Vercel and CloudFront each have their own. A refusal routed through Cloudflare without a Cloudflare page points to the site's own server. When nothing matches, the check says so.

From the response Common Crawl saved for getseedsrighthere.com on Jul 12, 2026, at 08:49 UTC. An excerpt; the shaded lines name Cloudflare.
  1. HTTP/1.1 403
  2. server: cloudflare
  3. cf-ray: a19ecb38ecd3f283-IAD
  4. …
  5. <title>Attention Required! | Cloudflare</title>
  6. …
  7. <h1 data-translate="block_headline">Sorry, you have been blocked</h1>

That is why the check says "Blocked by Cloudflare." The fix is in the site's Cloudflare settings, where AI bot blocking, Bot Fight Mode or a custom rule can turn CCBot away.

robots.txt as CCBot read it

CCBot read the robots.txt that Common Crawl saved during the crawl, and that copy may differ from the file on your site today.

The check reads it by the rules of the robots.txt standard, RFC 9309. Rules come in groups that open with User-agent lines. A group naming CCBot beats the group for all bots (User-agent: *), and CCBot then ignores the * group. Within the group that applies, the longest matching rule wins, and Allow wins a tie.

Example robots.txt, written for this page, not from a real site. CCBot follows lines 4 to 6, its own group.
  1. User-agent: *
  2. Disallow: /cart/
  3.  
  4. User-agent: CCBot
  5. Disallow: /private/
  6. Allow: /private/press/

What that file means for CCBot at three addresses:

/cart/
Allowed by robots.txt. Nothing in CCBot's group matches, and the rules for all bots don't apply to it.
/private/notes
Blocked by robots.txt. Disallow: /private/ matches.
/private/press/kit.pdf
Allowed by robots.txt. Both rules match, and the longer Allow wins.

robots.txt is one layer and the firewall is another, so the robots.txt result always names the layer. With no rule for CCBot, getseedsrighthere.com's rules for all bots allow its home page, yet its firewall said no in July:

  • Allowed by robots.txt.
  • But the server blocked it in the July 2026 crawl (HTTP 403 from Cloudflare).

The check also lists any sitemaps the file names.

The other AI bots

The check reads the same saved robots.txt for eleven other AI bots, with no extra lookups.

They are GPTBot, OAI-SearchBot and ChatGPT-User (OpenAI), ClaudeBot and Claude-SearchBot (Anthropic), Google-Extended, PerplexityBot, Applebot-Extended, Meta-ExternalAgent, Bytespider and Amazonbot, each read the same way as CCBot. The check also flags Cloudflare's managed robots.txt, where Cloudflare writes AI bot rules into the file, and shows any Content-Signal lines.

robots.txt only asks, though. A firewall can still refuse a bot the file allows, and a bot can ignore the file.

Why a saved copy can be nearly empty

When a page builds its text with JavaScript, the saved copy can be nearly empty.

CCBot saves the HTML your server sends and doesn't run scripts. If your words load later through JavaScript, the copy is a shell with a title and little text. The check warns when a copy has almost no words, and can show it with scripts off.

Limits of the check

Three limits of the check come up often.

Common Crawl's index is sometimes busy. It can answer slowly or with errors, and it slows down a connection that asks too much. The check retries each crawl up to three times; if two crawls in a row still fail, it stops and gives each row a Try again button. Waiting a minute usually clears it.

Being crawled is not the same as being in an AI model. Many AI training datasets start from these crawls, but their builders choose what to keep. A saved page is raw material; whether a model learned from it depends on those choices.

Whole site checks read up to 5,000 records per crawl, so very large sites hold more than the count shows.

What the checker stores

The checker stores none of your checks, because the check runs in your browser.

When you press Check, your browser sends the lookups straight to Common Crawl's index and archive, and the answers come back to your screen without passing through us. Common Crawl's servers answer those lookups, so they see what was asked, as with any request to them.

Like any website, the server that hosts this site keeps logs of the page addresses visited. A share link carries the address you checked in its own address, so opening one is logged like any other page visit.

See what Common Crawl kept of your site.

Check a page or site About this checker