CCBot: Common Crawl's web crawler
By David Quaid Published
CCBot is Common Crawl's web crawler, the bot that collects the pages behind the free, open web archive many AI models have been trained on. It obeys robots.txt, so blocking it takes two lines, but the problem I'd look for first is the opposite: a firewall turning CCBot away while robots.txt says come in.
What CCBot is
CCBot is the crawler that Common Crawl, a nonprofit, uses to build the web archive it gives away free, so when people search "cc bot" or "Common Crawl bot", this is it.
It crawls in monthly rounds (the September 2026 crawl ran September 4 to 17 and saved 2.17 billion pages), and nothing ranks off it directly: what comes out is a dataset and a link graph. More in our guide to what Common Crawl is.
CCBot user agent
The CCBot user agent string is CCBot/2.0 (https://commoncrawl.org/faq/), and the robots.txt token is CCBot.
Common Crawl's FAQ says the version number may change, so match on the token in filters and firewall rules, NOT the full string.
Fakes exist, so check the IP: the CCBot page says the real one runs on dedicated ranges with reverse DNS, where each IP resolves to a name under crawl.commoncrawl.org that points back to the same IP (IPv6 has no reverse DNS yet; use the JSON list of ranges instead). Anything else calling itself CCBot is a scraper in a costume.
How CCBot reads robots.txt
CCBot reads robots.txt before it fetches a page and obeys the group that names it, falling back to * only when no group does, which is the rule RFC 9309 set in 2022 (our checker applies the same rules).
Inside that group the longest matching rule wins, and Allow takes a tie.
User-agent: *
Disallow: /checkout/
User-agent: CCBot
Disallow: /blog/
Allow: /blog/guides/
CCBot obeys ONLY the second group, so /checkout/ is open to it, /blog/news is closed, and /blog/guides/ccbot is open because Allow: /blog/guides/ is the longer match. That first one trips people up: a CCBot group added to be friendly hands it everything you'd hidden from *.
How to block CCBot
To block CCBot from a whole site, add these two lines to the robots.txt at the root of the host (each subdomain needs its own file):
User-agent: CCBot
Disallow: /
To keep it out of one folder only:
User-agent: CCBot
Disallow: /members/
Allowing it takes nothing, because a URL no rule matches is allowed. If your * group shuts everyone out, give CCBot its own group and repeat the paths you still want private:
User-agent: CCBot
Allow: /
Disallow: /wp-admin/
For server load, CCBot obeys Crawl-delay (per Common Crawl's FAQ, though RFC 9309 doesn't define it): Crawl-delay: 2 in its group means one request every two seconds.
Should you block CCBot?
I wouldn't block CCBot, because blocking it takes your site out of the open datasets many AI models are trained on and out of Common Crawl's web graph, and it takes back nothing already saved.
Common Crawl's blog cites a 2024 Mozilla review where 64% of 47 LLMs used a filtered version of its data, and that post makes the obvious point that a site has to be in the crawl before it can be in the model.
The web graph is the part SEOs miss. Common Crawl turns the links CCBot finds into a host-level graph (245.8 million hosts in the latest release) with harmonic centrality and PageRank for each; block CCBot and your links never get in, and at best you're a dead end other sites link to.
The honest reason to block is not wanting your writing in training sets, and that's your call. It's a training-data decision, not a ranking one: AI answers cite what an assistant finds when it searches (the fan-out), which is why GEO is still SEO.
Why CCBot gets blocked when robots.txt allows it
CCBot gets blocked when robots.txt allows it because robots.txt is a request a polite crawler honors, while a firewall, CDN or security plugin (Cloudflare, Wordfence, Sucuri and the rest) answers first and never reads the file.
The answer is a 403 or a 429, and Common Crawl saves it like any other response, so the record holds the firewall's page instead of yours.
Cloudflare's own docs show what that record looks like. It carries a cf-ray header, which Cloudflare adds to every response it serves. Cloudflare's block page is a branded 403, which the docs tie to a WAF rule or security setting; Error 1020 (Access denied) means a firewall rule refused the request; a challenge page waits for the browser to run JavaScript, which CCBot doesn't, so the challenge itself gets saved; and rate limiting answers 429 by default. An unbranded 403 came from the server behind Cloudflare, which our checker reports as the site's own server.
The robots.txt request can get caught too: under RFC 9309 a 4xx there lets the crawler fetch anything, but a 5xx means complete disallow, so a firewall throwing 503s at CCBot has written Disallow: / for you. Blocking AI crawlers on purpose is another matter: how Cloudflare AI Crawl Control decides who gets in.
How to check if CCBot reached your site
To check if CCBot reached your site, look in two places: your logs, for what knocked, and Common Crawl's index, for what it kept.
Logs miss one thing: a request your CDN blocked never reaches your server, so a Cloudflare block only shows in Cloudflare's Security Events log.
For the index, run the address through Crawl Record, our free Common Crawl checker. It looks a page or a whole site up in the last 3, 6 or 12 crawls and gives a reason for every miss: robots.txt, a redirect, a page the crawl never reached, or a firewall, named (Cloudflare, Wordfence and ten others, or the site's own server). It also shows the saved robots.txt as CCBot and eleven other AI bots read it, plus the saved page.
What most people overlook about CCBot
What most people overlook about CCBot is that being allowed in doesn't mean being saved.
Common Crawl samples a random subset of each site, and the same blog post says harmonic centrality (a link-graph score) sets crawl priority, with higher-scoring sites crawled more often. So how much of your site lands in a crawl is an Authority question, as it is with Google. In more than 20 years of SEO I've found most indexing trouble is Authority trouble, and Common Crawl's crawler works the same way.
CCBot FAQ
Short answers to the CCBot questions that come up next.
Does CCBot run JavaScript?
CCBot doesn't run JavaScript or use cookies, according to Common Crawl's FAQ. A page that builds its text in the browser gets saved as the HTML the server sent first, often an empty shell, and Crawl Record warns when a saved copy is nearly empty.
How often does CCBot crawl a site?
How often CCBot crawls a site comes down to Common Crawl's schedule and your Authority. Crawls run about monthly (the August 2026 crawl ran August 7 to 20), and sites with higher harmonic centrality get crawled more often.
Is CCBot used to train ChatGPT?
CCBot data was used to train GPT-3, the OpenAI model that came before ChatGPT: the GPT-3 paper puts filtered Common Crawl, drawn from 41 monthly shards, at 60% of its training mix. For the models behind ChatGPT today nobody outside OpenAI can say, since the GPT-4 technical report stopped describing the dataset.
Does blocking CCBot remove pages it already saved?
Blocking CCBot doesn't remove pages it already saved. Common Crawl's FAQ describes the block in the future tense (the crawler stops from then on), nothing on its CCBot page, FAQ or opt-out post says older crawls get edited, and datasets already built from them keep what they took. Legal opt-out requests go in its public Opt-Out Ledger.
Does blocking CCBot affect Google rankings?
Blocking CCBot doesn't affect Google rankings, because a robots.txt group only binds the crawler it names, and Googlebot never reads a CCBot group. The only risk is a firewall rule written too wide, so if you block there, match the CCBot user agent alone.
References
- CCBot, Common Crawl. commoncrawl.org/ccbot
- Frequently Asked Questions, Common Crawl. commoncrawl.org/faq
- RFC 9309: Robots Exclusion Protocol, IETF (M. Koster, G. Illyes, H. Zeller, L. Sassman, September 2022). www.rfc-editor.org/rfc/rfc9309.html
- September 2026 Crawl Archive Now Available, Common Crawl (September 19, 2026). commoncrawl.org/blog/september-2026-crawl-archive-now-available
- August 2026 Crawl Archive Now Available, Common Crawl (August 24, 2026). commoncrawl.org/blog/august-2026-crawl-archive-now-available
- Host- and domain-level web graphs, July, August and September 2026 crawls, Common Crawl. data.commoncrawl.org/projects/hyperlinkgraph/cc-main-2026-jul-aug-sep/index.html
- How SEOs Are Using Common Crawl's Web Graph Data for AI Ranking Signals, Common Crawl (Stephen Burns, January 19, 2026). commoncrawl.org/blog/how-seos-are-using-common-crawls-web-graph-data-for-ai-ranking-signals
- Common Crawl Foundation Opt-Out Registry, Common Crawl (September 17, 2025). commoncrawl.org/blog/common-crawl-foundation-opt-out-registry
- Language Models are Few-Shot Learners, Brown et al., OpenAI (2020). arxiv.org/abs/2005.14165
- GPT-4 Technical Report, OpenAI (2023). arxiv.org/abs/2303.08774
- Error 403, Cloudflare Docs. developers.cloudflare.com/support/troubleshooting/http-status-codes/4xx-client-error/error-403/
- Error 1020, Cloudflare Docs. developers.cloudflare.com/support/troubleshooting/http-status-codes/cloudflare-1xxx-errors/error-1020/
- Interstitial Challenge Pages, Cloudflare Docs. developers.cloudflare.com/cloudflare-challenges/challenge-types/challenge-pages/
- Create a rate limiting rule in the dashboard, Cloudflare Docs.
- Troubleshooting a slow website, Cloudflare Docs.