What is Common Crawl?

By David Quaid Published

So, what is Common Crawl? It's a nonprofit that crawls the web about once a month and gives everything it saves away free, more than 300 billion pages since 2008, and it's the raw material behind the best-known open datasets AI models train on. You can search it by web address but never by keyword, and whether your pages get in comes down to the thing that decides rankings: links.

What Common Crawl is

Common Crawl is a 501(c)(3) nonprofit that crawls the public web and publishes what it collects as an open archive.

Gil Elbaz founded it in 2007 to give researchers and small companies the kind of web data only the big search engines had. The archive now runs past 10 petabytes.

It's an archive rather than a search engine, and it isn't the whole web either. Common Crawl calls its own dataset a sample, and Mozilla's researchers found it neither complete nor representative, with English pages and frequently linked domains over-represented.

How Common Crawl collects pages with CCBot

Common Crawl collects pages with CCBot, its own crawler, built on Apache Nutch and logged as CCBot/2.0.

Each month it draws URLs from its crawl database in an order set by link-graph scores and finds new ones through links on pages it already crawled and sitemaps listed in robots.txt. There's no submission form.

CCBot reads robots.txt first, fetches with a plain HTTP GET, follows up to four redirects and saves what the server returns. It doesn't run JavaScript or keep cookies, honors nofollow and Crawl-delay, and backs off on 429 and 5xx errors. It also revisits rarely changing pages less and less often, so a page missing from the newest crawl may sit in an older one. Our CCBot guide covers the user agent, IP ranges and robots.txt rules.

What's in the Common Crawl dataset

The Common Crawl dataset is a series of monthly crawls, each in three formats: WARC files with the raw HTTP responses (headers and HTML as your server sent them), WAT files with metadata and links, and WET files with plain text.

Each crawl also ships the robots.txt files CCBot fetched and its non-200 responses (404s, redirects and so on), which is how you tell a refused page from one that was never visited.

Common Crawl size is easiest to see crawl by crawl. The September 2026 crawl saved 2.17 billion pages (361.4 TiB uncompressed) from 40.4 million hosts. Here are the 12 newest, with dates from Common Crawl's announcement posts and page counts from its own crawl list:

Crawl IDCrawledPages saved
CC-MAIN-2026-39Sep 4 to 17, 20262,171,285,702
CC-MAIN-2026-34Aug 7 to 20, 20262,139,617,681
CC-MAIN-2026-30Jul 7 to 25, 20262,149,001,456
CC-MAIN-2026-25Jun 2 to 18, 20262,098,491,742
CC-MAIN-2026-21May 8 to 21, 20262,164,140,877
CC-MAIN-2026-17Apr 10 to 23, 20262,191,936,300
CC-MAIN-2026-12Mar 5 to 17, 20261,974,845,234
CC-MAIN-2026-08Feb 6 to 19, 20262,166,991,722
CC-MAIN-2026-04Jan 12 to 25, 20262,329,630,755
CC-MAIN-2025-51Dec 4 to 17, 20252,169,087,279
CC-MAIN-2025-47Nov 6 to 19, 20252,294,472,912
CC-MAIN-2025-43Oct 5 to 19, 20252,616,796,857

Each month got one crawl of about two weeks, and every crawl saved roughly 2.0 to 2.6 billion pages. The sites behind those pages shrank, though: the same announcements put the registered domains at 38.5 million in October 2025 and 33.2 million in September 2026, about 5.3 million fewer in a year. The posts don't say why, and I won't guess, but a place in the crawl is NOT a given.

How the Common Crawl index works

The Common Crawl index is a lookup table, rebuilt for every crawl, that maps each saved URL to the file, byte offset and length where its copy sits.

The CDXJ version, behind the commoncrawl index server at index.commoncrawl.org, gives one line per capture with the URL, time, HTTP status, content type and file location, enough to fetch that one page with a single range request. The URL Index (formerly the columnar index) holds the same data as Parquet files for bulk SQL.

The catch, spelled out in the CDXJ index documentation, is that no CDXJ index covers all the crawls. Each crawl gets its own, listed in a file called collinfo.json, so checking one page across a year through the API takes 12 lookups.

The Common Crawl API people ask about is the CDX API behind that server. It matches an exact URL, a path prefix, a host, or a domain with all its subdomains, one crawl per request:

https://index.commoncrawl.org/CC-MAIN-2026-39-index?url=example.com/*&output=json

The Common Crawl FAQ warns that the endpoint is heavily rate limited because it gets abused: a 503 means slow down, and a blocked IP should wait 24 hours.

Can you search Common Crawl?

You can search Common Crawl, but only by address, never by keyword.

The index is sorted by URL, so it answers two questions, one crawl at a time: did this crawl save this page, and which pages from this site did it save? Searching by topic means downloading the WET text files and building your own search engine.

The no-code way is Crawl Record, our free Common Crawl checker. It checks a page, a site, or a site and its subdomains across the newest 3, 6 or 12 crawls, and when a crawl missed you it says why: a firewall or bot blocker answered CCBot with an error (it names which one), robots.txt kept CCBot out, the address redirects, or the crawl didn't reach the page. It also shows the saved copy, and it stores nothing.

Who uses Common Crawl data for AI training

Common Crawl data is used for AI training in most language models whose builders publish their training data, almost always after heavy filtering.

The Mozilla Foundation's 2024 report found that at least 64% of 47 text-generating LLMs released from 2019 to October 2023 trained on a filtered version of Common Crawl, usually a cleaned copy someone else made. Common Crawl's own estimate (March 2025) is 70% to 90% of the training tokens behind nearly all large language models.

The GPT-3 paper shows what filtering means. OpenAI took 41 shards of monthly crawls from 2016 to 2019, 45TB of compressed text, and kept 570GB. That was 410 billion of the 499 billion tokens in its dataset table, yet OpenAI drew only 60% of its training examples from it, because it sampled smaller, higher-quality sources more often.

The best-known open training sets are Common Crawl underneath: Google's C4 (from a single crawl, to train T5), EleutherAI's Pile-CC, the Technology Innovation Institute's RefinedWeb and Hugging Face's FineWeb (more than 18.5 trillion English tokens from crawls going back to 2013).

Who funds Common Crawl?

The Elbaz Family Foundation funds Common Crawl, as it has since Gil Elbaz founded it, and the About page still names it as the primary funder.

In a November 2025 response to The Atlantic, the Common Crawl Foundation said that some AI companies are among its newer donors, that their gifts are a small fraction of its costs and are disclosed in its financial statements, and that no donor controls what it collects. Amazon hosts the archive free through its Open Data Sponsorship Program, and about 18 staff run it, led since 2023 by Rich Skrenta, who founded DMOZ (for those of us old enough to remember it).

Is Common Crawl free?

Common Crawl is free: you can download any crawl over HTTPS from data.commoncrawl.org without an AWS account, or process it inside AWS from the public S3 bucket in us-east-1.

You pay only for your own computing: September 2026's WARC files alone are 82.79 TiB compressed, and Athena queries are billed, though Common Crawl put a full scan of one crawl's URL Index at about $1.50 (September 2025).

Why Common Crawl matters for SEO and AI visibility

Common Crawl matters for SEO and AI visibility because it feeds the open datasets AI models learn from, and CCBot picks what goes in by ranking the link graph, which in my book is Authority by another name.

Each month Common Crawl publishes host- and domain-level web graphs ranked by PageRank and by harmonic centrality, a measure of how few link hops separate a site from the rest of the web. The release posts offer the ranks up for research, but Common Crawl's About page says CCBot uses harmonic centrality to prioritize URLs (its first web graph, in 2017, produced a ranked host list for expanding the crawl frontier), and Mozilla's interviews with its crawl engineer describe the scores deciding which domains get crawled and how many pages each gets. As Common Crawl's own blog put it in January 2026, "being in the crawl becomes a prerequisite for being in the model."

CCBot doesn't read your copy to decide if you deserve a place; it measures how close you sit to the well-linked core of the web, and a site can't vote itself in any more than an undergrad can confer a degree on himself. After more than 20 years in SEO, I'd call it what it is: an Authority problem, the same one behind most indexing trouble. Miss the crawl and you miss every open training set built on it.

One caveat, because evidence beats conjecture: Common Crawl shapes what a model knows before it searches. For fresh answers, an assistant rewrites the prompt into searches (the fan-out) and reads what a search index returns, which is why AI visibility is still SEO.

What to do, in order:

  1. Check whether recent crawls saved your key pages, and if not, why.
  2. Unblock CCBot, whether it's a forgotten robots.txt rule or a firewall returning 403 while robots.txt says yes.
  3. Put your content in the HTML, because CCBot won't run JavaScript.
  4. Earn links from well-connected sites, which is the same work that moves rankings.

Those steps are the Common Crawl part of an AI visibility audit.

Common Crawl FAQ

Here are short answers to the Common Crawl questions people ask next.

How often does Common Crawl crawl the web?

Common Crawl crawls the web about once a month, and each crawl runs for about two weeks. Pages that rarely change get revisited less often, so check several crawls, not just the newest.

How big is one Common Crawl crawl?

One Common Crawl crawl holds a little over two billion pages: September 2026's saved 2.17 billion (361.4 TiB uncompressed) from 33.2 million registered domains.

Does Common Crawl run JavaScript?

Common Crawl does not run JavaScript, and CCBot doesn't use cookies either. If your content only appears after scripts run, Common Crawl saves a near-empty shell, which Crawl Record flags when it shows the saved copy.

Can I remove my site from Common Crawl?

You can remove your site from Common Crawl going forward with two lines in robots.txt, which CCBot keeps rechecking:

User-agent: CCBot
Disallow: /

Pages already published stay in the archive files; for those, email info@commoncrawl.org, and Common Crawl filters the URLs from later crawls and its public indexes and lists legal requests in its Opt-Out Ledger. Copies in other datasets are up to their maintainers; FineWeb, for one, removed domains after a cease and desist notice in January 2025.

Does being in Common Crawl mean an AI model trained on my page?

Being in Common Crawl doesn't mean an AI model trained on your page, since builders filter hard first: OpenAI kept 570GB out of 45TB for GPT-3, and C4, Pile-CC, RefinedWeb and FineWeb each apply their own filters. The reverse holds, though: a page Common Crawl never saved can't appear in a dataset built only from Common Crawl.

References

  1. About, Common Crawl. commoncrawl.org/about
  2. Frequently Asked Questions, Common Crawl. commoncrawl.org/faq
  3. CCBot, Common Crawl.
  4. Get Started, Common Crawl.
  5. CDXJ Index, Common Crawl. commoncrawl.org/cdxj-index
  6. URL Index, Common Crawl.
  7. September 2026 Crawl Archive Now Available, Common Crawl (September 19, 2026). commoncrawl.org/blog/september-2026-crawl-archive-now-available
  8. Crawl list with per-crawl statistics (collinfo-full.json), Common Crawl.
  9. Crawl Archive Now Available posts, October 2025 to September 2026, Common Crawl blog.
  10. Data Sets Containing Robots.txt Files and Non-200 Responses, Common Crawl (September 16, 2016).
  11. Measuring Crawled Coverage of a Website in Common Crawl, Common Crawl (July 20, 2026).
  12. Host- and Domain-Level Web Graphs July, August, and September 2026, Common Crawl (September 22, 2026).
  13. Common Crawl's First In-House Web Graph, Common Crawl (May 22, 2017).
  14. How SEOs Are Using Common Crawl's Web Graph Data for AI Ranking Signals, Common Crawl (January 19, 2026). commoncrawl.org/blog/how-seos-are-using-common-crawls-web-graph-data-for-ai-ranking-signals
  15. Common Crawl Foundation Opt-Out Registry, Common Crawl (September 17, 2025).
  16. Setting the Record Straight: Common Crawl's Commitment to Transparency, Fair Use, and the Public Good, Common Crawl (November 4, 2025). commoncrawl.org/blog/setting-the-record-straight-common-crawls-commitment-to-transparency-fair-use-and-the-public-good
  17. Submission to the UK's Copyright and AI Consultation, Common Crawl (March 3, 2025).
  18. Terms of Use, Common Crawl.
  19. CDX Server API, pywb documentation, Webrecorder.
  20. Language Models are Few-Shot Learners (Brown et al., 2020), arXiv. arxiv.org/abs/2005.14165
  21. Training Data for the Price of a Sandwich: Common Crawl's Impact on Generative AI (Baack, February 6, 2024), Mozilla Foundation. www.mozillafoundation.org/en/research/library/generative-ai-training-data/common-crawl/
  22. C4 dataset card, Allen Institute for AI on Hugging Face.
  23. FineWeb dataset card and changelog, Hugging Face.
  24. Common Crawl Criticized for 'Quietly Funneling Paywalled Articles to AI Developers' (November 8, 2025), Slashdot.
  25. Common Crawl, Wikipedia.