Skip to content
Traffgate Review

Reading the web without overreading the name

AC-007 Archive Craft

Common Crawl, a Second Window Into Web History

The Wayback Machine is not the only public archive of the web. Learn what Common Crawl collects, how it differs, and when it helps your research.

A research setup for checking more than one public web archive.
A research setup for checking more than one public web archive.

Common Crawl is a free, public dataset of web pages, separate from the Wayback Machine. Its corpus contains petabytes of data collected regularly since 2008, stored on Amazon Web Services' Public Data Sets and on several academic cloud platforms, with free access through Amazon's infrastructure (Common Crawl, https://commoncrawl.org/overview). For a researcher checking whether a domain was archived anywhere, it is a second place to look, not a replacement for the first.

What Is Common Crawl, Exactly?

Common Crawl is a 501(c)(3) nonprofit, founded in 2007, that runs its own web crawler and republishes what it collects as an open dataset (Common Crawl, https://commoncrawl.org/). Unlike a browsable timeline of a single URL, Common Crawl releases its data in large periodic crawl batches meant mainly for bulk analysis; a URL Index exists for checking whether a given address appears in a batch, but it is not a page-by-page viewer. The organization describes its corpus as containing petabytes of data collected regularly since 2008 and made freely accessible (Common Crawl, https://commoncrawl.org/overview). That scale and that access model are the two things that set it apart from a single-page lookup tool.

How Is Common Crawl Different From the Wayback Machine?

Wayback Machine Common Crawl
Primary use Look up one URL, browse a timeline of its captures Download large batches of crawl data for analysis
Typical user A person checking a specific page's history A researcher or developer processing many pages at once
Access Free, browsable by URL at web.archive.org Free, hosted on public cloud data sets
Coverage Driven by its own crawling and user-submitted saves Driven by its own independent crawl, since 2008

Neither source crawls the entire web. Each has its own crawl priorities, so a page missing from one may or may not appear in the other. If you treat an absence in the Wayback Machine as final, you may be missing a capture that exists only in Common Crawl's batches, or in neither. See What Public Archives Do Not Capture for the broader limits that apply to both.

When Should You Check Common Crawl Instead of, or Alongside, the Wayback Machine?

Common Crawl is most useful when you need to confirm that a page existed at web scale around a certain period, or when you are working with a dataset of many pages rather than a single URL. It is less convenient for a quick, one-off check of a single page's timeline, where a browsable tool remains faster. The two are complementary: a single-page lookup first, and a bulk dataset check when the single lookup comes up empty or when you need to confirm a pattern across many pages rather than one.

What Should You Record When You Cite Common Crawl Data?

Note which crawl batch you used, since Common Crawl releases its data as separate periodic collections rather than one continuous archive. A page's presence in one batch does not confirm its state at any date outside that batch's window. Treat a Common Crawl record the same way this site treats any other archived capture: as evidence of a version collected at a known time, not as proof of the page's entire history. See Why Archive Timestamps Need Careful Reading for the same caution applied to capture dates in general.

A Short Routine for Using a Second Archive

  1. Run your normal Wayback Machine search first, by URL and by domain.
  2. If the page is missing, thin, or you need bulk confirmation across many URLs, check whether a Common Crawl batch from the relevant period covers it.
  3. Record which source and which batch or snapshot you used, since the two are not interchangeable.
  4. If both sources are silent on a page, record that as an absence, not as proof the page never existed. See What Missing Archive Pages Can Tell You for how to write that up responsibly.
  5. When citing Common Crawl data in your own notes, link to commoncrawl.org and the batch identifier you used, the same way you would cite an archived URL's capture date.

No public archive, including the Wayback Machine, captures the whole web. Knowing that a second, independently run project exists, collecting at web scale since 2008, gives you a real second chance to confirm or rule out a capture, as long as you keep its batch-based nature and its bulk-analysis purpose in mind rather than treating it like a single-page lookup tool.

Neighbouring entries

AC-004 Archive Craft

What Missing Archive Pages Can Tell You

No Wayback snapshots for a domain? Learn what that absence means, what it does not, and how to read the record fairly without inventing history.