Skip to main content

What is CCBot?

CCBot is the crawler of the non-profit Common Crawl Foundation, which builds a freely available archive of large parts of the public web. That archive is a widely used raw data source for research and for training language models.

Common Crawl runs neither a search engine nor a chatbot. It publishes its data openly, and many projects reuse it. That is why CCBot appears on almost every list of AI-relevant crawlers even though it does not produce answers itself.

Blocking CCBot works going forward: new crawls no longer capture the site. Whatever is already in earlier archives stays there. Common Crawl itself describes the block as a robots.txt group for CCBot with a disallow for everything.

Common Crawl also notes that other crawlers falsely identify themselves as CCBot. A log entry with this user agent is therefore not necessarily from Common Crawl. Keep that in mind if your firewall rules filter by name alone.

What it means for your website

Decide whether your content should end up in an open archive that others use for model training. If not, block CCBot in robots.txt. The deeploupe GPTBot check reads the CCBot rule and shows whether it applies. It does not fetch your page with the CCBot user agent. Keep in mind that a block only affects new crawls and does not remove versions of your site that are already archived.

Related terms

More on deeploupe

Sources

  1. Common Crawl: CCBot

Terms help you understand. Whether AI crawlers can reach your site is something you measure.

Check your website for freefree · no signup · no credit card

All terms in the GEO glossary