Common Crawl's web crawler that collects data to create an open repository of web crawl data.
CCBot (Common Crawl Bot) is the web crawler operated by Common Crawl, a nonprofit that maintains an open repository of web crawl data. Many AI companies and researchers use Common Crawl's dataset to train their models, including several large language models. While CCBot itself is not owned by an AI company, allowing it means your content may be included in datasets used for AI training by multiple organizations. The Common Crawl dataset is one of the largest publicly available web archives.
Qwairy can classify observed CCBot requests in connected server logs. The observation confirms a request, not inclusion in a particular archive, dataset, or downstream model.
Automated web client operated by an AI company for model training, AI search, or user-requested retrieval.
Text file placed at the root of a website to indicate to indexing robots which pages to explore or avoid.
OpenAI's web crawler that collects public web content to train future GPT models; site owners control its access via robots.txt.