Automated web client operated by an AI company for model training, AI search, or user-requested retrieval.
AI crawlers are automated clients that request web pages for several distinct purposes: model training, AI search indexing, or retrieval initiated by a user. Documented examples include GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, and PerplexityBot. A robots.txt control token is not necessarily an observable HTTP user agent: Google-Extended, for example, controls Google's AI use of content fetched by other crawlers and should not be reported as a separate visit. Allowing a crawler only permits access; it does not guarantee indexing, training use, retrieval, or citation.
Qwairy classifies documented, observable AI user agents in connected server logs. Crawler Analytics shows observed occurrence totals, retained frequency, requested page groups, HTTP responses, and coverage caveats without treating robots.txt control tokens as visits.
OpenAI's web crawler that collects public web content to train future GPT models; site owners control its access via robots.txt.
Anthropic's web crawler that collects public web content to train and improve Claude models; site owners control its access via robots.txt.
User-agent OpenAI uses when ChatGPT fetches webpages in real time during a conversation, distinct from the GPTBot training crawler.
OpenAI's search crawler that indexes the web to power ChatGPT Search, separate from GPTBot (training) and ChatGPT-User (live browsing).
Robots.txt control token letting publishers decide whether Google may use crawled content to train and ground its Gemini AI models.
Perplexity AI's web crawler used to index and retrieve content for real-time AI search responses.
Common Crawl's web crawler that collects data to create an open repository of web crawl data.
Text file placed at the root of a website to indicate to indexing robots which pages to explore or avoid.
Proposed file format offering a structured summary of a site's content to optimize its understanding by LLMs.
AI architecture that retrieves relevant information from external sources in real-time before generating responses.
Practice of systematically tracking where, how often, and in what context a brand or its content is cited in AI responses.
When an AI model generates factually incorrect, fabricated, or misleading information presented as truth.