AI Crawlers & Scraper Directory
Classify incoming automated traffic across foundation model scrapers, real-time AI search agents, and autonomous LLM web indexers.
Enterprise Crawler Directory
Tracium inspects HTTP request headers at the edge, classifying automated agents into verified foundation model scrapers, real-time search retrievers, and unverified crawlers:
| Crawler Identifier | Operator | Classification | Ingestion Purpose |
|---|---|---|---|
| GPTBot/1.2 | OpenAI | llm-training | Pre-training foundation dataset extraction |
| OAI-SearchBot/1.0 | OpenAI | ai-search | Real-time ChatGPT Search retrieval & citations |
| ClaudeBot/1.0 | Anthropic | llm-training | Claude model pre-training & retrieval synthesis |
| PerplexityBot/1.0 | Perplexity AI | ai-search | Live answer engine indexation & citations |
| Bytespider | ByteDance | aggressive-scraper | Doubao & TikTok LLM dataset scraping |
| Google-Extended | Google LLC | llm-training | Gemini and Vertex AI training corpora |
Crawl Ingestion Profiling
Every crawler request is recorded with microsecond timestamps, URI target (e.g. /sitemap.xml, /api/docs), payload byte size, response status code, and ASN origin network.
Robots.txt Directive Matrix (RFC 9309)
Govern AI search visibility independently from bulk model training scrapes using standardized robots.txt records:
# Section 1: Permit AI Search engines to index and cite content
User-agent: OAI-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
# Section 2: Prohibit foundation model pre-training ingestion
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Bytespider
Disallow: /
User-agent: CCBot
Disallow: /Origin Bandwidth & Concurrency Mitigation
Unregulated AI scrapers frequently spawn distributed crawler pools that exhaust origin server connection limits. Routing through Tracium's Edge Proxy caches static documentation responses at Cloudflare POPs, terminating 94%+ of crawler requests at the edge.