Tracium
v1.2 Telemetry

AI Crawlers & Scraper Directory

Classify incoming automated traffic across foundation model scrapers, real-time AI search agents, and autonomous LLM web indexers.

Robots Exclusion Standard (RFC 9309)User-Agent Signature EngineL7 Crawler Tagging

Enterprise Crawler Directory

Tracium inspects HTTP request headers at the edge, classifying automated agents into verified foundation model scrapers, real-time search retrievers, and unverified crawlers:

Crawler IdentifierOperatorClassificationIngestion Purpose
GPTBot/1.2OpenAIllm-trainingPre-training foundation dataset extraction
OAI-SearchBot/1.0OpenAIai-searchReal-time ChatGPT Search retrieval & citations
ClaudeBot/1.0Anthropicllm-trainingClaude model pre-training & retrieval synthesis
PerplexityBot/1.0Perplexity AIai-searchLive answer engine indexation & citations
BytespiderByteDanceaggressive-scraperDoubao & TikTok LLM dataset scraping
Google-ExtendedGoogle LLCllm-trainingGemini and Vertex AI training corpora

Crawl Ingestion Profiling

Every crawler request is recorded with microsecond timestamps, URI target (e.g. /sitemap.xml, /api/docs), payload byte size, response status code, and ASN origin network.

Robots.txt Directive Matrix (RFC 9309)

Govern AI search visibility independently from bulk model training scrapes using standardized robots.txt records:

/robots.txt
# Section 1: Permit AI Search engines to index and cite content
User-agent: OAI-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

# Section 2: Prohibit foundation model pre-training ingestion
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Bytespider
Disallow: /

User-agent: CCBot
Disallow: /

Origin Bandwidth & Concurrency Mitigation

Unregulated AI scrapers frequently spawn distributed crawler pools that exhaust origin server connection limits. Routing through Tracium's Edge Proxy caches static documentation responses at Cloudflare POPs, terminating 94%+ of crawler requests at the edge.