NeuralCrawl

About

What is NeuralCrawl?

NeuralCrawl was created by Elie Berreby to continuously archive the robots.txt files of major US & European companies, democratic governments and the largest social networks. Whenever a file changes, the new version is stored, diffed against the previous one, and the change is explained in plain language — with special attention to AI crawler policy (GPTBot, ClaudeBot, CCBot…).

According to Elie Berreby, the Robots Exclusion Protocol is the closest thing the web has to a public policy document: it records, in machine-readable form, which user-agent an organisation lets read its website. Watching the crawling directive evolve across the corporate world and government reveals how attitudes towards search engines, scrapers and AI training are shifting.

Neural Crawl is the visible layer of a larger crawler: a starting point for OSINT analysts, journalists and researchers who work with reproducible, citable evidence, not anecdotes. The monitoring pipeline runs deeper than what is published and will keep expanding.

Elie Berreby

Methodology

AI crawler registry (28 bots)

User-agentVendorPurpose
GPTBotOpenAILLM training crawler
ChatGPT-UserOpenAIUser-triggered browsing
OAI-SearchBotOpenAISearch indexing
ClaudeBotAnthropicLLM training crawler
Claude-UserAnthropicUser-triggered browsing
Claude-SearchBotAnthropicSearch indexing
anthropic-aiAnthropicLegacy crawler token
Claude-WebAnthropicLegacy crawler token
CCBotCommon CrawlOpen web corpus (used for LLM training)
Google-ExtendedGoogleGemini training opt-out token
Applebot-ExtendedAppleApple Intelligence training opt-out
PerplexityBotPerplexityAnswer-engine indexing
Perplexity-UserPerplexityUser-triggered browsing
BytespiderByteDanceLLM training crawler
AmazonbotAmazonAlexa / LLM crawler
FacebookBotMetaMeta AI crawler (legacy)
meta-externalagentMetaMeta AI training crawler
Meta-WebIndexerMetaMeta AI search indexing & citations
meta-externalfetcherMetaUser-triggered fetcher
cohere-aiCohereLLM training crawler
AI2BotAllen Institute for AIResearch crawler
DiffbotDiffbotStructured data extraction
omgiliWebz.ioData feeds resold for AI training
YouBotYou.comAnswer-engine indexing
DuckAssistBotDuckDuckGoDuckAssist answers
MistralAI-UserMistral AIUser-triggered browsing
PanguBotHuaweiLLM training crawler
TimpibotTimpiDecentralised index crawler

Using the site

Bulk data export and programmatic API access are not offered on the public site. Contact the creator for commercial licensing.