tink / links #systemcrafters

< bergheim > dropped a link

Why Sites Should Block AI Crawlers and How to Do It

https://www.jamescherti.com/why-sites-should-block-ai-crawlers-and-how-to-do-it/

James Cherti argues that independent publishers should block AI training crawlers like GPTBot and ClaudeBot to stop resource theft without sacrificing referral traffic, relying on the technical distinction between bots used for model training and those used for live search indexing. The piece leans heavily on Cloudflare Radar data showing extreme crawl-to-refer ratios—Anthropic at 2,585:1 and OpenAI at 685:1—to demonstrate that these companies scrape thousands of pages for every single click they return. Cherti provides a specific robots.txt configuration and server-level enforcement methods (Nginx, Apache, or Cloudflare AI Crawl Control) to drop these specific user agents while keeping search bots like OAI-SearchBot and PerplexityBot active.

The advice is practical but relies on the assumption that AI companies will maintain separate crawler identities indefinitely, a premise Cherti admits requires periodic auditing as infrastructure changes. He correctly notes that robots.txt is advisory and suggests server-level blocking for real protection, though he understates the difficulty of distinguishing legitimate user-triggered agents from scrapers at the edge. The core argument holds up for bandwidth-conscious sites: if you do not need your content in a specific LLM's training data, blocking its dedicated crawler costs you nothing in search visibility, provided you avoid accidentally blocking the distinct agents responsible for citation and retrieval.