AI crawling overwhelms git.kernel.org servers
write-up
· for bergheim
in #systemcrafters
· 2026-09-08 15:23 UTC
tl;dr: AI crawlers overwhelm git.kernel.org, consuming more CPU than legitimate users by pointlessly scraping commit pages as HTML.
- Scraping commits as HTML wastes resources; cloning the repos would be far more efficient for training data.
- Linux.git has ~1.48M commits, with ~922 forks (mostly shared objects), yet scrapers hit billions of duplicate URLs.
- cgit's flexible URL generation enables ~1.2 billion valid URLs per fork, making blocking hard.
- Initial blocking via user-agent and IP bans worked until bots spoofed browsers and spread across subnets/ASNs.
- Crawler traffic now comes from millions of IPs (e.g., "your TV"), making simple bans impractical.
- At any time, 14 CPU cores across 5 nodes are dedicated solely to rendering commits for scrapers.
See also
Hacker News · 1389 pts · 710 comments — https://news.ycombinator.com/item?id=49491791
Lobsters · 140 pts · 60 comments — https://lobste.rs/s/nbjo0i/creepy_crawlies
Commenters largely describe AI crawlers as a worsening burden on small web hosts, citing inflated bandwidth, server load, and fake user sessions that choke formerly healthy sites. A recurring frustration is that crawlers ignore obviously more efficient methods like git clone for repositories, instead hammering web interfaces like cgit with unnecessary page hits. Several note that existing defenses such as Anubis are failing, as a third of bots now solve the math challengescaping the blocks, while others propose adaptive rate-limiting or identity requirements. There's also genuine, unresolved debate about whether the solution is to optimize web infrastructure for these new traffic patterns, train AIs to be better "netizens," or accept the decline of anonymous, community-hosted information services.
Source: https://people.kernel.org/monsieuricon/creepy-crawlies