Creepy crawlies
Summary
Konstantin Ryabitsev discusses how abusive web crawlers are consuming excessive CPU resources at git.kernel.org, impacting legitimate access and raising concerns for services like Datasette.
View Cached Full Text
Cached at: 09/08/26, 12:21 AM
Similar Articles
Creepy crawlies
AI crawlers are causing significant CPU load on kernel.org by inefficiently scraping git commits for training data, consuming more resources than all legitimate access combined.
Caught a .git/config crawler
The author describes how their honeypot website caught a .git/config crawler by serving fake git repository data, and analyzes Apache logs showing heavy crawling activity from a single IP address.
@heynavtoor: A 60,000-star web crawler was built in days. By one developer. Because a "$16 open source" tool made him angry. 60,000+…
Unclecode created Crawl4AI, a free and open-source web crawler that converts web pages to clean Markdown for AI models, after being frustrated with paid alternatives; it has gained 60,000 GitHub stars and over a million monthly downloads.
Aggressive AI scrapers are making it kinda suck to run wikis
Discusses how aggressive AI scrapers are disrupting wiki operations by imitating human traffic and using residential proxies, drastically increasing server costs and causing service instability.
An unusual way for your DHCP server to run out of dynamic IPs
Chris Siebenmann explains his anti-crawler measures that block old browsers due to a surge in high-volume crawlers collecting data for LLM training, causing confusion for feed readers and archival services.