Amazonbot is finally respecting robots.txt
Summary
Amazonbot, Amazon's web crawling bot, now respects robots.txt directives, marking a change in its previous behavior.
View Cached Full Text
Cached at: 05/14/26, 09:26 PM
Similar Articles
Patreon stops asking AI bots not to scrape — and starts blocking them
Patreon has partnered with Cloudflare to actively block AI bots from scraping creator content for training, moving beyond the voluntary robots.txt approach. The move aims to give creators more control over how their work is used by AI companies.
Sites that block AI training crawlers mostly ignore the answer time bots
A study of robots.txt files from top 10,000 sites reveals that most block AI training crawlers like GPTBot while largely ignoring answer-time bots such as OAI-SearchBot, highlighting a blind spot in current web governance for AI.
Each AI agent crawls website completely differently. Here's what 3 mons of 11 million event logs actually show.
Analysis of 11 million crawler logs across 34 websites reveals distinct behaviors: GPTBot crawls relentlessly ignoring robots.txt, Google's bot checks rules frequently, ClaudeBot's crawling is rapidly accelerating, and Bytespider is the heaviest crawler. The findings suggest a shift from Google-centric SEO to optimizing for AI agent page selection.
How does AI follow ethical guidelines in Data Collection?
A commentary on the ethical challenges of AI agents ignoring website rules like robots.txt when generating scrapers, and the responsibility of AI providers to implement guardrails without hindering product usability.
Spent an afternoon making my site more AI friendly. The next day AI traffic went 12x
A developer optimized their website for AI bots by fixing robots.txt, adding llms.txt, improving semantic HTML, and more, resulting in a 12x increase in AI traffic the next day.