Notes about reading messages with the Python email packages

Hacker News Top News

Summary

A personal blog post explaining an anti-crawler browser blocking mechanism, including special notes for Inoreader, Feedly, Vivaldi, and archive.* users.

No content available
Original Article
View Cached Full Text

Cached at: 05/25/26, 12:46 PM

# You're using a too-old browser Source: [https://utcc.utoronto.ca/~cks/cspace-bad/cspace-old-browser.html](https://utcc.utoronto.ca/~cks/cspace-bad/cspace-old-browser.html) ## You're using a suspiciously old browser You're probably reading this page because you've attempted to access some part of[my blog \(Wandering Thoughts\)](https://utcc.utoronto.ca/~cks/space/blog/)or[CSpace](https://utcc.utoronto.ca/~cks/space/), the wiki thing it's part of\. Unfortunately you're using a browser version that my anti\-crawler precautions consider suspicious, most often because it's too old \(most often this applies to versions of Chrome\)\. Unfortunately, as of early 2025 there's a plague of high volume crawlers \(apparently in part to gather data for LLM training\) that use a variety of old browser user agents, especially Chrome user agents\. To reduce the load on[Wandering Thoughts](https://utcc.utoronto.ca/~cks/space/blog/)I'm experimenting with \(attempting to\) block all of them, and you've run into this\. If this is in error and you're using a current version of your browser of choice, you can contact me at[my current place at the university](https://www.cs.toronto.edu/~cks/)\(you should be able to work out the email address from that\)\. If possible, please let me know what browser you're using and so on, ideally with its exact User\-Agent string\. ## A special note to people using Inoreader \(the feed reader\) I am not blocking Inoreader's feed fetcher or considering it to be too old, and it routinely fetches feeds from me\. I don't know why Inoreader is showing you this page\. It is possible that they're periodically trying to fetch feeds or pages with an old browser HTTP User\-Agent \(or an actual old browser\) and taking the results of that fetch \(this page\) as what they should show people instead of the results of their syndication feed fetcher agents\. This is a bad mistake today;[the results of modern HTTP fetches depend partly on the HTTP User\-Agent used](https://utcc.utoronto.ca/~cks/space/blog/web/HTTPResultsAndUserAgents)\. ## A special note to people using Feedly \(the feed reader\) Much like Inoreader, Feedly is periodically fetching my syndication feeds with a fake, old browser HTTP User\-Agent header, which fails, and is then grimly latching on to the results for their actual feed fetching with their regular Feedly HTTP User\-Agent\. There is nothing I can do about this; you should contact Feedly support, if you can find them\. See[this comment of mine in Wandering Thoughts](https://utcc.utoronto.ca/~cks/space/blog/web/FeedReaderErrorsProblem?showcomments#cks-20260215115901)for more details\. ## A special note for people using Vivaldi Due to an ongoing attack, you may need to change[the "User Agent Brand Masking" setting](https://help.vivaldi.com/desktop/miscellaneous/user-agent-brand-masking/)so that your Vivaldi identifies itself as Vivaldi, instead of Google Chrome\. This applies to even the current version of Vivaldi\. ## A special note for people using archive\.\* You may be seeing this through archive\.today, archive\.ph, archive\.is, and so on\. Unfortunately, archive\.\* crawls pages to archive in a way that is impossible to distinguish from malicious actors\. They use old Chrome User\-Agent values, crawl from IP address blocks that are widely distributed and not clearly identified as theirs, and some of their IP addresses have falsified reverse DNS entries that claim they are googlebot IP addresses \(which is something that is normally done only by quite bad actors\)\. I suggest that you use archive\.org, which is a better behaved archival crawler and can crawl[my blog \(Wandering Thoughts\)](https://utcc.utoronto.ca/~cks/space/blog/)\. Chris Siebenmann, 2025\-02\-17

Similar Articles

Notes on using GNU Emacs' Tramp system in an unusual shell environment

Lobsters Hottest

The author explains that their blog is blocking requests from old or suspicious browser user agents to mitigate a surge in high-volume crawlers, likely for LLM training data. Specific instructions are provided for users of Vivaldi and Inoreader to adjust settings or report issues.

A modern feed reader (2024)

Lobsters Hottest

The author examines the decline of RSS feeds due to scraping and interference, arguing that modern feed readers must integrate alternative syndication methods to remain relevant.

How do you sieve/filter/manage your internet mail?

Lobsters Hottest

A discussion on lobste.rs asking for advice on managing email, filtering, and tooling, with a focus on FOSS solutions and workflows for handling high volumes of mailing lists and patches.

An Update on the scraper situation

Hacker News Top

An update on the escalating problem of AI scraper bots overwhelming websites, discussing residential proxy networks and their impact on the open web.

A peek into Reddit's anti-spam internals

Lobsters Hottest

A blog post revealing Reddit's anti-spam internals, exposed by a bug, detailing how Reddit's sitewide spam filters and moderation system work.