Tag
This paper assesses the prevalence of extremist speech in the training data of large language models, focusing on the Dolma corpus, and discusses implications for data curation and model safety.
An update on the escalating problem of AI scraper bots overwhelming websites, discussing residential proxy networks and their impact on the open web.
A personal blog post explaining an anti-crawler browser blocking mechanism, including special notes for Inoreader, Feedly, Vivaldi, and archive.* users.