llm-training-data

Tag

Cards List
#llm-training-data

Beyond the pale: Assessing prevalence and contents of extremist speech in LLM training data

arXiv cs.CL · 2026-08-18 Cached

This paper assesses the prevalence of extremist speech in the training data of large language models, focusing on the Dolma corpus, and discusses implications for data curation and model safety.

0 favorites 0 likes
#llm-training-data

An Update on the scraper situation

Hacker News Top · 2026-07-10 Cached

An update on the escalating problem of AI scraper bots overwhelming websites, discussing residential proxy networks and their impact on the open web.

0 favorites 0 likes
#llm-training-data

Notes about reading messages with the Python email packages

Hacker News Top · 2026-05-20 Cached

A personal blog post explaining an anti-crawler browser blocking mechanism, including special notes for Inoreader, Feedly, Vivaldi, and archive.* users.

0 favorites 0 likes
← Back to home

Submit Feedback