Tag
QuixiAI releases SYN-1B, a synthetic pretraining dataset designed to teach models rule tracking, belief updating, and evidence preservation over long contexts, available on Hugging Face.
This paper addresses the limitation of static surprisal in LLM-based scientific discovery by introducing evidence-informed non-stationary beliefs, and proposes belief-update filtering and diversity maximization to improve discovery, achieving 30.62% higher non-stationary surprisal across five domains.