Tag
This paper studies the trade-off in repeating high-quality domain data during LLM pretraining to maintain performance as models scale, finding that optimal repetition counts increase with model size and are negatively correlated with domain validation loss.
KletterMix is a high-quality German pretraining corpus built by translating a state-of-the-art English pretraining dataset into German while preserving structure and diversity. Controlled experiments show models trained on KletterMix achieve measurable improvements on German-language benchmarks.
A new AI model is being trained on over 100 trillion tokens, doubling the typical pretraining data size of 27-50 trillion tokens used by other models like Kimi, Mimo, and DeepSeek.
A unified survey of pretraining data exposure (PDE) in large language models, covering membership inference, data contamination, and security implications, with a review of attack and defense methods.