100 Trillion+ Pretraining data??? This is the largest data I've see a model being trained on.
Summary
A new AI model is being trained on over 100 trillion tokens, doubling the typical pretraining data size of 27-50 trillion tokens used by other models like Kimi, Mimo, and DeepSeek.
Similar Articles
Scalable and Efficient Joint Spiking Embedding Predictive Architecture for Large-Scale Dynamic Graphs
Proposes SG-JEPA, a joint spiking embedding predictive architecture for large-scale dynamic graphs that partitions nodes into context and target sets along the temporal dimension to learn predictive embeddings, achieving competitive performance on node classification while scaling to graphs with 13 million edges and avoiding complex self-supervised mechanisms.
Can Kimi K3 solve the same problems that Claude Fable can?
A discussion questioning whether open-source models like Kimi K3 or GLM can replicate the mathematical and cybersecurity problem-solving achievements recently demonstrated by closed-source models from OpenAI and Anthropic.
@BhavinJawade: ๐ข๐ป-๐ฝ๐ผ๐น๐ถ๐ฐ๐ ๐ฑ๐ถ๐๐๐ถ๐น๐น๐ฎ๐๐ถ๐ผ๐ป ๐ถ๐๐ป'๐ ๐ฎ ๐ณ๐ฟ๐ฒ๐ฒ-๐น๐๐ป๐ฐ๐ต On-policy distillation has become a defaultโฆ
Bhavin Jawade discusses several failure modes of on-policy distillation for training large language models, including early mistakes becoming uncorrectable, stronger teachers being worse, privileged information conditioning failing to transfer, and thinking collapse from dense supervision.
@jakevin7: Today I directly used GPT5.6-Sol to build Maka's official website. Found that Sol's frontend capability is still not good. It can only be said that compared to 5.5, there is improvement, but compared to other models' frontend capabilities, it's far behind.
User @jakevin7 tested GPT5.6-Sol to build Maka's official website and believes its frontend capability, though improved, is still far behind other models.
Which one should i buy? Claude, Cursor, or GPT?
A post asking for advice on whether to buy Claude, Cursor, or GPT, comparing these AI tools.