Instead of decentralized training effort we should build the “One dataset”

Reddit r/LocalLLaMA News

Summary

An argument that instead of focusing on decentralized training of LLMs, the open-source community should collaborate to create a large, high-quality pre-training dataset available on people's computers, similar to BitTorrent.

There are many threads here calling for united LLM training run of a new open model. Mainly, after govt. stunt of banning commercial frontier models. And also due to the lack of small-medium open-weight models releases lately. I genuinelly believe at some point we’ll have “SETI for LLM”. But not anytime soon, not this year. It requires a serious primary research of a training algorhytms over high latency network(s). What I believe be much more valuable, is to prepare a pre-training data for such future training run. It is much less “super-hard-skill” task. There can be clients invented (vibe engineered) similar to bittorrent downloaders that do scraping, cleaning and hosting (sharing) of the data from the Internet. A new global database with trillions of high quality tokens, openly available, hosted on people’s computers would represent a true message of open-source community to billionates stealing our data and VRAM. Let’s not dream about distributed LLM training on our home GPUs. We should focus on something more practical. The mere existence of a such dataset would accelerate the development of distri-train on its own.
Original Article

Similar Articles

LocalLLaMA crowdsourced coding dataset

Reddit r/LocalLLaMA

A community member proposes creating a crowdsourced coding dataset for local LLMs to enable collaborative model training and fine-tuning, addressing concerns about future availability of open-weight models.

Could AI training be decentralized like Bitcoin mining? [D]

Reddit r/MachineLearning

A discussion explores whether AI training could be decentralized like Bitcoin mining, with participants contributing GPU resources to train open-source models in exchange for tokens, raising questions about verification, fake gradients, and efficiency.