@charles_irl: did you know that MoEs like ChatGPT only use 10% of their brain

X AI KOLs Timeline News

Summary

A tweet notes that Mixture-of-Experts models like ChatGPT only use a fraction of their parameters at a time, humorously comparing it to the '10% of brain' myth.

did you know that MoEs like ChatGPT only use 10% of their brain
Original Article

Similar Articles

Are the rich RAM /poor GPU people wrong here?

Reddit r/LocalLLaMA

Discusses the trade-off between dense and Mixture-of-Experts (MoE) models for local AI, noting that high-RAM users have limited MoE options beyond Qwen 3.5 122B, and questioning if large GPU is the only viable path.

Mixture of Experts (MoEs) in Transformers

Hugging Face Blog

Hugging Face blog post explaining Mixture of Experts (MoEs) architecture in Transformers, covering the shift from dense to sparse models, weight loading optimizations, expert parallelism, and training techniques for MoE-based language models.