Tag
Testing a 1.56TB Mixture-of-Experts model on a 6GB RTX 4050 laptop, requiring patched memory streaming with NVMe to achieve 0.106 tokens/s decode speed.
A hackathon called 'Build Small' with a maximum of 32B parameters, designed to run on a laptop, has attracted sponsors including OpenAI, NVIDIA, OpenBMB, and Cohere, offering over $40k cash, RTX 5080s, and codex credits.