ssd-streaming

Tag

Cards List
#ssd-streaming

I run 35B–480B coding models on my 36 GB MacBook by streaming MoE experts from SSD — self-contained app, and I publish the benchmarks that *failed* too

Reddit r/LocalLLaMA · 2026-07-25

Slipstream streams MoE expert weights from SSD instead of RAM, enabling large coding models (35B–480B) on 36 GB MacBooks. Benchmarks show ~13–19 tok/s for 35B models and ~2.8 tok/s for 118B, with honest reporting of failed approaches.

0 favorites 0 likes
#ssd-streaming

@antirez: GLM 5.2, Q2_K routed experts (effectively ~2.6 bits) running with SSD streaming on an M5 Max 128GB computer.

X AI KOLs Following · 2026-06-28 Cached

GLM 5.2 model runs with Q2_K quantized routed experts (effective 2.6 bits) using SSD streaming on an M5 Max 128GB computer.

0 favorites 0 likes
#ssd-streaming

You can run Deepseek 4 flash on mac (M3 Max, 96gb)

Reddit r/LocalLLaMA · 2026-06-14

A guide on running DeepSeek 4 flash on a Mac M3 Max with 96GB RAM using Antirez's ds4 engine and SSD streaming, achieving ~12 tokens/second inference speed.

0 favorites 0 likes
#ssd-streaming

@antirez: DeepSeek v4 PRO running via SSD streaming on my 128GB MacBook m5 max. 1.6 trillion parameters.

X AI KOLs Timeline · 2026-06-04 Cached

DeepSeek v4 PRO, a 1.6 trillion parameter model, is running via SSD streaming on a 128GB MacBook m5 max, demonstrating local inference of a massive model.

0 favorites 0 likes
← Back to home

Submit Feedback