expert-streaming

Tag

Cards List
#expert-streaming

Qwen3.8-Flash-Next (95.5 GiB) on a 64GB Mac at ~27 tok/s, checkpoint + fork

Reddit r/LocalLLaMA · 4d ago

The author optimized the Qwen3.8-Flash-Next model to run on a 64GB Mac using expert streaming and other techniques, achieving ~27 tok/s by publishing a checkpoint and a llama.cpp fork.

0 favorites 0 likes
#expert-streaming

Faster than Light in Air: 8-22 tg/s Qwen3.8-Flash-Next (Q4/Q4ish) on a 32GB M4 MacBook Air

Reddit r/LocalLLaMA · 2026-09-10

Cherenkov is a new inference engine for Apple Silicon that enables efficient memory-constrained inference of large AI models like Qwen3.8-Flash-Next using predictive expert streaming and mixed-precision execution.

0 favorites 0 likes
#expert-streaming

@ErickSky: Forget about vLLM, llama.cpp, and expensive GPUs. [colibri] This runs GLM-5.2 (744B MoE) on ~25 GB of RAM with pure C a…

X AI KOLs Timeline · 2026-07-10 Cached

colibri is a pure C inference tool that runs the GLM-5.2 744B MoE model on ~25 GB RAM by streaming experts from disk, eliminating the need for expensive GPUs.

0 favorites 0 likes
← Back to home

Submit Feedback