Colibri streaming for Hy3 (Run Hy3 on 10GB (V)RAM)

Reddit r/artificial Tools

Summary

A port of Colibri streaming to work with Hy3, enabling the model to run on as little as 10GB of RAM/VRAM instead of the original 25GB.

Standing on the shoulders of giants, I vibe-coded a port of Colibri to work with Hy3 so you can run it on even smaller hardware specs (Colibri originally works with GLM 5.2 on 25GB, now you need no more than 10GB (even less actually)). Have a look and enjoy https://github.com/ErikTromp/colibri-hy3 PS. Use RAM instead of VRAM unless you have a lot of it. More means faster here.
Original Article

Similar Articles

@danveloper: Now everyone does it

X AI KOLs Timeline

Zane Chen demonstrates Colibri, which runs GLM-5.2 (744B MoE) on a laptop with 25GB RAM using pure C and CPU-only inference, by streaming experts from disk.

Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri

Hacker News Top

Lumabri lets users run huge mixture-of-experts models on a P2P swarm using the Colibri engine, allowing any machine to join and chat without downloading the full model up front. It is pure C, dependency-free, and works on CPU first with optional GPU acceleration.

Hy3 1Bit 89-93 GB

Reddit r/LocalLLaMA

Announcement of Hy3 1-bit quantized model with 89-93 GB memory footprint.