在 500MB 内存上运行 Gemma 4

Reddit r/LocalLLaMA 模型

摘要

讨论了如何在仅有 500MB 内存的设备上运行 Gemma 4,可能通过量化或其他优化技术实现。

https://x.com/i/status/2084656348617392261 https://www.reddit.com/r/LLMDevs/s/9oL5ogmE6s
查看原文

相似文章

Gemma 4 视觉

Reddit r/LocalLLaMA

Gemma 4 的视觉表现受默认 token 预算过低拖累;在 llama.cpp 中将 --image-max-tokens 提到 2240,可解锁顶尖 OCR 与细节识别,代价是额外占用约 14 GB 显存。