@lxfater: 64GB RAM, the 100-billion-parameter large model has already been deployed We already have the hammer, so where's the na…

X AI KOLs Timeline 工具

摘要

12GB 显存 + 64GB 内存即可运行千亿参数大模型:开源项目 Strata 跑 Qwen3.8-Flash-Next 的量化版,在 RTX 5070 上 Q2_0 版约 94 token/s、IQ3_S 版约 53 token/s,作者感叹锤子已备好,就看钉子(应用场景)在哪。

64GB RAM, the 100-billion-parameter large model has already been deployed We already have the hammer, so where's the nail?
查看原文
查看缓存全文

缓存时间: 2026/10/05 01:20

64GB RAM, the 100-billion-parameter large model has already been deployed We already have the hammer, so where’s the nail?

铁锤人 (@lxfater): 12GB 显存 + 64GB 内存,也能跑千亿模型了!!

这个开源项目叫 Strata,跑的是 Qwen3.8-Flash-Next 的量化版

作者的测试配置是 RTX 5070(12GB 显存)+ Ryzen 5 7600 + 64GB 内存,短聊天生成速度: Q2_0 版约 94 token/s IQ3_S 版约 53 token/s

但 94 token/s 对应的是压缩更狠的

相似文章

@Xudong07452910: Hacker News 上有一篇评论区火了的文章:Qwen 3.6 27B 是本地开发的理想选择。 核心发现是:密集参数模型、原生支持 256k 上下文,在 MacBook Max M5 上跑 Q8_0 量化版能达到 30 tokens/…

X AI KOLs Timeline

Qwen 3.6 27B is a dense 27B model that achieves impressive performance on local hardware with 256k context, running at 30 tokens/s on MacBook Max M5 and 50 tokens/s on RTX 5090, and is considered by some as the first local model with true general intelligence.