@no_stp_on_snek: 哦,不错。这个会很有趣
摘要
对 Unsloth AI 的回复,表达了对传闻中的 DeepSeek-V4-Flash 模型的兴奋,以及通过量化版本在本地运行它的可能性。
查看缓存全文
缓存时间: 2026/07/31 19:01
哦,不错。这个肯定有意思
Unsloth AI (@UnslothAI): 如果 DeepSeek-V4-Flash 既这么好又这么小,想象一下 DeepSeek-V4-Pro 会怎样!🤯
再想象一下在你自己设备上本地运行 Flash!🔥
我们迫不及待要制作量化版,让每个人都能在本地运行!!!
相似文章
unsloth/DeepSeek-V4-Flash-0731-GGUF
Unsloth 预告了即将在 Hugging Face 上发布的 DeepSeek V4 Flash GGUF 量化模型。
@no_stp_on_snek: 非常棒。对3090团队来说意义重大。而且TurboQuant+已经在很多推理引擎中实现了。
一条回复对Unsloth AI即将推出的Qwen3.8-27B模型表示赞赏,该模型将能在17GB内存/显存配置下运行,并指出TurboQuant+已集成到众多推理引擎中——这对RTX 3090用户来说是个好消息。
@no_stp_on_snek: 单个 Spark 能做的事情依然非常了不起。感谢 @NVIDIAAI
一位用户分享说,他们使用 antirez 的 DwarfStar-4 配置,在 DGX Spark 上复现了 DeepSeek-V4-Flash-0731 的运行,证实了该设备在单机上的出色性能。
@BrianRoemmele: Another day and another full frontier model running on your computer. Been teething DeepSeek V4 Flash on over 60 employ…
Brian Roemmele reports that DeepSeek V4 Flash (304B, 1M context) now runs locally on Apple Silicon via the ds4 engine, sharing GGUF quantized builds with a fresh imatrix. The Hugging Face repo provides installation instructions and notes that these files are ds4-specific, not for llama.cpp.
我简直不敢相信,我居然在家用PC上运行了前沿模型DeepSeek-V4-Flash-0731。太疯狂了!
一位用户对在配备24GB显存的中端Windows PC上通过量化运行前沿模型DeepSeek-V4-Flash-0731表示震惊,凸显了本地AI的飞速进步。