unsloth/DeepSeek-V4-Flash-0731-GGUF

Reddit r/LocalLLaMA 模型

摘要

Unsloth 预告了即将在 Hugging Face 上发布的 DeepSeek V4 Flash GGUF 量化模型。

https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF 它来了!😍
查看原文

相似文章

Unsloth MiniMax M3 GGUF

Reddit r/LocalLLaMA

Unsloth 正在将 MiniMax M3 模型的 GGUF 量化版本上传到 Hugging Face。

DeepSeek-V4-Flash-0731 unsloth gguf 在 A100 上

Reddit r/LocalLLaMA

展示了 DeepSeek-V4-Flash-0731 在单块 40GB A100 上以 unsloth GGUF 量化版本运行,速度为 17.7 tok/s,并将 6 个专家加载到显存中,从而实现了完整的智能体编码循环。

Deepseek V4 Flash 2位、3位和4位 GGUFs

Reddit r/LocalLLaMA

DeepSeek V4 Flash 的 2位、3位和4位精度 GGUF 量化版本,已在 Hugging Face 上发布,可用于 llama.cpp 和 Ollama 等工具的本地推理。

unsloth/Qwen3.6-27B-GGUF

Hugging Face Models Trending

Unsloth 发布了 Qwen3.6-27B 模型的 GGUF 量化版本,具备更强的智能体编程能力、工具调用功能,并支持 Unsloth Studio。

jabbatheduck/DeepSeek-v4-flash-mini · Hugging Face

Reddit r/LocalLLaMA

jabbatheduck 发布了 REAP 专家剪枝后的 DeepSeek-V4-Flash 检查点的 GGUF 量化版本,为消费级 GPU 上的内存受限推理进行了大幅压缩,同时保留了路由器和注意力的精度。