unsloth/DeepSeek-V4-Flash-0731-GGUF
摘要
Unsloth 预告了即将在 Hugging Face 上发布的 DeepSeek V4 Flash GGUF 量化模型。
https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF 它来了!😍
相似文章
Unsloth MiniMax M3 GGUF
Unsloth 正在将 MiniMax M3 模型的 GGUF 量化版本上传到 Hugging Face。
DeepSeek-V4-Flash-0731 unsloth gguf 在 A100 上
展示了 DeepSeek-V4-Flash-0731 在单块 40GB A100 上以 unsloth GGUF 量化版本运行,速度为 17.7 tok/s,并将 6 个专家加载到显存中,从而实现了完整的智能体编码循环。
Deepseek V4 Flash 2位、3位和4位 GGUFs
DeepSeek V4 Flash 的 2位、3位和4位精度 GGUF 量化版本,已在 Hugging Face 上发布,可用于 llama.cpp 和 Ollama 等工具的本地推理。
unsloth/Qwen3.6-27B-GGUF
Unsloth 发布了 Qwen3.6-27B 模型的 GGUF 量化版本,具备更强的智能体编程能力、工具调用功能,并支持 Unsloth Studio。
jabbatheduck/DeepSeek-v4-flash-mini · Hugging Face
jabbatheduck 发布了 REAP 专家剪枝后的 DeepSeek-V4-Flash 检查点的 GGUF 量化版本,为消费级 GPU 上的内存受限推理进行了大幅压缩,同时保留了路由器和注意力的精度。