Antirez Deepseek 4.1 flash gguf on HF

Reddit r/LocalLLaMA Models

Summary

Antirez has uploaded the quantized gguf version of Deepseek 4.1 flash model to Hugging Face, with partial availability and questions on usage.

Q2 is there and Q4 is uploading as I type. Has his github been updated yet? How do you run this? https://huggingface.co/antirez/deepseek-v4.1-flash-gguf/tree/main
Original Article

Similar Articles

Deepseek V4 Flash 2, 3 and 4 bits GGUFs

Reddit r/LocalLLaMA

GGUF quantizations of DeepSeek V4 Flash in 2-bit, 3-bit, and 4-bit precisions, made available on Hugging Face for local inference with tools like llama.cpp and Ollama.

antirez/deepseek-v4-gguf

Hugging Face Models Trending

Antirez released GGUF quantizations of DeepSeek V4 Flash specifically tailored for the DS4 inference engine, providing optimized configurations for different RAM sizes and enabling local execution of the large MoE model.

Bartowski has delivered DS4 GGUF

Reddit r/LocalLLaMA

Bartowski has released a GGUF quantized version of DeepSeek-V4-Flash, inviting comparison with Antirez's version.