Need Info on quality benchmarks to run on DeepSeek V3.2 different quant levels [D]
Summary
Developer seeks quality benchmarks to evaluate runtime quantization impact on DeepSeek V3.2 model performance.
Similar Articles
We quantized DeepSeek V4 0731 and benchmarked it against popular quants on 8× RTX 5090
The author quantizes DeepSeek V4 0731, fixing FP8 downconversion issues that skew baselines, and benchmarks 38 quant files on 8× RTX 5090 to show GPU-dependent results and file-size-based comparisons.
Deepseek V4 Flash 2-bit quant is the first model I can run locally that achieves 100% in this SQL benchmark
A user reports that DeepSeek V4 Flash, running as a 2-bit quantized GGUF on dual RTX 3080s, is the first local model to score 100% on a real-world SQL benchmark, matching frontier models like Opus 4.7 and GPT-5.5.
Qwen 3.8 27b vs Deepseek Flash
The post compares the open-source AI models Qwen 3.8 (27B) and Deepseek Flash, discussing benchmarks and seeking user experiences to evaluate their performance.
DeepSeek-V3 Technical Report
DeepSeek-V3 is a parameter-efficient Mixture-of-Experts language model with 671B total parameters, achieving strong performance comparable to leading closed-source models while requiring only 2.788M H800 GPU hours for training.
Planning to spend ~$100 benchmarking differnet Qwen3.8-27B quants and kv cache and looking for input before I start
The author plans to spend $100 on cloud GPUs to benchmark Qwen3.8-27B models with different quantization levels and KV cache settings, focusing on coding and agentic tasks to provide useful data for local AI enthusiasts, and is seeking community input on the methodology.