@superalesha: I sped up deepseek v4 flash by 29x on my 4x3090s !!! No, its not joke. 15 -> 443 t/s. a 23k prompt used to take 25 mins…

X AI KOLs Timeline News

Summary

A user achieved a 29x speedup for DeepSeek V4 Flash inference on 4x RTX 3090 GPUs by optimizing llama.cpp, reducing a 23k prompt from 25 minutes to 53 seconds.

I sped up deepseek v4 flash by 29x on my 4x3090s !!! No, its not joke. 15 -> 443 t/s. a 23k prompt used to take 25 mins. Now it takes 53 secs. 284b in 2bit, 87gb, barely squeezes into 96gb. Me and Fable 5 spent 4 days in llama.cpp. Fixed everything that was broken https://t.co/i4PEerevOo
Original Article
View Cached Full Text

Cached at: 07/10/26, 06:10 AM

I sped up deepseek v4 flash by 29x on my 4x3090s !!! No, its not joke. 15 -> 443 t/s. a 23k prompt used to take 25 mins. Now it takes 53 secs. 284b in 2bit, 87gb, barely squeezes into 96gb. Me and Fable 5 spent 4 days in llama.cpp. Fixed everything that was broken https://t.co/i4PEerevOo

Similar Articles

Deepseek V4 Flash running on RTX 5090 MoE

Reddit r/LocalLLaMA

User shares optimization benchmarks for DeepSeek-V4-Flash (Q2_K) running on an RTX 5090 using a fork of llama.cpp, achieving 21.3 tokens/s generation and 1 million context size.