@YRSM_Simon: Amazing!

X AI KOLs Timeline News

Summary

shi3z shares optimization experience for running DeepSeek v4.1 Flash on an A100 GPU without FP4 support, boosting inference speed from 33 tok/s to 673 tok/s, surpassing the official API speed.

Amazing!
Original Article
View Cached Full Text

Cached at: 09/14/26, 01:16 AM

Impressive

shi3z (@shi3z): Wrote over 10k characters | The story of getting DeepSeek v4.1 Flash, which I wanted to run, to work on an A100 that doesn’t support FP4, boosting it from 33tok/s to 673tok/s, and before I knew it, it was faster than the official API | shi3z @shi3z

Similar Articles

Deepseek V4 Flash running on RTX 5090 MoE

Reddit r/LocalLLaMA

User shares optimization benchmarks for DeepSeek-V4-Flash (Q2_K) running on an RTX 5090 using a fork of llama.cpp, achieving 21.3 tokens/s generation and 1 million context size.