api-speed

Tag

Cards List
#api-speed

@YRSM_Simon: Amazing!

X AI KOLs Timeline · yesterday Cached

shi3z shares optimization experience for running DeepSeek v4.1 Flash on an A100 GPU without FP4 support, boosting inference speed from 33 tok/s to 673 tok/s, surpassing the official API speed.

0 favorites 0 likes
← Back to home

Submit Feedback