PSA Update CUDA from 13.2 to 13.3 to solve DeepSeek V4 Flash 0731 Looping Problem!

Reddit r/LocalLLaMA News

Summary

A user reports that updating CUDA from 13.2 to 13.3 fixes a looping problem with DeepSeek V4 Flash 0731, making the model usable again for long coding tasks.

So one of yall mentioned that cuda 13.1 or 13.2 is broken for unsloth so I looked in to it, and they were right. I had 13.2 installed, after I switched to 13.3 no more looping!!! Before the cuda update, the model was literally unusable. A few minutes into the run it would start loop then the chat would start degrading. Anyway, if you had similar experience check your cuda version. Hopefully this will help some of you out. The model now runs great for long horizon coding tasks. It already found ways to improve my Qwen3.6 thinkingcap code, and I can visually see the improvement. DeepSeek v4 Flash 0731 is now my daily driver, until Qwen 3.8 27B comes out.
Original Article

Similar Articles

I have DeepSeek V4 Pro at home

Reddit r/LocalLLaMA

A user demonstrates successfully running the DeepSeek V4 Pro model on a local workstation using a modified llama.cpp CUDA repository, highlighting performance metrics and hardware requirements.