A lesson about retries, hidden in the DeepSeek-V4 paper

Reddit r/LocalLLaMA News

Summary

The article examines a warning in the DeepSeek-V4 paper that retrying interrupted LLM requests introduces length bias, and validates it by generating 100,000 poems, finding that retries can make responses shorter.

No content available
Original Article
View Cached Full Text

Cached at: 07/31/26, 04:55 PM

# A lesson about retries, hidden in the DeepSeek-V4 paper Source: [https://quesma.com/blog/hidden-lesson-deepseek-paper/](https://quesma.com/blog/hidden-lesson-deepseek-paper/) The DeepSeek\-V4 paper contains an unexpected lesson for anyone running LLM benchmarks: sometimes it’s not correct to retry failures\. So I decided to check it myself\. 100,000 AI poems later, here’s what I found\. ## DeepSeek’s warning Here’s the paragraph from the DeepSeek\-V4 paper that made me curious: > Importantly, it is mathematically incorrect to regenerate unfinished requests from scratch, as**this introduces length bias**\. Because shorter responses are more likely to survive interruption, regenerating from scratch makes the model more prone to producing shorter sequences whenever an interruption occurs\. From[DeepSeek\-V4 paper](https://arxiv.org/html/2606.19348v1#S5.SS2.SSS3.p3.1)Adding retries is an obvious way to fix issues with reliability\. But as DeepSeek noticed, longer requests have a higher chance of being interrupted\. When you retry a failed long request, there’s a chance you’ll get a short response as a replacement\. ## Retries make AI poems shorter ### Experiment Since statistical biases can often be hard to grasp, let’s see how this affects a real run\. I asked DeepSeek\-V4\-Flash to generate 100,000 poems, haikus, or other literary works for me: > Write a complete piece of literature in one randomly chosen form: a one\-line poem, a haiku, a novel chapter, or a limerick\. It cost me $6\.06 \(V4\-Flash is cheap\!\) and generated a large variety of responses: Response length histogramResponse lengths for 100,000 DeepSeek\-V4 Flash generations in 20\-word bins\. The final bin includes responses of 1,180 words or more\. The response count axis is logarithmic\. The average is 74 words\. Callouts identify the short\-form peak and the longer novel\-chapter peak\.0–19 words: 68,182 responses20–39 words: 18,770 responses40–59 words: 1,541 responses60–79 words: 19 responses80–99 words: 6 responses100–119 words: 4 responses120–139 words: 5 responses140–159 words: 20 responses160–179 words: 25 responses180–199 words: 39 responses200–219 words: 71 responses220–239 words: 104 responses240–259 words: 186 responses260–279 words: 237 responses280–299 words: 326 responses300–319 words: 392 responses320–339 words: 467 responses340–359 words: 507 responses360–379 words: 600 responses380–399 words: 622 responses400–419 words: 633 responses420–439 words: 590 responses440–459 words: 605 responses460–479 words: 525 responses480–499 words: 540 responses500–519 words: 468 responses520–539 words: 454 responses540–559 words: 441 responses560–579 words: 360 responses580–599 words: 357 responses600–619 words: 346 responses620–639 words: 266 responses640–659 words: 272 responses660–679 words: 217 responses680–699 words: 197 responses700–719 words: 212 responses720–739 words: 181 responses740–759 words: 175 responses760–779 words: 132 responses780–799 words: 130 responses800–819 words: 122 responses820–839 words: 109 responses840–859 words: 74 responses860–879 words: 74 responses880–899 words: 48 responses900–919 words: 51 responses920–939 words: 44 responses940–959 words: 37 responses960–979 words: 32 responses980–999 words: 25 responses1000–1019 words: 27 responses1020–1039 words: 20 responses1040–1059 words: 14 responses1060–1079 words: 15 responses1080–1099 words: 12 responses1100–1119 words: 6 responses1120–1139 words: 13 responses1140–1159 words: 7 responses1160–1179 words: 5 responses1,180\+ words: 41 responsesaverage: 74 wordsone\-line poems,haikuslong novelchapters0101001k10k100k03006009001,200\+wordscount \(log\) On average it generated a 74\-word text, but the results varied widely: - **73\.2% haikus**\. Very short: 15 words on average\. - **11\.5% novel chapters**\. 34 times longer: 503 words on average\. ### Simulating failures Now let’s simulate what would happen if the LLM inference were really unreliable and**10% of requests failed**by being interrupted randomly during generation\. I used a statistical model called[a Poisson process](https://en.wikipedia.org/wiki/Poisson_point_process)to model this behavior; an average time between interruptions of 31 seconds makes 10% of requests fail in my experiment\. As the DeepSeek paper noted, longer requests are affected more often\. You can use this formula for a Poisson process to calculate it: P\(failure during request\)=1−e−request time/mean time between interruptionsP\(\\text\{failure during request\}\) = 1 \- e^\{\-\\text\{request time\}/\\text\{mean time between interruptions\}\} For example, haikus take 2\.38 seconds on average to generate, so their failure rate is1−e−2\.38/31≈7\.4%1 \- e^\{\-2\.38/31\} \\approx 7\.4\\%\. But novel chapters are longer \(11\.52 seconds\), so their failure rate is higher:1−e−11\.52/31≈31\.0%1 \- e^\{\-11\.52/31\} \\approx 31\.0\\%\. You can see how the failure rate rises for longer requests: Response length histogramResponse lengths for 100,000 DeepSeek\-V4 Flash generations in 20\-word bins\. The final bin includes responses of 1,180 words or more\. The response count axis is logarithmic\. The red line shows the average Poisson failure probability in each response\-length bin on a linear percentage scale\. Its minimum and maximum are labeled\.0–19 words: 68,182 responses; 7\.0% failure rate20–39 words: 18,770 responses; 8\.7% failure rate40–59 words: 1,541 responses; 10\.3% failure rate60–79 words: 19 responses; 10\.7% failure rate80–99 words: 6 responses; 12\.5% failure rate100–119 words: 4 responses; 12\.7% failure rate120–139 words: 5 responses; 14\.3% failure rate140–159 words: 20 responses; 15\.0% failure rate160–179 words: 25 responses; 15\.6% failure rate180–199 words: 39 responses; 18\.2% failure rate200–219 words: 71 responses; 17\.7% failure rate220–239 words: 104 responses; 18\.9% failure rate240–259 words: 186 responses; 19\.5% failure rate260–279 words: 237 responses; 20\.8% failure rate280–299 words: 326 responses; 21\.7% failure rate300–319 words: 392 responses; 22\.9% failure rate320–339 words: 467 responses; 23\.5% failure rate340–359 words: 507 responses; 24\.5% failure rate360–379 words: 600 responses; 25\.4% failure rate380–399 words: 622 responses; 25\.9% failure rate400–419 words: 633 responses; 26\.8% failure rate420–439 words: 590 responses; 27\.7% failure rate440–459 words: 605 responses; 28\.5% failure rate460–479 words: 525 responses; 29\.4% failure rate480–499 words: 540 responses; 30\.1% failure rate500–519 words: 468 responses; 31\.0% failure rate520–539 words: 454 responses; 31\.5% failure rate540–559 words: 441 responses; 32\.3% failure rate560–579 words: 360 responses; 33\.3% failure rate580–599 words: 357 responses; 33\.7% failure rate600–619 words: 346 responses; 34\.6% failure rate620–639 words: 266 responses; 35\.2% failure rate640–659 words: 272 responses; 35\.6% failure rate660–679 words: 217 responses; 36\.7% failure rate680–699 words: 197 responses; 37\.2% failure rate700–719 words: 212 responses; 38\.0% failure rate720–739 words: 181 responses; 38\.8% failure rate740–759 words: 175 responses; 39\.7% failure rate760–779 words: 132 responses; 40\.6% failure rate780–799 words: 130 responses; 41\.1% failure rate800–819 words: 122 responses; 41\.9% failure rate820–839 words: 109 responses; 41\.5% failure rate840–859 words: 74 responses; 42\.8% failure rate860–879 words: 74 responses; 43\.6% failure rate880–899 words: 48 responses; 44\.1% failure rate900–919 words: 51 responses; 45\.1% failure rate920–939 words: 44 responses; 44\.7% failure rate940–959 words: 37 responses; 46\.1% failure rate960–979 words: 32 responses; 47\.5% failure rate980–999 words: 25 responses; 47\.5% failure rate1000–1019 words: 27 responses; 48\.1% failure rate1020–1039 words: 20 responses; 48\.2% failure rate1040–1059 words: 14 responses; 49\.2% failure rate1060–1079 words: 15 responses; 49\.2% failure rate1080–1099 words: 12 responses; 47\.8% failure rate1100–1119 words: 6 responses; 50\.8% failure rate1120–1139 words: 13 responses; 49\.6% failure rate1140–1159 words: 7 responses; 50\.2% failure rate1160–1179 words: 5 responses; 53\.4% failure rate1,180\+ words: 41 responses; 55\.4% failure rateFailure rate7\.0%55\.4%0%15%30%45%60%failure rate03006009001,200\+words ### Adding retries So let’s see what would happen if I added retries\. I artificially simulate failures and perform retries on the failed requests until they succeed\. After running the simulation, the resulting dataset looks noticeably different\! Just as the DeepSeek authors warned us: Without failuresFailures \+ retriesChangeAverage word count73\.5159\.40−19\.2%Responses ≥ 600 words2,9041,961−32\.5%Novel chapters11,4998,919−22\.4%The average response is now noticeably shorter and there are fewer longer\-form responses\. To better understand why it changed so much, let’s take a look at the requests that failed and what happened when they were retried: Failed requests and where their replacements land, in evenly spaced word bucketsThe 10,022 expected failed requests binned into evenly spaced 25\-word buckets, morphing into the lengths of their eventual replacements\. The count axis is logarithmic\. A red dashed outline preserves the failed\-request distribution while the animation loops\.0–24 words: 5,056 failed → 7,430 replacements land here25–49 words: 1,475 failed → 1,683 replacements land here50–74 words: 17 failed → 16 replacements land here75–99 words: 1 failed → 1 replacements land here100–124 words: 1 failed → 0 replacements land here125–149 words: 2 failed → 1 replacements land here150–174 words: 5 failed → 3 replacements land here175–199 words: 8 failed → 4 replacements land here200–224 words: 17 failed → 9 replacements land here225–249 words: 34 failed → 16 replacements land here250–274 words: 51 failed → 22 replacements land here275–299 words: 86 failed → 35 replacements land here300–324 words: 115 failed → 43 replacements land here325–349 words: 142 failed → 51 replacements land here350–374 words: 179 failed → 60 replacements land here375–399 words: 201 failed → 64 replacements land here400–424 words: 208 failed → 63 replacements land here425–449 words: 203 failed → 59 replacements land here450–474 words: 213 failed → 58 replacements land here475–499 words: 198 failed → 52 replacements land here500–524 words: 184 failed → 46 replacements land here525–549 words: 182 failed → 43 replacements land here550–574 words: 152 failed → 35 replacements land here575–599 words: 153 failed → 33 replacements land here600–624 words: 143 failed → 30 replacements land here625–649 words: 120 failed → 24 replacements land here650–674 words: 107 failed → 21 replacements land here675–699 words: 92 failed → 17 replacements land here700–724 words: 98 failed → 18 replacements land here725–749 words: 87 failed → 15 replacements land here750–774 words: 76 failed → 13 replacements land here775–799 words: 67 failed → 11 replacements land here800–824 words: 65 failed → 10 replacements land here825–849 words: 47 failed → 7 replacements land here850–874 words: 43 failed → 6 replacements land here875–899 words: 27 failed → 4 replacements land here900–924 words: 27 failed → 4 replacements land here925–949 words: 25 failed → 3 replacements land here950–974 words: 19 failed → 2 replacements land here975–999 words: 16 failed → 2 replacements land here1000–1024 words: 15 failed → 2 replacements land here1025–1049 words: 10 failed → 1 replacements land here1050–1074 words: 10 failed → 1 replacements land here1075–1099 words: 7 failed → 1 replacements land here1100–1124 words: 3 failed → 0 replacements land here1125–1149 words: 7 failed → 1 replacements land here1150–1174 words: 5 failed → 1 replacements land here1175–1199 words: 4 failed → 0 replacements land here1200\+ words: 19 failed → 2 replacements land hereFailed requestsAfter retries1101001k10k03006009001,200\+wordscount \(log\)After retries, \(would\-be\) long responses often become shorter responses\. The right\-hand side of the chart gets affected the most\. Looking at it with a statistics toolset, adding retries changed the final distribution of the data\. What we’re seeing here is a form of[selection bias](https://en.wikipedia.org/wiki/Selection_bias): the retried sample was not representative \- long responses were overrepresented in it\.[Survivorship bias](https://en.wikipedia.org/wiki/Survivorship_bias)is also a good perspective: the final dataset includes only responses that survived some filtering process \- in this case, interruptions that penalized long requests\. ## Conclusion Shorter poems might sound innocent, but the same problem could be dangerous in a real benchmark\. A longer request might be stuck in an endless reasoning loop or headed down the wrong path, and retrying it might give the model a second chance \- raising the score\. Armed with that knowledge, we have 3 ways to deal with the issue: - The DeepSeek authors architected their system to resume interrupted requests rather than regenerate them from scratch\. - Failures that happen truly randomly \(independently of request length\) can be safely retried\. - This problem can be important for benchmarks or research, but for most consumer\-facing applications the difference doesn’t really matter\. What started as a curious warning in the DeepSeek\-V4 paper became a surprisingly intuitive lesson in statistics for me\. Stay tuned for future posts and releases

Similar Articles

DeepSeek V4 paper full version is out, FP4 QAT details and stability tricks [D]

Reddit r/MachineLearning

DeepSeek released the full V4 paper detailing FP4 quantization-aware training, MoE training stability tricks (anticipatory routing and SwiGLU clamping), and a generative reward model for RLHF, achieving dramatic efficiency gains—V4-Flash uses only 10% of V3.2's FLOPs and 7% of its KV cache at 1M context length.

FlashMemory DeepSeek-V4 Retriever (GitHub Repo)

TLDR AI

Introduces FlashMemory DeepSeek-V4 Retriever, a lightweight model that sparsifies DeepSeek-V4's CSA KV-cache by predicting which chunks will be attended to next, keeping only ~10-15% on-device while matching full-attention performance.