Cached at:
07/31/26, 04:55 PM
# A lesson about retries, hidden in the DeepSeek-V4 paper
Source: [https://quesma.com/blog/hidden-lesson-deepseek-paper/](https://quesma.com/blog/hidden-lesson-deepseek-paper/)
The DeepSeek\-V4 paper contains an unexpected lesson for anyone running LLM benchmarks: sometimes it’s not correct to retry failures\.
So I decided to check it myself\. 100,000 AI poems later, here’s what I found\.
## DeepSeek’s warning
Here’s the paragraph from the DeepSeek\-V4 paper that made me curious:
> Importantly, it is mathematically incorrect to regenerate unfinished requests from scratch, as**this introduces length bias**\. Because shorter responses are more likely to survive interruption, regenerating from scratch makes the model more prone to producing shorter sequences whenever an interruption occurs\.
From[DeepSeek\-V4 paper](https://arxiv.org/html/2606.19348v1#S5.SS2.SSS3.p3.1)Adding retries is an obvious way to fix issues with reliability\. But as DeepSeek noticed, longer requests have a higher chance of being interrupted\. When you retry a failed long request, there’s a chance you’ll get a short response as a replacement\.
## Retries make AI poems shorter
### Experiment
Since statistical biases can often be hard to grasp, let’s see how this affects a real run\.
I asked DeepSeek\-V4\-Flash to generate 100,000 poems, haikus, or other literary works for me:
> Write a complete piece of literature in one randomly chosen form: a one\-line poem, a haiku, a novel chapter, or a limerick\.
It cost me $6\.06 \(V4\-Flash is cheap\!\) and generated a large variety of responses:
Response length histogramResponse lengths for 100,000 DeepSeek\-V4 Flash generations in 20\-word bins\. The final bin includes responses of 1,180 words or more\. The response count axis is logarithmic\. The average is 74 words\. Callouts identify the short\-form peak and the longer novel\-chapter peak\.0–19 words: 68,182 responses20–39 words: 18,770 responses40–59 words: 1,541 responses60–79 words: 19 responses80–99 words: 6 responses100–119 words: 4 responses120–139 words: 5 responses140–159 words: 20 responses160–179 words: 25 responses180–199 words: 39 responses200–219 words: 71 responses220–239 words: 104 responses240–259 words: 186 responses260–279 words: 237 responses280–299 words: 326 responses300–319 words: 392 responses320–339 words: 467 responses340–359 words: 507 responses360–379 words: 600 responses380–399 words: 622 responses400–419 words: 633 responses420–439 words: 590 responses440–459 words: 605 responses460–479 words: 525 responses480–499 words: 540 responses500–519 words: 468 responses520–539 words: 454 responses540–559 words: 441 responses560–579 words: 360 responses580–599 words: 357 responses600–619 words: 346 responses620–639 words: 266 responses640–659 words: 272 responses660–679 words: 217 responses680–699 words: 197 responses700–719 words: 212 responses720–739 words: 181 responses740–759 words: 175 responses760–779 words: 132 responses780–799 words: 130 responses800–819 words: 122 responses820–839 words: 109 responses840–859 words: 74 responses860–879 words: 74 responses880–899 words: 48 responses900–919 words: 51 responses920–939 words: 44 responses940–959 words: 37 responses960–979 words: 32 responses980–999 words: 25 responses1000–1019 words: 27 responses1020–1039 words: 20 responses1040–1059 words: 14 responses1060–1079 words: 15 responses1080–1099 words: 12 responses1100–1119 words: 6 responses1120–1139 words: 13 responses1140–1159 words: 7 responses1160–1179 words: 5 responses1,180\+ words: 41 responsesaverage: 74 wordsone\-line poems,haikuslong novelchapters0101001k10k100k03006009001,200\+wordscount \(log\)
On average it generated a 74\-word text, but the results varied widely:
- **73\.2% haikus**\. Very short: 15 words on average\.
- **11\.5% novel chapters**\. 34 times longer: 503 words on average\.
### Simulating failures
Now let’s simulate what would happen if the LLM inference were really unreliable and**10% of requests failed**by being interrupted randomly during generation\.
I used a statistical model called[a Poisson process](https://en.wikipedia.org/wiki/Poisson_point_process)to model this behavior; an average time between interruptions of 31 seconds makes 10% of requests fail in my experiment\.
As the DeepSeek paper noted, longer requests are affected more often\. You can use this formula for a Poisson process to calculate it:
P\(failure during request\)=1−e−request time/mean time between interruptionsP\(\\text\{failure during request\}\) = 1 \- e^\{\-\\text\{request time\}/\\text\{mean time between interruptions\}\}
For example, haikus take 2\.38 seconds on average to generate, so their failure rate is1−e−2\.38/31≈7\.4%1 \- e^\{\-2\.38/31\} \\approx 7\.4\\%\. But novel chapters are longer \(11\.52 seconds\), so their failure rate is higher:1−e−11\.52/31≈31\.0%1 \- e^\{\-11\.52/31\} \\approx 31\.0\\%\.
You can see how the failure rate rises for longer requests:
Response length histogramResponse lengths for 100,000 DeepSeek\-V4 Flash generations in 20\-word bins\. The final bin includes responses of 1,180 words or more\. The response count axis is logarithmic\. The red line shows the average Poisson failure probability in each response\-length bin on a linear percentage scale\. Its minimum and maximum are labeled\.0–19 words: 68,182 responses; 7\.0% failure rate20–39 words: 18,770 responses; 8\.7% failure rate40–59 words: 1,541 responses; 10\.3% failure rate60–79 words: 19 responses; 10\.7% failure rate80–99 words: 6 responses; 12\.5% failure rate100–119 words: 4 responses; 12\.7% failure rate120–139 words: 5 responses; 14\.3% failure rate140–159 words: 20 responses; 15\.0% failure rate160–179 words: 25 responses; 15\.6% failure rate180–199 words: 39 responses; 18\.2% failure rate200–219 words: 71 responses; 17\.7% failure rate220–239 words: 104 responses; 18\.9% failure rate240–259 words: 186 responses; 19\.5% failure rate260–279 words: 237 responses; 20\.8% failure rate280–299 words: 326 responses; 21\.7% failure rate300–319 words: 392 responses; 22\.9% failure rate320–339 words: 467 responses; 23\.5% failure rate340–359 words: 507 responses; 24\.5% failure rate360–379 words: 600 responses; 25\.4% failure rate380–399 words: 622 responses; 25\.9% failure rate400–419 words: 633 responses; 26\.8% failure rate420–439 words: 590 responses; 27\.7% failure rate440–459 words: 605 responses; 28\.5% failure rate460–479 words: 525 responses; 29\.4% failure rate480–499 words: 540 responses; 30\.1% failure rate500–519 words: 468 responses; 31\.0% failure rate520–539 words: 454 responses; 31\.5% failure rate540–559 words: 441 responses; 32\.3% failure rate560–579 words: 360 responses; 33\.3% failure rate580–599 words: 357 responses; 33\.7% failure rate600–619 words: 346 responses; 34\.6% failure rate620–639 words: 266 responses; 35\.2% failure rate640–659 words: 272 responses; 35\.6% failure rate660–679 words: 217 responses; 36\.7% failure rate680–699 words: 197 responses; 37\.2% failure rate700–719 words: 212 responses; 38\.0% failure rate720–739 words: 181 responses; 38\.8% failure rate740–759 words: 175 responses; 39\.7% failure rate760–779 words: 132 responses; 40\.6% failure rate780–799 words: 130 responses; 41\.1% failure rate800–819 words: 122 responses; 41\.9% failure rate820–839 words: 109 responses; 41\.5% failure rate840–859 words: 74 responses; 42\.8% failure rate860–879 words: 74 responses; 43\.6% failure rate880–899 words: 48 responses; 44\.1% failure rate900–919 words: 51 responses; 45\.1% failure rate920–939 words: 44 responses; 44\.7% failure rate940–959 words: 37 responses; 46\.1% failure rate960–979 words: 32 responses; 47\.5% failure rate980–999 words: 25 responses; 47\.5% failure rate1000–1019 words: 27 responses; 48\.1% failure rate1020–1039 words: 20 responses; 48\.2% failure rate1040–1059 words: 14 responses; 49\.2% failure rate1060–1079 words: 15 responses; 49\.2% failure rate1080–1099 words: 12 responses; 47\.8% failure rate1100–1119 words: 6 responses; 50\.8% failure rate1120–1139 words: 13 responses; 49\.6% failure rate1140–1159 words: 7 responses; 50\.2% failure rate1160–1179 words: 5 responses; 53\.4% failure rate1,180\+ words: 41 responses; 55\.4% failure rateFailure rate7\.0%55\.4%0%15%30%45%60%failure rate03006009001,200\+words
### Adding retries
So let’s see what would happen if I added retries\. I artificially simulate failures and perform retries on the failed requests until they succeed\.
After running the simulation, the resulting dataset looks noticeably different\! Just as the DeepSeek authors warned us:
Without failuresFailures \+ retriesChangeAverage word count73\.5159\.40−19\.2%Responses ≥ 600 words2,9041,961−32\.5%Novel chapters11,4998,919−22\.4%The average response is now noticeably shorter and there are fewer longer\-form responses\. To better understand why it changed so much, let’s take a look at the requests that failed and what happened when they were retried:
Failed requests and where their replacements land, in evenly spaced word bucketsThe 10,022 expected failed requests binned into evenly spaced 25\-word buckets, morphing into the lengths of their eventual replacements\. The count axis is logarithmic\. A red dashed outline preserves the failed\-request distribution while the animation loops\.0–24 words: 5,056 failed → 7,430 replacements land here25–49 words: 1,475 failed → 1,683 replacements land here50–74 words: 17 failed → 16 replacements land here75–99 words: 1 failed → 1 replacements land here100–124 words: 1 failed → 0 replacements land here125–149 words: 2 failed → 1 replacements land here150–174 words: 5 failed → 3 replacements land here175–199 words: 8 failed → 4 replacements land here200–224 words: 17 failed → 9 replacements land here225–249 words: 34 failed → 16 replacements land here250–274 words: 51 failed → 22 replacements land here275–299 words: 86 failed → 35 replacements land here300–324 words: 115 failed → 43 replacements land here325–349 words: 142 failed → 51 replacements land here350–374 words: 179 failed → 60 replacements land here375–399 words: 201 failed → 64 replacements land here400–424 words: 208 failed → 63 replacements land here425–449 words: 203 failed → 59 replacements land here450–474 words: 213 failed → 58 replacements land here475–499 words: 198 failed → 52 replacements land here500–524 words: 184 failed → 46 replacements land here525–549 words: 182 failed → 43 replacements land here550–574 words: 152 failed → 35 replacements land here575–599 words: 153 failed → 33 replacements land here600–624 words: 143 failed → 30 replacements land here625–649 words: 120 failed → 24 replacements land here650–674 words: 107 failed → 21 replacements land here675–699 words: 92 failed → 17 replacements land here700–724 words: 98 failed → 18 replacements land here725–749 words: 87 failed → 15 replacements land here750–774 words: 76 failed → 13 replacements land here775–799 words: 67 failed → 11 replacements land here800–824 words: 65 failed → 10 replacements land here825–849 words: 47 failed → 7 replacements land here850–874 words: 43 failed → 6 replacements land here875–899 words: 27 failed → 4 replacements land here900–924 words: 27 failed → 4 replacements land here925–949 words: 25 failed → 3 replacements land here950–974 words: 19 failed → 2 replacements land here975–999 words: 16 failed → 2 replacements land here1000–1024 words: 15 failed → 2 replacements land here1025–1049 words: 10 failed → 1 replacements land here1050–1074 words: 10 failed → 1 replacements land here1075–1099 words: 7 failed → 1 replacements land here1100–1124 words: 3 failed → 0 replacements land here1125–1149 words: 7 failed → 1 replacements land here1150–1174 words: 5 failed → 1 replacements land here1175–1199 words: 4 failed → 0 replacements land here1200\+ words: 19 failed → 2 replacements land hereFailed requestsAfter retries1101001k10k03006009001,200\+wordscount \(log\)After retries, \(would\-be\) long responses often become shorter responses\. The right\-hand side of the chart gets affected the most\.
Looking at it with a statistics toolset, adding retries changed the final distribution of the data\. What we’re seeing here is a form of[selection bias](https://en.wikipedia.org/wiki/Selection_bias): the retried sample was not representative \- long responses were overrepresented in it\.[Survivorship bias](https://en.wikipedia.org/wiki/Survivorship_bias)is also a good perspective: the final dataset includes only responses that survived some filtering process \- in this case, interruptions that penalized long requests\.
## Conclusion
Shorter poems might sound innocent, but the same problem could be dangerous in a real benchmark\. A longer request might be stuck in an endless reasoning loop or headed down the wrong path, and retrying it might give the model a second chance \- raising the score\.
Armed with that knowledge, we have 3 ways to deal with the issue:
- The DeepSeek authors architected their system to resume interrupted requests rather than regenerate them from scratch\.
- Failures that happen truly randomly \(independently of request length\) can be safely retried\.
- This problem can be important for benchmarks or research, but for most consumer\-facing applications the difference doesn’t really matter\.
What started as a curious warning in the DeepSeek\-V4 paper became a surprisingly intuitive lesson in statistics for me\.
Stay tuned for future posts and releases