When Does the Best Sampling Temperature Rise with the Budget? Sufficient Conditions for Pass@k
Summary
This paper provides a theoretical explanation for why the optimal sampling temperature for pass@k increases with the budget, deriving sufficient conditions and analyzing the empirical pattern without model training or queries.
View Cached Full Text
Cached at: 08/18/26, 10:23 AM
# When Does the Best Sampling Temperature Rise with the Budget? Sufficient Conditions for pass{@𝑘}
Source: [https://arxiv.org/html/2608.14665](https://arxiv.org/html/2608.14665)
\(July 2026\)
###### Abstract
The temperature that maximizes pass@kkis often low for a small sampling budget and higher for a large budget\. This pattern has been reported from Codex through recent multi\-sample inference studies\. It is not an algebraic property of pass@kk: asSlocum et al\. \[[12](https://arxiv.org/html/2608.14665#bib.bib12)\]also observe, for one fixed task the maximizing temperature is independent ofkk\.
Building on that fixed\-task observation and the prior qualitative hard/easy\-task explanation, we give a formal population\-level sufficient condition for the aggregate pattern\. For taskXX, letpt\(X\)p\_\{t\}\(X\)be one\-sample success probability at temperaturett, and define the conditional log\-success responsemt\(u\)=𝔼\[p˙t\(X\)∣pt\(X\)=u\]/um\_\{t\}\(u\)=\\mathbb\{E\}\[\\dot\{p\}\_\{t\}\(X\)\\mid p\_\{t\}\(X\)=u\]/u\. Ifmt\(u\)m\_\{t\}\(u\)is nonincreasing in current success probability, then the normalized temperature derivative of aggregate pass@kkis nondecreasing inkk\. Consequently, derivative signs are nested across budgets; if each temperature\-performance curve is strictly single\-peaked, its unique maximizer is nondecreasing inkk\. The proof identifies the mechanism as a monotone\-likelihood\-ratio power tilt toward lower\-success tasks\.
We derive a closed\-form two\-stratum phase diagram, including upward and downward regimes, and show that the marginal temperature derivative admits an exactBeta\(2,k\)\\mathrm\{Beta\}\(2,k\)kernel representation whose kernel concentrates at one\-sample success of order1/k1/k\. Interpreting that scale as task\-level localization additionally requires a regular, nonvanishing density–response factor near zero\. A signed\-moment representation yields diagnostic shape restrictions, while a short appendix records exact discrete refinements of the existing multi\-configuration allocation formulation\. No language model is trained, and no model query is used as an experimental measurement: the contribution is a conditional theory of an established empirical phenomenon, with assumptions that can be tested in future work\.
## 1Introduction
Sampling temperature mediates a familiar quality–exploration trade\-off in language\-model inference\. For pass@kk, only one accepted completion amongkkattempts is needed\. Empirically, the temperature that performs best often increases with the number of attempts\. The Codex study explicitly reported an optimum near0\.20\.2for pass@1 and0\.80\.8for pass@100 in one model, and plotted a rising upper hull across budgets\[[2](https://arxiv.org/html/2608.14665#bib.bib2)\]\. Later studies reproduced the qualitative small\-budget low\-temperature and large\-budget high\-temperature pattern across models and tasks\[[5](https://arxiv.org/html/2608.14665#bib.bib5),[3](https://arxiv.org/html/2608.14665#bib.bib3)\]\. Recent work also finds that different temperatures solve different subsets of problems and that a temperature portfolio can outperform a single temperature\[[15](https://arxiv.org/html/2608.14665#bib.bib15)\]\.
The standard explanation is that a larger budget rewards diversity\. That intuition is useful but incomplete under the conditional\-independent model used to define population pass@kk\. For a fixed task, pass@kkis a strictly increasing transformation of its one\-sample success probability\. Therefore, the temperature maximizing a fixed task cannot move withkk\.Slocum et al\. \[[12](https://arxiv.org/html/2608.14665#bib.bib12)\]state this observation explicitly and give the adjacent intuition that low\-success tasks receive larger marginal gains at high budgets\. Movement of the benchmark\-level optimum requires heterogeneous temperature responses across tasks, or a violation of conditional i\.i\.d\. sampling, fixed\-within\-schedule decoding, or perfect verification\.
This paper asks a narrow question:
> Under what task\-level response condition must the aggregate optimal temperature move weakly upward as the pass@kkbudget grows?
The answer uses a conditional log\-success response\. At temperaturett, tasks are ordered by current one\-sample successptp\_\{t\}\. If increasing temperature has a weakly larger proportional benefit on lower\-success tasks, then the pass@kkmarginal shifts monotonically in favor of higher temperature askkgrows\. The shift is not in the raw derivative magnitude; it is in a normalized derivative whose weighting law is ordered by monotone likelihood ratio\. This distinction prevents a common but false claim that∂tpass@k\\partial\_\{t\}\\operatorname\{pass\}@kitself must grow withkk\.
Against that prior background, our contributions are:
1. 1\.a sufficient\-condition theorem giving nested temperature\-derivative signs for every pair of budgets and, under strict single\-peakedness, nondecreasing optimal temperatures;
2. 2\.an exact two\-stratum affine model with a phase boundary, closed\-form optimizer, upward and downward regimes, and an explicit large\-budget limit;
3. 3\.aBeta\(2,k\)\\mathrm\{Beta\}\(2,k\)kernel representation and its conditional localization interpretation, a logit\-level interpretation of the response\-order condition, and signed\-moment diagnostics for future empirical tests\.
The mathematical tools—likelihood\-ratio order, covariance inequalities, monotone comparative statics, and the compact moment problem—are classical\[[9](https://arxiv.org/html/2608.14665#bib.bib9),[8](https://arxiv.org/html/2608.14665#bib.bib8),[10](https://arxiv.org/html/2608.14665#bib.bib10),[7](https://arxiv.org/html/2608.14665#bib.bib7)\]\. The contribution is their specialization to inference\-time temperature under pass@kk, not a new theorem about those tools\. We also do not claim to discover hard\-task reweighting: the factork\(1−p\)k−1k\(1\-p\)^\{k\-1\}and its implications for pass@kkoptimization already appear in recent work\[[1](https://arxiv.org/html/2608.14665#bib.bib1)\]\. Likewise, optimized mixtures of temperatures and other configurations are the subject of OSCA\[[16](https://arxiv.org/html/2608.14665#bib.bib16)\]\. Our allocation appendix is explicitly a discrete refinement of that framework\.
## 2Population pass@k
LetX∼μX\\sim\\mudenote a benchmark task, letI=\[t¯,t¯\]I=\[\\underline\{t\},\\overline\{t\}\]be a compact temperature interval, and let
pt\(X\)=Pr\{one completion is accepted∣X,t\}\.p\_\{t\}\(X\)=\\Pr\\\{\\text\{one completion is accepted\}\\mid X,t\\\}\.Conditional on\(X,t\)\(X,t\), thekkcompletions are independent and identically distributed, and the verifier is perfect for the binary success event\. Writeqt=1−ptq\_\{t\}=1\-p\_\{t\}\. Population pass@kkis
Ak\(t\)=𝔼\[1−qt\(X\)k\]\.A\_\{k\}\(t\)=\\mathbb\{E\}\\\!\\left\[1\-q\_\{t\}\(X\)^\{k\}\\right\]\.\(1\)This is the population target estimated by the usual unbiased finite\-sample pass@kkestimator\[[2](https://arxiv.org/html/2608.14665#bib.bib2)\]\.
For the differential results, fix an interior operating pointtt\. Assume thats↦ps\(X\)s\\mapsto p\_\{s\}\(X\)is differentiable almost surely in a neighborhood oftt, that𝔼\|p˙t\(X\)\|<∞\\mathbb\{E\}\|\\dot\{p\}\_\{t\}\(X\)\|<\\infty, that differentiation may pass through the expectation, and that0<pt\(X\)<10<p\_\{t\}\(X\)<1almost surely\. We state the differential theorems only at such points\. Boundary cases require separate one\-sided derivatives and dominated\-limit conditions\. We use a dot for∂t\\partial\_\{t\}\.
### 2\.1Background: fixed\-task budget invariance
The following elementary fact was stated explicitly bySlocum et al\. \[[12](https://arxiv.org/html/2608.14665#bib.bib12), Appendix C\]\. We include it as background because it separates a per\-task identity from the population comparative static developed below\.
###### Proposition 2\.1\(Background budget invariance for a fixed task\)\.
For every fixed taskxxand every integerk≥1k\\geq 1,
argmaxt∈I\{1−\(1−pt\(x\)\)k\}=argmaxt∈Ipt\(x\)\.\\operatorname\*\{arg\\,max\}\_\{t\\in I\}\\\{1\-\(1\-p\_\{t\}\(x\)\)^\{k\}\\\}=\\operatorname\*\{arg\\,max\}\_\{t\\in I\}p\_\{t\}\(x\)\.More generally, if two temperaturess,ts,tsatisfyps\(X\)≥pt\(X\)p\_\{s\}\(X\)\\geq p\_\{t\}\(X\)almost surely, thenAk\(s\)≥Ak\(t\)A\_\{k\}\(s\)\\geq A\_\{k\}\(t\)for everykk\.
###### Proof\.
Fork≥1k\\geq 1, the mapu↦1−\(1−u\)ku\\mapsto 1\-\(1\-u\)^\{k\}is strictly increasing on\[0,1\]\[0,1\]\. Apply it pointwise, and then take expectations for the second claim\. ∎
Thus a changing aggregate optimum cannot be obtained by saying only that largerkkrewards repeated draws\. Under the model, a fixed task’s output diversity matters for pass@kkonly through the total probability mass of accepted outputs,pt\(x\)p\_\{t\}\(x\)\. Budget dependence requires temperature rankings to cross across tasks\.
## 3A monotone\-temperature theorem
Differentiating \([1](https://arxiv.org/html/2608.14665#S2.E1)\) gives the known pass@kkweighting identity
Ak′\(t\)=k𝔼\[qtk−1p˙t\]\.A\_\{k\}^\{\\prime\}\(t\)=k\\,\\mathbb\{E\}\\\!\\left\[q\_\{t\}^\{k\-1\}\\dot\{p\}\_\{t\}\\right\]\.\(2\)Define the pointwise log\-success temperature response \(a semi\-elasticity\)
ηt\(X\)=∂tlogpt\(X\)=p˙t\(X\)pt\(X\)\\eta\_\{t\}\(X\)=\\partial\_\{t\}\\log p\_\{t\}\(X\)=\\frac\{\\dot\{p\}\_\{t\}\(X\)\}\{p\_\{t\}\(X\)\}and letP=pt\(X\)P=p\_\{t\}\(X\)\. Foru\>0u\>0, define the conditional log\-success response
mt\(u\)=𝔼\[p˙t\(X\)∣P=u\]u\.m\_\{t\}\(u\)=\\frac\{\\mathbb\{E\}\[\\dot\{p\}\_\{t\}\(X\)\\mid P=u\]\}\{u\}\.\(3\)This equals𝔼\[ηt\(X\)∣P=u\]\\mathbb\{E\}\[\\eta\_\{t\}\(X\)\\mid P=u\]wheneverηt\\eta\_\{t\}is integrable, while the displayed definition is well defined under the stated integrability ofp˙t\\dot\{p\}\_\{t\}\. Set
Zk,t=𝔼\[P\(1−P\)k−1\]Z\_\{k,t\}=\\mathbb\{E\}\[P\(1\-P\)^\{k\-1\}\]and, whenZk,t\>0Z\_\{k,t\}\>0, define a probability lawνk,t\\nu\_\{k,t\}on tasks by
dνk,tdμ\(X\)=pt\(X\)qt\(X\)k−1Zk,t\.\\frac\{\\,\\mathrm\{d\}\\nu\_\{k,t\}\}\{\\,\\mathrm\{d\}\\mu\}\(X\)=\\frac\{p\_\{t\}\(X\)q\_\{t\}\(X\)^\{k\-1\}\}\{Z\_\{k,t\}\}\.\(4\)Equations \([2](https://arxiv.org/html/2608.14665#S3.E2)\)–\([4](https://arxiv.org/html/2608.14665#S3.E4)\) imply
Ak′\(t\)=kZk,t𝔼νk,t\[mt\(P\)\]\.A\_\{k\}^\{\\prime\}\(t\)=kZ\_\{k,t\}\\,\\mathbb\{E\}\_\{\\nu\_\{k,t\}\}\[m\_\{t\}\(P\)\]\.\(5\)
###### Assumption 3\.1\(Decreasing conditional log\-success response\)\.
At every interiorttto which the standing differential assumptions apply, a version ofmt\(u\)m\_\{t\}\(u\)is nonincreasing inuuon the support ofpt\(X\)p\_\{t\}\(X\)\.
This is a hard\-task\-benefit condition: at the current operating point, raising temperature has a larger proportional response, or a smaller proportional cost, on lower\-success tasks\. It is a restriction on observable task\-level responses, not a consequence of pass@kk\. The conventional elasticity with respect to temperature istmt\(u\)t\\,m\_\{t\}\(u\); at any fixed positivett, this factor does not change the cross\-task order assumed here\.
###### Lemma 3\.2\(Budget induces a likelihood\-ratio tilt\)\.
For integersℓ\>k\\ell\>k,
dνℓ,tdνk,t\(P\)=\(1−P\)ℓ−k𝔼νk,t\[\(1−P\)ℓ−k\]\.\\frac\{\\,\\mathrm\{d\}\\nu\_\{\\ell,t\}\}\{\\,\\mathrm\{d\}\\nu\_\{k,t\}\}\(P\)=\\frac\{\(1\-P\)^\{\\ell\-k\}\}\{\\mathbb\{E\}\_\{\\nu\_\{k,t\}\}\[\(1\-P\)^\{\\ell\-k\}\]\}\.The likelihood ratio is nonincreasing inPP; henceνℓ,t\\nu\_\{\\ell,t\}is shifted toward lower one\-sample success in monotone likelihood\-ratio order\.
###### Proof\.
Divide the two densities in \([4](https://arxiv.org/html/2608.14665#S3.E4)\)\. The remaining factor is\(1−P\)ℓ−k\(1\-P\)^\{\\ell\-k\}, and the denominator normalizes it\. ∎
###### Theorem 3\.3\(Nested temperature\-derivative signs\)\.
At any interiorttsatisfying the standing differential assumptions and Assumption[3\.1](https://arxiv.org/html/2608.14665#S3.Thmtheorem1), for integersℓ\>k≥1\\ell\>k\\geq 1,
𝔼νℓ,t\[mt\(P\)\]≥𝔼νk,t\[mt\(P\)\]\.\\mathbb\{E\}\_\{\\nu\_\{\\ell,t\}\}\[m\_\{t\}\(P\)\]\\geq\\mathbb\{E\}\_\{\\nu\_\{k,t\}\}\[m\_\{t\}\(P\)\]\.\(6\)Consequently,
Ak′\(t\)≥0⟹Aℓ′\(t\)≥0\.A\_\{k\}^\{\\prime\}\(t\)\\geq 0\\quad\\Longrightarrow\\quad A\_\{\\ell\}^\{\\prime\}\(t\)\\geq 0\.\(7\)For adjacent budgets, the exact increment of the normalized score is
𝔼νk\+1,tmt−𝔼νk,tmt=Covνk,t\(mt\(P\),1−P\)𝔼νk,t\[1−P\]≥0\.\\mathbb\{E\}\_\{\\nu\_\{k\+1,t\}\}m\_\{t\}\-\\mathbb\{E\}\_\{\\nu\_\{k,t\}\}m\_\{t\}=\\frac\{\\operatorname\{Cov\}\_\{\\nu\_\{k,t\}\}\(m\_\{t\}\(P\),1\-P\)\}\{\\mathbb\{E\}\_\{\\nu\_\{k,t\}\}\[1\-P\]\}\\geq 0\.\(8\)
###### Proof\.
By Lemma[3\.2](https://arxiv.org/html/2608.14665#S3.Thmtheorem2),
𝔼νℓ,tmt=𝔼νk,t\[mt\(P\)\(1−P\)ℓ−k\]𝔼νk,t\[\(1−P\)ℓ−k\]\.\\mathbb\{E\}\_\{\\nu\_\{\\ell,t\}\}m\_\{t\}=\\frac\{\\mathbb\{E\}\_\{\\nu\_\{k,t\}\}\[m\_\{t\}\(P\)\(1\-P\)^\{\\ell\-k\}\]\}\{\\mathbb\{E\}\_\{\\nu\_\{k,t\}\}\[\(1\-P\)^\{\\ell\-k\}\]\}\.Bothmt\(P\)m\_\{t\}\(P\)and\(1−P\)ℓ−k\(1\-P\)^\{\\ell\-k\}are nonincreasing functions ofPP\. Their covariance is nonnegative by the elementary association inequality for comonotone functions, which proves \([6](https://arxiv.org/html/2608.14665#S3.E6)\)\. Equation \([5](https://arxiv.org/html/2608.14665#S3.E5)\) then gives \([7](https://arxiv.org/html/2608.14665#S3.E7)\)\. Takingℓ=k\+1\\ell=k\+1yields \([8](https://arxiv.org/html/2608.14665#S3.E8)\)\. ∎
###### Definition 3\.5\(Strict derivative single\-peakedness\)\.
A differentiable functionffonIIis strictly single\-peaked if it has a unique maximizert⋆t^\{\\star\},f′\(t\)\>0f^\{\\prime\}\(t\)\>0fort<t⋆t<t^\{\\star\}, andf′\(t\)<0f^\{\\prime\}\(t\)<0fort\>t⋆t\>t^\{\\star\}, using one\-sided derivatives at the boundary\.
###### Corollary 3\.6\(Budget\-monotone optimal temperature\)\.
Suppose the standing differential assumptions and Assumption[3\.1](https://arxiv.org/html/2608.14665#S3.Thmtheorem1)hold at every interiort∈It\\in I, and everyAkA\_\{k\}is strictly single\-peaked with unique maximizertkt\_\{k\}\. Then
tℓ≥tkwheneverℓ\>k\.t\_\{\\ell\}\\geq t\_\{k\}\\qquad\\text\{whenever \}\\ell\>k\.
###### Proof\.
Iftkt\_\{k\}is interior, thenAk′\(tk\)=0A\_\{k\}^\{\\prime\}\(t\_\{k\}\)=0, so Theorem[3\.3](https://arxiv.org/html/2608.14665#S3.Thmtheorem3)givesAℓ′\(tk\)≥0A\_\{\\ell\}^\{\\prime\}\(t\_\{k\}\)\\geq 0\. Strict single\-peakedness ofAℓA\_\{\\ell\}rules outtℓ<tkt\_\{\\ell\}<t\_\{k\}\. Iftk=t¯t\_\{k\}=\\underline\{t\}, the conclusion is immediate\. Iftk=t¯t\_\{k\}=\\overline\{t\}, thenAk′\(t\)\>0A\_\{k\}^\{\\prime\}\(t\)\>0throughout the interior; sign nesting givesAℓ′\(t\)≥0A\_\{\\ell\}^\{\\prime\}\(t\)\\geq 0there, forcingtℓ=t¯t\_\{\\ell\}=\\overline\{t\}\. ∎
Neither Assumption[3\.1](https://arxiv.org/html/2608.14665#S3.Thmtheorem1)nor single\-peakedness is automatic\. Without them, the optimal temperature may decrease, oscillate, or be nonunique\. The result is therefore a falsifiable sufficient condition, not a universal law\.
## 4A Beta\-kernel representation of the marginal decision
SupposeP=pt\(X\)P=p\_\{t\}\(X\)has a densityftf\_\{t\}on\(0,1\)\(0,1\)\. Conditioning \([2](https://arxiv.org/html/2608.14665#S3.E2)\) onPPgives
Ak′\(t\)=k∫01u\(1−u\)k−1mt\(u\)ft\(u\)du\.A\_\{k\}^\{\\prime\}\(t\)=k\\int\_\{0\}^\{1\}u\(1\-u\)^\{k\-1\}m\_\{t\}\(u\)f\_\{t\}\(u\)\\,\\mathrm\{d\}u\.The density ofZk∼Beta\(2,k\)Z\_\{k\}\\sim\\mathrm\{Beta\}\(2,k\)isk\(k\+1\)u\(1−u\)k−1k\(k\+1\)u\(1\-u\)^\{k\-1\}, so
Ak′\(t\)=1k\+1𝔼\[mt\(Zk\)ft\(Zk\)\]\.A\_\{k\}^\{\\prime\}\(t\)=\\frac\{1\}\{k\+1\}\\mathbb\{E\}\\\!\\left\[m\_\{t\}\(Z\_\{k\}\)f\_\{t\}\(Z\_\{k\}\)\\right\]\.\(9\)The kernel itself has mode1/k1/kfork\>1k\>1, mean2/\(k\+2\)2/\(k\+2\), and concentrates on one\-sample success probabilities of order1/k1/k\. Equation \([9](https://arxiv.org/html/2608.14665#S4.E9)\) averages the density–response factorht\(u\)=mt\(u\)ft\(u\)h\_\{t\}\(u\)=m\_\{t\}\(u\)f\_\{t\}\(u\)against that kernel\. Kernel concentration alone does not imply that the nonzero task\-level contribution is centered at the same scale\. Such an interpretation is justified, for example, whenhth\_\{t\}is continuous with a finite nonzero limit near zero\. Ifhth\_\{t\}vanishes or varies rapidly there, or if the success distribution is supported away from zero, the nonzero contribution can come from a different scale\. For a finite benchmark, the density representation itself is only an approximation or analogy\.
The same tilt underlies aggregate inference\-scaling laws\. A mixture of taskwise exponential failure curves can exhibit much slower aggregate decay when the success distribution has substantial mass near zero\[[11](https://arxiv.org/html/2608.14665#bib.bib11)\]\. Equation \([9](https://arxiv.org/html/2608.14665#S4.E9)\) asks a different, local question: which part of that distribution controls the response to a change in temperature?
## 5An exact two\-stratum phase diagram
Consider an easy stratum of massπ∈\(0,1\)\\pi\\in\(0,1\)and a hard stratum of mass1−π1\-\\pi, with affine temperature responses
pE\(t\)=e−at,pH\(t\)=h\+bt,p\_\{E\}\(t\)=e\-at,\\qquad p\_\{H\}\(t\)=h\+bt,\(10\)wherea,b\>0a,b\>0and both probabilities remain in\(0,1\)\(0,1\)onII\. PutA=1−eA=1\-e,B=1−hB=1\-h, and
C=πa\(1−π\)b\.C=\\frac\{\\pi a\}\{\(1\-\\pi\)b\}\.\(11\)The aggregate failure probability is
Fk\(t\)=1−Ak\(t\)=π\(A\+at\)k\+\(1−π\)\(B−bt\)k\.F\_\{k\}\(t\)=1\-A\_\{k\}\(t\)=\\pi\(A\+at\)^\{k\}\+\(1\-\\pi\)\(B\-bt\)^\{k\}\.
###### Proposition 5\.1\(Closed\-form optimizer and phase boundary\)\.
Fork=1k=1, a maximizer ist¯\\underline\{t\}whenC\>1C\>1,t¯\\overline\{t\}whenC<1C<1, and everyt∈It\\in IwhenC=1C=1\. Fork\>1k\>1,AkA\_\{k\}is strictly concave and has the unique maximizer
tk=clipI\(B−rkAb\+rka\),rk=C1/\(k−1\)\.t\_\{k\}=\\operatorname\{clip\}\_\{I\}\\\!\\left\(\\frac\{B\-r\_\{k\}A\}\{b\+r\_\{k\}a\}\\right\),\\qquad r\_\{k\}=C^\{1/\(k\-1\)\}\.\(12\)IfC\>1C\>1, the sequence\(tk\)\(t\_\{k\}\)is nondecreasing; ifC<1C<1, it is nonincreasing\. IfC=1C=1, it is constant fork\>1k\>1\. In every case,
limk→∞tk=clipI\(B−Aa\+b\),\\lim\_\{k\\to\\infty\}t\_\{k\}=\\operatorname\{clip\}\_\{I\}\\\!\\left\(\\frac\{B\-A\}\{a\+b\}\\right\),\(13\)the temperature at which the two strata have equal one\-sample success, clipped toII\.
###### Proof\.
Fork=1k=1,
A1′\(t\)=\(1−π\)b−πa,A\_\{1\}^\{\\prime\}\(t\)=\(1\-\\pi\)b\-\\pi a,which gives the boundary cases through \([11](https://arxiv.org/html/2608.14665#S5.E11)\)\. Fork\>1k\>1,
Ak′′\(t\)=−k\(k−1\)\[πa2\(A\+at\)k−2\+\(1−π\)b2\(B−bt\)k−2\]<0\.A\_\{k\}^\{\\prime\\prime\}\(t\)=\-k\(k\-1\)\\\!\\left\[\\pi a^\{2\}\(A\+at\)^\{k\-2\}\+\(1\-\\pi\)b^\{2\}\(B\-bt\)^\{k\-2\}\\right\]<0\.The unique unconstrained critical point solves
B−btA\+at=\(πa\(1−π\)b\)1/\(k−1\)=rk,\\frac\{B\-bt\}\{A\+at\}=\\left\(\\frac\{\\pi a\}\{\(1\-\\pi\)b\}\\right\)^\{1/\(k\-1\)\}=r\_\{k\},which yields \([12](https://arxiv.org/html/2608.14665#S5.E12)\); concavity justifies clipping\. WritingT\(r\)=\(B−rA\)/\(b\+ra\)T\(r\)=\(B\-rA\)/\(b\+ra\), we have
T′\(r\)=−Ab\+aB\(b\+ra\)2<0\.T^\{\\prime\}\(r\)=\-\\frac\{Ab\+aB\}\{\(b\+ra\)^\{2\}\}<0\.WhenC\>1C\>1,rk↓1r\_\{k\}\\downarrow 1, soT\(rk\)T\(r\_\{k\}\)increases; whenC<1C<1,rk↑1r\_\{k\}\\uparrow 1, so it decreases\. Clipping preserves either order\. Takingrk→1r\_\{k\}\\to 1proves \([13](https://arxiv.org/html/2608.14665#S5.E13)\)\. ∎
The ratioCCis a one\-shot phase boundary\. WhenC\>1C\>1, the population\-weighted easy\-stratum loss from raising temperature dominates atk=1k=1, so the optimum starts cold\. Askkincreases, failures on the hard stratum receive greater marginal weight and the optimum moves toward the difficulty\-crossing temperature\. WhenC<1C<1, the direction reverses\.
For the equal\-mass family
pE\(t\)=0\.60−0\.30t,pH\(t\)=0\.25\+0\.15t,t∈\[0,1\],p\_\{E\}\(t\)=0\.60\-0\.30t,\\qquad p\_\{H\}\(t\)=0\.25\+0\.15t,\\qquad t\\in\[0,1\],C=2C=2, and \([12](https://arxiv.org/html/2608.14665#S5.E12)\) gives
t1=t2=0,t3≈0\.321,t5≈0\.541,t10≈0\.671,t∞=79\.t\_\{1\}=t\_\{2\}=0,\\quad t\_\{3\}\\approx 0\.321,\\quad t\_\{5\}\\approx 0\.541,\\quad t\_\{10\}\\approx 0\.671,\\quad t\_\{\\infty\}=\\frac\{7\}\{9\}\.The reverse family
π=15,pE\(t\)=710−35t,pH\(t\)=15\+35t,t∈\[0,1\]\\pi=\\frac\{1\}\{5\},\\quad p\_\{E\}\(t\)=\\frac\{7\}\{10\}\-\\frac\{3\}\{5\}t,\\quad p\_\{H\}\(t\)=\\frac\{1\}\{5\}\+\\frac\{3\}\{5\}t,\\quad t\\in\[0,1\]hasC=1/4C=1/4, keeps all probabilities strictly between zero and one, and
t1=1,t2=2930,t3=1318,t∞=512\.t\_\{1\}=1,\\qquad t\_\{2\}=\\frac\{29\}\{30\},\\qquad t\_\{3\}=\\frac\{13\}\{18\},\\qquad t\_\{\\infty\}=\\frac\{5\}\{12\}\.This explicit downward sequence refutes any unconditional claim that the optimal temperature must rise with budget\. The original easy/hard ordering reverses att=5/12t=5/12, so Assumption[3\.1](https://arxiv.org/html/2608.14665#S3.Thmtheorem1)is not maintained over the relevant region\.
Figure 1:Analytical consequences of the theory\. Left: the exact upward\-regime optimizer from Proposition[5\.1](https://arxiv.org/html/2608.14665#S5.Thmtheorem1)\. Center:Beta\(2,k\)\\mathrm\{Beta\}\(2,k\)kernels whose own mass shifts to success probabilities of order1/k1/k; task\-level influence also depends on the density–response factor\. Right: the valid downward\-regime affine family, showing that upward movement is conditional rather than algebraic\.
## 6A completion\-level interpretation
Letπt\(y∣x\)\\pi\_\{t\}\(y\\mid x\)be a probability mass function on a fixed countable support, letCxC\_\{x\}be the set of accepted completions, and assume thatt↦πt\(⋅∣x\)t\\mapsto\\pi\_\{t\}\(\\cdot\\mid x\)is differentiable at the operating point as a map intoℓ1\\ell^\{1\}\. Also assumept\(x\)\>0p\_\{t\}\(x\)\>0andπt\(y∣x\)\>0\\pi\_\{t\}\(y\\mid x\)\>0on the declared support\. Summation overCxC\_\{x\}is a bounded linear functional onℓ1\\ell^\{1\}, so differentiation passes through the accepted\-set sum; theℓ1\\ell^\{1\}derivative also makes the score below integrable, both unconditionally and conditional onY∈CxY\\in C\_\{x\}\. Then
pt\(x\)=∑y∈Cxπt\(y∣x\)p\_\{t\}\(x\)=\\sum\_\{y\\in C\_\{x\}\}\\pi\_\{t\}\(y\\mid x\)and the score identity gives
∂tlogpt\(x\)=𝔼πt\[∂tlogπt\(Y∣x\)∣Y∈Cx\]\.\\partial\_\{t\}\\log p\_\{t\}\(x\)=\\mathbb\{E\}\_\{\\pi\_\{t\}\}\\\!\\left\[\\partial\_\{t\}\\log\\pi\_\{t\}\(Y\\mid x\)\\mid Y\\in C\_\{x\}\\right\]\.\(14\)For the idealized global Gibbs family att\>0t\>0, let the scoresxs\_\{x\}be independent ofttand write
πt\(y∣x\)=exp\(sx\(y\)/t\)Zx\(t\),Zx\(t\)=∑yexp\(sx\(y\)/t\)\.\\pi\_\{t\}\(y\\mid x\)=\\frac\{\\exp\(s\_\{x\}\(y\)/t\)\}\{Z\_\{x\}\(t\)\},\\qquad Z\_\{x\}\(t\)=\\sum\_\{y\}\\exp\(s\_\{x\}\(y\)/t\)\.Assume thatZx\(t\)Z\_\{x\}\(t\)is finite and differentiable in a neighborhood of the operating point, that differentiation may pass through the defining sums, and that𝔼t\|sx\(Y\)\|\\mathbb\{E\}\_\{t\}\|s\_\{x\}\(Y\)\|and𝔼t\[\|sx\(Y\)\|∣Y∈Cx\]\\mathbb\{E\}\_\{t\}\[\|s\_\{x\}\(Y\)\|\\mid Y\\in C\_\{x\}\]are finite\. Then
∂tlogpt\(x\)=𝔼t\[sx\(Y\)\]−𝔼t\[sx\(Y\)∣Y∈Cx\]t2\.\\partial\_\{t\}\\log p\_\{t\}\(x\)=\\frac\{\\mathbb\{E\}\_\{t\}\[s\_\{x\}\(Y\)\]\-\\mathbb\{E\}\_\{t\}\[s\_\{x\}\(Y\)\\mid Y\\in C\_\{x\}\]\}\{t^\{2\}\}\.\(15\)Higher temperature therefore helps when accepted completions occupy lower\-score modes than the current global average, and hurts when accepted completions already occupy the top modes\.
Autoregressive tokenwise temperature has an analogous identity when the summed token score is integrable and differentiation may pass through the path sum\. The derivative of a completion log probability is the sum, over prefixes, of the chosen\-token logit minus the prefix\-average logit, with the appropriate−1/t2\-1/t^\{2\}sign\. Averaging that score over accepted paths recovers \([14](https://arxiv.org/html/2608.14665#S6.E14)\)\. Assumption[3\.1](https://arxiv.org/html/2608.14665#S3.Thmtheorem1)therefore says, roughly, that accepted paths for currently low\-success tasks are deeper in the model’s score landscape than accepted paths for currently high\-success tasks\. This is a mechanism that can generate the condition, not a proof that real models always satisfy it\.
## 7Diagnostics for future tests
Fixtt, assume𝔼\|p˙t\|<∞\\mathbb\{E\}\|\\dot\{p\}\_\{t\}\|<\\infty, and define a finite signed measure on\[0,1\]\[0,1\]by
σt\(B\)=𝔼\[p˙t\(X\)1\{qt\(X\)∈B\}\]\.\\sigma\_\{t\}\(B\)=\\mathbb\{E\}\\\!\\left\[\\dot\{p\}\_\{t\}\(X\)\\,\\mathbf\{1\}\\\{q\_\{t\}\(X\)\\in B\\\}\\right\]\.Equation \([2](https://arxiv.org/html/2608.14665#S3.E2)\) implies
gj\(t\):=Aj\+1′\(t\)j\+1=∫01qjdσt\(q\),j=0,1,…\.g\_\{j\}\(t\):=\\frac\{A\_\{j\+1\}^\{\\prime\}\(t\)\}\{j\+1\}=\\int\_\{0\}^\{1\}q^\{j\}\\,\\mathrm\{d\}\\sigma\_\{t\}\(q\),\\qquad j=0,1,\\ldots\.\(16\)
###### Proposition 7\.1\(Signed\-moment identification and a shape test\)\.
The infinite exact sequence\{gj\(t\)\}j≥0\\\{g\_\{j\}\(t\)\\\}\_\{j\\geq 0\}uniquely determinesσt\\sigma\_\{t\}\. Ifp˙t\(X\)≥0\\dot\{p\}\_\{t\}\(X\)\\geq 0almost surely, then for all integersj,r≥0j,r\\geq 0,
\(−1\)rΔrgj\(t\)=𝔼\[qtjptrp˙t\]≥0\.\(\-1\)^\{r\}\\Delta^\{r\}g\_\{j\}\(t\)=\\mathbb\{E\}\[q\_\{t\}^\{j\}p\_\{t\}^\{r\}\\dot\{p\}\_\{t\}\]\\geq 0\.\(17\)
###### Proof\.
If two finite signed measures on a compact interval have the same moments, they integrate every polynomial equally\. Polynomial density inC\[0,1\]C\[0,1\]then implies equality of the measures\. Repeated forward differences in \([16](https://arxiv.org/html/2608.14665#S7.E16)\) multiply the integrand by\(q−1\)r=\(−p\)r\(q\-1\)^\{r\}=\(\-p\)^\{r\}, which proves \([17](https://arxiv.org/html/2608.14665#S7.E17)\)\. ∎
This is an application of the classical compact\-interval moment problem\[[7](https://arxiv.org/html/2608.14665#bib.bib7)\], not a claim of a new moment theorem\. For a general signed response measure, complete monotonicity need not hold\. Moreover, identification uses the entire infinite, noiseless sequence; inversion from finitely many noisy derivatives is ill\-posed\[[6](https://arxiv.org/html/2608.14665#bib.bib6)\]\. The useful output is therefore a set of diagnostics:
1. 1\.estimate taskwiseptp\_\{t\}and local log\-success responses at nearby temperatures;
2. 2\.test whether the conditional responsemt\(u\)m\_\{t\}\(u\)is nonincreasing, preferably with held\-out tasks and uncertainty bands;
3. 3\.separately test strict single\-peakedness of aggregatet↦Ak\(t\)t\\mapsto A\_\{k\}\(t\);
4. 4\.only when both conditions are supported, predict a nondecreasing sequence of optimal temperatures;
5. 5\.use violations of \([17](https://arxiv.org/html/2608.14665#S7.E17)\) to reject the stronger hypothesis that the temperature increase helps every task\.
These are preregisterable predictions for future empirical work\. They are not evidence for the present theorem\.
## 8Related work and claim boundaries
#### Temperature and multi\-sample inference\.
Chen et al\. \[[2](https://arxiv.org/html/2608.14665#bib.bib2)\]documented the increasing optimal\-temperature pattern for code generation\.Du et al\. \[[5](https://arxiv.org/html/2608.14665#bib.bib5)\]studied temperature selection for majority vote and best\-of\-NN, reporting lower temperatures at small samples, higher temperatures at larger samples, and roughly single\-peaked curves\. They also propose an entropy\-based selector\.Chow et al\. \[[3](https://arxiv.org/html/2608.14665#bib.bib3)\]reported temperature/sample co\-scaling and an easy/hard\-task interpretation\.Slocum et al\. \[[12](https://arxiv.org/html/2608.14665#bib.bib12), Appendix C\]explicitly state both fixed\-task budget invariance and the qualitative asymmetry by which low\-success tasks can matter more at high budgets\. Their variance interpretation does not follow from concavity alone: a mean\-preserving spread weakly lowers the expectation of a concave function\. Our response\-order assumption instead states the directional task heterogeneity needed for a rigorous ordering\.Dang et al\. \[[4](https://arxiv.org/html/2608.14665#bib.bib4)\]formalize a bias–variance view of pass@kkand show that temperature scaling trades bias against variance\. These works establish the empirical target and neighboring mechanisms; our contribution is the conditional derivative\-sign and optimizer ordering, not discovery of the pattern, fixed\-task invariance, or its first qualitative explanation\.
#### Pass@kkreweighting and inference scaling\.
Barakat et al\. \[[1](https://arxiv.org/html/2608.14665#bib.bib1)\]derive the samek\(1−p\)k−1k\(1\-p\)^\{k\-1\}prompt weight in policy gradients and study how hard\-prompt emphasis can conflict with pass@1 optimization\.Walder and Karkhanis \[[14](https://arxiv.org/html/2608.14665#bib.bib14)\]directly optimize pass@kk, derive gradient estimators, and report that larger target budgets prioritize harder problems\.Schaeffer et al\. \[[11](https://arxiv.org/html/2608.14665#bib.bib11)\]use the distribution of one\-attempt success to explain aggregate power\-law inference scaling\. We specialize the distributional view to a scalar inference\-time control and ask when its maximizers are ordered across budgets\.
#### Configuration portfolios\.
Zhang et al\. \[[16](https://arxiv.org/html/2608.14665#bib.bib16)\]formulate mixed allocation across configurations, including temperature, model, and language; establish the convex relaxed failure objective; and show that mixtures can beat one configuration\.Wu et al\. \[[15](https://arxiv.org/html/2608.14665#bib.bib15)\]empirically demonstrate complementary temperature\-specific task subsets\.Su et al\. \[[13](https://arxiv.org/html/2608.14665#bib.bib13)\]learn prompt\- and budget\-conditioned decoding policies, a broader adaptive setting than the fixed\-temperature schedules studied here\. Appendix[A](https://arxiv.org/html/2608.14665#A1)starts from the OSCA objective and records exact integer endpoint and finite\-type structural consequences\. We do not claim a new convex formulation or the first benefit of mixing or budget\-conditioned decoding\.
#### Priority statement\.
Slocum et al\. \[[12](https://arxiv.org/html/2608.14665#bib.bib12)\]own the fixed\-task invariance observation and a qualitative hard/easy\-task account of aggregate movement\. To our knowledge, a targeted primary\-source search through July 19, 2026 did not locate a prior result combining decreasing conditional log\-success response with derivative\-sign nesting and ordered inference\-time temperature maximizers\. We therefore make only that narrow sufficient\-condition claim\. The result is a pass@kk\-temperature specialization of classical likelihood\-ratio comparative statics; it is not the first theory of pass@kk, the first fixed\-task no\-go observation, the first hard\-task explanation, a new general monotone\-comparative\-statics theorem, or a universal law\.
## 9Limitations and conclusion
The analysis makes four substantive modeling commitments\. Samples are conditionally i\.i\.d\.; the temperature is held fixed within each homogeneous schedule; the verifier exactly recognizes success; and the benchmark average is the population of interest\. Correlation among attempts, adaptive decoding, verifier error, or distribution shift can all change the objective\. The response order and single\-peakedness must be checked rather than assumed\. Finite benchmarks can also produce grid and estimation noise that masks or reverses the population ordering\.
Within those boundaries, the result formalizes a precise sufficient condition beyond the prior fixed\-task observation\. Pass@kkdoes not make a fixed task prefer a different temperature\. It changes the benchmark\-level marginal by a power tilt toward tasks that still tend to fail\. When lower\-success tasks have larger conditional log\-success responses, that tilt creates nested derivative signs; with single\-peaked curves, the best temperature can only move upward\. The two\-stratum phase diagram shows both why the familiar pattern emerges and how it can reverse\. This turns a broad quality–diversity intuition into a testable conditional statement\.
## Tool\-use disclosure
OpenAI Codex was used to assist with literature search, algebraic and numerical checks, figure generation, and manuscript drafting\. It is not an author\. The named human author is responsible for independently verifying the claims, references, and final submitted text and for taking responsibility for the work\.
## References
- Barakat et al\. \[2026\]Anas Barakat, Souradip Chakraborty, Khushbu Pahwa, and Amrit Singh Bedi\.Why pass@k optimization can degrade pass@1: Prompt interference in LLM post\-training\.*arXiv preprint arXiv:2602\.21189*, 2026\.URL[https://arxiv\.org/abs/2602\.21189](https://arxiv.org/abs/2602.21189)\.
- Chen et al\. \[2021\]Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert\-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N\. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba\.Evaluating large language models trained on code\.*arXiv preprint arXiv:2107\.03374*, 2021\.URL[https://arxiv\.org/abs/2107\.03374](https://arxiv.org/abs/2107.03374)\.
- Chow et al\. \[2025\]Yinlam Chow, Guy Tennenholtz, Izzeddin Gur, Vincent Zhuang, Bo Dai, Aviral Kumar, Rishabh Agarwal, Sridhar Thiagarajan, Craig Boutilier, and Aleksandra Faust\.Inference\-aware fine\-tuning for best\-of\-n sampling in large language models\.In*International Conference on Learning Representations*, 2025\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2025/hash/c40bed606c51c8e827c1ba75aa2da054\-Abstract\-Conference\.html](https://proceedings.iclr.cc/paper_files/paper/2025/hash/c40bed606c51c8e827c1ba75aa2da054-Abstract-Conference.html)\.
- Dang et al\. \[2025\]Xingyu Dang, Christina Baek, Kaiyue Wen, Zico Kolter, and Aditi Raghunathan\.Weight ensembling improves reasoning in language models\.*arXiv preprint arXiv:2504\.10478*, 2025\.URL[https://arxiv\.org/abs/2504\.10478](https://arxiv.org/abs/2504.10478)\.COLM 2025\.
- Du et al\. \[2025\]Weihua Du, Yiming Yang, and Sean Welleck\.Optimizing temperature for language models with multi\-sample inference\.In*Proceedings of the 42nd International Conference on Machine Learning*, volume 267 of*Proceedings of Machine Learning Research*, pages 14648–14668\. PMLR, 2025\.URL[https://proceedings\.mlr\.press/v267/du25f\.html](https://proceedings.mlr.press/v267/du25f.html)\.
- Gerth et al\. \[2021\]Daniel Gerth, Bernd Hofmann, Christopher Hofmann, and Stefan Kindermann\.The hausdorff moment problem in the light of ill\-posedness of type I\.*arXiv preprint arXiv:2104\.06029*, 2021\.URL[https://arxiv\.org/abs/2104\.06029](https://arxiv.org/abs/2104.06029)\.
- Hildebrandt and Schoenberg \[1933\]T\. H\. Hildebrandt and I\. J\. Schoenberg\.On linear functional operations and the moment problem for a finite interval in one or several dimensions\.*Annals of Mathematics*, 34\(2\):317–328, 1933\.doi:10\.2307/1968205\.
- Karlin and Rubin \[1956\]Samuel Karlin and Herman Rubin\.The theory of decision procedures for distributions with monotone likelihood ratio\.*The Annals of Mathematical Statistics*, 27\(2\):272–299, 1956\.doi:10\.1214/aoms/1177728259\.
- Lehmann \[1955\]E\. L\. Lehmann\.Ordered families of distributions\.*The Annals of Mathematical Statistics*, 26\(3\):399–419, 1955\.doi:10\.1214/aoms/1177728487\.
- Milgrom and Shannon \[1994\]Paul Milgrom and Chris Shannon\.Monotone comparative statics\.*Econometrica*, 62\(1\):157–180, 1994\.doi:10\.2307/2951479\.
- Schaeffer et al\. \[2025\]Rylan Schaeffer, Joshua Kazdan, John Hughes, Jordan Juravsky, Sara Price, Aengus Lynch, Erik Jones, Robert Kirk, Azalia Mirhoseini, and Sanmi Koyejo\.How do large language monkeys get their power \(laws\)?In*Proceedings of the 42nd International Conference on Machine Learning*, volume 267 of*Proceedings of Machine Learning Research*, pages 53132–53176\. PMLR, 2025\.URL[https://proceedings\.mlr\.press/v267/schaeffer25a\.html](https://proceedings.mlr.press/v267/schaeffer25a.html)\.
- Slocum et al\. \[2025\]Stewart Slocum, Asher Parker\-Sartori, and Dylan Hadfield\-Menell\.Diverse preference learning for capabilities and alignment\.In*International Conference on Learning Representations*, 2025\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2025/hash/3df1eca840e82b11bbc33f68c773c38e\-Abstract\-Conference\.html](https://proceedings.iclr.cc/paper_files/paper/2025/hash/3df1eca840e82b11bbc33f68c773c38e-Abstract-Conference.html)\.
- Su et al\. \[2026\]Chloe H\. Su, Zhe Ye, Samuel Tenka, Aidan Yang, Soonho Kong, and Udaya Ghai\.Learning adaptive LLM decoding\.*arXiv preprint arXiv:2603\.09065*, 2026\.URL[https://arxiv\.org/abs/2603\.09065](https://arxiv.org/abs/2603.09065)\.
- Walder and Karkhanis \[2025\]Christian Walder and Deep Karkhanis\.Pass@K policy optimization: Solving harder reinforcement learning problems\.*arXiv preprint arXiv:2505\.15201*, 2025\.URL[https://arxiv\.org/abs/2505\.15201](https://arxiv.org/abs/2505.15201)\.
- Wu et al\. \[2025\]Yuheng Wu, Azalia Mirhoseini, and Thierry Tambe\.On the role of temperature sampling in test\-time scaling\.*arXiv preprint arXiv:2510\.02611*, 2025\.URL[https://arxiv\.org/abs/2510\.02611](https://arxiv.org/abs/2510.02611)\.
- Zhang et al\. \[2025\]Kexun Zhang, Shang Zhou, Danqing Wang, William Yang Wang, and Lei Li\.Scaling LLM inference efficiently with optimized sample compute allocation\.In*Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies*, pages 7959–7973\. Association for Computational Linguistics, 2025\.doi:10\.18653/v1/2025\.naacl\-long\.404\.URL[https://aclanthology\.org/2025\.naacl\-long\.404/](https://aclanthology.org/2025.naacl-long.404/)\.
## Appendix AExact refinements of configuration allocation
This appendix begins from the mixed\-configuration failure objective studied by OSCA\[[16](https://arxiv.org/html/2608.14665#bib.bib16)\]\. Its purpose is to make two discrete and asymptotic consequences explicit, not to reclaim the general formulation\.
### A\.1An exact integer criterion for two configurations
LetqA\(X\)q\_\{A\}\(X\)andqB\(X\)q\_\{B\}\(X\)be per\-sample failure probabilities for two configurations, and writepA=1−qAp\_\{A\}=1\-q\_\{A\}andpB=1−qBp\_\{B\}=1\-q\_\{B\}\. Assume that, conditional onXX, draws are independent across and within configurations\. Withnnsamples fromBBandk−nk\-nfromAA, define
Fk\(n\)=𝔼\[qA\(X\)k−nqB\(X\)n\],n=0,…,k\.F\_\{k\}\(n\)=\\mathbb\{E\}\[q\_\{A\}\(X\)^\{k\-n\}q\_\{B\}\(X\)^\{n\}\],\\qquad n=0,\\ldots,k\.\(18\)
###### Proposition A\.1\(Discrete convexity and strict\-mixing criterion\)\.
Forn=0,…,k−2n=0,\\ldots,k\-2,
Fk\(n\+2\)−2Fk\(n\+1\)\+Fk\(n\)=𝔼\[qAk−n−2qBn\(qA−qB\)2\]≥0\.F\_\{k\}\(n\+2\)\-2F\_\{k\}\(n\+1\)\+F\_\{k\}\(n\)=\\mathbb\{E\}\\\!\\left\[q\_\{A\}^\{k\-n\-2\}q\_\{B\}^\{n\}\(q\_\{A\}\-q\_\{B\}\)^\{2\}\\right\]\\geq 0\.\(19\)Fork≥2k\\geq 2, some interior integer schedule strictly beats both homogeneous endpoints if and only if
𝔼\[qAk−1\(pA−pB\)\]\\displaystyle\\mathbb\{E\}\[q\_\{A\}^\{k\-1\}\(p\_\{A\}\-p\_\{B\}\)\]<0,\\displaystyle<0,\(20\)𝔼\[qBk−1\(pA−pB\)\]\\displaystyle\\mathbb\{E\}\[q\_\{B\}^\{k\-1\}\(p\_\{A\}\-p\_\{B\}\)\]\>0\.\\displaystyle\>0\.\(21\)
###### Proof\.
Expanding the second difference of \([18](https://arxiv.org/html/2608.14665#A1.E18)\) gives \([19](https://arxiv.org/html/2608.14665#A1.E19)\)\. Hence the first differences form a nondecreasing sequence\. The left endpoint is not optimal exactly when
Fk\(1\)−Fk\(0\)=𝔼\[qAk−1\(qB−qA\)\]=𝔼\[qAk−1\(pA−pB\)\]<0\.F\_\{k\}\(1\)\-F\_\{k\}\(0\)=\\mathbb\{E\}\[q\_\{A\}^\{k\-1\}\(q\_\{B\}\-q\_\{A\}\)\]=\\mathbb\{E\}\[q\_\{A\}^\{k\-1\}\(p\_\{A\}\-p\_\{B\}\)\]<0\.Likewise, the right endpoint is not optimal exactly when
Fk\(k\)−Fk\(k−1\)=𝔼\[qBk−1\(pA−pB\)\]\>0\.F\_\{k\}\(k\)\-F\_\{k\}\(k\-1\)=\\mathbb\{E\}\[q\_\{B\}^\{k\-1\}\(p\_\{A\}\-p\_\{B\}\)\]\>0\.For a convex sequence, both strict endpoint conditions hold if and only if the minimum is attained strictly below both endpoints at an interior index\. ∎
### A\.2Finite task types and the large\-budget limit
Suppose there arerrtask types with masseswi\>0w\_\{i\}\>0summing to one, and configurationsj=1,…,mj=1,\\ldots,mhave failure probabilitiesqji∈\(0,1\]q\_\{ji\}\\in\(0,1\]\. Letaj∈ℝ\+ra\_\{j\}\\in\\mathbb\{R\}\_\{\+\}^\{r\}have coordinatesaji=−logqjia\_\{ji\}=\-\\log q\_\{ji\}\. For relaxed allocation proportionszzin the simplexΔm\\Delta\_\{m\},
Fk\(z\)=∑i=1rwiexp\(−k∑j=1mzjaji\)\.F\_\{k\}\(z\)=\\sum\_\{i=1\}^\{r\}w\_\{i\}\\exp\\\!\\left\(\-k\\sum\_\{j=1\}^\{m\}z\_\{j\}a\_\{ji\}\\right\)\.\(22\)
###### Proposition A\.2\(Sparse support, dominance, and maximin limit\)\.
There exists an optimizer of \([22](https://arxiv.org/html/2608.14665#A1.E22)\) supported on at mostrrconfigurations\. A configuration whose hazard vector is componentwise dominated by a convex combination of other hazard vectors can be removed without worsening the optimum\. Finally, with
ϕk\(z\)=−1klogFk\(z\),m\(z\)=mini∑jzjaji,wmin=miniwi,\\phi\_\{k\}\(z\)=\-\\frac\{1\}\{k\}\\log F\_\{k\}\(z\),\\qquad m\(z\)=\\min\_\{i\}\\sum\_\{j\}z\_\{j\}a\_\{ji\},\\qquad w\_\{\\min\}=\\min\_\{i\}w\_\{i\},the uniform bound
m\(z\)≤ϕk\(z\)≤m\(z\)\+log\(1/wmin\)km\(z\)\\leq\\phi\_\{k\}\(z\)\\leq m\(z\)\+\\frac\{\\log\(1/w\_\{\\min\}\)\}\{k\}\(23\)holds\. Every accumulation point of relaxed optimizers ask→∞k\\to\\inftymaximizesm\(z\)m\(z\)\.
###### Proof\.
Leth⋆=∑jzj⋆ajh^\{\\star\}=\\sum\_\{j\}z\_\{j\}^\{\\star\}a\_\{j\}be an optimal point in the convex hull of the hazard vectors, and let
G\(h\)=∑iwie−khi\.G\(h\)=\\sum\_\{i\}w\_\{i\}e^\{\-kh\_\{i\}\}\.First\-order optimality implies∇G\(h⋆\)⋅\(h−h⋆\)≥0\\nabla G\(h^\{\\star\}\)\\cdot\(h\-h^\{\\star\}\)\\geq 0for every feasiblehh\. Thush⋆h^\{\\star\}lies in the exposed face maximizing−∇G\(h⋆\)⋅h\-\\nabla G\(h^\{\\star\}\)\\cdot h\. The exposing normal is strictly positive and nonzero, so the face has affine dimension at mostr−1r\-1\. Carathéodory’s theorem within that face representsh⋆h^\{\\star\}using at mostrrhazard vectors\.
Ifaja\_\{j\}is componentwise dominated by a convex combination of the other vectors, replacing any mass onjjby that combination weakly increases every coordinate ofhh\. SinceGGis decreasing in every coordinate, the objective cannot worsen\.
Leti⋆i^\{\\star\}attainm\(z\)m\(z\)\. Since the weights sum to one,
wmine−km\(z\)≤Fk\(z\)≤e−km\(z\)\.w\_\{\\min\}e^\{\-km\(z\)\}\\leq F\_\{k\}\(z\)\\leq e^\{\-km\(z\)\}\.Taking−k−1log\-k^\{\-1\}\\loggives \([23](https://arxiv.org/html/2608.14665#A1.E23)\)\. The convergenceϕk→m\\phi\_\{k\}\\to mis uniform on the compact simplex, so every accumulation point of maximizers ofϕk\\phi\_\{k\}, equivalently minimizers ofFkF\_\{k\}, maximizesmm\. ∎
## Appendix BAdditional proof details
### B\.1Association inequality used in Theorem[3\.3](https://arxiv.org/html/2608.14665#S3.Thmtheorem3)
IfU,U′U,U^\{\\prime\}are independent with the same law andf,gf,gare both nonincreasing, then
2Cov\(f\(U\),g\(U\)\)=𝔼\[\(f\(U\)−f\(U′\)\)\(g\(U\)−g\(U′\)\)\]≥0\.2\\operatorname\{Cov\}\(f\(U\),g\(U\)\)=\\mathbb\{E\}\[\(f\(U\)\-f\(U^\{\\prime\}\)\)\(g\(U\)\-g\(U^\{\\prime\}\)\)\]\\geq 0\.Apply this withf=mtf=m\_\{t\}andg\(u\)=\(1−u\)ℓ−kg\(u\)=\(1\-u\)^\{\\ell\-k\}\.
### B\.2Boundary probability cases
The log\-success\-response formulation is cleanest when0<pt<10<p\_\{t\}<1\. If a task haspt=0p\_\{t\}=0at an interior temperature andptp\_\{t\}is differentiable and nonnegative in a neighborhood, thenp˙t=0\\dot\{p\}\_\{t\}=0, so the task does not contribute to \([2](https://arxiv.org/html/2608.14665#S3.E2)\) at that point\. Tasks withpt=1p\_\{t\}=1receive zero weight fork\>1k\>1\. Extensions can restrictνk,t\\nu\_\{k,t\}to positive\-weight tasks and take limits whenZk,t\>0Z\_\{k,t\}\>0, one\-sided derivatives exist, and dominated convergence justifies the limits\.
## Appendix CReproducibility
The accompanying standard\-library Python checks verify the closed\-form optimizers, upward and downward regimes, discrete\-convexity identities, finite\-difference diagnostics, and displayed numerical values\. The figure is generated from the analytical formulas\. These computations are regression checks; the general results rely on the proofs above\.Similar Articles
When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling
This paper identifies the 'modal ceiling' and 'correlation ceiling' in test-time scaling for reasoning models, showing that beyond a few dozen samples, additional sampling does not improve selection accuracy and can even harm it, highlighting the identifiability gap between generating and recognizing correct answers.
Forget Without Compromise: Nexus Sampling for Streaming KV-Cache Eviction Under Fixed Budgets
Introduces Nexus Sampling, a training-free KV-cache eviction method using weighted reservoir sampling instead of deterministic top-k, improving long-context LLM inference under fixed memory budgets, matching dense attention performance at 80% eviction.
Prof-K: Probabilistic One-Pass Filtering for Efficient Top-k Selection
Prof-K is a probabilistic one-pass filtering algorithm for fast, scalable top-k selection with correctness guarantees, achieving 1.5x–10x speedups over PyTorch topk and RadiK, especially in large-scale small-k regimes.
Sample Where You Struggle: Sharpening Base Model Reasoning via Entropy-Guided Power Sampling
This paper introduces Entropy-Guided Power Sampling (EGPS), a training-free and verifier-free sampler that improves the efficiency of power sampling for enhancing base language model reasoning. EGPS achieves up to 12.6x speedup over standard Metropolis-Hastings sampling while reaching best or tied-best accuracy on benchmarks like MATH500, HumanEval, and GPQA.
The conditional superiority of fast silicon sampling
This study evaluates silicon sampling methods, showing that fast modes outperform slow modes in efficiency and fidelity while highlighting limitations in accurately representing opinion variance.