Mentored Decoding:更快推理遇上Boosting
摘要
本文介绍了mentored decoding,这是一种有损推测解码的正式方法,通过允许与目标模型的受控偏差来提高语言模型的推理速度,将其与提升理论联系起来,并证明了优化的关键性质。
查看缓存全文
缓存时间: 2026/09/29 09:37
# Mentored Decoding: Faster Inference meets Boosting
Source: [https://arxiv.org/html/2609.30474](https://arxiv.org/html/2609.30474)
###### Abstract
Speculative decoding is a successful technique speeding up inference of a target autoregressive language model via a fast drafter model\. Lossy speculative decoding allows a drift with respect to the target to further improve speed\. Interestingly, it has been observed experimentally that the resulting model canalsobeat the targetquality\-wise\. Our paper formally proves how such a feat is possible with a formal approach to lossy speculative decoding calledmentored decoding\. To get there, we connect inference to a celebrated ML training theory,boosting, and proceed via the generalization of mentored decoding to the whole set offf\-divergences\. We uncover key properties of mentored decoding, among which \(i\) the particularly appealing geometric nature of the total variation case, \(ii\) simple approximations for anyff\-divergence in direct relation with boosting compliance, and \(iii\) adivergence independentO\(n\)O\(n\)space andO\(sort\(n\)\)O\(\\mathrm\{sort\}\(n\)\)time data structure built on drafter and target outputs, which allows to query the optimal parameters of the dual problem inO\(logn\)O\(\\log n\)time and constructing optimal mentored distributions inO\(n\)O\(n\)time for anyff\-divergence\.
## 1Introduction
Large language model \(LLM\) inference is often memory\-bandwidth bound during sequential token generation\.Speculative decoding\([Leviathan et al\., 2023](https://arxiv.org/html/2609.30474#bib.bib2);[Chen et al\., 2023](https://arxiv.org/html/2609.30474#bib.bib1)\)\(SD\) accelerates this process by using a smaller, faster draft model to propose candidate tokens, which are subsequently verified in parallel by the target model\. Crucially, SD constrains theoutput distributionto be the same as the target’s, which inherently constrains the overall acceptance probability as a tight function of drafter and target\. This fundamental limit is not an artifact of SD:[Sun et al\. \(2023\)](https://arxiv.org/html/2609.30474#bib.bib28)proved that SD achieves the optimal acceptance rate under exact target distribution matching\. To lift this cap,[Tran\-Thien \(2023\)](https://arxiv.org/html/2609.30474#bib.bib45)first framedmentored decoding\(MD\) as a constrained optimization problem maximizing draft acceptance subject to bounded divergence between target and the output\. The target, authorizing deviations with respect to its output as long as they do not substantially diverge, becomes thementorin MD and the final output, which mixes tokens from both models, is a composite output distribution from an ensemble model\. Initially,[Tran\-Thien \(2023\)](https://arxiv.org/html/2609.30474#bib.bib45)used the reverse Kullback\-Leibler divergence as divergence measure\. To the best of our knowledge, this was the first formal attempt to alleviate SD’s acceptance probability cap, even when heuristic proposals started in fact to flourish from the introduction oflenientSD in[Leviathan et al\. \(2023\)](https://arxiv.org/html/2609.30474#bib.bib2)\.
It is hard to exaggerate the experimental success of SD\([Kim et al\., 2023](https://arxiv.org/html/2609.30474#bib.bib39);[Cai et al\., 2024](https://arxiv.org/html/2609.30474#bib.bib32);[Li et al\., 2024b](https://arxiv.org/html/2609.30474#bib.bib33);[Fu et al\., 2024](https://arxiv.org/html/2609.30474#bib.bib40);[Yang et al\., 2023](https://arxiv.org/html/2609.30474#bib.bib42);[He et al\., 2024](https://arxiv.org/html/2609.30474#bib.bib41);[Wang et al\., 2025b](https://arxiv.org/html/2609.30474#bib.bib51);[Hao and Mou, 2026](https://arxiv.org/html/2609.30474#bib.bib48)\)\(and many others, See Section[2](https://arxiv.org/html/2609.30474#S2)\)\. Among the chorus of approval for speeding up inference, distinct voices later started to emerge, either on the fact that the drafter, even when smaller than the target, can occasionally produce high quality tokens that are then underutilized\([Liao et al\., 2025](https://arxiv.org/html/2609.30474#bib.bib52)\), or, more importantly, that thecombinationof models achieved in the composite output can in fact beat the target onquality metricsas well\([Qin et al\., 2025](https://arxiv.org/html/2609.30474#bib.bib53);[Li et al\., 2026a](https://arxiv.org/html/2609.30474#bib.bib4);[Zhong et al\., 2025](https://arxiv.org/html/2609.30474#bib.bib10)\)\. While the technical leads in the formal analysis of SD/MD inference speed\-up alone are already scarce\([Leviathan et al\., 2023](https://arxiv.org/html/2609.30474#bib.bib2);[Tran\-Thien, 2023](https://arxiv.org/html/2609.30474#bib.bib45);[Sun et al\., 2023](https://arxiv.org/html/2609.30474#bib.bib28);[Yin et al\., 2024](https://arxiv.org/html/2609.30474#bib.bib27);[Pankratov and Alistarh, 2026](https://arxiv.org/html/2609.30474#bib.bib9)\), there is to our knowledge no such analysis combining the possibility of speeding up inference to that of improving any quality metric on the output\. Our paperproposes the first analysis of this kind, on joint inference efficiency and model quality properties of mentored decoding as originally designed in[Tran\-Thien \(2023\)](https://arxiv.org/html/2609.30474#bib.bib45), demonstrating in particular how the MD setting achieves connections with one of machine learning \(ML\)’s seminal training framework especially suited to analyze the quality of model combinations: Boosting\([Schapire and Freund, 2012](https://arxiv.org/html/2609.30474#bib.bib23)\)\. Our contribution to get there is threefold: \(i\) we substantially improve the state of the art understanding of MD, \(ii\) we design and analyze a new boosting approach for the connection, and \(iii\) we design and analyze efficient algorithms to operate this connection on the MD side\. On improving MD understanding, we use as a warmup the particular case of the total variation divergence\.[Yin et al\. \(2024\)](https://arxiv.org/html/2609.30474#bib.bib27)partially covered the case but left aside the characterization of the set of optimal solutions\. It turns out that it has absolutely remarkable properties\. First, a deceptively simple geometric appeal: it is the intersection of thenn\-dimensional hyperrectangle defined by the drafter and target coordinates with the probability simplex and activating the divergence constraint\. Second, a remarkable extent: this set is big enough to contain the optimal solutions forallstrictly convex divergences\. Its properties bestow optimal solutions with unique appealing geometric and computational features, yielding extremely simple optimal solutions like the convex combination used in several papers\([Yin et al\., 2024](https://arxiv.org/html/2609.30474#bib.bib27);[Wang et al\., 2025b](https://arxiv.org/html/2609.30474#bib.bib51);[Zhong et al\., 2025](https://arxiv.org/html/2609.30474#bib.bib10)\)\. We then characterize the general solution for anyff\-divergence\. In particular, for any strictly convexff, the optimal mentored distribution is unique and takes an exceptionally simple, intuitive coordinate\-wise clamping form:π∗=max\{αq,min\{p,βq\}\}\\pi^\{\*\}=\\max\\\{\\alpha q,\\min\\\{p,\\beta q\\\}\\\}with0≤α<1<β0\\leq\\alpha<1<\\beta, tracing a one\-dimensional trajectory in the simplex connecting targetqqto drafterpp\. Remarkably, this trajectory is independent fromff\. Additionally, for any generatorffdifferentiable inz=1z=1, the curve giving the thresholdff\-divergence as a function of the optimal acceptance probability isalways of right\-derivative 0at speculative decoding’s "minimal" acceptance probability\. Hence, it is always possible to at least reasonably improve SD’s acceptance probability at negligible divergence cost to the target\. On the connection with Boosting, we first design a multiclass extension of the self\-normalized boosting algorithm of[Nock and Nielsen \(2007\)](https://arxiv.org/html/2609.30474#bib.bib12), simpler and more efficient than AdaBoost yet giving rates that compete with the state of the art\([Bartlett et al\., 1998](https://arxiv.org/html/2609.30474#bib.bib14)\)\. Mentored decoding being an inference technique, we develop two distinct paths connecting it with boosting\. The first path is general and relies on a novel use of boosting, showing how the compositeoutputsof mentored decoding "hides" a combination ofmodelsthat exhibits boosting properties\. Since the MD’s output depends on the drafter and target’s output distributions, we ultimately deliver boosting compliance forallrelated combinations of models, depending on these distributions and also on boosting’s key parameter: theedgeof the drafter and target’s last layers\. This makes it possible to evaluate how well drafter and target "complement" each other from the output quality’s standpoint, offering a concrete criterion to then select drafter and / or target from a pool of already available models – that now abound in repositories of public and private spaces\. Our second path connecting mentored decoding and boosting is specific to the total variation divergence, for which the conveniences of the set of optimal solutions make it possible to carve at reduced formal cost the distribution corresponding to the boosted ensemble of drafter and target directly in the optimal set of mentored decoding\. From the standpoint of algorithms, another remarkable invariant emerges at the level of generality of allff\-divergences: we show that there exists a simpledivergence independent"breakpoint" data structure of size≤n\\leq n\(=the vocabulary size\) which then allows to compute the optimal per\-token acceptance and resampling probabilities, foranyff; the computation of this data structure takesO\(sort\(n\)\)O\(\\mathrm\{sort\}\(n\)\), i\.e\. the complexity of sortingnnreals\. While solving a non\-linear constrained optimization problem per token might seem computationally demanding, the practical runtime overhead is in fact negligible\. First, our data structure reduces the optimization to a single pass query overO\(logn\)O\(\\log n\)precomputed breakpoints\. This can then be used to approximately find the optimal parameters inO\(1\)O\(1\)– i\.e\. with guarantees on the divergence –, and this can also be used to find theexactoptimal parameters for the dual problem inO\(1\)O\(1\)– i\.e\. minimize theff\-divergence subject to lowerbounded acceptance probability –\. Second, in modern LLM inference where top\-kktruncation is standard, the optimization domain reduces naturally fromnnto justkkcandidates\. Our paper is organized as follows: the next Section[2](https://arxiv.org/html/2609.30474#S2)summarizes related work\. Then, follow three key parts of our paper, organized so that readers familiar with only one of the two frameworks used \(lossy speculative decoding and boosting\) may easily process the part on which they are most familiar and then connect with the other one: Section[3](https://arxiv.org/html/2609.30474#S3)presents the main results on the mentored decoding side, Section[4](https://arxiv.org/html/2609.30474#S4)presents the boosting side and its connection to mentored decoding, finally Section[5](https://arxiv.org/html/2609.30474#S5)presents the algorithmic sides of the theory discussed\. A following Section[6](https://arxiv.org/html/2609.30474#S6)discusses additional topics related to mentored decoding and boosting, and a last Section[7](https://arxiv.org/html/2609.30474#S7)concludes with avenues for future research\. Our paper is self\-contained: all proofs are given either in the main body of the paper or in an Appendix starting page[VIII](https://arxiv.org/html/2609.30474#S8)\.
## 2Related Work
On the purespeculative decodingside, i\.e\.lossless decoding,Blockwise Parallel Decoding\([Stern et al\., 2018](https://arxiv.org/html/2609.30474#bib.bib25)\)pioneered interleaving fast draft sequence generation with parallel target verification to accelerate greedy sequence\-to\-sequence decoding\.[Xia et al\. \(2023\)](https://arxiv.org/html/2609.30474#bib.bib26)refined this approach and coined the termspeculative decoding, drawing analogy to speculative execution in computer architecture\.[Leviathan et al\. \(2023\)](https://arxiv.org/html/2609.30474#bib.bib2);[Chen et al\. \(2023\)](https://arxiv.org/html/2609.30474#bib.bib1)independently generalized the framework to multinomial sampling, establishing the standard rejection\-sampling formulation described in Section[3](https://arxiv.org/html/2609.30474#S3)\.[Sun et al\. \(2023\)](https://arxiv.org/html/2609.30474#bib.bib28)proved that this formulation achieves the optimal acceptance rate under exact target distribution matching\. Following its inception, speculative decoding has evolved along several dimensions\. A first one moved towards better aligned or faster draft models: Self\-speculative decoding\([Kim et al\., 2023](https://arxiv.org/html/2609.30474#bib.bib39);[Zhang et al\., 2024](https://arxiv.org/html/2609.30474#bib.bib37);[Liu et al\., 2024](https://arxiv.org/html/2609.30474#bib.bib38);[Gloeckle et al\., 2024](https://arxiv.org/html/2609.30474#bib.bib36);[Cai et al\., 2024](https://arxiv.org/html/2609.30474#bib.bib32)\)eliminates the need for a separate draft model by adding lightweight prediction heads, skipping transformer layers, pruning sub\-networks, or early\-exiting from the target model\.EAGLEand its variants\([Li et al\., 2024b](https://arxiv.org/html/2609.30474#bib.bib33);[Li et al\., 2024a](https://arxiv.org/html/2609.30474#bib.bib34);[Li et al\., 2026b](https://arxiv.org/html/2609.30474#bib.bib35)\)perform autoregressive drafting over target hidden feature representations rather than discrete tokens\. DistillSpec\([Zhou et al\., 2024](https://arxiv.org/html/2609.30474#bib.bib44)\)aligns draft models to target models during training by minimizingff\-divergence objectives\. Recently,DFlash\([Chen et al\., 2026](https://arxiv.org/html/2609.30474#bib.bib43)\)proposed non\-autoregressive block diffusion models for low\-latency draft prediction\. A second one moved towards model\-free drafting:[Fu et al\. \(2024\)](https://arxiv.org/html/2609.30474#bib.bib40)introducedLookahead Decoding, generating candidate n\-grams via parallel Jacobi fixed\-point iteration without relying on a draft model\.LLMA\([Yang et al\., 2023](https://arxiv.org/html/2609.30474#bib.bib42)\)copies recurring n\-gram patterns directly from the input prompt or reference documents, whileREST\([He et al\., 2024](https://arxiv.org/html/2609.30474#bib.bib41)\)retrieves candidate phrases from external datastores\. A third one moved towards multi\-draft and tree verification: Rather than proposing a single linear sequence of candidate tokens, tree\-based speculation generates candidate trees verified in parallel using tree\-attention masks\([Miao et al\., 2024](https://arxiv.org/html/2609.30474#bib.bib30);[Chen et al\., 2024](https://arxiv.org/html/2609.30474#bib.bib31);[Li et al\., 2024a](https://arxiv.org/html/2609.30474#bib.bib34)\)\.[Sun et al\. \(2023\)](https://arxiv.org/html/2609.30474#bib.bib28)proposedSpecTr, using optimal transport to verify multiple drafts\.[Hu et al\. \(2025\)](https://arxiv.org/html/2609.30474#bib.bib29)established that optimal multi\-draft speculative decoding \(MDSD\) reduces via total unimodularity to subset selection, introducingGreedy Draft Selectionas an efficient and theoretically grounded candidate selection strategy\. It has been observed that the performances of speculative decoding depend on many factors\([Liu et al\., 2026](https://arxiv.org/html/2609.30474#bib.bib11)\)\. In deep contrast with the work, essentially experimental, that flourished after the seminal work of[Leviathan et al\. \(2023\)](https://arxiv.org/html/2609.30474#bib.bib2);[Chen et al\. \(2023\)](https://arxiv.org/html/2609.30474#bib.bib1), the theory side of speculative decoding has remained in close contact with the seminal work, with essentially one exception digging in the expected number of tokens successfully predicted\([Pankratov and Alistarh, 2026](https://arxiv.org/html/2609.30474#bib.bib9)\)\.
Because lossless speculative decoding strictly preserves the target distribution, its acceptance rate is fundamentally bounded by the divergence between drafter and target\. To further increase throughput, several works have explored relaxing this exact\-matching constraint towardslossy speculative decoding\.[Xia et al\. \(2026\)](https://arxiv.org/html/2609.30474#bib.bib49)provide an empirical benchmark of most of these lossy decoding strategies\. Many approaches are fundamentally heuristic in nature\. In their seminal paper,[Leviathan et al\. \(2023\)](https://arxiv.org/html/2609.30474#bib.bib2)introducedLenient Speculative Decoding, making the per\-token acceptance probability dependent on a factor that skews it\. Subsequent work proposed various heuristic acceptance criteria:Typical Acceptance Sampling\([Cai et al\., 2024](https://arxiv.org/html/2609.30474#bib.bib32)\)accepts candidate tokens based on entropy heuristics;Fuzzy Speculative Decoding\([Holsman et al\., 2025](https://arxiv.org/html/2609.30474#bib.bib50)\)unconditionally accepts draft tokens whenever the step\-level divergence between drafter and target falls below a scalar thresholdTT;Speculative Contrastive Decoding\([Yuan et al\., 2024](https://arxiv.org/html/2609.30474#bib.bib46)\)incorporates contrastive penalties to steer generation away from draft errors; and[Narasimhan et al\. \(2025\)](https://arxiv.org/html/2609.30474#bib.bib47)adapt speculative verification to model cascading deferral rules\. Other work include using big models or more than two models\([Byun et al\., 2025](https://arxiv.org/html/2609.30474#bib.bib5);[Li et al\., 2026a](https://arxiv.org/html/2609.30474#bib.bib4)\), adding a linear head on top of the target called ajudge– being another example of last layer retraining –\([Bachmann et al\., 2025](https://arxiv.org/html/2609.30474#bib.bib7)\), completing the process with information from prefill\([Wang et al\., 2025a](https://arxiv.org/html/2609.30474#bib.bib6)\), completing the process with guessing appropriate draft length\([Zhang et al\., 2025](https://arxiv.org/html/2609.30474#bib.bib8)\), etc\.\([Holsman et al\., 2025](https://arxiv.org/html/2609.30474#bib.bib50)\)\.
The first workformalizingthe problem of mentored decoding as a constrained optimization problem maximizing draft acceptance subject to bounded \(reverse Kullback\-Leibler\) divergence is[Tran\-Thien \(2023\)](https://arxiv.org/html/2609.30474#bib.bib45)\.[Yin et al\. \(2024\)](https://arxiv.org/html/2609.30474#bib.bib27)analyzed the problem under Total Variation distance, characterizing the linear Pareto frontier\. Inspired by this result,DIVERSED\([Wang et al\., 2025b](https://arxiv.org/html/2609.30474#bib.bib51)\)introduced dynamic ensemble verification by sampling from a convex combination between drafter and target, like[Zhong et al\. \(2025\)](https://arxiv.org/html/2609.30474#bib.bib10)\. Under forward Kullback\-Leibler divergence,Cactus\([Hao and Mou, 2026](https://arxiv.org/html/2609.30474#bib.bib48)\)optimizes candidate acceptance via a second\-order Taylor approximation on the sampled token’s coordinate, though without globally controlling the divergence of the resulting joint output distribution\.
Finally, the papers observing that allowing some drift can beat the target’s own metrics are experimental\([Qin et al\., 2025](https://arxiv.org/html/2609.30474#bib.bib53);[Li et al\., 2026a](https://arxiv.org/html/2609.30474#bib.bib4);[Zhong et al\., 2025](https://arxiv.org/html/2609.30474#bib.bib10)\), as to our knowledge there is no formal work on the subject\.
## 3Mentored decoding
We first provide some definitions needed for this Section\.nnis the vocabulary size,Δn\\Delta\_\{n\}is thenn\-probability simplex,\[n\]=\.\{1,2,…,n\}\[n\]\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\\{1,2,\.\.\.,n\\\}\. Bold faces like𝒛\\bm\{z\}denote vectors, and their coordinates are denoted likeziz\_\{i\}\. Binary relations between vectors of the same dimension are coordinate\-wise:𝒂≤𝒃\\bm\{a\}\\leq\\bm\{b\}meansai≤bi,∀i∈\[n\]a\_\{i\}\\leq b\_\{i\},\\forall i\\in\[n\]\. For any𝒂,𝒃∈ℝn\\bm\{a\},\\bm\{b\}\\in\\mathbb\{R\}^\{n\}such that𝒂≤𝒃\\bm\{a\}\\leq\\bm\{b\}, we let\[𝒂,𝒃\]=\.∏i∈\[n\]\[ai,bi\]\[\\bm\{a\},\\bm\{b\}\]\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\prod\_\{i\\in\[n\]\}\[a\_\{i\},b\_\{i\}\]\. A prompt to the drafter and target models yields two distributions𝒑∈Δn\\bm\{p\}\\in\\Delta\_\{n\}\(drafter\) and𝒒∈Δn\\bm\{q\}\\in\\Delta\_\{n\}\(target\)\. Speculative and mentored decoding operate by generating multiple draft outputs and checking acceptance in parallel with the target\. Checking a token is probabilistic and relies on a vector of acceptance probabilities𝒓∈\[0,1\]n\\bm\{r\}\\in\[0,1\]^\{n\}; if rejected, a resampling distribution𝒔∈Δn\\bm\{s\}\\in\\Delta\_\{n\}resamples a new token\. Then, the same algorithm resumes until complete sequence generation\. We refer e\.g\. to[Leviathan et al\. \(2023\)](https://arxiv.org/html/2609.30474#bib.bib2);[Tran\-Thien \(2023\)](https://arxiv.org/html/2609.30474#bib.bib45)for more details on the algorithmic side\. We also define theff\-divergences between𝝅∈Δn\\bm\{\\pi\}\\in\\Delta\_\{n\}and𝒒∈Δn\\bm\{q\}\\in\\Delta\_\{n\}as\([Ali and Silvey, 1966](https://arxiv.org/html/2609.30474#bib.bib20);[Csiszár, 1963](https://arxiv.org/html/2609.30474#bib.bib21)\):
Df\(𝝅∥𝒒\)\\displaystyle D\_\{f\}\(\\bm\{\\pi\}\\\|\\bm\{q\}\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}∑iqif\(πiqi\),\\displaystyle\\sum\_\{i\}q\_\{i\}f\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\),\(1\)where thegenerator
f:ℝ\+→ℝis convex and such thatf\(1\)=0\.\\displaystyle f:\\mathbb\{R\}^\{\+\}\\rightarrow\\mathbb\{R\}\\mbox\{ is convex and such that \}f\(1\)=0\.\(2\)
### 3\.1One problem, two parameterizations
Without further ado, we define the core inference problem on which we focus\.
###### Definition 3\.1\.
For anyffas per \([2](https://arxiv.org/html/2609.30474#S3.E2)\),𝐩,𝐪∈Δn,D≥0\\bm\{p\},\\bm\{q\}\\in\\Delta\_\{n\},D\\geq 0, theff\-mentored decoding \(MD\) problem is defined as find
mdf2\(𝒑,𝒒;D\)=\.argmin𝒓∈\[0,1\]n,𝒔∈Δn−𝒑⊤𝒓s\.t\.Df\(𝒑⊙𝒓\+\(1−𝒑⊤𝒓\)⋅𝒔∥𝒒\)≤D,\\displaystyle\\textsc\{md\}^\{2\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\arg\\min\_\{\\bm\{r\}\\in\[0,1\]^\{n\},\\bm\{s\}\\in\\Delta\_\{n\}\}\-\\bm\{p\}^\{\\top\}\\bm\{r\}\\quad\\mbox\{s\.t\. \}D\_\{f\}\(\\bm\{p\}\\odot\\bm\{r\}\+\(1\-\\bm\{p\}^\{\\top\}\\bm\{r\}\)\\cdot\\bm\{s\}\\\|\\bm\{q\}\)\\leq D,\(ff\-MD\-2\)where⊙\\odotis Hadamard product\.
This problem was introduced by[Tran\-Thien \(2023\)](https://arxiv.org/html/2609.30474#bib.bib45)with the specific choicefrKL\(z\)=\.−ln\(z\)f\_\{\\mathrm\{rKL\}\}\(z\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\-\\ln\(z\), the reverse\-KL divergence\. Note that ifD=0D=0, the problem formalizes speculative decoding \(SD\)\. \([f \-MD\-2](https://arxiv.org/html/2609.30474#S3.Ex1)\) is a direct parameterization of MD: we directly seek the couple of vector of acceptance probabilities𝒓\\bm\{r\}and resampling distribution𝒔\\bm\{s\}\. A convenient result that we now state and prove is that this problem admits an equivalent parameterization with a single parameter, the mentored distribution itself\. Let us define it:
mdf1\(𝒑,𝒒;D\)=\.argmin𝝅∈ΔnDTV\(𝝅∥𝒑\)s\.t\.Df\(𝝅∥𝒒\)≤D,\\displaystyle\\textsc\{md\}^\{1\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\arg\\min\_\{\\bm\{\\pi\}\\in\\Delta\_\{n\}\}D\_\{\\mathrm\{TV\}\}\(\\bm\{\\pi\}\\\|\\bm\{p\}\)\\quad\\mbox\{s\.t\. \}D\_\{f\}\(\\bm\{\\pi\}\\\|\\bm\{q\}\)\\leq D,\(ff\-MD\-1\)whereDTVD\_\{\\mathrm\{TV\}\}denotes the total variation divergence, whose generator isfTV\(z\)=\.\|z−1\|/2f\_\{\\mathrm\{TV\}\}\(z\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\|z\-1\|/2\.
###### Lemma 3\.2\.
Let⊘\\oslashdenote the coordinate\-wise division\. For any𝛑∈mdf1\(𝐩,𝐪,D\)\\bm\{\\pi\}\\in\\textsc\{md\}^\{1\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\), we have\(𝐫,𝐬\)∈mdf2\(𝐩,𝐪,D\)\(\\bm\{r\},\\bm\{s\}\)\\in\\textsc\{md\}^\{2\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)where𝐫,𝐬\\bm\{r\},\\bm\{s\}are defined as:
𝒓=\.min\{𝟏,𝝅⊘𝒑\};𝒔=\.\{11−𝒑⊤𝒓⋅\(𝝅−𝒑⊙𝒓\)=11−𝒑⊤𝒓⋅max\{𝟎,𝝅−𝒑\}if𝒑⊤𝒓<1any𝒔∈Δnif𝒑⊤𝒓=1\.\\displaystyle\\bm\{r\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\min\\\{\\bm\{1\},\\bm\{\\pi\}\\oslash\\bm\{p\}\\\};\\quad\\bm\{s\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\left\\\{\\begin\{array\}\[\]\{ccl\}\\frac\{1\}\{1\-\\bm\{p\}^\{\\top\}\\bm\{r\}\}\\cdot\\left\(\\bm\{\\pi\}\-\\bm\{p\}\\odot\\bm\{r\}\\right\)=\\frac\{1\}\{1\-\\bm\{p\}^\{\\top\}\\bm\{r\}\}\\cdot\\max\\\{\\bm\{0\},\\bm\{\\pi\}\-\\bm\{p\}\\\}&\\mbox\{ if \}&\\bm\{p\}^\{\\top\}\\bm\{r\}<1\\\\ \\mbox\{any \}\\bm\{s\}\\in\\Delta\_\{n\}&\\mbox\{ if \}&\\bm\{p\}^\{\\top\}\\bm\{r\}=1\\end\{array\}\\right\.\.Respectively, for any\(𝐫,𝐬\)∈mdf2\(𝐩,𝐪,D\)\(\\bm\{r\},\\bm\{s\}\)\\in\\textsc\{md\}^\{2\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\), we have𝛑∈mdf1\(𝐩,𝐪,D\)\\bm\{\\pi\}\\in\\textsc\{md\}^\{1\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)with
𝝅\\displaystyle\\bm\{\\pi\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}𝒑⊙𝒓\+\(1−𝒑⊤𝒓\)⋅𝒔\.\\displaystyle\\bm\{p\}\\odot\\bm\{r\}\+\(1\-\\bm\{p\}^\{\\top\}\\bm\{r\}\)\\cdot\\bm\{s\}\.\(6\)Finally, the corresponding objective functions are related by𝐩⊤𝐫=1−DTV\(𝛑∥𝐩\)\\bm\{p\}^\{\\top\}\\bm\{r\}=1\-D\_\{\\mathrm\{TV\}\}\(\\bm\{\\pi\}\\\|\\bm\{p\}\)\.
Proof in Appendix, Section[VIII\.1](https://arxiv.org/html/2609.30474#S8.SS1)\. Lemma[3\.2](https://arxiv.org/html/2609.30474#S3.Thmtheorem2)allows us to work with whichever parameterization fits best to context; we denoteff\-MD as the general problem offfmentored decoding, with whichever parameterization\. The overall probability to accept a token in the MD setting is defined as
Pacc\(MD\)\\displaystyle P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(MD\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}𝒑⊤𝒓\.\\displaystyle\\bm\{p\}^\{\\top\}\\bm\{r\}\.The expression is the same for SD, only in this case we would have the "hidden" constraint boundDDto be zero\. We make the following assumptions regarding mentored decoding\.
###### Assumption 3\.3\.
We assume0<D<DTV\(𝐩∥𝐪\)0<D<D\_\{\\mathrm\{TV\}\}\(\\bm\{p\}\\\|\\bm\{q\}\),𝐩≠𝐪\\bm\{p\}\\neq\\bm\{q\},𝐩\>𝟎\\bm\{p\}\>\\bm\{0\}and𝐪\>𝟎\\bm\{q\}\>\\bm\{0\}\.
Note the weakness of those statements: if it does not hold that0<D<DTV\(𝒑∥𝒒\)0<D<D\_\{\\mathrm\{TV\}\}\(\\bm\{p\}\\\|\\bm\{q\}\), the problem is trivial \(DDbeing non\-negative, either𝝅=𝒑\\bm\{\\pi\}=\\bm\{p\}or𝝅=𝒒\\bm\{\\pi\}=\\bm\{q\}is optimal\); similarly ifpi=qip\_\{i\}=q\_\{i\}for somei∈\[n\]i\\in\[n\]then any optimum trivially meetsπi=pi=qi\\pi\_\{i\}=p\_\{i\}=q\_\{i\}so the coordinate can be dropped from solving \([f \-MD\-1](https://arxiv.org/html/2609.30474#S3.Ex2)\), notwithstanding a replacement of the unit mass constraint of the probability simplex by a1−qi1\-q\_\{i\}mass constraint\. Finally, the assumptions𝒑\>𝟎\\bm\{p\}\>\\bm\{0\}and𝒒\>𝟎\\bm\{q\}\>\\bm\{0\}are reasonable for LLMs, or any neural net architecture in which the last layer is a softmax\. Finally, note that we should theoretically add the technical assumption thatffbe proper in \([2](https://arxiv.org/html/2609.30474#S3.E2)\), but it is in fact always met for anyffrelevant to our context, and even more given Assumption[3\.3](https://arxiv.org/html/2609.30474#S3.Thmtheorem3)with which we can always restrict all of our analysis on a closed interval of the real line\.

Figure 1:Example of optimal solutions for the TV\-MD problem withn=3n=3tokens inΔ3\\Delta\_\{3\}, as intersection \(thick segment\) between the twoL1L\_\{1\}balls centered in𝒑\\bm\{p\}and𝒒\\bm\{q\}\(in colorredandblue\) whose radii defineDDand the optimal objective value for the total variation divergences in \([f \-MD\-1](https://arxiv.org/html/2609.30474#S3.Ex2)\)\. The shaded area is set\[min\{𝒑,𝒒\},max\{𝒑,𝒒\}\]∩Δ3\\left\[\\min\\\{\\bm\{p\},\\bm\{q\}\\\},\\max\\\{\\bm\{p\},\\bm\{q\}\\\}\\right\]\\cap\\Delta\_\{3\}\(see text\)\.
### 3\.2Warmup: the special case of the total variation divergence
The casef=fTVf=f\_\{\\mathrm\{TV\}\}inff\-MD is especially interesting: its set of optimal solutions has a beautiful geometric characterization and its proof is a few liner\. We denote it as TV\-MD\.
###### Theorem 3\.4\.
Under Assumption[3\.3](https://arxiv.org/html/2609.30474#S3.Thmtheorem3), we have
mdfTV1\(𝒑,𝒒,D\)\\displaystyle\\textsc\{md\}^\{1\}\_\{f\_\{\\mathrm\{TV\}\}\}\(\\bm\{p\},\\bm\{q\};D\)=\\displaystyle=Δn∩\{DTV\(\.∥𝒒\)=D\}∩\[min\{𝒑,𝒒\},max\{𝒑,𝒒\}\]\.\\displaystyle\\Delta\_\{n\}\\cap\\\{D\_\{\\mathrm\{TV\}\}\(\.\\\|\\bm\{q\}\)=D\\\}\\cap\\left\[\\min\\\{\\bm\{p\},\\bm\{q\}\\\},\\max\\\{\\bm\{p\},\\bm\{q\}\\\}\\right\]\.\(7\)
###### Proof\.
The TV divergence satisfies the triangle inequality, hence
DTV\(𝒒∥𝒑\)\\displaystyle D\_\{\\mathrm\{TV\}\}\(\\bm\{q\}\\\|\\bm\{p\}\)≤\\displaystyle\\leqDTV\(𝝅∥𝒑\)\+DTV\(𝝅∥𝒒\)\.\\displaystyle D\_\{\\mathrm\{TV\}\}\(\\bm\{\\pi\}\\\|\\bm\{p\}\)\+D\_\{\\mathrm\{TV\}\}\(\\bm\{\\pi\}\\\|\\bm\{q\}\)\.\(8\)The LHS is fixed and the objective to minimize isDTV\(𝝅∥𝒑\)D\_\{\\mathrm\{TV\}\}\(\\bm\{\\pi\}\\\|\\bm\{p\}\)\. Any𝝅∈Δn\\bm\{\\pi\}\\in\\Delta\_\{n\}can be formulated asπi=αipi\+\(1−αi\)qi,αi∈ℝ,∀i∈\[n\]\\pi\_\{i\}=\\alpha\_\{i\}p\_\{i\}\+\(1\-\\alpha\_\{i\}\)q\_\{i\},\\alpha\_\{i\}\\in\\mathbb\{R\},\\forall i\\in\[n\]since𝒑≠𝒒\\bm\{p\}\\neq\\bm\{q\}, which yields after simplification for the RHS of \([8](https://arxiv.org/html/2609.30474#S3.E8)\):
DTV\(𝝅∥𝒑\)\+DTV\(𝝅∥𝒒\)\\displaystyle D\_\{\\mathrm\{TV\}\}\(\\bm\{\\pi\}\\\|\\bm\{p\}\)\+D\_\{\\mathrm\{TV\}\}\(\\bm\{\\pi\}\\\|\\bm\{q\}\)=\\displaystyle=∑i∈\[n\]\(\|1−αi\|\+\|αi\|\)⋅\|pi−qi\|\.\\displaystyle\\sum\_\{i\\in\[n\]\}\(\|1\-\\alpha\_\{i\}\|\+\|\\alpha\_\{i\}\|\)\\cdot\|p\_\{i\}\-q\_\{i\}\|\.\(9\)Chooseαi=α∈\[0,1\],∀i∈\[n\]\\alpha\_\{i\}=\\alpha\\in\[0,1\],\\forall i\\in\[n\]: the RHS equalsDTV\(𝒒∥𝒑\)D\_\{\\mathrm\{TV\}\}\(\\bm\{q\}\\\|\\bm\{p\}\)and so \([8](https://arxiv.org/html/2609.30474#S3.E8)\) becomes an equality: this𝝅\\bm\{\\pi\}is optimal for TV\-MD if it maximizesDTV\(𝝅∥𝒒\)D\_\{\\mathrm\{TV\}\}\(\\bm\{\\pi\}\\\|\\bm\{q\}\)under the constraint, i\.e\. it makes it active asDTV\(𝝅∥𝒒\)=DD\_\{\\mathrm\{TV\}\}\(\\bm\{\\pi\}\\\|\\bm\{q\}\)=D, which happens for the choiceα=D/DTV\(𝒒∥𝒑\)\\alpha=D/D\_\{\\mathrm\{TV\}\}\(\\bm\{q\}\\\|\\bm\{p\}\)\. To be optimal thus generally requires to keep \(i\) the equality in \([8](https://arxiv.org/html/2609.30474#S3.E8)\) and \(ii\) the constraint active\. Sincef\(z\)=\.\|1−z\|\+\|z\|f\(z\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\|1\-z\|\+\|z\|satisfiesf\(\[0,1\]\)=\{1\}f\(\[0,1\]\)=\\\{1\\\}and is\>1\>1otherwise, optimality requiresαi∈\[0,1\],∀i∈\[n\]\\alpha\_\{i\}\\in\[0,1\],\\forall i\\in\[n\]in \([9](https://arxiv.org/html/2609.30474#S3.E9)\) and thusπi∈\[min\{pi,qi\},max\{pi,qi\}\],∀i∈\[n\]\\pi\_\{i\}\\in\[\\min\\\{p\_\{i\},q\_\{i\}\\\},\\max\\\{p\_\{i\},q\_\{i\}\\\}\],\\forall i\\in\[n\]\. The set of optimal solutions is thus\{𝝅∈\[min\{𝒑,𝒒\},max\{𝒑,𝒒\}\]∩Δn:DTV\(𝝅∥𝒒\)=D\}\\\{\\bm\{\\pi\}\\in\\left\[\\min\\\{\\bm\{p\},\\bm\{q\}\\\},\\max\\\{\\bm\{p\},\\bm\{q\}\\\}\\right\]\\cap\\Delta\_\{n\}:D\_\{\\mathrm\{TV\}\}\(\\bm\{\\pi\}\\\|\\bm\{q\}\)=D\\\}, as claimed\. ∎
It is worth mentioning that[Yin et al\. \(2024\)](https://arxiv.org/html/2609.30474#bib.bib27)analyzed the TV relaxation of SD\. While they provided a thorough analysis of the optimallosses, they did not provide the analytic form of the optimum, which has neat properties\. Indeed, the TV divergence constraints defineL1L\_\{1\}balls\. The optimal solution of \([7](https://arxiv.org/html/2609.30474#S3.E7)\) is the intersection between the simplex and two tangent balls, one whose radius depends on the optimal objective and one whose radius is parameterDD\. Figure[1](https://arxiv.org/html/2609.30474#S3.F1)exemplifies three such cases \(See also Figure[4](https://arxiv.org/html/2609.30474#S5.F4)for otherff\-MD cases\)\. The optimal set being this "big" and "nice" naturally opens the question as to whether such optimal solutions that speed up inference might in fact be grounded in a model producing𝝅\\bm\{\\pi\}that, since it is a function of the target and drafter models, couldcompete with or beatthe target model in terms ofquality\. We shall indeed give a formal positive answer in Section[4](https://arxiv.org/html/2609.30474#S4), but before, we address and solveff\-MD for a generalff\([2](https://arxiv.org/html/2609.30474#S3.E2)\)\.
### 3\.3ff\-mentored decoding: general solution

Figure 2:Illustration of the definition ofclampset\(z,𝔸,𝔹\)\\mathrm\{clampset\}\(z,\\mathbb\{A\},\\mathbb\{B\}\)\([10](https://arxiv.org/html/2609.30474#S3.E10)\) \(inred\) when𝔸<𝔹\\mathbb\{A\}<\\mathbb\{B\}are intervals \(figured by rectangles\) andzz\(thick vertical bar\) moves along the real line \(see text\)\.To step up to the general case, we need additional definitions\. Binary relations defined over sets of reals are true iff they hold for any applicable elements: for example,𝔸≤𝔹\\mathbb\{A\}\\leq\\mathbb\{B\}it true iffa≤b,∀a∈𝔸,b∈𝔹a\\leq b,\\forall a\\in\\mathbb\{A\},b\\in\\mathbb\{B\}\. For anyz∈ℝ,𝔸⊆ℝz\\in\\mathbb\{R\},\\mathbb\{A\}\\subseteq\\mathbb\{R\}, we let:
minset\(z,𝔸\)\\displaystyle\\mathrm\{minset\}\(z,\\mathbb\{A\}\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\{s∈𝔸:s≤z\},\\displaystyle\\\{s\\in\\mathbb\{A\}:s\\leq z\\\},maxset\(z,𝔸\)\\displaystyle\\mathrm\{maxset\}\(z,\\mathbb\{A\}\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\{s∈𝔸:s≥z\},\\displaystyle\\\{s\\in\\mathbb\{A\}:s\\geq z\\\},and for any𝔹⊂ℝ\\mathbb\{B\}\\subset\\mathbb\{R\}such that𝔸<𝔹\\mathbb\{A\}<\\mathbb\{B\},
clampset\(z,𝔸,𝔹\)=\.minset\(0,𝔹−z\)\+maxset\(0,𝔸−z\)\+⟦inf𝔸≤z≤sup𝔹⟧⋅\{z\},\\displaystyle\\mathrm\{clampset\}\(z,\\mathbb\{A\},\\mathbb\{B\}\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\mathrm\{minset\}\(0,\\mathbb\{B\}\-z\)\+\\mathrm\{maxset\}\(0,\\mathbb\{A\}\-z\)\+\\llbracket\\inf\\mathbb\{A\}\\leq z\\leq\\sup\\mathbb\{B\}\\rrbracket\\cdot\\\{z\\\},\(10\)where⟦\.⟧\\llbracket\.\\rrbracketdenotes Iverson’s bracket\([Knuth, 1992](https://arxiv.org/html/2609.30474#bib.bib13)\)\. Some properties are notable\.
###### Lemma 3\.5\.
For anyz∈ℝ,𝔸<𝔹⊂ℝz\\in\\mathbb\{R\},\\mathbb\{A\}<\\mathbb\{B\}\\subset\\mathbb\{R\},
\{min\{z,z′\}:z′∈clampset\(z,𝔸,𝔹\)\}\\displaystyle\\\{\\min\\\{z,z^\{\\prime\}\\\}:z^\{\\prime\}\\in\\mathrm\{clampset\}\(z,\\mathbb\{A\},\\mathbb\{B\}\)\\\}=\\displaystyle=\{z\}\+minset\(0,𝔹−z\),\\displaystyle\\\{z\\\}\+\\mathrm\{minset\}\(0,\\mathbb\{B\}\-z\),\(11\)\{max\{0,z′−z\}:z′∈clampset\(z,𝔸,𝔹\)\}\\displaystyle\\\{\\max\\\{0,z^\{\\prime\}\-z\\\}:z^\{\\prime\}\\in\\mathrm\{clampset\}\(z,\\mathbb\{A\},\\mathbb\{B\}\)\\\}=\\displaystyle=maxset\(0,𝔸−z\)\.\\displaystyle\\mathrm\{maxset\}\(0,\\mathbb\{A\}\-z\)\.\(12\)
###### Proof\.
In \([11](https://arxiv.org/html/2609.30474#S3.E11)\) the LHS iszzunless\{s∈𝔹:s≤z\}≠∅\\\{s\\in\\mathbb\{B\}:s\\leq z\\\}\\neq\\emptyset, in which case it is this set, i\.e\.minset\(z,𝔹\)\\mathrm\{minset\}\(z,\\mathbb\{B\}\)\. In \([12](https://arxiv.org/html/2609.30474#S3.E12)\) the LHS is00unless\{s∈𝔸:s−z≥0\}≠∅\\\{s\\in\\mathbb\{A\}:s\-z\\geq 0\\\}\\neq\\emptyset, in which case it is this set, i\.e\.maxset\(0,𝔸−z\)\\mathrm\{maxset\}\(0,\\mathbb\{A\}\-z\)\. ∎
#### Analysis forffconvex differentiable
we start with the case whereffis differentiable in \([2](https://arxiv.org/html/2609.30474#S3.E2)\), and later relax differentiability\.
###### Lemma 3\.6\.
Assumption[3\.3](https://arxiv.org/html/2609.30474#S3.Thmtheorem3)implies Slater’s constraint qualification satisfied for \([f \-MD\-1](https://arxiv.org/html/2609.30474#S3.Ex2)\) and \([f \-MD\-2](https://arxiv.org/html/2609.30474#S3.Ex1)\)\.
Proof in Appendix, Section[VIII\.2](https://arxiv.org/html/2609.30474#S8.SS2)\. So KKT conditions are necessary and sufficient for optimality in the study offf\-MD\. We now make a connection betweenmdf2\(𝒑,𝒒,D\)\\textsc\{md\}^\{2\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)and another problem, whose objective is the following one:
isf\(𝒑,𝒒;D\)=\.\{𝝅∈Δn:\{∃α<β∈Im\(−f′\):𝝅∈clampset\(𝒑,Lβ\(−f′\)⋅𝒒,Lα\(−f′\)⋅𝒒\)Df\(𝝅∥𝒒\)=D\},\\displaystyle\\textsc\{is\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\left\\\{\\bm\{\\pi\}\\in\\Delta\_\{n\}:\\left\\\{\\begin\{array\}\[\]\{l\}\\exists\\alpha<\\beta\\in\\mathrm\{Im\}\(\-f^\{\\prime\}\):\\bm\{\\pi\}\\in\\mathrm\{clampset\}\(\\bm\{p\},L\_\{\\beta\}\(\-f^\{\\prime\}\)\\cdot\\bm\{q\},L\_\{\\alpha\}\(\-f^\{\\prime\}\)\\cdot\\bm\{q\}\)\\\\ D\_\{f\}\(\\bm\{\\pi\}\\\|\\bm\{q\}\)=D\\end\{array\}\\right\.\\right\\\},whereLy\(g\)=\.\{z:g\(z\)=y\}L\_\{y\}\(g\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\\{z:g\(z\)=y\\\}denotes theyy\-level set of functiongg\. As we shall explain later when relaxing the differentiability assumption onff, the set described in \([3\.3](https://arxiv.org/html/2609.30474#S3.EGx14)\) is the same as in \([7](https://arxiv.org/html/2609.30474#S3.E7)\) when properly relaxing the derivative to the subdifferential\. In such a context, looking at Figure[1](https://arxiv.org/html/2609.30474#S3.F1)for an example, let us keep in mind for now that \([3\.3](https://arxiv.org/html/2609.30474#S3.EGx14)\) elicits isodivergence sets on the probability simplex which, just like \([7](https://arxiv.org/html/2609.30474#S3.E7)\) does for the case of the TV divergence, denote the set of optimal solutions that we seek forff\-MD\.
Without further ado, we state and prove the main Theorem that elicits the connections between \([f \-MD\-2](https://arxiv.org/html/2609.30474#S3.Ex1)\) and \([3\.3](https://arxiv.org/html/2609.30474#S3.EGx14)\)\.
###### Theorem 3\.7\.
Under Assumption[3\.3](https://arxiv.org/html/2609.30474#S3.Thmtheorem3), there exists a bijection betweenmdf2\(𝐩,𝐪,D\)\\textsc\{md\}^\{2\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)andisf\(𝐩,𝐪,D\)\\textsc\{is\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)\. Specifically,
- \(mdf2→isf\\textsc\{md\}^\{2\}\_\{f\}\\rightarrow\\textsc\{is\}\_\{f\}\)for any couple\(𝒓,𝒔\)∈mdf2\(𝒑,𝒒,D\)\(\\bm\{r\},\\bm\{s\}\)\\in\\textsc\{md\}^\{2\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\), we have𝝅=\.𝒑⊙𝒓\+\(1−𝒑⊤𝒓\)𝒔∈isf\(𝒑,𝒒,D\)\\bm\{\\pi\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\bm\{p\}\\odot\\bm\{r\}\+\(1\-\\bm\{p\}^\{\\top\}\\bm\{r\}\)\\bm\{s\}\\in\\textsc\{is\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)for the choices α\\displaystyle\\alpha=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}𝔼i∼𝒖\[\(−f′\)\(πiqi\)\],with𝒖=\.11−𝒑⊤𝒓⋅\(𝒑−𝒑⊙𝒓\)∈Δn\.\\displaystyle\\mathbb\{E\}\_\{i\\sim\\bm\{u\}\}\\left\[\(\-f^\{\\prime\}\)\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)\\right\],\\quad\\mbox\{ with \}\\bm\{u\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\frac\{1\}\{1\-\\bm\{p\}^\{\\top\}\\bm\{r\}\}\\cdot\(\\bm\{p\}\-\\bm\{p\}\\odot\\bm\{r\}\)\\in\\Delta\_\{n\}\.\(16\)β\\displaystyle\\beta=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}𝔼i∼𝒔\[\(−f′\)\(πiqi\)\]\\displaystyle\\mathbb\{E\}\_\{i\\sim\\bm\{s\}\}\\left\[\(\-f^\{\\prime\}\)\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)\\right\]\(17\)
- \(isf→mdf2\\textsc\{is\}\_\{f\}\\rightarrow\\textsc\{md\}^\{2\}\_\{f\}\)for any𝝅∈isf\(𝒑,𝒒,D\)\\bm\{\\pi\}\\in\\textsc\{is\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\), the couple\(𝒓,𝒔\)\(\\bm\{r\},\\bm\{s\}\)defined in \([3\.2](https://arxiv.org/html/2609.30474#S3.EGx5)\) is inmdf2\(𝒑,𝒒,D\)\\textsc\{md\}^\{2\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)\.
Proof in Appendix, Section[VIII\.3](https://arxiv.org/html/2609.30474#S8.SS3)\. Using Lemma[3\.2](https://arxiv.org/html/2609.30474#S3.Thmtheorem2), we get as immediate corollary another characterization ofmdf1\(𝒑,𝒒,D\)\\textsc\{md\}^\{1\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\):
mdf1\(𝒑,𝒒,D\)\\displaystyle\\textsc\{md\}^\{1\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)=\\displaystyle=isf\(𝒑,𝒒,D\),\\displaystyle\\textsc\{is\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\),which is not unreminiscent of the case of the total variation in \([7](https://arxiv.org/html/2609.30474#S3.E7)\) \(more on this later\)\. Finally, we have the following Lemma stating some important properties ofLαL\_\{\\alpha\}andLβL\_\{\\beta\}in \(mdf2→isf\\textsc\{md\}^\{2\}\_\{f\}\\rightarrow\\textsc\{is\}\_\{f\}\), whose proof is given in the proof of Theorem[3\.7](https://arxiv.org/html/2609.30474#S3.Thmtheorem7)\.
###### Lemma 3\.8\.
α,β\\alpha,\\betain \([16](https://arxiv.org/html/2609.30474#S3.E16)\), \([17](https://arxiv.org/html/2609.30474#S3.E17)\) satisfy:
Lβ\(−f′\)<Lα\(−f′\)\\displaystyle L\_\{\\beta\}\(\-f^\{\\prime\}\)<L\_\{\\alpha\}\(\-f^\{\\prime\}\);minLβ\(−f′\)≤1≤maxLα\(−f′\),\\displaystyle\\min L\_\{\\beta\}\(\-f^\{\\prime\}\)\\leq 1\\leq\\max L\_\{\\alpha\}\(\-f^\{\\prime\}\),\(18\)minipi/qi<maxLβ\(−f′\)\\displaystyle\\min\_\{i\}p\_\{i\}/q\_\{i\}<\\max L\_\{\\beta\}\(\-f^\{\\prime\}\);minLα\(−f′\)<maxipi/qi\.\\displaystyle\\min L\_\{\\alpha\}\(\-f^\{\\prime\}\)<\\max\_\{i\}p\_\{i\}/q\_\{i\}\.\(19\)
Note that \([19](https://arxiv.org/html/2609.30474#S3.E19)\) follows from the fact thatα,β\\alpha,\\betaare expectations in \([16](https://arxiv.org/html/2609.30474#S3.E16)\), \([17](https://arxiv.org/html/2609.30474#S3.E17)\)\. An additional important result is the following one, which states that𝒓\\bm\{r\}has full support\.
###### Lemma 3\.9\.
Under Assumption[3\.3](https://arxiv.org/html/2609.30474#S3.Thmtheorem3), any couple\(𝐫,𝐬\)∈mdf2\(𝐩,𝐪,D\)\(\\bm\{r\},\\bm\{s\}\)\\in\\textsc\{md\}^\{2\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)satisfies𝐫\>𝟎\\bm\{r\}\>\\bm\{0\}\.
Proof in Appendix, Section[VIII\.4](https://arxiv.org/html/2609.30474#S8.SS4)\.
#### Mentored decoding: analysis for generalff\([2](https://arxiv.org/html/2609.30474#S3.E2)\)
A simple trick allows to alleviate the differentiability condition onffand prove the result for any convexff\([2](https://arxiv.org/html/2609.30474#S3.E2)\), and it proceeds from the simple example of how Theorem[3\.7](https://arxiv.org/html/2609.30474#S3.Thmtheorem7)also covers the case of TV, whose generator isf\(z\)=\.\|z−1\|f\(z\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\|z\-1\|, non differentiable only inz=1z=1\. We first smooth the generator in an openδ\\delta\-neighborhood of 1, eventually with ayy\-translation of the graph to keepf\(1\)=0f\(1\)=0\. We want to preventLα\(−f′\),Lβ\(−f′\)L\_\{\\alpha\}\(\-f^\{\\prime\}\),L\_\{\\beta\}\(\-f^\{\\prime\}\)to be picked from this neighborhood, so we are going to tuneδ\>0\\delta\>0\. IfLα\(−f′\)L\_\{\\alpha\}\(\-f^\{\\prime\}\)is in, asδ↘0\\delta\\searrow 0, the objective converges to that of speculative decoding, and ifLβ\(−f′\)L\_\{\\beta\}\(\-f^\{\\prime\}\)is in, asδ↘0\\delta\\searrow 0, theff\-divergence value goes to 0\. So we can pickδ\>0\\delta\>0small enough forLα\(−f′\)L\_\{\\alpha\}\(\-f^\{\\prime\}\)to be out of the neighborhood \(objective small enough\) withLβ\(−f′\)L\_\{\\beta\}\(\-f^\{\\prime\}\)out of the neighborhood \(acceptable divergence\)\.
We thus end up with only one possible solution,Lα\(−f′\)=L−1\(−f′\)=\[1\+δ,\+∞\)L\_\{\\alpha\}\(\-f^\{\\prime\}\)=L\_\{\-1\}\(\-f^\{\\prime\}\)=\[1\+\\delta,\+\\infty\)andLβ\(−f′\)=L1\(−f′\)=\(−∞,1−δ\]L\_\{\\beta\}\(\-f^\{\\prime\}\)=L\_\{1\}\(\-f^\{\\prime\}\)=\(\-\\infty,1\-\\delta\]and thus according to \([3\.3](https://arxiv.org/html/2609.30474#S3.EGx14)\) all solutions of \(is\) satisfy
𝝅\\displaystyle\\bm\{\\pi\}∈\\displaystyle\\inclampset\(𝒑,\(−∞,1−δ\]⋅𝒒,⋅\[1\+δ,\+∞\)⋅𝒒\)\.\\displaystyle\\mathrm\{clampset\}\(\\bm\{p\},\(\-\\infty,1\-\\delta\]\\cdot\\bm\{q\},\\cdot\[1\+\\delta,\+\\infty\)\\cdot\\bm\{q\}\)\.\(20\)Note that∀i∈\[n\]\\forall i\\in\[n\], we have
clampset\(pi,\(−∞,1−δ\]⋅qi,⋅\[1\+δ,\+∞\)⋅qi\)\\displaystyle\\mathrm\{clampset\}\(p\_\{i\},\(\-\\infty,1\-\\delta\]\\cdot q\_\{i\},\\cdot\[1\+\\delta,\+\\infty\)\\cdot q\_\{i\}\)=\\displaystyle=\{\[pi,\(1−δ\)qi\]ifpi<\(1−δ\)qi\[\(1\+δ\)qi,pi\]ifpi\>\(1\+δ\)qipiifpi∈\[\(1−δ\)qi,\(1\+δ\)qi\]\.\\displaystyle\\left\\\{\\begin\{array\}\[\]\{ccl\}\\left\[p\_\{i\},\(1\-\\delta\)q\_\{i\}\\right\]&\\mbox\{ if \}&p\_\{i\}<\(1\-\\delta\)q\_\{i\}\\\\ \\left\[\(1\+\\delta\)q\_\{i\},p\_\{i\}\\right\]&\\mbox\{ if \}&p\_\{i\}\>\(1\+\\delta\)q\_\{i\}\\\\ p\_\{i\}&\\mbox\{ if \}&p\_\{i\}\\in\[\(1\-\\delta\)q\_\{i\},\(1\+\\delta\)q\_\{i\}\]\\end\{array\}\\right\.\.Becausepi≠qi,∀ip\_\{i\}\\neq q\_\{i\},\\forall i\(Assumption[3\.3](https://arxiv.org/html/2609.30474#S3.Thmtheorem3)\), we can further chooseδ\>0\\delta\>0small enough so that we always have
pi\\displaystyle p\_\{i\}∉\\displaystyle\\not\\in\[\(1−δ\)qi,\(1\+δ\)qi\],∀i\.\\displaystyle\[\(1\-\\delta\)q\_\{i\},\(1\+\\delta\)q\_\{i\}\],\\forall i\.\(22\)The set \([20](https://arxiv.org/html/2609.30474#S3.E20)\) simplifies as:
clampset\(𝒑,\(−∞,1−δ\]⋅𝒒,⋅\[1\+δ,\+∞\)⋅𝒒\)\\displaystyle\\mathrm\{clampset\}\(\\bm\{p\},\(\-\\infty,1\-\\delta\]\\cdot\\bm\{q\},\\cdot\[1\+\\delta,\+\\infty\)\\cdot\\bm\{q\}\)=\\displaystyle=\[min\{𝒑,\(1\+δ\)⋅𝒒\},max\{𝒑,\(1−δ\)⋅𝒒\}\],\\displaystyle\\left\[\\min\\\{\\bm\{p\},\(1\+\\delta\)\\cdot\\bm\{q\}\\\},\\max\\\{\\bm\{p\},\(1\-\\delta\)\\cdot\\bm\{q\}\\\}\\right\],and none of the intervals is empty thanks to \([22](https://arxiv.org/html/2609.30474#S3.E22)\)\. We thus get
isf\(𝒑,𝒒,D\)\\displaystyle\\textsc\{is\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)=\\displaystyle=\{𝝅∈\[min\{𝒑,\(1\+δ\)⋅𝒒\},max\{𝒑,\(1−δ\)⋅𝒒\}\]∩Δn∧Df\(𝝅∥𝒒\)=D\}\.\\displaystyle\\left\\\{\\bm\{\\pi\}\\in\\left\[\\min\\\{\\bm\{p\},\(1\+\\delta\)\\cdot\\bm\{q\}\\\},\\max\\\{\\bm\{p\},\(1\-\\delta\)\\cdot\\bm\{q\}\\\}\\right\]\\cap\\Delta\_\{n\}\\wedge D\_\{f\}\(\\bm\{\\pi\}\\\|\\bm\{q\}\)=D\\right\\\}\.\(23\)This set converges toisTV\(𝒑,𝒒,D\)\\textsc\{is\}\_\{TV\}\(\\bm\{p\},\\bm\{q\};D\)asδ→0\\delta\\rightarrow 0andffconverges to the generator of TV in anyℓp\\ell\_\{p\}norm in the interval\[minipi/qi,maxipi/qi\]\[\\min\_\{i\}p\_\{i\}/q\_\{i\},\\max\_\{i\}p\_\{i\}/q\_\{i\}\], and we check that the solution found in \([23](https://arxiv.org/html/2609.30474#S3.E23)\) matches \([7](https://arxiv.org/html/2609.30474#S3.E7)\)\.
Now, any convex function defined on an open convex set is differentiable anywhere except maybe on a set of measure zero\([Rockafellar, 1970](https://arxiv.org/html/2609.30474#bib.bib24), Theorem 25\.5\), so for any point of non differentiability of a generalff, our analysis above also holds for a sufficiently smallδ\>0\\delta\>0, for which, after passing to the limit withδ→0\\delta\\rightarrow 0, we get the proof thatisf\(𝒑,𝒒,D\)\\textsc\{is\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)in \([3\.3](https://arxiv.org/html/2609.30474#S3.EGx14)\) generalizes to
isf\(𝒑,𝒒,D\)\\displaystyle\\textsc\{is\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)=\.\\displaystyle\\hskip\-19\.91684pt\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\{𝝅∈Δn:\{∃g,h∈∂f,∃α∈Im\(−g\),∃β∈Im\(−h\):\{α<β𝝅∈clampset\(𝒑,Lβ\(−h\)⋅𝒒,Lα\(−g\)⋅𝒒\)Df\(𝝅∥𝒒\)=D\},\\displaystyle\\hskip\-19\.91684pt\\left\\\{\\bm\{\\pi\}\\in\\Delta\_\{n\}:\\left\\\{\\begin\{array\}\[\]\{l\}\\hskip\-8\.5359pt\\begin\{array\}\[\]\{l\}\\exists g,h\\in\\partial f,\\\\ \\exists\\alpha\\in\\mathrm\{Im\}\(\-g\),\\exists\\beta\\in\\mathrm\{Im\}\(\-h\)\\end\{array\}\\hskip\-8\.5359pt:\\left\\\{\\begin\{array\}\[\]\{l\}\\alpha<\\beta\\\\ \\bm\{\\pi\}\\in\\mathrm\{clampset\}\(\\bm\{p\},L\_\{\\beta\}\(\-h\)\\cdot\\bm\{q\},L\_\{\\alpha\}\(\-g\)\\cdot\\bm\{q\}\)\\end\{array\}\\right\.\\\\ \\hskip\-5\.69046ptD\_\{f\}\(\\bm\{\\pi\}\\\|\\bm\{q\}\)=D\\end\{array\}\\right\.\\hskip\-14\.22636pt\\right\\\},where∂f\\partial fis the subdifferential offf; \([16](https://arxiv.org/html/2609.30474#S3.E16)\), \([17](https://arxiv.org/html/2609.30474#S3.E17)\) become, forg,h∈∂fg,h\\in\\partial f,
α=\.𝔼i∼𝒖\[\(−g\)\(πiqi\)\]\\displaystyle\\alpha\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\mathbb\{E\}\_\{i\\sim\\bm\{u\}\}\\left\[\(\-g\)\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)\\right\],β=\.𝔼i∼𝒔\[\(−h\)\(πiqi\)\],\\displaystyle\\beta\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\mathbb\{E\}\_\{i\\sim\\bm\{s\}\}\\left\[\(\-h\)\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)\\right\],\(31\)and𝒖\\bm\{u\}does not change\. Since the level sets of the subdifferential of a strictly convex function are singletons, we immediately get the following Corollary as a consequence of \([3\.3](https://arxiv.org/html/2609.30474#S3.EGx23)\) and the definition ofclampset\\mathrm\{clampset\}\.
###### Corollary 3\.10\.
Under Assumption[3\.3](https://arxiv.org/html/2609.30474#S3.Thmtheorem3), supposeffin \([2](https://arxiv.org/html/2609.30474#S3.E2)\) is strictly convex\. Thenisf\(𝐩,𝐪,D\)\\textsc\{is\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)is a singleton consisting of the unique𝛑∈Δn\\bm\{\\pi\}\\in\\Delta\_\{n\}satisfying \(i\)Df\(𝛑∥𝐪\)=DD\_\{f\}\(\\bm\{\\pi\}\\\|\\bm\{q\}\)=Dand \(ii\)
𝝅\\displaystyle\\bm\{\\pi\}=\\displaystyle=max\{\(1−b\)⋅𝒒,min\{𝒑,\(1\+a\)⋅𝒒\}\}\\displaystyle\\max\\\{\(1\-b\)\\cdot\\bm\{q\},\\min\\\{\\bm\{p\},\(1\+a\)\\cdot\\bm\{q\}\\\}\\\}\(32\)for some uniquea\>0,b∈\(0,1\]a\>0,b\\in\(0,1\]\.
As a consequence of Theorem[3\.7](https://arxiv.org/html/2609.30474#S3.Thmtheorem7)and Lemma[3\.2](https://arxiv.org/html/2609.30474#S3.Thmtheorem2),𝝅\\bm\{\\pi\}in \([32](https://arxiv.org/html/2609.30474#S3.E32)\) also satisfies𝝅∈mdf1\(𝒑,𝒒,D\)\\bm\{\\pi\}\\in\\textsc\{md\}^\{1\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)\. Corollary[3\.10](https://arxiv.org/html/2609.30474#S3.Thmtheorem10)shows that the optimal solution of mentored decoding has a remarkable simple form in most cases\. We now show that there is in fact much more to this simple form\.
#### Simple approximations toff\-MD
In our path to join the properties of mentored decoding and boosting, we need an intermediate result of independent interest\. For any𝔸⊆ℝ,z∈ℝ\\mathbb\{A\}\\subseteq\\mathbb\{R\},z\\in\\mathbb\{R\}, we letz⋅𝔸=\.\{zz′:z′∈𝔸\}z\\cdot\\mathbb\{A\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\\{zz^\{\\prime\}:z^\{\\prime\}\\in\\mathbb\{A\}\\\}\. For anya,ba,b, let
𝔸\\displaystyle\\mathbb\{A\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\{i:pi\>\(1\+a\)⋅qi\},\\displaystyle\\\{i:p\_\{i\}\>\(1\+a\)\\cdot q\_\{i\}\\\},\(33\)𝕀\\displaystyle\\mathbb\{I\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\{i:pi∈qi⋅\[1−b,1\+a\]\},\\displaystyle\\\{i:p\_\{i\}\\in q\_\{i\}\\cdot\[1\-b,1\+a\]\\\},\(34\)𝔹\\displaystyle\\mathbb\{B\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\{i:pi<\(1−b\)⋅qi\}\.\\displaystyle\\\{i:p\_\{i\}<\(1\-b\)\\cdot q\_\{i\}\\\}\.\(35\)The dependence of𝔸,𝔹\\mathbb\{A\},\\mathbb\{B\}and𝕀\\mathbb\{I\}ona,ba,bis implicit for the sake of readability\. We now define an important set of couples of reals
###### Definition 3\.11\.
For any𝐩,𝐪\\bm\{p\},\\bm\{q\}output to the drafter and target, respectively, let𝒞\(𝐩,𝐪\)\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\)be the set of couples\(a,b\)\(a,b\)satisfying:
a\\displaystyle a∈\\displaystyle\\in\[0,maxipiqi−1\),\\displaystyle\\left\[0,\\max\_\{i\}\\frac\{p\_\{i\}\}\{q\_\{i\}\}\-1\\right\),\(36\)b\\displaystyle b∈\\displaystyle\\in\[0,1−minipiqi\),\\displaystyle\\left\[0,1\-\\min\_\{i\}\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\right\),\(37\)−aq\(𝔸\)\+bq\(𝔹\)\\displaystyle\-aq\(\\mathbb\{A\}\)\+bq\(\\mathbb\{B\}\)=\\displaystyle=p\(𝕀\)−q\(𝕀\),\\displaystyle p\(\\mathbb\{I\}\)\-q\(\\mathbb\{I\}\),\(38\)where we have letp\(𝕄\)=\.0\+∑i∈𝕄pip\(\\mathbb\{M\}\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}0\+\\sum\_\{i\\in\\mathbb\{M\}\}p\_\{i\}for any𝕄⊆\[n\]\\mathbb\{M\}\\subseteq\[n\]\(and similarly,q\(𝕄\)=\.0\+∑i∈𝕄qiq\(\\mathbb\{M\}\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}0\+\\sum\_\{i\\in\\mathbb\{M\}\}q\_\{i\}\)\.
Set𝒞\(𝒑,𝒒\)\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\)has important properties, that we now state\.
###### Theorem 3\.12\.
∀\(a,b\)∈𝒞\(𝒑,𝒒\)\\forall\(a,b\)\\in\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\), the choice
𝒓\\displaystyle\\bm\{r\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}min\{𝟏,\(1\+a\)⋅𝒒⊘𝒑\},\\displaystyle\\min\\\{\\bm\{1\},\(1\+a\)\\cdot\\bm\{q\}\\oslash\\bm\{p\}\\\},\(39\)𝒔\\displaystyle\\bm\{s\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\(1−𝒑⊤𝒓\)−1⋅max\{𝟎,\(1−b\)⋅𝒒−𝒑\}\\displaystyle\(1\-\\bm\{p\}^\{\\top\}\\bm\{r\}\)^\{\-1\}\\cdot\\max\\\{\\bm\{0\},\(1\-b\)\\cdot\\bm\{q\}\-\\bm\{p\}\\\}\(40\)has the properties that the corresponding mentored distribution𝛑=𝐩clamped to\[1−b,1\+a\]⋅𝐪∈Δn\\bm\{\\pi\}=\\bm\{p\}\\mbox\{ clamped to \}\[1\-b,1\+a\]\\cdot\\bm\{q\}\\in\\Delta\_\{n\}\([6](https://arxiv.org/html/2609.30474#S3.E6)\) and:
1. \(I\)The corresponding acceptance probability of mentored decoding,Pacc\(MD\)P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(MD\), satisfies Pacc\(MD\)\\displaystyle P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(MD\)=\\displaystyle=Pacc\(SD\)\+aq\(𝔸\)\+\(p\(𝕀\>1\)−q\(𝕀\>1\)\),\\displaystyle P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)\+aq\(\\mathbb\{A\}\)\+\(p\(\\mathbb\{I\}\_\{\>1\}\)\-q\(\\mathbb\{I\}\_\{\>1\}\)\),\(41\)wherePacc\(SD\)P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)is the acceptance probabilities of speculative decoding and 𝕀\>1\\displaystyle\\mathbb\{I\}\_\{\>1\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\{i:pi∈qi⋅\(1,1\+a\]\}\(⊆𝕀, with the conventionp\(∅\)=q\(∅\)=\.0\)\.\\displaystyle\\left\\\{i:p\_\{i\}\\in q\_\{i\}\\cdot\(1,1\+a\]\\right\\\}\\quad\\mbox\{\($\\subseteq\\mathbb\{I\}$, with the convention $p\(\\emptyset\)=q\(\\emptyset\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}0$\)\}\.
2. \(II\)for anyffas per \([2](https://arxiv.org/html/2609.30474#S3.E2)\),∃u∈\[1−b,1\+a\]\\exists u\\in\[1\-b,1\+a\]such that the choice\(𝒓,𝒔\)\(\\bm\{r\},\\bm\{s\}\)in \([39](https://arxiv.org/html/2609.30474#S3.E39)\), \([40](https://arxiv.org/html/2609.30474#S3.E40)\) satisfies\(𝒓,𝒔\)∈mdf2\(𝒑,𝒒,D\)\(\\bm\{r\},\\bm\{s\}\)\\in\\textsc\{md\}^\{2\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)\([f \-MD\-2](https://arxiv.org/html/2609.30474#S3.Ex1)\) for D\\displaystyle D=\\displaystyle=q\(𝔸\)⋅f\(1\+a\)\+q\(𝔹\)⋅f\(1−b\)\+q\(𝕀\)⋅f\(u\)\.\\displaystyle q\(\\mathbb\{A\}\)\\cdot f\(1\+a\)\+q\(\\mathbb\{B\}\)\\cdot f\(1\-b\)\+q\(\\mathbb\{I\}\)\\cdot f\(u\)\.\(42\)
Proof in Appendix, Section[VIII\.5](https://arxiv.org/html/2609.30474#S8.SS5)\. We check that Theorem[3\.12](https://arxiv.org/html/2609.30474#S3.Thmtheorem12)is optimal in the sense thata,b→0a,b\\rightarrow 0, we have the convergenceD→f\(1\)=0D\\rightarrow f\(1\)=0andPacc\(MD\)→Pacc\(SD\)P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(MD\)\\rightarrow P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\), since\(𝒓,𝒔\)\(\\bm\{r\},\\bm\{s\}\)converges towards the solution of speculative decoding\. By definition,p\(𝕀\>1\)\>q\(𝕀\>1\)p\(\\mathbb\{I\}\_\{\>1\}\)\>q\(\\mathbb\{I\}\_\{\>1\}\)if𝕀\>1≠∅\\mathbb\{I\}\_\{\>1\}\\neq\\emptysetand obviouslyaq\(𝔸\)≥0aq\(\\mathbb\{A\}\)\\geq 0, so both added terms in \([41](https://arxiv.org/html/2609.30474#S3.E41)\) contribute to havingPacc\(MD\)\>Pacc\(SD\)P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(MD\)\>P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)\.
Remember thatu∈\[1−b,1\+a\]u\\in\[1\-b,1\+a\]andf\(1\)=0f\(1\)=0\([2](https://arxiv.org/html/2609.30474#S3.E2)\) so the unknown term in \([42](https://arxiv.org/html/2609.30474#S3.E42)\) may be quite small depending on the choice offf\. We have already seen that TV is special inff\-divergences for mentored decoding: its set of solutions is a simple geometric problem, which, for any otherff, becomes substantially more involved\. It turns out that the TV divergence holds another singular property\.
#### One mentor to rule them all and the role of TV\-MD
Before tackling boosting, we show two important invariants\. First, under some lightweight conditions onff– satisfied in particular by all strictly convex generators –, all optimal solutions offf\-MD are also in the set of optimal solutions for the total variation divergence\. Second the set of optimal mentored distributions asDDranges as per Assumption[3\.3](https://arxiv.org/html/2609.30474#S3.Thmtheorem3)are the samefor any strictly convexff\.
###### Theorem 3\.13\.
Under assumption[3\.3](https://arxiv.org/html/2609.30474#S3.Thmtheorem3), for anyffas per \([2](https://arxiv.org/html/2609.30474#S3.E2)\) and any\(𝐫,𝐬\)∈mdf2\(𝐩,𝐪,D\)\(\\bm\{r\},\\bm\{s\}\)\\in\\textsc\{md\}^\{2\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)such thatα,β\\alpha,\\betain \([31](https://arxiv.org/html/2609.30474#S3.E31)\) satisfyLβ\(−h\)≤1≤Lα\(−g\)L\_\{\\beta\}\(\-h\)\\leq 1\\leq L\_\{\\alpha\}\(\-g\), there existsD′\>0D^\{\\prime\}\>0such that
\(𝒓,𝒔\)\\displaystyle\(\\bm\{r\},\\bm\{s\}\)∈\\displaystyle\\inmdfTV2\(𝒑,𝒒,D′\)\.\\displaystyle\\textsc\{md\}^\{2\}\_\{f\_\{\\mathrm\{TV\}\}\}\(\\bm\{p\},\\bm\{q\};D^\{\\prime\}\)\.Hence, any such optimal solution toff\-MD is also optimal for the total variation divergence\. Furthermore, for any strictly convex generatorsf,gf,g\([2](https://arxiv.org/html/2609.30474#S3.E2)\),
\{mdf2\(𝒑,𝒒,D\):Das per Assumption[3\.3](https://arxiv.org/html/2609.30474#S3.Thmtheorem3)\}=\{mdg2\(𝒑,𝒒,D\):Das per Assumption[3\.3](https://arxiv.org/html/2609.30474#S3.Thmtheorem3)\}\.\\displaystyle\\\{\\textsc\{md\}^\{2\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\):D\\mbox\{ as per Assumption \\ref\{assum\-mda\}\}\\\}=\\\{\\textsc\{md\}^\{2\}\_\{g\}\(\\bm\{p\},\\bm\{q\};D\):D\\mbox\{ as per Assumption \\ref\{assum\-mda\}\}\\\}\.\(43\)
Proof in Appendix, Section[VIII\.6](https://arxiv.org/html/2609.30474#S8.SS6)\. We stress the importance of these properties, both from the standpoint of finding optimal mentored distributions \(see also Section[5](https://arxiv.org/html/2609.30474#S5)\) and also for the particular case of the total variation, whose remarkable properties already included modeling the optimal rejection metric for SD\([Yin et al\., 2024](https://arxiv.org/html/2609.30474#bib.bib27), Theorem 2\)\.
## 4Mentored decoding meets boosting
In this Section, we connect mentored decoding as analyzed in Section[3](https://arxiv.org/html/2609.30474#S3)to one of ML’s most famous training framework, boosting\([Schapire and Freund, 2012](https://arxiv.org/html/2609.30474#bib.bib23)\)\. Our main boosting algorithm is different from the classical blueprint, so we shall have to introduce and analyze it first\. But before, we define the general boosting framework\. We have access to a training sample𝒮=\.\{wi,\(𝒙i,𝒚i\),𝒙i∈𝒳,𝒚i∈𝒴\}i∈\[m\]\\mathcal\{S\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\\{w\_\{i\},\(\\bm\{x\}\_\{i\},\\bm\{y\}\_\{i\}\),\\bm\{x\}\_\{i\}\\in\\mathcal\{X\},\\bm\{y\}\_\{i\}\\in\\mathcal\{Y\}\\\}\_\{i\\in\[m\]\}ofmmexamples\. Here,𝒳\\mathcal\{X\}is the set of all possible inputs of a LLM, including prompts, etc\.\. We adopt the lightweight approach of[Zhu et al\. \(2009\)](https://arxiv.org/html/2609.30474#bib.bib54)for𝒴\\mathcal\{Y\}\.𝒴=\.\{𝒚∈ℝn:𝟏⊤𝒚=0\}\\mathcal\{Y\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\\{\\bm\{y\}\\in\\mathbb\{R\}^\{n\}:\\bm\{1\}^\{\\top\}\\bm\{y\}=0\\\}and𝒚i\\bm\{y\}\_\{i\}has two possible coordinates,1/ni1/n\_\{i\}and−1/\(n−ni\)\-1/\(n\-n\_\{i\}\);yij=1/niy\_\{ij\}=1/n\_\{i\}iff tokenjjis a potential next token for𝒙i\\bm\{x\}\_\{i\}andnin\_\{i\}is the number of such potential next tokens\. We denote𝒴i=\.\{j:yij=1/ni\}\\mathcal\{Y\}\_\{i\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\\{j:y\_\{ij\}=1/n\_\{i\}\\\}and𝒴¯i=\.\[n\]\\𝒴i\\overline\{\\mathcal\{Y\}\}\_\{i\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\[n\]\\backslash\\mathcal\{Y\}\_\{i\}\. We assume without loss of generality that0<ni<n,∀i∈\[m\]0<n\_\{i\}<n,\\forall i\\in\[m\]so none of these sets is empty\. Finally,𝟎<𝒘∈Δm\\bm\{0\}<\\bm\{w\}\\in\\Delta\_\{m\}is the initial weight vector of the training sample, usually uniform\.
#### Boosting in our LLM context
Even when our embedding of boosting in mentored decoding shall be made with two models, one drafter and one target, we first develop a general theory for any number of such models\. Also, distinguishing drafters and targets makes no real sense for the general boosting theory we first develop, so let us assume first we have a sequence ofT\>1T\>1LLMs whose last layer \(real\) prediction is denoted𝒉t:𝒳→ℝn,t∈\[T\]\\bm\{h\}\_\{t\}:\\mathcal\{X\}\\rightarrow\\mathbb\{R\}^\{n\},t\\in\[T\]\. Note that we assume that these models are already available, which makes sense in the current state of LLMs, but we might as welltrainsequentiallyTTmodels as is usually the case in boosting\. The results we present here are oblivious to how the models are made available\.
#### Predictions
Should we use separately each of these models, the corresponding probability vectors𝒑t\\bm\{p\}\_\{t\}to predict the next token would be proportional toexp𝒉t\\exp\\bm\{h\}\_\{t\}\. In our case however and for technical reasons, we are going to renormalize𝒉t\\bm\{h\}\_\{t\}by a scalar positive constant computed from the training sample, thus playing no role in ranking probabilities\. Let
𝒑t\(𝒙\)\\displaystyle\\bm\{p\}\_\{t\}\(\\bm\{x\}\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}1Zt⋅exp\(1ht,∞⋅𝒉t\(𝒙\)\)∈Δn,withht,∞=\.maxj∈\[m\]‖𝒉t\(𝒙j\)‖∞,\\displaystyle\\frac\{1\}\{Z\_\{t\}\}\\cdot\\exp\\left\(\\frac\{1\}\{h\_\{t,\\infty\}\}\\cdot\\bm\{h\}\_\{t\}\(\\bm\{x\}\)\\right\)\\in\\Delta\_\{n\},\\quad\\mbox\{ with \}h\_\{t,\\infty\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\max\_\{j\\in\[m\]\}\\\|\\bm\{h\}\_\{t\}\(\\bm\{x\}\_\{j\}\)\\\|\_\{\\infty\},\(44\)whereZtZ\_\{t\}is used for normalization\. Importantly,ht,∞h\_\{t,\\infty\}is the maxL∞L\_\{\\infty\}norm of𝒉t\\bm\{h\}\_\{t\}on training: it is thus trivially computable and finite\. In boosting’s jargon, each such predictor is called aweakpredictor because boosting provides a way to craft an ensemble from each of them with rapidly improving quality even when each weak predictor is just slightly better than random guessing\. Boosting works by combining all last layers – or equivalently all theseTTprobability vectors – to get a boosted output𝒑T\(𝒙\)\\bm\{p\}\_\{T\}\(\\bm\{x\}\)\. The quality of𝒑T\(𝒙\)\\bm\{p\}\_\{T\}\(\\bm\{x\}\)is evaluated by comparing, for each training examplei∈\[m\]i\\in\[m\], output probabilities for its potential next tokens in𝒴i\\mathcal\{Y\}\_\{i\}to the other ones in𝒴¯i\\overline\{\\mathcal\{Y\}\}\_\{i\}\. Specifically, we want the coordinates of𝒑T\(𝒙\)\\bm\{p\}\_\{T\}\(\\bm\{x\}\)in𝒴i\\mathcal\{Y\}\_\{i\}to be large enough compared to those in𝒴¯i\\overline\{\\mathcal\{Y\}\}\_\{i\}, where comparisons use the geometric average of the corresponding sets\. The geometric average has the essential property to be zero\-attracting: for such successful examples, it will prevent in generalanycoordinate of𝒑T\(𝒙\)\\bm\{p\}\_\{T\}\(\\bm\{x\}\)in𝒴i\\mathcal\{Y\}\_\{i\}to be too close to zero\.
### 4\.1The boosting scheme for generalTT
Our boosting scheme relies on a substantial generalization of\([Nock and Nielsen, 2007](https://arxiv.org/html/2609.30474#bib.bib12)\)to the multiclass case and geared to the analysis of probabilities and not real valued predictions\. Define the sequence of weights𝒘t∈Δm,t∈\[T\]\\bm\{w\}\_\{t\}\\in\\Delta\_\{m\},t\\in\[T\]such that𝒘1=\.𝒘\>𝟎\\bm\{w\}\_\{1\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\bm\{w\}\>\\bm\{0\}is the weight vector in𝒮\\mathcal\{S\}and otherwise obeys the recurrence
w\(t\+1\)i\\displaystyle w\_\{\(t\+1\)i\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}wti⋅1−μt2ht,∞⋅𝒚i⊤𝒉t\(𝒙i\)1−μt2,t∈\[T\],i∈\[m\]\\displaystyle w\_\{ti\}\\cdot\\frac\{1\-\\frac\{\\mu\_\{t\}\}\{2h\_\{t,\\infty\}\}\\cdot\\bm\{y\}\_\{i\}^\{\\top\}\\bm\{h\}\_\{t\}\(\\bm\{x\}\_\{i\}\)\}\{1\-\\mu\_\{t\}^\{2\}\},t\\in\[T\],i\\in\[m\]\(45\)\(note that formula \([45](https://arxiv.org/html/2609.30474#S4.E45)\) is self\-normalized inΔm\\Delta\_\{m\}: there is no normalization coefficient as e\.g\. in AdaBoost\), where coefficientμt\\mu\_\{t\}is anedgedefined as
μt\\displaystyle\\mu\_\{t\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}12ht,∞⋅∑i∈\[m\]wti⋅𝒚i⊤𝒉t\(𝒙i\),t∈\[T\]\.\\displaystyle\\frac\{1\}\{2h\_\{t,\\infty\}\}\\cdot\\sum\_\{i\\in\[m\]\}w\_\{ti\}\\cdot\\bm\{y\}\_\{i\}^\{\\top\}\\bm\{h\}\_\{t\}\(\\bm\{x\}\_\{i\}\),t\\in\[T\]\.\(46\)Hölder’s inequality and the definition of𝒴\\mathcal\{Y\}imply\|𝒚i⊤𝒉t\(𝒙i\)\|≤‖𝒚i‖1⋅ht,∞\|\\bm\{y\}\_\{i\}^\{\\top\}\\bm\{h\}\_\{t\}\(\\bm\{x\}\_\{i\}\)\|\\leq\\\|\\bm\{y\}\_\{i\}\\\|\_\{1\}\\cdot h\_\{t,\\infty\}and‖𝒚i‖1=ni∗\(1/ni\)\+\(n−ni\)∗\(1/\(n−ni\)\)=2\\\|\\bm\{y\}\_\{i\}\\\|\_\{1\}=n\_\{i\}\*\(1/n\_\{i\}\)\+\(n\-n\_\{i\}\)\*\(1/\(n\-n\_\{i\}\)\)=2, so\|μt\|≤1\|\\mu\_\{t\}\|\\leq 1\. In fact, let us assume without loss of generality that\|μt\|<1\|\\mu\_\{t\}\|<1otherwise either𝒉t\(𝒙i\)\\bm\{h\}\_\{t\}\(\\bm\{x\}\_\{i\}\)or−𝒉t\(𝒙i\)\-\\bm\{h\}\_\{t\}\(\\bm\{x\}\_\{i\}\)has the same signs as𝒚i\\bm\{y\}\_\{i\}for alli∈\[m\]i\\in\[m\]and so we are guaranteedptj\(𝒙i\)\>ptk\(𝒙i\)p\_\{tj\}\(\\bm\{x\}\_\{i\}\)\>p\_\{tk\}\(\\bm\{x\}\_\{i\}\)for anyi∈\[m\],j∈𝒴i,k∈𝒴¯ii\\in\[m\],j\\in\\mathcal\{Y\}\_\{i\},k\\in\\overline\{\\mathcal\{Y\}\}\_\{i\}, which would defeat the purpose of boosting𝒉t\\bm\{h\}\_\{t\}\. Secondly, ifμt≤0\\mu\_\{t\}\\leq 0then by just flipping𝒉t→−𝒉t\\bm\{h\}\_\{t\}\\rightarrow\-\\bm\{h\}\_\{t\}, we get the newμt≥0\\mu\_\{t\}\\geq 0\. To summarize, we observe
μt\\displaystyle\\mu\_\{t\}∈\\displaystyle\\in\[0,1\),∀t∈\[T\]\.\\displaystyle\[0,1\),\\forall t\\in\[T\]\.The fact that our weight update does without normalization coefficient is a crucial differentiator with the AdaBoost lineage of boosting algorithms\([Bartlett et al\., 1998](https://arxiv.org/html/2609.30474#bib.bib14);[Schapire and Freund, 2012](https://arxiv.org/html/2609.30474#bib.bib23)\): it saves the algorithmic computation of the normalizing coefficient, and more importantly, the simple closed form of the weights shall be important for the analysis of boosting in the context of mentored decoding\. We now construct𝝅~T\(𝒙\)\\tilde\{\\bm\{\\pi\}\}\_\{T\}\(\\bm\{x\}\), the boosted output\. We voluntarily name it with the same symbol as the mentored distribution of mentored decoding\.
###### Definition 4\.1\.
Forμt\\mu\_\{t\}defined in \([46](https://arxiv.org/html/2609.30474#S4.E46)\), let
ct\\displaystyle c\_\{t\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}14⋅ln\(1\+μt1−μt\),t∈\[T\]\.\\displaystyle\\frac\{1\}\{4\}\\cdot\\ln\\left\(\\frac\{1\+\\mu\_\{t\}\}\{1\-\\mu\_\{t\}\}\\right\),t\\in\[T\]\.\(47\)The boosted output model𝛑~T:𝒳→Δm\\tilde\{\\bm\{\\pi\}\}\_\{T\}:\\mathcal\{X\}\\rightarrow\\Delta\_\{m\}is defined as:
𝝅~T\(𝒙\)\\displaystyle\\tilde\{\\bm\{\\pi\}\}\_\{T\}\(\\bm\{x\}\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}1ZT⋅∏t=1T\(𝒑t\(𝒙\)\)ct∑u∈\[T\]cu∈Δn,\\displaystyle\\frac\{1\}\{Z\_\{T\}\}\\cdot\\prod\_\{t=1\}^\{T\}\\left\(\\bm\{p\}\_\{t\}\(\\bm\{x\}\)\\right\)^\{\\frac\{c\_\{t\}\}\{\\sum\_\{u\\in\[T\]\}c\_\{u\}\}\}\\in\\Delta\_\{n\},\(48\)whereZTZ\_\{T\}is the normalization coefficient\.
Note that in the context of next token prediction, the full computation of \([48](https://arxiv.org/html/2609.30474#S4.E48)\) is optional\. In particular, we can always spare the computation ofZTZ\_\{T\}\. We now analyze the boosting abilities of𝝅~T\\tilde\{\\bm\{\\pi\}\}\_\{T\}\.
### 4\.2Boosting the individual predictions in𝝅~T\\tilde\{\\bm\{\\pi\}\}\_\{T\}: main theorem
For any set of non negative reals𝒜\\mathcal\{A\},𝒜¯G\\overline\{\\mathcal\{A\}\}^\{G\}denotes the geometric average with uniform weights of the elements of𝒜\\mathcal\{A\}: for example,\{1,2\}¯G=\.11/2⋅21/2=2\\overline\{\\\{1,2\\\}\}^\{G\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}1^\{1/2\}\\cdot 2^\{1/2\}=\\sqrt\{2\}\.
###### Theorem 4\.2\.
For anyT\>1T\>1, suppose without loss of generality that the sequenceμ1,μ2,…,μT\\mu\_\{1\},\\mu\_\{2\},\.\.\.,\\mu\_\{T\}is non\-negative and with expectation𝔼\[μ\]\>0\\mathbb\{E\}\[\\mu\]\>0\. Then𝛑~T\\tilde\{\\bm\{\\pi\}\}\_\{T\}in \([48](https://arxiv.org/html/2609.30474#S4.E48)\) satisfies
ℙi∼𝒘\[\{π~T,j\(𝒙i\),j∈𝒴i\}¯G≤ρ⋅\{π~T,j\(𝒙i\),j∈𝒴¯i\}¯G\]≤exp\(−16⋅∑t=1Tμt2\),∀ρ≤exp\(2Q3\),\\displaystyle\\mathbb\{P\}\_\{i\\sim\\bm\{w\}\}\\left\[\\overline\{\\\{\\tilde\{\\pi\}\_\{T,j\}\(\\bm\{x\}\_\{i\}\),j\\in\\mathcal\{Y\}\_\{i\}\\\}\}^\{G\}\\leq\\rho\\cdot\\overline\{\\\{\\tilde\{\\pi\}\_\{T,j\}\(\\bm\{x\}\_\{i\}\),j\\in\\overline\{\\mathcal\{Y\}\}\_\{i\}\\\}\}^\{G\}\\right\]\\leq\\exp\\left\(\-\\frac\{1\}\{6\}\\cdot\\sum\_\{t=1\}^\{T\}\\mu\_\{t\}^\{2\}\\right\),\\forall\\rho\\leq\\exp\\left\(\\frac\{2Q\}\{3\}\\right\),\(49\)withQ=\.𝔼\[μ\]\+𝕍\[μ\]𝔼\[μ\]Q\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\mathbb\{E\}\[\\mu\]\+\\frac\{\\mathbb\{V\}\[\\mu\]\}\{\\mathbb\{E\}\[\\mu\]\}and𝕍\[μ\]\\mathbb\{V\}\[\\mu\]is the variance of the sequenceμ1,μ2,…,μT\\mu\_\{1\},\\mu\_\{2\},\.\.\.,\\mu\_\{T\}\.
Proof in Appendix, Section[VIII\.7](https://arxiv.org/html/2609.30474#S8.SS7)\.
Hence, we are guaranteed that a rapidly growing proportion of training sample will have a geometric average of the probabilities for the true next tokens larger than the geometric average of the other "bad" tokens by a "margin" factorρ\>1\\rho\>1\. Note the quantitative advantage of the geometric average being zero\-attracting for those "good" examples: if the geometric average of the bad tokens is\>0\>0, then no coordinate in the good tokens can be zero\. We now summarize a more qualitative analysis based on Theorem[4\.2](https://arxiv.org/html/2609.30474#S4.Thmtheorem2)\.
###### Definition 4\.4\.
The boosting advantage of the sequence\{𝐡t\}t∈\[T\]\\\{\\bm\{h\}\_\{t\}\\\}\_\{t\\in\[T\]\}is the quantity
A\(\{𝒉t\}t∈\[T\]\)\\displaystyle A\\left\(\\\{\\bm\{h\}\_\{t\}\\\}\_\{t\\in\[T\]\}\\right\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}∑t=1Tμt2,\\displaystyle\\sum\_\{t=1\}^\{T\}\\mu\_\{t\}^\{2\},\(50\)whereμt\\mu\_\{t\}is defined in \([46](https://arxiv.org/html/2609.30474#S4.E46)\)\.
Introducing boosting’s so\-called Weak Learning Assumption\([Bartlett et al\., 1998](https://arxiv.org/html/2609.30474#bib.bib14);[Nock and Nielsen, 2007](https://arxiv.org/html/2609.30474#bib.bib12)\):
∃γ\>0:\|μt\|≥γ,∀t∈\[T\],\\displaystyle\\exists\\upgamma\>0:\|\\mu\_\{t\}\|\\geq\\upgamma,\\forall t\\in\[T\],\(WLA\)we get anΩ\(T\)\\Omega\(T\)boosting advantage:
A\(\{𝒉t\}t∈\[T\]\)\\displaystyle A\\left\(\\\{\\bm\{h\}\_\{t\}\\\}\_\{t\\in\[T\]\}\\right\)≥\\displaystyle\\geqγ2⋅T,\\displaystyle\\upgamma^\{2\}\\cdot T,\(51\)and so under \([WLA](https://arxiv.org/html/2609.30474#S4.Ex17)\), for anyε\>0\\varepsilon\>0, we have that a proportion≥1−ε\\geq 1\-\\varepsilonof the training examples observe\{π~T,j\(𝒙i\),j∈𝒴i\}¯G\>ρ⋅\{π~T,j\(𝒙i\),j∈𝒴¯i\}¯G\\overline\{\\\{\\tilde\{\\pi\}\_\{T,j\}\(\\bm\{x\}\_\{i\}\),j\\in\\mathcal\{Y\}\_\{i\}\\\}\}^\{G\}\>\\rho\\cdot\\overline\{\\\{\\tilde\{\\pi\}\_\{T,j\}\(\\bm\{x\}\_\{i\}\),j\\in\\overline\{\\mathcal\{Y\}\}\_\{i\}\\\}\}^\{G\}as soon as
T\\displaystyle T=\\displaystyle=⌈6γ2⋅log1ε⌉,\\displaystyle\\left\\lceil\\frac\{6\}\{\\upgamma^\{2\}\}\\cdot\\log\\frac\{1\}\{\\varepsilon\}\\right\\rceil,\(52\)and the largest possibleρ\\rhois≥exp\(2γ/3\)\\geq\\exp\(2\\upgamma/3\)\. Note also that if each𝒉t\\bm\{h\}\_\{t\}were to be chosen uniformly at random in a set of, say, unit\-L2L\_\{2\}norm predictors, then the expectation over randomness would give𝔼\[𝒚i⊤𝒉t\(𝒙i\)\]=0\\mathbb\{E\}\[\\bm\{y\}\_\{i\}^\{\\top\}\\bm\{h\}\_\{t\}\(\\bm\{x\}\_\{i\}\)\]=0for eachi∈\[m\]i\\in\[m\], which justifies the name weak predictors for our sequence of𝒉t\\bm\{h\}\_\{t\}as the \([WLA](https://arxiv.org/html/2609.30474#S4.Ex17)\) only requires them to slightly beat such a random performance\. Finally, in the context of LLMs, note also that \([49](https://arxiv.org/html/2609.30474#S4.E49)\) provides a simple way to cherry pick a subset of available pretrained models, by greedily picking the one maximizing\|μt\|\|\\mu\_\{t\}\|\*\*\*The greedy selection may not be optimal over all sequences of inclusion, see Section[6](https://arxiv.org/html/2609.30474#S6)\.\.
We now have the tools to connect boosting and mentored decoding\. We achieve this in two Subsections, first tackling the case of the total variation, and then the general case\. The way we fold boosting in is different in both cases\.
### 4\.3Mentored decoding and boosting: the case of total variation

Figure 3:Plots oflogt\\log\_\{t\}\(left\) andexpt\\exp\_\{t\}\(right\), fort=0,1/4,1/2,3/4,1t=0,1/4,1/2,3/4,1in black curves where thickness increases withtt\(see text\)\.Mentored decoding builds a mentored distribution𝝅\\bm\{\\pi\}that depend on the output𝒑,𝒒\\bm\{p\},\\bm\{q\}of the drafter and target\. From the boosting standpoint, which analyzes the composite / ensemblemodelproducing𝝅\\bm\{\\pi\}, we thus end up analyzing the boosting ability of potentiallyas many ensemble modelsas there can be foranyoutputs of the drafter and target\. The connection between mentored decoding and boosting is made by a combination of the models’ outputs specific to each output𝒑,𝒒\\bm\{p\},\\bm\{q\}, in such a way that it always yields guarantees on the exponential rate in \([49](https://arxiv.org/html/2609.30474#S4.E49)\) while being optimal from the mentored decoding problem \([f \-MD\-2](https://arxiv.org/html/2609.30474#S3.Ex1)\) \([f \-MD\-1](https://arxiv.org/html/2609.30474#S3.Ex2)\), and the key parameters of these two problems – the edges \([46](https://arxiv.org/html/2609.30474#S4.E46)\) for boosting, the divergence constraintDDfor mentored decoding – depend on a real parameter function of𝒑\\bm\{p\}and𝒒\\bm\{q\}, whose existence is guaranteed by Assumption[3\.3](https://arxiv.org/html/2609.30474#S3.Thmtheorem3)\. From now on, in the context of mentored decoding, the boosting setting corresponds to the specific case ofT=2T=2models\. This is obviously a very small number of models, but boosting has this property that the marginal improvement of the first few models due to boosting is usually dramatically larger than for the next ones: combining drafter and target models may be sufficient for the boosted model to be better than each of them\. Of course, our boosting setting may also apply to mentored decoding settings involving more than two models\.
#### Computation of the mentored distribution
The key non\-trivial constraint for any mentored distribution𝝅\\bm\{\\pi\}to be optimal forfTVf\_\{\\mathrm\{TV\}\}\-MD is to belong to the hyperrectangle defined by𝒑\\bm\{p\}and𝒒\\bm\{q\}\([7](https://arxiv.org/html/2609.30474#S3.E7)\)\. We analyze the construction of the boosted model in \([48](https://arxiv.org/html/2609.30474#S4.E48)\) and show how its output from𝒑\\bm\{p\}and𝒒\\bm\{q\}can be compliant with this constraint\. For the analysis, we introduce the tempered versions of log and exp\([Naudts, 2011](https://arxiv.org/html/2609.30474#bib.bib19), Chapter 7\):
logt\(z\)=\.11−t⋅\(z1−t−1\)\\displaystyle\\log\_\{t\}\(z\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\frac\{1\}\{1\-t\}\\cdot\\left\(z^\{1\-t\}\-1\\right\),expt\(z\)=\.\[1\+\(1−t\)z\]\+1/\(1−t\)\(\[z\]\+=\.max\{0,z\}\),\\displaystyle\\exp\_\{t\}\(z\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\left\[1\+\(1\-t\)z\\right\]^\{1/\(1\-t\)\}\_\{\+\}\\quad\(\[z\]\_\{\+\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\max\\\{0,z\\\}\),\(53\)where the caset=1t=1is the extension by continuity to thelog\\logandexp\\expfunctions, respectively \(see Figure[3](https://arxiv.org/html/2609.30474#S4.F3)for examples\)\. Our focus is essentially ont∈\(0,1\)t\\in\(0,1\), for which the concavity / convexity of functions is the same as fort=1t=1, see also[Amid et al\. \(2024\)](https://arxiv.org/html/2609.30474#bib.bib16);[Amid et al\. \(2023\)](https://arxiv.org/html/2609.30474#bib.bib17);[Nock et al\. \(2023\)](https://arxiv.org/html/2609.30474#bib.bib18);[Naudts \(2011\)](https://arxiv.org/html/2609.30474#bib.bib19)for further relevant properties\. The following Lemma is central to our analysis\. We let⟦\.⟧\\llbracket\.\\rrbracketbe Iverson’s bracket\([Knuth, 1992](https://arxiv.org/html/2609.30474#bib.bib13)\), i\.e\. the Boolean truth value of the predicate inside\.
###### Lemma 4\.5\.
For any𝐩,𝐪∈Δn\\bm\{p\},\\bm\{q\}\\in\\Delta\_\{n\}satisfying𝐩,𝐪\>𝟎\\bm\{p\},\\bm\{q\}\>\\bm\{0\}, any0≤α≤10\\leq\\alpha\\leq 1, denote
i∗\\displaystyle i^\{\*\}=\\displaystyle=argmaxi\(qipi\)⟦pi\>qi⟧−α\\displaystyle\\arg\\max\_\{i\}\\left\(\\frac\{q\_\{i\}\}\{p\_\{i\}\}\\right\)^\{\\llbracket p\_\{i\}\>q\_\{i\}\\rrbracket\-\\alpha\}\(54\)\(without loss of generality, this is a singleton\)\. Suppose the following holds:
expα𝔼i∼𝒑logαqipi\\displaystyle\\exp\_\{\\alpha\}\\mathbb\{E\}\_\{i\\sim\\bm\{p\}\}\\log\_\{\\alpha\}\\frac\{q\_\{i\}\}\{p\_\{i\}\}≥\\displaystyle\\geqmaximin\{qipi,piqi\}ifpi∗\>qi∗,\\displaystyle\\max\_\{i\}\\min\\left\\\{\\frac\{q\_\{i\}\}\{p\_\{i\}\},\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\right\\\}\\quad\\mbox\{ if $p\_\{i\_\{\*\}\}\>q\_\{i\_\{\*\}\}$\},\(55\)exp1−α𝔼i∼𝒒log1−αpiqi\\displaystyle\\exp\_\{1\-\\alpha\}\\mathbb\{E\}\_\{i\\sim\\bm\{q\}\}\\log\_\{1\-\\alpha\}\\frac\{p\_\{i\}\}\{q\_\{i\}\}≥\\displaystyle\\geqmaximin\{qipi,piqi\}ifpi∗<qi∗\.\\displaystyle\\max\_\{i\}\\min\\left\\\{\\frac\{q\_\{i\}\}\{p\_\{i\}\},\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\right\\\}\\quad\\mbox\{ if $p\_\{i\_\{\*\}\}<q\_\{i\_\{\*\}\}$\}\.\(56\)Then if we let
𝒗α\\displaystyle\\bm\{v\}\_\{\\alpha\}∝\\displaystyle\\propto𝒑α⊙𝒒1−α∈Δn\\displaystyle\\bm\{p\}^\{\\alpha\}\\odot\\bm\{q\}^\{1\-\\alpha\}\\in\\Delta\_\{n\}\(57\)the distribution obtained from the coordinate\-wise geometric average of𝐩\\bm\{p\}and𝐪\\bm\{q\}, the following holds:
𝒗α∈\[min\{𝒑,𝒒\},max\{𝒑,𝒒\}\]\.\\displaystyle\\bm\{v\}\_\{\\alpha\}\\in\[\\min\\\{\\bm\{p\},\\bm\{q\}\\\},\\max\\\{\\bm\{p\},\\bm\{q\}\\\}\]\.\(58\)
Proof in Appendix, Section[VIII\.8](https://arxiv.org/html/2609.30474#S8.SS8)\. Note that the Lemma is useful only if𝒑≠𝒒\\bm\{p\}\\neq\\bm\{q\}: otherwise, it can hold only when𝒑=𝒒\\bm\{p\}=\\bm\{q\}\. When𝒑≠𝒒\\bm\{p\}\\neq\\bm\{q\}and𝒑,𝒒\>𝟎\\bm\{p\},\\bm\{q\}\>\\bm\{0\}, it is easy to show thatlogt\\log\_\{t\}continuously converges toz↦z−1z\\mapsto z\-1ast→0t\\rightarrow 0so that the LHS of \([55](https://arxiv.org/html/2609.30474#S4.E55)\) \(resp\. \([56](https://arxiv.org/html/2609.30474#S4.E56)\)\) continuously converges to 1 asα→0\\alpha\\rightarrow 0\(resp\.α→1\\alpha\\rightarrow 1\), since we observeexpt\(0\)=1,∀t∈\[0,1\]\\exp\_\{t\}\(0\)=1,\\forall t\\in\[0,1\]\. Since the RHS are<1<1, \([55](https://arxiv.org/html/2609.30474#S4.E55)\) \(resp\. \([56](https://arxiv.org/html/2609.30474#S4.E56)\)\) necessarily holds for any0<α<α∗0<\\alpha<\\alpha\_\{\*\}\(resp1−α∗<α<11\-\\alpha\_\{\*\}<\\alpha<1\) for a small enoughα∗<1\\alpha\_\{\*\}<1\. So under our Assumption[3\.3](https://arxiv.org/html/2609.30474#S3.Thmtheorem3), mentored decoding for the output can be accompanied by a "qualitative" form of boosting for the models\. We now complete it with a quantitative one, first describing the ensemble model\.
#### The combination of drafter and target
We now define two key parameters to analyze the imbrication of boosting and mentored decoding\.
###### Definition 4\.6\.
For any𝐩\\bm\{p\},𝐪\\bm\{q\}complying with Assumption[3\.3](https://arxiv.org/html/2609.30474#S3.Thmtheorem3), letϱ𝐩𝐪,ε𝐩𝐪\\varrho\_\{\\bm\{p\}\\bm\{q\}\},\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}denote any reals such thatε𝐩𝐪∈\(0,1\],ϱ𝐩𝐪≥0,ε𝐩𝐪\(1\+ϱ𝐩𝐪\)≤1\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\in\(0,1\],\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\\geq 0,\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\(1\+\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\)\\leq 1and:
minimax\{qipi,piqi\}≥1\+ϱ𝒑𝒒,maxiqipi≤1ε𝒑𝒒,miniqipi≥ε𝒑𝒒\.\\displaystyle\\begin\{array\}\[\]\{ccc\}\\min\_\{i\}\\max\\left\\\{\\frac\{q\_\{i\}\}\{p\_\{i\}\},\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\right\\\}\\geq 1\+\\varrho\_\{\\bm\{p\}\\bm\{q\}\},&\\max\_\{i\}\\frac\{q\_\{i\}\}\{p\_\{i\}\}\\leq\\frac\{1\}\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\},&\\min\_\{i\}\\frac\{q\_\{i\}\}\{p\_\{i\}\}\\geq\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\.\\end\{array\}\(ED\)
While the two rightmost conditions bound the most dissimilar coordinates in𝒒\\bm\{q\}and𝒑\\bm\{p\}, the leftmost is a condition on the most similar one, all in term of density ratio\. For example, if𝒑=\.\(3/10,1/10,3/5\)\\bm\{p\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\(3/10,1/10,3/5\)and𝒒=\.\(2/5,1/5,2/5\)\\bm\{q\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\(2/5,1/5,2/5\), the coordinate realizing the leftmost to the rightmost condition in \([ED](https://arxiv.org/html/2609.30474#S4.Ex18)\) \(with equality\) will be the first, second and third respectively\. Intuitively, the "freedom" in the joint choice of𝒑\\bm\{p\}and𝒒\\bm\{q\}augments asε𝒑𝒒\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}andϱ𝒑𝒒\\varrho\_\{\\bm\{p\}\\bm\{q\}\}decrease: asϱ𝒑𝒒,ε𝒑𝒒→0\\varrho\_\{\\bm\{p\}\\bm\{q\}\},\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\rightarrow 0, the most similar coordinates can be as close to 1 as desired, the most dissimilar coordinates in terms of the ratioq\./p\.q\_\{\.\}/p\_\{\.\}can span as much asℝ\+\\mathbb\{R\}\_\{\+\}as desired\. In our context however, we can expect the opposite: drafter and target outputs should achieve some level of agreement in their outputs because they were trained to achieve some quality level in their predictions\. The more they would agree on𝒑\\bm\{p\}and𝒒\\bm\{q\}coordinate\-wise, the larger we can pickε𝒑𝒒\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\. We however need to eventually decreaseϱ𝒑𝒒\\varrho\_\{\\bm\{p\}\\bm\{q\}\}for \([ED](https://arxiv.org/html/2609.30474#S4.Ex18)\) to remain true, keeping in mind we must keepϱ𝒑𝒒\>0\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\>0because of Assumption[3\.3](https://arxiv.org/html/2609.30474#S3.Thmtheorem3)\.
From the boosting standpoint, we expect the target model to be better than the drafter from the edge standpoint, so the boosted ensemble first includes the target and then the drafter \(interestingly enough, this greedy strategy can prove suboptimal, see Section[6](https://arxiv.org/html/2609.30474#S6)\)\. While the first boosting coefficient strictly follows \([47](https://arxiv.org/html/2609.30474#S4.E47)\), the boosted coefficient of the drafter may be \(nonlinearly\)scaled downto ensure that the resulting mentored distribution𝝅\\bm\{\\pi\}is optimal for thefTVf\_\{\\mathrm\{TV\}\}\-MD problem\. This scaling depends onε𝒑𝒒\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}andϱ𝒑𝒒\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\. First,α\\alphain \([57](https://arxiv.org/html/2609.30474#S4.E57)\) is simply
α\\displaystyle\\alpha=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}c2c2\+ctarget,\\displaystyle\\frac\{c\_\{2\}\}\{c\_\{2\}\+c\_\{\\mbox\{\\tiny\{target\}\}\}\},\(60\)where bothccs are computed using \([47](https://arxiv.org/html/2609.30474#S4.E47)\)\. The edgeμtarget\\mu\_\{\\mbox\{\\tiny\{target\}\}\}is as in \([46](https://arxiv.org/html/2609.30474#S4.E46)\) and thus its computation does not depend on𝒑,𝒒\\bm\{p\},\\bm\{q\}\. Howeverμ2\\mu\_\{2\}used forc2c\_\{2\}isμdrafter\\mu\_\{\\mbox\{\\tiny\{drafter\}\}\}eventually clamped:
μ2\\displaystyle\\mu\_\{2\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}min\{μdrafter,\(1\+μtarget\)ϱ𝒑𝒒ε𝒑𝒒−\(1−μtarget\)ϱ𝒑𝒒ε𝒑𝒒\(1\+μtarget\)ϱ𝒑𝒒ε𝒑𝒒\+\(1−μtarget\)ϱ𝒑𝒒ε𝒑𝒒\},\\displaystyle\\min\\left\\\{\\mu\_\{\\mbox\{\\tiny\{drafter\}\}\},\\frac\{\\left\(1\+\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}\-\\left\(1\-\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}\}\{\\left\(1\+\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}\+\\left\(1\-\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}\}\\right\\\},\(61\)whereμdrafter\\mu\_\{\\mbox\{\\tiny\{drafter\}\}\}follows \([46](https://arxiv.org/html/2609.30474#S4.E46)\)\.
#### Main theorem
Armed with these definitions, we now prove the Theorem that bringsfTVf\_\{\\mathrm\{TV\}\}\-MD optimality and boosting\.
###### Theorem 4\.7\.
The computation ofα\\alphaas in \([60](https://arxiv.org/html/2609.30474#S4.E60)\) simultanously yields:
- •𝒗α\\bm\{v\}\_\{\\alpha\}in \([57](https://arxiv.org/html/2609.30474#S4.E57)\) is optimal for thefTVf\_\{\\mathrm\{TV\}\}\-MD problem for the choiceD=DfTV\(𝒗α∥𝒒\)D=D\_\{f\_\{\\mathrm\{TV\}\}\}\(\\bm\{v\}\_\{\\alpha\}\\\|\\bm\{q\}\);
- •the corresponding boosted model with coefficientsctarget,c2c\_\{\\mbox\{\\tiny\{target\}\}\},c\_\{2\}has boosted advantage satisfyingA\(\{𝒉1=\.𝒉target,𝒉2=\.𝒉drafter\}\)≥\(1\+ϱ𝒑𝒒2ε𝒑𝒒2\)⋅μtarget2A\\left\(\\\{\\bm\{h\}\_\{1\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\bm\{h\}\_\{\\mbox\{\\tiny\{target\}\}\},\\bm\{h\}\_\{2\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\bm\{h\}\_\{\\mbox\{\\tiny\{drafter\}\}\}\\\}\\right\)\\geq\(1\+\\varrho\_\{\\bm\{p\}\\bm\{q\}\}^\{2\}\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}^\{2\}\)\\cdot\\mu^\{2\}\_\{\\mbox\{\\tiny\{target\}\}\}\.
Finally, we observe
DfTV\(𝒗α∥𝒒\)\\displaystyle D\_\{f\_\{\\mathrm\{TV\}\}\}\(\\bm\{v\}\_\{\\alpha\}\\\|\\bm\{q\}\)≤\\displaystyle\\leqϱ𝒑𝒒ε𝒑𝒒\.\\displaystyle\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\.\(62\)
Proof in Appendix, Section[VIII\.9](https://arxiv.org/html/2609.30474#S8.SS9)\. Importantly, the boosting advantage can be free from any dependence inε𝒑𝒒\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}andϱ𝒑𝒒\\varrho\_\{\\bm\{p\}\\bm\{q\}\}if there is no clamping ofμ2\\mu\_\{2\}– in such a case, we getA=μtarget2\+μdrafter2A=\\mu^\{2\}\_\{\\mbox\{\\tiny\{target\}\}\}\+\\mu^\{2\}\_\{\\mbox\{\\tiny\{drafter\}\}\}\. In particular, quantityε𝒑𝒒\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}limits the quality of boosting and in the limit asε𝒑𝒒→0\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\rightarrow 0, the boosted guarantees of combining two models vanish and we can only guarantee a quality identical to the target model’s\. What we should expect in such a case, given that there is then a form of convergence of𝒗α\\bm\{v\}\_\{\\alpha\}to the target output𝒒\\bm\{q\}in this case from \([57](https://arxiv.org/html/2609.30474#S4.E57)\), is that the fate of boosting clearly becomes a blessing for the TV divergence, namely that the authorized boundDDalso converges to zero as a function ofε𝒑𝒒\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\. This is what \([62](https://arxiv.org/html/2609.30474#S4.E62)\) guarantees\.
### 4\.4Mentored decoding and boosting: general case
We now make use of the approximation toff\-MD in Theorem[3\.12](https://arxiv.org/html/2609.30474#S3.Thmtheorem12)for boosting\. Our path to get here is much different from Theorem[4\.7](https://arxiv.org/html/2609.30474#S4.Thmtheorem7): the case offTVf\_\{\\mathrm\{TV\}\}\-MD yields a huge set of optimal solutions which we showed can contain convenient boosting solutions as well\. For generalffhowever, the set of optimal solution is in general as small as a singleton \(ifffstrictly convex\)\. Instead of hammering boosting solutions in such a small set, we are going to show that theapproximatesolutions toff\-MD of Theorem[3\.12](https://arxiv.org/html/2609.30474#S3.Thmtheorem12)havede factonice boosting properties\. In other words, the coordinates of the mentored distribution for the potential next tokens cannot be "too small" with respect to other coordinates, that are associated to tokens that cannot be potential next tokens\. The amount by which both sets of coordinates compare to each other depends on parameters evaluated on a sample from the domain\.
###### Theorem 4\.8\.
For any drafter and target models, and any sample𝒮′=\.\{wi′,\(𝐱i,𝐲i\),𝐱i∈𝒳,𝐲i∈𝒴\}i∈\[m′\]\\mathcal\{S\}^\{\\prime\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\\{w^\{\\prime\}\_\{i\},\(\\bm\{x\}\_\{i\},\\bm\{y\}\_\{i\}\),\\bm\{x\}\_\{i\}\\in\\mathcal\{X\},\\bm\{y\}\_\{i\}\\in\\mathcal\{Y\}\\\}\_\{i\\in\[m^\{\\prime\}\]\}, denote respectively𝐩i,𝐪i\\bm\{p\}\_\{i\},\\bm\{q\}\_\{i\}the outputs of drafter and target on input𝐱i\\bm\{x\}\_\{i\}\. Suppose the edges of the target and drafter on𝒮′\\mathcal\{S\}^\{\\prime\}satisfyμtarget,μdrafter\>0\\mu\_\{\\mbox\{\\tiny\{target\}\}\},\\mu\_\{\\mbox\{\\tiny\{drafter\}\}\}\>0\([46](https://arxiv.org/html/2609.30474#S4.E46)\)\. For anyi∈\[m′\]i\\in\[m^\{\\prime\}\]and any\(ai,bi\)∈𝒞\(𝐩i,𝐪i\)\(a\_\{i\},b\_\{i\}\)\\in\\mathcal\{C\}\(\\bm\{p\}\_\{i\},\\bm\{q\}\_\{i\}\), let
𝝅i\\displaystyle\\bm\{\\pi\}\_\{i\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}𝒑iclamped to\[1−bi,1\+ai\]⋅𝒒i,∀i∈\[m′\]\\displaystyle\\bm\{p\}\_\{i\}\\mbox\{ clamped to \}\[1\-b\_\{i\},1\+a\_\{i\}\]\\cdot\\bm\{q\}\_\{i\},\\forall i\\in\[m^\{\\prime\}\]\(63\)be the mentored distribution defined from drafter and target via the respective𝐫i,𝐬i\\bm\{r\}\_\{i\},\\bm\{s\}\_\{i\}in \([39](https://arxiv.org/html/2609.30474#S3.E39)\), \([40](https://arxiv.org/html/2609.30474#S3.E40)\)\. Let
ki\\displaystyle k\_\{i\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}1\+ai1−bi\(≥1\)\.\\displaystyle\\sqrt\{\\frac\{1\+a\_\{i\}\}\{1\-b\_\{i\}\}\}\\quad\(\\geq 1\)\.\(64\)Then this mentored distribution satisfies
ℙi∼𝒘′\[\{πij:j∈𝒴i\}¯G≤ρ⋅min\{ki⋅ε𝒑i𝒒ic2c1\+c2,1ki\}2⋅\{πij:j∈𝒴¯i\}¯G\]\\displaystyle\\mathbb\{P\}\_\{i\\sim\\bm\{w\}^\{\\prime\}\}\\left\[\\overline\{\\\{\\pi\_\{ij\}:j\\in\\mathcal\{Y\}\_\{i\}\\\}\}^\{G\}\\leq\\rho\\cdot\\min\\left\\\{k\_\{i\}\\cdot\\varepsilon^\{\\frac\{c\_\{2\}\}\{c\_\{1\}\+c\_\{2\}\}\}\_\{\\bm\{p\}\_\{i\}\\bm\{q\}\_\{i\}\},\\frac\{1\}\{k\_\{i\}\}\\right\\\}^\{2\}\\cdot\\overline\{\\\{\\pi\_\{ij\}:j\\in\\overline\{\\mathcal\{Y\}\}\_\{i\}\\\}\}^\{G\}\\right\]≤\\displaystyle\\leqexp\(−μ12\+μ226\),\\displaystyle\\exp\\left\(\-\\frac\{\\mu^\{2\}\_\{1\}\+\\mu^\{2\}\_\{2\}\}\{6\}\\right\),\(65\)for anyρ≤exp\(2⋅\(μ12\+μ22\)3⋅\(μ1\+μ2\)\)\\rho\\leq\\exp\\left\(\\frac\{2\\cdot\(\\mu^\{2\}\_\{1\}\+\\mu^\{2\}\_\{2\}\)\}\{3\\cdot\(\\mu\_\{1\}\+\\mu\_\{2\}\)\}\\right\)\. Here,μ1=\.μtarget,μ2=\.μdrafter\\mu\_\{1\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\mu\_\{\\mbox\{\\tiny\{target\}\}\},\\mu\_\{2\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\mu\_\{\\mbox\{\\tiny\{drafter\}\}\},c1,c2c\_\{1\},c\_\{2\}are defined in \([47](https://arxiv.org/html/2609.30474#S4.E47)\) andε𝐩i𝐪i\\varepsilon\_\{\\bm\{p\}\_\{i\}\\bm\{q\}\_\{i\}\}is as in \([ED](https://arxiv.org/html/2609.30474#S4.Ex18)\)\.
Proof in Appendix, Section[VIII\.10](https://arxiv.org/html/2609.30474#S8.SS10)\. We have two important remarks regarding Theorem[4\.8](https://arxiv.org/html/2609.30474#S4.Thmtheorem8)\. First, anyff\-MD problem has optimal mentored distributions with the general form \([63](https://arxiv.org/html/2609.30474#S4.E63)\) \(Theorem[3\.12](https://arxiv.org/html/2609.30474#S3.Thmtheorem12)\), for any generatorff\([2](https://arxiv.org/html/2609.30474#S3.E2)\), so Theorem[4\.8](https://arxiv.org/html/2609.30474#S4.Thmtheorem8)applies to all instances offf\-MD\. Second, the Theorem holds foranysample𝒮′=\.\{wi′,\(𝒙i,𝒚i\),𝒙i∈𝒳,𝒚i∈𝒴\}i∈\[m′\]\\mathcal\{S\}^\{\\prime\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\\{w^\{\\prime\}\_\{i\},\(\\bm\{x\}\_\{i\},\\bm\{y\}\_\{i\}\),\\bm\{x\}\_\{i\}\\in\\mathcal\{X\},\\bm\{y\}\_\{i\}\\in\\mathcal\{Y\}\\\}\_\{i\\in\[m^\{\\prime\}\]\}for whichμtarget,μdrafter\>0\\mu\_\{\\mbox\{\\tiny\{target\}\}\},\\mu\_\{\\mbox\{\\tiny\{drafter\}\}\}\>0, which is arguably a very weak assumption – in fact weaker than the weak learning assumption \([WLA](https://arxiv.org/html/2609.30474#S4.Ex17)\)\. Given a domain for which a sample𝒮′\\mathcal\{S\}^\{\\prime\}is available, the Theorem can be used to get an indication of the general quality of mentored distributions obtained from Theorem[3\.12](https://arxiv.org/html/2609.30474#S3.Thmtheorem12)\. The Theorem also carries qualitative value if we consider that we can leave the boosting parameters implicit, since we do not need the boosting model\. In such a case, if the target model is substantially better than the drafter,c2/\(c1\+c2\)c\_\{2\}/\(c\_\{1\}\+c\_\{2\}\)is very small and we can reduce themin\\minto thekik\_\{i\}dependent part, making the factor of the geometric averageΩ\(ρ/ki2\)\\Omega\(\\rho/k\_\{i\}^\{2\}\), i\.e\. independent of the boosting parameters\. This makeskik\_\{i\}, which depends on the clamping in the mentored distribution, directly influence the quality of the coordinates for potential next token vs others\. We also shall see, on a toy simulation, thatkkcan stay very close to 1 even for a substantial increase of the acceptance probability compared toPacc\(SD\)P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)\(Table[5](https://arxiv.org/html/2609.30474#S5.T5)\)\.
## 5Algorithms and related properties
Algorithm 1UpdateCBreakpoints\(r,c,a,b,i,j,qA,qB,amax,bmax\)\(\\textsc\{r\},\\textsc\{c\},a,b,i,j,q\_\{A\},q\_\{B\},a\_\{\\max\},b\_\{\\max\}\)Input:sorted array
r=\.\{\(pi/qi,qi\)\}i=1n\\textsc\{r\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\\{\(p\_\{i\}/q\_\{i\},q\_\{i\}\)\\\}\_\{i=1\}^\{n\}in increasing order of the ratio
p\./q\.p\_\{\.\}/q\_\{\.\}; current set of breakpointsc, parameters
0≤a≤amax0\\leq a\\leq a\_\{\\max\}all in \([36](https://arxiv.org/html/2609.30474#S3.E36)\), parameters
0≤b≤bmax0\\leq b\\leq b\_\{\\max\}all in \([37](https://arxiv.org/html/2609.30474#S3.E37)\), indexes
i,j∈\[n\]i,j\\in\[n\], probabilities
qA,qB∈\[0,1\]q\_\{A\},q\_\{B\}\\in\[0,1\]; // Ratio values inrare called "ticks"
Step 1 :if
a≤amaxa\\leq a\_\{\\max\}and
b≤bmaxb\\leq b\_\{\\max\}then
c←c∪\{\(a,b\)\}\\textsc\{c\}\\leftarrow\\textsc\{c\}\\cup\\\{\(a,b\)\\\}; // add current breakpoint
Step 2 :if
a\>amaxa\>a\_\{\\max\}or
b\>bmaxb\>b\_\{\\max\}or
j=0j=0or
i=n−1i=n\-1then returnc;
Step 3 : // compute thresholds
τa←\(ri0−1−a\)⋅qA\\displaystyle\\tau\_\{a\}\\leftarrow\\left\(r\_\{i0\}\-1\-a\\right\)\\cdot q\_\{A\};τb←\(1−b−rj0\)⋅qB;\\displaystyle\\tau\_\{b\}\\leftarrow\\left\(1\-b\-r\_\{j0\}\\right\)\\cdot q\_\{B\};
Step 4 :if
τb<τa\\tau\_\{b\}<\\tau\_\{a\}then//
bbreaches a tick
4\.1 :
a←a\+τbqAa\\leftarrow a\+\\frac\{\\tau\_\{b\}\}\{q\_\{A\}\};
4\.2 :
b←1−rj0b\\leftarrow 1\-r\_\{j0\};
4\.3 :
qB←qB−rj1q\_\{B\}\\leftarrow q\_\{B\}\-r\_\{j1\};
4\.4 :
j←j−1j\\leftarrow j\-1;
Step 5 :else if
τb\>τa\\tau\_\{b\}\>\\tau\_\{a\}then//
aareaches a tick
5\.1 :
a←ri0−1a\\leftarrow r\_\{i0\}\-1;
5\.2 :
b←b\+τaqBb\\leftarrow b\+\\frac\{\\tau\_\{a\}\}\{q\_\{B\}\};
5\.3 :
qA←qA−ri1q\_\{A\}\\leftarrow q\_\{A\}\-r\_\{i1\};
5\.4 :
i←i\+1i\\leftarrow i\+1;
Step 6 :else//
aaand
bbreach a tick
6\.1 :
a←ri0−1a\\leftarrow r\_\{i0\}\-1;
6\.2 :
b←1−rj0b\\leftarrow 1\-r\_\{j0\};
6\.3 :
qA←qA−ri1q\_\{A\}\\leftarrow q\_\{A\}\-r\_\{i1\};
6\.4 :
qB←qB−rj1q\_\{B\}\\leftarrow q\_\{B\}\-r\_\{j1\};
6\.5 :
i←i\+1i\\leftarrow i\+1;
6\.6 :
j←j−1j\\leftarrow j\-1;
Step 9 :UpdateCBreakpoints\(r,c,a,b,i,j,qA,qB,amax,bmax\)\(\\textsc\{r\},\\textsc\{c\},a,b,i,j,q\_\{A\},q\_\{B\},a\_\{\\max\},b\_\{\\max\}\);
We now study the algorithmic side of the theory developed so far\. Note that the algorithmic efficiency of boosting to computeμdrafter\\mu\_\{\\mbox\{\\tiny\{drafter\}\}\}andμtarget\\mu\_\{\\mbox\{\\tiny\{target\}\}\}following \([46](https://arxiv.org/html/2609.30474#S4.E46)\) is orthogonal to the mentored decoding part and does not depart from boosting’s blueprint complexity, save of course the normalization of boosting’s distribution that we do not need to perform, unlike AdaBoost\. Onlyε𝒑𝒒,ϱ𝒑𝒒\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\},\\varrho\_\{\\bm\{p\}\\bm\{q\}\}in \([ED](https://arxiv.org/html/2609.30474#S4.Ex18)\) need to be computed in addition\. For any given𝒑,𝒒\\bm\{p\},\\bm\{q\}, the complexity isO\(n\)O\(n\)\. This is no more than the computation of the resampling probability \(Subsection[3\.1](https://arxiv.org/html/2609.30474#S3.SS1)\), yet it also gets in the computation of acceptance probabilities and thus brings an additional computation cost when accepting tokens\. This, of course, can be reduced, e\.g\. by quantization of the vectors\. We now investigate the mentored decoding side, which has several non\-trivial and very useful properties from an algorithmic standpoint\.
Algorithm 2CBreakpoints\(r,amax,bmax\)\(\\textsc\{r\},a\_\{\\max\},b\_\{\\max\}\)Input:sorted array
r=\.\{\(pi/qi,qi\)\}i=1n\\textsc\{r\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\\{\(p\_\{i\}/q\_\{i\},q\_\{i\}\)\\\}\_\{i=1\}^\{n\}in increasing order of the ratio
p\./q\.p\_\{\.\}/q\_\{\.\},
amax\>0a\_\{\\max\}\>0as in \([36](https://arxiv.org/html/2609.30474#S3.E36)\),
bmax\>0b\_\{\\max\}\>0as in \([37](https://arxiv.org/html/2609.30474#S3.E37)\);
Output:sequence of breakpoints
c\(𝒑,𝒒\)\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\);
Step 1 : compute initial values
1\.1 :
i←min\{k:rk0\>1\}i\\leftarrow\\min\\\{k:r\_\{k0\}\>1\\\}; //
rk0=\.r\_\{k0\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}coordinate 0 of couple
\#k\\\#kinr
1\.2 :
j←max\{k:rk0<1\}j\\leftarrow\\max\\\{k:r\_\{k0\}<1\\\};
1\.3 :
qA←∑k≥irk1q\_\{A\}\\leftarrow\\sum\_\{k\\geq i\}r\_\{k1\}; //
rk1=\.r\_\{k1\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}coordinate 1 of couple
\#k\\\#kinr
1\.4 :
qB←∑k≤jrk1q\_\{B\}\\leftarrow\\sum\_\{k\\leq j\}r\_\{k1\};
1\.5 :
a←0a\\leftarrow 0;
1\.6 :
b←0b\\leftarrow 0;
1\.7 :
c←∅\\textsc\{c\}\\leftarrow\\emptyset;
Step 2 :UpdateCBreakpoints\(r,c,a,b,i,j,qA,qB,amax,bmax\)\(\\textsc\{r\},\\textsc\{c\},a,b,i,j,q\_\{A\},q\_\{B\},a\_\{\\max\},b\_\{\\max\}\);
Step 3 :returnc; //
c\(𝒑,𝒒\)\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\)
Figure 4:In theΔ3\\Delta\_\{3\}simplex \(same setting as Figure[1](https://arxiv.org/html/2609.30474#S3.F1)\), we illustrate Theorem[3\.13](https://arxiv.org/html/2609.30474#S3.Thmtheorem13)and Lemma[5\.2](https://arxiv.org/html/2609.30474#S5.Thmtheorem2)\. The thick dark line is the set of mentored distributions built from queryingQueryCBreakpointsin Algorithm[3](https://arxiv.org/html/2609.30474#alg3)\. The bend represents a mentored distribution corresponding to a breakpoint inc\(𝒑,𝒒\)\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\)\. We display the couple\(a,b\)\(a,b\)corresponding to the mentored distribution𝝅\\bm\{\\pi\}\. Finally, for a set of fourff\-divergences \(rKL = reverse KL\), we display that𝝅\\bm\{\\pi\}is also the optimal solution in allmdf\(𝒑,𝒒;\.\)\\textsc\{md\}\_\{f\}\(\\bm\{p\},\\bm\{q\};\.\)because it is the intersection between the TV\-ball defining the optimal objective and theff\-divergence ball defining the constraint activated inff\-MD\.#### On computing and querying𝒞\(𝒑,𝒒\)\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\)in Definition[3\.11](https://arxiv.org/html/2609.30474#S3.Thmtheorem11)
Theorem[3\.12](https://arxiv.org/html/2609.30474#S3.Thmtheorem12)shows that𝒞\(𝒑,𝒒\)\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\)is key to solvingff\-MD\. Given outputs𝒑,𝒒\\bm\{p\},\\bm\{q\}of the drafter and target, we show how to build aO\(n\)O\(n\)\-sized data structure that we callbreakpoints,c\(𝒑,𝒒\)\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\), inO\(sort\(n\)\)O\(\\mathrm\{sort\}\(n\)\)time\. Such breakpoints are the cornerstone of our approach to get the desired elements of𝒞\(𝒑,𝒒\)\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\)\. Quite remarkably, the data structuredoes not depend onff, and can thus be used for any applicableff\([2](https://arxiv.org/html/2609.30474#S3.E2)\) afterwards\. The breakpoints are couples\(a,b\)∈𝒞\(𝒑,𝒒\)\(a,b\)\\in\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\)\. Apart from the speculative decoding solution for whicha=b=0a=b=0, all other couples ofc\(𝒑,𝒒\)\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\)satisfy the invariant that at least one ofaaandbbdepends on a ratiop\./q\.p\_\{\.\}/q\_\{\.\}\. Each of such ratios being uniquely present at the exclusion of at least one extreme ratio, the cardinalNNofc\(𝒑,𝒒\)\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\)satisfiesN≤nN\\leq n\.
Algorithm 3QueryCBreakpoints\(c,a~\)\(\\textsc\{c\},\\tilde\{a\}\)Input:breakpoints list
c=\.c\(𝒑,𝒒\)=\.\{\(ai,bi\):i∈\[N\]\}\\textsc\{c\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\\{\(a\_\{i\},b\_\{i\}\):i\\in\[N\]\\\}in strict increasing order of both coordinates, query
0≤a~≤aN0\\leq\\tilde\{a\}\\leq a\_\{N\};
Output:
b~\\tilde\{b\}such that
\(a~,b~\)∈𝒞\(𝒑,𝒒\)\(\\tilde\{a\},\\tilde\{b\}\)\\in\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\);
Step 1 : find
iisuch that
ai≤a~≤ai\+1a\_\{i\}\\leq\\tilde\{a\}\\leq a\_\{i\+1\};
Step 2: // compute
b~\\tilde\{b\}b~\\displaystyle\\tilde\{b\}←\\displaystyle\\leftarrowb\+\(a~−a\)⋅q\(𝔸i\)q\(𝔹i\);//𝔸i,𝔹iare as in \([33](https://arxiv.org/html/2609.30474#S3.E33)\), \([35](https://arxiv.org/html/2609.30474#S3.E35)\) for breakpoint\(ai,bi\)\\displaystyle b\+\\frac\{\(\\tilde\{a\}\-a\)\\cdot q\(\\mathbb\{A\}\_\{i\}\)\}\{q\(\\mathbb\{B\}\_\{i\}\)\};\\quad\\mbox\{// $\\mathbb\{A\}\_\{i\},\\mathbb\{B\}\_\{i\}$ are as in \\eqref\{defAbis\}, \\eqref\{defBbis\} for breakpoint $\(a\_\{i\},b\_\{i\}\)$\}\(66\)
Step 3 :return
b~\\tilde\{b\};
###### Theorem 5\.1\.
The set of breakpointsc\(𝐩,𝐪\)\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\)returned byCBreakpointsin Algorithm[2](https://arxiv.org/html/2609.30474#alg2)satisfyc\(𝐩,𝐪\)⊆𝒞\(𝐩,𝐪\)\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\)\\subseteq\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\)\.
Proof in Appendix, Section[VIII\.11](https://arxiv.org/html/2609.30474#S8.SS11)\. It is clear fromCBreakpointsthat all breakpoints returned are ordered in strictly increasing values of both coordinatesaaandbb\. So let us denotec\(𝒑,𝒒\)=\.\{\(ai,bi\):i∈\[N\]\}\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\\{\(a\_\{i\},b\_\{i\}\):i\\in\[N\]\\\}, indexing its elements to reflect the order\.
###### Lemma 5\.2\.
For any0≤a~≤aN0\\leq\\tilde\{a\}\\leq a\_\{N\},b~\\tilde\{b\}returned byQueryCBreakpointsin Algorithm[3](https://arxiv.org/html/2609.30474#alg3)satisfies\(a~,b~\)∈𝒞\(𝐩,𝐪\)\(\\tilde\{a\},\\tilde\{b\}\)\\in\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\)\.
Proof in Appendix, Section[VIII\.12](https://arxiv.org/html/2609.30474#S8.SS12)\. We follow with a series of fundamental properties, most of which follow directly from Lemma[5\.2](https://arxiv.org/html/2609.30474#S5.Thmtheorem2)\.
###### Lemma 5\.3\.
For anyaasatisfying \([36](https://arxiv.org/html/2609.30474#S3.E36)\), there existsbbsuch that\(a,b\)∈𝒞\(𝐩,𝐪\)\(a,b\)\\in\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\)\. Reciprocally, for anybbsatisfying \([37](https://arxiv.org/html/2609.30474#S3.E37)\), there existsaasuch that\(a,b\)∈𝒞\(𝐩,𝐪\)\(a,b\)\\in\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\)\. Finally, the set of points\(a~,b~\)\(\\tilde\{a\},\\tilde\{b\}\)from Lemma[5\.2](https://arxiv.org/html/2609.30474#S5.Thmtheorem2)is a continuous strictly increasing function fromℝ\+\\mathbb\{R\}\_\{\+\}toℝ\+\{\\mathbb\{R\}\}\_\{\+\}\.
Proof in Appendix, Section[VIII\.13](https://arxiv.org/html/2609.30474#S8.SS13)\. Wheneverffis strictly convex, the optimal mentored distribution is unique and thus Lemmata[5\.2](https://arxiv.org/html/2609.30474#S5.Thmtheorem2)and[5\.3](https://arxiv.org/html/2609.30474#S5.Thmtheorem3)guarantee the exhaustiveness of our algorithms\. Ifffis not strictly convex, such as for the total variation divergence \(see Figure[4](https://arxiv.org/html/2609.30474#S5.F4)\), our algorithms elicit one of many solutions\.
#### Finding optimal mentored distributions
There are two ways to use𝒞\(𝒑,𝒒\)\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\)for optimal mentored distributions\. The first tackles solutions offf\-MD in \([f \-MD\-1](https://arxiv.org/html/2609.30474#S3.Ex2)\), in two steps: first, we find the successive indexesiiandi\+1i\+1inc\(𝒑,𝒒\)\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\)whose mentored distributions𝝅i\\bm\{\\pi\}\_\{i\},𝝅i\+1\\bm\{\\pi\}\_\{i\+1\}satisfyD∈\[Df\(𝝅i∥𝒒\),Df\(𝝅i\+1∥𝒒\)\]D\\in\[D\_\{f\}\(\\bm\{\\pi\}\_\{i\}\\\|\\bm\{q\}\),D\_\{f\}\(\\bm\{\\pi\}\_\{i\+1\}\\\|\\bm\{q\}\)\]\. Then, if necessary, we queryQueryCBreakpointsfor a dichotomic search of the optimum sought to desired precision\. In all cases, the whole complexity isO\(nlogn\)O\(n\\log n\)\. This simple algorithm gets to the optimum at arbitrary desired precision\. However, there is amuchcheaper way to get approximate solutions with guarantees, and it relies on the fast that if in addition to storing couples\(a,b\)\(a,b\),CBreakpointsalso keeps track of the correspondingqA,qBq\_\{A\},q\_\{B\}computed in the algorithm, then any stored quadruple\(a,b,qA,qB\)\(a,b,q\_\{A\},q\_\{B\}\)allows to computeDDin \([42](https://arxiv.org/html/2609.30474#S3.E42)\) inO\(1\)O\(1\)for any desiredu∈\[1−b,1\+a\]u\\in\[1\-b,1\+a\]\. Hence if instead of computing actual values one relies on the corresponding upperboundD^\\hat\{D\}ofDDthat follows \(for a conservative approach to approximation\), the whole procedure described above drops in complexity fromO\(nlogn\)O\(n\\log n\)toO\(log\(n\)\)O\(\\log\(n\)\)to get to the targetD^\\hat\{D\}and the corresponding parametersa,ba,b\. This, of course, is subject to the usefulness of \([42](https://arxiv.org/html/2609.30474#S3.E42)\) for such a goal\. The remark after Theorem[3\.12](https://arxiv.org/html/2609.30474#S3.Thmtheorem12)applies: sinceu∈\[1−b,1\+a\]u\\in\[1\-b,1\+a\]andf\(1\)=0f\(1\)=0, since curve\{\(a~,b~\)\}\\\{\(\\tilde\{a\},\\tilde\{b\}\)\\\}obtained from Lemma[5\.3](https://arxiv.org/html/2609.30474#S5.Thmtheorem3)is continuous, strictly increasing and contains\(0,0\)\(0,0\), there is always an interval\[1−b,1\+a\]\[1\-b,1\+a\]for which such an approach is useful\. What we also establish below, from a toy simulation standpoint and severalff\-divergences, is that usefulness can extend to a substantial range of acceptance probabilities \(Table[3](https://arxiv.org/html/2609.30474#S5.T3)\)\.
The second way to use𝒞\(𝒑,𝒒\)\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\)consists in tackling the dual problem offf\-MD, which is also interesting, especially if the drafter is good enough that we can constrain on the acceptance probability instead of the divergence to target\. In this problem, subject to a minimal acceptance probabilityP≥Pacc\(SD\)P\\geq P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\), one is required to find a mentored distribution with minimalff\-divergence to the target\. This problem admits much cheaper routines than forff\-MD, namelyO\(logn\)O\(\\log n\)solution to compute the optimal couple\(a,b\)\(a,b\)\(and thusO\(n\)O\(n\)to compute each of the optimal𝝅,𝒓,𝒔\\bm\{\\pi\},\\bm\{r\},\\bm\{s\}\)\. It consists in sandwiching the soughtPPbetween those of two successive indexesiiandi\+1i\+1inc\(𝒑,𝒒\)\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\)– sayPiP\_\{i\}andPi\+1P\_\{i\+1\}–, and then doing a simple intrapolation on the respective parameters\(ai,bi\)\(a\_\{i\},b\_\{i\}\)and\(ai\+1,bi\+1\)\(a\_\{i\+1\},b\_\{i\+1\}\)based on solving forβ\\betathe convex combinationP=β⋅Pi\+\(1−β\)⋅Pi\+1P=\\beta\\cdot P\_\{i\}\+\(1\-\\beta\)\\cdot P\_\{i\+1\}\. The fact that we can bypass anyff\-divergence computation because the solution is invariant to the choice offfis a direct application of Theorem[3\.13](https://arxiv.org/html/2609.30474#S3.Thmtheorem13)\. Notice that achievingO\(logn\)O\(\\log n\)is without algorithmic frills, but allowing a few \(lookup tables, hashtables, etc\.\) allows to bring it down toO\(1\)O\(1\)\.
#### Guaranteed cheap solutions with betterPaccP\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}and small divergence
We state a fundamental property on the function giving the value of the divergence thersholdDDas a function of the optimal acceptance probabilityPacc\(MD\)P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(MD\)inmdf2\(𝒑,𝒒,D\)\\textsc\{md\}^\{2\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)\([f \-MD\-2](https://arxiv.org/html/2609.30474#S3.Ex1)\):
Df\(P\)\\displaystyle D\_\{f\}\(P\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}D:∃\(𝒓,𝒔\)∈mdf2\(𝒑,𝒒,D\)s\.t\.𝒑⊤𝒓=P\\displaystyle D:\\exists\(\\bm\{r\},\\bm\{s\}\)\\in\\textsc\{md\}^\{2\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)\\mbox\{ s\.t\. \}\\bm\{p\}^\{\\top\}\\bm\{r\}=P\(67\)\(parameters𝒑,𝒒\\bm\{p\},\\bm\{q\}are left implicit from context\)\. For any functionggfor which it exists,gr′g^\{\\prime\}\_\{r\}denotes the right derivative\.
###### Theorem 5\.4\.
For any convexff\([2](https://arxiv.org/html/2609.30474#S3.E2)\), the right derivative ofDf\(P\)D\_\{f\}\(P\)inPacc\(SD\)P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)exists and satisfies
\(Df\)r′\(Pacc\(SD\)\)\\displaystyle\(D\_\{f\}\)^\{\\prime\}\_\{r\}\(P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)\)=\\displaystyle=max∂f\(1\)−min∂f\(1\)\.\\displaystyle\\max\\partial f\(1\)\-\\min\\partial f\(1\)\.\(68\)Furthermore,Df\(P\)D\_\{f\}\(P\)is convex, strictly so iffffis strictly convex\.
Proof in Appendix, Section[VIII\.14](https://arxiv.org/html/2609.30474#S8.SS14)\. Supposeffdifferentiable inz=1z=1\. Then \([68](https://arxiv.org/html/2609.30474#S5.E68)\) crucially gives
\(Df\)r′\(Pacc\(SD\)\)\\displaystyle\(D\_\{f\}\)^\{\\prime\}\_\{r\}\(P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)\)=\\displaystyle=0,\\displaystyle 0,\(69\)and so foranysuch divergence, the neighborhood of the minimal acceptance probability will have divergence close to zero: depending on theff\-divergence, it may be possible to get a substantial increase of the acceptance probability at a low cost divergence\-wise\.
NameGeneratorf\(z\)f\(z\)CommentsKullback\-Leibler \(KL\)zlogzz\\log zreverse Kullback\-Leibler \(rKL\)−logz\-\\log zHellinger1−z1\-\\sqrt\{z\}Neyman\(z−1\)2\(z\-1\)^\{2\}reverseχ2\\chi^\{2\}Pearson\(z−1\)2/z\(z\-1\)^\{2\}/zχ2\\chi^\{2\}TV2TV\_\{2\}2⋅max\{fTV\(z\),2⋅\|z−1\|−1\}2\\cdot\\max\\left\\\{f\_\{\\mathrm\{TV\}\}\(z\),2\\cdot\|z\-1\|\-1\\right\\\}Amari\(α\\alpha\)\(zα−αz\+α−1\)/\(α\(α−1\)\)\(z^\{\\alpha\}\-\\alpha z\+\\alpha\-1\)/\(\\alpha\(\\alpha\-1\)\)α∈ℝ\\alpha\\in\\mathbb\{R\}Table 1:ff\-divergences used in our toy simulations; we recallfTV\(z\)=\.\|z−1\|/2f\_\{\\mathrm\{TV\}\}\(z\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\|z\-1\|/2\(see text\)\.
#### Toy simulation
We made a simple simulation of uniform𝒑,𝒒∈Δn\\bm\{p\},\\bm\{q\}\\in\\Delta\_\{n\}forn=100n=100\. Table[2](https://arxiv.org/html/2609.30474#S5.T2)presents results obtained on these distributions\. The left plot exemplifies Lemma[5\.3](https://arxiv.org/html/2609.30474#S5.Thmtheorem3)showing the strict monotonicity and continuity of the set of points returned byQueryCBreakpointsin Algorithm[3](https://arxiv.org/html/2609.30474#alg3)\. The right plot displays the correspondingff\-divergence as a function of the acceptance probability of mentored decoding,Pacc\(MD\)P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(MD\)\. From top to bottom in the legend, the generators of theff\-divergences are: Kullback\-Leibler, reverse Kullback\-Leibler, Hellinger, Neyman, Pearson, a scaling of the total variation,TV2TV\_\{2\}, which replaces parts of the TV divergence generator by steeper segments and half lines, and finally two instances of Amariα\\alpha\-divergences forα∈\{−1\.5,1\.5\}\\alpha\\in\\\{\-1\.5,1\.5\\\}\([Amari and Nagaoka, 2000](https://arxiv.org/html/2609.30474#bib.bib3)\)\(Table[1](https://arxiv.org/html/2609.30474#S5.T1)presents the associated generators\)\. This plots clearly exemplifies the importance of Theorem[5\.4](https://arxiv.org/html/2609.30474#S5.Thmtheorem4), as forallgenerators differentiable inz=1z=1,Pacc\(SD\)P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)can be increased by more than10%10\\%at negligible divergence cost\. For some divergences, such as theχ2\\chi^\{2\}, the divergence blows up at some point\. This, of course, ultimately depends on𝒑,𝒒\\bm\{p\},\\bm\{q\}andff\. Table[3](https://arxiv.org/html/2609.30474#S5.T3)takes all theff\-divergence curves and add the interval of possibleDDvalues of \([42](https://arxiv.org/html/2609.30474#S3.E42)\) in Theorem[3\.12](https://arxiv.org/html/2609.30474#S3.Thmtheorem12), in between the min and max values ofDDasuuranges in\[1−a,1\+b\]\[1\-a,1\+b\], for a range of acceptance probability for MD that ranges in betweenPacc\(SD\)P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)andPacc\(SD\)P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)plus 25%\\%\. Remark that the bound is quite crude for some divergences \(KL, rKL, Hellinger\) but can be quite informative on the trueDDfor the others \(Neyman, Pearson,TV2TV\_\{2\}, Amari’sα\\alpha\-divergence\) even for a substantial increase of the acceptance probability past SD’s\. This simple experiment demonstrates the potential usefulness of Theorem[3\.12](https://arxiv.org/html/2609.30474#S3.Thmtheorem12)for approximate solutions to \([f \-MD\-1](https://arxiv.org/html/2609.30474#S3.Ex2)\) as described above\.
Table[4](https://arxiv.org/html/2609.30474#S5.T4)further digs into the guarantees of Theorem[3\.12](https://arxiv.org/html/2609.30474#S3.Thmtheorem12), showing on this example how theff\-divergence boundDDand optimal acceptance probability of mentored decoding inmdf2\(𝒑,𝒒,D\)\\textsc\{md\}^\{2\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)\([f \-MD\-2](https://arxiv.org/html/2609.30474#S3.Ex1)\) vary as a function of parameters\(a~,b~\)\(\\tilde\{a\},\\tilde\{b\}\)in𝒞\(𝒑,𝒒\)\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\)extracted byQueryCBreakpointsin Algorithm[3](https://arxiv.org/html/2609.30474#alg3)\. Finally, Table[5](https://arxiv.org/html/2609.30474#S5.T5)computes the key coefficientkk\([64](https://arxiv.org/html/2609.30474#S4.E64)\) that governs our boosting bound in Theorem[4\.8](https://arxiv.org/html/2609.30474#S4.Thmtheorem8)\. It shows, in this simulated case, that one can easily increase the acceptance probability by more than10%10\\%and still keepkkvery close to 1, which is good news for the boosting bound \([65](https://arxiv.org/html/2609.30474#S4.E65)\)\. Interestingly also, the dependence ofkkinaais close to being linear\.
![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/sim-a-b.png)![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/sim-pacc-fdiv.png)Table 2:Left:curve of points\(a~,b~\)\(\\tilde\{a\},\\tilde\{b\}\)in𝒞\(𝒑,𝒒\)\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\)from Lemma[5\.3](https://arxiv.org/html/2609.30474#S5.Thmtheorem3), showing it is strictly increasing and continuous\.Right:curves showing the correspondingff\-divergence boundDDin \([f \-MD\-2](https://arxiv.org/html/2609.30474#S3.Ex1)\) \([f \-MD\-1](https://arxiv.org/html/2609.30474#S3.Ex2)\) as a function of the optimal acceptance probabilityPacc\(MD\)P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(MD\), for severalff\-divergences\. Remark that for all exceptTV2TV\_\{2\}, the right derivative atPacc\(SD\)P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)is indeed zero and the divergence stays close to 0 even past\>10%\>10\\%increase acceptance probabilityPacc\(MD\)P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(MD\)\(see text\)\.![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/new-thm312-KL.png)![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/new-thm312-rKL.png)![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/new-thm312-Hellinger.png)![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/new-thm312-Neyman.png)Kullback\-Leibler \(KL\)reverse KLHellingerNeyman![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/new-thm312-Pearson.png)![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/new-thm312-TV-2.png)![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/new-thm312-Amari-1-5.png)![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/new-thm312-Amari-m-1-5.png)PearsonTV2TV\_\{2\}Amari \(1\.51\.5\)Amari \(−1\.5\-1\.5\)Table 3:For each of theff\-divergence plots in Table[2](https://arxiv.org/html/2609.30474#S5.T2)\(Right\), we plot the divergence as a function ofPacc\(MD\)P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(MD\), and in filledgreenthe interval of min and max values corresponding to \([42](https://arxiv.org/html/2609.30474#S3.E42)\) in Theorem[3\.12](https://arxiv.org/html/2609.30474#S3.Thmtheorem12), over a range ofPacc\(MD\)P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(MD\)which coversPacc\(SD\)P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)\(reddot\) toPacc\(SD\)P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)increased by 25%\\%\(see text\)\.![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/sim-a-fdiv.png)![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/sim-b-fdiv.png)![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/sim-a-pacc.png)![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/sim-b-pacc.png)Table 4:Top:plots of theff\-divergence boundDDinmdf\(𝒑,𝒒,D\)\\textsc\{md\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)as a function ofaa\(Left\) andbb\(Right\), exemplifying Theorem[3\.12](https://arxiv.org/html/2609.30474#S3.Thmtheorem12)\.Bottom:plots of the optimal acceptance probabilityPacc\(MD\)P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(MD\)as a function ofaa\(Left\) andbb\(Right\), exemplifying Theorem[3\.12](https://arxiv.org/html/2609.30474#S3.Thmtheorem12)\(see text\)\.![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/sim-pacc-k.png)![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/sim-a-k.png)Table 5:Left:curve giving the boosting bound coefficientk=\.\(1\+a\)/\(1−b\)k\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\sqrt\{\(1\+a\)/\(1\-b\)\}\(Theorem[4\.8](https://arxiv.org/html/2609.30474#S4.Thmtheorem8)\) as a function of the acceptance probabilityPacc\(MD\)P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(MD\); remark thatkkremains close to11even for substantial increase of the acceptance probability compared to the speculative decoding solution\.Right:curve giving coefficientkkas a function ofaa\(see text\)\.We have also performed a second simulation in which the objective was to help visualize the mentored distribution\. In Table[6](https://arxiv.org/html/2609.30474#S5.T6), we have computed𝒑,𝒒∈Δn\\bm\{p\},\\bm\{q\}\\in\\Delta\_\{n\}forn=1000n=1000both following discretized Beta distributions, and ordered in thexxaxis the indexes in increasingq\./p\.q\_\{\.\}/p\_\{\.\}ratios\. Then, we have computed, from top\-left to bottom\-right, the mentored distribution \(thick purple curve\) by putting in evidence functioni↦min\{πi,pi\}i\\mapsto\\min\\\{\\pi\_\{i\},p\_\{i\}\\\}whose area \(purple\) is the acceptance probability and theff\-divergence is Kullback\-Leibler, for values ofb=\{0,0\.1,…,1\.0\}b=\\\{0,0\.1,\.\.\.,1\.0\\\}in Definition[3\.11](https://arxiv.org/html/2609.30474#S3.Thmtheorem11)\. We can clearly identify and track the three sets of indices defining sets \([33](https://arxiv.org/html/2609.30474#S3.E33)\), \([34](https://arxiv.org/html/2609.30474#S3.E34)\), \([35](https://arxiv.org/html/2609.30474#S3.E35)\)\.
![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/b-0.png)![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/b-01.png)![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/b-02.png)![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/b-03.png)b=0b=0\(Pacc=0\.219P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}=0\.219\)0\.1 \(0\.3070\.307\)0\.2 \(0\.3950\.395\)0\.3 \(0\.4830\.483\)![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/b-04.png)![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/b-05.png)![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/b-06.png)![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/b-07.png)b=0\.4b=0\.4\(Pacc=0\.567P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}=0\.567\)0\.5 \(0\.6510\.651\)0\.6 \(0\.7320\.732\)0\.7 \(0\.8110\.811\)![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/b-08.png)![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/b-09.png)![[Uncaptioned image]](https://arxiv.org/html/2609.30474v1/Figs/b-1.png)b=0\.8b=0\.8\(Pacc=0\.885P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}=0\.885\)0\.9 \(0\.9510\.951\)1 \(11\)Table 6:Mentored distribution \(thickpurplecurve\) and acceptance probability \(purplearea\) when𝒑\\bm\{p\}\(thinbluecurve\) and𝒒\\bm\{q\}\(thinredcurve\) are discretized from Beta distributions\. The values under each plot arebbandPacc\(MD\)P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(MD\)\(in parentheses, see text\)\.
## 6Discussion
In this section, we discuss the case where we want to enforce the decoding of only the top\-kktokens of the target, and two additional points on boosting\. The relevance of the discussion on boosting extends beyond the mentored decoding case\.
#### Efficient restriction to top\-kkdecoding
Suppose we maskn−kn\-ktokens on𝒒\\bm\{q\}by replacing the coordinates by00, e\.g\. to single out the top\-kkcoordinates of𝒒\\bm\{q\}, withk≥1k\\geq 1\. This is a practically relevant setting that clearly breaks Assumption[3\.3](https://arxiv.org/html/2609.30474#S3.Thmtheorem3)\. For the sake of readability, we assume the remaining token coordinates of𝒒\\bm\{q\}are renormalized in the simplex so we do not need to overload our MD problems with new parameters: what happens for the solutions of our mentored decoding problems \([f \-MD\-2](https://arxiv.org/html/2609.30474#S3.Ex1)\), \([f \-MD\-1](https://arxiv.org/html/2609.30474#S3.Ex2)\) ? Note that\(qi/πi\)⋅f\(πi/qi\)=f\(u\)/u\(q\_\{i\}/\\pi\_\{i\}\)\\cdot f\(\\pi\_\{i\}/q\_\{i\}\)=f\(u\)/uwithu=\.πi/qiu\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\pi\_\{i\}/q\_\{i\}\. Ifqi=0q\_\{i\}=0butπi≠0\\pi\_\{i\}\\neq 0, the expression takes the limitlimz→\+∞f\(z\)/z\\lim\_\{z\\rightarrow\+\\infty\}f\(z\)/z\. That is why we generalize the definition of aff\-divergence in \([1](https://arxiv.org/html/2609.30474#S3.E1)\) to the possibility that someqi=0q\_\{i\}=0by following[Csiszár \(1972\)](https://arxiv.org/html/2609.30474#bib.bib22)and letting
Df\(𝝅∥𝒒\)\\displaystyle D\_\{f\}\(\\bm\{\\pi\}\\\|\\bm\{q\}\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}∑i:qi\>0qif\(πiqi\)\+∑i:qi=0∧πi\>0πi⋅limz→\+∞f\(z\)z\.\\displaystyle\\sum\_\{i:q\_\{i\}\>0\}q\_\{i\}f\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)\+\\sum\_\{i:q\_\{i\}=0\\wedge\\pi\_\{i\}\>0\}\\pi\_\{i\}\\cdot\\lim\_\{z\\rightarrow\+\\infty\}\\frac\{f\(z\)\}\{z\}\.\(70\)For clarity, we explicitly discard the problematic cases where bothπi=qi=0\\pi\_\{i\}=q\_\{i\}=0: we treat it as the limit caselimz→1f\(z\)=f\(1\)=0\\lim\_\{z\\rightarrow 1\}f\(z\)=f\(1\)=0\(ffis convex, thus continuous\) per \([2](https://arxiv.org/html/2609.30474#S3.E2)\)\.
###### Lemma 6\.1\.
Suppose𝐪\\bm\{q\}has been masked to a set of1≤k<n1\\leq k<ncoordinates\. Then there always exists an optimal solution𝛑\\bm\{\\pi\}of \([f \-MD\-1](https://arxiv.org/html/2609.30474#S3.Ex2)\) whose support is thosekkcoordinates\.
Proof in Appendix, Section[VIII\.15](https://arxiv.org/html/2609.30474#S8.SS15)\. By Lemma[3\.2](https://arxiv.org/html/2609.30474#S3.Thmtheorem2), there also exists an optimal solution\(𝒓,𝒔\)∈mdf2\(𝒑,𝒒,D\)\(\\bm\{r\},\\bm\{s\}\)\\in\\textsc\{md\}^\{2\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)whose coordinates satisfyqi=0⇒ri=0∧si=0,∀i∈\[n\]q\_\{i\}=0\\Rightarrow r\_\{i\}=0\\wedge s\_\{i\}=0,\\forall i\\in\[n\]\. We stress the substantial practical importance of Lemma[6\.1](https://arxiv.org/html/2609.30474#S6.Thmtheorem1): when masking the target, instead of solving MD over the potentially huge set ofnncoordinates / tokens \(e\.g\.n≈105n\\approx 10^\{5\}\), we can restrict MD over the subset defining the mask \(e\.g\.k=16k=16\)\. Also, this does not affect the connection with boosting but must be appliedmutatis mutandisfor the parameters involved\.
The practical consequence of Lemma[6\.1](https://arxiv.org/html/2609.30474#S6.Thmtheorem1)is significant for modern LLM serving\. In production pipelines, target verification is often executed with top\-kktruncation\. When it is the case,identifying the top\-kktokens and renormalizing their probabilities is already performed by the baseline speculative decoding verification stage\.Consequently, mentored decoding incurszero additional overheadfor top\-kkextraction and the sorting complexity drops fromO\(nlogn\)O\(n\\log n\)toO\(klogk\)O\(k\\log k\)\. Hence, mentored decoding achieves higher acceptance rates with negligible wall\-clock overhead\.
#### Boosting the boosting advantages beyond \([WLA](https://arxiv.org/html/2609.30474#S4.Ex17)\)
The weak learning assumption has been instrumental in showing that boosting effectively works by amplifying the performances of models barely better than random\. Here, we show that, if instead of absolute performances we focus on relative performances of the weak models, i\.e\. correlations between each other, then there is a similar amplification framework which, instead of providing a boosting advantage linear in the number of modelsTT, gets a boosting advantage which is exponential inTT\. This result is facilitated in our case \(vs AdaBoost\) because weights in \([45](https://arxiv.org/html/2609.30474#S4.E45)\) are automatically normalized in the simplex\. Denote𝒉~t=\.\(1/\(2ht,∞\)\)⋅𝒉t\\tilde\{\\bm\{h\}\}\_\{t\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\(1/\(2h\_\{t,\\infty\}\)\)\\cdot\\bm\{h\}\_\{t\}, so we have\|𝒚i⊤𝒉~t\|≤1\|\\bm\{y\}\_\{i\}^\{\\top\}\\tilde\{\\bm\{h\}\}\_\{t\}\|\\leq 1fori∈\[m\]i\\in\[m\]\. For anyt≥1t\\geq 1and any sequence†††To spare notations, we write for exampleℐ=1,2,5\\mathcal\{I\}=1,2,5without other symbol\.of integersℐ≥t\\mathcal\{I\}\\geq t\(element\-wise\), we let
𝔼t\(ℐ\)\\displaystyle\\mathbb\{E\}\_\{t\}\(\\mathcal\{I\}\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}∑i∈\[m\]wti⋅∏j∈ℐ𝒚i⊤𝒉~j\(𝒙i\)\.\\displaystyle\\sum\_\{i\\in\[m\]\}w\_\{ti\}\\cdot\\prod\_\{j\\in\\mathcal\{I\}\}\\bm\{y\}\_\{i\}^\{\\top\}\\tilde\{\\bm\{h\}\}\_\{j\}\(\\bm\{x\}\_\{i\}\)\.\(71\)We can unravel its formula with respect to the weight index, from \([46](https://arxiv.org/html/2609.30474#S4.E46)\) and \([45](https://arxiv.org/html/2609.30474#S4.E45)\), into a very useful formula:
𝔼t\+1\(ℐ\)\\displaystyle\\mathbb\{E\}\_\{t\+1\}\(\\mathcal\{I\}\)=\\displaystyle=∑i∈\[m\]wti⋅1−μt⋅𝒚i⊤𝒉~t\(𝒙i\)1−μt2⋅∏j∈ℐ𝒚i⊤𝒉~j\(𝒙i\)\.\\displaystyle\\sum\_\{i\\in\[m\]\}w\_\{ti\}\\cdot\\frac\{1\-\\mu\_\{t\}\\cdot\\bm\{y\}\_\{i\}^\{\\top\}\\tilde\{\\bm\{h\}\}\_\{t\}\(\\bm\{x\}\_\{i\}\)\}\{1\-\\mu\_\{t\}^\{2\}\}\\cdot\\prod\_\{j\\in\\mathcal\{I\}\}\\bm\{y\}\_\{i\}^\{\\top\}\\tilde\{\\bm\{h\}\}\_\{j\}\(\\bm\{x\}\_\{i\}\)\.\(72\)=\\displaystyle=𝔼t\(ℐ\)−𝔼t\(t\)⋅𝔼t\(t,ℐ\)1−𝔼t2\(t\)\.\\displaystyle\\frac\{\\mathbb\{E\}\_\{t\}\(\\mathcal\{I\}\)\-\\mathbb\{E\}\_\{t\}\(t\)\\cdot\\mathbb\{E\}\_\{t\}\(t,\\mathcal\{I\}\)\}\{1\-\\mathbb\{E\}^\{2\}\_\{t\}\(t\)\}\.Suppose we replace the content of \([WLA](https://arxiv.org/html/2609.30474#S4.Ex17)\) by the following\. First, assumeμt\>0\\mu\_\{t\}\>0, a lightweight assumption since otherwise the chosen hypothesis does not perform better than random\. We also add assumptions fort≥1t\\geq 1that𝒉~t\+1\\tilde\{\\bm\{h\}\}\_\{t\+1\}is not too bad with respect to𝒉~t\\tilde\{\\bm\{h\}\}\_\{t\}while being different enough from𝒉~t\\tilde\{\\bm\{h\}\}\_\{t\}, which is especially relevant for large models\. These are grouped in a setting called Weak Correlation Assumption\.
###### Lemma 6\.2\.
LetT\>1T\>1be a number of boosting iterations and assume the following Weak Correlation Assumption holds:
∃β≥0,δ\>0:\{\(1\.\)μt\>0\(2\.\)𝔼t\(t\+1\)≥\(1−β\)μt\(3\.\)𝔼t\(t,t\+1\)≤1−β−\(1\+δ\)μt2,∀t=1,2,…T\.\\displaystyle\\exists\\beta\\geq 0,\\delta\>0:\\left\\\{\\begin\{array\}\[\]\{ll\}\(1\.\)&\\mu\_\{t\}\>0\\\\ \(2\.\)&\\mathbb\{E\}\_\{t\}\(t\+1\)\\geq\(1\-\\beta\)\\mu\_\{t\}\\\\ \(3\.\)&\\mathbb\{E\}\_\{t\}\(t,t\+1\)\\leq 1\-\\beta\-\(1\+\\delta\)\\mu\_\{t\}^\{2\}\\end\{array\}\\right\.,\\forall t=1,2,\.\.\.T\.\(WCA\)Then the boosting advantage grows exponentially withTTas:
A\(\{𝒉t\}t∈\[T\]\)\\displaystyle A\\left\(\\\{\\bm\{h\}\_\{t\}\\\}\_\{t\\in\[T\]\}\\right\)≥\\displaystyle\\geqμ12⋅\(1\+δ\)2T−12δ\+δ2\.\\displaystyle\\mu\_\{1\}^\{2\}\\cdot\\frac\{\(1\+\\delta\)^\{2T\}\-1\}\{2\\delta\+\\delta^\{2\}\}\.\(76\)
Proof in Appendix, Section[VIII\.16](https://arxiv.org/html/2609.30474#S8.SS16)\. \([WCA](https://arxiv.org/html/2609.30474#S6.Ex21)\) guarantees much better rates than \([WLA](https://arxiv.org/html/2609.30474#S4.Ex17)\),butit can typically hold for a much more limited number of iterations\. Indeed, the crux of \([WCA](https://arxiv.org/html/2609.30474#S6.Ex21)\) is to imply a geometric increase in edges,μt\+1≥\(1\+δ\)μt\\mu\_\{t\+1\}\\geq\(1\+\\delta\)\\mu\_\{t\}, and we obviously observeμt\+1≤1\\mu\_\{t\+1\}\\leq 1\. Yet, let us compare what the \([WLA](https://arxiv.org/html/2609.30474#S4.Ex17)\) and \([WCA](https://arxiv.org/html/2609.30474#S6.Ex21)\) can get after a maximal numberT∗T^\{\*\}of iterations for \([WCA](https://arxiv.org/html/2609.30474#S6.Ex21)\) to stand\. To simplify, assumeμ1≥γ\\mu\_\{1\}\\geq\\upgammaand pickδ=γ\\delta=\\upgammaof the \([WLA](https://arxiv.org/html/2609.30474#S4.Ex17)\)\. To ensureμt≤1,∀t\\mu\_\{t\}\\leq 1,\\forall t, we must haveT≤T∗=\.log\(1/γ\)/log\(1\+γ\)T\\leq T^\{\*\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\log\(1/\\upgamma\)/\\log\(1\+\\upgamma\)\.
After such a number of iterations, the boosting advantage in \([76](https://arxiv.org/html/2609.30474#S6.E76)\) satisfiesA\(\{𝒉t\}t∈\[T∗\]\)≥AWCAA\\left\(\\\{\\bm\{h\}\_\{t\}\\\}\_\{t\\in\[T^\{\*\}\]\}\\right\)\\geq A\_\{WCA\}with
AWCA\\displaystyle A\_\{WCA\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}γ2⋅\(1\+γ\)2γ2−12γ\+γ2=1\+2γ2γ\+γ2,\\displaystyle\\upgamma^\{2\}\\cdot\\frac\{\\frac\{\(1\+\\upgamma\)^\{2\}\}\{\\upgamma^\{2\}\}\-1\}\{2\\upgamma\+\\upgamma^\{2\}\}=\\frac\{1\+2\\upgamma\}\{2\\upgamma\+\\upgamma^\{2\}\},while the boosting advantage in \([51](https://arxiv.org/html/2609.30474#S4.E51)\) yields onlyA\(\{𝒉t\}t∈\[T∗\]\)≥AWLAA\\left\(\\\{\\bm\{h\}\_\{t\}\\\}\_\{t\\in\[T^\{\*\}\]\}\\right\)\\geq A\_\{WLA\}with
AWLA\\displaystyle A\_\{WLA\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}γ2⋅log\(1γ\)log\(1\+γ\),\\displaystyle\\upgamma^\{2\}\\cdot\\frac\{\\log\\left\(\\frac\{1\}\{\\upgamma\}\\right\)\}\{\\log\(1\+\\upgamma\)\},and it is not hard to show thatAWCA=Ω\(AWLA/γ\)A\_\{WCA\}=\\Omega\(A\_\{WLA\}/\\upgamma\), yielding a potential drop in the boosting rate dependenceO\(1/γ2\)O\(1/\\upgamma^\{2\}\)in \([52](https://arxiv.org/html/2609.30474#S4.E52)\) to the much more seldomO\(1/γ\)O\(1/\\upgamma\)under \([WCA](https://arxiv.org/html/2609.30474#S6.Ex21)\), which can be of independent interest in the context of boosting\([Alon et al\., 2023](https://arxiv.org/html/2609.30474#bib.bib15), Open problems\)but comes with the substantial caveat that the number of iterationsTTduring which the \([WCA](https://arxiv.org/html/2609.30474#S6.Ex21)\) can hold is substantially smaller than for \([WLA](https://arxiv.org/html/2609.30474#S4.Ex17)\)\.
#### Potential \(sub\)optimality of the greedy boosted sequence
The sequence of boosted classifiers is built iteratively, but the boosting advantage \([50](https://arxiv.org/html/2609.30474#S4.E50)\) is not invariant by permutation in the sequence\. Usually,𝒉t\\bm\{h\}\_\{t\}is the "best" classifier at iterationtt, say by maximizingμt\\mu\_\{t\}\. When we pick it from a pool of available classifiers, which is especially relevant in our case, a natural question comes as to whether this simple strategy always delivers the best boosting advantage\. A simple results shows that the greedy pick of the best classifier forμt\\mu\_\{t\}can, in a particular case highlighted below, lead to a suboptimal boosting advantage even from a very local standpoint, i\.e\. by just permuting two successive classifiers \(say𝒉t\\bm\{h\}\_\{t\}and𝒉t\+1\\bm\{h\}\_\{t\+1\}\) in the sequence\.
The reasoning is straightforward and comes directly from \([72](https://arxiv.org/html/2609.30474#S6.E72)\): sinceμt\+1=𝔼t\+1\(t\+1\)\\mu\_\{t\+1\}=\\mathbb\{E\}\_\{t\+1\}\(t\+1\)andμt=𝔼t\(t\)\\mu\_\{t\}=\\mathbb\{E\}\_\{t\}\(t\), we have:
μt\+1\\displaystyle\\mu\_\{t\+1\}=\\displaystyle=𝔼t\(t\+1\)−μt⋅𝔼t\(t,t\+1\)1−μt2\.\\displaystyle\\frac\{\\mathbb\{E\}\_\{t\}\(t\+1\)\-\\mu\_\{t\}\\cdot\\mathbb\{E\}\_\{t\}\(t,t\+1\)\}\{1\-\\mu\_\{t\}^\{2\}\}\.Denote𝒉d\\bm\{h\}\_\{d\}the hypothesis used inμt\+1\\mu\_\{t\+1\}and𝒉c\\bm\{h\}\_\{c\}the hypothesis used inμt\\mu\_\{t\}\. Under the best greedy fit scenario, we have
μt\>\|μ~t\|withμ~t=\.𝔼t\(d\)\\displaystyle\\mu\_\{t\}\>\|\\tilde\{\\mu\}\_\{t\}\|\\quad\\mbox\{ with \}\\tilde\{\\mu\}\_\{t\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\mathbb\{E\}\_\{t\}\(d\)\(77\)\(we remove the possibility of identity for simplicity\)\. The contribution to the boosting advantage of adding in this order𝒉c,𝒉d\\bm\{h\}\_\{c\},\\bm\{h\}\_\{d\}isμt2\+μt\+12\\mu^\{2\}\_\{t\}\+\\mu^\{2\}\_\{t\+1\}\. There is also the \(seemingly\) suboptimal scenario of preferring the sequence𝒉d,𝒉c\\bm\{h\}\_\{d\},\\bm\{h\}\_\{c\}, for an alternative contribution to the boosting advantageμ~t2\+μ~t\+12\\tilde\{\\mu\}^\{2\}\_\{t\}\+\\tilde\{\\mu\}^\{2\}\_\{t\+1\}withμ~t\+1=\.𝔼t\+1\(c\)\\tilde\{\\mu\}\_\{t\+1\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\mathbb\{E\}\_\{t\+1\}\(c\)\. Surprisingly perhaps, we show that there is a simple condition on thecovarianceof the two hypotheses such that the alternative scenario is strictly better than the best greedy fit\.
###### Lemma 6\.4\.
With the definition stated above, under the greedy choice condition \([77](https://arxiv.org/html/2609.30474#S6.E77)\), there exists0<ρ<0\.340<\\rho<0\.34depending onμt,μ~t\\mu\_\{t\},\\tilde\{\\mu\}\_\{t\}such thatμ~t2\+μ~t\+12\>μt2\+μt\+12\\tilde\{\\mu\}^\{2\}\_\{t\}\+\\tilde\{\\mu\}^\{2\}\_\{t\+1\}\>\\mu^\{2\}\_\{t\}\+\\mu^\{2\}\_\{t\+1\}iff one of the following holds:
- \(i\)μ~t\>0\\tilde\{\\mu\}\_\{t\}\>0and𝔼t\(c,d\)−𝔼t\(c\)𝔼t\(d\)∈\(0,ρ\)\\mathbb\{E\}\_\{t\}\(c,d\)\-\\mathbb\{E\}\_\{t\}\(c\)\\mathbb\{E\}\_\{t\}\(d\)\\in\(0,\\rho\), or
- \(ii\)μ~t<0\\tilde\{\\mu\}\_\{t\}<0and𝔼t\(c,d\)−𝔼t\(c\)𝔼t\(d\)∈\(−ρ,0\)\\mathbb\{E\}\_\{t\}\(c,d\)\-\\mathbb\{E\}\_\{t\}\(c\)\\mathbb\{E\}\_\{t\}\(d\)\\in\(\-\\rho,0\)\.
The proof, in Appendix, Section[VIII\.17](https://arxiv.org/html/2609.30474#S8.SS17), makesρ\\rhoexplicit\.
## 7Conclusion
There have been a number of recent approaches relaxing the key constraint of speculative decoding – that the output distribution be equal to the target’s\. While inference speedup was the original intent, a few recent papers also observed experimentally that the equivalent output model can sometimes beat the target when it comes to model quality\. In our paper, we have shown that such a remarkable feat – speeding up inference while getting a better model – is indeed possible\. Our two main bricks are mentored decoding as the formal setting authorizing deviations from the target, and boosting to evaluate the quality of the model produced by mentored decoding\. In the course of getting to this result, we derived several new key properties of the mentored decoding setting\. Among these, the particular geometric appeal of the total variation case is interesting for the variety of optimal solutions it supports, some of which are very convenient for boosting, but others might as well be relevant for other constraints\. We also reached an utterly simple approximation scheme of the optimal solutions for anyff\-divergence, also with interesting ties to boosting, which shows that there exists a data structure independent from the choice offf, but which, once computed, can be used for anyff\-divergence to get the two parameters to compute the optimal mentored decoding solution as fast as for speculative decoding\. Getting those parameters is done in logarithmic time via our breakpoint data structure\. While constructing the breakpoints involves sorting probability ratios, this overhead is practically negligible: under standard top\-kkdecoding, sorting operates over onlyk≪nk\\ll nelements already identified by the baseline pipeline, incurring negligible compute on accelerators\.
Another interesting avenue for future research relies on the boosting part of our paper\. The boosting part of our approach is efficient with respect to the canon of AdaBoost: it is self\-normalized and can be carried out with a bypass of boosting’s famous weight updates\. This latter property goes with computing linear correlation coefficients between models, which can be costly when the number of models increases, but at least for a few models it shows that the "architecture" of boosting does not necessarily need to be carved in the computation of the composite model producing the mentored distribution \(notwithstanding the risk of numerical approximation errors with weight updates in traditional \(Ada\)boosting\)\. Given the training cost of even the smallest LLM models, the LLM space – public or private – has plenty stored models for which boosting directly applies, but not the mentored decoding framework which originally applies to two models only\. Extending mentored decoding beyond the \(1 drafter, 1 target\) setting is an interesting question\.
## Acknowledgments
The authors thank Ariel Brand, Yishay Mansour, Nir Shabat and Ayala Shaubi\-Mann for early discussions on this material\.
## References
- Ali and Silvey \(1966\)S\. M\. Ali and S\. D\. SilveyA general class of coefficients of divergence of one distribution from another\.Journal of the Royal Statistical Society: Series B \(Methodological\)28\(1\),pp\. 131–142\.Cited by:[§3](https://arxiv.org/html/2609.30474#S3.p1.1)\.
- Alonet al\.\(2023\)N\. Alon, A\. Gonen, E\. Hazan, and S\. MoranBoosting simple learners\.TheoretiCS2\.External Links:[Link](https://doi.org/10.46298/theoretics.23.8),[Document](https://dx.doi.org/10.46298/THEORETICS.23.8)Cited by:[§6](https://arxiv.org/html/2609.30474#S6.SS0.SSS0.Px2.p3.3)\.
- Amari and Nagaoka \(2000\)S\. Amari and H\. NagaokaMethods of information geometry\.Oxford University Press\.Cited by:[§5](https://arxiv.org/html/2609.30474#S5.SS0.SSS0.Px4.p1.1)\.
- Amidet al\.\(2024\)E\. Amid, F\. Nielsen, R\. Nock, and M\. K\. WarmuthOptimal transport with tempered exponential measures\.InThirty\-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty\-Sixth Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2024, February 20\-27, 2024, Vancouver, Canada,M\. J\. Wooldridge, J\. G\. Dy, and S\. Natarajan \(Eds\.\),pp\. 10838–10846\.External Links:[Link](https://doi.org/10.1609/aaai.v38i10.28957),[Document](https://dx.doi.org/10.1609/AAAI.V38I10.28957)Cited by:[§4\.3](https://arxiv.org/html/2609.30474#S4.SS3.SSS0.Px1.p1.2),[§VIII\.8](https://arxiv.org/html/2609.30474#S8.SS8.p1.5)\.
- Amidet al\.\(2023\)E\. Amid, R\. Nock, and M\. K\. WarmuthClustering above exponential families with tempered exponential measures\.InInternational Conference on Artificial Intelligence and Statistics, 25\-27 April 2023, Palau de Congressos, Valencia, Spain,F\. J\. R\. Ruiz, J\. G\. Dy, and J\. van de Meent \(Eds\.\),Proceedings of Machine Learning Research, Vol\.206,pp\. 2994–3017\.External Links:[Link](https://proceedings.mlr.press/v206/amid23a.html)Cited by:[§4\.3](https://arxiv.org/html/2609.30474#S4.SS3.SSS0.Px1.p1.2),[§VIII\.8](https://arxiv.org/html/2609.30474#S8.SS8.p1.5)\.
- Bachmannet al\.\(2025\)G\. Bachmann, S\. Anagnostidis, A\. Pumarola, M\. Georgopoulos, A\. Sanakoyeu, Y\. Du, E\. Schönfeld, A\. Thabet, and J\. K\. KohlerJudge decoding: faster speculative sampling requires going beyond model alignment\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=mtSSFiqW6y)Cited by:[§2](https://arxiv.org/html/2609.30474#S2.p2.1)\.
- Bartlettet al\.\(1998\)P\. Bartlett, Y\. Freund, W\. S\. Lee, and R\. E\. SchapireBoosting the margin: a new explanation for the effectiveness of voting methods\.The Annals of Statistics26\(5\),pp\. 1651 – 1686\.External Links:[Document](https://dx.doi.org/10.1214/aos/1024691352),[Link](https://doi.org/10.1214/aos/1024691352)Cited by:[§1](https://arxiv.org/html/2609.30474#S1.p2.1),[§4\.1](https://arxiv.org/html/2609.30474#S4.SS1.p1.4),[§4\.2](https://arxiv.org/html/2609.30474#S4.SS2.p4.1),[Remark 4\.3](https://arxiv.org/html/2609.30474#S4.Thmtheorem3.p1.1.1)\.
- Byunet al\.\(2025\)S\. Byun, M\. Odema, J\. I\. Guack, B\. Lee, J\. Song, and W\. S\. Chung3\-model speculative decoding\.InNeurIPS 2025 Workshop on Structured Probabilistic Inference & Generative Modeling,External Links:[Link](https://openreview.net/forum?id=2e2RCF4Ncc)Cited by:[§2](https://arxiv.org/html/2609.30474#S2.p2.1)\.
- Caiet al\.\(2024\)T\. Cai, Y\. Li, Z\. Geng, H\. Peng, J\. D\. Lee, D\. Chen, and T\. DaoMedusa: simple llm inference acceleration framework with multiple decoding heads\.arXiv preprint arXiv:2401\.10774\.Cited by:[§1](https://arxiv.org/html/2609.30474#S1.p2.1),[§2](https://arxiv.org/html/2609.30474#S2.p1.1),[§2](https://arxiv.org/html/2609.30474#S2.p2.1)\.
- Chenet al\.\(2023\)C\. Chen, S\. Borgeaud, G\. Irving, J\. Lespiau, L\. Sifre, and J\. JumperAccelerating large language model decoding with speculative sampling\.External Links:2302\.01318Cited by:[§1](https://arxiv.org/html/2609.30474#S1.p1.1),[§2](https://arxiv.org/html/2609.30474#S2.p1.1)\.
- Chenet al\.\(2026\)J\. Chen, Y\. Liang, and Z\. LiuDFlash: block diffusion for flash speculative decoding\.arXiv preprint arXiv:2602\.06036\.Cited by:[§2](https://arxiv.org/html/2609.30474#S2.p1.1)\.
- Chenet al\.\(2024\)Z\. Chen, A\. May, R\. Svirschevski, Y\. Huang, M\. Ryabinin, Z\. Jia, and B\. ChenSequoia: scalable and robust speculative decoding\.Advances in Neural Information Processing Systems37,pp\. 129531–129563\.Cited by:[§2](https://arxiv.org/html/2609.30474#S2.p1.1)\.
- Csiszár \(1963\)I\. CsiszárEine informationstheoretische ungleichung und ihre anwendung auf den beweis der ergodizitat von Markoffschen ketten\.Magyar\. Tud\. Akad\. Mat\. Kutato Int\. Kozl\.8,pp\. 85–108\.Cited by:[§3](https://arxiv.org/html/2609.30474#S3.p1.1)\.
- Csiszár \(1972\)I\. CsiszárA class of measures of informativity of observation channels\.Periodica Mathematica Hungarica2,pp\. 191–213\.Cited by:[§6](https://arxiv.org/html/2609.30474#S6.SS0.SSS0.Px1.p1.1)\.
- Fuet al\.\(2024\)Y\. Fu, P\. Bailis, I\. Stoica, and H\. ZhangBreak the sequential dependency of llm inference using lookahead decoding\.arXiv preprint arXiv:2402\.02057\.Cited by:[§1](https://arxiv.org/html/2609.30474#S1.p2.1),[§2](https://arxiv.org/html/2609.30474#S2.p1.1)\.
- Gloeckleet al\.\(2024\)F\. Gloeckle, B\. Y\. Idrissi, B\. Rozière, D\. Lopez\-Paz, and G\. SynnaeveBetter & faster large language models via multi\-token prediction\.arXiv preprint arXiv:2404\.19737\.Cited by:[§2](https://arxiv.org/html/2609.30474#S2.p1.1)\.
- Hao and Mou \(2026\)Y\. Hao and L\. MouCactus: accelerating auto\-regressive decoding with constrained acceptance speculative sampling\.International Conference on Learning Representations \(ICLR\)\.Cited by:[§1](https://arxiv.org/html/2609.30474#S1.p2.1),[§2](https://arxiv.org/html/2609.30474#S2.p3.1)\.
- Heet al\.\(2024\)Z\. He, Z\. Zhong, T\. Cai, J\. Lee, and D\. HeRest: retrieval\-based speculative decoding\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 1582–1595\.Cited by:[§1](https://arxiv.org/html/2609.30474#S1.p2.1),[§2](https://arxiv.org/html/2609.30474#S2.p1.1)\.
- Holsmanet al\.\(2025\)M\. Holsman, Y\. Huang, and B\. DhingraFuzzy speculative decoding for a tunable accuracy\-runtime tradeoff\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 26257–26273\.Cited by:[§2](https://arxiv.org/html/2609.30474#S2.p2.1)\.
- Huet al\.\(2025\)Z\. Hu, T\. Zheng, V\. Viswanathan, Z\. Chen, R\. Rossi, Y\. Wu, D\. Manocha, and H\. HuangTowards optimal multi\-draft speculative decoding\.InInternational Conference on Learning Representations \(ICLR\),Vol\.2025,pp\. 3181–3203\.Cited by:[§2](https://arxiv.org/html/2609.30474#S2.p1.1)\.
- Kimet al\.\(2023\)S\. Kim, K\. Mangalam, S\. Moon, J\. Malik, M\. W\. Mahoney, A\. Gholami, and K\. KeutzerSpeculative decoding with big little decoder\.Advances in Neural Information Processing Systems36,pp\. 39236–39256\.Cited by:[§1](https://arxiv.org/html/2609.30474#S1.p2.1),[§2](https://arxiv.org/html/2609.30474#S2.p1.1)\.
- Knuth \(1992\)D\.\-E\. KnuthTwo notes on notation\.The American Mathematical Monthly99\(5\),pp\. 403–422\.Cited by:[§3\.3](https://arxiv.org/html/2609.30474#S3.SS3.p1.3),[§4\.3](https://arxiv.org/html/2609.30474#S4.SS3.SSS0.Px1.p1.2)\.
- Leviathanet al\.\(2023\)Y\. Leviathan, M\. Kalman, and Y\. MatiasFast inference from transformers via speculative decoding\.InInternational Conference on Machine Learning,pp\. 19274–19286\.Cited by:[§1](https://arxiv.org/html/2609.30474#S1.p1.1),[§1](https://arxiv.org/html/2609.30474#S1.p2.1),[§2](https://arxiv.org/html/2609.30474#S2.p1.1),[§2](https://arxiv.org/html/2609.30474#S2.p2.1),[§3](https://arxiv.org/html/2609.30474#S3.p1.1)\.
- Liet al\.\(2026a\)J\. Li, Y\. Xu, G\. Li, J\. Xu, S\. Yang, Y\. Zhang, X\. Yin, D\. Li, E\. C\. H\. Ngai, and E\. BarsoumBeyond the target: from imitation to collaboration in speculative decoding\.External Links:2605\.24793,[Link](https://arxiv.org/abs/2605.24793)Cited by:[§1](https://arxiv.org/html/2609.30474#S1.p2.1),[§2](https://arxiv.org/html/2609.30474#S2.p2.1),[§2](https://arxiv.org/html/2609.30474#S2.p4.1)\.
- Liet al\.\(2024a\)Y\. Li, F\. Wei, C\. Zhang, and H\. ZhangEagle\-2: faster inference of language models with dynamic draft trees\.InProceedings of the 2024 conference on empirical methods in natural language processing,pp\. 7421–7432\.Cited by:[§2](https://arxiv.org/html/2609.30474#S2.p1.1)\.
- Liet al\.\(2024b\)Y\. Li, F\. Wei, C\. Zhang, and H\. ZhangEAGLE: speculative sampling requires rethinking feature uncertainty\.International Conference on Machine Learning \(ICML\)\.Cited by:[§1](https://arxiv.org/html/2609.30474#S1.p2.1),[§2](https://arxiv.org/html/2609.30474#S2.p1.1)\.
- Liet al\.\(2026b\)Y\. Li, F\. Wei, C\. Zhang, and H\. ZhangEagle\-3: scaling up inference acceleration of large language models via training\-time test\.Advances in Neural Information Processing Systems38,pp\. 136737–136756\.Cited by:[§2](https://arxiv.org/html/2609.30474#S2.p1.1)\.
- Liaoet al\.\(2025\)B\. Liao, Y\. Xu, H\. Dong, J\. Li, C\. Monz, S\. Savarese, D\. Sahoo, and C\. XiongReward\-guided speculative decoding for efficient LLM reasoning\.InForty\-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13\-19, 2025,A\. Singh, M\. Fazel, D\. Hsu, S\. Lacoste\-Julien, F\. Berkenkamp, T\. Maharaj, K\. Wagstaff, and J\. Zhu \(Eds\.\),Proceedings of Machine Learning Research, Vol\.267\.External Links:[Link](https://proceedings.mlr.press/v267/liao25f.html)Cited by:[§1](https://arxiv.org/html/2609.30474#S1.p2.1)\.
- Liuet al\.\(2024\)F\. Liu, Y\. Tang, Z\. Liu, Y\. Ni, D\. Tang, K\. Han, and Y\. WangKangaroo: lossless self\-speculative decoding for accelerating llms via double early exiting\.Advances in Neural Information Processing Systems37,pp\. 11946–11965\.Cited by:[§2](https://arxiv.org/html/2609.30474#S2.p1.1)\.
- Liuet al\.\(2026\)X\. Liu, J\. Yu, J\. Park, I\. Stoica, and A\. CheungSpeculative decoding: performance or illusion?\.InNinth Conference on Machine Learning and Systems,External Links:[Link](https://openreview.net/forum?id=fzkqtezFEi)Cited by:[§2](https://arxiv.org/html/2609.30474#S2.p1.1)\.
- Miaoet al\.\(2024\)X\. Miao, G\. Oliaro, Z\. Zhang, X\. Cheng, Z\. Wang, Z\. Zhang, R\. Y\. Y\. Wong, A\. Zhu, L\. Yang, X\. Shi,et al\.Specinfer: accelerating large language model serving with tree\-based speculative inference and verification\.InProceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3,pp\. 932–949\.Cited by:[§2](https://arxiv.org/html/2609.30474#S2.p1.1)\.
- Narasimhanet al\.\(2025\)H\. Narasimhan, W\. Jitkrittum, A\. S\. Rawat, S\. Kim, N\. Gupta, A\. K\. Menon, and S\. KumarFaster cascades via speculative decoding\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 44949–44987\.Cited by:[§2](https://arxiv.org/html/2609.30474#S2.p2.1)\.
- Naudts \(2011\)J\. NaudtsGeneralized thermostatistics\.Springer\.Cited by:[§4\.3](https://arxiv.org/html/2609.30474#S4.SS3.SSS0.Px1.p1.1),[§4\.3](https://arxiv.org/html/2609.30474#S4.SS3.SSS0.Px1.p1.2),[§VIII\.8](https://arxiv.org/html/2609.30474#S8.SS8.p1.4)\.
- Nocket al\.\(2023\)R\. Nock, E\. Amid, and M\. K\. WarmuthBoosting with tempered exponential measures\.InAdvances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 \- 16, 2023,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),External Links:[Link](http://papers.nips.cc/paper/_files/paper/2023/hash/82d3258eb58ceac31744a88005b7ddef-Abstract-Conference.html)Cited by:[§4\.3](https://arxiv.org/html/2609.30474#S4.SS3.SSS0.Px1.p1.2),[§VIII\.8](https://arxiv.org/html/2609.30474#S8.SS8.p1.5)\.
- Nock and Nielsen \(2007\)R\. Nock and F\. NielsenAℝ\\mathbb\{R\}eal generalization of discrete AdaBoost\.Artif\. Intell\.171\(1\),pp\. 25–41\.External Links:[Link](https://doi.org/10.1016/j.artint.2006.10.014),[Document](https://dx.doi.org/10.1016/J.ARTINT.2006.10.014)Cited by:[§1](https://arxiv.org/html/2609.30474#S1.p2.1),[§4\.1](https://arxiv.org/html/2609.30474#S4.SS1.p1.1),[§4\.2](https://arxiv.org/html/2609.30474#S4.SS2.p4.1),[§VIII\.7](https://arxiv.org/html/2609.30474#S8.SS7.p4.1)\.
- Pankratov and Alistarh \(2026\)S\. Pankratov and D\. AlistarhSpeculative decoding speed\-of\-light: optimal lower bounds via branching random walks\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 6404–6418\.External Links:[Link](https://aclanthology.org/2026.eacl-long.301/),[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.301),ISBN 979\-8\-89176\-380\-7Cited by:[§1](https://arxiv.org/html/2609.30474#S1.p2.1),[§2](https://arxiv.org/html/2609.30474#S2.p1.1)\.
- Qinet al\.\(2025\)Z\. Qin, Z\. He, N\. Prakriya, J\. Cong, and Y\. SunDynamic\-width speculative beam decoding for LLM inference\.InThirty\-Ninth AAAI Conference on Artificial Intelligence, Thirty\-Seventh Conference on Innovative Applications of Artificial Intelligence, Fifteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2025, Philadelphia, PA, USA, February 25 \- March 4, 2025,T\. Walsh, J\. Shah, and Z\. Kolter \(Eds\.\),pp\. 25056–25064\.External Links:[Link](https://doi.org/10.1609/aaai.v39i23.34690),[Document](https://dx.doi.org/10.1609/AAAI.V39I23.34690)Cited by:[§1](https://arxiv.org/html/2609.30474#S1.p2.1),[§2](https://arxiv.org/html/2609.30474#S2.p4.1)\.
- Rockafellar \(1970\)R\. T\. RockafellarConvex Analysis\.Princeton University Press\.Cited by:[§3\.3](https://arxiv.org/html/2609.30474#S3.SS3.SSS0.Px2.p3.1),[§VIII\.14](https://arxiv.org/html/2609.30474#S8.SS14.p3.9)\.
- Schapire and Freund \(2012\)R\.\-E\. Schapire and Y\. FreundBoosting, foundations and algorithms\.MIT Press\.Cited by:[§1](https://arxiv.org/html/2609.30474#S1.p2.1),[§4\.1](https://arxiv.org/html/2609.30474#S4.SS1.p1.4),[§4](https://arxiv.org/html/2609.30474#S4.p1.1)\.
- Sternet al\.\(2018\)M\. Stern, N\. Shazeer, and J\. UszkoreitBlockwise parallel decoding for deep autoregressive models\.Advances in Neural Information Processing Systems31\.Cited by:[§2](https://arxiv.org/html/2609.30474#S2.p1.1)\.
- Sunet al\.\(2023\)Z\. Sun, A\. T\. Suresh, J\. H\. Ro, A\. Beirami, H\. Jain, and F\. YuSpecTr: fast speculative decoding via optimal transport\.Advances in Neural Information Processing Systems36,pp\. 30222–30242\.Cited by:[§1](https://arxiv.org/html/2609.30474#S1.p1.1),[§1](https://arxiv.org/html/2609.30474#S1.p2.1),[§2](https://arxiv.org/html/2609.30474#S2.p1.1)\.
- Tran\-Thien \(2023\)V\. Tran\-ThienAn optimal lossy variant of speculative decoding\.Note:[https://vivien000\.github\.io/blog/journal/a\-provably\-optimal\-lossy\-variant\-of\-speculative\-decoding\.html](https://vivien000.github.io/blog/journal/a-provably-optimal-lossy-variant-of-speculative-decoding.html)Cited by:[§1](https://arxiv.org/html/2609.30474#S1.p1.1),[§1](https://arxiv.org/html/2609.30474#S1.p2.1),[§2](https://arxiv.org/html/2609.30474#S2.p3.1),[§3\.1](https://arxiv.org/html/2609.30474#S3.SS1.p2.1),[§3](https://arxiv.org/html/2609.30474#S3.p1.1)\.
- Wanget al\.\(2025a\)J\. Wang, Z\. Tian, J\. Li, Q\. Xia, X\. Duan, Z\. Wang, B\. Huai, and M\. ZhangAlignment\-augmented speculative decoding with alignment sampling and conditional verification\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 6751–6763\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.343/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.343),ISBN 979\-8\-89176\-332\-6Cited by:[§2](https://arxiv.org/html/2609.30474#S2.p2.1)\.
- Wanget al\.\(2025b\)Z\. Wang, S\. R\. Kasa, A\. M\. S, S\. K\. Kasa, J\. Zou, N\. Jiang, S\. Negi, R\. Zhang, and Q\. SongDIVERSED: relaxed speculative decoding via dynamic ensemble verification\.InNeurIPS 2025 Workshop on Efficient Reasoning,External Links:[Link](https://openreview.net/forum?id=yrkf0GxTe7)Cited by:[§1](https://arxiv.org/html/2609.30474#S1.p2.1),[§2](https://arxiv.org/html/2609.30474#S2.p3.1)\.
- Xiaet al\.\(2026\)G\. Xia, L\. Ribar, and P\. BalancaA practical investigation of training\-free relaxed speculative decoding\.arXiv preprint arXiv:2607\.08690\.Cited by:[§2](https://arxiv.org/html/2609.30474#S2.p2.1)\.
- Xiaet al\.\(2023\)H\. Xia, T\. Ge, P\. Wang, S\. Chen, F\. Wei, and Z\. SuiSpeculative decoding: exploiting speculative execution for accelerating seq2seq generation\.Findings of the Association for Computational Linguistics: EMNLP 2023,pp\. 3909–3925\.Cited by:[§2](https://arxiv.org/html/2609.30474#S2.p1.1)\.
- Yanget al\.\(2023\)N\. Yang, T\. Ge, L\. Wang, B\. Jiao, D\. Jiang, L\. Yang, R\. Majumder, and F\. WeiInference with reference: lossless acceleration of large language models\.arXiv preprint arXiv:2304\.04487\.Cited by:[§1](https://arxiv.org/html/2609.30474#S1.p2.1),[§2](https://arxiv.org/html/2609.30474#S2.p1.1)\.
- Yinet al\.\(2024\)M\. Yin, M\. Chen, K\. Huang, and M\. WangA theoretical perspective for speculative decoding algorithm\.Advances in Neural Information Processing Systems37,pp\. 128082–128117\.Cited by:[§1](https://arxiv.org/html/2609.30474#S1.p2.1),[§2](https://arxiv.org/html/2609.30474#S2.p3.1),[§3\.2](https://arxiv.org/html/2609.30474#S3.SS2.p3.1),[§3\.3](https://arxiv.org/html/2609.30474#S3.SS3.SSS0.Px4.p2.1)\.
- Yuanet al\.\(2024\)H\. Yuan, K\. Lu, F\. Huang, Z\. Yuan, and C\. ZhouSpeculative contrastive decoding\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\),pp\. 56–64\.Cited by:[§2](https://arxiv.org/html/2609.30474#S2.p2.1)\.
- Zhanget al\.\(2024\)J\. Zhang, J\. Wang, H\. Li, L\. Shou, K\. Chen, G\. Chen, and S\. MehrotraDraft& verify: lossless large language model acceleration via self\-speculative decoding\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 11263–11282\.Cited by:[§2](https://arxiv.org/html/2609.30474#S2.p1.1)\.
- Zhanget al\.\(2025\)Z\. Zhang, J\. Xu, T\. Liang, X\. Chen, Z\. He, R\. Wang, and Z\. TuDraft model knows when to stop: self\-verification speculative decoding for long\-form generation\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 16685–16697\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.844/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.844),ISBN 979\-8\-89176\-332\-6Cited by:[§2](https://arxiv.org/html/2609.30474#S2.p2.1)\.
- Zhonget al\.\(2025\)M\. Zhong, N\. Teku, and R\. TandonSpeeding up speculative decoding via sequential approximate verification\.InES\-FoMo III: 3rd Workshop on Efficient Systems for Foundation Models,External Links:[Link](https://openreview.net/forum?id=Y4KcfotBkf)Cited by:[§1](https://arxiv.org/html/2609.30474#S1.p2.1),[§2](https://arxiv.org/html/2609.30474#S2.p3.1),[§2](https://arxiv.org/html/2609.30474#S2.p4.1)\.
- Zhouet al\.\(2024\)Y\. Zhou, K\. Lyu, A\. S\. Rawat, A\. K\. Menon, A\. Rostamizadeh, S\. Kumar, J\. Kagy, and R\. AgarwalDistillSpec: improving speculative decoding via knowledge distillation\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2609.30474#S2.p1.1)\.
- Zhuet al\.\(2009\)J\. Zhu, H\. Zou, S\. Rosset, and T\. HastieMulti\-class Adaboost\.Statistics and Its Interface2,pp\. 349–360\.Cited by:[§4](https://arxiv.org/html/2609.30474#S4.p1.1)\.
Appendix
This is the Appendix to paper "Mentored Decoding: Faster Inference meets Boosting"\. To differentiate with the numberings in the main file, the numbering of Theorems, etc\. is letter\-based \(A, B, …\)\.
## Table of contents
↪\\hookrightarrowProof of Lemma[3\.2](https://arxiv.org/html/2609.30474#S3.Thmtheorem2)
Pg[VIII\.1](https://arxiv.org/html/2609.30474#S8.SS1) ↪\\hookrightarrowProof of Lemma[3\.6](https://arxiv.org/html/2609.30474#S3.Thmtheorem6)
Pg[VIII\.2](https://arxiv.org/html/2609.30474#S8.SS2) ↪\\hookrightarrowProof of Theorem[3\.7](https://arxiv.org/html/2609.30474#S3.Thmtheorem7)and Lemma[3\.8](https://arxiv.org/html/2609.30474#S3.Thmtheorem8)
Pg[VIII\.3](https://arxiv.org/html/2609.30474#S8.SS3) ↪\\hookrightarrowProof of Lemma[3\.9](https://arxiv.org/html/2609.30474#S3.Thmtheorem9)
Pg[VIII\.4](https://arxiv.org/html/2609.30474#S8.SS4) ↪\\hookrightarrowProof of Theorem[3\.12](https://arxiv.org/html/2609.30474#S3.Thmtheorem12)
Pg[VIII\.5](https://arxiv.org/html/2609.30474#S8.SS5) ↪\\hookrightarrowProof of Theorem[3\.13](https://arxiv.org/html/2609.30474#S3.Thmtheorem13)
Pg[VIII\.6](https://arxiv.org/html/2609.30474#S8.SS6) ↪\\hookrightarrowProof of Theorem[4\.2](https://arxiv.org/html/2609.30474#S4.Thmtheorem2)
Pg[VIII\.7](https://arxiv.org/html/2609.30474#S8.SS7) ↪\\hookrightarrowProof of Lemma[4\.5](https://arxiv.org/html/2609.30474#S4.Thmtheorem5)
Pg[VIII\.8](https://arxiv.org/html/2609.30474#S8.SS8) ↪\\hookrightarrowProof of Theorem[4\.7](https://arxiv.org/html/2609.30474#S4.Thmtheorem7)
Pg[VIII\.9](https://arxiv.org/html/2609.30474#S8.SS9) ↪\\hookrightarrowProof of Theorem[4\.8](https://arxiv.org/html/2609.30474#S4.Thmtheorem8)
Pg[VIII\.10](https://arxiv.org/html/2609.30474#S8.SS10) ↪\\hookrightarrowProof of Theorem[5\.1](https://arxiv.org/html/2609.30474#S5.Thmtheorem1)
Pg[VIII\.11](https://arxiv.org/html/2609.30474#S8.SS11) ↪\\hookrightarrowProof of Lemma[5\.2](https://arxiv.org/html/2609.30474#S5.Thmtheorem2)
Pg[VIII\.12](https://arxiv.org/html/2609.30474#S8.SS12) ↪\\hookrightarrowProof of Lemma[5\.3](https://arxiv.org/html/2609.30474#S5.Thmtheorem3)
Pg[VIII\.13](https://arxiv.org/html/2609.30474#S8.SS13) ↪\\hookrightarrowProof of Theorem[5\.4](https://arxiv.org/html/2609.30474#S5.Thmtheorem4)
Pg[VIII\.14](https://arxiv.org/html/2609.30474#S8.SS14) ↪\\hookrightarrowProof of Lemma[6\.1](https://arxiv.org/html/2609.30474#S6.Thmtheorem1)
Pg[VIII\.15](https://arxiv.org/html/2609.30474#S8.SS15) ↪\\hookrightarrowProof of Lemma[6\.2](https://arxiv.org/html/2609.30474#S6.Thmtheorem2)
Pg[VIII\.16](https://arxiv.org/html/2609.30474#S8.SS16) ↪\\hookrightarrowProof of Lemma[6\.4](https://arxiv.org/html/2609.30474#S6.Thmtheorem4)
Pg[VIII\.17](https://arxiv.org/html/2609.30474#S8.SS17)
## VIIIProofs
### VIII\.1Proof of Lemma[3\.2](https://arxiv.org/html/2609.30474#S3.Thmtheorem2)
Suppose𝝅∈mdf1\(𝒑,𝒒,D\)\\bm\{\\pi\}\\in\\textsc\{md\}^\{1\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)\. Then with the choice𝒓=\.min\{𝟏,𝝅⊘𝒑\}∈\[0,1\]n\\bm\{r\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\min\\\{\\bm\{1\},\\bm\{\\pi\}\\oslash\\bm\{p\}\\\}\\in\[0,1\]^\{n\}\([3\.2](https://arxiv.org/html/2609.30474#S3.EGx5)\), we get𝒑⊤𝒓=𝟏⊤min\{𝝅,𝒑\}=1−DTV\(𝝅∥𝒑\)\\bm\{p\}^\{\\top\}\\bm\{r\}=\\bm\{1\}^\{\\top\}\\min\\\{\\bm\{\\pi\},\\bm\{p\}\\\}=1\-D\_\{\\mathrm\{TV\}\}\(\\bm\{\\pi\}\\\|\\bm\{p\}\)\. We also trivially have𝒔∈Δn\\bm\{s\}\\in\\Delta\_\{n\}so the couple\(𝒓,𝒔\)\(\\bm\{r\},\\bm\{s\}\)is feasible for \([f \-MD\-2](https://arxiv.org/html/2609.30474#S3.Ex1)\)\. Suppose it is not optimal and build𝝅′\\bm\{\\pi\}^\{\\prime\}from a better solution\(𝒓′,𝒔′\)\(\\bm\{r\}^\{\\prime\},\\bm\{s\}^\{\\prime\}\)– thus with𝒑⊤𝒓′\>𝒑⊤𝒓\\bm\{p\}^\{\\top\}\\bm\{r\}^\{\\prime\}\>\\bm\{p\}^\{\\top\}\\bm\{r\}– via \([6](https://arxiv.org/html/2609.30474#S3.E6)\)\. For anyi∈\[n\]i\\in\[n\], we havepiri′≤pip\_\{i\}r^\{\\prime\}\_\{i\}\\leq p\_\{i\}but alsopiri′≤πi′p\_\{i\}r^\{\\prime\}\_\{i\}\\leq\\pi^\{\\prime\}\_\{i\}because of \([6](https://arxiv.org/html/2609.30474#S3.E6)\)\. So we haveDTV\(𝝅′∥𝒑\)=1−𝟏⊤min\{𝝅′,𝒑\}≤1−𝒑⊤𝒓′<1−𝒑⊤𝒓=DTV\(𝝅∥𝒑\)D\_\{\\mathrm\{TV\}\}\(\\bm\{\\pi\}^\{\\prime\}\\\|\\bm\{p\}\)=1\-\\bm\{1\}^\{\\top\}\\min\\\{\\bm\{\\pi\}^\{\\prime\},\\bm\{p\}\\\}\\leq 1\-\\bm\{p\}^\{\\top\}\\bm\{r\}^\{\\prime\}<1\-\\bm\{p\}^\{\\top\}\\bm\{r\}=D\_\{\\mathrm\{TV\}\}\(\\bm\{\\pi\}\\\|\\bm\{p\}\), which contradicts the fact that𝝅∈mdf1\(𝒑,𝒒,D\)\\bm\{\\pi\}\\in\\textsc\{md\}^\{1\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)\. So we have\(𝒓,𝒔\)∈mdf2\(𝒑,𝒒,D\)\(\\bm\{r\},\\bm\{s\}\)\\in\\textsc\{md\}^\{2\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)\.
Respectively, suppose\(𝒓,𝒔\)∈mdf2\(𝒑,𝒒,D\)\(\\bm\{r\},\\bm\{s\}\)\\in\\textsc\{md\}^\{2\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)\. Clearly𝝅\\bm\{\\pi\}as per \([6](https://arxiv.org/html/2609.30474#S3.E6)\) is feasible for \([f \-MD\-1](https://arxiv.org/html/2609.30474#S3.Ex2)\)\. Suppose it is not optimal and build this time\(𝒓′,𝒔′\)\(\\bm\{r\}^\{\\prime\},\\bm\{s\}^\{\\prime\}\)from a better solution𝝅′\\bm\{\\pi\}^\{\\prime\}– thus withDTV\(𝝅′∥𝒑\)<DTV\(𝝅∥𝒑\)D\_\{\\mathrm\{TV\}\}\(\\bm\{\\pi\}^\{\\prime\}\\\|\\bm\{p\}\)<D\_\{\\mathrm\{TV\}\}\(\\bm\{\\pi\}\\\|\\bm\{p\}\)– via \([3\.2](https://arxiv.org/html/2609.30474#S3.EGx5)\)\. This time, we directly have from the construction of𝒓′\\bm\{r\}^\{\\prime\}the chain of \(in\)equalities1−𝒑⊤𝒓′=DTV\(𝝅′∥𝒑\)<DTV\(𝝅∥𝒑\)=1−𝒑⊤𝒓1\-\\bm\{p\}^\{\\top\}\\bm\{r\}^\{\\prime\}=D\_\{\\mathrm\{TV\}\}\(\\bm\{\\pi\}^\{\\prime\}\\\|\\bm\{p\}\)<D\_\{\\mathrm\{TV\}\}\(\\bm\{\\pi\}\\\|\\bm\{p\}\)=1\-\\bm\{p\}^\{\\top\}\\bm\{r\}, resulting in−𝒑⊤𝒓′<−𝒑⊤𝒓\-\\bm\{p\}^\{\\top\}\\bm\{r\}^\{\\prime\}<\-\\bm\{p\}^\{\\top\}\\bm\{r\}, a contradiction with the fact that\(𝒓,𝒔\)∈mdf2\(𝒑,𝒒,D\)\(\\bm\{r\},\\bm\{s\}\)\\in\\textsc\{md\}^\{2\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)\. So we have𝝅∈mdf1\(𝒑,𝒒,D\)\\bm\{\\pi\}\\in\\textsc\{md\}^\{1\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\), which ends the main part of the proof of Lemma[3\.2](https://arxiv.org/html/2609.30474#S3.Thmtheorem2)\. We easily check the two equivalent formulations for𝒔\\bm\{s\}sinceπi−piri=πi−min\{πi,pi\}=max\{0,πi−pi\},∀i∈\[n\]\\pi\_\{i\}\-p\_\{i\}r\_\{i\}=\\pi\_\{i\}\-\\min\\\{\\pi\_\{i\},p\_\{i\}\\\}=\\max\\\{0,\\pi\_\{i\}\-p\_\{i\}\\\},\\forall i\\in\[n\]\.
### VIII\.2Proof of Lemma[3\.6](https://arxiv.org/html/2609.30474#S3.Thmtheorem6)
We consider \([f \-MD\-1](https://arxiv.org/html/2609.30474#S3.Ex2)\) \(there is no difficulty in reparameterizing the proof using Lemma[3\.2](https://arxiv.org/html/2609.30474#S3.Thmtheorem2)for \([f \-MD\-2](https://arxiv.org/html/2609.30474#S3.Ex1)\)\)\. Sinceff\-divergences satisfy the identity of indiscernibles,𝒒\>𝟎\\bm\{q\}\>\\bm\{0\}implies the existence of𝒒≠𝝅~∈Δn\\bm\{q\}\\neq\\tilde\{\\bm\{\\pi\}\}\\in\\Delta\_\{n\}such that𝝅~\>0\\tilde\{\\bm\{\\pi\}\}\>0andDf\(𝝅~∥𝒒\)≤DD\_\{f\}\(\\tilde\{\\bm\{\\pi\}\}\\\|\\bm\{q\}\)\\leq D\(and we also have𝝅~≠𝒑\\tilde\{\\bm\{\\pi\}\}\\neq\\bm\{p\}\)\. We then check Slater’s constraint qualification by picking any0<δ<mini\{min\{pi,π~i\}/\(pi\+π~i\)\}0<\\delta<\\min\_\{i\}\\\{\\min\\\{p\_\{i\},\\tilde\{\\pi\}\_\{i\}\\\}/\(p\_\{i\}\+\\tilde\{\\pi\}\_\{i\}\)\\\}, and choosing
𝒕\\displaystyle\\bm\{t\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}δ⋅\(𝒑\+𝝅~\)\.\\displaystyle\\delta\\cdot\(\\bm\{p\}\+\\tilde\{\\bm\{\\pi\}\}\)\.\(78\)For this choice and that ofδ\\deltawe get𝒕<𝒑\\bm\{t\}<\\bm\{p\}and for the choice \(note thatδ<1/2\\delta<1/2\)
𝒔\\displaystyle\\bm\{s\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}1−δ1−2δ⋅𝝅~−δ1−2δ⋅𝒑,\\displaystyle\\frac\{1\-\\delta\}\{1\-2\\delta\}\\cdot\\tilde\{\\bm\{\\pi\}\}\-\\frac\{\\delta\}\{1\-2\\delta\}\\cdot\\bm\{p\},we have𝟏⊤𝒔=1\\bm\{1\}^\{\\top\}\\bm\{s\}=1but more importantly𝒔\>𝟎\\bm\{s\}\>\\bm\{0\}, so Slater’s constraint qualification are satisfied\.
### VIII\.3Proof of Theorem[3\.7](https://arxiv.org/html/2609.30474#S3.Thmtheorem7)and Lemma[3\.8](https://arxiv.org/html/2609.30474#S3.Thmtheorem8)
For readability reasons, we reparameterize \([f \-MD\-2](https://arxiv.org/html/2609.30474#S3.Ex1)\) as:
mdf′\(𝒑,𝒒;D\)=\.argmin𝟎≤𝒕≤𝒑,𝒔∈Δn−𝟏⊤𝒕s\.t\.Df\(𝒕\+\(1−𝟏⊤𝒕\)⋅𝒔∥𝒒\)≤D\.\\displaystyle\\textsc\{md\}^\{\\prime\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\arg\\min\_\{\\bm\{0\}\\leq\\bm\{t\}\\leq\\bm\{p\},\\bm\{s\}\\in\\Delta\_\{n\}\}\-\\bm\{1\}^\{\\top\}\\bm\{t\}\\quad\\mbox\{s\.t\. \}D\_\{f\}\(\\bm\{t\}\+\(1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\)\\cdot\\bm\{s\}\\\|\\bm\{q\}\)\\leq D\.\(79\)Mentored decoding’s𝒓\\bm\{r\}in \([f \-MD\-2](https://arxiv.org/html/2609.30474#S3.Ex1)\) is obtained as𝒓=\.𝒕⊘𝒑\\bm\{r\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\bm\{t\}\\oslash\\bm\{p\}\.
\(mdf′→isf\\textsc\{md\}^\{\\prime\}\_\{f\}\\rightarrow\\textsc\{is\}\_\{f\}\)We have the Lagrangian,
ℒ1\(𝒕,𝒔,μ,𝝌,𝝂\)\\displaystyle\\mathcal\{L\}\_\{1\}\(\\bm\{t\},\\bm\{s\};\\mu,\\bm\{\\chi\},\\bm\{\\nu\}\)=\\displaystyle=−𝟏⊤𝒕\+λ⋅\(Df\(𝒕\+\(1−𝟏⊤𝒕\)⋅𝒔∥𝒒\)−D\)\+μ⋅\(𝟏⊤𝒔−1\)\\displaystyle\-\\bm\{1\}^\{\\top\}\\bm\{t\}\+\\lambda\\cdot\(D\_\{f\}\(\\bm\{t\}\+\(1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\)\\cdot\\bm\{s\}\\\|\\bm\{q\}\)\-D\)\+\\mu\\cdot\(\\bm\{1\}^\{\\top\}\\bm\{s\}\-1\)\(80\)\+𝝌⊤−𝒔\+𝝂⊤\(𝒕−𝒑\)\.\\displaystyle\+\\bm\{\\chi\}^\{\\top\}\-\\bm\{s\}\+\\bm\{\\nu\}^\{\\top\}\(\\bm\{t\}\-\\bm\{p\}\)\.We have KKT the conditions \(using notations from \([f \-MD\-2](https://arxiv.org/html/2609.30474#S3.Ex1)\) and \([80](https://arxiv.org/html/2609.30474#S8.E80)\)\)
𝒕\\displaystyle\\bm\{t\}≤\\displaystyle\\leq𝒑,\\displaystyle\\bm\{p\},\(81\)𝒔\\displaystyle\\bm\{s\}≥\\displaystyle\\geq𝟎,\\displaystyle\\bm\{0\},\(82\)𝟏⊤𝒔\\displaystyle\\bm\{1\}^\{\\top\}\\bm\{s\}=\\displaystyle=1,\\displaystyle 1,\(83\)𝝌\\displaystyle\\bm\{\\chi\}≥\\displaystyle\\geq𝟎,\\displaystyle\\bm\{0\},\(84\)𝝂\\displaystyle\\bm\{\\nu\}≥\\displaystyle\\geq𝟎,\\displaystyle\\bm\{0\},\(85\)𝝌⊙𝒔\\displaystyle\\bm\{\\chi\}\\odot\\bm\{s\}=\\displaystyle=𝟎,\\displaystyle\\bm\{0\},\(86\)𝝂⊙\(𝒕−𝒑\)\\displaystyle\\bm\{\\nu\}\\odot\(\\bm\{t\}\-\\bm\{p\}\)=\\displaystyle=𝟎,\\displaystyle\\bm\{0\},\(87\)∇𝒕ℒ1=∇𝒔ℒ1\\displaystyle\\nabla\_\{\\bm\{t\}\}\\mathcal\{L\}\_\{1\}=\\nabla\_\{\\bm\{s\}\}\\mathcal\{L\}\_\{1\}=\\displaystyle=𝟎,\\displaystyle\\bm\{0\},\(88\)λ\\displaystyle\\lambda≥\\displaystyle\\geq0,\\displaystyle 0,\(89\)λ⋅\(Df\(𝒕\+\(1−𝟏⊤𝒕\)⋅𝒔∥𝒒\)−D\)\\displaystyle\\lambda\\cdot\(D\_\{f\}\(\\bm\{t\}\+\(1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\)\\cdot\\bm\{s\}\\\|\\bm\{q\}\)\-D\)=\\displaystyle=0,\\displaystyle 0,\(90\)Df\(𝒕\+\(1−𝟏⊤𝒕\)⋅𝒔∥𝒒\)\\displaystyle D\_\{f\}\(\\bm\{t\}\+\(1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\)\\cdot\\bm\{s\}\\\|\\bm\{q\}\)≤\\displaystyle\\leqD\.\\displaystyle D\.\(91\)Denote for short
πi\\displaystyle\\pi\_\{i\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}ti\+\(1−𝟏⊤𝒕\)⋅si\.\\displaystyle t\_\{i\}\+\(1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\)\\cdot s\_\{i\}\.\(92\)\([88](https://arxiv.org/html/2609.30474#S8.E88)\) is equivalent to:
∂ℒ1∂ti=λ⋅∑j\(−f′\)\(πjqj\)⋅sj−λ⋅\(−f′\)\(πiqi\)−1\+νi\\displaystyle\\frac\{\\partial\\mathcal\{L\}\_\{1\}\}\{\\partial t\_\{i\}\}=\\lambda\\cdot\\sum\_\{j\}\(\-f^\{\\prime\}\)\\left\(\\frac\{\\pi\_\{j\}\}\{q\_\{j\}\}\\right\)\\cdot s\_\{j\}\-\\lambda\\cdot\(\-f^\{\\prime\}\)\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)\-1\+\\nu\_\{i\}=\\displaystyle=0,∀i,\\displaystyle 0,\\forall i,\(93\)∂ℒ1∂si=μ−λ⋅\(−f′\)\(πiqi\)⋅\(1−𝟏⊤𝒕\)−χi\\displaystyle\\frac\{\\partial\\mathcal\{L\}\_\{1\}\}\{\\partial s\_\{i\}\}=\\mu\-\\lambda\\cdot\(\-f^\{\\prime\}\)\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)\\cdot\(1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\)\-\\chi\_\{i\}=\\displaystyle=0,∀i\.\\displaystyle 0,\\forall i\.\(94\)We have a first Lemma\.
###### Lemma A\.
IfD<Df\(𝐩∥𝐪\)D<D\_\{f\}\(\\bm\{p\}\\\|\\bm\{q\}\)then𝐭≠𝐩\\bm\{t\}\\neq\\bm\{p\}andλ\>0\\lambda\>0at the optimum\.
###### Proof\.
Proof immediate for𝒕≠𝒑\\bm\{t\}\\neq\\bm\{p\}because for𝒕=𝒑\\bm\{t\}=\\bm\{p\},Df\(𝝅∥𝒒\)=Df\(𝒑∥𝒒\)\>DD\_\{f\}\(\\bm\{\\pi\}\\\|\\bm\{q\}\)=D\_\{f\}\(\\bm\{p\}\\\|\\bm\{q\}\)\>D, not feasible in this case\. Ifλ=0\\lambda=0, we get from \([93](https://arxiv.org/html/2609.30474#S8.E93)\)νi=1≠0,∀i\\nu\_\{i\}=1\\neq 0,\\forall iand thus complementary slackness \([87](https://arxiv.org/html/2609.30474#S8.E87)\) imposes𝒕=𝒑\\bm\{t\}=\\bm\{p\}, impossible sinceD<Df\(𝒑∥𝒒\)D<D\_\{f\}\(\\bm\{p\}\\\|\\bm\{q\}\)\. ∎
We can thus reorganize \([93](https://arxiv.org/html/2609.30474#S8.E93)\) and \([94](https://arxiv.org/html/2609.30474#S8.E94)\) with the complementary slackness conditions \([86](https://arxiv.org/html/2609.30474#S8.E86)\), \([87](https://arxiv.org/html/2609.30474#S8.E87)\) to give \(⟦\.⟧\\llbracket\.\\rrbracketis Iverson’s bracket\):
\(−f′\)\(πiqi\)\\displaystyle\(\-f^\{\\prime\}\)\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)=\\displaystyle=α\+νiλ⋅⟦ti=pi⟧⏟≥0,∀i,withα=\.𝔼i∼𝒔\[\(−f′\)\(πiqi\)\]−1λ\.\\displaystyle\\alpha\+\\underbrace\{\\frac\{\\nu\_\{i\}\}\{\\lambda\}\\cdot\\color\[rgb\]\{1,0,0\}\{\\llbracket t\_\{i\}=p\_\{i\}\\rrbracket\}\}\_\{\\geq 0\},\\forall i,\\quad\\mbox\{with \}\\alpha\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\mathbb\{E\}\_\{i\\sim\\bm\{s\}\}\\left\[\(\-f^\{\\prime\}\)\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)\\right\]\-\\frac\{1\}\{\\lambda\}\.\(95\)\(−f′\)\(πiqi\)\\displaystyle\(\-f^\{\\prime\}\)\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)=\\displaystyle=β−χiλ⋅\(1−𝟏⊤𝒕\)⋅⟦si=0⟧⏟≥0,∀i,withβ=\.μλ⋅\(1−𝟏⊤𝒕\),\\displaystyle\\beta\-\\underbrace\{\\frac\{\\chi\_\{i\}\}\{\\lambda\\cdot\(1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\)\}\\cdot\\color\[rgb\]\{1,0,0\}\{\\llbracket s\_\{i\}=0\\rrbracket\}\}\_\{\\geq 0\},\\forall i,\\quad\\mbox\{with \}\\beta\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\frac\{\\mu\}\{\\lambda\\cdot\(1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\)\},\(96\)
###### Lemma B\.
At the optimum,α=β−\(1/λ\)\\alpha=\\beta\-\(1/\\lambda\); henceα<β\\alpha<\\beta\. Furthermore,α\\alphaandβ\\betaalso satisfy:
α\\displaystyle\\alpha=\\displaystyle=𝔼i∼𝒖\[\(−f′\)\(πiqi\)\],with𝒖=\.11−𝟏⊤𝒕⋅\(𝒑−𝒕\)∈Δn\.\\displaystyle\\mathbb\{E\}\_\{i\\sim\\bm\{u\}\}\\left\[\(\-f^\{\\prime\}\)\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)\\right\],\\quad\\mbox\{ with \}\\bm\{u\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\frac\{1\}\{1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\}\\cdot\(\\bm\{p\}\-\\bm\{t\}\)\\in\\Delta\_\{n\}\.\(97\)β\\displaystyle\\beta=\\displaystyle=𝔼i∼𝒔\[\(−f′\)\(πiqi\)\]\\displaystyle\\mathbb\{E\}\_\{i\\sim\\bm\{s\}\}\\left\[\(\-f^\{\\prime\}\)\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)\\right\]\(98\)
###### Proof\.
Sum \([96](https://arxiv.org/html/2609.30474#S8.E96)\) timessis\_\{i\}and we get
∑i\(−f′\)\(πiqi\)⋅si\\displaystyle\\sum\_\{i\}\(\-f^\{\\prime\}\)\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)\\cdot s\_\{i\}=\\displaystyle=β⋅∑isi⏟=1−∑iχiλ⋅\(1−𝟏⊤𝒕\)⋅⟦si=0⟧⋅si⏟=0,∀i=β,\\displaystyle\\beta\\cdot\\underbrace\{\\sum\_\{i\}s\_\{i\}\}\_\{=1\}\-\\sum\_\{i\}\\frac\{\\chi\_\{i\}\}\{\\lambda\\cdot\(1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\)\}\\cdot\\underbrace\{\{\\color\[rgb\]\{1,0,0\}\{\\llbracket s\_\{i\}=0\\rrbracket\\cdot s\_\{i\}\}\}\}\_\{=0,\\forall i\}=\\beta,and we reorganize using \([95](https://arxiv.org/html/2609.30474#S8.E95)\) to get
β\\displaystyle\\beta=\\displaystyle=𝔼i∼𝒔\[\(−f′\)\(πiqi\)\],\\displaystyle\\mathbb\{E\}\_\{i\\sim\\bm\{s\}\}\\left\[\(\-f^\{\\prime\}\)\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)\\right\],andα=β−\(1/λ\)\\alpha=\\beta\-\(1/\\lambda\)because of the definition ofα\\alphain \([95](https://arxiv.org/html/2609.30474#S8.E95)\)\. Now, sum \([95](https://arxiv.org/html/2609.30474#S8.E95)\) timespi−tip\_\{i\}\-t\_\{i\}and we get
∑i\(−f′\)\(πiqi\)⋅\(pi−ti\)\\displaystyle\\sum\_\{i\}\(\-f^\{\\prime\}\)\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)\\cdot\(p\_\{i\}\-t\_\{i\}\)=\\displaystyle=α⋅∑ipi−ti⏟=1−𝟏⊤𝒕−∑iνiλ⋅⟦ti=pi⟧⋅\(ti−pi\)⏟=0,∀i,\\displaystyle\\alpha\\cdot\\underbrace\{\\sum\_\{i\}p\_\{i\}\-t\_\{i\}\}\_\{=1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\}\-\\sum\_\{i\}\\frac\{\\nu\_\{i\}\}\{\\lambda\}\\cdot\\underbrace\{\\color\[rgb\]\{1,0,0\}\{\\llbracket t\_\{i\}=p\_\{i\}\\rrbracket\\cdot\(t\_\{i\}\-p\_\{i\}\)\}\}\_\{=0,\\forall i\},and rearrange to find the expression ofα\\alphain \([97](https://arxiv.org/html/2609.30474#S8.E97)\)\. ∎
Pick anyiisuch thatti<pit\_\{i\}<p\_\{i\}\. \([95](https://arxiv.org/html/2609.30474#S8.E95)\) yields\(−f′\)\(πiqi\)=α\(\-f^\{\\prime\}\)\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)=\\alphaand so, in any optimal solution,
ti<pi\\displaystyle t\_\{i\}<p\_\{i\}⇒\\displaystyle\\Rightarrowπi∈qi⋅Lα\(−f′\)\.\\displaystyle\\pi\_\{i\}\\in q\_\{i\}\\cdot L\_\{\\alpha\}\(\-f^\{\\prime\}\)\.\(99\)Pick anyiisuch thatsi\>0s\_\{i\}\>0\. \([96](https://arxiv.org/html/2609.30474#S8.E96)\) yields\(−f′\)\(πiqi\)=β\(\-f^\{\\prime\}\)\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)=\\betaand so, in any optimal solution,
si\>0\\displaystyle s\_\{i\}\>0⇒\\displaystyle\\Rightarrowπi∈qi⋅Lβ\(−f′\)\.\\displaystyle\\pi\_\{i\}\\in q\_\{i\}\\cdot L\_\{\\beta\}\(\-f^\{\\prime\}\)\.\(100\)All other cases must meetti=pit\_\{i\}=p\_\{i\}andsi=0s\_\{i\}=0, henceπi=pi\\pi\_\{i\}=p\_\{i\}\. Sinceα<β\\alpha<\\beta, we must haveLα\(−f′\)\>Lβ\(−f′\)L\_\{\\alpha\}\(\-f^\{\\prime\}\)\>L\_\{\\beta\}\(\-f^\{\\prime\}\)\(ffis convex\), so it is impossible that1\>Lα\(−f′\)1\>L\_\{\\alpha\}\(\-f^\{\\prime\}\)orLβ\(−f′\)\>1L\_\{\\beta\}\(\-f^\{\\prime\}\)\>1otherwise𝝅\\bm\{\\pi\}would not be a distribution\. We thus have simultaneously
Lβ\(−f′\)\\displaystyle L\_\{\\beta\}\(\-f^\{\\prime\}\)<\\displaystyle<Lα\(−f′\),\\displaystyle L\_\{\\alpha\}\(\-f^\{\\prime\}\),\(101\)1\\displaystyle 1≤\\displaystyle\\leqmaxLα\(−f′\),\\displaystyle\\max L\_\{\\alpha\}\(\-f^\{\\prime\}\),\(102\)minLβ\(−f′\)\\displaystyle\\min L\_\{\\beta\}\(\-f^\{\\prime\}\)≤\\displaystyle\\leq1\.\\displaystyle 1\.\(103\)We go back to \([99](https://arxiv.org/html/2609.30474#S8.E99)\): for anyiisuch thatti<pit\_\{i\}<p\_\{i\}and sinceα<β\\alpha<\\beta\([96](https://arxiv.org/html/2609.30474#S8.E96)\) yields that for all these indicessi=0s\_\{i\}=0soπi=ti<pi\\pi\_\{i\}=t\_\{i\}<p\_\{i\}\. Since otherwiseti=pit\_\{i\}=p\_\{i\}, we get that in all cases,
ti\\displaystyle t\_\{i\}∈\\displaystyle\\inminset\(pi,qi⋅Lα\(−f′\)\)=\{pi\}\+minset\(0,qi⋅Lα\(−f′\)−pi\)\\displaystyle\\mathrm\{minset\}\(p\_\{i\},q\_\{i\}\\cdot L\_\{\\alpha\}\(\-f^\{\\prime\}\)\)=\\\{p\_\{i\}\\\}\+\\mathrm\{minset\}\(0,q\_\{i\}\\cdot L\_\{\\alpha\}\(\-f^\{\\prime\}\)\-p\_\{i\}\)\(104\)and while this guarantees𝒕⪯𝒑\\bm\{t\}\\preceq\\bm\{p\}, we must also ensure𝒕≠𝒑\\bm\{t\}\\neq\\bm\{p\}\(Lemma[A](https://arxiv.org/html/2609.30474#S8.Thmtheorem1)\)\. We separately note that
minLα\(−f′\)\\displaystyle\\min L\_\{\\alpha\}\(\-f^\{\\prime\}\)<\\displaystyle<maxjpj/qj\\displaystyle\\max\_\{j\}p\_\{j\}/q\_\{j\}\(105\)otherwise the only feasible solution is𝒕=𝒑=𝝅\\bm\{t\}=\\bm\{p\}=\\bm\{\\pi\}, impossible \(Lemma[A](https://arxiv.org/html/2609.30474#S8.Thmtheorem1)\)\.
We go back to \([100](https://arxiv.org/html/2609.30474#S8.E100)\): for anyiisuch thatsi\>0s\_\{i\}\>0and sinceα<β\\alpha<\\beta\([95](https://arxiv.org/html/2609.30474#S8.E95)\) yields that for all these indicesti=pit\_\{i\}=p\_\{i\}\. Since otherwisesi=0s\_\{i\}=0, we get fromπi=\.ti\+si\(1−𝟏⊤𝒕\)\\pi\_\{i\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}t\_\{i\}\+s\_\{i\}\(1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\):
si\(1−𝟏⊤𝒕\)\\displaystyle s\_\{i\}\(1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\)∈\\displaystyle\\inmaxset\(0,qi⋅Lβ\(−f′\)−pi\)\.\\displaystyle\\mathrm\{maxset\}\(0,q\_\{i\}\\cdot L\_\{\\beta\}\(\-f^\{\\prime\}\)\-p\_\{i\}\)\.\(106\)We separately note that
maxLβ\(−f′\)\\displaystyle\\max L\_\{\\beta\}\(\-f^\{\\prime\}\)\>\\displaystyle\>minjpj/qj\\displaystyle\\min\_\{j\}p\_\{j\}/q\_\{j\}\(107\)otherwise𝒔=𝟎\\bm\{s\}=\\bm\{0\}, not admissible\. From \([101](https://arxiv.org/html/2609.30474#S8.E101)\), we getqi⋅Lβ\(−f′\)−pi<qi⋅Lα\(−f′\)−piq\_\{i\}\\cdot L\_\{\\beta\}\(\-f^\{\\prime\}\)\-p\_\{i\}<q\_\{i\}\\cdot L\_\{\\alpha\}\(\-f^\{\\prime\}\)\-p\_\{i\}and so we get the final expression for𝝅\\bm\{\\pi\}from its definition and \([104](https://arxiv.org/html/2609.30474#S8.E104)\), \([106](https://arxiv.org/html/2609.30474#S8.E106)\):
πi\\displaystyle\\pi\_\{i\}∈\\displaystyle\\in\{pi\}\+minset\(0,qi⋅Lα\(−f′\)−pi\)\+maxset\(0,qi⋅Lβ\(−f′\)−pi\)\\displaystyle\\\{p\_\{i\}\\\}\+\\mathrm\{minset\}\(0,q\_\{i\}\\cdot L\_\{\\alpha\}\(\-f^\{\\prime\}\)\-p\_\{i\}\)\+\\mathrm\{maxset\}\(0,q\_\{i\}\\cdot L\_\{\\beta\}\(\-f^\{\\prime\}\)\-p\_\{i\}\)=clampset\(pi,qi⋅Lβ\(−f′\),qi⋅Lα\(−f′\)\),\\displaystyle=\\mathrm\{clampset\}\(p\_\{i\},q\_\{i\}\\cdot L\_\{\\beta\}\(\-f^\{\\prime\}\),q\_\{i\}\\cdot L\_\{\\alpha\}\(\-f^\{\\prime\}\)\),where the equality comes by definition ofclampset\\mathrm\{clampset\}\. Thus, all optimal solutions satisfy
𝝅\\displaystyle\\bm\{\\pi\}∈\\displaystyle\\inclampset\(𝒑,Lβ\(−f′\)⋅𝒒,⋅Lα\(−f′\)⋅𝒒\)∩Δn,\\displaystyle\\mathrm\{clampset\}\(\\bm\{p\},L\_\{\\beta\}\(\-f^\{\\prime\}\)\\cdot\\bm\{q\},\\cdot L\_\{\\alpha\}\(\-f^\{\\prime\}\)\\cdot\\bm\{q\}\)\\cap\\Delta\_\{n\},\(108\)At this stage, we have explicitly satisfied KKT conditions \([81](https://arxiv.org/html/2609.30474#S8.E81)\), \([82](https://arxiv.org/html/2609.30474#S8.E82)\), \([83](https://arxiv.org/html/2609.30474#S8.E83)\), \([88](https://arxiv.org/html/2609.30474#S8.E88)\), \([89](https://arxiv.org/html/2609.30474#S8.E89)\)\.
To check \([84](https://arxiv.org/html/2609.30474#S8.E84)\) and \([86](https://arxiv.org/html/2609.30474#S8.E86)\), we compute
χi\\displaystyle\\chi\_\{i\}=\\displaystyle=\{λ⋅\(1−𝟏⊤𝒕\)⋅\(β−\(−f′\)\(πiqi\)\)ifti<pi∨πi=pi0otherwise\.\\displaystyle\\left\\\{\\begin\{array\}\[\]\{ccl\}\\lambda\\cdot\(1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\)\\cdot\\left\(\\beta\-\(\-f^\{\\prime\}\)\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)\\right\)&\\mbox\{ if \}&t\_\{i\}<p\_\{i\}\\vee\\pi\_\{i\}=p\_\{i\}\\\\ 0&\\lx@intercol\\mbox\{ otherwise \}\\hfil\\lx@intercol\\end\{array\}\\right\.\.Note that the first condition is equivalent toπi≤pi\\pi\_\{i\}\\leq p\_\{i\}\.
Suppose first thatti<pit\_\{i\}<p\_\{i\}\. Then we know thatsi=0s\_\{i\}=0so that \([VIII\.3](https://arxiv.org/html/2609.30474#S8.EGx92)\) is in fact \([96](https://arxiv.org/html/2609.30474#S8.E96)\)\. In this case, \([99](https://arxiv.org/html/2609.30474#S8.E99)\) yields the second identity in
χi\\displaystyle\\chi\_\{i\}=\\displaystyle=λ⋅\(1−𝟏⊤𝒕\)⋅\(β−\(−f′\)\(πiqi\)\)\\displaystyle\\lambda\\cdot\(1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\)\\cdot\\left\(\\beta\-\(\-f^\{\\prime\}\)\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)\\right\)\(112\)=\\displaystyle=λ⋅\(1−𝟏⊤𝒕\)⋅\(β−\(−f′\)\(Lα\(−f′\)\)\)\\displaystyle\\lambda\\cdot\(1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\)\\cdot\\left\(\\beta\-\(\-f^\{\\prime\}\)\\left\(L\_\{\\alpha\}\(\-f^\{\\prime\}\)\\right\)\\right\)=\\displaystyle=λ⋅\(1−𝟏⊤𝒕\)⋅\(β−α\)\\displaystyle\\lambda\\cdot\(1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\)\\cdot\\left\(\\beta\-\\alpha\\right\)=\\displaystyle=1−𝟏⊤𝒕≥0\.\\displaystyle 1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\\geq 0\.the penultimate identity comes from the definition ofLα\(−f′\)L\_\{\\alpha\}\(\-f^\{\\prime\}\)and the last one is Lemma[B](https://arxiv.org/html/2609.30474#S8.Thmtheorem2)\. Soti<pit\_\{i\}<p\_\{i\}impliesχi≥0\\chi\_\{i\}\\geq 0andχisi=0\\chi\_\{i\}s\_\{i\}=0\.
Suppose now thatπi=pi\(=ti\)\\pi\_\{i\}=p\_\{i\}\(=t\_\{i\}\)\. In this case, we have
\(−f′\)\(πiqi\)=\(−f′\)\(piqi\)∈\[α,β\],\\displaystyle\(\-f^\{\\prime\}\)\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)=\(\-f^\{\\prime\}\)\\left\(\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\right\)\\in\[\\alpha,\\beta\],Hence from \([112](https://arxiv.org/html/2609.30474#S8.Ex34)\)χi≥0\\chi\_\{i\}\\geq 0again while, since we also havesi=0s\_\{i\}=0, we haveχisi=0\\chi\_\{i\}s\_\{i\}=0\.
Finally, if¬\(ti<pi∨πi=pi\)≡πi\>pi\\neg\(t\_\{i\}<p\_\{i\}\\vee\\pi\_\{i\}=p\_\{i\}\)\\equiv\\pi\_\{i\}\>p\_\{i\}, we havesi\>0s\_\{i\}\>0butχisi=0\\chi\_\{i\}s\_\{i\}=0and stillχi≥0\\chi\_\{i\}\\geq 0\.
Hence, KKT \([84](https://arxiv.org/html/2609.30474#S8.E84)\) and \([86](https://arxiv.org/html/2609.30474#S8.E86)\) are satisfied\.
We now check KKT \([85](https://arxiv.org/html/2609.30474#S8.E85)\) and \([87](https://arxiv.org/html/2609.30474#S8.E87)\) and for that we let:
νi\\displaystyle\\nu\_\{i\}=\\displaystyle=\{λ⋅\(\(−f′\)\(πiqi\)−α\)ifπi≥pi0otherwise\.\\displaystyle\\left\\\{\\begin\{array\}\[\]\{ccl\}\\lambda\\cdot\\left\(\(\-f^\{\\prime\}\)\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)\-\\alpha\\right\)&\\mbox\{ if \}&\\pi\_\{i\}\\geq p\_\{i\}\\\\ 0&\\lx@intercol\\mbox\{ otherwise \}\\hfil\\lx@intercol\\end\{array\}\\right\.\.In the topmost case, it comes that because of \([108](https://arxiv.org/html/2609.30474#S8.E108)\) andffis convex \(thus−f′\-f^\{\\prime\}is non\-increasing\), we always have\(−f′\)\(πi/qi\)≥α\(\-f^\{\\prime\}\)\\left\(\\pi\_\{i\}/q\_\{i\}\\right\)\\geq\\alphaso𝝂≥𝟎\\bm\{\\nu\}\\geq\\bm\{0\}, which is \([85](https://arxiv.org/html/2609.30474#S8.E85)\)\. We also know thatti<pi⇒πi<pit\_\{i\}<p\_\{i\}\\Rightarrow\\pi\_\{i\}<p\_\{i\}, henceπi≥pi\\pi\_\{i\}\\geq p\_\{i\}impliesti=pit\_\{i\}=p\_\{i\}and KKT \([87](https://arxiv.org/html/2609.30474#S8.E87)\) is satisfied\.
At this stage, we have checked KKT conditions \([81](https://arxiv.org/html/2609.30474#S8.E81)\), \([82](https://arxiv.org/html/2609.30474#S8.E82)\), \([83](https://arxiv.org/html/2609.30474#S8.E83)\), \([84](https://arxiv.org/html/2609.30474#S8.E84)\), \([85](https://arxiv.org/html/2609.30474#S8.E85)\), \([86](https://arxiv.org/html/2609.30474#S8.E86)\), \([87](https://arxiv.org/html/2609.30474#S8.E87)\), \([88](https://arxiv.org/html/2609.30474#S8.E88)\), \([89](https://arxiv.org/html/2609.30474#S8.E89)\)\. To get optimality, we only need the last KKT conditions \([90](https://arxiv.org/html/2609.30474#S8.E90)\) and \([91](https://arxiv.org/html/2609.30474#S8.E91)\) to be satisfied and sinceλ\>0\\lambda\>0, this implies
Df\(𝝅∥𝒒\)\\displaystyle D\_\{f\}\(\\bm\{\\pi\}\\\|\\bm\{q\}\)=\\displaystyle=D\.\\displaystyle D\.Hence, via the construction \([104](https://arxiv.org/html/2609.30474#S8.E104)\) and[106](https://arxiv.org/html/2609.30474#S8.E106)and𝝅=\.𝒕\+\(1−𝟏⊤𝒕\)⋅𝒔\\bm\{\\pi\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\bm\{t\}\+\(1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\)\\cdot\\bm\{s\}, any optimal solution satisfies
\{𝝅∈clampset\(𝒑,Lβ\(−f′\)⋅𝒒,Lα\(−f′\)⋅𝒒\)∩ΔnDf\(𝝅∥𝒒\)=D\.\\displaystyle\\left\\\{\\begin\{array\}\[\]\{l\}\\bm\{\\pi\}\\in\\mathrm\{clampset\}\(\\bm\{p\},L\_\{\\beta\}\(\-f^\{\\prime\}\)\\cdot\\bm\{q\},L\_\{\\alpha\}\(\-f^\{\\prime\}\)\\cdot\\bm\{q\}\)\\cap\\Delta\_\{n\}\\\\ D\_\{f\}\(\\bm\{\\pi\}\\\|\\bm\{q\}\)=D\\end\{array\}\\right\.\.
This ends the proof of \(mdf′→isf\\textsc\{md\}^\{\\prime\}\_\{f\}\\rightarrow\\textsc\{is\}\_\{f\}\)\.
\(isf→mdf′\\textsc\{is\}\_\{f\}\\rightarrow\\textsc\{md\}^\{\\prime\}\_\{f\}\)Pick any𝝅\\bm\{\\pi\}satisfying \([VIII\.3](https://arxiv.org/html/2609.30474#S8.EGx97)\) forα<β∈Im\(−f′\)\\alpha<\\beta\\in\\mathrm\{Im\}\(\-f^\{\\prime\}\)\. We craft
𝒕\\displaystyle\\bm\{t\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}min\{𝒑,𝝅\},\\displaystyle\\min\\\{\\bm\{p\},\\bm\{\\pi\}\\\},\(117\)𝒔\\displaystyle\\bm\{s\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\(1−𝟏⊤𝒕\)−1⋅max\{𝟎,𝝅−𝒑\}\\displaystyle\(1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\)^\{\-1\}\\cdot\\max\\\{\\bm\{0\},\\bm\{\\pi\}\-\\bm\{p\}\\\}\(118\)\(Because of Assumption[3\.3](https://arxiv.org/html/2609.30474#S3.Thmtheorem3), we have𝒑⊤𝒓<1\\bm\{p\}^\{\\top\}\\bm\{r\}<1, so there is one choice only for𝒔\\bm\{s\}as per \([3\.2](https://arxiv.org/html/2609.30474#S3.EGx5)\)\)\. Lemma[A](https://arxiv.org/html/2609.30474#S8.Thmtheorem1)implies𝟏⊤𝒕<1\\bm\{1\}^\{\\top\}\\bm\{t\}<1so𝒔\\bm\{s\}is positive and finite, and1−𝟏⊤𝒕=∑iπi−min\{πi,pi\}=∑imax\{0,πi−pi\}1\-\\bm\{1\}^\{\\top\}\\bm\{t\}=\\sum\_\{i\}\\pi\_\{i\}\-\\min\\\{\\pi\_\{i\},p\_\{i\}\\\}=\\sum\_\{i\}\\max\\\{0,\\pi\_\{i\}\-p\_\{i\}\\\}, so𝒔∈Δn\\bm\{s\}\\in\\Delta\_\{n\}\.
At this point, we easily check KKT \([81](https://arxiv.org/html/2609.30474#S8.E81)\), \([82](https://arxiv.org/html/2609.30474#S8.E82)\), \([83](https://arxiv.org/html/2609.30474#S8.E83)\)\. Furthermore, we observe from \([11](https://arxiv.org/html/2609.30474#S3.E11)\)
\{min\{𝒑,𝝅\}:𝝅∈clampset\(𝒑,Lβ\(−f′\)⋅𝒒,Lα\(−f′\)⋅𝒒\)\}\\displaystyle\\\{\\min\\\{\\bm\{p\},\\bm\{\\pi\}\\\}:\\bm\{\\pi\}\\in\\mathrm\{clampset\}\(\\bm\{p\},L\_\{\\beta\}\(\-f^\{\\prime\}\)\\cdot\\bm\{q\},L\_\{\\alpha\}\(\-f^\{\\prime\}\)\\cdot\\bm\{q\}\)\\\}=\\displaystyle=minset\(𝒑,Lα\(−f′\)⋅𝒒\),\\displaystyle\\mathrm\{minset\}\(\\bm\{p\},L\_\{\\alpha\}\(\-f^\{\\prime\}\)\\cdot\\bm\{q\}\),which is just the way we built𝒕\\bm\{t\}in \([104](https://arxiv.org/html/2609.30474#S8.E104)\)\. Also,
\{max\{𝟎,𝝅−𝒑\}:𝝅∈clampset\(𝒑,Lβ\(−f′\)⋅𝒒,Lα\(−f′\)⋅𝒒\)\}\\displaystyle\\\{\\max\\\{\\bm\{0\},\\bm\{\\pi\}\-\\bm\{p\}\\\}:\\bm\{\\pi\}\\in\\mathrm\{clampset\}\(\\bm\{p\},L\_\{\\beta\}\(\-f^\{\\prime\}\)\\cdot\\bm\{q\},L\_\{\\alpha\}\(\-f^\{\\prime\}\)\\cdot\\bm\{q\}\)\\\}=\\displaystyle=maxset\(𝟎,Lβ\(−f′\)⋅𝒒−𝒑\)\\displaystyle\\mathrm\{maxset\}\(\\bm\{0\},L\_\{\\beta\}\(\-f^\{\\prime\}\)\\cdot\\bm\{q\}\-\\bm\{p\}\)from \([12](https://arxiv.org/html/2609.30474#S3.E12)\) which is just the way we built𝒔\\bm\{s\}from \([106](https://arxiv.org/html/2609.30474#S8.E106)\)\. So the way we craft𝒕\\bm\{t\}and𝒔\\bm\{s\}is the same as for the first step and ends the proof of Theorem[3\.7](https://arxiv.org/html/2609.30474#S3.Thmtheorem7)\.
### VIII\.4Proof of Lemma[3\.9](https://arxiv.org/html/2609.30474#S3.Thmtheorem9)
Sincepi\>0p\_\{i\}\>0,ripi≤0r\_\{i\}p\_\{i\}\\leq 0would imply0∈Lα\(−f′\)0\\in L\_\{\\alpha\}\(\-f^\{\\prime\}\)and thenα=β\\alpha=\\betasince we would be forced to also have0∈Lβ\(−f′\)0\\in L\_\{\\beta\}\(\-f^\{\\prime\}\)\([98](https://arxiv.org/html/2609.30474#S8.E98)\), which is impossible \(Lemma[B](https://arxiv.org/html/2609.30474#S8.Thmtheorem2)\)\.
### VIII\.5Proof of Theorem[3\.12](https://arxiv.org/html/2609.30474#S3.Thmtheorem12)
We show that all KKT conditions of \([f \-MD\-2](https://arxiv.org/html/2609.30474#S3.Ex1)\) are satisfied except eventually one, theff\-divergence constraint \([90](https://arxiv.org/html/2609.30474#S8.E90)\) and then compute bounds for the correspondingDD, so we start with the assumption thatffis differentiable and then remove it\. We reuse the KKT conditions in \([81](https://arxiv.org/html/2609.30474#S8.E81)\), \([82](https://arxiv.org/html/2609.30474#S8.E82)\), \([83](https://arxiv.org/html/2609.30474#S8.E83)\), \([84](https://arxiv.org/html/2609.30474#S8.E84)\), \([85](https://arxiv.org/html/2609.30474#S8.E85)\), \([86](https://arxiv.org/html/2609.30474#S8.E86)\), \([87](https://arxiv.org/html/2609.30474#S8.E87)\), \([88](https://arxiv.org/html/2609.30474#S8.E88)\), \([89](https://arxiv.org/html/2609.30474#S8.E89)\), \([90](https://arxiv.org/html/2609.30474#S8.E90)\)\. We consider𝝅\\bm\{\\pi\}defined by
πi=\.piifi∈𝕀,else\(1\+a\)qiifi∈𝔸,else\(1−b\)qiifi∈𝔹\.\\displaystyle\\pi\_\{i\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}p\_\{i\}\\mbox\{ if \}i\\in\\mathbb\{I\},\\mbox\{ else \}\(1\+a\)q\_\{i\}\\mbox\{ if \}i\\in\\mathbb\{A\},\\mbox\{ else \}\(1\-b\)q\_\{i\}\\mbox\{ if \}i\\in\\mathbb\{B\}\.\(119\)We then have𝝅=𝒕\+\(1−𝟏⊤𝒕\)⋅𝒔\\bm\{\\pi\}=\\bm\{t\}\+\(1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\)\\cdot\\bm\{s\}, letting𝒕=\.𝒑⊙𝒓\\bm\{t\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\bm\{p\}\\odot\\bm\{r\}, for the choices:
ti\\displaystyle t\_\{i\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}piifi∈𝔹∪𝕀elseti=\.\(1\+a\)qiifi∈𝔸,\\displaystyle p\_\{i\}\\mbox\{ if \}i\\in\\mathbb\{B\}\\cup\\mathbb\{I\}\\mbox\{ else \}t\_\{i\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\(1\+a\)q\_\{i\}\\mbox\{ if \}i\\in\\mathbb\{A\},\(120\)si\\displaystyle s\_\{i\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}0ifi∈𝔸∪𝕀elsesi=\.\(1−𝟏⊤𝒕\)−1⋅\(\(1−b\)qi−pi\)ifi∈𝔹\.\\displaystyle 0\\mbox\{ if \}i\\in\\mathbb\{A\}\\cup\\mathbb\{I\}\\mbox\{ else \}s\_\{i\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\(1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\)^\{\-1\}\\cdot\(\(1\-b\)q\_\{i\}\-p\_\{i\}\)\\mbox\{ if \}i\\in\\mathbb\{B\}\.\(121\)\(ifb≥1−minipi/qib\\geq 1\-\\min\_\{i\}p\_\{i\}/q\_\{i\}, we pick any distribution𝒔∈Δn\\bm\{s\}\\in\\Delta\_\{n\}since𝔹=∅\\mathbb\{B\}=\\emptyset\)\. KKT \([81](https://arxiv.org/html/2609.30474#S8.E81)\) holds because all indexes in𝔸\\mathbb\{A\}satisfiypi/qi\>1\+ap\_\{i\}/q\_\{i\}\>1\+a\. KKT \([82](https://arxiv.org/html/2609.30474#S8.E82)\) is satisfied because all indexes in𝔹\\mathbb\{B\}satisfypi/qi<1−bp\_\{i\}/q\_\{i\}<1\-b\. We note from the taxonomy \([33](https://arxiv.org/html/2609.30474#S3.E33)\), \([35](https://arxiv.org/html/2609.30474#S3.E35)\), \([34](https://arxiv.org/html/2609.30474#S3.E34)\),
𝟏⊤𝒕\\displaystyle\\bm\{1\}^\{\\top\}\\bm\{t\}=\\displaystyle=\(1\+a\)q\(𝔸\)\+p\(𝔹\)\+p\(𝕀\)=1−\(p\(𝔸\)−\(1\+a\)q\(𝔸\)\),\\displaystyle\(1\+a\)q\(\\mathbb\{A\}\)\+p\(\\mathbb\{B\}\)\+p\(\\mathbb\{I\}\)=1\-\(p\(\\mathbb\{A\}\)\-\(1\+a\)q\(\\mathbb\{A\}\)\),\(122\)and \([38](https://arxiv.org/html/2609.30474#S3.E38)\) also yields
𝟏⊤𝒕=1−\(\(1−b\)q\(𝔹\)−p\(𝔹\)\)\.\\displaystyle\\bm\{1\}^\{\\top\}\\bm\{t\}=1\-\(\(1\-b\)q\(\\mathbb\{B\}\)\-p\(\\mathbb\{B\}\)\)\.\(123\)We also have from the taxonomy\(1−b\)q\(𝔹\)=∑i∈𝔹πi=∑i∈𝔹pi\+\(1−𝟏⊤𝒕\)⋅si=p\(𝔹\)\+\(1−𝟏⊤𝒕\)⋅𝟏⊤𝒔\(1\-b\)q\(\\mathbb\{B\}\)=\\sum\_\{i\\in\\mathbb\{B\}\}\\pi\_\{i\}=\\sum\_\{i\\in\\mathbb\{B\}\}p\_\{i\}\+\(1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\)\\cdot s\_\{i\}=p\(\\mathbb\{B\}\)\+\(1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\)\\cdot\\bm\{1\}^\{\\top\}\\bm\{s\}which, via identification with \([123](https://arxiv.org/html/2609.30474#S8.E123)\), yields𝟏⊤𝒔=1\\bm\{1\}^\{\\top\}\\bm\{s\}=1because\(1−b\)q\(𝔹\)−p\(𝔹\)\>0\(1\-b\)q\(\\mathbb\{B\}\)\-p\(\\mathbb\{B\}\)\>0from the definition of𝔹\\mathbb\{B\}\. Hence KKT \([83](https://arxiv.org/html/2609.30474#S8.E83)\) holds\. We now compute𝝌\\bm\{\\chi\}and𝝂\\bm\{\\nu\}as
χi\\displaystyle\\chi\_\{i\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}λ\(1−𝟏⊤t\)⋅\(\(a\+b\)ifi∈𝔸,else\(b−1\+piqi\)ifi∈𝕀,else0ifi∈𝔹\),\\displaystyle\\lambda\(1\-\\bm\{1\}^\{\\top\}\{t\}\)\\cdot\\left\(\(a\+b\)\\mbox\{ if \}i\\in\\mathbb\{A\},\\mbox\{ else \}\\left\(b\-1\+\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\right\)\\mbox\{ if \}i\\in\\mathbb\{I\},\\mbox\{ else \}0\\mbox\{ if \}i\\in\\mathbb\{B\}\\right\),\(124\)νi\\displaystyle\\nu\_\{i\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}λ⋅\(\(a\+b\)ifi∈𝔹,else\(1−piqi\+a\)ifi∈𝕀,else0ifi∈𝔸\)\.\\displaystyle\\lambda\\cdot\\left\(\(a\+b\)\\mbox\{ if \}i\\in\\mathbb\{B\},\\mbox\{ else \}\\left\(1\-\\frac\{p\_\{i\}\}\{q\_\{i\}\}\+a\\right\)\\mbox\{ if \}i\\in\\mathbb\{I\},\\mbox\{ else \}0\\mbox\{ if \}i\\in\\mathbb\{A\}\\right\)\.\(125\)We easily check KKT \([84](https://arxiv.org/html/2609.30474#S8.E84)\) and \([85](https://arxiv.org/html/2609.30474#S8.E85)\) \(a,b≥0a,b\\geq 0and the taxonomy \([34](https://arxiv.org/html/2609.30474#S3.E34)\)\)\. \([86](https://arxiv.org/html/2609.30474#S8.E86)\) is checked from \([121](https://arxiv.org/html/2609.30474#S8.E121)\) and \([124](https://arxiv.org/html/2609.30474#S8.E124)\), \([87](https://arxiv.org/html/2609.30474#S8.E87)\) is checked from \([120](https://arxiv.org/html/2609.30474#S8.E120)\) and \([125](https://arxiv.org/html/2609.30474#S8.E125)\), and finally \([88](https://arxiv.org/html/2609.30474#S8.E88)\) is just \([96](https://arxiv.org/html/2609.30474#S8.E96)\) and \([95](https://arxiv.org/html/2609.30474#S8.E95)\)\. We finally letα=\.\(−f\)′\(1\+a\),β=\.\(−f′\)\(1−b\)\\alpha\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\(\-f\)^\{\\prime\}\(1\+a\),\\beta\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\(\-f^\{\\prime\}\)\(1\-b\)and
λ\\displaystyle\\lambda=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}1β−α\>0,\\displaystyle\\frac\{1\}\{\\beta\-\\alpha\}\>0,\(126\)so \([89](https://arxiv.org/html/2609.30474#S8.E89)\) is satisfied; there is only \([90](https://arxiv.org/html/2609.30474#S8.E90)\) which is eventually not satisfied\. Note that ifffis not differentiable, we just switch toα=\.\(−g\)\(1\+a\),β=\.\(−h\)\(1−b\)\\alpha\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\(\-g\)\(1\+a\),\\beta\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\(\-h\)\(1\-b\)forg,h∈∂fg,h\\in\\partial f\.
Hence,anyvaluesa,ba,bas in \([36](https://arxiv.org/html/2609.30474#S3.E36)\), \([37](https://arxiv.org/html/2609.30474#S3.E37)\) define the optimum of \([f \-MD\-2](https://arxiv.org/html/2609.30474#S3.Ex1)\) forsomeD~\\tilde\{D\}that we can compute:
D~=Df\(𝝅∥𝒒\)\\displaystyle\\tilde\{D\}=D\_\{f\}\(\\bm\{\\pi\}\\\|\\bm\{q\}\)=\\displaystyle=∑𝔸f\(πiqi\)⋅qi\+∑𝔹f\(πiqi\)⋅qi\+∑𝕀f\(πiqi\)⋅qi⏟=\.Df\(𝕀\)\\displaystyle\\sum\_\{\\mathbb\{A\}\}f\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)\\cdot q\_\{i\}\+\\sum\_\{\\mathbb\{B\}\}f\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)\\cdot q\_\{i\}\+\\underbrace\{\\sum\_\{\\mathbb\{I\}\}f\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)\\cdot q\_\{i\}\}\_\{\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}D\_\{f\}\(\\mathbb\{I\}\)\}=\\displaystyle=q\(𝔸\)⋅f\(1\+a\)\+q\(𝔹\)⋅f\(1−b\)\+Df\(𝕀\)\.\\displaystyle q\(\\mathbb\{A\}\)\\cdot f\(1\+a\)\+q\(\\mathbb\{B\}\)\\cdot f\(1\-b\)\+D\_\{f\}\(\\mathbb\{I\}\)\.Because of the definition of𝕀\\mathbb\{I\}in \([35](https://arxiv.org/html/2609.30474#S3.E35)\),𝝅\\bm\{\\pi\}in \([119](https://arxiv.org/html/2609.30474#S8.E119)\) and the fact thatffis convex \(therefore continuous\), we can upperboundDf\(𝕀\)D\_\{f\}\(\\mathbb\{I\}\)as
Df\(𝕀\)\\displaystyle D\_\{f\}\(\\mathbb\{I\}\)=\\displaystyle=∑i∈𝕀f\(πiqi\)⋅qi\\displaystyle\\sum\_\{i\\in\\mathbb\{I\}\}f\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)\\cdot q\_\{i\}=\\displaystyle=∑i∈𝕀f\(piqi\)⋅qi\\displaystyle\\sum\_\{i\\in\\mathbb\{I\}\}f\\left\(\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\right\)\\cdot q\_\{i\}=\\displaystyle=q\(𝕀\)⋅∑i∈𝕀qiq\(𝕀\)⋅f\(piqi\)=q\(𝕀\)⋅𝔼i∼𝒒\|𝕀\[f\(piqi\)\]\\displaystyle q\(\\mathbb\{I\}\)\\cdot\\sum\_\{i\\in\\mathbb\{I\}\}\\frac\{q\_\{i\}\}\{q\(\\mathbb\{I\}\)\}\\cdot f\\left\(\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\right\)=q\(\\mathbb\{I\}\)\\cdot\\mathbb\{E\}\_\{i\\sim\\bm\{q\}\_\{\|\\mathbb\{I\}\}\}\\left\[f\\left\(\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\right\)\\right\]=\\displaystyle=q\(𝕀\)⋅f\(u\)\\displaystyle q\(\\mathbb\{I\}\)\\cdot f\(u\)for someu∈\[1−b,1\+a\]u\\in\[1\-b,1\+a\]\(here,𝒒\|𝕀\\bm\{q\}\_\{\|\\mathbb\{I\}\}is𝒒\\bm\{q\}restricted to set𝕀\\mathbb\{I\}\)\. Summarizing,
∃u∈\[1−b,1\+a\]:D~\\displaystyle\\exists u\\in\[1\-b,1\+a\]:\\tilde\{D\}=\\displaystyle=q\(𝔸\)⋅f\(1\+a\)\+q\(𝔹\)⋅f\(1−b\)\+q\(𝕀\)⋅f\(u\)\.\\displaystyle q\(\\mathbb\{A\}\)\\cdot f\(1\+a\)\+q\(\\mathbb\{B\}\)\\cdot f\(1\-b\)\+q\(\\mathbb\{I\}\)\\cdot f\(u\)\.We finally compute the acceptance probability as
Pacc\(MD\)\\displaystyle P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(MD\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}𝟏⊤𝒕\\displaystyle\\bm\{1\}^\{\\top\}\\bm\{t\}=\\displaystyle=∑imin\{pi,\(1\+a\)qi\}\\displaystyle\\sum\_\{i\}\\min\\\{p\_\{i\},\(1\+a\)q\_\{i\}\\\}=\\displaystyle=Pacc\(SD\)\+aq\(𝔸\)\+\(p\(𝕀\>1\)−q\(𝕀\>1\)\),\\displaystyle P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)\+aq\(\\mathbb\{A\}\)\+\(p\(\\mathbb\{I\}\_\{\>1\}\)\-q\(\\mathbb\{I\}\_\{\>1\}\)\),where we have let
𝕀\>1=\.\{i:pi∈qi⋅\(1,1\+a\]\},\\displaystyle\\mathbb\{I\}\_\{\>1\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\left\\\{i:p\_\{i\}\\in q\_\{i\}\\cdot\(1,1\+a\]\\right\\\},\(127\)which ends the proof of Theorem[3\.12](https://arxiv.org/html/2609.30474#S3.Thmtheorem12)\.
### VIII\.6Proof of Theorem[3\.13](https://arxiv.org/html/2609.30474#S3.Thmtheorem13)
Letis~f\(𝒑,𝒒,D\)⊆isf\(𝒑,𝒒,D\)\\tilde\{\\textsc\{is\}\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)\\subseteq\\textsc\{is\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)be defined as:
is~f\(𝒑,𝒒,D\)\\displaystyle\\tilde\{\\textsc\{is\}\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\{𝝅∈Δn:\{∃Lβ\(−h\)≤1≤Lα\(−g\):𝝅∈clampset\(𝒑,Lβ\(−h\)⋅𝒒,Lα\(−g\)⋅𝒒\)Df\(𝝅∥𝒒\)=D\},\\displaystyle\\left\\\{\\bm\{\\pi\}\\in\\Delta\_\{n\}:\\left\\\{\\begin\{array\}\[\]\{l\}\\exists L\_\{\\beta\}\(\-h\)\\leq 1\\leq L\_\{\\alpha\}\(\-g\):\\bm\{\\pi\}\\in\\mathrm\{clampset\}\(\\bm\{p\},L\_\{\\beta\}\(\-h\)\\cdot\\bm\{q\},L\_\{\\alpha\}\(\-g\)\\cdot\\bm\{q\}\)\\\\ D\_\{f\}\(\\bm\{\\pi\}\\\|\\bm\{q\}\)=D\\end\{array\}\\right\.\\right\\\},with the additional constraintα<β\\alpha<\\beta\(we remindg,h∈∂fg,h\\in\\partial f, \([3\.3](https://arxiv.org/html/2609.30474#S3.EGx23)\)\)\. Using the definition of level sets andclampset\\mathrm\{clampset\}, we get that in this case and coordinate\-wise,
clampset\(pi,Lβ\(−h\)qi,Lα\(−g\)qi\)\\displaystyle\\mathrm\{clampset\}\(p\_\{i\},L\_\{\\beta\}\(\-h\)q\_\{i\},L\_\{\\alpha\}\(\-g\)q\_\{i\}\)=\\displaystyle=\{Lβ\(−h\)qiifpi<Lβ\(−h\)qi\(I\)\{z∈Lβ\(−h\)qi:z\>pi\}ifpi∈Lβ\(−h\)qi∧pi<maxLβ\(−h\)qi\(II\)piifpi∈\[maxLβ\(−h\)qi,minLα\(−g\)qi\]\(III\)\{z∈Lα\(−g\)qi:z<pi\}ifpi∈Lα\(−g\)qi∧pi\>minLα\(−g\)qi\(IV\)Lα\(−g\)qiifpi\>Lα\(−g\)qi\(V\)\\displaystyle\\left\\\{\\begin\{array\}\[\]\{ccll\}L\_\{\\beta\}\(\-h\)q\_\{i\}&\\mbox\{ if \}&p\_\{i\}<L\_\{\\beta\}\(\-h\)q\_\{i\}&\(I\)\\\\ \\\{z\\in L\_\{\\beta\}\(\-h\)q\_\{i\}:z\>p\_\{i\}\\\}&\\mbox\{ if \}&p\_\{i\}\\in L\_\{\\beta\}\(\-h\)q\_\{i\}\\wedge p\_\{i\}<\\max L\_\{\\beta\}\(\-h\)q\_\{i\}&\(II\)\\\\ p\_\{i\}&\\mbox\{ if \}&p\_\{i\}\\in\[\\max L\_\{\\beta\}\(\-h\)q\_\{i\},\\min L\_\{\\alpha\}\(\-g\)q\_\{i\}\]&\(III\)\\\\ \\\{z\\in L\_\{\\alpha\}\(\-g\)q\_\{i\}:z<p\_\{i\}\\\}&\\mbox\{ if \}&p\_\{i\}\\in L\_\{\\alpha\}\(\-g\)q\_\{i\}\\wedge p\_\{i\}\>\\min L\_\{\\alpha\}\(\-g\)q\_\{i\}&\(IV\)\\\\ L\_\{\\alpha\}\(\-g\)q\_\{i\}&\\mbox\{ if \}&p\_\{i\}\>L\_\{\\alpha\}\(\-g\)q\_\{i\}&\(V\)\\end\{array\}\\right\.We analyze case by case, noting thatLβ\(−h\)≤1L\_\{\\beta\}\(\-h\)\\leq 1impliesLβ\(−h\)qi≤qiL\_\{\\beta\}\(\-h\)q\_\{i\}\\leq q\_\{i\}, andLα\(−g\)≥1L\_\{\\alpha\}\(\-g\)\\geq 1impliesLα\(−g\)qi≥qiL\_\{\\alpha\}\(\-g\)q\_\{i\}\\geq q\_\{i\}:
- Case \(I\)Here,Lβ\(−h\)qi⊆\[pi,qi\]=\[min\{pi,qi\},max\{pi,qi\}\]L\_\{\\beta\}\(\-h\)q\_\{i\}\\subseteq\[p\_\{i\},q\_\{i\}\]=\[\\min\\\{p\_\{i\},q\_\{i\}\\\},\\max\\\{p\_\{i\},q\_\{i\}\\\}\];
- Case \(II\)is a subset of \(I\) still withpi<qip\_\{i\}<q\_\{i\};
- Case \(III\)in this case,\[min\{pi,qi\},max\{pi,qi\}\]=\[qi,pi\]\[\\min\\\{p\_\{i\},q\_\{i\}\\\},\\max\\\{p\_\{i\},q\_\{i\}\\\}\]=\[q\_\{i\},p\_\{i\}\]and we clearly havepi∈\[qi,pi\]p\_\{i\}\\in\[q\_\{i\},p\_\{i\}\];
- Case \(IV\)we observeqi<\{z∈Lα\(−g\)qi:z<pi\}<piq\_\{i\}<\\\{z\\in L\_\{\\alpha\}\(\-g\)q\_\{i\}:z<p\_\{i\}\\\}<p\_\{i\}so\{z∈Lα\(−g\)qi:z<pi\}⊂\[min\{pi,qi\},max\{pi,qi\}\]\\\{z\\in L\_\{\\alpha\}\(\-g\)q\_\{i\}:z<p\_\{i\}\\\}\\subset\[\\min\\\{p\_\{i\},q\_\{i\}\\\},\\max\\\{p\_\{i\},q\_\{i\}\\\}\];
- Case \(V\)we observe againqi<Lα\(−g\)qi<piq\_\{i\}<L\_\{\\alpha\}\(\-g\)q\_\{i\}<p\_\{i\}, so same conclusion as in \(IV\)\.
To summarize, we have shown that there existsD′D^\{\\prime\}such that
is~f\(𝒑,𝒒,D\)\\displaystyle\\tilde\{\\textsc\{is\}\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)⊆\\displaystyle\\subseteqisfTV\(𝒑,𝒒,D′\),\\displaystyle\\textsc\{is\}\_\{f\_\{\\mathrm\{TV\}\}\}\(\\bm\{p\},\\bm\{q\};D^\{\\prime\}\),Now denotemd~f2\(𝒑,𝒒,D\)\\tilde\{\\textsc\{md\}\}^\{2\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)the set of optimal solutions inmdf2\(𝒑,𝒒,D\)\\textsc\{md\}^\{2\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)havingLβ\(−h\)≤1≤Lα\(−g\)L\_\{\\beta\}\(\-h\)\\leq 1\\leq L\_\{\\alpha\}\(\-g\)withα,β\\alpha,\\betain \([16](https://arxiv.org/html/2609.30474#S3.E16)\), \([17](https://arxiv.org/html/2609.30474#S3.E17)\)\. What we have shown above make the following mappings connections, also using the bijection of Theorem[3\.7](https://arxiv.org/html/2609.30474#S3.Thmtheorem7):
is~f\(𝒑,𝒒,D\)→isfTV\(𝒑,𝒒,D′\)→mdfTV2\(𝒑,𝒒,D′\),\\displaystyle\\tilde\{\\textsc\{is\}\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)\\rightarrow\\textsc\{is\}\_\{f\_\{\\mathrm\{TV\}\}\}\(\\bm\{p\},\\bm\{q\};D^\{\\prime\}\)\\rightarrow\\textsc\{md\}^\{2\}\_\{f\_\{\\mathrm\{TV\}\}\}\(\\bm\{p\},\\bm\{q\};D^\{\\prime\}\),which shows the first part of Theorem[3\.13](https://arxiv.org/html/2609.30474#S3.Thmtheorem13)\.
The second part and \([43](https://arxiv.org/html/2609.30474#S3.E43)\) follows from the proof of Theorem[3\.12](https://arxiv.org/html/2609.30474#S3.Thmtheorem12)and the fact that \(i\) for strictly convex generators, level setsLβ,LαL\_\{\\beta\},L\_\{\\alpha\}are singletons and \(ii\) the set of couples𝒞\(𝒑,𝒒\)\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\)in Definition[3\.11](https://arxiv.org/html/2609.30474#S3.Thmtheorem11)define optimal solution irrespectively of the generator\.
### VIII\.7Proof of Theorem[4\.2](https://arxiv.org/html/2609.30474#S4.Thmtheorem2)
We first need two technical Lemmata\.
###### Lemma D\.
For anyx1,x2,…xT≥0x\_\{1\},x\_\{2\},\.\.\.x\_\{T\}\\geq 0and any realk≥1k\\geq 1, it holds that
∑i∈\[T\]xi2∑i∈\[T\]xi\\displaystyle\\frac\{\\sum\_\{i\\in\[T\]\}x\_\{i\}^\{2\}\}\{\\sum\_\{i\\in\[T\]\}x\_\{i\}\}≤\\displaystyle\\leq∑i∈\[T\]xi2k∑i∈\[T\]xi2k−1\.\\displaystyle\\frac\{\\sum\_\{i\\in\[T\]\}x\_\{i\}^\{2k\}\}\{\\sum\_\{i\\in\[T\]\}x\_\{i\}^\{2k\-1\}\}\.
###### Proof\.
We note that this is equivalent to showing
\(∑i∈\[T\]xi\)⋅\(∑i∈\[T\]xi2k\)−\(∑i∈\[T\]xi2\)⋅\(∑i∈\[T\]xi2k−1\)⏟=\.A\\displaystyle\\underbrace\{\\left\(\\sum\_\{i\\in\[T\]\}x\_\{i\}\\right\)\\cdot\\left\(\\sum\_\{i\\in\[T\]\}x\_\{i\}^\{2k\}\\right\)\-\\left\(\\sum\_\{i\\in\[T\]\}x\_\{i\}^\{2\}\\right\)\\cdot\\left\(\\sum\_\{i\\in\[T\]\}x\_\{i\}^\{2k\-1\}\\right\)\}\_\{\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}A\}≥\\displaystyle\\geq0\.\\displaystyle 0\.\(132\)We develop the sums inAAin two equivalent forms \(swapping indexes\):
A=∑i∈\[T\]∑j∈\[T\]xixj2k−xi2xj2k−1=∑i∈\[T\]∑j∈\[T\]xjxi2k−xj2xi2k−1\.\\displaystyle A=\\sum\_\{i\\in\[T\]\}\\sum\_\{j\\in\[T\]\}x\_\{i\}x\_\{j\}^\{2k\}\-x\_\{i\}^\{2\}x\_\{j\}^\{2k\-1\}=\\sum\_\{i\\in\[T\]\}\\sum\_\{j\\in\[T\]\}x\_\{j\}x\_\{i\}^\{2k\}\-x\_\{j\}^\{2\}x\_\{i\}^\{2k\-1\}\.We then writeAAas the arithmetic average of both expressions and factor:
2A\\displaystyle 2A=\\displaystyle=∑i∈\[T\]∑j∈\[T\]xixj2k−xi2xj2k−1\+xjxi2k−xj2xi2k−1\\displaystyle\\sum\_\{i\\in\[T\]\}\\sum\_\{j\\in\[T\]\}x\_\{i\}x\_\{j\}^\{2k\}\-x\_\{i\}^\{2\}x\_\{j\}^\{2k\-1\}\+x\_\{j\}x\_\{i\}^\{2k\}\-x\_\{j\}^\{2\}x\_\{i\}^\{2k\-1\}\(133\)=\\displaystyle=∑i∈\[T\]∑j∈\[T\]xixj⋅\(xj2k−1−xixj2k−2⏟\(xj−xi\)xj2k−2\+xi2k−1−xjxi2k−2⏟−\(xj−xi\)xi2k−2\)\\displaystyle\\sum\_\{i\\in\[T\]\}\\sum\_\{j\\in\[T\]\}x\_\{i\}x\_\{j\}\\cdot\(\\underbrace\{x\_\{j\}^\{2k\-1\}\-x\_\{i\}x\_\{j\}^\{2k\-2\}\}\_\{\(x\_\{j\}\-x\_\{i\}\)x\_\{j\}^\{2k\-2\}\}\+\\underbrace\{x\_\{i\}^\{2k\-1\}\-x\_\{j\}x\_\{i\}^\{2k\-2\}\}\_\{\-\(x\_\{j\}\-x\_\{i\}\)x\_\{i\}^\{2k\-2\}\}\)=\\displaystyle=∑i∈\[T\]∑j∈\[T\]xixj⋅\(xj−xi\)⋅\(xj2k−2−xi2k−2\)\.\\displaystyle\\sum\_\{i\\in\[T\]\}\\sum\_\{j\\in\[T\]\}x\_\{i\}x\_\{j\}\\cdot\(x\_\{j\}\-x\_\{i\}\)\\cdot\(x\_\{j\}^\{2k\-2\}\-x\_\{i\}^\{2k\-2\}\)\.Sincez↦zuz\\mapsto z^\{u\}is strictly increasing forz≥0z\\geq 0andu\>0u\>0, we get that fork≥1k\\geq 1we always have\(xj−xi\)⋅\(xj2k−2−xi2k−2\)≥0\(x\_\{j\}\-x\_\{i\}\)\\cdot\(x\_\{j\}^\{2k\-2\}\-x\_\{i\}^\{2k\-2\}\)\\geq 0for anyxi,xj≥0x\_\{i\},x\_\{j\}\\geq 0whilexixj≥0x\_\{i\}x\_\{j\}\\geq 0; hence, all terms in \([133](https://arxiv.org/html/2609.30474#S8.E133)\) are≥0\\geq 0, soA≥0A\\geq 0and the Lemma is proven\. ∎
###### Lemma E\.
for anyx1,x2,…xT∈\[0,1\)x\_\{1\},x\_\{2\},\.\.\.x\_\{T\}\\in\[0,1\)and any
y\\displaystyle y∈\\displaystyle\\in\[0,13⋅∑i∈\[T\]xi2∑i∈\[T\]xi\],\\displaystyle\\left\[0,\\frac\{1\}\{3\}\\cdot\\frac\{\\sum\_\{i\\in\[T\]\}x\_\{i\}^\{2\}\}\{\\sum\_\{i\\in\[T\]\}x\_\{i\}\}\\right\],\(134\)it holds that
\(∏i∈\[T\]\(1\+xi\)\)1\+y⋅\(∏i∈\[T\]\(1−xi\)\)1−y\\displaystyle\\left\(\\prod\_\{i\\in\[T\]\}\(1\+x\_\{i\}\)\\right\)^\{1\+y\}\\cdot\\left\(\\prod\_\{i\\in\[T\]\}\(1\-x\_\{i\}\)\\right\)^\{1\-y\}≤\\displaystyle\\leqexp\(−y⋅∑i∈\[T\]xi\)\.\\displaystyle\\exp\\left\(\-y\\cdot\\sum\_\{i\\in\[T\]\}x\_\{i\}\\right\)\.
###### Proof\.
Take the logs and reorganize: we want equivalently
y⋅∑i∈\[T\]xi\+∑i∈\[T\]log\(1−xi2\)\+y⋅∑i∈\[T\]log\(1\+xi1−xi\)\\displaystyle y\\cdot\\sum\_\{i\\in\[T\]\}x\_\{i\}\+\\sum\_\{i\\in\[T\]\}\\log\(1\-x\_\{i\}^\{2\}\)\+y\\cdot\\sum\_\{i\\in\[T\]\}\\log\\left\(\\frac\{1\+x\_\{i\}\}\{1\-x\_\{i\}\}\\right\)≤\\displaystyle\\leq0\.\\displaystyle 0\.Since\|xi\|<1,∀i\|x\_\{i\}\|<1,\\forall i, we consider the \(convergent\) Taylor\-MacLaurin serieslog\(1−x2\)=∑k≥1\(−x2k/k\)\\log\(1\-x^\{2\}\)=\\sum\_\{k\\geq 1\}\(\-x^\{2k\}/k\)andlog\(\(1\+x\)/\(1−x\)\)=∑k≥1\(2/\(2k−1\)\)⋅x2k−1\\log\(\(1\+x\)/\(1\-x\)\)=\\sum\_\{k\\geq 1\}\(2/\(2k\-1\)\)\\cdot x^\{2k\-1\}and plug them in the desired inequality and isolating the term fork=1k=1:
y⋅∑i∈\[T\]xi−∑i∈\[T\]xi2\+2y∑i∈\[T\]xi⏟=\.A\+∑k≥2∑i∈\[T\]\(22k−1⋅yxi2k−1−1k⋅xi2k\)⏟=\.Bk\\displaystyle\\underbrace\{y\\cdot\\sum\_\{i\\in\[T\]\}x\_\{i\}\-\\sum\_\{i\\in\[T\]\}x\_\{i\}^\{2\}\+2y\\sum\_\{i\\in\[T\]\}x\_\{i\}\}\_\{\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}A\}\+\\sum\_\{k\\geq 2\}\\underbrace\{\\sum\_\{i\\in\[T\]\}\\left\(\\frac\{2\}\{2k\-1\}\\cdot yx\_\{i\}^\{2k\-1\}\-\\frac\{1\}\{k\}\\cdot x\_\{i\}^\{2k\}\\right\)\}\_\{\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}B\_\{k\}\}≤\\displaystyle\\leq0\.\\displaystyle 0\.We now show that all ofAAandBk,k≥2B\_\{k\},k\\geq 2are non\-positive\. First, we factorAA:
A\\displaystyle A=\\displaystyle=\(3y−∑i∈\[T\]xi2∑i∈\[T\]xi\)⋅∑i∈\[T\]xi,\\displaystyle\\left\(3y\-\\frac\{\\sum\_\{i\\in\[T\]\}x\_\{i\}^\{2\}\}\{\\sum\_\{i\\in\[T\]\}x\_\{i\}\}\\right\)\\cdot\\sum\_\{i\\in\[T\]\}x\_\{i\},soA≤0A\\leq 0iff
y\\displaystyle y≤\\displaystyle\\leq13⋅∑i∈\[T\]xi2∑i∈\[T\]xi\.\\displaystyle\\frac\{1\}\{3\}\\cdot\\frac\{\\sum\_\{i\\in\[T\]\}x\_\{i\}^\{2\}\}\{\\sum\_\{i\\in\[T\]\}x\_\{i\}\}\.\(135\)and to have eachBk≤0B\_\{k\}\\leq 0we must observe equivalently
y\\displaystyle y≤\\displaystyle\\leq2k−12k⋅∑i∈\[T\]xi2k∑i∈\[T\]xi2k−1,∀k≥2,\\displaystyle\\frac\{2k\-1\}\{2k\}\\cdot\\frac\{\\sum\_\{i\\in\[T\]\}x\_\{i\}^\{2k\}\}\{\\sum\_\{i\\in\[T\]\}x\_\{i\}^\{2k\-1\}\},\\forall k\\geq 2,and from Lemma[D](https://arxiv.org/html/2609.30474#S8.Thmtheorem4)and the fact thatk↦1−\(1/\(2k\)\)k\\mapsto 1\-\(1/\(2k\)\)is strictly increasing, it is sufficient to require
y\\displaystyle y≤\\displaystyle\\leq12⋅∑i∈\[T\]xi2∑i∈\[T\]xi,\\displaystyle\\frac\{1\}\{2\}\\cdot\\frac\{\\sum\_\{i\\in\[T\]\}x\_\{i\}^\{2\}\}\{\\sum\_\{i\\in\[T\]\}x\_\{i\}\},and we observe that this is satisfied if \([135](https://arxiv.org/html/2609.30474#S8.E135)\) holds, which brings the statement of the Lemma\. ∎
We now embark on the proof of Theorem[4\.2](https://arxiv.org/html/2609.30474#S4.Thmtheorem2)\. We use the following inequality[Nock and Nielsen \(2007, Lemma 2\)](https://arxiv.org/html/2609.30474#bib.bib12):
1−ab\\displaystyle 1\-ab≥\\displaystyle\\geq1−a2⋅exp\(−b2⋅log\(1\+a1−a\)\),∀a,b∈\[−1,1\]\.\\displaystyle\\sqrt\{1\-a^\{2\}\}\\cdot\\exp\\left\(\-\\frac\{b\}\{2\}\\cdot\\log\\left\(\\frac\{1\+a\}\{1\-a\}\\right\)\\right\),\\forall a,b\\in\[\-1,1\]\.\(136\)Consider prediction forii\-th training example with associated next token vector𝒚i∈𝒴\\bm\{y\}\_\{i\}\\in\\mathcal\{Y\}, fixa=\.μt∈\[−1,1\]a\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\mu\_\{t\}\\in\[\-1,1\]and
b\\displaystyle b=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}𝒚i⊤𝒉t\(𝒙i\)2ht,∞\.\\displaystyle\\frac\{\\bm\{y\}\_\{i\}^\{\\top\}\\bm\{h\}\_\{t\}\(\\bm\{x\}\_\{i\}\)\}\{2h\_\{t,\\infty\}\}\.We note that Hölder’s inequality and the definition of𝒴\\mathcal\{Y\}imply\|𝒚i⊤𝒉t\(𝒙i\)\|≤‖𝒚i‖1⋅ht,∞\|\\bm\{y\}\_\{i\}^\{\\top\}\\bm\{h\}\_\{t\}\(\\bm\{x\}\_\{i\}\)\|\\leq\\\|\\bm\{y\}\_\{i\}\\\|\_\{1\}\\cdot h\_\{t,\\infty\}and‖𝒚i‖1=ni∗\(1/ni\)\+\(n−ni\)∗\(1/\(n−ni\)\)=2\\\|\\bm\{y\}\_\{i\}\\\|\_\{1\}=n\_\{i\}\*\(1/n\_\{i\}\)\+\(n\-n\_\{i\}\)\*\(1/\(n\-n\_\{i\}\)\)=2, so\|b\|≤1\|b\|\\leq 1\. \([136](https://arxiv.org/html/2609.30474#S8.E136)\) brings for these choices:
1−μt2ht,∞⋅𝒚i⊤𝒉t\(𝒙i\)\\displaystyle 1\-\\frac\{\\mu\_\{t\}\}\{2h\_\{t,\\infty\}\}\\cdot\\bm\{y\}\_\{i\}^\{\\top\}\\bm\{h\}\_\{t\}\(\\bm\{x\}\_\{i\}\)\(137\)≥\\displaystyle\\geq1−μt2⋅exp\(−𝒚i⊤𝒉t\(𝒙i\)4ht,∞ln\(1\+μt1−μt\)\)\\displaystyle\\sqrt\{1\-\\mu^\{2\}\_\{t\}\}\\cdot\\exp\\left\(\-\\frac\{\\bm\{y\}\_\{i\}^\{\\top\}\\bm\{h\}\_\{t\}\(\\bm\{x\}\_\{i\}\)\}\{4h\_\{t,\\infty\}\}\\ln\\left\(\\frac\{1\+\\mu\_\{t\}\}\{1\-\\mu\_\{t\}\}\\right\)\\right\)=1−μt2⋅exp\(−𝒚i⊤\(ct⋅𝒉t⋆\(𝒙i\)\)\),ct=\.14⋅ln\(1\+μt1−μt\),𝒉t⋆=\.1ht,∞⋅𝒉t\.\\displaystyle=\\sqrt\{1\-\\mu^\{2\}\_\{t\}\}\\cdot\\exp\\left\(\-\\bm\{y\}\_\{i\}^\{\\top\}\\left\(c\_\{t\}\\cdot\\bm\{h\}^\{\\star\}\_\{t\}\(\\bm\{x\}\_\{i\}\)\\right\)\\right\),\\\>c\_\{t\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\frac\{1\}\{4\}\\cdot\\ln\\left\(\\frac\{1\+\\mu\_\{t\}\}\{1\-\\mu\_\{t\}\}\\right\),\\\>\\bm\{h\}^\{\\star\}\_\{t\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\frac\{1\}\{h\_\{t,\\infty\}\}\\cdot\\bm\{h\}\_\{t\}\\\>\\\>\.Unraveling the weight update rule, we also obtain:
wT\+1,i⋅∏j=1T\(1−μt2\)\\displaystyle w\_\{T\+1,i\}\\cdot\\prod\_\{j=1\}^\{T\}\{\(1\-\\mu^\{2\}\_\{t\}\)\}=\\displaystyle=w1,i⋅∏j=1T\(1−μt2ht,∞⋅𝒚i⊤𝒉t\(𝒙i\)\)\.\\displaystyle w\_\{1,i\}\\cdot\\prod\_\{j=1\}^\{T\}\{\\left\(1\-\\frac\{\\mu\_\{t\}\}\{2h\_\{t,\\infty\}\}\\cdot\\bm\{y\}\_\{i\}^\{\\top\}\\bm\{h\}\_\{t\}\(\\bm\{x\}\_\{i\}\)\\right\)\}\\\>\\\>\.\(138\)UsingTTtimes \([137](https://arxiv.org/html/2609.30474#S8.E137)\) on the right\-hand side of \([138](https://arxiv.org/html/2609.30474#S8.E138)\) and simplifying yields:
w1,i⋅exp\(−𝒚i⊤∑t=1Tct⋅𝒉t⋆\(𝒙i\)\)\\displaystyle w\_\{1,i\}\\cdot\\exp\\left\(\-\\bm\{y\}\_\{i\}^\{\\top\}\\sum\_\{t=1\}^\{T\}c\_\{t\}\\cdot\\bm\{h\}^\{\\star\}\_\{t\}\(\\bm\{x\}\_\{i\}\)\\right\)≤\\displaystyle\\leqwT\+1,i⋅∏t=1T1−μt2,\\displaystyle w\_\{T\+1,i\}\\cdot\\prod\_\{t=1\}^\{T\}\{\\sqrt\{1\-\\mu^\{2\}\_\{t\}\}\}\\\>\\\>,\(139\)which we then sum fori∈\[m\]i\\in\[m\]and simplify \(𝒘\.∈Δn\\bm\{w\}\_\{\.\}\\in\\Delta\_\{n\}\):
∑i=1mw1,i⋅exp\(−𝒚i⊤∑t=1Tct⋅𝒉t⋆\(𝒙i\)\)\\displaystyle\\sum\_\{i=1\}^\{m\}w\_\{1,i\}\\cdot\\exp\\left\(\-\\bm\{y\}\_\{i\}^\{\\top\}\\sum\_\{t=1\}^\{T\}c\_\{t\}\\cdot\\bm\{h\}^\{\\star\}\_\{t\}\(\\bm\{x\}\_\{i\}\)\\right\)≤\\displaystyle\\leq\(∑i=1mwT\+1,i\)⋅∏t=1T1−μt2\\displaystyle\\left\(\\sum\_\{i=1\}^\{m\}w\_\{T\+1,i\}\\right\)\\cdot\\prod\_\{t=1\}^\{T\}\{\\sqrt\{1\-\\mu^\{2\}\_\{t\}\}\}\(140\)=∏t=1T1−μt2,\\displaystyle=\\prod\_\{t=1\}^\{T\}\{\\sqrt\{1\-\\mu^\{2\}\_\{t\}\}\}\\\>\\\>,Denote
𝑯T\(𝒙\)\\displaystyle\\bm\{H\}\_\{T\}\(\\bm\{x\}\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}1∑t=1Tct⋅∑t=1Tct⋅𝒉t⋆\(𝒙\)\.\\displaystyle\\frac\{1\}\{\\sum\_\{t=1\}^\{T\}c\_\{t\}\}\\cdot\\sum\_\{t=1\}^\{T\}c\_\{t\}\\cdot\\bm\{h\}^\{\\star\}\_\{t\}\(\\bm\{x\}\)\.We get from⟦q≤0⟧≤exp\(−q\),∀q\\llbracket q\\leq 0\\rrbracket\\leq\\exp\(\-q\),\\forall qthe first inequality and from \([140](https://arxiv.org/html/2609.30474#S8.E140)\) the last inequality, of
∑i=1mw1,i⋅⟦𝒚i⊤𝑯T\(𝒙i\)≤θ⟧\\displaystyle\\sum\_\{i=1\}^\{m\}w\_\{1,i\}\\cdot\\Biggl\\llbracket\\bm\{y\}\_\{i\}^\{\\top\}\\bm\{H\}\_\{T\}\(\\bm\{x\}\_\{i\}\)\\leq\\theta\\Biggr\\rrbracket=\\displaystyle=∑i=1mw1,i⋅⟦𝒚i⊤∑t=1Tct⋅𝒉t⋆\(𝒙\)−θ⋅∑t=1Tct≤0⟧\\displaystyle\\sum\_\{i=1\}^\{m\}w\_\{1,i\}\\cdot\\Biggl\\llbracket\\bm\{y\}\_\{i\}^\{\\top\}\\sum\_\{t=1\}^\{T\}c\_\{t\}\\cdot\\bm\{h\}^\{\\star\}\_\{t\}\(\\bm\{x\}\)\-\\theta\\cdot\\sum\_\{t=1\}^\{T\}c\_\{t\}\\leq 0\\Biggr\\rrbracket\(141\)≤\\displaystyle\\leq∑i=1mw1,i⋅exp\(−𝒚i⊤∑t=1Tct⋅𝒉t⋆\(𝒙\)\+θ⋅∑t=1Tct\)\\displaystyle\\sum\_\{i=1\}^\{m\}w\_\{1,i\}\\cdot\\exp\\left\(\-\\bm\{y\}\_\{i\}^\{\\top\}\\sum\_\{t=1\}^\{T\}c\_\{t\}\\cdot\\bm\{h\}^\{\\star\}\_\{t\}\(\\bm\{x\}\)\+\\theta\\cdot\\sum\_\{t=1\}^\{T\}c\_\{t\}\\right\)=exp\(θ⋅∑t=1Tct\)⋅∑i=1mw1,i⋅exp\(−𝒚i⊤∑t=1Tct⋅𝒉t⋆\(𝒙\)\)\\displaystyle=\\exp\\left\(\\theta\\cdot\\sum\_\{t=1\}^\{T\}c\_\{t\}\\right\)\\cdot\\sum\_\{i=1\}^\{m\}w\_\{1,i\}\\cdot\\exp\\left\(\-\\bm\{y\}\_\{i\}^\{\\top\}\\sum\_\{t=1\}^\{T\}c\_\{t\}\\cdot\\bm\{h\}^\{\\star\}\_\{t\}\(\\bm\{x\}\)\\right\)≤\\displaystyle\\leqexp\(θ⋅∑t=1Tct\)⋅∏t=1T1−μt2,∀θ∈ℝ\.\\displaystyle\\exp\\left\(\\theta\\cdot\\sum\_\{t=1\}^\{T\}c\_\{t\}\\right\)\\cdot\\prod\_\{t=1\}^\{T\}\{\\sqrt\{1\-\\mu^\{2\}\_\{t\}\}\},\\forall\\theta\\in\\mathbb\{R\}\.We simplify the RHS using the expression ofctc\_\{t\}in \([137](https://arxiv.org/html/2609.30474#S8.E137)\):
exp\(θ⋅∑t=1Tct\)⋅∏t=1T1−μt2\\displaystyle\\exp\\left\(\\theta\\cdot\\sum\_\{t=1\}^\{T\}c\_\{t\}\\right\)\\cdot\\prod\_\{t=1\}^\{T\}\{\\sqrt\{1\-\\mu^\{2\}\_\{t\}\}\}=\\displaystyle=∏t=1T\(1\+μt1−μt\)θ4⋅∏t=1T1−μt2\\displaystyle\\prod\_\{t=1\}^\{T\}\\left\(\\frac\{1\+\\mu\_\{t\}\}\{1\-\\mu\_\{t\}\}\\right\)^\{\\frac\{\\theta\}\{4\}\}\\cdot\\prod\_\{t=1\}^\{T\}\{\\sqrt\{1\-\\mu^\{2\}\_\{t\}\}\}=\\displaystyle=\(∏t=1T\(1\+μt\)\)1\+θ2⋅\(∏t=1T\(1−μt\)\)1−θ2,\\displaystyle\\sqrt\{\\left\(\\prod\_\{t=1\}^\{T\}\(1\+\\mu\_\{t\}\)\\right\)^\{1\+\\frac\{\\theta\}\{2\}\}\\cdot\\left\(\\prod\_\{t=1\}^\{T\}\(1\-\\mu\_\{t\}\)\\right\)^\{1\-\\frac\{\\theta\}\{2\}\}\},and we use Lemma[E](https://arxiv.org/html/2609.30474#S8.Thmtheorem5): assuming0≤μt<1,∀t0\\leq\\mu\_\{t\}<1,\\forall t\(note that we necessarily have\|μt\|≤1,∀t\|\\mu\_\{t\}\|\\leq 1,\\forall t\), we get
exp\(θ⋅∑t=1Tct\)⋅∏t=1T1−μt2\\displaystyle\\exp\\left\(\\theta\\cdot\\sum\_\{t=1\}^\{T\}c\_\{t\}\\right\)\\cdot\\prod\_\{t=1\}^\{T\}\{\\sqrt\{1\-\\mu^\{2\}\_\{t\}\}\}≤\\displaystyle\\leqexp\(−θ4⋅∑t=1Tμt\),∀θ∈\[0,23⋅∑t=1Tμt2∑t=1Tμt\),\\displaystyle\\exp\\left\(\-\\frac\{\\theta\}\{4\}\\cdot\\sum\_\{t=1\}^\{T\}\\mu\_\{t\}\\right\),\\forall\\theta\\in\\left\[0,\\frac\{2\}\{3\}\\cdot\\frac\{\\sum\_\{t=1\}^\{T\}\\mu\_\{t\}^\{2\}\}\{\\sum\_\{t=1\}^\{T\}\\mu\_\{t\}\}\\right\),which we connect to \([141](https://arxiv.org/html/2609.30474#S8.E141)\) and finally get
ℙi∼𝒘\[𝒚i⊤𝑯T\(𝒙i\)≤θ\]\\displaystyle\\mathbb\{P\}\_\{i\\sim\\bm\{w\}\}\\left\[\\bm\{y\}\_\{i\}^\{\\top\}\\bm\{H\}\_\{T\}\(\\bm\{x\}\_\{i\}\)\\leq\\theta\\right\]≤\\displaystyle\\leqexp\(−θ4⋅∑t=1Tμt\)\.\\displaystyle\\exp\\left\(\-\\frac\{\\theta\}\{4\}\\cdot\\sum\_\{t=1\}^\{T\}\\mu\_\{t\}\\right\)\.\(142\)We then remark that the LHS is a non\-decreasing function ofθ\\thetawhile the RHS is a strictly decreasing function ofθ\\theta, and so we get
∀θ≤23⋅∑t=1Tμt2∑t=1Tμt,ℙi∼𝒘\[𝒚i⊤𝑯T\(𝒙i\)≤θ\]\\displaystyle\\forall\\theta\\leq\\frac\{2\}\{3\}\\cdot\\frac\{\\sum\_\{t=1\}^\{T\}\\mu\_\{t\}^\{2\}\}\{\\sum\_\{t=1\}^\{T\}\\mu\_\{t\}\},\\quad\\mathbb\{P\}\_\{i\\sim\\bm\{w\}\}\\left\[\\bm\{y\}\_\{i\}^\{\\top\}\\bm\{H\}\_\{T\}\(\\bm\{x\}\_\{i\}\)\\leq\\theta\\right\]≤\\displaystyle\\leqexp\(−16⋅∑t=1Tμt2∑t=1Tμt⋅∑t=1Tμt\)\\displaystyle\\exp\\left\(\-\\frac\{1\}\{6\}\\cdot\\frac\{\\sum\_\{t=1\}^\{T\}\\mu\_\{t\}^\{2\}\}\{\\sum\_\{t=1\}^\{T\}\\mu\_\{t\}\}\\cdot\\sum\_\{t=1\}^\{T\}\\mu\_\{t\}\\right\)\(143\)=exp\(−16⋅∑t=1Tμt2\)\.\\displaystyle=\\exp\\left\(\-\\frac\{1\}\{6\}\\cdot\\sum\_\{t=1\}^\{T\}\\mu\_\{t\}^\{2\}\\right\)\.We finally process the event: we remark that
𝒚i⊤𝑯T\(𝒙i\)\\displaystyle\\bm\{y\}\_\{i\}^\{\\top\}\\bm\{H\}\_\{T\}\(\\bm\{x\}\_\{i\}\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}1∑t=1Tct⋅∑t=1Tct⋅𝒚i⊤𝒉t⋆\(𝒙i\)\\displaystyle\\frac\{1\}\{\\sum\_\{t=1\}^\{T\}c\_\{t\}\}\\cdot\\sum\_\{t=1\}^\{T\}c\_\{t\}\\cdot\\bm\{y\}\_\{i\}^\{\\top\}\\bm\{h\}^\{\\star\}\_\{t\}\(\\bm\{x\}\_\{i\}\)\(144\)=\\displaystyle=1∑t=1Tct⋅∑t=1Tct⋅1ht,∞⋅𝒚i⊤𝒉t\(𝒙i\)\\displaystyle\\frac\{1\}\{\\sum\_\{t=1\}^\{T\}c\_\{t\}\}\\cdot\\sum\_\{t=1\}^\{T\}c\_\{t\}\\cdot\\frac\{1\}\{h\_\{t,\\infty\}\}\\cdot\\bm\{y\}\_\{i\}^\{\\top\}\\bm\{h\}\_\{t\}\(\\bm\{x\}\_\{i\}\)=\\displaystyle=1∑t=1Tct⋅∑t=1T\(1ni⋅∑j∈𝒴ictht,jht,∞−1n−ni⋅∑j∈𝒴¯ictht,jht,∞\)\\displaystyle\\frac\{1\}\{\\sum\_\{t=1\}^\{T\}c\_\{t\}\}\\cdot\\sum\_\{t=1\}^\{T\}\\left\(\\frac\{1\}\{n\_\{i\}\}\\cdot\\sum\_\{j\\in\\mathcal\{Y\}\_\{i\}\}\\frac\{c\_\{t\}h\_\{t,j\}\}\{h\_\{t,\\infty\}\}\-\\frac\{1\}\{n\-n\_\{i\}\}\\cdot\\sum\_\{j\\in\\overline\{\\mathcal\{Y\}\}\_\{i\}\}\\frac\{c\_\{t\}h\_\{t,j\}\}\{h\_\{t,\\infty\}\}\\right\)where we have used the definition of𝒉t⋆\\bm\{h\}^\{\\star\}\_\{t\}and the fact that𝒚\.∈𝒴\\bm\{y\}\_\{\.\}\\in\\mathcal\{Y\}, reminding that𝒴i=\.\{j:yij=1/ni\}\\mathcal\{Y\}\_\{i\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\\{j:y\_\{ij\}=1/n\_\{i\}\\\}denotes the set of potential next tokens, while𝒴¯i=\.\{j:yij=−1/\(n−ni\)\}\\overline\{\\mathcal\{Y\}\}\_\{i\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\\{j:y\_\{ij\}=\-1/\(n\-n\_\{i\}\)\\\}denotes the rest of the tokens \(because of the definition of𝒴\\mathcal\{Y\}\)\. Recall that the mentored boosted probability vector is defined as
𝝅~T\(𝒙\)\\displaystyle\\tilde\{\\bm\{\\pi\}\}\_\{T\}\(\\bm\{x\}\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}1ZT⋅∏t=1T\(𝒑t\(𝒙\)\)ct∑u∈\[T\]cu,𝒑t\(𝒙\)=\.1Zt⋅exp\(1ht,∞⋅𝒉t\(𝒙\)\)\\displaystyle\\frac\{1\}\{Z\_\{T\}\}\\cdot\\prod\_\{t=1\}^\{T\}\\left\(\\bm\{p\}\_\{t\}\(\\bm\{x\}\)\\right\)^\{\\frac\{c\_\{t\}\}\{\\sum\_\{u\\in\[T\]\}c\_\{u\}\}\},\\quad\\bm\{p\}\_\{t\}\(\\bm\{x\}\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\frac\{1\}\{Z\_\{t\}\}\\cdot\\exp\\left\(\\frac\{1\}\{h\_\{t,\\infty\}\}\\cdot\\bm\{h\}\_\{t\}\(\\bm\{x\}\)\\right\)so that the event "𝒚i⊤𝑯T\(𝒙i\)≤θ\\bm\{y\}\_\{i\}^\{\\top\}\\bm\{H\}\_\{T\}\(\\bm\{x\}\_\{i\}\)\\leq\\theta" equivalently states, after taking exponentials and normalizing byZT⋅∏t=1TZtct∑u∈\[T\]cuZ\_\{T\}\\cdot\\prod\_\{t=1\}^\{T\}Z\_\{t\}^\{\\frac\{c\_\{t\}\}\{\\sum\_\{u\\in\[T\]\}c\_\{u\}\}\},
\{π~T,j\(𝒙i\),j∈𝒴i\}¯G\\displaystyle\\overline\{\\\{\\tilde\{\\pi\}\_\{T,j\}\(\\bm\{x\}\_\{i\}\),j\\in\\mathcal\{Y\}\_\{i\}\\\}\}^\{G\}≤\\displaystyle\\leq\{π~T,j\(𝒙i\),j∈𝒴¯i\}¯G⋅exp\(θ\)\\displaystyle\\overline\{\\\{\\tilde\{\\pi\}\_\{T,j\}\(\\bm\{x\}\_\{i\}\),j\\in\\overline\{\\mathcal\{Y\}\}\_\{i\}\\\}\}^\{G\}\\cdot\\exp\(\\theta\)whereπ~T,k\(𝒙\)\\tilde\{\\pi\}\_\{T,k\}\(\\bm\{x\}\)is coordinatekkin𝒑T\(𝒙\)\\bm\{p\}\_\{T\}\(\\bm\{x\}\)and for any set of non negative reals𝒜\\mathcal\{A\},𝒜¯G\\overline\{\\mathcal\{A\}\}^\{G\}denotes the geometric average of the elements of𝒜\\mathcal\{A\}\. There remains to put this event in \([143](https://arxiv.org/html/2609.30474#S8.E143)\) to get that∀ρ≤exp\(23⋅∑t=1Tμt2∑t=1Tμt\)\\forall\\rho\\leq\\exp\\left\(\\frac\{2\}\{3\}\\cdot\\frac\{\\sum\_\{t=1\}^\{T\}\\mu\_\{t\}^\{2\}\}\{\\sum\_\{t=1\}^\{T\}\\mu\_\{t\}\}\\right\),
ℙi∼𝒘\[\{π~T,j\(𝒙i\),j∈𝒴i\}¯G≤ρ⋅\{π~T,j\(𝒙i\),j∈𝒴¯i\}¯G\]\\displaystyle\\mathbb\{P\}\_\{i\\sim\\bm\{w\}\}\\left\[\\overline\{\\\{\\tilde\{\\pi\}\_\{T,j\}\(\\bm\{x\}\_\{i\}\),j\\in\\mathcal\{Y\}\_\{i\}\\\}\}^\{G\}\\leq\\rho\\cdot\\overline\{\\\{\\tilde\{\\pi\}\_\{T,j\}\(\\bm\{x\}\_\{i\}\),j\\in\\overline\{\\mathcal\{Y\}\}\_\{i\}\\\}\}^\{G\}\\right\]≤\\displaystyle\\leqexp\(−16⋅∑t=1Tμt2\)\.\\displaystyle\\exp\\left\(\-\\frac\{1\}\{6\}\\cdot\\sum\_\{t=1\}^\{T\}\\mu\_\{t\}^\{2\}\\right\)\.Finally, we remark that\(∑t=1Tμt2\)/∑t=1Tμt=𝔼\[μ\]\+𝕍\[μ\]/𝔼\[μ\]\(\\sum\_\{t=1\}^\{T\}\\mu\_\{t\}^\{2\}\)/\\sum\_\{t=1\}^\{T\}\\mu\_\{t\}=\\mathbb\{E\}\[\\mu\]\+\\mathbb\{V\}\[\\mu\]/\\mathbb\{E\}\[\\mu\], and conclude with the statement of the Theorem\.
### VIII\.8Proof of Lemma[4\.5](https://arxiv.org/html/2609.30474#S4.Thmtheorem5)
Coordinates of𝒗α\\bm\{v\}\_\{\\alpha\}arevα,i=piαqi1−α/Zv\_\{\\alpha,i\}=p\_\{i\}^\{\\alpha\}q\_\{i\}^\{1\-\\alpha\}/ZwithZ=∑j∈\[n\]pjαqj1−αZ=\\sum\_\{j\\in\[n\]\}p\_\{j\}^\{\\alpha\}q\_\{j\}^\{1\-\\alpha\}\. So we want
min\{pi,qi\}≤piαqi1−αZ≤max\{pi,qi\},∀i∈\[n\],\\displaystyle\\min\\\{p\_\{i\},q\_\{i\}\\\}\\leq\\frac\{p\_\{i\}^\{\\alpha\}q\_\{i\}^\{1\-\\alpha\}\}\{Z\}\\leq\\max\\\{p\_\{i\},q\_\{i\}\\\},\\forall i\\in\[n\],\(145\)The AGH inequality yieldsZ≤∑j∈\[n\]αpj\+\(1−α\)qj=1Z\\leq\\sum\_\{j\\in\[n\]\}\\alpha p\_\{j\}\+\(1\-\\alpha\)q\_\{j\}=1so to get \([145](https://arxiv.org/html/2609.30474#S8.E145)\) we only have to guaranteeZ≥piαqi1−α/max\{pi,qi\},∀i∈\[n\]Z\\geq p\_\{i\}^\{\\alpha\}q\_\{i\}^\{1\-\\alpha\}/\\max\\\{p\_\{i\},q\_\{i\}\\\},\\forall i\\in\[n\], or equivalently,
Z\\displaystyle Z≥\\displaystyle\\geqmaxi\(qipi\)⟦pi\>qi⟧−α\.\\displaystyle\\max\_\{i\}\\left\(\\frac\{q\_\{i\}\}\{p\_\{i\}\}\\right\)^\{\\llbracket p\_\{i\}\>q\_\{i\}\\rrbracket\-\\alpha\}\.\(146\)Introducingi∗∈\[n\]i\_\{\*\}\\in\[n\]the index realizing the max \(assuming it is unique for simplicity\), the RHS can be reformulated as \(we recall thatα∈\(0,1\)\\alpha\\in\(0,1\)\)
maxi\(qipi\)⟦pi\>qi⟧−α\\displaystyle\\max\_\{i\}\\left\(\\frac\{q\_\{i\}\}\{p\_\{i\}\}\\right\)^\{\\llbracket p\_\{i\}\>q\_\{i\}\\rrbracket\-\\alpha\}=\\displaystyle=maximin\{\(qipi\)1−α,\(piqi\)α\}\\displaystyle\\max\_\{i\}\\min\\left\\\{\\left\(\\frac\{q\_\{i\}\}\{p\_\{i\}\}\\right\)^\{1\-\\alpha\},\\left\(\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\right\)^\{\\alpha\}\\right\\\}=\\displaystyle=\(maximin\{qipi,piqi\}\)⟦pi∗\>qi∗⟧⋅\(1−α\)\+⟦pi∗<qi∗⟧⋅α\.\\displaystyle\\left\(\\max\_\{i\}\\min\\left\\\{\\frac\{q\_\{i\}\}\{p\_\{i\}\},\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\right\\\}\\right\)^\{\\llbracket p\_\{i\_\{\*\}\}\>q\_\{i\_\{\*\}\}\\rrbracket\\cdot\(1\-\\alpha\)\+\\llbracket p\_\{i\_\{\*\}\}<q\_\{i\_\{\*\}\}\\rrbracket\\cdot\\alpha\}\.We use the tempered logarithm and tempered exponential as\([Naudts, 2011](https://arxiv.org/html/2609.30474#bib.bib19), Chapter 7\):
logt\(z\)=\.11−t⋅\(z1−t−1\)\\displaystyle\\log\_\{t\}\(z\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\frac\{1\}\{1\-t\}\\cdot\\left\(z^\{1\-t\}\-1\\right\),expt\(z\)=\.\[1\+\(1−t\)z\]\+1/\(1−t\)\(\[z\]\+=\.max\{0,z\}\),\\displaystyle\\exp\_\{t\}\(z\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\left\[1\+\(1\-t\)z\\right\]^\{1/\(1\-t\)\}\_\{\+\}\\quad\(\[z\]\_\{\+\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\max\\\{0,z\\\}\),where the caset=1t=1is supposed to be the extension by continuity to thelog\\logandexp\\expfunctions, respectively\. We shall considert∈\(0,1\)t\\in\(0,1\), values for which the concavity / convexity of functions is the same as fort=1t=1, see also[Amid et al\. \(2024\)](https://arxiv.org/html/2609.30474#bib.bib16);[Amid et al\. \(2023\)](https://arxiv.org/html/2609.30474#bib.bib17);[Nock et al\. \(2023\)](https://arxiv.org/html/2609.30474#bib.bib18)\.
Suppose first thatpi∗\>qi∗p\_\{i\_\{\*\}\}\>q\_\{i\_\{\*\}\}\. Remark thatZ11−αZ^\{\\frac\{1\}\{1\-\\alpha\}\}can be conveniently rewritten as \(we recall thatα∈\(0,1\)\\alpha\\in\(0,1\)andZ∈\[0,1\]Z\\in\[0,1\]\)
Z11−α\\displaystyle Z^\{\\frac\{1\}\{1\-\\alpha\}\}=\\displaystyle=\[1\+\(1−α\)⋅\(11−α⋅∑i∈\[n\]piαqi1−α−pi\)\]11−α\\displaystyle\\left\[1\+\(1\-\\alpha\)\\cdot\\left\(\\frac\{1\}\{1\-\\alpha\}\\cdot\\sum\_\{i\\in\[n\]\}p\_\{i\}^\{\\alpha\}q\_\{i\}^\{1\-\\alpha\}\-p\_\{i\}\\right\)\\right\]^\{\\frac\{1\}\{1\-\\alpha\}\}=\\displaystyle=\[1\+\(1−α\)⋅\(∑i∈\[n\]pi⋅\{11−α⋅\(\(qipi\)1−α−1\)\}\)\]11−α\\displaystyle\\left\[1\+\(1\-\\alpha\)\\cdot\\left\(\\sum\_\{i\\in\[n\]\}p\_\{i\}\\cdot\\left\\\{\\frac\{1\}\{1\-\\alpha\}\\cdot\\left\(\\left\(\\frac\{q\_\{i\}\}\{p\_\{i\}\}\\right\)^\{1\-\\alpha\}\-1\\right\)\\right\\\}\\right\)\\right\]^\{\\frac\{1\}\{1\-\\alpha\}\}=\\displaystyle=expt𝔼i∼𝒑logtqipi,witht=\.α,\\displaystyle\\exp\_\{t\}\\mathbb\{E\}\_\{i\\sim\\bm\{p\}\}\\log\_\{t\}\\frac\{q\_\{i\}\}\{p\_\{i\}\},\\quad\\mbox\{ with $t\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\alpha$\},and ifpi∗<qi∗p\_\{i\_\{\*\}\}<q\_\{i\_\{\*\}\}, we rewriteZ1αZ^\{\\frac\{1\}\{\\alpha\}\}as
Z1α\\displaystyle Z^\{\\frac\{1\}\{\\alpha\}\}=\\displaystyle=\[1\+\(1−\(1−α\)\)⋅\(11−\(1−α\)⋅∑i∈\[n\]piαqi1−α−qi\)\]11−\(1−α\)\\displaystyle\\left\[1\+\(1\-\(1\-\\alpha\)\)\\cdot\\left\(\\frac\{1\}\{1\-\(1\-\\alpha\)\}\\cdot\\sum\_\{i\\in\[n\]\}p\_\{i\}^\{\\alpha\}q\_\{i\}^\{1\-\\alpha\}\-q\_\{i\}\\right\)\\right\]^\{\\frac\{1\}\{1\-\(1\-\\alpha\)\}\}=\\displaystyle=\[1\+\(1−\(1−α\)\)⋅\(∑i∈\[n\]qi⋅\{11−\(1−α\)⋅\(\(piqi\)1−\(1−α\)−1\)\}\)\]11−\(1−α\)\\displaystyle\\left\[1\+\(1\-\(1\-\\alpha\)\)\\cdot\\left\(\\sum\_\{i\\in\[n\]\}q\_\{i\}\\cdot\\left\\\{\\frac\{1\}\{1\-\(1\-\\alpha\)\}\\cdot\\left\(\\left\(\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\right\)^\{1\-\(1\-\\alpha\)\}\-1\\right\)\\right\\\}\\right\)\\right\]^\{\\frac\{1\}\{1\-\(1\-\\alpha\)\}\}=\\displaystyle=expt𝔼i∼𝒒logtpiqi,witht=\.1−α\.\\displaystyle\\exp\_\{t\}\\mathbb\{E\}\_\{i\\sim\\bm\{q\}\}\\log\_\{t\}\\frac\{p\_\{i\}\}\{q\_\{i\}\},\\quad\\mbox\{ with $t\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}1\-\\alpha$\}\.To summarize, \([146](https://arxiv.org/html/2609.30474#S8.E146)\) is equivalent to having
expα𝔼i∼𝒑logαqipi\\displaystyle\\exp\_\{\\alpha\}\\mathbb\{E\}\_\{i\\sim\\bm\{p\}\}\\log\_\{\\alpha\}\\frac\{q\_\{i\}\}\{p\_\{i\}\}≥\\displaystyle\\geqmaximin\{qipi,piqi\}ifpi∗\>qi∗,\\displaystyle\\max\_\{i\}\\min\\left\\\{\\frac\{q\_\{i\}\}\{p\_\{i\}\},\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\right\\\}\\quad\\mbox\{ if $p\_\{i\_\{\*\}\}\>q\_\{i\_\{\*\}\}$\},exp1−α𝔼i∼𝒒log1−αpiqi\\displaystyle\\exp\_\{1\-\\alpha\}\\mathbb\{E\}\_\{i\\sim\\bm\{q\}\}\\log\_\{1\-\\alpha\}\\frac\{p\_\{i\}\}\{q\_\{i\}\}≥\\displaystyle\\geqmaximin\{qipi,piqi\}ifpi∗<qi∗,\\displaystyle\\max\_\{i\}\\min\\left\\\{\\frac\{q\_\{i\}\}\{p\_\{i\}\},\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\right\\\}\\quad\\mbox\{ if $p\_\{i\_\{\*\}\}<q\_\{i\_\{\*\}\}$\},and provided these hold, we get
𝒗α∈\[min\{𝒑,𝒒\},max\{𝒑,𝒒\}\],\\displaystyle\\bm\{v\}\_\{\\alpha\}\\in\[\\min\\\{\\bm\{p\},\\bm\{q\}\\\},\\max\\\{\\bm\{p\},\\bm\{q\}\\\}\],\(147\)which is the statement of the Lemma\.
### VIII\.9Proof of Theorem[4\.7](https://arxiv.org/html/2609.30474#S4.Thmtheorem7)
We proceed in two steps, first showing the bullet elements of the statement, and then showing the bound onDfTV\(𝒗α∥𝒒\)D\_\{f\_\{\\mathrm\{TV\}\}\}\(\\bm\{v\}\_\{\\alpha\}\\\|\\bm\{q\}\)in \([62](https://arxiv.org/html/2609.30474#S4.E62)\)\. Our first step proof uses the following technical Lemma\.
###### Lemma F\.
letXXbe a random variable with values in an interval\[u,v\]\[u,v\]satisfyingu\>0u\>0and𝔼\[X\]=1\\mathbb\{E\}\[X\]=1\. Then for anyt∈\[0,1\]t\\in\[0,1\]it holds that
expt𝔼\[logtX\]\\displaystyle\\exp\_\{t\}\\mathbb\{E\}\[\\log\_\{t\}X\]≥\\displaystyle\\geq\(uv\)t\.\\displaystyle\\left\(\\frac\{u\}\{v\}\\right\)^\{t\}\.\(148\)
###### Proof\.
We first prove the result fort∈\(0,1\)t\\in\(0,1\)\. Sincelogt\\log\_\{t\}is concave for anyt∈\[0,1\]t\\in\[0,1\], it lies above any of its secants in the interval defined by the intersections, so we get
logtz\\displaystyle\\log\_\{t\}z≥\\displaystyle\\geqv−zv−u⋅logt\(u\)\+z−uv−u⋅logt\(v\),\\displaystyle\\frac\{v\-z\}\{v\-u\}\\cdot\\log\_\{t\}\(u\)\+\\frac\{z\-u\}\{v\-u\}\\cdot\\log\_\{t\}\(v\),which yields, since𝔼\[X\]=1\\mathbb\{E\}\[X\]=1,
𝔼\[logtX\]\\displaystyle\\mathbb\{E\}\[\\log\_\{t\}X\]≥\\displaystyle\\geqv−1v−u⋅logt\(u\)\+1−uv−u⋅logt\(v\),\\displaystyle\\frac\{v\-1\}\{v\-u\}\\cdot\\log\_\{t\}\(u\)\+\\frac\{1\-u\}\{v\-u\}\\cdot\\log\_\{t\}\(v\),and lettingp=\.\(v−1\)/\(v−u\)∈\[0,1\]p\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\(v\-1\)/\(v\-u\)\\in\[0,1\], yields the equivalent inequality after expressinglogt\\log\_\{t\}:
𝔼\[logtX\]\\displaystyle\\mathbb\{E\}\[\\log\_\{t\}X\]≥\\displaystyle\\geqpu1−t\+\(1−p\)v1−t−11−t,\\displaystyle\\frac\{pu^\{1\-t\}\+\(1\-p\)v^\{1\-t\}\-1\}\{1\-t\},and sinceexpt\\exp\_\{t\}is non\-decreasing for anyt∈\[0,1\]t\\in\[0,1\], ensures \([148](https://arxiv.org/html/2609.30474#S8.E148)\) provided the sufficient condition holds:expt\(pu1−t\+\(1−p\)v1−t−11−t\)≥\(uv\)t\\exp\_\{t\}\\left\(\\frac\{pu^\{1\-t\}\+\(1\-p\)v^\{1\-t\}\-1\}\{1\-t\}\\right\)\\geq\\left\(\\frac\{u\}\{v\}\\right\)^\{t\}, which simplifies to checking the condition:
pu1−t\+\(1−p\)v1−t\\displaystyle pu^\{1\-t\}\+\(1\-p\)v^\{1\-t\}≥\\displaystyle\\geq\(uv\)t\(1−t\)\.\\displaystyle\\left\(\\frac\{u\}\{v\}\\right\)^\{t\(1\-t\)\}\.\(149\)Sinceu≤𝔼\[X\]=1≤vu\\leq\\mathbb\{E\}\[X\]=1\\leq v, it is enough to check this inequality for any0<u<1,v\>10<u<1,v\>1withp=\.\(v−1\)/\(v−u\)∈\[0,1\]p\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\(v\-1\)/\(v\-u\)\\in\[0,1\]\. Let us simplify it once more\. Definex=\.v/u≥1x\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}v/u\\geq 1\. The RHS of \([149](https://arxiv.org/html/2609.30474#S8.E149)\) only depends onxxandtt, and it turns out the LHS simplifies:
pu1−t\+\(1−p\)v1−t\\displaystyle pu^\{1\-t\}\+\(1\-p\)v^\{1\-t\}=\\displaystyle=u1−t⋅\(p\+\(1−p\)x1−t\)\\displaystyle u^\{1\-t\}\\cdot\\left\(p\+\(1\-p\)x^\{1\-t\}\\right\)=\\displaystyle=p\+\(1−p\)x1−t\(p\+\(1−p\)x\)1−t,\\displaystyle\\frac\{p\+\(1\-p\)x^\{1\-t\}\}\{\(p\+\(1\-p\)x\)^\{1\-t\}\},sincepu\+\(1−p\)v=1pu\+\(1\-p\)v=1andv=uxv=uxwhich yieldsu=1/\(p\+\(1−p\)x\)u=1/\(p\+\(1\-p\)x\)\. Using these expressions depending onp,x,tp,x,tand taking logs in \([149](https://arxiv.org/html/2609.30474#S8.E149)\), we want to show
log\(p\+\(1−p\)x1−t\)−\(1−t\)log\(p\+\(1−p\)x\)⏟=\.gx\(p\)\\displaystyle\\underbrace\{\\log\(p\+\(1\-p\)x^\{1\-t\}\)\-\(1\-t\)\\log\(p\+\(1\-p\)x\)\}\_\{\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}g\_\{x\}\(p\)\}≥\\displaystyle\\geq−t\(1−t\)log\(x\)\.\\displaystyle\-t\(1\-t\)\\log\(x\)\.\(150\)To show this, let us first analyzegx\(p\)g\_\{x\}\(p\)\. We have
gx′\(p\)\\displaystyle g^\{\\prime\}\_\{x\}\(p\)=\\displaystyle=\(1−t\)\(x−1\)p\+\(1−p\)x−x1−t−1p\+\(1−p\)x1−t,\\displaystyle\\frac\{\(1\-t\)\(x\-1\)\}\{p\+\(1\-p\)x\}\-\\frac\{x^\{1\-t\}\-1\}\{p\+\(1\-p\)x^\{1\-t\}\},and we obtaingx′\(p\)≤0g^\{\\prime\}\_\{x\}\(p\)\\leq 0iff
p\\displaystyle p≤\\displaystyle\\leqQ=\.\(1−t\)x1−t\+tx2−t−xt\(x−1\)\(x1−t−1\),\\displaystyle Q\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\frac\{\(1\-t\)x^\{1\-t\}\+tx^\{2\-t\}\-x\}\{t\(x\-1\)\(x^\{1\-t\}\-1\)\},and we remark that
Q\\displaystyle Q=\\displaystyle=1\+x1−t\+\(t\+1\)x−tt\(x−1\)\(x1−t−1\)\\displaystyle 1\+\\frac\{x^\{1\-t\}\+\(t\+1\)x\-t\}\{t\(x\-1\)\(x^\{1\-t\}\-1\)\}≥\\displaystyle\\geq1,\\displaystyle 1,sincex≥1,t∈\[0,1\]x\\geq 1,t\\in\[0,1\], so the minimum ofgx\(p\)g\_\{x\}\(p\)forp∈\[0,1\]p\\in\[0,1\]is obtained atgx\(1\)=0g\_\{x\}\(1\)=0, showinggx\(p\)≥0,∀x≥1,t∈\[0,1\]g\_\{x\}\(p\)\\geq 0,\\forall x\\geq 1,t\\in\[0,1\]\. Since the RHS of \([150](https://arxiv.org/html/2609.30474#S8.E150)\) is≤0\\leq 0under these conditions, \([150](https://arxiv.org/html/2609.30474#S8.E150)\) is proven and so is Lemma[F](https://arxiv.org/html/2609.30474#S8.Thmtheorem6)fort∈\(0,1\)t\\in\(0,1\)\. We then remark the continuity of both functions in \([148](https://arxiv.org/html/2609.30474#S8.E148)\) fort∈\[0,1\]t\\in\[0,1\]so taking the limits in0\+0^\{\+\}and1−1^\{\-\}completes the proof\. ∎
We now shift to analyzing boostingunder the constraintthat we must keep \([147](https://arxiv.org/html/2609.30474#S8.E147)\) true\. Our first step consists of showing constraints on the exponentα\\alpha, and for this we distinguish two cases\.
Case 1:we first assumepi∗\>qi∗p\_\{i\_\{\*\}\}\>q\_\{i\_\{\*\}\}, so we work with the constraint
expα𝔼i∼𝒑logαqipi\\displaystyle\\exp\_\{\\alpha\}\\mathbb\{E\}\_\{i\\sim\\bm\{p\}\}\\log\_\{\\alpha\}\\frac\{q\_\{i\}\}\{p\_\{i\}\}≥\\displaystyle\\geqmaximin\{qipi,piqi\},\\displaystyle\\max\_\{i\}\\min\\left\\\{\\frac\{q\_\{i\}\}\{p\_\{i\}\},\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\right\\\},which, from Lemma[F](https://arxiv.org/html/2609.30474#S8.Thmtheorem6), holds if we haveαlog\(miniqipimaxiqipi\)≥log\(maximin\{qipi,piqi\}\)\\alpha\\log\\left\(\\frac\{\\min\_\{i\}\\frac\{q\_\{i\}\}\{p\_\{i\}\}\}\{\\max\_\{i\}\\frac\{q\_\{i\}\}\{p\_\{i\}\}\}\\right\)\\geq\\log\\left\(\\max\_\{i\}\\min\\left\\\{\\frac\{q\_\{i\}\}\{p\_\{i\}\},\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\right\\\}\\right\), that is,
α\\displaystyle\\alpha≤\\displaystyle\\leqlog\(minimax\{qipi,piqi\}\)log\(maxiqipiminiqipi\)\.\\displaystyle\\frac\{\\log\\left\(\\min\_\{i\}\\max\\left\\\{\\frac\{q\_\{i\}\}\{p\_\{i\}\},\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\right\\\}\\right\)\}\{\\log\\left\(\\frac\{\\max\_\{i\}\\frac\{q\_\{i\}\}\{p\_\{i\}\}\}\{\\min\_\{i\}\\frac\{q\_\{i\}\}\{p\_\{i\}\}\}\\right\)\}\.\(151\)This provides our first constraint onα\\alpha\. We move to the alternative one\.
Case 2:if insteadpi∗<qi∗p\_\{i\_\{\*\}\}<q\_\{i\_\{\*\}\}, we want from Lemma[4\.5](https://arxiv.org/html/2609.30474#S4.Thmtheorem5)
exp1−α𝔼i∼𝒒log1−αpiqi\\displaystyle\\exp\_\{1\-\\alpha\}\\mathbb\{E\}\_\{i\\sim\\bm\{q\}\}\\log\_\{1\-\\alpha\}\\frac\{p\_\{i\}\}\{q\_\{i\}\}≥\\displaystyle\\geqmaximin\{qipi,piqi\},\\displaystyle\\max\_\{i\}\\min\\left\\\{\\frac\{q\_\{i\}\}\{p\_\{i\}\},\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\right\\\},and so we want from Lemma[F](https://arxiv.org/html/2609.30474#S8.Thmtheorem6)
1−α\\displaystyle 1\-\\alpha≤\\displaystyle\\leqlog\(minimax\{qipi,piqi\}\)log\(maxiqipiminiqipi\)\.\\displaystyle\\frac\{\\log\\left\(\\min\_\{i\}\\max\\left\\\{\\frac\{q\_\{i\}\}\{p\_\{i\}\},\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\right\\\}\\right\)\}\{\\log\\left\(\\frac\{\\max\_\{i\}\\frac\{q\_\{i\}\}\{p\_\{i\}\}\}\{\\min\_\{i\}\\frac\{q\_\{i\}\}\{p\_\{i\}\}\}\\right\)\}\.\(152\)
At this point, if we can provide boosting coefficientsc1=f\(μ1\)c\_\{1\}=f\(\\mu\_\{1\}\)andc2=f\(μ2\)c\_\{2\}=f\(\\mu\_\{2\}\)whereffis given in \([47](https://arxiv.org/html/2609.30474#S4.E47)\) \(main file\), such that \(a\)
α\\displaystyle\\alpha=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}cucu\+cv\\displaystyle\\frac\{c\_\{u\}\}\{c\_\{u\}\+c\_\{v\}\}\(153\)satisfies whichever \([151](https://arxiv.org/html/2609.30474#S8.E151)\) or \([152](https://arxiv.org/html/2609.30474#S8.E152)\) is relevant \(where distinctu,vu,vare in\{1,2\}\\\{1,2\\\}\),and\(b\) the associated boosting advantage is large enough to show that the combination of two models does satisfy the exponential decrease associated in \([49](https://arxiv.org/html/2609.30474#S4.E49)\) \(main file\), then we are done: in all cases, the boosted solution is also optimal for the TV\-MD problem\. What we need to do is find the sequence of models \(among drafter and target\) to include in the boosted model, and find the edgesμ1\\mu\_\{1\}andμ2\\mu\_\{2\}such that the related boosting advantage is large enough\.
In the context of boosting, we want the best guarantee from the boosting advantage standpoint, so let us assume that the first model we include is the target model, so thatμ1=μtarget\\mu\_\{1\}=\\mu\_\{\\mbox\{\\tiny\{target\}\}\}, and show how to collapse both cases above as a single one that controlsμ2\\mu\_\{2\}as an eventual clamping ofμdrafter\\mu\_\{\\mbox\{\\tiny\{drafter\}\}\}\. Ifpi∗\>qi∗p\_\{i\_\{\*\}\}\>q\_\{i\_\{\*\}\}, we fixu=1,v=2u=1,v=2in \([153](https://arxiv.org/html/2609.30474#S8.E153)\) and need to show \([151](https://arxiv.org/html/2609.30474#S8.E151)\) forα=c1/\(c1\+c2\)\\alpha=c\_\{1\}/\(c\_\{1\}\+c\_\{2\}\)\. Otherwise ifpi∗<qi∗p\_\{i\_\{\*\}\}<q\_\{i\_\{\*\}\}, we permuteu=2,v=1u=2,v=1in \([153](https://arxiv.org/html/2609.30474#S8.E153)\) so that1−α=c1/\(c1\+c2\)1\-\\alpha=c\_\{1\}/\(c\_\{1\}\+c\_\{2\}\)and \([152](https://arxiv.org/html/2609.30474#S8.E152)\) is the same condition as for the first case\.
We now include \([ED](https://arxiv.org/html/2609.30474#S4.Ex18)\)\. The RHS of \([151](https://arxiv.org/html/2609.30474#S8.E151)\) is≥log\(1\+ϱ𝒑𝒒\)/log\(1/ε𝒑𝒒2\)\\geq\\log\(1\+\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\)/\\log\(1/\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}^\{2\}\)and sinceα=c2/\(c1\+c2\)\\alpha=c\_\{2\}/\(c\_\{1\}\+c\_\{2\}\), \([151](https://arxiv.org/html/2609.30474#S8.E151)\) is equivalent to having
c2\\displaystyle c\_\{2\}≤\\displaystyle\\leqlog\(1\+ϱ𝒑𝒒\)log\(1/ε𝒑𝒒2\)−log\(1\+ϱ𝒑𝒒\)⋅c1,\\displaystyle\\frac\{\\log\(1\+\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\)\}\{\\log\(1/\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}^\{2\}\)\-\\log\(1\+\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\)\}\\cdot c\_\{1\},which, predictably, prevents a too large boosting coefficient for the drafter model\. We now need another technical Lemma
###### Lemma G\.
For anyz,xz,xsuch thatz∈\(0,1\]z\\in\(0,1\],x≥0x\\geq 0andz\(1\+x\)≤1z\(1\+x\)\\leq 1,
ln\(1\+x\)ln\(1z2\)−ln\(1\+x\)\\displaystyle\\frac\{\\ln\(1\+x\)\}\{\\ln\\left\(\\frac\{1\}\{z^\{2\}\}\\right\)\-\\ln\(1\+x\)\}≥\\displaystyle\\geqz⋅x\.\\displaystyle z\\cdot x\.\(154\)
###### Proof\.
The denominator in the LHS of \([154](https://arxiv.org/html/2609.30474#S8.E154)\) being non negative, we reformulate the inequality asfx\(z\)=\.\(1\+zx\)ln\(1\+x\)\+2zxln\(z\)≥0f\_\{x\}\(z\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\(1\+zx\)\\ln\(1\+x\)\+2zx\\ln\(z\)\\geq 0, where we treatxxas a constant\. We easily get thatfx\(z\)f\_\{x\}\(z\)is strictly convex and has a global minimum atz∗=\.1/\(e1\+x\)z\_\{\*\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}1/\(e\\sqrt\{1\+x\}\)\. If the minimum falls in the set of constraints for the value ofxx\(we can show that this happens iffx≤e2−1x\\leq e^\{2\}\-1\), we get that
fx\(z∗\)\\displaystyle f\_\{x\}\(z\_\{\*\}\)=\\displaystyle=ln\(1\+x\)−2xe1\+x,\\displaystyle\\ln\(1\+x\)\-\\frac\{2x\}\{e\\sqrt\{1\+x\}\},and so, lettingy=\.1\+xy\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\sqrt\{1\+x\}, we need to showeylny−y2\+1≥0,∀y∈\[1,e\]ey\\ln y\-y^\{2\}\+1\\geq 0,\\forall y\\in\[1,e\], which is easily checked \(the function is 0 iny=1y=1and its derivative is≥0\\geq 0on\[1,e\]\[1,e\]\)\. A similar proof holds ifz∗z\_\{\*\}is not in the set of constraints\. ∎
Lemma[G](https://arxiv.org/html/2609.30474#S8.Thmtheorem7)yields, sinceε𝒑𝒒∈\(0,1\]\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\in\(0,1\],ϱ𝒑𝒒≥0\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\\geq 0andε𝒑𝒒\(1\+ϱ𝒑𝒒\)≤1\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\(1\+\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\)\\leq 1,
log\(1\+ϱ𝒑𝒒\)log\(1/ε𝒑𝒒2\)−log\(1\+ϱ𝒑𝒒\)\\displaystyle\\frac\{\\log\(1\+\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\)\}\{\\log\(1/\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}^\{2\}\)\-\\log\(1\+\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\)\}≥\\displaystyle\\geqϱ𝒑𝒒⋅ε𝒑𝒒,\\displaystyle\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\\cdot\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\},so we can simplify the requirement toc2≤ϱ𝒑𝒒ε𝒑𝒒c1c\_\{2\}\\leq\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}c\_\{1\}, which equivalently reads
1\+μ21−μ2\\displaystyle\\frac\{1\+\\mu\_\{2\}\}\{1\-\\mu\_\{2\}\}≤\\displaystyle\\leq\(1\+μtarget1−μtarget\)ϱ𝒑𝒒ε𝒑𝒒,\\displaystyle\\left\(\\frac\{1\+\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\}\{1\-\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\}\\right\)^\{\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\},\(155\)and thus yields
μ2\\displaystyle\\mu\_\{2\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}min\{μdrafter,\(1\+μtarget1−μtarget\)ϱ𝒑𝒒ε𝒑𝒒−1\(1\+μtarget1−μtarget\)ϱ𝒑𝒒ε𝒑𝒒\+1\}\.\\displaystyle\\min\\left\\\{\\mu\_\{\\mbox\{\\tiny\{drafter\}\}\},\\frac\{\\left\(\\frac\{1\+\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\}\{1\-\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\}\\right\)^\{\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}\-1\}\{\\left\(\\frac\{1\+\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\}\{1\-\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\}\\right\)^\{\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}\+1\}\\right\\\}\.\(156\)Note that this biases the computation of boosting coefficients but since we do not include further models aftert=2t=2, this does not change the analysis of the convergence, and the current analysis in the proof of Theorem[4\.2](https://arxiv.org/html/2609.30474#S4.Thmtheorem2)accommodates for the case where thelastboosting coefficient is eventually reduced\.
We now need another simple technical Lemma\.
###### Lemma H\.
Letffbe convex such thatf\(0\)=0f\(0\)=0\. Then for anyz∈domfz\\in\\mathrm\{dom\}fand anyt∈\[0,1\]t\\in\[0,1\],t⋅f\(z\)≥f\(t⋅z\)t\\cdot f\(z\)\\geq f\(t\\cdot z\)\.
###### Proof\.
Fort∈\(0,1\]t\\in\(0,1\], we just remark that sinceffis convex,f\(tz\+\(1−t\)⋅0\)≤t⋅f\(z\)\+\(1−t\)⋅f\(0\)f\(tz\+\(1\-t\)\\cdot 0\)\\leq t\\cdot f\(z\)\+\(1\-t\)\\cdot f\(0\), which, sincef\(0\)=0f\(0\)=0, simplifies into the Lemma’s statement\. The caset=0t=0is immediate\. ∎
Now, pick
f\(z\)\\displaystyle f\(z\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}log\(1\+z1−z\)\\displaystyle\\log\\left\(\\frac\{1\+z\}\{1\-z\}\\right\)restricted to\[0,1\]\[0,1\], in which it is convex\. Using Lemma[H](https://arxiv.org/html/2609.30474#S8.Thmtheorem8), we can expand the inequalityt⋅log\(1\+z1−z\)≥log\(1\+tz1−tz\)t\\cdot\\log\\left\(\\frac\{1\+z\}\{1\-z\}\\right\)\\geq\\log\\left\(\\frac\{1\+tz\}\{1\-tz\}\\right\)\(for anyz,t∈\[0,1\]z,t\\in\[0,1\]\) into the equivalent one in which we substitutet=\.ϱ𝒑𝒒ε𝒑𝒒t\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}:
\(\(1\+z1−z\)ϱ𝒑𝒒ε𝒑𝒒−1\(1\+z1−z\)ϱ𝒑𝒒ε𝒑𝒒\+1\)2\\displaystyle\\left\(\\frac\{\\left\(\\frac\{1\+z\}\{1\-z\}\\right\)^\{\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}\-1\}\{\\left\(\\frac\{1\+z\}\{1\-z\}\\right\)^\{\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}\+1\}\\right\)^\{2\}≥\\displaystyle\\geqϱ𝒑𝒒2ε𝒑𝒒2z2,∀z∈\[0,1\],\\displaystyle\\varrho\_\{\\bm\{p\}\\bm\{q\}\}^\{2\}\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}^\{2\}z^\{2\},\\forall z\\in\[0,1\],so getting back to \([156](https://arxiv.org/html/2609.30474#S8.E156)\), this translates into a guaranteed boosting advantage
A\(\{𝒉1=\.𝒉target,𝒉2=\.𝒉drafter\}\)\\displaystyle A\\left\(\\\{\\bm\{h\}\_\{1\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\bm\{h\}\_\{\\mbox\{\\tiny\{target\}\}\},\\bm\{h\}\_\{2\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\bm\{h\}\_\{\\mbox\{\\tiny\{drafter\}\}\}\\\}\\right\)≥\\displaystyle\\geqμtarget2\+ϱ𝒑𝒒2ε𝒑𝒒2μtarget2\\displaystyle\\mu^\{2\}\_\{\\mbox\{\\tiny\{target\}\}\}\+\\varrho\_\{\\bm\{p\}\\bm\{q\}\}^\{2\}\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}^\{2\}\\mu^\{2\}\_\{\\mbox\{\\tiny\{target\}\}\}=\(1\+ϱ𝒑𝒒2ε𝒑𝒒2\)⋅μtarget2\.\\displaystyle=\(1\+\\varrho\_\{\\bm\{p\}\\bm\{q\}\}^\{2\}\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}^\{2\}\)\\cdot\\mu^\{2\}\_\{\\mbox\{\\tiny\{target\}\}\}\.This guarantee holds for theε𝒑𝒒\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}corresponding to the current outputs of the drafter and target models\. We just need to replace it by the minimalϱ𝒑𝒒ε𝒑𝒒\>0\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\>0that would be satisfied for any outputs, and we get the bullet statements of Theorem[4\.7](https://arxiv.org/html/2609.30474#S4.Thmtheorem7)\.
We now proceed to showing the bound onDfTV\(𝒗α∥𝒒\)D\_\{f\_\{\\mathrm\{TV\}\}\}\(\\bm\{v\}\_\{\\alpha\}\\\|\\bm\{q\}\)in \([62](https://arxiv.org/html/2609.30474#S4.E62)\)\. Let us rewrite the TV distance:
DfTV\(𝒗α∥𝒒\)\\displaystyle D\_\{f\_\{\\mathrm\{TV\}\}\}\(\\bm\{v\}\_\{\\alpha\}\\\|\\bm\{q\}\)=\\displaystyle=12∑i∈n\|piαqi1−αZ−qi\|\\displaystyle\\frac\{1\}\{2\}\\sum\_\{i\\in n\}\\left\|\\frac\{p\_\{i\}^\{\\alpha\}q\_\{i\}^\{1\-\\alpha\}\}\{Z\}\-q\_\{i\}\\right\|=\\displaystyle=12∑i∈nqi⋅\|1Z⋅piαqiα−1\|\\displaystyle\\frac\{1\}\{2\}\\sum\_\{i\\in n\}q\_\{i\}\\cdot\\left\|\\frac\{1\}\{Z\}\\cdot\\frac\{p\_\{i\}^\{\\alpha\}\}\{q\_\{i\}^\{\\alpha\}\}\-1\\right\|=\\displaystyle=12⋅𝔼\[\|XαZ−1\|\],\\displaystyle\\frac\{1\}\{2\}\\cdot\\mathbb\{E\}\\left\[\\left\|\\frac\{X^\{\\alpha\}\}\{Z\}\-1\\right\|\\right\],whereXXis a random variable taking valuepi/qip\_\{i\}/q\_\{i\}with probabilityqiq\_\{i\}, hence satisfying𝔼\[X\]=1\\mathbb\{E\}\[X\]=1\. Since𝒑,𝒒\\bm\{p\},\\bm\{q\}obey \([ED](https://arxiv.org/html/2609.30474#S4.Ex18)\) andz↦zαz\\mapsto z^\{\\alpha\}is strictly monotonic forα∈\(0,1\]\\alpha\\in\(0,1\], the TV is maximal iff all values taken byXXare at the boundary of its range, i\.e\. eitherε𝒑𝒒≤1\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\leq 1or1/ε𝒑𝒒≥11/\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\geq 1\. But𝔼\[X\]=1\\mathbb\{E\}\[X\]=1and so the total mass atε𝒑𝒒\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\},με𝒑𝒒\\mu\_\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}and the total mass at1/ε𝒑𝒒1/\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\},μ1/ε𝒑𝒒\\mu\_\{1/\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}satisfy \(a\)μ1/ε𝒑𝒒\+με𝒑𝒒=1\\mu\_\{1/\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}\+\\mu\_\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}=1and \(b\)\(1/ε𝒑𝒒\)⋅μ1/ε𝒑𝒒\+ε𝒑𝒒⋅με𝒑𝒒=1\(1/\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\)\\cdot\\mu\_\{1/\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}\+\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\cdot\\mu\_\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}=1, a system whose solution isμ1/ε𝒑𝒒=ε𝒑𝒒/\(1\+ε𝒑𝒒\),με𝒑𝒒=1/\(1\+ε𝒑𝒒\)\\mu\_\{1/\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}=\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}/\(1\+\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\),\\mu\_\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}=1/\(1\+\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\), giving the normalization coefficient
Z\\displaystyle Z=\\displaystyle=με𝒑𝒒⋅ε𝒑𝒒α\+μ1/ε𝒑𝒒⋅\(1ε𝒑𝒒\)α=ε𝒑𝒒α\+ε𝒑𝒒1−α1\+ε𝒑𝒒\\displaystyle\\mu\_\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}\\cdot\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}^\{\\alpha\}\+\\mu\_\{1/\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}\\cdot\\left\(\\frac\{1\}\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}\\right\)^\{\\alpha\}=\\frac\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}^\{\\alpha\}\+\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}^\{1\-\\alpha\}\}\{1\+\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}and yielding the upperbound
DfTV\(𝒗α∥𝒒\)\\displaystyle D\_\{f\_\{\\mathrm\{TV\}\}\}\(\\bm\{v\}\_\{\\alpha\}\\\|\\bm\{q\}\)≤\\displaystyle\\leq12⋅\(με𝒑𝒒⋅\(1−ε𝒑𝒒αZ\)\+μ1/ε𝒑𝒒⋅\(ε𝒑𝒒−αZ−1\)\)\\displaystyle\\frac\{1\}\{2\}\\cdot\\left\(\\mu\_\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}\\cdot\\left\(1\-\\frac\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}^\{\\alpha\}\}\{Z\}\\right\)\+\\mu\_\{1/\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}\\cdot\\left\(\\frac\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}^\{\-\\alpha\}\}\{Z\}\-1\\right\)\\right\)\(157\)=1−ε𝒑𝒒2⋅\(1\+ε𝒑𝒒\)\+ε𝒑𝒒1−α−ε𝒑𝒒α2⋅\(ε𝒑𝒒1−α\+ε𝒑𝒒α\)\.\\displaystyle=\\frac\{1\-\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}\{2\\cdot\(1\+\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\)\}\+\\frac\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}^\{1\-\\alpha\}\-\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}^\{\\alpha\}\}\{2\\cdot\(\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}^\{1\-\\alpha\}\+\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}^\{\\alpha\}\)\}\.Now, remark that if we pick
μ2\\displaystyle\\mu\_\{2\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\(1\+μtarget\)ε𝒑𝒒ϱ𝒑𝒒−\(1−μtarget\)ε𝒑𝒒ϱ𝒑𝒒\(1\+μtarget\)ε𝒑𝒒ϱ𝒑𝒒\+\(1−μtarget\)ε𝒑𝒒ϱ𝒑𝒒,\\displaystyle\\frac\{\\left\(1\+\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\-\\left\(1\-\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\}\{\\left\(1\+\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\+\\left\(1\-\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\},\(158\)thenα\\alphasimplifies to
α\\displaystyle\\alpha=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}c2c2\+ctarget\\displaystyle\\frac\{c\_\{2\}\}\{c\_\{2\}\+c\_\{\\mbox\{\\tiny\{target\}\}\}\}=\\displaystyle=14⋅log\(1\+\(1\+μtarget\)ε𝒑𝒒ϱ𝒑𝒒−\(1−μtarget\)ε𝒑𝒒ϱ𝒑𝒒\(1\+μtarget\)ε𝒑𝒒ϱ𝒑𝒒\+\(1−μtarget\)ε𝒑𝒒ϱ𝒑𝒒1−\(1\+μtarget\)ε𝒑𝒒ϱ𝒑𝒒−\(1−μtarget\)ε𝒑𝒒ϱ𝒑𝒒\(1\+μtarget\)ε𝒑𝒒ϱ𝒑𝒒\+\(1−μtarget\)ε𝒑𝒒ϱ𝒑𝒒\)14⋅log\(1\+\(1\+μtarget\)ε𝒑𝒒ϱ𝒑𝒒−\(1−μtarget\)ε𝒑𝒒ϱ𝒑𝒒\(1\+μtarget\)ε𝒑𝒒ϱ𝒑𝒒\+\(1−μtarget\)ε𝒑𝒒ϱ𝒑𝒒1−\(1\+μtarget\)ε𝒑𝒒ϱ𝒑𝒒−\(1−μtarget\)ε𝒑𝒒ϱ𝒑𝒒\(1\+μtarget\)ε𝒑𝒒ϱ𝒑𝒒\+\(1−μtarget\)ε𝒑𝒒ϱ𝒑𝒒\)\+14⋅log\(1\+μtarget1−μtarget\)\\displaystyle\\frac\{\\frac\{1\}\{4\}\\cdot\\log\\left\(\\frac\{1\+\\frac\{\\left\(1\+\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\-\\left\(1\-\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\}\{\\left\(1\+\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\+\\left\(1\-\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\}\}\{1\-\\frac\{\\left\(1\+\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\-\\left\(1\-\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\}\{\\left\(1\+\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\+\\left\(1\-\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\}\}\\right\)\}\{\\frac\{1\}\{4\}\\cdot\\log\\left\(\\frac\{1\+\\frac\{\\left\(1\+\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\-\\left\(1\-\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\}\{\\left\(1\+\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\+\\left\(1\-\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\}\}\{1\-\\frac\{\\left\(1\+\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\-\\left\(1\-\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\}\{\\left\(1\+\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\+\\left\(1\-\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\\right\)^\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\}\}\\right\)\+\\frac\{1\}\{4\}\\cdot\\log\\left\(\\frac\{1\+\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\}\{1\-\\mu\_\{\\mbox\{\\tiny\{target\}\}\}\}\\right\)\}=\\displaystyle=ε𝒑𝒒ϱ𝒑𝒒1\+ε𝒑𝒒ϱ𝒑𝒒,\\displaystyle\\frac\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\{1\+\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\},so that \([157](https://arxiv.org/html/2609.30474#S8.E157)\) becomes an upperbound depending solely onε𝒑𝒒\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}:
DfTV\(𝒗α∥𝒒\)\\displaystyle D\_\{f\_\{\\mathrm\{TV\}\}\}\(\\bm\{v\}\_\{\\alpha\}\\\|\\bm\{q\}\)≤\\displaystyle\\leq1−ε𝒑𝒒2⋅\(1\+ε𝒑𝒒\)\+exp\(11\+ε𝒑𝒒ϱ𝒑𝒒⋅logε𝒑𝒒\)−exp\(ε𝒑𝒒ϱ𝒑𝒒1\+ε𝒑𝒒ϱ𝒑𝒒⋅logε𝒑𝒒\)2⋅\(exp\(11\+ε𝒑𝒒ϱ𝒑𝒒⋅logε𝒑𝒒\)\+exp\(ε𝒑𝒒ϱ𝒑𝒒1\+ε𝒑𝒒ϱ𝒑𝒒⋅logε𝒑𝒒\)\)\.\\displaystyle\\frac\{1\-\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\}\{2\\cdot\(1\+\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\)\}\+\\frac\{\\exp\\left\(\\frac\{1\}\{1\+\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\\cdot\\log\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\right\)\-\\exp\\left\(\\frac\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\{1\+\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\\cdot\\log\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\right\)\}\{2\\cdot\\left\(\\exp\\left\(\\frac\{1\}\{1\+\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\\cdot\\log\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\right\)\+\\exp\\left\(\\frac\{\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\{1\+\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\}\\cdot\\log\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\right\)\\right\)\}\.\(159\)Some tedious calculation allow to show that the RHS is≤ϱ𝒑𝒒ε𝒑𝒒\\leq\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}for anyε𝒑𝒒∈\[0,1\],ϱ𝒑𝒒≥0\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}\\in\[0,1\],\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\\geq 0, which givesDfTV\(𝒗α∥𝒒\)≤ϱ𝒑𝒒ε𝒑𝒒D\_\{f\_\{\\mathrm\{TV\}\}\}\(\\bm\{v\}\_\{\\alpha\}\\\|\\bm\{q\}\)\\leq\\varrho\_\{\\bm\{p\}\\bm\{q\}\}\\varepsilon\_\{\\bm\{p\}\\bm\{q\}\}, as claimed\.
### VIII\.10Proof of Theorem[4\.8](https://arxiv.org/html/2609.30474#S4.Thmtheorem8)
We recall some notations: for anyi∈\[m\]i\\in\[m\],𝝅i\\bm\{\\pi\}\_\{i\}denotes the mentored distribution for example\#i\\\#iin𝒮\\mathcal\{S\}following the mentored decoding setting:
πij\\displaystyle\\pi\_\{ij\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\{\(1−bi\)⋅qijifpij<\(1−bi\)qij\(Case \(i\)\)pijifpij∈\[1−bi,1\+ai\]⋅qij\(Case \(ii\)\)\(1\+ai\)⋅qijifpij\>\(1\+ai\)qij\(Case \(iii\)\),j∈\[n\]\\displaystyle\\left\\\{\\begin\{array\}\[\]\{rcll\}\(1\-b\_\{i\}\)\\cdot q\_\{ij\}&\\mbox\{if\}&p\_\{ij\}<\(1\-b\_\{i\}\)q\_\{ij\}&\\mbox\{\(Case \(i\)\)\}\\\\ p\_\{ij\}&\\mbox\{if\}&p\_\{ij\}\\in\[1\-b\_\{i\},1\+a\_\{i\}\]\\cdot q\_\{ij\}&\\mbox\{\(Case \(ii\)\)\}\\\\ \(1\+a\_\{i\}\)\\cdot q\_\{ij\}&\\mbox\{if\}&p\_\{ij\}\>\(1\+a\_\{i\}\)q\_\{ij\}&\\mbox\{\(Case \(iii\)\)\}\\end\{array\}\\right\.,j\\in\[n\]and
π~ij\\displaystyle\\tilde\{\\pi\}\_\{ij\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}pijαqij1−αZ,j∈\[n\]\\displaystyle\\frac\{p\_\{ij\}^\{\\alpha\}q\_\{ij\}^\{1\-\\alpha\}\}\{Z\},j\\in\[n\]denote the boosting distribution coordinates following the boosting combination setup in \([48](https://arxiv.org/html/2609.30474#S4.E48)\)\. Notice that under Assumption[3\.3](https://arxiv.org/html/2609.30474#S3.Thmtheorem3),𝝅~i\>𝟎,∀i∈\[m\]\\tilde\{\\bm\{\\pi\}\}\_\{i\}\>\\bm\{0\},\\forall i\\in\[m\]\. Let us say that example\#i\\\#iisρ\\rho\-good iff the event in \([49](https://arxiv.org/html/2609.30474#S4.E49)\) \(Theorem[4\.2](https://arxiv.org/html/2609.30474#S4.Thmtheorem2)\) is false: in such a case, boosting guarantees
\{π~ij:j∈𝒴i\}¯G\{π~ij:j∈𝒴¯i\}G¯\\displaystyle\\frac\{\\overline\{\\\{\\tilde\{\\pi\}\_\{ij\}:j\\in\\mathcal\{Y\}\_\{i\}\\\}\}^\{G\}\}\{\\overline\{\\\{\\tilde\{\\pi\}\_\{ij\}:j\\in\\overline\{\\mathcal\{Y\}\}\_\{i\}\\\}^\{G\}\}\}≥\\displaystyle\\geqρ\.\\displaystyle\\rho\.\(164\)We recall that𝒜¯G\\overline\{\\mathcal\{A\}\}^\{G\}denotes the geometric average of the elements of set𝒜\\mathcal\{A\}\. Denore for shortεi=\.ε𝒑i𝒒i\\varepsilon\_\{i\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\varepsilon\_\{\\bm\{p\}\_\{i\}\\bm\{q\}\_\{i\}\}in \([ED](https://arxiv.org/html/2609.30474#S4.Ex18)\)\. We now combine \([ED](https://arxiv.org/html/2609.30474#S4.Ex18)\) with \([VIII\.10](https://arxiv.org/html/2609.30474#S8.EGx182)\) to find intervalsπ~ij∈𝕀\(πij\)\\tilde\{\\pi\}\_\{ij\}\\in\\mathbb\{I\}\(\\pi\_\{ij\}\)to transform \([164](https://arxiv.org/html/2609.30474#S8.E164)\) in an inequality involving only the mentored distribution𝝅i\\bm\{\\pi\}\_\{i\}\. To simplify notations, we drop indexiito focus on coordinatejjonly\.
Case \(i\)Let us start with Case \(i\) \([VIII\.10](https://arxiv.org/html/2609.30474#S8.EGx182)\)\. Here,pj<\(1−bi\)qjp\_\{j\}<\(1\-b\_\{i\}\)q\_\{j\}, but \([ED](https://arxiv.org/html/2609.30474#S4.Ex18)\) guaranteespj≥εiqjp\_\{j\}\\geq\\varepsilon\_\{i\}q\_\{j\}, so to get Case \(i\), we must have
εi\\displaystyle\\varepsilon\_\{i\}<\\displaystyle<1−bi\.\\displaystyle 1\-b\_\{i\}\.\(165\)Provided this holds, we also knowπj=\(1−bi\)⋅qj\\pi\_\{j\}=\(1\-b\_\{i\}\)\\cdot q\_\{j\}, so the boosting coordinateπ~j\\tilde\{\\pi\}\_\{j\}satisfiesπ~j<\(1−bi\)αqj/Z=πj/\(Z\(1−bi\)1−α\)\\tilde\{\\pi\}\_\{j\}<\(1\-b\_\{i\}\)^\{\\alpha\}q\_\{j\}/Z=\\pi\_\{j\}/\(Z\(1\-b\_\{i\}\)^\{1\-\\alpha\}\)\. On the other hand, \([ED](https://arxiv.org/html/2609.30474#S4.Ex18)\) guaranteespj≥εiqjp\_\{j\}\\geq\\varepsilon\_\{i\}q\_\{j\}, which yields a lowerboundπ~j≥\(εiα/Z\)⋅qi=\(εiα/\(Z\(1−bi\)\)\)⋅πj\\tilde\{\\pi\}\_\{j\}\\geq\(\\varepsilon\_\{i\}^\{\\alpha\}/Z\)\\cdot q\_\{i\}=\(\\varepsilon\_\{i\}^\{\\alpha\}/\(Z\(1\-b\_\{i\}\)\)\)\\cdot\\pi\_\{j\}, and thus we overall obtain
π~j∈πj\(1−bi\)Z⋅\[εiα,\(1−bi\)α\)ifpj∈qj⋅\[εi,1−bi\)\.\\displaystyle\\tilde\{\\pi\}\_\{j\}\\in\\frac\{\\pi\_\{j\}\}\{\(1\-b\_\{i\}\)Z\}\\cdot\\left\[\\varepsilon\_\{i\}^\{\\alpha\},\(1\-b\_\{i\}\)^\{\\alpha\}\\right\)\\quad\\mbox\{ if \}p\_\{j\}\\in q\_\{j\}\\cdot\[\\varepsilon\_\{i\},1\-b\_\{i\}\)\.\(166\)
Case \(iii\)Now, we havepj\>\(1\+ai\)qjp\_\{j\}\>\(1\+a\_\{i\}\)q\_\{j\}but \([ED](https://arxiv.org/html/2609.30474#S4.Ex18)\) guaranteespj≤qj/εip\_\{j\}\\leq q\_\{j\}/\\varepsilon\_\{i\}, so to get Case \(iii\), we must have
εi\\displaystyle\\varepsilon\_\{i\}<\\displaystyle<11\+ai\.\\displaystyle\\frac\{1\}\{1\+a\_\{i\}\}\.\(167\)Provided this holds, we also knowπj=\(1\+ai\)⋅qj\\pi\_\{j\}=\(1\+a\_\{i\}\)\\cdot q\_\{j\}, so the boosting coordinateπ~j\\tilde\{\\pi\}\_\{j\}satisfiesπ~j\>\(1\+ai\)αqj/Z=πj/\(Z\(1\+ai\)1−α\)\\tilde\{\\pi\}\_\{j\}\>\(1\+a\_\{i\}\)^\{\\alpha\}q\_\{j\}/Z=\\pi\_\{j\}/\(Z\(1\+a\_\{i\}\)^\{1\-\\alpha\}\)\. On the other hand, \([ED](https://arxiv.org/html/2609.30474#S4.Ex18)\) guaranteespj≤qj/εip\_\{j\}\\leq q\_\{j\}/\\varepsilon\_\{i\}, which yields an upperboundπ~j≤\(1/\(Zεiα\)\)⋅qi=\(1/\(Z\(1\+ai\)εiα\)\)⋅πj\\tilde\{\\pi\}\_\{j\}\\leq\(1/\(Z\\varepsilon\_\{i\}^\{\\alpha\}\)\)\\cdot q\_\{i\}=\(1/\(Z\(1\+a\_\{i\}\)\\varepsilon\_\{i\}^\{\\alpha\}\)\)\\cdot\\pi\_\{j\}, and thus we overall obtain
π~j∈πj\(1\+ai\)Z⋅\(\(1\+ai\)α,1εiα\]ifpj∈qj⋅\(1\+ai,1/εi\]\.\\displaystyle\\tilde\{\\pi\}\_\{j\}\\in\\frac\{\\pi\_\{j\}\}\{\(1\+a\_\{i\}\)Z\}\\cdot\\left\(\(1\+a\_\{i\}\)^\{\\alpha\},\\frac\{1\}\{\\varepsilon\_\{i\}^\{\\alpha\}\}\\right\]\\quad\\mbox\{ if \}p\_\{j\}\\in q\_\{j\}\\cdot\(1\+a\_\{i\},1/\\varepsilon\_\{i\}\]\.\(168\)
Case \(ii\)Now, we have simultaneously\(1−bi\)qj≤pj≤\(1\+ai\)qj\(1\-b\_\{i\}\)q\_\{j\}\\leq p\_\{j\}\\leq\(1\+a\_\{i\}\)q\_\{j\}, but \([ED](https://arxiv.org/html/2609.30474#S4.Ex18)\) guaranteesεiqj≤pj≤qj/εj\\varepsilon\_\{i\}q\_\{j\}\\leq p\_\{j\}\\leq q\_\{j\}/\\varepsilon\_\{j\}so to get Case \(ii\), the intervals must have a non\-empty intersection and we must haveεi≤1/\(1−bi\)\\varepsilon\_\{i\}\\leq 1/\(1\-b\_\{i\}\)orεi≤1\+ai\\varepsilon\_\{i\}\\leq 1\+a\_\{i\}– which always holds sincebi∈\[0,1\],ai≥0b\_\{i\}\\in\[0,1\],a\_\{i\}\\geq 0–\. Sinceπj=pj\\pi\_\{j\}=p\_\{j\}, we now have the direct expression
π~j\\displaystyle\\tilde\{\\pi\}\_\{j\}=\\displaystyle=pjαqj1−αZ\\displaystyle\\frac\{p\_\{j\}^\{\\alpha\}q\_\{j\}^\{1\-\\alpha\}\}\{Z\}=\\displaystyle=πj⋅1Z⋅\(qjpj\)α,\\displaystyle\\pi\_\{j\}\\cdot\\frac\{1\}\{Z\}\\cdot\\left\(\\frac\{q\_\{j\}\}\{p\_\{j\}\}\\right\)^\{\\alpha\},and we can check that with the inequalities above, we get
π~j∈πjZ⋅\[1\(1\+ai\)α,1\(1−bi\)α\]ifpj∈qj⋅\[1−bi,1\+ai\]\.\\displaystyle\\tilde\{\\pi\}\_\{j\}\\in\\frac\{\\pi\_\{j\}\}\{Z\}\\cdot\\left\[\\frac\{1\}\{\(1\+a\_\{i\}\)^\{\\alpha\}\},\\frac\{1\}\{\(1\-b\_\{i\}\)^\{\\alpha\}\}\\right\]\\quad\\mbox\{ if \}p\_\{j\}\\in q\_\{j\}\\cdot\[1\-b\_\{i\},1\+a\_\{i\}\]\.\(169\)Using \([166](https://arxiv.org/html/2609.30474#S8.E166)\), \([168](https://arxiv.org/html/2609.30474#S8.E168)\), \([169](https://arxiv.org/html/2609.30474#S8.E169)\), we obtain the upperbound,
\{π~ij:j∈𝒴i\}¯G\{π~ij:j∈𝒴¯i\}¯G\\displaystyle\\frac\{\\overline\{\\\{\\tilde\{\\pi\}\_\{ij\}:j\\in\\mathcal\{Y\}\_\{i\}\\\}\}^\{G\}\}\{\\overline\{\\\{\\tilde\{\\pi\}\_\{ij\}:j\\in\\overline\{\\mathcal\{Y\}\}\_\{i\}\\\}\}^\{G\}\}≤\\displaystyle\\leq\{πij:j∈𝒴i\}¯G\{πij:j∈𝒴¯i\}¯G⋅max\{1\(1−bi\)1−α,1\(1−bi\)α,1\(1\+ai\)⋅εiα\}min\{1\(1\+ai\)1−α,1\(1\+ai\)α,εiα1−bi\}⏟=\.R−1\.\\displaystyle\\frac\{\\overline\{\\\{\\pi\_\{ij\}:j\\in\\mathcal\{Y\}\_\{i\}\\\}\}^\{G\}\}\{\\overline\{\\\{\\pi\_\{ij\}:j\\in\\overline\{\\mathcal\{Y\}\}\_\{i\}\\\}\}^\{G\}\}\\cdot\\underbrace\{\\frac\{\\max\\left\\\{\\frac\{1\}\{\(1\-b\_\{i\}\)^\{1\-\\alpha\}\},\\frac\{1\}\{\(1\-b\_\{i\}\)^\{\\alpha\}\},\\frac\{1\}\{\(1\+a\_\{i\}\)\\cdot\\varepsilon\_\{i\}^\{\\alpha\}\}\\right\\\}\}\{\\min\\left\\\{\\frac\{1\}\{\(1\+a\_\{i\}\)^\{1\-\\alpha\}\},\\frac\{1\}\{\(1\+a\_\{i\}\)^\{\\alpha\}\},\\frac\{\\varepsilon\_\{i\}^\{\\alpha\}\}\{1\-b\_\{i\}\}\\right\\\}\}\}\_\{\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}R^\{\-1\}\}\.\(170\)We now need to find a lowerbound forRR\. We distinguish two cases:
Case \(A\):εiα<\(1−bi\)/\(1\+ai\)\\varepsilon^\{\\alpha\}\_\{i\}<\(1\-b\_\{i\}\)/\(1\+a\_\{i\}\)\. Thus,εiα/\(1−bi\)<1/\(1\+ai\)≤1\\varepsilon^\{\\alpha\}\_\{i\}/\(1\-b\_\{i\}\)<1/\(1\+a\_\{i\}\)\\leq 1sinceai≤0a\_\{i\}\\leq 0, and since0≤α≤10\\leq\\alpha\\leq 1, the numerator ofRRsatisfies
min\{1\(1\+ai\)1−α,1\(1\+ai\)α,εiα1−bi\}\\displaystyle\\min\\left\\\{\\frac\{1\}\{\(1\+a\_\{i\}\)^\{1\-\\alpha\}\},\\frac\{1\}\{\(1\+a\_\{i\}\)^\{\\alpha\}\},\\frac\{\\varepsilon\_\{i\}^\{\\alpha\}\}\{1\-b\_\{i\}\}\\right\\\}≥\\displaystyle\\geqεiα1−bi;\\displaystyle\\frac\{\\varepsilon^\{\\alpha\}\_\{i\}\}\{1\-b\_\{i\}\};We also get1≤1/\(1−bi\)≤1/\(\(1\+ai\)εiα\)1\\leq 1/\(1\-b\_\{i\}\)\\leq 1/\(\(1\+a\_\{i\}\)\\varepsilon^\{\\alpha\}\_\{i\}\)sincebi∈\[0,1\]b\_\{i\}\\in\[0,1\], and since0≤α≤10\\leq\\alpha\\leq 1, the denominator ofRRsatisfies
max\{1\(1−bi\)1−α,1\(1−bi\)α,1\(1\+ai\)⋅εiα\}\\displaystyle\\max\\left\\\{\\frac\{1\}\{\(1\-b\_\{i\}\)^\{1\-\\alpha\}\},\\frac\{1\}\{\(1\-b\_\{i\}\)^\{\\alpha\}\},\\frac\{1\}\{\(1\+a\_\{i\}\)\\cdot\\varepsilon\_\{i\}^\{\\alpha\}\}\\right\\\}≤\\displaystyle\\leq1\(1\+ai\)εiα,\\displaystyle\\frac\{1\}\{\(1\+a\_\{i\}\)\\varepsilon^\{\\alpha\}\_\{i\}\},and finally
R\\displaystyle R≥\\displaystyle\\geq1\+ai1−bi⋅εi2α\.\\displaystyle\\frac\{1\+a\_\{i\}\}\{1\-b\_\{i\}\}\\cdot\\varepsilon^\{2\\alpha\}\_\{i\}\.\(171\)Case \(B\):εiα≥\(1−bi\)/\(1\+ai\)\\varepsilon^\{\\alpha\}\_\{i\}\\geq\(1\-b\_\{i\}\)/\(1\+a\_\{i\}\)\. Thusεiα/\(1−bi\)≥1/\(1\+ai\)\\varepsilon^\{\\alpha\}\_\{i\}/\(1\-b\_\{i\}\)\\geq 1/\(1\+a\_\{i\}\)and the numerator ofRRsatisfies
min\{1\(1\+ai\)1−α,1\(1\+ai\)α,εiα1−bi\}\\displaystyle\\min\\left\\\{\\frac\{1\}\{\(1\+a\_\{i\}\)^\{1\-\\alpha\}\},\\frac\{1\}\{\(1\+a\_\{i\}\)^\{\\alpha\}\},\\frac\{\\varepsilon\_\{i\}^\{\\alpha\}\}\{1\-b\_\{i\}\}\\right\\\}≥\\displaystyle\\geqmin\{1\(1\+ai\)1−α,1\(1\+ai\)α,11\+ai\}\\displaystyle\\min\\left\\\{\\frac\{1\}\{\(1\+a\_\{i\}\)^\{1\-\\alpha\}\},\\frac\{1\}\{\(1\+a\_\{i\}\)^\{\\alpha\}\},\\frac\{1\}\{1\+a\_\{i\}\}\\right\\\}=11\+ai\\displaystyle=\\frac\{1\}\{1\+a\_\{i\}\}since0≤α≤10\\leq\\alpha\\leq 1\. Similarly for the denominator, since1/\(1−bi\)≥1/\(\(1\+ai\)εiα\)1/\(1\-b\_\{i\}\)\\geq 1/\(\(1\+a\_\{i\}\)\\varepsilon^\{\\alpha\}\_\{i\}\), we observe
max\{1\(1−bi\)1−α,1\(1−bi\)α,1\(1\+ai\)⋅εiα\}\\displaystyle\\max\\left\\\{\\frac\{1\}\{\(1\-b\_\{i\}\)^\{1\-\\alpha\}\},\\frac\{1\}\{\(1\-b\_\{i\}\)^\{\\alpha\}\},\\frac\{1\}\{\(1\+a\_\{i\}\)\\cdot\\varepsilon\_\{i\}^\{\\alpha\}\}\\right\\\}≤\\displaystyle\\leqmax\{1\(1−bi\)1−α,1\(1−bi\)α,11−bi\}\\displaystyle\\max\\left\\\{\\frac\{1\}\{\(1\-b\_\{i\}\)^\{1\-\\alpha\}\},\\frac\{1\}\{\(1\-b\_\{i\}\)^\{\\alpha\}\},\\frac\{1\}\{1\-b\_\{i\}\}\\right\\\}=11−bi,\\displaystyle=\\frac\{1\}\{1\-b\_\{i\}\},and finally
R\\displaystyle R≥\\displaystyle\\geq1−bi1\+ai\.\\displaystyle\\frac\{1\-b\_\{i\}\}\{1\+a\_\{i\}\}\.\(172\)and we finally check that \([171](https://arxiv.org/html/2609.30474#S8.E171)\) and \([172](https://arxiv.org/html/2609.30474#S8.E172)\) can be folded into one:
R\\displaystyle R≥\\displaystyle\\geq1\+ai1−bi⋅min\{εiα,1−bi1\+ai\}2\\displaystyle\\frac\{1\+a\_\{i\}\}\{1\-b\_\{i\}\}\\cdot\\min\\left\\\{\\varepsilon^\{\\alpha\}\_\{i\},\\frac\{1\-b\_\{i\}\}\{1\+a\_\{i\}\}\\right\\\}^\{2\}=min\{εiα⋅1\+ai1−bi,1−bi1\+ai\}2\\displaystyle=\\min\\left\\\{\\varepsilon^\{\\alpha\}\_\{i\}\\cdot\\sqrt\{\\frac\{1\+a\_\{i\}\}\{1\-b\_\{i\}\}\},\\sqrt\{\\frac\{1\-b\_\{i\}\}\{1\+a\_\{i\}\}\}\\right\\\}^\{2\}We can then simplify \([170](https://arxiv.org/html/2609.30474#S8.E170)\) into a more readable uperbound:
\{π~ij:j∈𝒴i\}¯G\{π~ij:j∈𝒴¯i\}¯G\\displaystyle\\frac\{\\overline\{\\\{\\tilde\{\\pi\}\_\{ij\}:j\\in\\mathcal\{Y\}\_\{i\}\\\}\}^\{G\}\}\{\\overline\{\\\{\\tilde\{\\pi\}\_\{ij\}:j\\in\\overline\{\\mathcal\{Y\}\}\_\{i\}\\\}\}^\{G\}\}≤\\displaystyle\\leq\{πij:j∈𝒴i\}¯G\{πij:j∈𝒴¯i\}¯G⋅1min\{εiα⋅1\+ai1−bi,1−bi1\+ai\}2,\\displaystyle\\frac\{\\overline\{\\\{\\pi\_\{ij\}:j\\in\\mathcal\{Y\}\_\{i\}\\\}\}^\{G\}\}\{\\overline\{\\\{\\pi\_\{ij\}:j\\in\\overline\{\\mathcal\{Y\}\}\_\{i\}\\\}\}^\{G\}\}\\cdot\\frac\{1\}\{\\min\\left\\\{\\varepsilon^\{\\alpha\}\_\{i\}\\cdot\\sqrt\{\\frac\{1\+a\_\{i\}\}\{1\-b\_\{i\}\}\},\\sqrt\{\\frac\{1\-b\_\{i\}\}\{1\+a\_\{i\}\}\}\\right\\\}^\{2\}\},from which it comes that, for anyi∈\[m\]i\\in\[m\]andρ≥0\\rho\\geq 0, if
\{πij:j∈𝒴i\}¯G\\displaystyle\\overline\{\\\{\\pi\_\{ij\}:j\\in\\mathcal\{Y\}\_\{i\}\\\}\}^\{G\}≤\\displaystyle\\leqρ⋅min\{εiα⋅1\+ai1−bi,1−bi1\+ai\}2⋅\{πij:j∈𝒴¯i\}¯G,\\displaystyle\\rho\\cdot\\min\\left\\\{\\varepsilon^\{\\alpha\}\_\{i\}\\cdot\\sqrt\{\\frac\{1\+a\_\{i\}\}\{1\-b\_\{i\}\}\},\\sqrt\{\\frac\{1\-b\_\{i\}\}\{1\+a\_\{i\}\}\}\\right\\\}^\{2\}\\cdot\\overline\{\\\{\\pi\_\{ij\}:j\\in\\overline\{\\mathcal\{Y\}\}\_\{i\}\\\}\}^\{G\},then
\{π~ij:j∈𝒴i\}¯G\\displaystyle\\overline\{\\\{\\tilde\{\\pi\}\_\{ij\}:j\\in\\mathcal\{Y\}\_\{i\}\\\}\}^\{G\}≤\\displaystyle\\leqρ⋅\{π~ij:j∈𝒴¯i\}¯G,\\displaystyle\\rho\\cdot\\overline\{\\\{\\tilde\{\\pi\}\_\{ij\}:j\\in\\overline\{\\mathcal\{Y\}\}\_\{i\}\\\}\}^\{G\},and yields to the statement of Theorem[4\.8](https://arxiv.org/html/2609.30474#S4.Thmtheorem8)via Theorem[4\.2](https://arxiv.org/html/2609.30474#S4.Thmtheorem2)\.
### VIII\.11Proof of Theorem[5\.1](https://arxiv.org/html/2609.30474#S5.Thmtheorem1)
Without loss of generality, we assume all ratiosp\./q\.p\_\{\.\}/q\_\{\.\}are distinct \(otherwise, we group thepps andqqs\) and indices are ordered in increasing ratio\. Let us index in\{1,2,…\}\\\{1,2,\.\.\.\\\}the breakpoints in the order they are put incby algorithmUpdateCBreakpoints\.
Clearly, the list of breakpoints built byUpdateCBreakpointsis built in strictly increasing order ofaaandbb\. Clearly also, the first breakpoint,\(a,b\)=\(0,0\)\(a,b\)=\(0,0\)is in𝒞\(𝒑,𝒒\)\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\)becauseaq\(𝔸\)=0=0\+1−1=bq\(𝔹\)\+q\(𝕀\)−p\(𝕀\)aq\(\\mathbb\{A\}\)=0=0\+1\-1=bq\(\\mathbb\{B\}\)\+q\(\\mathbb\{I\}\)\-p\(\\mathbb\{I\}\)so all \([36](https://arxiv.org/html/2609.30474#S3.E36)\), \([37](https://arxiv.org/html/2609.30474#S3.E37)\) and \([38](https://arxiv.org/html/2609.30474#S3.E38)\) are satisfied\. Now take any such breakpoint\(a,b\)∈𝒞\(𝒑,𝒒\)\(a,b\)\\in\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\)\. Let
i\\displaystyle i=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}min𝔸,\\displaystyle\\min\\mathbb\{A\},j\\displaystyle j=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}max𝔹\.\\displaystyle\\max\\mathbb\{B\}\.
We have three cases:
Case 1: suppose that the current breakpoint satisfies
\(1−b−pjqj\)⋅q\(𝔹\)\\displaystyle\\left\(1\-b\-\\frac\{p\_\{j\}\}\{q\_\{j\}\}\\right\)\\cdot q\(\\mathbb\{B\}\)<\\displaystyle<\(piqi−1−a\)⋅q\(𝔸\)\.\\displaystyle\\left\(\\frac\{p\_\{i\}\}\{q\_\{i\}\}\-1\-a\\right\)\\cdot q\(\\mathbb\{A\}\)\.\(173\)Letb′=\.1−pj/qjb^\{\\prime\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}1\-p\_\{j\}/q\_\{j\}andΔb=\.b′−b\>0\\Delta\_\{b\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}b^\{\\prime\}\-b\>0\(by definition of𝔹\\mathbb\{B\}\([35](https://arxiv.org/html/2609.30474#S3.E35)\)\)\. Reformulate \([38](https://arxiv.org/html/2609.30474#S3.E38)\) as
aq\(𝔸\)\\displaystyle aq\(\\mathbb\{A\}\)=\\displaystyle=bq\(𝔹\)\+q\(𝕀\)−p\(𝕀\),\\displaystyle bq\(\\mathbb\{B\}\)\+q\(\\mathbb\{I\}\)\-p\(\\mathbb\{I\}\),\(174\)and rewrite the RHS:
bq\(𝔹\)\+q\(𝕀\)−p\(𝕀\)\\displaystyle bq\(\\mathbb\{B\}\)\+q\(\\mathbb\{I\}\)\-p\(\\mathbb\{I\}\)\(175\)=\\displaystyle=\(1−Δb−pjqj\)\(q\(𝔹\\\{j\}\)\+qj\)\+q\(𝕀\)−p\(𝕀\)\\displaystyle\\left\(1\-\\Delta\_\{b\}\-\\frac\{p\_\{j\}\}\{q\_\{j\}\}\\right\)\(q\(\\mathbb\{B\}\\backslash\\\{j\\\}\)\+q\_\{j\}\)\+q\(\\mathbb\{I\}\)\-p\(\\mathbb\{I\}\)=\\displaystyle=\(1−pjqj\)q\(𝔹\\\{j\}\)\+q\(𝕀\)−p\(𝕀\)\+\(qj−pj\)−Δb⋅\(q\(𝔹\\\{j\}\)\+qj\)\\displaystyle\\left\(1\-\\frac\{p\_\{j\}\}\{q\_\{j\}\}\\right\)q\(\\mathbb\{B\}\\backslash\\\{j\\\}\)\+q\(\\mathbb\{I\}\)\-p\(\\mathbb\{I\}\)\+\(q\_\{j\}\-p\_\{j\}\)\-\\Delta\_\{b\}\\cdot\(q\(\\mathbb\{B\}\\backslash\\\{j\\\}\)\+q\_\{j\}\)=\\displaystyle=\(1−pjqj\)q\(𝔹\\\{j\}\)\+q\(𝕀∪\{j\}\)−p\(𝕀∪\{j\}\)⏟=\.Ra−Δb⋅q\(𝔹\)\.\\displaystyle\\underbrace\{\\left\(1\-\\frac\{p\_\{j\}\}\{q\_\{j\}\}\\right\)q\(\\mathbb\{B\}\\backslash\\\{j\\\}\)\+q\(\\mathbb\{I\}\\cup\\\{j\\\}\)\-p\(\\mathbb\{I\}\\cup\\\{j\\\}\)\}\_\{\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}R\_\{a\}\}\-\\Delta\_\{b\}\\cdot q\(\\mathbb\{B\}\)\.Note thatRaR\_\{a\}is the RHS of \([174](https://arxiv.org/html/2609.30474#S8.E174)\) for a new solution\(a′,b′\)∈c\(𝒑,𝒒\)\(a^\{\\prime\},b^\{\\prime\}\)\\in\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\)whereb′b^\{\\prime\}has already been defined anda′=\.a\+δaa^\{\\prime\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}a\+\\delta\_\{a\}is such that
\(a\+δa\)q\(𝔸\)\\displaystyle\(a\+\\delta\_\{a\}\)q\(\\mathbb\{A\}\)=\\displaystyle=Ra,\\displaystyle R\_\{a\},\(176\)which means we need to guarantee that there is no change in𝔸\\mathbb\{A\}in the process of moving frombbtob′b^\{\\prime\}, i\.e\.a\+δa<pi/qi−1a\+\\delta\_\{a\}<p\_\{i\}/q\_\{i\}\-1\. ReplacingRaR\_\{a\}in \([175](https://arxiv.org/html/2609.30474#S8.E175)\) by its expression in \([176](https://arxiv.org/html/2609.30474#S8.E176)\) and using \([174](https://arxiv.org/html/2609.30474#S8.E174)\) yields the sufficient conditions for\(a′,b′\)∈c\(𝒑,𝒒\)\(a^\{\\prime\},b^\{\\prime\}\)\\in\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\):
δa\\displaystyle\\delta\_\{a\}=\\displaystyle=Δb⋅q\(𝔹\)q\(𝔸\)=\(1−pjqj−b\)⋅q\(𝔹\)q\(𝔸\),\\displaystyle\\Delta\_\{b\}\\cdot\\frac\{q\(\\mathbb\{B\}\)\}\{q\(\\mathbb\{A\}\)\}=\\left\(1\-\\frac\{p\_\{j\}\}\{q\_\{j\}\}\-b\\right\)\\cdot\\frac\{q\(\\mathbb\{B\}\)\}\{q\(\\mathbb\{A\}\)\},\(177\)a\+δa\\displaystyle a\+\\delta\_\{a\}<\\displaystyle<piqi−1\.\\displaystyle\\frac\{p\_\{i\}\}\{q\_\{i\}\}\-1\.\(178\)and we check that the inequality is \([173](https://arxiv.org/html/2609.30474#S8.E173)\)\. To summarize, if\(a,b\)∈c\(𝒑,𝒒\)\(a,b\)\\in\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\)and \([173](https://arxiv.org/html/2609.30474#S8.E173)\) holds, then the new breakpoint
\(a′,b′\)\\displaystyle\(a^\{\\prime\},b^\{\\prime\}\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\(a\+\(1−pjqj−b\)⋅q\(𝔹\)q\(𝔸\),1−pjqj\)\\displaystyle\\left\(a\+\\left\(1\-\\frac\{p\_\{j\}\}\{q\_\{j\}\}\-b\\right\)\\cdot\\frac\{q\(\\mathbb\{B\}\)\}\{q\(\\mathbb\{A\}\)\},1\-\\frac\{p\_\{j\}\}\{q\_\{j\}\}\\right\)\(179\)is inc\(𝒑,𝒒\)\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\)\. Also𝔸\\mathbb\{A\}does not change but we have the updates𝔹←𝔹\\\{j\}\\mathbb\{B\}\\leftarrow\\mathbb\{B\}\\backslash\\\{j\\\}\(one index less\) and𝕀←𝕀∪\{j\}\\mathbb\{I\}\\leftarrow\\mathbb\{I\}\\cup\\\{j\\\}\.
Case 2: suppose that the current breakpoint satisfies
\(piqi−1−a\)⋅q\(𝔸\)\\displaystyle\\left\(\\frac\{p\_\{i\}\}\{q\_\{i\}\}\-1\-a\\right\)\\cdot q\(\\mathbb\{A\}\)<\\displaystyle<\(1−b−pjqj\)⋅q\(𝔹\)\.\\displaystyle\\left\(1\-b\-\\frac\{p\_\{j\}\}\{q\_\{j\}\}\\right\)\\cdot q\(\\mathbb\{B\}\)\.\(180\)Leta′=\.pi/qi−1a^\{\\prime\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}p\_\{i\}/q\_\{i\}\-1andΔa=\.a′−a\>0\\Delta\_\{a\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}a^\{\\prime\}\-a\>0\(by definition of𝔸\\mathbb\{A\}\([33](https://arxiv.org/html/2609.30474#S3.E33)\)\)\. Reformulate \([38](https://arxiv.org/html/2609.30474#S3.E38)\) as
bq\(𝔹\)\\displaystyle bq\(\\mathbb\{B\}\)=\\displaystyle=aq\(𝔸\)−q\(𝕀\)\+p\(𝕀\),\\displaystyle aq\(\\mathbb\{A\}\)\-q\(\\mathbb\{I\}\)\+p\(\\mathbb\{I\}\),\(181\)and rewrite the RHS:
aq\(𝔸\)−q\(𝕀\)\+p\(𝕀\)\\displaystyle aq\(\\mathbb\{A\}\)\-q\(\\mathbb\{I\}\)\+p\(\\mathbb\{I\}\)\(182\)=\\displaystyle=\(piqi−1−Δa\)\(q\(𝔸\\\{i\}\)\+qi\)−q\(𝕀\)\+p\(𝕀\)\\displaystyle\\left\(\\frac\{p\_\{i\}\}\{q\_\{i\}\}\-1\-\\Delta\_\{a\}\\right\)\(q\(\\mathbb\{A\}\\backslash\\\{i\\\}\)\+q\_\{i\}\)\-q\(\\mathbb\{I\}\)\+p\(\\mathbb\{I\}\)=\\displaystyle=\(piqi−1\)q\(𝔸\\\{i\}\)−q\(𝕀\)\+p\(𝕀\)\+\(pi−qi\)−Δa⋅\(q\(𝔸\\\{i\}\)\+qi\)\\displaystyle\\left\(\\frac\{p\_\{i\}\}\{q\_\{i\}\}\-1\\right\)q\(\\mathbb\{A\}\\backslash\\\{i\\\}\)\-q\(\\mathbb\{I\}\)\+p\(\\mathbb\{I\}\)\+\(p\_\{i\}\-q\_\{i\}\)\-\\Delta\_\{a\}\\cdot\(q\(\\mathbb\{A\}\\backslash\\\{i\\\}\)\+q\_\{i\}\)=\\displaystyle=\(piqi−1\)q\(𝔸\\\{i\}\)−q\(𝕀∪\{i\}\)\+p\(𝕀∪\{i\}\)⏟=\.Rb−Δa⋅q\(𝔸\)\.\\displaystyle\\underbrace\{\\left\(\\frac\{p\_\{i\}\}\{q\_\{i\}\}\-1\\right\)q\(\\mathbb\{A\}\\backslash\\\{i\\\}\)\-q\(\\mathbb\{I\}\\cup\\\{i\\\}\)\+p\(\\mathbb\{I\}\\cup\\\{i\\\}\)\}\_\{\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}R\_\{b\}\}\-\\Delta\_\{a\}\\cdot q\(\\mathbb\{A\}\)\.Note thatRbR\_\{b\}is the RHS of \([181](https://arxiv.org/html/2609.30474#S8.E181)\) for a new solution\(a′,b′\)∈c\(𝒑,𝒒\)\(a^\{\\prime\},b^\{\\prime\}\)\\in\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\)wherea′a^\{\\prime\}has already been defined andb′=\.b\+δbb^\{\\prime\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}b\+\\delta\_\{b\}is such that
\(b\+δb\)q\(𝔹\)\\displaystyle\(b\+\\delta\_\{b\}\)q\(\\mathbb\{B\}\)=\\displaystyle=Rb,\\displaystyle R\_\{b\},\(183\)which means we need to guarantee that there is no change in𝔹\\mathbb\{B\}in the process of moving fromaatoa′a^\{\\prime\}, i\.e\.b\+δb<1−pj/qjb\+\\delta\_\{b\}<1\-p\_\{j\}/q\_\{j\}\. ReplacingRbR\_\{b\}in \([182](https://arxiv.org/html/2609.30474#S8.E182)\) by its expression in \([183](https://arxiv.org/html/2609.30474#S8.E183)\) and using \([181](https://arxiv.org/html/2609.30474#S8.E181)\) yields the sufficient conditions for\(a′,b′\)∈c\(𝒑,𝒒\)\(a^\{\\prime\},b^\{\\prime\}\)\\in\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\):
δb\\displaystyle\\delta\_\{b\}=\\displaystyle=Δa⋅q\(𝔸\)q\(𝔹\)=\(piqi−1−a\)⋅q\(𝔸\)q\(𝔹\),\\displaystyle\\Delta\_\{a\}\\cdot\\frac\{q\(\\mathbb\{A\}\)\}\{q\(\\mathbb\{B\}\)\}=\\left\(\\frac\{p\_\{i\}\}\{q\_\{i\}\}\-1\-a\\right\)\\cdot\\frac\{q\(\\mathbb\{A\}\)\}\{q\(\\mathbb\{B\}\)\},\(184\)b\+δb\\displaystyle b\+\\delta\_\{b\}<\\displaystyle<1−pjqj\.\\displaystyle 1\-\\frac\{p\_\{j\}\}\{q\_\{j\}\}\.\(185\)and we check that the inequality is \([180](https://arxiv.org/html/2609.30474#S8.E180)\)\. To summarize, if\(a,b\)∈c\(𝒑,𝒒\)\(a,b\)\\in\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\)and \([180](https://arxiv.org/html/2609.30474#S8.E180)\) holds, then the new breakpoint
\(a′,b′\)\\displaystyle\(a^\{\\prime\},b^\{\\prime\}\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\(piqi−1,b\+\(piqi−1−a\)⋅q\(𝔸\)q\(𝔹\)\)\\displaystyle\\left\(\\frac\{p\_\{i\}\}\{q\_\{i\}\}\-1,b\+\\left\(\\frac\{p\_\{i\}\}\{q\_\{i\}\}\-1\-a\\right\)\\cdot\\frac\{q\(\\mathbb\{A\}\)\}\{q\(\\mathbb\{B\}\)\}\\right\)is inc\(𝒑,𝒒\)\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\)\. Also𝔹\\mathbb\{B\}does not change but we have the updates𝔸←𝔸\\\{i\}\\mathbb\{A\}\\leftarrow\\mathbb\{A\}\\backslash\\\{i\\\}\(one index less\) and𝕀←𝕀∪\{i\}\\mathbb\{I\}\\leftarrow\\mathbb\{I\}\\cup\\\{i\\\}\.
Case 3: suppose that the current breakpoint satisfies
\(piqi−1−a\)⋅q\(𝔸\)\\displaystyle\\left\(\\frac\{p\_\{i\}\}\{q\_\{i\}\}\-1\-a\\right\)\\cdot q\(\\mathbb\{A\}\)=\\displaystyle=\(1−b−pjqj\)⋅q\(𝔹\)\.\\displaystyle\\left\(1\-b\-\\frac\{p\_\{j\}\}\{q\_\{j\}\}\\right\)\\cdot q\(\\mathbb\{B\}\)\.\(186\)We now work with the following equivalent to \([38](https://arxiv.org/html/2609.30474#S3.E38)\):
−aq\(𝔸\)\+bq\(𝔹\)−p\(𝕀\)\+q\(𝕀\)\\displaystyle\-aq\(\\mathbb\{A\}\)\+bq\(\\mathbb\{B\}\)\-p\(\\mathbb\{I\}\)\+q\(\\mathbb\{I\}\)=\\displaystyle=0,\\displaystyle 0,\(187\)and we now considera′=\.pi/qi−1,b′=\.pj/qja^\{\\prime\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}p\_\{i\}/q\_\{i\}\-1,b^\{\\prime\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}p\_\{j\}/q\_\{j\}simultaneously\. Rewrite the LHS of \([187](https://arxiv.org/html/2609.30474#S8.E187)\) using both \([175](https://arxiv.org/html/2609.30474#S8.E175)\) and \([182](https://arxiv.org/html/2609.30474#S8.E182)\) with their notations as
−aq\(𝔸\)\+bq\(𝔹\)−p\(𝕀\)\+q\(𝕀\)\\displaystyle\-aq\(\\mathbb\{A\}\)\+bq\(\\mathbb\{B\}\)\-p\(\\mathbb\{I\}\)\+q\(\\mathbb\{I\}\)\(188\)=\\displaystyle=−\(piqi−1\)q\(𝔸\\\{i\}\)−p\(𝕀\)\+q\(𝕀\)−pi\+qi\+Δa⋅q\(𝔸\)\\displaystyle\-\\left\(\\frac\{p\_\{i\}\}\{q\_\{i\}\}\-1\\right\)q\(\\mathbb\{A\}\\backslash\\\{i\\\}\)\-p\(\\mathbb\{I\}\)\+q\(\\mathbb\{I\}\)\-p\_\{i\}\+q\_\{i\}\+\\Delta\_\{a\}\\cdot q\(\\mathbb\{A\}\)\+\(1−pjqj\)q\(𝔹\\\{j\}\)\+qj−pj−Δb⋅q\(𝔹\),\\displaystyle\+\\left\(1\-\\frac\{p\_\{j\}\}\{q\_\{j\}\}\\right\)q\(\\mathbb\{B\}\\backslash\\\{j\\\}\)\+q\_\{j\}\-p\_\{j\}\-\\Delta\_\{b\}\\cdot q\(\\mathbb\{B\}\),and we check that \([186](https://arxiv.org/html/2609.30474#S8.E186)\) impliesΔa⋅q\(𝔸\)=Δb⋅q\(𝔹\)\\Delta\_\{a\}\\cdot q\(\\mathbb\{A\}\)=\\Delta\_\{b\}\\cdot q\(\\mathbb\{B\}\), so with
\(a′,b′\)\\displaystyle\(a^\{\\prime\},b^\{\\prime\}\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\(piqi−1,1−pjqj\)\\displaystyle\\left\(\\frac\{p\_\{i\}\}\{q\_\{i\}\}\-1,1\-\\frac\{p\_\{j\}\}\{q\_\{j\}\}\\right\)we check that \([188](https://arxiv.org/html/2609.30474#S8.E188)\) becomes
−aq\(𝔸\)\+bq\(𝔹\)−p\(𝕀\)\+q\(𝕀\)\\displaystyle\-aq\(\\mathbb\{A\}\)\+bq\(\\mathbb\{B\}\)\-p\(\\mathbb\{I\}\)\+q\(\\mathbb\{I\}\)=\\displaystyle=−a′q\(𝔸′\)\+b′q\(𝔹′\)−p\(𝕀′\)\+q\(𝕀′\),\\displaystyle\-a^\{\\prime\}q\(\\mathbb\{A\}^\{\\prime\}\)\+b^\{\\prime\}q\(\\mathbb\{B\}^\{\\prime\}\)\-p\(\\mathbb\{I\}^\{\\prime\}\)\+q\(\\mathbb\{I\}^\{\\prime\}\),\(189\)with𝔸′=\.𝔸\\\{i\}\\mathbb\{A\}^\{\\prime\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\mathbb\{A\}\\backslash\\\{i\\\},𝔹′=\.𝔹\\\{j\}\\mathbb\{B\}^\{\\prime\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\mathbb\{B\}\\backslash\\\{j\\\}and𝕀′=\.𝕀∪\{j,i\}\\mathbb\{I\}^\{\\prime\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\mathbb\{I\}\\cup\\\{j,i\\\}the updates to make, and since \([189](https://arxiv.org/html/2609.30474#S8.E189)\) is 0 from \([187](https://arxiv.org/html/2609.30474#S8.E187)\),\(a′,b′\)∈c\(𝒑,𝒒\)\(a^\{\\prime\},b^\{\\prime\}\)\\in\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\)\. This achieves the proof of Theorem[5\.1](https://arxiv.org/html/2609.30474#S5.Thmtheorem1)\.
### VIII\.12Proof of Lemma[5\.2](https://arxiv.org/html/2609.30474#S5.Thmtheorem2)
We prove Lemma[5\.2](https://arxiv.org/html/2609.30474#S5.Thmtheorem2)by using the proof of Theorem[5\.1](https://arxiv.org/html/2609.30474#S5.Thmtheorem1)in Section[VIII\.11](https://arxiv.org/html/2609.30474#S8.SS11)\. If we are not on a breakpoint \(otherwise, the algorithm returns the breakpoint\), we have two cases:
Case 1: We querya~\\tilde\{a\}and ask forb~\\tilde\{b\}such that\(a~,b~\)∈𝒞\(𝒑,𝒒\)\(\\tilde\{a\},\\tilde\{b\}\)\\in\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\)\. Suppose we havea<a~<a′a<\\tilde\{a\}<a^\{\\prime\}for two consecutive breakpoints\(a,b\)\(a,b\)and\(a′,b′\)\(a^\{\\prime\},b^\{\\prime\}\)inc\(𝒑,𝒒\)\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\)\. We have three subcases, where indexesi,ji,jare defined in the proof of Theorem[5\.1](https://arxiv.org/html/2609.30474#S5.Thmtheorem1):
Subcase 1\.1: we have
\(a′,b′\)\\displaystyle\(a^\{\\prime\},b^\{\\prime\}\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\(a\+\(1−pjqj−b\)⋅q\(𝔹\)q\(𝔸\),1−pjqj\),\\displaystyle\\left\(a\+\\left\(1\-\\frac\{p\_\{j\}\}\{q\_\{j\}\}\-b\\right\)\\cdot\\frac\{q\(\\mathbb\{B\}\)\}\{q\(\\mathbb\{A\}\)\},1\-\\frac\{p\_\{j\}\}\{q\_\{j\}\}\\right\),so we reuse the proof of Theorem[5\.1](https://arxiv.org/html/2609.30474#S5.Thmtheorem1), Case 1\. Remark that as long asδb′=ε⋅Δb\\delta^\{\\prime\}\_\{b\}=\\varepsilon\\cdot\\Delta\_\{b\}forb′=\.b\+δb′b^\{\\prime\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}b\+\\delta^\{\\prime\}\_\{b\}with0<ε<10<\\varepsilon<1, the RHS of \([174](https://arxiv.org/html/2609.30474#S8.E174)\) is
bq\(𝔹\)\+q\(𝕀\)−p\(𝕀\)\\displaystyle bq\(\\mathbb\{B\}\)\+q\(\\mathbb\{I\}\)\-p\(\\mathbb\{I\}\)=\\displaystyle=b′q\(𝔹\)\+q\(𝕀\)−p\(𝕀\)−ε⋅Δb⋅q\(𝔹\)\.\\displaystyle b^\{\\prime\}q\(\\mathbb\{B\}\)\+q\(\\mathbb\{I\}\)\-p\(\\mathbb\{I\}\)\-\\varepsilon\\cdot\\Delta\_\{b\}\\cdot q\(\\mathbb\{B\}\)\.while the LHS becomes\(a′−δa′\)q\(𝔸\)=−δa′q\(𝔸\)\+a′q\(𝔸\)\(a^\{\\prime\}\-\\delta^\{\\prime\}\_\{a\}\)q\(\\mathbb\{A\}\)=\-\\delta^\{\\prime\}\_\{a\}q\(\\mathbb\{A\}\)\+a^\{\\prime\}q\(\\mathbb\{A\}\)\. Sinceε<1\\varepsilon<1,𝔸,𝔹,𝕀\\mathbb\{A\},\\mathbb\{B\},\\mathbb\{I\}do not change and if we ensure \(i\)δa′≤δa\\delta^\{\\prime\}\_\{a\}\\leq\\delta\_\{a\}\([177](https://arxiv.org/html/2609.30474#S8.E177)\) and \(ii\)δa′q\(𝔸\)=ε⋅Δb⋅q\(𝔹\)\\delta^\{\\prime\}\_\{a\}q\(\\mathbb\{A\}\)=\\varepsilon\\cdot\\Delta\_\{b\}\\cdot q\(\\mathbb\{B\}\), yielding
δa′\\displaystyle\\delta^\{\\prime\}\_\{a\}=\\displaystyle=ε⋅Δb⋅q\(𝔹\)q\(𝔸\),\\displaystyle\\varepsilon\\cdot\\Delta\_\{b\}\\cdot\\frac\{q\(\\mathbb\{B\}\)\}\{q\(\\mathbb\{A\}\)\},and we checkδa′≤δa\\delta^\{\\prime\}\_\{a\}\\leq\\delta\_\{a\}\. Solving forε\\varepsilonwhileδa′=\.a~−a\\delta^\{\\prime\}\_\{a\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\tilde\{a\}\-ayieldsε=\(a~−a\)q\(𝔸\)/\(Δb⋅q\(𝔹\)\)\\varepsilon=\(\\tilde\{a\}\-a\)q\(\\mathbb\{A\}\)/\(\\Delta\_\{b\}\\cdot q\(\\mathbb\{B\}\)\)and the solution
\(a~,b~=\.b\+\(a~−a\)⋅q\(𝔸\)q\(𝔹\)\)∈𝒞\(𝒑,𝒒\),\\displaystyle\\left\(\\tilde\{a\},\\tilde\{b\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}b\+\(\\tilde\{a\}\-a\)\\cdot\\frac\{q\(\\mathbb\{A\}\)\}\{q\(\\mathbb\{B\}\)\}\\right\)\\in\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\),where𝔸,𝔹\\mathbb\{A\},\\mathbb\{B\}are associated to breakpoint\(a,b\)\(a,b\)\.
Subcase 1\.2: we have
\(a′,b′\)\\displaystyle\(a^\{\\prime\},b^\{\\prime\}\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\(piqi−1,b\+\(piqi−1−a\)⋅q\(𝔸\)q\(𝔹\)\)\\displaystyle\\left\(\\frac\{p\_\{i\}\}\{q\_\{i\}\}\-1,b\+\\left\(\\frac\{p\_\{i\}\}\{q\_\{i\}\}\-1\-a\\right\)\\cdot\\frac\{q\(\\mathbb\{A\}\)\}\{q\(\\mathbb\{B\}\)\}\\right\)so we reuse the proof of Theorem[5\.1](https://arxiv.org/html/2609.30474#S5.Thmtheorem1), Case 2\. Remark that as long asδb′=ε⋅Δaq\(𝔸\)/q\(𝔹\)\\delta^\{\\prime\}\_\{b\}=\\varepsilon\\cdot\\Delta\_\{a\}q\(\\mathbb\{A\}\)/q\(\\mathbb\{B\}\)forb′=\.b\+δb′b^\{\\prime\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}b\+\\delta^\{\\prime\}\_\{b\}with0<ε<10<\\varepsilon<1, the RHS of \([181](https://arxiv.org/html/2609.30474#S8.E181)\) is
aq\(𝔸\)−q\(𝕀\)\+p\(𝕀\)\\displaystyle aq\(\\mathbb\{A\}\)\-q\(\\mathbb\{I\}\)\+p\(\\mathbb\{I\}\)=\\displaystyle=a′q\(𝔸\)−q\(𝕀\)\+p\(𝕀\)−ε⋅Δa⋅q\(𝔸\)\.\\displaystyle a^\{\\prime\}q\(\\mathbb\{A\}\)\-q\(\\mathbb\{I\}\)\+p\(\\mathbb\{I\}\)\-\\varepsilon\\cdot\\Delta\_\{a\}\\cdot q\(\\mathbb\{A\}\)\.while the LHS becomes\(b′−δb′\)q\(𝔹\)=−δb′q\(𝔹\)\+b′q\(𝔹\)\(b^\{\\prime\}\-\\delta^\{\\prime\}\_\{b\}\)q\(\\mathbb\{B\}\)=\-\\delta^\{\\prime\}\_\{b\}q\(\\mathbb\{B\}\)\+b^\{\\prime\}q\(\\mathbb\{B\}\)\. Sinceε<1\\varepsilon<1,𝔸,𝔹,𝕀\\mathbb\{A\},\\mathbb\{B\},\\mathbb\{I\}do not change and if we ensure \(i\)δb′≤δb\\delta^\{\\prime\}\_\{b\}\\leq\\delta\_\{b\}\([184](https://arxiv.org/html/2609.30474#S8.E184)\) and \(ii\)δb′q\(𝔹\)=ε⋅Δa⋅q\(𝔸\)\\delta^\{\\prime\}\_\{b\}q\(\\mathbb\{B\}\)=\\varepsilon\\cdot\\Delta\_\{a\}\\cdot q\(\\mathbb\{A\}\), yielding
δb′\\displaystyle\\delta^\{\\prime\}\_\{b\}=\\displaystyle=ε⋅Δa⋅q\(𝔸\)q\(𝔹\),\\displaystyle\\varepsilon\\cdot\\Delta\_\{a\}\\cdot\\frac\{q\(\\mathbb\{A\}\)\}\{q\(\\mathbb\{B\}\)\},and we checkδb′≤δb\\delta^\{\\prime\}\_\{b\}\\leq\\delta\_\{b\}\. This time,Δa\\Delta\_\{a\}mapsaato the next breakpoint in the proof of Theorem[5\.1](https://arxiv.org/html/2609.30474#S5.Thmtheorem1)so we haveε⋅Δa=a~−a\\varepsilon\\cdot\\Delta\_\{a\}=\\tilde\{a\}\-aand we get that the solution
\(a~,b~=\.b\+\(a~−a\)⋅q\(𝔸\)q\(𝔹\)\)∈𝒞\(𝒑,𝒒\),\\displaystyle\\left\(\\tilde\{a\},\\tilde\{b\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}b\+\(\\tilde\{a\}\-a\)\\cdot\\frac\{q\(\\mathbb\{A\}\)\}\{q\(\\mathbb\{B\}\)\}\\right\)\\in\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\),where𝔸,𝔹\\mathbb\{A\},\\mathbb\{B\}are associated to breakpoint\(a,b\)\(a,b\)\.
Subcase 1\.3is Theorem[5\.1](https://arxiv.org/html/2609.30474#S5.Thmtheorem1), Case 3, and yields the same solution as the two preceding cases\.
Case 2: We queryb~\\tilde\{b\}and ask fora~\\tilde\{a\}such that\(a~,b~\)∈𝒞\(𝒑,𝒒\)\(\\tilde\{a\},\\tilde\{b\}\)\\in\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\)\. This time, we suppose we haveb<b~<b′b<\\tilde\{b\}<b^\{\\prime\}for two consecutive breakpoints\(a,b\)\(a,b\)and\(a′,b′\)\(a^\{\\prime\},b^\{\\prime\}\)inc\(𝒑,𝒒\)\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\)\. Sparing all the computation, we get this time
\(a~=\.a\+\(b~−b\)⋅q\(𝔹\)q\(𝔸\),b~\)∈𝒞\(𝒑,𝒒\),\\displaystyle\\left\(\\tilde\{a\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}a\+\(\\tilde\{b\}\-b\)\\cdot\\frac\{q\(\\mathbb\{B\}\)\}\{q\(\\mathbb\{A\}\)\},\\tilde\{b\}\\right\)\\in\\mathcal\{C\}\(\\bm\{p\},\\bm\{q\}\),where𝔸,𝔹\\mathbb\{A\},\\mathbb\{B\}are associated to breakpoint\(a,b\)\(a,b\)\. This ends the proof of Lemma[5\.2](https://arxiv.org/html/2609.30474#S5.Thmtheorem2)\.
### VIII\.13Proof of Lemma[5\.3](https://arxiv.org/html/2609.30474#S5.Thmtheorem3)
All properties except continuity are immediate consequences of Lemma[5\.2](https://arxiv.org/html/2609.30474#S5.Thmtheorem2)\. For the continuity part, pick anyaasatisfying \([36](https://arxiv.org/html/2609.30474#S3.E36)\)\. Forb=0b=0, we necessarily have−aq\(𝔸\)\+bq\(𝔹\)=−aq\(𝔸\)<p\(𝕀\)−q\(𝕀\)\-aq\(\\mathbb\{A\}\)\+bq\(\\mathbb\{B\}\)=\-aq\(\\mathbb\{A\}\)<p\(\\mathbb\{I\}\)\-q\(\\mathbb\{I\}\)sincep\(𝕀\)−q\(𝕀\)=∑i:qi≤pi≤\(1\+a\)qipi−qi≥0p\(\\mathbb\{I\}\)\-q\(\\mathbb\{I\}\)=\\sum\_\{i:q\_\{i\}\\leq p\_\{i\}\\leq\(1\+a\)q\_\{i\}\}p\_\{i\}\-q\_\{i\}\\geq 0\. Elements in𝔸\\mathbb\{A\}havepi\>\(1\+a\)qip\_\{i\}\>\(1\+a\)q\_\{i\}, which yields after summing and taking negation−aq\(𝔸\)\>q\(𝔸\)−p\(𝔸\)\-aq\(\\mathbb\{A\}\)\>q\(\\mathbb\{A\}\)\-p\(\\mathbb\{A\}\)so for the other extreme case,b=1b=1, since we have in this casep\(𝕀\)−q\(𝕀\)=\(1−p\(𝔸\)\)−\(1−q\(𝔸\)\)=q\(𝔸\)−p\(𝔸\)p\(\\mathbb\{I\}\)\-q\(\\mathbb\{I\}\)=\(1\-p\(\\mathbb\{A\}\)\)\-\(1\-q\(\\mathbb\{A\}\)\)=q\(\\mathbb\{A\}\)\-p\(\\mathbb\{A\}\), we observe−aq\(𝔸\)\+0\>p\(𝔸\)−q\(𝔸\)=p\(𝕀\)−q\(𝕀\)\-aq\(\\mathbb\{A\}\)\+0\>p\(\\mathbb\{A\}\)\-q\(\\mathbb\{A\}\)=p\(\\mathbb\{I\}\)\-q\(\\mathbb\{I\}\)\. Now, forb∈\(0,1\)b\\in\(0,1\), reformulate \([38](https://arxiv.org/html/2609.30474#S3.E38)\) as
aq\(𝔸\)\\displaystyle aq\(\\mathbb\{A\}\)=\\displaystyle=bq\(𝔹\)\+q\(𝕀\)−p\(𝕀\),\\displaystyle bq\(\\mathbb\{B\}\)\+q\(\\mathbb\{I\}\)\-p\(\\mathbb\{I\}\),\(190\)and remark that having chosenaa, the LHS is fixed\. Suppose the currentbbis at\(1−b\)=\(pi/qi\)\+δ\(1\-b\)=\(p\_\{i\}/q\_\{i\}\)\+\\deltawithδ\>0\\delta\>0, meaningi∈𝔹i\\in\\mathbb\{B\}, and suppose no other element of𝔹\\mathbb\{B\}is closer\. Supposeδ\\deltasmall enough so that when we increasebb,q\(𝕀\)−p\(𝕀\)q\(\\mathbb\{I\}\)\-p\(\\mathbb\{I\}\)does not change\. In this case, the RHS of \([190](https://arxiv.org/html/2609.30474#S8.E190)\) equals
bq\(𝔹\)\+q\(𝕀\)−p\(𝕀\)\\displaystyle bq\(\\mathbb\{B\}\)\+q\(\\mathbb\{I\}\)\-p\(\\mathbb\{I\}\)=\\displaystyle=\(1−δ−piqi\)\(q\(𝔹\\\{i\}\)\+qi\)\+q\(𝕀\)−p\(𝕀\)\\displaystyle\\left\(1\-\\delta\-\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\right\)\(q\(\\mathbb\{B\}\\backslash\\\{i\\\}\)\+q\_\{i\}\)\+q\(\\mathbb\{I\}\)\-p\(\\mathbb\{I\}\)=\\displaystyle=\(1−piqi\)q\(𝔹\\\{i\}\)\+q\(𝕀\)−p\(𝕀\)\+\(qi−pi\)−δ⋅\(q\(𝔹\\\{i\}\)\+qi\)\\displaystyle\\left\(1\-\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\right\)q\(\\mathbb\{B\}\\backslash\\\{i\\\}\)\+q\(\\mathbb\{I\}\)\-p\(\\mathbb\{I\}\)\+\(q\_\{i\}\-p\_\{i\}\)\-\\delta\\cdot\(q\(\\mathbb\{B\}\\backslash\\\{i\\\}\)\+q\_\{i\}\)=\\displaystyle=\(1−piqi\)q\(𝔹\\\{i\}\)\+q\(𝕀∪\{i\}\)−p\(𝕀∪\{i\}\)−δ⋅\(q\(𝔹\\\{i\}\)\+qi\),\\displaystyle\\left\(1\-\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\right\)q\(\\mathbb\{B\}\\backslash\\\{i\\\}\)\+q\(\\mathbb\{I\}\\cup\\\{i\\\}\)\-p\(\\mathbb\{I\}\\cup\\\{i\\\}\)\-\\delta\\cdot\(q\(\\mathbb\{B\}\\backslash\\\{i\\\}\)\+q\_\{i\}\),and asbbcontinuously increases further whileδ\>0\\delta\>0decreases, the RHS increases, but the limit of the RHS asδ→0\+\\delta\\rightarrow 0^\{\+\}is\(1−\(pi/qi\)\)q\(𝔹\\\{i\}\)\+q\(𝕀∪\{i\}\)−p\(𝕀∪\{i\}\)\\left\(1\-\(p\_\{i\}/q\_\{i\}\)\\right\)q\(\\mathbb\{B\}\\backslash\\\{i\\\}\)\+q\(\\mathbb\{I\}\\cup\\\{i\\\}\)\-p\(\\mathbb\{I\}\\cup\\\{i\\\}\), which, in fact, is the RHS of \([190](https://arxiv.org/html/2609.30474#S8.E190)\) asiigoes from𝔹\\mathbb\{B\}to𝕀\\mathbb\{I\}\. What we just showed is that the RHS of \([190](https://arxiv.org/html/2609.30474#S8.E190)\) has continuous variations withb∈\[0,1\]b\\in\[0,1\]\. Sincec\(𝒑,𝒒\)=\.\{\(ai,bi\):i∈\[N\]\}\\textsc\{c\}\(\\bm\{p\},\\bm\{q\}\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\\{\(a\_\{i\},b\_\{i\}\):i\\in\[N\]\\\}has its elements indexed in strictly increasing values of bothaaandbbwith\(a1,b1\)=\(0,0\)\(a\_\{1\},b\_\{1\}\)=\(0,0\), this achieves the proof of the Lemma \(choosingbbfirst yields the same proof\)\.
### VIII\.14Proof of Theorem[5\.4](https://arxiv.org/html/2609.30474#S5.Thmtheorem4)
As it is formulated, the problem can be conveniently solved by addressing the dual of \([f \-MD\-2](https://arxiv.org/html/2609.30474#S3.Ex1)\):
dmdf2\(𝒑,𝒒,D\)\\displaystyle\\textsc\{dmd\}^\{2\}\_\{f\}\(\\bm\{p\},\\bm\{q\};D\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}argmin𝒓∈\[0,1\]n,𝒔∈ΔnDf\(𝒑⊙𝒓\+\(1−𝒑⊤𝒓\)⋅𝒔∥𝒒\)s\.t\.𝒑⊤𝒓≥P,\\displaystyle\\arg\\min\_\{\\bm\{r\}\\in\[0,1\]^\{n\},\\bm\{s\}\\in\\Delta\_\{n\}\}D\_\{f\}\(\\bm\{p\}\\odot\\bm\{r\}\+\(1\-\\bm\{p\}^\{\\top\}\\bm\{r\}\)\\cdot\\bm\{s\}\\\|\\bm\{q\}\)\\quad\\mbox\{s\.t\. \}\\bm\{p\}^\{\\top\}\\bm\{r\}\\geq P,forP\>Pacc\(SD\)P\>P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)\.
We first prove the \(strict\) convexity part\. For any given outputs𝒑,𝒒\\bm\{p\},\\bm\{q\}of the drafter and target, consider three distinct acceptance probabilitiesP\.P\_\{\.\}, forPb=\.γPa\+\(1−γ\)PcP\_\{b\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\gamma P\_\{a\}\+\(1\-\\gamma\)P\_\{c\}andPa≤Pb≤PcP\_\{a\}\\leq P\_\{b\}\\leq P\_\{c\}\. Denote𝒓a,𝒔a\\bm\{r\}\_\{a\},\\bm\{s\}\_\{a\};𝒓b,𝒔b\\bm\{r\}\_\{b\},\\bm\{s\}\_\{b\}and𝒓c,𝒔c\\bm\{r\}\_\{c\},\\bm\{s\}\_\{c\}the respective optimal solutions components\. We need to show
Df\(𝒑⊙𝒓b\+\(1−𝒑⊤𝒓b\)⋅𝒔b∥𝒒\)\\displaystyle D\_\{f\}\(\\bm\{p\}\\odot\\bm\{r\}\_\{b\}\+\(1\-\\bm\{p\}^\{\\top\}\\bm\{r\}\_\{b\}\)\\cdot\\bm\{s\}\_\{b\}\\\|\\bm\{q\}\)≤\\displaystyle\\leqγ⋅Df\(𝒑⊙𝒓a\+\(1−𝒑⊤𝒓a\)⋅𝒔a∥𝒒\)\\displaystyle\\gamma\\cdot D\_\{f\}\(\\bm\{p\}\\odot\\bm\{r\}\_\{a\}\+\(1\-\\bm\{p\}^\{\\top\}\\bm\{r\}\_\{a\}\)\\cdot\\bm\{s\}\_\{a\}\\\|\\bm\{q\}\)\(191\)\+\(1−γ\)⋅Df\(𝒑⊙𝒓c\+\(1−𝒑⊤𝒓c\)⋅𝒔c∥𝒒\),\\displaystyle\+\(1\-\\gamma\)\\cdot D\_\{f\}\(\\bm\{p\}\\odot\\bm\{r\}\_\{c\}\+\(1\-\\bm\{p\}^\{\\top\}\\bm\{r\}\_\{c\}\)\\cdot\\bm\{s\}\_\{c\}\\\|\\bm\{q\}\),\(192\)and this will hold if we can find afeasiblesolution forPbP\_\{b\}that can be put in the LHS \(the corresponding optimal one cannot increase the divergence by definition\)\. First, pick
𝒓~\\displaystyle\\tilde\{\\bm\{r\}\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}γ⋅𝒓a\+\(1−γ\)⋅𝒓c\.\\displaystyle\\gamma\\cdot\\bm\{r\}\_\{a\}\+\(1\-\\gamma\)\\cdot\\bm\{r\}\_\{c\}\.\(193\)Since𝒑⊤𝒓a≥Pa\\bm\{p\}^\{\\top\}\\bm\{r\}\_\{a\}\\geq P\_\{a\}and𝒑⊤𝒓c≥Pc\\bm\{p\}^\{\\top\}\\bm\{r\}\_\{c\}\\geq P\_\{c\}, we have𝒑⊤𝒓~≥Pb\\bm\{p\}^\{\\top\}\\tilde\{\\bm\{r\}\}\\geq P\_\{b\}and also𝒓~∈\[0,1\]n\\tilde\{\\bm\{r\}\}\\in\[0,1\]^\{n\}\. To find the corresponding feasible𝒔~∈Δn\\tilde\{\\bm\{s\}\}\\in\\Delta\_\{n\}, we want it to satisfy
\(1−𝒑⊤𝒓~\)⋅𝒔~=γ\(1−𝒑⊤𝒓a\)⋅𝒔a\+\(1−γ\)\(1−𝒑⊤𝒓c\)⋅𝒔c,\\displaystyle\(1\-\\bm\{p\}^\{\\top\}\\tilde\{\\bm\{r\}\}\)\\cdot\\tilde\{\\bm\{s\}\}=\\gamma\(1\-\\bm\{p\}^\{\\top\}\\bm\{r\}\_\{a\}\)\\cdot\\bm\{s\}\_\{a\}\+\(1\-\\gamma\)\(1\-\\bm\{p\}^\{\\top\}\\bm\{r\}\_\{c\}\)\\cdot\\bm\{s\}\_\{c\},but since1−𝒑⊤𝒓~=γ\(1−𝒑⊤𝒓a\)\+\(1−γ\)\(1−𝒑⊤𝒓c\)1\-\\bm\{p\}^\{\\top\}\\tilde\{\\bm\{r\}\}=\\gamma\(1\-\\bm\{p\}^\{\\top\}\\bm\{r\}\_\{a\}\)\+\(1\-\\gamma\)\(1\-\\bm\{p\}^\{\\top\}\\bm\{r\}\_\{c\}\), we may choose𝒔~\\tilde\{\\bm\{s\}\}as:
𝒔~\\displaystyle\\tilde\{\\bm\{s\}\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}ε⋅𝒔a\+\(1−ε\)⋅𝒔c,ε=\.γ⋅1−𝒑⊤𝒓a1−𝒑⊤𝒓~,\\displaystyle\\varepsilon\\cdot\\bm\{s\}\_\{a\}\+\(1\-\\varepsilon\)\\cdot\\bm\{s\}\_\{c\},\\quad\\varepsilon\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\gamma\\cdot\\frac\{1\-\\bm\{p\}^\{\\top\}\\bm\{r\}\_\{a\}\}\{1\-\\bm\{p\}^\{\\top\}\\tilde\{\\bm\{r\}\}\},\(194\)and we have𝒔~∈Δn\\tilde\{\\bm\{s\}\}\\in\\Delta\_\{n\}\.𝒓~,𝒔~\\tilde\{\\bm\{r\}\},\\tilde\{\\bm\{s\}\}as in \([193](https://arxiv.org/html/2609.30474#S8.E193)\), \([194](https://arxiv.org/html/2609.30474#S8.E194)\) is feasible and because of the convexity offf, we get the inequality \(strict ifγ≠0,1\\gamma\\neq 0,1andffis strictly convex\) in
Df\(𝒑⊙𝒓~\+\(1−𝒑⊤𝒓~\)⋅𝒔~∥𝒒\)\\displaystyle D\_\{f\}\(\\bm\{p\}\\odot\\tilde\{\\bm\{r\}\}\+\(1\-\\bm\{p\}^\{\\top\}\\tilde\{\\bm\{r\}\}\)\\cdot\\tilde\{\\bm\{s\}\}\\\|\\bm\{q\}\)=\\displaystyle=Df\(γ⋅\(𝒑⊙𝒓a\+\(1−𝒑⊤𝒓a\)⋅𝒔a\)\+\(1−γ\)⋅\(𝒑⊙𝒓c\+\(1−𝒑⊤𝒓c\)⋅𝒔c\)∥𝒒\)\\displaystyle D\_\{f\}\\left\(\\gamma\\cdot\(\\bm\{p\}\\odot\\bm\{r\}\_\{a\}\+\(1\-\\bm\{p\}^\{\\top\}\\bm\{r\}\_\{a\}\)\\cdot\\bm\{s\}\_\{a\}\)\+\(1\-\\gamma\)\\cdot\(\\bm\{p\}\\odot\\bm\{r\}\_\{c\}\+\(1\-\\bm\{p\}^\{\\top\}\\bm\{r\}\_\{c\}\)\\cdot\\bm\{s\}\_\{c\}\)\\\|\\bm\{q\}\\right\)≤\\displaystyle\\leqγ⋅Df\(𝒑⊙𝒓a\+\(1−𝒑⊤𝒓a\)⋅𝒔a∥𝒒\)\+\(1−γ\)⋅Df\(𝒑⊙𝒓c\+\(1−𝒑⊤𝒓c\)⋅𝒔c∥𝒒\),\\displaystyle\\gamma\\cdot D\_\{f\}\(\\bm\{p\}\\odot\\bm\{r\}\_\{a\}\+\(1\-\\bm\{p\}^\{\\top\}\\bm\{r\}\_\{a\}\)\\cdot\\bm\{s\}\_\{a\}\\\|\\bm\{q\}\)\+\(1\-\\gamma\)\\cdot D\_\{f\}\(\\bm\{p\}\\odot\\bm\{r\}\_\{c\}\+\(1\-\\bm\{p\}^\{\\top\}\\bm\{r\}\_\{c\}\)\\cdot\\bm\{s\}\_\{c\}\\\|\\bm\{q\}\),and of courseDf\(𝒑⊙𝒓b\+\(1−𝒑⊤𝒓b\)⋅𝒔b∥𝒒\)≤Df\(𝒑⊙𝒓~\+\(1−𝒑⊤𝒓~\)⋅𝒔~∥𝒒\)D\_\{f\}\(\\bm\{p\}\\odot\\bm\{r\}\_\{b\}\+\(1\-\\bm\{p\}^\{\\top\}\\bm\{r\}\_\{b\}\)\\cdot\\bm\{s\}\_\{b\}\\\|\\bm\{q\}\)\\leq D\_\{f\}\(\\bm\{p\}\\odot\\tilde\{\\bm\{r\}\}\+\(1\-\\bm\{p\}^\{\\top\}\\tilde\{\\bm\{r\}\}\)\\cdot\\tilde\{\\bm\{s\}\}\\\|\\bm\{q\}\)\(the optimum cannot be worse than any feasible solution\), which shows \([192](https://arxiv.org/html/2609.30474#S8.E192)\) and ends the proof of the \(strict\) convexity part\.
Let us now tackle the right derivative part\. Without loss of generality, indexes are ordered such thatpi\+1/qi\+1≥pi/qi,∀ip\_\{i\+1\}/q\_\{i\+1\}\\geq p\_\{i\}/q\_\{i\},\\forall iand all ratios are distinct \(any of thenndistinct ratio values is called a "tick"\)\. For a current optimal solution given byzβ∈\(minipi/qi,1\],zα∈\[1,maxipi/qi\)z\_\{\\beta\}\\in\(\\min\_\{i\}p\_\{i\}/q\_\{i\},1\],z\_\{\\alpha\}\\in\[1,\\max\_\{i\}p\_\{i\}/q\_\{i\}\), the corresponding value of theff\-divergence is:
Df\(𝝅∥𝒒\)\\displaystyle D\_\{f\}\(\\bm\{\\pi\}\\\|\\bm\{q\}\)=\\displaystyle=∑𝕀βqif\(zβ\)\+∑𝕀qif\(piqi\)\+∑𝕀αqif\(zα\),\\displaystyle\\sum\_\{\\mathbb\{I\}\_\{\\beta\}\}q\_\{i\}f\\left\(z\_\{\\beta\}\\right\)\+\\sum\_\{\\mathbb\{I\}\}q\_\{i\}f\\left\(\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\right\)\+\\sum\_\{\\mathbb\{I\}\_\{\\alpha\}\}q\_\{i\}f\\left\(z\_\{\\alpha\}\\right\),\(195\)with the three sets of ticks
𝕀β=\.\{i:pi/qi<zβ\};𝕀=\.\{i:zβ≤pi/qi≤zα\};𝕀α=\.\{i:zα<pi/qi\}\.\\displaystyle\\mathbb\{I\}\_\{\\beta\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\\{i:p\_\{i\}/q\_\{i\}<z\_\{\\beta\}\\\}\\quad;\\quad\\mathbb\{I\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\\{i:z\_\{\\beta\}\\leq p\_\{i\}/q\_\{i\}\\leq z\_\{\\alpha\}\\\}\\quad;\\quad\\mathbb\{I\}\_\{\\alpha\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\\{i:z\_\{\\alpha\}<p\_\{i\}/q\_\{i\}\\\}\.\(196\)Suppose we pickδβ\>0,δα\>0\\delta\_\{\\beta\}\>0,\\delta\_\{\\alpha\}\>0such that the choice
zβ′\\displaystyle z^\{\\prime\}\_\{\\beta\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}zβ−δβ,\\displaystyle z\_\{\\beta\}\-\\delta\_\{\\beta\},\(197\)zα′\\displaystyle z^\{\\prime\}\_\{\\alpha\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}zα\+δα\\displaystyle z\_\{\\alpha\}\+\\delta\_\{\\alpha\}\(198\)satisfies
- \(i\)it yields a new optimal solution,
- \(ii\)sets𝕀β,𝕀,𝕀α\\mathbb\{I\}\_\{\\beta\},\\mathbb\{I\},\\mathbb\{I\}\_\{\\alpha\}do not change\.
Let us compute the variation of the acceptance probabilityPPand the variation of theff\-divergence as a function of this variation\. We get forPP,
P\(zα′\)=P\(zα\)\+δαq\(𝕀α\)\\displaystyle P\(z^\{\\prime\}\_\{\\alpha\}\)=P\(z\_\{\\alpha\}\)\+\\delta\_\{\\alpha\}q\(\\mathbb\{I\}\_\{\\alpha\}\);P\(zβ′\)=P\(zβ\)\+δβq\(𝕀β\),\\displaystyle P\(z^\{\\prime\}\_\{\\beta\}\)=P\(z\_\{\\beta\}\)\+\\delta\_\{\\beta\}q\(\\mathbb\{I\}\_\{\\beta\}\),\(199\)\(whereq\(𝕌⊆\[n\]\)=\.∑i∈𝕌qiq\(\\mathbb\{U\}\\subseteq\[n\]\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\sum\_\{i\\in\\mathbb\{U\}\}q\_\{i\}\) so for \(i\) to hold we must have the relationship betweenδα\\delta\_\{\\alpha\}andδβ\\delta\_\{\\beta\}
δαq\(𝕀α\)\\displaystyle\\delta\_\{\\alpha\}q\(\\mathbb\{I\}\_\{\\alpha\}\)=\\displaystyle=δβq\(𝕀β\)\.\\displaystyle\\delta\_\{\\beta\}q\(\\mathbb\{I\}\_\{\\beta\}\)\.\(200\)\(and we also need to assumeP\(zα′\)=P\(zβ′\)<1P\(z^\{\\prime\}\_\{\\alpha\}\)=P\(z^\{\\prime\}\_\{\\beta\}\)<1\)\. RecallPacc\(SD\)=\.∑imin\{pi,qi\}P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\sum\_\{i\}\\min\\\{p\_\{i\},q\_\{i\}\\\}and denoteDf\(P\)D\_\{f\}\(P\)the optimal value of theff\-divergence for the requestedPP\. Under Assumption[3\.3](https://arxiv.org/html/2609.30474#S3.Thmtheorem3),pi/qi≠1,∀i∈\[n\]p\_\{i\}/q\_\{i\}\\neq 1,\\forall i\\in\[n\]\. To compute the right derivative atPacc\(SD\)P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\),\(Df\)r′\(Pacc\(SD\)\)\(D\_\{f\}\)^\{\\prime\}\_\{r\}\(P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)\), we start fromzα=zβ=1z\_\{\\alpha\}=z\_\{\\beta\}=1\(Note that𝕀β∪𝕀α=\[n\]\\mathbb\{I\}\_\{\\beta\}\\cup\\mathbb\{I\}\_\{\\alpha\}=\[n\]in \([196](https://arxiv.org/html/2609.30474#S8.E196)\)\) and then compute a variation \(zα′=zα\+δαz^\{\\prime\}\_\{\\alpha\}=z\_\{\\alpha\}\+\\delta\_\{\\alpha\}, negative forzβ′=zβ−δβz^\{\\prime\}\_\{\\beta\}=z\_\{\\beta\}\-\\delta\_\{\\beta\},δα,δβ\>0\\delta\_\{\\alpha\},\\delta\_\{\\beta\}\>0\), small enough not to change𝕀β\\mathbb\{I\}\_\{\\beta\}and𝕀α\\mathbb\{I\}\_\{\\alpha\}\. We know from \([199](https://arxiv.org/html/2609.30474#S8.E199)\) that we must have the relationshipδβqβ=δα\(1−qβ\)\\delta\_\{\\beta\}q\_\{\\beta\}=\\delta\_\{\\alpha\}\(1\-q\_\{\\beta\}\)withqβ=\.q\(𝕀β\)q\_\{\\beta\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}q\(\\mathbb\{I\}\_\{\\beta\}\)\([199](https://arxiv.org/html/2609.30474#S8.E199)\) for optimality to hold for a new valueP′P^\{\\prime\}of the acceptance probability\. Keepingf\(1\)f\(1\)\(=0\) in expressions for clarity, we compute the ratio
Df\(P′\)−Df\(Pacc\(SD\)\)P′−Pacc\(SD\)\\displaystyle\\frac\{D\_\{f\}\(P^\{\\prime\}\)\-D\_\{f\}\(P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)\)\}\{P^\{\\prime\}\-P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)\}=\\displaystyle=∑𝕀βqif\(1−δβ\)\+∑𝕀αqif\(1\+δα\)−f\(1\)Pacc\(SD\)\+δα\(1−qβ\)−Pacc\(SD\)\\displaystyle\\frac\{\\sum\_\{\\mathbb\{I\}\_\{\\beta\}\}q\_\{i\}f\\left\(1\-\\delta\_\{\\beta\}\\right\)\+\\sum\_\{\\mathbb\{I\}\_\{\\alpha\}\}q\_\{i\}f\\left\(1\+\\delta\_\{\\alpha\}\\right\)\-f\(1\)\}\{P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)\+\\delta\_\{\\alpha\}\(1\-q\_\{\\beta\}\)\-P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)\}=\\displaystyle=qβ⋅\(f\(1−δα\(1−qβ\)qβ\)−f\(1\)\)δα\(1−qβ\)\+\(1−qβ\)⋅\(f\(1\+δα\)−f\(1\)\)δα\(1−qβ\)\\displaystyle\\frac\{q\_\{\\beta\}\\cdot\\left\(f\\left\(1\-\\frac\{\\delta\_\{\\alpha\}\(1\-q\_\{\\beta\}\)\}\{q\_\{\\beta\}\}\\right\)\-f\(1\)\\right\)\}\{\\delta\_\{\\alpha\}\(1\-q\_\{\\beta\}\)\}\+\\frac\{\(1\-q\_\{\\beta\}\)\\cdot\\left\(f\\left\(1\+\\delta\_\{\\alpha\}\\right\)\-f\(1\)\\right\)\}\{\\delta\_\{\\alpha\}\(1\-q\_\{\\beta\}\)\}=\\displaystyle=−f\(1\)−f\(1−δ~α\)δ~α\+f\(1\+δα\)−f\(1\)δα\\displaystyle\-\\frac\{f\(1\)\-f\(1\-\\tilde\{\\delta\}\_\{\\alpha\}\)\}\{\\tilde\{\\delta\}\_\{\\alpha\}\}\+\\frac\{f\\left\(1\+\\delta\_\{\\alpha\}\\right\)\-f\(1\)\}\{\\delta\_\{\\alpha\}\}whereδ~α=\.δα\(1−qβ\)/qβ\\tilde\{\\delta\}\_\{\\alpha\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\delta\_\{\\alpha\}\(1\-q\_\{\\beta\}\)/q\_\{\\beta\}\. Now we pass to the limit:
\(Df\)r′\(Pacc\(SD\)\)\\displaystyle\(D\_\{f\}\)^\{\\prime\}\_\{r\}\(P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}limP′↘Pacc\(SD\)Df\(P′\)−Df\(Pacc\(SD\)\)P′−Pacc\(SD\)\\displaystyle\\lim\_\{P^\{\\prime\}\\searrow P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)\}\\frac\{D\_\{f\}\(P^\{\\prime\}\)\-D\_\{f\}\(P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)\)\}\{P^\{\\prime\}\-P\_\{\\hskip\-2\.84544pt\\mbox\{\\tiny acc\}\}\(SD\)\}\(201\)=\\displaystyle=−limδ~α↘0f\(1\)−f\(1−δ~α\)δ~α\+limδα↘0f\(1\+δα\)−f\(1\)δα\\displaystyle\-\\lim\_\{\\tilde\{\\delta\}\_\{\\alpha\}\\searrow 0\}\\frac\{f\(1\)\-f\(1\-\\tilde\{\\delta\}\_\{\\alpha\}\)\}\{\\tilde\{\\delta\}\_\{\\alpha\}\}\+\\lim\_\{\\delta\_\{\\alpha\}\\searrow 0\}\\frac\{f\\left\(1\+\\delta\_\{\\alpha\}\\right\)\-f\(1\)\}\{\\delta\_\{\\alpha\}\}=\\displaystyle=−fl′\(1\)\+fr′\(1\)\\displaystyle\-f^\{\\prime\}\_\{l\}\(1\)\+f^\{\\prime\}\_\{r\}\(1\)wherefl′f^\{\\prime\}\_\{l\}is the left derivative andfr′f^\{\\prime\}\_\{r\}the right derivative, that must exist becauseffis convex over an open set so differentiable anywhere except maybe on a set of measure zero\([Rockafellar, 1970](https://arxiv.org/html/2609.30474#bib.bib24), Theorem 25\.5\), and any point of non\-differentiabilityzzhas a subdifferential which is exactly\[fl′\(z\),fr′\(z\)\]\[f^\{\\prime\}\_\{l\}\(z\),f^\{\\prime\}\_\{r\}\(z\)\], yielding for \([201](https://arxiv.org/html/2609.30474#S8.E201)\)−fl′\(1\)\+fr′\(1\)=max∂f\(1\)−min∂f\(1\)\-f^\{\\prime\}\_\{l\}\(1\)\+f^\{\\prime\}\_\{r\}\(1\)=\\max\\partial f\(1\)\-\\min\\partial f\(1\), as claimed\.
### VIII\.15Proof of Lemma[6\.1](https://arxiv.org/html/2609.30474#S6.Thmtheorem1)
Without loss of generality, we assume masking𝒒\\bm\{q\}is accompanied by a renormalization of the top\-kkcoordinates for the sake of the proof\. We immediately remark that iflimz→\+∞f\(z\)/z=\+∞\\lim\_\{z\\rightarrow\+\\infty\}f\(z\)/z=\+\\infty, any solution whereqi=0q\_\{i\}=0andπi\>0\\pi\_\{i\}\>0enforcesDf\(𝝅∥𝒒\)=\+∞\>DD\_\{f\}\(\\bm\{\\pi\}\\\|\\bm\{q\}\)=\+\\infty\>Dand is thus not feasible, so the Lemma is proven\. Supposelimz→\+∞f\(z\)/z=ℓ≪\+∞\\lim\_\{z\\rightarrow\+\\infty\}f\(z\)/z=\\ell\\ll\+\\infty\. DenoteΠ=\.∑i∈ℐπi\\Pi\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\sum\_\{i\\in\\mathcal\{I\}\}\\pi\_\{i\}withℐ=\.\{i:qi=0∧πi\>0\}\\mathcal\{I\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\\{i:q\_\{i\}=0\\wedge\\pi\_\{i\}\>0\\\}, so that
Df\(𝝅∥𝒒\)\\displaystyle D\_\{f\}\(\\bm\{\\pi\}\\\|\\bm\{q\}\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}∑i∉ℐqif\(πiqi\)\+ℓ⋅Π\.\\displaystyle\\sum\_\{i\\not\\in\\mathcal\{I\}\}q\_\{i\}f\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)\+\\ell\\cdot\\Pi\.\(202\)𝒒\\bm\{q\}being renormalized over the top\-kkcoordinates, Lemma[3\.6](https://arxiv.org/html/2609.30474#S3.Thmtheorem6)still applies\. We show how the proof of Theorem[3\.7](https://arxiv.org/html/2609.30474#S3.Thmtheorem7)adapts via a simple change of parameter in the computation ofα\\alphaandβ\\betain Lemma[B](https://arxiv.org/html/2609.30474#S8.Thmtheorem2)\. \([88](https://arxiv.org/html/2609.30474#S8.E88)\) now becomes:
∂ℒ1∂ti=λ⋅∑jCj⋅sj−λ⋅Ci−1\+νi\\displaystyle\\frac\{\\partial\\mathcal\{L\}\_\{1\}\}\{\\partial t\_\{i\}\}=\\lambda\\cdot\\sum\_\{j\}C\_\{j\}\\cdot s\_\{j\}\-\\lambda\\cdot C\_\{i\}\-1\+\\nu\_\{i\}=\\displaystyle=0,∀i,\\displaystyle 0,\\forall i,∂ℒ1∂si=μ−λ⋅Ci⋅\(1−𝟏⊤𝒕\)−χi\\displaystyle\\frac\{\\partial\\mathcal\{L\}\_\{1\}\}\{\\partial s\_\{i\}\}=\\mu\-\\lambda\\cdot C\_\{i\}\\cdot\(1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\)\-\\chi\_\{i\}=\\displaystyle=0,∀i,\\displaystyle 0,\\forall i,with
Ci\\displaystyle C\_\{i\}=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\(−f′\)\(πiqi\)⋅⟦i∉ℐ⟧−⟦i∈ℐ⟧,\\displaystyle\(\-f^\{\\prime\}\)\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)\\cdot\\llbracket i\\not\\in\\mathcal\{I\}\\rrbracket\-\\llbracket i\\in\\mathcal\{I\}\\rrbracket,so we just have to replaceα\\alphaandβ\\betaby
α\\displaystyle\\alpha=\\displaystyle=𝔼i∼𝒖\[Ci\],with𝒖=\.11−𝟏⊤𝒕⋅\(𝒑−𝒕\)∈Δn,\\displaystyle\\mathbb\{E\}\_\{i\\sim\\bm\{u\}\}\\left\[C\_\{i\}\\right\],\\quad\\mbox\{ with \}\\bm\{u\}\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\frac\{1\}\{1\-\\bm\{1\}^\{\\top\}\\bm\{t\}\}\\cdot\(\\bm\{p\}\-\\bm\{t\}\)\\in\\Delta\_\{n\},β\\displaystyle\\beta=\\displaystyle=𝔼i∼𝒔\[Ci\],\\displaystyle\\mathbb\{E\}\_\{i\\sim\\bm\{s\}\}\\left\[C\_\{i\}\\right\],and the rest of the proof of Theorem[3\.7](https://arxiv.org/html/2609.30474#S3.Thmtheorem7)follows\. Notice however that
α\\displaystyle\\alpha=\\displaystyle=\(1−u\(ℐ\)\)⋅𝔼i∼𝒖~\[\(−f′\)\(πiqi\)\]−u\(ℐ\),\\displaystyle\(1\-u\(\\mathcal\{I\}\)\)\\cdot\\mathbb\{E\}\_\{i\\sim\\tilde\{\\bm\{u\}\}\}\\left\[\(\-f^\{\\prime\}\)\\left\(\\frac\{\\pi\_\{i\}\}\{q\_\{i\}\}\\right\)\\right\]\-u\(\\mathcal\{I\}\),whereu\(ℐ\)=\.∑i∈ℐuiu\(\\mathcal\{I\}\)\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\sum\_\{i\\in\\mathcal\{I\}\}u\_\{i\}and𝒖~\\tilde\{\\bm\{u\}\}is distribution𝒖\\bm\{u\}masked and renormalized to support in the top\-kkcoordinates\. A similar reformulation holds forβ\\beta\. So instead ofLα\(−g\)L\_\{\\alpha\}\(\-g\)in \([3\.3](https://arxiv.org/html/2609.30474#S3.EGx23)\), we have to considerLα\(−γg−\(1−γ\)\)L\_\{\\alpha\}\(\-\\gamma g\-\(1\-\\gamma\)\)for someγ∈\[0,1\]\\gamma\\in\[0,1\], which slightly extends the possible range of values to search in, and the same happens forLβ\(−h\)L\_\{\\beta\}\(\-h\)\. In the end,clampset\(𝒑,Lβ\(−h\)⋅𝒒,Lα\(−g\)⋅𝒒\)\\mathrm\{clampset\}\(\\bm\{p\},L\_\{\\beta\}\(\-h\)\\cdot\\bm\{q\},L\_\{\\alpha\}\(\-g\)\\cdot\\bm\{q\}\)in \([3\.3](https://arxiv.org/html/2609.30474#S3.EGx23)\) becomesclampset\(𝒑,Lβ\(−γ′h−\(1−γ′\)\)⋅𝒒,Lα\(−γg−\(1−γ\)\)⋅𝒒\)\\mathrm\{clampset\}\(\\bm\{p\},L\_\{\\beta\}\(\-\\gamma^\{\\prime\}h\-\(1\-\\gamma^\{\\prime\}\)\)\\cdot\\bm\{q\},L\_\{\\alpha\}\(\-\\gamma g\-\(1\-\\gamma\)\)\\cdot\\bm\{q\}\)for someγ,γ′∈\[0,1\]\\gamma,\\gamma^\{\\prime\}\\in\[0,1\]\. The set of optimal solutions looks more complicated but keeps the fundamental property that coordinateiiis necessarily00ifi∈ℐi\\in\\mathcal\{I\}, and we do not look for all optimal solutions but just for one whose support coincides with the mask\. Ifffis strictly convex, we still observe that the optimal solution𝝅\\bm\{\\pi\}of \([f \-MD\-1](https://arxiv.org/html/2609.30474#S3.Ex2)\) has the same support as the top\-kkcoordinates of𝒒\\bm\{q\}\. Ifffis not, we easily check thatthere existsoptimal solutions𝝅\\bm\{\\pi\}of \([f \-MD\-1](https://arxiv.org/html/2609.30474#S3.Ex2)\) with the same support as the top\-kkcoordinates of𝒒\\bm\{q\}\. This ends the proof of Lemma[6\.1](https://arxiv.org/html/2609.30474#S6.Thmtheorem1)\.
### VIII\.16Proof of Lemma[6\.2](https://arxiv.org/html/2609.30474#S6.Thmtheorem2)
Fix anyt∈\{2,…,T−1\}t\\in\\\{2,\.\.\.,T\-1\\\}\. Sinceμt\+1=𝔼t\+1\(t\+1\)\\mu\_\{t\+1\}=\\mathbb\{E\}\_\{t\+1\}\(t\+1\), note the dependence betweenμt\+1=𝔼t\+1\(t\+1\)\\mu\_\{t\+1\}=\\mathbb\{E\}\_\{t\+1\}\(t\+1\)andμt=𝔼t\(t−1\)\\mu\_\{t\}=\\mathbb\{E\}\_\{t\}\(t\-1\):
μt\+1\\displaystyle\\mu\_\{t\+1\}=\\displaystyle=𝔼t\(t\+1\)−μt⋅𝔼t\(t,t\+1\)1−μt2\.\\displaystyle\\frac\{\\mathbb\{E\}\_\{t\}\(t\+1\)\-\\mu\_\{t\}\\cdot\\mathbb\{E\}\_\{t\}\(t,t\+1\)\}\{1\-\\mu\_\{t\}^\{2\}\}\.\(203\)We now want a geometric progression on edges:
μt\+1\\displaystyle\\mu\_\{t\+1\}≥\\displaystyle\\geq\(1\+δ\)μt\.\\displaystyle\(1\+\\delta\)\\mu\_\{t\}\.\(204\)This is equivalent, from \([203](https://arxiv.org/html/2609.30474#S8.E203)\), to requesting
μt⋅𝔼t\(t,t\+1\)\\displaystyle\\mu\_\{t\}\\cdot\\mathbb\{E\}\_\{t\}\(t,t\+1\)≤\\displaystyle\\leq𝔼t\(t\+1\)−\(1\+δ\)⋅μt\(1−μt2\),\\displaystyle\\mathbb\{E\}\_\{t\}\(t\+1\)\-\(1\+\\delta\)\\cdot\\mu\_\{t\}\(1\-\\mu\_\{t\}^\{2\}\),and since we assume𝔼t\(t\+1\)≥\(1−β\)μt\\mathbb\{E\}\_\{t\}\(t\+1\)\\geq\(1\-\\beta\)\\mu\_\{t\}\(2\.\), it is sufficient to requestμt⋅𝔼t\(t,t\+1\)≤\(1−β\)μt−\(1\+δ\)⋅μt\(1−μt2\)\\mu\_\{t\}\\cdot\\mathbb\{E\}\_\{t\}\(t,t\+1\)\\leq\(1\-\\beta\)\\mu\_\{t\}\-\(1\+\\delta\)\\cdot\\mu\_\{t\}\(1\-\\mu\_\{t\}^\{2\}\), which after simplification \(μt\>0\\mu\_\{t\}\>0\(1\.\)\) gives the sufficient condition
𝔼t\(t,t\+1\)\\displaystyle\\mathbb\{E\}\_\{t\}\(t,t\+1\)≤\\displaystyle\\leq\(1\+δ\)⋅\(1−μt2\)−\(β\+δ\)\\displaystyle\(1\+\\delta\)\\cdot\(1\-\\mu\_\{t\}^\{2\}\)\-\(\\beta\+\\delta\)=1−β−\(1\+δ\)μt2,\\displaystyle=1\-\\beta\-\(1\+\\delta\)\\mu\_\{t\}^\{2\},which is \(3\.\)\. We then get for the boosting advantage:
A\(\{𝒉t\}t∈\[T\]\)\\displaystyle A\\left\(\\\{\\bm\{h\}\_\{t\}\\\}\_\{t\\in\[T\]\}\\right\)=\\displaystyle=∑t=1Tμt2\\displaystyle\\sum\_\{t=1\}^\{T\}\\mu\_\{t\}^\{2\}≥\\displaystyle\\geqμ12⋅∑t=1T\(1\+δ\)2t\\displaystyle\\mu\_\{1\}^\{2\}\\cdot\\sum\_\{t=1\}^\{T\}\(1\+\\delta\)^\{2t\}=μ12⋅\(1\+δ\)2T−12δ\+δ2,\\displaystyle=\\mu\_\{1\}^\{2\}\\cdot\\frac\{\(1\+\\delta\)^\{2T\}\-1\}\{2\\delta\+\\delta^\{2\}\},as claimed\.
### VIII\.17Proof of Lemma[6\.4](https://arxiv.org/html/2609.30474#S6.Thmtheorem4)
Denote for shorta=\.μta\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\mu\_\{t\}andb=\.μ~tb\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\tilde\{\\mu\}\_\{t\}\. Remark that
B\(c,d\)\\displaystyle B\(c,d\)=\.\\displaystyle\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}μt2\+μt\+12\\displaystyle\\mu^\{2\}\_\{t\}\+\\mu^\{2\}\_\{t\+1\}=\\displaystyle=μt2\+\(𝔼t\(d\)−μt𝔼t\(c,d\)1−μt2\)2\\displaystyle\\mu^\{2\}\_\{t\}\+\\left\(\\frac\{\\mathbb\{E\}\_\{t\}\(d\)\-\\mu\_\{t\}\\mathbb\{E\}\_\{t\}\(c,d\)\}\{1\-\\mu\_\{t\}^\{2\}\}\\right\)^\{2\}=\\displaystyle=μt2\+\(μ~t−μt𝔼t\(c,d\)1−μt2\)2\\displaystyle\\mu^\{2\}\_\{t\}\+\\left\(\\frac\{\\tilde\{\\mu\}\_\{t\}\-\\mu\_\{t\}\\mathbb\{E\}\_\{t\}\(c,d\)\}\{1\-\\mu\_\{t\}^\{2\}\}\\right\)^\{2\}=\\displaystyle=a2\+\(b−a⋅𝔼t\(c,d\)1−a2\)2,\\displaystyle a^\{2\}\+\\left\(\\frac\{b\-a\\cdot\\mathbb\{E\}\_\{t\}\(c,d\)\}\{1\-a^\{2\}\}\\right\)^\{2\},and similarly
B\(d,c\)\\displaystyle B\(d,c\)=\\displaystyle=b2\+\(a−b⋅𝔼t\(c,d\)1−b2\)2\.\\displaystyle b^\{2\}\+\\left\(\\frac\{a\-b\\cdot\\mathbb\{E\}\_\{t\}\(c,d\)\}\{1\-b^\{2\}\}\\right\)^\{2\}\.LetC=\.𝔼t\(c,d\)−abC\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\mathbb\{E\}\_\{t\}\(c,d\)\-abbe the covariance between the sequence of𝒚i⊤𝒉~c\(𝒙i\)\\bm\{y\}\_\{i\}^\{\\top\}\\tilde\{\\bm\{h\}\}\_\{c\}\(\\bm\{x\}\_\{i\}\)and𝒚i⊤𝒉~d\(𝒙i\)\\bm\{y\}\_\{i\}^\{\\top\}\\tilde\{\\bm\{h\}\}\_\{d\}\(\\bm\{x\}\_\{i\}\)computed using𝒘t\\bm\{w\}\_\{t\}\. We remark the simplification:
\(b−a⋅𝔼t\(c,d\)1−a2\)2\\displaystyle\\left\(\\frac\{b\-a\\cdot\\mathbb\{E\}\_\{t\}\(c,d\)\}\{1\-a^\{2\}\}\\right\)^\{2\}=\\displaystyle=\(b−aC−a2b1−a2\)2\\displaystyle\\left\(\\frac\{b\-aC\-a^\{2\}b\}\{1\-a^\{2\}\}\\right\)^\{2\}=\\displaystyle=\(b−aC1−a2\)2\\displaystyle\\left\(b\-\\frac\{aC\}\{1\-a^\{2\}\}\\right\)^\{2\}and similarly\(a−b⋅𝔼t\(c,d\)1−b2\)2=\(a−bC1−b2\)2\\left\(\\frac\{a\-b\\cdot\\mathbb\{E\}\_\{t\}\(c,d\)\}\{1\-b^\{2\}\}\\right\)^\{2\}=\\left\(a\-\\frac\{bC\}\{1\-b^\{2\}\}\\right\)^\{2\}\. We compute the difference and get after factoring
OPENB\(c,d\)−B\(d,c\)\)\\displaystyle B\(c,d\)\-B\(d,c\)\)=\\displaystyle=a2−b2\+\(b−aC1−a2\)2−\(a−bC1−b2\)2\\displaystyle a^\{2\}\-b^\{2\}\+\\left\(b\-\\frac\{aC\}\{1\-a^\{2\}\}\\right\)^\{2\}\-\\left\(a\-\\frac\{bC\}\{1\-b^\{2\}\}\\right\)^\{2\}=\\displaystyle=1\(1−a2b2\)\(1−a2\)2\(1−b2\)2⋅\(a2−b2\)⏟=\.A⋅C⋅\(C−2ab\(1−a2\)\(1−b2\)\(1−a2b2\)\)⏟=\.D\.\\displaystyle\\underbrace\{\\frac\{1\}\{\(1\-a^\{2\}b^\{2\}\)\(1\-a^\{2\}\)^\{2\}\(1\-b^\{2\}\)^\{2\}\}\\cdot\(a^\{2\}\-b^\{2\}\)\}\_\{\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}A\}\\cdot\\underbrace\{C\\cdot\\left\(C\-\\frac\{2ab\(1\-a^\{2\}\)\(1\-b^\{2\}\)\}\{\(1\-a^\{2\}b^\{2\}\)\}\\right\)\}\_\{\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}D\}\.Under the greedy fit scenario,A\>0A\>0so the difference is<0<0iffD<0D<0\. We also remark
0≤ρ=\.2a\|b\|\(1−a2\)\(1−b2\)\(1−a2b2\)≤1,∀a∈\(0,1\],b∈\[−1,1\],\\displaystyle 0\\leq\\rho\\stackrel\{\{\\scriptstyle\\mathrm\{\.\}\}\}\{\{=\}\}\\frac\{2a\|b\|\(1\-a^\{2\}\)\(1\-b^\{2\}\)\}\{\(1\-a^\{2\}b^\{2\}\)\}\\leq 1,\\forall a\\in\(0,1\],b\\in\[\-1,1\],and with a bit more analytical analysis, we get that the upperbound can be replaced by6−42<0\.356\-4\\sqrt\{2\}<0\.35\. We know thata\>0a\>0, so ifb\>0b\>0thenD<0D<0iffC∈\(0,ρ\)C\\in\(0,\\rho\)while ifb<0b<0, thenD<0D<0iffC∈\(−ρ,0\)C\\in\(\-\\rho,0\), as claimed\.相似文章
什么是推测性解码?(在paperswithco.de上热门)[R]
推测性解码是一种推理优化技术,它使用快速草稿模型提出未来 token,并由较大模型并行验证,从而提高 LLM 的生成速度。文章强调了它在 Papers with Code 上的热门状态,以及最近的 SGLang 博客文章,该文章介绍了使用 DFlash 模型实现的最先进延迟。
Mistletoe:针对推测解码的隐蔽加速崩溃攻击
本文识别了基于模型的推测解码在大语言模型中的新漏洞:微小扰动可以在不影响输出质量的情况下降低草稿令牌接受率,从而使加速效果崩溃。作者提出了Mistletoe攻击,该攻击联合优化退化与语义保持,展示了在各种系统上显著的加速降低效果。
跨语言的推测解码
本文比较了三种策略以提高非英语语言的推测解码效率,发现任务特定蒸馏能提高接受率但泛化性差,而n-gram草稿模型尽管接受率较低,却能提供持续的加速。
Speculative Refinement: 一种混合自回归扩散解码策略及其在不同基准测试中的行为表现
介绍了 Speculative Refinement (SpecRef),一种无需训练的混合解码策略,它通过熵引导的选择性掩码,从自回归草稿中热启动掩码扩散语言模型。在六个基准测试上的评估表明,代码基准测试混淆了结构发现与逻辑正确性,识别出了一种精炼张力现象,并显示评估协议可能产生不同的模型排名。
减少草稿,增加检索:用于推测解码的混合树构建
Graft 是一个无需训练的框架,通过结合剪枝与检索来增强推测解码,从而提高接受率和推理速度。在短上下文基准测试中,其加速比最高可达5.41倍,在Qwen3-235B上相比EAGLE-3的提升最高可达21.8%。