理解离策略与同策略蒸馏:两类不同训练目标的故事

arXiv cs.LG 论文

摘要

这篇 UCLA 论文从理论上分析了同策略蒸馏(OPD)与离策略/SFT 式蒸馏的区别,展示了前向 KL 与反向 KL 目标如何分别将多个教师模型聚合为算术混合与几何混合,建立了对数级遗憾界,并解释了 OPD 的优势与脆弱性(例如对低概率教师的敏感性以及错误前缀偏差)。

arXiv:2609.38666v1 Announce Type: new Abstract: On-policy distillation (OPD) learns from teacher feedback on student-generated responses and has shown promise in reducing forgetting relative to supervised fine-tuning (SFT). However, its benefits and fragility remain incompletely understood. We study sequential distillation from multiple teachers, where the student minimizes its average divergence from the teachers. Forward Kullback--Leibler (KL) divergence yields a weighted arithmetic mixture, while reverse KL yields a normalized weighted geometric aggregate. We develop algorithms that learn these targets under off-policy and on-policy feedback, respectively, establishing logarithmic regret bounds in the tabular setting and extending the analysis to function approximation. By analyzing these aggregation targets, we identify mechanisms that help explain both the benefits and fragility of OPD. Relative to forward KL, reverse KL can better retain a confident expert's preferences under uninformative feedback, but is more sensitive to teachers that assign very low probabilities to correct responses. Its token-level conditionals also reveal a dependence on continuation distributions that can favor incorrect prefixes over long horizons.
查看原文
查看缓存全文

缓存时间: 2026/10/02 09:51

# A Tale of Distinct Training Objectives
Source: [https://arxiv.org/html/2609.38666](https://arxiv.org/html/2609.38666)
## Understanding Off\- vs On\-Policy Distillation: A Tale of Distinct Training ObjectivesThanks:Department of Computer Science, University of California, Los Angeles, CA 90095, USA; e\-mail:hyzhao@cs\.ucla\.eduThanks:Department of Computer Science, University of California, Los Angeles, CA 90095, USA; e\-mail:qgu@cs\.ucla\.edu

Qiwei Di Xuheng Li Kaixuan Ji Chenggong Zhang††thanks:Department of Computer Science, University of California, Los Angeles, CA 90095, USA; e\-mail:qiwei2000@cs\.ucla\.edu††thanks:Department of Computer Science, University of California, Los Angeles, CA 90095, USA; e\-mail:xuheng\.li@cs\.ucla\.edu††thanks:Department of Computer Science, University of California, Los Angeles, CA 90095, USA; e\-mail:kaixuanji@cs\.ucla\.edu††thanks:Department of Computer Science, University of California, Los Angeles, CA 90095, USA; e\-mail:chenggong61@g\.ucla\.edu

###### Abstract

On\-policy distillation \(OPD\) learns from teacher feedback on student\-generated responses and has shown promise in reducing forgetting relative to supervised fine\-tuning \(SFT\)\. However, its benefits and fragility remain incompletely understood\. We study sequential distillation from multiple teachers, where the student minimizes its average divergence from the teachers\. Forward Kullback–Leibler \(KL\) divergence yields a weighted arithmetic mixture, while reverse KL yields a normalized weighted geometric aggregate\. We develop algorithms that learn these targets under off\-policy and on\-policy feedback, respectively, establishing logarithmic regret bounds in the tabular setting and extending the analysis to function approximation\. By analyzing these aggregation targets, we identify mechanisms that help explain both the benefits and fragility of OPD\. Relative to forward KL, reverse KL can better retain a confident expert’s preferences under uninformative feedback, but is more sensitive to teachers that assign very low probabilities to correct responses\. Its token\-level conditionals also reveal a dependence on continuation distributions that can favor incorrect prefixes over long horizons\.

## 1Introduction

Knowledge distillation transfers capabilities from teacher models to a student language model\. A conventional approach is supervised fine\-tuning \(SFT\) on teacher\-generated responses, where the student learns to predict tokens along trajectories supplied by the teacher\. On\-policy distillation \(OPD\)\([Agarwal et al\., 2024](https://arxiv.org/html/2609.38666#bib.bib14);[Lu and Lab, 2025](https://arxiv.org/html/2609.38666#bib.bib15)\)instead samples responses from the current student and obtains teacher logits along the resulting trajectories\. This allows the student to learn from teacher feedback on its own behavior, an approach that has been adopted in the post\-training of large language models\([Yang et al\., 2025](https://arxiv.org/html/2609.38666#bib.bib19);[Xiao et al\., 2026](https://arxiv.org/html/2609.38666#bib.bib20);[Zeng et al\., 2026](https://arxiv.org/html/2609.38666#bib.bib21)\)\. Related studies show that on\-policy reinforcement learning can reduce forgetting relative to SFT\([Shenfeld et al\., 2026b](https://arxiv.org/html/2609.38666#bib.bib16);[Chen et al\., 2025](https://arxiv.org/html/2609.38666#bib.bib17)\)and, in some settings, recover capabilities lost during SFT\([Jin et al\., 2025](https://arxiv.org/html/2609.38666#bib.bib18)\)\.

Despite these advantages, OPD can also be fragile\. It may fail under teacher–student mismatch and suffer from training instability over long generation horizons\([Li et al\., 2026](https://arxiv.org/html/2609.38666#bib.bib10)\)\. However, a theoretical understanding of when OPD offers advantages over off\-policy methods such as SFT and what limits those advantages remains to be established\. Recent work has begun to clarify the conditions under which these advantages arise\. For example,[Sriraman et al\. \(2026\)](https://arxiv.org/html/2609.38666#bib.bib22)show that in a noisy\-expert model, matching the clean expert’s return can require exponentially many offline samples in the horizon, whereas an on\-policy algorithm achieves polynomial dependence when the corruption is known\.[Viano et al\. \(2026\)](https://arxiv.org/html/2609.38666#bib.bib23)show that on\-policy value\-based imitation can efficiently match the expert’s return under realizability of its action\-value function, without requiring policy realizability\. Together, these works establish conditions under which on\-policy interaction offers advantages over offline learning\. However, they do not directly explain the fragility observed in practical OPD, leaving a gap in our understanding of when and why it succeeds or fails\.

This naturally raises a question:*What are the pros and cons caused by the distinction between off\-policy and on\-policy feedback in knowledge distillation?*

As a first step toward answering this question, we consider the setting where the student distills sequential data from multiple teachers, and the goal of the student is to minimize its average divergence from all teachers\. Within this framework, we connect off\- and on\-policy distillation to forward\- and reverse\-Kullback–Leibler \(KL\) objectives, respectively\. We then derive closed\-form solutions for both objectives, offering a new perspective on the differences between SFT and OPD through the aggregation targets they induce\. Our main contributions are listed as follows:

- •We propose a sequential distillation framework in which the student minimizes its average divergence from all teachers, and characterize the resulting aggregation targets under forward and reverse KL divergence\. Specifically, forward KL yields a weightedarithmeticmixture of the teacher distributions, while reverse KL yields a normalized weightedgeometricmean\. We also derive the token\-level conditionals of both targets for autoregressive generation\.
- •We develop algorithms that learn the forward\-KL target from off\-policy data collection and the reverse\-KL target from teacher logits on student\-generated responses\. In the tabular setting, we establishO~​\(log⁡T\)\\widetilde\{O\}\(\\log T\)regret bounds for both protocols, with regret defined as the cumulative divergence between the student’s policy at each round and the corresponding aggregation target\. We further extend the logarithmic regret to the setting with general function approximation\. As a result, these guarantees connect off\-policy distillation to the forward\-KL aggregation objective and on\-policy distillation to the reverse\-KL aggregation objective\.
- •We compare the two targets to examine the benefits and fragility of reverse\-KL aggregation relative to forward\-KL aggregation\. First, we consider a setting with one confident expert teacher and several uninformative teachers and show that reverse\-KL aggregation can retain a higher probability of the correct response than forward\-KL aggregation\. We also show that a misleading teacher can suppress correct responses under reverse KL, whereas forward KL preserves the expert’s weighted contribution\. Finally, by examining the token\-level conditionals, we show that over long horizons, reverse\-KL aggregation can favor an incorrect prefix even while preserving the expert’s ranking of complete responses\. These results offer possible explanations for the benefits and fragility of OPD observed in practice\.

Notation\.For a positive integernn, let\[n\]:=\{1,…,n\}\[n\]:=\\\{1,\\ldots,n\\\}\. For a finite set𝒮\\mathcal\{S\}, letΔ⁡\(𝒮\)\\Delta\(\\mathcal\{S\}\)denote the set of probability distributions on𝒮\\mathcal\{S\}\. We write𝟙⁡\(E\)\\ind\(E\)for the indicator of an eventEE, and use𝔼\\mathbb\{E\},ℙ\\mathbb\{P\}, andVar\\Varfor expectation, probability, and variance, respectively\. For distributionsp,q∈Δ⁡\(𝒮\)p,q\\in\\Delta\(\\mathcal\{S\}\), the Kullback–Leibler divergence isKL\(p∥q\):=∑z∈𝒮p\(z\)log\[p\(z\)/q\(z\)\]\\text\{KL\}\(p\\\|q\):=\\sum\_\{z\\in\\mathcal\{S\}\}p\(z\)\\log\[p\(z\)/q\(z\)\], where terms withp⁡\(z\)=0p\(z\)=0are zero, andKL\(p∥q\)=\+∞\\text\{KL\}\(p\\\|q\)=\+\\inftyifp⁡\(z\)\>0=q⁡\(z\)p\(z\)\>0=q\(z\)for somezz\. We useO⁡\(⋅\)O\(\\cdot\)for asymptotic upper bounds andO~​\(⋅\)\\widetilde\{O\}\(\\cdot\)to hide logarithmic factors exceptlog⁡T\\log T\. The notationa≲ba\\lesssim bmeansa≤C​ba\\leq Cbfor a universal constantC\>0C\>0\.

## 2Related Work

#### Variants of on\-policy distillation\.

On\-policy distillation trains a language model by having it generate responses and using a teacher’s next\-token distributions to supervise those responses\([Agarwal et al\., 2024](https://arxiv.org/html/2609.38666#bib.bib14);[Gu et al\., 2024](https://arxiv.org/html/2609.38666#bib.bib24)\)\. Following the empirical success of OPD, subsequent work has explored improvements to its training objectives, the use of teacher feedback, and the construction of teachers\. To improve the training objective, existing methods introduce skew KL or contrastive losses\([Ko et al\., 2024](https://arxiv.org/html/2609.38666#bib.bib25);[Ko et al\., 2025](https://arxiv.org/html/2609.38666#bib.bib43)\), combine forward and reverse KL using hybrid objectives or teacher–student agreement\([Zhu et al\., 2026](https://arxiv.org/html/2609.38666#bib.bib27);[Jin et al\., 2026](https://arxiv.org/html/2609.38666#bib.bib28);[Xing et al\., 2026](https://arxiv.org/html/2609.38666#bib.bib31)\), or construct intermediate distributions between the teacher and student as training targets\([Jang et al\., 2026](https://arxiv.org/html/2609.38666#bib.bib29);[Xie et al\., 2026](https://arxiv.org/html/2609.38666#bib.bib30)\)\. To make better use of teacher feedback, other methods use response outcomes to weight or calibrate supervision\([Zheng et al\., 2026](https://arxiv.org/html/2609.38666#bib.bib35);[Hou et al\., 2026](https://arxiv.org/html/2609.38666#bib.bib36)\), reduce variance or reshape distillation rewards\([Oh et al\., 2026](https://arxiv.org/html/2609.38666#bib.bib33);[Zhao et al\., 2026a](https://arxiv.org/html/2609.38666#bib.bib34)\), and stabilize training through policy constraints and mixed rollout sources\([Luo et al\., 2026](https://arxiv.org/html/2609.38666#bib.bib32)\)\. Data collection is also adapted through selective teacher intervention during student generation\([Xu et al\., 2025](https://arxiv.org/html/2609.38666#bib.bib40)\)and early termination of student rollouts to reduce generation costs and avoid unreliable teacher feedback\([Ziheng et al\., 2026](https://arxiv.org/html/2609.38666#bib.bib11);[Xin et al\., 2026](https://arxiv.org/html/2609.38666#bib.bib37);[Zhang et al\., 2026](https://arxiv.org/html/2609.38666#bib.bib41)\)\. Beyond these changes to objectives and training, the source of supervision has expanded from a separate teacher model to self\-teachers conditioned on reasoning traces, demonstrations, additional context, or environmental feedback\([Zhao et al\., 2026e](https://arxiv.org/html/2609.38666#bib.bib38);[Shenfeld et al\., 2026a](https://arxiv.org/html/2609.38666#bib.bib3);[Ye et al\., 2026](https://arxiv.org/html/2609.38666#bib.bib39);[Hübotter et al\., 2026](https://arxiv.org/html/2609.38666#bib.bib6);[Liu et al\., 2026](https://arxiv.org/html/2609.38666#bib.bib7)\)\. Multiple teachers can also supervise a shared student, whether they are independently trained domain experts\([Ma et al\., 2026a](https://arxiv.org/html/2609.38666#bib.bib8)\)or constructed through task\-specific soft prompts\([Ma et al\., 2026b](https://arxiv.org/html/2609.38666#bib.bib4)\)\. These extensions motivate studying how distillation combines knowledge from different teachers\. We refer readers to[Song and Zheng \(2026\)](https://arxiv.org/html/2609.38666#bib.bib44)for a comprehensive survey of OPD methods\.

#### Understanding on\-policy distillation\.

Recent studies examine both the advantages and the failure modes of on\-policy distillation\. For capability retention during RL fine\-tuning,[Shenfeld et al\. \(2026b\)](https://arxiv.org/html/2609.38666#bib.bib16)relate reduced forgetting to an implicit bias toward solutions close to the initial policy in KL divergence, while[Chen et al\. \(2025\)](https://arxiv.org/html/2609.38666#bib.bib17)emphasize on\-policy data and mode\-seeking behavior\. Empirical studies also show that stronger teachers and denser supervision do not necessarily yield better distillation\.[Li et al\. \(2026\)](https://arxiv.org/html/2609.38666#bib.bib10)find that OPD gains depend on compatible reasoning patterns and the teacher offering capabilities beyond those already acquired by the student\. They further observe that teacher guidance becomes less effective on longer student\-generated prefixes, with training instability emerging at later tokens and spreading to earlier positions\.[Wang et al\. \(2026\)](https://arxiv.org/html/2609.38666#bib.bib9)identify settings where teacher–student mismatch causes incorrect responses to receive higher average token\-level rewards than correct ones\.[Fu et al\. \(2026\)](https://arxiv.org/html/2609.38666#bib.bib42)identify three sources of instability in sampled\-token OPD: predominantly negative token\-level rewards that make learning sensitive to a small set of positively rewarded tokens; unreliable teacher feedback that can reinforce repetitive or meaningless continuations on student\-generated prefixes; and tokenizer or special\-token mismatches that penalize semantically valid outputs\. The theoretical understanding of OPD is relatively limited\.[Sriraman et al\. \(2026\)](https://arxiv.org/html/2609.38666#bib.bib22)establish an exponential separation in horizon dependence between offline learning and an on\-policy algorithm with known expert corruption, and[Viano et al\. \(2026\)](https://arxiv.org/html/2609.38666#bib.bib23)obtain efficient imitation under expert action\-value realizability without policy realizability\. Our contribution is to characterize their distinct optimal targets when aggregating multiple teachers and connect those targets to learning guarantees\. We then analyze how teacher confidence, misleading feedback, and autoregressive continuations affect the expert’s preferences at these targets, providing possible mechanisms for both capability retention and fragility\.

#### Relevant RL theory\.

To turn our perspective on distinct aggregation objectives into formal learning guarantees, we rely on the established RL theory, including supervised learning, KL\-regularized bandits, and reinforcement learning with function approximation\. Our analysis of the off\-policy distillation protocol builds on the classical theory of behavior cloning, which aims to learn a policy from expert demonstrations\. Classical analyses characterize how prediction errors compound under distribution shift\([Syed and Schapire, 2010](https://arxiv.org/html/2609.38666#bib.bib45);[Ross and Bagnell, 2010](https://arxiv.org/html/2609.38666#bib.bib46)\)\. DAgger addresses this issue by collecting expert feedback on learner\-visited states, reducing imitation learning to no\-regret online learning\([Ross et al\., 2011](https://arxiv.org/html/2609.38666#bib.bib47)\)\. Subsequent work further characterizes the statistical limits of imitation learning and the benefits of interaction\([Rajaraman et al\., 2020](https://arxiv.org/html/2609.38666#bib.bib48);[Rajaraman et al\., 2021](https://arxiv.org/html/2609.38666#bib.bib49)\)\. More recently,[Foster et al\. \(2024\)](https://arxiv.org/html/2609.38666#bib.bib2)establish guarantees for behavior cloning with logarithmic loss under realizability, highlighting the importance of the loss function and policy\-class complexity\. Our forward\-KL guarantees draw on log\-loss prediction methods, combining smoothed probability estimation in the tabular setting and exponentially weighted policy aggregation under function approximation with a new one\-sided bounded martingale concentration inequality\.

For on\-policy distillation, our analysis draws on prior work establishing logarithmic regret for KL\-regularized bandits and reinforcement learning\.[Zhao et al\. \(2026b\)](https://arxiv.org/html/2609.38666#bib.bib50)establishO⁡\(1/ϵ\)O\(1/\\epsilon\)sample complexity in a sufficiently small\-error regime under reference\-policy coverage, and later[Zhao et al\. \(2025\)](https://arxiv.org/html/2609.38666#bib.bib51)extend the results to obtain logarithmic regret in online contextual bandits and RL\. For multi\-armed bandits,[Ji et al\. \(2026b\)](https://arxiv.org/html/2609.38666#bib.bib52)sharpen the dependence on the number of arms and regularization strength through nearly matching regret upper and lower bounds\. Similar fast\-rate guarantees have also been established in the offline setting for multi\-armed bandits\([Ji et al\., 2026a](https://arxiv.org/html/2609.38666#bib.bib57)\), or with sufficient coverage conditions under KL and more general divergences\([Zhao et al\., 2026d](https://arxiv.org/html/2609.38666#bib.bib53);[Zhao et al\., 2026c](https://arxiv.org/html/2609.38666#bib.bib58)\)\. This line of work has since been extended to multiple reference models\([Aminian et al\., 2026](https://arxiv.org/html/2609.38666#bib.bib59)\), general preference models\([Wu et al\., 2026](https://arxiv.org/html/2609.38666#bib.bib60)\), game\-theoretic settings\([Nayak et al\., 2025](https://arxiv.org/html/2609.38666#bib.bib61)\), differential privacy\([Wu et al\., 2025](https://arxiv.org/html/2609.38666#bib.bib62)\), and misspecified models\([Hong et al\., 2026](https://arxiv.org/html/2609.38666#bib.bib54)\)\. We adapt these techniques to sequential distillation by treating teacher log\-probability ratios as rewards and accounting for dependence among rollouts from the same sampled teacher\. This yields logarithmic regret guarantees for learning the reverse\-KL aggregation target from on\-policy teacher feedback\.

While we have established the connection between off\- and on\-policy distillation and their respective aggregation targets in the tabular setting, we seek to extend this connection to more general settings, drawing inspiration from prior work on general function approximation\. In contextual bandits,[Russo and Van Roy \(2013\)](https://arxiv.org/html/2609.38666#bib.bib26)introduce eluder dimension to quantify how past observations constrain predictions at unobserved actions\. Subsequent work extends eluder\-based analyses to model estimation through value\-targeted regression\([Ayoub et al\., 2020](https://arxiv.org/html/2609.38666#bib.bib55)\)and general value\-function approximation\([Wang et al\., 2020](https://arxiv.org/html/2609.38666#bib.bib56)\)\. We use a generalized notion of eluder dimension developed in later work\([Agarwal et al\., 2023](https://arxiv.org/html/2609.38666#bib.bib63);[Zhao et al\., 2024](https://arxiv.org/html/2609.38666#bib.bib64);[Di et al\., 2024](https://arxiv.org/html/2609.38666#bib.bib65)\)and adapt it to the autoregressive setting\.

## 3Preliminaries

We study sequential distillation, in which a single student learns from multiple teachers over successive training rounds with the goal of aggregating their knowledge\.

Let𝒳\\mathcal\{X\}be a finite context space of sizeS:=\|𝒳\|S:=\|\\mathcal\{X\}\|,𝒜\\mathcal\{A\}a finite token space of sizeA:=\|𝒜\|A:=\|\\mathcal\{A\}\|, andHHthe generation horizon\. We write𝒴:=𝒜H\\mathcal\{Y\}:=\\mathcal\{A\}^\{H\}for the response space\. Each responsey=\(a1,…,aH\)∈𝒴y=\(a\_\{1\},\\ldots,a\_\{H\}\)\\in\\mathcal\{Y\}is generated autoregressively, and we denote its length\-\(h−1\)\(h\-1\)prefix byy<h:=\(a1,…,ah−1\)y\_\{<h\}:=\(a\_\{1\},\\ldots,a\_\{h\-1\}\)\. We consider a sequential distillation setting with a finite collection of teacher models indexed byℐ\\mathcal\{I\}, where\|ℐ\|=I\|\\mathcal\{I\}\|=I\. For eachi∈ℐi\\in\\mathcal\{I\}, letρi∈Δ⁡\(𝒳\)\\rho\_\{i\}\\in\\Delta\(\\mathcal\{X\}\)be its context distribution and letpip\_\{i\}be an autoregressive teacher policy such that for anyh∈\[H\]h\\in\[H\], the token conditionals are written as

pi\(⋅∣x,y<h\)∈Δ\(𝒜\)\.p\_\{i\}\(\\cdot\\mid x,y\_\{<h\}\)\\in\\Delta\(\\mathcal\{A\}\)\.The induced sequence\-level policy is

pi​\(y∣x\)=∏h=1Hpi​\(ah∣x,y<h\)\.p\_\{i\}\(y\\mid x\)=\\prod\_\{h=1\}^\{H\}p\_\{i\}\(a\_\{h\}\\mid x,y\_\{<h\}\)\.At roundtt, the student policyπt\\pi\_\{t\}is determined by the observations from previous rounds\. A latent teacher indexiti\_\{t\}is then drawn uniformly fromℐ\\mathcal\{I\}, independently of previous rounds\. Conditional oniti\_\{t\}, we drawmmcontexts independently fromρit\\rho\_\{i\_\{t\}\}\. The entire batch shares this teacher: teacher responses and feedback are generated according topitp\_\{i\_\{t\}\}\. We consider two distillation protocols:

- •\(Off\-policy Distillation\) We samplemmcontexts\{xt,j\}∼ρit\\\{x\_\{t,j\}\\\}\\sim\\rho\_\{i\_\{t\}\}and full teacher rolloutsyt,j∼pit\(⋅\|xt,j\)y\_\{t,j\}\\sim p\_\{i\_\{t\}\}\(\\cdot\|x\_\{t,j\}\)\. If we further have access to the teacher logits, at each teacher\-generated prefixyt,j,<hy\_\{t,j,<h\}, we additionally observe the full logit vectorZt,j,h∈ℝ𝒜Z\_\{t,j,h\}\\in\\mathbb\{R\}^\{\\mathcal\{A\}\}\. The corresponding token log\-probabilities are obtained by applying log\-softmax: log⁡pit​\(a∣xt,j,yt,j,<h\)\\displaystyle\\log p\_\{i\_\{t\}\}\(a\\mid x\_\{t,j\},y\_\{t,j,<h\}\)=Zt,j,h\(a\)−log∑b∈𝒜eZt,j,h​\(b\),a∈𝒜\.\\displaystyle=Z\_\{t,j,h\}\(a\)\-\\log\\sum\_\{b\\in\\mathcal\{A\}\}e^\{Z\_\{t,j,h\}\(b\)\},\\qquad a\\in\\mathcal\{A\}\.
- •\(On\-policy Distillation\) We samplemmcontexts\{xt,j\}∼ρit\\\{x\_\{t,j\}\\\}\\sim\\rho\_\{i\_\{t\}\}, use the current student policyπt\\pi\_\{t\}to roll out withyt,j∼πt\(⋅\|xt,j\)y\_\{t,j\}\\sim\\pi\_\{t\}\(\\cdot\|x\_\{t,j\}\)token by token, and receive the per\-token teacher log\-probabilities lt,j,h=log⁡pit​\(at,j,h\|xt,j,yt,j,<h\),h∈\[H\]\.\\displaystyle l\_\{t,j,h\}=\\log p\_\{i\_\{t\}\}\(a\_\{t,j,h\}\\,\|\\,x\_\{t,j\},\\,y\_\{t,j,<h\}\),\\qquad h\\in\[H\]\.

Rather than matching any single teacher selected in the current round, our goal is to learn a single autoregressive student policy that aggregates the behavior of all teacher models\. This objective is designed to avoid catastrophic forgetting: sequentially adapting the student to the current teacher may reduce its loss on the corresponding task while degrading its performance on teachers encountered previously\. To formalize the desired aggregate policy, let

D:Δ⁡\(𝒴\)×Δ⁡\(𝒴\)→ℝ≥0∪\{\+∞\}D:\\Delta\(\\mathcal\{Y\}\)\\times\\Delta\(\\mathcal\{Y\}\)\\rightarrow\\mathbb\{R\}\_\{\\geq 0\}\\cup\\\{\+\\infty\\\}be a divergence between distributions\. Given the collection of teacher models\{\(ρi,pi\)\}i∈ℐ\\\{\(\\rho\_\{i\},p\_\{i\}\)\\\}\_\{i\\in\\mathcal\{I\}\}, we define the target policy by

π∗∈argminπ:𝒳→Δ⁡\(𝒴\)∑i∈ℐ𝔼x∼ρi\[D\(π\(⋅∣x\)∥pi\(⋅∣x\)\)\]\.\\displaystyle\\pi^\{\*\}\\in\\mathop\{\\mathrm\{argmin\}\}\_\{\\pi:\\mathcal\{X\}\\rightarrow\\Delta\(\\mathcal\{Y\}\)\}\\sum\_\{i\\in\\mathcal\{I\}\}\\mathbb\{E\}\_\{x\\sim\\rho\_\{i\}\}\\big\[D\\big\(\\pi\(\\cdot\\mid x\)\\,\\big\\\|\\,p\_\{i\}\(\\cdot\\mid x\)\\big\)\\big\]\.\(3\.1\)Here,π\(⋅∣x\)\\pi\(\\cdot\\mid x\)andpi\(⋅∣x\)p\_\{i\}\(\\cdot\\mid x\)denote the induced distributions over complete responses in𝒴=𝒜H\\mathcal\{Y\}=\\mathcal\{A\}^\{H\}\.

Specifically, whenDDis the reverse KL or forward KL divergence, the objective can be decomposed into a sum of token\-level divergences using the chain rule\. In this work, we will mainly focus on the cases of forward KL\(D\(p∥q\)=KL\(q∥p\)\)\(D\(p\\\|q\)=\\text\{KL\}\(q\\\|p\)\)and reverse KL\(D\(p∥q\)=KL\(p∥q\)\)\(D\(p\\\|q\)=\\text\{KL\}\(p\\\|q\)\)\.

#### Regret\.

Since teacher indices are sampled uniformly, the marginal context distribution is the uniform mixture of the teacher context distributions:

ρ¯​\(x\):=1I​∑i∈ℐρi​\(x\)\.\\displaystyle\\bar\{\\rho\}\(x\):=\\frac\{1\}\{I\}\\sum\_\{i\\in\\mathcal\{I\}\}\\rho\_\{i\}\(x\)\.For each choice ofDD, letπ∗\\pi^\{\*\}denote the corresponding minimizer of \([3\.1](https://arxiv.org/html/2609.38666#S3.E1)\)\. We define regret as the cumulative divergence of the student policies\{πt\}t=1T\\\{\\pi\_\{t\}\\\}\_\{t=1\}^\{T\}from this target, averaged over the marginal context distribution:

Regret\(T\):=∑t=1T𝔼x∼ρ¯\[D\(πt\(⋅∣x\)∥π∗\(⋅∣x\)\)\]\.\\displaystyle\\operatorname\{Regret\}\(T\):=\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\_\{x\\sim\\bar\{\\rho\}\}\\left\[D\\bigl\(\\pi\_\{t\}\(\\cdot\\mid x\)\\,\\big\\\|\\,\\pi^\{\*\}\(\\cdot\\mid x\)\\bigr\)\\right\]\.\(3\.2\)SFT has a well\-known forward\-KL interpretation, while reverse KL is commonly used in OPD\. Motivated by these connections, we takeD\(p∥q\)=KL\(q∥p\)D\(p\\\|q\)=\\text\{KL\}\(q\\\|p\)for off\-policy distillation andD\(p∥q\)=KL\(p∥q\)D\(p\\\|q\)=\\text\{KL\}\(p\\\|q\)for on\-policy distillation\.

## 4Off\-policy Distillation & Forward KL

In this section, we study how off\-policy feedback can be used to learn the forward\-KL aggregate\. We first characterize this target, then construct a tabular algorithm using teacher\-generated responses, with or without full teacher logits, and establish logarithmic regret bounds\. As the first step, the following theorem gives the target distribution and its token conditionals\.

###### Theorem 4\.1\(Forward KL\)\.

IfD\(p∥q\)=KL\(q∥p\)D\(p\\\|q\)=\\text\{KL\}\(q\\\|p\), a minimizer of \([3\.1](https://arxiv.org/html/2609.38666#S3.E1)\) is given at the sequence level, for eachxxwithρ¯​\(x\)\>0\\bar\{\\rho\}\(x\)\>0, by

πF∗​\(y\|x\)=∑i∈ℐwi​\(x\)​pi​\(y\|x\),\\displaystyle\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(y\|x\)=\\sum\_\{i\\in\\mathcal\{I\}\}w\_\{i\}\(x\)p\_\{i\}\(y\|x\),\(4\.1\)wherewi​\(x\)=ρi​\(x\)/\[∑j∈ℐρj​\(x\)\]w\_\{i\}\(x\)=\\rho\_\{i\}\(x\)/\[\\sum\_\{j\\in\\mathcal\{I\}\}\\rho\_\{j\}\(x\)\]\. Moreover, for anyh∈\[H\]h\\in\[H\], supposing∑j∈ℐρj​\(x\)​pj​\(y<h∣x\)\>0\\sum\_\{j\\in\\mathcal\{I\}\}\\rho\_\{j\}\(x\)p\_\{j\}\(y\_\{<h\}\\mid x\)\>0, the token\-level distributions are given by

πF∗​\(a\|x,y<h\)=∑i∈ℐwi​\(x,y<h\)​pi​\(a\|x,y<h\),wi​\(x,y<h\):=ρi​\(x\)​pi​\(y<h\|x\)∑j∈ℐρj​\(x\)​pj​\(y<h\|x\)\.\\displaystyle\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,y\_\{<h\}\)=\\sum\_\{i\\in\\mathcal\{I\}\}w\_\{i\}\(x,y\_\{<h\}\)\\,p\_\{i\}\(a\|x,y\_\{<h\}\),\\qquad w\_\{i\}\(x,y\_\{<h\}\):=\\frac\{\\rho\_\{i\}\(x\)\\,p\_\{i\}\(y\_\{<h\}\|x\)\}\{\\sum\_\{j\\in\\mathcal\{I\}\}\\rho\_\{j\}\(x\)\\,p\_\{j\}\(y\_\{<h\}\|x\)\}\.\(4\.2\)

### 4\.1Algorithm Design

The posterior weights above also determine the next\-token distribution in teacher\-generated data\. For a teacher\-generated rollout\(xt,j,yt,j\)\(x\_\{t,j\},y\_\{t,j\}\), letu∈𝒜h−1u\\in\\mathcal\{A\}^\{h\-1\}denote a prefix\. Conditioning on\(xt,j,yt,j,<h\)=\(x,u\)\(x\_\{t,j\},y\_\{t,j,<h\}\)=\(x,u\)gives teacher weightswi​\(x,u\)w\_\{i\}\(x,u\)\. Hence, the next\-token distribution is

ℙ⁡\(at,j,h=a∣xt,j=x,yt,j,<h=u\)\\displaystyle\\mathbb\{P\}\\bigl\(a\_\{t,j,h\}=a\\mid x\_\{t,j\}=x,\\,y\_\{t,j,<h\}=u\\bigr\)=∑i∈ℐwi​\(x,u\)​pi​\(a\|x,u\)=πF∗​\(a\|x,u\)\.\\displaystyle=\\sum\_\{i\\in\\mathcal\{I\}\}w\_\{i\}\(x,u\)p\_\{i\}\(a\|x,u\)=\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\.Thus, at every visited prefix, the observed next token is a sample from the corresponding token conditional of the forward\-KL target, which suggests estimating each target conditional from the token frequencies observed at the corresponding context–prefix pair\. When full logits are available, each sampled token can instead be replaced by the teacher’s conditional probability vector\. To express both updates in a common form, define

qt,j,h​\(a\):=\{𝟙⁡\(at,j,h=a\)w/o​logit,exp⁡\(Zt,j,h​\(a\)\)∑b∈𝒜exp⁡\(Zt,j,h​\(b\)\),w/logit\.\\displaystyle q\_\{t,j,h\}\(a\):=\\begin\{cases\}\\ind\(a\_\{t,j,h\}=a\)&\\mathrm\{w/o\\ logit\},\\\\ \\frac\{\\exp\(Z\_\{t,j,h\}\(a\)\)\}\{\\sum\_\{b\\in\\mathcal\{A\}\}\\exp\(Z\_\{t,j,h\}\(b\)\)\},&\\mathrm\{w/\\ logit\}\.\\end\{cases\}We then define the per\-round normalized counts as

ct​\(x,u,a\):=1m​∑j=1m𝟙⁡\(xt,j=x,yt,j,<h=u\)​qt,j,h​\(a\)\.\\displaystyle c\_\{t\}\(x,u,a\):=\\frac\{1\}\{m\}\\sum\_\{j=1\}^\{m\}\\ind\(x\_\{t,j\}=x,\\,y\_\{t,j,<h\}=u\)\\,q\_\{t,j,h\}\(a\)\.\(4\.3\)Here, the token positionh=\|u\|\+1h=\|u\|\+1is determined by the prefixuuand is hidden from the notationct​\(x,u,a\)c\_\{t\}\(x,u,a\)\. The algorithm accumulates these counts after each round according to

Wt​\(x,u,a\)\\displaystyle W\_\{t\}\(x,u,a\):=∑s=1tcs​\(x,u,a\),Wt​\(x,u\):=∑a∈𝒜Wt​\(x,u,a\),\\displaystyle:=\\sum\_\{s=1\}^\{t\}c\_\{s\}\(x,u,a\),\\ W\_\{t\}\(x,u\):=\\sum\_\{a\\in\\mathcal\{A\}\}W\_\{t\}\(x,u,a\),and outputs the conditional policy

π^t\+1​\(a\|x,u\):=Wt​\(x,u,a\)\+1/2Wt​\(x,u\)\+A/2\.\\displaystyle\\widehat\{\\pi\}\_\{t\+1\}\(a\|x,u\):=\\frac\{W\_\{t\}\(x,u,a\)\+1/2\}\{W\_\{t\}\(x,u\)\+A/2\}\.\(4\.4\)Here we add a\(1/2\)\(1/2\)\-smoothing to each accumulated weight, following the spirit of the classical Krichevsky–Trofimov estimator\([Krichevsky and Trofimov, 1981](https://arxiv.org/html/2609.38666#bib.bib12)\), which prevents zero predicted probabilities and assigns the uniform distribution to unvisited states\.

Algorithm 1Per\-Prefix Forward KL1:Input:Context set

𝒳\\mathcal\{X\}, token set

𝒜\\mathcal\{A\}, horizon

HH\.

2:Initialize

W0​\(x,u,a\)=0W\_\{0\}\(x,u,a\)=0; equivalently,

π^1​\(a\|x,u\)=1/A\\widehat\{\\pi\}\_\{1\}\(a\|x,u\)=1/Aat every active prefix\.

3:for

t=1,…,Tt=1,\\ldots,Tdo

4:Observe the batch of teacher rollouts

\{\(xt,j,yt,j\)\}j=1m\\\{\(x\_\{t,j\},y\_\{t,j\}\)\\\}\_\{j=1\}^\{m\}and, when available, the teacher logits

\{Zt,j,h\}j,h\\\{Z\_\{t,j,h\}\\\}\_\{j,h\}\.

5:For every level\-

hhstate

\(x,u\)\(x,u\)and token

aa, set

ct​\(x,u,a\)c\_\{t\}\(x,u,a\)as in \([4\.3](https://arxiv.org/html/2609.38666#S4.E3)\)\.

6:Update the weight

Wt​\(x,u,a\)=Wt−1​\(x,u,a\)\+ct​\(x,u,a\)W\_\{t\}\(x,u,a\)=W\_\{t\-1\}\(x,u,a\)\+c\_\{t\}\(x,u,a\)and

Wt​\(x,u\)=∑aWt​\(x,u,a\)W\_\{t\}\(x,u\)=\\sum\_\{a\}W\_\{t\}\(x,u,a\)\.

7:Let

π^t\+1​\(a\|x,u\)\\widehat\{\\pi\}\_\{t\+1\}\(a\|x,u\)be as defined in \([4\.4](https://arxiv.org/html/2609.38666#S4.E4)\)\.

8:endfor

9:Output

\{π^t\}t=1T\\\{\\widehat\{\\pi\}\_\{t\}\\\}\_\{t=1\}^\{T\}\.

### 4\.2Theoretical Guarantees

We now bound the cumulative forward\-KL divergence from the targetπF∗\\pi\_\{\\mathrm\{F\}\}^\{\*\}to the policies produced by Algorithm[1](https://arxiv.org/html/2609.38666#alg1)\. The following theorem gives logarithmic regret bounds under both feedback protocols, with expectation and probability taken over the teachers and rollouts sampled during training\.

###### Theorem 4\.3\.

LetT≥2T\\geq 2, and define the regret in \([3\.2](https://arxiv.org/html/2609.38666#S3.E2)\) withD\(p∥q\)=KL\(q∥p\)D\(p\\\|q\)=\\text\{KL\}\(q\\\|p\)\. Under the off\-policy distillation protocol, both with and without access to full teacher logits, Algorithm[1](https://arxiv.org/html/2609.38666#alg1)satisfies

𝔼⁡\[Regret⁡\(T\)\]\\displaystyle\\mathbb\{E\}\\big\[\\operatorname\{Regret\}\(T\)\\big\]≤4​S​AH​log⁡T\.\\displaystyle\\leq 4SA^\{H\}\\log T\.Moreover, for anyδ∈\(0,1\)\\delta\\in\(0,1\), with probability at least1−δ1\-\\delta, Algorithm[1](https://arxiv.org/html/2609.38666#alg1)satisfies

Regret⁡\(T\)≤8​S​AH​log​T\+4​H​\(1\+log⁡\(2​T\+A\)\)​log​1δ\.\\displaystyle\\operatorname\{Regret\}\(T\)\\leq 8SA^\{H\}\\log T\+4H\\bigl\(1\+\\log\(2T\+A\)\\bigr\)\\log\\frac\{1\}\{\\delta\}\.

For fixedSS,AA, andHH, Theorem[4\.3](https://arxiv.org/html/2609.38666#S4.Thmtheorem3)implies that the forward\-KL divergence fromπF∗\\pi\_\{\\mathrm\{F\}\}^\{\*\}toπ^t\\widehat\{\\pi\}\_\{t\}, averaged over contexts and training rounds, decays asO⁡\(log⁡T/T\)O\(\\log T/T\)in expectation and with high probability\. Thus, off\-policy feedback allows the student to learn the forward\-KL aggregate in this average\-divergence sense\. The factorS​AHSA^\{H\}reflects the size of the effective state–action space: at stephh, each context–prefix pair\(x,u\)∈𝒳×𝒜h−1\(x,u\)\\in\\mathcal\{X\}\\times\\mathcal\{A\}^\{h\-1\}acts as a distinct state, yieldingS​AhSA^\{h\}state–action pairs\. The exponential dependence onHHcomes from the growth of the effective state space itself and the bound scales linearly with the total number of state–action pairs across all steps\.

## 5On\-policy Distillation & Reverse KL

In this section, we study how on\-policy feedback can be used to learn the reverse\-KL aggregate\. With the same structure as the last section, we first characterize this target, then construct a tabular algorithm using teacher log\-probabilities evaluated along student\-generated responses\. The following theorem gives the target distribution and its token conditionals\.

###### Theorem 5\.1\(Reverse KL\)\.

Letwi​\(x\)=ρi​\(x\)/\[∑j∈ℐρj​\(x\)\]w\_\{i\}\(x\)=\\rho\_\{i\}\(x\)/\[\\sum\_\{j\\in\\mathcal\{I\}\}\\rho\_\{j\}\(x\)\]\. Define the unnormalized token\-level geometric mean

gh​\(a\|x,y<h\)\\displaystyle g\_\{h\}\(a\|x,y\_\{<h\}\):=∏i∈ℐpi​\(a\|x,y<h\)wi​\(x\)\.\\displaystyle:=\\prod\_\{i\\in\\mathcal\{I\}\}p\_\{i\}\(a\|x,y\_\{<h\}\)^\{w\_\{i\}\(x\)\}\.We define the value function iteratively backward as follows: for anyx∈𝒳x\\in\\mathcal\{X\},y∈𝒜Hy\\in\\mathcal\{A\}^\{H\},

VH\+1​\(x,y\):=1,Vh​\(x,y<h\)\\displaystyle V\_\{H\+1\}\(x,y\):=1,\\qquad V\_\{h\}\(x,y\_\{<h\}\):=∑a∈𝒜gh\(a\|x,y<h\)Vh\+1\(x,\(y<h,a\)\),h=H,…,1\.\\displaystyle:=\\sum\_\{a\\in\\mathcal\{A\}\}g\_\{h\}\(a\|x,y\_\{<h\}\)\\,V\_\{h\+1\}\\big\(x,\(y\_\{<h\},a\)\\big\),\\qquad h=H,\\ldots,1\.Then the sequence\-level reverse\-KL minimizer is unique and is given by

πR∗​\(y\|x\)=∏i∈ℐpi​\(y\|x\)wi​\(x\)/V1​\(x\)\.\\displaystyle\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(y\|x\)=\\prod\_\{i\\in\\mathcal\{I\}\}p\_\{i\}\(y\|x\)^\{w\_\{i\}\(x\)\}/V\_\{1\}\(x\)\.\(5\.1\)Moreover, we assume that at every prefixy<hy\_\{<h\},Vh​\(x,y<h\)\>0V\_\{h\}\(x,y\_\{<h\}\)\>0\. Then, the token\-level distribution is unique, i\.e\.,

πR∗​\(a\|x,y<h\)=gh​\(a\|x,y<h\)⋅Vh\+1​\(x,\(y<h,a\)\)/Vh​\(x,y<h\)\.\\displaystyle\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(a\|x,y\_\{<h\}\)=g\_\{h\}\(a\|x,y\_\{<h\}\)\\cdot V\_\{h\+1\}\\big\(x,\(y\_\{<h\},a\)\\big\)/V\_\{h\}\(x,y\_\{<h\}\)\.\(5\.2\)

### 5\.1Algorithm Design

To control the estimation error of teacher feedback, we assume that teacher token log\-probabilities differ from those of a fixed reference policy by a bounded amount\.

###### Assumption 5\.3\.

LetB\>0B\>0\. There exists a fixed full\-support reference policyπref\\pi\_\{\\text\{ref\}\}such that, for everyi∈ℐi\\in\\mathcal\{I\},h∈\[H\]h\\in\[H\], and\(x,u,a\)∈𝒳×𝒜h−1×𝒜\(x,u,a\)\\in\\mathcal\{X\}\\times\\mathcal\{A\}^\{h\-1\}\\times\\mathcal\{A\},

\|log⁡\[pi​\(a\|x,u\)/πref​\(a\|x,u\)\]\|≤B\.\\displaystyle\\big\|\\log\[\{p\_\{i\}\(a\|x,u\)\}/\{\\pi\_\{\\text\{ref\}\}\(a\|x,u\)\}\]\\big\|\\leq B\.

No such assumption is required for our forward\-KL analysis, since the feedback consists of token indicators or teacher probabilities, both bounded in\[0,1\]\[0,1\]\. Here, we instead estimate teacher log\-probabilities, which can be unbounded below\. Assumption[5\.3](https://arxiv.org/html/2609.38666#S5.Thmtheorem3)imposes a two\-sided bound on the token\-level probability ratios between each teacher and the reference policy\.

Letℱt\\mathcal\{F\}\_\{t\}denote the observations available before roundtt, so thatπt\\pi\_\{t\}isℱt\\mathcal\{F\}\_\{t\}\-measurable\. For a student\-generated rollout\(xt,j,yt,j\)\(x\_\{t,j\},y\_\{t,j\}\), letu∈𝒜h−1u\\in\\mathcal\{A\}^\{h\-1\}denote a prefix\. At any visited context–prefix–token triple\(x,u,a\)\(x,u,a\), Bayes’ rule gives

ℙ⁡\(it=i∣ℱt,xt,j=x,yt,j,<h=u,at,j,h=a\)=ρi​\(x\)​πt​\(u\|x\)​πt​\(a\|x,u\)∑k∈ℐρk​\(x\)​πt​\(u\|x\)​πt​\(a\|x,u\)=wi​\(x\)\.\\displaystyle\\mathbb\{P\}\\bigl\(i\_\{t\}=i\\mid\\mathcal\{F\}\_\{t\},\\,x\_\{t,j\}=x,\\,y\_\{t,j,<h\}=u,\\,a\_\{t,j,h\}=a\\bigr\)=\\frac\{\\rho\_\{i\}\(x\)\\pi\_\{t\}\(u\|x\)\\pi\_\{t\}\(a\|x,u\)\}\{\\sum\_\{k\\in\\mathcal\{I\}\}\\rho\_\{k\}\(x\)\\pi\_\{t\}\(u\|x\)\\pi\_\{t\}\(a\|x,u\)\}=w\_\{i\}\(x\)\.Given the context and past observations, the prefix and token are generated by the student policy, which is shared across teacher indices\. Their probabilities therefore cancel, leaving the teacher weights unchanged\. Recall that the observed feedback islt,j,h=log⁡pit​\(a\|x,u\)l\_\{t,j,h\}=\\log p\_\{i\_\{t\}\}\(a\|x,u\)\. Its conditional mean satisfies

𝔼\[lt,j,h\|ℱt,xt,j=x,yt,j,<h=u,at,j,h=a\]\\displaystyle\\mathbb\{E\}\\big\[l\_\{t,j,h\}\\,\\big\|\\,\\mathcal\{F\}\_\{t\},\\,x\_\{t,j\}=x,\\,y\_\{t,j,<h\}=u,\\,a\_\{t,j,h\}=a\\big\]=∑i∈ℐwi​\(x\)​log⁡pi​\(a\|x,u\)=log⁡gh​\(a\|x,u\)\.\\displaystyle=\\sum\_\{i\\in\\mathcal\{I\}\}w\_\{i\}\(x\)\\log p\_\{i\}\(a\|x,u\)=\\log g\_\{h\}\(a\|x,u\)\.Visits to\(x,u,a\)\(x,u,a\)thus provide unbiased observations of the log\-geometric scorelog⁡gh​\(a\|x,u\)\\log g\_\{h\}\(a\|x,u\)\. Motivated by this, we estimate this score by averaging the feedback collected at the same triple\. Define the cumulative visit countNtN\_\{t\}, the cumulative feedbackStS\_\{t\}, and the empirical meanl¯t\\bar\{l\}\_\{t\}as

Nt​\(x,u,a\)\\displaystyle N\_\{t\}\(x,u,a\):=∑s=1t∑j=1m𝟙⁡\(xs,j=x,ys,j,<h=u,as,j,h=a\),\\displaystyle:=\\sum\_\{s=1\}^\{t\}\\sum\_\{j=1\}^\{m\}\\ind\\bigl\(x\_\{s,j\}=x,\\,y\_\{s,j,<h\}=u,\\,a\_\{s,j,h\}=a\\bigr\),St​\(x,u,a\)\\displaystyle S\_\{t\}\(x,u,a\):=∑s=1t∑j=1mls,j,h​𝟙⁡\(xs,j=x,ys,j,<h=u,as,j,h=a\),\\displaystyle:=\\sum\_\{s=1\}^\{t\}\\sum\_\{j=1\}^\{m\}l\_\{s,j,h\}\\ind\\bigl\(x\_\{s,j\}=x,\\,y\_\{s,j,<h\}=u,\\,a\_\{s,j,h\}=a\\bigr\),l¯t​\(x,u,a\)\\displaystyle\\bar\{l\}\_\{t\}\(x,u,a\):=St​\(x,u,a\)/\(Nt​\(x,u,a\)∨1\)\.\\displaystyle:=\{S\_\{t\}\(x,u,a\)\}/\\big\(\{N\_\{t\}\(x,u,a\)\\vee 1\}\\big\)\.\(5\.3\)To account for uncertainty in these empirical means, we construct optimistic estimates of the log\-geometric scores\. The following lemma bounds the estimation error at visited triples\.

###### Lemma 5\.4\.

Under Assumption[5\.3](https://arxiv.org/html/2609.38666#S5.Thmtheorem3), fix anyδ∈\(0,1\)\\delta\\in\(0,1\)\. With probability at least1−2​δ1\-2\\delta, the following inequality holds for everyt∈\[T\]t\\in\[T\]and every\(x,u,a\)\(x,u,a\)withNt​\(x,u,a\)≥1N\_\{t\}\(x,u,a\)\\geq 1simultaneously,

\|l¯t​\(x,u,a\)−log⁡gh​\(a\|x,u\)\|\\displaystyle\\big\|\\bar\{l\}\_\{t\}\(x,u,a\)\-\\log g\_\{h\}\(a\|x,u\)\\big\|≤βt​\(x,u,a\),\\displaystyle\\leq\\beta\_\{t\}\(x,u,a\),whereβt​\(x,u,a\):=O~​\(B​m​H/\(Nt​\(x,u,a\)∨1\)\+B​m​H/\(Nt​\(x,u,a\)∨1\)\)\\beta\_\{t\}\(x,u,a\):=\\widetilde\{O\}\\big\(B\\sqrt\{mH/\(N\_\{t\}\(x,u,a\)\\vee 1\)\}\+BmH/\(N\_\{t\}\(x,u,a\)\\vee 1\)\\big\)\.

As a result, we can define the optimistic estimate

l^t​\(x,u,a\):=\{l¯t​\(x,u,a\)\+min⁡\{βt​\(x,u,a\),2​B\}Nt​\(x,u,a\)≥1,log⁡πref​\(a\|x,u\)\+BNt​\(x,u,a\)=0\.\\displaystyle\\widehat\{l\}\_\{t\}\(x,u,a\):=\\begin\{cases\}\\bar\{l\}\_\{t\}\(x,u,a\)\+\\min\\big\\\{\\beta\_\{t\}\(x,u,a\),2B\\big\\\}&N\_\{t\}\(x,u,a\)\\geq 1,\\\\ \\log\\pi\_\{\\text\{ref\}\}\(a\|x,u\)\+B&N\_\{t\}\(x,u,a\)=0\.\\end\{cases\}\(5\.4\)At visited triples, Lemma[5\.4](https://arxiv.org/html/2609.38666#S5.Thmtheorem4)ensures thatl^t​\(x,u,a\)≥log⁡gh​\(a\|x,u\)\\widehat\{l\}\_\{t\}\(x,u,a\)\\geq\\log g\_\{h\}\(a\|x,u\)on the confidence event\. At unvisited triples, this inequality follows directly from Assumption[5\.3](https://arxiv.org/html/2609.38666#S5.Thmtheorem3)\. We then substitute these optimistic scores into the backward recursion defining the reverse\-KL target in \([5\.2](https://arxiv.org/html/2609.38666#S5.E2)\):

g^t,h​\(a\|x,u\)\\displaystyle\\widehat\{g\}\_\{t,h\}\(a\|x,u\):=exp⁡\(l^t​\(x,u,a\)\),V^t,H\+1​\(x,y\):=1,\\displaystyle:=\\exp\\bigl\(\\widehat\{l\}\_\{t\}\(x,u,a\)\\bigr\),\\ \\widehat\{V\}\_\{t,H\+1\}\(x,y\):=1,V^t,h​\(x,u\)\\displaystyle\\widehat\{V\}\_\{t,h\}\(x,u\):=∑a∈𝒜g^t,h\(a\|x,u\)V^t,h\+1\(x,\(u,a\)\),h=H,…,1\.\\displaystyle:=\\sum\_\{a\\in\\mathcal\{A\}\}\\widehat\{g\}\_\{t,h\}\(a\|x,u\)\\widehat\{V\}\_\{t,h\+1\}\\bigl\(x,\(u,a\)\\bigr\),\\qquad h=H,\\ldots,1\.Therefore, for anyhh, any contextxxand prefixy<h=uy\_\{<h\}=u, the output policy for the next token is defined as

π^t\+1​\(a\|x,u\):=g^t,h​\(a\|x,u\)​V^t,h\+1​\(x,\(u,a\)\)/V^t,h​\(x,u\)\.\\displaystyle\\widehat\{\\pi\}\_\{t\+1\}\(a\|x,u\):=\{\\widehat\{g\}\_\{t,h\}\(a\|x,u\)\\widehat\{V\}\_\{t,h\+1\}\\bigl\(x,\(u,a\)\\bigr\)\}/\{\\widehat\{V\}\_\{t,h\}\(x,u\)\}\.\(5\.5\)The algorithm is outlined in Algorithm[2](https://arxiv.org/html/2609.38666#alg2)\.

Algorithm 2Per\-Prefix Reverse KL1:Input:Context set

𝒳\\mathcal\{X\}, token set

𝒜\\mathcal\{A\},

HH,

BB,

δ\\delta, and

TT\.

2:Initialize

π^1​\(a\|x,u\)=1/A\\widehat\{\\pi\}\_\{1\}\(a\|x,u\)=1/Afor any

\(x,u,a\)\(x,u,a\)\.

3:for

t=1,…,Tt=1,\\ldots,Tdo

4:Observe the contexts

\{xt,j\}j=1m\\\{x\_\{t,j\}\\\}\_\{j=1\}^\{m\}generated from

ρit\\rho\_\{i\_\{t\}\}\.

5:For each

jj, generate

yt,j∼π^t\(⋅\|xt,j\)y\_\{t,j\}\\sim\\widehat\{\\pi\}\_\{t\}\(\\cdot\|x\_\{t,j\}\)and observe

\{lt,j,h\}h=1H\\\{l\_\{t,j,h\}\\\}\_\{h=1\}^\{H\}\.

6:For each triple

\(x,u,a\)\(x,u,a\), update

Nt​\(x,u,a\)N\_\{t\}\(x,u,a\),

St​\(x,u,a\)S\_\{t\}\(x,u,a\), and

l¯t​\(x,u,a\)\\bar\{l\}\_\{t\}\(x,u,a\)as in \([5\.3](https://arxiv.org/html/2609.38666#S5.E3)\)\.

7:Construct the optimistic estimate

l^t\\widehat\{l\}\_\{t\}according to \([5\.4](https://arxiv.org/html/2609.38666#S5.E4)\)\.

8:Define

π^t\+1\\widehat\{\\pi\}\_\{t\+1\}according to \([5\.5](https://arxiv.org/html/2609.38666#S5.E5)\)\.

9:endfor

10:Output

\{π^t\}t=1T\\\{\\widehat\{\\pi\}\_\{t\}\\\}\_\{t=1\}^\{T\}\.

### 5\.2Theoretical Guarantees

We now bound the regret induced by Algorithm[2](https://arxiv.org/html/2609.38666#alg2)\.

###### Theorem 5\.5\.

LetT≥2T\\geq 2, and define the regret in \([3\.2](https://arxiv.org/html/2609.38666#S3.E2)\) withD\(p∥q\)=KL\(p∥q\)D\(p\\\|q\)=\\text\{KL\}\(p\\\|q\)\. Under the on\-policy distillation protocol and Assumption[5\.3](https://arxiv.org/html/2609.38666#S5.Thmtheorem3), for anyδ∈\(0,1/3\)\\delta\\in\(0,1/3\), with probability at least1−3​δ1\-3\\delta, Algorithm[2](https://arxiv.org/html/2609.38666#alg2)satisfies

Regret⁡\(T\)\\displaystyle\\operatorname\{Regret\}\(T\)≤O~​\(B2​H2​S​AH​log⁡T\)\.\\displaystyle\\leq\\widetilde\{O\}\\big\(B^\{2\}H^\{2\}SA^\{H\}\\log T\\big\)\.

For fixedSS,AA,HH,BB, andδ\\delta, Theorem[5\.5](https://arxiv.org/html/2609.38666#S5.Thmtheorem5)implies that the reverse\-KL divergence fromπ^t\\widehat\{\\pi\}\_\{t\}toπR∗\\pi\_\{\\mathrm\{R\}\}^\{\*\}, averaged over contexts and training rounds, decays asO~​\(log⁡T/T\)\\widetilde\{O\}\(\\log T/T\)with high probability\. Thus, on\-policy feedback allows the student to learn the reverse\-KL aggregate in this average\-divergence sense\. As in the off\-policy setting, the factorS​AHSA^\{H\}reflects the size of the effective state–action space\. For the detailed proof, see Appendix[C](https://arxiv.org/html/2609.38666#A3)\.

## 6Comparison of Distinct Aggregation Objectives

The preceding sections establish learning guarantees that connect off\-policy distillation to forward\-KL aggregation and on\-policy distillation to reverse\-KL aggregation\. We now examine the differences between these two forms of aggregation and their implications for combining the capabilities of multiple teachers in a single policy\.

### 6\.1Uninformative Teachers

Consider a contextxxthat falls within teacherkk’s area of expertise but is unfamiliar to the remaining teachers\. We represent this difference in coverage by assuming thatρk​\(x\)\\rho\_\{k\}\(x\)is large relative to∑i≠kρi​\(x\)\\sum\_\{i\\neq k\}\\rho\_\{i\}\(x\)\. The expert therefore receives a large aggregation weightα:=wk​\(x\)∈\(0,1\)\\alpha:=w\_\{k\}\(x\)\\in\(0,1\)\.

Writep\(⋅\|x\):=pk\(⋅\|x\)p\(\\cdot\|x\):=p\_\{k\}\(\\cdot\|x\)andN:=\|𝒴\|≥2N:=\|\\mathcal\{Y\}\|\\geq 2, and letq:=p⁡\(y⋆\|x\)q:=p\(y^\{\\star\}\|x\)denote the expert’s probability of a desirable responsey⋆y^\{\\star\}\. A value ofqqclose to one means that the expert already produces the correct response with high probability\. We ask whether the aggregate can maintain this high probability when the other teachers are uninformative\. To represent the remaining teachers’ lack of information atxx, we model their response distributions as uniform:pi​\(y\|x\)=1/Np\_\{i\}\(y\|x\)=1/Nfor everyi≠ki\\neq kandy∈𝒴y\\in\\mathcal\{Y\}\. These teachers have total weight1−α1\-\\alphaand are equivalent, under either objective, to a single uniform teacher with this weight\. The two aggregation targets are therefore

πF∗​\(y\|x\)\\displaystyle\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(y\|x\)=α​p​\(y\|x\)\+1−αN,πR∗​\(y\|x\)=p​\(y\|x\)α∑z∈𝒴p​\(z\|x\)α\.\\displaystyle=\\alpha p\(y\|x\)\+\\frac\{1\-\\alpha\}\{N\},\\quad\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(y\|x\)=\\frac\{p\(y\|x\)^\{\\alpha\}\}\{\\sum\_\{z\\in\\mathcal\{Y\}\}p\(z\|x\)^\{\\alpha\}\}\.For a concrete comparison, supposey⋆y^\{\\star\}is the only correct response andq∈\(1/N,1\)q\\in\(1/N,1\), with the expert assigning uniform probability to each incorrect response\. The following proposition shows that, for fixed teacher weights and response space, the comparison is determined by a confidence threshold\.

###### Proposition 6\.1\.

FixN≥2N\\geq 2andα∈\(0,1\)\\alpha\\in\(0,1\), and consider the teacher distributions above\. IfN≥3N\\geq 3, there exists a uniqueqc=qc​\(N,α\)∈\(1/N,1\)q\_\{\\mathrm\{c\}\}=q\_\{\\mathrm\{c\}\}\(N,\\alpha\)\\in\(1/N,1\)such that, for everyq∈\(1/N,1\)q\\in\(1/N,1\),

πR∗\(y⋆\|x\)\>πF∗\(y⋆\|x\)⟺q\>qc\.\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(y^\{\\star\}\|x\)\>\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(y^\{\\star\}\|x\)\\quad\\Longleftrightarrow\\quad q\>q\_\{\\mathrm\{c\}\}\.

The proof is given in Appendix[E](https://arxiv.org/html/2609.38666#A5)\. Figure[1](https://arxiv.org/html/2609.38666#S6.F1)\(a\) plots the two correct\-response probabilities forq\>0\.5q\>0\.5, withN=10N=10andα=0\.9\\alpha=0\.9fixed\.

Figure 1:Forward\- and reverse\-KL aggregation with expert weightα=0\.9\\alpha=0\.9\. Panels \(a,b\) useN=10N=10responses, with each teacher distributing its remaining probability uniformly over incorrect responses\. \(a\) The other teachers are uniform, and expert confidenceqqvaries above0\.50\.5; the targets cross atqc≈0\.7617q\_\{\\mathrm\{c\}\}\\approx 0\.7617\. \(b\) The expert’s probability is fixed atq=0\.99q=0\.99, while the other teacher’s correct\-response probabilityε\\varepsilonvaries on a logarithmic axis\. \(c\) For binary responses, the horizonHHvaries while all token probabilities remain fixed\. The expert chooses the correct first tokenaawith probability0\.990\.99; subsequent tokens have probabilities\(0\.99,0\.01\)\(0\.99,0\.01\)afteraaand\(1/2,1/2\)\(1/2,1/2\)afterbb\. The other teacher is uniform at every prefix\.
### 6\.2Misleading Teachers

The preceding examples model teachers that are unfamiliar with a context as uniform\. In this section, we consider another scenario where teachers may be misleading\. As the small aggregation weights of some teachers reflect their limited coverage of the context, their predictions on such unfamiliar contexts may be highly inaccurate\. Thus, they may assign extremely low probability to the correct response\. We show that forward aggregation preserves a high probability of the correct response when the expert is highly confident and receives most of the weight\. Under reverse aggregation, however, even a teacher with a small weight can drive this probability arbitrarily close to zero by assigning sufficiently low probability to the correct response\.

Figure[1](https://arxiv.org/html/2609.38666#S6.F1)\(b\) illustrates this effect by varying only the misleading teacher’s correct\-response probabilityε\\varepsilon, while fixing the expert’s probability atq=0\.99q=0\.99and the teacher weights at0\.90\.9and0\.10\.1\. Asε\\varepsilondecreases toward zero, the reverse target approaches zero, whereas the forward target remains bounded below by the expert’s weighted contribution\. The following proposition extends this comparison to a set of correct responses\.

###### Proposition 6\.3\.

Let∅≠E⊊𝒴\\varnothing\\neq E\\subsetneq\\mathcal\{Y\}be a set of correct responses\. For any teacherkk, the forward target satisfies

πF∗​\(E\|x\)≥wk​pk​\(E\|x\)\.\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(E\|x\)\\geq w\_\{k\}p\_\{k\}\(E\|x\)\.In contrast, for anyα,q,η∈\(0,1\)\\alpha,q,\\eta\\in\(0,1\), there exist two teachers with weightsw1=αw\_\{1\}=\\alphaandw2=1−αw\_\{2\}=1\-\\alpha, each assigning positive probability to every response, such that

p1​\(E\|x\)=q,πR∗​\(E\|x\)<η\.p\_\{1\}\(E\|x\)=q,\\qquad\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(E\|x\)<\\eta\.

Intuitively, forward aggregation is a weighted sum of nonnegative probabilities, so other teachers cannot cancel the expert’s contribution to correct responses\. Reverse aggregation instead averages log\-probabilities, which can be arbitrarily negative\. Consequently, a teacher that assigns sufficiently low probability to correct responses relative to incorrect ones can override the expert’s preference, even with a small aggregation weight\.

### 6\.3Long Horizons

The preceding comparisons focus on the probabilities of complete responses\. During autoregressive generation, however, reverse aggregation can favor a prefix that the expert considers less likely, even while preserving the expert’s ranking of complete responses\. This can steer generation toward an incorrect trajectory\.

Consider an expert teacherp1p\_\{1\}with weightα∈\(0,1\)\\alpha\\in\(0,1\)and a uniform teacherp2p\_\{2\}with weightβ=1−α\\beta=1\-\\alpha\. Sincep2​\(y\|x\)=A−Hp\_\{2\}\(y\|x\)=A^\{\-H\}is the same for every complete response, its contribution to the geometric aggregate cancels upon normalization, leaving

πR∗​\(y\|x\)=p1​\(y\|x\)α/\[∑z∈𝒴p1​\(z\|x\)α\]\.\\displaystyle\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(y\|x\)=\{p\_\{1\}\(y\|x\)^\{\\alpha\}\}/\\big\[\{\\textstyle\\sum\_\{z\\in\\mathcal\{Y\}\}p\_\{1\}\(z\|x\)^\{\\alpha\}\}\\big\]\.Raising probabilities to the powerα<1\\alpha<1preserves the ranking of complete responses while increasing the relative weight of less likely responses\. Token\-level probabilities, however, also depend on a sum over possible continuations, as shown by \([5\.2](https://arxiv.org/html/2609.38666#S5.E2)\)\. For a prefixuuof lengthh−1h\-1and a candidate tokencc,

πR∗​\(c\|x,u\)∝p1​\(c\|x,u\)α​∑v∈𝒜H−hp1​\(v\|x,u,c\)α\.\\displaystyle\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(c\|x,u\)\\propto p\_\{1\}\(c\|x,u\)^\{\\alpha\}\\sum\_\{v\\in\\mathcal\{A\}^\{H\-h\}\}p\_\{1\}\(v\|x,u,c\)^\{\\alpha\}\.Sinceα<1\\alpha<1, spreading probability mass more evenly across continuations increases this sum\. The continuation factor can therefore favor tokens with more diffuse continuations over those whose continuations are concentrated on a few responses\.

To further illustrate this effect, consider responses of lengthHHover𝒜=\{a,b\}\\mathcal\{A\}=\\\{a,b\\\}, whereaais the correct first token\. The expertp1p\_\{1\}has weightα=0\.9\\alpha=0\.9and selectsaaat the first step with probabilityr=0\.99r=0\.99\. The other teacherp2p\_\{2\}has weightβ=0\.1\\beta=0\.1and assigns probability1/21/2to each token at every prefix\.

At each of the remainingH−1H\-1positions, the expert’s token distribution depends only on the first token\. After an initialaa, it assigns probability0\.990\.99toaaand0\.010\.01tobb, regardless of the intervening tokens\. After an initialbb, it assigns probability1/21/2to each token\. Thus, the expert’s continuations are concentrated afteraaand uniform afterbb, while all token probabilities remain strictly positive\. Under forward aggregation, the probability of choosingaaat the first step is the weighted average of the teachers’ corresponding probabilities and is therefore independent of the continuation distributions and the horizon\. Figure[1](https://arxiv.org/html/2609.38666#S6.F1)\(c\) plots the probability of choosing the correct first tokenaaasHHvaries from11to4040, keeping the teacher weights and token\-level probabilities fixed\.

More generally, the following proposition shows that this reversal can occur even when the uniform teacher receives an arbitrarily small weight\.

###### Proposition 6\.5\.

Fixβ∈\(0,1/2\)\\beta\\in\(0,1/2\)andr∈\(1/2,1\)r\\in\(1/2,1\)\. For every sufficiently large horizonHH, there exist two teacher policiesp1p\_\{1\}andp2p\_\{2\}on𝒴=\{a,b\}H\\mathcal\{Y\}=\\\{a,b\\\}^\{H\}, each assigning positive probability to every response, such thatp1​\(a\|x\)=rp\_\{1\}\(a\|x\)=randp2p\_\{2\}is uniform at every prefix\. With aggregation weightsw1​\(x\)=1−βw\_\{1\}\(x\)=1\-\\betaandw2​\(x\)=βw\_\{2\}\(x\)=\\beta, their forward and reverse targets satisfy

πF∗​\(a\|x\)\>πF∗​\(b\|x\),πR∗​\(a\|x\)<πR∗​\(b\|x\)\.\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x\)\>\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(b\|x\),\\qquad\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(a\|x\)<\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(b\|x\)\.

## 7Conclusion

We studied off\- and on\-policy distillation in a sequential multi\-teacher framework\. By characterizing the forward\- and reverse\-KL targets and establishing learning guarantees, we connected the two feedback protocols to distinct aggregation objectives\. Our comparisons show that reverse\-KL aggregation can better preserve a confident expert’s preference in the presence of uniform teachers, but can suppress correct responses under misleading feedback or favor incorrect prefixes over long horizons\. These results offer possible explanations for the benefits and fragility of OPD relative to SFT, and provide a perspective on the difference between SFT and OPD through the lens of distinct aggregation objectives\.

### AI use statement

This work builds on the authors’ research intuition and ideas\. We used AI models, including GPT 5\.6 Sol, GPT 6 Astra, Claude Fable 5, and Grok 4\.6, to help articulate and summarize these ideas and to draft and polish the manuscript\. These models also assisted in proving several key technical lemmas and developing proof strategies\. The authors verified the proofs and reviewed the AI\-assisted writing, and take full responsibility for the correctness and final content of the paper\.

## Appendix AProof of the KL Aggregation Theorems

In this section, we prove Theorems[4\.1](https://arxiv.org/html/2609.38666#S4.Thmtheorem1)and[5\.1](https://arxiv.org/html/2609.38666#S5.Thmtheorem1)\. We first reduce the objective to a separate optimization problem at each context\. Recall that

ρ¯​\(x\):=1I​∑i∈ℐρi​\(x\),wi​\(x\):=ρi​\(x\)∑j∈ℐρj​\(x\)\\displaystyle\\bar\{\\rho\}\(x\):=\\frac\{1\}\{I\}\\sum\_\{i\\in\\mathcal\{I\}\}\\rho\_\{i\}\(x\),\\qquad w\_\{i\}\(x\):=\\frac\{\\rho\_\{i\}\(x\)\}\{\\sum\_\{j\\in\\mathcal\{I\}\}\\rho\_\{j\}\(x\)\}for everyxxwithρ¯​\(x\)\>0\\bar\{\\rho\}\(x\)\>0\. Define

ℐx\+\\displaystyle\\mathcal\{I\}\_\{x\}^\{\+\}:=\{i∈ℐ:wi\(x\)\>0\},LD,x\(π\):=∑i∈ℐx\+wi\(x\)D\(π\(⋅∣x\)∥pi\(⋅∣x\)\)\.\\displaystyle:=\\\{i\\in\\mathcal\{I\}:w\_\{i\}\(x\)\>0\\\},\\ L\_\{D,x\}\(\\pi\):=\\sum\_\{i\\in\\mathcal\{I\}\_\{x\}^\{\+\}\}w\_\{i\}\(x\)D\\bigl\(\\pi\(\\cdot\\mid x\)\\,\\big\\\|\\,p\_\{i\}\(\\cdot\\mid x\)\\bigr\)\.Usingρi​\(x\)=I​ρ¯​\(x\)​wi​\(x\)\\rho\_\{i\}\(x\)=I\\bar\{\\rho\}\(x\)w\_\{i\}\(x\), we obtain

∑i∈ℐ𝔼x∼ρi\[D\(π\(⋅∣x\)∥pi\(⋅∣x\)\)\]=I∑x:ρ¯​\(x\)\>0ρ¯\(x\)LD,x\(π\)\.\\displaystyle\\sum\_\{i\\in\\mathcal\{I\}\}\\mathbb\{E\}\_\{x\\sim\\rho\_\{i\}\}\\bigl\[D\\bigl\(\\pi\(\\cdot\\mid x\)\\,\\big\\\|\\,p\_\{i\}\(\\cdot\\mid x\)\\bigr\)\\bigr\]=I\\sum\_\{x:\\bar\{\\rho\}\(x\)\>0\}\\bar\{\\rho\}\(x\)L\_\{D,x\}\(\\pi\)\.\(A\.1\)The optimization domain in \([3\.1](https://arxiv.org/html/2609.38666#S3.E1)\) allows a separate distribution inΔ⁡\(𝒴\)\\Delta\(\\mathcal\{Y\}\)at every context\. Therefore, it suffices to minimizeLD,xL\_\{D,x\}at each context withρ¯​\(x\)\>0\\bar\{\\rho\}\(x\)\>0\. Contexts withρ¯​\(x\)=0\\bar\{\\rho\}\(x\)=0do not affect the objective, and their policies may be chosen arbitrarily\. Throughout the proofs, zero\-weight teachers are omitted from geometric products\. We use the conventions0​log⁡\(0/q\)=00\\log\(0/q\)=0forq≥0q\\geq 0andp​log⁡\(p/0\)=\+∞p\\log\(p/0\)=\+\\inftyforp\>0p\>0\.

### A\.1Proof of Theorem[4\.1](https://arxiv.org/html/2609.38666#S4.Thmtheorem1)

###### Proof of Theorem[4\.1](https://arxiv.org/html/2609.38666#S4.Thmtheorem1)\.

Fix a contextxxwithρ¯​\(x\)\>0\\bar\{\\rho\}\(x\)\>0\. For forward KL, the contextwise objective in \([A\.1](https://arxiv.org/html/2609.38666#A1.E1)\) is

LF,x\(π\):=∑i∈ℐx\+wi\(x\)KL\(pi\(⋅∣x\)∥π\(⋅∣x\)\)\.\\displaystyle L\_\{\\mathrm\{F\},x\}\(\\pi\):=\\sum\_\{i\\in\\mathcal\{I\}\_\{x\}^\{\+\}\}w\_\{i\}\(x\)\\text\{KL\}\\bigl\(p\_\{i\}\(\\cdot\\mid x\)\\,\\big\\\|\\,\\pi\(\\cdot\\mid x\)\\bigr\)\.We first identify its minimizing sequence distribution, and then derive the corresponding token conditionals\.

Sequence\-level minimizer\.Define the candidate distribution

qF​\(y∣x\):=∑i∈ℐx\+wi​\(x\)​pi​\(y∣x\),y∈𝒴\.\\displaystyle q\_\{\\mathrm\{F\}\}\(y\\mid x\):=\\sum\_\{i\\in\\mathcal\{I\}\_\{x\}^\{\+\}\}w\_\{i\}\(x\)p\_\{i\}\(y\\mid x\),\\qquad y\\in\\mathcal\{Y\}\.Since the weights are nonnegative and sum to one,qF\(⋅∣x\)∈Δ\(𝒴\)q\_\{\\mathrm\{F\}\}\(\\cdot\\mid x\)\\in\\Delta\(\\mathcal\{Y\}\)\. Let

𝒮F​\(x\)\\displaystyle\\mathcal\{S\}\_\{\\mathrm\{F\}\}\(x\):=\{y∈𝒴:qF​\(y∣x\)\>0\},\\displaystyle:=\\\{y\\in\\mathcal\{Y\}:q\_\{\\mathrm\{F\}\}\(y\\mid x\)\>0\\\},CF​\(x\)\\displaystyle C\_\{\\mathrm\{F\}\}\(x\):=∑i∈ℐx\+wi\(x\)KL\(pi\(⋅∣x\)∥qF\(⋅∣x\)\)\.\\displaystyle:=\\sum\_\{i\\in\\mathcal\{I\}\_\{x\}^\{\+\}\}w\_\{i\}\(x\)\\text\{KL\}\\bigl\(p\_\{i\}\(\\cdot\\mid x\)\\,\\big\\\|\\,q\_\{\\mathrm\{F\}\}\(\\cdot\\mid x\)\\bigr\)\.For everyi∈ℐx\+i\\in\\mathcal\{I\}\_\{x\}^\{\+\}, we haveqF​\(y∣x\)≥wi​\(x\)​pi​\(y∣x\)q\_\{\\mathrm\{F\}\}\(y\\mid x\)\\geq w\_\{i\}\(x\)p\_\{i\}\(y\\mid x\)\. Consequently,CF​\(x\)C\_\{\\mathrm\{F\}\}\(x\)is finite and does not depend onπ\\pi\. For a policy that is positive on𝒮F​\(x\)\\mathcal\{S\}\_\{\\mathrm\{F\}\}\(x\), expanding the KL divergences gives

LF,x​\(π\)−CF​\(x\)\\displaystyle L\_\{\\mathrm\{F\},x\}\(\\pi\)\-C\_\{\\mathrm\{F\}\}\(x\)=∑i∈ℐx\+wi​\(x\)​∑y∈𝒮F​\(x\)pi​\(y∣x\)​log⁡qF​\(y∣x\)π⁡\(y∣x\)\\displaystyle=\\sum\_\{i\\in\\mathcal\{I\}\_\{x\}^\{\+\}\}w\_\{i\}\(x\)\\sum\_\{y\\in\\mathcal\{S\}\_\{\\mathrm\{F\}\}\(x\)\}p\_\{i\}\(y\\mid x\)\\log\\frac\{q\_\{\\mathrm\{F\}\}\(y\\mid x\)\}\{\\pi\(y\\mid x\)\}=∑y∈𝒮F​\(x\)qF​\(y∣x\)​log⁡qF​\(y∣x\)π⁡\(y∣x\)\\displaystyle=\\sum\_\{y\\in\\mathcal\{S\}\_\{\\mathrm\{F\}\}\(x\)\}q\_\{\\mathrm\{F\}\}\(y\\mid x\)\\log\\frac\{q\_\{\\mathrm\{F\}\}\(y\\mid x\)\}\{\\pi\(y\\mid x\)\}=KL\(qF\(⋅∣x\)∥π\(⋅∣x\)\),\\displaystyle=\\text\{KL\}\\bigl\(q\_\{\\mathrm\{F\}\}\(\\cdot\\mid x\)\\,\\big\\\|\\,\\pi\(\\cdot\\mid x\)\\bigr\),\(A\.2\)where the second equality uses the definition ofqFq\_\{\\mathrm\{F\}\}\. Ifπ\\piassigns zero probability to somey∈𝒮F​\(x\)y\\in\\mathcal\{S\}\_\{\\mathrm\{F\}\}\(x\), at least one positive\-weight teacher assigns positive probability to that response\. In that case, both sides of \([A\.2](https://arxiv.org/html/2609.38666#A1.E2)\) are\+∞\+\\infty, so the identity remains valid in the extended\-real sense\.

By nonnegativity of KL divergence, the right\-hand side of \([A\.2](https://arxiv.org/html/2609.38666#A1.E2)\) is minimized at zero, with equality if and only ifπ\(⋅∣x\)=qF\(⋅∣x\)\\pi\(\\cdot\\mid x\)=q\_\{\\mathrm\{F\}\}\(\\cdot\\mid x\)\. Thus, the minimizing sequence distribution is unique at this context and satisfies

πF∗​\(y∣x\)=qF​\(y∣x\)=∑i∈ℐwi​\(x\)​pi​\(y∣x\),\\displaystyle\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(y\\mid x\)=q\_\{\\mathrm\{F\}\}\(y\\mid x\)=\\sum\_\{i\\in\\mathcal\{I\}\}w\_\{i\}\(x\)p\_\{i\}\(y\\mid x\),which proves \([4\.1](https://arxiv.org/html/2609.38666#S4.E1)\)\.

Token\-level conditionals\.Fixh∈\[H\]h\\in\[H\]andu∈𝒜h−1u\\in\\mathcal\{A\}^\{h\-1\}\. Marginalizing the sequence distribution over all continuations gives

πF∗​\(u∣x\)\\displaystyle\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(u\\mid x\)=∑i∈ℐwi​\(x\)​pi​\(u∣x\),\\displaystyle=\\sum\_\{i\\in\\mathcal\{I\}\}w\_\{i\}\(x\)p\_\{i\}\(u\\mid x\),πF∗​\(\(u,a\)∣x\)\\displaystyle\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(\(u,a\)\\mid x\)=∑i∈ℐwi​\(x\)​pi​\(u∣x\)​pi​\(a∣x,u\)\.\\displaystyle=\\sum\_\{i\\in\\mathcal\{I\}\}w\_\{i\}\(x\)p\_\{i\}\(u\\mid x\)p\_\{i\}\(a\\mid x,u\)\.\(A\.3\)The first equality follows from linearity of marginalization; the second additionally uses the autoregressive factorization of each teacher\. Suppose that∑iρi​\(x\)​pi​\(u∣x\)\>0\\sum\_\{i\}\\rho\_\{i\}\(x\)p\_\{i\}\(u\\mid x\)\>0\. ThenπF∗​\(u∣x\)\>0\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(u\\mid x\)\>0, and taking the ratio of the two prefix probabilities yields

πF∗​\(a∣x,u\)\\displaystyle\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\\mid x,u\)=πF∗​\(\(u,a\)∣x\)πF∗​\(u∣x\)\\displaystyle=\\frac\{\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(\(u,a\)\\mid x\)\}\{\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(u\\mid x\)\}=∑i∈ℐwi​\(x\)​pi​\(u∣x\)​pi​\(a∣x,u\)∑j∈ℐwj​\(x\)​pj​\(u∣x\)\\displaystyle=\\frac\{\\sum\_\{i\\in\\mathcal\{I\}\}w\_\{i\}\(x\)p\_\{i\}\(u\\mid x\)p\_\{i\}\(a\\mid x,u\)\}\{\\sum\_\{j\\in\\mathcal\{I\}\}w\_\{j\}\(x\)p\_\{j\}\(u\\mid x\)\}=∑i∈ℐρi​\(x\)​pi​\(u∣x\)∑j∈ℐρj​\(x\)​pj​\(u∣x\)​pi​\(a∣x,u\),\\displaystyle=\\sum\_\{i\\in\\mathcal\{I\}\}\\frac\{\\rho\_\{i\}\(x\)p\_\{i\}\(u\\mid x\)\}\{\\sum\_\{j\\in\\mathcal\{I\}\}\\rho\_\{j\}\(x\)p\_\{j\}\(u\\mid x\)\}p\_\{i\}\(a\\mid x,u\),\(A\.4\)where the last equality substitutes the definition ofwi​\(x\)w\_\{i\}\(x\)and cancels its common normalization factor\. This is exactly \([4\.2](https://arxiv.org/html/2609.38666#S4.E2)\)\. At a prefix withπF∗​\(u∣x\)=0\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(u\\mid x\)=0, the conditional may be chosen arbitrarily without changing the sequence distribution\. Sincexxwas arbitrary, combining the contextwise minimizers using \([A\.1](https://arxiv.org/html/2609.38666#A1.E1)\) completes the proof\. ∎

### A\.2Proof of Theorem[5\.1](https://arxiv.org/html/2609.38666#S5.Thmtheorem1)

###### Proof of Theorem[5\.1](https://arxiv.org/html/2609.38666#S5.Thmtheorem1)\.

Fix a contextxxwithρ¯​\(x\)\>0\\bar\{\\rho\}\(x\)\>0\. The contextwise reverse\-KL objective is

LR,x\(π\):=∑i∈ℐx\+wi\(x\)KL\(π\(⋅∣x\)∥pi\(⋅∣x\)\)\.\\displaystyle L\_\{\\mathrm\{R\},x\}\(\\pi\):=\\sum\_\{i\\in\\mathcal\{I\}\_\{x\}^\{\+\}\}w\_\{i\}\(x\)\\text\{KL\}\\bigl\(\\pi\(\\cdot\\mid x\)\\,\\big\\\|\\,p\_\{i\}\(\\cdot\\mid x\)\\bigr\)\.Define the unnormalized sequence score and its total mass by

Gx​\(y\):=∏i∈ℐx\+pi​\(y∣x\)wi​\(x\),Zx:=∑y∈𝒴Gx​\(y\)\.\\displaystyle G\_\{x\}\(y\):=\\prod\_\{i\\in\\mathcal\{I\}\_\{x\}^\{\+\}\}p\_\{i\}\(y\\mid x\)^\{w\_\{i\}\(x\)\},\\qquad Z\_\{x\}:=\\sum\_\{y\\in\\mathcal\{Y\}\}G\_\{x\}\(y\)\.We first show thatZx=V1​\(x\)Z\_\{x\}=V\_\{1\}\(x\), then identify the minimizing sequence distribution and its token conditionals\.

Backward recursion and normalization\.By the autoregressive factorization of the teachers,

Gx​\(y\)\\displaystyle G\_\{x\}\(y\)=∏i∈ℐx\+\(∏h=1Hpi​\(ah∣x,y<h\)\)wi​\(x\)\\displaystyle=\\prod\_\{i\\in\\mathcal\{I\}\_\{x\}^\{\+\}\}\\left\(\\prod\_\{h=1\}^\{H\}p\_\{i\}\(a\_\{h\}\\mid x,y\_\{<h\}\)\\right\)^\{w\_\{i\}\(x\)\}=∏h=1H∏i∈ℐx\+pi​\(ah∣x,y<h\)wi​\(x\)\\displaystyle=\\prod\_\{h=1\}^\{H\}\\prod\_\{i\\in\\mathcal\{I\}\_\{x\}^\{\+\}\}p\_\{i\}\(a\_\{h\}\\mid x,y\_\{<h\}\)^\{w\_\{i\}\(x\)\}=∏h=1Hgh​\(ah∣x,y<h\)\.\\displaystyle=\\prod\_\{h=1\}^\{H\}g\_\{h\}\(a\_\{h\}\\mid x,y\_\{<h\}\)\.\(A\.5\)For everyh∈\[H\+1\]h\\in\[H\+1\]andu∈𝒜h−1u\\in\\mathcal\{A\}^\{h\-1\}, we claim that

Vh​\(x,u\)=∑ah,…,aH∈𝒜∏k=hHgk​\(ak∣x,y<k\),y=\(u,ah,…,aH\)\.\\displaystyle V\_\{h\}\(x,u\)=\\sum\_\{a\_\{h\},\\ldots,a\_\{H\}\\in\\mathcal\{A\}\}\\prod\_\{k=h\}^\{H\}g\_\{k\}\(a\_\{k\}\\mid x,y\_\{<k\}\),\\qquad y=\(u,a\_\{h\},\\ldots,a\_\{H\}\)\.\(A\.6\)Ath=H\+1h=H\+1, the sum contains one empty continuation and its empty product is one, so the identity agrees withVH\+1​\(x,y\)=1V\_\{H\+1\}\(x,y\)=1\. Suppose that it holds at levelh\+1h\+1\. Using the defining recursion forVhV\_\{h\}, we obtain

Vh​\(x,u\)\\displaystyle V\_\{h\}\(x,u\)=∑ah∈𝒜gh​\(ah∣x,u\)​Vh\+1​\(x,\(u,ah\)\)\\displaystyle=\\sum\_\{a\_\{h\}\\in\\mathcal\{A\}\}g\_\{h\}\(a\_\{h\}\\mid x,u\)V\_\{h\+1\}\(x,\(u,a\_\{h\}\)\)=∑ah∈𝒜gh​\(ah∣x,u\)​∑ah\+1,…,aH∈𝒜∏k=h\+1Hgk​\(ak∣x,y<k\)\\displaystyle=\\sum\_\{a\_\{h\}\\in\\mathcal\{A\}\}g\_\{h\}\(a\_\{h\}\\mid x,u\)\\sum\_\{a\_\{h\+1\},\\ldots,a\_\{H\}\\in\\mathcal\{A\}\}\\prod\_\{k=h\+1\}^\{H\}g\_\{k\}\(a\_\{k\}\\mid x,y\_\{<k\}\)=∑ah,…,aH∈𝒜∏k=hHgk​\(ak∣x,y<k\),\\displaystyle=\\sum\_\{a\_\{h\},\\ldots,a\_\{H\}\\in\\mathcal\{A\}\}\\prod\_\{k=h\}^\{H\}g\_\{k\}\(a\_\{k\}\\mid x,y\_\{<k\}\),where the second equality uses the induction hypothesis\. This proves \([A\.6](https://arxiv.org/html/2609.38666#A1.E6)\) by backward induction\. At the empty prefix, it gives

V1​\(x\)=∑y∈𝒴Gx​\(y\)=Zx\.\\displaystyle V\_\{1\}\(x\)=\\sum\_\{y\\in\\mathcal\{Y\}\}G\_\{x\}\(y\)=Z\_\{x\}\.\(A\.7\)In particular, the positivity conditionV1​\(x\)\>0V\_\{1\}\(x\)\>0impliesZx\>0Z\_\{x\}\>0\.

Sequence\-level minimizer\.Define

qR​\(y∣x\):=Gx​\(y\)Zx,𝒮R​\(x\):=\{y∈𝒴:Gx​\(y\)\>0\}\.\\displaystyle q\_\{\\mathrm\{R\}\}\(y\\mid x\):=\\frac\{G\_\{x\}\(y\)\}\{Z\_\{x\}\},\\qquad\\mathcal\{S\}\_\{\\mathrm\{R\}\}\(x\):=\\\{y\\in\\mathcal\{Y\}:G\_\{x\}\(y\)\>0\\\}\.By \([A\.7](https://arxiv.org/html/2609.38666#A1.E7)\),qR\(⋅∣x\)q\_\{\\mathrm\{R\}\}\(\\cdot\\mid x\)is a probability distribution\. The set𝒮R​\(x\)\\mathcal\{S\}\_\{\\mathrm\{R\}\}\(x\)is the common support of all positive\-weight teachers\. Any policy assigning positive mass outside this set has infinite reverse KL to at least one such teacher, and therefore incurs infinite objective value\. For a policy supported on𝒮R​\(x\)\\mathcal\{S\}\_\{\\mathrm\{R\}\}\(x\), we have

LR,x​\(π\)\\displaystyle L\_\{\\mathrm\{R\},x\}\(\\pi\)=∑y∈𝒮R​\(x\)π⁡\(y∣x\)​\[log⁡π⁡\(y∣x\)−∑i∈ℐx\+wi​\(x\)​log⁡pi​\(y∣x\)\]\\displaystyle=\\sum\_\{y\\in\\mathcal\{S\}\_\{\\mathrm\{R\}\}\(x\)\}\\pi\(y\\mid x\)\\left\[\\log\\pi\(y\\mid x\)\-\\sum\_\{i\\in\\mathcal\{I\}\_\{x\}^\{\+\}\}w\_\{i\}\(x\)\\log p\_\{i\}\(y\\mid x\)\\right\]=∑y∈𝒮R​\(x\)π⁡\(y∣x\)​log⁡π⁡\(y∣x\)Gx​\(y\)\\displaystyle=\\sum\_\{y\\in\\mathcal\{S\}\_\{\\mathrm\{R\}\}\(x\)\}\\pi\(y\\mid x\)\\log\\frac\{\\pi\(y\\mid x\)\}\{G\_\{x\}\(y\)\}=∑y∈𝒮R​\(x\)π⁡\(y∣x\)​log⁡π⁡\(y∣x\)qR​\(y∣x\)−log⁡Zx\\displaystyle=\\sum\_\{y\\in\\mathcal\{S\}\_\{\\mathrm\{R\}\}\(x\)\}\\pi\(y\\mid x\)\\log\\frac\{\\pi\(y\\mid x\)\}\{q\_\{\\mathrm\{R\}\}\(y\\mid x\)\}\-\\log Z\_\{x\}=KL\(π\(⋅∣x\)∥qR\(⋅∣x\)\)−logV1\(x\)\.\\displaystyle=\\text\{KL\}\\bigl\(\\pi\(\\cdot\\mid x\)\\,\\big\\\|\\,q\_\{\\mathrm\{R\}\}\(\\cdot\\mid x\)\\bigr\)\-\\log V\_\{1\}\(x\)\.\(A\.8\)Here, the second equality uses the definition ofGxG\_\{x\}, the third usesGx=Zx​qRG\_\{x\}=Z\_\{x\}q\_\{\\mathrm\{R\}\}and∑yπ⁡\(y∣x\)=1\\sum\_\{y\}\\pi\(y\\mid x\)=1, and the last uses \([A\.7](https://arxiv.org/html/2609.38666#A1.E7)\)\. The same identity holds in the extended\-real sense whenπ\\piputs positive mass outside𝒮R​\(x\)\\mathcal\{S\}\_\{\\mathrm\{R\}\}\(x\), since both sides are then infinite\.

Since−log⁡V1​\(x\)\-\\log V\_\{1\}\(x\)does not depend onπ\\pi, nonnegativity of KL divergence shows that the unique minimizing sequence distribution is

πR∗​\(y∣x\)=qR​\(y∣x\)=∏i∈ℐx\+pi​\(y∣x\)wi​\(x\)V1​\(x\),\\displaystyle\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(y\\mid x\)=q\_\{\\mathrm\{R\}\}\(y\\mid x\)=\\frac\{\\prod\_\{i\\in\\mathcal\{I\}\_\{x\}^\{\+\}\}p\_\{i\}\(y\\mid x\)^\{w\_\{i\}\(x\)\}\}\{V\_\{1\}\(x\)\},which proves \([5\.1](https://arxiv.org/html/2609.38666#S5.E1)\)\.

Token\-level conditionals\.Fixh∈\[H\]h\\in\[H\]and a prefixu=\(u1,…,uh−1\)u=\(u\_\{1\},\\ldots,u\_\{h\-1\}\), and let

G<h​\(x,u\):=∏k=1h−1gk​\(uk∣x,u<k\),\\displaystyle G\_\{<h\}\(x,u\):=\\prod\_\{k=1\}^\{h\-1\}g\_\{k\}\(u\_\{k\}\\mid x,u\_\{<k\}\),where the empty product is one\. Combining \([A\.5](https://arxiv.org/html/2609.38666#A1.E5)\) and \([A\.6](https://arxiv.org/html/2609.38666#A1.E6)\), and summing over all continuations, gives

πR∗​\(u∣x\)\\displaystyle\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(u\\mid x\)=G<h​\(x,u\)​Vh​\(x,u\)V1​\(x\),\\displaystyle=\\frac\{G\_\{<h\}\(x,u\)V\_\{h\}\(x,u\)\}\{V\_\{1\}\(x\)\},πR∗​\(\(u,a\)∣x\)\\displaystyle\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(\(u,a\)\\mid x\)=G<h​\(x,u\)​gh​\(a∣x,u\)​Vh\+1​\(x,\(u,a\)\)V1​\(x\)\.\\displaystyle=\\frac\{G\_\{<h\}\(x,u\)g\_\{h\}\(a\\mid x,u\)V\_\{h\+1\}\(x,\(u,a\)\)\}\{V\_\{1\}\(x\)\}\.\(A\.9\)IfπR∗​\(u∣x\)\>0\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(u\\mid x\)\>0, then bothG<h​\(x,u\)G\_\{<h\}\(x,u\)andVh​\(x,u\)V\_\{h\}\(x,u\)are positive\. Taking the ratio of the two prefix probabilities yields

πR∗​\(a∣x,u\)\\displaystyle\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(a\\mid x,u\)=πR∗​\(\(u,a\)∣x\)πR∗​\(u∣x\)\\displaystyle=\\frac\{\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(\(u,a\)\\mid x\)\}\{\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(u\\mid x\)\}=gh​\(a∣x,u\)​Vh\+1​\(x,\(u,a\)\)Vh​\(x,u\),\\displaystyle=\\frac\{g\_\{h\}\(a\\mid x,u\)V\_\{h\+1\}\(x,\(u,a\)\)\}\{V\_\{h\}\(x,u\)\},\(A\.10\)which is \([5\.2](https://arxiv.org/html/2609.38666#S5.E2)\)\. These conditionals are uniquely determined at every positive\-probability prefix\.

At a null prefix withVh​\(x,u\)\>0V\_\{h\}\(x,u\)\>0, the right\-hand side of \([A\.10](https://arxiv.org/html/2609.38666#A1.E10)\) still defines a valid conditional, because its numerator sums toVh​\(x,u\)V\_\{h\}\(x,u\)\. WhenVh​\(x,u\)=0V\_\{h\}\(x,u\)=0, choose an arbitrary conditional instead\. To verify that these choices induce the minimizing sequence law, take anyy∈𝒮R​\(x\)y\\in\\mathcal\{S\}\_\{\\mathrm\{R\}\}\(x\)\. All factors and continuation values along this response are positive, so the selected conditionals telescope:

∏h=1Hgh​\(ah∣x,y<h\)​Vh\+1​\(x,y≤h\)Vh​\(x,y<h\)\\displaystyle\\prod\_\{h=1\}^\{H\}\\frac\{g\_\{h\}\(a\_\{h\}\\mid x,y\_\{<h\}\)V\_\{h\+1\}\(x,y\_\{\\leq h\}\)\}\{V\_\{h\}\(x,y\_\{<h\}\)\}=∏h=1Hgh​\(ah∣x,y<h\)V1​\(x\)\\displaystyle=\\frac\{\\prod\_\{h=1\}^\{H\}g\_\{h\}\(a\_\{h\}\\mid x,y\_\{<h\}\)\}\{V\_\{1\}\(x\)\}=Gx​\(y\)Zx=πR∗​\(y∣x\),\\displaystyle=\\frac\{G\_\{x\}\(y\)\}\{Z\_\{x\}\}=\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(y\\mid x\),wherey≤h:=\(a1,…,ah\)y\_\{\\leq h\}:=\(a\_\{1\},\\ldots,a\_\{h\}\)and the first equality usesVH\+1​\(x,y\)=1V\_\{H\+1\}\(x,y\)=1\. These probabilities already sum to one over𝒮R​\(x\)\\mathcal\{S\}\_\{\\mathrm\{R\}\}\(x\)\. Since every selected token conditional is normalized, the induced autoregressive law assigns zero mass to all other responses and equalsπR∗\(⋅∣x\)\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(\\cdot\\mid x\)\. Thus, uniqueness concerns the sequence distribution and its conditionals at positive\-probability prefixes, not arbitrary conditionals at null prefixes\. Applying \([A\.1](https://arxiv.org/html/2609.38666#A1.E1)\) across contexts completes the proof\. ∎

## Appendix BProof of Theorem[4\.3](https://arxiv.org/html/2609.38666#S4.Thmtheorem3)

###### Proof\.

In this section, we prove Theorem[4\.3](https://arxiv.org/html/2609.38666#S4.Thmtheorem3)\. Letℱt:=σ\(\{xs,j,ys,j,qs,j,h\(a\):1≤s<t,j∈\[m\],h∈\[H\],a∈𝒜\}\)\\mathcal\{F\}\_\{t\}:=\\sigma\\big\(\\\{x\_\{s,j\},\\,y\_\{s,j\},\\,q\_\{s,j,h\}\(a\):1\\leq s<t,\\,j\\in\[m\],\\,h\\in\[H\],\\,a\\in\\mathcal\{A\}\\\}\\big\)\. Then,π^t\\widehat\{\\pi\}\_\{t\}isℱt\\mathcal\{F\}\_\{t\}\-measurable\.

Fix a levelhh, a state\(x,u\)∈𝒳×𝒜h−1\(x,u\)\\in\\mathcal\{X\}\\times\\mathcal\{A\}^\{h\-1\}, and a tokena∈𝒜a\\in\\mathcal\{A\}\. Under either feedback model,

𝔼\[qt,j,h\(a\)\|ℱt,it=i,xt,j=x,yt,j,<h=u\]=pi\(a\|x,u\)\.\\displaystyle\\mathbb\{E\}\\left\[q\_\{t,j,h\}\(a\)\\,\\middle\|\\,\\mathcal\{F\}\_\{t\},\\,i\_\{t\}=i,\\,x\_\{t,j\}=x,\\,y\_\{t,j,<h\}=u\\right\]=p\_\{i\}\(a\|x,u\)\.\(B\.1\)Indeed, without logits,qt,j,h​\(a\)=𝟙⁡\(at,j,h=a\)q\_\{t,j,h\}\(a\)=\\ind\(a\_\{t,j,h\}=a\), whose conditional mean ispi​\(a\|x,u\)p\_\{i\}\(a\|x,u\)\. With logits,qt,j,h​\(a\)=pi​\(a\|x,u\)q\_\{t,j,h\}\(a\)=p\_\{i\}\(a\|x,u\)is observed directly\.

Conditioned onit=ii\_\{t\}=i, using the independence of samples between each round, we have

ℙ\(xt,j=x,yt,j,<h=u\|it=i,ℱt\)=ρi\(x\)pi\(u\|x\)\.\\displaystyle\\mathbb\{P\}\\big\(x\_\{t,j\}=x,\\,y\_\{t,j,<h\}=u\|i\_\{t\}=i,\\mathcal\{F\}\_\{t\}\\big\)=\\rho\_\{i\}\(x\)p\_\{i\}\(u\|x\)\.\(B\.2\)Combining \([B\.1](https://arxiv.org/html/2609.38666#A2.E1)\) and \([B\.2](https://arxiv.org/html/2609.38666#A2.E2)\), and using thatiti\_\{t\}is uniform overℐ\\mathcal\{I\}, we have

𝔼⁡\[𝟙⁡\(xt,j=x,yt,j,<h=u\)​qt,j,h​\(a\)\|ℱt\]\\displaystyle\\mathbb\{E\}\\big\[\\ind\(x\_\{t,j\}=x,\\,y\_\{t,j,<h\}=u\)q\_\{t,j,h\}\(a\)\\,\\big\|\\,\\mathcal\{F\}\_\{t\}\\big\]=∑i∈ℐℙ\(it=i∣ℱt\)𝔼\[𝟙\(xt,j=x,yt,j,<h=u\)qt,j,h\(a\)\|ℱt,it=i\]\\displaystyle\\quad=\\sum\_\{i\\in\\mathcal\{I\}\}\\mathbb\{P\}\(i\_\{t\}=i\\mid\\mathcal\{F\}\_\{t\}\)\\,\\mathbb\{E\}\\left\[\\ind\(x\_\{t,j\}=x,\\,y\_\{t,j,<h\}=u\)q\_\{t,j,h\}\(a\)\\,\\middle\|\\,\\mathcal\{F\}\_\{t\},i\_\{t\}=i\\right\]=1I∑i∈ℐℙ\(xt,j=x,yt,j,<h=u\|ℱt,it=i\)\\displaystyle\\quad=\\frac\{1\}\{I\}\\sum\_\{i\\in\\mathcal\{I\}\}\\mathbb\{P\}\\big\(x\_\{t,j\}=x,\\,y\_\{t,j,<h\}=u\\big\|\\mathcal\{F\}\_\{t\},i\_\{t\}=i\\big\)⋅𝔼\[qt,j,h\(a\)\|ℱt,it=i,xt,j=x,yt,j,<h=u\]\\displaystyle\\qquad\\cdot\\mathbb\{E\}\\big\[q\_\{t,j,h\}\(a\)\\big\|\\mathcal\{F\}\_\{t\},i\_\{t\}=i,\\,x\_\{t,j\}=x,\\,y\_\{t,j,<h\}=u\\big\]=1I​∑i∈ℐρi​\(x\)​pi​\(u\|x\)​pi​\(a\|x,u\)\.\\displaystyle\\quad=\\frac\{1\}\{I\}\\sum\_\{i\\in\\mathcal\{I\}\}\\rho\_\{i\}\(x\)p\_\{i\}\(u\|x\)p\_\{i\}\(a\|x,u\)\.\(B\.3\)The right\-hand side is precisely the joint prefix–token probability under the forward target \([4\.2](https://arxiv.org/html/2609.38666#S4.E2)\)\. More explicitly,

1I​∑i∈ℐρi​\(x\)​pi​\(u\|x\)​pi​\(a\|x,u\)=ρ¯​\(x\)​πF∗​\(u\|x\)​πF∗​\(a\|x,u\)\.\\displaystyle\\frac\{1\}\{I\}\\sum\_\{i\\in\\mathcal\{I\}\}\\rho\_\{i\}\(x\)p\_\{i\}\(u\|x\)p\_\{i\}\(a\|x,u\)=\\bar\{\\rho\}\(x\)\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(u\|x\)\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\.\(B\.4\)To see this identity, first note that

ρ¯​\(x\)​πF∗​\(u\|x\)\\displaystyle\\bar\{\\rho\}\(x\)\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(u\|x\)=\(1I​∑k∈ℐρk​\(x\)\)​\(∑i∈ℐρi​\(x\)∑k∈ℐρk​\(x\)​pi​\(u\|x\)\)\\displaystyle=\\bigg\(\\frac\{1\}\{I\}\\sum\_\{k\\in\\mathcal\{I\}\}\\rho\_\{k\}\(x\)\\bigg\)\\left\(\\sum\_\{i\\in\\mathcal\{I\}\}\\frac\{\\rho\_\{i\}\(x\)\}\{\\sum\_\{k\\in\\mathcal\{I\}\}\\rho\_\{k\}\(x\)\}p\_\{i\}\(u\|x\)\\right\)=1I​∑i∈ℐρi​\(x\)​pi​\(u\|x\)\.\\displaystyle=\\frac\{1\}\{I\}\\sum\_\{i\\in\\mathcal\{I\}\}\\rho\_\{i\}\(x\)p\_\{i\}\(u\|x\)\.\(B\.5\)Moreover, conditioned on the prefix\(x,u\)\(x,u\), the posterior weight of teacheriiis

wi​\(x,u\)=ρi​\(x\)​pi​\(u\|x\)∑k∈ℐρk​\(x\)​pk​\(u\|x\)\.\\displaystyle w\_\{i\}\(x,u\)=\\frac\{\\rho\_\{i\}\(x\)p\_\{i\}\(u\|x\)\}\{\\sum\_\{k\\in\\mathcal\{I\}\}\\rho\_\{k\}\(x\)p\_\{k\}\(u\|x\)\}\.Therefore,

πF∗​\(a\|x,u\)\\displaystyle\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)=∑i∈ℐwi​\(x,u\)​pi​\(a\|x,u\)\\displaystyle=\\sum\_\{i\\in\\mathcal\{I\}\}w\_\{i\}\(x,u\)p\_\{i\}\(a\|x,u\)=∑i∈ℐρi​\(x\)​pi​\(u\|x\)​pi​\(a\|x,u\)∑k∈ℐρk​\(x\)​pk​\(u\|x\)\.\\displaystyle=\\frac\{\\sum\_\{i\\in\\mathcal\{I\}\}\\rho\_\{i\}\(x\)p\_\{i\}\(u\|x\)p\_\{i\}\(a\|x,u\)\}\{\\sum\_\{k\\in\\mathcal\{I\}\}\\rho\_\{k\}\(x\)p\_\{k\}\(u\|x\)\}\.\(B\.6\)Multiplying \([B\.5](https://arxiv.org/html/2609.38666#A2.E5)\) and \([B\.6](https://arxiv.org/html/2609.38666#A2.E6)\), the prefix normalizing factor cancels, giving

ρ¯​\(x\)​πF∗​\(u\|x\)​πF∗​\(a\|x,u\)=1I​∑i∈ℐρi​\(x\)​pi​\(u\|x\)​pi​\(a\|x,u\)\.\\displaystyle\\bar\{\\rho\}\(x\)\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(u\|x\)\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)=\\frac\{1\}\{I\}\\sum\_\{i\\in\\mathcal\{I\}\}\\rho\_\{i\}\(x\)p\_\{i\}\(u\|x\)p\_\{i\}\(a\|x,u\)\.Recall that

ct​\(x,u,a\):=1m​∑j=1m𝟙⁡\(xt,j=x,yt,j,<h=u\)​qt,j,h​\(a\)\.\\displaystyle c\_\{t\}\(x,u,a\):=\\frac\{1\}\{m\}\\sum\_\{j=1\}^\{m\}\\ind\(x\_\{t,j\}=x,\\,y\_\{t,j,<h\}=u\)\\,q\_\{t,j,h\}\(a\)\.Averaging \([B\.3](https://arxiv.org/html/2609.38666#A2.E3)\) over themmrollouts therefore yields

𝔼⁡\[ct​\(x,u,a\)\|ℱt\]=ρ¯​\(x\)​πF∗​\(u\|x\)​πF∗​\(a\|x,u\)\.\\displaystyle\\mathbb\{E\}\\left\[c\_\{t\}\(x,u,a\)\\,\\middle\|\\,\\mathcal\{F\}\_\{t\}\\right\]=\\bar\{\\rho\}\(x\)\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(u\|x\)\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\.\(B\.7\)Only linearity of expectation is used here\. In particular, themmrollouts need not be independent after marginalizing over their shared teacher indexiti\_\{t\}\.

For any autoregressive policyπ\\pi, define its round\-ttempirical loss by

ℓt\(π\):=−∑h=1H∑x∈𝒳∑u∈𝒜h−1∑a∈𝒜ct\(x,u,a\)logπ\(a\|x,u\)\.\\displaystyle\\ell\_\{t\}\(\\pi\):=\-\\sum\_\{h=1\}^\{H\}\\sum\_\{x\\in\\mathcal\{X\}\}\\sum\_\{u\\in\\mathcal\{A\}^\{h\-1\}\}\\sum\_\{a\\in\\mathcal\{A\}\}c\_\{t\}\(x,u,a\)\\log\\pi\(a\|x,u\)\.\(B\.8\)DefineL\(π\):=𝔼x∼ρ¯,y∼πF∗\(⋅\|x\)\[−logπ\(y\|x\)\]L\(\\pi\):=\\mathbb\{E\}\_\{x\\sim\\bar\{\\rho\},y\\sim\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(\\cdot\|x\)\}\\big\[\-\\log\\pi\(y\|x\)\\big\]\. Using the autoregressive representation ofπ\\pi, we have

−logπ\(y\|x\)=−∑h=1Hlogπ\(ah\|x,y<h\)\.\\displaystyle\-\\log\\pi\(y\|x\)=\-\\sum\_\{h=1\}^\{H\}\\log\\pi\(a\_\{h\}\|x,y\_\{<h\}\)\.Consequently, using the linearity of expectation,L⁡\(π\)L\(\\pi\)can be expressed as

L\(π\)=−∑h=1H∑x,u,aρ¯\(x\)πF∗\(u\|x\)πF∗\(a\|x,u\)logπ\(a\|x,u\)\.\\displaystyle L\(\\pi\)=\-\\sum\_\{h=1\}^\{H\}\\sum\_\{x,u,a\}\\bar\{\\rho\}\(x\)\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(u\|x\)\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\\log\\pi\(a\|x,u\)\.Substituting \([B\.7](https://arxiv.org/html/2609.38666#A2.E7)\) into \([B\.8](https://arxiv.org/html/2609.38666#A2.E8)\) now gives, for everyℱt\\mathcal\{F\}\_\{t\}\-measurable policyπ\\pi,

𝔼⁡\[ℓt​\(π\)\|ℱt\]=L⁡\(π\)\.\\displaystyle\\mathbb\{E\}\\left\[\\ell\_\{t\}\(\\pi\)\\,\\middle\|\\,\\mathcal\{F\}\_\{t\}\\right\]=L\(\\pi\)\.\(B\.9\)Moreover, we have

L⁡\(π\)−L⁡\(πF∗\)\\displaystyle L\(\\pi\)\-L\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)=𝔼x∼ρ¯,y∼πF∗\(⋅\|x\)\[logπF∗​\(y\|x\)π⁡\(y\|x\)\]\\displaystyle=\\mathbb\{E\}\_\{x\\sim\\bar\{\\rho\},y\\sim\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(\\cdot\|x\)\}\\bigg\[\\log\\frac\{\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(y\|x\)\}\{\\pi\(y\|x\)\}\\bigg\]=𝔼x∼ρ¯KL\[πF∗\(⋅\|x\)∥π\(⋅\|x\)\]\.\\displaystyle=\\mathbb\{E\}\_\{x\\sim\\bar\{\\rho\}\}\\text\{KL\}\\big\[\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(\\cdot\|x\)\\\|\\pi\(\\cdot\|x\)\\big\]\.\(B\.10\)Applying \([B\.9](https://arxiv.org/html/2609.38666#A2.E9)\) and \([B\.10](https://arxiv.org/html/2609.38666#A2.E10)\) toπ^t\\widehat\{\\pi\}\_\{t\}andπF∗\\pi\_\{\\mathrm\{F\}\}^\{\*\}and the tower property, we have

𝔼⁡\[Regret⁡\(T\)\]\\displaystyle\\mathbb\{E\}\\big\[\\operatorname\{Regret\}\(T\)\\big\]=𝔼⁡\[∑t=1T\[L⁡\(π^t\)−L⁡\(πF∗\)\]\]\\displaystyle=\\mathbb\{E\}\\bigg\[\\sum\_\{t=1\}^\{T\}\\big\[L\(\\widehat\{\\pi\}\_\{t\}\)\-L\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\\big\]\\bigg\]=𝔼⁡\[∑t=1T\[ℓt​\(π^t\)−ℓt​\(πF∗\)\]\]\.\\displaystyle=\\mathbb\{E\}\\bigg\[\\sum\_\{t=1\}^\{T\}\\big\[\\ell\_\{t\}\(\\widehat\{\\pi\}\_\{t\}\)\-\\ell\_\{t\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\\big\]\\bigg\]\.\(B\.11\)Fixh∈\[H\]h\\in\[H\], define

ℓt,h\(π\):=−∑x∈𝒳∑u∈𝒜h−1∑a∈𝒜ct\(x,u,a\)logπ\(a\|x,u\)\.\\displaystyle\\ell\_\{t,h\}\(\\pi\):=\-\\sum\_\{x\\in\\mathcal\{X\}\}\\sum\_\{u\\in\\mathcal\{A\}^\{h\-1\}\}\\sum\_\{a\\in\\mathcal\{A\}\}c\_\{t\}\(x,u,a\)\\log\\pi\(a\|x,u\)\.Then,ℓt​\(π\)=∑hℓt,h​\(π\)\\ell\_\{t\}\(\\pi\)=\\sum\_\{h\}\\ell\_\{t,h\}\(\\pi\)\. We consider∑t=1T\[ℓt,h​\(π^t\)−ℓt,h​\(πF∗\)\]\\sum\_\{t=1\}^\{T\}\\big\[\\ell\_\{t,h\}\(\\widehat\{\\pi\}\_\{t\}\)\-\\ell\_\{t,h\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\\big\]\. Recall thatWt​\(x,u,a\)=∑i=1tci​\(x,u,a\)W\_\{t\}\(x,u,a\)=\\sum\_\{i=1\}^\{t\}c\_\{i\}\(x,u,a\),Wt​\(x,u\)=∑aWt​\(x,u,a\)W\_\{t\}\(x,u\)=\\sum\_\{a\}W\_\{t\}\(x,u,a\)\. For a fixed state\(x,u\)\(x,u\), we omit\(x,u\)\(x,u\)and use the shorthand notationWt​\(a\)W\_\{t\}\(a\)andWtW\_\{t\}when it will not cause any confusion\. Similarly, we writect​\(a\)=ct​\(x,u,a\)c\_\{t\}\(a\)=c\_\{t\}\(x,u,a\)and definect:=∑act​\(a\)c\_\{t\}:=\\sum\_\{a\}c\_\{t\}\(a\)\. Thus,Wt=∑i=1tciW\_\{t\}=\\sum\_\{i=1\}^\{t\}c\_\{i\}\.

Using the definition ofct​\(x,u,a\)c\_\{t\}\(x,u,a\)in \([4\.3](https://arxiv.org/html/2609.38666#S4.E3)\), we have

ct=1m​∑j=1m𝟙⁡\(xt,j=x,yt,j,<h=u\)​∑aqt,j,h​\(a\)\.\\displaystyle c\_\{t\}=\\frac\{1\}\{m\}\\sum\_\{j=1\}^\{m\}\\ind\(x\_\{t,j\}=x,\\,y\_\{t,j,<h\}=u\)\\,\\sum\_\{a\}q\_\{t,j,h\}\(a\)\.Note that in both cases,∑aqt,j,h​\(a\)=1\\sum\_\{a\}q\_\{t,j,h\}\(a\)=1\. Thus, we have

ct\\displaystyle c\_\{t\}=1m​∑j=1m𝟙⁡\(xt,j=x,yt,j,<h=u\)≤1\.\\displaystyle=\\frac\{1\}\{m\}\\sum\_\{j=1\}^\{m\}\\ind\(x\_\{t,j\}=x,\\,y\_\{t,j,<h\}=u\)\\leq 1\.In Algorithm[1](https://arxiv.org/html/2609.38666#alg1), the policy at context\(x,u\)\(x,u\)is defined as

π^t​\(a\|x,u\)=Wt−1​\(a\)\+1/2Wt−1\+A/2\.\\displaystyle\\widehat\{\\pi\}\_\{t\}\(a\|x,u\)=\\frac\{W\_\{t\-1\}\(a\)\+1/2\}\{W\_\{t\-1\}\+A/2\}\.For simplicity, we write the action set as𝒜=\{1,2,…,A\}\\mathcal\{A\}=\\\{1,2,\\ldots,A\\\}\. Let𝐐t\\mathbf\{Q\}\_\{t\}be a randomAA\-dim probability vector with Dirichlet distribution

𝐐t∼Dir⁡\(Wt−1​\(1\)\+12,…,Wt−1​\(A\)\+12\)\.\\displaystyle\\mathbf\{Q\}\_\{t\}\\sim\\operatorname\{Dir\}\\bigg\(W\_\{t\-1\}\(1\)\+\\frac\{1\}\{2\},\\ldots,W\_\{t\-1\}\(A\)\+\\frac\{1\}\{2\}\\bigg\)\.The goal of this construction is the following fact: the mean of𝐐t\\mathbf\{Q\}\_\{t\}is exactly the output policy:

𝔼q∼𝐐t​\[q⁡\(a\)\]=Wt−1​\(a\)\+1/2Wt−1\+A/2=π^t​\(a\|x,u\)\.\\displaystyle\\mathbb\{E\}\_\{q\\sim\\mathbf\{Q\}\_\{t\}\}\[q\(a\)\]=\\frac\{W\_\{t\-1\}\(a\)\+1/2\}\{W\_\{t\-1\}\+A/2\}=\\widehat\{\\pi\}\_\{t\}\(a\|x,u\)\.For a positive vectorv=\(v1,…,vA\)v=\(v\_\{1\},\\ldots,v\_\{A\}\), define the multivariate beta function by

B⁡\(v\):=∏a=1AΓ⁡\(va\)Γ⁡\(∑a=1Ava\),\\displaystyle\\mathrm\{B\}\(v\):=\\frac\{\\prod\_\{a=1\}^\{A\}\\Gamma\(v\_\{a\}\)\}\{\\Gamma\(\\sum\_\{a=1\}^\{A\}v\_\{a\}\)\},\(B\.12\)whereΓ\\Gammais the gamma function\. The beta function is the normalizing constant of the Dirichlet distribution\. Indeed, for every positive vectorvv, the density of the Dirichlet distribution𝐐v\\mathbf\{Q\}\_\{v\}on the simplexΔA:=\{q∈ℝ\+A:∑aq⁡\(a\)=1\}\\Delta\_\{A\}:=\\\{q\\in\\mathbb\{R\}\_\{\+\}^\{A\}:\\sum\_\{a\}q\(a\)=1\\\}is

fv​\(q\)=1B⁡\(v\)​∏a=1Aq​\(a\)va−1,q∈ΔA\.\\displaystyle f\_\{v\}\(q\)=\\frac\{1\}\{\\mathrm\{B\}\(v\)\}\\prod\_\{a=1\}^\{A\}q\(a\)^\{v\_\{a\}\-1\},\\qquad q\\in\\Delta\_\{A\}\.Equivalently,

∫ΔA∏a=1Aq​\(a\)va−1​𝑑q=B⁡\(v\)\.\\displaystyle\\int\_\{\\Delta\_\{A\}\}\\prod\_\{a=1\}^\{A\}q\(a\)^\{v\_\{a\}\-1\}\\,\\mathrm\{d\}q=\\mathrm\{B\}\(v\)\.More generally, for every nonnegative vectorc∈ℝ\+Ac\\in\\mathbb\{R\}\_\{\+\}^\{A\},

𝔼q∼𝐐v​\[∏a=1Aq​\(a\)ca\]\\displaystyle\\mathbb\{E\}\_\{q\\sim\\mathbf\{Q\}\_\{v\}\}\\bigg\[\\prod\_\{a=1\}^\{A\}q\(a\)^\{c\_\{a\}\}\\bigg\]=1B⁡\(v\)​∫ΔA∏a=1Aq​\(a\)ca\+va−1​𝑑q\\displaystyle=\\frac\{1\}\{\\mathrm\{B\}\(v\)\}\\int\_\{\\Delta\_\{A\}\}\\prod\_\{a=1\}^\{A\}q\(a\)^\{c\_\{a\}\+v\_\{a\}\-1\}\\mathrm\{d\}q=B⁡\(v\+c\)B⁡\(v\)\.\\displaystyle=\\frac\{\\mathrm\{B\}\(v\+c\)\}\{\\mathrm\{B\}\(v\)\}\.Therefore, we have

𝔼q∼𝐐t​\[∏a∈𝒜q​\(a\)ct​\(a\)\]\\displaystyle\\mathbb\{E\}\_\{q\\sim\\mathbf\{Q\}\_\{t\}\}\\bigg\[\\prod\_\{a\\in\\mathcal\{A\}\}q\(a\)^\{c\_\{t\}\(a\)\}\\bigg\]=B⁡\(Wt−1​\(1\)\+ct​\(1\)\+12,…,Wt−1​\(A\)\+ct​\(A\)\+12\)B⁡\(Wt−1​\(1\)\+12,…,Wt−1​\(A\)\+12\)\\displaystyle=\\frac\{\\mathrm\{B\}\\big\(W\_\{t\-1\}\(1\)\+c\_\{t\}\(1\)\+\\frac\{1\}\{2\},\\ldots,W\_\{t\-1\}\(A\)\+c\_\{t\}\(A\)\+\\frac\{1\}\{2\}\\big\)\}\{\\mathrm\{B\}\\big\(W\_\{t\-1\}\(1\)\+\\frac\{1\}\{2\},\\ldots,W\_\{t\-1\}\(A\)\+\\frac\{1\}\{2\}\\big\)\}=B⁡\(Wt​\(1\)\+12,…,Wt​\(A\)\+12\)B⁡\(Wt−1​\(1\)\+12,…,Wt−1​\(A\)\+12\),\\displaystyle=\\frac\{\\mathrm\{B\}\\big\(W\_\{t\}\(1\)\+\\frac\{1\}\{2\},\\ldots,W\_\{t\}\(A\)\+\\frac\{1\}\{2\}\\big\)\}\{\\mathrm\{B\}\\big\(W\_\{t\-1\}\(1\)\+\\frac\{1\}\{2\},\\ldots,W\_\{t\-1\}\(A\)\+\\frac\{1\}\{2\}\\big\)\},\(B\.13\)where we useWt​\(a\)=Wt−1​\(a\)\+ct​\(a\)W\_\{t\}\(a\)=W\_\{t\-1\}\(a\)\+c\_\{t\}\(a\)\. Let

αt:=∑a∈𝒜ct​\(a\)≤1,𝒜t\+:=\{a∈𝒜:ct​\(a\)\>0\}\.\\displaystyle\\alpha\_\{t\}:=\\sum\_\{a\\in\\mathcal\{A\}\}c\_\{t\}\(a\)\\leq 1,\\qquad\\mathcal\{A\}\_\{t\}^\{\+\}:=\\\{a\\in\\mathcal\{A\}:c\_\{t\}\(a\)\>0\\\}\.For everya∈𝒜t\+a\\in\\mathcal\{A\}\_\{t\}^\{\+\}, define the Hölder exponentra:=1/ct​\(a\)r\_\{a\}:=\{1\}/\{c\_\{t\}\(a\)\}\. Ifαt<1\\alpha\_\{t\}<1, introduce one additional exponentr0:=1/\(1−αt\)r\_\{0\}:=1/\(1\-\\alpha\_\{t\}\)\. These exponents satisfy

1r0\+∑a∈𝒜t\+1ra=\(1−αt\)\+∑a∈𝒜t\+ct​\(a\)=1\.\\displaystyle\\frac\{1\}\{r\_\{0\}\}\+\\sum\_\{a\\in\\mathcal\{A\}\_\{t\}^\{\+\}\}\\frac\{1\}\{r\_\{a\}\}=\(1\-\\alpha\_\{t\}\)\+\\sum\_\{a\\in\\mathcal\{A\}\_\{t\}^\{\+\}\}c\_\{t\}\(a\)=1\.Using the generalized Hölder’s inequality, we have

𝔼q∼𝐐t​\[∏a∈𝒜q​\(a\)ct​\(a\)\]\\displaystyle\\mathbb\{E\}\_\{q\\sim\\mathbf\{Q\}\_\{t\}\}\\bigg\[\\prod\_\{a\\in\\mathcal\{A\}\}q\(a\)^\{c\_\{t\}\(a\)\}\\bigg\]≤𝔼​\[1r0\]1r0​∏a∈𝒜t\+𝔼q∼𝐐t​\[\[q​\(a\)ct​\(a\)\]ra\]1ra\\displaystyle\\leq\\mathbb\{E\}\[1^\{r\_\{0\}\}\]^\{\\frac\{1\}\{r\_\{0\}\}\}\\prod\_\{a\\in\\mathcal\{A\}\_\{t\}^\{\+\}\}\\mathbb\{E\}\_\{q\\sim\\mathbf\{Q\}\_\{t\}\}\\bigg\[\\Big\[q\(a\)^\{c\_\{t\}\(a\)\}\\Big\]^\{r\_\{a\}\}\\bigg\]^\{\\frac\{1\}\{r\_\{a\}\}\}=∏a∈𝒜t\+𝔼q∼𝐐t​\[q⁡\(a\)\]ct​\(a\)=∏a∈𝒜π^t​\(a\|x,u\)ct​\(a\),\\displaystyle=\\prod\_\{a\\in\\mathcal\{A\}\_\{t\}^\{\+\}\}\\mathbb\{E\}\_\{q\\sim\\mathbf\{Q\}\_\{t\}\}\\big\[q\(a\)\\big\]^\{c\_\{t\}\(a\)\}=\\prod\_\{a\\in\\mathcal\{A\}\}\\widehat\{\\pi\}\_\{t\}\(a\|x,u\)^\{c\_\{t\}\(a\)\},\(B\.14\)where we used𝔼q∼𝐐t​\[q⁡\(a\)\]=π^t​\(a\|x,u\)\\mathbb\{E\}\_\{q\\sim\\mathbf\{Q\}\_\{t\}\}\[q\(a\)\]=\\widehat\{\\pi\}\_\{t\}\(a\|x,u\)\. Ifαt=1\\alpha\_\{t\}=1, we can skipr0r\_\{0\}and use the same generalized Hölder’s inequality\.

Combining \([B\.13](https://arxiv.org/html/2609.38666#A2.E13)\) and \([B\.14](https://arxiv.org/html/2609.38666#A2.E14)\), multiplying overtt, and using the telescoping argument, we have

∏t=1T∏a∈𝒜π^t​\(a\|x,u\)ct​\(a\)\\displaystyle\\prod\_\{t=1\}^\{T\}\\prod\_\{a\\in\\mathcal\{A\}\}\\widehat\{\\pi\}\_\{t\}\(a\|x,u\)^\{c\_\{t\}\(a\)\}≥∏t=1TB​\(Wt​\(⋅\)\+12​𝟏\)B​\(Wt−1​\(⋅\)\+12​𝟏\)\\displaystyle\\geq\\prod\_\{t=1\}^\{T\}\\frac\{\\mathrm\{B\}\\big\(W\_\{t\}\(\\cdot\)\+\\frac\{1\}\{2\}\\mathbf\{1\}\\big\)\}\{\\mathrm\{B\}\\big\(W\_\{t\-1\}\(\\cdot\)\+\\frac\{1\}\{2\}\\mathbf\{1\}\\big\)\}=B​\(WT​\(⋅\)\+12​𝟏\)B⁡\(12​𝟏\)\.\\displaystyle=\\frac\{\\mathrm\{B\}\\big\(W\_\{T\}\(\\cdot\)\+\\frac\{1\}\{2\}\\mathbf\{1\}\\big\)\}\{\\mathrm\{B\}\\big\(\\frac\{1\}\{2\}\\mathbf\{1\}\\big\)\}\.Taking the negative logarithm, we have

−∑t=1T∑a∈𝒜ct\(a\)logπ^t\(a\|x,u\)≤logB⁡\(12​𝟏\)B​\(WT​\(⋅\)\+12​𝟏\)\.\\displaystyle\-\\sum\_\{t=1\}^\{T\}\\sum\_\{a\\in\\mathcal\{A\}\}c\_\{t\}\(a\)\\log\\widehat\{\\pi\}\_\{t\}\(a\|x,u\)\\leq\\log\\frac\{\\mathrm\{B\}\(\\frac\{1\}\{2\}\\mathbf\{1\}\)\}\{\\mathrm\{B\}\\big\(W\_\{T\}\(\\cdot\)\+\\frac\{1\}\{2\}\\mathbf\{1\}\\big\)\}\.\(B\.15\)Next, we consider an arbitrary policyπ\(⋅\|x,u\)\\pi\(\\cdot\|x,u\)\. First supposeWT\>0W\_\{T\}\>0\. By the nonnegativity of KL divergence,

∑a∈𝒜WT​\(a\)​log⁡π⁡\(a\|x,u\)\\displaystyle\\sum\_\{a\\in\\mathcal\{A\}\}W\_\{T\}\(a\)\\log\\pi\(a\|x,u\)=∑a∈𝒜WT​\(a\)​log⁡WT​\(a\)WT−WT​∑a∈𝒜WT​\(a\)WT​log⁡WT​\(a\)WT⋅π⁡\(a\|x,u\)\\displaystyle=\\sum\_\{a\\in\\mathcal\{A\}\}W\_\{T\}\(a\)\\log\\frac\{W\_\{T\}\(a\)\}\{W\_\{T\}\}\-W\_\{T\}\\sum\_\{a\\in\\mathcal\{A\}\}\\frac\{W\_\{T\}\(a\)\}\{W\_\{T\}\}\\log\\frac\{W\_\{T\}\(a\)\}\{W\_\{T\}\\cdot\\pi\(a\|x,u\)\}=∑a∈𝒜WT\(a\)logWT​\(a\)WT−WTKL\(WT​\(⋅\)WT∥π\(⋅\|x,u\)\)\\displaystyle=\\sum\_\{a\\in\\mathcal\{A\}\}W\_\{T\}\(a\)\\log\\frac\{W\_\{T\}\(a\)\}\{W\_\{T\}\}\-W\_\{T\}\\text\{KL\}\\bigg\(\\frac\{W\_\{T\}\(\\cdot\)\}\{W\_\{T\}\}\\bigg\\\|\\pi\(\\cdot\|x,u\)\\bigg\)≤∑a∈𝒜WT​\(a\)​log⁡WT​\(a\)WT,\\displaystyle\\leq\\sum\_\{a\\in\\mathcal\{A\}\}W\_\{T\}\(a\)\\log\\frac\{W\_\{T\}\(a\)\}\{W\_\{T\}\},\(B\.16\)where we set0​log⁡0:=00\\log 0:=0\. Summing \([B\.15](https://arxiv.org/html/2609.38666#A2.E15)\) and \([B\.16](https://arxiv.org/html/2609.38666#A2.E16)\), we have

∑t=1T∑a∈𝒜ct​\(a\)​log⁡π⁡\(a\|z\)π^t​\(a\|z\)\\displaystyle\\sum\_\{t=1\}^\{T\}\\sum\_\{a\\in\\mathcal\{A\}\}c\_\{t\}\(a\)\\log\\frac\{\\pi\(a\|z\)\}\{\\widehat\{\\pi\}\_\{t\}\(a\|z\)\}≤log⁡B⁡\(12​𝟏\)B​\(WT​\(⋅\)\+12​𝟏\)\+∑a∈𝒜WT​\(a\)​log⁡WT​\(a\)WT=:ℛ⁡\(WT​\(⋅\)\)\.\\displaystyle\\leq\\log\\frac\{\\mathrm\{B\}\(\\frac\{1\}\{2\}\\mathbf\{1\}\)\}\{\\mathrm\{B\}\(W\_\{T\}\(\\cdot\)\+\\frac\{1\}\{2\}\\mathbf\{1\}\)\}\+\\sum\_\{a\\in\\mathcal\{A\}\}W\_\{T\}\(a\)\\log\\frac\{W\_\{T\}\(a\)\}\{W\_\{T\}\}=:\\mathcal\{R\}\\big\(W\_\{T\}\(\\cdot\)\\big\)\.WhenWT=0W\_\{T\}=0, letℛ​\(WT​\(⋅\)\)=0\\mathcal\{R\}\\big\(W\_\{T\}\(\\cdot\)\\big\)=0\. The following lemma boundsℛ​\(WT​\(⋅\)\)\\mathcal\{R\}\\big\(W\_\{T\}\(\\cdot\)\\big\)using Stirling’s formula\.

###### Lemma B\.1\.

LetWT​\(a\)≥0W\_\{T\}\(a\)\\geq 0for everya∈𝒜a\\in\\mathcal\{A\}, and let

WT:=∑a∈𝒜WT​\(a\)\.\\displaystyle W\_\{T\}:=\\sum\_\{a\\in\\mathcal\{A\}\}W\_\{T\}\(a\)\.Then,ℛ​\(WT​\(⋅\)\)\\mathcal\{R\}\\big\(W\_\{T\}\(\\cdot\)\\big\)satisfies

ℛ⁡\(WT​\(⋅\)\)≤A​log⁡\(WT\+1\)\+3​A\.\\displaystyle\\mathcal\{R\}\\bigl\(W\_\{T\}\(\\cdot\)\\bigr\)\\leq A\\log\(W\_\{T\}\+1\)\+3A\.

We defer the proof of Lemma[B\.1](https://arxiv.org/html/2609.38666#A2.Thmtheorem1)to Appendix[D](https://arxiv.org/html/2609.38666#A4)\. Chooseπ=πF∗\\pi=\\pi\_\{\\mathrm\{F\}\}^\{\*\}\. For any\(x,u\)∈𝒳×𝒜h−1\(x,u\)\\in\\mathcal\{X\}\\times\\mathcal\{A\}^\{h\-1\}, we have

∑t=1T∑a∈𝒜ct​\(x,u,a\)​log⁡πF∗​\(a\|x,u\)π^t​\(a\|x,u\)≤A​log⁡\(WT​\(x,u\)\+1\)\+3​A\.\\displaystyle\\sum\_\{t=1\}^\{T\}\\sum\_\{a\\in\\mathcal\{A\}\}c\_\{t\}\(x,u,a\)\\log\\frac\{\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\}\{\\widehat\{\\pi\}\_\{t\}\(a\|x,u\)\}\\leq A\\log\\bigl\(W\_\{T\}\(x,u\)\+1\\bigr\)\+3A\.Therefore, we have

∑t=1T\[ℓt,h​\(π^t\)−ℓt,h​\(πF∗\)\]\\displaystyle\\sum\_\{t=1\}^\{T\}\\big\[\\ell\_\{t,h\}\(\\widehat\{\\pi\}\_\{t\}\)\-\\ell\_\{t,h\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\\big\]=∑x∈𝒳∑u∈𝒜h−1∑t=1T∑a∈𝒜ct​\(x,u,a\)​log⁡πF∗​\(a\|x,u\)π^t​\(a\|x,u\)\\displaystyle=\\sum\_\{x\\in\\mathcal\{X\}\}\\sum\_\{u\\in\\mathcal\{A\}^\{h\-1\}\}\\sum\_\{t=1\}^\{T\}\\sum\_\{a\\in\\mathcal\{A\}\}c\_\{t\}\(x,u,a\)\\log\\frac\{\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\}\{\\widehat\{\\pi\}\_\{t\}\(a\|x,u\)\}≤∑\(x,u\)∈𝒳×𝒜h−1\[A​log⁡\(WT​\(x,u\)\+1\)\+3​A\]\.\\displaystyle\\leq\\sum\_\{\(x,u\)\\in\\mathcal\{X\}\\times\\mathcal\{A\}^\{h\-1\}\}\\Big\[A\\log\\bigl\(W\_\{T\}\(x,u\)\+1\\bigr\)\+3A\\Big\]\.For every fixed levelhhand roundtt, it is easy to check that

∑\(x,u\)∈𝒳×𝒜h−1∑a∈𝒜ct​\(x,u,a\)=1\.\\displaystyle\\sum\_\{\(x,u\)\\in\\mathcal\{X\}\\times\\mathcal\{A\}^\{h\-1\}\}\\sum\_\{a\\in\\mathcal\{A\}\}c\_\{t\}\(x,u,a\)=1\.Consequently,

∑\(x,u\)∈𝒳×𝒜h−1WT​\(x,u\)=T,\\displaystyle\\sum\_\{\(x,u\)\\in\\mathcal\{X\}\\times\\mathcal\{A\}^\{h\-1\}\}W\_\{T\}\(x,u\)=T,and henceWT​\(x,u\)≤TW\_\{T\}\(x,u\)\\leq Tfor any\(x,u\)\(x,u\)\. This leads to

∑t=1T\[ℓt,h​\(π^t\)−ℓt,h​\(πF∗\)\]\\displaystyle\\sum\_\{t=1\}^\{T\}\\big\[\\ell\_\{t,h\}\(\\widehat\{\\pi\}\_\{t\}\)\-\\ell\_\{t,h\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\\big\]≤2​S​Ah​log⁡T\.\\displaystyle\\leq 2SA^\{h\}\\log T\.\(B\.17\)Finally, summing \([B\.17](https://arxiv.org/html/2609.38666#A2.E17)\) overh∈\[H\]h\\in\[H\], we have

∑t=1T\[ℓt​\(π^t\)−ℓt​\(πF∗\)\]\\displaystyle\\sum\_\{t=1\}^\{T\}\\big\[\\ell\_\{t\}\(\\widehat\{\\pi\}\_\{t\}\)\-\\ell\_\{t\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\\big\]≤2​S​log⁡T​∑h=1HAh\\displaystyle\\leq 2S\\log T\\sum\_\{h=1\}^\{H\}A^\{h\}=2​S⋅log⁡T​A⁡\(AH−1\)A−1\\displaystyle=2S\\cdot\\log T\\frac\{A\(A^\{H\}\-1\)\}\{A\-1\}≤4​S​AH​log⁡T\.\\displaystyle\\leq 4SA^\{H\}\\log T\.\(B\.18\)Using \([B\.11](https://arxiv.org/html/2609.38666#A2.E11)\), we obtain

𝔼⁡\[Regret⁡\(T\)\]\\displaystyle\\mathbb\{E\}\\big\[\\operatorname\{Regret\}\(T\)\\big\]≤4​S​AH​log⁡T\.\\displaystyle\\leq 4SA^\{H\}\\log T\.We next consider the high\-probability regret bound\. Define

Xt\\displaystyle X\_\{t\}:=ℓt​\(πF∗\)−ℓt​\(π^t\),\\displaystyle:=\\ell\_\{t\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\-\\ell\_\{t\}\(\\widehat\{\\pi\}\_\{t\}\),rt\\displaystyle r\_\{t\}:=−𝔼\[Xt∣ℱt\]=𝔼x∼ρ¯KL\(πF∗\(⋅\|x\)∥π^t\(⋅\|x\)\)\.\\displaystyle:=\-\\mathbb\{E\}\[X\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]=\\mathbb\{E\}\_\{x\\sim\\bar\{\\rho\}\}\\text\{KL\}\\big\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(\\cdot\|x\)\\\|\\widehat\{\\pi\}\_\{t\}\(\\cdot\|x\)\\big\)\.Thus,Regret⁡\(T\)=∑t=1Trt\\operatorname\{Regret\}\(T\)=\\sum\_\{t=1\}^\{T\}r\_\{t\}\. DefineM0=0M\_\{0\}=0, and

Mt:=∑s=1t\(Xs\+rs\),t∈\[T\]\.\\displaystyle M\_\{t\}:=\\sum\_\{s=1\}^\{t\}\(X\_\{s\}\+r\_\{s\}\),\\qquad t\\in\[T\]\.Sincert=−𝔼⁡\[Xt∣ℱt\]r\_\{t\}=\-\\mathbb\{E\}\[X\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\], we have𝔼⁡\[Mt−Mt−1\|ℱt\]=0\\mathbb\{E\}\[M\_\{t\}\-M\_\{t\-1\}\|\\mathcal\{F\}\_\{t\}\]=0\. Thus,\{Mt\}t=0T\\\{M\_\{t\}\\\}\_\{t=0\}^\{T\}is a martingale with respect to\{ℱt\+1\}t=0T\\\{\\mathcal\{F\}\_\{t\+1\}\\\}\_\{t=0\}^\{T\}\.

Our goal is to controlMTM\_\{T\}\. Using the definition ofℓt\\ell\_\{t\},

Xt\\displaystyle X\_\{t\}=∑h=1H∑\(x,u,a\)ct​\(x,u,a\)​log⁡π^t​\(a\|x,u\)πF∗​\(a\|x,u\)\.\\displaystyle=\\sum\_\{h=1\}^\{H\}\\sum\_\{\(x,u,a\)\}c\_\{t\}\(x,u,a\)\\log\\frac\{\\widehat\{\\pi\}\_\{t\}\(a\|x,u\)\}\{\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\}\.First note that in Algorithm[1](https://arxiv.org/html/2609.38666#alg1),

π^t​\(a\|x,u\)\\displaystyle\\widehat\{\\pi\}\_\{t\}\(a\|x,u\)=Wt−1​\(x,u,a\)\+1/2Wt−1​\(x,u\)\+A/2\\displaystyle=\\frac\{W\_\{t\-1\}\(x,u,a\)\+1/2\}\{W\_\{t\-1\}\(x,u\)\+A/2\}≥12​\(t−1\)\+A≥12​T\+A\.\\displaystyle\\geq\\frac\{1\}\{2\(t\-1\)\+A\}\\geq\\frac\{1\}\{2T\+A\}\.SinceπF∗​\(a\|x,u\)≤1\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\\leq 1, it follows that

XtH≥−LT,LT:=log⁡\(2​T\+A\)\.\\displaystyle\\frac\{X\_\{t\}\}\{H\}\\geq\-L\_\{T\},\\qquad L\_\{T\}:=\\log\(2T\+A\)\.\(B\.19\)Here we use the following equation

∑\(x,u,a\)ct​\(x,u,a\)=1\.\\displaystyle\\sum\_\{\(x,u,a\)\}c\_\{t\}\(x,u,a\)=1\.Moreover, we can see

∑h=1H∑\(x,u,a\)ct​\(x,u,a\)H=1\.\\displaystyle\\sum\_\{h=1\}^\{H\}\\sum\_\{\(x,u,a\)\}\\frac\{c\_\{t\}\(x,u,a\)\}\{H\}=1\.Using the AM\-GM inequality, we have

exp⁡\(XtH\)\\displaystyle\\exp\\bigg\(\\frac\{X\_\{t\}\}\{H\}\\bigg\)=∏h=1H∏\(x,u,a\)\(π^t​\(a\|x,u\)πF∗​\(a\|x,u\)\)ct​\(x,u,a\)/H\\displaystyle=\\prod\_\{h=1\}^\{H\}\\prod\_\{\(x,u,a\)\}\\left\(\\frac\{\\widehat\{\\pi\}\_\{t\}\(a\|x,u\)\}\{\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\}\\right\)^\{c\_\{t\}\(x,u,a\)/H\}≤1H​∑h=1H∑\(x,u,a\)ct​\(x,u,a\)​π^t​\(a\|x,u\)πF∗​\(a\|x,u\)\.\\displaystyle\\leq\\frac\{1\}\{H\}\\sum\_\{h=1\}^\{H\}\\sum\_\{\(x,u,a\)\}c\_\{t\}\(x,u,a\)\\frac\{\\widehat\{\\pi\}\_\{t\}\(a\|x,u\)\}\{\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\}\.Taking conditional expectation overℱt\\mathcal\{F\}\_\{t\}and using

𝔼⁡\[ct​\(x,u,a\)∣ℱt\]=ρ¯​\(x\)​πF∗​\(u\|x\)​πF∗​\(a\|x,u\),\\displaystyle\\mathbb\{E\}\[c\_\{t\}\(x,u,a\)\\mid\\mathcal\{F\}\_\{t\}\]=\\bar\{\\rho\}\(x\)\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(u\|x\)\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\),we obtain

𝔼⁡\[exp⁡\(XtH\)\|ℱt\]\\displaystyle\\mathbb\{E\}\\bigg\[\\exp\\bigg\(\\frac\{X\_\{t\}\}\{H\}\\bigg\)\\bigg\|\\mathcal\{F\}\_\{t\}\\bigg\]≤1H​∑h=1H∑\(x,u,a\)ρ¯​\(x\)​πF∗​\(u\|x\)​πF∗​\(a\|x,u\)​π^t​\(a\|x,u\)πF∗​\(a\|x,u\)\\displaystyle\\leq\\frac\{1\}\{H\}\\sum\_\{h=1\}^\{H\}\\sum\_\{\(x,u,a\)\}\\bar\{\\rho\}\(x\)\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(u\|x\)\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\\frac\{\\widehat\{\\pi\}\_\{t\}\(a\|x,u\)\}\{\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\}=1H​∑h=1H∑\(x,u\)ρ¯​\(x\)​πF∗​\(u\|x\)​∑aπ^t​\(a\|x,u\)\\displaystyle=\\frac\{1\}\{H\}\\sum\_\{h=1\}^\{H\}\\sum\_\{\(x,u\)\}\\bar\{\\rho\}\(x\)\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(u\|x\)\\sum\_\{a\}\\widehat\{\\pi\}\_\{t\}\(a\|x,u\)≤1H​∑h=1H∑x∈𝒳ρ¯​\(x\)​∑u∈𝒜h−1πF∗​\(u\|x\)\\displaystyle\\leq\\frac\{1\}\{H\}\\sum\_\{h=1\}^\{H\}\\sum\_\{x\\in\\mathcal\{X\}\}\\bar\{\\rho\}\(x\)\\sum\_\{u\\in\\mathcal\{A\}^\{h\-1\}\}\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(u\|x\)=1\.\\displaystyle=1\.\(B\.20\)The following lemma shows that a one\-sided lower bound and an exponential central condition provide exponential control of deviations from the conditional mean\.

###### Lemma B\.2\.

Letℱ\\mathcal\{F\}be aσ\\sigma\-algebra, and letYYbe a random variable satisfying, almost surely,

Y≥−L,𝔼⁡\[eY∣ℱ\]≤1,\\displaystyle Y\\geq\-L,\\qquad\\mathbb\{E\}\[e^\{Y\}\\mid\\mathcal\{F\}\]\\leq 1,for someL≥0L\\geq 0\. Defineμ:=−𝔼⁡\[Y∣ℱ\]\\mu:=\-\\mathbb\{E\}\[Y\\mid\\mathcal\{F\}\]\. Thenμ∈\[0,L\]\\mu\\in\[0,L\], and, for everyθ∈\[0,1\]\\theta\\in\[0,1\],

log⁡𝔼⁡\[eθ⁡\(Y\+μ\)\|ℱ\]≤\(1\+L\)​θ2​μ\.\\displaystyle\\log\\mathbb\{E\}\\Big\[e^\{\\theta\(Y\+\\mu\)\}\\Big\|\\mathcal\{F\}\\Big\]\\leq\(1\+L\)\\theta^\{2\}\\mu\.

Applying Lemma[B\.2](https://arxiv.org/html/2609.38666#A2.Thmtheorem2)toY=Xt/HY=X\_\{t\}/H,L=LTL=L\_\{T\}, we have for everyλ∈\[0,1/H\]\\lambda\\in\[0,1/H\]

log⁡𝔼⁡\[eλ⁡\(Xt\+rt\)\|ℱt\]≤H⁡\(1\+LT\)​λ2​rt\.\\displaystyle\\log\\mathbb\{E\}\\left\[e^\{\\lambda\(X\_\{t\}\+r\_\{t\}\)\}\\,\\middle\|\\,\\mathcal\{F\}\_\{t\}\\right\]\\leq H\(1\+L\_\{T\}\)\\lambda^\{2\}r\_\{t\}\.\(B\.21\)Define

ℰt​\(λ\):=exp⁡\(λ​∑s=1t\(Xs\+rs\)−H⁡\(1\+LT\)​λ2​∑s=1trs\)\.\\displaystyle\\mathcal\{E\}\_\{t\}\(\\lambda\):=\\exp\\bigg\(\\lambda\\sum\_\{s=1\}^\{t\}\(X\_\{s\}\+r\_\{s\}\)\-H\(1\+L\_\{T\}\)\\lambda^\{2\}\\sum\_\{s=1\}^\{t\}r\_\{s\}\\bigg\)\.Then, \([B\.21](https://arxiv.org/html/2609.38666#A2.E21)\) implies that\{ℰs​\(λ\)\}s=0T\\\{\\mathcal\{E\}\_\{s\}\(\\lambda\)\\\}\_\{s=0\}^\{T\}is a nonnegative supermartingale withℰ0​\(λ\)=1\\mathcal\{E\}\_\{0\}\(\\lambda\)=1\. As a result,𝔼⁡\[ℰT​\(λ\)\]≤𝔼⁡\[ℰ0​\(λ\)\]=1\\mathbb\{E\}\[\\mathcal\{E\}\_\{T\}\(\\lambda\)\]\\leq\\mathbb\{E\}\[\\mathcal\{E\}\_\{0\}\(\\lambda\)\]=1\. Moreover, we have

ℰT​\(λ\)=exp⁡\(λ​MT−H⁡\(1\+LT\)​λ2​Regret⁡\(T\)\)\.\\displaystyle\\mathcal\{E\}\_\{T\}\(\\lambda\)=\\exp\\Big\(\\lambda M\_\{T\}\-H\(1\+L\_\{T\}\)\\lambda^\{2\}\\operatorname\{Regret\}\(T\)\\Big\)\.Therefore, for anyC\>0C\>0, by Markov’s inequality, we have

ℙ\[λMT−H\(1\+LT\)λ2Regret\(T\)≥C\]≤𝔼​\[ℰT​\(λ\)\]eC≤1eC\.\\displaystyle\\mathbb\{P\}\\Big\[\\lambda M\_\{T\}\-H\(1\+L\_\{T\}\)\\lambda^\{2\}\\operatorname\{Regret\}\(T\)\\geq C\\Big\]\\leq\\frac\{\\mathbb\{E\}\[\\mathcal\{E\}\_\{T\}\(\\lambda\)\]\}\{e^\{C\}\}\\leq\\frac\{1\}\{e^\{C\}\}\.SetC=log⁡\(1/δ\)C=\\log\(1/\\delta\)\. Now we obtain with probability at least1−δ1\-\\delta,

MT=∑t=1T\(Xt\+rt\)≤H⁡\(1\+LT\)​λ​Regret⁡\(T\)\+log⁡\(1/δ\)λ\.\\displaystyle M\_\{T\}=\\sum\_\{t=1\}^\{T\}\(X\_\{t\}\+r\_\{t\}\)\\leq H\(1\+L\_\{T\}\)\\lambda\\operatorname\{Regret\}\(T\)\+\\frac\{\\log\(1/\\delta\)\}\{\\lambda\}\.\(B\.22\)
On the other hand, \([B\.18](https://arxiv.org/html/2609.38666#A2.E18)\) gives

−∑t=1TXt=∑t=1T\[ℓt\(π^t\)−ℓt\(πF∗\)\]≤4SAHlogT\.\\displaystyle\-\\sum\_\{t=1\}^\{T\}X\_\{t\}=\\sum\_\{t=1\}^\{T\}\\left\[\\ell\_\{t\}\(\\widehat\{\\pi\}\_\{t\}\)\-\\ell\_\{t\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\\right\]\\leq 4SA^\{H\}\\log T\.Moreover, we have

Regret\(T\)=−∑t=1TXt\+∑t=1T\(Xt\+rt\)\.\\displaystyle\\operatorname\{Regret\}\(T\)=\-\\sum\_\{t=1\}^\{T\}X\_\{t\}\+\\sum\_\{t=1\}^\{T\}\(X\_\{t\}\+r\_\{t\}\)\.Therefore, with probability at least1−δ1\-\\delta, the following inequality holds:

Regret⁡\(T\)≤4​S​AH​log⁡T\+H⁡\(1\+LT\)​λ​Regret⁡\(T\)\+log⁡\(1/δ\)λ\.\\displaystyle\\operatorname\{Regret\}\(T\)\\leq 4SA^\{H\}\\log T\+H\(1\+L\_\{T\}\)\\lambda\\operatorname\{Regret\}\(T\)\+\\frac\{\\log\(1/\\delta\)\}\{\\lambda\}\.\(B\.23\)Choose

λ=12​H​\(1\+log⁡\(2​T\+A\)\)≤1H\.\\displaystyle\\lambda=\\frac\{1\}\{2H\(1\+\\log\(2T\+A\)\)\}\\leq\\frac\{1\}\{H\}\.We haveH⁡\(1\+LT\)​λ=1/2H\(1\+L\_\{T\}\)\\lambda=1/2\. Rearranging \([B\.23](https://arxiv.org/html/2609.38666#A2.E23)\) yields

Regret⁡\(T\)≤8​S​AH​log​T\+4​H​\(1\+log⁡\(2​T\+A\)\)​log​1δ\.\\displaystyle\\operatorname\{Regret\}\(T\)\\leq 8SA^\{H\}\\log T\+4H\\bigl\(1\+\\log\(2T\+A\)\\bigr\)\\log\\frac\{1\}\{\\delta\}\.∎

## Appendix CProof of Theorem[5\.5](https://arxiv.org/html/2609.38666#S5.Thmtheorem5)

###### Proof of Theorem[5\.5](https://arxiv.org/html/2609.38666#S5.Thmtheorem5)\.

We first control the error of the optimistic token\-score estimates\. Letℱt\\mathcal\{F\}\_\{t\}denote the history before roundtt\. In this way, bothl^t−1\\widehat\{l\}\_\{t\-1\}andπ^t\\widehat\{\\pi\}\_\{t\}areℱt\\mathcal\{F\}\_\{t\}\-measurable\.

Letℰ\\mathcal\{E\}be the event in Lemma[5\.4](https://arxiv.org/html/2609.38666#S5.Thmtheorem4), which satisfiesℙ⁡\(ℰ\)≥1−2​δ\\mathbb\{P\}\(\\mathcal\{E\}\)\\geq 1\-2\\delta\. On this event, simultaneously for every roundssand every visited triple\(x,u,a\)\(x,u,a\),

\|l¯s​\(x,u,a\)−log⁡gh​\(a\|x,u\)\|≤βs​\(x,u,a\)\.\\displaystyle\\left\|\\bar\{l\}\_\{s\}\(x,u,a\)\-\\log g\_\{h\}\(a\|x,u\)\\right\|\\leq\\beta\_\{s\}\(x,u,a\)\.\(C\.1\)Moreover, Assumption[5\.3](https://arxiv.org/html/2609.38666#S5.Thmtheorem3)and the definition ofghg\_\{h\}imply that for every teacherii,

\|log⁡pi​\(a\|x,u\)−log⁡gh​\(a\|x,u\)\|\\displaystyle\\big\|\\log p\_\{i\}\(a\|x,u\)\-\\log g\_\{h\}\(a\|x,u\)\\big\|=\|log⁡pi​\(a\|x,u\)πref​\(a\|x,u\)−∑k∈ℐwk​\(x\)​log⁡pk​\(a\|x,u\)πref​\(a\|x,u\)\|≤2​B\.\\displaystyle=\\bigg\|\\log\\frac\{p\_\{i\}\(a\|x,u\)\}\{\\pi\_\{\\text\{ref\}\}\(a\|x,u\)\}\-\\sum\_\{k\\in\\mathcal\{I\}\}w\_\{k\}\(x\)\\log\\frac\{p\_\{k\}\(a\|x,u\)\}\{\\pi\_\{\\text\{ref\}\}\(a\|x,u\)\}\\bigg\|\\leq 2B\.Sincel¯s​\(x,u,a\)\\bar\{l\}\_\{s\}\(x,u,a\)is an average of the observed teacher log\-probabilities at this triple, it follows that

\|l¯s​\(x,u,a\)−log⁡gh​\(a\|x,u\)\|≤2​B\.\\displaystyle\\big\|\\bar\{l\}\_\{s\}\(x,u,a\)\-\\log g\_\{h\}\(a\|x,u\)\\big\|\\leq 2B\.\(C\.2\)Combining \([C\.1](https://arxiv.org/html/2609.38666#A3.E1)\) and \([C\.2](https://arxiv.org/html/2609.38666#A3.E2)\), we obtain

\|l¯s​\(x,u,a\)−log⁡gh​\(a\|x,u\)\|≤min⁡\{βs​\(x,u,a\),2​B\}\.\\displaystyle\\big\|\\bar\{l\}\_\{s\}\(x,u,a\)\-\\log g\_\{h\}\(a\|x,u\)\\big\|\\leq\\min\\big\\\{\\beta\_\{s\}\(x,u,a\),2B\\big\\\}\.Consequently, on the eventℰ\\mathcal\{E\}, the optimistic estimatel^s=l¯s\+min⁡\{βs,2​B\}\\widehat\{l\}\_\{s\}=\\bar\{l\}\_\{s\}\+\\min\\\{\\beta\_\{s\},2B\\\}satisfies

0≤l^s​\(x,u,a\)−log⁡gh​\(a\|x,u\)≤min⁡\{2​βs​\(x,u,a\),4​B\}\.\\displaystyle 0\\leq\\widehat\{l\}\_\{s\}\(x,u,a\)\-\\log g\_\{h\}\(a\|x,u\)\\leq\\min\\big\\\{2\\beta\_\{s\}\(x,u,a\),4B\\big\\\}\.For an unvisited triple, the algorithm sets

l^s​\(x,u,a\)=log⁡πref​\(a\|x,u\)\+B\.\\displaystyle\\widehat\{l\}\_\{s\}\(x,u,a\)=\\log\\pi\_\{\\text\{ref\}\}\(a\|x,u\)\+B\.Assumption[5\.3](https://arxiv.org/html/2609.38666#S5.Thmtheorem3)implies

−B≤log⁡gh​\(a\|x,u\)−log⁡πref​\(a\|x,u\)≤B\.\\displaystyle\-B\\leq\\log g\_\{h\}\(a\|x,u\)\-\\log\\pi\_\{\\text\{ref\}\}\(a\|x,u\)\\leq B\.Therefore,

0≤l^s​\(x,u,a\)−log⁡gh​\(a\|x,u\)≤2​B\.\\displaystyle 0\\leq\\widehat\{l\}\_\{s\}\(x,u,a\)\-\\log g\_\{h\}\(a\|x,u\)\\leq 2B\.We next consider the output policy\. Multiplying the token conditionals defined by the backward recursion gives

π^t​\(y\|x\)\\displaystyle\\widehat\{\\pi\}\_\{t\}\(y\|x\)=∏h=1Hexp⁡\(l^t−1​\(x,y<h,ah\)\)​V^t−1,h\+1​\(x,y≤h\)V^t−1,h​\(x,y<h\)\\displaystyle=\\prod\_\{h=1\}^\{H\}\\frac\{\\exp\\bigl\(\\widehat\{l\}\_\{t\-1\}\(x,y\_\{<h\},a\_\{h\}\)\\bigr\)\\widehat\{V\}\_\{t\-1,h\+1\}\(x,y\_\{\\leq h\}\)\}\{\\widehat\{V\}\_\{t\-1,h\}\(x,y\_\{<h\}\)\}=exp⁡\(∑h=1Hl^t−1​\(x,y<h,ah\)\)V^t−1,1​\(x\),\\displaystyle=\\frac\{\\exp\\left\(\\sum\_\{h=1\}^\{H\}\\widehat\{l\}\_\{t\-1\}\(x,y\_\{<h\},a\_\{h\}\)\\right\)\}\{\\widehat\{V\}\_\{t\-1,1\}\(x\)\},\(C\.3\)where the second equation holds by canceling out the telescoping terms andV^t−1,H\+1=1\\widehat\{V\}\_\{t\-1,H\+1\}=1\. Comparing this expression with the reverse target in \([5\.1](https://arxiv.org/html/2609.38666#S5.E1)\), we obtain

log⁡π^t​\(y\|x\)πR∗​\(y\|x\)\\displaystyle\\log\\frac\{\\widehat\{\\pi\}\_\{t\}\(y\|x\)\}\{\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(y\|x\)\}=∑h=1H\[l^t−1​\(x,y<h,ah\)−log⁡gh​\(ah\|x,y<h\)\]\+log⁡V1​\(x\)V^t−1,1​\(x\)\.\\displaystyle=\\sum\_\{h=1\}^\{H\}\\Big\[\\widehat\{l\}\_\{t\-1\}\(x,y\_\{<h\},a\_\{h\}\)\-\\log g\_\{h\}\(a\_\{h\}\|x,y\_\{<h\}\)\\Big\]\+\\log\\frac\{V\_\{1\}\(x\)\}\{\\widehat\{V\}\_\{t\-1,1\}\(x\)\}\.\(C\.4\)Moreover, the ratio of the normalizing constants can be expressed as an expectation underπ^t\\widehat\{\\pi\}\_\{t\}\. To see this, first expand the expectation:

𝔼y∼π^t\(⋅\|x\)exp\(−∑h=1H\[l^t−1\(x,y<h,ah\)−loggh\(ah\|x,y<h\)\]\)\\displaystyle\\mathbb\{E\}\_\{y\\sim\\widehat\{\\pi\}\_\{t\}\(\\cdot\|x\)\}\\exp\\bigg\(\-\\sum\_\{h=1\}^\{H\}\\Big\[\\widehat\{l\}\_\{t\-1\}\(x,y\_\{<h\},a\_\{h\}\)\-\\log g\_\{h\}\(a\_\{h\}\|x,y\_\{<h\}\)\\Big\]\\bigg\)=∑y∈𝒴π^t\(y\|x\)exp\(−∑h=1Hl^t−1\(x,y<h,ah\)\+∑h=1Hloggh\(ah\|x,y<h\)\)\.\\displaystyle\\quad=\\sum\_\{y\\in\\mathcal\{Y\}\}\\widehat\{\\pi\}\_\{t\}\(y\|x\)\\exp\\bigg\(\-\\sum\_\{h=1\}^\{H\}\\widehat\{l\}\_\{t\-1\}\(x,y\_\{<h\},a\_\{h\}\)\+\\sum\_\{h=1\}^\{H\}\\log g\_\{h\}\(a\_\{h\}\|x,y\_\{<h\}\)\\bigg\)\.\(C\.5\)Substituting \([C\.3](https://arxiv.org/html/2609.38666#A3.E3)\) into \([C\.5](https://arxiv.org/html/2609.38666#A3.E5)\), the right\-hand side becomes

∑y∈𝒴exp⁡\(∑h=1Hl^t−1​\(x,y<h,ah\)\)V^t−1,1​\(x\)exp\(−∑h=1Hl^t−1\(x,y<h,ah\)\+∑h=1Hloggh\(ah\|x,y<h\)\)\\displaystyle\\sum\_\{y\\in\\mathcal\{Y\}\}\\frac\{\\exp\\big\(\\sum\_\{h=1\}^\{H\}\\widehat\{l\}\_\{t\-1\}\(x,y\_\{<h\},a\_\{h\}\)\\big\)\}\{\\widehat\{V\}\_\{t\-1,1\}\(x\)\}\\exp\\bigg\(\-\\sum\_\{h=1\}^\{H\}\\widehat\{l\}\_\{t\-1\}\(x,y\_\{<h\},a\_\{h\}\)\+\\sum\_\{h=1\}^\{H\}\\log g\_\{h\}\(a\_\{h\}\|x,y\_\{<h\}\)\\bigg\)=1V^t−1,1​\(x\)​∑y∈𝒴exp⁡\(∑h=1Hlog⁡gh​\(ah\|x,y<h\)\)\\displaystyle\\quad=\\frac\{1\}\{\\widehat\{V\}\_\{t\-1,1\}\(x\)\}\\sum\_\{y\\in\\mathcal\{Y\}\}\\exp\\bigg\(\\sum\_\{h=1\}^\{H\}\\log g\_\{h\}\(a\_\{h\}\|x,y\_\{<h\}\)\\bigg\)=1V^t−1,1​\(x\)​∑y∈𝒴∏h=1Hgh​\(ah\|x,y<h\)\.\\displaystyle\\quad=\\frac\{1\}\{\\widehat\{V\}\_\{t\-1,1\}\(x\)\}\\sum\_\{y\\in\\mathcal\{Y\}\}\\prod\_\{h=1\}^\{H\}g\_\{h\}\(a\_\{h\}\|x,y\_\{<h\}\)\.Finally, recall that in Theorem[5\.1](https://arxiv.org/html/2609.38666#S5.Thmtheorem1), we have seen

V1​\(x\)=∑y∈𝒴∏h=1Hgh​\(ah\|x,y<h\)\.\\displaystyle V\_\{1\}\(x\)=\\sum\_\{y\\in\\mathcal\{Y\}\}\\prod\_\{h=1\}^\{H\}g\_\{h\}\(a\_\{h\}\|x,y\_\{<h\}\)\.Consequently,

𝔼y∼π^t\(⋅\|x\)exp\(−∑h=1H\[l^t−1\(x,y<h,ah\)−loggh\(ah\|x,y<h\)\]\)=V1​\(x\)V^t−1,1​\(x\)\.\\displaystyle\\mathbb\{E\}\_\{y\\sim\\widehat\{\\pi\}\_\{t\}\(\\cdot\|x\)\}\\exp\\bigg\(\-\\sum\_\{h=1\}^\{H\}\\Big\[\\widehat\{l\}\_\{t\-1\}\(x,y\_\{<h\},a\_\{h\}\)\-\\log g\_\{h\}\(a\_\{h\}\|x,y\_\{<h\}\)\\Big\]\\bigg\)=\\frac\{V\_\{1\}\(x\)\}\{\\widehat\{V\}\_\{t\-1,1\}\(x\)\}\.\(C\.6\)Taking expectation overπ^t\(⋅\|x\)\\widehat\{\\pi\}\_\{t\}\(\\cdot\|x\)and substituting \([C\.6](https://arxiv.org/html/2609.38666#A3.E6)\) into \([C\.4](https://arxiv.org/html/2609.38666#A3.E4)\) yields

KL\(π^t\(⋅\|x\)∥πR∗\(⋅\|x\)\)=𝔼y∼π^t\(⋅\|x\)\[∑h=1H\[l^t−1\(x,y<h,ah\)−loggh\(ah\|x,y<h\)\]\]\\displaystyle\\text\{KL\}\\bigl\(\\widehat\{\\pi\}\_\{t\}\(\\cdot\|x\)\\,\\big\\\|\\,\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(\\cdot\|x\)\\bigr\)=\\mathbb\{E\}\_\{y\\sim\\widehat\{\\pi\}\_\{t\}\(\\cdot\|x\)\}\\bigg\[\\sum\_\{h=1\}^\{H\}\\Big\[\\widehat\{l\}\_\{t\-1\}\(x,y\_\{<h\},a\_\{h\}\)\-\\log g\_\{h\}\(a\_\{h\}\|x,y\_\{<h\}\)\\Big\]\\bigg\]\+log𝔼y∼π^t\(⋅\|x\)exp\(−∑h=1H\[l^t−1\(x,y<h,ah\)−loggh\(ah\|x,y<h\)\]\)\.\\displaystyle\\quad\+\\log\\mathbb\{E\}\_\{y\\sim\\widehat\{\\pi\}\_\{t\}\(\\cdot\|x\)\}\\exp\\bigg\(\-\\sum\_\{h=1\}^\{H\}\\Big\[\\widehat\{l\}\_\{t\-1\}\(x,y\_\{<h\},a\_\{h\}\)\-\\log g\_\{h\}\(a\_\{h\}\|x,y\_\{<h\}\)\\Big\]\\bigg\)\.\(C\.7\)On the eventℰ\\mathcal\{E\},l^t−1​\(x,y<h,ah\)≥log⁡gh​\(ah\|x,y<h\)\\widehat\{l\}\_\{t\-1\}\(x,y\_\{<h\},a\_\{h\}\)\\geq\\log g\_\{h\}\(a\_\{h\}\|x,y\_\{<h\}\)for anyt,ht,h\. We can therefore apply the basic inequalitye−v≤1−v\+v2/2e^\{\-v\}\\leq 1\-v\+v^\{2\}/2forv≥0v\\geq 0\. LetV=∑h=1H\[l^t−1​\(x,y<h,ah\)−log⁡gh​\(ah\|x,y<h\)\]V=\\sum\_\{h=1\}^\{H\}\\big\[\\widehat\{l\}\_\{t\-1\}\(x,y\_\{<h\},a\_\{h\}\)\-\\log g\_\{h\}\(a\_\{h\}\|x,y\_\{<h\}\)\\big\]\. The right\-hand side can be bounded as

𝔼⁡\[V\]\+log⁡𝔼​exp⁡\(−V\)\\displaystyle\\mathbb\{E\}\[V\]\+\\log\\mathbb\{E\}\\exp\(\-V\)≤𝔼⁡\[V\]\+log⁡\(1−𝔼⁡\[V\]\+𝔼⁡\[V2\]/2\)\\displaystyle\\leq\\mathbb\{E\}\[V\]\+\\log\\big\(1\-\\mathbb\{E\}\[V\]\+\\mathbb\{E\}\[V^\{2\}\]/2\\big\)≤𝔼⁡\[V\]−𝔼⁡\[V\]\+12​𝔼​\[V2\]\\displaystyle\\leq\\mathbb\{E\}\[V\]\-\\mathbb\{E\}\[V\]\+\\frac\{1\}\{2\}\\mathbb\{E\}\[V^\{2\}\]=12​𝔼​\[V2\],\\displaystyle=\\frac\{1\}\{2\}\\mathbb\{E\}\[V^\{2\}\],where we uselog⁡x≤x−1\\log x\\leq x\-1for anyx\>0x\>0\. As a result, we have

KL\(π^t\(⋅\|x\)∥πR∗\(⋅\|x\)\)\\displaystyle\\text\{KL\}\\bigl\(\\widehat\{\\pi\}\_\{t\}\(\\cdot\|x\)\\,\\big\\\|\\,\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(\\cdot\|x\)\\bigr\)≤12𝔼y∼π^t\(⋅\|x\)\[\(∑h=1H\[l^t−1\(x,y<h,ah\)−loggh\(ah\|x,y<h\)\]\)2\]\\displaystyle\\leq\\frac\{1\}\{2\}\\mathbb\{E\}\_\{y\\sim\\widehat\{\\pi\}\_\{t\}\(\\cdot\|x\)\}\\bigg\[\\bigg\(\\sum\_\{h=1\}^\{H\}\\Big\[\\widehat\{l\}\_\{t\-1\}\(x,y\_\{<h\},a\_\{h\}\)\-\\log g\_\{h\}\(a\_\{h\}\|x,y\_\{<h\}\)\\Big\]\\bigg\)^\{2\}\\bigg\]≤H2𝔼y∼π^t\(⋅\|x\)\[∑h=1H\[l^t−1\(x,y<h,ah\)−loggh\(ah\|x,y<h\)\]2\],\\displaystyle\\quad\\leq\\frac\{H\}\{2\}\\mathbb\{E\}\_\{y\\sim\\widehat\{\\pi\}\_\{t\}\(\\cdot\|x\)\}\\bigg\[\\sum\_\{h=1\}^\{H\}\\Big\[\\widehat\{l\}\_\{t\-1\}\(x,y\_\{<h\},a\_\{h\}\)\-\\log g\_\{h\}\(a\_\{h\}\|x,y\_\{<h\}\)\\Big\]^\{2\}\\bigg\],\(C\.8\)where the last inequality holds due to the Cauchy–Schwarz inequality\.

We next relate the expected squared errors in \([C\.8](https://arxiv.org/html/2609.38666#A3.E8)\) to those evaluated on the observed rollouts\. Define

Xt:=H2​m​∑j=1m∑h=1H\[l^t−1​\(xt,j,yt,j,<h,at,j,h\)−log⁡gh​\(at,j,h\|xt,j,yt,j,<h\)\]2\.\\displaystyle X\_\{t\}:=\\frac\{H\}\{2m\}\\sum\_\{j=1\}^\{m\}\\sum\_\{h=1\}^\{H\}\\Big\[\\widehat\{l\}\_\{t\-1\}\(x\_\{t,j\},y\_\{t,j,<h\},a\_\{t,j,h\}\)\-\\log g\_\{h\}\(a\_\{t,j,h\}\|x\_\{t,j\},y\_\{t,j,<h\}\)\\Big\]^\{2\}\.Conditionally onℱt\\mathcal\{F\}\_\{t\}, each rollout has joint distribution

ℙ⁡\(xt,j=x,yt,j=y∣ℱt\)\\displaystyle\\mathbb\{P\}\(x\_\{t,j\}=x,y\_\{t,j\}=y\\mid\\mathcal\{F\}\_\{t\}\)=1I​∑i∈ℐρi​\(x\)​π^t​\(y\|x\)\\displaystyle=\\frac\{1\}\{I\}\\sum\_\{i\\in\\mathcal\{I\}\}\\rho\_\{i\}\(x\)\\widehat\{\\pi\}\_\{t\}\(y\|x\)=ρ¯​\(x\)​π^t​\(y\|x\)\.\\displaystyle=\\bar\{\\rho\}\(x\)\\widehat\{\\pi\}\_\{t\}\(y\|x\)\.Sincel^t−1\\widehat\{l\}\_\{t\-1\}andπ^t\\widehat\{\\pi\}\_\{t\}areℱt\\mathcal\{F\}\_\{t\}\-measurable, linearity of conditional expectation therefore gives

𝔼⁡\[Xt∣ℱt\]\\displaystyle\\mathbb\{E\}\[X\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]=H2​m∑j=1m𝔼x∼ρ¯𝔼y∼π^t\(⋅\|x\)\[∑h=1H\[l^t−1\(x,y<h,ah\)−loggh\(ah\|x,y<h\)\]2\]\\displaystyle=\\frac\{H\}\{2m\}\\sum\_\{j=1\}^\{m\}\\mathbb\{E\}\_\{x\\sim\\bar\{\\rho\}\}\\mathbb\{E\}\_\{y\\sim\\widehat\{\\pi\}\_\{t\}\(\\cdot\|x\)\}\\bigg\[\\sum\_\{h=1\}^\{H\}\\Big\[\\widehat\{l\}\_\{t\-1\}\(x,y\_\{<h\},a\_\{h\}\)\-\\log g\_\{h\}\(a\_\{h\}\|x,y\_\{<h\}\)\\Big\]^\{2\}\\bigg\]=H2𝔼x∼ρ¯𝔼y∼π^t\(⋅\|x\)\[∑h=1H\[l^t−1\(x,y<h,ah\)−loggh\(ah\|x,y<h\)\]2\]\.\\displaystyle=\\frac\{H\}\{2\}\\mathbb\{E\}\_\{x\\sim\\bar\{\\rho\}\}\\mathbb\{E\}\_\{y\\sim\\widehat\{\\pi\}\_\{t\}\(\\cdot\|x\)\}\\bigg\[\\sum\_\{h=1\}^\{H\}\\Big\[\\widehat\{l\}\_\{t\-1\}\(x,y\_\{<h\},a\_\{h\}\)\-\\log g\_\{h\}\(a\_\{h\}\|x,y\_\{<h\}\)\\Big\]^\{2\}\\bigg\]\.\(C\.9\)Note that this does not require independence among the rollouts within a round\. On the eventℰ\\mathcal\{E\}, combining \([C\.9](https://arxiv.org/html/2609.38666#A3.E9)\) with \([C\.8](https://arxiv.org/html/2609.38666#A3.E8)\), we have

Regret⁡\(T\)\\displaystyle\\operatorname\{Regret\}\(T\)=∑t=1T𝔼x∼ρ¯KL\(π^t\(⋅\|x\)∥πR∗\(⋅\|x\)\)≤∑t=1T𝔼\[Xt∣ℱt\]\.\\displaystyle=\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\_\{x\\sim\\bar\{\\rho\}\}\\text\{KL\}\\bigl\(\\widehat\{\\pi\}\_\{t\}\(\\cdot\|x\)\\,\\big\\\|\\,\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(\\cdot\|x\)\\bigr\)\\leq\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\[X\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]\.\(C\.10\)To apply Lemma[I\.6](https://arxiv.org/html/2609.38666#A9.Thmtheorem6), we need a bound onXtX\_\{t\}\. At every previously visited triple, \([C\.2](https://arxiv.org/html/2609.38666#A3.E2)\) and the definition ofl^t−1\\widehat\{l\}\_\{t\-1\}give

\|l^t−1​\(x,u,a\)−log⁡gh​\(a\|x,u\)\|\\displaystyle\\big\|\\widehat\{l\}\_\{t\-1\}\(x,u,a\)\-\\log g\_\{h\}\(a\|x,u\)\\big\|≤\|l¯t−1​\(x,u,a\)−log⁡gh​\(a\|x,u\)\|\+min⁡\{βt−1​\(x,u,a\),2​B\}\\displaystyle\\leq\\big\|\\bar\{l\}\_\{t\-1\}\(x,u,a\)\-\\log g\_\{h\}\(a\|x,u\)\\big\|\+\\min\\\{\\beta\_\{t\-1\}\(x,u,a\),2B\\\}≤4​B\.\\displaystyle\\leq 4B\.At an unvisited triple,l^t−1​\(x,u,a\)=log⁡πref​\(a\|x,u\)\+B\\widehat\{l\}\_\{t\-1\}\(x,u,a\)=\\log\\pi\_\{\\text\{ref\}\}\(a\|x,u\)\+Bindicates\|l^t−1​\(x,u,a\)−log⁡gh​\(a\|x,u\)\|∈\[0,2​B\]\\big\|\\widehat\{l\}\_\{t\-1\}\(x,u,a\)\-\\log g\_\{h\}\(a\|x,u\)\\big\|\\in\[0,2B\]\. Consequently, in both cases,

0≤Xt≤H2​m⋅m​H⋅\(4​B\)2=8​H2​B2\.\\displaystyle 0\\leq X\_\{t\}\\leq\\frac\{H\}\{2m\}\\cdot mH\\cdot\(4B\)^\{2\}=8H^\{2\}B^\{2\}\.Moreover,XtX\_\{t\}isℱt\+1\\mathcal\{F\}\_\{t\+1\}\-measurable\. Applying Lemma[I\.6](https://arxiv.org/html/2609.38666#A9.Thmtheorem6), with probability at least1−δ1\-\\delta, the following inequality holds

∑t=1T𝔼⁡\[Xt∣ℱt\]≤2​∑t=1TXt\+64​H2​B2​log⁡2δ\.\\displaystyle\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\[X\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]\\leq 2\\sum\_\{t=1\}^\{T\}X\_\{t\}\+64H^\{2\}B^\{2\}\\log\\frac\{2\}\{\\delta\}\.\(C\.11\)Substituting \([C\.11](https://arxiv.org/html/2609.38666#A3.E11)\) into \([C\.10](https://arxiv.org/html/2609.38666#A3.E10)\) and taking a union bound, we conclude that, with probability at least1−3​δ1\-3\\delta,

Regret⁡\(T\)\\displaystyle\\operatorname\{Regret\}\(T\)≤Hm​∑t=1T∑j=1m∑h=1H\[l^t−1​\(xt,j,yt,j,<h,at,j,h\)−log⁡gh​\(at,j,h\|xt,j,yt,j,<h\)\]2⏟I\\displaystyle\\leq\\frac\{H\}\{m\}\\underbrace\{\\sum\_\{t=1\}^\{T\}\\sum\_\{j=1\}^\{m\}\\sum\_\{h=1\}^\{H\}\\Big\[\\widehat\{l\}\_\{t\-1\}\(x\_\{t,j\},y\_\{t,j,<h\},a\_\{t,j,h\}\)\-\\log g\_\{h\}\(a\_\{t,j,h\}\|x\_\{t,j\},y\_\{t,j,<h\}\)\\Big\]^\{2\}\}\_\{I\}\+64​H2​B2​log⁡2δ\.\\displaystyle\\qquad\+64H^\{2\}B^\{2\}\\log\\frac\{2\}\{\\delta\}\.\(C\.12\)It remains to boundIIfor every data trajectory onℰ\\mathcal\{E\}\. Regrouping the observations by their corresponding triples gives

I=∑h=1H∑\(x,u,a\)∑t=1T\[Nt​\(x,u,a\)−Nt−1​\(x,u,a\)\]​\[l^t−1​\(x,u,a\)−log⁡gh​\(a\|x,u\)\]2\.\\displaystyle I=\\sum\_\{h=1\}^\{H\}\\sum\_\{\(x,u,a\)\}\\sum\_\{t=1\}^\{T\}\\big\[N\_\{t\}\(x,u,a\)\-N\_\{t\-1\}\(x,u,a\)\\big\]\\Big\[\\widehat\{l\}\_\{t\-1\}\(x,u,a\)\-\\log g\_\{h\}\(a\|x,u\)\\Big\]^\{2\}\.HereNt​\(x,u,a\)−Nt−1​\(x,u,a\)N\_\{t\}\(x,u,a\)\-N\_\{t\-1\}\(x,u,a\)counts the visits to this triple in roundtt, each of which contributes the same squared error\. Since there aremmsamples per round,

0≤Nt​\(x,u,a\)−Nt−1​\(x,u,a\)≤m,NT​\(x,u,a\)≤m​T\.\\displaystyle 0\\leq N\_\{t\}\(x,u,a\)\-N\_\{t\-1\}\(x,u,a\)\\leq m,\\qquad N\_\{T\}\(x,u,a\)\\leq mT\.For each fixed triple\(x,u,a\)\(x,u,a\), we split the inner sum according to whetherNt−1​\(x,u,a\)=0N\_\{t\-1\}\(x,u,a\)=0orNt−1​\(x,u,a\)≥1N\_\{t\-1\}\(x,u,a\)\\geq 1\. In the former case, only the first round that visits this triple can contribute, with at mostmmvisits\. Hence,

∑t∈\[T\]:Nt−1​\(x,u,a\)=0\[Nt\(x,u,a\)−Nt−1\(x,u,a\)\]\[l^t−1\(x,u,a\)−loggh\(a\|x,u\)\]2≤4mB2\.\\displaystyle\\sum\_\{\\begin\{subarray\}\{c\}t\\in\[T\]:\\\\ N\_\{t\-1\}\(x,u,a\)=0\\end\{subarray\}\}\\big\[N\_\{t\}\(x,u,a\)\-N\_\{t\-1\}\(x,u,a\)\\big\]\\Big\[\\widehat\{l\}\_\{t\-1\}\(x,u,a\)\-\\log g\_\{h\}\(a\|x,u\)\\Big\]^\{2\}\\leq 4mB^\{2\}\.Summing this bound over allK:=S​∑h=1HAhK:=S\\sum\_\{h=1\}^\{H\}A^\{h\}triples and applying Lemma[5\.4](https://arxiv.org/html/2609.38666#S5.Thmtheorem4)gives

I≤4​m​B2​K\\displaystyle I\\leq 4mB^\{2\}K\+∑h=1H∑\(x,u,a\)∑t∈\[T\]:Nt−1​\(x,u,a\)≥1\[Nt\(x,u,a\)−Nt−1\(x,u,a\)\]⋅min\{16B2,4βt−1\(x,u,a\)2\}\.\\displaystyle\\qquad\+\\sum\_\{h=1\}^\{H\}\\sum\_\{\(x,u,a\)\}\\sum\_\{\\begin\{subarray\}\{c\}t\\in\[T\]:\\\\ N\_\{t\-1\}\(x,u,a\)\\geq 1\\end\{subarray\}\}\\big\[N\_\{t\}\(x,u,a\)\-N\_\{t\-1\}\(x,u,a\)\\big\]\\cdot\\min\\Big\\\{16B^\{2\},\\,4\\beta\_\{t\-1\}\(x,u,a\)^\{2\}\\Big\\\}\.\(C\.13\)Recall that

βt​\(x,u,a\)=O~​\(B​m​HNt​\(x,u,a\)\+B​m​HNt​\(x,u,a\)\)\.\\displaystyle\\beta\_\{t\}\(x,u,a\)=\\widetilde\{O\}\\bigg\(B\\sqrt\{\\frac\{mH\}\{N\_\{t\}\(x,u,a\)\}\}\+\\frac\{BmH\}\{N\_\{t\}\(x,u,a\)\}\\bigg\)\.ForNt−1​\(x,u,a\)≥1N\_\{t\-1\}\(x,u,a\)\\geq 1, squaring the confidence radius gives

4​βt−1​\(x,u,a\)2≤O~​\(B2​m​HNt−1​\(x,u,a\)\+B2​m2​H2Nt−1​\(x,u,a\)2\),\\displaystyle 4\\beta\_\{t\-1\}\(x,u,a\)^\{2\}\\leq\\widetilde\{O\}\\left\(\\frac\{B^\{2\}mH\}\{N\_\{t\-1\}\(x,u,a\)\}\+\\frac\{B^\{2\}m^\{2\}H^\{2\}\}\{N\_\{t\-1\}\(x,u,a\)^\{2\}\}\\right\),where we used\(v\+w\)2≤2​v2\+2​w2\(v\+w\)^\{2\}\\leq 2v^\{2\}\+2w^\{2\}\. Combining this with the constant bound16​B216B^\{2\}, we obtain

min⁡\{16​B2,4​βt−1​\(x,u,a\)2\}\\displaystyle\\min\\Big\\\{16B^\{2\},\\,4\\beta\_\{t\-1\}\(x,u,a\)^\{2\}\\Big\\\}≤O~​\(B2​min⁡\{1,m​HNt−1​\(x,u,a\)\+m2​H2Nt−1​\(x,u,a\)2\}\)\\displaystyle\\leq\\widetilde\{O\}\\bigg\(B^\{2\}\\min\\bigg\\\{1,\\,\\frac\{mH\}\{N\_\{t\-1\}\(x,u,a\)\}\+\\frac\{m^\{2\}H^\{2\}\}\{N\_\{t\-1\}\(x,u,a\)^\{2\}\}\\bigg\\\}\\bigg\)≤O~​\(B2​min⁡\{1,m​HNt−1​\(x,u,a\)\}\),\\displaystyle\\qquad\\leq\\widetilde\{O\}\\bigg\(B^\{2\}\\min\\bigg\\\{1,\\frac\{mH\}\{N\_\{t\-1\}\(x,u,a\)\}\\bigg\\\}\\bigg\),where we usemin⁡\{1,v\+v2\}≤2​min⁡\{1,v\}\\min\\\{1,v\+v^\{2\}\\\}\\leq 2\\min\\\{1,v\\\}sincev2≤vv^\{2\}\\leq vwhen0≤v≤10\\leq v\\leq 1\.

Fix a triple\(x,u,a\)\(x,u,a\)withNT​\(x,u,a\)≥1N\_\{T\}\(x,u,a\)\\geq 1\. Order its visits chronologically, breaking ties within each round arbitrarily\. If itskk\-th visit occurs in roundtt, then

k≤Nt−1​\(x,u,a\)\+m≤Nt−1​\(x,u,a\)\+m​H,\\displaystyle k\\leq N\_\{t\-1\}\(x,u,a\)\+m\\leq N\_\{t\-1\}\(x,u,a\)\+mH,because at mostmmvisits occur in that round andH≥1H\\geq 1\. For each visit included in the remaining sum,Nt−1​\(x,u,a\)≥1N\_\{t\-1\}\(x,u,a\)\\geq 1, and hence

min⁡\{1,m​HNt−1​\(x,u,a\)\}\\displaystyle\\min\\left\\\{1,\\frac\{mH\}\{N\_\{t\-1\}\(x,u,a\)\}\\right\\\}≤2​m​HNt−1​\(x,u,a\)\+m​H\\displaystyle\\leq\\frac\{2mH\}\{N\_\{t\-1\}\(x,u,a\)\+mH\}≤2​m​Hk\.\\displaystyle\\leq\\frac\{2mH\}\{k\}\.The first inequality usesmin⁡\{1,v\}≤2​v/\(1\+v\)\\min\\\{1,v\\\}\\leq 2v/\(1\+v\)forv≥0v\\geq 0\. SinceNt​\(x,u,a\)−Nt−1​\(x,u,a\)N\_\{t\}\(x,u,a\)\-N\_\{t\-1\}\(x,u,a\)counts the visits in roundtt, the following summation counts all the visits to\(x,u,a\)\(x,u,a\)and thus can be indexed bykk, leading to

∑t∈\[T\]:Nt−1​\(x,u,a\)≥1\[Nt\(x,u,a\)−Nt−1\(x,u,a\)\]min\{1,m​HNt−1​\(x,u,a\)\}\\displaystyle\\sum\_\{\\begin\{subarray\}\{c\}t\\in\[T\]:\\\\ N\_\{t\-1\}\(x,u,a\)\\geq 1\\end\{subarray\}\}\\big\[N\_\{t\}\(x,u,a\)\-N\_\{t\-1\}\(x,u,a\)\\big\]\\min\\bigg\\\{1,\\frac\{mH\}\{N\_\{t\-1\}\(x,u,a\)\}\\bigg\\\}≤2​m​H​∑k=1NT​\(x,u,a\)1k\\displaystyle\\qquad\\leq 2mH\\sum\_\{k=1\}^\{N\_\{T\}\(x,u,a\)\}\\frac\{1\}\{k\}≤2​m​H​\[1\+log⁡NT​\(x,u,a\)\]\\displaystyle\\qquad\\leq 2mH\\bigl\[1\+\\log N\_\{T\}\(x,u,a\)\\bigr\]≤2​m​H​log⁡\(e​m​T\),\\displaystyle\\qquad\\leq 2mH\\log\(emT\),\(C\.14\)where we use∑k=1T1/k≤1\+log⁡T\\sum\_\{k=1\}^\{T\}1/k\\leq 1\+\\log TandNT​\(x,u,a\)≤m​TN\_\{T\}\(x,u,a\)\\leq mT\.

Substituting \([C\.14](https://arxiv.org/html/2609.38666#A3.E14)\) into \([C\.13](https://arxiv.org/html/2609.38666#A3.E13)\) and summing over all\(x,u,a\)\(x,u,a\)triples, we have

I\\displaystyle I≤4​m​B2​K\+O~​\(B2​m​H​K​log⁡\(e​m​T\)\)=O~​\(B2​m​H​K​log⁡T\)=O~​\(B2​m​S​H​AH​log⁡T\)\.\\displaystyle\\leq 4mB^\{2\}K\+\\widetilde\{O\}\\bigl\(B^\{2\}mHK\\log\(emT\)\\bigr\)=\\widetilde\{O\}\\bigl\(B^\{2\}mHK\\log T\\bigr\)=\\widetilde\{O\}\\bigl\(B^\{2\}mSHA^\{H\}\\log T\\bigr\)\.In conclusion, \([C\.12](https://arxiv.org/html/2609.38666#A3.E12)\) gives that with probability at least1−3​δ1\-3\\delta, the following inequality holds:

Regret⁡\(T\)\\displaystyle\\operatorname\{Regret\}\(T\)≤Hm​I\+O~​\(B2​H2\)\\displaystyle\\leq\\frac\{H\}\{m\}I\+\\widetilde\{O\}\(B^\{2\}H^\{2\}\)≤O~​\(B2​H2​S​AH​log⁡T\)\.\\displaystyle\\leq\\widetilde\{O\}\\big\(B^\{2\}H^\{2\}SA^\{H\}\\log T\\big\)\.This completes the proof of Theorem[5\.5](https://arxiv.org/html/2609.38666#S5.Thmtheorem5)\. ∎

## Appendix DProof of Lemmas Used in Appendices[B](https://arxiv.org/html/2609.38666#A2)and[C](https://arxiv.org/html/2609.38666#A3)

### D\.1Proof of Lemma[B\.1](https://arxiv.org/html/2609.38666#A2.Thmtheorem1)

###### Proof of Lemma[B\.1](https://arxiv.org/html/2609.38666#A2.Thmtheorem1)\.

WhenWT=0W\_\{T\}=0, the inequality holds naturally\. We therefore assume throughout the rest of the proof thatWT\>0W\_\{T\}\>0\. IfA=1A=1, thenB⁡\(v\)=Γ⁡\(v\)Γ⁡\(v\)=1\\mathrm\{B\}\(v\)=\\frac\{\\Gamma\(v\)\}\{\\Gamma\(v\)\}=1for everyv\>0v\>0, and

WT​\(1\)​log⁡WT​\(1\)WT=0\.\\displaystyle W\_\{T\}\(1\)\\log\\frac\{W\_\{T\}\(1\)\}\{W\_\{T\}\}=0\.Thus,ℛ​\(WT​\(⋅\)\)=0\\mathcal\{R\}\(W\_\{T\}\(\\cdot\)\)=0\. It remains to considerA≥2A\\geq 2\.

By the definition of the beta function \([B\.12](https://arxiv.org/html/2609.38666#A2.E12)\), we have

−log⁡B⁡\(WT​\(⋅\)\+12​𝟏\)\\displaystyle\-\\log\\mathrm\{B\}\\Big\(W\_\{T\}\(\\cdot\)\+\\frac\{1\}\{2\}\\mathbf\{1\}\\Big\)=log⁡Γ⁡\(WT\+A2\)−∑a∈𝒜log⁡Γ⁡\(WT​\(a\)\+12\)\.\\displaystyle=\\log\\Gamma\\Big\(W\_\{T\}\+\\frac\{A\}\{2\}\\Big\)\-\\sum\_\{a\\in\\mathcal\{A\}\}\\log\\Gamma\\Big\(W\_\{T\}\(a\)\+\\frac\{1\}\{2\}\\Big\)\.\(D\.1\)Using Lemma[I\.8](https://arxiv.org/html/2609.38666#A9.Thmtheorem8)withs=WT\+A/2s=W\_\{T\}\+A/2, we have

log⁡Γ⁡\(WT\+A2\)\\displaystyle\\log\\Gamma\\Big\(W\_\{T\}\+\\frac\{A\}\{2\}\\Big\)≤\(WT\+A−12\)​log⁡\(WT\+A2\)−\(WT\+A2\)\\displaystyle\\leq\\Big\(W\_\{T\}\+\\frac\{A\-1\}\{2\}\\Big\)\\log\\Big\(W\_\{T\}\+\\frac\{A\}\{2\}\\Big\)\-\\Big\(W\_\{T\}\+\\frac\{A\}\{2\}\\Big\)\+12​log⁡\(2​π\)\+112​\(WT\+A/2\)\.\\displaystyle\\qquad\+\\frac\{1\}\{2\}\\log\(2\\pi\)\+\\frac\{1\}\{12\(W\_\{T\}\+A/2\)\}\.\(D\.2\)For everya∈𝒜a\\in\\mathcal\{A\}, applying Lemma[I\.8](https://arxiv.org/html/2609.38666#A9.Thmtheorem8)withs=WT​\(a\)\+1/2s=W\_\{T\}\(a\)\+1/2, we have

log⁡Γ⁡\(WT​\(a\)\+12\)\\displaystyle\\log\\Gamma\\Big\(W\_\{T\}\(a\)\+\\frac\{1\}\{2\}\\Big\)≥WT​\(a\)​log⁡\(WT​\(a\)\+12\)−\(WT​\(a\)\+12\)\+12​log⁡\(2​π\)\.\\displaystyle\\geq W\_\{T\}\(a\)\\log\\Big\(W\_\{T\}\(a\)\+\\frac\{1\}\{2\}\\Big\)\-\\Big\(W\_\{T\}\(a\)\+\\frac\{1\}\{2\}\\Big\)\+\\frac\{1\}\{2\}\\log\(2\\pi\)\.\(D\.3\)Substituting \([D\.2](https://arxiv.org/html/2609.38666#A4.E2)\) and \([D\.3](https://arxiv.org/html/2609.38666#A4.E3)\) into \([D\.1](https://arxiv.org/html/2609.38666#A4.E1)\), we obtain

−logB\(WT\(⋅\)\+12𝟏\)≤−∑a∈𝒜WT\(a\)log\(WT\(a\)\+12\)\\displaystyle\-\\log\\mathrm\{B\}\\Big\(W\_\{T\}\(\\cdot\)\+\\frac\{1\}\{2\}\\mathbf\{1\}\\Big\)\\leq\-\\sum\_\{a\\in\\mathcal\{A\}\}W\_\{T\}\(a\)\\log\\Big\(W\_\{T\}\(a\)\+\\frac\{1\}\{2\}\\Big\)\+\(WT\+A−12\)​log⁡\(WT\+A2\)−A−12​log⁡\(2​π\)\+112​\(WT\+A/2\),\\displaystyle\\quad\+\\Big\(W\_\{T\}\+\\frac\{A\-1\}\{2\}\\Big\)\\log\\left\(W\_\{T\}\+\\frac\{A\}\{2\}\\right\)\-\\frac\{A\-1\}\{2\}\\log\(2\\pi\)\+\\frac\{1\}\{12\(W\_\{T\}\+A/2\)\},where we use∑a∈𝒜WT​\(a\)=WT\\sum\_\{a\\in\\mathcal\{A\}\}W\_\{T\}\(a\)=W\_\{T\}\. Recall that

ℛ​\(WT​\(⋅\)\)\\displaystyle\\mathcal\{R\}\\bigl\(W\_\{T\}\(\\cdot\)\\bigr\)=log⁡B⁡\(12​𝟏\)−log⁡B⁡\(WT​\(⋅\)\+12​𝟏\)\+∑a∈𝒜WT​\(a\)​log​WT​\(a\)WT\.\\displaystyle=\\log\\mathrm\{B\}\\Big\(\\frac\{1\}\{2\}\\mathbf\{1\}\\Big\)\-\\log\\mathrm\{B\}\\Big\(W\_\{T\}\(\\cdot\)\+\\frac\{1\}\{2\}\\mathbf\{1\}\\Big\)\+\\sum\_\{a\\in\\mathcal\{A\}\}W\_\{T\}\(a\)\\log\\frac\{W\_\{T\}\(a\)\}\{W\_\{T\}\}\.Therefore, we have

ℛ⁡\(WT​\(⋅\)\)≤log⁡B⁡\(12​𝟏\)\+∑a∈𝒜WT​\(a\)​log⁡WT​\(a\)WT​\(a\)\+1/2⏟I1\\displaystyle\\mathcal\{R\}\\bigl\(W\_\{T\}\(\\cdot\)\\bigr\)\\leq\\log\\mathrm\{B\}\\Big\(\\frac\{1\}\{2\}\\mathbf\{1\}\\Big\)\+\\underbrace\{\\sum\_\{a\\in\\mathcal\{A\}\}W\_\{T\}\(a\)\\log\\frac\{W\_\{T\}\(a\)\}\{W\_\{T\}\(a\)\+1/2\}\}\_\{I\_\{1\}\}\+WT​log⁡WT\+A/2WT⏟I2\+A−12​log⁡\(WT\+A2\)⏟I3−A−12​log⁡\(2​π\)\+112​\(WT\+A/2\)\.\\displaystyle\\quad\+\\underbrace\{W\_\{T\}\\log\\frac\{W\_\{T\}\+A/2\}\{W\_\{T\}\}\}\_\{I\_\{2\}\}\+\\underbrace\{\\frac\{A\-1\}\{2\}\\log\\Big\(W\_\{T\}\+\\frac\{A\}\{2\}\\Big\)\}\_\{I\_\{3\}\}\-\\frac\{A\-1\}\{2\}\\log\(2\\pi\)\+\\frac\{1\}\{12\(W\_\{T\}\+A/2\)\}\.\(D\.4\)Since0≤WT​\(a\)/\[WT​\(a\)\+1/2\]≤10\\leq\{W\_\{T\}\(a\)\}/\{\[W\_\{T\}\(a\)\+1/2\]\}\\leq 1,I1≤0I\_\{1\}\\leq 0\. ForI2I\_\{2\}, usinglog⁡\(1\+s\)≤s\\log\(1\+s\)\\leq sfor everys≥0s\\geq 0, we have

I2\\displaystyle I\_\{2\}=WT​log⁡\(1\+A2​WT\)≤A2\.\\displaystyle=W\_\{T\}\\log\\left\(1\+\\frac\{A\}\{2W\_\{T\}\}\\right\)\\leq\\frac\{A\}\{2\}\.ForI3I\_\{3\}, we haveWT\+A2≤\(WT\+1\)​\(1\+A/2\)W\_\{T\}\+\\frac\{A\}\{2\}\\leq\(W\_\{T\}\+1\)\\big\(1\+A/2\\big\)\. Thus,

WT\+A2≤\(WT\+1\)​\(1\+A2\),\\displaystyle W\_\{T\}\+\\frac\{A\}\{2\}\\leq\(W\_\{T\}\+1\)\\Big\(1\+\\frac\{A\}\{2\}\\Big\),and henceI3≤\(A−1\)​log⁡\(WT\+1\)/2\+\(A−1\)​log⁡\(1\+A/2\)/2I\_\{3\}\\leq\(A\-1\)\\log\(W\_\{T\}\+1\)/2\+\(A\-1\)\\log\(1\+A/2\)/2\. As a result, \([D\.4](https://arxiv.org/html/2609.38666#A4.E4)\) indicates

ℛ​\(WT​\(⋅\)\)\\displaystyle\\mathcal\{R\}\\bigl\(W\_\{T\}\(\\cdot\)\\bigr\)≤A−12​log⁡\(WT\+1\)\+log⁡B⁡\(12​𝟏\)\+A2\+A−12​log⁡\(1\+A2\)\+112\.\\displaystyle\\leq\\frac\{A\-1\}\{2\}\\log\(W\_\{T\}\+1\)\+\\log\\mathrm\{B\}\\Big\(\\frac\{1\}\{2\}\\mathbf\{1\}\\Big\)\+\\frac\{A\}\{2\}\+\\frac\{A\-1\}\{2\}\\log\\Big\(1\+\\frac\{A\}\{2\}\\Big\)\+\\frac\{1\}\{12\}\.\(D\.5\)Finally, by definition,

log⁡B⁡\(12​𝟏\)=A2​log⁡π−log⁡Γ⁡\(A2\)\.\\displaystyle\\log\\mathrm\{B\}\\Big\(\\frac\{1\}\{2\}\\mathbf\{1\}\\Big\)=\\frac\{A\}\{2\}\\log\\pi\-\\log\\Gamma\\Big\(\\frac\{A\}\{2\}\\Big\)\.Applying Lemma[I\.8](https://arxiv.org/html/2609.38666#A9.Thmtheorem8)withs=A/2s=A/2, we have

log⁡Γ⁡\(A2\)\\displaystyle\\log\\Gamma\\Big\(\\frac\{A\}\{2\}\\Big\)≥A−12​log⁡A2−A2\+12​log⁡\(2​π\)\.\\displaystyle\\geq\\frac\{A\-1\}\{2\}\\log\\frac\{A\}\{2\}\-\\frac\{A\}\{2\}\+\\frac\{1\}\{2\}\\log\(2\\pi\)\.SinceA≥2A\\geq 2, we have

log⁡B⁡\(12​𝟏\)≤A2​\(log⁡π\+1\)−A−12​log​A2\.\\displaystyle\\log\\mathrm\{B\}\\Big\(\\frac\{1\}\{2\}\\mathbf\{1\}\\Big\)\\leq\\frac\{A\}\{2\}\(\\log\\pi\+1\)\-\\frac\{A\-1\}\{2\}\\log\\frac\{A\}\{2\}\.\(D\.6\)Substituting \([D\.6](https://arxiv.org/html/2609.38666#A4.E6)\) into \([D\.5](https://arxiv.org/html/2609.38666#A4.E5)\), we have

ℛ​\(WT​\(⋅\)\)\\displaystyle\\mathcal\{R\}\\bigl\(W\_\{T\}\(\\cdot\)\\bigr\)≤A−12​log⁡\(WT\+1\)\+A2\+A−12​log⁡\(1\+A2\)\+112\\displaystyle\\leq\\frac\{A\-1\}\{2\}\\log\(W\_\{T\}\+1\)\+\\frac\{A\}\{2\}\+\\frac\{A\-1\}\{2\}\\log\\Big\(1\+\\frac\{A\}\{2\}\\Big\)\+\\frac\{1\}\{12\}\+A2​\(log⁡π\+1\)−A−12​log⁡A2\\displaystyle\\qquad\+\\frac\{A\}\{2\}\(\\log\\pi\+1\)\-\\frac\{A\-1\}\{2\}\\log\\frac\{A\}\{2\}≤A​log⁡\(WT\+1\)\+3​A,\\displaystyle\\leq A\\log\(W\_\{T\}\+1\)\+3A,where we uselog⁡\(1\+A/2\)−log⁡\(A/2\)≤2/A\\log\(1\+A/2\)\-\\log\(A/2\)\\leq 2/A\. This completes the proof of Lemma[B\.1](https://arxiv.org/html/2609.38666#A2.Thmtheorem1)\. ∎

### D\.2Proof of Lemma[B\.2](https://arxiv.org/html/2609.38666#A2.Thmtheorem2)

###### Proof of Lemma[B\.2](https://arxiv.org/html/2609.38666#A2.Thmtheorem2)\.

By conditional Jensen’s inequality,

exp⁡\(𝔼⁡\[Y∣ℱ\]\)≤𝔼⁡\[eY∣ℱ\]≤1\.\\displaystyle\\exp\\bigl\(\\mathbb\{E\}\[Y\\mid\\mathcal\{F\}\]\\bigr\)\\leq\\mathbb\{E\}\[e^\{Y\}\\mid\\mathcal\{F\}\]\\leq 1\.Therefore,𝔼⁡\[Y∣ℱ\]≤0\\mathbb\{E\}\[Y\\mid\\mathcal\{F\}\]\\leq 0and henceμ≥0\\mu\\geq 0\. Moreover,Y≥−LY\\geq\-Limplies𝔼⁡\[Y∣ℱ\]≥−L\\mathbb\{E\}\[Y\\mid\\mathcal\{F\}\]\\geq\-L, soμ≤L\\mu\\leq L\.

We first show that, for everyy≥−Ly\\geq\-Landθ∈\[0,1\]\\theta\\in\[0,1\],

eθ​y−1−θ​y≤\(1\+L\)​θ2​\(ey−1−y\)\.\\displaystyle e^\{\\theta y\}\-1\-\\theta y\\leq\(1\+L\)\\theta^\{2\}\(e^\{y\}\-1\-y\)\.\(D\.7\)First, suppose thaty≥0y\\geq 0\. Expanding the exponential using Taylor’s series gives

eθ​y−1−θ​y\\displaystyle e^\{\\theta y\}\-1\-\\theta y=∑k=2∞θk​ykk\!\\displaystyle=\\sum\_\{k=2\}^\{\\infty\}\\frac\{\\theta^\{k\}y^\{k\}\}\{k\!\}≤θ2​∑k=2∞ykk\!\\displaystyle\\leq\\theta^\{2\}\\sum\_\{k=2\}^\{\\infty\}\\frac\{y^\{k\}\}\{k\!\}=θ2​\(ey−1−y\),\\displaystyle=\\theta^\{2\}\(e^\{y\}\-1\-y\),where we usedθk≤θ2\\theta^\{k\}\\leq\\theta^\{2\}for everyk≥2k\\geq 2\.

It remains to show \([D\.7](https://arxiv.org/html/2609.38666#A4.E7)\) when−L≤y≤0\-L\\leq y\\leq 0\. Lety=−sy=\-sfor somes∈\[0,L\]s\\in\[0,L\]\. Then, on the one hand, the left\-hand side of \([D\.7](https://arxiv.org/html/2609.38666#A4.E7)\) satisfies

e−θ​s−1\+θ​s≤θ2​s22\.\\displaystyle e^\{\-\\theta s\}\-1\+\\theta s\\leq\\frac\{\\theta^\{2\}s^\{2\}\}\{2\}\.On the other hand, the right\-hand side of \([D\.7](https://arxiv.org/html/2609.38666#A4.E7)\) satisfies:

\(1\+L\)​θ2​\(e−s−1\+s\)\\displaystyle\(1\+L\)\\theta^\{2\}\(e^\{\-s\}\-1\+s\)=\(1\+L\)​θ2​∫0s\(1−e−v\)​𝑑v\\displaystyle=\(1\+L\)\\theta^\{2\}\\int\_\{0\}^\{s\}\(1\-e^\{\-v\}\)\\,\\mathrm\{d\}v≥\(1\+L\)​θ2​∫0sv1\+v​𝑑v\\displaystyle\\geq\(1\+L\)\\theta^\{2\}\\int\_\{0\}^\{s\}\\frac\{v\}\{1\+v\}\\,\\mathrm\{d\}v≥\(1\+L\)​θ2​∫0sv1\+L​𝑑v\\displaystyle\\geq\(1\+L\)\\theta^\{2\}\\int\_\{0\}^\{s\}\\frac\{v\}\{1\+L\}\\,\\mathrm\{d\}v=θ2​s22,\\displaystyle=\\frac\{\\theta^\{2\}s^\{2\}\}\{2\},where the first inequality holds due toex≥1\+xe^\{x\}\\geq 1\+x, and thuse−x≤1/\(1\+x\)e^\{\-x\}\\leq 1/\(1\+x\)wheneverx≥0x\\geq 0\. The second inequality holds due tov≤s≤Lv\\leq s\\leq L\. Therefore, we have proved \([D\.7](https://arxiv.org/html/2609.38666#A4.E7)\) fory∈\[−L,0\]y\\in\[\-L,0\]\.

Using \([D\.7](https://arxiv.org/html/2609.38666#A4.E7)\), we obtain

𝔼⁡\[eθ​Y∣ℱ\]\\displaystyle\\mathbb\{E\}\[e^\{\\theta Y\}\\mid\\mathcal\{F\}\]=1\+θ​𝔼​\[Y∣ℱ\]\+𝔼⁡\[eθ​Y−1−θ​Y∣ℱ\]\\displaystyle=1\+\\theta\\mathbb\{E\}\[Y\\mid\\mathcal\{F\}\]\+\\mathbb\{E\}\[e^\{\\theta Y\}\-1\-\\theta Y\\mid\\mathcal\{F\}\]≤1−θ​μ\+\(1\+L\)​θ2​𝔼​\[eY−1−Y∣ℱ\]\.\\displaystyle\\leq 1\-\\theta\\mu\+\(1\+L\)\\theta^\{2\}\\mathbb\{E\}\[e^\{Y\}\-1\-Y\\mid\\mathcal\{F\}\]\.\(D\.8\)Using𝔼⁡\[eY∣ℱ\]≤1\\mathbb\{E\}\[e^\{Y\}\\mid\\mathcal\{F\}\]\\leq 1, we have

𝔼⁡\[eY−1−Y∣ℱ\]\\displaystyle\\mathbb\{E\}\[e^\{Y\}\-1\-Y\\mid\\mathcal\{F\}\]=𝔼⁡\[eY∣ℱ\]−1−𝔼⁡\[Y∣ℱ\]\\displaystyle=\\mathbb\{E\}\[e^\{Y\}\\mid\\mathcal\{F\}\]\-1\-\\mathbb\{E\}\[Y\\mid\\mathcal\{F\}\]≤−𝔼⁡\[Y∣ℱ\]\\displaystyle\\leq\-\\mathbb\{E\}\[Y\\mid\\mathcal\{F\}\]=μ\.\\displaystyle=\\mu\.Therefore, \([D\.8](https://arxiv.org/html/2609.38666#A4.E8)\) gives

𝔼⁡\[eθ​Y∣ℱ\]≤1−θ​μ\+\(1\+L\)​θ2​μ\.\\displaystyle\\mathbb\{E\}\[e^\{\\theta Y\}\\mid\\mathcal\{F\}\]\\leq 1\-\\theta\\mu\+\(1\+L\)\\theta^\{2\}\\mu\.Consequently,

log⁡𝔼⁡\[eθ⁡\(Y\+μ\)\|ℱ\]\\displaystyle\\log\\mathbb\{E\}\\left\[e^\{\\theta\(Y\+\\mu\)\}\\,\\middle\|\\,\\mathcal\{F\}\\right\]=θ​μ\+log⁡𝔼⁡\[eθ​Y∣ℱ\]\\displaystyle=\\theta\\mu\+\\log\\mathbb\{E\}\[e^\{\\theta Y\}\\mid\\mathcal\{F\}\]≤θ​μ\+log⁡\(1−θ​μ\+\(1\+L\)​θ2​μ\)\\displaystyle\\leq\\theta\\mu\+\\log\\left\(1\-\\theta\\mu\+\(1\+L\)\\theta^\{2\}\\mu\\right\)≤θ​μ−θ​μ\+\(1\+L\)​θ2​μ\\displaystyle\\leq\\theta\\mu\-\\theta\\mu\+\(1\+L\)\\theta^\{2\}\\mu=\(1\+L\)​θ2​μ,\\displaystyle=\(1\+L\)\\theta^\{2\}\\mu,where the last inequality useslog⁡\(1\+s\)≤s\\log\(1\+s\)\\leq s\. This completes the proof of Lemma[B\.2](https://arxiv.org/html/2609.38666#A2.Thmtheorem2)\. ∎

### D\.3Proof of Lemma[5\.4](https://arxiv.org/html/2609.38666#S5.Thmtheorem4)

###### Proof\.

Fix a levelhhand a triplez=\(x,u,a\)z=\(x,u,a\)\. Letℱt\\mathcal\{F\}\_\{t\}denote the history before roundtt\. Define

Yt​\(z\):=1m​∑j=1m𝟙⁡\(xt,j=x,yt,j,<h=u,at,j,h=a\)​\[lt,j,h−log⁡gh​\(a\|x,u\)\]\.\\displaystyle Y\_\{t\}\(z\):=\\frac\{1\}\{m\}\\sum\_\{j=1\}^\{m\}\\ind\\bigl\(x\_\{t,j\}=x,\\,y\_\{t,j,<h\}=u,\\,a\_\{t,j,h\}=a\\bigr\)\\bigl\[l\_\{t,j,h\}\-\\log g\_\{h\}\(a\|x,u\)\\bigr\]\.By the definitions ofStS\_\{t\}andNtN\_\{t\}, we have

∑s=1tYs​\(z\)=St​\(x,u,a\)−Nt​\(x,u,a\)​log⁡gh​\(a\|x,u\)m\.\\displaystyle\\sum\_\{s=1\}^\{t\}Y\_\{s\}\(z\)=\\frac\{S\_\{t\}\(x,u,a\)\-N\_\{t\}\(x,u,a\)\\log g\_\{h\}\(a\|x,u\)\}\{m\}\.\(D\.9\)Under the on\-policy protocol, conditionally onℱt\\mathcal\{F\}\_\{t\}, the teacher index is uniform overℐ\\mathcal\{I\}, the context followsρit\\rho\_\{i\_\{t\}\}, and the response is generated byπ^t\\widehat\{\\pi\}\_\{t\}\. Therefore,

𝔼⁡\[Yt​\(z\)∣ℱt\]\\displaystyle\\mathbb\{E\}\[Y\_\{t\}\(z\)\\mid\\mathcal\{F\}\_\{t\}\]=1m​∑j=1m1I​∑i∈ℐρi​\(x\)​π^t​\(u\|x\)​π^t​\(a\|x,u\)​\[log⁡pi​\(a\|x,u\)−log⁡gh​\(a\|x,u\)\]\\displaystyle=\\frac\{1\}\{m\}\\sum\_\{j=1\}^\{m\}\\frac\{1\}\{I\}\\sum\_\{i\\in\\mathcal\{I\}\}\\rho\_\{i\}\(x\)\\widehat\{\\pi\}\_\{t\}\(u\|x\)\\widehat\{\\pi\}\_\{t\}\(a\|x,u\)\\bigl\[\\log p\_\{i\}\(a\|x,u\)\-\\log g\_\{h\}\(a\|x,u\)\\bigr\]=ρ¯​\(x\)​π^t​\(u\|x\)​π^t​\(a\|x,u\)​∑i∈ℐwi​\(x\)​\[log⁡pi​\(a\|x,u\)−log⁡gh​\(a\|x,u\)\]\\displaystyle=\\bar\{\\rho\}\(x\)\\widehat\{\\pi\}\_\{t\}\(u\|x\)\\widehat\{\\pi\}\_\{t\}\(a\|x,u\)\\sum\_\{i\\in\\mathcal\{I\}\}w\_\{i\}\(x\)\\bigl\[\\log p\_\{i\}\(a\|x,u\)\-\\log g\_\{h\}\(a\|x,u\)\\bigr\]=0,\\displaystyle=0,where the second equality usesρi​\(x\)/I=ρ¯​\(x\)​wi​\(x\)\\rho\_\{i\}\(x\)/I=\\bar\{\\rho\}\(x\)w\_\{i\}\(x\)\. The last equality holds due to∑iwi​\(x\)=1\\sum\_\{i\}w\_\{i\}\(x\)=1andlog⁡gh​\(a\|x,u\)=∑iwi​\(x\)​log⁡pi​\(a\|x,u\)\\log g\_\{h\}\(a\|x,u\)=\\sum\_\{i\}w\_\{i\}\(x\)\\log p\_\{i\}\(a\|x,u\)\. Using Assumption[5\.3](https://arxiv.org/html/2609.38666#S5.Thmtheorem3), for every teacherii,

\|log⁡pi​\(a\|x,u\)−log⁡gh​\(a\|x,u\)\|\\displaystyle\\big\|\\log p\_\{i\}\(a\|x,u\)\-\\log g\_\{h\}\(a\|x,u\)\\big\|=\|log⁡pi​\(a\|x,u\)πref​\(a\|x,u\)−∑k∈ℐwk​\(x\)​log⁡pk​\(a\|x,u\)πref​\(a\|x,u\)\|\\displaystyle=\\bigg\|\\log\\frac\{p\_\{i\}\(a\|x,u\)\}\{\\pi\_\{\\text\{ref\}\}\(a\|x,u\)\}\-\\sum\_\{k\\in\\mathcal\{I\}\}w\_\{k\}\(x\)\\log\\frac\{p\_\{k\}\(a\|x,u\)\}\{\\pi\_\{\\text\{ref\}\}\(a\|x,u\)\}\\bigg\|≤B\+∑k∈ℐwk​\(x\)​B=2​B\.\\displaystyle\\leq B\+\\sum\_\{k\\in\\mathcal\{I\}\}w\_\{k\}\(x\)B=2B\.When𝟙⁡\(xt,j=x,yt,j,<h=u,at,j,h=a\)≠0\\ind\\bigl\(x\_\{t,j\}=x,\\,y\_\{t,j,<h\}=u,\\,a\_\{t,j,h\}=a\\bigr\)\\neq 0,lt,j,h=log⁡pi​\(a\|x,u\)l\_\{t,j,h\}=\\log p\_\{i\}\(a\|x,u\)for somei∈ℐi\\in\\mathcal\{I\}\. Thus,

\|lt,j,h−log⁡gh​\(a\|x,u\)\|≤2​B\.\\displaystyle\\big\|l\_\{t,j,h\}\-\\log g\_\{h\}\(a\|x,u\)\\big\|\\leq 2B\.In the definition ofYt​\(z\)Y\_\{t\}\(z\), at mostmmsummands can be nonzero\. Thus,\|Yt​\(z\)\|≤2​B​m/m=2​B\|Y\_\{t\}\(z\)\|\\leq 2Bm/m=2B\. Moreover, Jensen’s inequality gives

Yt​\(z\)2\\displaystyle Y\_\{t\}\(z\)^\{2\}≤1m​∑j=1m𝟙⁡\(xt,j=x,yt,j,<h=u,at,j,h=a\)​\[lt,j,h−log⁡gh​\(a\|x,u\)\]2\\displaystyle\\leq\\frac\{1\}\{m\}\\sum\_\{j=1\}^\{m\}\\ind\\bigl\(x\_\{t,j\}=x,\\,y\_\{t,j,<h\}=u,\\,a\_\{t,j,h\}=a\\bigr\)\\bigl\[l\_\{t,j,h\}\-\\log g\_\{h\}\(a\|x,u\)\\bigr\]^\{2\}≤4​B2m​∑j=1m𝟙⁡\(xt,j=x,yt,j,<h=u,at,j,h=a\)\.\\displaystyle\\leq\\frac\{4B^\{2\}\}\{m\}\\sum\_\{j=1\}^\{m\}\\ind\\bigl\(x\_\{t,j\}=x,\\,y\_\{t,j,<h\}=u,\\,a\_\{t,j,h\}=a\\bigr\)\.Taking the conditional expectation, we obtain

Var⁡\(Yt​\(z\)∣ℱt\)\\displaystyle\\Var\(Y\_\{t\}\(z\)\\mid\\mathcal\{F\}\_\{t\}\)=𝔼⁡\[Yt​\(z\)2∣ℱt\]\\displaystyle=\\mathbb\{E\}\[Y\_\{t\}\(z\)^\{2\}\\mid\\mathcal\{F\}\_\{t\}\]≤4​B2m​𝔼​\[∑j=1m𝟙⁡\(xt,j=x,yt,j,<h=u,at,j,h=a\)\|ℱt\]\.\\displaystyle\\leq\\frac\{4B^\{2\}\}\{m\}\\mathbb\{E\}\\bigg\[\\sum\_\{j=1\}^\{m\}\\ind\\bigl\(x\_\{t,j\}=x,\\,y\_\{t,j,<h\}=u,\\,a\_\{t,j,h\}=a\\bigr\)\\bigg\|\\,\\mathcal\{F\}\_\{t\}\\bigg\]\.\(D\.10\)DefineVt​\(z\):=∑s=1tVar⁡\(Ys​\(z\)∣ℱs\)V\_\{t\}\(z\):=\\sum\_\{s=1\}^\{t\}\\Var\(Y\_\{s\}\(z\)\\mid\\mathcal\{F\}\_\{s\}\)\. Since𝔼⁡\[Ys​\(z\)∣ℱs\]=0\\mathbb\{E\}\[Y\_\{s\}\(z\)\\mid\\mathcal\{F\}\_\{s\}\]=0and\|Ys​\(z\)\|≤2​B\|Y\_\{s\}\(z\)\|\\leq 2B, Lemma[I\.4](https://arxiv.org/html/2609.38666#A9.Thmtheorem4)implies that both

\(∑s=1tYs\(z\),Vt\(z\)\)and\(−∑s=1tYs\(z\),Vt\(z\)\)\\displaystyle\\bigg\(\\sum\_\{s=1\}^\{t\}Y\_\{s\}\(z\),\\,V\_\{t\}\(z\)\\bigg\)\\quad\\text\{and\}\\quad\\bigg\(\-\\sum\_\{s=1\}^\{t\}Y\_\{s\}\(z\),\\,V\_\{t\}\(z\)\\bigg\)are sub\-gamma processes with parameter2​B/32B/3\. LetK:=S​∑h=1HAhK:=S\\sum\_\{h=1\}^\{H\}A^\{h\}\. Then, applying Lemma[I\.5](https://arxiv.org/html/2609.38666#A9.Thmtheorem5)withρ=2​B\\rho=2B, with probability at least1−δ/\(2​K\)1\-\\delta/\(2K\), we have

∑s=1tYs​\(z\)≤4​Vt​\(z\)​log⁡\(2​Ht​K/δ\)\+883​B​log⁡\(2​Ht​K/δ\),\\displaystyle\\sum\_\{s=1\}^\{t\}Y\_\{s\}\(z\)\\leq 4\\sqrt\{V\_\{t\}\(z\)\\log\(2H\_\{t\}K/\\delta\)\}\+\\frac\{88\}\{3\}B\\log\(2H\_\{t\}K/\\delta\),\(D\.11\)where we can chooseHt=e\+log⁡\(1\+T\)H\_\{t\}=e\+\\log\(1\+T\)sinceVt​\(z\)≤4​B2​TV\_\{t\}\(z\)\\leq 4B^\{2\}T\. Applying the same inequality to−Ys​\(z\)\-Y\_\{s\}\(z\), and taking the union bound, we obtain that with probability at least1−δ1\-\\delta, the following inequality holds for anyhhand anyz=\(x,u,a\)z=\(x,u,a\):

\|∑s=1tYs​\(z\)\|≤4​Vt​\(z\)​log⁡\(2​Ht​K/δ\)\+883​B​log⁡\(2​Ht​K/δ\)\.\\displaystyle\\bigg\|\\sum\_\{s=1\}^\{t\}Y\_\{s\}\(z\)\\bigg\|\\leq 4\\sqrt\{V\_\{t\}\(z\)\\log\(2H\_\{t\}K/\\delta\)\}\+\\frac\{88\}\{3\}B\\log\(2H\_\{t\}K/\\delta\)\.Since

1m​∑j=1m𝟙⁡\(xs,j=x,ys,j,<h=u,as,j,h=a\)∈\[0,1\],\\displaystyle\\frac\{1\}\{m\}\\sum\_\{j=1\}^\{m\}\\ind\\bigl\(x\_\{s,j\}=x,\\,y\_\{s,j,<h\}=u,\\,a\_\{s,j,h\}=a\\bigr\)\\in\[0,1\],we can apply Lemma[I\.6](https://arxiv.org/html/2609.38666#A9.Thmtheorem6)and obtain that for anyhhandz=\(x,u,a\)z=\(x,u,a\), with probability at least1−δ/K1\-\\delta/K, the following inequality holds simultaneously for allt∈\[T\]t\\in\[T\]:

∑s=1t𝔼⁡\[1m​∑j=1m𝟙⁡\(xs,j=x,ys,j,<h=u,as,j,h=a\)\|ℱs\]\\displaystyle\\sum\_\{s=1\}^\{t\}\\mathbb\{E\}\\bigg\[\\frac\{1\}\{m\}\\sum\_\{j=1\}^\{m\}\\ind\\bigl\(x\_\{s,j\}=x,\\,y\_\{s,j,<h\}=u,\\,a\_\{s,j,h\}=a\\bigr\)\\,\\bigg\|\\,\\mathcal\{F\}\_\{s\}\\bigg\]≤2​∑s=1t1m​∑j=1m𝟙⁡\(xs,j=x,ys,j,<h=u,as,j,h=a\)\+8​log⁡\(2​K/δ\)\.\\displaystyle\\qquad\\leq 2\\sum\_\{s=1\}^\{t\}\\frac\{1\}\{m\}\\sum\_\{j=1\}^\{m\}\\ind\\bigl\(x\_\{s,j\}=x,\\,y\_\{s,j,<h\}=u,\\,a\_\{s,j,h\}=a\\bigr\)\+8\\log\(2K/\\delta\)\.\(D\.12\)Taking a union bound overhhand\(x,u,a\)\(x,u,a\), the inequality holds for allhhand\(x,u,a\)\(x,u,a\)simultaneously\. Combining \([D\.10](https://arxiv.org/html/2609.38666#A4.E10)\) and \([D\.12](https://arxiv.org/html/2609.38666#A4.E12)\), we have

Vt​\(z\)\\displaystyle V\_\{t\}\(z\)≤8​B2m​∑s=1t∑j=1m𝟙⁡\(xs,j=x,ys,j,<h=u,as,j,h=a\)\+32​B2​log⁡\(2​K/δ\)\\displaystyle\\leq\\frac\{8B^\{2\}\}\{m\}\\sum\_\{s=1\}^\{t\}\\sum\_\{j=1\}^\{m\}\\ind\\bigl\(x\_\{s,j\}=x,\\,y\_\{s,j,<h\}=u,\\,a\_\{s,j,h\}=a\\bigr\)\+32B^\{2\}\\log\(\{2K\}/\{\\delta\}\)=8​B2​Nt​\(x,u,a\)m\+32​B2​log⁡\(2​K/δ\)\.\\displaystyle=\\frac\{8B^\{2\}N\_\{t\}\(x,u,a\)\}\{m\}\+32B^\{2\}\\log\(\{2K\}/\{\\delta\}\)\.\(D\.13\)Substituting \([D\.13](https://arxiv.org/html/2609.38666#A4.E13)\) into \([D\.11](https://arxiv.org/html/2609.38666#A4.E11)\), we have, with probability at least1−2​δ1\-2\\delta, simultaneously for everyt∈\[T\]t\\in\[T\],hh, andz=\(x,u,a\)z=\(x,u,a\), the following inequality holds

\|∑s=1tYs​\(z\)\|\\displaystyle\\bigg\|\\sum\_\{s=1\}^\{t\}Y\_\{s\}\(z\)\\bigg\|≤4​\(8​B2​Nt​\(x,u,a\)m\+32​B2​log⁡2​Kδ\)​log⁡2​Ht​Kδ\+883​B​log⁡2​Ht​Kδ\\displaystyle\\leq 4\\sqrt\{\\Big\(\\frac\{8B^\{2\}N\_\{t\}\(x,u,a\)\}\{m\}\+32B^\{2\}\\log\\frac\{2K\}\{\\delta\}\\Big\)\\log\\frac\{2H\_\{t\}K\}\{\\delta\}\}\+\\frac\{88\}\{3\}B\\log\\frac\{2H\_\{t\}K\}\{\\delta\}≲B​Nt​\(x,u,a\)m​log⁡2​Ht​Kδ\+B​log⁡2​Ht​Kδ,\\displaystyle\\lesssim B\\sqrt\{\\frac\{N\_\{t\}\(x,u,a\)\}\{m\}\\log\\frac\{2H\_\{t\}K\}\{\\delta\}\}\+B\\log\\frac\{2H\_\{t\}K\}\{\\delta\},where we usev\+w≤v\+w\\sqrt\{v\+w\}\\leq\\sqrt\{v\}\+\\sqrt\{w\}andlog⁡\(2​K/δ\)≤log⁡\(2​Ht​K/δ\)\\log\(2K/\\delta\)\\leq\\log\(2H\_\{t\}K/\\delta\)\.

For every triple withNt​\(x,u,a\)≥1N\_\{t\}\(x,u,a\)\\geq 1, \([D\.9](https://arxiv.org/html/2609.38666#A4.E9)\) gives

\|l¯t​\(x,u,a\)−log⁡gh​\(a\|x,u\)\|\\displaystyle\\big\|\\bar\{l\}\_\{t\}\(x,u,a\)\-\\log g\_\{h\}\(a\|x,u\)\\big\|=mNt​\(x,u,a\)​\|∑s=1tYs​\(z\)\|\\displaystyle=\\frac\{m\}\{N\_\{t\}\(x,u,a\)\}\\bigg\|\\sum\_\{s=1\}^\{t\}Y\_\{s\}\(z\)\\bigg\|≲B​mNt​\(x,u,a\)​log⁡2​Ht​Kδ\+B​mNt​\(x,u,a\)​log⁡2​Ht​Kδ\.\\displaystyle\\lesssim B\\sqrt\{\\frac\{m\}\{N\_\{t\}\(x,u,a\)\}\\log\\frac\{2H\_\{t\}K\}\{\\delta\}\}\+\\frac\{Bm\}\{N\_\{t\}\(x,u,a\)\}\\log\\frac\{2H\_\{t\}K\}\{\\delta\}\.Finally, sinceK≤S​H​AHK\\leq SHA^\{H\}andHt=e\+log⁡\(1\+T\)H\_\{t\}=e\+\\log\(1\+T\),

log⁡2​Ht​Kδ≤H​log⁡A\+log⁡\[2​S​H​\(e\+log⁡\(1\+T\)\)δ\]\.\\displaystyle\\log\\frac\{2H\_\{t\}K\}\{\\delta\}\\leq H\\log A\+\\log\\bigg\[\\frac\{2SH\\bigl\(e\+\\log\(1\+T\)\\bigr\)\}\{\\delta\}\\bigg\]\.Hiding the logarithmic factors, we can define

βt​\(x,u,a\)=O~​\(B​m​HNt​\(x,u,a\)\+B​m​HNt​\(x,u,a\)\),\\displaystyle\\beta\_\{t\}\(x,u,a\)=\\widetilde\{O\}\\bigg\(B\\sqrt\{\\frac\{mH\}\{N\_\{t\}\(x,u,a\)\}\}\+\\frac\{BmH\}\{N\_\{t\}\(x,u,a\)\}\\bigg\),and conclude that

\|l¯t​\(x,u,a\)−log⁡gh​\(a\|x,u\)\|≤βt​\(x,u,a\)\.\\displaystyle\\big\|\\bar\{l\}\_\{t\}\(x,u,a\)\-\\log g\_\{h\}\(a\|x,u\)\\big\|\\leq\\beta\_\{t\}\(x,u,a\)\.This completes the proof of Lemma[5\.4](https://arxiv.org/html/2609.38666#S5.Thmtheorem4)\. ∎

## Appendix EMissing Proof in Section[6](https://arxiv.org/html/2609.38666#S6)

###### Proof of Proposition[6\.1](https://arxiv.org/html/2609.38666#S6.Thmtheorem1)\.

FixN≥2N\\geq 2andα∈\(0,1\)\\alpha\\in\(0,1\), and letb:=1/Nb:=1/N\. Forq∈\(b,1\)q\\in\(b,1\), write the correct\-response probabilities as

F⁡\(q\)\\displaystyle F\(q\):=πF∗​\(y⋆\|x\)=α​q\+\(1−α\)​b,\\displaystyle:=\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(y^\{\\star\}\|x\)=\\alpha q\+\(1\-\\alpha\)b,R⁡\(q\)\\displaystyle R\(q\):=πR∗​\(y⋆\|x\)=qαqα\+\(N−1\)1−α​\(1−q\)α\.\\displaystyle:=\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(y^\{\\star\}\|x\)=\\frac\{q^\{\\alpha\}\}\{q^\{\\alpha\}\+\(N\-1\)^\{1\-\\alpha\}\(1\-q\)^\{\\alpha\}\}\.Both probabilities lie in\(0,1\)\(0,1\)\. Since the log\-odds function is strictly increasing,F⁡\(q\)−R⁡\(q\)F\(q\)\-R\(q\)has the same sign as

G⁡\(q\)\\displaystyle G\(q\):=log⁡F⁡\(q\)1−F⁡\(q\)−log⁡R⁡\(q\)1−R⁡\(q\)\\displaystyle:=\\log\\frac\{F\(q\)\}\{1\-F\(q\)\}\-\\log\\frac\{R\(q\)\}\{1\-R\(q\)\}=log⁡F⁡\(q\)1−F⁡\(q\)−α​log⁡q1−q\+\(1−α\)​log⁡\(N−1\)\.\\displaystyle=\\log\\frac\{F\(q\)\}\{1\-F\(q\)\}\-\\alpha\\log\\frac\{q\}\{1\-q\}\+\(1\-\\alpha\)\\log\(N\-1\)\.The functionGGextends continuously toq=bq=b, withG⁡\(b\)=0G\(b\)=0\. Differentiating gives

G′​\(q\)\\displaystyle G^\{\\prime\}\(q\)=αF​\(q\)​\(1−F​\(q\)\)−αq⁡\(1−q\)\\displaystyle=\\frac\{\\alpha\}\{F\(q\)\(1\-F\(q\)\)\}\-\\frac\{\\alpha\}\{q\(1\-q\)\}=α⁡\(q−F⁡\(q\)\)​\(1−q−F⁡\(q\)\)F⁡\(q\)​\(1−F⁡\(q\)\)​q​\(1−q\)\.\\displaystyle=\\frac\{\\alpha\(q\-F\(q\)\)\(1\-q\-F\(q\)\)\}\{F\(q\)\(1\-F\(q\)\)q\(1\-q\)\}\.Forq\>bq\>b, we haveq−F⁡\(q\)=\(1−α\)​\(q−b\)\>0q\-F\(q\)=\(1\-\\alpha\)\(q\-b\)\>0\. Thus, the sign ofG′​\(q\)G^\{\\prime\}\(q\)is the sign of

1−q−F⁡\(q\)=1−\(1\+α\)​q−\(1−α\)​b,1\-q\-F\(q\)=1\-\(1\+\\alpha\)q\-\(1\-\\alpha\)b,which vanishes atqm:=\[1−\(1−α\)​b\]/\(1\+α\)q\_\{\\mathrm\{m\}\}:=\[1\-\(1\-\\alpha\)b\]/\(1\+\\alpha\)\. IfN≥3N\\geq 3, thenb<qm<1b<q\_\{\\mathrm\{m\}\}<1, soGGis strictly increasing on\(b,qm\)\(b,q\_\{\\mathrm\{m\}\}\)and strictly decreasing on\(qm,1\)\(q\_\{\\mathrm\{m\}\},1\)\. In particular,G⁡\(qm\)\>0G\(q\_\{\\mathrm\{m\}\}\)\>0\. Moreover,F⁡\(q\)→α\+\(1−α\)​b∈\(0,1\)F\(q\)\\to\\alpha\+\(1\-\\alpha\)b\\in\(0,1\)asq↑1q\\uparrow 1, and henceG⁡\(q\)→−∞G\(q\)\\to\-\\infty\. The intermediate value theorem and strict decrease on\(qm,1\)\(q\_\{\\mathrm\{m\}\},1\)imply that there is exactly one zeroqc∈\(qm,1\)q\_\{\\mathrm\{c\}\}\\in\(q\_\{\\mathrm\{m\}\},1\)\. We haveG⁡\(q\)\>0G\(q\)\>0forb<q<qcb<q<q\_\{\\mathrm\{c\}\}andG⁡\(q\)<0G\(q\)<0forqc<q<1q\_\{\\mathrm\{c\}\}<q<1, giving the claimed ordering\.

IfN=2N=2, thenqm=b=1/2q\_\{\\mathrm\{m\}\}=b=1/2, soG′​\(q\)<0G^\{\\prime\}\(q\)<0throughout\(1/2,1\)\(1/2,1\)\. SinceG⁡\(b\)=0G\(b\)=0, this givesG⁡\(q\)<0G\(q\)<0and thereforeR⁡\(q\)\>F⁡\(q\)R\(q\)\>F\(q\)on that interval\. Finally, both targets are below the expert probability forq\>bq\>b: we haveF⁡\(q\)<qF\(q\)<q, and

R⁡\(q\)1−R⁡\(q\)=q1−q​\(1−q\(N−1\)​q\)1−α<q1−q,\\frac\{R\(q\)\}\{1\-R\(q\)\}=\\frac\{q\}\{1\-q\}\\left\(\\frac\{1\-q\}\{\(N\-1\)q\}\\right\)^\{1\-\\alpha\}<\\frac\{q\}\{1\-q\},which impliesR⁡\(q\)<qR\(q\)<q\. ∎

###### Proof of Proposition[6\.3](https://arxiv.org/html/2609.38666#S6.Thmtheorem3)\.

Fix a contextxxand suppress its dependence in the notation\. Since the forward target is the weighted arithmetic mixture of the teachers, summing overEEgives

πF∗​\(E\)=∑iwi​pi​\(E\)≥wk​pk​\(E\),\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(E\)=\\sum\_\{i\}w\_\{i\}p\_\{i\}\(E\)\\geq w\_\{k\}p\_\{k\}\(E\),where the inequality follows from the nonnegativity of every summand\.

For the reverse target, fixα,q,η∈\(0,1\)\\alpha,q,\\eta\\in\(0,1\)and writem:=\|E\|m:=\|E\|andn:=\|𝒴∖E\|n:=\|\\mathcal\{Y\}\\setminus E\|\. Bothmmandnnare positive\. Forε∈\(0,1\)\\varepsilon\\in\(0,1\), define two teachers by

p1​\(y\)=\{q/m,y∈E,\(1−q\)/n,y∉E,p2​\(y\)=\{ε/m,y∈E,\(1−ε\)/n,y∉E\.p\_\{1\}\(y\)=\\begin\{cases\}q/m,&y\\in E,\\\\ \(1\-q\)/n,&y\\notin E,\\end\{cases\}\\qquad p\_\{2\}\(y\)=\\begin\{cases\}\\varepsilon/m,&y\\in E,\\\\ \(1\-\\varepsilon\)/n,&y\\notin E\.\\end\{cases\}These distributions have full support, andp1​\(E\)=qp\_\{1\}\(E\)=q\. With weightsα\\alphaand1−α1\-\\alpha, the unnormalized geometric masses on the two sets are

∑y∈Ep1​\(y\)α​p2​\(y\)1−α\\displaystyle\\sum\_\{y\\in E\}p\_\{1\}\(y\)^\{\\alpha\}p\_\{2\}\(y\)^\{1\-\\alpha\}=qα​ε1−α,\\displaystyle=q^\{\\alpha\}\\varepsilon^\{1\-\\alpha\},∑y∉Ep1​\(y\)α​p2​\(y\)1−α\\displaystyle\\sum\_\{y\\notin E\}p\_\{1\}\(y\)^\{\\alpha\}p\_\{2\}\(y\)^\{1\-\\alpha\}=\(1−q\)α​\(1−ε\)1−α\.\\displaystyle=\(1\-q\)^\{\\alpha\}\(1\-\\varepsilon\)^\{1\-\\alpha\}\.Consequently,

πR∗​\(E\)=qα​ε1−αqα​ε1−α\+\(1−q\)α​\(1−ε\)1−α\.\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(E\)=\\frac\{q^\{\\alpha\}\\varepsilon^\{1\-\\alpha\}\}\{q^\{\\alpha\}\\varepsilon^\{1\-\\alpha\}\+\(1\-q\)^\{\\alpha\}\(1\-\\varepsilon\)^\{1\-\\alpha\}\}\.Asε↓0\\varepsilon\\downarrow 0, the numerator tends to zero and the denominator tends to\(1−q\)α\>0\(1\-q\)^\{\\alpha\}\>0\. HenceπR∗​\(E\)→0\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(E\)\\to 0, and a sufficiently small positiveε\\varepsilonensuresπR∗​\(E\)<η\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(E\)<\\eta\. This choice retains full support for both teachers and proves the claim\. ∎

###### Proof of Proposition[6\.5](https://arxiv.org/html/2609.38666#S6.Thmtheorem5)\.

Fixβ∈\(0,1/2\)\\beta\\in\(0,1/2\)andr∈\(1/2,1\)r\\in\(1/2,1\), and setα:=1−β\\alpha:=1\-\\beta\. We construct the expert using token probabilities that do not depend on the horizon\. Fixδ∈\(0,1/2\)\\delta\\in\(0,1/2\), letp1​\(a\|x\)=rp\_\{1\}\(a\|x\)=r, and, at every subsequent position, set

p1​\(a\|x,u\)=\{1−δ,if the first token ofuisa,1/2,if the first token ofuisb\.p\_\{1\}\(a\|x,u\)=\\begin\{cases\}1\-\\delta,&\\text\{if the first token of $u$ is $a$\},\\\\ 1/2,&\\text\{if the first token of $u$ is $b$\}\.\\end\{cases\}The probability ofbbis the complementary probability at every prefix\. Letp2p\_\{2\}assign probability1/21/2to each token at every prefix\. Both teachers assign positive probability to every response in\{a,b\}H\\\{a,b\\\}^\{H\}for every finiteHH\.

Marginalizing the forward mixture over the lastH−1H\-1tokens yields

πF∗​\(a\|x\)=α​r\+β2=12\+α⁡\(r−12\)\>12\.\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x\)=\\alpha r\+\\frac\{\\beta\}\{2\}=\\frac\{1\}\{2\}\+\\alpha\\left\(r\-\\frac\{1\}\{2\}\\right\)\>\\frac\{1\}\{2\}\.ThusπF∗​\(a\|x\)\>πF∗​\(b\|x\)\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x\)\>\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(b\|x\)for everyHH\.

For the reverse target, the uniform teacher contributes the same factor2−β​H2^\{\-\\beta H\}to every complete response, so

πR∗​\(y\|x\)=p1​\(y\|x\)α∑z∈\{a,b\}Hp1​\(z\|x\)α\.\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(y\|x\)=\\frac\{p\_\{1\}\(y\|x\)^\{\\alpha\}\}\{\\sum\_\{z\\in\\\{a,b\\\}^\{H\}\}p\_\{1\}\(z\|x\)^\{\\alpha\}\}\.DefineCα:=\(1−δ\)α\+δαC\_\{\\alpha\}:=\(1\-\\delta\)^\{\\alpha\}\+\\delta^\{\\alpha\}andDα:=21−αD\_\{\\alpha\}:=2^\{1\-\\alpha\}\. Summing the powered expert probabilities over continuations gives

∑v∈\{a,b\}H−1p1​\(a​v\|x\)α\\displaystyle\\sum\_\{v\\in\\\{a,b\\\}^\{H\-1\}\}p\_\{1\}\(av\|x\)^\{\\alpha\}=rα​CαH−1,\\displaystyle=r^\{\\alpha\}C\_\{\\alpha\}^\{H\-1\},∑v∈\{a,b\}H−1p1​\(b​v\|x\)α\\displaystyle\\sum\_\{v\\in\\\{a,b\\\}^\{H\-1\}\}p\_\{1\}\(bv\|x\)^\{\\alpha\}=\(1−r\)α​DαH−1\.\\displaystyle=\(1\-r\)^\{\\alpha\}D\_\{\\alpha\}^\{H\-1\}\.Since0<α<10<\\alpha<1, the functiont↦tαt\\mapsto t^\{\\alpha\}is strictly concave\. Asδ≠1/2\\delta\\neq 1/2, strict concavity implies

Cα=\(1−δ\)α\+δα<2​\(12\)α=Dα\.C\_\{\\alpha\}=\(1\-\\delta\)^\{\\alpha\}\+\\delta^\{\\alpha\}<2\\left\(\\frac\{1\}\{2\}\\right\)^\{\\alpha\}=D\_\{\\alpha\}\.The reverse target’s first\-token odds therefore satisfy

πR∗​\(a\|x\)πR∗​\(b\|x\)=\(r1−r\)α​\(CαDα\)H−1\.\\frac\{\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(a\|x\)\}\{\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(b\|x\)\}=\\left\(\\frac\{r\}\{1\-r\}\\right\)^\{\\alpha\}\\left\(\\frac\{C\_\{\\alpha\}\}\{D\_\{\\alpha\}\}\\right\)^\{H\-1\}\.In particular, these odds are strictly less than one whenever

H\>1\+α​log⁡\(r/\(1−r\)\)log⁡\(Dα/Cα\)\.H\>1\+\\frac\{\\alpha\\log\(r/\(1\-r\)\)\}\{\\log\(D\_\{\\alpha\}/C\_\{\\alpha\}\)\}\.The denominator is strictly positive, so this bound is finite for every fixedβ\\betaandrrin the stated ranges\. The same construction thus satisfiesπR∗​\(a\|x\)<πR∗​\(b\|x\)\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(a\|x\)<\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(b\|x\)for every sufficiently large integerHH, completing the proof\. ∎

## Appendix FParameter Sensitivity of Aggregation Targets

We extend the examples in Section[6](https://arxiv.org/html/2609.38666#S6)by varying the teacher weights, expert confidence, response\-space size, and continuation distributions\. All curves are deterministic evaluations of the closed\-form targets, rather than results of training a student\. Each comparison holds the other parameters fixed\. The vertical axes show the full probability range, and dashed lines indicate the expert’s probability\.

#### Uninformative teachers\.

The expert has weightα\\alpha, assigns probabilityqqto the correct response, and distributes its remaining mass uniformly over theN−1N\-1incorrect responses\. The other teachers are uniform and have total weight1−α1\-\\alpha\. WritingF⁡\(q\):=πF∗​\(y⋆\|x\)F\(q\):=\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(y^\{\\star\}\|x\)andR⁡\(q\):=πR∗​\(y⋆\|x\)R\(q\):=\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(y^\{\\star\}\|x\), we evaluate

F⁡\(q\)=α​q\+1−αN,R⁡\(q\)=qαqα\+\(N−1\)1−α​\(1−q\)α\.F\(q\)=\\alpha q\+\\frac\{1\-\\alpha\}\{N\},\\qquad R\(q\)=\\frac\{q^\{\\alpha\}\}\{q^\{\\alpha\}\+\(N\-1\)^\{1\-\\alpha\}\(1\-q\)^\{\\alpha\}\}\.HereNNcounts responses; splitting the uniform teachers’ total weight among more identical teachers leaves both targets unchanged\.

Figure 2:Sensitivity to expert weight and response\-space size with uninformative teachers\. Panels \(a–c\) varyα∈\{0\.5,0\.9,0\.99\}\\alpha\\in\\\{0\.5,0\.9,0\.99\\\}atN=10N=10\. Panels \(d–f\), together with \(b\), compareN∈\{2,10,100,1000\}N\\in\\\{2,10,100,1000\\\}atα=0\.9\\alpha=0\.9\. Each curve variesqqfrom1/N1/Ntoward11\. Dotted vertical lines mark the interior crossing when it exists\.Figure[2](https://arxiv.org/html/2609.38666#A6.F2)shows that the region in which reverse aggregation retains more correct\-response probability depends on both parameters\. For the weights shown atN=10N=10, increasingα\\alphalowers the confidence threshold, although the targets become close as both approach the expert\. Atα=0\.9\\alpha=0\.9, increasingNNfrom1010to10001000raises the threshold, narrowing the range ofqqwith a reverse advantage\. The binary case differs: reverse aggregation is larger throughoutq∈\(1/2,1\)q\\in\(1/2,1\), with no interior crossing\.

#### Misleading teachers\.

We fixN=10N=10and replace the uniform teacher by a teacher of weight1−α1\-\\alphathat assigns probabilityε\\varepsilonto the correct response\. Both teachers distribute their remaining probability uniformly over the incorrect responses\. The targets are

F⁡\(ε\)=α​q\+\(1−α\)​ε,R⁡\(ε\)=qα​ε1−αqα​ε1−α\+\(1−q\)α​\(1−ε\)1−α\.F\(\\varepsilon\)=\\alpha q\+\(1\-\\alpha\)\\varepsilon,\\qquad R\(\\varepsilon\)=\\frac\{q^\{\\alpha\}\\varepsilon^\{1\-\\alpha\}\}\{q^\{\\alpha\}\\varepsilon^\{1\-\\alpha\}\+\(1\-q\)^\{\\alpha\}\(1\-\\varepsilon\)^\{1\-\\alpha\}\}\.Figure[3](https://arxiv.org/html/2609.38666#A6.F3)varies the expert weight at fixedq=0\.99q=0\.99, and its confidence at fixedα=0\.9\\alpha=0\.9\. The horizontal coordinate islog10⁡ε\\log\_\{10\}\\varepsilon; we evaluate the reverse target in the log domain to retain accuracy for extremely small probabilities\.

Figure 3:Sensitivity to expert weight and confidence with misleading teachers\. Panels \(a–c\) fixq=0\.99q=0\.99and varyα\\alpha\. Panels \(d–f\), together with \(b\), fixα=0\.9\\alpha=0\.9and compareq∈\{0\.6,0\.9,0\.99,0\.999\}q\\in\\\{0\.6,0\.9,0\.99,0\.999\\\}\. Horizontal ranges differ:log10⁡ε∈\[−10,−1\]\\log\_\{10\}\\varepsilon\\in\[\-10,\-1\]in \(a\),\[−600,−1\]\[\-600,\-1\]in \(c\), and\[−60,−1\]\[\-60,\-1\]otherwise\. Moving left makes the other teacher more misleading\.Higher expert weight or confidence shifts the decline in reverse probability to smaller values ofε\\varepsilon\. This protection can be substantial: atq=0\.99q=0\.99, the reverse probability equals1/21/2at approximatelyε=10−18\\varepsilon=10^\{\-18\}forα=0\.9\\alpha=0\.9, compared with10−19810^\{\-198\}forα=0\.99\\alpha=0\.99\. Nevertheless, every fixedα<1\\alpha<1permits suppression asε↓0\\varepsilon\\downarrow 0, whereas forward aggregation remains aboveα​q\\alpha q\. The extreme ranges illustrate the absence of a uniform positive lower bound for reverse aggregation; they do not estimate how often such predictions arise in practice\.

#### Long horizons\.

The expert selects the correct first tokenaawith probabilityrr\. At every subsequent position, its probabilities for\(a,b\)\(a,b\)are\(1−δ,δ\)\(1\-\\delta,\\delta\)after an initialaaand\(1/2,1/2\)\(1/2,1/2\)after an initialbb\. The other teacher is uniform at every prefix\. With expert weightα\\alpha, letCα=\(1−δ\)α\+δαC\_\{\\alpha\}=\(1\-\\delta\)^\{\\alpha\}\+\\delta^\{\\alpha\}andDα=21−αD\_\{\\alpha\}=2^\{1\-\\alpha\}\. The first\-token probabilities are

FH=α​r\+1−α2,RH=\[1\+\(1−rr\)α​\(DαCα\)H−1\]−1\.F\_\{H\}=\\alpha r\+\\frac\{1\-\\alpha\}\{2\},\\qquad R\_\{H\}=\\left\[1\+\\left\(\\frac\{1\-r\}\{r\}\\right\)^\{\\alpha\}\\left\(\\frac\{D\_\{\\alpha\}\}\{C\_\{\\alpha\}\}\\right\)^\{H\-1\}\\right\]^\{\-1\}\.Figure[4](https://arxiv.org/html/2609.38666#A6.F4)variesα\\alpha,rr, andδ\\deltaseparately around\(α,r,δ\)=\(0\.9,0\.99,0\.01\)\(\\alpha,r,\\delta\)=\(0\.9,0\.99,0\.01\), evaluating every integerHHfrom11to10001000\.

Figure 4:Sensitivity of first\-token probabilities to the horizon\. Top row: varyα\\alphawithr=0\.99r=0\.99andδ=0\.01\\delta=0\.01\. Middle row: varyrrwithα=0\.9\\alpha=0\.9andδ=0\.01\\delta=0\.01\. Bottom row: varyδ∈\{0\.01,0\.1,0\.3\}\\delta\\in\\\{0\.01,0\.1,0\.3\\\}withα=0\.9\\alpha=0\.9andr=0\.99r=0\.99\. The baseline is repeated in \(b\), \(f\), and \(g\)\. The horizontal axis is logarithmic; curves connect values computed at integer horizons\.Increasing the expert’s weight or first\-token confidence delays the reversal in the displayed settings\. Atr=0\.99r=0\.99andδ=0\.01\\delta=0\.01, the first horizon withRH<1/2R\_\{H\}<1/2is1010,6868, or717717forα=0\.5\\alpha=0\.5,0\.90\.9, or0\.990\.99, respectively\. Making the continuations afteraamore diffuse also delays reversal\. Forδ=0\.3\\delta=0\.3, the reverse probability first falls below1/21/2atH=555H=555, compared withH=68H=68atδ=0\.01\\delta=0\.01\. Forward aggregation is independent ofHHin every panel\. Thus, in this construction, the deterioration arises from the difference between the continuation distributions, while expert confidence and weight determine how long the initial preference persists\.

## Appendix GForward KL with Function Approximation

The tabular estimator learns a separate conditional distribution at each prefix\. Function approximation allows these conditionals to share a common representation\. LetΠ=\{πθ:θ∈Θ\}\\Pi=\\\{\\pi\_\{\\theta\}:\\theta\\in\\Theta\\\}be a finite class of autoregressive policies, where

πθ​\(y\|x\)=∏h=1Hπθ​\(ah\|x,y<h\)\.\\displaystyle\\pi\_\{\\theta\}\(y\|x\)=\\prod\_\{h=1\}^\{H\}\\pi\_\{\\theta\}\(a\_\{h\}\|x,y\_\{<h\}\)\.We make the standard realizability assumption for the function class\.

###### Assumption G\.1\.

There existsθ∗∈Θ\\theta^\{\*\}\\in\\Thetasuch that, for every contextxxand every responsey∈𝒴y\\in\\mathcal\{Y\},

πθ∗​\(y\|x\)=πF∗​\(y\|x\)\.\\displaystyle\\pi\_\{\\theta^\{\*\}\}\(y\|x\)=\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(y\|x\)\.

### G\.1Algorithm Design

We aggregate the candidate policies at the token level, using sampled tokens when logits are unavailable and full teacher probabilities otherwise\. Recall the token feedbackqt,j,h​\(a\)q\_\{t,j,h\}\(a\)defined in Section[4](https://arxiv.org/html/2609.38666#S4):

qt,j,h​\(a\):=\{𝟙⁡\(at,j,h=a\)w/o​logit,exp⁡\(Zt,j,h​\(a\)\)∑b∈𝒜exp⁡\(Zt,j,h​\(b\)\),w/logit\.\\displaystyle q\_\{t,j,h\}\(a\):=\\begin\{cases\}\\ind\(a\_\{t,j,h\}=a\)&\\mathrm\{w/o\\ logit\},\\\\ \\frac\{\\exp\(Z\_\{t,j,h\}\(a\)\)\}\{\\sum\_\{b\\in\\mathcal\{A\}\}\\exp\(Z\_\{t,j,h\}\(b\)\)\},&\\mathrm\{w/\\ logit\}\.\\end\{cases\}For eachθ∈Θ\\theta\\in\\Theta, define its empirical loss in roundttby

ℓt\(πθ\):=−1m∑j=1m∑h=1H∑a∈𝒜qt,j,h\(a\)logπθ\(a\|xt,j,yt,j,<h\)\.\\displaystyle\\ell\_\{t\}\(\\pi\_\{\\theta\}\):=\-\\frac\{1\}\{m\}\\sum\_\{j=1\}^\{m\}\\sum\_\{h=1\}^\{H\}\\sum\_\{a\\in\\mathcal\{A\}\}q\_\{t,j,h\}\(a\)\\log\\pi\_\{\\theta\}\(a\|x\_\{t,j\},y\_\{t,j,<h\}\)\.\(G\.1\)The main idea of the algorithm is to maintain a distributionpt​\(θ\)p\_\{t\}\(\\theta\)over the parameter spaceΘ\\Theta\. More specifically, choose a prior distributionp0p\_\{0\}onΘ\\Thetawithp0​\(θ\)\>0p\_\{0\}\(\\theta\)\>0for everyθ∈Θ\\theta\\in\\Theta, and initializeL0​\(θ\)=0L\_\{0\}\(\\theta\)=0\. At the beginning of roundtt, for any prefix\(x,u\)\(x,u\), the learner uses

π^t​\(a\|x,u\):=∑θ∈Θpt−1​\(θ\)​πθ​\(a\|x,u\)\.\\displaystyle\\widehat\{\\pi\}\_\{t\}\(a\|x,u\):=\\sum\_\{\\theta\\in\\Theta\}p\_\{t\-1\}\(\\theta\)\\pi\_\{\\theta\}\(a\|x,u\)\.\(G\.2\)After observing the feedback, the learner updates

pt​\(θ\)\\displaystyle p\_\{t\}\(\\theta\):=pt−1\(θ\)exp\(−ℓt\(πθ\)/H\)∑θ′∈Θpt−1\(θ′\)exp\(−ℓt\(πθ′\)/H\)\.\\displaystyle:=\\frac\{p\_\{t\-1\}\(\\theta\)\\exp\\bigl\(\-\\ell\_\{t\}\(\\pi\_\{\\theta\}\)/H\\bigr\)\}\{\\sum\_\{\\theta^\{\\prime\}\\in\\Theta\}p\_\{t\-1\}\(\\theta^\{\\prime\}\)\\exp\\bigl\(\-\\ell\_\{t\}\(\\pi\_\{\\theta^\{\\prime\}\}\)/H\\bigr\)\}\.\(G\.3\)The factor1/H1/Hnormalizes the sum of theHHtoken losses in each rollout\. Intuitively, the update favors candidates with smaller cumulative empirical losses, relative to their prior weights\.

Algorithm 3Forward KL with Function Approximation1:Input:Finite policy class

Π=\{πθ:θ∈Θ\}\\Pi=\\\{\\pi\_\{\\theta\}:\\theta\\in\\Theta\\\}, prior

p0p\_\{0\}with

p0​\(θ\)\>0p\_\{0\}\(\\theta\)\>0for every

θ\\theta, number of rounds

TT\.

2:for

t=1,…,Tt=1,\\ldots,Tdo

3:For any prefix

\(x,u\)\(x,u\), construct the policy

π^t​\(a\|x,u\)\\widehat\{\\pi\}\_\{t\}\(a\|x,u\)as in \([G\.2](https://arxiv.org/html/2609.38666#A7.E2)\)\.

4:Observe the batch of teacher\-generated rollouts

\{\(xt,j,yt,j\)\}j=1m\\\{\(x\_\{t,j\},y\_\{t,j\}\)\\\}\_\{j=1\}^\{m\}and, when available, the teacher logits

\{Zt,j,h\}j,h\\\{Z\_\{t,j,h\}\\\}\_\{j,h\}\.

5:For every

θ∈Θ\\theta\\in\\Theta, compute

ℓt​\(πθ\)\\ell\_\{t\}\(\\pi\_\{\\theta\}\)as in \([G\.1](https://arxiv.org/html/2609.38666#A7.E1)\)

6:For every

θ∈Θ\\theta\\in\\Theta, update the mixture weights

pt​\(θ\)p\_\{t\}\(\\theta\)as in \([G\.3](https://arxiv.org/html/2609.38666#A7.E3)\)

7:endfor

8:Output:

\{π^t\}t=1T\\\{\\widehat\{\\pi\}\_\{t\}\\\}\_\{t=1\}^\{T\}\.

This exponential reweighting resembles the posterior updates used in Thompson sampling\. Unlike Thompson sampling, however, our algorithm predicts with a mixture of the candidate policies rather than a single sampled candidate\. Evaluating this mixture at each decoding step introduces large computational cost, which is a key limitation of our forward\-KL algorithm with function approximation\.

### G\.2Theoretical Guarantee

In this section, we first work on a finite function classΘ\\Theta\. We will apply the results to the setting with finite covering number later\. The following theorem bounds the regret by the prior weight of a good candidate and its approximation error\. Under realizability, the approximation term vanishes\.

###### Theorem G\.2\.

LetT≥2T\\geq 2\. Suppose the regret in \([3\.2](https://arxiv.org/html/2609.38666#S3.E2)\) is defined withD\(p∥q\)=KL\(q∥p\)D\(p\\\|q\)=\\text\{KL\}\(q\\\|p\)\. Under off\-policy distillation protocol and the function approximation setting, both with and without access to teacher logits, Algorithm[3](https://arxiv.org/html/2609.38666#alg3)satisfies

𝔼⁡\[Regret⁡\(T\)\]\\displaystyle\\mathbb\{E\}\\big\[\\operatorname\{Regret\}\(T\)\\big\]≤infθ∈Θ\{Hlog1p0​\(θ\)\+T𝔼x∼ρ¯KL\(πF∗\(⋅\|x\)∥πθ\(⋅\|x\)\)\},\\displaystyle\\leq\\inf\_\{\\theta\\in\\Theta\}\\bigg\\\{H\\log\\frac\{1\}\{p\_\{0\}\(\\theta\)\}\+T\\mathbb\{E\}\_\{x\\sim\\bar\{\\rho\}\}\\text\{KL\}\\bigl\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(\\cdot\|x\)\\,\\big\\\|\\,\\pi\_\{\\theta\}\(\\cdot\|x\)\\bigr\)\\bigg\\\},provided the right\-hand side is finite\. In particular, under Assumption[G\.1](https://arxiv.org/html/2609.38666#A7.Thmtheorem1), for the optimal parameterθ∗\\theta^\{\*\}, Algorithm[3](https://arxiv.org/html/2609.38666#alg3)satisfies

𝔼⁡\[Regret⁡\(T\)\]≤H​log⁡1p0​\(θ∗\)\.\\displaystyle\\mathbb\{E\}\\big\[\\operatorname\{Regret\}\(T\)\\big\]\\leq H\\log\\frac\{1\}\{p\_\{0\}\(\\theta^\{\*\}\)\}\.Moreover, for anyδ\>0\\delta\>0, under Assumption[G\.1](https://arxiv.org/html/2609.38666#A7.Thmtheorem1), the following inequality holds with probability at least1−δ1\-\\delta:

Regret⁡\(T\)≤2​H​log⁡1p0​\(θ∗\)\+4​H​\(1\+log⁡2δ​p0​\(θ∗\)\)​log​2δ\.\\displaystyle\\operatorname\{Regret\}\(T\)\\leq 2H\\log\\frac\{1\}\{p\_\{0\}\(\\theta^\{\*\}\)\}\+4H\\bigg\(1\+\\log\\frac\{2\}\{\\delta p\_\{0\}\(\\theta^\{\*\}\)\}\\bigg\)\\log\\frac\{2\}\{\\delta\}\.

###### Proof\.

Letℱt\\mathcal\{F\}\_\{t\}be theσ\\sigma\-algebra generated by all rollouts and feedback observed before roundtt:

ℱt:=σ\(\(xs,j,ys,j\),qs,j,h\(a\):s<t,j∈\[m\],h∈\[H\],a∈𝒜\)\.\\displaystyle\\mathcal\{F\}\_\{t\}:=\\sigma\\Big\(\(x\_\{s,j\},y\_\{s,j\}\),q\_\{s,j,h\}\(a\):s<t,\\ j\\in\[m\],\\ h\\in\[H\],\\ a\\in\\mathcal\{A\}\\Big\)\.Thus, bothpt−1p\_\{t\-1\}andπ^t\\widehat\{\\pi\}\_\{t\}areℱt\\mathcal\{F\}\_\{t\}\-measurable\. Recall the feedback countsct​\(x,u,a\)c\_\{t\}\(x,u,a\)defined in \([4\.3](https://arxiv.org/html/2609.38666#S4.E3)\)\. For any policyπ\\pi, \([G\.1](https://arxiv.org/html/2609.38666#A7.E1)\) can be rewritten as

ℓt\(π\)=−∑h=1H∑\(x,u,a\)ct\(x,u,a\)logπ\(a\|x,u\)\.\\displaystyle\\ell\_\{t\}\(\\pi\)=\-\\sum\_\{h=1\}^\{H\}\\sum\_\{\(x,u,a\)\}c\_\{t\}\(x,u,a\)\\log\\pi\(a\|x,u\)\.\(G\.4\)Easy to see, the following equation holds

∑\(x,u,a\)ct​\(x,u,a\)=1\\displaystyle\\sum\_\{\(x,u,a\)\}c\_\{t\}\(x,u,a\)=1Therefore, we have

∑h=1H∑\(x,u,a\)ct​\(x,u,a\)H=1\.\\displaystyle\\sum\_\{h=1\}^\{H\}\\sum\_\{\(x,u,a\)\}\\frac\{c\_\{t\}\(x,u,a\)\}\{H\}=1\.\(G\.5\)For anyt≥1t\\geq 1, define the normalizing constant

Ct:=∑θ∈Θp0\(θ\)exp\(−1H∑s=1tℓs\(πθ\)\),\\displaystyle C\_\{t\}:=\\sum\_\{\\theta\\in\\Theta\}p\_\{0\}\(\\theta\)\\exp\\bigg\(\-\\frac\{1\}\{H\}\\sum\_\{s=1\}^\{t\}\\ell\_\{s\}\(\\pi\_\{\\theta\}\)\\bigg\),andC0=1C\_\{0\}=1\. Direct iteration of \([G\.3](https://arxiv.org/html/2609.38666#A7.E3)\) gives

pt​\(θ\)=p0\(θ\)exp\(−H−1∑s=1tℓs\(πθ\)\)Ct\.\\displaystyle p\_\{t\}\(\\theta\)=\\frac\{p\_\{0\}\(\\theta\)\\exp\\big\(\-H^\{\-1\}\\sum\_\{s=1\}^\{t\}\\ell\_\{s\}\(\\pi\_\{\\theta\}\)\\big\)\}\{C\_\{t\}\}\.Moreover, we have

CtCt−1\\displaystyle\\frac\{C\_\{t\}\}\{C\_\{t\-1\}\}=1Ct−1∑θ′∈Θp0\(θ′\)exp\(−1H∑s=1t−1ℓs\(πθ′\)−ℓt​\(πθ′\)H\)\\displaystyle=\\frac\{1\}\{C\_\{t\-1\}\}\\sum\_\{\\theta^\{\\prime\}\\in\\Theta\}p\_\{0\}\(\\theta^\{\\prime\}\)\\exp\\bigg\(\-\\frac\{1\}\{H\}\\sum\_\{s=1\}^\{t\-1\}\\ell\_\{s\}\(\\pi\_\{\\theta^\{\\prime\}\}\)\-\\frac\{\\ell\_\{t\}\(\\pi\_\{\\theta^\{\\prime\}\}\)\}\{H\}\\bigg\)=∑θ′∈Θp0\(θ′\)exp\(−H−1∑s=1t−1ℓs\(πθ′\)\)Ct−1⏟=pt−1​\(θ′\)exp\(−ℓt\(πθ′\)/H\)\\displaystyle=\\sum\_\{\\theta^\{\\prime\}\\in\\Theta\}\\underbrace\{\\frac\{p\_\{0\}\(\\theta^\{\\prime\}\)\\exp\\big\(\-H^\{\-1\}\\sum\_\{s=1\}^\{t\-1\}\\ell\_\{s\}\(\\pi\_\{\\theta^\{\\prime\}\}\)\\big\)\}\{C\_\{t\-1\}\}\}\_\{=\\,p\_\{t\-1\}\(\\theta^\{\\prime\}\)\}\\exp\\bigl\(\-\\ell\_\{t\}\(\\pi\_\{\\theta^\{\\prime\}\}\)/H\\bigr\)=∑θ′∈Θpt−1\(θ′\)exp\(−ℓt\(πθ′\)/H\)\.\\displaystyle=\\sum\_\{\\theta^\{\\prime\}\\in\\Theta\}p\_\{t\-1\}\(\\theta^\{\\prime\}\)\\exp\\bigl\(\-\\ell\_\{t\}\(\\pi\_\{\\theta^\{\\prime\}\}\)/H\\bigr\)\.\(G\.6\)The definition ofℓt\\ell\_\{t\}in \([G\.4](https://arxiv.org/html/2609.38666#A7.E4)\) gives

exp\(−ℓt\(πθ′\)/H\)\\displaystyle\\exp\\bigl\(\-\\ell\_\{t\}\(\\pi\_\{\\theta^\{\\prime\}\}\)/H\\bigr\)=exp⁡\(∑h=1H∑\(x,u,a\)ct​\(x,u,a\)H​log⁡πθ′​\(a\|x,u\)\)\\displaystyle=\\exp\\bigg\(\\sum\_\{h=1\}^\{H\}\\sum\_\{\(x,u,a\)\}\\frac\{c\_\{t\}\(x,u,a\)\}\{H\}\\log\\pi\_\{\\theta^\{\\prime\}\}\(a\|x,u\)\\bigg\)=∏h=1H∏\(x,u,a\)πθ′​\(a\|x,u\)ct​\(x,u,a\)/H\.\\displaystyle=\\prod\_\{h=1\}^\{H\}\\prod\_\{\(x,u,a\)\}\\pi\_\{\\theta^\{\\prime\}\}\(a\|x,u\)^\{c\_\{t\}\(x,u,a\)/H\}\.\(G\.7\)Substituting \([G\.7](https://arxiv.org/html/2609.38666#A7.E7)\) into \([G\.6](https://arxiv.org/html/2609.38666#A7.E6)\), we have

CtCt−1=∑θ′∈Θpt−1​\(θ′\)​∏h=1H∏\(x,u,a\)πθ′​\(a\|x,u\)ct​\(x,u,a\)/H\.\\displaystyle\\frac\{C\_\{t\}\}\{C\_\{t\-1\}\}=\\sum\_\{\\theta^\{\\prime\}\\in\\Theta\}p\_\{t\-1\}\(\\theta^\{\\prime\}\)\\prod\_\{h=1\}^\{H\}\\prod\_\{\(x,u,a\)\}\\pi\_\{\\theta^\{\\prime\}\}\(a\|x,u\)^\{c\_\{t\}\(x,u,a\)/H\}\.\(G\.8\)Using the generalized Hölder’s inequality, we have

CtCt−1\\displaystyle\\frac\{C\_\{t\}\}\{C\_\{t\-1\}\}≤∏h=1H∏\(x,u,a\)\(∑θ′∈Θpt−1​\(θ′\)​πθ′​\(a\|x,u\)\)ct​\(x,u,a\)/H\\displaystyle\\leq\\prod\_\{h=1\}^\{H\}\\prod\_\{\(x,u,a\)\}\\bigg\(\\sum\_\{\\theta^\{\\prime\}\\in\\Theta\}p\_\{t\-1\}\(\\theta^\{\\prime\}\)\\pi\_\{\\theta^\{\\prime\}\}\(a\|x,u\)\\bigg\)^\{c\_\{t\}\(x,u,a\)/H\}=∏h=1H∏\(x,u,a\)π^t​\(a\|x,u\)ct​\(x,u,a\)/H\\displaystyle=\\prod\_\{h=1\}^\{H\}\\prod\_\{\(x,u,a\)\}\\widehat\{\\pi\}\_\{t\}\(a\|x,u\)^\{c\_\{t\}\(x,u,a\)/H\}=exp\(−ℓt\(π^t\)/H\),\\displaystyle=\\exp\\bigl\(\-\\ell\_\{t\}\(\\widehat\{\\pi\}\_\{t\}\)/H\\bigr\),where the last equation holds using \([G\.4](https://arxiv.org/html/2609.38666#A7.E4)\)\.

Taking the negative logarithm, and summing overt∈\[T\]t\\in\[T\]yields

∑t=1Tℓt​\(π^t\)\\displaystyle\\sum\_\{t=1\}^\{T\}\\ell\_\{t\}\(\\widehat\{\\pi\}\_\{t\}\)≤−H∑t=1TlogCtCt−1=−HlogCT,\\displaystyle\\leq\-H\\sum\_\{t=1\}^\{T\}\\log\\frac\{C\_\{t\}\}\{C\_\{t\-1\}\}=\-H\\log C\_\{T\},\(G\.9\)where the equality usesC0=1C\_\{0\}=1\.

Now fix any comparator policyπθ\\pi\_\{\\theta\}such thatθ∈Θ\\theta\\in\\Theta\. Easy to see

CT\\displaystyle C\_\{T\}=∑θ′p0\(θ′\)exp\(−1H∑t=1Tℓt\(πθ′\)\)\\displaystyle=\\sum\_\{\\theta^\{\\prime\}\}p\_\{0\}\(\\theta^\{\\prime\}\)\\exp\\bigg\(\-\\frac\{1\}\{H\}\\sum\_\{t=1\}^\{T\}\\ell\_\{t\}\(\\pi\_\{\\theta^\{\\prime\}\}\)\\bigg\)≥p0\(θ\)exp\(−1H∑t=1Tℓt\(πθ\)\)\.\\displaystyle\\geq p\_\{0\}\(\\theta\)\\exp\\bigg\(\-\\frac\{1\}\{H\}\\sum\_\{t=1\}^\{T\}\\ell\_\{t\}\(\\pi\_\{\\theta\}\)\\bigg\)\.Therefore, we have

−∑t=1Tℓt\(πθ\)≤HlogCT\+Hlog1p0​\(θ\)\.\\displaystyle\-\\sum\_\{t=1\}^\{T\}\\ell\_\{t\}\(\\pi\_\{\\theta\}\)\\leq H\\log C\_\{T\}\+H\\log\\frac\{1\}\{p\_\{0\}\(\\theta\)\}\.\(G\.10\)Summing up \([G\.9](https://arxiv.org/html/2609.38666#A7.E9)\) and \([G\.10](https://arxiv.org/html/2609.38666#A7.E10)\), we have

∑t=1T\[ℓt​\(π^t\)−ℓt​\(πθ\)\]≤H​log⁡1p0​\(θ\)\.\\displaystyle\\sum\_\{t=1\}^\{T\}\\big\[\\ell\_\{t\}\(\\widehat\{\\pi\}\_\{t\}\)\-\\ell\_\{t\}\(\\pi\_\{\\theta\}\)\\big\]\\leq H\\log\\frac\{1\}\{p\_\{0\}\(\\theta\)\}\.\(G\.11\)As a result,

∑t=1T\[ℓt​\(π^t\)−ℓt​\(πF∗\)\]\\displaystyle\\sum\_\{t=1\}^\{T\}\\big\[\\ell\_\{t\}\(\\widehat\{\\pi\}\_\{t\}\)\-\\ell\_\{t\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\\big\]=∑t=1T\[ℓt​\(π^t\)−ℓt​\(πθ\)\]\+∑t=1T\[ℓt​\(πθ\)−ℓt​\(πF∗\)\]\\displaystyle=\\sum\_\{t=1\}^\{T\}\\big\[\\ell\_\{t\}\(\\widehat\{\\pi\}\_\{t\}\)\-\\ell\_\{t\}\(\\pi\_\{\\theta\}\)\\big\]\+\\sum\_\{t=1\}^\{T\}\\big\[\\ell\_\{t\}\(\\pi\_\{\\theta\}\)\-\\ell\_\{t\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\\big\]≤H​log⁡1p0​\(θ\)\+∑t=1T\[ℓt​\(πθ\)−ℓt​\(πF∗\)\]\.\\displaystyle\\qquad\\leq H\\log\\frac\{1\}\{p\_\{0\}\(\\theta\)\}\+\\sum\_\{t=1\}^\{T\}\\big\[\\ell\_\{t\}\(\\pi\_\{\\theta\}\)\-\\ell\_\{t\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\\big\]\.\(G\.12\)Using the same argument as in Appendix[B](https://arxiv.org/html/2609.38666#A2), we have

𝔼⁡\[ct​\(x,u,a\)∣ℱt\]\\displaystyle\\mathbb\{E\}\[c\_\{t\}\(x,u,a\)\\mid\\mathcal\{F\}\_\{t\}\]=1I​∑i∈ℐρi​\(x\)​pi​\(u\|x\)​pi​\(a\|x,u\)\\displaystyle=\\frac\{1\}\{I\}\\sum\_\{i\\in\\mathcal\{I\}\}\\rho\_\{i\}\(x\)p\_\{i\}\(u\|x\)p\_\{i\}\(a\|x,u\)=ρ¯​\(x\)​πF∗​\(u\|x\)​πF∗​\(a\|x,u\)\.\\displaystyle=\\bar\{\\rho\}\(x\)\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(u\|x\)\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\.Consequently, using \([G\.4](https://arxiv.org/html/2609.38666#A7.E4)\), for anyℱt\\mathcal\{F\}\_\{t\}\-measurable policyπ\\piwith finite loss,

𝔼⁡\[ℓt​\(π\)−ℓt​\(πF∗\)∣ℱt\]\\displaystyle\\mathbb\{E\}\\big\[\\ell\_\{t\}\(\\pi\)\-\\ell\_\{t\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\\mid\\mathcal\{F\}\_\{t\}\\big\]=∑h=1H∑\(x,u,a\)𝔼⁡\[ct​\(x,u,a\)∣ℱt\]​log⁡πF∗​\(a\|x,u\)π⁡\(a\|x,u\)\\displaystyle\\quad=\\sum\_\{h=1\}^\{H\}\\sum\_\{\(x,u,a\)\}\\mathbb\{E\}\[c\_\{t\}\(x,u,a\)\\mid\\mathcal\{F\}\_\{t\}\]\\log\\frac\{\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\}\{\\pi\(a\|x,u\)\}=∑h=1H∑\(x,u\)ρ¯\(x\)πF∗\(u\|x\)KL\(πF∗\(⋅\|x,u\)∥π\(⋅\|x,u\)\)\\displaystyle\\quad=\\sum\_\{h=1\}^\{H\}\\sum\_\{\(x,u\)\}\\bar\{\\rho\}\(x\)\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(u\|x\)\\text\{KL\}\\big\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(\\cdot\|x,u\)\\,\\big\\\|\\,\\pi\(\\cdot\|x,u\)\\big\)=𝔼x∼ρ¯KL\(πF∗\(⋅\|x\)∥π\(⋅\|x\)\)\.\\displaystyle\\quad=\\mathbb\{E\}\_\{x\\sim\\bar\{\\rho\}\}\\text\{KL\}\\big\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(\\cdot\|x\)\\,\\big\\\|\\,\\pi\(\\cdot\|x\)\\big\)\.\(G\.13\)The last equality follows from the chain rule for KL divergence\.

Taking expectations in \([G\.12](https://arxiv.org/html/2609.38666#A7.E12)\) and applying \([G\.13](https://arxiv.org/html/2609.38666#A7.E13)\) to bothπ^t\\widehat\{\\pi\}\_\{t\}andπθ\\pi\_\{\\theta\}, we conclude that

𝔼⁡\[Regret⁡\(T\)\]\\displaystyle\\mathbb\{E\}\[\\operatorname\{Regret\}\(T\)\]≤Hlog1p0​\(θ\)\+∑t=1T𝔼x∼ρ¯KL\(πF∗\(⋅\|x\)∥πθ\(⋅\|x\)\)\\displaystyle\\leq H\\log\\frac\{1\}\{p\_\{0\}\(\\theta\)\}\+\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\_\{x\\sim\\bar\{\\rho\}\}\\text\{KL\}\\big\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(\\cdot\|x\)\\,\\big\\\|\\,\\pi\_\{\\theta\}\(\\cdot\|x\)\\big\)=Hlog1p0​\(θ\)\+T𝔼x∼ρ¯KL\(πF∗\(⋅\|x\)∥πθ\(⋅\|x\)\)\.\\displaystyle=H\\log\\frac\{1\}\{p\_\{0\}\(\\theta\)\}\+T\\mathbb\{E\}\_\{x\\sim\\bar\{\\rho\}\}\\text\{KL\}\\big\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(\\cdot\|x\)\\,\\big\\\|\\,\\pi\_\{\\theta\}\(\\cdot\|x\)\\big\)\.Since the inequality holds for every fixedθ\\theta, taking the infimum gives

𝔼\[Regret\(T\)\]≤infθ∈Θ\{Hlog1p0​\(θ\)\+T𝔼x∼ρ¯KL\(πF∗\(⋅\|x\)∥πθ\(⋅\|x\)\)\}\.\\displaystyle\\mathbb\{E\}\[\\operatorname\{Regret\}\(T\)\]\\leq\\inf\_\{\\theta\\in\\Theta\}\\bigg\\\{H\\log\\frac\{1\}\{p\_\{0\}\(\\theta\)\}\+T\\mathbb\{E\}\_\{x\\sim\\bar\{\\rho\}\}\\text\{KL\}\\big\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(\\cdot\|x\)\\,\\big\\\|\\,\\pi\_\{\\theta\}\(\\cdot\|x\)\\big\)\\bigg\\\}\.\(G\.14\)

#### High\-probability bound under realizability\.

Supposeπθ∗=πF∗\\pi\_\{\\theta^\{\*\}\}=\\pi\_\{\\mathrm\{F\}\}^\{\*\}for someθ∗∈Θ\\theta^\{\*\}\\in\\Theta\. As in Appendix[B](https://arxiv.org/html/2609.38666#A2), define

Xt\\displaystyle X\_\{t\}:=ℓt​\(πF∗\)−ℓt​\(π^t\),\\displaystyle:=\\ell\_\{t\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\-\\ell\_\{t\}\(\\widehat\{\\pi\}\_\{t\}\),rt\\displaystyle r\_\{t\}:=−𝔼⁡\[Xt∣ℱt\]\.\\displaystyle:=\-\\mathbb\{E\}\[X\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]\.By \([G\.13](https://arxiv.org/html/2609.38666#A7.E13)\),Regret⁡\(T\)=∑t=1Trt\\operatorname\{Regret\}\(T\)=\\sum\_\{t=1\}^\{T\}r\_\{t\}\. The argument establishing \([B\.20](https://arxiv.org/html/2609.38666#A2.E20)\) applies to anyℱt\\mathcal\{F\}\_\{t\}\-measurable policyπ\\pi, since it only uses policy normalization and the conditional mean of the feedback counts\. Therefore,

𝔼⁡\[exp⁡\(XtH\)\|ℱt\]=𝔼⁡\[exp⁡\(ℓt​\(πF∗\)−ℓt​\(π^t\)H\)\|ℱt\]≤1\.\\displaystyle\\mathbb\{E\}\\Big\[\\exp\\Big\(\\frac\{X\_\{t\}\}\{H\}\\Big\)\\Big\|\\mathcal\{F\}\_\{t\}\\Big\]=\\mathbb\{E\}\\bigg\[\\exp\\bigg\(\\frac\{\\ell\_\{t\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\-\\ell\_\{t\}\(\\widehat\{\\pi\}\_\{t\}\)\}\{H\}\\bigg\)\\,\\bigg\|\\,\\mathcal\{F\}\_\{t\}\\bigg\]\\leq 1\.\(G\.15\)
To replace the lower bound provided by KT smoothing in the tabular proof, we control the mixture weight ofθ∗\\theta^\{\*\}\. DefineW0=1W\_\{0\}=1and

Wt\\displaystyle W\_\{t\}:=Ct​exp⁡\(1H​∑s=1tℓs​\(πF∗\)\)=p0​\(θ∗\)pt​\(θ∗\)\.\\displaystyle:=C\_\{t\}\\exp\\bigg\(\\frac\{1\}\{H\}\\sum\_\{s=1\}^\{t\}\\ell\_\{s\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\\bigg\)=\\frac\{p\_\{0\}\(\\theta^\{\*\}\)\}\{p\_\{t\}\(\\theta^\{\*\}\)\}\.The equality follows from the formula forpt​\(θ∗\)p\_\{t\}\(\\theta^\{\*\}\)and realizability\. By \([G\.6](https://arxiv.org/html/2609.38666#A7.E6)\),

WtWt−1\\displaystyle\\frac\{W\_\{t\}\}\{W\_\{t\-1\}\}=CtCt−1​exp⁡\(ℓt​\(πF∗\)/H\)\\displaystyle=\\frac\{C\_\{t\}\}\{C\_\{t\-1\}\}\\exp\\bigl\(\\ell\_\{t\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)/H\\bigr\)=∑θ∈Θpt−1​\(θ\)​exp⁡\(ℓt​\(πF∗\)−ℓt​\(πθ\)H\),\\displaystyle=\\sum\_\{\\theta\\in\\Theta\}p\_\{t\-1\}\(\\theta\)\\exp\\bigg\(\\frac\{\\ell\_\{t\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\-\\ell\_\{t\}\(\\pi\_\{\\theta\}\)\}\{H\}\\bigg\),where we apply \([G\.6](https://arxiv.org/html/2609.38666#A7.E6)\)\. SinceWt−1W\_\{t\-1\}andpt−1p\_\{t\-1\}areℱt\\mathcal\{F\}\_\{t\}\-measurable, \([G\.15](https://arxiv.org/html/2609.38666#A7.E15)\) implies

𝔼⁡\[Wt∣ℱt\]\\displaystyle\\mathbb\{E\}\[W\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]=Wt−1​∑θ∈Θpt−1​\(θ\)​𝔼​\[exp⁡\(ℓt​\(πF∗\)−ℓt​\(πθ\)H\)\|ℱt\]\\displaystyle=W\_\{t\-1\}\\sum\_\{\\theta\\in\\Theta\}p\_\{t\-1\}\(\\theta\)\\mathbb\{E\}\\bigg\[\\exp\\bigg\(\\frac\{\\ell\_\{t\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\-\\ell\_\{t\}\(\\pi\_\{\\theta\}\)\}\{H\}\\bigg\)\\,\\bigg\|\\,\\mathcal\{F\}\_\{t\}\\bigg\]≤Wt−1,\\displaystyle\\leq W\_\{t\-1\},where the last inequality holds due to \([G\.15](https://arxiv.org/html/2609.38666#A7.E15)\)\. Thus,WtW\_\{t\}is a nonnegative supermartingale with respect to\{ℱt\+1\}t=0T\\\{\\mathcal\{F\}\_\{t\+1\}\\\}\_\{t=0\}^\{T\}\. Using Ville’s inequality \(Lemma[I\.9](https://arxiv.org/html/2609.38666#A9.Thmtheorem9)\), we have

ℙ⁡\(max0≤t≤T⁡Wt≤2δ\)≥1−δ2\.\\displaystyle\\mathbb\{P\}\\bigg\(\\max\_\{0\\leq t\\leq T\}W\_\{t\}\\leq\\frac\{2\}\{\\delta\}\\bigg\)\\geq 1\-\\frac\{\\delta\}\{2\}\.\(G\.16\)Define

Jt\\displaystyle J\_\{t\}:=𝟙⁡\(max0≤s<t⁡Ws≤2δ\),L:=log⁡2δ​p0​\(θ∗\)\.\\displaystyle:=\\ind\\bigg\(\\max\_\{0\\leq s<t\}W\_\{s\}\\leq\\frac\{2\}\{\\delta\}\\bigg\),\\qquad L:=\\log\\frac\{2\}\{\\delta p\_\{0\}\(\\theta^\{\*\}\)\}\.WheneverJt=1J\_\{t\}=1, the token\-level mixture satisfies

π^t​\(a\|x,u\)πF∗​\(a\|x,u\)≥pt−1​\(θ∗\)=p0​\(θ∗\)Wt−1≥e−L,\\displaystyle\\frac\{\\widehat\{\\pi\}\_\{t\}\(a\|x,u\)\}\{\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\}\\geq p\_\{t\-1\}\(\\theta^\{\*\}\)=\\frac\{p\_\{0\}\(\\theta^\{\*\}\)\}\{W\_\{t\-1\}\}\\geq e^\{\-L\},where the first inequality holds due toπ^t​\(a\|x,u\):=∑θpt−1​\(θ\)​πθ​\(a\|x,u\)≥pt−1​\(θ∗\)​πF∗​\(a\|x,u\)\\widehat\{\\pi\}\_\{t\}\(a\|x,u\):=\\sum\_\{\\theta\}p\_\{t\-1\}\(\\theta\)\\pi\_\{\\theta\}\(a\|x,u\)\\geq p\_\{t\-1\}\(\\theta^\{\*\}\)\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\. The second inequality holds due to \([G\.16](https://arxiv.org/html/2609.38666#A7.E16)\)\. Thus,

log⁡π^t​\(a\|x,u\)πF∗​\(a\|x,u\)≥−L\\displaystyle\\log\\frac\{\\widehat\{\\pi\}\_\{t\}\(a\|x,u\)\}\{\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\}\\geq\-Lfor every\(x,u,a\)\(x,u,a\)\. Using the definition ofXtX\_\{t\}and the loss representation, we obtain

XtH\\displaystyle\\frac\{X\_\{t\}\}\{H\}=ℓt​\(πF∗\)−ℓt​\(π^t\)H\\displaystyle=\\frac\{\\ell\_\{t\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\-\\ell\_\{t\}\(\\widehat\{\\pi\}\_\{t\}\)\}\{H\}=∑h=1H∑\(x,u,a\)ct​\(x,u,a\)H​log⁡π^t​\(a\|x,u\)πF∗​\(a\|x,u\)\\displaystyle=\\sum\_\{h=1\}^\{H\}\\sum\_\{\(x,u,a\)\}\\frac\{c\_\{t\}\(x,u,a\)\}\{H\}\\log\\frac\{\\widehat\{\\pi\}\_\{t\}\(a\|x,u\)\}\{\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\}≥−L∑h=1H∑\(x,u,a\)ct​\(x,u,a\)H\\displaystyle\\geq\-L\\sum\_\{h=1\}^\{H\}\\sum\_\{\(x,u,a\)\}\\frac\{c\_\{t\}\(x,u,a\)\}\{H\}=−L\.\\displaystyle=\-L\.The inequality usesct​\(x,u,a\)≥0c\_\{t\}\(x,u,a\)\\geq 0, and the last equality follows from \([G\.5](https://arxiv.org/html/2609.38666#A7.E5)\)\.

WhenJt=0J\_\{t\}=0, we instead haveJt​Xt/H=0≥−LJ\_\{t\}X\_\{t\}/H=0\\geq\-L, sinceL≥0L\\geq 0\. Combining the two cases gives

Jt​XtH≥−L\.\\displaystyle\\frac\{J\_\{t\}X\_\{t\}\}\{H\}\\geq\-L\.Moreover, sinceJtJ\_\{t\}isℱt\\mathcal\{F\}\_\{t\}\-measurable,

𝔼⁡\[eJt​Xt/H∣ℱt\]\\displaystyle\\mathbb\{E\}\[e^\{J\_\{t\}X\_\{t\}/H\}\\mid\\mathcal\{F\}\_\{t\}\]=1∗𝟙⁡\(Jt=0\)\+𝟙⁡\(Jt=1\)​𝔼​\[eXt/H∣ℱt\]\\displaystyle=1\*\\ind\(J\_\{t\}=0\)\+\\ind\(J\_\{t\}=1\)\\mathbb\{E\}\[e^\{X\_\{t\}/H\}\\mid\\mathcal\{F\}\_\{t\}\]=1−Jt\+Jt​𝔼​\[eXt/H∣ℱt\]\\displaystyle=1\-J\_\{t\}\+J\_\{t\}\\mathbb\{E\}\[e^\{X\_\{t\}/H\}\\mid\\mathcal\{F\}\_\{t\}\]≤1,\\displaystyle\\leq 1,where the last inequality holds due to \([G\.15](https://arxiv.org/html/2609.38666#A7.E15)\)\.

Hence, Lemma[B\.2](https://arxiv.org/html/2609.38666#A2.Thmtheorem2)and the argument leading to \([B\.22](https://arxiv.org/html/2609.38666#A2.E22)\) apply toJt​XtJ\_\{t\}X\_\{t\}, whose conditional means are−Jt​rt\-J\_\{t\}r\_\{t\}\. ReplacingLTL\_\{T\}byLLand the failure probability byδ/2\\delta/2, we obtain, for any fixedλ∈\(0,1/H\]\\lambda\\in\(0,1/H\], with probability at least1−δ/21\-\\delta/2,

∑t=1TJt​\(Xt\+rt\)≤H⁡\(1\+L\)​λ​∑t=1TJt​rt\+log⁡\(2/δ\)λ\.\\displaystyle\\sum\_\{t=1\}^\{T\}J\_\{t\}\(X\_\{t\}\+r\_\{t\}\)\\leq H\(1\+L\)\\lambda\\sum\_\{t=1\}^\{T\}J\_\{t\}r\_\{t\}\+\\frac\{\\log\(2/\\delta\)\}\{\\lambda\}\.
Chooseλ=\[2​H​\(1\+L\)\]−1\\lambda=\[2H\(1\+L\)\]^\{\-1\}\. Conditioned on the event in \([G\.16](https://arxiv.org/html/2609.38666#A7.E16)\), and taking a union bound, we conclude that, with probability at least1−δ1\-\\delta,

Regret\(T\)≤−∑t=1TXt\+12Regret\(T\)\+2H\(1\+L\)log2δ\.\\displaystyle\\operatorname\{Regret\}\(T\)\\leq\-\\sum\_\{t=1\}^\{T\}X\_\{t\}\+\\frac\{1\}\{2\}\\operatorname\{Regret\}\(T\)\+2H\(1\+L\)\\log\\frac\{2\}\{\\delta\}\.Finally, \([G\.11](https://arxiv.org/html/2609.38666#A7.E11)\) withθ=θ∗\\theta=\\theta^\{\*\}gives

−∑t=1TXt≤Hlog1p0​\(θ∗\)\.\\displaystyle\-\\sum\_\{t=1\}^\{T\}X\_\{t\}\\leq H\\log\\frac\{1\}\{p\_\{0\}\(\\theta^\{\*\}\)\}\.Substituting this bound and rearranging yields

Regret⁡\(T\)≤2​H​log⁡1p0​\(θ∗\)\+4​H​\(1\+log⁡2δ​p0​\(θ∗\)\)​log​2δ\.\\displaystyle\\operatorname\{Regret\}\(T\)\\leq 2H\\log\\frac\{1\}\{p\_\{0\}\(\\theta^\{\*\}\)\}\+4H\\left\(1\+\\log\\frac\{2\}\{\\delta p\_\{0\}\(\\theta^\{\*\}\)\}\\right\)\\log\\frac\{2\}\{\\delta\}\.∎

### G\.3Function Class with Finite Covering Number

LetΠ=\{πθ:θ∈Θ\}\\Pi=\\\{\\pi\_\{\\theta\}:\\theta\\in\\Theta\\\}be a class of autoregressive policies containingπF∗\\pi\_\{\\mathrm\{F\}\}^\{\*\}\. Forε\>0\\varepsilon\>0, we define theε\\varepsilon\-covering number of the token log\-probability classlog⁡Π\\log\\Pias follows\.

###### Definition G\.3\(Covering number\)\.

The covering number𝒩\(ε,logΠ,∥⋅∥∞\)\\mathcal\{N\}\(\\varepsilon,\\log\\Pi,\\\|\\cdot\\\|\_\{\\infty\}\)is the smallest cardinality of a finite subsetΠε⊆Π\\Pi\_\{\\varepsilon\}\\subseteq\\Pisuch that, for everyπ∈Π\\pi\\in\\Pi, there existsπ~∈Πε\\widetilde\{\\pi\}\\in\\Pi\_\{\\varepsilon\}satisfying

‖log⁡π−log⁡π~‖∞\\displaystyle\\\|\\log\\pi\-\\log\\widetilde\{\\pi\}\\\|\_\{\\infty\}:=suph∈\[H\],x∈𝒳u∈𝒜h−1,a∈𝒜\|log⁡π⁡\(a\|x,u\)−log⁡π~​\(a\|x,u\)\|≤ε\.\\displaystyle:=\\sup\_\{\\begin\{subarray\}\{c\}h\\in\[H\],\\,x\\in\\mathcal\{X\}\\\\ u\\in\\mathcal\{A\}^\{h\-1\},\\,a\\in\\mathcal\{A\}\\end\{subarray\}\}\\big\|\\log\\pi\(a\|x,u\)\-\\log\\widetilde\{\\pi\}\(a\|x,u\)\\big\|\\leq\\varepsilon\.

When the function classΠ\\Pihas finite covering numberNε:=𝒩\(ε,logΠ,∥⋅∥∞\)<∞N\_\{\\varepsilon\}:=\\mathcal\{N\}\(\\varepsilon,\\log\\Pi,\\\|\\cdot\\\|\_\{\\infty\}\)<\\infty, we can choose anε\\varepsilon\-coverΠε⊆Π\\Pi\_\{\\varepsilon\}\\subseteq\\Piof cardinalityNεN\_\{\\varepsilon\}, and run Algorithm[3](https://arxiv.org/html/2609.38666#alg3)onΠε\\Pi\_\{\\varepsilon\}with a prior distributionp0p\_\{0\}\. Then, we have the following regret guarantee:

###### Theorem G\.4\.

LetΠ\\Pibe a class of autoregressive policies containingπF∗\\pi\_\{\\mathrm\{F\}\}^\{\*\}\. Fixε\>0\\varepsilon\>0such that

Nε:=𝒩\(ε,logΠ,∥⋅∥∞\)<∞\.\\displaystyle N\_\{\\varepsilon\}:=\\mathcal\{N\}\(\\varepsilon,\\log\\Pi,\\\|\\cdot\\\|\_\{\\infty\}\)<\\infty\.Let the algorithm be described as above andp0p\_\{0\}be the uniform prior\. Under off\-policy feedback, both with and without teacher logits, the expected regret can be bounded by

𝔼⁡\[Regret⁡\(T\)\]≤H​log⁡Nε\+T​H​ε\.\\displaystyle\\mathbb\{E\}\[\\operatorname\{Regret\}\(T\)\]\\leq H\\log N\_\{\\varepsilon\}\+TH\\varepsilon\.Moreover, for everyδ∈\(0,1\)\\delta\\in\(0,1\), with probability at least1−δ1\-\\delta,

Regret⁡\(T\)\\displaystyle\\operatorname\{Regret\}\(T\)≤2​H​\(log⁡Nε\+T​ε\)\+4​H​\(1\+log⁡Nε\+T​ε\+log⁡2δ\)​log​2δ\.\\displaystyle\\leq 2H\\bigl\(\\log N\_\{\\varepsilon\}\+T\\varepsilon\\bigr\)\+4H\\Big\(1\+\\log N\_\{\\varepsilon\}\+T\\varepsilon\+\\log\\frac\{2\}\{\\delta\}\\Big\)\\log\\frac\{2\}\{\\delta\}\.

###### Proof\.

SinceπF∗∈Π\\pi\_\{\\mathrm\{F\}\}^\{\*\}\\in\\Pi, the definition of theε\\varepsilon\-cover ensures that there existsπθ¯∈Πε\\pi\_\{\\bar\{\\theta\}\}\\in\\Pi\_\{\\varepsilon\}such that

‖log⁡πF∗−log⁡πθ¯‖∞≤ε\.\\displaystyle\\\|\\log\\pi\_\{\\mathrm\{F\}\}^\{\*\}\-\\log\\pi\_\{\\bar\{\\theta\}\}\\\|\_\{\\infty\}\\leq\\varepsilon\.\(G\.17\)In particular, for every prefix–token triple\(x,u,a\)\(x,u,a\),

log⁡πF∗​\(a\|x,u\)πθ¯​\(a\|x,u\)≤ε\.\\displaystyle\\log\\frac\{\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\}\{\\pi\_\{\\bar\{\\theta\}\}\(a\|x,u\)\}\\leq\\varepsilon\.Using the autoregressive factorization, for every complete responsey=\(a1,…,aH\)y=\(a\_\{1\},\\ldots,a\_\{H\}\), we obtain

log⁡πF∗​\(y\|x\)πθ¯​\(y\|x\)\\displaystyle\\log\\frac\{\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(y\|x\)\}\{\\pi\_\{\\bar\{\\theta\}\}\(y\|x\)\}=∑h=1Hlog⁡πF∗​\(ah\|x,y<h\)πθ¯​\(ah\|x,y<h\)\\displaystyle=\\sum\_\{h=1\}^\{H\}\\log\\frac\{\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\_\{h\}\|x,y\_\{<h\}\)\}\{\\pi\_\{\\bar\{\\theta\}\}\(a\_\{h\}\|x,y\_\{<h\}\)\}≤H​ε\.\\displaystyle\\leq H\\varepsilon\.Taking expectations overx∼ρ¯x\\sim\\bar\{\\rho\}andy∼πF∗\(⋅\|x\)y\\sim\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(\\cdot\|x\)therefore gives

𝔼x∼ρ¯KL\(πF∗\(⋅\|x\)∥πθ¯\(⋅\|x\)\)≤Hε\.\\displaystyle\\mathbb\{E\}\_\{x\\sim\\bar\{\\rho\}\}\\text\{KL\}\\big\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(\\cdot\|x\)\\,\\big\\\|\\,\\pi\_\{\\bar\{\\theta\}\}\(\\cdot\|x\)\\big\)\\leq H\\varepsilon\.\(G\.18\)
The uniform prior on the cover assignsp0​\(θ¯\)=1/Nεp\_\{0\}\(\\bar\{\\theta\}\)=1/N\_\{\\varepsilon\}\. Applying Theorem[G\.2](https://arxiv.org/html/2609.38666#A7.Thmtheorem2), we have

𝔼⁡\[Regret⁡\(T\)\]\\displaystyle\\mathbb\{E\}\[\\operatorname\{Regret\}\(T\)\]≤Hlog1p0​\(θ¯\)\+T𝔼x∼ρ¯KL\(πF∗\(⋅\|x\)∥πθ¯\(⋅\|x\)\)\\displaystyle\\leq H\\log\\frac\{1\}\{p\_\{0\}\(\\bar\{\\theta\}\)\}\+T\\mathbb\{E\}\_\{x\\sim\\bar\{\\rho\}\}\\text\{KL\}\\big\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(\\cdot\|x\)\\,\\big\\\|\\,\\pi\_\{\\bar\{\\theta\}\}\(\\cdot\|x\)\\big\)≤H​log⁡Nε\+T​H​ε,\\displaystyle\\leq H\\log N\_\{\\varepsilon\}\+TH\\varepsilon,where we use thatp0p\_\{0\}is uniform and \([G\.18](https://arxiv.org/html/2609.38666#A7.E18)\)\. This has proved the expected regret bound\.

For the high\-probability bound, letΘε\\Theta\_\{\\varepsilon\}index theNεN\_\{\\varepsilon\}policies inΠε\\Pi\_\{\\varepsilon\}, withp0​\(θ\)=1/Nεp\_\{0\}\(\\theta\)=1/N\_\{\\varepsilon\}for everyθ∈Θε\\theta\\in\\Theta\_\{\\varepsilon\}\. We use the same filtrationℱt\\mathcal\{F\}\_\{t\}as in the finite\-class analysis and define

Xt\\displaystyle X\_\{t\}:=ℓt​\(πF∗\)−ℓt​\(π^t\),\\displaystyle:=\\ell\_\{t\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\-\\ell\_\{t\}\(\\widehat\{\\pi\}\_\{t\}\),rt\\displaystyle r\_\{t\}:=−𝔼⁡\[Xt∣ℱt\]\.\\displaystyle:=\-\\mathbb\{E\}\[X\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]\.By \([G\.13](https://arxiv.org/html/2609.38666#A7.E13)\),Regret⁡\(T\)=∑t=1Trt\\operatorname\{Regret\}\(T\)=\\sum\_\{t=1\}^\{T\}r\_\{t\}\. For the policyπθ¯\\pi\_\{\\bar\{\\theta\}\}in \([G\.17](https://arxiv.org/html/2609.38666#A7.E17)\),

ℓt​\(πθ¯\)−ℓt​\(πF∗\)\\displaystyle\\ell\_\{t\}\(\\pi\_\{\\bar\{\\theta\}\}\)\-\\ell\_\{t\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)=∑h=1H∑\(x,u,a\)ct​\(x,u,a\)​log⁡πF∗​\(a\|x,u\)πθ¯​\(a\|x,u\)\\displaystyle=\\sum\_\{h=1\}^\{H\}\\sum\_\{\(x,u,a\)\}c\_\{t\}\(x,u,a\)\\log\\frac\{\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\}\{\\pi\_\{\\bar\{\\theta\}\}\(a\|x,u\)\}≤ε​∑h=1H∑\(x,u,a\)ct​\(x,u,a\)=H​ε\.\\displaystyle\\leq\\varepsilon\\sum\_\{h=1\}^\{H\}\\sum\_\{\(x,u,a\)\}c\_\{t\}\(x,u,a\)=H\\varepsilon\.\(G\.19\)Combining \([G\.19](https://arxiv.org/html/2609.38666#A7.E19)\) with \([G\.11](https://arxiv.org/html/2609.38666#A7.E11)\) gives

−∑t=1TXt\\displaystyle\-\\sum\_\{t=1\}^\{T\}X\_\{t\}=∑t=1T\[ℓt​\(π^t\)−ℓt​\(πF∗\)\]\\displaystyle=\\sum\_\{t=1\}^\{T\}\[\\ell\_\{t\}\(\\widehat\{\\pi\}\_\{t\}\)\-\\ell\_\{t\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\]=∑t=1T\[ℓt​\(π^t\)−ℓt​\(πθ¯\)\]\+∑t=1T\[ℓt​\(πθ¯\)−ℓt​\(πF∗\)\]\\displaystyle=\\sum\_\{t=1\}^\{T\}\[\\ell\_\{t\}\(\\widehat\{\\pi\}\_\{t\}\)\-\\ell\_\{t\}\(\\pi\_\{\\bar\{\\theta\}\}\)\]\+\\sum\_\{t=1\}^\{T\}\[\\ell\_\{t\}\(\\pi\_\{\\bar\{\\theta\}\}\)\-\\ell\_\{t\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\]≤H​log⁡Nε\+∑t=1T\[ℓt​\(πθ¯\)−ℓt​\(πF∗\)\]\\displaystyle\\leq H\\log N\_\{\\varepsilon\}\+\\sum\_\{t=1\}^\{T\}\[\\ell\_\{t\}\(\\pi\_\{\\bar\{\\theta\}\}\)\-\\ell\_\{t\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\]≤H​log⁡Nε\+T​H​ε\.\\displaystyle\\leq H\\log N\_\{\\varepsilon\}\+TH\\varepsilon\.\(G\.20\)DefineCtC\_\{t\}andWtW\_\{t\}the same as in the last section, that is,

Ct\\displaystyle C\_\{t\}:=∑θ∈Θεp0\(θ\)exp\(−1H∑s=1tℓs\(πθ\)\)=1Nε∑θ∈Θεexp\(−1H∑s=1tℓs\(πθ\)\),\\displaystyle:=\\sum\_\{\\theta\\in\\Theta\_\{\\varepsilon\}\}p\_\{0\}\(\\theta\)\\exp\\bigg\(\-\\frac\{1\}\{H\}\\sum\_\{s=1\}^\{t\}\\ell\_\{s\}\(\\pi\_\{\\theta\}\)\\bigg\)=\\frac\{1\}\{N\_\{\\varepsilon\}\}\\sum\_\{\\theta\\in\\Theta\_\{\\varepsilon\}\}\\exp\\bigg\(\-\\frac\{1\}\{H\}\\sum\_\{s=1\}^\{t\}\\ell\_\{s\}\(\\pi\_\{\\theta\}\)\\bigg\),Wt\\displaystyle W\_\{t\}:=Ct​exp⁡\(1H​∑s=1tℓs​\(πF∗\)\),\\displaystyle:=C\_\{t\}\\exp\\bigg\(\\frac\{1\}\{H\}\\sum\_\{s=1\}^\{t\}\\ell\_\{s\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\\bigg\),withC0=W0=1C\_\{0\}=W\_\{0\}=1\. With a similar argument to \([B\.20](https://arxiv.org/html/2609.38666#A2.E20)\), we have

𝔼⁡\[exp⁡\(ℓt​\(πF∗\)−ℓt​\(π^t\)H\)\|ℱt\]≤1\.\\displaystyle\\mathbb\{E\}\\bigg\[\\exp\\bigg\(\\frac\{\\ell\_\{t\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\-\\ell\_\{t\}\(\\widehat\{\\pi\}\_\{t\}\)\}\{H\}\\bigg\)\\,\\bigg\|\\,\\mathcal\{F\}\_\{t\}\\bigg\]\\leq 1\.\(G\.21\)Using the normalizer\-ratio identity \([G\.6](https://arxiv.org/html/2609.38666#A7.E6)\), we have

WtWt−1\\displaystyle\\frac\{W\_\{t\}\}\{W\_\{t\-1\}\}=CtCt−1​exp⁡\(ℓt​\(πF∗\)/H\)\\displaystyle=\\frac\{C\_\{t\}\}\{C\_\{t\-1\}\}\\exp\\bigl\(\\ell\_\{t\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)/H\\bigr\)=∑θ∈Θεpt−1​\(θ\)​exp⁡\(ℓt​\(πF∗\)−ℓt​\(πθ\)H\),\\displaystyle=\\sum\_\{\\theta\\in\\Theta\_\{\\varepsilon\}\}p\_\{t\-1\}\(\\theta\)\\exp\\bigg\(\\frac\{\\ell\_\{t\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\-\\ell\_\{t\}\(\\pi\_\{\\theta\}\)\}\{H\}\\bigg\),where the last equality holds due to the same argument as in \([G\.6](https://arxiv.org/html/2609.38666#A7.E6)\)\. We also have

pt​\(θ¯\)\\displaystyle p\_\{t\}\(\\bar\{\\theta\}\)=Nε−1exp\(−H−1∑s=1tℓs\(πθ¯\)\)Ct\\displaystyle=\\frac\{N\_\{\\varepsilon\}^\{\-1\}\\exp\\big\(\-H^\{\-1\}\\sum\_\{s=1\}^\{t\}\\ell\_\{s\}\(\\pi\_\{\\bar\{\\theta\}\}\)\\big\)\}\{C\_\{t\}\}=1Nε​Wtexp\(−1H∑s=1t\[ℓs\(πθ¯\)−ℓs\(πF∗\)\]\)\\displaystyle=\\frac\{1\}\{N\_\{\\varepsilon\}W\_\{t\}\}\\exp\\bigg\(\-\\frac\{1\}\{H\}\\sum\_\{s=1\}^\{t\}\[\\ell\_\{s\}\(\\pi\_\{\\bar\{\\theta\}\}\)\-\\ell\_\{s\}\(\\pi\_\{\\mathrm\{F\}\}^\{\*\}\)\]\\bigg\)≥e−t​εNε​Wt,\\displaystyle\\geq\\frac\{e^\{\-t\\varepsilon\}\}\{N\_\{\\varepsilon\}W\_\{t\}\},where the inequality uses \([G\.19](https://arxiv.org/html/2609.38666#A7.E19)\)\. Furthermore, \([G\.17](https://arxiv.org/html/2609.38666#A7.E17)\) impliesπθ¯​\(a\|x,u\)≥e−ε​πF∗​\(a\|x,u\)\\pi\_\{\\bar\{\\theta\}\}\(a\|x,u\)\\geq e^\{\-\\varepsilon\}\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\. Therefore, the token\-level mixture satisfies

π^t​\(a\|x,u\)πF∗​\(a\|x,u\)\\displaystyle\\frac\{\\widehat\{\\pi\}\_\{t\}\(a\|x,u\)\}\{\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\}≥pt−1​\(θ¯\)​πθ¯​\(a\|x,u\)πF∗​\(a\|x,u\)\\displaystyle\\geq p\_\{t\-1\}\(\\bar\{\\theta\}\)\\frac\{\\pi\_\{\\bar\{\\theta\}\}\(a\|x,u\)\}\{\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\}≥e−t​εNε​Wt−1\.\\displaystyle\\geq\\frac\{e^\{\-t\\varepsilon\}\}\{N\_\{\\varepsilon\}W\_\{t\-1\}\}\.Using the same concentration argument as in the last section, we can consider the high\-probability event

\{max0≤s<tWs≤2δ\}\.\\displaystyle\\bigg\\\{\\max\_\{0\\leq s<t\}W\_\{s\}\\leq\\frac\{2\}\{\\delta\}\\bigg\\\}\.On this event, we have

π^t​\(a\|x,u\)πF∗​\(a\|x,u\)≥δ​e−T​ε2​Nε=e−L,\\displaystyle\\frac\{\\widehat\{\\pi\}\_\{t\}\(a\|x,u\)\}\{\\pi\_\{\\mathrm\{F\}\}^\{\*\}\(a\|x,u\)\}\\geq\\frac\{\\delta e^\{\-T\\varepsilon\}\}\{2N\_\{\\varepsilon\}\}=e^\{\-L\},where we defineL:=log⁡\(2​Nε/δ\)\+T​εL:=\\log\(\{2N\_\{\\varepsilon\}\}/\{\\delta\}\)\+T\\varepsilon\. Repeating the stopped\-process argument from the proof of Theorem[G\.2](https://arxiv.org/html/2609.38666#A7.Thmtheorem2), we obtain, with probability at least1−δ1\-\\delta,

Regret⁡\(T\)\\displaystyle\\operatorname\{Regret\}\(T\)≤2​H​\(log⁡Nε\+T​ε\)\+4​H​\(1\+log⁡Nε\+T​ε\+log⁡2δ\)​log​2δ\.\\displaystyle\\leq 2H\\bigl\(\\log N\_\{\\varepsilon\}\+T\\varepsilon\\bigr\)\+4H\\bigg\(1\+\\log N\_\{\\varepsilon\}\+T\\varepsilon\+\\log\\frac\{2\}\{\\delta\}\\bigg\)\\log\\frac\{2\}\{\\delta\}\.This proves Theorem[G\.4](https://arxiv.org/html/2609.38666#A7.Thmtheorem4)\. ∎

###### Example G\.5\.

SupposeΘ⊆ℝd\\Theta\\subseteq\\mathbb\{R\}^\{d\}is a nonempty compact set contained in a Euclidean ball of radiusRR, whered≥1d\\geq 1\. Assume thatπθ∗=πF∗\\pi\_\{\\theta^\{\*\}\}=\\pi\_\{\\mathrm\{F\}\}^\{\*\}for someθ∗∈Θ\\theta^\{\*\}\\in\\Theta\. Moreover, suppose that, for someL0\>0L\_\{0\}\>0,

\|log⁡πθ​\(a\|x,u\)−log⁡πθ′​\(a\|x,u\)\|≤L0​‖θ−θ′‖2\\displaystyle\\big\|\\log\\pi\_\{\\theta\}\(a\|x,u\)\-\\log\\pi\_\{\\theta^\{\\prime\}\}\(a\|x,u\)\\big\|\\leq L\_\{0\}\\\|\\theta\-\\theta^\{\\prime\}\\\|\_\{2\}for everyθ,θ′∈Θ\\theta,\\theta^\{\\prime\}\\in\\Theta,h∈\[H\]h\\in\[H\], and\(x,u,a\)∈𝒳×𝒜h−1×𝒜\(x,u,a\)\\in\\mathcal\{X\}\\times\\mathcal\{A\}^\{h\-1\}\\times\\mathcal\{A\}\. Then, a standard Euclidean covering argument gives

𝒩\(ε,logΠ,∥⋅∥∞\)≤\(1\+2​R​L0ε\)d\.\\displaystyle\\mathcal\{N\}\(\\varepsilon,\\log\\Pi,\\\|\\cdot\\\|\_\{\\infty\}\)\\leq\\left\(1\+\\frac\{2RL\_\{0\}\}\{\\varepsilon\}\\right\)^\{d\}\.For everyT≥2T\\geq 2, choose such a cover withε=1/T\\varepsilon=1/T\. By Theorem[G\.4](https://arxiv.org/html/2609.38666#A7.Thmtheorem4), the expected regret satisfies

𝔼⁡\[Regret⁡\(T\)\]≤H⁡\[d​log⁡\(1\+2​R​L0​T\)\+1\]\.\\displaystyle\\mathbb\{E\}\\big\[\\operatorname\{Regret\}\(T\)\\big\]\\leq H\\left\[d\\log\\bigl\(1\+2RL\_\{0\}T\\bigr\)\+1\\right\]\.Moreover, with probability at least1−δ1\-\\delta, we have

Regret⁡\(T\)≤O~​\(d​H​log⁡T\)\.\\displaystyle\\operatorname\{Regret\}\(T\)\\leq\\widetilde\{O\}\(dH\\log T\)\.

## Appendix HReverse KL with Function Approximation

We now study on\-policy distillation with function approximation\. Fix anyh∈\[H\]h\\in\[H\]\. For any prefixx∈𝒳x\\in\\mathcal\{X\},u∈𝒜h−1u\\in\\mathcal\{A\}^\{h\-1\},a∈𝒜a\\in\\mathcal\{A\}, following the notation in Theorem[5\.1](https://arxiv.org/html/2609.38666#S5.Thmtheorem1), we define the target as

f∗​\(h,x,u,a\):=log⁡gh​\(a\|x,u\)=∑i∈ℐwi​\(x\)​log⁡pi​\(a\|x,u\)\.\\displaystyle f^\{\*\}\(h,x,u,a\):=\\log g\_\{h\}\(a\|x,u\)=\\sum\_\{i\\in\\mathcal\{I\}\}w\_\{i\}\(x\)\\log p\_\{i\}\(a\|x,u\)\.Specifically, we define

𝒵:=⋃h=1H\(\{h\}×𝒳×𝒜h−1×𝒜\),\\displaystyle\\mathcal\{Z\}:=\\bigcup\_\{h=1\}^\{H\}\\bigl\(\\\{h\\\}\\times\\mathcal\{X\}\\times\\mathcal\{A\}^\{h\-1\}\\times\\mathcal\{A\}\\bigr\),and writezt,j,h:=\(h,xt,j,yt,j,<h,at,j,h\)z\_\{t,j,h\}:=\(h,x\_\{t,j\},y\_\{t,j,<h\},a\_\{t,j,h\}\)for the query associated with thehh\-th token of rolloutjjin roundtt\.

Throughout this section, we retain Assumption[5\.3](https://arxiv.org/html/2609.38666#S5.Thmtheorem3), with reference policyπref\\pi\_\{\\text\{ref\}\}and boundB\>0B\>0\. Moreover, letℱ\\mathcal\{F\}be a finite class of functionsf:𝒵→ℝf:\\mathcal\{Z\}\\to\\mathbb\{R\}\. We further impose the following realizability assumption\.

###### Assumption H\.1\.

f∗∈ℱf^\{\*\}\\in\\mathcal\{F\}\.

To measure how past observations constrain predictions at a new query, we introduce a new version of generalized Eluder dimension adapted to autoregressive trajectory batches\. The underlying uncertainty compares the disagreement between two candidate score functions at the current query with their cumulative squared disagreement on previously observed data\. Since the estimator is updated only after each round, this comparison uses observations from completed rounds, without incorporating feedback from the current batch\.

###### Definition H\.2\.

Fixλ\>0\\lambda\>0\. For a sequence of trajectory batches, write

Zt:=\(zt,j,h\)j∈\[m\],h∈\[H\],Z<t:=\(Z1,…,Zt−1\)\.\\displaystyle Z\_\{t\}:=\(z\_\{t,j,h\}\)\_\{j\\in\[m\],\\,h\\in\[H\]\},\\qquad Z\_\{<t\}:=\(Z\_\{1\},\\ldots,Z\_\{t\-1\}\)\.For any queryz∈𝒵z\\in\\mathcal\{Z\}, define

Dℱ2​\(z,Z<t\):=supf1,f2∈ℱ\(f1​\(z\)−f2​\(z\)\)2λ\+∑s=1t−11m​H​∑j=1m∑h=1H\(f1​\(zs,j,h\)−f2​\(zs,j,h\)\)2\.\\displaystyle D\_\{\\mathcal\{F\}\}^\{2\}\(z;Z\_\{<t\}\):=\\sup\_\{f\_\{1\},f\_\{2\}\\in\\mathcal\{F\}\}\\frac\{\\bigl\(f\_\{1\}\(z\)\-f\_\{2\}\(z\)\\bigr\)^\{2\}\}\{\\lambda\+\\displaystyle\\sum\_\{s=1\}^\{t\-1\}\\frac\{1\}\{mH\}\\sum\_\{j=1\}^\{m\}\\sum\_\{h=1\}^\{H\}\\bigl\(f\_\{1\}\(z\_\{s,j,h\}\)\-f\_\{2\}\(z\_\{s,j,h\}\)\\bigr\)^\{2\}\}\.\(H\.1\)The batched generalized Eluder dimension is

dimT\(ℱ;λ\):=supZ1:T∑t=1T1m​H∑j=1m∑h=1Hmin\{1,Dℱ2\(zt,j,h;Z<t\)\},\\displaystyle\\dim\_\{T\}\(\\mathcal\{F\};\\lambda\):=\\sup\_\{Z\_\{1:T\}\}\\sum\_\{t=1\}^\{T\}\\frac\{1\}\{mH\}\\sum\_\{j=1\}^\{m\}\\sum\_\{h=1\}^\{H\}\\min\\left\\\{1,\\,D\_\{\\mathcal\{F\}\}^\{2\}\(z\_\{t,j,h\};Z\_\{<t\}\)\\right\\\},\(H\.2\)where the supremum is over all sequences ofTTtrajectory batches, each containingmmrollouts of horizonHH\.

### H\.1Algorithm Design

Recall that, for every visited queryzt,j,h=\(h,x,u,a\)z\_\{t,j,h\}=\(h,x,u,a\),

𝔼\[lt,j,h\|ℱt,zt,j,h=\(h,x,u,a\)\]\\displaystyle\\mathbb\{E\}\\big\[l\_\{t,j,h\}\\,\\big\|\\,\\mathcal\{F\}\_\{t\},\\,z\_\{t,j,h\}=\(h,x,u,a\)\\big\]=∑i∈ℐwi​\(x\)​log⁡pi​\(a\|x,u\)\\displaystyle=\\sum\_\{i\\in\\mathcal\{I\}\}w\_\{i\}\(x\)\\log p\_\{i\}\(a\|x,u\)=log⁡gh​\(a\|x,u\)=f∗​\(h,x,u,a\),\\displaystyle=\\log g\_\{h\}\(a\|x,u\)=f^\{\*\}\(h,x,u,a\),\(H\.3\)whereℱt\\mathcal\{F\}\_\{t\}denotes the history before roundtt\. This identity motivates estimatingf∗f^\{\*\}by least squares\. For eachf∈ℱf\\in\\mathcal\{F\}, define

ℓt​\(f\)\\displaystyle\\ell\_\{t\}\(f\):=1m​H​∑j=1m∑h=1H\(f⁡\(zt,j,h\)−lt,j,h\)2\.\\displaystyle:=\\frac\{1\}\{mH\}\\sum\_\{j=1\}^\{m\}\\sum\_\{h=1\}^\{H\}\\bigl\(f\(z\_\{t,j,h\}\)\-l\_\{t,j,h\}\\bigr\)^\{2\}\.\(H\.4\)Moreover, we define

Lt​\(f\)\\displaystyle L\_\{t\}\(f\):=∑s=1tℓs​\(f\),L0​\(f\):=0\.\\displaystyle:=\\sum\_\{s=1\}^\{t\}\\ell\_\{s\}\(f\),\\qquad L\_\{0\}\(f\):=0\.At the beginning of roundtt, the learner computes

f¯t−1∈argminf∈ℱLt−1​\(f\)\.\\displaystyle\\bar\{f\}\_\{t\-1\}\\in\\mathop\{\\mathrm\{argmin\}\}\_\{f\\in\\mathcal\{F\}\}L\_\{t\-1\}\(f\)\.\(H\.5\)Using the principle of optimism for online reinforcement learning, we therefore define a bonus using the batched generalized Eluder dimension, i\.e\., for anyz∈𝒵z\\in\\mathcal\{Z\}, the bonus function is defined as

bt−1​\(z\):=\(4​β\+λ\)​Dℱ2​\(z,Z<t\),\\displaystyle b\_\{t\-1\}\(z\):=\\sqrt\{\(4\\beta\+\\lambda\)D\_\{\\mathcal\{F\}\}^\{2\}\(z;Z\_\{<t\}\)\},\(H\.6\)whereλ\>0\\lambda\>0andβ\>0\\beta\>0are parameters to be specified later\. Given the reference policyπref\\pi\_\{\\text\{ref\}\}and the boundBBfrom Assumption[5\.3](https://arxiv.org/html/2609.38666#S5.Thmtheorem3), as it implies

\|f∗​\(h,x,u,a\)−log⁡πref​\(a\|x,u\)\|≤B,\\displaystyle\\big\|f^\{\*\}\(h,x,u,a\)\-\\log\\pi\_\{\\text\{ref\}\}\(a\|x,u\)\\big\|\\leq B,we define the optimistic score by

f^t−1​\(h,x,u,a\):=min⁡\{f¯t−1​\(h,x,u,a\)\+bt−1​\(h,x,u,a\),log⁡πref​\(a\|x,u\)\+B\}\.\\displaystyle\\widehat\{f\}\_\{t\-1\}\(h,x,u,a\):=\\min\\Big\\\{\\bar\{f\}\_\{t\-1\}\(h,x,u,a\)\+b\_\{t\-1\}\(h,x,u,a\),\\,\\log\\pi\_\{\\text\{ref\}\}\(a\|x,u\)\+B\\Big\\\}\.\(H\.7\)The following lemma establishes pointwise optimism and bounds the error of the optimistic score\.

###### Lemma H\.3\.

Under Assumptions[5\.3](https://arxiv.org/html/2609.38666#S5.Thmtheorem3)and[H\.1](https://arxiv.org/html/2609.38666#A8.Thmtheorem1), fixδ∈\(0,1\)\\delta\\in\(0,1\)and choose

β=32​B2​log⁡\|ℱ\|δ,λ\>0\.\\displaystyle\\beta=32B^\{2\}\\log\\frac\{\|\\mathcal\{F\}\|\}\{\\delta\},\\qquad\\lambda\>0\.Then, with probability at least1−δ1\-\\delta, simultaneously for everyt∈\[T\]t\\in\[T\]andz∈𝒵z\\in\\mathcal\{Z\},

\|f¯t−1​\(z\)−f∗​\(z\)\|\\displaystyle\\big\|\\bar\{f\}\_\{t\-1\}\(z\)\-f^\{\*\}\(z\)\\big\|≤bt−1​\(z\)\.\\displaystyle\\leq b\_\{t\-1\}\(z\)\.Furthermore, the following inequality holds

0≤f^t−1​\(z\)−f∗​\(z\)\\displaystyle 0\\leq\\widehat\{f\}\_\{t\-1\}\(z\)\-f^\{\*\}\(z\)≤min⁡\{2​bt−1​\(z\),2​B\}\.\\displaystyle\\leq\\min\\bigl\\\{2b\_\{t\-1\}\(z\),2B\\bigr\\\}\.

To construct the output policy, we apply the backward propagation as in the tabular setting to the estimated scoresf^t\\widehat\{f\}\_\{t\}\. More specifically, setV^t−1,H\+1​\(x,y\)=1\\widehat\{V\}\_\{t\-1,H\+1\}\(x,y\)=1and recursively define

V^t−1,h​\(x,u\)\\displaystyle\\widehat\{V\}\_\{t\-1,h\}\(x,u\):=∑a∈𝒜exp⁡\(f^t−1​\(h,x,u,a\)\)​V^t−1,h\+1​\(x,\(u,a\)\),\\displaystyle:=\\sum\_\{a\\in\\mathcal\{A\}\}\\exp\\bigl\(\\widehat\{f\}\_\{t\-1\}\(h,x,u,a\)\\bigr\)\\widehat\{V\}\_\{t\-1,h\+1\}\(x,\(u,a\)\),\(H\.8\)forh=H,…,1h=H,\\ldots,1\. We then output

π^t​\(a\|x,u\):=exp⁡\(f^t−1​\(h,x,u,a\)\)​V^t−1,h\+1​\(x,\(u,a\)\)V^t−1,h​\(x,u\)\.\\displaystyle\\widehat\{\\pi\}\_\{t\}\(a\|x,u\):=\\frac\{\\exp\\bigl\(\\widehat\{f\}\_\{t\-1\}\(h,x,u,a\)\\bigr\)\\widehat\{V\}\_\{t\-1,h\+1\}\(x,\(u,a\)\)\}\{\\widehat\{V\}\_\{t\-1,h\}\(x,u\)\}\.\(H\.9\)This procedure is described in Algorithm[4](https://arxiv.org/html/2609.38666#alg4)\.

Algorithm 4Optimistic Reverse KL with Function Approximation1:Input:Function class

ℱ\\mathcal\{F\}, reference policy

πref\\pi\_\{\\text\{ref\}\}, bound

BB, horizon

HH, batch size

mm, number of rounds

TT, and parameters

λ,β\>0\\lambda,\\beta\>0\.

2:Initialize

L0​\(f\)=0L\_\{0\}\(f\)=0for every

f∈ℱf\\in\\mathcal\{F\}\.

3:for

t=1,…,Tt=1,\\ldots,Tdo

4:Compute the regression estimate

f¯t−1\\bar\{f\}\_\{t\-1\}using \([H\.5](https://arxiv.org/html/2609.38666#A8.E5)\)\.

5:Compute

bt−1b\_\{t\-1\}and

f^t−1\\widehat\{f\}\_\{t\-1\}using \([H\.6](https://arxiv.org/html/2609.38666#A8.E6)\) and \([H\.7](https://arxiv.org/html/2609.38666#A8.E7)\)\.

6:Set

V^t−1,H\+1​\(x,y\)=1\\widehat\{V\}\_\{t\-1,H\+1\}\(x,y\)=1and compute the backward recursion \([H\.8](https://arxiv.org/html/2609.38666#A8.E8)\)\.

7:Construct

π^t\\widehat\{\\pi\}\_\{t\}using \([H\.9](https://arxiv.org/html/2609.38666#A8.E9)\)\.

8:Receive contexts

\{xt,j\}j=1m\\\{x\_\{t,j\}\\\}\_\{j=1\}^\{m\}according to the on\-policy protocol and generate

yt,j∼π^t\(⋅\|xt,j\)y\_\{t,j\}\\sim\\widehat\{\\pi\}\_\{t\}\(\\cdot\|x\_\{t,j\}\)\.

9:Observe teacher feedback

\{lt,j,h\}j∈\[m\],h∈\[H\]\\\{l\_\{t,j,h\}\\\}\_\{j\\in\[m\],\\,h\\in\[H\]\}\.

10:Compute

ℓt​\(f\)\\ell\_\{t\}\(f\)using \([H\.4](https://arxiv.org/html/2609.38666#A8.E4)\) and update

Lt​\(f\)=Lt−1​\(f\)\+ℓt​\(f\)L\_\{t\}\(f\)=L\_\{t\-1\}\(f\)\+\\ell\_\{t\}\(f\)for every

f∈ℱf\\in\\mathcal\{F\}\.

11:endfor

12:Output:

\{π^t\}t=1T\\\{\\widehat\{\\pi\}\_\{t\}\\\}\_\{t=1\}^\{T\}\.

### H\.2Theoretical Guarantee

Given Lemma[H\.3](https://arxiv.org/html/2609.38666#A8.Thmtheorem3), we have the following guarantee on the regret of Algorithm[4](https://arxiv.org/html/2609.38666#alg4)\.

###### Theorem H\.4\.

Under Assumptions[5\.3](https://arxiv.org/html/2609.38666#S5.Thmtheorem3)and[H\.1](https://arxiv.org/html/2609.38666#A8.Thmtheorem1), fixδ∈\(0,1\)\\delta\\in\(0,1\)and chooseλ\>0\\lambda\>0and

β≥32​B2​log⁡2​\|ℱ\|δ\\displaystyle\\beta\\geq 32B^\{2\}\\log\\frac\{2\|\\mathcal\{F\}\|\}\{\\delta\}Then, with probability at least1−δ1\-\\delta, Algorithm[4](https://arxiv.org/html/2609.38666#alg4)satisfies

Regret⁡\(T\)\\displaystyle\\operatorname\{Regret\}\(T\):=∑t=1T𝔼x∼ρ¯KL\(π^t\(⋅\|x\)∥πR∗\(⋅\|x\)\)\\displaystyle:=\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\_\{x\\sim\\bar\{\\rho\}\}\\text\{KL\}\\bigl\(\\widehat\{\\pi\}\_\{t\}\(\\cdot\|x\)\\,\\big\\\|\\,\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(\\cdot\|x\)\\bigr\)≤4​H2​\(4​β\+λ\)​dimT\(ℱ;λ\)\+16​H2​B2​log⁡4δ\.\\displaystyle\\leq 4H^\{2\}\(4\\beta\+\\lambda\)\\dim\_\{T\}\(\\mathcal\{F\};\\lambda\)\+16H^\{2\}B^\{2\}\\log\\frac\{4\}\{\\delta\}\.In particular, if we chooseλ=B2\\lambda=B^\{2\}andβ=32​B2​log⁡\(2​\|ℱ\|/δ\)\\beta=32B^\{2\}\\log\(2\|\\mathcal\{F\}\|/\\delta\), we have

Regret⁡\(T\)≤O⁡\(B2​H2​dimT\(ℱ;B2\)​log⁡2​\|ℱ\|δ\)\.\\displaystyle\\operatorname\{Regret\}\(T\)\\leq O\\bigg\(B^\{2\}H^\{2\}\\dim\_\{T\}\(\\mathcal\{F\};B^\{2\}\)\\log\\frac\{2\|\\mathcal\{F\}\|\}\{\\delta\}\\bigg\)\.

###### Proof\.

Letℱt\\mathcal\{F\}\_\{t\}denote the history before roundtt\. In particular,f¯t−1\\bar\{f\}\_\{t\-1\},bt−1b\_\{t\-1\}, andπ^t\\widehat\{\\pi\}\_\{t\}areℱt\\mathcal\{F\}\_\{t\}\-measurable\. Applying Lemma[H\.3](https://arxiv.org/html/2609.38666#A8.Thmtheorem3), we have with probability at least1−δ/21\-\\delta/2, the following inequality holds simultaneously for everyt∈\[T\]t\\in\[T\]andz∈𝒵z\\in\\mathcal\{Z\}

0≤f^t−1​\(z\)−f∗​\(z\)≤min⁡\{2​bt−1​\(z\),2​B\},\\displaystyle 0\\leq\\widehat\{f\}\_\{t\-1\}\(z\)\-f^\{\*\}\(z\)\\leq\\min\\\{2b\_\{t\-1\}\(z\),2B\\\},\(H\.10\)whenβ≥32​B2​log⁡\(2​\|ℱ\|/δ\)\\beta\\geq 32B^\{2\}\\log\(\{2\|\\mathcal\{F\}\|\}/\{\\delta\}\)\. Letℰ\\mathcal\{E\}be this high\-probability event\. For everyttandzz, define

ψt​\(z\)\\displaystyle\\psi\_\{t\}\(z\):=min⁡\{4​bt−1​\(z\)2,4​B2\}\.\\displaystyle:=\\min\\\{4b\_\{t\-1\}\(z\)^\{2\},4B^\{2\}\\\}\.\(H\.11\)Then we have0≤ψt​\(z\)≤4​B20\\leq\\psi\_\{t\}\(z\)\\leq 4B^\{2\}\. Onℰ\\mathcal\{E\}, \([H\.10](https://arxiv.org/html/2609.38666#A8.E10)\) further gives

\[f^t−1​\(z\)−f∗​\(z\)\]2≤ψt​\(z\)\.\\displaystyle\\big\[\\widehat\{f\}\_\{t\-1\}\(z\)\-f^\{\*\}\(z\)\\big\]^\{2\}\\leq\\psi\_\{t\}\(z\)\.\(H\.12\)Fix a contextxxand a responseyy\. Write

vt​\(x,y\):=∑h=1H\[f^t−1​\(h,x,y<h,ah\)−f∗​\(h,x,y<h,ah\)\]\.\\displaystyle v\_\{t\}\(x,y\):=\\sum\_\{h=1\}^\{H\}\\Big\[\\widehat\{f\}\_\{t\-1\}\(h,x,y\_\{<h\},a\_\{h\}\)\-f^\{\*\}\(h,x,y\_\{<h\},a\_\{h\}\)\\Big\]\.Onℰ\\mathcal\{E\}, every summand is nonnegative, sovt​\(x,y\)≥0v\_\{t\}\(x,y\)\\geq 0\. Recall that fory=\(a1,…,aH\)y=\(a\_\{1\},\\ldots,a\_\{H\}\), the backward propagation gives

π^t​\(y\|x\)\\displaystyle\\widehat\{\\pi\}\_\{t\}\(y\|x\)=∏h=1Hef^t−1​\(h,x,y<h,ah\)​V^t−1,h\+1​\(x,y≤h\)V^t−1,h​\(x,y<h\)\\displaystyle=\\prod\_\{h=1\}^\{H\}\\frac\{e^\{\\widehat\{f\}\_\{t\-1\}\(h,x,y\_\{<h\},a\_\{h\}\)\}\\widehat\{V\}\_\{t\-1,h\+1\}\(x,y\_\{\\leq h\}\)\}\{\\widehat\{V\}\_\{t\-1,h\}\(x,y\_\{<h\}\)\}=exp⁡\(∑h=1Hf^t−1​\(h,x,y<h,ah\)\)V^t−1,1​\(x\),\\displaystyle=\\frac\{\\exp\\big\(\\sum\_\{h=1\}^\{H\}\\widehat\{f\}\_\{t\-1\}\(h,x,y\_\{<h\},a\_\{h\}\)\\big\)\}\{\\widehat\{V\}\_\{t\-1,1\}\(x\)\},where we usedV^t−1,H\+1​\(x,y\)=1\\widehat\{V\}\_\{t\-1,H\+1\}\(x,y\)=1\. Similarly, the reverse target satisfies

πR∗​\(y\|x\)=exp⁡\(∑h=1Hf∗​\(h,x,y<h,ah\)\)V1​\(x\),\\displaystyle\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(y\|x\)=\\frac\{\\exp\\big\(\\sum\_\{h=1\}^\{H\}f^\{\*\}\(h,x,y\_\{<h\},a\_\{h\}\)\\big\)\}\{V\_\{1\}\(x\)\},whereV1​\(x\)=∑yexp⁡\(∑h=1Hf∗​\(h,x,y<h,ah\)\)V\_\{1\}\(x\)=\\sum\_\{y\}\\exp\\big\(\\sum\_\{h=1\}^\{H\}f^\{\*\}\(h,x,y\_\{<h\},a\_\{h\}\)\\big\)\. Taking the logarithm of their ratio therefore yields

log⁡π^t​\(y\|x\)πR∗​\(y\|x\)\\displaystyle\\log\\frac\{\\widehat\{\\pi\}\_\{t\}\(y\|x\)\}\{\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(y\|x\)\}=∑h=1H\[f^t−1​\(h,x,y<h,ah\)−f∗​\(h,x,y<h,ah\)\]\\displaystyle=\\sum\_\{h=1\}^\{H\}\\big\[\\widehat\{f\}\_\{t\-1\}\(h,x,y\_\{<h\},a\_\{h\}\)\-f^\{\*\}\(h,x,y\_\{<h\},a\_\{h\}\)\\big\]\+log⁡V1​\(x\)−log⁡V^t−1,1​\(x\)\\displaystyle\\qquad\+\\log V\_\{1\}\(x\)\-\\log\\widehat\{V\}\_\{t\-1,1\}\(x\)=vt​\(x,y\)\+log⁡V1​\(x\)V^t−1,1​\(x\)\.\\displaystyle=v\_\{t\}\(x,y\)\+\\log\\frac\{V\_\{1\}\(x\)\}\{\\widehat\{V\}\_\{t\-1,1\}\(x\)\}\.\(H\.13\)To express the ratio of normalizing constants as an expectation, expand the expectation over complete responses:

𝔼y∼π^t\(⋅\|x\)\[e−vt​\(x,y\)\]\\displaystyle\\mathbb\{E\}\_\{y\\sim\\widehat\{\\pi\}\_\{t\}\(\\cdot\|x\)\}\\big\[e^\{\-v\_\{t\}\(x,y\)\}\\big\]=∑y∈𝒴π^t​\(y\|x\)​e−vt​\(x,y\)\\displaystyle=\\sum\_\{y\\in\\mathcal\{Y\}\}\\widehat\{\\pi\}\_\{t\}\(y\|x\)e^\{\-v\_\{t\}\(x,y\)\}=∑y∈𝒴exp⁡\(∑h=1Hf^t−1​\(h,x,y<h,ah\)\)V^t−1,1​\(x\)​exp⁡\(∑h=1H\[f∗​\(h,x,y<h,ah\)−f^t−1​\(h,x,y<h,ah\)\]\)\\displaystyle=\\sum\_\{y\\in\\mathcal\{Y\}\}\\frac\{\\exp\\big\(\\sum\_\{h=1\}^\{H\}\\widehat\{f\}\_\{t\-1\}\(h,x,y\_\{<h\},a\_\{h\}\)\\big\)\}\{\\widehat\{V\}\_\{t\-1,1\}\(x\)\}\\exp\\bigg\(\\sum\_\{h=1\}^\{H\}\\big\[f^\{\*\}\(h,x,y\_\{<h\},a\_\{h\}\)\-\\widehat\{f\}\_\{t\-1\}\(h,x,y\_\{<h\},a\_\{h\}\)\\big\]\\bigg\)=1V^t−1,1​\(x\)​∑y∈𝒴exp⁡\(∑h=1Hf∗​\(h,x,y<h,ah\)\)\\displaystyle=\\frac\{1\}\{\\widehat\{V\}\_\{t\-1,1\}\(x\)\}\\sum\_\{y\\in\\mathcal\{Y\}\}\\exp\\bigg\(\\sum\_\{h=1\}^\{H\}f^\{\*\}\(h,x,y\_\{<h\},a\_\{h\}\)\\bigg\)=V1​\(x\)V^t−1,1​\(x\)\.\\displaystyle=\\frac\{V\_\{1\}\(x\)\}\{\\widehat\{V\}\_\{t\-1,1\}\(x\)\}\.\(H\.14\)The third equality follows by cancellation of thef^t−1\\widehat\{f\}\_\{t\-1\}terms in the exponent\. The last equality uses the definition ofV1​\(x\)V\_\{1\}\(x\)\.

Taking expectation overπ^t\(⋅\|x\)\\widehat\{\\pi\}\_\{t\}\(\\cdot\|x\)in \([H\.13](https://arxiv.org/html/2609.38666#A8.E13)\) and considering \([H\.14](https://arxiv.org/html/2609.38666#A8.E14)\), we have

KL\(π^t\(⋅\|x\)∥πR∗\(⋅\|x\)\)\\displaystyle\\text\{KL\}\\bigl\(\\widehat\{\\pi\}\_\{t\}\(\\cdot\|x\)\\,\\big\\\|\\,\\pi\_\{\\mathrm\{R\}\}^\{\*\}\(\\cdot\|x\)\\bigr\)=𝔼π^t​\[vt​\(x,y\)\]\+log⁡𝔼π^t​\[e−vt​\(x,y\)\]\\displaystyle=\\mathbb\{E\}\_\{\\widehat\{\\pi\}\_\{t\}\}\[v\_\{t\}\(x,y\)\]\+\\log\\mathbb\{E\}\_\{\\widehat\{\\pi\}\_\{t\}\}\[e^\{\-v\_\{t\}\(x,y\)\}\]≤𝔼π^t​\[vt​\(x,y\)\]\+𝔼π^t​\[e−vt​\(x,y\)\]−1\\displaystyle\\leq\\mathbb\{E\}\_\{\\widehat\{\\pi\}\_\{t\}\}\[v\_\{t\}\(x,y\)\]\+\\mathbb\{E\}\_\{\\widehat\{\\pi\}\_\{t\}\}\[e^\{\-v\_\{t\}\(x,y\)\}\]\-1≤12​𝔼π^t​\[vt​\(x,y\)2\]\.\\displaystyle\\leq\\frac\{1\}\{2\}\\mathbb\{E\}\_\{\\widehat\{\\pi\}\_\{t\}\}\[v\_\{t\}\(x,y\)^\{2\}\]\.\(H\.15\)Here the first inequality useslog⁡u≤u−1\\log u\\leq u\-1foru\>0u\>0, and the second usese−v≤1−v\+v2/2e^\{\-v\}\\leq 1\-v\+v^\{2\}/2forv≥0v\\geq 0\. All expectations in this display are conditional onxx\.

By the Cauchy–Schwarz inequality and \([H\.12](https://arxiv.org/html/2609.38666#A8.E12)\),

vt​\(x,y\)2\\displaystyle v\_\{t\}\(x,y\)^\{2\}≤H​∑h=1H\[f^t−1​\(h,x,y<h,ah\)−f∗​\(h,x,y<h,ah\)\]2\\displaystyle\\leq H\\sum\_\{h=1\}^\{H\}\\big\[\\widehat\{f\}\_\{t\-1\}\(h,x,y\_\{<h\},a\_\{h\}\)\-f^\{\*\}\(h,x,y\_\{<h\},a\_\{h\}\)\\big\]^\{2\}≤H​∑h=1Hψt​\(h,x,y<h,ah\)\.\\displaystyle\\leq H\\sum\_\{h=1\}^\{H\}\\psi\_\{t\}\(h,x,y\_\{<h\},a\_\{h\}\)\.Consequently, onℰ\\mathcal\{E\},

Regret⁡\(T\)\\displaystyle\\operatorname\{Regret\}\(T\)≤H2∑t=1T𝔼x∼ρ¯,y∼π^t\(⋅\|x\)\[∑h=1Hψt\(h,x,y<h,ah\)\]\.\\displaystyle\\leq\\frac\{H\}\{2\}\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\_\{x\\sim\\bar\{\\rho\},\\,y\\sim\\widehat\{\\pi\}\_\{t\}\(\\cdot\|x\)\}\\bigg\[\\sum\_\{h=1\}^\{H\}\\psi\_\{t\}\(h,x,y\_\{<h\},a\_\{h\}\)\\bigg\]\.\(H\.16\)Define

Xt:=1m​H​∑j=1m∑h=1Hψt​\(zt,j,h\)\.\\displaystyle X\_\{t\}:=\\frac\{1\}\{mH\}\\sum\_\{j=1\}^\{m\}\\sum\_\{h=1\}^\{H\}\\psi\_\{t\}\(z\_\{t,j,h\}\)\.ThenXtX\_\{t\}isℱt\+1\\mathcal\{F\}\_\{t\+1\}\-measurable and satisfies0≤Xt≤4​B20\\leq X\_\{t\}\\leq 4B^\{2\}almost surely\. Under the on\-policy protocol, each rollout has conditional marginal distribution

ℙ⁡\(xt,j=x,yt,j=y∣ℱt\)\\displaystyle\\mathbb\{P\}\\bigl\(x\_\{t,j\}=x,\\,y\_\{t,j\}=y\\mid\\mathcal\{F\}\_\{t\}\\bigr\)=1I​∑i∈ℐρi​\(x\)​π^t​\(y\|x\)\\displaystyle=\\frac\{1\}\{I\}\\sum\_\{i\\in\\mathcal\{I\}\}\\rho\_\{i\}\(x\)\\widehat\{\\pi\}\_\{t\}\(y\|x\)=ρ¯​\(x\)​π^t​\(y\|x\)\.\\displaystyle=\\bar\{\\rho\}\(x\)\\widehat\{\\pi\}\_\{t\}\(y\|x\)\.Sinceψt\\psi\_\{t\}is determined before roundtt, this implies

𝔼\[Xt∣ℱt\]=1H𝔼x∼ρ¯,y∼π^t\(⋅\|x\)\[∑h=1Hψt\(h,x,y<h,ah\)\]\.\\displaystyle\\mathbb\{E\}\[X\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]=\\frac\{1\}\{H\}\\mathbb\{E\}\_\{x\\sim\\bar\{\\rho\},\\,y\\sim\\widehat\{\\pi\}\_\{t\}\(\\cdot\|x\)\}\\bigg\[\\sum\_\{h=1\}^\{H\}\\psi\_\{t\}\(h,x,y\_\{<h\},a\_\{h\}\)\\bigg\]\.\(H\.17\)This identity uses only the marginal law of each rollout, not independence within the batch\. Combining \([H\.16](https://arxiv.org/html/2609.38666#A8.E16)\) and \([H\.17](https://arxiv.org/html/2609.38666#A8.E17)\), we obtain, onℰ\\mathcal\{E\},

Regret⁡\(T\)≤H22​∑t=1T𝔼⁡\[Xt∣ℱt\]\.\\displaystyle\\operatorname\{Regret\}\(T\)\\leq\\frac\{H^\{2\}\}\{2\}\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\[X\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]\.\(H\.18\)Apply Lemma[I\.6](https://arxiv.org/html/2609.38666#A9.Thmtheorem6)to\{Xt\}t=1T\\\{X\_\{t\}\\\}\_\{t=1\}^\{T\}, with range bound4​B24B^\{2\}and failure probabilityδ/2\\delta/2\. With probability at least1−δ/21\-\\delta/2,

∑t=1T𝔼⁡\[Xt∣ℱt\]≤2​∑t=1TXt\+32​B2​log⁡4δ\.\\displaystyle\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\[X\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]\\leq 2\\sum\_\{t=1\}^\{T\}X\_\{t\}\+32B^\{2\}\\log\\frac\{4\}\{\\delta\}\.\(H\.19\)Taking a union bound, we conclude that, with probability at least1−δ1\-\\delta,

Regret⁡\(T\)≤H2​∑t=1TXt\+16​H2​B2​log⁡4δ\.\\displaystyle\\operatorname\{Regret\}\(T\)\\leq H^\{2\}\\sum\_\{t=1\}^\{T\}X\_\{t\}\+16H^\{2\}B^\{2\}\\log\\frac\{4\}\{\\delta\}\.\(H\.20\)Finally, by the definition of the bonus function, we have

ψt​\(z\)=4​min⁡\{\(4​β\+λ\)​Dℱ2​\(z,Z<t\),B2\}\.\\displaystyle\\psi\_\{t\}\(z\)=4\\min\\big\\\{\(4\\beta\+\\lambda\)D\_\{\\mathcal\{F\}\}^\{2\}\(z;Z\_\{<t\}\),B^\{2\}\\big\\\}\.Our choice ofβ\\betaensures4​β\+λ≥B24\\beta\+\\lambda\\geq B^\{2\}\. Therefore,

ψt​\(z\)≤4​\(4​β\+λ\)​min⁡\{1,Dℱ2​\(z,Z<t\)\}\.\\displaystyle\\psi\_\{t\}\(z\)\\leq 4\(4\\beta\+\\lambda\)\\min\\big\\\{1,D\_\{\\mathcal\{F\}\}^\{2\}\(z;Z\_\{<t\}\)\\big\\\}\.\(H\.21\)Indeed, whenDℱ2​\(z,Z<t\)≤1D\_\{\\mathcal\{F\}\}^\{2\}\(z;Z\_\{<t\}\)\\leq 1, we use the first term in the minimum definingψt\\psi\_\{t\}\. Otherwise, we useψt​\(z\)≤4​B2≤4​\(4​β\+λ\)\\psi\_\{t\}\(z\)\\leq 4B^\{2\}\\leq 4\(4\\beta\+\\lambda\)\.

Summing \([H\.21](https://arxiv.org/html/2609.38666#A8.E21)\) over the observed queries gives

∑t=1TXt\\displaystyle\\sum\_\{t=1\}^\{T\}X\_\{t\}=∑t=1T1m​H​∑j=1m∑h=1Hψt​\(zt,j,h\)\\displaystyle=\\sum\_\{t=1\}^\{T\}\\frac\{1\}\{mH\}\\sum\_\{j=1\}^\{m\}\\sum\_\{h=1\}^\{H\}\\psi\_\{t\}\(z\_\{t,j,h\}\)≤4​\(4​β\+λ\)​∑t=1T1m​H​∑j=1m∑h=1Hmin⁡\{1,Dℱ2​\(zt,j,h,Z<t\)\}\\displaystyle\\leq 4\(4\\beta\+\\lambda\)\\sum\_\{t=1\}^\{T\}\\frac\{1\}\{mH\}\\sum\_\{j=1\}^\{m\}\\sum\_\{h=1\}^\{H\}\\min\\big\\\{1,D\_\{\\mathcal\{F\}\}^\{2\}\(z\_\{t,j,h\};Z\_\{<t\}\)\\big\\\}≤4​\(4​β\+λ\)​dimT\(ℱ;λ\)\.\\displaystyle\\leq 4\(4\\beta\+\\lambda\)\\dim\_\{T\}\(\\mathcal\{F\};\\lambda\)\.The last inequality follows from the definition of the generalized Eluder dimension, since the observed batches form a valid sequence of autoregressive trajectories\. Substituting this inequality into \([H\.20](https://arxiv.org/html/2609.38666#A8.E20)\) completes the proof of Theorem[H\.4](https://arxiv.org/html/2609.38666#A8.Thmtheorem4)\. ∎

### H\.3Proof of Lemma[H\.3](https://arxiv.org/html/2609.38666#A8.Thmtheorem3)

###### Proof\.

Letℱt\\mathcal\{F\}\_\{t\}denote the history before roundtt\. Fixf∈ℱf\\in\\mathcal\{F\}, and define

Δt,j,h​\(f\)\\displaystyle\\Delta\_\{t,j,h\}\(f\):=f⁡\(zt,j,h\)−f∗​\(zt,j,h\),\\displaystyle:=f\(z\_\{t,j,h\}\)\-f^\{\*\}\(z\_\{t,j,h\}\),ξt,j,h\\displaystyle\\xi\_\{t,j,h\}:=lt,j,h−f∗​\(zt,j,h\)\.\\displaystyle:=l\_\{t,j,h\}\-f^\{\*\}\(z\_\{t,j,h\}\)\.The conditional\-mean identity gives

𝔼\[ξt,j,h∣ℱt,zt,j,h\]=0\.\\displaystyle\\mathbb\{E\}\[\\xi\_\{t,j,h\}\\mid\\mathcal\{F\}\_\{t\},z\_\{t,j,h\}\]=0\.\(H\.22\)Moreover, conditioned onzt,j,h=z=\(h,x,u,a\)z\_\{t,j,h\}=z=\(h,x,u,a\), Assumption[5\.3](https://arxiv.org/html/2609.38666#S5.Thmtheorem3)implies

ξt,j,h∈\[log⁡πref​\(a\|x,u\)−B−f∗​\(z\),log⁡πref​\(a\|x,u\)\+B−f∗​\(z\)\]\.\\displaystyle\\xi\_\{t,j,h\}\\in\\big\[\\log\\pi\_\{\\text\{ref\}\}\(a\|x,u\)\-B\-f^\{\*\}\(z\),\\,\\log\\pi\_\{\\text\{ref\}\}\(a\|x,u\)\+B\-f^\{\*\}\(z\)\\big\]\.We use the abbreviationΔ:=Δt,j,h​\(f\),ξ:=ξt,j,h\\Delta:=\\Delta\_\{t,j,h\}\(f\),\\xi:=\\xi\_\{t,j,h\}\. Then, applying Assumption[5\.3](https://arxiv.org/html/2609.38666#S5.Thmtheorem3)tof∗f^\{\*\}and observing that it is an average over all the teachers, we have

𝔼\[ξ∣ℱt,zt,j,h\]=0,\|ξ\|≤2B\.\\displaystyle\\mathbb\{E\}\[\\xi\\mid\\mathcal\{F\}\_\{t\},z\_\{t,j,h\}\]=0,\\qquad\|\\xi\|\\leq 2B\.Setη=1/\(32​B2\)\\eta=1/\(32B^\{2\}\)\. We claim that

𝔼\[eη⁡\(2​Δ​ξ−Δ2/2\)∣ℱt,zt,j,h\]≤1\.\\displaystyle\\mathbb\{E\}\\big\[e^\{\\eta\(2\\Delta\\xi\-\\Delta^\{2\}/2\)\}\\mid\\mathcal\{F\}\_\{t\},z\_\{t,j,h\}\\big\]\\leq 1\.\(H\.23\)In fact, if\|Δ\|≥8​B\|\\Delta\|\\geq 8B, then

2​Δ​ξ−Δ22≤4​B​\|Δ\|−Δ22≤0\.\\displaystyle 2\\Delta\\xi\-\\frac\{\\Delta^\{2\}\}\{2\}\\leq 4B\|\\Delta\|\-\\frac\{\\Delta^\{2\}\}\{2\}\\leq 0\.Thus, \([H\.23](https://arxiv.org/html/2609.38666#A8.E23)\) holds in this case\. Otherwise, if\|Δ\|<8​B\|\\Delta\|<8B, then

\|2​η​Δ​ξ\|≤2​η⋅8​B⋅2​B=1\.\\displaystyle\|2\\eta\\Delta\\xi\|\\leq 2\\eta\\cdot 8B\\cdot 2B=1\.Usingeu≤1\+u\+u2e^\{u\}\\leq 1\+u\+u^\{2\}for\|u\|≤1\|u\|\\leq 1, we obtain

𝔼\[e2​η​Δ​ξ∣ℱt,zt,j,h\]\\displaystyle\\mathbb\{E\}\[e^\{2\\eta\\Delta\\xi\}\\mid\\mathcal\{F\}\_\{t\},z\_\{t,j,h\}\]≤1\+2ηΔ𝔼\[ξ∣ℱt,zt,j,h\]\+4η2Δ2𝔼\[ξ2∣ℱt,zt,j,h\]\\displaystyle\\leq 1\+2\\eta\\Delta\\mathbb\{E\}\[\\xi\\mid\\mathcal\{F\}\_\{t\},z\_\{t,j,h\}\]\+4\\eta^\{2\}\\Delta^\{2\}\\mathbb\{E\}\[\\xi^\{2\}\\mid\\mathcal\{F\}\_\{t\},z\_\{t,j,h\}\]≤1\+16​η2​B2​Δ2\\displaystyle\\leq 1\+16\\eta^\{2\}B^\{2\}\\Delta^\{2\}≤exp⁡\(16​η2​B2​Δ2\)\\displaystyle\\leq\\exp\(16\\eta^\{2\}B^\{2\}\\Delta^\{2\}\)=exp⁡\(η​Δ2/2\)\.\\displaystyle=\\exp\(\\eta\\Delta^\{2\}/2\)\.The second inequality uses the zero conditional mean andξ2≤4​B2\\xi^\{2\}\\leq 4B^\{2\}\. The third uses1\+v≤ev1\+v\\leq e^\{v\}, and the last equality follows from our choice ofη\\eta\. Multiplying both sides bye−ηΔ2/2e^\{\-\\eta\\Delta^\{2\}/2\}proves \([H\.23](https://arxiv.org/html/2609.38666#A8.E23)\)\.

To handle the dependence within a round, apply Jensen’s inequality to the exponential function:

exp⁡\(ηm​H​∑j=1m∑h=1H\[2​Δt,j,h​\(f\)​ξt,j,h−12​Δt,j,h​\(f\)2\]\)\\displaystyle\\exp\\bigg\(\\frac\{\\eta\}\{mH\}\\sum\_\{j=1\}^\{m\}\\sum\_\{h=1\}^\{H\}\\bigg\[2\\Delta\_\{t,j,h\}\(f\)\\xi\_\{t,j,h\}\-\\frac\{1\}\{2\}\\Delta\_\{t,j,h\}\(f\)^\{2\}\\bigg\]\\bigg\)≤1m​H​∑j=1m∑h=1Hexp⁡\(η⁡\[2​Δt,j,h​\(f\)​ξt,j,h−12​Δt,j,h​\(f\)2\]\)\.\\displaystyle\\qquad\\leq\\frac\{1\}\{mH\}\\sum\_\{j=1\}^\{m\}\\sum\_\{h=1\}^\{H\}\\exp\\bigg\(\\eta\\bigg\[2\\Delta\_\{t,j,h\}\(f\)\\xi\_\{t,j,h\}\-\\frac\{1\}\{2\}\\Delta\_\{t,j,h\}\(f\)^\{2\}\\bigg\]\\bigg\)\.Taking conditional expectations givenℱt\\mathcal\{F\}\_\{t\}and applying \([H\.23](https://arxiv.org/html/2609.38666#A8.E23)\) to each summand gives

𝔼⁡\[exp⁡\(ηm​H​∑j=1m∑h=1H\[2​Δt,j,h​\(f\)​ξt,j,h−12​Δt,j,h​\(f\)2\]\)\|ℱt\]≤1\.\\displaystyle\\mathbb\{E\}\\bigg\[\\exp\\bigg\(\\frac\{\\eta\}\{mH\}\\sum\_\{j=1\}^\{m\}\\sum\_\{h=1\}^\{H\}\\bigg\[2\\Delta\_\{t,j,h\}\(f\)\\xi\_\{t,j,h\}\-\\frac\{1\}\{2\}\\Delta\_\{t,j,h\}\(f\)^\{2\}\\bigg\]\\bigg\)\\,\\bigg\|\\,\\mathcal\{F\}\_\{t\}\\bigg\]\\leq 1\.\(H\.24\)We do not condition on the entire batch or assume independence among its observations\.

DefineM0​\(f\)=1M\_\{0\}\(f\)=1and

Mt​\(f\):=exp⁡\(η​∑s=1t1m​H​∑j=1m∑h=1H\[2​Δs,j,h​\(f\)​ξs,j,h−12​Δs,j,h​\(f\)2\]\)\.\\displaystyle M\_\{t\}\(f\):=\\exp\\bigg\(\\eta\\sum\_\{s=1\}^\{t\}\\frac\{1\}\{mH\}\\sum\_\{j=1\}^\{m\}\\sum\_\{h=1\}^\{H\}\\bigg\[2\\Delta\_\{s,j,h\}\(f\)\\xi\_\{s,j,h\}\-\\frac\{1\}\{2\}\\Delta\_\{s,j,h\}\(f\)^\{2\}\\bigg\]\\bigg\)\.Equation \([H\.24](https://arxiv.org/html/2609.38666#A8.E24)\) implies

𝔼⁡\[Mt​\(f\)∣ℱt\]≤Mt−1​\(f\)\.\\displaystyle\\mathbb\{E\}\[M\_\{t\}\(f\)\\mid\\mathcal\{F\}\_\{t\}\]\\leq M\_\{t\-1\}\(f\)\.Thus,\{Mt​\(f\)\}t=0T\\\{M\_\{t\}\(f\)\\\}\_\{t=0\}^\{T\}is a nonnegative supermartingale\. By Ville’s inequality \(Lemma[I\.9](https://arxiv.org/html/2609.38666#A9.Thmtheorem9)\),

ℙ⁡\(max0≤t≤T⁡Mt​\(f\)\>\|ℱ\|δ\)≤δ\|ℱ\|\.\\displaystyle\\mathbb\{P\}\\bigg\(\\max\_\{0\\leq t\\leq T\}M\_\{t\}\(f\)\>\\frac\{\|\\mathcal\{F\}\|\}\{\\delta\}\\bigg\)\\leq\\frac\{\\delta\}\{\|\\mathcal\{F\}\|\}\.Taking a union bound overf∈ℱf\\in\\mathcal\{F\}, we obtain an eventℰ\\mathcal\{E\}withℙ⁡\(ℰ\)≥1−δ\\mathbb\{P\}\(\\mathcal\{E\}\)\\geq 1\-\\deltaon which, simultaneously for everyf∈ℱf\\in\\mathcal\{F\}andt∈\{0,…,T\}t\\in\\\{0,\\ldots,T\\\},

∑s=1t1m​H​∑j=1m∑h=1H\[2​Δs,j,h​\(f\)​ξs,j,h−12​Δs,j,h​\(f\)2\]\\displaystyle\\sum\_\{s=1\}^\{t\}\\frac\{1\}\{mH\}\\sum\_\{j=1\}^\{m\}\\sum\_\{h=1\}^\{H\}\\bigg\[2\\Delta\_\{s,j,h\}\(f\)\\xi\_\{s,j,h\}\-\\frac\{1\}\{2\}\\Delta\_\{s,j,h\}\(f\)^\{2\}\\bigg\]≤1η​log⁡\|ℱ\|δ\\displaystyle\\leq\\frac\{1\}\{\\eta\}\\log\\frac\{\|\\mathcal\{F\}\|\}\{\\delta\}=32​B2​log⁡\|ℱ\|δ\\displaystyle=32B^\{2\}\\log\\frac\{\|\\mathcal\{F\}\|\}\{\\delta\}≤β\.\\displaystyle\\leq\\beta\.\(H\.25\)Expanding the difference of squared losses gives

\(f⁡\(zs,j,h\)−ls,j,h\)2−\(f∗​\(zs,j,h\)−ls,j,h\)2=Δs,j,h​\(f\)2−2​Δs,j,h​\(f\)​ξs,j,h\.\\displaystyle\(f\(z\_\{s,j,h\}\)\-l\_\{s,j,h\}\)^\{2\}\-\(f^\{\*\}\(z\_\{s,j,h\}\)\-l\_\{s,j,h\}\)^\{2\}=\\Delta\_\{s,j,h\}\(f\)^\{2\}\-2\\Delta\_\{s,j,h\}\(f\)\\xi\_\{s,j,h\}\.Therefore, onℰ\\mathcal\{E\},

Lt​\(f\)−Lt​\(f∗\)\\displaystyle L\_\{t\}\(f\)\-L\_\{t\}\(f^\{\*\}\)=12​∑s=1t1m​H​∑j=1m∑h=1HΔs,j,h​\(f\)2\\displaystyle=\\frac\{1\}\{2\}\\sum\_\{s=1\}^\{t\}\\frac\{1\}\{mH\}\\sum\_\{j=1\}^\{m\}\\sum\_\{h=1\}^\{H\}\\Delta\_\{s,j,h\}\(f\)^\{2\}−∑s=1t1m​H∑j=1m∑h=1H\[2Δs,j,h\(f\)ξs,j,h−12Δs,j,h\(f\)2\]\\displaystyle\\quad\-\\sum\_\{s=1\}^\{t\}\\frac\{1\}\{mH\}\\sum\_\{j=1\}^\{m\}\\sum\_\{h=1\}^\{H\}\\bigg\[2\\Delta\_\{s,j,h\}\(f\)\\xi\_\{s,j,h\}\-\\frac\{1\}\{2\}\\Delta\_\{s,j,h\}\(f\)^\{2\}\\bigg\]≥12​∑s=1t1m​H​∑j=1m∑h=1HΔs,j,h​\(f\)2−β,\\displaystyle\\geq\\frac\{1\}\{2\}\\sum\_\{s=1\}^\{t\}\\frac\{1\}\{mH\}\\sum\_\{j=1\}^\{m\}\\sum\_\{h=1\}^\{H\}\\Delta\_\{s,j,h\}\(f\)^\{2\}\-\\beta,\(H\.26\)where the last inequality holds due to \([H\.25](https://arxiv.org/html/2609.38666#A8.E25)\)\. It holds for anyf∈ℱf\\in\\mathcal\{F\}, so it also applies tof¯t−1\\bar\{f\}\_\{t\-1\}\. Moreover, by \([H\.5](https://arxiv.org/html/2609.38666#A8.E5)\) and the realizability assumption, we have

Lt−1​\(f¯t−1\)≤Lt−1​\(f∗\)\.\\displaystyle L\_\{t\-1\}\(\\bar\{f\}\_\{t\-1\}\)\\leq L\_\{t\-1\}\(f^\{\*\}\)\.Applying \([H\.26](https://arxiv.org/html/2609.38666#A8.E26)\) at timet−1t\-1consequently gives

∑s<t1m​H​∑j=1m∑h=1H\[f¯t−1​\(zs,j,h\)−f∗​\(zs,j,h\)\]2≤2​β\.\\displaystyle\\sum\_\{s<t\}\\frac\{1\}\{mH\}\\sum\_\{j=1\}^\{m\}\\sum\_\{h=1\}^\{H\}\\big\[\\bar\{f\}\_\{t\-1\}\(z\_\{s,j,h\}\)\-f^\{\*\}\(z\_\{s,j,h\}\)\\big\]^\{2\}\\leq 2\\beta\.\(H\.27\)Since bothf¯t−1\\bar\{f\}\_\{t\-1\}andf∗f^\{\*\}belong toℱ\\mathcal\{F\}, the definition ofDℱ2D\_\{\\mathcal\{F\}\}^\{2\}implies, for everyz∈𝒵z\\in\\mathcal\{Z\},

\|f¯t−1​\(z\)−f∗​\(z\)\|2\\displaystyle\\big\|\\bar\{f\}\_\{t\-1\}\(z\)\-f^\{\*\}\(z\)\\big\|^\{2\}≤Dℱ2​\(z,Z<t\)​\(λ\+∑s<t1m​H​∑j=1m∑h=1H\[f¯t−1​\(zs,j,h\)−f∗​\(zs,j,h\)\]2\)\\displaystyle\\quad\\leq D\_\{\\mathcal\{F\}\}^\{2\}\(z;Z\_\{<t\}\)\\bigg\(\\lambda\+\\sum\_\{s<t\}\\frac\{1\}\{mH\}\\sum\_\{j=1\}^\{m\}\\sum\_\{h=1\}^\{H\}\\big\[\\bar\{f\}\_\{t\-1\}\(z\_\{s,j,h\}\)\-f^\{\*\}\(z\_\{s,j,h\}\)\\big\]^\{2\}\\bigg\)≤\(λ\+2​β\)​Dℱ2​\(z,Z<t\)\\displaystyle\\quad\\leq\(\\lambda\+2\\beta\)D\_\{\\mathcal\{F\}\}^\{2\}\(z;Z\_\{<t\}\)≤bt−1​\(z\)2\.\\displaystyle\\quad\\leq b\_\{t\-1\}\(z\)^\{2\}\.Taking square roots proves the pointwise confidence bound\.

Finally, forz=\(h,x,u,a\)z=\(h,x,u,a\), Assumption[5\.3](https://arxiv.org/html/2609.38666#S5.Thmtheorem3)and the definition off∗f^\{\*\}give

log⁡πref​\(a\|x,u\)−B≤f∗​\(z\)≤log⁡πref​\(a\|x,u\)\+B\.\\displaystyle\\log\\pi\_\{\\text\{ref\}\}\(a\|x,u\)\-B\\leq f^\{\*\}\(z\)\\leq\\log\\pi\_\{\\text\{ref\}\}\(a\|x,u\)\+B\.Onℰ\\mathcal\{E\}, the confidence bound also gives

f¯t−1​\(z\)\+bt−1​\(z\)≥f∗​\(z\)\.\\displaystyle\\bar\{f\}\_\{t\-1\}\(z\)\+b\_\{t\-1\}\(z\)\\geq f^\{\*\}\(z\)\.Therefore,f^t−1​\(z\)≥f∗​\(z\)\\widehat\{f\}\_\{t\-1\}\(z\)\\geq f^\{\*\}\(z\), which proves optimism\. Moreover,

f^t−1​\(z\)−f∗​\(z\)\\displaystyle\\widehat\{f\}\_\{t\-1\}\(z\)\-f^\{\*\}\(z\)=min⁡\{f¯t−1​\(z\)−f∗​\(z\)\+bt−1​\(z\),log⁡πref​\(a\|x,u\)\+B−f∗​\(z\)\}\\displaystyle=\\min\\big\\\{\\bar\{f\}\_\{t\-1\}\(z\)\-f^\{\*\}\(z\)\+b\_\{t\-1\}\(z\),\\,\\log\\pi\_\{\\text\{ref\}\}\(a\|x,u\)\+B\-f^\{\*\}\(z\)\\big\\\}≤min⁡\{2​bt−1​\(z\),2​B\}\.\\displaystyle\\leq\\min\\\{2b\_\{t\-1\}\(z\),2B\\\}\.This completes the proof of Lemma[H\.3](https://arxiv.org/html/2609.38666#A8.Thmtheorem3)\. ∎

## Appendix IAuxiliary Lemmas

### I\.1Uniform Bernstein Concentration in[Bakhtiari et al\. \(2026\)](https://arxiv.org/html/2609.38666#bib.bib1)

First, we list several key results from[Bakhtiari et al\. \(2026\)](https://arxiv.org/html/2609.38666#bib.bib1)that are used in our analysis\. To start with, we present several definitions\.

###### Definition I\.1\(CGF\-like, Definition 4[Bakhtiari et al\. 2026](https://arxiv.org/html/2609.38666#bib.bib1)\)\.

A twice differentiable functionψ:\[0,λmax\)→ℝ\+\\psi:\[0,\\lambda\_\{\\max\}\)\\to\\mathbb\{R\}\_\{\+\}is called CGF\-like if it is convex, non\-negative, and satisfiesψ⁡\(0\)=ψ′​\(0\)=0\\psi\(0\)=\\psi^\{\\prime\}\(0\)=0\.

###### Definition I\.2\(sub\-ψ\\psiprocess, Definition 5[Bakhtiari et al\. 2026](https://arxiv.org/html/2609.38666#bib.bib1)\)\.

Let\{ℱt\}t∈\[T\]\\\{\\mathcal\{F\}\_\{t\}\\\}\_\{t\\in\[T\]\}be a filtration,ψ:\[0,λmax\)→ℝ\+\\psi:\[0,\\lambda\_\{\\max\}\)\\to\\mathbb\{R\}\_\{\+\}be a CGF\-like function\. Let\{St\}t∈\[T\]\\\{S\_\{t\}\\\}\_\{t\\in\[T\]\}be a real\-valued stochastic process adapted to\{ℱt\}\\\{\\mathcal\{F\}\_\{t\}\\\},\{Vt\}t∈\[T\]\\\{V\_\{t\}\\\}\_\{t\\in\[T\]\}be another process adapted to\{ℱt\}\\\{\\mathcal\{F\}\_\{t\}\\\}taking values inℝ\+\\mathbb\{R\}\_\{\+\}\. We say that\{\(St,Vt\)\}t∈\[T\]\\\{\(S\_\{t\},V\_\{t\}\)\\\}\_\{t\\in\[T\]\}is a sub\-ψ\\psiprocess if for allt∈\[T\]t\\in\[T\]andλ∈\[0,λmax\)\\lambda\\in\[0,\\lambda\_\{\\max\}\), there exists anℱ\\mathcal\{F\}\-adapted supermartingale\{Lt​\(λ\)\}t∈\[T\]\\\{L\_\{t\}\(\\lambda\)\\\}\_\{t\\in\[T\]\}such that

Mt​\(λ\):=exp⁡\(λ​St−∑s=1tψ⁡\(λ\)​Vs\)≤Lt​\(λ\),a\.s\.\\displaystyle M\_\{t\}\(\\lambda\):=\\exp\\Big\(\\lambda S\_\{t\}\-\\sum\_\{s=1\}^\{t\}\\psi\(\\lambda\)V\_\{s\}\\Big\)\\leq L\_\{t\}\(\\lambda\),\\quad\\text\{a\.s\.\}

###### Definition I\.3\(Definition 6[Bakhtiari et al\. 2026](https://arxiv.org/html/2609.38666#bib.bib1)\)\.

A random process\{\(St,Vt\)\}t∈\[T\]\\\{\(S\_\{t\},V\_\{t\}\)\\\}\_\{t\\in\[T\]\}is sub\-gamma with parameterv\>0v\>0if it is sub\-ψ\\psiwith the CGF\-like functionψ⁡\(λ\)=λ22​\(1−v​λ\)\\psi\(\\lambda\)=\\frac\{\\lambda^\{2\}\}\{2\(1\-v\\lambda\)\}for somev\>0v\>0\.

###### Lemma I\.4\(Proposition 11[Bakhtiari et al\. 2026](https://arxiv.org/html/2609.38666#bib.bib1)\)\.

Let\{ℱt\}t∈\[T\]\\\{\\mathcal\{F\}\_\{t\}\\\}\_\{t\\in\[T\]\}be a filtration and let\(Xt\)t∈\[T\]\(X\_\{t\}\)\_\{t\\in\[T\]\}be a real\-valued stochastic process adapted toℱ\\mathcal\{F\}\. Assume that there exists a constantb\>0b\>0such that for allt∈\[T\]t\\in\[T\],

Xt≤𝔼⁡\[Xt\|ℱt−1\]\+b​a\.s\.\\displaystyle X\_\{t\}\\leq\\mathbb\{E\}\[X\_\{t\}\|\\mathcal\{F\}\_\{t\-1\}\]\+b\\ a\.s\.Then, for

St=∑s=1t\(Xs−𝔼⁡\[Xs\|ℱs−1\]\),Vt=∑s=1tVar⁡\(Xs\|ℱs−1\),\\displaystyle S\_\{t\}=\\sum\_\{s=1\}^\{t\}\(X\_\{s\}\-\\mathbb\{E\}\[X\_\{s\}\|\\mathcal\{F\}\_\{s\-1\}\]\),\\quad V\_\{t\}=\\sum\_\{s=1\}^\{t\}\\Var\(X\_\{s\}\|\\mathcal\{F\}\_\{s\-1\}\),the process\{\(St,Vt\)\}t∈\[T\]\\\{\(S\_\{t\},V\_\{t\}\)\\\}\_\{t\\in\[T\]\}is sub\-gamma with parameterb/3b/3\.

###### Lemma I\.5\(Theorem 10[Bakhtiari et al\. 2026](https://arxiv.org/html/2609.38666#bib.bib1)\)\.

For any sub\-gamma process\{\(St,Vt\)\}t∈\[T\]\\\{\(S\_\{t\},V\_\{t\}\)\\\}\_\{t\\in\[T\]\}with parameterv\>0v\>0, and anyρ\>0\\rho\>0,δ∈\(0,1\)\\delta\\in\(0,1\), with probability at least1−δ1\-\\delta, we have for allt∈\[T\]t\\in\[T\],

St≤4​Vt​log⁡\(Ht/δ\)\+11​\(v\+ρ\)​log⁡\(Ht/δ\),\\displaystyle S\_\{t\}\\leq 4\\sqrt\{V\_\{t\}\\log\(H\_\{t\}/\\delta\)\}\+11\(v\+\\rho\)\\log\(H\_\{t\}/\\delta\),whereHt=log⁡\(1\+Vt/ρ2\)\+eH\_\{t\}=\\log\(1\+V\_\{t\}/\\rho^\{2\}\)\+e\.

### I\.2Other Auxiliary Lemmas

###### Lemma I\.6\(Lemma B\.3,[Foster et al\. 2024](https://arxiv.org/html/2609.38666#bib.bib2)\)\.

Let\{𝑿t\}t∈\[T\]\\\{\\bm\{X\}\_\{t\}\\\}\_\{t\\in\[T\]\}be a sequence of non\-negative random variables adapted to a filtration\{ℱt\}t∈\[T\]\\\{\\mathcal\{F\}\_\{t\}\\\}\_\{t\\in\[T\]\}\. Assume that0≤𝑿t≤R0\\leq\\bm\{X\}\_\{t\}\\leq Ralmost surely for allt∈\[T\]t\\in\[T\]and some constantR\>0R\>0\. Then, for anyδ∈\(0,1\)\\delta\\in\(0,1\)and allT′≤TT^\{\\prime\}\\leq T, with probability at least1−δ1\-\\delta, we have

∑t=1T′𝔼⁡\[𝑿t\|ℱt−1\]\\displaystyle\\sum\_\{t=1\}^\{T^\{\\prime\}\}\\mathbb\{E\}\[\\bm\{X\}\_\{t\}\|\\mathcal\{F\}\_\{t\-1\}\]≤2​∑t=1T′𝑿t\+8​R​log⁡\(2/δ\),\\displaystyle\\leq 2\\sum\_\{t=1\}^\{T^\{\\prime\}\}\\bm\{X\}\_\{t\}\+8R\\log\(2/\\delta\),∑t=1T′𝑿t\\displaystyle\\sum\_\{t=1\}^\{T^\{\\prime\}\}\\bm\{X\}\_\{t\}≤32​∑t=1T′𝔼⁡\[𝑿t\|ℱt−1\]\+4​R​log⁡\(2/δ\)\.\\displaystyle\\leq\\frac\{3\}\{2\}\\sum\_\{t=1\}^\{T^\{\\prime\}\}\\mathbb\{E\}\[\\bm\{X\}\_\{t\}\|\\mathcal\{F\}\_\{t\-1\}\]\+4R\\log\(2/\\delta\)\.

###### Lemma I\.7\.

Letμf​\(⋅\)∝exp⁡\(f⁡\(⋅\)\)\\mu\_\{f\}\(\\cdot\)\\propto\\exp\(f\(\\cdot\)\),μg​\(⋅\)∝exp⁡\(g⁡\(⋅\)\)\\mu\_\{g\}\(\\cdot\)\\propto\\exp\(g\(\\cdot\)\)be two probability distributions\. Then,

KL\(μf∥μg\)=𝔼μf\[f−g\]\+log\[𝔼μf\[exp\(g−f\)\]\]\.\\displaystyle\\text\{KL\}\(\\mu\_\{f\}\\\|\\mu\_\{g\}\)=\\mathbb\{E\}\_\{\\mu\_\{f\}\}\[f\-g\]\+\\log\\big\[\\mathbb\{E\}\_\{\\mu\_\{f\}\}\[\\exp\(g\-f\)\]\\big\]\.

###### Proof\.

LetZf=∑xexp⁡\(f⁡\(x\)\)Z\_\{f\}=\\sum\_\{x\}\\exp\(f\(x\)\),Zg=∑xexp⁡\(g⁡\(x\)\)Z\_\{g\}=\\sum\_\{x\}\\exp\(g\(x\)\)be the normalization constants\. Then,

μf​\(x\)=exp⁡\(f⁡\(x\)\)Zf,μg​\(x\)=exp⁡\(g⁡\(x\)\)Zg\.\\displaystyle\\mu\_\{f\}\(x\)=\\frac\{\\exp\(f\(x\)\)\}\{Z\_\{f\}\},\\ \\mu\_\{g\}\(x\)=\\frac\{\\exp\(g\(x\)\)\}\{Z\_\{g\}\}\.The KL\-divergence is equal to

KL\(μf∥μg\)\\displaystyle\\text\{KL\}\(\\mu\_\{f\}\\\|\\mu\_\{g\}\)=∑xμf​\(x\)​log⁡μf​\(x\)μg​\(x\)\\displaystyle=\\sum\_\{x\}\\mu\_\{f\}\(x\)\\log\\frac\{\\mu\_\{f\}\(x\)\}\{\\mu\_\{g\}\(x\)\}=∑xμf​\(x\)​\[f⁡\(x\)−g⁡\(x\)\+log⁡\(Zg/Zf\)\]\\displaystyle=\\sum\_\{x\}\\mu\_\{f\}\(x\)\\big\[f\(x\)\-g\(x\)\+\\log\(Z\_\{g\}/Z\_\{f\}\)\\big\]=𝔼μf​\[f−g\]\+log⁡\(Zg/Zf\)\.\\displaystyle=\\mathbb\{E\}\_\{\\mu\_\{f\}\}\[f\-g\]\+\\log\(Z\_\{g\}/Z\_\{f\}\)\.Moreover,

Zg/Zf\\displaystyle Z\_\{g\}/Z\_\{f\}=∑xexp⁡\(g⁡\(x\)\)∑xexp⁡\(f⁡\(x\)\)\\displaystyle=\\frac\{\\sum\_\{x\}\\exp\(g\(x\)\)\}\{\\sum\_\{x\}\\exp\(f\(x\)\)\}=∑xexp⁡\(f⁡\(x\)\)∑yexp⁡\(f⁡\(y\)\)⋅exp⁡\(g⁡\(x\)−f⁡\(x\)\)\\displaystyle=\\sum\_\{x\}\\frac\{\\exp\(f\(x\)\)\}\{\\sum\_\{y\}\\exp\(f\(y\)\)\}\\cdot\\exp\(g\(x\)\-f\(x\)\)=𝔼μf​\[exp⁡\(g−f\)\]\.\\displaystyle=\\mathbb\{E\}\_\{\\mu\_\{f\}\}\[\\exp\(g\-f\)\]\.Thus, we complete the proof\. ∎

###### Lemma I\.8\(Stirling’s inequality\)\.

LetΓ⁡\(⋅\)\\Gamma\(\\cdot\)be the gamma function\. For anys\>0s\>0, the following inequality holds

\(s−12\)​log⁡s−s\+12​log⁡\(2​π\)\\displaystyle\\Big\(s\-\\frac\{1\}\{2\}\\Big\)\\log s\-s\+\\frac\{1\}\{2\}\\log\(2\\pi\)≤log⁡Γ⁡\(s\)\\displaystyle\\leq\\log\\Gamma\(s\)≤\(s−12\)​log⁡s−s\+12​log⁡\(2​π\)\+112​s\.\\displaystyle\\leq\\Big\(s\-\\frac\{1\}\{2\}\\Big\)\\log s\-s\+\\frac\{1\}\{2\}\\log\(2\\pi\)\+\\frac\{1\}\{12s\}\.

###### Lemma I\.9\(Ville’s inequality\)\.

Let\{Xt\}t=0T\\\{X\_\{t\}\\\}\_\{t=0\}^\{T\}be a non\-negative supermartingale\. Then, for anyv\>0v\>0,

ℙ\[max0≤t≤TXt≥v\]≤𝔼⁡\[X0\]v\.\\displaystyle\\mathbb\{P\}\\Big\[\\max\_\{0\\leq t\\leq T\}X\_\{t\}\\geq v\\Big\]\\leq\\frac\{\\mathbb\{E\}\[X\_\{0\}\]\}\{v\}\.

###### Proof\.

Please refer to[Durrett \(2019\)](https://arxiv.org/html/2609.38666#bib.bib13)\. ∎

## References

- Agarwalet al\.\(2023\)A\. Agarwal, Y\. Jin, and T\. ZhangVOQQl: towards optimal regret in model\-free rl with nonlinear function approximation\.InThe Thirty Sixth Annual Conference on Learning Theory,pp\. 987–1063\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px3.p3.1)\.
- Agarwalet al\.\(2024\)R\. Agarwal, N\. Vieillard, Y\. Zhou, P\. Stanczyk, S\. Ramos Garea, M\. Geist, and O\. BachemOn\-policy distillation of language models: learning from self\-generated mistakes\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 21246–21263\.Cited by:[§1](https://arxiv.org/html/2609.38666#S1.p1.1),[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1)\.
- Aminianet al\.\(2026\)G\. Aminian, A\. R\. Asadi, I\. Shenfeld, and Y\. MrouehKl\-regularized rlhf with multiple reference models: exact solutions and sample complexity\.Advances in Neural Information Processing Systems38,pp\. 101117–101151\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px3.p2.1)\.
- Ayoubet al\.\(2020\)A\. Ayoub, Z\. Jia, C\. Szepesvari, M\. Wang, and L\. YangModel\-based reinforcement learning with value\-targeted regression\.InInternational Conference on Machine Learning,pp\. 463–474\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px3.p3.1)\.
- Bakhtiariet al\.\(2026\)A\. Bakhtiari, A\. Ayoub, S\. Robertson, D\. Janz, and C\. SzepesváriEluder dimension: localise it\!\.arXiv preprint arXiv:2601\.09825\.Cited by:[§I\.1](https://arxiv.org/html/2609.38666#A9.SS1),[§I\.1](https://arxiv.org/html/2609.38666#A9.SS1.p1.1),[Definition I\.1](https://arxiv.org/html/2609.38666#A9.Thmtheorem1),[Definition I\.2](https://arxiv.org/html/2609.38666#A9.Thmtheorem2),[Definition I\.3](https://arxiv.org/html/2609.38666#A9.Thmtheorem3),[Lemma I\.4](https://arxiv.org/html/2609.38666#A9.Thmtheorem4),[Lemma I\.5](https://arxiv.org/html/2609.38666#A9.Thmtheorem5)\.
- Chenet al\.\(2025\)H\. Chen, N\. Razin, K\. Narasimhan, and D\. ChenRetaining by doing: the role of on\-policy data in mitigating forgetting\.arXiv preprint arXiv:2510\.18874\.Cited by:[§1](https://arxiv.org/html/2609.38666#S1.p1.1),[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px2.p1.1)\.
- Diet al\.\(2024\)Q\. Di, H\. Zhao, J\. He, and Q\. GuPessimistic nonlinear least\-squares value iteration for offline reinforcement learning\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 4377–4410\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px3.p3.1)\.
- Durrett \(2019\)R\. DurrettProbability: theory and examples\.Vol\.49,Cambridge university press\.Cited by:[§I\.2](https://arxiv.org/html/2609.38666#A9.SS2.p2.1.1)\.
- Fosteret al\.\(2024\)D\. J\. Foster, A\. Block, and D\. MisraIs behavior cloning all you need? understanding horizon in imitation learning\.Advances in Neural Information Processing Systems37,pp\. 120602–120666\.Cited by:[Lemma I\.6](https://arxiv.org/html/2609.38666#A9.Thmtheorem6),[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px3.p1.1)\.
- Fuet al\.\(2026\)Y\. Fu, H\. Huang, K\. Jiang, J\. Liu, Z\. Jiang, Y\. Zhu, and D\. ZhaoRevisiting on\-policy distillation: empirical failure modes and simple fixes\.arXiv preprint arXiv:2603\.25562\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px2.p1.1)\.
- Guet al\.\(2024\)Y\. Gu, L\. Dong, F\. Wei, and M\. HuangMinillm: knowledge distillation of large language models\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 32694–32717\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1)\.
- Honget al\.\(2026\)H\. Hong, Z\. Wang, Q\. Gu, and H\. WangOnline kl\-regularized reinforcement learning with function approximation under misspecification\.arXiv preprint arXiv:2606\.06053\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px3.p2.1)\.
- Houet al\.\(2026\)W\. Hou, S\. Peng, W\. Wang, Z\. Ruan, Y\. Zhang, Z\. Zhou, M\. Gao, Y\. Chen, K\. Wang, H\. Yang,et al\.Uni\-opd: unifying on\-policy distillation with a dual\-perspective recipe\.arXiv preprint arXiv:2605\.03677\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1)\.
- Hübotteret al\.\(2026\)J\. Hübotter, F\. Lübeck, L\. Behric, A\. Baumann, M\. Bagatella, D\. Marta, I\. Hakimi, I\. Shenfeld, T\. K\. Buening, C\. Guestrin,et al\.Reinforcement learning via self\-distillation\.arXiv preprint arXiv:2601\.20802\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1),[Remark 6\.2](https://arxiv.org/html/2609.38666#S6.Thmtheorem2.p1.1)\.
- Janget al\.\(2026\)I\. Jang, J\. Yeom, J\. Yeo, H\. Lim, and T\. KimStable on\-policy distillation through adaptive target reformulation\.InFindings of the Association for Computational Linguistics: ACL 2026,pp\. 42217–42227\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1)\.
- Jiet al\.\(2026a\)K\. Ji, Q\. Di, H\. Zhao, Q\. Zhao, and Q\. GuOn the optimal sample complexity of offline multi\-armed bandits with kl regularization\.arXiv preprint arXiv:2605\.02141\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px3.p2.1)\.
- Jiet al\.\(2026b\)K\. Ji, Q\. Zhao, H\. Zhao, Q\. Di, and Q\. GuNear\-optimal regret for kl\-regularized multi\-armed bandits\.arXiv preprint arXiv:2603\.02155\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px3.p2.1)\.
- Jinet al\.\(2025\)H\. Jin, S\. Luan, T\. Ni, S\. Lyu, G\. Rabusseau, R\. Rabbany, D\. Precup, and M\. HamdaqaRl fine\-tuning heals ood forgetting in sft\.arXiv preprint arXiv:2509\.12235\.Cited by:[§1](https://arxiv.org/html/2609.38666#S1.p1.1)\.
- Jinet al\.\(2026\)W\. Jin, T\. Min, Y\. Yang, D\. Wei, Y\. Zhou, S\. R\. Kadhe, N\. Baracaldo, and K\. LeeEntropy\-aware on\-policy distillation of language models\.arXiv preprint arXiv:2603\.07079\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1)\.
- Koet al\.\(2025\)J\. Ko, T\. Chen, S\. Kim, T\. Ding, L\. Liang, I\. Zharkov, and S\. YunDistillm\-2: a contrastive approach boosts the distillation of llms\.arXiv preprint arXiv:2503\.07067\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1)\.
- Koet al\.\(2024\)J\. Ko, S\. Kim, T\. Chen, and S\. YunDistillm: towards streamlined distillation for large language models\.arXiv preprint arXiv:2402\.03898\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1)\.
- Krichevsky and Trofimov \(1981\)R\. Krichevsky and V\. TrofimovThe performance of universal encoding\.IEEE Transactions on Information Theory27\(2\),pp\. 199–207\.Cited by:[§4\.1](https://arxiv.org/html/2609.38666#S4.SS1.p1.6)\.
- Liet al\.\(2026\)Y\. Li, Y\. Zuo, B\. He, J\. Zhang, C\. Xiao, C\. Qian, T\. Yu, H\. Gao, W\. Yang, Z\. Liu,et al\.Rethinking on\-policy distillation of large language models: phenomenology, mechanism, and recipe\.arXiv preprint arXiv:2604\.13016\.Cited by:[§1](https://arxiv.org/html/2609.38666#S1.p2.1),[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px2.p1.1),[Remark 6\.4](https://arxiv.org/html/2609.38666#S6.Thmtheorem4.p1.1),[Remark 6\.6](https://arxiv.org/html/2609.38666#S6.Thmtheorem6.p1.1)\.
- Liuet al\.\(2026\)Y\. Liu, S\. Zhang, Y\. Zhang, and Q\. GuSelf\-distilled policy gradient\.arXiv preprint arXiv:2606\.04036\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1)\.
- Lu and Lab \(2025\)K\. Lu and T\. M\. LabOn\-policy distillation\.Thinking Machines Lab: Connectionism\.Note:https://thinkingmachines\.ai/blog/on\-policy\-distillationExternal Links:[Document](https://dx.doi.org/10.64434/tml.20251026)Cited by:[§1](https://arxiv.org/html/2609.38666#S1.p1.1)\.
- Luoet al\.\(2026\)F\. Luo, Y\. Chuang, G\. Wang, Z\. Xu, X\. Han, T\. Zhang, and V\. BravermanDemystifying opd: length inflation and stabilization strategies for large language models\.arXiv preprint arXiv:2604\.08527\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1)\.
- Maet al\.\(2026a\)W\. Ma, J\. Wei, L\. Zhao, H\. Zhang, B\. Xiao, L\. Li, Q\. Yang, B\. Gao, Y\. Wang, R\. Li,et al\.Mopd: multi\-teacher on\-policy distillation for capability integration in llm post\-training\.arXiv preprint arXiv:2606\.30406\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1),[Remark 6\.4](https://arxiv.org/html/2609.38666#S6.Thmtheorem4.p1.1)\.
- Maet al\.\(2026b\)Y\. Ma, Z\. Zhu, M\. Jiang, and C\. XiaoOne student, many teachers: multi\-task on\-policy distillation via soft\-prompt privileged context\.arXiv preprint arXiv:2607\.18293\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1),[Remark 6\.2](https://arxiv.org/html/2609.38666#S6.Thmtheorem2.p1.1)\.
- Nayaket al\.\(2025\)A\. Nayak, T\. Yang, O\. Yagan, G\. Joshi, and Y\. ChiAchieving logarithmic regret in kl\-regularized zero\-sum markov games\.arXiv preprint arXiv:2510\.13060\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px3.p2.1)\.
- Ohet al\.\(2026\)M\. Oh, S\. Song, G\. Choi, Y\. Choi, and Y\. JoKL for a kl: on\-policy distillation with control variate baseline\.arXiv preprint arXiv:2605\.07865\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1)\.
- Rajaramanet al\.\(2021\)N\. Rajaraman, Y\. Han, L\. Yang, J\. Liu, J\. Jiao, and K\. RamchandranOn the value of interaction and function approximation in imitation learning\.Advances in Neural Information Processing Systems34,pp\. 1325–1336\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px3.p1.1)\.
- Rajaramanet al\.\(2020\)N\. Rajaraman, L\. Yang, J\. Jiao, and K\. RamchandranToward the fundamental limits of imitation learning\.Advances in Neural Information Processing Systems33,pp\. 2914–2924\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px3.p1.1)\.
- Ross and Bagnell \(2010\)S\. Ross and D\. BagnellEfficient reductions for imitation learning\.InProceedings of the thirteenth international conference on artificial intelligence and statistics,pp\. 661–668\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px3.p1.1)\.
- Rosset al\.\(2011\)S\. Ross, G\. Gordon, and D\. BagnellA reduction of imitation learning and structured prediction to no\-regret online learning\.InProceedings of the fourteenth international conference on artificial intelligence and statistics,pp\. 627–635\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px3.p1.1)\.
- Russo and Van Roy \(2013\)D\. Russo and B\. Van RoyEluder dimension and the sample complexity of optimistic exploration\.Advances in Neural Information Processing Systems26\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px3.p3.1)\.
- Shenfeldet al\.\(2026a\)I\. Shenfeld, M\. Damani, J\. Hübotter, and P\. AgrawalSelf\-distillation enables continual learning\.arXiv preprint arXiv:2601\.19897\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1),[Remark 6\.2](https://arxiv.org/html/2609.38666#S6.Thmtheorem2.p1.1)\.
- Shenfeldet al\.\(2026b\)I\. Shenfeld, J\. Pari, and P\. AgrawalRl’s razor: why online reinforcement learning forgets less\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 59839–59864\.Cited by:[§1](https://arxiv.org/html/2609.38666#S1.p1.1),[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px2.p1.1)\.
- Song and Zheng \(2026\)M\. Song and M\. ZhengA survey of on\-policy distillation for large language models\.arXiv preprint arXiv:2604\.00626\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1)\.
- Sriramanet al\.\(2026\)V\. Sriraman, P\. Liu, D\. Hsu, and A\. BlockBehavior cloning is not all you need: the optimality of on\-policy distillation for noisy expert feedback\.arXiv preprint arXiv:2606\.30923\.Cited by:[§1](https://arxiv.org/html/2609.38666#S1.p2.1),[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px2.p1.1)\.
- Syed and Schapire \(2010\)U\. Syed and R\. SchapireA reduction from apprenticeship learning to classification\.Advances in neural information processing systems23\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px3.p1.1)\.
- Vianoet al\.\(2026\)L\. Viano, A\. Moulin, A\. Huang, V\. Cevher, P\. Amortila, and D\. J\. FosterWhen does on\-policy interaction help? representational tradeoffs in value\-based imitation learning\.arXiv preprint arXiv:2607\.29617\.Cited by:[§1](https://arxiv.org/html/2609.38666#S1.p2.1),[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px2.p1.1)\.
- Wanget al\.\(2026\)R\. Wang, H\. Wang, Y\. Chen, B\. Xue, T\. Fang, W\. Yu, and K\. WongDemystifying on\-policy distillation: roles, pathologies, and regulations\.arXiv preprint arXiv:2607\.13399\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px2.p1.1),[Remark 6\.4](https://arxiv.org/html/2609.38666#S6.Thmtheorem4.p1.1)\.
- Wanget al\.\(2020\)R\. Wang, R\. R\. Salakhutdinov, and L\. YangReinforcement learning with general value function approximation: provably efficient approach via bounded eluder dimension\.Advances in Neural Information Processing Systems33,pp\. 6123–6135\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px3.p3.1)\.
- Wuet al\.\(2026\)D\. Wu, C\. Shi, J\. Yang, and C\. ShenGreedy sampling is provably efficient for rlhf\.Advances in Neural Information Processing Systems38,pp\. 108198–108232\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px3.p2.1)\.
- Wuet al\.\(2025\)Y\. Wu, R\. Thareja, P\. Vepakomma, and F\. OrabonaOffline and online kl\-regularized rlhf under differential privacy\.arXiv preprint arXiv:2510\.13512\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px3.p2.1)\.
- Xiaoet al\.\(2026\)B\. Xiao, B\. Xia, B\. Yang, B\. Gao, B\. Shen, C\. Zhang, C\. He, C\. Lou, F\. Luo, G\. Wang,et al\.Mimo\-v2\-flash technical report\.arXiv preprint arXiv:2601\.02780\.Cited by:[§1](https://arxiv.org/html/2609.38666#S1.p1.1)\.
- Xieet al\.\(2026\)Z\. Xie, L\. L\. Zhang, Z\. Xie, and M\. YangTrust region policy distillation\.arXiv preprint arXiv:2607\.04751\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1)\.
- Xinet al\.\(2026\)H\. Xin, A\. Zhao, Y\. Sun, J\. Li, X\. Shen, and H\. XiongEscaping the kl agreement trap in on\-policy distillation\.arXiv preprint arXiv:2606\.09471\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1)\.
- Xinget al\.\(2026\)X\. Xing, H\. Wang, B\. Gao, Z\. Li, and Y\. TangTrust region on\-policy distillation\.arXiv preprint arXiv:2606\.01249\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1)\.
- Xuet al\.\(2025\)W\. Xu, R\. Han, Z\. Wang, L\. Le, D\. Madeka, L\. Li, W\. Wang, R\. Agarwal, C\. Lee, and T\. PfisterSpeculative knowledge distillation: bridging the teacher\-student gap through interleaved sampling\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 64616–64646\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1)\.
- Yanget al\.\(2025\)A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§1](https://arxiv.org/html/2609.38666#S1.p1.1)\.
- Yeet al\.\(2026\)T\. Ye, L\. Dong, X\. Wu, S\. Huang, and F\. WeiOn\-policy context distillation for language models\.arXiv preprint arXiv:2602\.12275\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1)\.
- Zenget al\.\(2026\)A\. Zeng, X\. Lv, Z\. Hou, Z\. Du, Q\. Zheng, B\. Chen, D\. Yin, C\. Ge, C\. Huang, C\. Xie,et al\.Glm\-5: from vibe coding to agentic engineering\.arXiv preprint arXiv:2602\.15763\.Cited by:[§1](https://arxiv.org/html/2609.38666#S1.p1.1)\.
- Zhanget al\.\(2026\)D\. Zhang, Z\. Yang, S\. Janghorbani, J\. Han, A\. Ressler II, Q\. Qian, G\. D\. Lyng, S\. S\. Batra, and R\. E\. TillmanFast and effective on\-policy distillation from reasoning prefixes\.InFindings of the Association for Computational Linguistics: ACL 2026,pp\. 25553–25569\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1)\.
- Zhaoet al\.\(2026a\)A\. Zhao, J\. Tong, Y\. Fan, P\. Nie, W\. Li, and X\. ShenPowerOPD: stabilizing on\-policy distillation with bounded power transformation\.arXiv preprint arXiv:2606\.17199\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1)\.
- Zhaoet al\.\(2024\)H\. Zhao, J\. He, and Q\. GuA nearly optimal and low\-switching algorithm for reinforcement learning with general function approximation\.Advances in Neural Information Processing Systems37,pp\. 94684–94735\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px3.p3.1)\.
- Zhaoet al\.\(2026b\)H\. Zhao, C\. Ye, Q\. Gu, and T\. ZhangSharp analysis for kl\-regularized contextual bandits and rlhf\.Advances in Neural Information Processing Systems38,pp\. 107964–108002\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px3.p2.1)\.
- Zhaoet al\.\(2025\)H\. Zhao, C\. Ye, W\. Xiong, Q\. Gu, and T\. ZhangLogarithmic regret for online kl\-regularized reinforcement learning\.arXiv preprint arXiv:2502\.07460\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px3.p2.1)\.
- Zhaoet al\.\(2026c\)Q\. Zhao, K\. Ji, H\. Zhao, and Q\. GuFast rates for offline contextual bandits with forward\-kl regularization under single\-policy concentrability\.arXiv preprint arXiv:2605\.09214\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px3.p2.1)\.
- Zhaoet al\.\(2026d\)Q\. Zhao, K\. Ji, H\. Zhao, T\. Zhang, and Q\. GuTowards a sharp analysis of offline policy learning forff\-divergence\-regularized contextual bandits\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 117176–117210\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px3.p2.1)\.
- Zhaoet al\.\(2026e\)S\. Zhao, Z\. Xie, M\. Liu, J\. Huang, G\. Pang, F\. Chen, and A\. GroverSelf\-distilled reasoner: on\-policy self\-distillation for large language models\.arXiv preprint arXiv:2601\.18734\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1)\.
- Zhenget al\.\(2026\)B\. Zheng, X\. Ma, Y\. Liang, J\. Ruan, X\. Fu, K\. Lin, B\. Zhu, K\. Zeng, and X\. CaiScope: signal\-calibrated on\-policy distillation enhancement with dual\-path adaptive weighting\.arXiv preprint arXiv:2604\.10688\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1)\.
- Zhonget al\.\(2026\)Z\. Zhong, H\. Yan, J\. Li, J\. He, T\. Zhang, and H\. LiVla\-opd: bridging offline sft and online rl for vision\-language\-action models via on\-policy distillation\.arXiv preprint arXiv:2603\.26666\.Cited by:[Remark 6\.2](https://arxiv.org/html/2609.38666#S6.Thmtheorem2.p1.1)\.
- Zhuet al\.\(2026\)W\. Zhu, R\. Xie, R\. Wang, and P\. LiuHybrid policy distillation for llms\.arXiv preprint arXiv:2604\.20244\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1)\.
- Zihenget al\.\(2026\)Z\. Ziheng, J\. Li, H\. Tang, Y\. N\. Wu, and D\. TerzopoulosLess is more: early stopping rollout for on\-policy distillation\.arXiv preprint arXiv:2605\.27028\.Cited by:[§2](https://arxiv.org/html/2609.38666#S2.SS0.SSS0.Px1.p1.1),[Remark 6\.6](https://arxiv.org/html/2609.38666#S6.Thmtheorem6.p1.1)\.

相似文章

同策略蒸馏(5分钟阅读)

TLDR AI

本文引入同策略蒸馏,通过在教师提供的token级KL正则化下,在学生自身轨迹上训练学生模型,解决训练-推理分布不匹配问题,统一了前向KL、反向KL和JSD损失,其中反向KL更适用于较小的学生模型。

揭秘 On-Policy Distillation:角色、病理与调控

Hugging Face Daily Papers

本文系统研究了LLM后训练中的on-policy distillation,阐明了其作为探索催化剂的作用,并识别了Student-Teacher Mismatch和Length Exploitation等病理现象,提出了轻量级信号调控方法。

On-Policy 蒸馏中的 Off-Policy 教师问题

Hugging Face Daily Papers

该论文指出了 On-Policy 蒸馏中存在的一种 off-policy 不对称性:教师必须监督来自学生自身生成的、其未训练过的前缀。为此,论文提出了 SCOUT——一个通过可验证奖励的 RL 持续适应教师的协同训练框架,在不同配置、模型规模和推理领域上均一致地提升了蒸馏效果。

OPRD:在策略表示蒸馏

Hugging Face Daily Papers

OPRD提出了一种新的知识蒸馏方法,该方法在策略部署期间跨层对齐学生和教师的隐藏状态,消除了来自词空间KL估计的采样方差。实验表明,OPRD在数学推理基准(AIME 2024/2025、AIMO)上优于输出空间基线,同时速度快1.44倍,内存使用减少54%。