Online Learning on Hidden-Convex Losses via Algorithmic Equivalence: Optimal Regret, Geometric Barrier, and Bandit Feedback

arXiv cs.LG Papers

Summary

This paper proves that online gradient descent achieves optimal √T regret for hidden-convex losses under a Hessian compatibility condition, resolving open questions in adversarial online learning. It also extends results to one-point bandit feedback with a T^{3/4} expected regret bound.

arXiv:2605.26373v1 Announce Type: new Abstract: We study adversarial online learning with hidden-convex losses, i.e., nonconvex losses that become convex after a nonlinear reparameterization. Ghai, Lu and Hazan (2022) proved that, under geometric and smoothness assumptions, online gradient descent (OGD) on such nonconvex losses approximately simulates online mirror descent (OMD) on the underlying convex losses with a suitable regularizer, yielding $\mathcal{O}(T^{2/3})$ regret. They left open whether the optimal $\Theta(\sqrt{T})$ regret from online convex optimization can be recovered in this hidden-convex setting. We answer this question affirmatively. More specifically, via a sharper discrete-time algorithmic equivalence argument, we prove that OGD achieves $\mathcal{O}(\sqrt{T})$ regret under the same assumptions, matching the optimal worst-case rate for adversarial online convex optimization. We also address another open question of Ghai, Lu and Hazan (2022) by clarifying the geometry required for this algorithmic equivalence. We replace the diagonal-Jacobian sufficient condition with a necessary-and-sufficient Hessian compatibility condition, thereby expanding the class of admissible reparameterizations. We complement our tight regret bound with a lower bound showing that the Hessian compatibility assumption is essential for OGD; when it fails, we construct a smooth reparameterization and an adversarial sequence of hidden-convex losses for which OGD suffers $\Omega(T)$ regret. Finally, we extend our analysis to one-point bandit feedback and prove a $\mathcal{O}(T^{3/4})$ expected regret bound for bandit OGD with spherical smoothing, matching its classical rate on convex losses.
Original Article
View Cached Full Text

Cached at: 05/27/26, 09:09 AM

# Online Learning on Hidden-Convex Losses via Algorithmic Equivalence: Optimal Regret, Geometric Barrier, and Bandit Feedback
Source: [https://arxiv.org/html/2605.26373](https://arxiv.org/html/2605.26373)
Anas Barakat1Andreas Kontogiannis2,5Vasilis Pollatos3,5 Ioannis Panageas4Antonios Varvitsiotis1,5,6

###### Abstract

We study adversarial online learning with hidden\-convex losses, i\.e\., nonconvex losses that become convex after a nonlinear reparameterization\.Ghaiet al\.\([2022](https://arxiv.org/html/2605.26373#bib.bib46)\)proved that, under geometric and smoothness assumptions, online gradient descent \(OGD\) on such nonconvex losses approximately simulates online mirror descent \(OMD\) on the underlying convex losses with a suitable regularizer, yielding𝒪​\(T2/3\)\\mathcal\{O\}\(T^\{2/3\}\)regret\. They left open whether the optimalΘ​\(T\)\\Theta\(\\sqrt\{T\}\)regret from online convex optimization can be recovered in this hidden\-convex setting\. We answer this question affirmatively\. More specifically, via a sharper discrete\-time algorithmic equivalence argument, we prove that OGD achieves𝒪​\(T\)\\mathcal\{O\}\(\\sqrt\{T\}\)regret under the same assumptions, matching the optimal worst\-case rate for adversarial online convex optimization\. We also address another open question of\(Ghaiet al\.,[2022](https://arxiv.org/html/2605.26373#bib.bib46)\)by clarifying the geometry required for this algorithmic equivalence\. We replace the diagonal\-Jacobian sufficient condition with a necessary\-and\-sufficient Hessian compatibility condition, thereby expanding the class of admissible reparameterizations\. We complement our tight regret bound with a lower bound showing that the Hessian compatibility assumption is essential for OGD; when it fails, we construct a smooth reparameterization and an adversarial sequence of hidden\-convex losses for which OGD suffersΩ​\(T\)\\Omega\(T\)regret\. Finally, we extend our analysis to one\-point bandit feedback and prove a𝒪​\(T3/4\)\\mathcal\{O\}\(T^\{3/4\}\)expected regret bound for bandit OGD with spherical smoothing, matching its classical rate on convex losses\.

00footnotetext:Affiliations:1Singapore University of Technology and Design2National Technical University of Athens3National and Kapodistrian University of Athens4University of California, Irvine5Archimedes, Athena Research Center, Greece6National University of Singapore, Centre for Quantum Technologies00footnotetext:Contact:barakat9anas@gmail\.com, andr\.kontog@gmail\.com, vaspoll97@gmail\.com,
ipanagea@ics\.uci\.edu,antonios@sutd\.edu\.sg\.## 1Introduction

Online convex optimization \(OCO\) provides a theoretical framework for adversarial sequential decision\-making: for convex Lipschitz losses, simple first\-order methods such as online gradient descent \(OGD\) achieve the optimalΘ​\(T\)\\Theta\(\\sqrt\{T\}\)regret rate\(Zinkevich,[2003](https://arxiv.org/html/2605.26373#bib.bib27); Hazan,[2016](https://arxiv.org/html/2605.26373#bib.bib39); Shalev\-Shwartz,[2012](https://arxiv.org/html/2605.26373#bib.bib38)\)\. Beyond convexity, the situation is substantially more delicate\. For arbitrary nonconvex losses, sublinear regret against the best fixed decision in hindsight is generally unattainable by simple first\-order methods without additional structure or oracle access\(Kricheneet al\.,[2015](https://arxiv.org/html/2605.26373#bib.bib35); Agarwalet al\.,[2019](https://arxiv.org/html/2605.26373#bib.bib34); Suggala and Netrapalli,[2020](https://arxiv.org/html/2605.26373#bib.bib36); Héliouet al\.,[2020](https://arxiv.org/html/2605.26373#bib.bib33)\)\. For instance, it is known that no deterministic algorithm can achieve sublinear regret in online nonconvex learning in general\(Suggala and Netrapalli,[2020](https://arxiv.org/html/2605.26373#bib.bib36), Proposition 3\)\. This motivates the search for structured classes of nonconvex online problems that retain enough geometry to permit convex\-like guarantees with simple first\-order algorithms\.

A prominent form of benign nonconvexity ishidden convexity\(Ben\-Tal and Teboulle,[1996](https://arxiv.org/html/2605.26373#bib.bib16); Liet al\.,[2005](https://arxiv.org/html/2605.26373#bib.bib15); Fatkhullinet al\.,[2025a](https://arxiv.org/html/2605.26373#bib.bib42); Levinet al\.,[2025](https://arxiv.org/html/2605.26373#bib.bib29); Gorissenet al\.,[2026](https://arxiv.org/html/2605.26373#bib.bib14)\)\. This structural property appears in several modern optimization problems, including neural\-network training\(Wanget al\.,[2022](https://arxiv.org/html/2605.26373#bib.bib18); Patel and Vlatakis\-Gkaragkounis,[2025](https://arxiv.org/html/2605.26373#bib.bib22); Ergen and Pilanci,[2025](https://arxiv.org/html/2605.26373#bib.bib1); Zeger and Pilanci,[2026](https://arxiv.org/html/2605.26373#bib.bib17)\), reinforcement learning\(Hazanet al\.,[2019](https://arxiv.org/html/2605.26373#bib.bib3); Zhanget al\.,[2020](https://arxiv.org/html/2605.26373#bib.bib2); Zahavyet al\.,[2021](https://arxiv.org/html/2605.26373#bib.bib5); Barakatet al\.,[2023](https://arxiv.org/html/2605.26373#bib.bib7),[2025](https://arxiv.org/html/2605.26373#bib.bib8); D’Orazioet al\.,[2025](https://arxiv.org/html/2605.26373#bib.bib20)\), and nonconvex games\(Vlatakis\-Gkaragkouniset al\.,[2019](https://arxiv.org/html/2605.26373#bib.bib23); Mladenovicet al\.,[2022](https://arxiv.org/html/2605.26373#bib.bib9); Sakoset al\.,[2023](https://arxiv.org/html/2605.26373#bib.bib24); Kalogianniset al\.,[2024](https://arxiv.org/html/2605.26373#bib.bib21); Gempet al\.,[2025](https://arxiv.org/html/2605.26373#bib.bib12); Kalogianniset al\.,[2025](https://arxiv.org/html/2605.26373#bib.bib11); Barakatet al\.,[2026](https://arxiv.org/html/2605.26373#bib.bib10)\)\. In these settings, the objective is nonconvex in the algorithm’s native parameters but convex after an appropriate nonlinear change of variables\. This structure has been exploited in deterministic and stochastic optimization to obtain global convergence guarantees for nonconvex problems\(Chenet al\.,[2025](https://arxiv.org/html/2605.26373#bib.bib40); Fatkhullinet al\.,[2025a](https://arxiv.org/html/2605.26373#bib.bib42),[b](https://arxiv.org/html/2605.26373#bib.bib41); Bhaskaraet al\.,[2025](https://arxiv.org/html/2605.26373#bib.bib13)\)\. In the online setting, however, the learner faces a distinct challenge: the loss functions which are revealed sequentially over time may even be chosen adversarially and the goal becomes to achieve low regret against the best fixed comparator in hindsight rather than convergence to the minimizer of a fixed function\.

Ghaiet al\.\([2022](https://arxiv.org/html/2605.26373#bib.bib46)\)initiated the study of such a structured nonconvex online learning setting through algorithmic equivalence\. They showed that, under suitable geometric and smoothness assumptions on the reparameterization, OGD applied to the nonconvex losses approximately simulates OMD applied to the corresponding convex losses in the reparameterized space\. This connection led to a𝒪​\(T2/3\)\\mathcal\{O\}\(T^\{2/3\}\)regret bound\. Their result left open whether the slower\-than\-convex rate is inherent to the nonconvex online learning setting\. Since the identity reparameterization recovers ordinary OCO, the best possible regret one could hope for in this class with exact gradient feedback isΘ​\(T\)\\Theta\(\\sqrt\{T\}\)\. This raises the central question:

> When does OGD recover the optimal regret under exact gradient feedback for nonconvex hidden\-convex losses, and does the same algorithmic equivalence principle extend to the bandit feedback setting?

In this paper, we answer the exact gradient feedback question affirmatively and show that the framework extends to one\-point bandit feedback\. Specifically, we prove a𝒪​\(T\)\\mathcal\{O\}\(\\sqrt\{T\}\)regret for OGD, discuss the required compatibility conditions needed for the algorithmic equivalence, show a linear\-regret lower bound without it, and sublinear expected regret guarantees in the bandit setting\. We elaborate on our contributions in more details in the next section\.

### 1\.1Main contributions

We study online nonconvex learning for losses that are convex under a nonlinear reparameterization which isunknownto the learner\. Our results address several open questions formulated byGhaiet al\.\([2022](https://arxiv.org/html/2605.26373#bib.bib46)\), refining and extending their algorithmic equivalence analysis between OGD and OMD\. Our main contributions are as follows:

- •Optimal regret under exact gradient feedback\.In the exact gradient setting, we prove that OGD achieves𝒪​\(T\)\\mathcal\{O\}\(\\sqrt\{T\}\)regret for hidden\-convex losses \(Theorem[1](https://arxiv.org/html/2605.26373#Thmtheorem1)\) under the geometric compatibility and smoothness assumptions considered inGhaiet al\.\([2022](https://arxiv.org/html/2605.26373#bib.bib46)\)\. This improves their𝒪​\(T2/3\)\\mathcal\{O\}\(T^\{2/3\}\)bound and matches the optimal regret for adversarial onlineconvexoptimization\. Since classical OCO is recovered by the identity reparameterization, the rate is optimal for this problem class in general\. Our analysis is based on a sharper algorithmic equivalence between OGD and online mirror descent \(OMD\), improving the perturbation analysis underlying the previous𝒪​\(T2/3\)\\mathcal\{O\}\(T^\{2/3\}\)guarantee\. Specifically, our proof shows that OGD on the nonconvex losses tracks OMD in the convex parameterization with sufficiently small discretization error to preserve the classicalT\\sqrt\{T\}regret\.
- •Geometry of the reparameterization\.Prior work\(Ghaiet al\.,[2022](https://arxiv.org/html/2605.26373#bib.bib46)\)imposed a diagonal\-Jacobian condition to guarantee the Hessian compatibility needed for the OGD–OMD equivalence\. We relax this structural assumption and give necessary and sufficient conditions for compatibility \(Proposition[1](https://arxiv.org/html/2605.26373#Thmproposition1)\), thereby enlarging the class of reparameterizations covered by the analysis\.
- •Geometric barrier for OGD\.We show that the Hessian compatibility condition is not merely a proof artifact\. We prove that there exists a smooth reparameterization violating compatibility and an adversarial sequence of hidden\-convex losses for which OGD incursΩ​\(T\)\\Omega\(T\)regret \(Theorem[2](https://arxiv.org/html/2605.26373#Thmtheorem2)\)\.
- •Bandit feedback\.We extend our analysis to the one\-point bandit setting, where the learner only observes a single loss value per round\. Against oblivious adversaries, we prove an expected regret bound of𝒪​\(T3/4\)\\mathcal\{O\}\(T^\{3/4\}\)under standard additional Lipschitzness and boundedness assumptions \(Theorem[3](https://arxiv.org/html/2605.26373#Thmtheorem3)\)\. Notably, our result matches the best known regret rate of OGD for convex losses in the bandit setting\(Flaxmanet al\.,[2005](https://arxiv.org/html/2605.26373#bib.bib43)\)\.

Together, our results establish new regret guarantees for OGD for a structured class ofonline nonconvexoptimization problems, extending classical results from onlineconvexoptimization\.

### 1\.2Related Work

Online nonconvex optimization\.Online convex optimization is a standard framework for adversarial sequential decision\-making; see, e\.g\.,Hazan \([2016](https://arxiv.org/html/2605.26373#bib.bib39)\); Shalev\-Shwartz \([2012](https://arxiv.org/html/2605.26373#bib.bib38)\); Orabona \([2019](https://arxiv.org/html/2605.26373#bib.bib37)\)\. A separate line of work studies online learning with nonconvex losses\(Kricheneet al\.,[2015](https://arxiv.org/html/2605.26373#bib.bib35); Agarwalet al\.,[2019](https://arxiv.org/html/2605.26373#bib.bib34); Suggala and Netrapalli,[2020](https://arxiv.org/html/2605.26373#bib.bib36); Héliouet al\.,[2020](https://arxiv.org/html/2605.26373#bib.bib33)\)\. Sublinear global regret in such settings typically requires access to strong oracles, such as sampling oracles that allow continuous\-domain exponential weights\(Maillard and Munos,[2010](https://arxiv.org/html/2605.26373#bib.bib32); Kricheneet al\.,[2015](https://arxiv.org/html/2605.26373#bib.bib35)\), or offline optimization oracles for the nonconvex losses\(Agarwalet al\.,[2019](https://arxiv.org/html/2605.26373#bib.bib34); Suggala and Netrapalli,[2020](https://arxiv.org/html/2605.26373#bib.bib36)\)\. Our work takes a different route: rather than assuming oracle access for general nonconvex losses, we exploit hidden convexity of the loss sequence and analyze first\-order algorithms through the algorithmic equivalence between OGD in the nonconvex parameterization and OMD in the convex parameterization\(Amid and Warmuth,[2020](https://arxiv.org/html/2605.26373#bib.bib58); Ghaiet al\.,[2022](https://arxiv.org/html/2605.26373#bib.bib46); Liet al\.,[2022](https://arxiv.org/html/2605.26373#bib.bib28)\)\. A few works in the literature also consider structured nonconvex losses with special structure such as the composition of a non\-increasing function with a linear function\(Zhanget al\.,[2015](https://arxiv.org/html/2605.26373#bib.bib44)\), or weaker notions of convexity such as strictly locally quasi\-convexity\(Hazanet al\.,[2015](https://arxiv.org/html/2605.26373#bib.bib25)\)and weak pseudo\-convexity\(Gaoet al\.,[2018](https://arxiv.org/html/2605.26373#bib.bib26)\)\.

Hidden\-convex optimization\.Hidden convexity has recently been studied in several offline deterministic and stochastic optimization settings\(Chenet al\.,[2025](https://arxiv.org/html/2605.26373#bib.bib40); Fatkhullinet al\.,[2025a](https://arxiv.org/html/2605.26373#bib.bib42),[b](https://arxiv.org/html/2605.26373#bib.bib41); Levinet al\.,[2025](https://arxiv.org/html/2605.26373#bib.bib29)\)\. These works consider nonconvex problems that admit a convex reformulation under a nonlinear transformation, often under assumptions weaker than those used in the online algorithmic\-equivalence framework\. Their focus, however, is offline optimization: the objective is fixed, or arises as an expectation over a fixed data\-generating distribution\. By contrast, we study an adversarial online setting in which the loss functions may change arbitrarily over time, and the performance criterion is regret rather than convergence to a global optimizer\. This distinction requires a different analysis, since the learner must compete with the best fixed decision in hindsight while only observing feedback sequentially\.

Online learning under bandit feedback\.The literature on online bandit learning is extensive; we refer toLattimore and Szepesvári \([2020](https://arxiv.org/html/2605.26373#bib.bib31)\); Lattimore \([2024](https://arxiv.org/html/2605.26373#bib.bib30)\)for modern treatments\. For bandit online convex optimization,Flaxmanet al\.\([2005](https://arxiv.org/html/2605.26373#bib.bib43)\)introduced a one\-point gradient estimator based on spherical smoothing, showing that gradient\-based online methods can be implemented when only function values are observed\. The bandit setting has also been studied for general Lipschitz nonconvex losses and for special nonconvex classes\(Zhanget al\.,[2015](https://arxiv.org/html/2605.26373#bib.bib44); Gaoet al\.,[2018](https://arxiv.org/html/2605.26373#bib.bib26)\)\. Our bandit results are different in focus: we consider hidden\-convex losses and use smoothing together with the OGD–OMD equivalence to transfer regret guarantees from the convex parameterization to the original nonconvex one\. Thus, while our algorithmic template is inspired by bandit OCO, the main challenge is controlling the interaction between bandit gradient estimation, nonlinear reparameterization, and the approximation error in algorithmic equivalence\.

## 2Online Learning on Hidden\-Convex Losses

In this section, we introduce the onlinehidden\-convex optimization \(OHCO\) problem which extends the celebrated online convex optimization \(OCO\) setting\(Shalev\-Shwartz,[2012](https://arxiv.org/html/2605.26373#bib.bib38); Hazanet al\.,[2015](https://arxiv.org/html/2605.26373#bib.bib25); Orabona,[2019](https://arxiv.org/html/2605.26373#bib.bib37)\)\.

Hidden\-convex losses\.Let𝒳,𝒴⊂ℝd\{\\mathcal\{X\}\},\{\\mathcal\{Y\}\}\\subset\\mathbb\{R\}^\{d\}be convex compact sets\. Consider a sequence of loss functions\(ℓt\)\(\\ell\_\{t\}\)where for anyt≥1,t~\\geq~1,ℓt:𝒳→ℝ\\ell\_\{t\}:\{\\mathcal\{X\}\}\\to\\mathbb\{R\}has the following hidden\-convexity structure:

ℓt​\(x\)=ht​\(q​\(x\)\),∀x∈𝒳,\\ell\_\{t\}\(x\)=h\_\{t\}\(q\(x\)\)\\,,\\quad\\forall x\\in\{\\mathcal\{X\}\}\\,,\(1\)whereq:𝒳→𝒴q:\{\\mathcal\{X\}\}\\to\{\\mathcal\{Y\}\}is a smooth bijective reparameterization function andht:𝒴→ℝh\_\{t\}:\{\\mathcal\{Y\}\}\\to\\mathbb\{R\}is a convex function\. In this work, we will focus on a subclass of hidden\-convex functions satisfying some geometric and smoothness assumptions made inGhaiet al\.\([2022](https://arxiv.org/html/2605.26373#bib.bib46)\)\. We defer the statement of these assumptions onqqand\(ℓt\)\(\\ell\_\{t\}\)to Sections[3](https://arxiv.org/html/2605.26373#S3)\-[5](https://arxiv.org/html/2605.26373#S5)\. Whileℓt\\ell\_\{t\}is nonconvex in general, it is convex under the reparameterizationqqand we say then that the functionℓt\\ell\_\{t\}ishidden\-convex\. When the reparameterizationqqis simply the identity, the functionℓt\\ell\_\{t\}is a \(truly\) convex function\.

Online regret minimization\.We are interested in the online optimization setting where a learner seeks to minimize an unknown sequence of hidden\-convex functions\(ℓt\)\(\\ell\_\{t\}\)by choosing a sequence\(xt\)\(x\_\{t\}\)where eachxtx\_\{t\}is based only on the information observed before roundtt\. At each roundtt, a learner choosesxt∈𝒳x\_\{t\}\\in\{\\mathcal\{X\}\}using an online algorithm𝒜\\mathcal\{A\}, then an adversary reveals a loss functionℓt\\ell\_\{t\}which is hidden\-convex \(see \([1](https://arxiv.org/html/2605.26373#S2.E1)\)\) and the player suffers the lossℓt​\(xt\)\\ell\_\{t\}\(x\_\{t\}\)\. The goal of the learner is to achieve a good performance with respect to the best possible fixed decision in hindsight, minimizing the \(external\) regret of the algorithm𝒜\\mathcal\{A\}defined for any time horizonT≥1T~\\geq~1as follows:

RT:=∑t=1Tℓt​\(xt\)−minx∈𝒳​∑t=1Tℓt​\(x\)\.R\_\{T\}:=\\sum\_\{t=1\}^\{T\}\\ell\_\{t\}\(x\_\{t\}\)\-\\min\_\{x\\in\{\\mathcal\{X\}\}\}\\sum\_\{t=1\}^\{T\}\\ell\_\{t\}\(x\)\\,\.Sinceq:𝒳→𝒴q:\{\\mathcal\{X\}\}\\to\{\\mathcal\{Y\}\}is bijective, the regret can equivalently be written in hidden coordinates asRT=∑t=1Tht​\(zt\)−minz∈𝒴​∑t=1Tht​\(z\)R\_\{T\}=\\sum\_\{t=1\}^\{T\}h\_\{t\}\(z\_\{t\}\)\-\\min\_\{z\\in\{\\mathcal\{Y\}\}\}\\sum\_\{t=1\}^\{T\}h\_\{t\}\(z\)wherezt=q​\(xt\)\.z\_\{t\}=q\(x\_\{t\}\)\\,\.Unlike in the classical OCO setting, the sequence\(ℓt\)\(\\ell\_\{t\}\)is not supposed to be convex\. We recover the standard OCO setting by choosingqqto be the identity function in \([1](https://arxiv.org/html/2605.26373#S2.E1)\)\.

Information settings\.The reparameterization functionqqis unknown to the learner who cannot run an online algorithm on the sequence of losses\(ht\)\(h\_\{t\}\)and recover a decision in the original space𝒳\{\\mathcal\{X\}\}by invertingqq\. Preconditioning by the Jacobian ofqq\(which is unknown\) is not possible either\. We consider two different settings depending on the feedback available to the learner:

1. \(i\)Exact Gradient Feedback:In this setting, we suppose that the loss functionℓt\\ell\_\{t\}is differentiable and the learner has access to the gradient∇ℓt​\(xt\)\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\(x\_\{t\}\)of the loss functionℓt\\ell\_\{t\}at each time step of the online interaction\.
2. \(ii\)Bandit Feedback:In this setting, the learner has only access to a single function value call to the loss functionℓt\\ell\_\{t\}at each round\.

Online Gradient Descent\.In this work, we will focus on the most famous and classical algorithm to address this problem: Online Gradient Descent \(OGD\)\(Zinkevich,[2003](https://arxiv.org/html/2605.26373#bib.bib27)\)\. In the exact gradient feedback setting, the learner performs the following update for all time stepstt:

xt\+1=Π𝒳​\(xt−η​∇ℓt​\(xt\)\),x\_\{t\+1\}=\\Pi\_\{\{\\mathcal\{X\}\}\}\(x\_\{t\}\-\\eta\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\(x\_\{t\}\)\)\\,,\(OGD\)whereη\\etais a positive step size andΠ𝒳\\Pi\_\{\{\\mathcal\{X\}\}\}is the Euclidean projection over the convex compact set𝒳\{\\mathcal\{X\}\}\. In the bandit setting, the learner will use the one\-point estimate of the gradient resulting in the bandit gradient descent algorithm proposed byFlaxmanet al\.\([2005](https://arxiv.org/html/2605.26373#bib.bib43)\)\. We will discuss this case in more details in Section[5](https://arxiv.org/html/2605.26373#S5)\.

We conclude this section by introducing some general notation that will be used throughout the paper\.

Additional notation\.For a nonzero integerdd, we denote by\[d\]:=\{1,…,d\}\.\[d\]:=\\\{1,\\dots,d\\\}\\,\.For any functionf:ℝd→ℝdf:\\mathbb\{R\}^\{d\}\\rightarrow\\mathbb\{R\}^\{d\}, we denote byJf​\(x\)∈ℝd×dJ\_\{f\}\(x\)\\in\\mathbb\{R\}^\{d\\times d\}the Jacobian offfatxx\. Given a strictly convex and differentiable functionR:ℝd→ℝR:\\mathbb\{R\}^\{d\}\\rightarrow\\mathbb\{R\}, the Bregman divergence induced byRRis defined asDR​\(x∥y\):=R​\(x\)−R​\(y\)−∇R​\(y\)⊤​\(x−y\)\.D\_\{R\}\(x\\\|y\):=R\(x\)\-R\(y\)\-\\nabla\\mkern\-2\.5muR\(y\)^\{\\top\}\(x\-y\)~\.For a closed convex set𝒦\\mathcal\{K\}and a strictly convex regularizerRR, we useΠ𝒦R​\(x\):=arg​miny∈𝒦⁡DR​\(y∥x\)\\Pi^\{R\}\_\{\\mathcal\{K\}\}\(x\):=\\operatorname\*\{arg\\,min\}\_\{y\\in\\mathcal\{K\}\}D\_\{R\}\(y\\\|x\)to denote the Bregman projection, and we use the shorthand notationΠ𝒦:=Π𝒦∥⋅∥2\\Pi\_\{\\mathcal\{K\}\}:=\\Pi^\{\\\|\\cdot\\\|^\{2\}\}\_\{\\mathcal\{K\}\}for the Euclidean projection\. Given a positive\-definite matrixM∈ℝd×dM\\in\\mathbb\{R\}^\{d\\times d\}, we define the norm‖x‖M:=x⊤​M​x\\\|x\\\|\_\{M\}:=\\sqrt\{x^\{\\top\}Mx\}and we denote byM−⊤M^\{\-\\top\}the inverse transpose matrix\. We use the notationBp:=\{x∈ℝd:‖x‖p≤1\}B\_\{p\}:=\\\{x\\in\\mathbb\{R\}^\{d\}:\\\|x\\\|\_\{p\}\\leq 1\\\}for anℓp\\ell\_\{p\}ball andBp\+B^\{\+\}\_\{p\}to denote the intersection of theℓp\\ell\_\{p\}ball and the positive orthant\.

## 3Regret Analysis of OGD for OHCO under Exact Gradient Feedback

In this section, we establish a𝒪​\(T\)\\mathcal\{O\}\(\\sqrt\{T\}\)regret bound for OGD in the OHCO setting described in Section[2](https://arxiv.org/html/2605.26373#S2)under exact gradient feedback, i\.e\. when the learner has access to gradients of the loss functionsℓt\\ell\_\{t\}\. We start by introducing our main assumptions in Section[3\.1](https://arxiv.org/html/2605.26373#S3.SS1)before stating our first main result in Section[3\.2](https://arxiv.org/html/2605.26373#S3.SS2)and providing a proof sketch in Section[3\.3](https://arxiv.org/html/2605.26373#S3.SS3)\.

### 3\.1Assumptions

We consider the same set of assumptions on the reparameterization and the loss functions as inGhaiet al\.\([2022](https://arxiv.org/html/2605.26373#bib.bib46)\)\.

###### Assumption 1\(Hessian\-compatible reparameterization\)\.

The mapq:𝒳→𝒴q:\{\\mathcal\{X\}\}\\to\{\\mathcal\{Y\}\}is a smooth bijection onto the convex set𝒴\{\\mathcal\{Y\}\}and there exists a twice continuously differentiable, strictly convex regularizerR:𝒴→ℝR:\{\\mathcal\{Y\}\}\\to\\mathbb\{R\}such that, for allx∈𝒳x\\in\{\\mathcal\{X\}\}, withz=q​\(x\)z=q\(x\),\[∇2R​\(z\)\]−1=Jq​\(x\)​Jq​\(x\)⊤\.\\big\[\\nabla\\mkern\-2\.5mu^\{2\}R\(z\)\\big\]^\{\-1\}=J\_\{q\}\(x\)J\_\{q\}\(x\)^\{\\top\}\\,\.

Assumption[1](https://arxiv.org/html/2605.26373#Thmassumption1)is a crucial assumption which restricts the class of parameterization functionsqq\. Examples of such parameterizations include the quadratic reparameterization with the entropy regularizer, the exponential reparameterization with the log\-barrier regularizer and the power reparameterization with tempered Bregman divergences\(Amidet al\.,[2019](https://arxiv.org/html/2605.26373#bib.bib45)\)\. We refer the reader to Section 3\.1 inGhaiet al\.\([2022](https://arxiv.org/html/2605.26373#bib.bib46)\)for a more detailed discussion on these examples\. We will elaborate more on Assumption[1](https://arxiv.org/html/2605.26373#Thmassumption1)and its usefulness in Section[3\.3](https://arxiv.org/html/2605.26373#S3.SS3)when discussing the proof of our main results in the exact gradient feedback setting, and in Section[4](https://arxiv.org/html/2605.26373#S4)where we will discuss both verifiable necessary and sufficient conditions under which it is satisfied, and its necessity\.

###### Assumption 2\(Regularizer and reparameterization smoothness\)\.

There existsG\>1G\>1such that:

1. \(i\)The mapqqisGG\-Lipschitz on𝒳\{\\mathcal\{X\}\}, and the 1st and 2nd derivatives ofq−1q^\{\-1\}are bounded byGG\.
2. \(ii\)The regularizerRRis11\-strongly convex and smooth with 1st and 3rd derivatives bounded byGG\.
3. \(iii\)For everyz∈𝒴z\\in\{\\mathcal\{Y\}\}, the mapw↦DR​\(z∥w\)w\\mapsto D\_\{R\}\(z\\\|w\)isGG\-Lipschitz\.

Our last assumption asks that the loss functionsℓt\\ell\_\{t\}have uniformly bounded gradients and that the diameter of𝒳\{\\mathcal\{X\}\}is bounded\. These assumptions are also standard in OCO\.

###### Assumption 3\(Boundedness of gradients and diameter\)\.

There existsGF\>0G\_\{F\}\>0s\.t\.‖∇ℓt​\(x\)‖2≤GF\\\|\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\(x\)\\\|\_\{2\}\\leq G\_\{F\}for allx∈𝒳\.x\\in\{\\mathcal\{X\}\}\\,\.Moreover, there existsD1\>0D\_\{1\}\>0such thatsupz,z′∈𝒴DR​\(z∥z′\)≤D1\\sup\_\{z,z^\{\\prime\}\\in\{\\mathcal\{Y\}\}\}D\_\{R\}\(z\\\|z^\{\\prime\}\)\\leq D\_\{1\}\.

### 3\.2Regret Bound

We are now ready to state the main result of this section\.

###### Theorem 1\(OGD regret under full\-information feedback\)\. Let Assumptions[1](https://arxiv.org/html/2605.26373#Thmassumption1)\-[3](https://arxiv.org/html/2605.26373#Thmassumption3)hold and letT≥1T~\\geq~1\. Settingη=D17​G6​GF3​T\\eta=\\sqrt\{\\frac\{D\_\{1\}\}\{7G^\{6\}G\_\{F\}^\{3\}T\}\}, the regret of \([OGD](https://arxiv.org/html/2605.26373#S2.Ex2)\) run on the sequence of hidden\-convex losses\(ℓt\)\(\\ell\_\{t\}\)defined in \([1](https://arxiv.org/html/2605.26373#S2.E1)\) using stepsizeη\\etais bounded as follows:RT≤7​D1​G6​GF3​TR\_\{T\}~\\leq~\\sqrt\{7D\_\{1\}G^\{6\}G\_\{F\}^\{3\}T\}\.

This regret bound improves over the𝒪​\(T2/3\)\\mathcal\{O\}\(T^\{2/3\}\)regret bound established inGhaiet al\.\([2022](https://arxiv.org/html/2605.26373#bib.bib46)\)\. Under Assumptions[1](https://arxiv.org/html/2605.26373#Thmassumption1)to[3](https://arxiv.org/html/2605.26373#Thmassumption3), this result forhidden\-convexlosses matches the standard𝒪​\(T\)\\mathcal\{O\}\(\\sqrt\{T\}\)regret bound of OGD in onlineconvexoptimization\.

### 3\.3Proof Sketch of Theorem[1](https://arxiv.org/html/2605.26373#Thmtheorem1)

In this section, we provide a proof sketch of Theorem[1](https://arxiv.org/html/2605.26373#Thmtheorem1)\. We provide a complete proof in Section[A](https://arxiv.org/html/2605.26373#A1)\. The main strategy of the proof is inspired byGhaiet al\.\([2022](https://arxiv.org/html/2605.26373#bib.bib46)\)which shows that OGD applied to non\-convex functions is an approximation of online mirror descent applied to convex functions under reparameterization\. We briefly recall the intuition behind this idea for the convenience of the reader to motivate Assumption[1](https://arxiv.org/html/2605.26373#Thmassumption1)before discussing our new key technical steps driving our tighter analysis\.

Implicit OMD via OGD reparameterization\.In continuous time,Amid and Warmuth \([2020](https://arxiv.org/html/2605.26373#bib.bib58)\)showed that the mirror\-flow trajectory in the hidden space and the gradient\-flow trajectory in the original space coincide exactly under the reparameterization\. More precisely, ifz​\(t\)=q​\(x​\(t\)\)z\(t\)=q\(x\(t\)\), they showed that both trajectories of the following ordinary differential equations:

dd​t​∇R​\(z​\(t\)\)=−∇h​\(z​\(t\)\),d​x​\(t\)d​t=−∇\(h∘q\)⁡\(x​\(t\)\)\\frac\{d\}\{dt\}\\nabla\\mkern\-2\.5muR\(z\(t\)\)=\-\\nabla\\mkern\-2\.5muh\(z\(t\)\)\\,,\\quad\\frac\{dx\(t\)\}\{dt\}=\-\\nabla\\mkern\-2\.5mu\(h\\circ q\)\(x\(t\)\)\(2\)coincide if the mirror descent regularizerRRand the reparameterization functionqqsatisfy\[∇2R​\(q​\(x\)\)\]−1=Jq​\(x\)​Jq​\(x\)⊤\[\\nabla\\mkern\-2\.5mu^\{2\}R\(q\(x\)\)\]^\{\-1\}=J\_\{q\}\(x\)J\_\{q\}\(x\)^\{\\top\}\. In discrete time, OGD iterates on the losses\(ℓt\)\(\\ell\_\{t\}\)are given by \([OGD](https://arxiv.org/html/2605.26373#S2.Ex2)\) and Online Mirror Descent \(OMD\) iterates on the convex losses\(ht\)\(h\_\{t\}\)are defined as follows:

yt\+1=arg⁡miny∈𝒴⁡\{⟨∇ht​\(yt\),y−yt⟩\+1η​DR​\(y∥yt\)\},y\_\{t\+1\}=\\arg\\min\_\{y\\in\{\\mathcal\{Y\}\}\}\\left\\\{\\langle\\nabla\\mkern\-2\.5muh\_\{t\}\(y\_\{t\}\),y\-y\_\{t\}\\rangle\+\\frac\{1\}\{\\eta\}D\_\{R\}\(y\\\|y\_\{t\}\)\\right\\\}\\,,\(OMD\)withy0=q​\(x0\)\.y\_\{0\}=q\(x\_\{0\}\)\\,\.In discrete time, the exact equivalence in the continuous dynamics in \([2](https://arxiv.org/html/2605.26373#S3.E2)\) no longer holds because of the discretization error\. Starting from a coupled stateyt=q​\(xt\)y\_\{t\}=q\(x\_\{t\}\), one step of OGD on\(ℓt\)\(\\ell\_\{t\}\)in the original space producesxt\+1x\_\{t\+1\}and hencezt\+1:=q​\(xt\+1\)z\_\{t\+1\}:=q\(x\_\{t\+1\}\), while one step of OMD on\(ht\)\(h\_\{t\}\)in the hidden space producesyt\+1y\_\{t\+1\}\.Ghaiet al\.\([2022](https://arxiv.org/html/2605.26373#bib.bib46)\)showed that these two hidden\-space iterates areO​\(η3/2\)O\(\\eta^\{3/2\}\)\-close\.

Our key improvement\.To obtain Theorem[1](https://arxiv.org/html/2605.26373#Thmtheorem1), our key contribution is to show a sharperO​\(η2\)O\(\\eta^\{2\}\)estimate of the distance between two hidden\-space iterates, which is small enough to preserve the classicalT\\sqrt\{T\}regret rate after applying a perturbed\-OMD regret bound\.

###### Lemma 1\(One\-step OGD\-OMD coupling\)\.

Suppose Assumptions[1](https://arxiv.org/html/2605.26373#Thmassumption1)\-[3](https://arxiv.org/html/2605.26373#Thmassumption3)hold\. If the ghost OMD iterateyty\_\{t\}is initialized at the image of the OGD iterateztz\_\{t\}, i\.e\.,yt=zt=q​\(xt\)y\_\{t\}=z\_\{t\}=q\(x\_\{t\}\),xt\+1x\_\{t\+1\}is computed via one\-step \([OGD](https://arxiv.org/html/2605.26373#S2.Ex2)\) andyt\+1y\_\{t\+1\}via one\-step \([OMD](https://arxiv.org/html/2605.26373#S3.Ex3)\), then we have‖zt\+1−yt\+1‖2≤6​G5​GF3​η2\.\\\|z\_\{t\+1\}\-y\_\{t\+1\}\\\|\_\{2\}~\\leq~6G^\{5\}G\_\{F\}^\{3\}\\eta^\{2\}\\,\.

The proof of Lemma[1](https://arxiv.org/html/2605.26373#Thmlemma1)is the main new step\. Starting from the coupled stateyt=zt=q​\(xt\)y\_\{t\}=z\_\{t\}=q\(x\_\{t\}\), we first localize bothyt\+1y\_\{t\+1\}andzt\+1z\_\{t\+1\}in anO​\(η\)O\(\\eta\)\-neighborhood𝒴t\{\\mathcal\{Y\}\}\_\{t\}ofyty\_\{t\}\. On this local set, both updates can be expressed as minimizers of the same strongly convex quadratic model, up to different higher\-order perturbations:

yt\+1\\displaystyle y\_\{t\+1\}=arg⁡miny∈𝒴t⁡\{Φt​\(y\)\+εtOMD​\(y\)\},zt\+1=arg⁡miny∈𝒴t⁡\{Φt​\(y\)\+εtOGD​\(y\)\},\\displaystyle=\\arg\\min\_\{y\\in\{\\mathcal\{Y\}\}\_\{t\}\}\\left\\\{\\Phi\_\{t\}\(y\)\+\\varepsilon\_\{t\}^\{\\rm OMD\}\(y\)\\right\\\},\\qquad z\_\{t\+1\}=\\arg\\min\_\{y\\in\{\\mathcal\{Y\}\}\_\{t\}\}\\left\\\{\\Phi\_\{t\}\(y\)\+\\varepsilon\_\{t\}^\{\\rm OGD\}\(y\)\\right\\\},Φt​\(y\):=⟨∇ht​\(yt\),y−yt⟩\+12​η​‖y−yt‖∇2R​\(yt\)2\.\\Phi\_\{t\}\(y\):=\\langle\\nabla\\mkern\-2\.5muh\_\{t\}\(y\_\{t\}\),y\-y\_\{t\}\\rangle\+\\frac\{1\}\{2\\eta\}\\\|y\-y\_\{t\}\\\|\_\{\\nabla\\mkern\-2\.5mu^\{2\}R\(y\_\{t\}\)\}^\{2\}\.HereεtOMD\\varepsilon\_\{t\}^\{\\rm OMD\}is the Taylor remainder obtained by replacingDR​\(y∥yt\)D\_\{R\}\(y\\\|y\_\{t\}\)with its second\-order expansion atyty\_\{t\}, whileεtOGD\\varepsilon\_\{t\}^\{\\rm OGD\}is the Taylor remainder obtained after changing variablesy=q​\(x\)y=q\(x\)in the OGD update and expandingq−1q^\{\-1\}atyty\_\{t\}\. The leading quadratic terms agree exactly because Assumption[1](https://arxiv.org/html/2605.26373#Thmassumption1)identifies the OGD Euclidean metric inxx\-space with the mirror metric∇2R​\(yt\)\\nabla\\mkern\-2\.5mu^\{2\}R\(y\_\{t\}\)inyy\-space\.

The crucial departure fromGhaiet al\.\([2022](https://arxiv.org/html/2605.26373#bib.bib46)\)is how we compare these two perturbed minimization problems\. Rather than controlling the sup\-norm difference between the perturbed objectives and then translating this into a distance between minimizers, we compare the first\-order optimality conditions directly\. SinceΦt\\Phi\_\{t\}is1/η1/\\eta\-strongly convex and the normal cone of𝒴t\{\\mathcal\{Y\}\}\_\{t\}is monotone, Lemma[4](https://arxiv.org/html/2605.26373#Thmlemma4)in App\.[A\.1](https://arxiv.org/html/2605.26373#A1.SS1)gives

‖yt\+1−zt\+1‖≤η​‖∇εtOGD​\(zt\+1\)−∇εtOMD​\(yt\+1\)‖\.\\\|y\_\{t\+1\}\-z\_\{t\+1\}\\\|\\leq\\eta\\left\\\|\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\rm OGD\}\(z\_\{t\+1\}\)\-\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\rm OMD\}\(y\_\{t\+1\}\)\\right\\\|\.Since both iterates lie in an𝒪​\(η\)\\mathcal\{O\}\(\\eta\)\-neighborhood ofyty\_\{t\}, we prove thats‖∇εtOGD​\(zt\+1\)‖=𝒪​\(η\)\\\|\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\rm OGD\}\(z\_\{t\+1\}\)\\\|=\\mathcal\{O\}\(\\eta\)and‖∇εtOMD​\(yt\+1\)‖=𝒪​\(η\)\\\|\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\rm OMD\}\(y\_\{t\+1\}\)\\\|=\\mathcal\{O\}\(\\eta\)\(in Lemma[6](https://arxiv.org/html/2605.26373#Thmlemma6)\-[7](https://arxiv.org/html/2605.26373#Thmlemma7)in App\.[A\.1](https://arxiv.org/html/2605.26373#A1.SS1)\)\. Hence‖yt\+1−zt\+1‖=𝒪​\(η2\)\.\\\|y\_\{t\+1\}\-z\_\{t\+1\}\\\|=\\mathcal\{O\}\(\\eta^\{2\}\)\\,\.This sharper one\-step coupling improves the𝒪​\(η3/2\)\\mathcal\{O\}\(\\eta^\{3/2\}\)discrepancy used inGhaiet al\.\([2022](https://arxiv.org/html/2605.26373#bib.bib46)\)to𝒪​\(η2\)\\mathcal\{O\}\(\\eta^\{2\}\)\. Finally, the mapped OGD iterateszt=q​\(xt\)z\_\{t\}=q\(x\_\{t\}\)can be viewed as OMD iterates perturbed by errorsrt\+1=zt\+1−yt\+1r\_\{t\+1\}=z\_\{t\+1\}\-y\_\{t\+1\}s\.t\.‖rt\+1‖≤Cη:=6​G5​GF3​η2\\\|r\_\{t\+1\}\\\|\\leq C\_\{\\eta\}:=6G^\{5\}G\_\{F\}^\{3\}\\eta^\{2\}\. Applying the perturbed\-OMD regret bound \(Lemma[8](https://arxiv.org/html/2605.26373#Thmlemma8)\) yields:

RT≤Cη​T​Gη\+D1η\+η​G2​T2,R\_\{T\}\\leq\\frac\{C\_\{\\eta\}TG\}\{\\eta\}\+\\frac\{D\_\{1\}\}\{\\eta\}\+\\frac\{\\eta G^\{2\}T\}\{2\}\\,,which concludes the proof by optimizing the step sizeη\\eta\. See App\.[A](https://arxiv.org/html/2605.26373#A1)for the full proof\.

## 4Discussion of the Hessian\-Compatibility Hidden Geometry Assumption[1](https://arxiv.org/html/2605.26373#Thmassumption1)

Our analysis relies on Assumption[1](https://arxiv.org/html/2605.26373#Thmassumption1)which allows to obtain an implicit OMD reparameterization and exploit the connection with OMD analysis over convex functions to obtain our tight regret analysis\. In this section, we discuss this assumption in more depth\. First, we answer an open question inGhaiet al\.\([2022](https://arxiv.org/html/2605.26373#bib.bib46)\)regarding whether we can relax their sufficient diagonal Jacobian condition \(Assumption 4 therein\) to verify Assumption[1](https://arxiv.org/html/2605.26373#Thmassumption1)\. In section[4\.1](https://arxiv.org/html/2605.26373#S4.SS1), we provide necessary and sufficient conditions under which Assumption[1](https://arxiv.org/html/2605.26373#Thmassumption1)is satisfied\. Second, we address the question of the necessity of this Hessian\-compatibility hidden\-geometry assumption, which is also mentioned as an open question inGhaiet al\.\([2022](https://arxiv.org/html/2605.26373#bib.bib46)\)\. In Section[4\.2](https://arxiv.org/html/2605.26373#S4.SS2), we show that Assumption[1](https://arxiv.org/html/2605.26373#Thmassumption1)is necessary to obtain sublinear regret for OGD\.

### 4\.1Relaxation of the diagonal Jacobian assumption inGhaiet al\.\([2022](https://arxiv.org/html/2605.26373#bib.bib46)\)

Define the inverse metric fieldM:𝒴→ℝd×dM:\{\\mathcal\{Y\}\}\\to\\mathbb\{R\}^\{d\\times d\}by

M​\(z\):=\[Jq​\(q−1​\(z\)\)​Jq​\(q−1​\(z\)\)⊤\]−1\.M\(z\):=\\left\[J\_\{q\}\(q^\{\-1\}\(z\)\)J\_\{q\}\(q^\{\-1\}\(z\)\)^\{\\top\}\\right\]^\{\-1\}\.\(3\)Recall that Assumption[1](https://arxiv.org/html/2605.26373#Thmassumption1)states that there exists a scalar functionR:𝒴→ℝR:\{\\mathcal\{Y\}\}\\to\\mathbb\{R\}s\.t\.M​\(z\)=∇2R​\(z\),M\(z\)=\\nabla\\mkern\-2\.5mu^\{2\}R\(z\),for allz∈𝒴\.z\\in\{\\mathcal\{Y\}\}\.Theorem 7 inGhaiet al\.\([2022](https://arxiv.org/html/2605.26373#bib.bib46)\)showed that ifJqJ\_\{q\}is diagonal then there exists suchRR\. They used this diagonal separable structure to construct a regularizerRRexplicitly as a solution to second\-order ODEs with coefficients depending on the parameterization functionqqand its derivatives\. In the following result, we relax this condition and provide more general conditions under which the matrix fieldMMis a Hessian field\.

###### Proposition 1\. Ifqqis aC3C^\{3\}diffeomorphism and the matrix fieldMMdefined in \([3](https://arxiv.org/html/2605.26373#S4.E3)\) satisfies:∂zkMi​j​\(z\)=∂zjMi​k​\(z\),∀i,j,k∈\[d\],∀z∈𝒴,\\partial\_\{z\_\{k\}\}M\_\{ij\}\(z\)=\\partial\_\{z\_\{j\}\}M\_\{ik\}\(z\)\\,,\\quad\\forall i,j,k\\in\[d\],\\quad\\forall z\\in\{\\mathcal\{Y\}\}\\,,where𝒴\{\\mathcal\{Y\}\}is a convex domain, then there exists a scalar functionRRs\.t\.M=∇2RM=\\nabla\\mkern\-2\.5mu^\{2\}R, which is unique up to an affine function\. Moreover, the functionRRis also strongly convex on𝒴\{\\mathcal\{Y\}\}\.

This result follows as a consequence of the fundamental result from differential analysis stating that a vector field with matching cross\-partial coordinate function derivatives is a conservative field, i\.e\., a gradient field deriving from a scalar potential\. See Proposition[3](https://arxiv.org/html/2605.26373#Thmproposition3)for a precise statement of that result and a complete proof of Proposition[1](https://arxiv.org/html/2605.26373#Thmproposition1)\.

We provide two examples where the Jacobian of the parameterizationqqis not diagonal in general, yet the matrix fieldMMdefined in \([3](https://arxiv.org/html/2605.26373#S4.E3)\) satisfies the cross\-derivatives conditions∂zkMi​j=∂zjMi​k\\partial\_\{z\_\{k\}\}M\_\{ij\}=\\partial\_\{z\_\{j\}\}M\_\{ik\}of Proposition[1](https://arxiv.org/html/2605.26373#Thmproposition1)\. Moreover, the compatibility Assumption[1](https://arxiv.org/html/2605.26373#Thmassumption1)is satisfied with explicit regularizersRR\. Detailed derivations for both examples can be found in App\.[B\.2](https://arxiv.org/html/2605.26373#A2.SS2)\.

###### Example 1\(Affine mixing of a separable parameterization\)\.

Letq​\(x\)=A​s​\(x\)\+bq\(x\)=As\(x\)\+bands​\(x\)=\(s1​\(x1\),…,sd​\(xd\)\)s\(x\)=\\bigl\(s\_\{1\}\(x\_\{1\}\),\\dots,s\_\{d\}\(x\_\{d\}\)\\bigr\)whereA∈ℝd×dA\\in\\mathbb\{R\}^\{d\\times d\}is invertible,b∈ℝdb\\in\\mathbb\{R\}^\{d\}, and eachsi:ℝ→ℝs\_\{i\}:\\mathbb\{R\}\\to\\mathbb\{R\}is aC2C^\{2\}diffeomorphism\. ThenJq​\(x\)=A​D​\(x\)J\_\{q\}\(x\)=A\\,D\(x\)withD​\(x\):=diag⁡\(s1′​\(x1\),…,sd′​\(xd\)\)D\(x\):=\\operatorname\{diag\}\\bigl\(s\_\{1\}^\{\\prime\}\(x\_\{1\}\),\\dots,s\_\{d\}^\{\\prime\}\(x\_\{d\}\)\\bigr\)\. HenceM​\(z\)=A−T​D​\(x\)−2​A−1,z=q​\(x\)\.M\(z\)=A^\{\-T\}D\(x\)^\{\-2\}A^\{\-1\}\\,,z=q\(x\)\.Moreover,MMis explicitly the Hessian of the following regularizer:R​\(z\):=∑i=1dri​\(\(A−1​\(z−b\)\)i\)R\(z\):=\\sum\_\{i=1\}^\{d\}r\_\{i\}\\bigl\(\(A^\{\-1\}\(z\-b\)\)\_\{i\}\\bigr\)andri′′​\(t\)=\(si′​\(si−1​\(t\)\)\)−2\.r\_\{i\}^\{\\prime\\prime\}\(t\)=\\bigl\(s\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\-1\}\(t\)\)\\bigr\)^\{\-2\}\\,\.For instance, ifsi​\(ui\)=exp⁡\(ui\)s\_\{i\}\(u\_\{i\}\)=\\exp\(u\_\{i\}\)for alli∈\[d\],i\\in\[d\],thenri​\(t\)=−log⁡tr\_\{i\}\(t\)=\-\\log tsatisfies the conditions\.

###### Example 2\(Rank\-one nonlinear mixing\)\.

Leta∈ℝda\\in\\mathbb\{R\}^\{d\}and letq​\(x\)=x\+h​\(a⊤​x\)​aq\(x\)=x\+h\(a^\{\\top\}x\)\\,awhereh:ℝ→ℝh:\\mathbb\{R\}\\to\\mathbb\{R\}isC2C^\{2\}\. We haveJq​\(x\)=I\+h′​\(a⊤​x\)​a​a⊤\.J\_\{q\}\(x\)=I\+h^\{\\prime\}\(a^\{\\top\}x\)\\,aa^\{\\top\}\.This Jacobian is typically non\-diagonal whenever at least two components ofaaare nonzero\.

### 4\.2OGD can incur linear regret without Hessian\-compatible hidden geometry

In this section, we discuss how crucial is Assumption[1](https://arxiv.org/html/2605.26373#Thmassumption1)to obtain sublinear regret for OGD in our nonconvex online learning setting\. Perhaps surprisingly, we show that OGD can incur linear regret if the Hessian compatibility assumption does not hold; i\.e\., if the metric induced by the symmetric positive definite matrixMMin \([3](https://arxiv.org/html/2605.26373#S4.E3)\) does not derive from a Hessian metric such thatM​\(z\)=∇2R​\(z\)M\(z\)=\\nabla\\mkern\-2\.5mu^\{2\}R\(z\)for some regularization functionRR\.

###### Theorem 2\(Linear regret without Hessian\-compatible geometry\)\. There exist a smooth reparameterizationqqwith uniformly positive definite induced metricJq​\(x\)​Jq​\(x\)⊤⪰μ​IJ\_\{q\}\(x\)J\_\{q\}\(x\)^\{\\top\}\\succeq\\mu Iwhose inverse metric field is not the Hessian of any regularizer, and an adaptive sequence of time\-varying convex hidden losseshth\_\{t\}such that OGD run on the original lossesℓt=ht∘q\\ell\_\{t\}=h\_\{t\}\\circ qsuffers linear regret:RT=Ω​\(T\)R\_\{T\}=\\Omega\(T\)\.

The theorem shows that smoothness and uniform positive definiteness of the induced hidden geometry are not sufficient for sublinear regret of OGD\. This means that Assumption[1](https://arxiv.org/html/2605.26373#Thmassumption1)is not merely a technical condition for our analysis of OGD in Theorem[1](https://arxiv.org/html/2605.26373#Thmtheorem1)\. The obstruction is geometric: when the metric induced by the reparameterization is not Hessian\-compatible, OGD may fail to achieve sublinear regret even when the reparameterization is smooth and its induced metric is uniformly positive definite\. The hidden\-space dynamics of OGD need not arise from an integrable potential\. An adversary can then exploit this fact by forcing the iterates to cycle around a loop with nonzero circulation and force linear regret\. We elaborate more on this in the next proof sketch\.

Proof sketch of Theorem[2](https://arxiv.org/html/2605.26373#Thmtheorem2)\.We provide a brief proof sketch of the result to highlight the intuitions\. An extended proof sketch with all the steps can be found in App\.[B\.4](https://arxiv.org/html/2605.26373#A2.SS4)and the full proof is provided in App\.[B\.5](https://arxiv.org/html/2605.26373#A2.SS5)\. The construction uses the two\-dimensional reparameterizationq​\(x1,x2\)=ex1​\(cos⁡x2,sin⁡x2\)q\(x\_\{1\},x\_\{2\}\)=e^\{x\_\{1\}\}\(\\cos x\_\{2\},\\sin x\_\{2\}\)on a compact rectangular domain\. It can be easily verified that its induced metric is uniformly positive definite, but the corresponding inverse metric in hidden coordinates isM​\(z\)=\[Jq​\(x\)​Jq​\(x\)⊤\]−1=1‖z‖2​I2M\(z\)=\[J\_\{q\}\(x\)J\_\{q\}\(x\)^\{\\top\}\]^\{\-1\}=\\frac\{1\}\{\\\|z\\\|^\{2\}\}I\_\{2\}wherez=q​\(x\)z=q\(x\), which is not Hessian\-compatible in the sense needed for the OGD\-OMD equivalence\. We consider linear hidden lossesℓt​\(x\)=ht​\(q​\(x\)\)\\ell\_\{t\}\(x\)=h\_\{t\}\(q\(x\)\)withht​\(z\)=⟨st,z⟩h\_\{t\}\(z\)=\\langle s\_\{t\},z\\rangle\. Writingzt=q​\(xt\)z\_\{t\}=q\(x\_\{t\}\), the OGD update in the original coordinates induces the hidden\-space dynamics:

zt\+1=zt−η​M​\(zt\)−1​st\+O​\(η2\)\.z\_\{t\+1\}=z\_\{t\}\-\\eta M\(z\_\{t\}\)^\{\-1\}s\_\{t\}\+O\(\\eta^\{2\}\)\.\(4\)The adversary choosesst=−M​\(zt\)​vts\_\{t\}=\-M\(z\_\{t\}\)v\_\{t\}, wherevtv\_\{t\}is a unit coordinate direction\. Thus the OGD iterates in the hidden space satisfy:

zt\+1=zt\+η​vt\+O​\(η2\),z\_\{t\+1\}=z\_\{t\}\+\\eta v\_\{t\}\+O\(\\eta^\{2\}\),so the adversary can make the trajectory approximately follow any axis\-aligned polygonal path\. Fixing a comparatoru∈𝒳u\\in\{\\mathcal\{X\}\}and writingzu=q​\(u\)∈𝒴z\_\{u\}=q\(u\)\\in\{\\mathcal\{Y\}\}, the instantaneous regret can be expressed as:

ℓt​\(xt\)−ℓt​\(u\)=−⟨Fu​\(zt\),vt⟩,Fu​\(z\):=M​\(z\)​\(z−zu\)\.\\ell\_\{t\}\(x\_\{t\}\)\-\\ell\_\{t\}\(u\)=\-\\langle F\_\{u\}\(z\_\{t\}\),v\_\{t\}\\rangle,\\qquad F\_\{u\}\(z\):=M\(z\)\(z\-z\_\{u\}\)\.Using the hidden\-space dynamics \([4](https://arxiv.org/html/2605.26373#S4.E4)\), it follows that:

ℓt​\(xt\)−ℓt​\(u\)=−1η​⟨Fu​\(zt\),zt\+1−zt⟩\+O​\(η\)\.\\ell\_\{t\}\(x\_\{t\}\)\-\\ell\_\{t\}\(u\)=\-\\frac\{1\}\{\\eta\}\\langle F\_\{u\}\(z\_\{t\}\),z\_\{t\+1\}\-z\_\{t\}\\rangle\+O\(\\eta\)\.Hence cumulative regret along a trajectory is, up to lower\-order terms, a rescaled discrete line integral of the vector fieldFuF\_\{u\}\. For an appropriate comparatoruu, the vector fieldFuF\_\{u\}is not conservative\. Equivalently, its curl is negative on a small rectangleRRin hidden space, so the counter\-clockwise circulation satisfies∮∂RFu​\(z\)⋅𝑑z<0\.\\oint\_\{\\partial R\}F\_\{u\}\(z\)\\cdot dz<0\.The adversary chooses the directionsvt∈\{±e1,±e2\}v\_\{t\}\\in\\\{\\pm e\_\{1\},\\pm e\_\{2\}\\\}\(wheree1=\(1,0\)e\_\{1\}=\(1,0\)ande2=\(0,1\)e\_\{2\}=\(0,1\)\) so that the hidden iterates move once around this rectangle\. The Riemann\-sum approximation of the line integral then gives regret of order1/η1/\\etaover one cycle, while the cycle itself lastsΘ​\(1/η\)\\Theta\(1/\\eta\)rounds\. Thus each cycle contributes a positive constant amount of regret per round\. Repeating the same cycle for the horizonTTyieldsRT=Ω​\(T\)R\_\{T\}=\\Omega\(T\)\. The full construction and error estimates are given in App\.[B\.3](https://arxiv.org/html/2605.26373#A2.SS3)\.

Open question\.The lower bound of Theorem[2](https://arxiv.org/html/2605.26373#Thmtheorem2)only holds for OGD\. Whether Assumption[1](https://arxiv.org/html/2605.26373#Thmassumption1)is necessary to obtain sublinear regret in general, foranyonline algorithm, remains open\.

## 5Regret Analysis of OGD for OHCO under Bandit Feedback

### 5\.1Expected Regret Bound in the Bandit Setting

We now consider the one\-point bandit feedback, where the learner observes only one function value per round\. We analyze the classical Bandit OGD \(BOGD\) algorithm with spherical smoothing \(Algorithm[1](https://arxiv.org/html/2605.26373#alg1),Flaxmanet al\.\([2005](https://arxiv.org/html/2605.26373#bib.bib43)\)\) in the hidden\-convex setting, against an oblivious adversary\. LetUs​pU\_\{sp\}denote the uniform distribution over the Euclidean unit sphere\. For an exploration radiusδ\>0\\delta\>0, the algorithm maintains its iterates in the shrunken domain𝒳δ:=\(1−δ\)​𝒳\{\\mathcal\{X\}\}\_\{\\delta\}:=\(1\-\\delta\)\{\\mathcal\{X\}\}, so that the perturbed query points remain feasible \(under the standard normalization0∈𝒳0\\in\{\\mathcal\{X\}\}andB2⊆𝒳B\_\{2\}\\subseteq\{\\mathcal\{X\}\}as inFlaxmanet al\.\([2005](https://arxiv.org/html/2605.26373#bib.bib43)\)\)\. We denote by𝒴δ:=q​\(𝒳δ\)\{\\mathcal\{Y\}\}\_\{\\delta\}:=q\(\{\\mathcal\{X\}\}\_\{\\delta\}\)the corresponding hidden\-space domain, which is only used in the analysis to compare the mapped BOGD iterates with mirror descent on the convex losses\.

Algorithm 1Bandit OGD \(BOGD\)\(Flaxmanet al\.,[2005](https://arxiv.org/html/2605.26373#bib.bib43)\)1:Input:

x1∈𝒳δx\_\{1\}\\in\{\\mathcal\{X\}\}\_\{\\delta\}
2:for

t=1,…,Tt=1,\\ldots,Tdo

3:Sample a random direction

ζt∼Us​p\\zeta\_\{t\}\\sim U\_\{sp\}
4:Choose

x^t=xt\+δ​ζt\\hat\{x\}\_\{t\}=\{\\color\[rgb\]\{0,0,1\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0,1\}x\_\{t\}\}\+\\delta\\zeta\_\{t\}
5:Observe the bandit loss

ℓt​\(x^t\)\\ell\_\{t\}\(\\hat\{x\}\_\{t\}\)
6:Construct

gt=dδ​ℓt​\(x^t\)​ζt\{\\color\[rgb\]\{0\.29296875,0\.8046875,0\.29296875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.29296875,0\.8046875,0\.29296875\}g\_\{t\}\}=\\frac\{d\}\{\\delta\}\\ell\_\{t\}\(\\hat\{x\}\_\{t\}\)\\zeta\_\{t\}
7:Update

xt\+1=Π𝒳δ​\(xt−η​gt\)x\_\{t\+1\}=\\Pi\_\{\{\\mathcal\{X\}\}\_\{\\delta\}\}\(\{\\color\[rgb\]\{0,0,1\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0,1\}x\_\{t\}\}\-\\eta\{\\color\[rgb\]\{0\.29296875,0\.8046875,0\.29296875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.29296875,0\.8046875,0\.29296875\}g\_\{t\}\}\)
8:endfor

Before stating the result, we introduce a few extra standard assumptions on the smoothness and boundedness of the losses\(ℓt\)\(\\ell\_\{t\}\), and the diameter of𝒳\{\\mathcal\{X\}\}\.

###### Assumption 4\(Loss boundedness\)\.

There existsM\>0M\>0s\.t\.\|ℓt​\(x\)\|≤M\|\\ell\_\{t\}\(x\)\|\\leq M, for allx∈𝒳,t≥1x\\in\{\\mathcal\{X\}\},t~\\geq~1\.

###### Assumption 5\(Smoothness\)\.

For anytt, the loss functionℓt\\ell\_\{t\}is smooth, i\.e\., there exists a constantH\>0H\>0, such that for anyx,x′∈𝒳x,x^\{\\prime\}\\in\{\\mathcal\{X\}\}and anyt≥1t~\\geq~1,‖∇ℓt​\(x\)−∇ℓt​\(x′\)‖≤H​‖x−x′‖\.\\\|\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\(x\)\-\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\(x^\{\\prime\}\)\\\|~\\leq~H\\\|x\-x^\{\\prime\}\\\|\.

###### Assumption 6\(Boundedness of𝒳\{\\mathcal\{X\}\}\)\.

There exists a constantD\>0D\>0, such that for anyx∈𝒳x\\in\{\\mathcal\{X\}\}, it holds that‖x‖2≤D\\\|x\\\|\_\{2\}\\leq D, and for anyx1,x2∈𝒳x\_\{1\},x\_\{2\}\\in\{\\mathcal\{X\}\}, it holds that‖x1−x2‖2≤D\\\|x\_\{1\}\-x\_\{2\}\\\|\_\{2\}\\leq D\.

In the following theorem, we present the no\-regret guarantees of Bandit OGD\.

###### Theorem 3\(Expected regret under bandit feedback\)\. Let Assumptions[1](https://arxiv.org/html/2605.26373#Thmassumption1)to[6](https://arxiv.org/html/2605.26373#Thmassumption6)hold\. Then, the expected regret of BOGD \(Algorithm[1](https://arxiv.org/html/2605.26373#alg1)\) is bounded as follows:𝔼​\[RT\]≤d​M​G3​\(GF​D\+G2​H​D\+GF\)​\(3​D11/2\+D13/2\)⋅T3/4,\\mathbb\{E\}\[\\text\{R\}\_\{T\}\]~\\leq~\\sqrt\{dMG^\{3\}\(G\_\{F\}D\+G^\{2\}HD\+G\_\{F\}\)\(3D\_\{1\}^\{1/2\}\+D\_\{1\}^\{3/2\}\)\}\\,\\cdot T^\{3/4\}\\,,when using the stepsizeη=\(4​D13K12​K2​T3\)1/4\\eta=\\left\(\\frac\{4D\_\{1\}^\{3\}\}\{K\_\{1\}^\{2\}K\_\{2\}T^\{3\}\}\\right\)^\{1/4\}, and the exploration radiusδ=\(4​K2​D1K12​T\)1/4\\delta=\\left\(\\frac\{4K\_\{2\}D\_\{1\}\}\{K\_\{1\}^\{2\}T\}\\right\)^\{1/4\}, whereK1:=GF​D\+G2​H​D\+GFK\_\{1\}:=G\_\{F\}D\+G^\{2\}HD\+G\_\{F\}andK2:=6​d2​G6​M2K\_\{2\}:=6d^\{2\}G^\{6\}M^\{2\}\.

Our result implies that, despite hidden convexity, OGD with spherical smoothing achieves the same best known regret rate as in the convex bandit case\(Flaxmanet al\.,[2005](https://arxiv.org/html/2605.26373#bib.bib43)\)\.

### 5\.2Proof Sketch of Theorem[3](https://arxiv.org/html/2605.26373#Thmtheorem3)

We give a proof sketch of Theorem[3](https://arxiv.org/html/2605.26373#Thmtheorem3); the complete proof is in App\.[C](https://arxiv.org/html/2605.26373#A3)\. The proof follows the same high\-level strategy as in the exact gradient setting with a few similar steps to the analysis of OBGD in OCO: after mapping the BOGD iterates to the hidden space,zt=q​\(xt\)z\_\{t\}=q\(x\_\{t\}\), we compare them to mirror\-descent steps on the convex losseshth\_\{t\}\. The bandit setting introduces three additional errors: the domain must be shrunk to allow perturbations, the one\-point estimator is biased for the true gradient, and the algorithm suffers the loss at the perturbed pointx^t=xt\+δ​ζt\\hat\{x\}\_\{t\}=x\_\{t\}\+\\delta\\zeta\_\{t\}, rather than atxtx\_\{t\}\. To define the ghost OMD iterate in the hidden space \(which is only used in the analysis and never computed\), we introduce the hidden\-space gradient estimatorg~t:=Jq​\(xt\)−⊤​gt\\tilde\{g\}\_\{t\}:=J\_\{q\}\(x\_\{t\}\)^\{\-\\top\}g\_\{t\}which is the natural hidden\-space analogue of the bandit estimatorgtg\_\{t\}by the chain rule:∇ht​\(zt\)=Jq​\(xt\)−⊤​∇ℓt​\(xt\),zt=q​\(xt\)\.\\nabla\\mkern\-2\.5muh\_\{t\}\(z\_\{t\}\)=J\_\{q\}\(x\_\{t\}\)^\{\-\\top\}\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\(x\_\{t\}\),z\_\{t\}=q\(x\_\{t\}\)\.Note that it is not computed by the algorithm but introduced only to analyze the mapped BOGD trajectory as a perturbed OMD trajectory in𝒴δ=q​\(𝒳δ\)\{\\mathcal\{Y\}\}\_\{\\delta\}=q\(\{\\mathcal\{X\}\}\_\{\\delta\}\)\. Since\|ℓt\|≤M\|\\ell\_\{t\}\|\\leq Mand‖ζt‖=1\\\|\\zeta\_\{t\}\\\|=1, we have‖gt‖≤d​M/δ\\\|g\_\{t\}\\\|~\\leq~dM/\\delta\. Letz⋆∈arg⁡minz∈𝒴​∑t=1Tht​\(z\)z^\{\\star\}\\in\\arg\\min\_\{z\\in\{\\mathcal\{Y\}\}\}\\sum\_\{t=1\}^\{T\}h\_\{t\}\(z\)andzδ⋆∈arg⁡minz∈𝒴δ​∑t=1Tht​\(z\)z^\{\\star\}\_\{\\delta\}\\in\\arg\\min\_\{z\\in\{\\mathcal\{Y\}\}\_\{\\delta\}\}\\sum\_\{t=1\}^\{T\}h\_\{t\}\(z\)be two comparators, and letz^t=q​\(x^t\)\\hat\{z\}\_\{t\}=q\(\\hat\{x\}\_\{t\}\)\. The starting point is the following regret decomposition:

𝔼​\[RT\]≤𝔼​\[∑t=1T⟨g~t,zt−zδ⋆⟩\]⏟:RT\(1\)perturbed OMD regretfor random linear losses\+∑t=1Tht​\(zδ⋆\)−ht​\(z⋆\)⏟:RT\(2\)domainshrinkage error\+𝔼​\[∑t=1TBt\]⏟RT\(3\): gradientestimation bias\+𝔼​\[∑t=1Tht​\(z^t\)−ht​\(zt\)\]⏟RT\(4\):smoothing error,\\mathbb\{E\}\[R\_\{T\}\]\\leq\\underbrace\{\\mathbb\{E\}\\left\[\\sum\_\{t=1\}^\{T\}\\langle\\tilde\{g\}\_\{t\},z\_\{t\}\-z^\{\\star\}\_\{\\delta\}\\rangle\\right\]\}\_\{\\text\{\\noindent\\hbox\{\}\\hfill\{\{\\hbox\{\\begin\{tabular\}\[c\]\{@\{\}c@\{\}\}$R\_\{T\}^\{\(1\)\}:$ perturbed OMD regret\\\\ for random linear losses\\end\{tabular\}\}\}\}\\hfill\\hbox\{\}\}\}\+\\underbrace\{\\sum\_\{t=1\}^\{T\}h\_\{t\}\(z^\{\\star\}\_\{\\delta\}\)\-h\_\{t\}\(z^\{\\star\}\)\}\_\{\\text\{\\noindent\\hbox\{\}\\hfill\{\{\\hbox\{\\begin\{tabular\}\[c\]\{@\{\}c@\{\}\}$R\_\{T\}^\{\(2\)\}:$ domain\\\\ shrinkage error\\end\{tabular\}\}\}\}\\hfill\\hbox\{\}\}\}\+\\underbrace\{\\mathbb\{E\}\\left\[\\sum\_\{t=1\}^\{T\}B\_\{t\}\\right\]\}\_\{\\text\{\\noindent\\hbox\{\}\\hfill\{\{\\hbox\{\\begin\{tabular\}\[c\]\{@\{\}c@\{\}\}$R\_\{T\}^\{\(3\)\}$: gradient\\\\ estimation bias\\end\{tabular\}\}\}\}\\hfill\\hbox\{\}\}\}\+\\underbrace\{\\mathbb\{E\}\\left\[\\sum\_\{t=1\}^\{T\}h\_\{t\}\(\\hat\{z\}\_\{t\}\)\-h\_\{t\}\(z\_\{t\}\)\\right\]\}\_\{R\_\{T\}^\{\(4\)\}:\\ \\text\{smoothing error\}\},whereBt:=⟨bt,zδ⋆−zt⟩B\_\{t\}:=\\langle b\_\{t\},z^\{\\star\}\_\{\\delta\}\-z\_\{t\}\\rangleandbt:=𝔼​\[g~t∣ℱt−1\]−∇ht​\(zt\)\.b\_\{t\}:=\\mathbb\{E\}\[\\tilde\{g\}\_\{t\}\\mid\\mathcal\{F\}\_\{t\-1\}\]\-\\nabla\\mkern\-2\.5muh\_\{t\}\(z\_\{t\}\)\\,\.The four terms isolate the distinct sources of error in the bandit setting\. The shrinkage term is controlled by comparing optimization over𝒳\{\\mathcal\{X\}\}and𝒳δ=\(1−δ\)​𝒳\{\\mathcal\{X\}\}\_\{\\delta\}=\(1\-\\delta\)\{\\mathcal\{X\}\}, givingRT\(2\)≤δ​L​D​TR\_\{T\}^\{\(2\)\}\\leq\\delta LDT\. The smoothing term is controlled directly by Lipschitzness ofℓt\\ell\_\{t\}:\|RT\(4\)\|≤δ​L​T\.\|R\_\{T\}^\{\(4\)\}\|~\\leq~\\delta LT\.The bias term comes from the fact that the one\-point spherical estimator is unbiased for the gradient of the ball\-smoothed loss, not for∇ℓt​\(xt\)\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\(x\_\{t\}\)\. Using smoothness ofℓt\\ell\_\{t\}and the chain\-rule transformation above yields‖bt‖≤δ​G​H\\\|b\_\{t\}\\\|\\leq\\delta GHand‖zt−zδ⋆‖≤G​D\\\|z\_\{t\}\-z^\{\\star\}\_\{\\delta\}\\\|\\leq GD, and hence\|RT\(3\)\|≤δ​G2​H​D​T\.\|R\_\{T\}^\{\(3\)\}\|\\leq\\delta G^\{2\}HDT\\,\.It remains to controlRT\(1\)R\_\{T\}^\{\(1\)\}, which is the term most closely related to the exact gradient setting proof\. For each round, define the ghost OMD iterate, only used in the analysis:

yt\+1=arg⁡miny∈𝒴δ⁡\{⟨g~t,y−zt⟩\+1η​DR​\(y∥zt\)\}\.y\_\{t\+1\}=\\arg\\min\_\{y\\in\{\\mathcal\{Y\}\}\_\{\\delta\}\}\\left\\\{\\langle\\tilde\{g\}\_\{t\},y\-z\_\{t\}\\rangle\+\\frac\{1\}\{\\eta\}D\_\{R\}\(y\\\|z\_\{t\}\)\\right\\\}\.The actual mapped BOGD iterate satisfieszt\+1=q​\(xt\+1\)=yt\+1\+rt\+1z\_\{t\+1\}=q\(x\_\{t\+1\}\)=y\_\{t\+1\}\+r\_\{t\+1\}\. As in the exact gradient feedback setting, the OGD–OMD one\-step coupling applies pathwise because the proof only requires a bounded vectorgtg\_\{t\}, not an exact gradient\. WithG^F=d​M/δ\\hat\{G\}\_\{F\}=dM/\\delta, Corollary[5](https://arxiv.org/html/2605.26373#Thmtheorem5)gives‖rt\+1‖≤C^η:=G5​\(5​G^F2\+G^F3​η\)​η2\.\\\|r\_\{t\+1\}\\\|~\\leq~\\hat\{C\}\_\{\\eta\}:=G^\{5\}\\left\(5\\hat\{G\}\_\{F\}^\{2\}\+\\hat\{G\}\_\{F\}^\{3\}\\eta\\right\)\\eta^\{2\}\.Similarly to the exact setting, the perturbed OMD bound for the random linear losses⟨g~t,⋅⟩\\langle\\tilde\{g\}\_\{t\},\\cdot\\ranglethen givesRT\(1\)≤D1η\+η​T2​G^F2\+G​C^η​Tη,R\_\{T\}^\{\(1\)\}\\leq\\frac\{D\_\{1\}\}\{\\eta\}\+\\frac\{\\eta T\}\{2\}\\hat\{G\}\_\{F\}^\{2\}\+\\frac\{G\\hat\{C\}\_\{\\eta\}T\}\{\\eta\}\\,,whereG^F2\\hat\{G\}\_\{F\}^\{2\}andC^η\\hat\{C\}\_\{\\eta\}replaceGFG\_\{F\}andCηC\_\{\\eta\}\. Combining the four bounds yields:𝔼​\[RT\]≤K1​δ​T\+D1η\+K2​η​Tδ2\+K3​η2​Tδ3\.\\mathbb\{E\}\[R\_\{T\}\]\\leq K\_\{1\}\\delta T\+\\frac\{D\_\{1\}\}\{\\eta\}\+K\_\{2\}\\frac\{\\eta T\}\{\\delta^\{2\}\}\+K\_\{3\}\\frac\{\\eta^\{2\}T\}\{\\delta^\{3\}\}\\,\.Balancing the first three terms gives the values ofη,δ\\eta,\\deltareported in Theorem[3](https://arxiv.org/html/2605.26373#Thmtheorem3)and the desired𝒪​\(T3/4\)\\mathcal\{O\}\(T^\{3/4\}\)bound follows\.

## 6Conclusion and Future Work

In this paper, we studied online nonconvex learning for a structured class of hidden\-convex losses\. We showed that OGD achieves𝒪​\(T\)\\mathcal\{O\}\(\\sqrt\{T\}\)regret in the exact\-gradient setting, sharpened the geometric conditions required for the underlying algorithmic equivalence, and proved that OGD can suffer linear regret when these conditions fail\. We also extended the framework to one\-point bandit feedback, establishing sublinear expected regret guarantees against oblivious adversaries\.

Several questions remain open\. First, our lower bound shows that OGD can suffer linear regret without Hessian compatibility, but it does not rule out sublinear regret for other algorithms\. An important direction is therefore to determine whether online hidden\-convex optimization admits sublinear, or even𝒪​\(T\)\\mathcal\{O\}\(\\sqrt\{T\}\), regret without this geometric assumption\. Second, our bandit guarantees leave open the possibility of sharper rates, high\-probability bounds, and algorithms matching the best\-known guarantees from bandit online convex optimization\. More broadly, it would be interesting to understand for which classes of hidden\-convex functions online algorithms incur a regret that matches that of their convex counterparts\.

## Acknowledgements

Ioannis Panageas is supported by NSF grant CCF\- 2454115\. This work is supported by the MOE Tier 2 Grant \(MOE\-T2EP20223\-0018\), the CQT\+\+ Core Research Funding Grant \(SUTD\) \(RS\-NRCQT\-00002\), the National Research Foundation Singapore and DSO National Laboratories under the AI Singapore Programme \(Award Number: AISG2\-RP\-2020\-016\), and partially by Project MIS 5154714 of the National Recovery and Resilience Plan, Greece 2\.0, funded by the European Union under the NextGenerationEU Program\.

## References

- Learning in non\-convex games with an optimization oracle\.InConference on Learning Theory,pp\. 18–29\.Cited by:[§1\.2](https://arxiv.org/html/2605.26373#S1.SS2.p1.1),[§1](https://arxiv.org/html/2605.26373#S1.p1.1)\.
- E\. Amid, M\. K\. Warmuth, R\. Anil, and T\. Koren \(2019\)Robust bi\-tempered logistic loss based on bregman divergences\.Advances in Neural Information Processing Systems32\.Cited by:[§3\.1](https://arxiv.org/html/2605.26373#S3.SS1.p2.1)\.
- E\. Amid and M\. K\. Warmuth \(2020\)Reparameterizing mirror descent as gradient descent\.Advances in Neural Information Processing Systems33,pp\. 8430–8439\.Cited by:[§1\.2](https://arxiv.org/html/2605.26373#S1.SS2.p1.1),[§3\.3](https://arxiv.org/html/2605.26373#S3.SS3.p2.1)\.
- A\. Barakat, S\. Chakraborty, P\. Yu, P\. Tokekar, and A\. S\. Bedi \(2025\)On the global optimality of policy gradient methods in general utility reinforcement learning\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,Cited by:[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- A\. Barakat, I\. Fatkhullin, and N\. He \(2023\)Reinforcement learning with general utilities: simpler variance reduction and large state\-action space\.InInternational Conference on Machine Learning,Cited by:[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- A\. Barakat, I\. Panageas, and A\. Varvitsiotis \(2026\)Convex markov games and beyond: new proof of existence, characterization and learning algorithms for nash equilibria\.InInternational Conference on Artificial Intelligence and Statistics,Cited by:[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- A\. Ben\-Tal and M\. Teboulle \(1996\)Hidden convexity in some nonconvex quadratically constrained quadratic programming\.Mathematical Programming72\(1\),pp\. 51–63\.Cited by:[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- A\. Bhaskara, A\. Cutkosky, R\. Kumar, and M\. Purohit \(2025\)Descent with misaligned gradients and applications to hidden convexity\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- X\. Chen, N\. He, Y\. Hu, and Z\. Ye \(2025\)Efficient algorithms for a class of stochastic hidden convex optimization and its applications in network revenue management\.Operations Research73\(2\),pp\. 704–719\.Cited by:[§1\.2](https://arxiv.org/html/2605.26373#S1.SS2.p2.1),[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- R\. D’Orazio, D\. Vucetic, Z\. Liu, J\. L\. Kim, I\. Mitliagkas, and G\. Gidel \(2025\)Solving hidden monotone variational inequalities with surrogate losses\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- T\. Ergen and M\. Pilanci \(2025\)The convex landscape of neural networks: characterizing global optima and stationary points via lasso models\.IEEE Transactions on Information Theory71\(5\),pp\. 3854–3870\.Cited by:[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- I\. Fatkhullin, N\. He, and Y\. Hu \(2025a\)Stochastic optimization under hidden convexity\.SIAM Journal on Optimization35\(4\),pp\. 2544–2571\.Cited by:[§1\.2](https://arxiv.org/html/2605.26373#S1.SS2.p2.1),[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- I\. Fatkhullin, N\. He, G\. Lan, and F\. Wolf \(2025b\)Global solutions to non\-convex functional constrained problems with hidden convexity\.arXiv preprint arXiv:2511\.10626\.Cited by:[§1\.2](https://arxiv.org/html/2605.26373#S1.SS2.p2.1),[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- A\. D\. Flaxman, A\. T\. Kalai, and H\. B\. McMahan \(2005\)Online convex optimization in the bandit setting: gradient descent without a gradient\.InProceedings of the Sixteenth Annual ACM\-SIAM Symposium on Discrete Algorithms,SODA ’05,USA,pp\. 385–394\.Cited by:[Appendix C](https://arxiv.org/html/2605.26373#A3.7.p4.7),[4th item](https://arxiv.org/html/2605.26373#S1.I1.i4.p1.1),[§1\.2](https://arxiv.org/html/2605.26373#S1.SS2.p3.1),[§2](https://arxiv.org/html/2605.26373#S2.p5.4),[§5\.1](https://arxiv.org/html/2605.26373#S5.SS1.p1.6),[§5\.1](https://arxiv.org/html/2605.26373#S5.SS1.p5.1),[Algorithm 1](https://arxiv.org/html/2605.26373#alg1)\.
- X\. Gao, X\. Li, and S\. Zhang \(2018\)Online learning with non\-convex losses and non\-stationary regret\.InInternational Conference on Artificial Intelligence and Statistics,Cited by:[§1\.2](https://arxiv.org/html/2605.26373#S1.SS2.p1.1),[§1\.2](https://arxiv.org/html/2605.26373#S1.SS2.p3.1)\.
- I\. Gemp, A\. A\. Haupt, L\. Marris, S\. Liu, and G\. Piliouras \(2025\)Convex Markov games: a new frontier for multi\-agent reinforcement learning\.InInternational Conference on Machine Learning,Cited by:[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- U\. Ghai, Z\. Lu, and E\. Hazan \(2022\)Non\-convex online learning via algorithmic equivalence\.Advances in Neural Information Processing Systems35,pp\. 22161–22172\.Cited by:[§A\.1](https://arxiv.org/html/2605.26373#A1.SS1.14.p5.2),[§A\.2](https://arxiv.org/html/2605.26373#A1.SS2.p2.1),[1st item](https://arxiv.org/html/2605.26373#S1.I1.i1.p1.4),[2nd item](https://arxiv.org/html/2605.26373#S1.I1.i2.p1.1),[§1\.1](https://arxiv.org/html/2605.26373#S1.SS1.p1.1),[§1\.2](https://arxiv.org/html/2605.26373#S1.SS2.p1.1),[§1](https://arxiv.org/html/2605.26373#S1.p3.2),[§2](https://arxiv.org/html/2605.26373#S2.p2.13),[§3\.1](https://arxiv.org/html/2605.26373#S3.SS1.p1.1),[§3\.1](https://arxiv.org/html/2605.26373#S3.SS1.p2.1),[§3\.2](https://arxiv.org/html/2605.26373#S3.SS2.p3.2),[§3\.3](https://arxiv.org/html/2605.26373#S3.SS3.p1.1),[§3\.3](https://arxiv.org/html/2605.26373#S3.SS3.p2.14),[§3\.3](https://arxiv.org/html/2605.26373#S3.SS3.p5.13),[§3\.3](https://arxiv.org/html/2605.26373#S3.SS3.p5.3),[§4\.1](https://arxiv.org/html/2605.26373#S4.SS1),[§4\.1](https://arxiv.org/html/2605.26373#S4.SS1.p1.9),[§4](https://arxiv.org/html/2605.26373#S4.p1.1),[Lemma 8](https://arxiv.org/html/2605.26373#Thmlemma8)\.
- B\. L\. Gorissen, D\. den Hertog, and M\. Reusken \(2026\)Hidden convexity in a class of optimization problems with bilinear terms\.Operations Research74\(2\),pp\. 1126–1152\.Cited by:[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- E\. Hazan, S\. Kakade, K\. Singh, and A\. Van Soest \(2019\)Provably efficient maximum entropy exploration\.InInternational Conference on Machine Learning,Cited by:[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- E\. Hazan, K\. Levy, and S\. Shalev\-Shwartz \(2015\)Beyond convexity: stochastic quasi\-convex optimization\.Advances in Neural Information Processing Systems\.Cited by:[§1\.2](https://arxiv.org/html/2605.26373#S1.SS2.p1.1),[§2](https://arxiv.org/html/2605.26373#S2.p1.1)\.
- E\. Hazan \(2016\)Introduction to online convex optimization\.Foundations and Trends in Optimization2\(3\-4\),pp\. 157–325\.Cited by:[§1\.2](https://arxiv.org/html/2605.26373#S1.SS2.p1.1),[§1](https://arxiv.org/html/2605.26373#S1.p1.1)\.
- A\. Héliou, M\. Martin, P\. Mertikopoulos, and T\. Rahier \(2020\)Online non\-convex optimization with imperfect feedback\.Advances in Neural Information Processing Systems33,pp\. 17224–17235\.Cited by:[§1\.2](https://arxiv.org/html/2605.26373#S1.SS2.p1.1),[§1](https://arxiv.org/html/2605.26373#S1.p1.1)\.
- F\. Kalogiannis, E\. Vlatakis\-Gkaragkounis, I\. Gemp, and G\. Piliouras \(2025\)Solving zero\-sum convex Markov games\.InInternational Conference on Machine Learning,Cited by:[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- F\. Kalogiannis, J\. Yan, and I\. Panageas \(2024\)Learning equilibria in adversarial team markov games: a nonconvex\-hidden\-concave min\-max optimization problem\.Advances in Neural Information Processing Systems37,pp\. 92832–92890\.Cited by:[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- W\. Krichene, M\. Balandat, C\. Tomlin, and A\. Bayen \(2015\)The hedge algorithm on a continuum\.InInternational Conference on Machine Learning,Cited by:[§1\.2](https://arxiv.org/html/2605.26373#S1.SS2.p1.1),[§1](https://arxiv.org/html/2605.26373#S1.p1.1)\.
- T\. Lattimore and C\. Szepesvári \(2020\)Bandit algorithms\.Cambridge University Press\.Cited by:[§1\.2](https://arxiv.org/html/2605.26373#S1.SS2.p3.1)\.
- T\. Lattimore \(2024\)Bandit convex optimisation\.arXiv preprint arXiv:2402\.06535\.Cited by:[§1\.2](https://arxiv.org/html/2605.26373#S1.SS2.p3.1)\.
- E\. Levin, J\. Kileel, and N\. Boumal \(2025\)The effect of smooth parametrizations on nonconvex optimization landscapes\.Mathematical Programming209\(1\),pp\. 63–111\.Cited by:[§1\.2](https://arxiv.org/html/2605.26373#S1.SS2.p2.1),[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- D\. Li, Z\. Wu, H\. Joseph Lee, X\. Yang, and L\. Zhang \(2005\)Hidden convex minimization\.Journal of Global Optimization31\(2\),pp\. 211–233\.Cited by:[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- Z\. Li, T\. Wang, J\. D\. Lee, and S\. Arora \(2022\)Implicit bias of gradient descent on reparametrized models: on equivalence to mirror descent\.Advances in Neural Information Processing Systems35,pp\. 34626–34640\.Cited by:[§1\.2](https://arxiv.org/html/2605.26373#S1.SS2.p1.1)\.
- O\. Maillard and R\. Munos \(2010\)Online learning in adversarial lipschitz environments\.InJoint european conference on machine learning and knowledge discovery in databases,pp\. 305–320\.Cited by:[§1\.2](https://arxiv.org/html/2605.26373#S1.SS2.p1.1)\.
- A\. Mladenovic, I\. Sakos, G\. Gidel, and G\. Piliouras \(2022\)Generalized natural gradient flows in hidden convex\-concave games and GANs\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- F\. Orabona \(2019\)A modern introduction to online learning\.arXiv preprint arXiv:1912\.13213\.Cited by:[§1\.2](https://arxiv.org/html/2605.26373#S1.SS2.p1.1),[§2](https://arxiv.org/html/2605.26373#S2.p1.1)\.
- D\. Patel and E\. Vlatakis\-Gkaragkounis \(2025\)Solving neural min\-max games: the role of architecture, initialization & dynamics\.InAnnual Conference on Neural Information Processing Systems,Cited by:[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- I\. Sakos, E\. Vlatakis\-Gkaragkounis, P\. Mertikopoulos, and G\. Piliouras \(2023\)Exploiting hidden structures in non\-convex games for convergence to nash equilibrium\.Advances in Neural Information Processing Systems\.Cited by:[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- S\. Shalev\-Shwartz \(2012\)Online learning and online convex optimization\.Foundations and Trends in Machine Learning4\(2\),pp\. 107–194\.Cited by:[§1\.2](https://arxiv.org/html/2605.26373#S1.SS2.p1.1),[§1](https://arxiv.org/html/2605.26373#S1.p1.1),[§2](https://arxiv.org/html/2605.26373#S2.p1.1)\.
- A\. S\. Suggala and P\. Netrapalli \(2020\)Online non\-convex learning: following the perturbed leader is optimal\.InAlgorithmic Learning Theory,pp\. 845–861\.Cited by:[§1\.2](https://arxiv.org/html/2605.26373#S1.SS2.p1.1),[§1](https://arxiv.org/html/2605.26373#S1.p1.1)\.
- E\. Vlatakis\-Gkaragkounis, L\. Flokas, and G\. Piliouras \(2019\)Poincaré recurrence, cycles and spurious equilibria in gradient\-descent\-ascent for non\-convex non\-concave zero\-sum games\.Advances in Neural Information Processing Systems32\.Cited by:[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- Y\. Wang, J\. Lacotte, and M\. Pilanci \(2022\)The hidden convex optimization landscape of regularized two\-layer reLU networks: an exact characterization of optimal solutions\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- T\. Zahavy, B\. O’Donoghue, G\. Desjardins, and S\. Singh \(2021\)Reward is enough for convex mdps\.InAdvances in Neural Information Processing Systems,Cited by:[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- E\. Zeger and M\. Pilanci \(2026\)Unveiling hidden convexity in deep learning: a sparse signal processing perspective\.arXiv preprint arXiv:2603\.23831\.Cited by:[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- J\. Zhang, A\. Koppel, A\. S\. Bedi, C\. Szepesvari, and M\. Wang \(2020\)Variational policy gradient method for reinforcement learning with general utilities\.InAdvances in Neural Information Processing Systems,Cited by:[§1](https://arxiv.org/html/2605.26373#S1.p2.1)\.
- L\. Zhang, T\. Yang, R\. Jin, and Z\. Zhou \(2015\)Online bandit learning for a special class of non\-convex losses\.Proceedings of the AAAI Conference on Artificial Intelligence29\(1\)\.Cited by:[§1\.2](https://arxiv.org/html/2605.26373#S1.SS2.p1.1),[§1\.2](https://arxiv.org/html/2605.26373#S1.SS2.p3.1)\.
- M\. Zinkevich \(2003\)Online convex programming and generalized infinitesimal gradient ascent\.InInternational Conference on Machine Learning,pp\. 928–936\.Cited by:[§1](https://arxiv.org/html/2605.26373#S1.p1.1),[§2](https://arxiv.org/html/2605.26373#S2.p5.1)\.

###### Contents

1. [1Introduction](https://arxiv.org/html/2605.26373#S1)1. [1\.1Main contributions](https://arxiv.org/html/2605.26373#S1.SS1) 2. [1\.2Related Work](https://arxiv.org/html/2605.26373#S1.SS2)
2. [2Online Learning on Hidden\-Convex Losses](https://arxiv.org/html/2605.26373#S2)
3. [3Regret Analysis of OGD for OHCO under Exact Gradient Feedback](https://arxiv.org/html/2605.26373#S3)1. [3\.1Assumptions](https://arxiv.org/html/2605.26373#S3.SS1) 2. [3\.2Regret Bound](https://arxiv.org/html/2605.26373#S3.SS2) 3. [3\.3Proof Sketch of Theorem1](https://arxiv.org/html/2605.26373#S3.SS3)
4. [4Discussion of the Hessian\-Compatibility Hidden Geometry Assumption1](https://arxiv.org/html/2605.26373#S4)1. [4\.1Relaxation of the diagonal Jacobian assumption inGhaiet al\.\(2022\)](https://arxiv.org/html/2605.26373#S4.SS1) 2. [4\.2OGD can incur linear regret without Hessian\-compatible hidden geometry](https://arxiv.org/html/2605.26373#S4.SS2)
5. [5Regret Analysis of OGD for OHCO under Bandit Feedback](https://arxiv.org/html/2605.26373#S5)1. [5\.1Expected Regret Bound in the Bandit Setting](https://arxiv.org/html/2605.26373#S5.SS1) 2. [5\.2Proof Sketch of Theorem3](https://arxiv.org/html/2605.26373#S5.SS2)
6. [6Conclusion and Future Work](https://arxiv.org/html/2605.26373#S6)
7. [References](https://arxiv.org/html/2605.26373#bib)
8. [AProofs for Section3: Proof of Theorem1](https://arxiv.org/html/2605.26373#A1)1. [A\.1Proof of Lemma1](https://arxiv.org/html/2605.26373#A1.SS1) 2. [A\.2Proof of Theorem1using Lemma1](https://arxiv.org/html/2605.26373#A1.SS2)
9. [BProofs for Section4](https://arxiv.org/html/2605.26373#A2)1. [B\.1Proof of Proposition1](https://arxiv.org/html/2605.26373#A2.SS1) 2. [B\.2Detailed derivations for examples](https://arxiv.org/html/2605.26373#A2.SS2) 3. [B\.3Proof of Theorem2: Linear Regret Lower Bound for OGD](https://arxiv.org/html/2605.26373#A2.SS3) 4. [B\.4Detailed proof sketch of Theorem2](https://arxiv.org/html/2605.26373#A2.SS4) 5. [B\.5Proof of Theorem2](https://arxiv.org/html/2605.26373#A2.SS5)
10. [CProof of Theorem3: Expected Regret Bound in the Bandit Setting](https://arxiv.org/html/2605.26373#A3)

## Appendix AProofs for Section[3](https://arxiv.org/html/2605.26373#S3): Proof of Theorem[1](https://arxiv.org/html/2605.26373#Thmtheorem1)

In this section, we provide a complete proof of Theorem[1](https://arxiv.org/html/2605.26373#Thmtheorem1)\. The proof strategy consist in showing that the OGD iterates mapped to𝒴\{\\mathcal\{Y\}\}viaqqand the OMD iterates are close to each other in the hidden space𝒴\{\\mathcal\{Y\}\}after a single update step starting from the same initial point\. Then we can view the OGD update as a perturbed version of OMD, and combine it with the fact that the OMD algorithm can tolerate bounded noise per trial\.

### A\.1Proof of Lemma[1](https://arxiv.org/html/2605.26373#Thmlemma1)

We begin with the following key lemma showing that the updatesyt\+1y\_\{t\+1\}andzt\+1=q​\(xt\+1\)z\_\{t\+1\}=q\(x\_\{t\+1\}\)created by OMD and OGD respectively, are close to each other when starting from the same initial pointyt=zt=q​\(xt\)y\_\{t\}=z\_\{t\}=q\(x\_\{t\}\)\. We show that the two iterates are minimizers of approximately the same strongly\-convex objective, hence the minimizers must be close\. In preparation of the bandit setting, we prove a general result that holds pathwise when we use any bounded vector instead of the exact gradients of both the original loss functionℓt\\ell\_\{t\}and the hidden functionht\.h\_\{t\}\.

###### Proposition 2\. Suppose Assumptions[1](https://arxiv.org/html/2605.26373#Thmassumption1)\-[3](https://arxiv.org/html/2605.26373#Thmassumption3)hold and assumeyt=q​\(xt\)y\_\{t\}=q\(x\_\{t\}\)\. Consider the following OMD and OGD\-like update rules:yt\+1\\displaystyle y\_\{t\+1\}=arg​miny∈𝒴⁡g~t⊤​\(y−yt\)\+1η​DR​\(y∥yt\),\\displaystyle=\\operatorname\*\{arg\\,min\}\_\{y\\in\{\\mathcal\{Y\}\}\}\\tilde\{g\}\_\{t\}^\{\\top\}\(y\-y\_\{t\}\)\+\\frac\{1\}\{\\eta\}D\_\{R\}\(y\\\|y\_\{t\}\)\\,,\(5\)xt\+1\\displaystyle x\_\{t\+1\}=arg​minx∈𝒳⁡gt⊤​\(x−xt\)\+12​η​‖x−xt‖22,\\displaystyle=\\operatorname\*\{arg\\,min\}\_\{x\\in\{\\mathcal\{X\}\}\}g\_\{t\}^\{\\top\}\(x\-x\_\{t\}\)\+\\frac\{1\}\{2\\eta\}\\\|x\-x\_\{t\}\\\|\_\{2\}^\{2\}\\,,\(6\)whereg~t=Jq​\(xt\)−⊤​gt\\tilde\{g\}\_\{t\}=J\_\{q\}\(x\_\{t\}\)^\{\-\\top\}g\_\{t\}andgt∈ℝdg\_\{t\}\\in\\mathbb\{R\}^\{d\}is any vector111g~t\\tilde\{g\}\_\{t\}is used in the analysis but never actually computed, in contrast togtg\_\{t\}which is computed and used in BOGD\. Note that we suppose that we do not have access to the parameterizationqq\.\. If in addition there existsG^F\>0\\hat\{G\}\_\{F\}\>0such that‖gt‖≤G^F\\\|g\_\{t\}\\\|~\\leq~\\hat\{G\}\_\{F\}for anytt, then,‖yt\+1−q​\(xt\+1\)‖2≤G5​\(5​G^F2\+G^F3​η\)​η2\.\\\|y\_\{t\+1\}\-q\(x\_\{t\+1\}\)\\\|\_\{2\}~\\leq~G^\{5\}\(5\\hat\{G\}\_\{F\}^\{2\}\+\\hat\{G\}\_\{F\}^\{3\}\\eta\)\\eta^\{2\}\\,\.\(7\)

Before providing the proof, we state an immediate corollary of this result in the exact gradient information setting of Theorem[1](https://arxiv.org/html/2605.26373#Thmtheorem1)\.

###### Corollary 4\(Proposition[2](https://arxiv.org/html/2605.26373#Thmproposition2)in the exact gradient setting\)\. In Proposition[2](https://arxiv.org/html/2605.26373#Thmproposition2), setgt=∇ℓt​\(xt\)g\_\{t\}=\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\(x\_\{t\}\)andg~t=∇ht​\(zt\)\\tilde\{g\}\_\{t\}=\\nabla\\mkern\-2\.5muh\_\{t\}\(z\_\{t\}\)in \([5](https://arxiv.org/html/2605.26373#A1.E5)\)\-\([6](https://arxiv.org/html/2605.26373#A1.E6)\), then the approximation inequality \([7](https://arxiv.org/html/2605.26373#A1.E7)\) holds withG^F=GF\\hat\{G\}\_\{F\}=G\_\{F\}whereGFG\_\{F\}is the uniform loss gradient \(∇ℓt\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\) bound in Assumption[3](https://arxiv.org/html/2605.26373#Thmassumption3), i\.e\.,‖yt\+1−q​\(xt\+1\)‖2≤6​G5​GF3​η2\.\\\|y\_\{t\+1\}\-q\(x\_\{t\+1\}\)\\\|\_\{2\}~\\leq~6G^\{5\}G\_\{F\}^\{3\}\\eta^\{2\}\\,\.

###### Proof of Proposition[2](https://arxiv.org/html/2605.26373#Thmproposition2)\.

Note first that‖g~t‖≤G~F:=G​G^F\.\\\|\\tilde\{g\}\_\{t\}\\\|~\\leq~\\tilde\{G\}\_\{F\}:=G\\hat\{G\}\_\{F\}\\,\.Both objectives in \([5](https://arxiv.org/html/2605.26373#A1.E5)\)\-\([6](https://arxiv.org/html/2605.26373#A1.E6)\) can be written as the sum of a linear function and a strongly convex function\. By 1\-strong convexity ofRRand uniform boundedness ofgtg\_\{t\}byG^F\\hat\{G\}\_\{F\}, we have for anyy∈𝒴y\\in\{\\mathcal\{Y\}\},

g~t⊤​\(y−yt\)\+1η​DR​\(y∥yt\)≥−G~F​‖y−yt‖\+12​η​‖y−yt‖2\.\\tilde\{g\}\_\{t\}^\{\\top\}\(y\-y\_\{t\}\)\+\\frac\{1\}\{\\eta\}D\_\{R\}\(y\\\|y\_\{t\}\)~\\geq~\-\\tilde\{G\}\_\{F\}\\\|y\-y\_\{t\}\\\|\+\\frac\{1\}\{2\\eta\}\\\|y\-y\_\{t\}\\\|^\{2\}\\,\.Therefore, if‖y−yt‖2\>2​η​G~F\\\|y\-y\_\{t\}\\\|\_\{2\}\>2\\eta\\tilde\{G\}\_\{F\}, then the objective takes a positive value which cannot be minimal sincey=yty=y\_\{t\}gives zero objective value in \([5](https://arxiv.org/html/2605.26373#A1.E5)\)\. Define the set:

𝒴r,yt=𝒦∩\{y∈ℝd:‖y−yt‖2≤r\}\.\{\\mathcal\{Y\}\}\_\{r,y\_\{t\}\}=\\mathcal\{K\}\\cap\\\{y\\in\\mathbb\{R\}^\{d\}:\\\|y\-y\_\{t\}\\\|\_\{2\}\\leq r\\\}\\,\.Then it follows that the OMD iterate in \([5](https://arxiv.org/html/2605.26373#A1.E5)\) satisfies:

yt\+1∈𝒴t:=𝒴2​η​G~F,yt\.y\_\{t\+1\}\\in\{\\mathcal\{Y\}\}\_\{t\}:=\{\\mathcal\{Y\}\}\_\{2\\eta\\tilde\{G\}\_\{F\},y\_\{t\}\}\\,\.\(8\)As we want to localize both iterates \([5](https://arxiv.org/html/2605.26373#A1.E5)\)\-\([6](https://arxiv.org/html/2605.26373#A1.E6)\) in a set of size controlled byη\\eta, we now show that we also haveq​\(xt\+1\)∈𝒴tq\(x\_\{t\+1\}\)\\in\{\\mathcal\{Y\}\}\_\{t\}whenyt=q​\(xt\)y\_\{t\}=q\(x\_\{t\}\)\. Observe for this that:

‖q​\(xt\+1\)−yt‖=‖q​\(xt\+1\)−q​\(xt\)‖≤G​‖xt\+1−xt‖≤η​G​‖gt‖≤η​G​G^F≤2​η​G​G^F=2​η​G~F\.\\\|q\(x\_\{t\+1\}\)\-y\_\{t\}\\\|=\\\|q\(x\_\{t\+1\}\)\-q\(x\_\{t\}\)\\\|~\\leq~G\\\|x\_\{t\+1\}\-x\_\{t\}\\\|~\\leq~\\eta G\\\|g\_\{t\}\\\|~\\leq~\\eta G\\hat\{G\}\_\{F\}~\\leq~2\\eta G\\hat\{G\}\_\{F\}=2\\eta\\tilde\{G\}\_\{F\}\\,\.We now write both update rules under a similar form\. Define the functionΦt:𝒴→ℝ\\Phi\_\{t\}:\{\\mathcal\{Y\}\}\\to\\mathbb\{R\}by:

Φt​\(y\):=g~t⊤​\(y−yt\)\+12​η​‖y−yt‖∇2R​\(yt\)2\.\\Phi\_\{t\}\(y\):=\\tilde\{g\}\_\{t\}^\{\\top\}\(y\-y\_\{t\}\)\+\\frac\{1\}\{2\\eta\}\\\|y\-y\_\{t\}\\\|\_\{\\nabla\\mkern\-2\.5mu^\{2\}R\(y\_\{t\}\)\}^\{2\}\\,\.\(9\)Note thatΦt\\Phi\_\{t\}is1η\\frac\{1\}\{\\eta\}\-strongly convex since∇2Φt​\(x\)=1η​∇2R​\(xt\)⪰1η​I\\nabla\\mkern\-2\.5mu^\{2\}\\Phi\_\{t\}\(x\)=\\frac\{1\}\{\\eta\}\\nabla\\mkern\-2\.5mu^\{2\}R\(x\_\{t\}\)\\succeq\\frac\{1\}\{\\eta\}IasRRis supposed to be11\-strongly convex\.

###### Lemma 2\. For anyt≥1,t~\\geq~1,the iterates of \([OMD](https://arxiv.org/html/2605.26373#S3.Ex3)\)\-\([5](https://arxiv.org/html/2605.26373#A1.E5)\) can be rewritten as follows:yt\+1=arg​miny∈𝒴t⁡\{Φt​\(y\)\+εtOMD​\(y\)\},εtOMD​\(y\):=1η​εt​\(y\),y\_\{t\+1\}=\\operatorname\*\{arg\\,min\}\_\{y\\in\{\\mathcal\{Y\}\}\_\{t\}\}\\left\\\{\\Phi\_\{t\}\(y\)\+\\varepsilon\_\{t\}^\{\\mathrm\{OMD\}\}\(y\)\\right\\\}\\,,\\quad\\varepsilon\_\{t\}^\{\\mathrm\{OMD\}\}\(y\):=\\frac\{1\}\{\\eta\}\\varepsilon\_\{t\}\(y\)\\,,whereεt​\(y\):=DR​\(y∥yt\)−12​‖y−yt‖∇2R​\(yt\)2\.\\varepsilon\_\{t\}\(y\):=D\_\{R\}\(y\\\|y\_\{t\}\)\-\\frac\{1\}\{2\}\\\|y\-y\_\{t\}\\\|\_\{\\nabla\\mkern\-2\.5mu^\{2\}R\(y\_\{t\}\)\}^\{2\}\\,\.Moreover, fory∈𝒴ty\\in\{\\mathcal\{Y\}\}\_\{t\},\|εt​\(y\)\|≤43​G​G~F3​η3\.\|\\varepsilon\_\{t\}\(y\)\|~\\leq~\\frac\{4\}\{3\}G\\tilde\{G\}\_\{F\}^\{3\}\\eta^\{3\}\\,\.

###### Proof\.

We have proved in \([8](https://arxiv.org/html/2605.26373#A1.E8)\) thatyt\+1∈𝒴ty\_\{t\+1\}\\in\{\\mathcal\{Y\}\}\_\{t\}\. Hence, we can rewrite \([5](https://arxiv.org/html/2605.26373#A1.E5)\) by replacing𝒴\{\\mathcal\{Y\}\}by𝒴t\{\\mathcal\{Y\}\}\_\{t\}:

yt\+1=arg​miny∈𝒴t⁡g~t⊤​\(y−yt\)\+1η​DR​\(y∥yt\)\.y\_\{t\+1\}=\\operatorname\*\{arg\\,min\}\_\{y\\in\{\\mathcal\{Y\}\}\_\{t\}\}\\tilde\{g\}\_\{t\}^\{\\top\}\(y\-y\_\{t\}\)\+\\frac\{1\}\{\\eta\}D\_\{R\}\(y\\\|y\_\{t\}\)\.\(10\)
Letht:=y−yth\_\{t\}:=y\-y\_\{t\}\. By Taylor’s theorem with integral remainder, we have:

DR​\(y∥yt\)\\displaystyle D\_\{R\}\(y\\\|y\_\{t\}\)=R​\(y\)−R​\(yt\)−∇R​\(yt\)⊤​\(y−yt\)\\displaystyle=R\(y\)\-R\(y\_\{t\}\)\-\\nabla\\mkern\-2\.5muR\(y\_\{t\}\)^\{\\top\}\(y\-y\_\{t\}\)=12​ht⊤​∇2R​\(yt\)​ht\+εt​\(y\)\\displaystyle=\\frac\{1\}\{2\}h\_\{t\}^\{\\top\}\\nabla\\mkern\-2\.5mu^\{2\}R\(y\_\{t\}\)h\_\{t\}\+\\varepsilon\_\{t\}\(y\)=12​‖ht‖∇2R​\(yt\)2\+εt​\(y\),\\displaystyle=\\frac\{1\}\{2\}\\\|h\_\{t\}\\\|\_\{\\nabla\\mkern\-2\.5mu^\{2\}R\(y\_\{t\}\)\}^\{2\}\+\\varepsilon\_\{t\}\(y\),\(11\)whereεt​\(y\)\\varepsilon\_\{t\}\(y\)is the integral remainder defined as follows:

εt​\(y\):=12​∫01\(1−s\)2​∇3R​\(yt\+s​ht\)​\[ht,ht,ht\]​𝑑s,\\varepsilon\_\{t\}\(y\):=\\frac\{1\}\{2\}\\int\_\{0\}^\{1\}\(1\-s\)^\{2\}\\nabla\\mkern\-2\.5mu^\{3\}R\(y\_\{t\}\+sh\_\{t\}\)\[h\_\{t\},h\_\{t\},h\_\{t\}\]\\,ds\\,,where∇3R\\nabla\\mkern\-2\.5mu^\{3\}Ris the third\-order derivative tensor ofRR\. Since this third\-order is supposed to be uniformly bounded byGGoverz∈𝒴z\\in\{\\mathcal\{Y\}\}by Assumption[2](https://arxiv.org/html/2605.26373#Thmassumption2), we have:

\|εt​\(y\)\|≤G6​‖ht‖23\.\|\\varepsilon\_\{t\}\(y\)\|\\leq\\frac\{G\}\{6\}\\\|h\_\{t\}\\\|\_\{2\}^\{3\}\.In particular, fory∈𝒴ty\\in\\mathcal\{Y\}\_\{t\}, i\.e\.‖y−yt‖2≤2​η​G~F\\\|y\-y\_\{t\}\\\|\_\{2\}\\leq 2\\eta\\tilde\{G\}\_\{F\}, we obtain the following estimate:

\|εt​\(y\)\|≤G6​\(2​η​G~F\)3=43​G​G~F3​η3\.\|\\varepsilon\_\{t\}\(y\)\|~\\leq~\\frac\{G\}\{6\}\(2\\eta\\tilde\{G\}\_\{F\}\)^\{3\}=\\frac\{4\}\{3\}G\\tilde\{G\}\_\{F\}^\{3\}\\eta^\{3\}\\,\.Combining the Bregman approximation of \([A\.1](https://arxiv.org/html/2605.26373#A1.Ex19)\) with \([10](https://arxiv.org/html/2605.26373#A1.E10)\), we obtain the desired identity using the definition of the strongly\-convex functionΦt\\Phi\_\{t\}introduced in \([9](https://arxiv.org/html/2605.26373#A1.E9)\)\. ∎

In the following lemma, we rewrite the OGD iterates mapped to𝒴\{\\mathcal\{Y\}\}viaqqunder a similar form to the update rule of OMD obtained in Lemma[2](https://arxiv.org/html/2605.26373#Thmlemma2)\.

###### Lemma 3\. If\(xt\)\(x\_\{t\}\)is the sequence of \([OGD](https://arxiv.org/html/2605.26373#S2.Ex2)\) iterates andyt=q​\(xt\)y\_\{t\}=q\(x\_\{t\}\), then for anyt≥1,t~\\geq~1,zt\+1=q​\(xt\+1\)\\displaystyle z\_\{t\+1\}=q\(x\_\{t\+1\}\)=arg​miny∈𝒴t⁡\{Φt​\(y\)\+εtOGD​\(y\)\},\\displaystyle=\\operatorname\*\{arg\\,min\}\_\{y\\in\{\\mathcal\{Y\}\}\_\{t\}\}\\left\\\{\\Phi\_\{t\}\(y\)\+\\varepsilon\_\{t\}^\{\\mathrm\{OGD\}\}\(y\)\\right\\\}\\,,\(12\)εtOGD​\(y\)\\displaystyle\\varepsilon\_\{t\}^\{\\mathrm\{OGD\}\}\(y\):=ε~t​\(y\)\+1η​\(12​‖εtq​\(y\)‖22\+⟨Jq​\(xt\)−1​\(y−yt\),εtq​\(y\)⟩\),\\displaystyle:=\\tilde\{\\varepsilon\}\_\{t\}\(y\)\+\\frac\{1\}\{\\eta\}\\left\(\\frac\{1\}\{2\}\\\|\\varepsilon\_\{t\}^\{q\}\(y\)\\\|\_\{2\}^\{2\}\+\\big\\langle J\_\{q\}\(x\_\{t\}\)^\{\-1\}\(y\-y\_\{t\}\),\\,\\varepsilon\_\{t\}^\{q\}\(y\)\\big\\rangle\\right\)\\,,\(13\)whereε~t​\(y\):=gt⊤​εtq​\(y\)\\tilde\{\\varepsilon\}\_\{t\}\(y\):=g\_\{t\}^\{\\top\}\\varepsilon\_\{t\}^\{q\}\(y\)andεtq​\(y\):=q−1​\(y\)−q−1​\(yt\)−Jq−1​\(xt\)​\(y−yt\)\.\\varepsilon\_\{t\}^\{q\}\(y\):=q^\{\-1\}\(y\)\-q^\{\-1\}\(y\_\{t\}\)\-J\_\{q\}^\{\-1\}\(x\_\{t\}\)\(y\-y\_\{t\}\)\\,\.Moreover, we have‖εtq​\(y\)‖2≤2​G​G~F2​η2\\\|\\varepsilon\_\{t\}^\{q\}\(y\)\\\|\_\{2\}~\\leq~2G\\tilde\{G\}\_\{F\}^\{2\}\\eta^\{2\}and\|ε~t​\(y\)\|≤2​G~F3​η2\|\\tilde\{\\varepsilon\}\_\{t\}\(y\)\|~\\leq~2\\tilde\{G\}\_\{F\}^\{3\}\\eta^\{2\}

###### Proof\.

Recall that the \([OGD](https://arxiv.org/html/2605.26373#S2.Ex2)\) update \([6](https://arxiv.org/html/2605.26373#A1.E6)\) in the original space𝒳\{\\mathcal\{X\}\}can be written as follows:

xt\+1=arg​minx∈𝒳⁡\{gt⊤​\(x−xt\)\+12​η​‖x−xt‖22\}\.x\_\{t\+1\}=\\operatorname\*\{arg\\,min\}\_\{x\\in\{\\mathcal\{X\}\}\}\\left\\\{g\_\{t\}^\{\\top\}\(x\-x\_\{t\}\)\+\\frac\{1\}\{2\\eta\}\\\|x\-x\_\{t\}\\\|\_\{2\}^\{2\}\\right\\\}\.\(14\)
Recall from \([9](https://arxiv.org/html/2605.26373#A1.E9)\) that for anyy∈𝒴y\\in\{\\mathcal\{Y\}\}:

Φt​\(y\):=g~t⊤​\(y−yt\)\+12​η​‖y−yt‖∇2R​\(yt\)2\.\\Phi\_\{t\}\(y\):=\\tilde\{g\}\_\{t\}^\{\\top\}\(y\-y\_\{t\}\)\+\\frac\{1\}\{2\\eta\}\\\|y\-y\_\{t\}\\\|\_\{\\nabla\\mkern\-2\.5mu^\{2\}R\(y\_\{t\}\)\}^\{2\}\\,\.To prove \([12](https://arxiv.org/html/2605.26373#A1.E12)\), we relate the linear termgt⊤​\(x−xt\)g\_\{t\}^\{\\top\}\(x\-x\_\{t\}\)tog~t⊤​\(y−yt\)\\tilde\{g\}\_\{t\}^\{\\top\}\(y\-y\_\{t\}\)and the quadratic term12​η​‖x−xt‖22\\frac\{1\}\{2\\eta\}\\\|x\-x\_\{t\}\\\|\_\{2\}^\{2\}to12​η​‖y−yt‖∇2R​\(yt\)2\\frac\{1\}\{2\\eta\}\\\|y\-y\_\{t\}\\\|\_\{\\nabla\\mkern\-2\.5mu^\{2\}R\(y\_\{t\}\)\}^\{2\}respectively in\(i\)and\(ii\)below\.

\(i\) Relatinggt⊤​\(x−xt\)g\_\{t\}^\{\\top\}\(x\-x\_\{t\}\)tog~t⊤​\(y−yt\)\.\\tilde\{g\}\_\{t\}^\{\\top\}\(y\-y\_\{t\}\)\.Recalling thatgt=Jq​\(xt\)⊤​g~tg\_\{t\}=J\_\{q\}\(x\_\{t\}\)^\{\\top\}\\tilde\{g\}\_\{t\}, we have:

gt⊤​\(x−xt\)\\displaystyle g\_\{t\}^\{\\top\}\(x\-x\_\{t\}\)=g~t⊤​Jq​\(xt\)​\(x−xt\)\\displaystyle=\\tilde\{g\}\_\{t\}^\{\\top\}J\_\{q\}\(x\_\{t\}\)\(x\-x\_\{t\}\)=g~t⊤​Jq​\(xt\)​\(q−1​\(y\)−q−1​\(yt\)\),\\displaystyle=\\tilde\{g\}\_\{t\}^\{\\top\}J\_\{q\}\(x\_\{t\}\)\(q^\{\-1\}\(y\)\-q^\{\-1\}\(y\_\{t\}\)\)\\,,\(15\)where the last equality follows from using the identitiesy=q​\(x\)y=q\(x\)andyt=q​\(xt\)\.y\_\{t\}=q\(x\_\{t\}\)\\,\.

Under Assumption[2](https://arxiv.org/html/2605.26373#Thmassumption2), applying Taylor’s theorem toq−1q^\{\-1\}aroundyt=q​\(xt\)y\_\{t\}=q\(x\_\{t\}\)gives, fory=q​\(x\)y=q\(x\),

q−1​\(y\)−q−1​\(yt\)=Jq−1​\(yt\)​\(y−yt\)\+εtq​\(y\),q^\{\-1\}\(y\)\-q^\{\-1\}\(y\_\{t\}\)=J\_\{q^\{\-1\}\}\(y\_\{t\}\)\(y\-y\_\{t\}\)\+\\varepsilon\_\{t\}^\{q\}\(y\)\\,,\(16\)where the remainder is given by:

εtq​\(y\)=∫01\(1−s\)​\(y−yt\)⊤​∇2q−1​\(yt\+s​\(y−yt\)\)​\(y−yt\)​𝑑s\.\\varepsilon\_\{t\}^\{q\}\(y\)=\\int\_\{0\}^\{1\}\(1\-s\)\\,\(y\-y\_\{t\}\)^\{\\top\}\\nabla\\mkern\-2\.5mu^\{2\}q^\{\-1\}\(y\_\{t\}\+s\(y\-y\_\{t\}\)\)\(y\-y\_\{t\}\)\\,ds\.\(17\)SinceJq−1​\(yt\)=Jq​\(xt\)−1J\_\{q^\{\-1\}\}\(y\_\{t\}\)=J\_\{q\}\(x\_\{t\}\)^\{\-1\}, it follows that:

q−1​\(y\)−q−1​\(yt\)=Jq​\(xt\)−1​\(y−yt\)\+εtq​\(y\)\.q^\{\-1\}\(y\)\-q^\{\-1\}\(y\_\{t\}\)=J\_\{q\}\(x\_\{t\}\)^\{\-1\}\(y\-y\_\{t\}\)\+\\varepsilon\_\{t\}^\{q\}\(y\)\.\(18\)Since the second derivative ofq−1q^\{\-1\}is bounded byGG, we have:

‖εtq​\(y\)‖2≤G2​‖y−yt‖22\.\\\|\\varepsilon\_\{t\}^\{q\}\(y\)\\\|\_\{2\}\\leq\\frac\{G\}\{2\}\\\|y\-y\_\{t\}\\\|\_\{2\}^\{2\}\.In particular, fory∈𝒴t=𝒴2​η​G~F,yty\\in\{\\mathcal\{Y\}\}\_\{t\}=\\mathcal\{Y\}\_\{2\\eta\\tilde\{G\}\_\{F\},y\_\{t\}\}, i\.e\.y∈𝒴y\\in\{\\mathcal\{Y\}\}s\.t\.‖y−yt‖2≤2​η​G~F\\\|y\-y\_\{t\}\\\|\_\{2\}\\leq 2\\eta\\tilde\{G\}\_\{F\}, we obtain:

‖εtq​\(y\)‖2≤G2​\(2​η​G~F\)2=2​G​G~F2​η2\.\\\|\\varepsilon\_\{t\}^\{q\}\(y\)\\\|\_\{2\}~\\leq~\\frac\{G\}\{2\}\(2\\eta\\tilde\{G\}\_\{F\}\)^\{2\}=2G\\tilde\{G\}\_\{F\}^\{2\}\\eta^\{2\}\\,\.\(19\)
Using \([18](https://arxiv.org/html/2605.26373#A1.E18)\) in \([A\.1](https://arxiv.org/html/2605.26373#A1.Ex25)\) yields:

gt⊤​\(x−xt\)\\displaystyle g\_\{t\}^\{\\top\}\(x\-x\_\{t\}\)=g~t⊤​Jq​\(xt\)​\(q−1​\(y\)−q−1​\(yt\)\)\\displaystyle=\\tilde\{g\}\_\{t\}^\{\\top\}J\_\{q\}\(x\_\{t\}\)\(q^\{\-1\}\(y\)\-q^\{\-1\}\(y\_\{t\}\)\)=g~t⊤​\(y−yt\)\+g~t⊤​Jq​\(xt\)​εtq​\(y\)\\displaystyle=\\tilde\{g\}\_\{t\}^\{\\top\}\(y\-y\_\{t\}\)\+\\tilde\{g\}\_\{t\}^\{\\top\}J\_\{q\}\(x\_\{t\}\)\\,\\varepsilon\_\{t\}^\{q\}\(y\)=g~t⊤​\(y−yt\)\+ε~t​\(y\),\\displaystyle=\\tilde\{g\}\_\{t\}^\{\\top\}\(y\-y\_\{t\}\)\+\\tilde\{\\varepsilon\}\_\{t\}\(y\),\(20\)whereε~t​\(y\):=g~t⊤​Jq​\(xt\)​εtq​\(y\)=gt⊤​εtq​\(y\)\\tilde\{\\varepsilon\}\_\{t\}\(y\):=\\tilde\{g\}\_\{t\}^\{\\top\}J\_\{q\}\(x\_\{t\}\)\\,\\varepsilon\_\{t\}^\{q\}\(y\)=g\_\{t\}^\{\\top\}\\varepsilon\_\{t\}^\{q\}\(y\)\. Therefore, using \([19](https://arxiv.org/html/2605.26373#A1.E19)\), we have for anyy∈𝒴t,y\\in\{\\mathcal\{Y\}\}\_\{t\},

\|ε~t​\(y\)\|≤‖gt‖⋅‖εtq​\(y\)‖≤G​\(2​G^F​G~F2​η2\)=2​G~F3​η2,\|\\tilde\{\\varepsilon\}\_\{t\}\(y\)\|~\\leq~\\\|g\_\{t\}\\\|\\cdot\\\|\\varepsilon\_\{t\}^\{q\}\(y\)\\\|~\\leq~G\(2\\hat\{G\}\_\{F\}\\tilde\{G\}\_\{F\}^\{2\}\\eta^\{2\}\)=2\\tilde\{G\}\_\{F\}^\{3\}\\eta^\{2\}\\,,where the last equality stems from recalling thatG~F=G​G^F\.\\tilde\{G\}\_\{F\}=G\\hat\{G\}\_\{F\}\\,\.

\(ii\) Relating12​η​‖x−xt‖22\\frac\{1\}\{2\\eta\}\\\|x\-x\_\{t\}\\\|\_\{2\}^\{2\}to12​η​‖y−yt‖∇2R​\(yt\)2\.\\frac\{1\}\{2\\eta\}\\\|y\-y\_\{t\}\\\|\_\{\\nabla\\mkern\-2\.5mu^\{2\}R\(y\_\{t\}\)\}^\{2\}\.For anyx∈𝒳x\\in\{\\mathcal\{X\}\},y=q​\(x\),y=q\(x\),given thatyt=q​\(xt\)y\_\{t\}=q\(x\_\{t\}\),

12​η​‖x−xt‖22\\displaystyle\\frac\{1\}\{2\\eta\}\\\|x\-x\_\{t\}\\\|\_\{2\}^\{2\}=12​η​‖q−1​\(y\)−q−1​\(yt\)‖22\\displaystyle=\\frac\{1\}\{2\\eta\}\\\|q^\{\-1\}\(y\)\-q^\{\-1\}\(y\_\{t\}\)\\\|\_\{2\}^\{2\}=12​η​‖Jq​\(xt\)−1​\(y−yt\)\+εtq​\(y\)‖22\\displaystyle=\\frac\{1\}\{2\\eta\}\\\|J\_\{q\}\(x\_\{t\}\)^\{\-1\}\(y\-y\_\{t\}\)\+\\varepsilon\_\{t\}^\{q\}\(y\)\\\|\_\{2\}^\{2\}=12​η​‖Jq​\(xt\)−1​\(y−yt\)‖22\+1η​\(12​‖εtq​\(y\)‖22\+⟨Jq​\(xt\)−1​\(y−yt\),εtq​\(y\)⟩\)\.\\displaystyle=\\frac\{1\}\{2\\eta\}\\\|J\_\{q\}\(x\_\{t\}\)^\{\-1\}\(y\-y\_\{t\}\)\\\|\_\{2\}^\{2\}\+\\frac\{1\}\{\\eta\}\\left\(\\frac\{1\}\{2\}\\\|\\varepsilon\_\{t\}^\{q\}\(y\)\\\|\_\{2\}^\{2\}\+\\big\\langle J\_\{q\}\(x\_\{t\}\)^\{\-1\}\(y\-y\_\{t\}\),\\,\\varepsilon\_\{t\}^\{q\}\(y\)\\big\\rangle\\right\)\.\(21\)Using Assumption[1](https://arxiv.org/html/2605.26373#Thmassumption1)relatingJqJ\_\{q\}and∇2R\\nabla\\mkern\-2\.5mu^\{2\}Rand the initial couplingyt=zt=q​\(xt\)y\_\{t\}=z\_\{t\}=q\(x\_\{t\}\), we have:

\[∇2R​\(yt\)\]−1=Jq​\(xt\)​Jq​\(xt\)⊤,∇2R​\(yt\)=\[Jq​\(xt\)⊤\]−1​Jq​\(xt\)−1\.\\big\[\\nabla\\mkern\-2\.5mu^\{2\}R\(y\_\{t\}\)\\big\]^\{\-1\}=J\_\{q\}\(x\_\{t\}\)J\_\{q\}\(x\_\{t\}\)^\{\\top\},\\qquad\\nabla\\mkern\-2\.5mu^\{2\}R\(y\_\{t\}\)=\[J\_\{q\}\(x\_\{t\}\)^\{\\top\}\]^\{\-1\}J\_\{q\}\(x\_\{t\}\)^\{\-1\}\\,\.Therefore we can write:

‖Jq​\(xt\)−1​\(y−yt\)‖22\\displaystyle\\\|J\_\{q\}\(x\_\{t\}\)^\{\-1\}\(y\-y\_\{t\}\)\\\|\_\{2\}^\{2\}=\(y−yt\)⊤​\[Jq​\(xt\)−1\]⊤​Jq​\(xt\)−1​\(y−yt\)\\displaystyle=\(y\-y\_\{t\}\)^\{\\top\}\[J\_\{q\}\(x\_\{t\}\)^\{\-1\}\]^\{\\top\}J\_\{q\}\(x\_\{t\}\)^\{\-1\}\(y\-y\_\{t\}\)=\(y−yt\)⊤​∇2R​\(yt\)​\(y−yt\)\\displaystyle=\(y\-y\_\{t\}\)^\{\\top\}\\nabla\\mkern\-2\.5mu^\{2\}R\(y\_\{t\}\)\(y\-y\_\{t\}\)=‖y−yt‖∇2R​\(yt\)2\.\\displaystyle=\\\|y\-y\_\{t\}\\\|^\{2\}\_\{\\nabla\\mkern\-2\.5mu^\{2\}R\(y\_\{t\}\)\}\.\(22\)Combining \([A\.1](https://arxiv.org/html/2605.26373#A1.Ex30)\) and \([A\.1](https://arxiv.org/html/2605.26373#A1.Ex33)\) yields:

12​η​‖x−xt‖22=12​η​‖y−yt‖∇2R​\(yt\)2\+1η​\(12​‖εtq​\(y\)‖22\+⟨Jq​\(xt\)−1​\(y−yt\),εtq​\(y\)⟩\)\.\\frac\{1\}\{2\\eta\}\\\|x\-x\_\{t\}\\\|\_\{2\}^\{2\}=\\frac\{1\}\{2\\eta\}\\\|y\-y\_\{t\}\\\|^\{2\}\_\{\\nabla\\mkern\-2\.5mu^\{2\}R\(y\_\{t\}\)\}\+\\frac\{1\}\{\\eta\}\\left\(\\frac\{1\}\{2\}\\\|\\varepsilon\_\{t\}^\{q\}\(y\)\\\|\_\{2\}^\{2\}\+\\big\\langle J\_\{q\}\(x\_\{t\}\)^\{\-1\}\(y\-y\_\{t\}\),\\,\\varepsilon\_\{t\}^\{q\}\(y\)\\big\\rangle\\right\)\\,\.
Given the OGD update rule \([14](https://arxiv.org/html/2605.26373#A1.E14)\), the change of variabley=q​\(x\)y=q\(x\)and \([A\.1](https://arxiv.org/html/2605.26373#A1.Ex27)\), \([A\.1](https://arxiv.org/html/2605.26373#A1.Ex35)\) proven in\(i\)and\(ii\)respectively, we have shown that

q​\(xt\+1\)=arg​miny∈𝒴t⁡\{Φt​\(y\)\+εtOGD​\(y\)\},q\(x\_\{t\+1\}\)=\\operatorname\*\{arg\\,min\}\_\{y\\in\{\\mathcal\{Y\}\}\_\{t\}\}\\left\\\{\\Phi\_\{t\}\(y\)\+\\varepsilon\_\{t\}^\{\\mathrm\{OGD\}\}\(y\)\\right\\\},where the error termεtOGD​\(y\)\\varepsilon\_\{t\}^\{\\mathrm\{OGD\}\}\(y\)is defined as follows:

εtOGD​\(y\):=ε~t​\(y\)\+1η​\(12​‖εtq​\(y\)‖22\+⟨Jq​\(xt\)−1​\(y−yt\),εtq​\(y\)⟩\)\.\\varepsilon\_\{t\}^\{\\mathrm\{OGD\}\}\(y\):=\\tilde\{\\varepsilon\}\_\{t\}\(y\)\+\\frac\{1\}\{\\eta\}\\left\(\\frac\{1\}\{2\}\\\|\\varepsilon\_\{t\}^\{q\}\(y\)\\\|\_\{2\}^\{2\}\+\\big\\langle J\_\{q\}\(x\_\{t\}\)^\{\-1\}\(y\-y\_\{t\}\),\\,\\varepsilon\_\{t\}^\{q\}\(y\)\\big\\rangle\\right\)\.\(23\)∎

So far, we have shown in Lemmas[2](https://arxiv.org/html/2605.26373#Thmlemma2)and[3](https://arxiv.org/html/2605.26373#Thmlemma3)respectively the following similar update rule forms:

yt\+1\\displaystyle y\_\{t\+1\}=arg​miny∈𝒴t⁡\{Φt​\(y\)\+εtOMD​\(y\)\},\\displaystyle=\\operatorname\*\{arg\\,min\}\_\{y\\in\{\\mathcal\{Y\}\}\_\{t\}\}\\\{\\Phi\_\{t\}\(y\)\+\\varepsilon\_\{t\}^\{\\mathrm\{OMD\}\}\(y\)\\\},\(24\)zt\+1=q​\(xt\+1\)\\displaystyle z\_\{t\+1\}=q\(x\_\{t\+1\}\)=arg​miny∈𝒴t⁡\{Φt​\(y\)\+εtOGD​\(y\)\}\.\\displaystyle=\\operatorname\*\{arg\\,min\}\_\{y\\in\{\\mathcal\{Y\}\}\_\{t\}\}\\\{\\Phi\_\{t\}\(y\)\+\\varepsilon\_\{t\}^\{\\mathrm\{OGD\}\}\(y\)\\\}\.\(25\)In the next lemma, we use first\-order conditions together with strong convexity of the functionΦt\\Phi\_\{t\}to show that the distance between the minimizers can be upperbounded byη\\etatimes the distance between the gradients of the errors in each one of the updates \(for OMD and OGD\)\. Note that our proof technique here is different from the strategy adopted inGhaiet al\.\[[2022](https://arxiv.org/html/2605.26373#bib.bib46)\]\(Lemma 4\) as we do not control the distance between objectives and translate it to the distance between minimizers using strong convexity\.

###### Lemma 4\. For anyt≥1t~\\geq~1, ifyt=q​\(xt\)y\_\{t\}=q\(x\_\{t\}\), then‖yt\+1−q​\(xt\+1\)‖≤η​‖∇εtOGD​\(q​\(xt\+1\)\)−∇εtOMD​\(yt\+1\)‖\.\\\|y\_\{t\+1\}\-q\(x\_\{t\+1\}\)\\\|\\leq\\eta\\,\\left\\\|\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\mathrm\{OGD\}\}\(q\(x\_\{t\+1\}\)\)\-\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\mathrm\{OMD\}\}\(y\_\{t\+1\}\)\\right\\\|\\,\.

To prove this result, we will need the following technical lemma stating monotonicity of normal cones of nonempty closed convex sets, which is a standard result in convex analysis\.

###### Lemma 5\(Monotonicity of the normal cone\)\.

LetK⊂ℝdK\\subset\\mathbb\{R\}^\{d\}be a nonempty closed convex set\. For anyx,y∈Kx,y\\in K, let

NK​\(x\):=\{v∈ℝd:⟨v,z−x⟩≤0,∀z∈K\}N\_\{K\}\(x\):=\\\{v\\in\\mathbb\{R\}^\{d\}:\\langle v,z\-x\\rangle\\leq 0\\,,\\quad\\forall z\\in K\\\}denote the normal cone ofKKatxx\. Then, for anyvx∈NK​\(x\)v\_\{x\}\\in N\_\{K\}\(x\)andvy∈NK​\(y\)v\_\{y\}\\in N\_\{K\}\(y\),

⟨vx−vy,x−y⟩≥0\.\\langle v\_\{x\}\-v\_\{y\},x\-y\\rangle\\geq 0\.

###### Proof\.

Sincevx∈NK​\(x\)v\_\{x\}\\in N\_\{K\}\(x\)andy∈Ky\\in K, the definition of the normal cone gives⟨vx,y−x⟩≤0,\\langle v\_\{x\},y\-x\\rangle\\leq 0,or equivalently,⟨vx,x−y⟩≥0\.\\langle v\_\{x\},x\-y\\rangle\\geq 0\.Similarly, sincevy∈NK​\(y\)v\_\{y\}\\in N\_\{K\}\(y\)andx∈Kx\\in K, we have⟨vy,x−y⟩≤0\.\\langle v\_\{y\},x\-y\\rangle\\leq 0\.Therefore, it follows from the above inequalities that:

⟨vx−vy,x−y⟩=⟨vx,x−y⟩−⟨vy,x−y⟩≥0\.\\langle v\_\{x\}\-v\_\{y\},x\-y\\rangle=\\langle v\_\{x\},x\-y\\rangle\-\\langle v\_\{y\},x\-y\\rangle\\geq 0\\,\.∎

###### Proof of Lemma[4](https://arxiv.org/html/2605.26373#Thmlemma4)\.

Letδt\+1:=yt\+1−q​\(xt\+1\)\\delta\_\{t\+1\}:=y\_\{t\+1\}\-q\(x\_\{t\+1\}\)as a shorthand notation\. From first\-order optimality conditions for \([24](https://arxiv.org/html/2605.26373#A1.E24)\) and \([25](https://arxiv.org/html/2605.26373#A1.E25)\), there existvt\+1OMD∈N𝒴t​\(yt\+1\),vt\+1OGD∈N𝒴t​\(q​\(xt\+1\)\)v\_\{t\+1\}^\{\\mathrm\{OMD\}\}\\in N\_\{\{\\mathcal\{Y\}\}\_\{t\}\}\(y\_\{t\+1\}\),v\_\{t\+1\}^\{\\mathrm\{OGD\}\}\\in N\_\{\{\\mathcal\{Y\}\}\_\{t\}\}\(q\(x\_\{t\+1\}\)\)s\.t\.

∇Φt​\(yt\+1\)\+∇εtOMD​\(yt\+1\)\+vt\+1OMD\\displaystyle\\nabla\\mkern\-2\.5mu\\Phi\_\{t\}\(y\_\{t\+1\}\)\+\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\mathrm\{OMD\}\}\(y\_\{t\+1\}\)\+v\_\{t\+1\}^\{\\mathrm\{OMD\}\}=0,\\displaystyle=0\\,,∇Φt​\(q​\(xt\+1\)\)\+∇εtOGD​\(q​\(xt\+1\)\)\+vt\+1OGD\\displaystyle\\nabla\\mkern\-2\.5mu\\Phi\_\{t\}\(q\(x\_\{t\+1\}\)\)\+\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\mathrm\{OGD\}\}\(q\(x\_\{t\+1\}\)\)\+v\_\{t\+1\}^\{\\mathrm\{OGD\}\}=0\.\\displaystyle=0\\,\.Subtracting both identities and taking inner product withδt\+1\\delta\_\{t\+1\}gives

⟨∇Φt​\(yt\+1\)−∇Φt​\(q​\(xt\+1\)\),δt\+1⟩\+⟨∇εtOMD​\(yt\+1\)−∇εtOGD​\(q​\(xt\+1\)\),δt\+1⟩\+⟨vt\+1OMD−vt\+1OGD,δt\+1⟩=0\.\\big\\langle\\nabla\\mkern\-2\.5mu\\Phi\_\{t\}\(y\_\{t\+1\}\)\-\\nabla\\mkern\-2\.5mu\\Phi\_\{t\}\(q\(x\_\{t\+1\}\)\),\\,\\delta\_\{t\+1\}\\big\\rangle\+\\big\\langle\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\mathrm\{OMD\}\}\(y\_\{t\+1\}\)\-\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\mathrm\{OGD\}\}\(q\(x\_\{t\+1\}\)\),\\,\\delta\_\{t\+1\}\\big\\rangle\\\\ \+\\big\\langle v\_\{t\+1\}^\{\\mathrm\{OMD\}\}\-v\_\{t\+1\}^\{\\mathrm\{OGD\}\},\\,\\delta\_\{t\+1\}\\big\\rangle=0\.\(26\)By1η\\frac\{1\}\{\\eta\}\-strong convexity ofΦt\\Phi\_\{t\}, we have:

⟨∇Φt​\(yt\+1\)−∇Φt​\(q​\(xt\+1\)\),δt\+1⟩≥1η​‖δt\+1‖2\.\\big\\langle\\nabla\\mkern\-2\.5mu\\Phi\_\{t\}\(y\_\{t\+1\}\)\-\\nabla\\mkern\-2\.5mu\\Phi\_\{t\}\(q\(x\_\{t\+1\}\)\),\\,\\delta\_\{t\+1\}\\big\\rangle\\geq\\frac\{1\}\{\\eta\}\\\|\\delta\_\{t\+1\}\\\|^\{2\}\.\(27\)By monotonicity of normal cones of convex sets \(Lemma[5](https://arxiv.org/html/2605.26373#Thmlemma5)above\), we can also control the term involving the normal cone directions as follows:

⟨vt\+1OMD−vt\+1OGD,δt\+1⟩≥0\.\\big\\langle v\_\{t\+1\}^\{\\mathrm\{OMD\}\}\-v\_\{t\+1\}^\{\\mathrm\{OGD\}\},\\,\\delta\_\{t\+1\}\\big\\rangle\\geq 0\.\(28\)Plugging \([27](https://arxiv.org/html/2605.26373#A1.E27)\) and \([28](https://arxiv.org/html/2605.26373#A1.E28)\) into \([26](https://arxiv.org/html/2605.26373#A1.E26)\), we obtain:

1η​‖δt\+1‖2≤‖∇εtOGD​\(q​\(xt\+1\)\)−∇εtOMD​\(yt\+1\)‖⋅‖δt\+1‖\.\\frac\{1\}\{\\eta\}\\\|\\delta\_\{t\+1\}\\\|^\{2\}\\leq\\left\\\|\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\mathrm\{OGD\}\}\(q\(x\_\{t\+1\}\)\)\-\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\mathrm\{OMD\}\}\(y\_\{t\+1\}\)\\right\\\|\\cdot\\\|\\delta\_\{t\+1\}\\\|\.Therefore, it follows that the distanceδt\+1\\delta\_\{t\+1\}between the two minimizers is upperbounded as follows:

‖δt\+1‖≤η​‖∇εtOGD​\(q​\(xt\+1\)\)−∇εtOMD​\(yt\+1\)‖\.\\\|\\delta\_\{t\+1\}\\\|\\leq\\eta\\,\\left\\\|\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\mathrm\{OGD\}\}\(q\(x\_\{t\+1\}\)\)\-\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\mathrm\{OMD\}\}\(y\_\{t\+1\}\)\\right\\\|\.∎

It remains to bound the last norm in the right\-hand side of Lemma[4](https://arxiv.org/html/2605.26373#Thmlemma4)and show it is also of orderη\\eta\. We show that each one of the norms of the gradients‖∇εtOMD​\(yt\+1\)‖\\\|\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\mathrm\{OMD\}\}\(y\_\{t\+1\}\)\\\|and‖∇εtOGD​\(q​\(xt\+1\)\)‖\\\|\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\mathrm\{OGD\}\}\(q\(x\_\{t\+1\}\)\)\\\|are of orderη\\etain the two following lemmas\.

###### Lemma 6\. ‖∇εtOMD​\(yt\+1\)‖≤2​G​G~F2​η\.\\\|\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\mathrm\{OMD\}\}\(y\_\{t\+1\}\)\\\|~\\leq~2G\\tilde\{G\}\_\{F\}^\{2\}\\eta\\,\.

###### Proof\.

Recall thatεtOMD​\(y\)=1η​εt​\(y\),\\varepsilon\_\{t\}^\{\\mathrm\{OMD\}\}\(y\)=\\frac\{1\}\{\\eta\}\\varepsilon\_\{t\}\(y\),where:

εt​\(y\)=DR​\(y∥yt\)−12​\(y−yt\)⊤​∇2R​\(yt\)​\(y−yt\)\.\\varepsilon\_\{t\}\(y\)=D\_\{R\}\(y\\\|y\_\{t\}\)\-\\frac\{1\}\{2\}\(y\-y\_\{t\}\)^\{\\top\}\\nabla\\mkern\-2\.5mu^\{2\}R\(y\_\{t\}\)\(y\-y\_\{t\}\)\\,\.Differentiating the above error function w\.r\.t\.yyyields:

∇εt​\(y\)=∇R​\(y\)−∇R​\(yt\)−∇2R​\(yt\)​\(y−yt\)\.\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}\(y\)=\\nabla\\mkern\-2\.5muR\(y\)\-\\nabla\\mkern\-2\.5muR\(y\_\{t\}\)\-\\nabla\\mkern\-2\.5mu^\{2\}R\(y\_\{t\}\)\(y\-y\_\{t\}\)\.Using third\-order smoothness ofRR\(bounded third derivative\) as supposed in Assumption[2](https://arxiv.org/html/2605.26373#Thmassumption2),

‖∇εt​\(yt\+1\)‖≤G2​‖yt\+1−yt‖2\.\\\|\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}\(y\_\{t\+1\}\)\\\|\\leq\\frac\{G\}\{2\}\\\|y\_\{t\+1\}\-y\_\{t\}\\\|^\{2\}\.\(29\)Sinceyt\+1∈𝒴ty\_\{t\+1\}\\in\{\\mathcal\{Y\}\}\_\{t\}, we have‖yt\+1−yt‖≤2​η​G~F\\\|y\_\{t\+1\}\-y\_\{t\}\\\|~\\leq~2\\eta\\tilde\{G\}\_\{F\}and it follows from \([29](https://arxiv.org/html/2605.26373#A1.E29)\) that:

‖∇εt​\(yt\+1\)‖≤G2​\(2​η​G~F\)2=2​G​G~F2​η2\.\\\|\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}\(y\_\{t\+1\}\)\\\|~\\leq~\\frac\{G\}\{2\}\(2\\eta\\tilde\{G\}\_\{F\}\)^\{2\}=2G\\tilde\{G\}\_\{F\}^\{2\}\\eta^\{2\}\\,\.SinceεtOMD​\(y\)=1η​εt​\(y\),\\varepsilon\_\{t\}^\{\\mathrm\{OMD\}\}\(y\)=\\frac\{1\}\{\\eta\}\\varepsilon\_\{t\}\(y\),we conclude that:

‖∇εtOMD​\(yt\+1\)‖=1η​‖∇εt​\(yt\+1\)‖≤2​G​G~F2​η\.\\\|\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\mathrm\{OMD\}\}\(y\_\{t\+1\}\)\\\|=\\frac\{1\}\{\\eta\}\\\|\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}\(y\_\{t\+1\}\)\\\|~\\leq~2G\\tilde\{G\}\_\{F\}^\{2\}\\eta\\,\.∎

###### Lemma 7\. ‖∇εtOGD​\(q​\(xt\+1\)\)‖≤G5​\(3​G^F2\+G^F3​η\)​η\.\\\|\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\mathrm\{OGD\}\}\(q\(x\_\{t\+1\}\)\)\\\|~\\leq~G^\{5\}\(3\\hat\{G\}\_\{F\}^\{2\}\+\\hat\{G\}\_\{F\}^\{3\}\\eta\)\\eta\.

###### Proof\.

Recall from \([23](https://arxiv.org/html/2605.26373#A1.E23)\) that:

εtOGD​\(y\)=ε~t​\(y\)\+1η​\(12​‖εtq​\(y\)‖22\+⟨Jq​\(xt\)−1​\(y−yt\),εtq​\(y\)⟩\)\.\\varepsilon\_\{t\}^\{\\mathrm\{OGD\}\}\(y\)=\\tilde\{\\varepsilon\}\_\{t\}\(y\)\+\\frac\{1\}\{\\eta\}\\left\(\\frac\{1\}\{2\}\\\|\\varepsilon\_\{t\}^\{q\}\(y\)\\\|\_\{2\}^\{2\}\+\\big\\langle J\_\{q\}\(x\_\{t\}\)^\{\-1\}\(y\-y\_\{t\}\),\\,\\varepsilon\_\{t\}^\{q\}\(y\)\\big\\rangle\\right\)\.Differentiating this error term w\.r\.t\.yygives:

∇εtOGD​\(y\)=∇ε~t​\(y\)\+1η​Jεtq​\(y\)⊤​εtq​\(y\)\+1η​\(\[Jq​\(xt\)−1\]⊤​εtq​\(y\)\+Jεtq​\(y\)⊤​Jq​\(xt\)−1​\(y−yt\)\)\.\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\mathrm\{OGD\}\}\(y\)=\\nabla\\mkern\-2\.5mu\\tilde\{\\varepsilon\}\_\{t\}\(y\)\+\\frac\{1\}\{\\eta\}J\_\{\\varepsilon\_\{t\}^\{q\}\}\(y\)^\{\\top\}\\varepsilon\_\{t\}^\{q\}\(y\)\+\\frac\{1\}\{\\eta\}\\Big\(\[J\_\{q\}\(x\_\{t\}\)^\{\-1\}\]^\{\\top\}\\varepsilon\_\{t\}^\{q\}\(y\)\+J\_\{\\varepsilon\_\{t\}^\{q\}\}\(y\)^\{\\top\}J\_\{q\}\(x\_\{t\}\)^\{\-1\}\(y\-y\_\{t\}\)\\Big\)\.\(30\)Sinceε~t​\(y\):=gt⊤​εtq​\(y\)\\tilde\{\\varepsilon\}\_\{t\}\(y\):=g\_\{t\}^\{\\top\}\\varepsilon\_\{t\}^\{q\}\(y\), we have:

∇ε~t​\(y\)=gt⊤​Jεtq​\(y\),‖∇ε~t​\(y\)‖≤G^F​‖Jεtq​\(y\)‖\.\\nabla\\mkern\-2\.5mu\\tilde\{\\varepsilon\}\_\{t\}\(y\)=g\_\{t\}^\{\\top\}J\_\{\\varepsilon\_\{t\}^\{q\}\}\(y\)\\,,\\quad\\\|\\nabla\\mkern\-2\.5mu\\tilde\{\\varepsilon\}\_\{t\}\(y\)\\\|~\\leq~\\hat\{G\}\_\{F\}\\\|J\_\{\\varepsilon\_\{t\}^\{q\}\}\(y\)\\\|\\,\.
We now prove the following estimates:

1. \(i\)‖q​\(xt\+1\)−yt‖≤G​G^F​η,\\\|q\(x\_\{t\+1\}\)\-y\_\{t\}\\\|~\\leq~G\\hat\{G\}\_\{F\}\\eta\\,,
2. \(ii\)‖Jεtq​\(q​\(xt\+1\)\)‖≤G2​G^F​η,\\\|J\_\{\\varepsilon\_\{t\}^\{q\}\}\(q\(x\_\{t\+1\}\)\)\\\|~\\leq~G^\{2\}\\hat\{G\}\_\{F\}\\eta,
3. \(iii\)‖εtq​\(q​\(xt\+1\)\)‖≤G3​G^F22​η2\.\\\|\\varepsilon\_\{t\}^\{q\}\(q\(x\_\{t\+1\}\)\)\\\|~\\leq~\\frac\{G^\{3\}\\hat\{G\}\_\{F\}^\{2\}\}\{2\}\\eta^\{2\}\\,\.

Proof of[\(i\)](https://arxiv.org/html/2605.26373#A1.I1.i1)\.We control this first error term as follows:

‖q​\(xt\+1\)−yt‖\\displaystyle\\\|q\(x\_\{t\+1\}\)\-y\_\{t\}\\\|=‖q​\(xt\+1\)−q​\(xt\)‖\\displaystyle=\\\|q\(x\_\{t\+1\}\)\-q\(x\_\{t\}\)\\\|\(sinceyt=q​\(xt\)y\_\{t\}=q\(x\_\{t\}\)\)≤G​‖xt\+1−xt‖\\displaystyle~\\leq~G\\\|x\_\{t\+1\}\-x\_\{t\}\\\|\(qqisGG\-Lipschitz by Assumption[2](https://arxiv.org/html/2605.26373#Thmassumption2)\)=G​‖Π𝒳δ​\(xt−η​gt\)−xt‖\\displaystyle=G\\\|\\Pi\_\{\{\\mathcal\{X\}\}\_\{\\delta\}\}\(x\_\{t\}\-\\eta g\_\{t\}\)\-x\_\{t\}\\\|\(by definition ofxt\+1x\_\{t\+1\}in \([OGD](https://arxiv.org/html/2605.26373#S2.Ex2)\) or BOGD\)=G​‖Π𝒳δ​\(xt−η​gt\)−Π𝒳δ​\(xt\)‖\\displaystyle=G\\\|\\Pi\_\{\{\\mathcal\{X\}\}\_\{\\delta\}\}\(x\_\{t\}\-\\eta g\_\{t\}\)\-\\Pi\_\{\{\\mathcal\{X\}\}\_\{\\delta\}\}\(x\_\{t\}\)\\\|\(sincext∈𝒳δx\_\{t\}\\in\{\\mathcal\{X\}\}\_\{\\delta\},Π𝒳δ​\(xt\)=xt\\Pi\_\{\{\\mathcal\{X\}\}\_\{\\delta\}\}\(x\_\{t\}\)=x\_\{t\}\)≤η​G​‖gt‖\\displaystyle~\\leq~\\eta G\\\|g\_\{t\}\\\|\(by nonexpansiveness of the projection\)≤η​G​G^F=η​G~F\.\\displaystyle~\\leq~\\eta G\\hat\{G\}\_\{F\}=\\eta\\tilde\{G\}\_\{F\}\\,\.\(using Assumption[3](https://arxiv.org/html/2605.26373#Thmassumption3)\)\(31\)
Proof of[\(ii\)](https://arxiv.org/html/2605.26373#A1.I1.i2)\.Recall from \([16](https://arxiv.org/html/2605.26373#A1.E16)\) that for anyy=q​\(x\),x∈𝒳,y=q\(x\),x\\in\{\\mathcal\{X\}\}\\,,

εtq​\(y\)=q−1​\(y\)−q−1​\(yt\)−Jq−1​\(yt\)​\(y−yt\)\.\\varepsilon\_\{t\}^\{q\}\(y\)=q^\{\-1\}\(y\)\-q^\{\-1\}\(y\_\{t\}\)\-J\_\{q^\{\-1\}\}\(y\_\{t\}\)\(y\-y\_\{t\}\)\\,\.\(32\)Differentiating w\.r\.t\.yyyields:

Jεtq​\(q​\(xt\+1\)\)=Jq−1​\(q​\(xt\+1\)\)−Jq−1​\(yt\)\.J\_\{\\varepsilon\_\{t\}^\{q\}\}\(q\(x\_\{t\+1\}\)\)=J\_\{q^\{\-1\}\}\(q\(x\_\{t\+1\}\)\)\-J\_\{q^\{\-1\}\}\(y\_\{t\}\)\\,\.Using Assumption[2](https://arxiv.org/html/2605.26373#Thmassumption2), we obtain:

‖Jεtq​\(q​\(xt\+1\)\)‖=‖Jq−1​\(q​\(xt\+1\)\)−Jq−1​\(yt\)‖≤G​‖q​\(xt\+1\)−yt‖≤η​G2​G^F,\\\|J\_\{\\varepsilon\_\{t\}^\{q\}\}\(q\(x\_\{t\+1\}\)\)\\\|=\\\|J\_\{q^\{\-1\}\}\(q\(x\_\{t\+1\}\)\)\-J\_\{q^\{\-1\}\}\(y\_\{t\}\)\\\|~\\leq~G\\\|q\(x\_\{t\+1\}\)\-y\_\{t\}\\\|~\\leq~\\eta G^\{2\}\\hat\{G\}\_\{F\}\\,,where the last step uses the first estimate[\(i\)](https://arxiv.org/html/2605.26373#A1.I1.i1)proved above\.

Proof of[\(iii\)](https://arxiv.org/html/2605.26373#A1.I1.i3)\.Recalling \([32](https://arxiv.org/html/2605.26373#A1.E32)\) and using boundedness of the second\-order derivative ofq−1q^\{\-1\}\(Assumption[2](https://arxiv.org/html/2605.26373#Thmassumption2)\), we immediately obtain:

‖εtq​\(q​\(xt\+1\)\)‖≤G2⋅‖q​\(xt\+1\)−yt‖2≤G3​G^F22​η2,\\\|\\varepsilon\_\{t\}^\{q\}\(q\(x\_\{t\+1\}\)\)\\\|~\\leq~\\frac\{G\}\{2\}\\cdot\\\|q\(x\_\{t\+1\}\)\-y\_\{t\}\\\|^\{2\}~\\leq~\\frac\{G^\{3\}\\hat\{G\}\_\{F\}^\{2\}\}\{2\}\\eta^\{2\}\\,,where the last estimate follows from using the first proven estimate[\(i\)](https://arxiv.org/html/2605.26373#A1.I1.i1):‖q​\(xt\+1\)−yt‖≤G​G^F​η\.\\\|q\(x\_\{t\+1\}\)\-y\_\{t\}\\\|~\\leq~G\\hat\{G\}\_\{F\}\\eta\\,\.

We conclude the proof of Lemma[7](https://arxiv.org/html/2605.26373#Thmlemma7)by using all the estimates[\(i\)](https://arxiv.org/html/2605.26373#A1.I1.i1)to[\(iii\)](https://arxiv.org/html/2605.26373#A1.I1.i3)to bound the norm of the gradient∇εtOGD​\(x\)\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\mathrm\{OGD\}\}\(x\)in \([30](https://arxiv.org/html/2605.26373#A1.E30)\) as follows:

‖∇εtOGD​\(q​\(xt\+1\)\)‖\\displaystyle\\\|\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\mathrm\{OGD\}\}\(q\(x\_\{t\+1\}\)\)\\\|≤G^F​‖Jεtq​\(q​\(xt\+1\)\)‖\+1η​‖Jεtq​\(q​\(xt\+1\)\)‖⋅‖εtq​\(q​\(xt\+1\)\)‖\\displaystyle~\\leq~\\hat\{G\}\_\{F\}\\\|J\_\{\\varepsilon\_\{t\}^\{q\}\}\(q\(x\_\{t\+1\}\)\)\\\|\+\\frac\{1\}\{\\eta\}\\\|J\_\{\\varepsilon\_\{t\}^\{q\}\}\(q\(x\_\{t\+1\}\)\)\\\|\\cdot\\\|\\varepsilon\_\{t\}^\{q\}\(q\(x\_\{t\+1\}\)\)\\\|\+1η​G​‖εtq​\(q​\(xt\+1\)\)‖\+1η​G​‖Jεtq​\(q​\(xt\+1\)\)‖⋅‖q​\(xt\+1\)−yt‖\\displaystyle\+\\frac\{1\}\{\\eta\}G\\\|\\varepsilon\_\{t\}^\{q\}\(q\(x\_\{t\+1\}\)\)\\\|\+\\frac\{1\}\{\\eta\}G\\\|J\_\{\\varepsilon\_\{t\}^\{q\}\}\(q\(x\_\{t\+1\}\)\)\\\|\\cdot\\\|q\(x\_\{t\+1\}\)\-y\_\{t\}\\\|≤G2​G^F2​η\+12​G5​G^F3​η2\+12​G4​G^F2​η\+G4​G^F2​η\\displaystyle~\\leq~G^\{2\}\\hat\{G\}\_\{F\}^\{2\}\\eta\+\\frac\{1\}\{2\}G^\{5\}\\hat\{G\}\_\{F\}^\{3\}\\eta^\{2\}\+\\frac\{1\}\{2\}G^\{4\}\\hat\{G\}\_\{F\}^\{2\}\\eta\+G^\{4\}\\hat\{G\}\_\{F\}^\{2\}\\eta\\,≤G5​\(3​G^F2\+G^F3​η\)​η,\\displaystyle~\\leq~G^\{5\}\(3\\hat\{G\}\_\{F\}^\{2\}\+\\hat\{G\}\_\{F\}^\{3\}\\eta\)\\eta\\,,\(33\)where the last inequality follows from using the assumptionG≥1G~\\geq~1which is without loss of generality \(otherwise replace the constant by 1\)\. ∎

End of Proof of Proposition[2](https://arxiv.org/html/2605.26373#Thmproposition2)\.Using Lemma[4](https://arxiv.org/html/2605.26373#Thmlemma4), we have:

‖yt\+1−q​\(xt\+1\)‖≤η​‖∇εtOGD​\(q​\(xt\+1\)\)−∇εtOMD​\(yt\+1\)‖\.\\\|y\_\{t\+1\}\-q\(x\_\{t\+1\}\)\\\|\\leq\\eta\\,\\left\\\|\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\mathrm\{OGD\}\}\(q\(x\_\{t\+1\}\)\)\-\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\mathrm\{OMD\}\}\(y\_\{t\+1\}\)\\right\\\|\\,\.Then it follows from using Lemma[6](https://arxiv.org/html/2605.26373#Thmlemma6)and Lemma[7](https://arxiv.org/html/2605.26373#Thmlemma7)that:

‖yt\+1−q​\(xt\+1\)‖=‖δt\+1‖\\displaystyle\\\|y\_\{t\+1\}\-q\(x\_\{t\+1\}\)\\\|=\\\|\\delta\_\{t\+1\}\\\|≤η⋅\(‖∇εtOGD​\(q​\(xt\+1\)\)‖\+‖∇εtOMD​\(yt\+1\)‖\)\\displaystyle\\leq\\eta\\cdot\\left\(\\\|\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\mathrm\{OGD\}\}\(q\(x\_\{t\+1\}\)\)\\\|\+\\\|\\nabla\\mkern\-2\.5mu\\varepsilon\_\{t\}^\{\\mathrm\{OMD\}\}\(y\_\{t\+1\}\)\\\|\\right\)≤\(G5​\(3​G^F2\+G^F3​η\)\+2​G​G~F2\)​η2\\displaystyle~\\leq~\(G^\{5\}\(3\\hat\{G\}\_\{F\}^\{2\}\+\\hat\{G\}\_\{F\}^\{3\}\\eta\)\+2G\\tilde\{G\}\_\{F\}^\{2\}\)\\eta^\{2\}=\(G5​\(3​G^F2\+G^F3​η\)\+2​G3​G^F2\)​η2\\displaystyle=\(G^\{5\}\(3\\hat\{G\}\_\{F\}^\{2\}\+\\hat\{G\}\_\{F\}^\{3\}\\eta\)\+2G^\{3\}\\hat\{G\}\_\{F\}^\{2\}\)\\eta^\{2\}≤G5​\(5​G^F2\+G^F3​η\)​η2,\\displaystyle~\\leq~G^\{5\}\(5\\hat\{G\}\_\{F\}^\{2\}\+\\hat\{G\}\_\{F\}^\{3\}\\eta\)\\eta^\{2\}\\,,\(34\)which concludes the proof\. ∎

### A\.2Proof of Theorem[1](https://arxiv.org/html/2605.26373#Thmtheorem1)using Lemma[1](https://arxiv.org/html/2605.26373#Thmlemma1)

The following lemma controls the regret of an approximate OMD algorithm\.

###### Lemma 8\(Lemma 6 inGhaiet al\.\[[2022](https://arxiv.org/html/2605.26373#bib.bib46)\]\)\.

Suppose Assumptions[2](https://arxiv.org/html/2605.26373#Thmassumption2)and[3](https://arxiv.org/html/2605.26373#Thmassumption3)hold and let Algorithm𝒜\\mathcal\{A\}update the sequence\(zt\)\(z\_\{t\}\)as follows:

zt\+1=rt\+1\+arg​miny∈𝒴⁡\{∇ht​\(yt\)⊤​\(y−yt\)\+1η​DR​\(y∥yt\)\},z\_\{t\+1\}=r\_\{t\+1\}\+\\operatorname\*\{arg\\,min\}\_\{y\\in\{\\mathcal\{Y\}\}\}\\left\\\{\\nabla\\mkern\-2\.5muh\_\{t\}\(y\_\{t\}\)^\{\\top\}\(y\-y\_\{t\}\)\+\\frac\{1\}\{\\eta\}D\_\{R\}\(y\\\|y\_\{t\}\)\\right\\\}\\,,where‖rt\+1‖2≤Cη\\\|r\_\{t\+1\}\\\|\_\{2\}\\leq C\_\{\\eta\}\. Then the regret of Algorithm𝒜\\mathcal\{A\}is upper bounded as follows:

RT​\(𝒜\)≤Cη​T​Gη\+D1η\+η​G2​T2\.R\_\{T\}\(\\mathcal\{A\}\)\\leq\\frac\{C\_\{\\eta\}TG\}\{\\eta\}\+\\frac\{D\_\{1\}\}\{\\eta\}\+\\frac\{\\eta G^\{2\}T\}\{2\}\\,\.

End of Proof of Theorem[1](https://arxiv.org/html/2605.26373#Thmtheorem1)\.The rest of the proof follows the same lines as the proof ofGhaiet al\.\[[2022](https://arxiv.org/html/2605.26373#bib.bib46)\]using our tighter estimate obtained in Corollary[4](https://arxiv.org/html/2605.26373#Thmtheorem4)\. We provide a full proof for completeness\.

###### Proof\.

The \([OGD](https://arxiv.org/html/2605.26373#S2.Ex2)\) iterates mapped to the hidden space𝒴\{\\mathcal\{Y\}\}can be seen as a perturbed version of OMD since:

zt\+1=q​\(xt\+1\)=q​\(xt\+1\)−yt\+1\+yt\+1=rt\+1\+arg​miny∈𝒴⁡∇⁡ht​\(yt\)⊤​\(y−yt\)\+1η​DR​\(y∥yt\),z\_\{t\+1\}=q\(x\_\{t\+1\}\)=q\(x\_\{t\+1\}\)\-y\_\{t\+1\}\+y\_\{t\+1\}=r\_\{t\+1\}\+\\operatorname\*\{arg\\,min\}\_\{y\\in\{\\mathcal\{Y\}\}\}\\nabla\\mkern\-2\.5muh\_\{t\}\(y\_\{t\}\)^\{\\top\}\(y\-y\_\{t\}\)\+\\frac\{1\}\{\\eta\}D\_\{R\}\(y\\\|y\_\{t\}\)\\,,wherert\+1:=q​\(xt\+1\)−yt\+1\.r\_\{t\+1\}:=q\(x\_\{t\+1\}\)\-y\_\{t\+1\}\\,\.It follows from Corollary[4](https://arxiv.org/html/2605.26373#Thmtheorem4)that‖rt\+1‖≤Cη:=6​G5​GF3​η2\.\\\|r\_\{t\+1\}\\\|~\\leq~C\_\{\\eta\}:=6G^\{5\}G\_\{F\}^\{3\}\\eta^\{2\}\\,\.Using now Lemma[8](https://arxiv.org/html/2605.26373#Thmlemma8), we obtain the following regret bound for the \([OGD](https://arxiv.org/html/2605.26373#S2.Ex2)\) iterates\(zt\)\(z\_\{t\}\)in the hidden space𝒴\{\\mathcal\{Y\}\}:

RT=∑t=1Tht​\(zt\)−minz∈𝒴​∑t=1Tht​\(z\)≤Cη​T​Gη\+D1η\+η​G2​T2≤D1η\+7​G6​GF3​η​T\.R\_\{T\}=\\sum\_\{t=1\}^\{T\}h\_\{t\}\(z\_\{t\}\)\-\\min\_\{z\\in\{\\mathcal\{Y\}\}\}\\sum\_\{t=1\}^\{T\}h\_\{t\}\(z\)\\leq\\frac\{C\_\{\\eta\}TG\}\{\\eta\}\+\\frac\{D\_\{1\}\}\{\\eta\}\+\\frac\{\\eta G^\{2\}T\}\{2\}~\\leq~\\frac\{D\_\{1\}\}\{\\eta\}\+7G^\{6\}G\_\{F\}^\{3\}\\eta T\\,\.Optimizing the stepsizeη\\etato minimize the upperbound yields:

RT≤7​D1​G6​GF3​T,R\_\{T\}~\\leq~\\sqrt\{7D\_\{1\}G^\{6\}G\_\{F\}^\{3\}T\}\\,,with stepsizeη=D17​G6​GF3​T\.\\eta=\\sqrt\{\\frac\{D\_\{1\}\}\{7G^\{6\}G\_\{F\}^\{3\}T\}\}\\,\.∎

## Appendix BProofs for Section[4](https://arxiv.org/html/2605.26373#S4)

### B\.1Proof of Proposition[1](https://arxiv.org/html/2605.26373#Thmproposition1)

The proof relies on the following classical result from vector calculus\.

###### Proposition 3\. Let𝒰⊆ℝd\\mathcal\{U\}\\subseteq\\mathbb\{R\}^\{d\}be convex222A simple connected domain suffices here\. Convexity is enough for our purpose\.and letF:𝒰→ℝdF:\\mathcal\{U\}\\to\\mathbb\{R\}^\{d\}be aC1C^\{1\}vector field\. DenotingF=\(F1,⋯,Fd\)F=\(F\_\{1\},\\cdots,F\_\{d\}\)using its coordinate functions, if∂xjFk​\(x\)=∂xkFj​\(x\)\\partial\_\{x\_\{j\}\}F\_\{k\}\(x\)=\\partial\_\{x\_\{k\}\}F\_\{j\}\(x\)for allj,k∈\[d\]j,k\\in\[d\]and allx∈𝒰x\\in\\mathcal\{U\}, then there exists a scalar functionϕ:ℝd→ℝ\\phi:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}such thatF=∇ϕ\.F=\\nabla\\mkern\-2\.5mu\\phi\\,\.

###### Proof of Prop\.[1](https://arxiv.org/html/2605.26373#Thmproposition1)\.

For eachi∈\[d\]i\\in\[d\], define the vector fieldM\(i\)M^\{\(i\)\}of theiith row of the matrix fieldMM:

M\(i\)​\(x\):=\(Mi,1​\(x\),⋯,Mi,d​\(x\)\)\.M^\{\(i\)\}\(x\):=\(M\_\{i,1\}\(x\),\\cdots,M\_\{i,d\}\(x\)\)\\,\.The cross\-derivatives conditions∂xkMi​j​\(x\)=∂xjMi​k​\(x\),∀i,j,k∈\[d\],∀x∈K,\\partial\_\{x\_\{k\}\}M\_\{ij\}\(x\)=\\partial\_\{x\_\{j\}\}M\_\{ik\}\(x\)\\,,\\forall i,j,k\\in\[d\],\\forall x\\in K\\,,guarantee thatM\(i\)M^\{\(i\)\}has also matching cross\-partial derivatives in the sense of Proposition[3](https://arxiv.org/html/2605.26373#Thmproposition3)\. Therefore, it follows from Proposition[3](https://arxiv.org/html/2605.26373#Thmproposition3)that there exists a scalar functiongi:ℝd→ℝg\_\{i\}:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}such that∇gi=F\(i\)\\nabla\\mkern\-2\.5mug\_\{i\}=F^\{\(i\)\}for alli∈\[d\]\.i\\in\[d\]\\,\.

We obtain a vector fieldg=\(g1,⋯,gd\)g=\(g\_\{1\},\\cdots,g\_\{d\}\)whose JacobianJgJ\_\{g\}coincides withMM, i\.e\.Jg=MJ\_\{g\}=MsinceM\(i\)M^\{\(i\)\}are the rows of the matrix fieldMM\. Then, observe thatMMis symmetric becauseJq​Jq⊤J\_\{q\}J\_\{q\}^\{\\top\}is symmetric and its inverse is also symmetric\. Hence, the JacobianJg=MJ\_\{g\}=Mis also symmetric and it follows thatggitself has matching cross partial derivatives\. Applying Proposition[3](https://arxiv.org/html/2605.26373#Thmproposition3)again toggyields the existence of a scalar functionRRsuch thatg=∇R\.g=\\nabla\\mkern\-2\.5muR\\,\.We have proved thatM=JgM=J\_\{g\}andg=∇Rg=\\nabla\\mkern\-2\.5muRwhich implies thatM=∇2RM=\\nabla\\mkern\-2\.5mu^\{2\}R, concluding the proof\. ∎

### B\.2Detailed derivations for examples

#### Example 1: Affine mixing of a separable reparameterization\.

Let

q​\(x\)=A​s​\(x\)\+b,s​\(x\)=\(s1​\(x1\),…,sd​\(xd\)\),q\(x\)=A\\,s\(x\)\+b,\\qquad s\(x\)=\\bigl\(s\_\{1\}\(x\_\{1\}\),\\dots,s\_\{d\}\(x\_\{d\}\)\\bigr\),whereA∈ℝd×dA\\in\\mathbb\{R\}^\{d\\times d\}is invertible,b∈ℝdb\\in\\mathbb\{R\}^\{d\}, and eachsi:ℝ→ℝs\_\{i\}:\\mathbb\{R\}\\to\\mathbb\{R\}is aC2C^\{2\}diffeomorphism\. Writing

D​\(x\):=diag⁡\(s1′​\(x1\),…,sd′​\(xd\)\),D\(x\):=\\operatorname\{diag\}\\bigl\(s\_\{1\}^\{\\prime\}\(x\_\{1\}\),\\dots,s\_\{d\}^\{\\prime\}\(x\_\{d\}\)\\bigr\),it follows that the Jacobian ofqqatxxand its outer product can be written as:

Jq​\(x\)=A​D​\(x\),Jq​\(x\)​Jq​\(x\)⊤=A​D​\(x\)2​A⊤\.J\_\{q\}\(x\)=AD\(x\),\\qquad J\_\{q\}\(x\)J\_\{q\}\(x\)^\{\\top\}=AD\(x\)^\{2\}A^\{\\top\}\.Therefore, forz=q​\(x\)z=q\(x\),

M​\(z\)=\[Jq​\(q−1​\(z\)\)​Jq​\(q−1​\(z\)\)⊤\]−1=A−T​D​\(x\)−2​A−1,x=q−1​\(z\)\.M\(z\)=\\left\[J\_\{q\}\(q^\{\-1\}\(z\)\)J\_\{q\}\(q^\{\-1\}\(z\)\)^\{\\top\}\\right\]^\{\-1\}=A^\{\-T\}D\(x\)^\{\-2\}A^\{\-1\},\\qquad x=q^\{\-1\}\(z\)\.
Letξ:=A−1​\(z−b\)\.\\xi:=A^\{\-1\}\(z\-b\)\.Sincez=A​s​\(x\)\+bz=As\(x\)\+b, we haveξ=s​\(x\)\\xi=s\(x\), hencexi=si−1​\(ξi\)x\_\{i\}=s\_\{i\}^\{\-1\}\(\\xi\_\{i\}\)\. Define

mi​\(ξi\):=1\(si′​\(si−1​\(ξi\)\)\)2\.m\_\{i\}\(\\xi\_\{i\}\):=\\frac\{1\}\{\\bigl\(s\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\-1\}\(\\xi\_\{i\}\)\)\\bigr\)^\{2\}\}\.Using this notation, we can rewriteM​\(z\)M\(z\)as follows:

M​\(z\)=A−T​diag⁡\(m1​\(ξ1\),…,md​\(ξd\)\)​A−1,ξ=A−1​\(z−b\)\.M\(z\)=A^\{\-T\}\\operatorname\{diag\}\\bigl\(m\_\{1\}\(\\xi\_\{1\}\),\\dots,m\_\{d\}\(\\xi\_\{d\}\)\\bigr\)A^\{\-1\},\\qquad\\xi=A^\{\-1\}\(z\-b\)\.
This matrix field is explicitly Hessian\. Indeed, definerir\_\{i\}by

ri′′​\(t\)=mi​\(t\)=1\(si′​\(si−1​\(t\)\)\)2,r\_\{i\}^\{\\prime\\prime\}\(t\)=m\_\{i\}\(t\)=\\frac\{1\}\{\\bigl\(s\_\{i\}^\{\\prime\}\(s\_\{i\}^\{\-1\}\(t\)\)\\bigr\)^\{2\}\},and set

R​\(z\):=∑i=1dri​\(\(A−1​\(z−b\)\)i\)\.R\(z\):=\\sum\_\{i=1\}^\{d\}r\_\{i\}\\bigl\(\(A^\{\-1\}\(z\-b\)\)\_\{i\}\\bigr\)\.SinceA−1​\(z−b\)A^\{\-1\}\(z\-b\)is affine inzz,

∇2R​\(z\)=A−T​diag⁡\(r1′′​\(ξ1\),…,rd′′​\(ξd\)\)​A−1=M​\(z\)\.\\nabla\\mkern\-2\.5mu^\{2\}R\(z\)=A^\{\-T\}\\operatorname\{diag\}\\bigl\(r\_\{1\}^\{\\prime\\prime\}\(\\xi\_\{1\}\),\\dots,r\_\{d\}^\{\\prime\\prime\}\(\\xi\_\{d\}\)\\bigr\)A^\{\-1\}=M\(z\)\.ThusMMsatisfies the Hessian\-compatibility condition\. In particular, the compatibility holds althoughJq​\(x\)=A​D​\(x\)J\_\{q\}\(x\)=AD\(x\)is not diagonal unlessAAis diagonal\.

#### Example 2: Rank\-one nonlinear mixing\.

Fixa∈ℝd∖\{0\}a\\in\\mathbb\{R\}^\{d\}\\setminus\\\{0\\\}and let

q​\(x\)=x\+h​\(a⊤​x\)​a,q\(x\)=x\+h\(a^\{\\top\}x\)a,whereh:ℝ→ℝh:\\mathbb\{R\}\\to\\mathbb\{R\}isC2C^\{2\}\. Writingr:=a⊤​xr:=a^\{\\top\}xands:=‖a‖2s:=\\\|a\\\|^\{2\}, we have

Jq​\(x\)=I\+h′​\(r\)​a​a⊤\.J\_\{q\}\(x\)=I\+h^\{\\prime\}\(r\)aa^\{\\top\}\.This Jacobian is generally non\-diagonal whenever at least two components ofaaare nonzero\.

The matrixJq​\(x\)J\_\{q\}\(x\)acts as multiplication by1\+s​h′​\(r\)1\+sh^\{\\prime\}\(r\)in the directionaa, and as the identity ona⟂a^\{\\perp\}\. Assume1\+s​h′​\(r\)≠01\+sh^\{\\prime\}\(r\)\\neq 0on the domain, so thatqqis locally invertible\. ThenJq​\(x\)​Jq​\(x\)⊤J\_\{q\}\(x\)J\_\{q\}\(x\)^\{\\top\}has eigenvalue\(1\+s​h′​\(r\)\)2\(1\+sh^\{\\prime\}\(r\)\)^\{2\}alongaaand eigenvalue11ona⟂a^\{\\perp\}\. Consequently, the inverse metric field has the form

M​\(z\)=I\+β​\(a⊤​z\)​a​a⊤,z=q​\(x\),M\(z\)=I\+\\beta\(a^\{\\top\}z\)\\,aa^\{\\top\},\\qquad z=q\(x\),for a scalar functionβ\\beta\. To identify it, note that

a⊤​z=a⊤​q​\(x\)=r\+s​h​\(r\)\.a^\{\\top\}z=a^\{\\top\}q\(x\)=r\+sh\(r\)\.Letg​\(r\):=r\+s​h​\(r\)g\(r\):=r\+sh\(r\), and assumeggis invertible on the relevant interval\. Then

r=g−1​\(a⊤​z\),r=g^\{\-1\}\(a^\{\\top\}z\),and the eigenvalue ofM​\(z\)M\(z\)alongaagives

1\+s​β​\(a⊤​z\)=\(1\+s​h′​\(r\)\)−2\.1\+s\\,\\beta\(a^\{\\top\}z\)=\\bigl\(1\+sh^\{\\prime\}\(r\)\\bigr\)^\{\-2\}\.Thus we can deduce an expression forβ\\betaas follows:

β​\(t\)=\(1\+s​h′​\(g−1​\(t\)\)\)−2−1s\.\\beta\(t\)=\\frac\{\\bigl\(1\+sh^\{\\prime\}\(g^\{\-1\}\(t\)\)\\bigr\)^\{\-2\}\-1\}\{s\}\.
AgainMMis explicitly Hessian\. Letψ:ℝ→ℝ\\psi:\\mathbb\{R\}\\to\\mathbb\{R\}satisfyψ′′​\(t\)=β​\(t\)\\psi^\{\\prime\\prime\}\(t\)=\\beta\(t\), and define:

R​\(z\):=12​‖z‖2\+ψ​\(a⊤​z\)\.R\(z\):=\\frac\{1\}\{2\}\\\|z\\\|^\{2\}\+\\psi\(a^\{\\top\}z\)\.Then, it follows from the above definition that the Hessian can be written as follows:

∇2R​\(z\)=I\+ψ′′​\(a⊤​z\)​a​a⊤=I\+β​\(a⊤​z\)​a​a⊤=M​\(z\)\.\\nabla\\mkern\-2\.5mu^\{2\}R\(z\)=I\+\\psi^\{\\prime\\prime\}\(a^\{\\top\}z\)aa^\{\\top\}=I\+\\beta\(a^\{\\top\}z\)aa^\{\\top\}=M\(z\)\.Hence the Hessian\-compatibility condition holds, despite the generally non\-diagonal JacobianJq​\(x\)J\_\{q\}\(x\)\.

### B\.3Proof of Theorem[2](https://arxiv.org/html/2605.26373#Thmtheorem2): Linear Regret Lower Bound for OGD

### B\.4Detailed proof sketch of Theorem[2](https://arxiv.org/html/2605.26373#Thmtheorem2)

The proof consists of several steps\. We give an overview of each one of them below\. The complete proof with all the details can be found in Section[B\.3](https://arxiv.org/html/2605.26373#A2.SS3)\.

Step 1\. Parameterizationqqand induced metric\. Let𝒳:=\[0,log⁡3\]×\[π6,π3\]⊂ℝ2\.\\mathcal\{X\}:=\[0,\\log 3\]\\times\[\\frac\{\\pi\}\{6\},\\frac\{\\pi\}\{3\}\]\\subset\\mathbb\{R\}^\{2\}\.Define the parameterization functionq:𝒳→ℝ2q:\\mathcal\{X\}\\to\\mathbb\{R\}^\{2\}for anyx=\(x1,x2\)∈𝒳x=\(x\_\{1\},x\_\{2\}\)\\in\\mathcal\{X\}as follows:q​\(x1,x2\):=ex1​\(cos⁡x2,sin⁡x2\)\.q\(x\_\{1\},x\_\{2\}\):=e^\{x\_\{1\}\}\(\\cos x\_\{2\},\\sin x\_\{2\}\)\\,\.One can then easily verify thatJq​\(x\)​Jq​\(x\)⊤=e2​x1​I2J\_\{q\}\(x\)J\_\{q\}\(x\)^\{\\top\}=e^\{2x\_\{1\}\}I\_\{2\}whereI2I\_\{2\}is the2×22\\times 2identity matrix\. Moreover, sincex1≥0x\_\{1\}~\\geq~0,e2​x1≥1e^\{2x\_\{1\}\}~\\geq~1and henceJq​\(x\)​Jq​\(x\)⊤⪰I2J\_\{q\}\(x\)J\_\{q\}\(x\)^\{\\top\}\\succeq I\_\{2\}\.z:=q​\(x\)z:=q\(x\), we have‖z‖=‖q​\(x\)‖=ex1\\\|z\\\|=\\\|q\(x\)\\\|=e^\{x\_\{1\}\}and thereforeG​\(z\):=M​\(z\)−1=1‖z‖2​I2\.G\(z\):=M\(z\)^\{\-1\}=\\frac\{1\}\{\\\|z\\\|^\{2\}\}I\_\{2\}\\,\.

Step 2: OGD dynamics in the hidden space\.To construct our adversarial sequence of losses, we first write the OGD dynamics in the hidden space𝒴:=q​\(𝒳\)\.\{\\mathcal\{Y\}\}:=q\(\\mathcal\{X\}\)\.We will consider linear functionsht:𝒴→ℝh\_\{t\}:\{\\mathcal\{Y\}\}\\to\\mathbb\{R\}defined for everyz∈𝒵z\\in\\mathcal\{Z\}byht​\(z\)=⟨st,z⟩h\_\{t\}\(z\)=\\langle s\_\{t\},z\\ranglewhere\(st\)\(s\_\{t\}\)is an adversarial sequence that will be specified later in an adversarial way\. It follows that for anyx∈𝒳x\\in\\mathcal\{X\},ℓt​\(x\)=ht​\(q​\(x\)\)=⟨st,q​\(x\)⟩,\\ell\_\{t\}\(x\)=h\_\{t\}\(q\(x\)\)=\\langle s\_\{t\},q\(x\)\\rangle\\,,and∇ℓt​\(x\)=Jq​\(x\)⊤​st\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\(x\)=J\_\{q\}\(x\)^\{\\top\}s\_\{t\}\. The OGD update rule on the sequence\(ft\)\(f\_\{t\}\)can then be written as follows:

xt\+1=xt−η​∇ℓt​\(xt\)=xt−η​Jq​\(xt\)⊤​st\.x\_\{t\+1\}=x\_\{t\}\-\\eta\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\(x\_\{t\}\)=x\_\{t\}\-\\eta J\_\{q\}\(x\_\{t\}\)^\{\\top\}s\_\{t\}\\,\.\(35\)Now defining the sequence\(zt\)\(z\_\{t\}\)byzt:=q​\(xt\)z\_\{t\}:=q\(x\_\{t\}\)for anyttin the hidden space𝒵\\mathcal\{Z\}, we obtain the following using Taylor’s theorem with exact remainder:

zt\+1=q​\(xt\+1\)=q​\(xt\)\+Jq​\(xt\)​\(xt\+1−xt\)\+rt=zt−η​G​\(zt\)−1​st\+rt,z\_\{t\+1\}=q\(x\_\{t\+1\}\)=q\(x\_\{t\}\)\+J\_\{q\}\(x\_\{t\}\)\(x\_\{t\+1\}\-x\_\{t\}\)\+r\_\{t\}=z\_\{t\}\-\\eta G\(z\_\{t\}\)^\{\-1\}s\_\{t\}\+r\_\{t\}\\,,\(36\)where‖rt‖=𝒪​\(η2\)\.\\\|r\_\{t\}\\\|=\\mathcal\{O\}\(\\eta^\{2\}\)\\,\.To simplify the OGD dynamics in the hidden space, we choosest=−G​\(zt\)​vts\_\{t\}=\-G\(z\_\{t\}\)v\_\{t\}where\(vt\)\(v\_\{t\}\)is an adversarial sequence that will be chosen later\. Given this choice, we obtain the following OGD dynamics in the hidden space:

zt\+1=zt\+η​vt\+rt\.z\_\{t\+1\}=z\_\{t\}\+\\eta v\_\{t\}\+r\_\{t\}\\,\.\(37\)
Step 3\. Regret as an approximate discrete path sum\.For a fixed comparatoru∈𝒳u\\in\\mathcal\{X\}, letzu:=q​\(u\)\.z\_\{u\}:=q\(u\)\.Then we have by settingst=−G​\(zt\)​vts\_\{t\}=\-G\(z\_\{t\}\)v\_\{t\}:

ℓt​\(xt\)−ℓt​\(u\)=ht​\(zt\)−ht​\(zu\)=⟨st,zt−zu⟩=−⟨G​\(zt\)​vt,zt−zu⟩=−⟨G​\(zt\)​\(zt−zu\),vt⟩,\\ell\_\{t\}\(x\_\{t\}\)\-\\ell\_\{t\}\(u\)=h\_\{t\}\(z\_\{t\}\)\-h\_\{t\}\(z\_\{u\}\)=\\langle s\_\{t\},z\_\{t\}\-z\_\{u\}\\rangle=\-\\langle G\(z\_\{t\}\)v\_\{t\},z\_\{t\}\-z\_\{u\}\\rangle=\-\\langle G\(z\_\{t\}\)\(z\_\{t\}\-z\_\{u\}\),v\_\{t\}\\rangle\\,,where the last step uses symmetry ofG​\(zt\)G\(z\_\{t\}\)\. Define the vector fieldFu:𝒵→ℝ2F\_\{u\}:\\mathcal\{Z\}\\to\\mathbb\{R\}^\{2\}for anyz∈𝒵z\\in\\mathcal\{Z\}by:

Fu​\(z\):=G​\(z\)​\(z−zu\)\.F\_\{u\}\(z\):=G\(z\)\(z\-z\_\{u\}\)\\,\.\(38\)Using the above with the hidden\-space dynamics \([37](https://arxiv.org/html/2605.26373#A2.E37)\) and the estimate‖rt‖=𝒪​\(η2\)\\\|r\_\{t\}\\\|=\\mathcal\{O\}\(\\eta^\{2\}\), we obtain:

ℓt​\(xt\)−ℓt​\(u\)=−1η​⟨Fu​\(zt\),zt\+1−zt⟩\+𝒪​\(η\)\.\\displaystyle\\ell\_\{t\}\(x\_\{t\}\)\-\\ell\_\{t\}\(u\)=\-\\frac\{1\}\{\\eta\}\\langle F\_\{u\}\(z\_\{t\}\),z\_\{t\+1\}\-z\_\{t\}\\rangle\+\\mathcal\{O\}\(\\eta\)\\,\.\(39\)
Step 4: The vector fieldFuF\_\{u\}is not conservative\.We show in this step that the vector fieldFuF\_\{u\}defined in \([38](https://arxiv.org/html/2605.26373#A2.E38)\) is not a gradient field\. In particular, we will show the existence of a rectangleR⊂ℝ2R\\subset\\mathbb\{R\}^\{2\}in the hidden space over which the path integral is negative\. Set the comparatoru=\(0,π6\)∈𝒳u=\(0,\\frac\{\\pi\}\{6\}\)\\in\\mathcal\{X\}to show thatFuF\_\{u\}is not a gradient field\. Note that it is enough to consider a single comparatoruuin order to show the linear regret bound\. Denoting the coordinate functions ofFu​\(z\)F\_\{u\}\(z\)byPu​\(z\)P\_\{u\}\(z\)andQu​\(z\)Q\_\{u\}\(z\)respectively, it is easy to verify atz=\(1,1\)z=\(1,1\)for instance that∂z1Qu​\(z\)−∂z2Pu​\(z\)<0\\partial\_\{z\_\{1\}\}Q\_\{u\}\(z\)\-\\partial\_\{z\_\{2\}\}P\_\{u\}\(z\)<0\. By continuity of\(z1,z2\)↦∂z1Qu​\(z\)−∂z2Pu​\(z\)\(z\_\{1\},z\_\{2\}\)\\mapsto\\partial\_\{z\_\{1\}\}Q\_\{u\}\(z\)\-\\partial\_\{z\_\{2\}\}P\_\{u\}\(z\), there exists a rectangleR:=\[a,a\+α\]×\[b,b\+β\]⊂q​\(𝒳\)R:=\[a,a\+\\alpha\]\\times\[b,b\+\\beta\]\\subset q\(\\mathcal\{X\}\)for someα,β\>0\\alpha,\\beta\>0anda,b∈q​\(𝒳\)a,b\\in q\(\\mathcal\{X\}\)such that: there exists a constantγ\>0\\gamma\>0such that for allz∈Rz\\in R,∂z1Qu​\(z\)−∂z2Pu​\(z\)≤−γ\\partial\_\{z\_\{1\}\}Q\_\{u\}\(z\)\-\\partial\_\{z\_\{2\}\}P\_\{u\}\(z\)~\\leq~\-\\gamma\. We denote byA:=\(a,b\),B:=\(a\+α,B\),C:=\(a\+α,b\+β\)A:=\(a,b\),B:=\(a\+\\alpha,B\),C:=\(a\+\\alpha,b\+\\beta\)andD:=\(a,b\+β\)D:=\(a,b\+\\beta\)its corners and byIRI\_\{R\}the continuous oriented integral over the edges of the rectangleRRin the counter\-clockwise orientation \(A→B→C→DA\\to B\\to C\\to D\)\. It follows then that:

IR=∫aa\+α∫bb\+β\(∂z1Qu−∂z2Pu\)​\(x,y\)​𝑑x​𝑑y≤−γ​α​β<0\.I\_\{R\}=\\int\_\{a\}^\{a\+\\alpha\}\\int\_\{b\}^\{b\+\\beta\}\(\\partial\_\{z\_\{1\}\}Q\_\{u\}\-\\partial\_\{z\_\{2\}\}P\_\{u\}\)\(x,y\)dxdy~\\leq~\-\\gamma\\alpha\\beta<0\\,\.\(40\)
Step 5: Definition of the adversarial sequence\(st\)\(s\_\{t\}\)\.Recall that we have setst=−G​\(zt\)​vts\_\{t\}=\-G\(z\_\{t\}\)v\_\{t\}\. We now choose the sequence\(vt\)\(v\_\{t\}\)in an adaptive adversarial way that makes the hidden iterates move counter\-clockwise around the rectangleRR\. More precisely, we choosevt∈\{e1,e2,−e1,−e2\}v\_\{t\}\\in\\\{e\_\{1\},e\_\{2\},\-e\_\{1\},\-e\_\{2\}\\\}wheree1=\(1,0\)⊤,e2=\(0,1\)⊤e\_\{1\}=\(1,0\)^\{\\top\},e\_\{2\}=\(0,1\)^\{\\top\}in the following way, defined periodically for anytt:vt=e1v\_\{t\}=e\_\{1\}alongA​BAB,e2e\_\{2\}alongB​CBC,−e1\-e\_\{1\}alongC​DCDand−e2\-e\_\{2\}alongD​ADA\. Letn1:=⌊αη⌋,n2:=⌊βη⌋n\_\{1\}:=\\lfloor\\frac\{\\alpha\}\{\\eta\}\\rfloor,n\_\{2\}:=\\lfloor\\frac\{\\beta\}\{\\eta\}\\rfloor\. Then the full cycle around the rectangleRRconsists ofN=2​n1\+2​n2=2​\(α\+β\)η\+𝒪​\(1\)N=2n\_\{1\}\+2n\_\{2\}=\\frac\{2\(\\alpha\+\\beta\)\}\{\\eta\}\+\\mathcal\{O\}\(1\)steps\. Thenht​\(z\)=⟨st,z⟩h\_\{t\}\(z\)=\\langle s\_\{t\},z\\rangleat roundttwithst=−G​\(zt\)​vt\.s\_\{t\}=\-G\(z\_\{t\}\)v\_\{t\}\\,\.This givesℓt​\(x\)=ht​\(q​\(x\)\)=⟨st,q​\(x\)⟩\\ell\_\{t\}\(x\)=h\_\{t\}\(q\(x\)\)=\\langle s\_\{t\},q\(x\)\\rangleand defines entirely the sequence of adversarial loss functions over which OGD is run\.

Step 6: Riemann sum estimate of the path sum in regret in \([51](https://arxiv.org/html/2605.26373#A2.E51)\) using the rectangle integralIRI\_\{R\}\.Our goal in this step is to relate the path sum appearing in the regret decomposition in \([51](https://arxiv.org/html/2605.26373#A2.E51)\) and the negative rectangle integralIRI\_\{R\}in \([58](https://arxiv.org/html/2605.26373#A2.E58)\)\. Specifically we show that:

∑t=1N⟨Fu​\(zt\),zt\+1−zt⟩=IR\+𝒪​\(η\),\\sum\_\{t=1\}^\{N\}\\langle F\_\{u\}\(z\_\{t\}\),z\_\{t\+1\}\-z\_\{t\}\\rangle=I\_\{R\}\+\\mathcal\{O\}\(\\eta\)\\,,\(41\)whereN=2​\(n1\+n2\)N=2\(n\_\{1\}\+n\_\{2\}\)is the number of steps in one cycle of the iterates around the rectangleRR

Step 7: Regret over one cycle of lengthNNaround the rectangleRR\.We show that one full cycle results in regret of order1η\\frac\{1\}\{\\eta\}\. Since one cycle lastsN=Θ​\(1η\)N=\\Theta\(\\frac\{1\}\{\\eta\}\)rounds, this gives a positive constant regret per round \(on average\)\. Summing up \([50](https://arxiv.org/html/2605.26373#A2.E50)\) over one cycle ofNNsteps yields:

∑t=1Nℓt​\(xt\)−ℓt​\(u\)=−1η​∑t=1N⟨Fu​\(zt\),zt\+1−zt⟩\+𝒪​\(N​η\)=−1η​IR\+𝒪​\(1\),\\sum\_\{t=1\}^\{N\}\\ell\_\{t\}\(x\_\{t\}\)\-\\ell\_\{t\}\(u\)=\-\\frac\{1\}\{\\eta\}\\sum\_\{t=1\}^\{N\}\\langle F\_\{u\}\(z\_\{t\}\),z\_\{t\+1\}\-z\_\{t\}\\rangle\+\\mathcal\{O\}\(N\\eta\)=\-\\frac\{1\}\{\\eta\}I\_\{R\}\+\\mathcal\{O\}\(1\)\\,,\(42\)where the first equality uses the fact that one cycle hasN=Θ​\(1η\)N=\\Theta\(\\frac\{1\}\{\\eta\}\)steps and the second identity follows from using \([41](https://arxiv.org/html/2605.26373#A2.E41)\)\. SinceIR≤−γ​α​β<0I\_\{R\}~\\leq~\-\\gamma\\alpha\\beta<0, it follows that:

∑t=1Nℓt​\(xt\)−ℓt​\(u\)≥\+γ​α​βη\+𝒪​\(1\)≥c0η,\\sum\_\{t=1\}^\{N\}\\ell\_\{t\}\(x\_\{t\}\)\-\\ell\_\{t\}\(u\)~\\geq~\+\\frac\{\\gamma\\alpha\\beta\}\{\\eta\}\+\\mathcal\{O\}\(1\)~\\geq~\\frac\{c\_\{0\}\}\{\\eta\}\\,,\(43\)for somec0\>0c\_\{0\}\>0and forη∈\(0,η0\]\\eta\\in\(0,\\eta\_\{0\}\]for someη0\>0\\eta\_\{0\}\>0\.

Step 8: Linear regret lower bound by repeating cycles\.We conclude the proof of Theorem[2](https://arxiv.org/html/2605.26373#Thmtheorem2)by repeating the previous cycle over the time horizonTT\. LetK:=⌊TN⌋K:=\\lfloor\\frac\{T\}\{N\}\\rfloorforT≥NT~\\geq~Nbe the number of complete cycles up to timeTT\. Since each cycle contributes at leastc0η\\frac\{c\_\{0\}\}\{\\eta\}regret, summing up these contributions over theKKcycles concludes the proof\.

### B\.5Proof of Theorem[2](https://arxiv.org/html/2605.26373#Thmtheorem2)

Step 1: Definition of the parameterizationqqand induced metric\.We start by defining the parameterization functionqqto prove the result\.

Let𝒳:=\[0,log⁡3\]×\[π6,π3\]⊂ℝ2\.\\mathcal\{X\}:=\[0,\\log 3\]\\times\[\\frac\{\\pi\}\{6\},\\frac\{\\pi\}\{3\}\]\\subset\\mathbb\{R\}^\{2\}\.Define the parameterization functionq:𝒳→ℝ2q:\\mathcal\{X\}\\to\\mathbb\{R\}^\{2\}for anyx=\(x1,x2\)∈𝒳x=\(x\_\{1\},x\_\{2\}\)\\in\\mathcal\{X\}as follows:

q​\(x1,x2\):=ex1​\(cos⁡x2,sin⁡x2\)\.q\(x\_\{1\},x\_\{2\}\):=e^\{x\_\{1\}\}\(\\cos x\_\{2\},\\sin x\_\{2\}\)\\,\.Then it follows that the Jacobian ofqqat anyx=\(x1,x2\)∈𝒳x=\(x\_\{1\},x\_\{2\}\)\\in\\mathcal\{X\}is given by:

Jq​\(x\)=\(∂q1​\(x\)∂x1∂q1​\(x\)∂x2∂q2​\(x\)∂x1∂q2​\(x\)∂x2\)=ex1​\(cos⁡x2−sin⁡x2sin⁡x2cos⁡x2\)\.J\_\{q\}\(x\)=\\begin\{pmatrix\}\\frac\{\\partial q\_\{1\}\(x\)\}\{\\partial x\_\{1\}\}&\\frac\{\\partial q\_\{1\}\(x\)\}\{\\partial x\_\{2\}\}\\\\ \\frac\{\\partial q\_\{2\}\(x\)\}\{\\partial x\_\{1\}\}&\\frac\{\\partial q\_\{2\}\(x\)\}\{\\partial x\_\{2\}\}\\end\{pmatrix\}=e^\{x\_\{1\}\}\\begin\{pmatrix\}\\cos x\_\{2\}&\-\\sin x\_\{2\}\\\\ \\sin x\_\{2\}&\\cos x\_\{2\}\\end\{pmatrix\}\\,\.Therefore, we obtain:

Jq​\(x\)​Jq​\(x\)⊤=e2​x1​I2,J\_\{q\}\(x\)J\_\{q\}\(x\)^\{\\top\}=e^\{2x\_\{1\}\}I\_\{2\}\\,,whereI2I\_\{2\}is the2×22\\times 2identity matrix\. Sincex1≥0x\_\{1\}~\\geq~0,e2​x1≥1e^\{2x\_\{1\}\}~\\geq~1and henceJq​\(x\)​Jq​\(x\)⊤⪰I2J\_\{q\}\(x\)J\_\{q\}\(x\)^\{\\top\}\\succeq I\_\{2\}\. Moreover, settingz:=q​\(x\)z:=q\(x\), we have‖z‖=‖q​\(x\)‖=ex1\\\|z\\\|=\\\|q\(x\)\\\|=e^\{x\_\{1\}\}and therefore:

Jq​\(x\)​Jq​\(x\)⊤=‖z‖2​I2\.J\_\{q\}\(x\)J\_\{q\}\(x\)^\{\\top\}=\\\|z\\\|^\{2\}I\_\{2\}\\,\.We now introduce additional notation which will be useful in the rest of the proof\. Define for anyz∈q​\(𝒳\)z\\in q\(\\mathcal\{X\}\),

M​\(z\)\\displaystyle M\(z\):=Jq​\(x\)​Jq​\(x\)⊤=‖z‖2​I2,\\displaystyle:=J\_\{q\}\(x\)J\_\{q\}\(x\)^\{\\top\}=\\\|z\\\|^\{2\}I\_\{2\}\\,,G​\(z\)\\displaystyle G\(z\):=M​\(z\)−1=1‖z‖2​I2\.\\displaystyle:=M\(z\)^\{\-1\}=\\frac\{1\}\{\\\|z\\\|^\{2\}\}I\_\{2\}\\,\.
Step 2: OGD dynamics in the hidden space\.To construct our adversarial sequence of losses and give more insight about the idea of the proof, we first write the OGD dynamics in the hidden space𝒴:=q​\(𝒳\)\.\{\\mathcal\{Y\}\}:=q\(\\mathcal\{X\}\)\.

We will consider linear functionsht:𝒴→ℝh\_\{t\}:\{\\mathcal\{Y\}\}\\to\\mathbb\{R\}defined for everyz∈𝒴z\\in\{\\mathcal\{Y\}\}byht​\(z\)=⟨st,z⟩h\_\{t\}\(z\)=\\langle s\_\{t\},z\\ranglewhere\(st\)\(s\_\{t\}\)is an adversarial sequence that will be specified later in an adversarial way\. It follows that for anyx∈𝒳x\\in\\mathcal\{X\},

ℓt​\(x\)=ht​\(q​\(x\)\)=⟨st,q​\(x\)⟩,∇ℓt​\(x\)=Jq​\(x\)⊤​st\.\\ell\_\{t\}\(x\)=h\_\{t\}\(q\(x\)\)=\\langle s\_\{t\},q\(x\)\\rangle\\,,\\quad\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\(x\)=J\_\{q\}\(x\)^\{\\top\}s\_\{t\}\.
We first record the local hidden\-space dynamics in the case where the Euclidean projection is inactive\. For \(projected\) OGD on the sequence of function\(ℓt\)\(\\ell\_\{t\}\), the update rule is:

xt\+1=Π𝒳​\(xt−η​∇ℓt​\(xt\)\)=Π𝒳​\(xt−η​Jq​\(xt\)⊤​st\)\.x\_\{t\+1\}=\\Pi\_\{\\mathcal\{X\}\}\\bigl\(x\_\{t\}\-\\eta\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\(x\_\{t\}\)\\bigr\)=\\Pi\_\{\\mathcal\{X\}\}\\bigl\(x\_\{t\}\-\\eta J\_\{q\}\(x\_\{t\}\)^\{\\top\}s\_\{t\}\\bigr\)\.Thus, on any round for which the pre\-projection pointx~t\+1:=xt−η​Jq​\(xt\)⊤​st\\widetilde\{x\}\_\{t\+1\}:=x\_\{t\}\-\\eta J\_\{q\}\(x\_\{t\}\)^\{\\top\}s\_\{t\}belongs to𝒳\\mathcal\{X\}, the projection is inactive and

xt\+1=xt−η​∇ℓt​\(xt\)=xt−η​Jq​\(xt\)⊤​st\.x\_\{t\+1\}=x\_\{t\}\-\\eta\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\(x\_\{t\}\)=x\_\{t\}\-\\eta J\_\{q\}\(x\_\{t\}\)^\{\\top\}s\_\{t\}\\,\.\(44\)We will verify below \(in step 5\), after constructing the adversarial cycle, that the iterates remain a positive distance away from the boundary∂𝒳\\partial\\mathcal\{X\}\. Consequently, for sufficiently smallη\\eta, the pre\-projection points remain in𝒳\\mathcal\{X\}, so the projection is indeed inactive throughout the constructed trajectory\.

Now defining the sequence\(zt\)\(z\_\{t\}\)byzt:=q​\(xt\)z\_\{t\}:=q\(x\_\{t\}\)for anyttin hidden space𝒴\{\\mathcal\{Y\}\}, we obtain the following using Taylor’s theorem with exact remainder:

zt\+1=q​\(xt\+1\)=q​\(xt\)\+Jq​\(xt\)​\(xt\+1−xt\)\+rt,z\_\{t\+1\}=q\(x\_\{t\+1\}\)=q\(x\_\{t\}\)\+J\_\{q\}\(x\_\{t\}\)\(x\_\{t\+1\}\-x\_\{t\}\)\+r\_\{t\}\\,,\(45\)wherertr\_\{t\}is a Taylor remainder which can be bounded as follows by global boundedness of the Hessian ofqq\(recall here that the space𝒳\\mathcal\{X\}is compact\):

‖rt‖=𝒪​\(‖xt\+1−xt‖2\)=𝒪​\(η2​‖∇ℓt​\(xt\)‖2\)=𝒪​\(η2\),\\\|r\_\{t\}\\\|=\\mathcal\{O\}\(\\\|x\_\{t\+1\}\-x\_\{t\}\\\|^\{2\}\)=\\mathcal\{O\}\(\\eta^\{2\}\\\|\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\(x\_\{t\}\)\\\|^\{2\}\)=\\mathcal\{O\}\(\\eta^\{2\}\)\\,,where the last step follows from uniform boundedness of the gradients offtf\_\{t\}on the compact space𝒳\\mathcal\{X\}asqqis smooth and\(st\)\(s\_\{t\}\)will be chosen as a bounded sequence\.

Plugging the OGD update rule of\(xt\)\(x\_\{t\}\)\(see \([44](https://arxiv.org/html/2605.26373#A2.E44)\)\) in \([45](https://arxiv.org/html/2605.26373#A2.E45)\), we obtain:

zt\+1\\displaystyle z\_\{t\+1\}=zt−η​Jq​\(xt\)​∇ℓt​\(xt\)\+rt\\displaystyle=z\_\{t\}\-\\eta J\_\{q\}\(x\_\{t\}\)\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\(x\_\{t\}\)\+r\_\{t\}=zt−η​Jq​\(xt\)​Jq​\(xt\)⊤​st\+rt\\displaystyle=z\_\{t\}\-\\eta J\_\{q\}\(x\_\{t\}\)J\_\{q\}\(x\_\{t\}\)^\{\\top\}s\_\{t\}\+r\_\{t\}=zt−M​\(zt\)​st\+rt\.\\displaystyle=z\_\{t\}\-M\(z\_\{t\}\)s\_\{t\}\+r\_\{t\}\\,\.To simplify the OGD dynamics in the hidden space, we choosest=−G​\(zt\)​vt=−M​\(zt\)−1​vts\_\{t\}=\-G\(z\_\{t\}\)v\_\{t\}=\-M\(z\_\{t\}\)^\{\-1\}v\_\{t\}where\(vt\)\(v\_\{t\}\)is an adversarial sequence that will be chosen later\. Given this choice, we obtain the following OGD dynamics in the hidden space:

zt\+1=zt\+η​vt\+rt,z\_\{t\+1\}=z\_\{t\}\+\\eta v\_\{t\}\+r\_\{t\}\\,,\(46\)where‖rt‖=𝒪​\(η2\)\\\|r\_\{t\}\\\|=\\mathcal\{O\}\(\\eta^\{2\}\)as previously mentioned\.

Step 3: Regret as an approximate discrete path sum\.In this step, we write the regret under a specific path sum form for some fieldFuF\_\{u\}depending on a fixed comparatoru∈𝒳u\\in\\mathcal\{X\}\. The main idea of the proof will then be to show thatFuF\_\{u\}is not a gradient field and we will choose the sequence\(vt\)\(v\_\{t\}\)in an appropriate way to exploit this fact\. Non\-vanishing regret will be accumulated along the path due some cycling behavior ofFuF\_\{u\}\.

For a fixed comparatoru∈𝒳u\\in\\mathcal\{X\}, letzu:=q​\(u\)\.z\_\{u\}:=q\(u\)\.Then we have

ℓt​\(xt\)−ℓt​\(u\)\\displaystyle\\ell\_\{t\}\(x\_\{t\}\)\-\\ell\_\{t\}\(u\)=ht​\(q​\(xt\)\)−ht​\(q​\(u\)\)\\displaystyle=h\_\{t\}\(q\(x\_\{t\}\)\)\-h\_\{t\}\(q\(u\)\)=ht​\(zt\)−ht​\(zu\)\\displaystyle=h\_\{t\}\(z\_\{t\}\)\-h\_\{t\}\(z\_\{u\}\)=⟨st,zt−zu⟩\\displaystyle=\\langle s\_\{t\},z\_\{t\}\-z\_\{u\}\\rangle=−⟨G​\(zt\)​vt,zt−zu⟩\\displaystyle=\-\\langle G\(z\_\{t\}\)v\_\{t\},z\_\{t\}\-z\_\{u\}\\rangle\(st=−G​\(zt\)​vts\_\{t\}=\-G\(z\_\{t\}\)v\_\{t\}\)=−⟨G​\(zt\)​zt−zu,vt⟩\\displaystyle=\-\\langle G\(z\_\{t\}\)z\_\{t\}\-z\_\{u\},v\_\{t\}\\rangle\(by symmetry ofG​\(zt\)\)\.\\displaystyle\\hfill\\text\{\(by symmetry of $G\(z\_\{t\}\)$\)\}\.\(47\)Now define the vector fieldFu:𝒵→ℝ2F\_\{u\}:\\mathcal\{Z\}\\to\\mathbb\{R\}^\{2\}for anyz∈𝒵z\\in\\mathcal\{Z\}by:

Fu​\(z\):=G​\(z\)​\(z−zu\)\.F\_\{u\}\(z\):=G\(z\)\(z\-z\_\{u\}\)\\,\.\(48\)Using \([47](https://arxiv.org/html/2605.26373#A2.E47)\) together with the OGD dynamics \([46](https://arxiv.org/html/2605.26373#A2.E46)\) in the hidden space, we obtain:

ℓt​\(xt\)−ℓt​\(u\)\\displaystyle\\ell\_\{t\}\(x\_\{t\}\)\-\\ell\_\{t\}\(u\)=−⟨Fu​\(zt\),vt⟩\\displaystyle=\-\\langle F\_\{u\}\(z\_\{t\}\),v\_\{t\}\\rangle=−1η​⟨Fu​\(zt\),zt\+1−zt−rt⟩\\displaystyle=\-\\frac\{1\}\{\\eta\}\\langle F\_\{u\}\(z\_\{t\}\),z\_\{t\+1\}\-z\_\{t\}\-r\_\{t\}\\rangle=−1η​⟨Fu​\(zt\),zt\+1−zt⟩\+1η​⟨Fu​\(zt\),rt⟩\.\\displaystyle=\-\\frac\{1\}\{\\eta\}\\langle F\_\{u\}\(z\_\{t\}\),z\_\{t\+1\}\-z\_\{t\}\\rangle\+\\frac\{1\}\{\\eta\}\\langle F\_\{u\}\(z\_\{t\}\),r\_\{t\}\\rangle\\,\.\(49\)SinceFuF\_\{u\}is bounded and‖rt‖=𝒪​\(η2\)\\\|r\_\{t\}\\\|=\\mathcal\{O\}\(\\eta^\{2\}\), we have:

\|1η​⟨Fu​\(zt\),rt⟩\|=𝒪​\(1η​‖rt‖\)=𝒪​\(η\)\.\\left\|\\frac\{1\}\{\\eta\}\\langle F\_\{u\}\(z\_\{t\}\),r\_\{t\}\\rangle\\right\|=\\mathcal\{O\}\\left\(\\frac\{1\}\{\\eta\}\\\|r\_\{t\}\\\|\\right\)=\\mathcal\{O\}\(\\eta\)\\,\.Overall, we have

ℓt​\(xt\)−ℓt​\(u\)=−1η​⟨Fu​\(zt\),zt\+1−zt⟩\+𝒪​\(η\)\.\\ell\_\{t\}\(x\_\{t\}\)\-\\ell\_\{t\}\(u\)=\-\\frac\{1\}\{\\eta\}\\langle F\_\{u\}\(z\_\{t\}\),z\_\{t\+1\}\-z\_\{t\}\\rangle\+\\mathcal\{O\}\(\\eta\)\\,\.\(50\)We have written regret as a path sum of the vector fieldFuF\_\{u\}up to an error of orderη​𝒯\\mathcal\{\\eta T\}:

RT:=∑t=1Tℓt​\(xt\)−ℓt​\(u\)=−1η​∑t=1T⟨Fu​\(zt\),zt\+1−zt⟩\+𝒪​\(η​T\)\.\\displaystyle R\_\{T\}:=\\sum\_\{t=1\}^\{T\}\\ell\_\{t\}\(x\_\{t\}\)\-\\ell\_\{t\}\(u\)=\-\\frac\{1\}\{\\eta\}\\sum\_\{t=1\}^\{T\}\\langle F\_\{u\}\(z\_\{t\}\),z\_\{t\+1\}\-z\_\{t\}\\rangle\+\\mathcal\{O\}\(\\eta T\)\\,\.\(51\)
Step 4: The vector fieldFuF\_\{u\}is not conservative\.We show in this step that the vector fieldFuF\_\{u\}defined in \([48](https://arxiv.org/html/2605.26373#A2.E48)\) is not a gradient field\. In particular, we will show the existence of a rectangle in the hidden space over which the path integral is negative\.

Set the comparatoru=\(0,π6\)∈𝒳u=\(0,\\frac\{\\pi\}\{6\}\)\\in\\mathcal\{X\}to show thatFuF\_\{u\}is not a gradient field\. Note that it will be enough for our purpose to consider a single comparatoruuto show our linear regret bound\.

Thenzu=q​\(u\)=\(cos⁡π6,sin⁡π6\)=\(32,12\)z\_\{u\}=q\(u\)=\(\\cos\\frac\{\\pi\}\{6\},\\sin\\frac\{\\pi\}\{6\}\)=\(\\frac\{\\sqrt\{3\}\}\{2\},\\frac\{1\}\{2\}\)\. Recalling the definition ofFuF\_\{u\}in \([48](https://arxiv.org/html/2605.26373#A2.E48)\), we have for anyz=\(z1,z2\)∈Q​\(𝒳\)z=\(z\_\{1\},z\_\{2\}\)\\in Q\(\\mathcal\{X\}\),

Fu​\(z\)=G​\(z\)​\(z−zu\)=1‖z‖2​\(z−zu\)=1z12\+z22​\(z−zu\)\.F\_\{u\}\(z\)=G\(z\)\(z\-z\_\{u\}\)=\\frac\{1\}\{\\\|z\\\|^\{2\}\}\(z\-z\_\{u\}\)=\\frac\{1\}\{z\_\{1\}^\{2\}\+z\_\{2\}^\{2\}\}\(z\-z\_\{u\}\)\\,\.
Denote the coordinate functions ofFu​\(z\)F\_\{u\}\(z\)asPu​\(z\)P\_\{u\}\(z\)andQu​\(z\)Q\_\{u\}\(z\)\. Then for anyz=\(z1,z2\)∈q​\(𝒳\)z=\(z\_\{1\},z\_\{2\}\)\\in q\(\\mathcal\{X\}\),

Fu​\(z\)=\(Pu​\(z\),Qu​\(z\)\)=\(z1−zu,1z12\+z22,z2−zu,2z12\+z22\)\.F\_\{u\}\(z\)=\(P\_\{u\}\(z\),Q\_\{u\}\(z\)\)=\\left\(\\frac\{z\_\{1\}\-z\_\{u,1\}\}\{z\_\{1\}^\{2\}\+z\_\{2\}^\{2\}\},\\frac\{z\_\{2\}\-z\_\{u,2\}\}\{z\_\{1\}^\{2\}\+z\_\{2\}^\{2\}\}\\right\)\\,\.We now compute the difference of partial derivatives to obtain:

∂z1Qu​\(z\)−∂z2Pu​\(z\)=2​\(z1​zu,2−z2​zu1\)\(z12\+z22\)2\.\\partial\_\{z\_\{1\}\}Q\_\{u\}\(z\)\-\\partial\_\{z\_\{2\}\}P\_\{u\}\(z\)=\\frac\{2\(z\_\{1\}z\_\{u,2\}\-z\_\{2\}z\_\{u\_\{1\}\}\)\}\{\(z\_\{1\}^\{2\}\+z\_\{2\}^\{2\}\)^\{2\}\}\\,\.Atz=\(1,1\)z=\(1,1\), we have:

∂z1Qu​\(z\)−∂z2Pu​\(z\)=12​\(12−32\)<0\.\\partial\_\{z\_\{1\}\}Q\_\{u\}\(z\)\-\\partial\_\{z\_\{2\}\}P\_\{u\}\(z\)=\\frac\{1\}\{2\}\\left\(\\frac\{1\}\{2\}\-\\frac\{\\sqrt\{3\}\}\{2\}\\right\)<0\\,\.
Since\(1,1\)=q​\(log⁡2,π/4\)∈q​\(int⁡𝒳\)\(1,1\)=q\(\\log\\sqrt\{2\},\\pi/4\)\\in q\(\\operatorname\{int\}\\mathcal\{X\}\), and sincez↦∂z1Qu​\(z\)−∂z2Pu​\(z\)z\\mapsto\\partial\_\{z\_\{1\}\}Q\_\{u\}\(z\)\-\\partial\_\{z\_\{2\}\}P\_\{u\}\(z\)is continuous and strictly negative atz=\(1,1\)z=\(1,1\), there existα,β\>0\\alpha,\\beta\>0,a,b∈ℝa,b\\in\\mathbb\{R\}, andγ\>0\\gamma\>0such that:

\(1,1\)∈R:=\[a,a\+α\]×\[b,b\+β\]⊂q​\(int⁡𝒳\),\(1,1\)\\in R:=\[a,a\+\\alpha\]\\times\[b,b\+\\beta\]\\subset q\(\\operatorname\{int\}\\mathcal\{X\}\),\(52\)whereRRis a rectangle, and for allz∈Rz\\in R,

∂z1Qu​\(z\)−∂z2Pu​\(z\)≤−γ\.\\partial\_\{z\_\{1\}\}Q\_\{u\}\(z\)\-\\partial\_\{z\_\{2\}\}P\_\{u\}\(z\)\\leq\-\\gamma\.\(53\)Note here that we are making sure the rectangleRRis inq​\(int⁡𝒳\)q\(\\operatorname\{int\}\\mathcal\{X\}\)is order to be able to ignore the projection in OGD\. Moreover, by shrinkingRRif necessary, we may assume that there existsρz\>0\\rho\_\{z\}\>0such that the compact set:

Kz:=\{z∈ℝ2:dist⁡\(z,R\)≤ρz\}\.K\_\{z\}:=\\\{z\\in\\mathbb\{R\}^\{2\}:\\operatorname\{dist\}\(z,R\)\\leq\\rho\_\{z\}\\\}\.\(54\)satisfiesKz⊂q​\(int⁡𝒳\)K\_\{z\}\\subset q\(\\operatorname\{int\}\\mathcal\{X\}\)\. This neighborhood is a buffer around the rectangle which ensures that the actual perturbed discrete trajectoryztz\_\{t\}\(induced by the adversarial losses we will define\) remains insideq​\(int⁡𝒳\)q\(\\operatorname\{int\}\\mathcal\{X\}\), not just the ideal sequence which goes exactly around the rectangle without errors\. This will be used in steps 5 and 6 below\.

Let the corners of the rectangleRRbe denoted as follows:

D\\displaystyle D:=\(a,b\+β\),\\displaystyle:=\(a,b\+\\beta\)\\,,C:=\(a\+α,b\+β\),\\displaystyle C:=\(a\+\\alpha,b\+\\beta\)\\,,A\\displaystyle A:=\(a,b\),\\displaystyle:=\(a,b\)\\,,B:=\(a\+α,B\)\.\\displaystyle B:=\(a\+\\alpha,B\)\\,\.
We consider the counter\-clockwise orientation:

A→B→C→D\.A\\to B\\to C\\to D\\,\.Given this orientation, we define the following continuous edge integral:

IR:=∫ABPu​\(z\)​𝑑z1\+∫BCQu​\(z\)​𝑑z2\+∫CDPu​\(z\)​𝑑z1\+∫DAQu​\(z\)​𝑑z2\.I\_\{R\}:=\\int\_\{A\}^\{B\}P\_\{u\}\(z\)dz\_\{1\}\+\\int\_\{B\}^\{C\}Q\_\{u\}\(z\)dz\_\{2\}\+\\int\_\{C\}^\{D\}P\_\{u\}\(z\)dz\_\{1\}\+\\int\_\{D\}^\{A\}Q\_\{u\}\(z\)dz\_\{2\}\\,\.\(55\)
The above integral can be rewritten as follows using differential calculus:

IR=∫aa\+α∫bb\+β\(∂z1Qu−∂z2Pu\)​\(x,y\)​𝑑x​𝑑y\.I\_\{R\}=\\int\_\{a\}^\{a\+\\alpha\}\\int\_\{b\}^\{b\+\\beta\}\(\\partial\_\{z\_\{1\}\}Q\_\{u\}\-\\partial\_\{z\_\{2\}\}P\_\{u\}\)\(x,y\)dxdy\\,\.\(56\)The reader familiar with differential calculus may immediately see that the result holds\. We provide an elementary proof for of this identity below for completeness:

###### Proof\.

\(of \([56](https://arxiv.org/html/2605.26373#A2.E56)\)\) By definition of the continuous edge integral and the coordinate functions, we have:

IR\\displaystyle I\_\{R\}=∫aa\+αPu​\(x,b\)​𝑑x\+∫bb\+βQu​\(a\+α,y\)​𝑑y−∫aa\+αPu​\(x,b\+β\)​𝑑x−∫bb\+βQu​\(a,y\)​𝑑y\\displaystyle=\\int\_\{a\}^\{a\+\\alpha\}P\_\{u\}\(x,b\)dx\+\\int\_\{b\}^\{b\+\\beta\}Q\_\{u\}\(a\+\\alpha,y\)dy\-\\int\_\{a\}^\{a\+\\alpha\}P\_\{u\}\(x,b\+\\beta\)dx\-\\int\_\{b\}^\{b\+\\beta\}Q\_\{u\}\(a,y\)dy=∫aa\+α\(Pu​\(x,b\)−Pu​\(x,b\+β\)\)​𝑑x\+∫bb\+β\(Qu​\(a\+α,y\)−Qu​\(a,y\)\)​𝑑y\\displaystyle=\\int\_\{a\}^\{a\+\\alpha\}\(P\_\{u\}\(x,b\)\-P\_\{u\}\(x,b\+\\beta\)\)dx\+\\int\_\{b\}^\{b\+\\beta\}\(Q\_\{u\}\(a\+\\alpha,y\)\-Q\_\{u\}\(a,y\)\)dy=−∫aa\+α∫bb\+β∂z2Pu​\(x,y\)​d​y​d​x\+∫bb\+β∫aa\+α∂z1Qu​\(x,y\)​d​x​d​y,\\displaystyle=\-\\int\_\{a\}^\{a\+\\alpha\}\\int\_\{b\}^\{b\+\\beta\}\\partial\_\{z\_\{2\}\}P\_\{u\}\(x,y\)dydx\+\\int\_\{b\}^\{b\+\\beta\}\\int\_\{a\}^\{a\+\\alpha\}\\partial\_\{z\_\{1\}\}Q\_\{u\}\(x,y\)dxdy\\,,\(57\)where the third equality follows from the one\-dimensional fundamental theorem of calculus\. ∎

Using \([53](https://arxiv.org/html/2605.26373#A2.E53)\), we obtain that:

IR≤∫aa\+α∫bb\+β\(−γ\)​𝑑x​𝑑y=−γ​α​β<0\.I\_\{R\}~\\leq~\\int\_\{a\}^\{a\+\\alpha\}\\int\_\{b\}^\{b\+\\beta\}\(\-\\gamma\)dxdy=\-\\gamma\\alpha\\beta<0\\,\.\(58\)
Step 5: Definition of the adversarial sequence\(st\)\(s\_\{t\}\)via\(vt\)\(v\_\{t\}\)\.We choose the sequence\(vt\)\(v\_\{t\}\)in an adaptive adversarial way that makes the hidden iterates move counter\-clockwise around the rectangleRRdefined in \([52](https://arxiv.org/html/2605.26373#A2.E52)\)\.

More precisely, we choosevt∈\{e1,e2,−e1,−e2\}v\_\{t\}\\in\\\{e\_\{1\},e\_\{2\},\-e\_\{1\},\-e\_\{2\}\\\}wheree1=\(1,0\)⊤,e2=\(0,1\)⊤e\_\{1\}=\(1,0\)^\{\\top\},e\_\{2\}=\(0,1\)^\{\\top\}in the following way, defined periodically for anytt:

- •along the edgeA​BAB:vt=e1v\_\{t\}=e\_\{1\},
- •along the edgeB​CBC:vt=e2v\_\{t\}=e\_\{2\},
- •along the edgeC​DCD:vt=−e1v\_\{t\}=\-e\_\{1\},
- •along the edgeD​ADA:vt=−e2v\_\{t\}=\-e\_\{2\}\.

Letn1:=⌊αη⌋,n2:=⌊βη⌋n\_\{1\}:=\\lfloor\\frac\{\\alpha\}\{\\eta\}\\rfloor,n\_\{2\}:=\\lfloor\\frac\{\\beta\}\{\\eta\}\\rfloor\. Then the full cycle around the rectangleRRconsists ofN=2​n1\+2​n2=2​\(α\+β\)η\+𝒪​\(1\)N=2n\_\{1\}\+2n\_\{2\}=\\frac\{2\(\\alpha\+\\beta\)\}\{\\eta\}\+\\mathcal\{O\}\(1\)steps\. Thenht​\(z\)=⟨st,z⟩h\_\{t\}\(z\)=\\langle s\_\{t\},z\\rangleat roundttwithst=−G​\(zt\)​vt\.s\_\{t\}=\-G\(z\_\{t\}\)v\_\{t\}\\,\.This givesℓt​\(x\)=ht​\(q​\(x\)\)=⟨st,q​\(x\)⟩\\ell\_\{t\}\(x\)=h\_\{t\}\(q\(x\)\)=\\langle s\_\{t\},q\(x\)\\rangleand defines entirely the sequence of adversarial loss functions over which OGD is run\.

We next verify that, for this adversarial sequence and for sufficiently smallη\\eta, the Euclidean projection in projected OGD is never active\. This allows us to use the same local hidden\-space expansion as in the unconstrained update, i\.e\., it remains to justify the claim made in Step 2 that the Euclidean projection is inactive along the constructed trajectory\. SinceKz⊂q​\(int⁡𝒳\)K\_\{z\}\\subset q\(\\operatorname\{int\}\\mathcal\{X\}\), the setKx:=q−1​\(Kz\)K\_\{x\}:=q^\{\-1\}\(K\_\{z\}\)is a compact subset ofint⁡𝒳\\operatorname\{int\}\\mathcal\{X\}, using continuity ofq−1q^\{\-1\}\. Thereforeρx:=dist⁡\(Kx,∂𝒳\)\>0\.\\rho\_\{x\}:=\\operatorname\{dist\}\(K\_\{x\},\\partial\\mathcal\{X\}\)\>0\.Moreover, onKzK\_\{z\}, the adversarial vectors are uniformly bounded\. Indeed,vt∈\{±e1,±e2\}v\_\{t\}\\in\\\{\\pm e\_\{1\},\\pm e\_\{2\}\\\},st=−M​\(zt\)​vts\_\{t\}=\-M\(z\_\{t\}\)v\_\{t\}, andMMis continuous on the compact setKzK\_\{z\}\. Hence there existsCs<∞C\_\{s\}<\\inftysuch that‖st‖≤Cs\\\|s\_\{t\}\\\|\\leq C\_\{s\}wheneverzt∈Kzz\_\{t\}\\in K\_\{z\}\. SinceJqJ\_\{q\}is continuous on the compact setKxK\_\{x\}, there existsCJ<∞C\_\{J\}<\\inftysuch that‖Jq​\(x\)⊤‖≤CJ\\\|J\_\{q\}\(x\)^\{\\top\}\\\|\\leq C\_\{J\}for allx∈Kxx\\in K\_\{x\}\. Thus, wheneverzt∈Kzz\_\{t\}\\in K\_\{z\}andxt=q−1​\(zt\)x\_\{t\}=q^\{\-1\}\(z\_\{t\}\),‖η​Jq​\(xt\)⊤​st‖≤η​CJ​Cs\.\\\|\\eta J\_\{q\}\(x\_\{t\}\)^\{\\top\}s\_\{t\}\\\|\\leq\\eta C\_\{J\}C\_\{s\}\.Choosingη≤ρx/\(2​CJ​Cs\)\\eta\\leq\\rho\_\{x\}/\(2C\_\{J\}C\_\{s\}\), we getx~t\+1=xt−η​Jq​\(xt\)⊤​st∈𝒳\.\\widetilde\{x\}\_\{t\+1\}=x\_\{t\}\-\\eta J\_\{q\}\(x\_\{t\}\)^\{\\top\}s\_\{t\}\\in\\mathcal\{X\}\.ThereforeΠ𝒳​\(x~t\+1\)=x~t\+1\.\\Pi\_\{\\mathcal\{X\}\}\(\\widetilde\{x\}\_\{t\+1\}\)=\\widetilde\{x\}\_\{t\+1\}\.So projected OGD coincides with the unconstrained update on every round for whichzt∈Kzz\_\{t\}\\in K\_\{z\}\.

Step 6: Riemann sum estimate of the path sum in regret in \([51](https://arxiv.org/html/2605.26373#A2.E51)\) using the rectangle integralIRI\_\{R\}\.Our goal in this section is to relate the path sum appearing in the regret decomposition in \([51](https://arxiv.org/html/2605.26373#A2.E51)\) and the negative rectangle integralIRI\_\{R\}in \([58](https://arxiv.org/html/2605.26373#A2.E58)\)\. For this we will rely on an edge integral approximation lemma\. We present the result for the horizontal edgeA​BAB\. Similar results hold for the other edges of the rectangle with minor modifications\. Recall the OGD dynamics in the hidden space\. For anyt≥0,t~\\geq~0,

zt\+1=zt\+η​vt\+rt\.z\_\{t\+1\}=z\_\{t\}\+\\eta v\_\{t\}\+r\_\{t\}\\,\.
###### Lemma 9\(Horizontal edge approximation lemma\)\. LetF=\(P,Q\)F=\(P,Q\)beC1C^\{1\}on a compact neighborhood of the horizontal segmentE:=\{\(x,b\):x∈\[a,a\+α\]\}\.E:=\\\{\(x,b\):x\\in\[a,a\+\\alpha\]\\\}\.Suppose that a discrete sequencez0,…,znz\_\{0\},\\dots,z\_\{n\}forn=⌊αη⌋n=\\lfloor\\frac\{\\alpha\}\{\\eta\}\\rfloorsatisfies for all0≤k≤n−10~\\leq~k~\\leq~n\-1,zk\+1=zk\+η​e1\+rk,z\_\{k\+1\}=z\_\{k\}\+\\eta e\_\{1\}\+r\_\{k\}\\,,\(59\)where‖rk‖=O​\(η2\)\\\|r\_\{k\}\\\|=O\(\\eta^\{2\}\)withz0=\(a,b\)z\_\{0\}=\(a,b\)\. Then we have:∑k=0n−1⟨F​\(zk\),zk\+1−zk⟩=∫aa\+αP​\(x,b\)​𝑑x\+𝒪​\(η\)\.\\sum\_\{k=0\}^\{n\-1\}\\langle F\(z\_\{k\}\),z\_\{k\+1\}\-z\_\{k\}\\rangle=\\int\_\{a\}^\{a\+\\alpha\}P\(x,b\)dx\+\\mathcal\{O\}\(\\eta\)\\,\.

###### Proof\.

First, define the following auxiliary grid on the segment\[a,a\+α\]\[a,a\+\\alpha\],

z¯k:=\(a\+k​η,b\),k=0,…,n\.\\bar\{z\}\_\{k\}:=\(a\+k\\eta,b\),\\quad k=0,\\dots,n\\,\.Then it follows that:

zk=z¯k\+∑j=0n−1rj\.z\_\{k\}=\\bar\{z\}\_\{k\}\+\\sum\_\{j=0\}^\{n\-1\}r\_\{j\}\\,\.and therefore we have for any1≤k≤n1~\\leq~k~\\leq~n,

‖zk−z¯k‖=‖∑j=0k−1rj‖=𝒪​\(n​η2\)=𝒪​\(α​η\),\\\|z\_\{k\}\-\\bar\{z\}\_\{k\}\\\|=\\left\\\|\\sum\_\{j=0\}^\{k\-1\}r\_\{j\}\\right\\\|=\\mathcal\{O\}\(n\\eta^\{2\}\)=\\mathcal\{O\}\(\\alpha\\eta\)\\,,\(60\)where the last estimate follows from recalling thatn=⌊αη⌋\.n=\\lfloor\\frac\{\\alpha\}\{\\eta\}\\rfloor\\,\.This shows thatzkz\_\{k\}stays𝒪​\(η\)\\mathcal\{O\}\(\\eta\)\-close toz¯k\.\\bar\{z\}\_\{k\}\\,\.

Now we can estimate the path sum in \([51](https://arxiv.org/html/2605.26373#A2.E51)\) as follows:

∑k=0n−1⟨F​\(zk\),zk\+1−zk⟩\\displaystyle\\sum\_\{k=0\}^\{n\-1\}\\langle F\(z\_\{k\}\),z\_\{k\+1\}\-z\_\{k\}\\rangle=∑k=0n−1⟨F​\(zk\),η​e1\+rk⟩\\displaystyle=\\sum\_\{k=0\}^\{n\-1\}\\langle F\(z\_\{k\}\),\\eta e\_\{1\}\+r\_\{k\}\\rangle=η​∑k=0n−1P​\(zk\)\+∑k=0n−1⟨F​\(zk\),rk⟩\\displaystyle=\\eta\\sum\_\{k=0\}^\{n\-1\}P\(z\_\{k\}\)\+\\sum\_\{k=0\}^\{n\-1\}\\langle F\(z\_\{k\}\),r\_\{k\}\\rangle=η​∑k=0n−1P​\(z¯k\)\+η​η​∑k=0n−1P​\(zk\)−P​\(z¯k\)\+∑k=0n−1⟨F​\(zk\),rk⟩,\\displaystyle=\\eta\\sum\_\{k=0\}^\{n\-1\}P\(\\bar\{z\}\_\{k\}\)\+\\eta\\eta\\sum\_\{k=0\}^\{n\-1\}P\(z\_\{k\}\)\-P\(\\bar\{z\}\_\{k\}\)\+\\sum\_\{k=0\}^\{n\-1\}\\langle F\(z\_\{k\}\),r\_\{k\}\\rangle\\,,\(61\)where the first identity follows from using \([59](https://arxiv.org/html/2605.26373#A2.E59)\), the second one follows from recalling thatPPis the first coordinate function ofFF\.

We estimate each one of the three terms above separately\.

Last term in \([B\.5](https://arxiv.org/html/2605.26373#A2.Ex123)\)\.AsFFis bounded by continuity ofFFon the cimpact neighborhood, we obtain:

∑k=0n−1⟨F​\(zk\),rk⟩=𝒪​\(n⋅max0≤k≤n−1⁡‖rk‖\)=𝒪​\(n​η2\)=𝒪​\(η\),\\sum\_\{k=0\}^\{n\-1\}\\langle F\(z\_\{k\}\),r\_\{k\}\\rangle=\\mathcal\{O\}\\left\(n\\cdot\\max\_\{0~\\leq~k~\\leq~n\-1\}\\\|r\_\{k\}\\\|\\right\)=\\mathcal\{O\}\(n\\eta^\{2\}\)=\\mathcal\{O\}\(\\eta\)\\,,\(62\)where the last step uses again the fact thatn=⌊αη⌋\.n=\\lfloor\\frac\{\\alpha\}\{\\eta\}\\rfloor\.

Second term in \([B\.5](https://arxiv.org/html/2605.26373#A2.Ex123)\)\.Using Lipschitzness ofPPand recalling thatmax0≤k≤n−1⁡‖zk−z¯k‖=𝒪​\(η\)\\max\_\{0~\\leq~k~\\leq~n\-1\}\\\|z\_\{k\}\-\\bar\{z\}\_\{k\}\\\|=\\mathcal\{O\}\(\\eta\)proved in \([60](https://arxiv.org/html/2605.26373#A2.E60)\), we haveP​\(zk\)−P​\(z¯k\)=𝒪​\(η\)P\(z\_\{k\}\)\-P\(\\bar\{z\}\_\{k\}\)=\\mathcal\{O\}\(\\eta\)and it follows that:

η​∑k=0n−1P​\(zk\)−P​\(z¯k\)=𝒪​\(n​η2\)=𝒪​\(η\)\.\\eta\\sum\_\{k=0\}^\{n\-1\}P\(z\_\{k\}\)\-P\(\\bar\{z\}\_\{k\}\)=\\mathcal\{O\}\(n\\eta^\{2\}\)=\\mathcal\{O\}\(\\eta\)\\,\.\(63\)
First term in \([B\.5](https://arxiv.org/html/2605.26373#A2.Ex123)\)\.Denotingx¯k=a\+k​η\\bar\{x\}\_\{k\}=a\+k\\etafor0≤k≤n−10~\\leq~k~\\leq~n\-1, we can write:

η​∑k=0n−1P​\(z¯k\)\\displaystyle\\eta\\sum\_\{k=0\}^\{n\-1\}P\(\\bar\{z\}\_\{k\}\)=∑k=0n−1∫x¯kx¯k\+1P​\(z¯k\)​𝑑x\\displaystyle=\\sum\_\{k=0\}^\{n\-1\}\\int\_\{\\bar\{x\}\_\{k\}\}^\{\\bar\{x\}\_\{k\+1\}\}P\(\\bar\{z\}\_\{k\}\)dx=∑k=0n−1∫x¯kx¯k\+1P​\(x,b\)​𝑑x\+∑k=0n−1∫x¯kx¯k\+1P​\(z¯k\)−P​\(x,b\)​d​x\\displaystyle=\\sum\_\{k=0\}^\{n\-1\}\\int\_\{\\bar\{x\}\_\{k\}\}^\{\\bar\{x\}\_\{k\+1\}\}P\(x,b\)dx\+\\sum\_\{k=0\}^\{n\-1\}\\int\_\{\\bar\{x\}\_\{k\}\}^\{\\bar\{x\}\_\{k\+1\}\}P\(\\bar\{z\}\_\{k\}\)\-P\(x,b\)dx=∫0a\+n​ηP​\(x,b\)​𝑑x\+Rn,\\displaystyle=\\int\_\{0\}^\{a\+n\\eta\}P\(x,b\)dx\+R\_\{n\}\\,,\(64\)where the residualRnR\_\{n\}is defined as follows:

Rn:=∑k=0n−1∫x¯kx¯k\+1P​\(z¯k\)−P​\(x,b\)​d​x\.R\_\{n\}:=\\sum\_\{k=0\}^\{n\-1\}\\int\_\{\\bar\{x\}\_\{k\}\}^\{\\bar\{x\}\_\{k\+1\}\}P\(\\bar\{z\}\_\{k\}\)\-P\(x,b\)dx\\,\.
Using again Lipschitzness ofPP, the above residual can be bounded as follows:

\|Rn\|\\displaystyle\|R\_\{n\}\|≤∑k=0n−1∫x¯kx¯k\+1\|P​\(z¯k\)−P​\(x,b\)\|​𝑑x\\displaystyle~\\leq~\\sum\_\{k=0\}^\{n\-1\}\\int\_\{\\bar\{x\}\_\{k\}\}^\{\\bar\{x\}\_\{k\+1\}\}\|P\(\\bar\{z\}\_\{k\}\)\-P\(x,b\)\|\\,dx=𝒪​\(∑k=0n−1∫x¯kx¯k\+1\|x¯k−x\|​𝑑x\)\\displaystyle=\\mathcal\{O\}\\left\(\\sum\_\{k=0\}^\{n\-1\}\\int\_\{\\bar\{x\}\_\{k\}\}^\{\\bar\{x\}\_\{k\+1\}\}\|\\bar\{x\}\_\{k\}\-x\|\\,dx\\right\)=𝒪​\(n⋅max0≤k≤n−1⁡\|x¯k\+1−x¯k\|2\)\\displaystyle=\\mathcal\{O\}\\left\(n\\cdot\\max\_\{0~\\leq~k~\\leq~n\-1\}\|\\bar\{x\}\_\{k\+1\}\-\\bar\{x\}\_\{k\}\|^\{2\}\\right\)=𝒪​\(n​η2\)\\displaystyle=\\mathcal\{O\}\(n\\eta^\{2\}\)=𝒪​\(η\)\.\\displaystyle=\\mathcal\{O\}\(\\eta\)\\,\.\(65\)Now, observe that:

∫aa\+αP​\(x,b\)​𝑑x=∫aa\+n​ηP​\(x,b\)​𝑑x\+∫a\+n​ηa\+αP​\(x,b\)​𝑑x\.\\int\_\{a\}^\{a\+\\alpha\}P\(x,b\)dx=\\int\_\{a\}^\{a\+n\\eta\}P\(x,b\)dx\+\\int\_\{a\+n\\eta\}^\{a\+\\alpha\}P\(x,b\)dx\\,\.Hence, since0≤α−n​η≤η0~\\leq~\\alpha\-n\\eta~\\leq~\\etaby definition ofn=⌊αη⌋n=\\lfloor\\frac\{\\alpha\}\{\\eta\}\\rfloorand sinceP​\(x,b\)P\(x,b\)is bounded over the compact neighborhood of the horizontal segment, we obtain:

∫aa\+αP​\(x,b\)​𝑑x=∫aa\+n​ηP​\(x,b\)​𝑑x\+𝒪​\(η\)\.\\int\_\{a\}^\{a\+\\alpha\}P\(x,b\)dx=\\int\_\{a\}^\{a\+n\\eta\}P\(x,b\)dx\+\\mathcal\{O\}\(\\eta\)\\,\.\(66\)Combining all the above estimates \(\([62](https://arxiv.org/html/2605.26373#A2.E62)\), \([63](https://arxiv.org/html/2605.26373#A2.E63)\), \([B\.5](https://arxiv.org/html/2605.26373#A2.Ex125)\), \([B\.5](https://arxiv.org/html/2605.26373#A2.Ex128)\) and \([66](https://arxiv.org/html/2605.26373#A2.E66)\)\) in the main term \([B\.5](https://arxiv.org/html/2605.26373#A2.Ex123)\), we obtain the following estimate:

∑k=0n−1⟨F​\(zk\),zk\+1−zk⟩=∫aa\+αP​\(x,b\)​𝑑x\+𝒪​\(η\),\\sum\_\{k=0\}^\{n\-1\}\\langle F\(z\_\{k\}\),z\_\{k\+1\}\-z\_\{k\}\\rangle=\\int\_\{a\}^\{a\+\\alpha\}P\(x,b\)dx\+\\mathcal\{O\}\(\\eta\)\\,,concluding the proof of Lemma[9](https://arxiv.org/html/2605.26373#Thmlemma9)\. ∎

The previous estimates \(Lemma[9](https://arxiv.org/html/2605.26373#Thmlemma9)\) also imply that the actual discrete trajectory\(zt\)\(z\_\{t\}\)remains inKzK\_\{z\}as defined in \([54](https://arxiv.org/html/2605.26373#A2.E54)\)\. Indeed, along each edge the one\-step error is𝒪​\(η2\)\\mathcal\{O\}\(\\eta^\{2\}\), and each edge lasts𝒪​\(1/η\)\\mathcal\{O\}\(1/\\eta\)rounds, so the accumulated deviation from the intended rectangular path is𝒪​\(η\)\\mathcal\{O\}\(\\eta\)\. Choosingη\\etasufficiently small so that this deviation is at mostρz\\rho\_\{z\}, the iterates remain inKzK\_\{z\}throughout the cycle\. Consequently, by the conditional projection\-inactivity argument above \(see end of step 5\), the projection is inactive at every round of the cycle\.

Step 7: Regret over one cycle of lengthNNaround the rectangleRR\.Applying Lemma[9](https://arxiv.org/html/2605.26373#Thmlemma9)to the edge AB, and similarly to BC, CD, DA \(which follow using similar proofs with minor adjustments, e\.g\. replacinge1e\_\{1\}by−e1\-e\_\{1\},e2e\_\{2\}and−e2\-e\_\{2\}and the coordinate functionPuP\_\{u\}byQuQ\_\{u\}\), we sum all the estimates to obtain:

∑t=1N⟨Fu​\(zt\),zt\+1−zt⟩=IR\+𝒪​\(η\),\\sum\_\{t=1\}^\{N\}\\langle F\_\{u\}\(z\_\{t\}\),z\_\{t\+1\}\-z\_\{t\}\\rangle=I\_\{R\}\+\\mathcal\{O\}\(\\eta\)\\,,whereN=2​\(n1\+n2\)N=2\(n\_\{1\}\+n\_\{2\}\)is the number of steps in one cycle of the iterates around the rectangleR=A​B​C​D\.R=ABCD\.

Summing up \([50](https://arxiv.org/html/2605.26373#A2.E50)\) over one cycle ofNNsteps yields:

∑t=1Nℓt​\(xt\)−ℓt​\(u\)=−1η​∑t=1N⟨Fu​\(zt\),zt\+1−zt⟩\+𝒪​\(N​η\)\.\\sum\_\{t=1\}^\{N\}\\ell\_\{t\}\(x\_\{t\}\)\-\\ell\_\{t\}\(u\)=\-\\frac\{1\}\{\\eta\}\\sum\_\{t=1\}^\{N\}\\langle F\_\{u\}\(z\_\{t\}\),z\_\{t\+1\}\-z\_\{t\}\\rangle\+\\mathcal\{O\}\(N\\eta\)\\,\.\(67\)
Since one cycle hasN=Θ​\(1η\)N=\\Theta\(\\frac\{1\}\{\\eta\}\)steps, we haveN​η=𝒪​\(1\)N\\eta=\\mathcal\{O\}\(1\)\. Combining this with \([67](https://arxiv.org/html/2605.26373#A2.E67)\), we obtain:

∑t=1Nℓt​\(xt\)−ℓt​\(u\)=−1η​IR\+𝒪​\(1\)\.\\sum\_\{t=1\}^\{N\}\\ell\_\{t\}\(x\_\{t\}\)\-\\ell\_\{t\}\(u\)=\-\\frac\{1\}\{\\eta\}I\_\{R\}\+\\mathcal\{O\}\(1\)\\,\.Recalling from \([58](https://arxiv.org/html/2605.26373#A2.E58)\) that we have proved thatIR≤−γ​α​β<0I\_\{R\}~\\leq~\-\\gamma\\alpha\\beta<0, it follows that:

∑t=1Nℓt​\(xt\)−ℓt​\(u\)≥\+γ​α​βη\+𝒪​\(1\)\.\\sum\_\{t=1\}^\{N\}\\ell\_\{t\}\(x\_\{t\}\)\-\\ell\_\{t\}\(u\)~\\geq~\+\\frac\{\\gamma\\alpha\\beta\}\{\\eta\}\+\\mathcal\{O\}\(1\)\\,\.
Asγ​α​βη→\+∞\\frac\{\\gamma\\alpha\\beta\}\{\\eta\}\\to\+\\inftywhenη→0\\eta\\to 0, there existsη0\>0\\eta\_\{0\}\>0andc0\>0c\_\{0\}\>0such that for allη∈\(0,η0\]\\eta\\in\(0,\\eta\_\{0\}\],

∑t=1Nℓt​\(xt\)−ℓt​\(u\)≥c0η\.\\sum\_\{t=1\}^\{N\}\\ell\_\{t\}\(x\_\{t\}\)\-\\ell\_\{t\}\(u\)~\\geq~\\frac\{c\_\{0\}\}\{\\eta\}\\,\.\(68\)Therefore, we have shown that one full cycle results in regret of order1η\\frac\{1\}\{\\eta\}\. Since one cycle lastsN=Θ​\(1η\)N=\\Theta\(\\frac\{1\}\{\\eta\}\)rounds, this gives a positive constant regret per round \(on average\)\.

Step 8: Linear regret lower bound by repeating cycles\.We conclude the proof of Theorem[2](https://arxiv.org/html/2605.26373#Thmtheorem2)by repeating the previous cycle over the time horizonTT\.

LetK:=⌊TN⌋K:=\\lfloor\\frac\{T\}\{N\}\\rfloorforT≥NT~\\geq~Nbe the number of complete cycles up to timeTT\. We know from \([68](https://arxiv.org/html/2605.26373#A2.E68)\) in the previous step that each cycle contributes at leastc0η\\frac\{c\_\{0\}\}\{\\eta\}regret\. Thus, we have:

∑t=1Tℓt​\(xt\)−ℓt​\(u\)≥K​c0η−𝒪​\(1\)\.\\sum\_\{t=1\}^\{T\}\\ell\_\{t\}\(x\_\{t\}\)\-\\ell\_\{t\}\(u\)~\\geq~K\\frac\{c\_\{0\}\}\{\\eta\}\-\\mathcal\{O\}\(1\)\\,\.SinceN≤c2ηN~\\leq~\\frac\{c\_\{2\}\}\{\\eta\}for some constantc2\>0c\_\{2\}\>0, we have:

K≥TN−1≥η​Tc2−1\.K~\\geq~\\frac\{T\}\{N\}\-1~\\geq~\\eta\\frac\{T\}\{c\_\{2\}\}\-1\\,\.Substituting this in the previous lower bound yields:

∑t=1Tℓt​\(xt\)−ℓt​\(u\)≥\(η​Tc2−1\)​c0η−𝒪​\(1\)=c0c2​T−c0η−𝒪​\(1\)=c0c2​T−𝒪​\(1\),\\sum\_\{t=1\}^\{T\}\\ell\_\{t\}\(x\_\{t\}\)\-\\ell\_\{t\}\(u\)~\\geq~\\left\(\\frac\{\\eta T\}\{c\_\{2\}\}\-1\\right\)\\frac\{c\_\{0\}\}\{\\eta\}\-\\mathcal\{O\}\(1\)=\\frac\{c\_\{0\}\}\{c\_\{2\}\}T\-\\frac\{c\_\{0\}\}\{\\eta\}\-\\mathcal\{O\}\(1\)=\\frac\{c\_\{0\}\}\{c\_\{2\}\}T\-\\mathcal\{O\}\(1\)\\,,which concludes the proof of Theorem[2](https://arxiv.org/html/2605.26373#Thmtheorem2)\.

## Appendix CProof of Theorem[3](https://arxiv.org/html/2605.26373#Thmtheorem3): Expected Regret Bound in the Bandit Setting

Recall first that the regret of bandit OGD \(Algorithm[1](https://arxiv.org/html/2605.26373#alg1)\) is defined for any time horizonT≥1T~\\geq~1by:

RT:=∑t=1Tℓt​\(x^t\)−minx∈𝒳​∑t=1Tℓt​\(x\)=∑t=1Tht​\(z^t\)−minz∈𝒴​∑t=1Tht​\(z\),R\_\{T\}:=\\sum\_\{t=1\}^\{T\}\\ell\_\{t\}\(\\hat\{x\}\_\{t\}\)\-\\min\_\{x\\in\{\\mathcal\{X\}\}\}\\sum\_\{t=1\}^\{T\}\\ell\_\{t\}\(x\)\\,=\\sum\_\{t=1\}^\{T\}h\_\{t\}\(\\hat\{z\}\_\{t\}\)\-\\min\_\{z\\in\{\\mathcal\{Y\}\}\}\\sum\_\{t=1\}^\{T\}h\_\{t\}\(z\)\\,,\(69\)where the last equality follows from using the notationz^t:=q​\(x^t\)\\hat\{z\}\_\{t\}:=q\(\\hat\{x\}\_\{t\}\)for anyt≥1t~\\geq~1and recalling thatℓt=ht∘q\\ell\_\{t\}=h\_\{t\}\\circ qwithqqbeing a diffeomorphism\.

To prove Theorem[3](https://arxiv.org/html/2605.26373#Thmtheorem3), we first decompose the expected regret into four different error terms\.

###### Lemma 10\(Expected regret bound decomposition\)\. Fix a time horizonT≥1T~\\geq~1\. Letz⋆∈arg​minz∈𝒴​∑t=1Tht​\(z\),zδ⋆∈arg​minz∈𝒴δ​∑t=1Tht​\(z\)z^\{\\star\}\\in\\operatorname\*\{arg\\,min\}\_\{z\\in\{\\mathcal\{Y\}\}\}\\sum\_\{t=1\}^\{T\}h\_\{t\}\(z\),z^\{\\star\}\_\{\\delta\}\\in\\operatorname\*\{arg\\,min\}\_\{z\\in\{\\mathcal\{Y\}\}\_\{\\delta\}\}\\sum\_\{t=1\}^\{T\}h\_\{t\}\(z\)\\,and letℱt:=σ​\(ζ1,…,ζt\)\\mathcal\{F\}\_\{t\}:=\\sigma\(\\zeta\_\{1\},\\dots,\\zeta\_\{t\}\)be the filtration ofσ\\sigma\-algebras induced by the random variables\(ζt\)\(\\zeta\_\{t\}\)used in bandit OGD\. Then the expected regret𝔼​\[RT\]\\mathbb\{E\}\[R\_\{T\}\]of bandit OGD \(Algorithm[1](https://arxiv.org/html/2605.26373#alg1)\) can be upper\-bounded as follows:𝔼​\[RT\]≤RT\(1\)\+RT\(2\)\+RT\(3\)\+RT\(4\),\\mathbb\{E\}\[R\_\{T\}\]~\\leq~R\_\{T\}^\{\(1\)\}\+R\_\{T\}^\{\(2\)\}\+R\_\{T\}^\{\(3\)\}\+R\_\{T\}^\{\(4\)\}\\,,where each error term in the right\-hand side above is defined as follows:RT\(1\)\\displaystyle R\_\{T\}^\{\(1\)\}:=𝔼​\[∑t=1T⟨g~t,zt−zδ⋆⟩\],\\displaystyle:=\\mathbb\{E\}\\left\[\\sum\_\{t=1\}^\{T\}\\langle\\tilde\{g\}\_\{t\},z\_\{t\}\-z^\{\\star\}\_\{\\delta\}\\rangle\\right\]\\,,\(perturbed OMD regret for stochastic linearized losses\)\(70\)RT\(2\)\\displaystyle R\_\{T\}^\{\(2\)\}:=∑t=1Tht​\(zδ⋆\)−ht​\(z⋆\),\\displaystyle:=\\sum\_\{t=1\}^\{T\}h\_\{t\}\(z^\{\\star\}\_\{\\delta\}\)\-h\_\{t\}\(z^\{\\star\}\)\\,,\(domain shrinkage error\)\(71\)RT\(3\)\\displaystyle R\_\{T\}^\{\(3\)\}:=𝔼​\[∑t=1T⟨bt,zδ⋆−zt⟩\],\\displaystyle:=\\mathbb\{E\}\\left\[\\sum\_\{t=1\}^\{T\}\\langle b\_\{t\},z^\{\\star\}\_\{\\delta\}\-z\_\{t\}\\rangle\\right\]\\,,\(cumulative gradient estimation bias\)\(72\)bt\\displaystyle b\_\{t\}:=𝔼​\[g~t\|ℱt−1\]−∇ht​\(zt\),\\displaystyle:=\\mathbb\{E\}\[\\tilde\{g\}\_\{t\}\|\\mathcal\{F\}\_\{t\-1\}\]\-\\nabla\\mkern\-2\.5muh\_\{t\}\(z\_\{t\}\)\\,,\(gradient estimation bias\)\(73\)RT\(4\)\\displaystyle R\_\{T\}^\{\(4\)\}:=𝔼​\[∑t=1Tht​\(z^t\)−ht​\(zt\)\]\\displaystyle:=\\mathbb\{E\}\\left\[\\sum\_\{t=1\}^\{T\}h\_\{t\}\(\\hat\{z\}\_\{t\}\)\-h\_\{t\}\(z\_\{t\}\)\\right\]\(smoothing error\)\.\\displaystyle\\hfill\\text\{\(smoothing error\)\}\\,\.\(74\)

###### Proof\.

Using the regret definition \([69](https://arxiv.org/html/2605.26373#A3.E69)\), we have the following decomposition:

𝔼​\[RT\]=𝔼​\[∑t=1Tht​\(z^t\)−ht​\(zδ⋆\)\]\+∑t=1Tht​\(zδ⋆\)−ht​\(z⋆\)=𝔼​\[∑t=1Tht​\(z^t\)−ht​\(zδ⋆\)\]\+RT\(2\)\.\\mathbb\{E\}\[R\_\{T\}\]=\\mathbb\{E\}\\left\[\\sum\_\{t=1\}^\{T\}h\_\{t\}\(\\hat\{z\}\_\{t\}\)\-h\_\{t\}\(z^\{\\star\}\_\{\\delta\}\)\\right\]\+\\sum\_\{t=1\}^\{T\}h\_\{t\}\(z^\{\\star\}\_\{\\delta\}\)\-h\_\{t\}\(z^\{\\star\}\)=\\mathbb\{E\}\\left\[\\sum\_\{t=1\}^\{T\}h\_\{t\}\(\\hat\{z\}\_\{t\}\)\-h\_\{t\}\(z^\{\\star\}\_\{\\delta\}\)\\right\]\+R\_\{T\}^\{\(2\)\}\\,\.\(75\)To control regret against the shrunk comparatorzδ⋆z\_\{\\delta\}^\{\\star\}, we start by decomposing the instantaneous regret\. Sinceztz\_\{t\}isℱt−1\\mathcal\{F\}\_\{t\-1\}\-measurable, we have:

𝔼​\[ht​\(z^t\)−ht​\(zδ⋆\)\|ℱt−1\]=𝔼​\[ht​\(z^t\)−ht​\(zt\)\|ℱt−1\]\+ht​\(zt\)−ht​\(zδ⋆\)\.\\mathbb\{E\}\[h\_\{t\}\(\\hat\{z\}\_\{t\}\)\-h\_\{t\}\(z^\{\\star\}\_\{\\delta\}\)\|\\mathcal\{F\}\_\{t\-1\}\]=\\mathbb\{E\}\[h\_\{t\}\(\\hat\{z\}\_\{t\}\)\-h\_\{t\}\(z\_\{t\}\)\|\\mathcal\{F\}\_\{t\-1\}\]\+h\_\{t\}\(z\_\{t\}\)\-h\_\{t\}\(z\_\{\\delta\}^\{\\star\}\)\\,\.\(76\)Then by convexity ofhth\_\{t\}, we have:

ht​\(zt\)−ht​\(zδ⋆\)≤⟨∇ht​\(zt\),zt−zδ⋆⟩\.h\_\{t\}\(z\_\{t\}\)\-h\_\{t\}\(z\_\{\\delta\}^\{\\star\}\)~\\leq~\\langle\\nabla\\mkern\-2\.5muh\_\{t\}\(z\_\{t\}\),z\_\{t\}\-z\_\{\\delta\}^\{\\star\}\\rangle\\,\.\(77\)Recall now that we denote bybtb\_\{t\}the bias due to one\-point smoothing defined by:

bt=𝔼​\[g~t\|ℱt−1\]−∇ht​\(zt\)\.b\_\{t\}=\\mathbb\{E\}\[\\tilde\{g\}\_\{t\}\|\\mathcal\{F\}\_\{t\-1\}\]\-\\nabla\\mkern\-2\.5muh\_\{t\}\(z\_\{t\}\)\\,\.Note here thatg~t:=Jq​\(xt\)−⊤​gt\\tilde\{g\}\_\{t\}:=J\_\{q\}\(x\_\{t\}\)^\{\-\\top\}g\_\{t\}is is not computed and just used here in the analysis whereasgtg\_\{t\}is the one\-point estimator actually computed in BOGD\. It follows then that:

⟨∇ht​\(zt\),zt−zδ⋆⟩=⟨𝔼​\[g~t\|ℱt−1\],zt−zδ⋆⟩−⟨bt,zt−zδ⋆⟩\.\\langle\\nabla\\mkern\-2\.5muh\_\{t\}\(z\_\{t\}\),z\_\{t\}\-z\_\{\\delta\}^\{\\star\}\\rangle=\\langle\\mathbb\{E\}\[\\tilde\{g\}\_\{t\}\|\\mathcal\{F\}\_\{t\-1\}\],z\_\{t\}\-z\_\{\\delta\}^\{\\star\}\\rangle\-\\langle b\_\{t\},z\_\{t\}\-z\_\{\\delta\}^\{\\star\}\\rangle\\,\.\(78\)Combining \([76](https://arxiv.org/html/2605.26373#A3.E76)\), \([77](https://arxiv.org/html/2605.26373#A3.E77)\) and \([78](https://arxiv.org/html/2605.26373#A3.E78)\), taking expectations and noting thatzt−zδ⋆z\_\{t\}\-z\_\{\\delta\}^\{\\star\}isℱt−1\\mathcal\{F\}\_\{t\-1\}\-measurable, we obtain:

𝔼​\[ht​\(zt\)−ht​\(zδ⋆\)\]≤𝔼​\[⟨g~t,zt−zδ⋆⟩\]−𝔼​\[⟨bt,zt−zδ⋆⟩\]\+𝔼​\[ht​\(z^t\)−ht​\(zt\)\]\.\\mathbb\{E\}\[h\_\{t\}\(z\_\{t\}\)\-h\_\{t\}\(z\_\{\\delta\}^\{\\star\}\)\]~\\leq~\\mathbb\{E\}\[\\langle\\tilde\{g\}\_\{t\},z\_\{t\}\-z\_\{\\delta\}^\{\\star\}\\rangle\]\-\\mathbb\{E\}\[\\langle b\_\{t\},z\_\{t\}\-z\_\{\\delta\}^\{\\star\}\\rangle\]\+\\mathbb\{E\}\[h\_\{t\}\(\\hat\{z\}\_\{t\}\)\-h\_\{t\}\(z\_\{t\}\)\]\\,\.Summing the above inequality overttyields:

𝔼​\[∑t=1Tht​\(zt\)−ht​\(zδ⋆\)\]\\displaystyle\\mathbb\{E\}\\left\[\\sum\_\{t=1\}^\{T\}h\_\{t\}\(z\_\{t\}\)\-h\_\{t\}\(z\_\{\\delta\}^\{\\star\}\)\\right\]≤𝔼​\[∑t=1T⟨g~t,zt−zδ⋆⟩\]\+𝔼​\[∑t=1T⟨bt,zδ⋆−zt⟩\]\+𝔼​\[∑t=1Tht​\(z^t\)−ht​\(zt\)\]\\displaystyle~\\leq~\\mathbb\{E\}\\left\[\\sum\_\{t=1\}^\{T\}\\langle\\tilde\{g\}\_\{t\},z\_\{t\}\-z\_\{\\delta\}^\{\\star\}\\rangle\\right\]\+\\mathbb\{E\}\\left\[\\sum\_\{t=1\}^\{T\}\\langle b\_\{t\},z^\{\\star\}\_\{\\delta\}\-z\_\{t\}\\rangle\\right\]\+\\mathbb\{E\}\\left\[\\sum\_\{t=1\}^\{T\}h\_\{t\}\(\\hat\{z\}\_\{t\}\)\-h\_\{t\}\(z\_\{t\}\)\\right\]=RT\(1\)\+RT\(3\)\+RT\(4\)\.\\displaystyle=R\_\{T\}^\{\(1\)\}\+R\_\{T\}^\{\(3\)\}\+R\_\{T\}^\{\(4\)\}\.\(79\)Finally, combining \([C](https://arxiv.org/html/2605.26373#A3.Ex143)\) with \([75](https://arxiv.org/html/2605.26373#A3.E75)\) gives the desired bound:

𝔼​\[RT\]≤RT\(1\)\+RT\(2\)\+RT\(3\)\+RT\(4\)\.\\mathbb\{E\}\[R\_\{T\}\]~\\leq~R\_\{T\}^\{\(1\)\}\+R\_\{T\}^\{\(2\)\}\+R\_\{T\}^\{\(3\)\}\+R\_\{T\}^\{\(4\)\}\\,\.∎

Given the regret decomposition of Lemma[10](https://arxiv.org/html/2605.26373#Thmlemma10), we control each of the error termsRT\(i\),i∈\{1,2,3,4\}R\_\{T\}^\{\(i\)\},i\\in\\\{1,2,3,4\\\}defined in \([70](https://arxiv.org/html/2605.26373#A3.E70)\) to \([74](https://arxiv.org/html/2605.26373#A3.E74)\) respectively, in a dedicated lemma\.

###### Lemma 11\(Domain shrinkage error\)\. LetT≥1T~\\geq~1, letz⋆∈arg​minz∈𝒴​∑t=1Tht​\(z\)z^\{\\star\}\\in\\operatorname\*\{arg\\,min\}\_\{z\\in\{\\mathcal\{Y\}\}\}\\sum\_\{t=1\}^\{T\}h\_\{t\}\(z\), andzδ⋆∈arg​minz∈𝒴δ​∑t=1Tht​\(z\)\.z^\{\\star\}\_\{\\delta\}\\in\\operatorname\*\{arg\\,min\}\_\{z\\in\{\\mathcal\{Y\}\}\_\{\\delta\}\}\\sum\_\{t=1\}^\{T\}h\_\{t\}\(z\)\\,\.Then we have:RT\(2\)=∑t=1Tht​\(zδ⋆\)−ht​\(z⋆\)≤δ​GF​D​T,R\_\{T\}^\{\(2\)\}=\\sum\_\{t=1\}^\{T\}h\_\{t\}\(z\_\{\\delta\}^\{\\star\}\)\-h\_\{t\}\(z^\{\\star\}\)~\\leq~\\delta G\_\{F\}DT\\,,where we recall thatδ\\deltais the smoothing parameter in the one\-point zeroth\-order gradient estimator,GFG\_\{F\}is the uniform Lipschitz constant ofℓt\\ell\_\{t\}andDDis the finite diameter of the compact set𝒳\{\\mathcal\{X\}\}\.

###### Proof\.

Since𝒴δ=q​\(𝒳δ\)\{\\mathcal\{Y\}\}\_\{\\delta\}=q\(\{\\mathcal\{X\}\}\_\{\\delta\}\),qqis a diffeomorphism,𝒳δ=\(1−δ\)​𝒳\{\\mathcal\{X\}\}\_\{\\delta\}=\(1\-\\delta\)\{\\mathcal\{X\}\}andℓt=ht∘q\\ell\_\{t\}=h\_\{t\}\\circ q, the comparator term over the shrunk set𝒳δ\{\\mathcal\{X\}\}\_\{\\delta\}can be rewritten as follows:

minz∈𝒴δ​∑t=1Tht​\(z\)=minx∈𝒳δ​∑t=1Tℓt​\(x\)=minx∈𝒳​∑t=1Tℓt​\(\(1−δ\)​x\)\.\\min\_\{z\\in\{\\mathcal\{Y\}\}\_\{\\delta\}\}\\sum\_\{t=1\}^\{T\}h\_\{t\}\(z\)=\\min\_\{x\\in\{\\mathcal\{X\}\}\_\{\\delta\}\}\\sum\_\{t=1\}^\{T\}\\ell\_\{t\}\(x\)=\\min\_\{x\\in\{\\mathcal\{X\}\}\}\\sum\_\{t=1\}^\{T\}\\ell\_\{t\}\(\(1\-\\delta\)x\)\\,\.\(80\)Then for anyx∈𝒳,x\\in\{\\mathcal\{X\}\},we have:

∑t=1Tℓt​\(\(1−δ\)​x\)≤∑t=1T\|ℓt​\(\(1−δ\)​x\)−ℓt​\(x\)\|\+∑t=1Tℓt​\(x\)≤δ​GF​D​T\+∑t=1Tht​\(q​\(x\)\),\\sum\_\{t=1\}^\{T\}\\ell\_\{t\}\(\(1\-\\delta\)x\)~\\leq~\\sum\_\{t=1\}^\{T\}\|\\ell\_\{t\}\(\(1\-\\delta\)x\)\-\\ell\_\{t\}\(x\)\|\+\\sum\_\{t=1\}^\{T\}\\ell\_\{t\}\(x\)~\\leq~\\delta G\_\{F\}DT\+\\sum\_\{t=1\}^\{T\}h\_\{t\}\(q\(x\)\)\\,,where the last inequality uses uniformGFG\_\{F\}\-Lipschitzness of the loss functionsℓt\\ell\_\{t\}and boundedness of the set𝒳\{\\mathcal\{X\}\}which has finite diameterDD\.

Taking the minimum overx∈𝒳x\\in\{\\mathcal\{X\}\}, using \([80](https://arxiv.org/html/2605.26373#A3.E80)\) and recalling thatz⋆∈arg​minz∈𝒴​∑t=1Tht​\(z\)z^\{\\star\}\\in\\operatorname\*\{arg\\,min\}\_\{z\\in\{\\mathcal\{Y\}\}\}\\sum\_\{t=1\}^\{T\}h\_\{t\}\(z\)andzδ⋆∈arg​minz∈𝒴δ​∑t=1Tht​\(z\)z^\{\\star\}\_\{\\delta\}\\in\\operatorname\*\{arg\\,min\}\_\{z\\in\{\\mathcal\{Y\}\}\_\{\\delta\}\}\\sum\_\{t=1\}^\{T\}h\_\{t\}\(z\)yields the desired inequality:

RT\(2\)=∑t=1Tht​\(zδ⋆\)−ht​\(z⋆\)≤δ​GF​D​T\.R\_\{T\}^\{\(2\)\}=\\sum\_\{t=1\}^\{T\}h\_\{t\}\(z\_\{\\delta\}^\{\\star\}\)\-h\_\{t\}\(z^\{\\star\}\)~\\leq~\\delta G\_\{F\}DT\\,\.∎

###### Lemma 12\(Gradient estimation bias\)\. For anyT≥1T~\\geq~1, we have:\|RT\(3\)\|=\|𝔼​\[∑t=1T⟨bt,zδ⋆−zt⟩\]\|≤δ​G2​H​D​T,\|R\_\{T\}^\{\(3\)\}\|=\\left\|\\mathbb\{E\}\\left\[\\sum\_\{t=1\}^\{T\}\\langle b\_\{t\},z\_\{\\delta\}^\{\\star\}\-z\_\{t\}\\rangle\\right\]\\right\|~\\leq~\\delta G^\{2\}HDT\\,,where we recall thatHHis the uniform smoothness constant of the lossℓt\.\\ell\_\{t\}\.

###### Proof\.

First, using the Cauchy\-Schwartz inequality, it follows that:

\|RT\(3\)\|=\|𝔼​\[∑t=1T⟨bt,zδ⋆−zt⟩\]\|≤𝔼​\[∑t=1T‖bt‖⋅‖zt−zδ⋆‖\]\.\|R\_\{T\}^\{\(3\)\}\|=\\left\|\\mathbb\{E\}\\left\[\\sum\_\{t=1\}^\{T\}\\langle b\_\{t\},z\_\{\\delta\}^\{\\star\}\-z\_\{t\}\\rangle\\right\]\\right\|~\\leq~\\mathbb\{E\}\\left\[\\sum\_\{t=1\}^\{T\}\\\|b\_\{t\}\\\|\\cdot\\\|z\_\{t\}\-z\_\{\\delta\}^\{\\star\}\\\|\\right\]\\,\.\(81\)
We control each one of the terms in the right\-hand side separately\.

Sincezt=q​\(xt\)z\_\{t\}=q\(x\_\{t\}\)andzδ⋆=q​\(xδ⋆\)z\_\{\\delta\}^\{\\star\}=q\(x\_\{\\delta\}^\{\\star\}\)for somext∈𝒳x\_\{t\}\\in\{\\mathcal\{X\}\}andxδ⋆∈𝒳δ⊂𝒳x\_\{\\delta\}^\{\\star\}\\in\{\\mathcal\{X\}\}\_\{\\delta\}\\subset\{\\mathcal\{X\}\}, we have:

‖zt−zδ⋆‖=‖q​\(xt\)−q​\(xδ⋆\)‖≤G​‖xt−xδ⋆‖≤G​D,\\\|z\_\{t\}\-z\_\{\\delta\}^\{\\star\}\\\|=\\\|q\(x\_\{t\}\)\-q\(x\_\{\\delta\}^\{\\star\}\)\\\|~\\leq~G\\\|x\_\{t\}\-x\_\{\\delta\}^\{\\star\}\\\|~\\leq~GD\\,,\(82\)where the first inequality follows from usingGG\-Lipschitzness ofqq\(Assumption[2](https://arxiv.org/html/2605.26373#Thmassumption2)\) and the second one uses the definition of the diameter of𝒳\{\\mathcal\{X\}\}supposed to be finite\.

We now control the norm of the bias‖bt‖\\\|b\_\{t\}\\\|for anytt\. Define for anyx∈𝒳x\\in\{\\mathcal\{X\}\}the smoothed loss:

ℓ¯t​\(x\):=𝔼v∼𝒰​\(𝔹\)​\[ℓt​\(x\+δ​v\)\]\\bar\{\\ell\}\_\{t\}\(x\):=\\mathbb\{E\}\_\{v\\sim\\mathcal\{U\}\(\\mathbb\{B\}\)\}\[\\ell\_\{t\}\(x\+\\delta v\)\]\\,wherevvis a random variable with uniform distribution𝒰​\(𝔹\)\\mathcal\{U\}\(\\mathbb\{B\}\)on the unit ball𝔹\\mathbb\{B\}\. It is known fromFlaxmanet al\.\[[2005](https://arxiv.org/html/2605.26373#bib.bib43), Lemma 1\]that the conditional expectation of the spherical estimatorgtg\_\{t\}is equal to the gradient of the ball\-smoothed loss:

𝔼​\[gt\|ℱt−1\]=∇ℓ¯t​\(xt\)\.\\mathbb\{E\}\[g\_\{t\}\|\\mathcal\{F\}\_\{t\-1\}\]=\\nabla\\mkern\-2\.5mu\\bar\{\\ell\}\_\{t\}\(x\_\{t\}\)\\,\.The \(hidden\-space ghost\) gradient estimator of∇ht​\(zt\)\\nabla\\mkern\-2\.5muh\_\{t\}\(z\_\{t\}\)is

g~t=Jq​\(xt\)−⊤​gt\.\\tilde\{g\}\_\{t\}=J\_\{q\}\(x\_\{t\}\)^\{\-\\top\}g\_\{t\}\\,\.\(Recall that∇ℓt​\(xt\)=Jq​\(xt\)⊤​∇ht​\(zt\)\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\(x\_\{t\}\)=J\_\{q\}\(x\_\{t\}\)^\{\\top\}\\nabla\\mkern\-2\.5muh\_\{t\}\(z\_\{t\}\)by the chain rule\.\) Sincextx\_\{t\}isℱt−1\\mathcal\{F\}\_\{t\-1\}\-measurable, it follows that:

𝔼​\[g~t\|ℱt−1\]=Jq​\(xt\)−⊤​𝔼​\[gt\|ℱt−1\]=Jq​\(xt\)−⊤​∇ℓ¯t​\(xt\)\.\\mathbb\{E\}\[\\tilde\{g\}\_\{t\}\|\\mathcal\{F\}\_\{t\-1\}\]=J\_\{q\}\(x\_\{t\}\)^\{\-\\top\}\\mathbb\{E\}\[g\_\{t\}\|\\mathcal\{F\}\_\{t\-1\}\]=J\_\{q\}\(x\_\{t\}\)^\{\-\\top\}\\nabla\\mkern\-2\.5mu\\bar\{\\ell\}\_\{t\}\(x\_\{t\}\)\\,\.\(83\)Moreover, we have by the chain rule∇ℓt​\(xt\)=Jq​\(xt\)⊤​∇ht​\(zt\)\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\(x\_\{t\}\)=J\_\{q\}\(x\_\{t\}\)^\{\\top\}\\nabla\\mkern\-2\.5muh\_\{t\}\(z\_\{t\}\)which implies that∇ht​\(zt\)=Jq​\(xt\)−⊤​∇ℓt​\(xt\)\.\\nabla\\mkern\-2\.5muh\_\{t\}\(z\_\{t\}\)=J\_\{q\}\(x\_\{t\}\)^\{\-\\top\}\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\(x\_\{t\}\)\.Therefore, using this identity together with \([83](https://arxiv.org/html/2605.26373#A3.E83)\), we can express the biasbtb\_\{t\}as follows:

bt\\displaystyle b\_\{t\}=𝔼​\[g~t\|ℱt−1\]−∇ht​\(zt\)\\displaystyle=\\mathbb\{E\}\[\\tilde\{g\}\_\{t\}\|\\mathcal\{F\}\_\{t\-1\}\]\-\\nabla\\mkern\-2\.5muh\_\{t\}\(z\_\{t\}\)=Jq​\(xt\)−⊤​\(∇ℓ¯t​\(xt\)−∇ℓt​\(xt\)\)\\displaystyle=J\_\{q\}\(x\_\{t\}\)^\{\-\\top\}\(\\nabla\\mkern\-2\.5mu\\bar\{\\ell\}\_\{t\}\(x\_\{t\}\)\-\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\(x\_\{t\}\)\)=Jq​\(xt\)−⊤​𝔼v∼𝒰​\(𝔹\)​\[∇ℓt​\(xt\+δ​v\)−∇ℓt​\(xt\)\]\.\\displaystyle=J\_\{q\}\(x\_\{t\}\)^\{\-\\top\}\\mathbb\{E\}\_\{v\\sim\\mathcal\{U\}\(\\mathbb\{B\}\)\}\[\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\(x\_\{t\}\+\\delta v\)\-\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\(x\_\{t\}\)\]\\,\.Taking the norm, using boundedness of the Jacobian ofqq\(Assumption[2](https://arxiv.org/html/2605.26373#Thmassumption2)\) and smoothness of the loss functionℓt\\ell\_\{t\}\(Assumption[5](https://arxiv.org/html/2605.26373#Thmassumption5)\), we have:

‖bt‖2≤‖Jq​\(xt\)−⊤‖​𝔼v∼𝒰​\(𝔹\)​\[‖∇ℓt​\(xt\+δ​v\)−∇ℓt​\(xt\)‖\]≤δ​G​H​‖v‖2=δ​G​H\.\\\|b\_\{t\}\\\|\_\{2\}~\\leq~\\\|J\_\{q\}\(x\_\{t\}\)^\{\-\\top\}\\\|\\mathbb\{E\}\_\{v\\sim\\mathcal\{U\}\(\\mathbb\{B\}\)\}\[\\\|\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\(x\_\{t\}\+\\delta v\)\-\\nabla\\mkern\-2\.5mu\\ell\_\{t\}\(x\_\{t\}\)\\\|\]~\\leq~\\delta GH\\\|v\\\|\_\{2\}=\\delta GH\\,\.\(84\)Combining \([84](https://arxiv.org/html/2605.26373#A3.E84)\) and \([82](https://arxiv.org/html/2605.26373#A3.E82)\) in \([81](https://arxiv.org/html/2605.26373#A3.E81)\), we obtain the desired inequality:

\|RT\(3\)\|≤δ​G2​H​D​T\.\|R\_\{T\}^\{\(3\)\}\|~\\leq~\\delta G^\{2\}HDT\\,\.∎

###### Lemma 13\(Smoothing error\)\. For anyT≥1T~\\geq~1, we have:\|RT\(4\)\|≤\|𝔼​\[∑t=1Tht​\(z^t\)−ht​\(zt\)\]\|≤δ​GF​T\.\|R\_\{T\}^\{\(4\)\}\|~\\leq~\\left\|\\mathbb\{E\}\\left\[\\sum\_\{t=1\}^\{T\}h\_\{t\}\(\\hat\{z\}\_\{t\}\)\-h\_\{t\}\(z\_\{t\}\)\\right\]\\right\|~\\leq~\\delta G\_\{F\}T\\,\.

###### Proof\.

The desired bound follows from the following:

\|RT\(4\)\|\\displaystyle\|R\_\{T\}^\{\(4\)\}\|=\|𝔼​\[∑t=1Tht​\(z^t\)−ht​\(zt\)\]\|\\displaystyle=\\left\|\\mathbb\{E\}\\left\[\\sum\_\{t=1\}^\{T\}h\_\{t\}\(\\hat\{z\}\_\{t\}\)\-h\_\{t\}\(z\_\{t\}\)\\right\]\\right\|=\|𝔼​\[∑t=1Tℓt​\(x^t\)−ℓt​\(xt\)\]\|\\displaystyle=\\left\|\\mathbb\{E\}\\left\[\\sum\_\{t=1\}^\{T\}\\ell\_\{t\}\(\\hat\{x\}\_\{t\}\)\-\\ell\_\{t\}\(x\_\{t\}\)\\right\]\\right\|\(ℓt=ht∘q\\ell\_\{t\}=h\_\{t\}\\circ q,z^t=q​\(x^t\),zt=q​\(xt\)\\hat\{z\}\_\{t\}=q\(\\hat\{x\}\_\{t\}\),z\_\{t\}=q\(x\_\{t\}\)\)≤GF​𝔼​\[∑t=1T‖x^t−xt‖\]\\displaystyle~\\leq~G\_\{F\}\\mathbb\{E\}\\left\[\\sum\_\{t=1\}^\{T\}\\\|\\hat\{x\}\_\{t\}\-x\_\{t\}\\\|\\right\]\(byGFG\_\{F\}\-Lipschitzness ofℓt\\ell\_\{t\}\(Assumption[3](https://arxiv.org/html/2605.26373#Thmassumption3)\)\)=GF​𝔼​\[∑t=1Tδ​‖ζt‖\]\\displaystyle=G\_\{F\}\\mathbb\{E\}\\left\[\\sum\_\{t=1\}^\{T\}\\delta\\\|\\zeta\_\{t\}\\\|\\right\]\(x^t=xt\+δ​ζt\\hat\{x\}\_\{t\}=x\_\{t\}\+\\delta\\zeta\_\{t\}in BOGD \(Algorithm[1](https://arxiv.org/html/2605.26373#alg1)\)\)=δ​GF​T\\displaystyle=\\delta G\_\{F\}T\(‖ζt‖=1\)\.\\displaystyle\\hfill\\text\{\($\\\|\\zeta\_\{t\}\\\|=1$\)\}\\,\.\(85\)∎

###### Lemma 14\(Perturbed ghost OMD regret for random linear losses\)\. Letz1∈𝒴δ\.z\_\{1\}\\in\{\\mathcal\{Y\}\}\_\{\\delta\}\.Suppose the sequence\(zt\)\(z\_\{t\}\)is defined as:zt\+1=rt\+1\+yt\+1,yt\+1=arg​miny∈𝒴δ\{g~t⊤\(y−zt\)\+1ηDR\(y\|\|zt\)\}\.z\_\{t\+1\}=r\_\{t\+1\}\+y\_\{t\+1\}\\,,\\quad y\_\{t\+1\}=\\operatorname\*\{arg\\,min\}\_\{y\\in\{\\mathcal\{Y\}\}\_\{\\delta\}\}\\\{\\tilde\{g\}\_\{t\}^\{\\top\}\(y\-z\_\{t\}\)\+\\frac\{1\}\{\\eta\}D\_\{R\}\(y\|\|z\_\{t\}\)\\\}\\,\.Then for everyzδ⋆∈𝒴δ,z\_\{\\delta\}^\{\\star\}\\in\{\\mathcal\{Y\}\}\_\{\\delta\},∑t=1T⟨g~t,zt−zδ⋆⟩≤DR\(zδ⋆\|\|z1\)η\+η2​∑t=1T‖g~t‖22\+Gη​∑t=1T‖rt\+1‖2\.\\sum\_\{t=1\}^\{T\}\\langle\\tilde\{g\}\_\{t\},z\_\{t\}\-z\_\{\\delta\}^\{\\star\}\\rangle~\\leq~\\frac\{D\_\{R\}\(z\_\{\\delta\}^\{\\star\}\|\|z\_\{1\}\)\}\{\\eta\}\+\\frac\{\\eta\}\{2\}\\sum\_\{t=1\}^\{T\}\\\|\\tilde\{g\}\_\{t\}\\\|\_\{2\}^\{2\}\+\\frac\{G\}\{\\eta\}\\sum\_\{t=1\}^\{T\}\\\|r\_\{t\+1\}\\\|\_\{2\}\\,\.If in addition‖g~t‖≤G~F\\\|\\tilde\{g\}\_\{t\}\\\|~\\leq~\\tilde\{G\}\_\{F\}and‖rt\+1‖≤C^η\\\|r\_\{t\+1\}\\\|~\\leq~\\hat\{C\}\_\{\\eta\}, then:RT\(1\)=𝔼​\[∑t=1T⟨g~t,zt−zδ⋆⟩\]≤D1η\+η​G~F2​T2\+C^η​G​Tη\.R\_\{T\}^\{\(1\)\}=\\mathbb\{E\}\\left\[\\sum\_\{t=1\}^\{T\}\\langle\\tilde\{g\}\_\{t\},z\_\{t\}\-z\_\{\\delta\}^\{\\star\}\\rangle\\right\]~\\leq~\\frac\{D\_\{1\}\}\{\\eta\}\+\\frac\{\\eta\\tilde\{G\}\_\{F\}^\{2\}T\}\{2\}\+\\frac\{\\hat\{C\}\_\{\\eta\}GT\}\{\\eta\}\\,\.\(86\)

###### Proof\.

First, we prove the following inequality which follows from using the first\-order optimality condition \(from the ghost OMD update\) together with Bregman three\-point identity\. For any comparatorzδ⋆∈𝒴δz\_\{\\delta\}^\{\\star\}\\in\{\\mathcal\{Y\}\}\_\{\\delta\}, we have:

⟨g~t,zt−zδ⋆⟩≤DR\(zδ⋆\|\|zt\)−DR\(zδ⋆\|\|yt\+1\)η\+η2​‖gt‖22\.\\langle\\tilde\{g\}\_\{t\},z\_\{t\}\-z\_\{\\delta\}^\{\\star\}\\rangle~\\leq~\\frac\{D\_\{R\}\(z\_\{\\delta\}^\{\\star\}\|\|z\_\{t\}\)\-D\_\{R\}\(z\_\{\\delta\}^\{\\star\}\|\|y\_\{t\+1\}\)\}\{\\eta\}\+\\frac\{\\eta\}\{2\}\\\|g\_\{t\}\\\|\_\{2\}^\{2\}\\,\.\(87\)
We now provide the proof of \([87](https://arxiv.org/html/2605.26373#A3.E87)\)\. By definition ofyt\+1y\_\{t\+1\}, using the first\-order optimality condition gives for everyzδ⋆∈𝒴δz\_\{\\delta\}^\{\\star\}\\in\{\\mathcal\{Y\}\}\_\{\\delta\},

⟨g~t\+1η​\(∇R​\(yt\+1\)−∇R​\(zt\)\),zδ⋆−yt\+1⟩≥0\.\\left\\langle\\tilde\{g\}\_\{t\}\+\\frac\{1\}\{\\eta\}\(\\nabla\\mkern\-2\.5muR\(y\_\{t\+1\}\)\-\\nabla\\mkern\-2\.5muR\(z\_\{t\}\)\),z\_\{\\delta\}^\{\\star\}\-y\_\{t\+1\}\\right\\rangle~\\geq~0\\,\.Rearranging this inequality, we obtain:

⟨g~t,yt\+1−zδ⋆⟩≤1η​⟨∇R​\(yt\+1\)−∇R​\(zt\),zδ⋆−yt\+1⟩\.\\langle\\tilde\{g\}\_\{t\},y\_\{t\+1\}\-z\_\{\\delta\}^\{\\star\}\\rangle~\\leq~\\frac\{1\}\{\\eta\}\\langle\\nabla\\mkern\-2\.5muR\(y\_\{t\+1\}\)\-\\nabla\\mkern\-2\.5muR\(z\_\{t\}\),z\_\{\\delta\}^\{\\star\}\-y\_\{t\+1\}\\rangle\\,\.\(88\)Using the three\-point identity, we have:

⟨∇R\(yt\+1\)−∇R\(zt\),zδ⋆−yt\+1⟩=DR\(zδ⋆\|\|zt\)−DR\(zδ⋆\|\|yt\+1\)−DR\(yt\+1\|\|zt\)\.\\langle\\nabla\\mkern\-2\.5muR\(y\_\{t\+1\}\)\-\\nabla\\mkern\-2\.5muR\(z\_\{t\}\),z\_\{\\delta\}^\{\\star\}\-y\_\{t\+1\}\\rangle=D\_\{R\}\(z\_\{\\delta\}^\{\\star\}\|\|z\_\{t\}\)\-D\_\{R\}\(z\_\{\\delta\}^\{\\star\}\|\|y\_\{t\+1\}\)\-D\_\{R\}\(y\_\{t\+1\}\|\|z\_\{t\}\)\\,\.\(89\)Combining \([88](https://arxiv.org/html/2605.26373#A3.E88)\) and \([89](https://arxiv.org/html/2605.26373#A3.E89)\) yields:

⟨g~t,yt\+1−zδ⋆⟩≤1η\[DR\(zδ⋆\|\|zt\)−DR\(zδ⋆\|\|yt\+1\)−DR\(yt\+1\|\|zt\)\]\.\\langle\\tilde\{g\}\_\{t\},y\_\{t\+1\}\-z\_\{\\delta\}^\{\\star\}\\rangle~\\leq~\\frac\{1\}\{\\eta\}\\left\[D\_\{R\}\(z\_\{\\delta\}^\{\\star\}\|\|z\_\{t\}\)\-D\_\{R\}\(z\_\{\\delta\}^\{\\star\}\|\|y\_\{t\+1\}\)\-D\_\{R\}\(y\_\{t\+1\}\|\|z\_\{t\}\)\\right\]\\,\.Using the above inequality, we obtain:

⟨g~t,zt−zδ⋆⟩\\displaystyle\\langle\\tilde\{g\}\_\{t\},z\_\{t\}\-z\_\{\\delta\}^\{\\star\}\\rangle=⟨g~t,zt−yt\+1⟩\+⟨g~t,yt\+1−zδ⋆⟩\\displaystyle=\\langle\\tilde\{g\}\_\{t\},z\_\{t\}\-y\_\{t\+1\}\\rangle\+\\langle\\tilde\{g\}\_\{t\},y\_\{t\+1\}\-z\_\{\\delta\}^\{\\star\}\\rangle≤⟨g~t,zt−yt\+1⟩\+1η\[DR\(zδ⋆\|\|zt\)−DR\(zδ⋆\|\|yt\+1\)−DR\(yt\+1\|\|zt\)\]\\displaystyle~\\leq~\\langle\\tilde\{g\}\_\{t\},z\_\{t\}\-y\_\{t\+1\}\\rangle\+\\frac\{1\}\{\\eta\}\\left\[D\_\{R\}\(z\_\{\\delta\}^\{\\star\}\|\|z\_\{t\}\)\-D\_\{R\}\(z\_\{\\delta\}^\{\\star\}\|\|y\_\{t\+1\}\)\-D\_\{R\}\(y\_\{t\+1\}\|\|z\_\{t\}\)\\right\]=⟨g~t,zt−yt\+1⟩−1ηDR\(yt\+1\|\|zt\)\+1η\[DR\(zδ⋆\|\|zt\)−DR\(zδ⋆\|\|yt\+1\)\]\.\\displaystyle=\\langle\\tilde\{g\}\_\{t\},z\_\{t\}\-y\_\{t\+1\}\\rangle\-\\frac\{1\}\{\\eta\}D\_\{R\}\(y\_\{t\+1\}\|\|z\_\{t\}\)\+\\frac\{1\}\{\\eta\}\\left\[D\_\{R\}\(z\_\{\\delta\}^\{\\star\}\|\|z\_\{t\}\)\-D\_\{R\}\(z\_\{\\delta\}^\{\\star\}\|\|y\_\{t\+1\}\)\\right\]\\,\.Using11\-strong convexity ofRRand Cauchy\-Schwarz inequality yields:

⟨g~t,zt−yt\+1⟩−1ηDR\(yt\+1\|\|zt\)≤∥g~t∥2⋅∥zt−yt\+1∥2−1η∥zt−yt\+1∥22\.\\langle\\tilde\{g\}\_\{t\},z\_\{t\}\-y\_\{t\+1\}\\rangle\-\\frac\{1\}\{\\eta\}D\_\{R\}\(y\_\{t\+1\}\|\|z\_\{t\}\)~\\leq~\\\|\\tilde\{g\}\_\{t\}\\\|\_\{2\}\\cdot\\\|z\_\{t\}\-y\_\{t\+1\}\\\|\_\{2\}\-\\frac\{1\}\{\\eta\}\\\|z\_\{t\}\-y\_\{t\+1\}\\\|\_\{2\}^\{2\}\\,\.Note now that for anya≥0,r≥0,a​r−12​η​r2≤η2​a2\.a~\\geq~0,r~\\geq~0,ar\-\\frac\{1\}\{2\\eta\}r^\{2\}~\\leq~\\frac\{\\eta\}\{2\}a^\{2\}\\,\.Hence witha=‖g~t‖2a=\\\|\\tilde\{g\}\_\{t\}\\\|\_\{2\}andr=‖zt−yt\+1‖r=\\\|z\_\{t\}\-y\_\{t\+1\}\\\|, we obtain:

⟨g~t,zt−yt\+1⟩−1ηDR\(yt\+1\|\|zt\)≤η2∥g~t∥22\.\\langle\\tilde\{g\}\_\{t\},z\_\{t\}\-y\_\{t\+1\}\\rangle\-\\frac\{1\}\{\\eta\}D\_\{R\}\(y\_\{t\+1\}\|\|z\_\{t\}\)~\\leq~\\frac\{\\eta\}\{2\}\\\|\\tilde\{g\}\_\{t\}\\\|\_\{2\}^\{2\}\\,\.We have proved that:

⟨g~t,zt−zδ⋆⟩≤DR\(zδ⋆\|\|zt\)−DR\(zδ⋆\|\|yt\+1\)η=η2​‖g~t‖22\.\\langle\\tilde\{g\}\_\{t\},z\_\{t\}\-z\_\{\\delta\}^\{\\star\}\\rangle~\\leq~\\frac\{D\_\{R\}\(z\_\{\\delta\}^\{\\star\}\|\|z\_\{t\}\)\-D\_\{R\}\(z\_\{\\delta\}^\{\\star\}\|\|y\_\{t\+1\}\)\}\{\\eta\}=\\frac\{\\eta\}\{2\}\\\|\\tilde\{g\}\_\{t\}\\\|\_\{2\}^\{2\}\\,\.Using the previous inequality, we have:

⟨g~t,zt−zδ⋆⟩≤DR\(zδ⋆\|\|zt\)−DR\(zδ⋆\|\|zt\+1\)η\+DR\(zδ⋆\|\|zt\+1\)−DR\(zδ⋆\|\|yt\+1\)η\+η2​‖g~2‖22\.\\langle\\tilde\{g\}\_\{t\},z\_\{t\}\-z\_\{\\delta\}^\{\\star\}\\rangle~\\leq~\\frac\{D\_\{R\}\(z\_\{\\delta\}^\{\\star\}\|\|z\_\{t\}\)\-D\_\{R\}\(z\_\{\\delta\}^\{\\star\}\|\|z\_\{t\+1\}\)\}\{\\eta\}\+\\frac\{D\_\{R\}\(z\_\{\\delta\}^\{\\star\}\|\|z\_\{t\+1\}\)\-D\_\{R\}\(z\_\{\\delta\}^\{\\star\}\|\|y\_\{t\+1\}\)\}\{\\eta\}\+\\frac\{\\eta\}\{2\}\\\|\\tilde\{g\}\_\{2\}\\\|\_\{2\}^\{2\}\\,\.ByGG\-Lipschitzness ofz↦DR\(zδ⋆\|\|z\)z\\mapsto D\_\{R\}\(z\_\{\\delta\}^\{\\star\}\|\|z\), we have:

DR\(zδ⋆\|\|zt\+1\)−DR\(zδ⋆\|\|yt\+1\)≤G∥zt\+1−yt\+1∥,D\_\{R\}\(z\_\{\\delta\}^\{\\star\}\|\|z\_\{t\+1\}\)\-D\_\{R\}\(z\_\{\\delta\}^\{\\star\}\|\|y\_\{t\+1\}\)~\\leq~G\\\|z\_\{t\+1\}\-y\_\{t\+1\}\\\|\\,,which implies that:

⟨g~t,zt−zδ⋆⟩≤DR\(zδ⋆\|\|zt\)−DR\(zδ⋆\|\|zt\+1\)η\+Gη​‖zt\+1−yt\+1‖\+η2​‖g~2‖22\.\\langle\\tilde\{g\}\_\{t\},z\_\{t\}\-z\_\{\\delta\}^\{\\star\}\\rangle~\\leq~\\frac\{D\_\{R\}\(z\_\{\\delta\}^\{\\star\}\|\|z\_\{t\}\)\-D\_\{R\}\(z\_\{\\delta\}^\{\\star\}\|\|z\_\{t\+1\}\)\}\{\\eta\}\+\\frac\{G\}\{\\eta\}\\\|z\_\{t\+1\}\-y\_\{t\+1\}\\\|\+\\frac\{\\eta\}\{2\}\\\|\\tilde\{g\}\_\{2\}\\\|\_\{2\}^\{2\}\\,\.Summing up the above inequality, the first term telescopes and we obtain:

∑t=1T⟨g~t,zt−zδ⋆⟩≤DR\(zδ⋆\|\|z1\)η\+η2​∑t=1T‖g~t‖22\+Gη​∑t=1T‖zt\+1−yt\+1‖2,\\sum\_\{t=1\}^\{T\}\\langle\\tilde\{g\}\_\{t\},z\_\{t\}\-z\_\{\\delta\}^\{\\star\}\\rangle~\\leq~\\frac\{D\_\{R\}\(z\_\{\\delta\}^\{\\star\}\|\|z\_\{1\}\)\}\{\\eta\}\+\\frac\{\\eta\}\{2\}\\sum\_\{t=1\}^\{T\}\\\|\\tilde\{g\}\_\{t\}\\\|\_\{2\}^\{2\}\+\\frac\{G\}\{\\eta\}\\sum\_\{t=1\}^\{T\}\\\|z\_\{t\+1\}\-y\_\{t\+1\}\\\|\_\{2\}\\,,which concludes the proof upon recalling the notationrt\+1=zt\+1−yt\+1\.r\_\{t\+1\}=z\_\{t\+1\}\-y\_\{t\+1\}\\,\.As for the second bound, it follows immediately from using‖g~t‖≤G~F\\\|\\tilde\{g\}\_\{t\}\\\|~\\leq~\\tilde\{G\}\_\{F\}and‖rt\+1‖≤Capp\\\|r\_\{t\+1\}\\\|~\\leq~C\_\{\\text\{app\}\}\. ∎

In the next lemma, we show how to control the error sequence\(rt\)\(r\_\{t\}\)in Lemma[14](https://arxiv.org/html/2605.26373#Thmlemma14)by providing a precise constantC^η\\hat\{C\}\_\{\\eta\}which quantifies the approximation error in the bandit setting, as a corollary of Proposition[2](https://arxiv.org/html/2605.26373#Thmproposition2)\.

###### Lemma 15\(Gradient estimators norm bounds\)\. Suppose thatℓt\\ell\_\{t\}is uniformly bounded byMM, i\.e\.\|ℓt​\(x\)\|≤M,∀x∈𝒳\|\\ell\_\{t\}\(x\)\|~\\leq~M,\\forall x\\in\{\\mathcal\{X\}\}, and the inverse Jacobian ofqqis bounded byGGat any pointx∈𝒳x\\in\{\\mathcal\{X\}\}\(as in Assumption[2](https://arxiv.org/html/2605.26373#Thmassumption2)\)\. Then for anyt≥1t~\\geq~1, we have:‖gt‖\\displaystyle\\\|g\_\{t\}\\\|≤G^F:=d​Mδ,‖g~t‖≤G~F:=d​M​Gδ\.\\displaystyle~\\leq~\\hat\{G\}\_\{F\}:=\\frac\{dM\}\{\\delta\}\\,,\\quad\\\|\\tilde\{g\}\_\{t\}\\\|~\\leq~\\tilde\{G\}\_\{F\}:=\\frac\{dMG\}\{\\delta\}\\,\.\(90\)

###### Proof\.

Recall from BOGD \(Algorithm[1](https://arxiv.org/html/2605.26373#alg1)\) thatgt=dδ​ℓt​\(x^t\)​ζtg\_\{t\}=\\frac\{d\}\{\\delta\}\\ell\_\{t\}\(\\hat\{x\}\_\{t\}\)\\zeta\_\{t\}\. Then, since‖ζt‖=1\\\|\\zeta\_\{t\}\\\|=1, we have:

‖gt‖=dδ​\|ℓt​\(x^t\)\|≤d​Mδ=G^F\.\\\|g\_\{t\}\\\|=\\frac\{d\}\{\\delta\}\|\\ell\_\{t\}\(\\hat\{x\}\_\{t\}\)\|~\\leq~\\frac\{dM\}\{\\delta\}=\\hat\{G\}\_\{F\}\\,\.Then, sinceg~t=Jq​\(xt\)−⊤​gt\\tilde\{g\}\_\{t\}=J\_\{q\}\(x\_\{t\}\)^\{\-\\top\}g\_\{t\}, we have:

‖g~t‖≤G​‖gt‖≤d​M​Gδ=G~F\.\\\|\\tilde\{g\}\_\{t\}\\\|~\\leq~G\\\|g\_\{t\}\\\|~\\leq~\\frac\{dMG\}\{\\delta\}=\\tilde\{G\}\_\{F\}\\,\.∎

###### Corollary 5\(Proposition[2](https://arxiv.org/html/2605.26373#Thmproposition2)in the bandit setting\)\. In Proposition[2](https://arxiv.org/html/2605.26373#Thmproposition2), letgtg\_\{t\}be the gradient estimator used in BOGD and recall thatg~t=Jq​\(xt\)−⊤​gt\\tilde\{g\}\_\{t\}=J\_\{q\}\(x\_\{t\}\)^\{\-\\top\}g\_\{t\}in \([5](https://arxiv.org/html/2605.26373#A1.E5)\)\-\([6](https://arxiv.org/html/2605.26373#A1.E6)\), then the approximation inequality \([7](https://arxiv.org/html/2605.26373#A1.E7)\) holds withG^F=d​M/δ\\hat\{G\}\_\{F\}=dM/\\deltaas defined in Lemma[15](https://arxiv.org/html/2605.26373#Thmlemma15), i\.e\., for anyt≥1t~\\geq~1, ifyt=q​\(xt\)y\_\{t\}=q\(x\_\{t\}\),‖yt\+1−q​\(xt\+1\)‖2≤C^η:=G5​\(5​G^F2\+G^F3​η\)​η2\.\\\|y\_\{t\+1\}\-q\(x\_\{t\+1\}\)\\\|\_\{2\}~\\leq~\\hat\{C\}\_\{\\eta\}:=G^\{5\}\(5\\hat\{G\}\_\{F\}^\{2\}\+\\hat\{G\}\_\{F\}^\{3\}\\eta\)\\eta^\{2\}\\,\.

End of Proof of Theorem[3](https://arxiv.org/html/2605.26373#Thmtheorem3)\.We are ready to conclude the proof of the theorem using the results established so far\. First, it follows from Lemma[10](https://arxiv.org/html/2605.26373#Thmlemma10)that:

𝔼​\[RT\]≤RT\(1\)\+RT\(2\)\+RT\(3\)\+RT\(4\)\.\\mathbb\{E\}\[R\_\{T\}\]~\\leq~R\_\{T\}^\{\(1\)\}\+R\_\{T\}^\{\(2\)\}\+R\_\{T\}^\{\(3\)\}\+R\_\{T\}^\{\(4\)\}\\,\.\(91\)We have proved a bound for each one of the error terms in the above bound:

1. \(i\)By Lemma[11](https://arxiv.org/html/2605.26373#Thmlemma11), we haveRT\(2\)≤δ​GF​D​T\.R\_\{T\}^\{\(2\)\}~\\leq~\\delta G\_\{F\}DT\\,\.
2. \(ii\)By Lemma[12](https://arxiv.org/html/2605.26373#Thmlemma12), we get:\|RT\(3\)\|≤δ​G2​H​D​T\.\|R\_\{T\}^\{\(3\)\}\|~\\leq~\\delta G^\{2\}HDT\\,\.
3. \(iii\)By Lemma[13](https://arxiv.org/html/2605.26373#Thmlemma13), we obtain:\|RT\(4\)\|≤δ​GF​T\.\|R\_\{T\}^\{\(4\)\}\|~\\leq~\\delta G\_\{F\}T\\,\.
4. \(iv\)Lemma[14](https://arxiv.org/html/2605.26373#Thmlemma14)\-\([86](https://arxiv.org/html/2605.26373#A3.E86)\) together with Corollary[5](https://arxiv.org/html/2605.26373#Thmtheorem5)yield: RT\(1\)≤D1η\+η​G~F2​T2\+C^η​G​Tη≤D1η\+\(6​G^F2\+G^F3​η\)​G6​η​T,R\_\{T\}^\{\(1\)\}~\\leq~\\frac\{D\_\{1\}\}\{\\eta\}\+\\frac\{\\eta\\tilde\{G\}\_\{F\}^\{2\}T\}\{2\}\+\\frac\{\\hat\{C\}\_\{\\eta\}GT\}\{\\eta\}~\\leq~\\frac\{D\_\{1\}\}\{\\eta\}\+\(6\\hat\{G\}\_\{F\}^\{2\}\+\\hat\{G\}\_\{F\}^\{3\}\\eta\)G^\{6\}\\eta T\\,,where the last inequality follows from usingC^η=G5​\(5​G^F2\+G^F3​η\)​η2\\hat\{C\}\_\{\\eta\}=G^\{5\}\(5\\hat\{G\}\_\{F\}^\{2\}\+\\hat\{G\}\_\{F\}^\{3\}\\eta\)\\eta^\{2\}\(Corollary[5](https://arxiv.org/html/2605.26373#Thmtheorem5)\) and upper bounding the resulting constant under the assumptionsG\>1G\>1\(which is without loss of generality\), similarly to the deterministic setting\. Note here thatG^F=d​Mδ\\hat\{G\}\_\{F\}=\\frac\{dM\}\{\\delta\}as shown in Lemma[15](https://arxiv.org/html/2605.26373#Thmlemma15)\.

Combining all the above bounds in \([91](https://arxiv.org/html/2605.26373#A3.E91)\), we obtain:

𝔼​\[RT\]≤K1​δ​T\+D1η\+K2​η​Tδ2\+K3​η2​Tδ3,\\mathbb\{E\}\[R\_\{T\}\]~\\leq~K\_\{1\}\\delta T\+\\frac\{D\_\{1\}\}\{\\eta\}\+K\_\{2\}\\frac\{\\eta T\}\{\\delta^\{2\}\}\+K\_\{3\}\\frac\{\\eta^\{2\}T\}\{\\delta^\{3\}\}\\,,\(92\)whereK1:=GF​D\+G2​H​D\+GF,K2:=6​d2​G6​M2,K3=d3​G6​M3\.K\_\{1\}:=G\_\{F\}D\+G^\{2\}HD\+G\_\{F\},K\_\{2\}:=6d^\{2\}G^\{6\}M^\{2\},K\_\{3\}=d^\{3\}G^\{6\}M^\{3\}\\,\.

We now derive a choice of\(η,δ\)\(\\eta,\\delta\)which gives an optimal dependence on the time horizonTT\. A lower bound on the upper regret bound \(inη,δ\\eta,\\delta\) is given by dropping the last positive error termK3​η2​Tδ3\\frac\{K\_\{3\}\\eta^\{2\}T\}\{\\delta^\{3\}\}\(which as we will see will contribute to a faster regret rate and will hence not be the leading term\)\. Define then for any\(η,δ\)\(\\eta,\\delta\),

f​\(η,δ\):=K1​δ​T\+D1η\+K2​η​Tδ2,f\(\\eta,\\delta\):=K\_\{1\}\\delta T\+\\frac\{D\_\{1\}\}\{\\eta\}\+K\_\{2\}\\frac\{\\eta T\}\{\\delta^\{2\}\}\\,,which consists of the three first terms in the upper bound \([92](https://arxiv.org/html/2605.26373#A3.E92)\)\. We now optimize this function w\.r\.t\(η,δ\)\(\\eta,\\delta\)\. By the first order optimality conditions, the optimal pair\(η⋆,δ⋆\)\(\\eta\_\{\\star\},\\delta\_\{\\star\}\)minimizingf​\(η,δ\)f\(\\eta,\\delta\)satisfies:

∂f∂η​\(η⋆,δ⋆\)=−D1η2\+K2​Tδ⋆2=0,∂f∂δ​\(η⋆,δ⋆\)=K1​T−2​K2​η⋆​Tδ⋆3\.\\frac\{\\partial f\}\{\\partial\\eta\}\(\\eta\_\{\\star\},\\delta\_\{\\star\}\)=\-\\frac\{D\_\{1\}\}\{\\eta^\{2\}\}\+\\frac\{K\_\{2\}T\}\{\\delta\_\{\\star\}^\{2\}\}=0\\,,\\quad\\frac\{\\partial f\}\{\\partial\\delta\}\(\\eta\_\{\\star\},\\delta\_\{\\star\}\)=K\_\{1\}T\-\\frac\{2K\_\{2\}\\eta\_\{\\star\}T\}\{\\delta\_\{\\star\}^\{3\}\}\\,\.Rewriting these conditions yields:

η⋆=D1​δ⋆2K2​T,η⋆=K1​δ⋆32​K2\.\\eta\_\{\\star\}=\\sqrt\{\\frac\{D\_\{1\}\\delta\_\{\\star\}^\{2\}\}\{K\_\{2\}T\}\}\\,,\\quad\\eta\_\{\\star\}=\\frac\{K\_\{1\}\\delta\_\{\\star\}^\{3\}\}\{2K\_\{2\}\}\\,\.Solving forδ⋆\\delta\_\{\\star\}as a function ofTT\(eliminatingη\\eta\) and then plugging backδ⋆\\delta\_\{\\star\}in the first expression ofη⋆\\eta\_\{\\star\}yields:

δ⋆=\(4​K2​D1K12​T\)1/4,η⋆=\(4​D13K12​K2​T3\)1/4\.\\delta\_\{\\star\}=\\left\(\\frac\{4K\_\{2\}D\_\{1\}\}\{K\_\{1\}^\{2\}T\}\\right\)^\{1/4\}\\,,\\quad\\eta\_\{\\star\}=\\left\(\\frac\{4D\_\{1\}^\{3\}\}\{K\_\{1\}^\{2\}K\_\{2\}T^\{3\}\}\\right\)^\{1/4\}\\,\.Now, plugging back these values ofη⋆\\eta\_\{\\star\}andδ⋆\\delta\_\{\\star\}in \([92](https://arxiv.org/html/2605.26373#A3.E92)\), we obtain:

𝔼​\[RT\]≤\(K12​K2​D1\)1/4​T3/4\+\(K12​D13​K344​K25\)​T1/4\.\\mathbb\{E\}\[R\_\{T\}\]~\\leq~\(K\_\{1\}^\{2\}K\_\{2\}D\_\{1\}\)^\{1/4\}T^\{3/4\}\+\\left\(\\frac\{K\_\{1\}^\{2\}D\_\{1\}^\{3\}K\_\{3\}^\{4\}\}\{4K\_\{2\}^\{5\}\}\\right\)T^\{1/4\}\\,\.Recalling thatK1=GF​D\+G2​H​D\+GF,K2=6​d2​G6​M2,K3=d3​G6​M3K\_\{1\}=G\_\{F\}D\+G^\{2\}HD\+G\_\{F\},K\_\{2\}=6d^\{2\}G^\{6\}M^\{2\},K\_\{3\}=d^\{3\}G^\{6\}M^\{3\}, we bound each one of the terms in the bound above to obtain:

𝔼​\[RT\]\\displaystyle\\mathbb\{E\}\[R\_\{T\}\]≤3​d​M​G3​\(GF​D\+G2​H​D\+GF\)​D11/2​T3/4\+d​M​G3​\(GF​D\+G2​H​D\+GF\)​D13/2​T1/4\\displaystyle~\\leq~\\sqrt\{3dMG^\{3\}\(G\_\{F\}D\+G^\{2\}HD\+G\_\{F\}\)D\_\{1\}^\{1/2\}\}T^\{3/4\}\+\\sqrt\{dMG^\{3\}\(G\_\{F\}D\+G^\{2\}HD\+G\_\{F\}\)D\_\{1\}^\{3/2\}\}T^\{1/4\}≤d​M​G3​\(GF​D\+G2​H​D\+GF\)​\(3​D11/2\+D13/2\)​T3/4,\\displaystyle~\\leq~\\sqrt\{dMG^\{3\}\(G\_\{F\}D\+G^\{2\}HD\+G\_\{F\}\)\(3D\_\{1\}^\{1/2\}\+D\_\{1\}^\{3/2\}\)\}T^\{3/4\}\\,,\(93\)which concludes the proof of Theorem[3](https://arxiv.org/html/2605.26373#Thmtheorem3)\.

Similar Articles