Learning When to Act: Communication-Efficient Reinforcement Learning via Run-Time Assurance
Summary
This paper presents a framework (CARE) that jointly learns control inputs and communication-efficient timing decisions under a pointwise Lyapunov safety shield, achieving higher inter-sample intervals than classical methods on inverted pendulum, cart-pole, and planar quadrotor systems.
View Cached Full Text
Cached at: 05/14/26, 06:16 AM
# Communication-Efficient Reinforcement Learning via Run-Time Assurance
Source: [https://arxiv.org/html/2605.12561](https://arxiv.org/html/2605.12561)
Adam Haroon1,3Erick J\. Rodríguez\-Seda2Cody Fleming1Tristan Schuler3 1Department of Mechanical Engineering, Iowa State University, Ames, IA, USA 2Department of Weapons, Robotics, and Control Engineering, United States Naval Academy, Annapolis, MD, USA 3Navy Center for Applied Research in Artificial Intelligence \(NCARAI\), U\.S\. Naval Research Laboratory, Washington, D\.C\., USA
###### Abstract
Safe reinforcement learning \(RL\) typically asks*what*an agent should do\. We ask*when*it needs to act, and show that a single policy can jointly learn control inputs and communication\-efficient timing decisions under a pointwise Lyapunov safety shield\. We scope the framework to stabilization around a known equilibrium, where CARE\-based LQR backups, Lyapunov certificates, and classical Lyapunov\-STC are well defined, enabling a clean comparison against the analytical baseline\. A run\-time assurance \(RTA\) layer overrides the policy pointwise via a one\-step\-ahead Lyapunov prediction and a precomputed LQR backup, providing a strictly stronger guarantee than constrained MDP methods that enforce safety only in expectation\. On an inverted pendulum, cart–pole, and planar quadrotor, the learned policy achieves1\.91×1\.91\\times,1\.45×1\.45\\times, and3\.51×3\.51\\timeshigher mean inter\-sample interval \(MSI\) than a classical Lyapunov\-triggered baseline; a fixed LQR controller at the same average rate is unstable on all three plants, showing that adaptive timing, not a lower average rate, is what makes sparsity safe\. A CARE\-derived Lyapunov reward transfers across environments without redesign, with a single weightwcw\_\{c\}controlling the stability–communication tradeoff; ablations confirm the RTA shield is essential, with its removal reducing MSI by1\.271\.27–1\.84×1\.84\\timesand degrading state norms\. A preference\-conditioned extension recovers the full tradeoff frontier from a single model at211\\tfrac\{2\}\{11\}of training compute, and SAC experiments confirm the results are algorithm\-agnostic across discrete and continuous domains\. A 12\-state 3D quadrotor case study extends the framework to higher\-dimensional systems where classical STC design is analytically intractable: a SAC agent reaches MSI=0\.302=0\.302s \(94%94\\%ofτmax\\tau\_\{\\max\}\) at0%0\\%RTA, while classical Lyapunov\-STC remains pinned atτmin\\tau\_\{\\min\}and a fixed\-rate LQR controller at the same average interval crashes within two control updates\. Robustness to±30%\\pm 30\\%plant\-mass variation and additive disturbances confirms graceful degradation, with the RTA absorbing what the learned policy cannot\.
## 1Introduction
Safe reinforcement learning \(RL\) has made substantial progress on*what*an agent should do, but has largely ignored the question of*when*it needs to act\. Fixed\-rate control executes at every timestep regardless of whether the state has changed meaningfully; in safety\-critical cyber\-physical systems, this is expensive, consuming sensing, computation, and bandwidth on every update\. Lifting this assumption without sacrificing safety is the central problem we address\.
The control theory literature studies this as self\-triggered control \(STC\) and event\-triggered control \(ETC\)\[[16](https://arxiv.org/html/2605.12561#bib.bib1),[20](https://arxiv.org/html/2605.12561#bib.bib2),[6](https://arxiv.org/html/2605.12561#bib.bib3),[13](https://arxiv.org/html/2605.12561#bib.bib4),[30](https://arxiv.org/html/2605.12561#bib.bib5)\]: classical analytical methods determine the next update time from a Lyapunov or model\-based triggering rule, providing strong stability guarantees but becoming difficult to apply to nonlinear, underactuated, or high\-dimensional dynamics where the linearized one\-step prediction leaves significant communication savings on the table\. Learning\-based approaches relax this conservatism by approaching the admissibility boundary empirically\[[7](https://arxiv.org/html/2605.12561#bib.bib6),[28](https://arxiv.org/html/2605.12561#bib.bib7),[11](https://arxiv.org/html/2605.12561#bib.bib8),[29](https://arxiv.org/html/2605.12561#bib.bib9),[3](https://arxiv.org/html/2605.12561#bib.bib11),[24](https://arxiv.org/html/2605.12561#bib.bib10),[27](https://arxiv.org/html/2605.12561#bib.bib12)\], but existing formulations either omit formal safety guarantees on the triggering decision\[[7](https://arxiv.org/html/2605.12561#bib.bib6),[28](https://arxiv.org/html/2605.12561#bib.bib7)\], rely on model\-based learning of the dynamics\[[11](https://arxiv.org/html/2605.12561#bib.bib8)\], or address an interaction\-cost continuous\-time setting without a hard timing\-safety mechanism\[[27](https://arxiv.org/html/2605.12561#bib.bib12)\]\. Our framing differs on multiple axes: we operate in discrete\-time STC whereτk\\tau\_\{k\}is selected at each decision instant, with both discrete\-action \(DQN\) and continuous\-action \(SAC\) agents; we are model\-free with a precomputed Linear\-Quadratic Regulator \(LQR\) backup controller; and we provide*pointwise*Lyapunov\-decreasing safety via a hard RTA override, a formal guarantee that expectation\-level constraint methods such as Lagrangian\-relaxation\[[5](https://arxiv.org/html/2605.12561#bib.bib20),[2](https://arxiv.org/html/2605.12561#bib.bib21)\]cannot provide by construction \(Section[5\.5](https://arxiv.org/html/2605.12561#S5.SS5)\)\.
\\begin\{overpic\}\[width=255\.8342pt,grid=false\]\{figures/RTA\_Blank\.png\} \\put\(16\.0,20\.0\)\{\\makebox\(0\.0,0\.0\)\[\]\{Backup\}\} \\put\(16\.0,15\.0\)\{\\makebox\(0\.0,0\.0\)\[\]\{Controller\}\} \\put\(51\.0,17\.0\)\{\\makebox\(0\.0,0\.0\)\[\]\{RL Agent\}\} \\put\(85\.0,56\.0\)\{\\makebox\(0\.0,0\.0\)\[\]\{Environment\}\} \\put\(53\.0,56\.0\)\{\\makebox\(0\.0,0\.0\)\[\]\{ZOH\}\} \\put\(16\.0,58\.0\)\{\\makebox\(0\.0,0\.0\)\[\]\{RTA\}\} \\put\(86\.0,40\.0\)\{\\footnotesize$\\bm\{r\}\(t\),\\bm\{o\}\(t\)$\} \\put\(61\.0,57\.5\)\{\\footnotesize$\\bm\{u\}\(t\)$\} \\put\(72\.0,2\.5\)\{\\footnotesize$r\_\{k\},\\bm\{o\}\_\{k\}$\} \\put\(36\.0,2\.0\)\{\\footnotesize$\\bm\{x\}\_\{k\}$\} \\put\(36\.0,36\.0\)\{\\footnotesize$\\bm\{u\}\_\{k\},\\tau\_\{k\}$\} \\put\(18\.0,30\.0\)\{\\footnotesize$\\bm\{u\}\_\{k\},\\tau\_\{\\min\}$\} \\put\(37\.0,57\.0\)\{\\footnotesize$\\bm\{u\}\_\{k\}$\} \\put\(68\.0,35\.0\)\{\\footnotesize$\\tau\_\{k\}$\} \\put\(11\.0,32\.0\)\{\\rotatebox\{90\.0\}\{\\tiny$\|\\hat\{q\}\_\{k\+1\}\|\>\\theta\_\{\\mathrm\{RTA\}\}$\}\} \\end\{overpic\}Figure 1:Proposed self\-triggered control framework\. The RL agent outputs a joint action\(𝒖k,τk\)\(\\bm\{u\}\_\{k\},\\tau\_\{k\}\)\. The RTA evaluates the one\-step\-ahead safety predicate\|q^k\+1\|\>θRTA\|\\hat\{q\}\_\{k\+1\}\|\>\\theta\_\{\\mathrm\{RTA\}\}on the linearized prediction of the next state and, if violated, substitutes the LQR backup atτmin\\tau\_\{\\min\}\. The selected input is held by a zero\-order hold \(ZOH\) until the next sampling instanttk\+τkt\_\{k\}\+\\tau\_\{k\}\.We propose a unified RL framework \(Fig\.[1](https://arxiv.org/html/2605.12561#S1.F1)\) that jointly selects the control input and the next inter\-sample interval, supervised by an RTA layer acting as a safety shield\[[17](https://arxiv.org/html/2605.12561#bib.bib15),[19](https://arxiv.org/html/2605.12561#bib.bib17),[25](https://arxiv.org/html/2605.12561#bib.bib19)\]\. The RTA uses a precomputed LQR backup and a one\-step\-ahead safety prediction to override unsafe actions, with each backup intervention provably unsaturated and Lyapunov\-decreasing\[[4](https://arxiv.org/html/2605.12561#bib.bib13),[21](https://arxiv.org/html/2605.12561#bib.bib14),[10](https://arxiv.org/html/2605.12561#bib.bib18)\]\. Episodes are time\-bounded \(not step\-bounded\), so a larger MSI directly reduces the number of agent decisions per episode of fixed duration: the operational quantity that communication efficiency targets\. The RTA differs fundamentally from constrained MDP methods\[[5](https://arxiv.org/html/2605.12561#bib.bib20),[2](https://arxiv.org/html/2605.12561#bib.bib21),[14](https://arxiv.org/html/2605.12561#bib.bib22),[9](https://arxiv.org/html/2605.12561#bib.bib23),[12](https://arxiv.org/html/2605.12561#bib.bib24)\], which enforce safety only in expectation; we demonstrate empirically that a Lagrangian\-DQN baseline using the identical constraint predicate achieves2121–52%52\\%lower MSI than RL\-STC while accumulating up to35%35\\%hard safety violations at the best checkpoint \(Section[5\.5](https://arxiv.org/html/2605.12561#S5.SS5)\)\. The framework is scoped to stabilization around a known equilibrium where CARE, Lyapunov analysis, and classical STC are well defined, enabling clean comparison; we evaluate on three lower\-dimensional plants \(inverted pendulum, CartPole, planar quadrotor\) plus a 12\-state 3D quadrotor case study \(Section[6](https://arxiv.org/html/2605.12561#S6)\) where analytical STC design is intractable\.
#### Main contributions\.
\(1\)A joint RL formulation that simultaneously learns control inputs and inter\-sample intervals under a hard pointwise safety certificate \(Proposition[1](https://arxiv.org/html/2605.12561#Thmproposition1)\); unlike prior shielded RL that enforces safety at fixed time steps, our shield acts on the timing decisionτk\\tau\_\{k\}itself\. A single weightwcw\_\{c\}traces the full stability/communication tradeoff frontier\.\(2\)A CARE\-derived Lyapunov reward that transfers across SISO and MIMO environments without redesign\.\(3\)Systematic empirical validation on three lower\-dimensional plants:1\.451\.45–3\.51×3\.51\\timesMSI gains over Classical Lyapunov\-STC, ablations isolating the RTA’s role, a Lagrangian\-DQN comparison, a preference\-conditioned extension at211\\tfrac\{2\}\{11\}compute, SAC confirming algorithm\-agnosticism, and robustness under±30%\\pm 30\\%mass variation and additive disturbances\.\(4\)A 12\-state 3D quadrotor case study showing the framework scales to higher\-dimensional MIMO systems where analytical STC design is intractable and where Proposition[1](https://arxiv.org/html/2605.12561#Thmproposition1)’s formal certificate does not extend at the chosenτmin\\tau\_\{\\min\}, testing the framework’s empirical reach beyond where the certificate applies \(MSI=0\.302s=0\.302\\,\\mathrm\{s\},94%94\\%ofτmax\\tau\_\{\\max\},0%0\\%RTA; Classical STC pinned atτmin\\tau\_\{\\min\}; B2 unstable within two control updates\)\.
## 2Problem Formulation
### 2\.1Evaluation Scope
The framework targets stabilization around a known equilibrium\. Three structural conditions are required for the safety certificates developed below: \(i\) a valid local linearization, \(ii\) a positive\-definite quadratic Lyapunov function, and \(iii\) accuracy of the one\-step\-ahead prediction in \([2](https://arxiv.org/html/2605.12561#S2.E2)\)\. Stabilization around a known equilibrium guarantees all three: the linearization at the equilibrium is well posed, the CARE solution yieldsV\(x\)=x⊤PxV\(x\)=x^\{\\top\}Pxthat is positive\-definite about that point, and the prediction is accurate within its neighborhood\. Propositions[1](https://arxiv.org/html/2605.12561#Thmproposition1)and[2](https://arxiv.org/html/2605.12561#Thmproposition2)therefore hold under this scope\.
### 2\.2Self\-Triggered Control
Consider a continuous\-time nonlinear plant𝒙˙\(t\)=f\(𝒙\(t\),𝒖\(t\)\)\\dot\{\\bm\{x\}\}\(t\)=f\(\\bm\{x\}\(t\),\\bm\{u\}\(t\)\), where𝒙∈ℝn\\bm\{x\}\\in\\mathbb\{R\}^\{n\}is the state and𝒖∈ℝm\\bm\{u\}\\in\\mathbb\{R\}^\{m\}is a piecewise\-constant control input\. In self\-triggered control the controller executes at sampling instants\{t0,t1,…\}\\\{t\_\{0\},t\_\{1\},\\ldots\\\}determined online\. At each instanttkt\_\{k\}the controller observes𝒙k≜𝒙\(tk\)\\bm\{x\}\_\{k\}\\triangleq\\bm\{x\}\(t\_\{k\}\), selects𝒖k\\bm\{u\}\_\{k\}held constant over\[tk,tk\+1\)\[t\_\{k\},t\_\{k\+1\}\), and decides
tk\+1=tk\+τk,τk∈𝒯=\{τmin,2τmin,…,Nτmin\},t\_\{k\+1\}=t\_\{k\}\+\\tau\_\{k\},\\qquad\\tau\_\{k\}\\in\\mathcal\{T\}=\\\{\\tau\_\{\\min\},2\\tau\_\{\\min\},\\ldots,N\\tau\_\{\\min\}\\\},\(1\)whereNNis the cardinality of𝒯\\mathcal\{T\}; for consistency we chooseN=8N=8across all experiments\. The mean sampling interval \(MSI\) is the causalnn\-point moving averageMSIk=\[\(n−1\)MSIk−1\+τk\]/n\\mathrm\{MSI\}\_\{k\}=\[\(n\{\-\}1\)\\mathrm\{MSI\}\_\{k\-1\}\+\\tau\_\{k\}\]/n, initialized atMSI0=τmin\\mathrm\{MSI\}\_\{0\}=\\tau\_\{\\min\}; in what follows we taken=5n=5\. A larger MSI corresponds to sparser communication; the objective is to maximize MSI subject to closed\-loop stability\. Episodes are time\-bounded by a fixed simulation horizonTmaxT\_\{\\max\}rather than a fixed step count, so a larger MSI directly reduces the number of agent decisions per episode\. This reflects what communication efficiency means in practice: fewer control updates per unit of operating time, and therefore lower demand on sensing, computation, and network bandwidth\.
### 2\.3Run\-Time Assurance
Run\-time assurance augments the RL policy with a precomputed LQR backup \(Section[3\.1](https://arxiv.org/html/2605.12561#S3.SS1)\)\. At each step the linearized one\-step\-ahead prediction of the safety\-critical scalarqk=𝒄⊤𝒙kq\_\{k\}=\\bm\{c\}^\{\\top\}\\bm\{x\}\_\{k\}is
q^k\+1=qk\+τkq˙k\+12τk2q¨lin\(𝒙k,𝒖k\)\.\\hat\{q\}\_\{k\+1\}=q\_\{k\}\+\\tau\_\{k\}\\dot\{q\}\_\{k\}\+\\tfrac\{1\}\{2\}\\tau\_\{k\}^\{2\}\\,\\ddot\{q\}\_\{\\mathrm\{lin\}\}\(\\bm\{x\}\_\{k\},\\bm\{u\}\_\{k\}\)\.\(2\)If\|q^k\+1\|\>θRTA\|\\hat\{q\}\_\{k\+1\}\|\>\\theta\_\{\\mathrm\{RTA\}\}\(or a position bound is exceeded\), the RL action is overridden:
𝒖k←clipcw\(−K𝒙k,−𝒖max,\+𝒖max\),τk←τmin\.\\bm\{u\}\_\{k\}\\leftarrow\\mathrm\{clip\}\_\{\\mathrm\{cw\}\}\(\-K\\bm\{x\}\_\{k\},\-\\bm\{u\}\_\{\\max\},\+\\bm\{u\}\_\{\\max\}\),\\quad\\tau\_\{k\}\\leftarrow\\tau\_\{\\min\}\.\(3\)The thresholdθRTA\\theta\_\{\\mathrm\{RTA\}\}is set strictly below the LQR saturation angleθsat=umax/\|Kθ\|\\theta\_\{\\mathrm\{sat\}\}=u\_\{\\max\}/\|K\_\{\\theta\}\|so that the backup retains full authority\. The role of RL is not to guarantee safety, but to maximize performance within the safe set defined by the RTA layer, discovering non\-conservative inter\-sample intervals that would be difficult to obtain analytically\.
###### Proposition 1\(ZOH Lyapunov Decrease for the Linearized Backup\)\.
Let
M\(τ\)≜eAτ−\(∫0τeAs𝑑s\)BKM\(\\tau\)\\triangleq e^\{A\\tau\}\-\\biggl\(\\\!\\int\_\{0\}^\{\\tau\}\\\!e^\{As\}\\,ds\\biggr\)BK\(4\)denote the ZOH\-discretized closed\-loop transition matrix that arises when the backup command𝐮k=−K𝐱k\\bm\{u\}\_\{k\}=\-K\\bm\{x\}\_\{k\}is held constant on\[tk,tk\+τ\)\[t\_\{k\},\\,t\_\{k\}\+\\tau\), and suppose this command is component\-wise unsaturated at𝐱k\\bm\{x\}\_\{k\}\(\|\[−K𝐱k\]i\|<umax,i\|\[\-K\\bm\{x\}\_\{k\}\]\_\{i\}\|<u\_\{\\max,i\}for allii\)\. Then𝐱k\+1=M\(τmin\)𝐱k\\bm\{x\}\_\{k\+1\}=M\(\\tau\_\{\\min\}\)\\,\\bm\{x\}\_\{k\}on the linearized plant, andV\(𝐱k\+1\)≤V\(𝐱k\)V\(\\bm\{x\}\_\{k\+1\}\)\\leq V\(\\bm\{x\}\_\{k\}\)holds for all𝐱k\\bm\{x\}\_\{k\}if and only if
Mdisc≜P−M\(τmin\)⊤PM\(τmin\)⪰0\.M\_\{\\mathrm\{disc\}\}\\triangleq P\-M\(\\tau\_\{\\min\}\)^\{\\top\}P\\,M\(\\tau\_\{\\min\}\)\\;\\succeq\\;0\.Numerical verification \(Table[4](https://arxiv.org/html/2605.12561#A1.T4)\) confirmsMdisc≻0M\_\{\\mathrm\{disc\}\}\\succ 0for the Pendulum, CartPole, and planar Quadrotor environments\. The Quadrotor3D plant fails this verification at its chosenτmin\\tau\_\{\\min\}, retained deliberately for the case study \(Section[6](https://arxiv.org/html/2605.12561#S6)\) which probes empirical behavior in the regime where the certificate does not extend\.
*Proof and an extended scope discussion are in Appendix[A](https://arxiv.org/html/2605.12561#A1)\.*
## 3Method
### 3\.1CARE\-Based Lyapunov Function
For each environment we linearize the dynamics around the target equilibrium to obtain𝒙˙=A𝒙\+B𝒖\\dot\{\\bm\{x\}\}=A\\bm\{x\}\+B\\bm\{u\}\. The Lyapunov functionV\(𝒙\)=𝒙⊤P𝒙V\(\\bm\{x\}\)=\\bm\{x\}^\{\\top\}P\\bm\{x\}is obtained by solving the Continuous\-time Algebraic Riccati Equation \(CARE\)
A⊤P\+PA−PBR−1B⊤P\+Q=0,A^\{\\top\}P\+PA\-PBR^\{\-1\}B^\{\\top\}P\+Q=0,\(5\)withQ≻0Q\\succ 0,R≻0R\\succ 0\. The optimal LQR gain isK=R−1B⊤PK=R^\{\-1\}B^\{\\top\}Pand the closed\-loop matrixAcl=A−BKA\_\{\\mathrm\{cl\}\}=A\-BKis Hurwitz\. By the CARE identity,V˙\(𝒙\)≤−λV\(𝒙\)\\dot\{V\}\(\\bm\{x\}\)\\leq\-\\lambda\\,V\(\\bm\{x\}\)along the linear closed\-loop trajectory, whereλ=λmin\(Q\+K⊤RK,P\)\\lambda=\\lambda\_\{\\min\}\(Q\+K^\{\\top\}RK,\\,P\)is the minimum generalized eigenvalue\[[8](https://arxiv.org/html/2605.12561#bib.bib30)\], yielding a tighter bound than the standard estimate\. The Lyapunov value is normalized byVscale=tr\(P\)/nV\_\{\\mathrm\{scale\}\}=\\operatorname\{tr\}\(P\)/n\.
###### Proposition 2\(Admissible Inter\-Sample Interval\)\.
For the linear closed\-loop system withP≻0P\\succ 0solving \([5](https://arxiv.org/html/2605.12561#S3.E5)\), for any𝐱k≠𝟎\\bm\{x\}\_\{k\}\\neq\\bm\{0\}there existsτ∗\(𝐱k\)\>0\\tau^\{\*\}\(\\bm\{x\}\_\{k\}\)\>0such thatV\(𝐱k\+1\)<V\(𝐱k\)V\(\\bm\{x\}\_\{k\+1\}\)<V\(\\bm\{x\}\_\{k\}\)for allτk∈\(0,τ∗\(𝐱k\)\)\\tau\_\{k\}\\in\(0,\\tau^\{\*\}\(\\bm\{x\}\_\{k\}\)\), where𝐱k\+1=M\(τk\)𝐱k\\bm\{x\}\_\{k\+1\}=M\(\\tau\_\{k\}\)\\,\\bm\{x\}\_\{k\}is the ZOH solution under𝐮k=−K𝐱k\\bm\{u\}\_\{k\}=\-K\\bm\{x\}\_\{k\}withM\(⋅\)M\(\\cdot\)from \([4](https://arxiv.org/html/2605.12561#S2.E4)\)\. On any compact annular regionc≤V\(𝐱k\)≤c′c\\leq V\(\\bm\{x\}\_\{k\}\)\\leq c^\{\\prime\}, a uniform constantτ\+\>0\\tau^\{\+\}\>0independent of𝐱k\\bm\{x\}\_\{k\}exists with this property \(Corollary[A\.3](https://arxiv.org/html/2605.12561#A1.SS3)\);τmin\\tau\_\{\\min\}is verified numerically to satisfyτmin≤τ\+\\tau\_\{\\min\}\\leq\\tau^\{\+\}viaMdisc≻0M\_\{\\mathrm\{disc\}\}\\succ 0\(Table[4](https://arxiv.org/html/2605.12561#A1.T4)\)\.
*Proofs of Proposition[2](https://arxiv.org/html/2605.12561#Thmproposition2)and Corollary[A\.3](https://arxiv.org/html/2605.12561#A1.SS3)are in Appendix[A](https://arxiv.org/html/2605.12561#A1)\. Corollary[A\.4](https://arxiv.org/html/2605.12561#A1.SS4)\(Appendix[A](https://arxiv.org/html/2605.12561#A1)\) extends these results to the nonlinear plant with an explicit stability radiusr∗r^\{\*\}\.*
### 3\.2Action Space, Observations, and Reward
The agent selects a discrete action from a set of\(τ,𝒖\)\(\\tau,\\bm\{u\}\)tuples\. For SISO environments𝒜=𝒯×𝒰\\mathcal\{A\}=\\mathcal\{T\}\\times\\mathcal\{U\}; for the MIMO quadrotor𝒜=𝒯×𝒰δF×𝒰M\\mathcal\{A\}=\\mathcal\{T\}\\times\\mathcal\{U\}\_\{\\delta F\}\\times\\mathcal\{U\}\_\{M\}\. The agent observes𝒐k=\[𝒙k⊤,MSIk,bk\]⊤\\bm\{o\}\_\{k\}=\[\\bm\{x\}\_\{k\}^\{\\top\},\\mathrm\{MSI\}\_\{k\},b\_\{k\}\]^\{\\top\}, wherebk∈\{0,1\}b\_\{k\}\\in\\\{0,1\\\}flags whether the RTA was active on the previous step; includingMSIk\\mathrm\{MSI\}\_\{k\}enables credit assignment for the communication reward\.
The per\-step reward is
rk=rstab\+1−V\(𝒙k\+1\)Vscale⏟stability\+wc\(MSIk−τminτmax−τmin\)2⏟communication\+100rsafe⏟safety,r\_\{k\}=\\underbrace\{r\_\{\\mathrm\{stab\}\}\+1\-\\frac\{V\(\\bm\{x\}\_\{k\+1\}\)\}\{V\_\{\\mathrm\{scale\}\}\}\}\_\{\\text\{stability\}\}\+\\underbrace\{w\_\{c\}\\\!\\left\(\\frac\{\\mathrm\{MSI\}\_\{k\}\-\\tau\_\{\\min\}\}\{\\tau\_\{\\max\}\-\\tau\_\{\\min\}\}\\right\)^\{\\\!2\}\}\_\{\\text\{communication\}\}\+\\underbrace\{100\\,r\_\{\\mathrm\{safe\}\}\}\_\{\\text\{safety\}\},\(6\)whererstab=\+1r\_\{\\mathrm\{stab\}\}=\+1ifV\(𝒙k\+1\)≤V\(𝒙k\)e−λτkV\(\\bm\{x\}\_\{k\+1\}\)\\leq V\(\\bm\{x\}\_\{k\}\)e^\{\-\\lambda\\tau\_\{k\}\}\(with a near\-origin guardV\(𝒙k\)<14VscaleV\(\\bm\{x\}\_\{k\}\)<\\frac\{1\}\{4\}V\_\{\\mathrm\{scale\}\}preventing penalization of residual oscillations\) else−1\-1; the graded1−V\(𝒙k\+1\)/Vscale1\-V\(\\bm\{x\}\_\{k\+1\}\)/V\_\{\\mathrm\{scale\}\}term provides a signal proportional to the absolute Lyapunov value;rsafe∈\{−1,0\}r\_\{\\mathrm\{safe\}\}\\in\\\{\-1,0\\\}flags RTA overrides; and a terminal penalty of−1000\-1000applies on constraint\-violation termination\. The weightwcw\_\{c\}controls the stability/communication tradeoff; we sweepwc∈\{0\.25,0\.5,1,2,4,6,8,10,12,14,16\}w\_\{c\}\\in\\\{0\.25,0\.5,1,2,4,6,8,10,12,14,16\\\}\.
### 3\.3DQN Training and Preference\-Conditioned Extension
A DQN \(256→128→128256\{\\to\}128\{\\to\}128ReLU\) is trained for1M1\\,\\mathrm\{M\}steps using Stable\-Baselines3\[[22](https://arxiv.org/html/2605.12561#bib.bib27)\]with best\-model checkpointing at the per\-step average\-reward peak\. Per\-step rather than total\-episode reward is the natural metric since episodes are time\-bounded: a longer MSI yields fewer steps per fixed\-duration episode\. Per\-step reward thus decouples policy quality fromτk\\tau\_\{k\}choice\. To avoid 11 separate per\-wcw\_\{c\}sweeps, a*preference\-conditioned DQN*\[[1](https://arxiv.org/html/2605.12561#bib.bib31),[31](https://arxiv.org/html/2605.12561#bib.bib32)\]sampleswcw\_\{c\}uniformly at each episode reset from the same grid and appends a log\-normalizedw^c∈\[0,1\]\\hat\{w\}\_\{c\}\\in\[0,1\]to the observation; at deploymentwcw\_\{c\}is fixed per\-episode, so a single2M2\\,\\mathrm\{M\}\-step model recovers the per\-wcw\_\{c\}frontier within11–8%8\\%at211\\tfrac\{2\}\{11\}of total compute\.
## 4Environments
Four environments of increasing complexity span SISO and MIMO underactuated dynamics: the inverted pendulum \(2 states, 1 input; Gymnasium Pendulum\-v1\), CartPole \(4, 1; Gymnasium CartPole\-v1\[[26](https://arxiv.org/html/2605.12561#bib.bib26)\]\), the planar quadrotor \(6, 2; coupled MIMO hover with indirect horizontal control through tilt\), and Quadrotor3D \(12, 4; 6\-DOF rigid\-body quadrotor\) used as the higher\-dimensional case study\. Each plant exhibits a distinct constraint geometry: the pendulum has the narrowest safety margin \(θsat−θRTA=1\.9∘\\theta\_\{\\mathrm\{sat\}\}\-\\theta\_\{\\mathrm\{RTA\}\}=1\.9^\{\\circ\}\); CartPole has a tight termination angle \(12∘12^\{\\circ\}\) and largeVscale≈56\.6V\_\{\\mathrm\{scale\}\}\\approx 56\.6that weakens the graded stability term; the planar quadrotor yields the largest MSI gain over Classical STC \(3\.51×3\.51\\times\) despite a9∘9^\{\\circ\}margin; Quadrotor3D’s5,0005\{,\}000\-action discrete space rules out DQN, motivating SAC on the continuous Box action space, and is also deliberately chosen as a regime where Proposition[1](https://arxiv.org/html/2605.12561#Thmproposition1)’s formal certificate does not extend at the chosenτmin\\tau\_\{\\min\}, enabling us to probe the framework’s empirical reach beyond where the certificate applies \(Section[6](https://arxiv.org/html/2605.12561#S6)\)\. All integrate atΔt=0\.001s\\Delta t=0\.001\\,\\mathrm\{s\}\. Full dynamics, parameters \(Table[5](https://arxiv.org/html/2605.12561#A2.T5)\), and DQN/SAC hyperparameters are in Appendices[B](https://arxiv.org/html/2605.12561#A2)and[C](https://arxiv.org/html/2605.12561#A3); compute resources and wall\-clock times are in Appendix[D](https://arxiv.org/html/2605.12561#A4)\.
## 5Results
Each environment is trained for all 11wcw\_\{c\}values; 100 deterministic evaluation episodes are run per checkpoint\. We report MSI \(mean±\\pmstd\), RTA activation rate \(RTA%\), andL2L\_\{2\}state normsPi=∫0Tsi2\(t\)𝑑tP\_\{i\}=\\sqrt\{\\int\_\{0\}^\{T\}s\_\{i\}^\{2\}\(t\)\\,dt\}\. Full per\-weight sweep tables for DQN, SAC, and Preference\-Conditioned DQN appear in Appendix[H](https://arxiv.org/html/2605.12561#A8)\.
### 5\.1DQN Results and Algorithm Comparison
Table[1](https://arxiv.org/html/2605.12561#S5.T1)summarizes the best\-model checkpoint MSI at the best\-checkpoint peakwcw\_\{c\}for each environment and algorithm\. The DQNwcw\_\{c\}sweeps and their corresponding ablation/Lagrangian comparisons are aggregated across33independent training seeds \(Appendix[E](https://arxiv.org/html/2605.12561#A5)\); all other rows use the standard single\-seed protocol\. The pendulum saturates nearτmax\\tau\_\{\\max\}fromwc=6w\_\{c\}=6onwards\. CartPole reaches96%96\\%ofτmax\\tau\_\{\\max\}atwc=16w\_\{c\}=16with0\.04%0\.04\\%RTA, despite a non\-monotonic final\-model trend caused by the largeVscaleV\_\{\\mathrm\{scale\}\}saturating the graded stability term\. The quadrotor reaches88%88\\%ofτmax\\tau\_\{\\max\}, reflecting the harder coupled hover dynamics\. SAC achieves0%0\\%RTA across every weight and environment in both the final model and best checkpoint, with no high\-wcw\_\{c\}collapses, confirming that DQN collapses are a property of the discrete\-action optimizer rather than the framework\. The preference\-conditioned DQN recovers the per\-wcw\_\{c\}frontier within11–8%8\\%from a single2M2\\,\\mathrm\{M\}\-step model, confirming that the full tradeoff frontier can be recovered at a fraction of total training compute\.
#### Off\-policy methods are required in practice\.
PPO\[[23](https://arxiv.org/html/2605.12561#bib.bib33)\]on the Pendulum either collapses toτmin\\tau\_\{\\min\}or thrashes against the RTA shield at everywcw\_\{c\}\(best MSI≈0\.09s\\approx 0\.09\\,\\mathrm\{s\}vs\. DQN’s0\.397±0\.001s0\.397\\pm 0\.001\\,\\mathrm\{s\}\); the rare\-transition−100\-100RTA and−1000\-1000terminal penalties are inherently easier for replay\-based off\-policy methods to leverage\. The framework is algorithm\-agnostic in principle, but practical convergence within standard training budgets favors off\-policy methods \(full sweep: Table[19](https://arxiv.org/html/2605.12561#A8.T19), Appendix[H](https://arxiv.org/html/2605.12561#A8)\)\.
Table 1:Best\-model checkpoint MSI \(s\) at the best\-checkpoint peakwcw\_\{c\}per algorithm\. DQN values are 3\-seed mean±\\pmstd \(Appendix[E](https://arxiv.org/html/2605.12561#A5)\); SAC and Pref\-DQN values are single\-seed \(seed 0\)\. All three algorithms substantially exceed the Classical STC baseline and a fixed LQR at the same average rate is unstable on all three plants \(B2, Table[2](https://arxiv.org/html/2605.12561#S5.T2)\)\. Full per\-weight sweep tables are in Appendix[H](https://arxiv.org/html/2605.12561#A8)\.
### 5\.2Cross\-Environment Analysis
Figure[2](https://arxiv.org/html/2605.12561#S5.F2)summarizes the best\-model MSI vs\.wcw\_\{c\}; training dynamics \(MSI, RTA activation, and episode reward curves\) are in Appendix[K](https://arxiv.org/html/2605.12561#A11)\(Fig\.[4](https://arxiv.org/html/2605.12561#A11.F4)\)\. Three patterns emerge across DQN, SAC, and Preference\-Conditioned DQN\.\(1\)Higherwcw\_\{c\}accelerates MSI exploration: low\-wcw\_\{c\}policies converge toMSI≈τmin\\mathrm\{MSI\}\\approx\\tau\_\{\\min\}; highwcw\_\{c\}pushesτk\\tau\_\{k\}towardτ∗\(𝒙k\)\\tau^\{\*\}\(\\bm\{x\}\_\{k\}\)\(Proposition[2](https://arxiv.org/html/2605.12561#Thmproposition2)\), and all three methods exceed Classical Lyapunov\-STC at moderate\-to\-highwcw\_\{c\}\.\(2\)At highwcw\_\{c\}the optimizer overshoots admissibility, triggering sustained RTA intervention and reward degradation; the best\-model checkpoint captures the MSI peak before this onset, with the canonical ablationwcw\_\{c\}rising with complexity:wc=8w\_\{c\}=8for Pendulum \(anchored within the saturation plateau,≥96%\\geq 96\\%ofτmax\\tau\_\{\\max\}\) andwc=16w\_\{c\}=16for CartPole and Quadrotor \(strict best\-checkpoint peak\)\.\(3\)Final\-model degradation depends on plant geometry: Pendulum and Quadrotor tolerate highwcw\_\{c\}, while CartPole’s largeVscale≈56\.6V\_\{\\mathrm\{scale\}\}\\approx 56\.6and12∘12^\{\\circ\}termination margin produce non\-monotonic final\-model MSI; the best\-model checkpoint is especially critical there, yielding a clean tradeoff \(0\.308±0\.014s0\.308\\pm 0\.014\\,\\mathrm\{s\},96%96\\%ofτmax\\tau\_\{\\max\}\)\. Detailed per\-environment analyses are in Appendix[L](https://arxiv.org/html/2605.12561#A12)\.
Figure 2:Best\-model MSI vs\.wcw\_\{c\}for DQN, SAC, and Preference\-Conditioned DQN across all three environments \(seed\-0 curves; multi\-seed DQN aggregates are in Tables[10](https://arxiv.org/html/2605.12561#A8.T10)–[12](https://arxiv.org/html/2605.12561#A8.T12)\)\. Reference lines markτmin\\tau\_\{\\min\}\(dotted\),τmax\\tau\_\{\\max\}\(dashed\), and Classical Lyapunov\-STC baseline B3 \(dash\-dot green\)\. All algorithms exceed B3 at moderate\-to\-highwcw\_\{c\}; the gap above B3 represents communication efficiency that conservative analytical methods leave on the table\. DQN saturates nearτmax\\tau\_\{\\max\}for Pendulum and CartPole; SAC converges more reliably but peaks slightly lower; Pref\-DQN matches DQN within11–8%8\\%from a single model\.
### 5\.3Baseline Comparisons
Three baselines isolate the sources of communication efficiency \(Neval=100N\_\{\\mathrm\{eval\}\}=100each\)\.
B1 – Fixed LQR atτmin\\tau\_\{\\min\}:Always stable, never communication\-efficient\. Establishes the hard safety floor\.
B2 – Fixed LQR atτmatch\\tau\_\{\\mathrm\{match\}\}:The same LQR controller at a fixed interval equal to the seed\-0 best\-checkpoint MSI:0\.397s0\.397\\,\\mathrm\{s\}\(Pendulum\),0\.317s0\.317\\,\\mathrm\{s\}\(CartPole\),0\.290s0\.290\\,\\mathrm\{s\}\(Quadrotor\)\.*This baseline is unstable on all three plants*\(mean episode lengths1\.11\.1–2\.4s2\.4\\,\\mathrm\{s\}\), proving that adaptive timing, not merely a reduced average rate, is what makes sparsity safe: each plant’s unstable mode has a time constant shorter thanτmatch\\tau\_\{\\mathrm\{match\}\}\(e\.g\.,2l/\(3g\)≈0\.258s\\sqrt\{2l/\(3g\)\}\\approx 0\.258\\,\\mathrm\{s\}for the Pendulum’s rod dynamics\), so the fixed\-rate ZOH closed loop diverges\. The RL policy avoids this by sampling fast near the unstable manifold and stretchingτk\\tau\_\{k\}near equilibrium \(Proposition[2](https://arxiv.org/html/2605.12561#Thmproposition2)\)\.
B3 – Classical Lyapunov\-STC:At each step the LQR control is applied andτk\\tau\_\{k\}is chosen greedily as the largest value in𝒯\\mathcal\{T\}satisfyingV\(𝒙~\(τk\)\)≤V\(𝒙k\)e−λτkV\(\\tilde\{\\bm\{x\}\}\(\\tau\_\{k\}\)\)\\leq V\(\\bm\{x\}\_\{k\}\)e^\{\-\\lambda\\tau\_\{k\}\}\. This baseline maintains stability but achieves1\.91×1\.91\\times,1\.45×1\.45\\times, and3\.51×3\.51\\timeslower MSI than RL\-STC on Pendulum, CartPole, and Quadrotor respectively \(Table[2](https://arxiv.org/html/2605.12561#S5.T2)\); the gap is largest for the Quadrotor as the tightθ˙\\dot\{\\theta\}coupling pins the trigger nearτmin\\tau\_\{\\min\}\. The RL policy discovers that longer intervals are safe on average even where the instantaneous linearized Lyapunov bound is violated in the prediction\.
Table 2:Baseline comparison \(Neval=100N\_\{\\mathrm\{eval\}\}=100\)\. RL\-STC values are seed 0 to align with theτmatch\\tau\_\{\\mathrm\{match\}\}used for B2 \(the seed\-0 best\-checkpoint MSI; multi\-seed canonical bests differ by less than theτ\\tau\-grid spacing per Appendix[F](https://arxiv.org/html/2605.12561#A6)\)\.Italics= system failure \(ep\. length<3s\{<\}3\\,\\mathrm\{s\}\); “–” = norms not meaningful for failed episodes\.The RL\-STC policy simultaneously achieves the high inter\-sample sparsity of B2*and*the full\-episode stability of B1, demonstrating that the joint learning of control input and timing is necessary to unlock the communication efficiency that neither fixed\-rate nor greedy trigger\-based controllers can deliver\. All reportedPiP\_\{i\}values correspond to episodes that completed the full50s50\\,\\mathrm\{s\}horizon, confirming that larger excursions reflect higher\-MSI trajectories rather than unstable operation\.
### 5\.4Ablation Study
#### Ablation A – Removing the RTA Shield\.
With the RTA override and penalty disabled \(wc=8/16/16w\_\{c\}=8/16/16, canonical per Section[5\.2](https://arxiv.org/html/2605.12561#S5.SS2)\), best\-checkpoint MSI drops by1\.271\.27–1\.84×1\.84\\times\(Pendulum:0\.385→0\.266s0\.385\\to 0\.266\\,\\mathrm\{s\}; CartPole:0\.308→0\.167s0\.308\\to 0\.167\\,\\mathrm\{s\}; Quadrotor:0\.281→0\.222s0\.281\\to 0\.222\\,\\mathrm\{s\}, Table[3](https://arxiv.org/html/2605.12561#S5.T3)\), and state norms degrade substantially on CartPole and Quadrotor \(P1\(x\)P\_\{1\}\(x\):2\.95→5\.812\.95\\to 5\.81on CartPole,2\.33→6\.992\.33\\to 6\.99on Quadrotor;P3\(θ\)P\_\{3\}\(\\theta\)on Quadrotor degrades2\.30×2\.30\\times\)\. The Lyapunov reward alone cannot replicate the RTA’s role: the shield shapes the feasible policy space during training and provides a hard safety floor at deployment that the learned policy can leverage\.
Table 3:Ablation A – No RTA vs\. full method \(Neval=100N\_\{\\mathrm\{eval\}\}=100, best\-model checkpoint, 3 seeds per row; MSI as mean±\\pmstd across the 3 per\-seed mean values withddof=1\\mathrm\{ddof\}=1,PiP\_\{i\}norms as mean only;wc=8/16/16w\_\{c\}=8/16/16for Pendulum/CartPole/Quadrotor\)\.Figure[3](https://arxiv.org/html/2605.12561#S5.F3)provides geometric evidence: the RTA\-enabled policy \(top\) occupies a compact high\-reward region with1\.271\.27–1\.84×1\.84\\timesfewer control steps per episode \(matching the inverse MSI ratios\), while the no\-RTA policy \(bottom\) visits a much wider region at substantially lower per\-step rewards, isolating the RTA as a geometric constraint that enables long\-interval strategies the Lyapunov reward alone cannot enforce\.
Figure 3:Per\-step reward distribution \(100 episodes; seed\-0 trajectories per Appendix[E](https://arxiv.org/html/2605.12561#A5)\): RL\-STC \(top\) vs\. no\-RTA \(bottom\)\. Pendulum axes are\(θ,θ˙\)\(\\theta,\\dot\{\\theta\}\); CartPole and Quadrotor use a shared t\-SNE embedding fitted jointly on both policies’ trajectories\. Color encodes per\-step reward \(green high, red low\)\. The RTA\-enabled policy occupies a compact high\-reward region with1\.271\.27–1\.84×1\.84\\timesfewer steps per episode; the no\-RTA policy visits a wider region at lower rewards\.
#### Ablation B – Fixedτ\\tau\.
Fixingτ\\tauat the nearest grid value toτmatch\\tau\_\{\\mathrm\{match\}\}\(with RTA active\) and learning only the control input degenerates to LQR atτmin\\tau\_\{\\min\}on Pendulum \(RTA fires on80\.6%80\.6\\%of steps\), reaches a near\-tie on CartPole only becauseτmatch=τmax\\tau\_\{\\mathrm\{match\}\}=\\tau\_\{\\max\}there, and gives up a4\.3%4\.3\\%MSI advantage on Quadrotor\. Joint optimization of\(𝒖k,τk\)\(\\bm\{u\}\_\{k\},\\tau\_\{k\}\)is not separable; full results in Appendix[F](https://arxiv.org/html/2605.12561#A6)\.
### 5\.5Comparison with Lagrangian Safe RL
A Lagrangian\-DQN baseline replaces the hard RTA override with a soft penaltyrkLag=rk−λtckr\_\{k\}^\{\\mathrm\{Lag\}\}=r\_\{k\}\-\\lambda\_\{t\}c\_\{k\}on the same one\-step\-ahead predicate \(λt\\lambda\_\{t\}updated via dual gradient ascent\)\. Across the three lower\-dimensional environments, Lagrangian\-DQN achieves2121–52%52\\%lower best\-checkpoint MSI than RL\-STC and accumulates up to35\.40%35\.40\\%hard violations on Quadrotor, while RL\-STC has zero hard violations by construction; final\-model Lagrangian deteriorates further \(e\.g\.,43\.49±48\.55%43\.49\\pm 48\.55\\%hard violations on Pendulum,59\.35%59\.35\\%on Quadrotor\) asλ\\lambdafails to hold the constraint across seeds\. The hard override decouples safety from inter\-sample interval selection during training, enabling long\-interval commitment without the hedging that a soft penalty forces\. Full per\-environment results are in Appendix[G](https://arxiv.org/html/2605.12561#A7), Table[9](https://arxiv.org/html/2605.12561#A7.T9)\.
### 5\.6Robustness
Best\-model policies are evaluated under±30%\\pm 30\\%mass mismatch and under additive constant, periodic, and impulse disturbances injected directly into the plant \(KK,PP, and RTA thresholds fixed at nominal\)\.*Classical Lyapunov\-STC fails on Pendulum and CartPole at0\.7×0\.7\\timesmass*, the faster actual dynamics exceeding the nominal prediction’s stability limit; RL\-STC remains stable across all scales by trading MSI for safety\. At0\.7×0\.7\\timesCartPole mass MSI drops37%37\\%to0\.201s0\.201\\,\\mathrm\{s\}with13\.91%13\.91\\%RTA activation and recovers fully at1\.3×1\.3\\times; domain randomization over±40%\\pm 40\\%mass eliminates the sensitivity at negligible cost to nominal performance\. Under disturbances, RTA activation grows monotonically with amplitude while every episode completes safely \(the filter absorbs what the policy cannot\), and RL\-STC maintains higher MSI than Classical STC under all conditions; the Quadrotor is structurally decoupled from thrust\-channel disturbances, with MSI varying by at most0\.002s0\.002\\,\\mathrm\{s\}across all conditions\. Full mass\-mismatch and disturbance tables \(Appendix[J](https://arxiv.org/html/2605.12561#A10), Tables[26](https://arxiv.org/html/2605.12561#A10.T26)and[27](https://arxiv.org/html/2605.12561#A10.T27)\) and per\-channel analysis \(Appendix[M](https://arxiv.org/html/2605.12561#A13)\) are in the supplementary material\.
## 6Case Study: Scaling to Higher\-Dimensional Systems
The three lower\-dimensional environments share discrete action spaces of168168–792792actions that DQN can adequately cover\. Quadrotor3D is qualitatively different: its5,0005\{,\}000\-action discrete space leaves DQN with under0\.50\.5samples per action, rendering discrete learning intractable\. We adopt SAC\[[15](https://arxiv.org/html/2605.12561#bib.bib28)\]on the equivalent continuous Box action space, keeping the reward, RTA, and Lyapunov construction identical\. Training extends to2M2\\,\\mathrm\{M\}steps with the policy widened to256→128→128256\{\\to\}128\{\\to\}128, and the sweep extends towc∈\{0\.25,1,4,8,16,24,32,40,48,56,64\}w\_\{c\}\\in\\\{0\.25,1,4,8,16,24,32,40,48,56,64\\\}\.
Pointwise safety in our framework decomposes into a*deployment\-time*override and a*training\-time*shaping signal \(the RTA penalty cascading to the terminal penalty when the held backup propagates state out of bounds\)\. The case study deliberately probes the regime where the deployment certificate does not extend: atτmin=0\.04s\\tau\_\{\\min\}=0\.04\\,\\mathrm\{s\}the ZOH\-discretized closed loop has spectral radius1\.191\.19\(τc≈0\.037s\\tau\_\{c\}\\approx 0\.037\\,\\mathrm\{s\}\), mirroring the conditions in which analytical STC design becomes intractable on higher\-dimensional systems\. The certificate could be restored locally by a smallerτmin\\tau\_\{\\min\}givingτc\>τmin\\tau\_\{c\}\>\\tau\_\{\\min\}, a discrete\-LQR redesign matched toτmin\\tau\_\{\\min\}, or a non\-quadratic Lyapunov certificate; we instead test whether the training\-time shaping alone produces a policy that respects the safety predicate\. The trained policy maintains0%0\\%RTA across the nominal sweep,±30%\\pm 30\\%mass perturbation, and the disturbance suite reported below, a property of the trained policy on the tested distribution rather than a formal guarantee for states or conditions outside it\.
#### SAC sweep and Pref\-SAC\.
The sweep reveals a two\-phase dynamic:wc≤16w\_\{c\}\\leq 16keeps both final and best nearτmin\\tau\_\{\\min\};wc≥40w\_\{c\}\\geq 40saturates nearτmax=0\.32s\\tau\_\{\\max\}=0\.32\\,\\mathrm\{s\}at0%0\\%RTA\. We selectwc=48w\_\{c\}=48, where the best\-model checkpoint achieves MSI=0\.302s=0\.302\\,\\mathrm\{s\}\(94\.3%94\.3\\%ofτmax\\tau\_\{\\max\},0%0\\%RTA\) with low attitude norms \(P3\(φ\)≈0\.04P\_\{3\}\(\\varphi\)\\approx 0\.04–0\.050\.05\)\. A single preference\-conditioned SAC model \(4M4\\,\\mathrm\{M\}steps,wcw\_\{c\}sampled per episode\) reaches peak MSI0\.239s0\.239\\,\\mathrm\{s\}\(79%79\\%of the per\-wcw\_\{c\}SAC result\) at≤0\.01%\\leq 0\.01\\%RTA across the full grid, recovering the frontier from one model\.
#### Baselines, ablations, and Lagrangian\-SAC\.
Classical Lyapunov\-STC remains pinned atτmin\\tau\_\{\\min\}throughout: the one\-step\-ahead trigger is maximally conservative on the 12\-state system, firing at every step\. Baseline 2 \(LQR atτmatch=0\.302s\\tau\_\{\\mathrm\{match\}\}=0\.302\\,\\mathrm\{s\}\) crashes in0\.39±0\.14s0\.39\\pm 0\.14\\,\\mathrm\{s\}, fewer than two steps, the most dramatic confirmation that adaptive timing is a prerequisite\. RL\-STC achieves7\.5×7\.5\\timeslonger inter\-sample intervals than B1/B3 with attitude\-rate normP4\(p\)=0\.190P\_\{4\}\(p\)=0\.190vs\.7\.237\.23\. Removing the RTA \(Ablation A\) maintains MSI≈0\.310s\\approx 0\.310\\,\\mathrm\{s\}but roughly doubles attitude norms \(P3\(φ\)P\_\{3\}\(\\varphi\):0\.044→0\.0850\.044\\to 0\.085\)\. Lagrangian\-SAC converges to MSI≈0\.311\\approx 0\.311–0\.312s0\.312\\,\\mathrm\{s\}with zero hard violations, matching RL\-STC empirically; SAC’s entropy regularization keeps the policy away from the constraint\. Neither method holds a pointwise certificate on Q3D, but RL\-STC retains a constructive path to Proposition[1](https://arxiv.org/html/2605.12561#Thmproposition1)via discrete\-LQR redesign or a smallerτmin\\tau\_\{\\min\}\(Section[7](https://arxiv.org/html/2605.12561#S7)\), while Lagrangian\-SAC admits no analytical extension\.
#### Robustness\.
The nominal RL\-STC policy shows*zero*degradation under±30%\\pm 30\\%mass perturbation \(MSI and RTA unchanged at0\.302s0\.302\\,\\mathrm\{s\}and0%0\\%\), while Classical STC remains pinned atτmin\\tau\_\{\\min\}throughout\. It also shows complete insensitivity to additive thrust disturbances \(constant up to1\.0N1\.0\\,\\mathrm\{N\}, periodic up to1\.5N1\.5\\,\\mathrm\{N\}at2Hz2\\,\\mathrm\{Hz\}, impulses up to2\.0N2\.0\\,\\mathrm\{N\}\): MSI remains0\.302s0\.302\\,\\mathrm\{s\}with0%0\\%RTA across all seven conditions, mirroring the structural decoupling in the 2D Quadrotor \(the angular RTA trigger monitors tilt while thrust disturbances perturb the altitude channel\)\. Full Q3D sweep, ablation, Lagrangian, mismatch, and disturbance tables are in Appendix[I](https://arxiv.org/html/2605.12561#A9)\.
## 7Limitations and Future Work
The RTA’s linearized one\-step\-ahead prediction grows conservative far from the operating point, the direct mechanism behind the empirical MSI ceiling and high\-wcw\_\{c\}DQN collapses; a learned residual dynamics term could raise this ceiling without sacrificing the hard safety floor\. The framework is scoped to stabilization around a known equilibrium so the CARE solution, quadraticVV, and linearized one\-step prediction remain well posed; non\-equilibrium tasks \(tracking, navigation, manipulation\) and purely black\-box dynamics lie outside this scope\. Control barrier functions\[[14](https://arxiv.org/html/2605.12561#bib.bib22)\]offer a natural extension but require a different backup construction\. Scaling beyond the 12\-state Quadrotor3D to higher\-dimensional systems, contact\-rich dynamics, and real\-world platforms remains the primary direction for future work\.
## 8Conclusion
We have presented a unified RL framework for communication\-efficient self\-triggered control with run\-time assurance, evaluated across four environments from a 2\-state pendulum to a 12\-state 3D quadrotor case study\. The CARE\-derived Lyapunov reward transfers without redesign, and Proposition[2](https://arxiv.org/html/2605.12561#Thmproposition2)provides a formal basis for the observed MSI ceiling\. Three findings establish the framework’s value:\(i\)classical baselines fail \(B2 unstable on all three lower\-dimensional plants and crashes within two control updates on Quadrotor3D; B3 leaves a1\.45–3\.51×1\.45\\text\{\-\-\}3\.51\\timesMSI gap and remains pinned atτmin\\tau\_\{\\min\}on Quadrotor3D\);\(ii\)on plants whereMdisc⪰0M\_\{\\mathrm\{disc\}\}\\succeq 0, the RTA hard override delivers Lyapunov\-decrease safety by construction regardless of algorithm, in contrast to Lagrangian\-DQN which produces21–52%21\\text\{\-\-\}52\\%lower MSI with up to35%35\\%hard violations;\(iii\)ablations confirm RTA removal degrades MSI by1\.27–1\.84×1\.27\\text\{\-\-\}1\.84\\timesand that joint optimization of\(𝒖k,τk\)\(\\bm\{u\}\_\{k\},\\tau\_\{k\}\)is not separable\. SAC and a preference\-conditioned extension confirm algorithm\-agnostic gains and frontier recovery at211\\tfrac\{2\}\{11\}of total compute\. The Quadrotor3D case study reaches MSI=0\.302s=0\.302\\,\\mathrm\{s\}\(94%94\\%ofτmax\\tau\_\{\\max\}\) with zero RTA activation, extending the framework to systems where analytical STC design is intractable and where the formal certificate does not extend at the chosenτmin\\tau\_\{\\min\}, and robustness experiments across all four plants confirm graceful degradation while every episode completes safely\. The question of*when*to act is as consequential as*what*to do, and answering it safely and efficiently is tractable via the joint optimization this framework provides\.
## Acknowledgments and Disclosure of Funding
A\. Haroon conducted part of this work as an NREIP intern at the U\.S\. Naval Research Laboratory\. The views expressed are those of the authors and do not reflect the official policy or position of the U\.S\. Naval Academy, Department of the Navy, Department of War, or U\.S\. Government\.
## References
- \[1\]A\. Abels, D\. Roijers, T\. Lenaerts, A\. Nowé, and D\. Steckelmacher\(2019\)Dynamic weights in multi\-objective deep reinforcement learning\.InInternational Conference on Machine Learning,pp\. 11–20\.Cited by:[§3\.3](https://arxiv.org/html/2605.12561#S3.SS3.p1.12)\.
- \[2\]J\. Achiam, D\. Held, A\. Tamar, and P\. Abbeel\(2017\)Constrained policy optimization\.InInternational Conference on Machine Learning,pp\. 22–31\.Cited by:[§1](https://arxiv.org/html/2605.12561#S1.p2.1),[§1](https://arxiv.org/html/2605.12561#S1.p3.3)\.
- \[3\]S\. Aggarwal, D\. Maity, and T\. Başar\(2025\)InterQ: a dqn framework for optimal intermittent control\.IEEE Control Systems Letters\.Cited by:[§1](https://arxiv.org/html/2605.12561#S1.p2.1)\.
- \[4\]M\. Alshiekh, R\. Bloem, R\. Ehlers, B\. Könighofer, S\. Niekum, and U\. Topcu\(2018\)Safe reinforcement learning via shielding\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.32\.Cited by:[§1](https://arxiv.org/html/2605.12561#S1.p3.3)\.
- \[5\]E\. Altman\(2021\)Constrained markov decision processes\.Routledge\.Cited by:[§1](https://arxiv.org/html/2605.12561#S1.p2.1),[§1](https://arxiv.org/html/2605.12561#S1.p3.3)\.
- \[6\]A\. Anta and P\. Tabuada\(2010\)To sample or not to sample: self\-triggered control for nonlinear systems\.IEEE Transactions on Automatic Control55\(9\),pp\. 2030–2042\.Cited by:[§1](https://arxiv.org/html/2605.12561#S1.p2.1)\.
- \[7\]D\. Baumann, J\. Zhu, G\. Martius, and S\. Trimpe\(2018\)Deep reinforcement learning for event\-triggered control\.In2018 IEEE Conference on Decision and Control \(CDC\),pp\. 943–950\.Cited by:[§1](https://arxiv.org/html/2605.12561#S1.p2.1)\.
- \[8\]S\. Boyd, L\. El Ghaoui, E\. Feron, and V\. Balakrishnan\(1994\)Linear matrix inequalities in system and control theory\.SIAM\.Cited by:[§3\.1](https://arxiv.org/html/2605.12561#S3.SS1.p1.9)\.
- \[9\]L\. Brunke, M\. Greeff, A\. W\. Hall, Z\. Yuan, S\. Zhou, J\. Panerati, and A\. P\. Schoellig\(2022\)Safe learning in robotics: from learning\-based control to safe reinforcement learning\.Annual Review of Control, Robotics, and Autonomous Systems5\(1\),pp\. 411–444\.Cited by:[§1](https://arxiv.org/html/2605.12561#S1.p3.3)\.
- \[10\]K\. Dunlap, M\. Mote, K\. Delsing, and K\. L\. Hobbs\(2023\)Run time assured reinforcement learning for safe satellite docking\.Journal of Aerospace Information Systems20\(1\),pp\. 25–36\.Cited by:[§1](https://arxiv.org/html/2605.12561#S1.p3.3)\.
- \[11\]N\. Funk, D\. Baumann, V\. Berenz, and S\. Trimpe\(2021\)Learning event\-triggered control from data through joint optimization\.IFAC Journal of Systems and Control16,pp\. 100144\.Cited by:[§1](https://arxiv.org/html/2605.12561#S1.p2.1)\.
- \[12\]J\. Garcıa and F\. Fernández\(2015\)A comprehensive survey on safe reinforcement learning\.Journal of Machine Learning Research16\(1\),pp\. 1437–1480\.Cited by:[§1](https://arxiv.org/html/2605.12561#S1.p3.3)\.
- \[13\]T\. Gommans, D\. Antunes, T\. Donkers, P\. Tabuada, and M\. Heemels\(2014\)Self\-triggered linear quadratic control\.Automatica50\(4\),pp\. 1279–1287\.Cited by:[§1](https://arxiv.org/html/2605.12561#S1.p2.1)\.
- \[14\]S\. Gu, L\. Yang, Y\. Du, G\. Chen, F\. Walter, J\. Wang, and A\. Knoll\(2024\)A review of safe reinforcement learning: methods, theories, and applications\.IEEE Transactions on Pattern Analysis and Machine Intelligence46\(12\),pp\. 11216–11235\.Cited by:[§1](https://arxiv.org/html/2605.12561#S1.p3.3),[§7](https://arxiv.org/html/2605.12561#S7.p1.2)\.
- \[15\]T\. Haarnoja, A\. Zhou, P\. Abbeel, and S\. Levine\(2018\)Soft actor\-critic: off\-policy maximum entropy deep reinforcement learning with a stochastic actor\.InInternational conference on machine learning,pp\. 1861–1870\.Cited by:[§6](https://arxiv.org/html/2605.12561#S6.p1.7)\.
- \[16\]W\. P\. Heemels, K\. H\. Johansson, and P\. Tabuada\(2012\)An introduction to event\-triggered and self\-triggered control\.In2012 IEEE 51st IEEE Conference on Decision and Control \(CDC\),pp\. 3270–3285\.Cited by:[§1](https://arxiv.org/html/2605.12561#S1.p2.1)\.
- \[17\]K\. L\. Hobbs, M\. L\. Mote, M\. C\. Abate, S\. D\. Coogan, and E\. M\. Feron\(2023\)Runtime assurance for safety\-critical systems: an introduction to safety filtering approaches for complex control systems\.IEEE Control Systems Magazine43\(2\),pp\. 28–65\.Cited by:[§1](https://arxiv.org/html/2605.12561#S1.p3.3)\.
- \[18\]H\. K\. Khalil and J\. W\. Grizzle\(2002\)Nonlinear systems\.Vol\.3,Prentice Hall Upper Saddle River, NJ\.Cited by:[Remark 2](https://arxiv.org/html/2605.12561#Thmremark2.p1.8.8)\.
- \[19\]C\. Lazarus, J\. G\. Lopez, and M\. J\. Kochenderfer\(2020\)Runtime safety assurance using reinforcement learning\.In2020 AIAA/IEEE 39th Digital Avionics Systems Conference \(DASC\),pp\. 1–9\.Cited by:[§1](https://arxiv.org/html/2605.12561#S1.p3.3)\.
- \[20\]M\. Mazo and P\. Tabuada\(2011\)Decentralized event\-triggered control over wireless sensor/actuator networks\.IEEE Transactions on Automatic Control56\(10\),pp\. 2456–2461\.Cited by:[§1](https://arxiv.org/html/2605.12561#S1.p2.1)\.
- \[21\]K\. Miller, C\. K\. Zeitler, W\. Shen, K\. Hobbs, J\. Schierman, M\. Viswanathan, and S\. Mitra\(2024\)Optimal runtime assurance via reinforcement learning\.In2024 ACM/IEEE 15th International Conference on Cyber\-Physical Systems \(ICCPS\),pp\. 67–76\.Cited by:[§1](https://arxiv.org/html/2605.12561#S1.p3.3)\.
- \[22\]A\. Raffin, A\. Hill, A\. Gleave, A\. Kanervisto, M\. Ernestus, and N\. Dormann\(2021\)Stable\-baselines3: reliable reinforcement learning implementations\.Vol\.22\.Cited by:[§3\.3](https://arxiv.org/html/2605.12561#S3.SS3.p1.12)\.
- \[23\]J\. Schulman, F\. Wolski, P\. Dhariwal, A\. Radford, and O\. Klimov\(2017\)Proximal policy optimization algorithms\.arXiv preprint arXiv:1707\.06347\.Cited by:[§5\.1](https://arxiv.org/html/2605.12561#S5.SS1.SSS0.Px1.p1.6)\.
- \[24\]L\. Sedghi, Z\. Ijaz, M\. Noor\-A\-Rahim, K\. Witheephanich, and D\. Pesch\(2022\)Machine learning in event\-triggered control: recent advances and open issues\.IEEE Access10,pp\. 74671–74690\.Cited by:[§1](https://arxiv.org/html/2605.12561#S1.p2.1)\.
- \[25\]D\. Seto, B\. Krogh, L\. Sha, and A\. Chutinan\(1998\)The simplex architecture for safe online control system upgrades\.InProceedings of the 1998 American Control Conference\. ACC \(IEEE Cat\. No\. 98CH36207\),Vol\.6,pp\. 3504–3508\.Cited by:[§1](https://arxiv.org/html/2605.12561#S1.p3.3)\.
- \[26\]M\. Towers, A\. Kwiatkowski, J\. Terry, J\. U\. Balis, G\. De Cola, T\. Deleu, M\. Goulão, A\. Kallinteris, M\. Krimmel, A\. KG,et al\.\(2024\)Gymnasium: a standard interface for reinforcement learning environments\.arXiv preprint arXiv:2407\.17032\.Cited by:[§B\.2](https://arxiv.org/html/2605.12561#A2.SS2.p1.20),[§4](https://arxiv.org/html/2605.12561#S4.p1.8)\.
- \[27\]L\. Treven, B\. Sukhija, Y\. As, F\. Dörfler, and A\. Krause\(2024\)When to sense and control? a time\-adaptive approach for continuous\-time rl\.Advances in Neural Information Processing Systems37,pp\. 63654–63685\.Cited by:[§1](https://arxiv.org/html/2605.12561#S1.p2.1)\.
- \[28\]H\. Wan, H\. R\. Karimi, X\. Luan, and F\. Liu\(2023\)Model\-free self\-triggered control based on deep reinforcement learning for unknown nonlinear systems\.International Journal of Robust and Nonlinear Control33\(3\),pp\. 2238–2250\.Cited by:[§1](https://arxiv.org/html/2605.12561#S1.p2.1)\.
- \[29\]R\. Wang, I\. Takeuchi, and K\. Kashima\(2021\)Deep reinforcement learning for continuous\-time self\-triggered control\.IFAC\-PapersOnLine54\(14\),pp\. 203–208\.Cited by:[§1](https://arxiv.org/html/2605.12561#S1.p2.1)\.
- \[30\]X\. Wang and M\. D\. Lemmon\(2010\)Event\-triggering in distributed networked control systems\.IEEE Transactions on Automatic Control56\(3\),pp\. 586–601\.Cited by:[§1](https://arxiv.org/html/2605.12561#S1.p2.1)\.
- \[31\]R\. Yang, X\. Sun, and K\. Narasimhan\(2019\)A generalized algorithm for multi\-objective reinforcement learning and policy adaptation\.Advances in Neural Information Processing Systems32\.Cited by:[§3\.3](https://arxiv.org/html/2605.12561#S3.SS3.p1.12)\.
## Appendix AProofs
### A\.1Proof of Proposition[1](https://arxiv.org/html/2605.12561#Thmproposition1)
Under ZOH with𝒖k=−K𝒙k\\bm\{u\}\_\{k\}=\-K\\bm\{x\}\_\{k\}held constant on\[tk,tk\+τmin\)\[t\_\{k\},\\,t\_\{k\}\+\\tau\_\{\\min\}\), the linear plant𝒙˙=A𝒙\+B𝒖\\dot\{\\bm\{x\}\}=A\\bm\{x\}\+B\\bm\{u\}integrates to
𝒙\(tk\+τmin\)=eAτmin𝒙k\+\(∫0τmineAs𝑑s\)B𝒖k=M\(τmin\)𝒙k,\\bm\{x\}\(t\_\{k\}\+\\tau\_\{\\min\}\)=e^\{A\\tau\_\{\\min\}\}\\bm\{x\}\_\{k\}\+\\biggl\(\\\!\\int\_\{0\}^\{\\tau\_\{\\min\}\}\\\!e^\{As\}\\,ds\\biggr\)B\\,\\bm\{u\}\_\{k\}=M\(\\tau\_\{\\min\}\)\\,\\bm\{x\}\_\{k\},so𝒙k\+1=M\(τmin\)𝒙k\\bm\{x\}\_\{k\+1\}=M\(\\tau\_\{\\min\}\)\\bm\{x\}\_\{k\}\. Substituting intoV\(𝒙\)=𝒙⊤P𝒙V\(\\bm\{x\}\)=\\bm\{x\}^\{\\top\}P\\bm\{x\}givesV\(𝒙k\+1\)=𝒙k⊤M\(τmin\)⊤PM\(τmin\)𝒙kV\(\\bm\{x\}\_\{k\+1\}\)=\\bm\{x\}\_\{k\}^\{\\top\}M\(\\tau\_\{\\min\}\)^\{\\top\}P\\,M\(\\tau\_\{\\min\}\)\\,\\bm\{x\}\_\{k\}, soV\(𝒙k\+1\)≤V\(𝒙k\)V\(\\bm\{x\}\_\{k\+1\}\)\\leq V\(\\bm\{x\}\_\{k\}\)for all𝒙k\\bm\{x\}\_\{k\}is equivalent toMdisc⪰0M\_\{\\mathrm\{disc\}\}\\succeq 0\. Table[4](https://arxiv.org/html/2605.12561#A1.T4)reportsλmin\(Mdisc\)\>0\\lambda\_\{\\min\}\(M\_\{\\mathrm\{disc\}\}\)\>0for the three lower\-dimensional environments, directly verifying this condition\.□\\square
### A\.2Proof of Proposition[2](https://arxiv.org/html/2605.12561#Thmproposition2)
Under ZOH with𝒖k=−K𝒙k\\bm\{u\}\_\{k\}=\-K\\bm\{x\}\_\{k\}held constant on\[tk,tk\+τ\)\[t\_\{k\},\\,t\_\{k\}\+\\tau\), the linear closed\-loop solution is𝒙\(τ\)=M\(τ\)𝒙k\\bm\{x\}\(\\tau\)=M\(\\tau\)\\,\\bm\{x\}\_\{k\}withM\(⋅\)M\(\\cdot\)from \([4](https://arxiv.org/html/2605.12561#S2.E4)\)\. Define the Lyapunov differenceΔV\(τ\)≜V\(M\(τ\)𝒙k\)−V\(𝒙k\)\\Delta V\(\\tau\)\\triangleq V\(M\(\\tau\)\\bm\{x\}\_\{k\}\)\-V\(\\bm\{x\}\_\{k\}\);M\(0\)=IM\(0\)=IgivesΔV\(0\)=0\\Delta V\(0\)=0\. Differentiating,
ddτΔV\(τ\)=𝒙k⊤\(M˙\(τ\)⊤PM\(τ\)\+M\(τ\)⊤PM˙\(τ\)\)𝒙k\.\\frac\{d\}\{d\\tau\}\\Delta V\(\\tau\)=\\bm\{x\}\_\{k\}^\{\\top\}\\\!\\bigl\(\\dot\{M\}\(\\tau\)^\{\\top\}P\\,M\(\\tau\)\+M\(\\tau\)^\{\\top\}P\\,\\dot\{M\}\(\\tau\)\\bigr\)\\,\\bm\{x\}\_\{k\}\.Differentiating \([4](https://arxiv.org/html/2605.12561#S2.E4)\) givesM˙\(τ\)=AeAτ−eAτBK=eAτ\(A−BK\)=eAτAcl\\dot\{M\}\(\\tau\)=Ae^\{A\\tau\}\-e^\{A\\tau\}BK=e^\{A\\tau\}\(A\-BK\)=e^\{A\\tau\}A\_\{\\mathrm\{cl\}\}, soM˙\(0\)=Acl\\dot\{M\}\(0\)=A\_\{\\mathrm\{cl\}\}\. Substituting atτ=0\\tau=0and applying the CARE identityAcl⊤P\+PAcl=−\(Q\+K⊤RK\)≺0A\_\{\\mathrm\{cl\}\}^\{\\top\}P\+PA\_\{\\mathrm\{cl\}\}=\-\(Q\+K^\{\\top\}RK\)\\prec 0,
ddτΔV\(τ\)\|τ=0=𝒙k⊤\(Acl⊤P\+PAcl\)𝒙k=−𝒙k⊤\(Q\+K⊤RK\)𝒙k<0\\left\.\\frac\{d\}\{d\\tau\}\\Delta V\(\\tau\)\\right\|\_\{\\tau=0\}=\\bm\{x\}\_\{k\}^\{\\top\}\(A\_\{\\mathrm\{cl\}\}^\{\\top\}P\+PA\_\{\\mathrm\{cl\}\}\)\\,\\bm\{x\}\_\{k\}=\-\\bm\{x\}\_\{k\}^\{\\top\}\(Q\+K^\{\\top\}RK\)\\,\\bm\{x\}\_\{k\}<0for all𝒙k≠𝟎\\bm\{x\}\_\{k\}\\neq\\bm\{0\}\. By continuity ofΔV\(τ\)\\Delta V\(\\tau\)inτ\\tau, there existsτ∗\(𝒙k\)\>0\\tau^\{\*\}\(\\bm\{x\}\_\{k\}\)\>0such thatΔV\(τ\)<0\\Delta V\(\\tau\)<0for allτ∈\(0,τ∗\(𝒙k\)\)\\tau\\in\(0,\\tau^\{\*\}\(\\bm\{x\}\_\{k\}\)\)\.□\\square
### A\.3Corollary[A\.3](https://arxiv.org/html/2605.12561#A1.SS3): State\-Independent Admissible Interval
###### Corollary 1\.
For any0<c≤c′0<c\\leq c^\{\\prime\}, let𝒜c,c′=\{𝐱k:c≤V\(𝐱k\)≤c′\}\\mathcal\{A\}\_\{c,c^\{\\prime\}\}=\\\{\\bm\{x\}\_\{k\}:c\\leq V\(\\bm\{x\}\_\{k\}\)\\leq c^\{\\prime\}\\\}\. There exists a constantτ\+\>0\\tau^\{\+\}\>0, independent of𝐱k\\bm\{x\}\_\{k\}, such thatΔV\(τ\)<0\\Delta V\(\\tau\)<0for allτ∈\(0,τ\+\)\\tau\\in\(0,\\tau^\{\+\}\)and all𝐱k∈𝒜c,c′\\bm\{x\}\_\{k\}\\in\\mathcal\{A\}\_\{c,c^\{\\prime\}\}\. In particular,τmin\\tau\_\{\\min\}serves as such aτ\+\\tau^\{\+\}, verified numerically viaMdisc≻0M\_\{\\mathrm\{disc\}\}\\succ 0\(Table[4](https://arxiv.org/html/2605.12561#A1.T4)\)\.
###### Proof\.
LetMQ≜Q\+K⊤RK≻0M\_\{Q\}\\triangleq Q\+K^\{\\top\}RK\\succ 0\(distinct fromM\(τ\)M\(\\tau\)in Eq\. \([4](https://arxiv.org/html/2605.12561#S2.E4)\)\); from Proposition[2](https://arxiv.org/html/2605.12561#Thmproposition2)’s derivative computation,dΔV/dτ\|τ=0=−𝒙k⊤MQ𝒙kd\\Delta V/d\\tau\|\_\{\\tau=0\}=\-\\bm\{x\}\_\{k\}^\{\\top\}M\_\{Q\}\\bm\{x\}\_\{k\}\. Onc≤V\(𝒙k\)≤c′c\\leq V\(\\bm\{x\}\_\{k\}\)\\leq c^\{\\prime\},‖𝒙k‖2≥c/λmax\(P\)\\\|\\bm\{x\}\_\{k\}\\\|^\{2\}\\geq c/\\lambda\_\{\\max\}\(P\), so𝒙k⊤MQ𝒙k≥λmin\(MQ\)c/λmax\(P\)\>0\\bm\{x\}\_\{k\}^\{\\top\}M\_\{Q\}\\bm\{x\}\_\{k\}\\geq\\lambda\_\{\\min\}\(M\_\{Q\}\)c/\\lambda\_\{\\max\}\(P\)\>0uniformly\. SinceΔV\(0\)=0\\Delta V\(0\)=0anddΔV/dτ\|τ=0≤−λmin\(MQ\)c/λmax\(P\)<0d\\Delta V/d\\tau\|\_\{\\tau=0\}\\leq\-\\lambda\_\{\\min\}\(M\_\{Q\}\)c/\\lambda\_\{\\max\}\(P\)<0uniformly, joint continuity ofΔV\(τ,𝒙k\)\\Delta V\(\\tau,\\bm\{x\}\_\{k\}\)implies the existence of a uniformτ\+\>0\\tau^\{\+\}\>0\. ∎
### A\.4Corollary[A\.4](https://arxiv.org/html/2605.12561#A1.SS4): Nonlinear Admissible Region
###### Corollary 2\.
Letδ\(𝐱\)=f\(𝐱\)−\(A𝐱\+B𝐮k\)\\delta\(\\bm\{x\}\)=f\(\\bm\{x\}\)\-\(A\\bm\{x\}\+B\\bm\{u\}\_\{k\}\)denote the nonlinear residual\. Sinceδ\(𝟎\)=0\\delta\(\\bm\{0\}\)=0and∇δ\(𝟎\)=0\\nabla\\delta\(\\bm\{0\}\)=0, there existsLδ\>0L\_\{\\delta\}\>0such that‖δ\(𝐱\)‖≤Lδ‖𝐱‖2\\\|\\delta\(\\bm\{x\}\)\\\|\\leq L\_\{\\delta\}\\\|\\bm\{x\}\\\|^\{2\}near the origin\. WithMQ≜Q\+K⊤RK≻0M\_\{Q\}\\triangleq Q\+K^\{\\top\}RK\\succ 0as in the proof of Corollary[A\.3](https://arxiv.org/html/2605.12561#A1.SS3), the true Lyapunov derivative satisfies
V˙\(𝒙\)≤−λmin\(MQ\)‖𝒙‖2\+2λmax\(P\)Lδ‖𝒙‖3\.\\dot\{V\}\(\\bm\{x\}\)\\leq\-\\lambda\_\{\\min\}\(M\_\{Q\}\)\\\|\\bm\{x\}\\\|^\{2\}\+2\\lambda\_\{\\max\}\(P\)L\_\{\\delta\}\\\|\\bm\{x\}\\\|^\{3\}\.HenceV˙\(𝐱\)<0\\dot\{V\}\(\\bm\{x\}\)<0for all𝐱≠𝟎\\bm\{x\}\\neq\\bm\{0\}with‖𝐱‖<r∗≜λmin\(MQ\)/\[2λmax\(P\)Lδ\]\\\|\\bm\{x\}\\\|<r^\{\*\}\\triangleq\\lambda\_\{\\min\}\(M\_\{Q\}\)/\[2\\lambda\_\{\\max\}\(P\)L\_\{\\delta\}\]\. Withinℬ\(r∗\)\\mathcal\{B\}\(r^\{\*\}\), LaSalle’s invariance principle implies asymptotic stability of the origin\.
###### Proof\.
Write𝒙˙=Acl𝒙\+δ\(𝒙\)\\dot\{\\bm\{x\}\}=A\_\{\\mathrm\{cl\}\}\\bm\{x\}\+\\delta\(\\bm\{x\}\)\. ThenV˙\(𝒙\)=𝒙⊤\(Acl⊤P\+PAcl\)𝒙\+2𝒙⊤Pδ\(𝒙\)≤−𝒙⊤MQ𝒙\+2‖𝒙‖λmax\(P\)‖δ\(𝒙\)‖\\dot\{V\}\(\\bm\{x\}\)=\\bm\{x\}^\{\\top\}\(A\_\{\\mathrm\{cl\}\}^\{\\top\}P\+PA\_\{\\mathrm\{cl\}\}\)\\bm\{x\}\+2\\bm\{x\}^\{\\top\}P\\delta\(\\bm\{x\}\)\\leq\-\\bm\{x\}^\{\\top\}M\_\{Q\}\\bm\{x\}\+2\\\|\\bm\{x\}\\\|\\lambda\_\{\\max\}\(P\)\\\|\\delta\(\\bm\{x\}\)\\\|\. Using𝒙⊤MQ𝒙≥λmin\(MQ\)‖𝒙‖2\\bm\{x\}^\{\\top\}M\_\{Q\}\\bm\{x\}\\geq\\lambda\_\{\\min\}\(M\_\{Q\}\)\\\|\\bm\{x\}\\\|^\{2\}and‖δ\(𝒙\)‖≤Lδ‖𝒙‖2\\\|\\delta\(\\bm\{x\}\)\\\|\\leq L\_\{\\delta\}\\\|\\bm\{x\}\\\|^\{2\}yields the stated bound\. Setting the right\-hand side negative yieldsr∗r^\{\*\}\. ∎
Table 4:Theoretical quantities per environment\.MQ≜Q\+K⊤RKM\_\{Q\}\\triangleq Q\+K^\{\\top\}RKis the CARE\-quadratic matrix \(distinct from the ZOH transitionM\(τ\)M\(\\tau\)in Eq\. \([4](https://arxiv.org/html/2605.12561#S2.E4)\)\)\.r∗r^\{\*\}is a conservative lower bound on the true region of attraction \(Corollary[A\.4](https://arxiv.org/html/2605.12561#A1.SS4)\)\. The marginθsat−θRTA\\theta\_\{\\mathrm\{sat\}\}\-\\theta\_\{\\mathrm\{RTA\}\}keeps the angle channel of the backup unsaturated\.λmin\(Mdisc\)\\lambda\_\{\\min\}\(M\_\{\\mathrm\{disc\}\}\)is the minimum eigenvalue ofMdisc=P−M\(τmin\)⊤PM\(τmin\)M\_\{\\mathrm\{disc\}\}=P\-M\(\\tau\_\{\\min\}\)^\{\\top\}P\\,M\(\\tau\_\{\\min\}\); positivity verifies Proposition[1](https://arxiv.org/html/2605.12561#Thmproposition1)numerically\. “–” indicates that the ZOH\-discretized backup atτmin\\tau\_\{\\min\}is not Lyapunov\-decreasing \(see Section[6](https://arxiv.org/html/2605.12561#S6)\)\. Angles in degrees;r∗r^\{\*\}in‖𝒙‖\\\|\\bm\{x\}\\\|\.
## Appendix BEnvironment Details
All environments integrate the nonlinear plant atΔt=0\.001s\\Delta t=0\.001\\,\\mathrm\{s\}\.
Table 5:Environment parameters \(ordered by complexity\)\.### B\.1Pendulum \(Gymnasium Pendulum\-v1 Physics\)
State𝒙=\[θ,θ˙\]⊤\\bm\{x\}=\[\\theta,\\dot\{\\theta\}\]^\{\\top\},θ=0\\theta=0at upright equilibrium\. Equation of motion:θ¨=\(3g/2l\)sinθ\+\(3/ml2\)u\\ddot\{\\theta\}=\(3g/2l\)\\sin\\theta\+\(3/ml^\{2\}\)u,m=1kgm=1\\,\\mathrm\{kg\},l=1ml=1\\,\\mathrm\{m\},g=10m/s2g=10\\,\\mathrm\{m/s\}^\{2\},\|u\|≤2Nm\|u\|\\leq 2\\,\\mathrm\{Nm\},\|θ˙\|≤8rad/s\|\\dot\{\\theta\}\|\\leq 8\\,\\mathrm\{rad/s\}\. Linearization aroundθ=0\\theta=0:A=\[0,1;15,0\]A=\[0,1;\\,15,0\],B=\[0;3\]B=\[0;\\,3\]\. Design matricesQ=diag\(10,1\)Q=\\operatorname\{diag\}\(10,1\),R=1R=1yieldK≈\[10\.92,2\.88\]K\\approx\[10\.92,\\,2\.88\],λ≈6\.23\\lambda\\approx 6\.23,Vscale≈8\.99V\_\{\\mathrm\{scale\}\}\\approx 8\.99\. RTA thresholdθRTA=0\.15rad\(≈8\.6∘\)\\theta\_\{\\mathrm\{RTA\}\}=0\.15\\,\\mathrm\{rad\}\\;\(\\approx 8\.6^\{\\circ\}\), saturation angleθsat≈10\.5∘\\theta\_\{\\mathrm\{sat\}\}\\approx 10\.5^\{\\circ\}\. Episodes terminate at\|θ\|\>60∘\|\\theta\|\>60^\{\\circ\}ort≥50st\\geq 50\\,\\mathrm\{s\}\. Initial conditions:θ0∈\[−0\.1,0\.1\]rad\\theta\_\{0\}\\in\[\-0\.1,0\.1\]\\,\\mathrm\{rad\},θ˙0∈\[−0\.5,0\.5\]rad/s\\dot\{\\theta\}\_\{0\}\\in\[\-0\.5,0\.5\]\\,\\mathrm\{rad/s\}\. Action set:\|𝒯\|=8\|\\mathcal\{T\}\|=8,\|𝒰\|=21\|\\mathcal\{U\}\|=21,\|𝒜\|=168\|\\mathcal\{A\}\|=168\.
### B\.2CartPole \(Gymnasium CartPole\-v1 Physics\)
State𝒙=\[x,x˙,θ,θ˙\]⊤\\bm\{x\}=\[x,\\dot\{x\},\\theta,\\dot\{\\theta\}\]^\{\\top\}\. Equations of motion:
θ¨\\displaystyle\\ddot\{\\theta\}=gsinθ−cosθ\(F\+mpLθ˙2sinθ\)/mtL\(4/3−mpcos2θ/mt\),x¨=F\+mpL\(θ˙2sinθ−θ¨cosθ\)mt,\\displaystyle=\\frac\{g\\sin\\theta\-\\cos\\theta\(F\+m\_\{p\}L\\dot\{\\theta\}^\{2\}\\sin\\theta\)/m\_\{t\}\}\{L\(4/3\-m\_\{p\}\\cos^\{2\}\\\!\\theta/m\_\{t\}\)\},\\quad\\ddot\{x\}=\\frac\{F\+m\_\{p\}L\(\\dot\{\\theta\}^\{2\}\\sin\\theta\-\\ddot\{\\theta\}\\cos\\theta\)\}\{m\_\{t\}\},mc=1\.0kgm\_\{c\}=1\.0\\,\\mathrm\{kg\},mp=0\.1kgm\_\{p\}=0\.1\\,\\mathrm\{kg\},mt=1\.1kgm\_\{t\}=1\.1\\,\\mathrm\{kg\},L=0\.5mL=0\.5\\,\\mathrm\{m\},g=9\.8m/s2g=9\.8\\,\\mathrm\{m/s\}^\{2\},\|F\|≤20N\|F\|\\leq 20\\,\\mathrm\{N\}\. Letd=L\(4/3−mp/mt\)d=L\(4/3\-m\_\{p\}/m\_\{t\}\); linearizing around𝒙⋆=𝟎\\bm\{x\}^\{\\star\}=\\bm\{0\}:
A=\[010000−mpLg/\(mtd\)0000100g/d0\],B=\[01/mt\+mpL/\(mt2d\)0−1/\(mtd\)\]\.A=\\begin\{bmatrix\}0&1&0&0\\\\ 0&0&\-m\_\{p\}Lg/\(m\_\{t\}d\)&0\\\\ 0&0&0&1\\\\ 0&0&g/d&0\\end\{bmatrix\},\\;B=\\begin\{bmatrix\}0\\\\ 1/m\_\{t\}\+m\_\{p\}L/\(m\_\{t\}^\{2\}d\)\\\\ 0\\\\ \-1/\(m\_\{t\}d\)\\end\{bmatrix\}\.Design matricesQ=diag\(6,1,11\.5,5\)Q=\\operatorname\{diag\}\(6,1,11\.5,5\),R=1R=1\. Saturation angleθsat≈29\.9∘\\theta\_\{\\mathrm\{sat\}\}\\approx 29\.9^\{\\circ\}; RTA thresholdθRTA=12∘\\theta\_\{\\mathrm\{RTA\}\}=12^\{\\circ\}; additional trigger\|xk\|≥1\.92m\|x\_\{k\}\|\\geq 1\.92\\,\\mathrm\{m\}\. Episodes terminate at\|θ\|\>12∘\|\\theta\|\>12^\{\\circ\},\|x\|\>2\.4m\|x\|\>2\.4\\,\\mathrm\{m\}, ort≥50st\\geq 50\\,\\mathrm\{s\}\[[26](https://arxiv.org/html/2605.12561#bib.bib26)\]\. Action set:\|𝒯\|=8\|\\mathcal\{T\}\|=8,\|𝒰\|=41\|\\mathcal\{U\}\|=41,\|𝒜\|=328\|\\mathcal\{A\}\|=328\.
### B\.3Planar Quadrotor \(Hover Stabilization\)
State𝒙=\[x,z,θ,x˙,z˙,θ˙\]⊤\\bm\{x\}=\[x,z,\\theta,\\dot\{x\},\\dot\{z\},\\dot\{\\theta\}\]^\{\\top\}, inputs𝒖=\[δF,M\]⊤\\bm\{u\}=\[\\delta F,M\]^\{\\top\}\(thrust deviation, pitching moment\):x¨=−\(F/m\)sinθ\\ddot\{x\}=\-\(F/m\)\\sin\\theta,z¨=\(F/m\)cosθ−g\\ddot\{z\}=\(F/m\)\\cos\\theta\-g,θ¨=M/I\\ddot\{\\theta\}=M/I,F=mg\+δFF=mg\+\\delta F,m=1kgm=1\\,\\mathrm\{kg\},I=0\.05kgm2I=0\.05\\,\\mathrm\{kg\\,m\}^\{2\},g=9\.81m/s2g=9\.81\\,\\mathrm\{m/s\}^\{2\},\|δF\|≤5N\|\\delta F\|\\leq 5\\,\\mathrm\{N\},\|M\|≤1Nm\|M\|\\leq 1\\,\\mathrm\{Nm\}\. Linearization around hover:
A=\[00010000001000000100−g000000000000000\],B=\[000000001/m001/I\]\.A=\\begin\{bmatrix\}0&0&0&1&0&0\\\\ 0&0&0&0&1&0\\\\ 0&0&0&0&0&1\\\\ 0&0&\{\-g\}&0&0&0\\\\ 0&0&0&0&0&0\\\\ 0&0&0&0&0&0\\end\{bmatrix\},\\quad B=\\begin\{bmatrix\}0&0\\\\ 0&0\\\\ 0&0\\\\ 0&0\\\\ 1/m&0\\\\ 0&1/I\\end\{bmatrix\}\.Design matricesQ=diag\(2,2,10,1,1,5\)Q=\\operatorname\{diag\}\(2,2,10,1,1,5\),R=diag\(0\.1,5\)R=\\operatorname\{diag\}\(0\.1,5\)\. Moment saturation angleθsat≈11\.96∘\\theta\_\{\\mathrm\{sat\}\}\\approx 11\.96^\{\\circ\}; RTA thresholdθRTA≈9\.57∘\\theta\_\{\\mathrm\{RTA\}\}\\approx 9\.57^\{\\circ\}\(80%80\\%ofθsat\\theta\_\{\\mathrm\{sat\}\}\); additional trigger\|xk\|,\|zk\|≥2\.0m\|x\_\{k\}\|,\|z\_\{k\}\|\\geq 2\.0\\,\\mathrm\{m\}\. Episodes terminate at\|θ\|\>30∘\|\\theta\|\>30^\{\\circ\},\|x\|\|x\|or\|z\|\>2\.5m\|z\|\>2\.5\\,\\mathrm\{m\}, ort≥50st\\geq 50\\,\\mathrm\{s\}\. Initial conditions:x0,z0∈\[−0\.3,0\.3\]mx\_\{0\},z\_\{0\}\\in\[\-0\.3,0\.3\]\\,\\mathrm\{m\},θ0∈\[−0\.1,0\.1\]rad\\theta\_\{0\}\\in\[\-0\.1,0\.1\]\\,\\mathrm\{rad\},x˙0,z˙0,θ˙0∈\[−0\.3,0\.3\]\\dot\{x\}\_\{0\},\\dot\{z\}\_\{0\},\\dot\{\\theta\}\_\{0\}\\in\[\-0\.3,0\.3\]\. Action set:\|𝒯\|=8\|\\mathcal\{T\}\|=8,\|𝒰δF\|=11\|\\mathcal\{U\}\_\{\\delta F\}\|=11,\|𝒰M\|=9\|\\mathcal\{U\}\_\{M\}\|=9,\|𝒜\|=792\|\\mathcal\{A\}\|=792\.
### B\.4Quadrotor3D \(6\-DOF Hover Stabilization\)
State𝒙=\[px,py,pz,φ,θ,ψ,vx,vy,vz,p,q,r\]⊤∈ℝ12\\bm\{x\}=\[p\_\{x\},p\_\{y\},p\_\{z\},\\varphi,\\theta,\\psi,v\_\{x\},v\_\{y\},v\_\{z\},p,q,r\]^\{\\top\}\\in\\mathbb\{R\}^\{12\}, inputs𝒖=\[δF,τφ,τθ,τψ\]⊤∈ℝ4\\bm\{u\}=\[\\delta F,\\tau\_\{\\varphi\},\\tau\_\{\\theta\},\\tau\_\{\\psi\}\]^\{\\top\}\\in\\mathbb\{R\}^\{4\}, whereδF=F−mg\\delta F=F\-mgis the thrust deviation from hover and\(τφ,τθ,τψ\)\(\\tau\_\{\\varphi\},\\tau\_\{\\theta\},\\tau\_\{\\psi\}\)are body\-axis torques\. The nonlinear equations of motion are the standard rigid\-body model𝒑˙inert=R\(φ,θ,ψ\)𝒗body\\dot\{\\bm\{p\}\}\_\{\\mathrm\{inert\}\}=R\(\\varphi,\\theta,\\psi\)\\bm\{v\}\_\{\\mathrm\{body\}\},𝚯˙=W\(φ,θ\)𝝎body\\dot\{\\bm\{\\Theta\}\}=W\(\\varphi,\\theta\)\\bm\{\\omega\}\_\{\\mathrm\{body\}\},𝒗˙body=R⊤𝒈\+\(F/m\)𝒆z−𝝎×𝒗\\dot\{\\bm\{v\}\}\_\{\\mathrm\{body\}\}=R^\{\\top\}\\bm\{g\}\+\(F/m\)\\bm\{e\}\_\{z\}\-\\bm\{\\omega\}\\\!\\times\\\!\\bm\{v\},I𝝎˙=𝝉−𝝎×I𝝎I\\dot\{\\bm\{\\omega\}\}=\\bm\{\\tau\}\-\\bm\{\\omega\}\\\!\\times\\\!I\\bm\{\\omega\}, withR\(⋅\)R\(\\cdot\)the body\-to\-inertial rotation matrix,W\(⋅\)W\(\\cdot\)mapping body angular rates to Euler\-angle rates, andI=diag\(Ixx,Iyy,Izz\)I=\\operatorname\{diag\}\(I\_\{xx\},I\_\{yy\},I\_\{zz\}\)\. Physical parameters:m=1kgm=1\\,\\mathrm\{kg\},Ixx=Iyy=0\.02kgm2I\_\{xx\}=I\_\{yy\}=0\.02\\,\\mathrm\{kg\\,m\}^\{2\},Izz=0\.04kgm2I\_\{zz\}=0\.04\\,\\mathrm\{kg\\,m\}^\{2\},g=9\.81m/s2g=9\.81\\,\\mathrm\{m/s\}^\{2\}; integrated with RK4 atdt=0\.001s\\mathrm\{d\}t=0\.001\\,\\mathrm\{s\}\. Actuator limits:\|δF\|≤5N\|\\delta F\|\\leq 5\\,\\mathrm\{N\},\|τφ\|,\|τθ\|≤1Nm\|\\tau\_\{\\varphi\}\|,\|\\tau\_\{\\theta\}\|\\leq 1\\,\\mathrm\{Nm\},\|τψ\|≤0\.5Nm\|\\tau\_\{\\psi\}\|\\leq 0\.5\\,\\mathrm\{Nm\}\. Linearizing around hover gives anAAmatrix in which translational and rotational channels decouple, with horizontal positions controlled indirectly via tilt \(v˙x≈gθ\\dot\{v\}\_\{x\}\\approx g\\theta,v˙y≈−gφ\\dot\{v\}\_\{y\}\\approx\-g\\varphi\)\. Design matricesQ=diag\(2,2,2,10,10,1,1,1,1,5,5,1\)Q=\\operatorname\{diag\}\(2,2,2,10,10,1,1,1,1,5,5,1\)andR=diag\(0\.1,5,5,10\)R=\\operatorname\{diag\}\(0\.1,5,5,10\)yield\|Kτφ,φ\|≈4\.55Nm/rad\|K\_\{\\tau\_\{\\varphi\},\\varphi\}\|\\approx 4\.55\\,\\mathrm\{Nm/rad\}, givingθsat≈12\.60∘\\theta\_\{\\mathrm\{sat\}\}\\approx 12\.60^\{\\circ\}and the RTA thresholdθRTA≈10\.08∘\\theta\_\{\\mathrm\{RTA\}\}\\approx 10\.08^\{\\circ\}\(80%80\\%ofθsat\\theta\_\{\\mathrm\{sat\}\}\); the one\-step\-ahead predicate is evaluated independently forφ^k\+1\\hat\{\\varphi\}\_\{k\+1\}andθ^k\+1\\hat\{\\theta\}\_\{k\+1\}\. Additional trigger:\|px\|,\|py\|,\|pz\|≥2\.0m\|p\_\{x\}\|,\|p\_\{y\}\|,\|p\_\{z\}\|\\geq 2\.0\\,\\mathrm\{m\}\. Episodes terminate at\|φ\|\|\\varphi\|or\|θ\|\>30∘\|\\theta\|\>30^\{\\circ\},\|ψ\|\>90∘\|\\psi\|\>90^\{\\circ\},\|px\|,\|py\|,\|pz\|\>2\.5m\|p\_\{x\}\|,\|p\_\{y\}\|,\|p\_\{z\}\|\>2\.5\\,\\mathrm\{m\}, ort≥50st\\geq 50\\,\\mathrm\{s\}\. Initial conditions:px,0,py,0,pz,0∈\[−0\.3,0\.3\]mp\_\{x,0\},p\_\{y,0\},p\_\{z,0\}\\in\[\-0\.3,0\.3\]\\,\\mathrm\{m\},φ0,θ0,ψ0∈\[−0\.1,0\.1\]rad\\varphi\_\{0\},\\theta\_\{0\},\\psi\_\{0\}\\in\[\-0\.1,0\.1\]\\,\\mathrm\{rad\},vx,0,vy,0,vz,0∈\[−0\.3,0\.3\]m/sv\_\{x,0\},v\_\{y,0\},v\_\{z,0\}\\in\[\-0\.3,0\.3\]\\,\\mathrm\{m/s\},p0,q0,r0∈\[−0\.1,0\.1\]rad/sp\_\{0\},q\_\{0\},r\_\{0\}\\in\[\-0\.1,0\.1\]\\,\\mathrm\{rad/s\}\. Action set:\|𝒯\|=8\|\\mathcal\{T\}\|=8,\|𝒰δF\|=\|𝒰τφ\|=\|𝒰τθ\|=\|𝒰τψ\|=5\|\\mathcal\{U\}\_\{\\delta F\}\|=\|\\mathcal\{U\}\_\{\\tau\_\{\\varphi\}\}\|=\|\\mathcal\{U\}\_\{\\tau\_\{\\theta\}\}\|=\|\\mathcal\{U\}\_\{\\tau\_\{\\psi\}\}\|=5, giving\|𝒜\|=5,000\|\\mathcal\{A\}\|=5\{,\}000for the discrete \(DQN\) variant; the SAC variant adopted in Section[6](https://arxiv.org/html/2605.12561#S6)uses the equivalent continuous Box action space\.
## Appendix CHyperparameters
Table 6:DQN hyperparameters \(identical across all environments\)\.Table 7:SAC hyperparameters \(identical across all lower\-dimensional environments; Quadrotor3D uses a wider256→128→128256\{\\to\}128\{\\to\}128network and2M2\\,\\mathrm\{M\}training steps as discussed in Section[6](https://arxiv.org/html/2605.12561#S6)\)\.
## Appendix DCompute Resources
DQN and SAC training budgets are reported in Tables[6](https://arxiv.org/html/2605.12561#A3.T6)and[7](https://arxiv.org/html/2605.12561#A3.T7):1M1\\,\\mathrm\{M\}steps for DQN/SAC on the lower\-dimensional environments,2M2\\,\\mathrm\{M\}steps for preference\-conditioned DQN,2M2\\,\\mathrm\{M\}steps for SAC and PPO on Quadrotor3D, and4M4\\,\\mathrm\{M\}steps for preference\-conditioned SAC on Quadrotor3D\. All experiments were run on a single workstation \(Intel Core i9\-13900KF, 24 cores / 32 threads; NVIDIA RTX 4090, 24 GB; 64 GB RAM; Ubuntu 22\.04\) via Stable\-Baselines3\. Wall\-clock times for an 11\-pointwcw\_\{c\}sweep run as 11 parallel processes are approximately11hour for the lower\-dimensional DQN sweeps \(1M1\\,\\mathrm\{M\}steps each\),66–77hours for the lower\-dimensional SAC sweeps \(1M1\\,\\mathrm\{M\}steps each\), and1313–1515hours for the Quadrotor3D SAC sweep \(2M2\\,\\mathrm\{M\}steps each\)\.
## Appendix EMulti\-Seed Evaluation Methodology
For the lower\-dimensional DQNwcw\_\{c\}sweeps \(Tables[10](https://arxiv.org/html/2605.12561#A8.T10),[11](https://arxiv.org/html/2605.12561#A8.T11),[12](https://arxiv.org/html/2605.12561#A8.T12)\) and their corresponding No\-RTA ablation \(Table[3](https://arxiv.org/html/2605.12561#S5.T3)\) and Lagrangian\-DQN comparison \(Table[9](https://arxiv.org/html/2605.12561#A7.T9)\), we report results aggregated across the33independent training seeds\{0,1,2\}\\\{0,1,2\\\}\. Note that seed 0 is included in this set, so the seed\-0 results referenced for single\-seed experiments correspond to one of the three seeds in the multi\-seed aggregates\. For each seed we run the standardNeval=100N\_\{\\mathrm\{eval\}\}=100deterministic evaluation; the per\-seed mean is the average over those100100episodes\. MSI and RTA\-rate columns are then reported as the mean±\\pmstandard deviation across the33per\-seed mean values, usingddof=1\\mathrm\{ddof\}=1\(unbiased estimator\)\. ThePiP\_\{i\}state\-norm columns report the mean across the33per\-seed means only \(no std\), to keep the tables compact\. All other experiments in the paper \(SAC sweeps, Preference\-Conditioned DQN, baselines, Ablation B, robustness studies, and the Quadrotor3D case study\) use the standard single\-seed evaluation protocol onseed 0, with std columns reflecting per\-episode variance across the100100evaluation episodes\. Figures and tables that report RL\-STC results without an explicit±\\pmacross\-seed annotation are accordingly seed\-0 results unless labeled otherwise; we call this out per\-table where relevant\.
## Appendix FAblation B: Fixed\-τ\\tauComparison
τ\\tauis fixed at the nearest grid value toτmatch\\tau\_\{\\mathrm\{match\}\}\(the RL\-STC best\-checkpoint MSI for each environment\); the agent optimizes only the control input\. The RTA shield remains active\.τmatch\\tau\_\{\\mathrm\{match\}\}is taken from the seed\-0 best\-checkpoint MSI, which sets the nearest grid value used for the fixed\-τ\\taucomparison; seed 0 is one of the three seeds in the multi\-seed evaluation \(Appendix[E](https://arxiv.org/html/2605.12561#A5)\), and its grid value is unchanged under multi\-seed re\-evaluation since the multi\-seed mean differs from seed 0 by less than the grid spacing\. The comparison below is reported on seed 0 throughout, keeping fixed\-τ\\tauand RL\-STC at the same seed\.
ForPendulum, fixingτ=0\.40s\\tau=0\.40\\,\\mathrm\{s\}triggers RTA on80\.6%80\.6\\%of steps, the shield resetsτ←τmin\\tau\\leftarrow\\tau\_\{\\min\}on the majority of steps, degenerating the effective policy to LQR atτmin\\tau\_\{\\min\}regardless of what the RL agent learns\. This is the clearest demonstration that joint optimization of\(𝒖k,τk\)\(\\bm\{u\}\_\{k\},\\tau\_\{k\}\)is not separable: it is precisely the learned adaptive inter\-sample interval that allows the full policy to sustain an MSI of0\.396s0\.396\\,\\mathrm\{s\}\. ForCartPole,τmatch=τmax\\tau\_\{\\mathrm\{match\}\}=\\tau\_\{\\max\}, so fixed\-τ\\tauachieves essentially the same MSI as RL\-STC \(0\.3200\.320vs\.0\.317s0\.317\\,\\mathrm\{s\}\) with no RTA, but this presupposes the stability/communication tradeoff point that RL\-STC had to discover; for a plant where this does not coincide withτmax\\tau\_\{\\max\}, RL\-STC eliminates the search by incorporating interval selection into optimization\. ForQuadrotor, adaptive timing recovers a4\.3%4\.3\\%MSI advantage \(0\.2900\.290vs\.0\.278s0\.278\\,\\mathrm\{s\}\) with lowerP1P\_\{1\}; the higherP4P\_\{4\}\(5\.6645\.664vs\.4\.3424\.342\) reflects angular velocity excursion during higher\-MSI hover corrections, not reduced stability\.
Table 8:Ablation B: Fixedτ\\tauvs\. full method \(best\-model checkpoint,wc=8/16/16w\_\{c\}=8/16/16, seed 0\)\.
## Appendix GLagrangian\-DQN Detailed Results
The Lagrangian multiplierλt\\lambda\_\{t\}is updated via projected dual gradient ascentλt\+1=clip\(λt\+αλ\(gt−ϵ\),0,λmax\)\\lambda\_\{t\+1\}=\\mathrm\{clip\}\(\\lambda\_\{t\}\+\\alpha\_\{\\lambda\}\(g\_\{t\}\-\\epsilon\),0,\\lambda\_\{\\max\}\)with budgetϵ=0\.01\\epsilon=0\.01,αλ=0\.01\\alpha\_\{\\lambda\}=0\.01\. The critical difference from RL\-STC is what happens when the predicate fires: RL\-STC overrides the action so Hard Viol\.=0\.0=0\.0by construction; Lagrangian\-DQN applies only a soft penalty, so predicted events propagate into hard violations at rates that vary with plant complexity \(Table[9](https://arxiv.org/html/2605.12561#A7.T9)\)\. ForPendulum, Lagrangian achieves21%21\\%lower MSI \(0\.3040\.304vs\.0\.385s0\.385\\,\\mathrm\{s\}\) with7\.95%7\.95\\%hard violations, and the final model deteriorates to43\.49±48\.55%43\.49\\pm 48\.55\\%hard violations asλ\\lambdafails to hold across seeds\. ForCartPole, near\-zero best\-checkpoint hard violations come at52%52\\%lower MSI \(0\.1490\.149vs\.0\.308s0\.308\\,\\mathrm\{s\}\); without a hard override, the policy must select shorter intervals to avoid violations, and final\-model episodes terminate early at high rates, a failure invisible to the expectation\-level constraint signal\. ForQuadrotor,33%33\\%lower MSI \(0\.1880\.188vs\.0\.281s0\.281\\,\\mathrm\{s\}\) with35\.40%35\.40\\%hard violations, deteriorating to59\.35%59\.35\\%in the final model\.
Table 9:Lagrangian\-DQN vs\. RL\-STC \(Neval=100N\_\{\\mathrm\{eval\}\}=100, 3 seeds per row;wc=8/16/16w\_\{c\}=8/16/16for Pendulum/CartPole/Quadrotor\)\. MSI and the safety\-event columns are reported as mean±\\pmstd across the 3 per\-seed mean values \(ddof=1\\mathrm\{ddof\}=1\)\. “Pred\. Safety \(%\)” measures the same one\-step\-ahead predicate \(\|q^k\+1\|\>θRTA\|\\hat\{q\}\_\{k\+1\}\|\>\\theta\_\{\\mathrm\{RTA\}\}\) for both methods: for RL\-STC it is the RTA activation rate \(predicate fires, action overridden, state violation prevented\); for Lagrangian\-DQN it is the constraint\-violation rate \(predicate fires, soft penalty only, no override\)\. “Hard Viol\. \(%\)” = fraction of executed timesteps where\|θk\|\>θRTA\|\\theta\_\{k\}\|\>\\theta\_\{\\mathrm\{RTA\}\}in the actual trajectory\. RL\-STC Hard Viol\.=0\.0=0\.0by construction of the pointwise override\.
## Appendix HFull Sweep Tables
DQN sweeps \(Tables[10](https://arxiv.org/html/2605.12561#A8.T10)–[12](https://arxiv.org/html/2605.12561#A8.T12)\) are 3\-seed aggregates per Appendix[E](https://arxiv.org/html/2605.12561#A5); all other sweeps in this section \(SAC, Pref\-DQN, PPO, and Quadrotor3D\) are single\-seed \(seed 0\), with±\\pmvalues in those tables reflecting per\-episode variance across the100100evaluation episodes\.
### H\.1DQNwcw\_\{c\}Sweep
Table 10:Pendulum – DQNwcw\_\{c\}sweep \(Neval=100N\_\{\\mathrm\{eval\}\}=100, 3 seeds per row\)\. MSI and RTA are reported as mean±\\pmstd across the 3 per\-seed mean values \(ddof=1\\mathrm\{ddof\}=1\);PiP\_\{i\}norms are mean only\. Bold marks the highest MSI in each section\.Table 11:CartPole – DQNwcw\_\{c\}sweep \(Neval=100N\_\{\\mathrm\{eval\}\}=100, 3 seeds per row\)\. MSI and RTA are reported as mean±\\pmstd across the 3 per\-seed mean values \(ddof=1\\mathrm\{ddof\}=1\);PiP\_\{i\}norms are mean only\. Bold marks the highest MSI in each section\.Table 12:Quadrotor – DQNwcw\_\{c\}sweep \(Neval=100N\_\{\\mathrm\{eval\}\}=100, 3 seeds per row\)\. MSI and RTA are reported as mean±\\pmstd across the 3 per\-seed mean values \(ddof=1\\mathrm\{ddof\}=1\);PiP\_\{i\}norms are mean only\. Bold marks the highest MSI in each section\.
### H\.2SACwcw\_\{c\}Sweep
SAC achieves0%0\\%RTA across every weight and environment in both the final model and best checkpoint, with no high\-wcw\_\{c\}collapses\. For the Quadrotor, the SAC best checkpoint atwc=16w\_\{c\}=16reaches0\.312s0\.312\\,\\mathrm\{s\}versus DQN’s multi\-seed best of0\.281±0\.018s0\.281\\pm 0\.018\\,\\mathrm\{s\}, while maintaining0%0\\%RTA\. Across all three environments, the RTA shield and Lyapunov reward remain effective under a fundamentally different optimization algorithm, confirming that communication efficiency gains are a property of the framework, not of DQN\.
Table 13:SAC – Pendulumwcw\_\{c\}sweep \(Neval=100N\_\{\\mathrm\{eval\}\}=100, 1 M steps\)\.Table 14:SAC – CartPolewcw\_\{c\}sweep \(Neval=100N\_\{\\mathrm\{eval\}\}=100, 1 M steps\)\.Table 15:SAC – Quadrotorwcw\_\{c\}sweep \(Neval=100N\_\{\\mathrm\{eval\}\}=100, 1 M steps\)\.
### H\.3Preference\-Conditioned DQN
The preference\-conditioned DQN is trained once per environment for2M2\\,\\mathrm\{M\}steps and evaluated at the same 11\-pointwcw\_\{c\}sweep\. For Pendulum, the best\-model checkpoint achieves0\.393s0\.393\\,\\mathrm\{s\}atwc=16w\_\{c\}=16, within1%1\\%of standard DQN’s multi\-seed best \(0\.397±0\.001s0\.397\\pm 0\.001\\,\\mathrm\{s\}\) at211\\tfrac\{2\}\{11\}of compute\. For CartPole, the best\-model checkpoint plateaus near0\.3130\.313–0\.316s0\.316\\,\\mathrm\{s\}abovewc=10w\_\{c\}=10, matching or slightly exceeding standard DQN \(0\.308±0\.014s0\.308\\pm 0\.014\\,\\mathrm\{s\}\)\. For Quadrotor, the best checkpoint increases from0\.102s0\.102\\,\\mathrm\{s\}to0\.266s0\.266\\,\\mathrm\{s\}, within5%5\\%of standard DQN \(0\.281±0\.018s0\.281\\pm 0\.018\\,\\mathrm\{s\}\); the harder coupled hover dynamics remain challenging for a single shared policy at extremewcw\_\{c\}\.
Table 16:Preference\-Conditioned DQN – Pendulumwcw\_\{c\}sweep \(Neval=100N\_\{\\mathrm\{eval\}\}=100, 2 M steps, single model\)\.Table 17:Preference\-Conditioned DQN – CartPolewcw\_\{c\}sweep \(Neval=100N\_\{\\mathrm\{eval\}\}=100, 2 M steps, single model\)\.Table 18:Preference\-Conditioned DQN – Quadrotorwcw\_\{c\}sweep \(Neval=100N\_\{\\mathrm\{eval\}\}=100, 2 M steps, single model\)\.
### H\.4PPOwcw\_\{c\}Sweep \(Pendulum\)
Table[19](https://arxiv.org/html/2605.12561#A8.T19)reports the full PPO sweep referenced from Section[5\.1](https://arxiv.org/html/2605.12561#S5.SS1.SSS0.Px1)\. No PPO checkpoint achieves communication efficiency comparable to DQN \(0\.396s0\.396\\,\\mathrm\{s\}atwc=8w\_\{c\}=8\); runs that hold0%0\\%RTA collapse toτmin=0\.05s\\tau\_\{\\min\}=0\.05\\,\\mathrm\{s\}, while runs that explore longer intervals incur4545–80%80\\%RTA activation\.
Table 19:PPO – Pendulumwcw\_\{c\}sweep \(Neval=100N\_\{\\mathrm\{eval\}\}=100, 2 M steps\)\. Bold marks the highest MSI in each section\.
### H\.5Quadrotor3Dwcw\_\{c\}Sweeps
Tables[20](https://arxiv.org/html/2605.12561#A8.T20)and[21](https://arxiv.org/html/2605.12561#A8.T21)report the per\-wcw\_\{c\}SAC sweep and the preference\-conditioned SAC sweep on Quadrotor3D referenced from Section[6](https://arxiv.org/html/2605.12561#S6)\.
Table 20:SAC – Quadrotor3Dwcw\_\{c\}sweep \(Neval=100N\_\{\\mathrm\{eval\}\}=100, 2 M steps\)\. Bold marks the best row in each section\.Table 21:Preference\-Conditioned SAC – Quadrotor3Dwcw\_\{c\}sweep \(Neval=100N\_\{\\mathrm\{eval\}\}=100, 4 M steps, single model, discrete sampling\)\. Bold marks the best row in each section\.
## Appendix IQuadrotor3D Ablations, Lagrangian Comparison, and Robustness
Tables[22](https://arxiv.org/html/2605.12561#A9.T22)and[23](https://arxiv.org/html/2605.12561#A9.T23)report the Quadrotor3D baseline/ablation comparison and the Lagrangian\-SAC comparison referenced from Section[6](https://arxiv.org/html/2605.12561#S6); mass\-mismatch and disturbance results are in Tables[24](https://arxiv.org/html/2605.12561#A9.T24)–[25](https://arxiv.org/html/2605.12561#A9.T25)\. All Q3D results are single\-seed \(seed 0\) per Appendix[E](https://arxiv.org/html/2605.12561#A5), with±\\pmvalues reflecting per\-episode variance across the100100evaluation episodes\.
Table 22:Quadrotor3D baselines and Ablations A & B atwc=48w\_\{c\}=48\(Neval=100N\_\{\\mathrm\{eval\}\}=100, best\-model checkpoint\)\. RTA % is0by construction for all LQR\-based baselines\.Italics= system failure \(ep\. length<3s\{<\}3\\,\\mathrm\{s\}\); “–” = norms not meaningful for failed episodes\.Table 23:Lagrangian\-SAC vs\. RL\-STC for Quadrotor3D \(wc=48w\_\{c\}=48,Neval=100N\_\{\\mathrm\{eval\}\}=100\)\. RL\-STC Hard Viol\.=0\.0=0\.0by construction\.Table 24:Quadrotor3D mass\-mismatch robustness \(Neval=100N\_\{\\mathrm\{eval\}\}=100, best\-model checkpoint,wc=48w\_\{c\}=48\)\. DR = domain\-randomized training \(±40%\\pm 40\\%mass\)\.Table 25:Quadrotor3D disturbance robustness \(Neval=100N\_\{\\mathrm\{eval\}\}=100, best\-model checkpoint,wc=48w\_\{c\}=48\)\. Disturbance applied as additive thrust deviation \(δF\\delta F, N\)\.ConditionRL\-STC MSI \(s\)Cl\.\-STC MSI \(s\)RTA \(%\)No disturbance0\.302±0\.0010\.302\\pm 0\.0010\.040±0\.0000\.040\\pm 0\.0000\.00Constant0\.5N0\.5\\,\\mathrm\{N\}0\.302±0\.0010\.302\\pm 0\.0010\.040±0\.0000\.040\\pm 0\.0000\.001\.0N1\.0\\,\\mathrm\{N\}0\.302±0\.0010\.302\\pm 0\.0010\.040±0\.0000\.040\\pm 0\.0000\.00Periodic0\.8N0\.8\\,\\mathrm\{N\},1Hz1\\,\\mathrm\{Hz\}0\.302±0\.0010\.302\\pm 0\.0010\.040±0\.0000\.040\\pm 0\.0000\.001\.5N1\.5\\,\\mathrm\{N\},2Hz2\\,\\mathrm\{Hz\}0\.301±0\.0010\.301\\pm 0\.0010\.040±0\.0000\.040\\pm 0\.0000\.00Impulse \(p=0\.05p=0\.05, random sign\)1\.0N1\.0\\,\\mathrm\{N\}0\.302±0\.0010\.302\\pm 0\.0010\.040±0\.0000\.040\\pm 0\.0000\.002\.0N2\.0\\,\\mathrm\{N\}0\.302±0\.0010\.302\\pm 0\.0010\.040±0\.0000\.040\\pm 0\.0000\.00
## Appendix JRobustness Tables
Tables[26](https://arxiv.org/html/2605.12561#A10.T26)and[27](https://arxiv.org/html/2605.12561#A10.T27)report the full model\-mismatch and disturbance\-robustness results\. Detailed discussion of these results is in Appendix[M](https://arxiv.org/html/2605.12561#A13)\.
Table 26:Model\-mismatch robustness: mass scaled to0\.7×0\.7\\timesand1\.3×1\.3\\timesnominal \(Neval=100N\_\{\\mathrm\{eval\}\}=100, seed 0 for RL\-STC and RL\-STC\+DR\)\. DR = domain randomization \(mass fromUniform\[0\.6,1\.4\]×\\mathrm\{Uniform\}\[0\.6,1\.4\]\\timesnominal at each episode reset;KK,PP, and RTA thresholds fixed at nominal\)\. Best\-model checkpoint reported for all variants\.Italics= system failure \(mean episode length<5s\{<\}5\\,\\mathrm\{s\}\); “–” = RTA not applicable\.Table 27:Disturbance robustness: RL\-STC MSI, Classical STC MSI, and RTA activation under constant, periodic, and impulse disturbances \(Neval=100N\_\{\\mathrm\{eval\}\}=100, best\-model checkpoint, seed 0 for RL\-STC,wc=8/16/16w\_\{c\}=8/16/16\)\.KK,PP, and RTA thresholds fixed at nominal\. Pendulum: torque \(umax=2Nmu\_\{\\max\}=2\\,\\mathrm\{Nm\}\); CartPole: force \(Fmax=20NF\_\{\\max\}=20\\,\\mathrm\{N\}\); Quadrotor: thrust deviation \(δFmax=5N\\delta F\_\{\\max\}=5\\,\\mathrm\{N\}\)\. Impulse per RL step \(p=0\.05p=0\.05, random sign\)\.
## Appendix KTraining Curves
Figure[4](https://arxiv.org/html/2605.12561#A11.F4)shows DQN training curves \(MSI, RTA activation rate, and per\-step reward; see Section[3\.3](https://arxiv.org/html/2605.12561#S3.SS3)for the per\-step checkpoint rationale\) for fivewcw\_\{c\}values across the three lower\-dimensional environments\. The MSI peak followed by reward collapse at highwcw\_\{c\}directly motivates the best\-model checkpoint strategy discussed in Section[5\.2](https://arxiv.org/html/2605.12561#S5.SS2)\. SAC Quadrotor3D training curves are in Figure[5](https://arxiv.org/html/2605.12561#A11.F5); SAC exhibits substantially smoother MSI and reward trajectories than DQN with no high\-wcw\_\{c\}reward collapse, consistent with the algorithmic stability advantage that motivates the continuous\-action variant for the higher\-dimensional case study\.



Figure 4:DQN training curves for fivewcw\_\{c\}values across the three lower\-dimensional environments \(seed\-0 runs from the multi\-seed sweep\)\.\(top\)Mean inter\-sample interval; dashed line shows the Classical STC baseline\.\(middle\)RTA activation rate \(%\)\.\(bottom\)Per\-step reward\. Faint traces are raw per\-episode data; bold lines are EMA\-smoothed \(α=0\.06\\alpha=0\.06,≈17\{\\approx\}17\-episode window\)\. Higherwcw\_\{c\}accelerates exploration of the sparse\-sampling regime; the MSI peak before reward collapse motivates the best\-model checkpoint strategy\.Figure 5:SAC training curves for six representativewcw\_\{c\}values on Quadrotor3D \(2M2\\,\\mathrm\{M\}steps, seed 0\)\.\(left\)Mean inter\-sample interval; dashed line shows Classical STC pinned atτmin\\tau\_\{\\min\}\.\(center\)RTA activation rate \(%\)\.\(right\)Per\-step reward\. Faint traces are raw per\-episode data; bold lines are EMA\-smoothed \(α=0\.06\\alpha=0\.06,≈17\{\\approx\}17\-episode window\)\. The two\-phase dynamic from Section[6](https://arxiv.org/html/2605.12561#S6)is visible:wc≤16w\_\{c\}\\leq 16keeps MSI nearτmin\\tau\_\{\\min\}during training, whilewc≥40w\_\{c\}\\geq 40saturates nearτmax\\tau\_\{\\max\}at0%0\\%RTA\.
## Appendix LDetailed Per\-Environment Analysis
### L\.1DQN: Per\-Environment Discussion
#### Pendulum\.
The final\-model MSI generally increases withwcw\_\{c\}, peaking atwc=10w\_\{c\}=10\(0\.375±0\.030s0\.375\\pm 0\.030\\,\\mathrm\{s\},0\.08%0\.08\\%RTA\), with the across\-seed std reflecting training sensitivity at individual weights\. The best\-model checkpoint saturates nearτmax=0\.40s\\tau\_\{\\max\}=0\.40\\,\\mathrm\{s\}fromwc=6w\_\{c\}=6onwards withwc∈\{10,12,16\}w\_\{c\}\\in\\\{10,12,16\\\}all achieving MSI≥0\.39s\\geq 0\.39\\,\\mathrm\{s\}andwc=16w\_\{c\}=16tyingwc=10w\_\{c\}=10at the peak \(0\.397±0\.001s0\.397\\pm 0\.001\\,\\mathrm\{s\}\)\. At the canonicalwc=8w\_\{c\}=8used in ablations, the best checkpoint reaches0\.385±0\.014s0\.385\\pm 0\.014\\,\\mathrm\{s\}\(≥96%\\geq 96\\%ofτmax\\tau\_\{\\max\}\) with0\.37±0\.24%0\.37\\pm 0\.24\\%RTA\. The plateau is stable throughwc=16w\_\{c\}=16with low RTA across seeds \(≤0\.5%\\leq 0\.5\\%forwc≥6w\_\{c\}\\geq 6\), demonstrating that pendulum dynamics fully saturate the sampling capacity of theτ\\taugrid at moderatewcw\_\{c\}\.
#### CartPole\.
Two structural properties limit how high the final\-model MSI can grow\. First, CartPole’s CARE solution yields a largeVscale=56\.6V\_\{\\mathrm\{scale\}\}=56\.6\(versus8\.998\.99for the Pendulum\), because theθ\\thetadiagonal entry ofPPis approximately196196, a consequence of the high LQR cost of the indirectxx\-control path throughθ\\theta\. Typical episode Lyapunov values are therefore a small fraction ofVscaleV\_\{\\mathrm\{scale\}\}, saturating the graded stability term1−V\(𝒙k\+1\)/Vscale1\-V\(\\bm\{x\}\_\{k\+1\}\)/V\_\{\\mathrm\{scale\}\}near\+1\+1throughout training, so the communication reward becomes the dominant differentiating signal even at lowwcw\_\{c\}\. Second, the tight termination angle \(θterm=12∘\\theta\_\{\\mathrm\{term\}\}=12^\{\\circ\}vs\.60∘60^\{\\circ\}for the Pendulum\) means recovery from large inter\-sample intervals carries higher penalty\. Despite these constraints, the best\-model checkpoint reveals a clean monotonic stability–communication tradeoff: MSI increases steadily from0\.146±0\.068s0\.146\\pm 0\.068\\,\\mathrm\{s\}atwc=0\.25w\_\{c\}=0\.25to0\.308±0\.014s0\.308\\pm 0\.014\\,\\mathrm\{s\}atwc=16w\_\{c\}=16with0\.04%0\.04\\%RTA activity,96%96\\%ofτmax\\tau\_\{\\max\}\.
#### Quadrotor\.
The final\-model trend is intermediate between Pendulum and CartPole:wc=14w\_\{c\}=14yields the highest final\-model MSI \(0\.203±0\.008s0\.203\\pm 0\.008\\,\\mathrm\{s\},0\.40%0\.40\\%RTA\), whilewc=16w\_\{c\}=16shows12\.23%12\.23\\%RTA andwc=12w\_\{c\}=12shows7\.02%7\.02\\%RTA as the policy overshoots the admissible inter\-sample interval at the highest weights\. The best\-model checkpoint improves monotonically through the sweep, withwc=16w\_\{c\}=16achieving0\.281±0\.018s0\.281\\pm 0\.018\\,\\mathrm\{s\}at1\.96%1\.96\\%RTA,88%88\\%ofτmax\\tau\_\{\\max\}, compared to96%96\\%for Pendulum and CartPole, reflecting the harder coupled hover dynamics\. The elevatedP4\(θ˙\)P\_\{4\}\(\\dot\{\\theta\}\)values \(3\.63\.6–12\.612\.6\) across all weights are a structural property of the underactuated dynamics: to correct horizontal position error the policy must tilt the vehicle and return it to level, necessarily generating non\-zeroθ˙\\dot\{\\theta\}throughout the episode\. This is not a sign of instability but an inherent consequence of the indirect coupling betweenMMandxx\.
### L\.2SAC: Per\-Environment Discussion
SAC reaches near\-peak MSI for the Pendulum bywc=6w\_\{c\}=6and plateaus around0\.330\.33–0\.37s0\.37\\,\\mathrm\{s\}, matching DQN’s best checkpoint with zero RTA everywhere\. The DQN final model collapses at highwcw\_\{c\}; SAC’s final model remains stable, suggesting the continuous policy landscape is easier to optimize without collapsing into high\-RTA regimes\.
For CartPole, SAC eliminates the high\-RTA collapses that plague DQN \(38\.6%38\.6\\%atwc=0\.5w\_\{c\}=0\.5,24\.2%24\.2\\%atwc=10w\_\{c\}=10\)\. All final models are stable and MSI increases monotonically withwcw\_\{c\}, reaching0\.227s0\.227\\,\\mathrm\{s\}atwc=16w\_\{c\}=16\. The best checkpoint \(0\.258s0\.258\\,\\mathrm\{s\}\) falls short of DQN’s multi\-seed best \(0\.308±0\.014s0\.308\\pm 0\.014\\,\\mathrm\{s\}\), indicating that DQN transiently discovers sparser policies but cannot sustain them; SAC converges more reliably but has not yet matched DQN’s peak at1M1\\,\\mathrm\{M\}steps\.
For the Quadrotor, SAC matches DQN closely at moderate weights and surpasses it at highwcw\_\{c\}: the SAC best checkpoint atwc=16w\_\{c\}=16reaches0\.312s0\.312\\,\\mathrm\{s\}versus DQN’s0\.281±0\.018s0\.281\\pm 0\.018\\,\\mathrm\{s\}, while maintaining0%0\\%RTA\. The final model also increases monotonically \(0\.310s0\.310\\,\\mathrm\{s\}atwc=16w\_\{c\}=16\), in contrast to DQN’s12\.23%12\.23\\%RTA at that weight\. SAC’s stability advantage at highwcw\_\{c\}further validates the benefit of continuous\-action methods when DQN’s discrete policy degrades\.
### L\.3Preference\-Conditioned DQN: Per\-Environment Discussion
For Pendulum, the preference\-conditioned policy delivers a clean monotonic stability–communication tradeoff from a single2M2\\,\\mathrm\{M\}\-step training run\. The final model reaches0\.371s0\.371\\,\\mathrm\{s\}atwc=16w\_\{c\}=16with1\.6%1\.6\\%RTA and exhibits no high\-wcw\_\{c\}collapses, in contrast to the standard DQN final model\. The best\-model checkpoint achieves0\.393s0\.393\\,\\mathrm\{s\}, matching the standard DQN best checkpoint \(0\.397s0\.397\\,\\mathrm\{s\}\) within1%1\\%while using only211\\tfrac\{2\}\{11\}of the total training compute\.
For CartPole, the preference\-conditioned policy keeps RTA below0\.77%0\.77\\%in the final model across allwcw\_\{c\}values\. The best\-model checkpoint plateaus near0\.3130\.313–0\.316s0\.316\\,\\mathrm\{s\}abovewc=10w\_\{c\}=10, matching or slightly exceeding the multi\-seed standard DQN best \(0\.308±0\.014s0\.308\\pm 0\.014\\,\\mathrm\{s\}\)\. A single2M2\\,\\mathrm\{M\}\-step run thus produces a clean monotonic tradeoff that matches the per\-wcw\_\{c\}DQN ceiling at a fraction of the training compute\.
For the Quadrotor, the best checkpoint increases monotonically from0\.102s0\.102\\,\\mathrm\{s\}to0\.266s0\.266\\,\\mathrm\{s\}, within5%5\\%of multi\-seed standard DQN \(0\.281±0\.018s0\.281\\pm 0\.018\\,\\mathrm\{s\}\)\. The final model carries elevated RTA abovewc=8w\_\{c\}=8\(1111–16%16\\%\), indicating that the harder coupled hover dynamics remain challenging for a single shared policy at extreme communication weights\. The best checkpoint maintains≤2%\\leq 2\\%RTA throughout the sweep, confirming that checkpoint selection is particularly important for the quadrotor at highwcw\_\{c\}\.
## Appendix MDetailed Robustness Analysis
### M\.1Model Mismatch: Detailed Discussion
#### Pendulum degrades gracefully\.
The Pendulum’s safety margin is only1\.9∘1\.9^\{\\circ\}\. At0\.7×0\.7\\timesmass the pendulum swings faster for a given torque input, so the nominal one\-step\-ahead prediction systematically underestimates true angular velocity, causing RTA activation on30\.48±24\.80%30\.48\\pm 24\.80\\%of steps\. The large standard deviation \(versus4\.99±1\.68%4\.99\\pm 1\.68\\%at1\.3×1\.3\\times\) indicates bimodal behavior: episodes that begin near the upright position are handled safely, while those with larger initial angles trigger persistent RTA engagement\. Crucially, the system remains stable in all 100 episodes: the shield forcesτk←τmin\\tau\_\{k\}\\leftarrow\\tau\_\{\\min\}on affected steps, cutting achieved MSI to0\.246s0\.246\\,\\mathrm\{s\}but preventing divergence\. At1\.3×1\.3\\timesmass the slower dynamics increase the margin available to the nominal prediction, reducing RTA to4\.99%4\.99\\%and recovering most of the MSI \(0\.349s0\.349\\,\\mathrm\{s\}vs\.0\.396s0\.396\\,\\mathrm\{s\}nominal\)\.
#### Quadrotor is robust across the tested mass range\.
The Quadrotor is largely unaffected across both scaling factors \(MSI within2\.5%2\.5\\%of nominal, RTA at or below0\.2%0\.2\\%\), consistent with its9\.0∘9\.0^\{\\circ\}safety margin absorbing the mismatch without triggering additional interventions\.
#### Baseline 2 boundary case\.
At1\.3×1\.3\\timesmass, the fixed\-LQR controller atτmatch\\tau\_\{\\mathrm\{match\}\}succeeds for both Pendulum \(mean episode length50\.3s50\.3\\,\\mathrm\{s\}\) and CartPole \(50\.2s50\.2\\,\\mathrm\{s\}\), whereas both fail at nominal mass \(Table[2](https://arxiv.org/html/2605.12561#S5.T2)\)\. The heavier plant has lower acceleration per unit input, shifting the closed\-loop eigenvalues of the ZOH discretization at that fixed interval back into the stable region\. This confirms that the nominal failure of B2 is a genuine dynamical instability rather than a conservative evaluation artifact, and that the RL policy’s adaptive inter\-sample interval is what allows it to succeed where fixed\-rate LQR cannot\.
#### Domain randomization\.
Training with mass sampled fromUniform\[0\.6,1\.4\]×\\mathrm\{Uniform\}\[0\.6,1\.4\]\\timesat each episode reset substantially reduces Pendulum sensitivity: RTA activation at0\.7×0\.7\\timesmass drops from30\.48±24\.80%30\.48\\pm 24\.80\\%\(nominal training\) to0\.00%0\.00\\%\(DR\), confirming that the agent has internalized faster dynamics through training diversity\. For CartPole, DR reduces the0\.7×0\.7\\timesRTA activation from13\.91%13\.91\\%to0\.16%0\.16\\%and recovers MSI to0\.296s0\.296\\,\\mathrm\{s\}, essentially matching the non\-DR RL\-STC \(0\.317s0\.317\\,\\mathrm\{s\}\) while eliminating mass sensitivity\. For Quadrotor, already robust under nominal training, DR provides negligible additional benefit\. The DR models impose minimal performance cost at the best\-model checkpoint: nominal MSI values of0\.389s0\.389\\,\\mathrm\{s\},0\.319s0\.319\\,\\mathrm\{s\}, and0\.290s0\.290\\,\\mathrm\{s\}match the non\-DR references closely\. The benefit of DR scales with the magnitude of perturbation: larger mass deviations increase the one\-step\-ahead prediction error, causing more RTA activations and greater MSI degradation for the nominal policy; the DR policy, trained across the full deviation range, retains appropriate timing conservatism for those regimes and is correspondingly less affected\.
The natural resolution to mass sensitivity would be to include plant mass in the observation, allowing the policy to condition its timing on the current dynamics\. We deliberately exclude this because our goal is to evaluate a single fixed policy across mass variation, matching the deployment scenario where mass is unknown\. Adding mass to the observation would require online identification of a vector quantity \(for MIMO systems such as the Quadrotor, mass affects multiple actuator channels simultaneously\) and constitutes a separate problem\. The DR results therefore represent an inherent robustness–performance tradeoff: meaningful sensitivity reduction for environments with genuine mass fragility \(Pendulum and CartPole\) at negligible cost to nominal performance\.
### M\.2Disturbance Robustness: Detailed Discussion
#### RL\-STC dominates Classical STC even without disturbances\.
The no\-disturbance row of Table[27](https://arxiv.org/html/2605.12561#A10.T27)establishes the communication\-efficiency gap before any disturbance is applied: RL\-STC achieves0\.389s0\.389\\,\\mathrm\{s\}vs\.0\.202s0\.202\\,\\mathrm\{s\}for Pendulum \(\+93%93\\%\),0\.319s0\.319\\,\\mathrm\{s\}vs\.0\.212s0\.212\\,\\mathrm\{s\}for CartPole \(\+50%50\\%\), and0\.290s0\.290\\,\\mathrm\{s\}vs\.0\.080s0\.080\\,\\mathrm\{s\}for Quadrotor \(\+263%263\\%\)\. Classical STC’s conservative Lyapunov trigger consistently rejectsτ\\tauvalues the RL policy has learned to exploit safely, fixing MSI well below theτmax\\tau\_\{\\max\}ceiling in all three environments\.
#### RTA acts as a graduated shield under growing disturbances\.
For RL\-STC, RTA activation grows monotonically with disturbance amplitude while every episode completes safely: the safety filter absorbs what the policy cannot\. For the Pendulum, a constant0\.5Nm0\.5\\,\\mathrm\{Nm\}torque \(25%25\\%ofumaxu\_\{\\max\}\) drives RTA activation to88\.54%88\.54\\%and collapses MSI toτmin=0\.05s\\tau\_\{\\min\}=0\.05\\,\\mathrm\{s\}; at the milder0\.2Nm0\.2\\,\\mathrm\{Nm\}level, MSI falls36%36\\%\(from0\.3890\.389to0\.249s0\.249\\,\\mathrm\{s\}\) with26%26\\%activation\. CartPole is substantially more robust: even the strongest constant force tested \(2\.0N2\.0\\,\\mathrm\{N\},10%10\\%ofFmaxF\_\{\\max\}\) raises RTA to only10\.11±17\.22%10\.11\\pm 17\.22\\%and reduces MSI by38%38\\%\(0\.319→0\.199s0\.319\\to 0\.199\\,\\mathrm\{s\}\)\. The Quadrotor shows near\-complete insensitivity: MSI varies by at most0\.002s0\.002\\,\\mathrm\{s\}and RTA stays within0\.810\.81–1\.13%1\.13\\%across all conditions, because thrust disturbances perturb the altitude channel while the angular RTA trigger monitors tilt, leaving the timing policy structurally decoupled from the disturbance\. Classical STC’sτ\\tau\-selection uses only the undisturbed nominal model; disturbances enter only indirectly through the next state𝒙k\+1\\bm\{x\}\_\{k\+1\}\. For the Pendulum this makes Cl\.\-STC MSI almost completely flat \(≈0\.200s\{\\approx\}0\.200\\,\\mathrm\{s\}\) across all conditions, because disturbances are invisible to the trigger\. Despite maintaining higher MSI than RL\-STC under strong Pendulum disturbances, state quality is worse: under0\.5Nm0\.5\\,\\mathrm\{Nm\}constant torque, Cl\.\-STC achievesP3=0\.594P\_\{3\}=0\.594vs\. RL\-STC’sP3=0\.490P\_\{3\}=0\.490\(RL\-STC’s88%88\\%RTA activation imposes frequent LQR corrections that tighten the trajectory\)\. Under the0\.5Nm0\.5\\,\\mathrm\{Nm\}2Hz2\\,\\mathrm\{Hz\}periodic case, Cl\.\-STC’s angular\-velocity norm explodes toP4=3\.26P\_\{4\}=3\.26vs\. RL\-STC’sP4=1\.36P\_\{4\}=1\.36; the Lyapunov trigger selectsτ=0\.20s\\tau=0\.20\\,\\mathrm\{s\}throughout, oblivious to the resonant forcing\.
For CartPole under constant force, Cl\.\-STC MSI drops significantly \(0\.212→0\.123s0\.212\\to 0\.123\\,\\mathrm\{s\}at1\.0N1\.0\\,\\mathrm\{N\}\) as the drifting cart state makes the nominal Lyapunov condition more demanding; yet despite communicating more frequently, the cart\-position normP1P\_\{1\}worsens dramatically \(0\.065→2\.951m⋅s0\.065\\to 2\.951\\,\\mathrm\{m\\cdot s\}\)\. RL\-STC simultaneously maintains higher MSI \(0\.284s0\.284\\,\\mathrm\{s\}\) and better position tracking \(P1=3\.345m⋅sP\_\{1\}=3\.345\\,\\mathrm\{m\\cdot s\}with near\-zero RTA activation\), showing the learned policy handles constant bias more gracefully than the greedy Lyapunov trigger\. For the Quadrotor, Cl\.\-STC is already at0\.080s0\.080\\,\\mathrm\{s\}in the undisturbed case; disturbances move it only slightly \(0\.080→0\.087s0\.080\\to 0\.087\\,\\mathrm\{s\}\), mirroring the structural insensitivity of RL\-STC but at a communication cost3\.6×3\.6\\timeshigher\.
#### Impulse disturbances reveal bimodal RL\-STC behavior\.
Random impulse kicks \(p=0\.05p=0\.05per RL step\) produce strikingly different responses\. The Pendulum is highly sensitive: a0\.5Nm0\.5\\,\\mathrm\{Nm\}kick collapses RL\-STC MSI to0\.252±0\.063s0\.252\\pm 0\.063\\,\\mathrm\{s\}and raises RTA activation to38\.38±18\.59%38\.38\\pm 18\.59\\%\. The large standard deviation reflects bimodal episode behavior: some episodes receive well\-timed kicks that sustain RTA engagement, while others see none\. CartPole and Quadrotor are resilient: the largest impulses produce≤1%\\leq 1\\%RTA activation and≤2%\\leq 2\\%MSI change, for the same structural reasons described above \(angular RTA trigger is decoupled from the thrust/force disturbance channel for the Quadrotor\)\. Classical STC is insensitive to impulses in all environments, since a rare kick shifts the state only marginally relative to the Lyapunov basin\.
\(a\)Pendulum
\(b\)CartPole
\(c\)Quadrotor
Figure 6:Representative episode trajectories on the three lower\-dimensional environments \(seed 0, up to 1200 steps\)\. Each sub\-figure:leftRL\-STC \(ours\);centerRL\-STC without RTA \(Ablation A\);rightClassical Lyapunov\-STC\. Top row: state trajectories with±θRTA\\pm\\theta\_\{\\mathrm\{RTA\}\}band\. Bottom row: inter\-sample intervalsτk\\tau\_\{k\}; red stems = RTA\-active steps; dashed line = MSI\. The Quadrotor sub\-figure highlights the dramatic MSI gap: the RL policy sustains long inter\-sample intervals near hover while the classical trigger remains nearτmin\\tau\_\{\\min\}throughout\.Similar Articles
Safe Continual Reinforcement Learning under Nonstationarity via Adaptive Safety Constraints
Proposes LILAC+, a framework for safe continual reinforcement learning under nonstationarity that uses three adaptive safety mechanisms: context-based safety constraints, adaptation-speed constraints, and budget-to-state safety enforcement. Evaluations in simulated driving environments show reduced safety violations under distribution shift while maintaining competitive performance.
Uncertainty-Aware and Temporally Regulated Expert Advice in Reinforcement Learning for Autonomous Driving
This paper proposes an uncertainty-aware reinforcement learning framework for autonomous driving that uses expert advice guided by adaptive uncertainty thresholds and a commitment-cooldown strategy to improve safety and efficiency. Experiments in the CARLA simulator show a 5-7% success improvement over the IQN baseline.
From Cumulative Constraints to Adaptive Runtime Safety Control for Nonstationary Reinforcement Learning
Proposes CPSS, a runtime safety mechanism that converts cumulative cost constraints into adaptive state-level thresholds for safe reinforcement learning in nonstationary environments, demonstrating reduced violations in highway merging scenarios.
Distributionally Robust and Safe Imitation Learning
This paper proposes a unified imitation learning framework using Taylor Series Imitation Learning and distributionally robust adaptive control to address both policy-induced and uncertainty-induced distribution shifts, with a UAV case study demonstrating safety under uncertainty.
Adjustment Speed as a Safety Constraint for Nonstationary Reinforcement Learning
This paper proposes adjustment speed as a safety constraint for nonstationary reinforcement learning, defining safety in terms of adaptation feasibility and using representation learning with context forecasts to proactively regulate behavior when predicted adaptation demand exceeds the system's achievable capacity.