PowerZooJax:一个基于 JAX 的电力系统强化学习基准平台

arXiv cs.AI 论文

摘要

来自南洋理工大学、康奈尔大学和布里斯托大学的研究者提出了 PowerZooJax,这是一个面向电力系统运行强化学习的开源 JAX 基准套件,提供了五个带约束的 MDP 任务(发电、输电、配电、分布式能源资源、数据中心微电网),可完全在 GPU 上运行,相比基于 CPU 的仿真器实现了大幅加速。

arXiv:2609.36052v1 Announce Type: new Abstract: Power system operation is a safety-critical sequential decision-making problem, making it a natural testbed for reinforcement learning (RL). However, existing RL environments for power systems are often narrow in scope and computationally limited by CPU-based simulation workflows, making large-scale evaluation difficult. We introduce PowerZooJax, a JAX-based benchmark suite for RL in power system operation. It provides five constrained Markov decision process tasks spanning generation, transmission, distribution, distributed energy resources, and data center microgrid. By rewriting power flow, economic dispatch, market clearing, and device dynamics as JAX computation graphs, PowerZooJax keeps the entire training and evaluation loop on the GPU. Experiments show substantial speedups over CPU-based simulations and demonstrate standardized evaluation of policy returns, safety violations, and out-of-distribution stress conditions. Our open-source benchmark is available at: https://github.com/powerzoojax/PowerZooJax.
查看原文
查看缓存全文

缓存时间: 2026/09/30 09:40

# PowerZooJax: A JAX-based Power System Benchmark for Reinforcement Learning
Source: [https://arxiv.org/html/2609.36052](https://arxiv.org/html/2609.36052)
Zhanhua Pan, Xiao Liu, Zhilong Cao, Jianhong Wang, Dawei QiuNanyang Technological University, Singapore\. Cornell University, USA\. University of Bristol, UK\.††thanks:Corresponding author:dawei\.qiu@ntu\.edu\.sg\.†These authors contributed equally\.‡Co\-supervisors\.

###### Abstract

Power system operation is a safety\-critical sequential decision\-making problem, making it a natural testbed for reinforcement learning \(RL\)\. However, existing RL environments for power systems are often narrow in scope and computationally limited by CPU\-based simulation workflows, making large\-scale evaluation difficult\. We introduce PowerZooJax, a JAX\-based benchmark suite for RL in power system operation\. It provides five constrained Markov decision process tasks spanning generation, transmission, distribution, distributed energy resources, and data center microgrid\. By rewriting power flow, economic dispatch, market clearing, and device dynamics as JAX computation graphs, PowerZooJax keeps the entire training and evaluation loop on the GPU\. Experiments show substantial speedups over CPU\-based simulations and demonstrate standardized evaluation of policy returns, safety violations, and out\-of\-distribution stress conditions\. Our open\-source benchmark is available at:[https://github\.com/powerzoojax/PowerZooJax](https://github.com/powerzoojax/PowerZooJax)\.

![Refer to caption](https://arxiv.org/html/2609.36052v1/PowerZooJax.png)Figure 1:Execution paradigm of PowerZooJax\.## 1Introduction

Power systems are undergoing a fundamental transformation driven by the 3Ds: decarbonization, decentralization, and digitalization\[[1](https://arxiv.org/html/2609.36052#bib.bib19)\]\. Traditionally, power systems were operated around centralized generation, long\-distance transmission, and largely passive demand\. This paradigm is increasingly challenged by the growing integration of renewable generation, distributed energy resources \(DERs\), battery storage, electric vehicle \(EV\) charging, flexible loads, and emerging critical infrastructure such as data centers\[[2](https://arxiv.org/html/2609.36052#bib.bib11)\]\. Furthermore, the rise of artificial intelligence \(AI\) reinforces the central role of power systems: AI development depends on reliable, affordable, and sustainable electricity supply\[[3](https://arxiv.org/html/2609.36052#bib.bib13)\], while AI\-driven data centers are becoming a major source of new electricity demand\[[4](https://arxiv.org/html/2609.36052#bib.bib14)\]\.

Operating power systems requires sequential decision\-making under uncertainty while satisfying strict physical, operational, and security constraints\[[5](https://arxiv.org/html/2609.36052#bib.bib42),[6](https://arxiv.org/html/2609.36052#bib.bib37)\]\. This structure makes power system operation a natural domain formulated as a constrained Markov decision process \(CMDP\) to study reinforcement learning \(RL\) algorithms\. As national critical infrastructure, the power system is one of the most complex and safety\-critical physical systems, where operational failures can disrupt economic activity and public services\[[7](https://arxiv.org/html/2609.36052#bib.bib38),[8](https://arxiv.org/html/2609.36052#bib.bib15)\]\. System operators must continuously balance supply and demand, manage constrained power flows, maintain voltage and power quality, coordinate DERs, and ensure the reliable operation of local microgrids\[[9](https://arxiv.org/html/2609.36052#bib.bib1)\]\.

On the other hand, applying RL to power system operation remains challenging from both modeling and computational perspectives\. Specifically, power systems rely on different physical approximations and optimization formulations, ranging from DC optimal power flow \(DC OPF\) in transmission networks to AC optimal power flow \(AC OPF\) in distribution networks and device\-level DER coordination at the grid edge\. These differences motivate benchmarks that preserve task\-specific power system physics while providing a consistent interface and standardized evaluation protocol for RL algorithms\. Beyond these domain\-specific challenges, existing power system RL environments are often limited by their computational structure\. Most environments run power flow calculations and system dynamics through CPU\-based simulators, while RL policies are commonly trained on GPUs\. This separation makes large\-scale training inefficient: rollouts are constrained by CPU simulation speed, and repeated CPU\-GPU communication introduces additional overhead\. As a result, evaluation over tremendous conditions, random seeds and trajectories will strongly delay the procedure of developing RL algorithms for real\-world power system applications\.

To mitigate this issue, we aim to develop a more efficient benchmark for evaluating RL algorithms in power system operation\. Inspired by the recent proliferation of JAX\-based RL environments\[[10](https://arxiv.org/html/2609.36052#bib.bib18),[11](https://arxiv.org/html/2609.36052#bib.bib22)\], we introduce PowerZooJax, a novel power system RL benchmark fully implemented in JAX\[[12](https://arxiv.org/html/2609.36052#bib.bib21)\], so it enables compiled and vectorized rollouts completely on GPUs\. This design reduces the CPU\-GPU communication overhead present in many existing power system RL benchmarks\[[13](https://arxiv.org/html/2609.36052#bib.bib29),[14](https://arxiv.org/html/2609.36052#bib.bib41),[15](https://arxiv.org/html/2609.36052#bib.bib35)\]\.

The main contributions of this work are threefold\. First, we present a unified benchmark scope that covers five major tasks in daily power system operation, enabling RL algorithms to be evaluated across the entire power system decision\-making problems\. Second, we formulate these operational problems under a consistent CMDP interface, enabling systematic comparison of RL algorithms across five tasks with different physical models, operational constraints, and objectives\. Third, we provide JAX\-based benchmark environments, baseline RL algorithms, and evaluation metrics that support scalable and reproducible assessment of RL\-based power system control methods\.

## 2Related Work

Power System RL Benchmarks\.Existing power system RL benchmarks cover different parts of the power grid\. For grid operation, MAPDN\[[16](https://arxiv.org/html/2609.36052#bib.bib43)\]studies voltage control, ANDES\-Gym\[[17](https://arxiv.org/html/2609.36052#bib.bib2)\]supports dynamic power system simulation, PowerGridworld\[[18](https://arxiv.org/html/2609.36052#bib.bib28)\]examines distribution system control, and RL2Grid\[[15](https://arxiv.org/html/2609.36052#bib.bib35)\]and its multi\-agent extension MARL2Grid\-TR\[[19](https://arxiv.org/html/2609.36052#bib.bib24)\]target transmission topology and redispatching control\. At the end\-use DER level, CityLearn v2\[[20](https://arxiv.org/html/2609.36052#bib.bib6)\], GridLearn\[[21](https://arxiv.org/html/2609.36052#bib.bib17)\], and SustainGym\[[14](https://arxiv.org/html/2609.36052#bib.bib41)\]provide environments for flexible consumption, building energy management, demand response, and grid\-interactive energy communities\. At the microgrid level, pymgrid\[[22](https://arxiv.org/html/2609.36052#bib.bib31)\]and CommonPower\[[23](https://arxiv.org/html/2609.36052#bib.bib7)\]focus on local generation, storage, and load scheduling\. Collectively, these benchmarks demonstrate that power systems provide realistic sequential decision\-making problems with uncertainty, physical constraints, and operational objectives\.

However, most existing benchmarks focus on a specific operational task\. Many also rely on CPU\-based solvers, making large\-scale parallel rollouts costly: MARL2Grid\-TR, for example, reports roughly 120,000 CPU hours to train and evaluate its baselines\. In addition, they often lack standardized task protocols, violation accounting, and evaluation records, which makes cross\-layer comparison and independent verification difficult\. PowerZooJax addresses these limitations through a JAX\-based architecture that provides physically grounded tasks across generation, transmission, distribution, end\-use DER coordination, and microgrid operation\. It further offers fixed task protocols, user\-defined reward functions and safety costs, and reproducible evaluation records for scalable and standardized RL evaluation\.

JAX\-based RL Benchmarks\.Recent years have seen a growing number of RL benchmarks implemented in JAX, motivated by the need for fast, vectorized, and accelerator\-compatible environment simulation\. Brax\[[24](https://arxiv.org/html/2609.36052#bib.bib5)\]is an early and influential JAX\-based benchmark for physics\-based RL, demonstrating the benefits of differentiable and massively parallel simulation\. Gymnax\[[10](https://arxiv.org/html/2609.36052#bib.bib18)\]further re\-implements classic Gym control tasks in JAX, providing lightweight environments for benchmarking RL algorithms\. Jumanji\[[25](https://arxiv.org/html/2609.36052#bib.bib4)\]extends this direction to combinatorial optimization and structured decision\-making problems, while Pgx\[[26](https://arxiv.org/html/2609.36052#bib.bib27)\]focuses on classic board games\.

Beyond single\-agent RL, JAX\-based benchmarks have also been developed for multi\-agent and long\-horizon settings\. JaxMARL\[[11](https://arxiv.org/html/2609.36052#bib.bib22)\]re\-implements a collection of popular multi\-agent RL environments in JAX, enabling accelerated simulation for MARL research\. More recently, Craftax\[[27](https://arxiv.org/html/2609.36052#bib.bib10)\]builds on ideas from MiniHack and Crafter to provide a fast grid\-world benchmark for studying long\-horizon exploration and generalization of RL algorithms\. Octax\[[28](https://arxiv.org/html/2609.36052#bib.bib32)\]re\-implements a suite of CHIP\-8 games in JAX for benchmarking different aspects of capabilities of RL algorithms\.

PowerZooJax belongs to this growing family of JAX\-based RL benchmarks, but targets a distinct and underrepresented domain: real\-world power system operation\. Unlike existing JAX benchmarks that primarily focus on classic control, combinatorial optimization, or games, PowerZooJax provides tasks grounded in practical power system operation while exposing core RL challenges such as partial observability, mixed action spaces, heterogeneous agents, long\-horizon scheduling, and multi\-objective control\. In this sense, PowerZooJax complements existing JAX\-based benchmarks by extending accelerator\-compatible RL evaluation to large\-scale, safety\-critical power system operation\.

## 3Background and Problem Setting

The daily operation of a modern power system can be decomposed into five tasks, as illustrated in Figure[2](https://arxiv.org/html/2609.36052#S3.F2): generation, transmission, distribution, end\-use DER, and data center microgrid\. These tasks span dispatching conventional and renewable generation resources to meet time\-varying demand, operating the high\-voltage transmission network via DC OPF, operating the lower\-voltage distribution network via AC OPF, coordinating user\-side DER flexibility, and managing local generation, battery storage, backup resources, and critical loads in a data center microgrid\.

![Refer to caption](https://arxiv.org/html/2609.36052v1/power_grid_model.png)Figure 2:Overview of a modern power system model spanning generation, transmission, distribution, end\-use DERs, and a data center microgrid\.Despite their different physical scales and modeling details, these five tasks share a common sequential decision\-making structure under uncertainties and safety constraints, which can be formulated as constrained Markov decision processes \(CMDPs\)\[[29](https://arxiv.org/html/2609.36052#bib.bib34)\]\. At each step, the system state may include: demand, renewable availability, generator status, network conditions, voltage profiles, storage states of charge, and microgrid operating modes\. Each agent’s observation space is varied among different tasks\. Agents select actions from task\-specific action spaces, such as generation dispatch, power flow control, DER scheduling, storage charging or discharging, and data center workload scheduling\. The system then evolves according to the physical power flow dynamics, device constraints, and stochastic exogenous factors \(e\.g\. load variation and renewable uncertainty\)\. Reward functions are designed by encoding operational objectives, including cost minimization, carbon reduction, and service continuity, while safety costs capture constraint violations such as power imbalance, voltage violations, and line overloads\. The following five paragraphs provide an overview of the CMDP formulation for each task; the formal CMDP tuple, notation, and full details on state and action spaces, transition dynamics, reward functions, and safety costs are deferred to Appendix[A](https://arxiv.org/html/2609.36052#A1)and Appendix[C](https://arxiv.org/html/2609.36052#A3)\.

GenCos \(Generation Companies\) — competitive market bidding\.Five generation companies bid into the wholesale electricity market on the IEEE 5\-bus transmission networkcase5over a 24\-hour horizon at 30\-minute resolution, using Great Britain \(GB\) demand profiles111[https://www\.neso\.energy/data\-portal/historic\-demand\-data](https://www.neso.energy/data-portal/historic-demand-data)National Energy System Operator \(NESO\)\.\. Each agent operating a generation company observes its own generation status and public market price signals, submits a piecewise\-linear bid curve, and receives realized profit as the reward after security constrained economic dispatch \(SCED\)\-based market clearing\. Thermal line overloads are recorded as the safety cost\. This is a competitive multi\-agent RL task with partial observability and strategic interactions\.

TSO \(Transmission System Operator\) — transmission unit commitment\.A transmission system operator commits and dispatches 54 generators on the IEEE 118\-bus transmission networkcase118using GB demand profiles over a 24\-hour horizon at 30\-minute resolution\. The task enforces ramping limit, minimum up/down\-time, and DC OPF feasibility constraints at each step\. The transmission system operator agent observes the full system state and receives negative operating cost as the reward, while thermal line overloads and reserve shortfalls are recorded as two separate safety costs\. This is a single\-agent RL task with a large mixed continuous\-binary action space\.

DSO \(Distribution System Operator\) — distribution demand response\.A distribution system operator coordinates six flexible loads on the IEEE 33\-bus distribution networkcase33bw, using real Ausgrid distribution demand profiles222[https://www\.ausgrid\.com\.au/Industry/Our\-Research/Data\-to\-share/](https://www.ausgrid.com.au/Industry/Our-Research/Data-to-share/), Ausgrid\.over a 24\-hour horizon at 30\-minute resolution\. The distribution system operator agent observes the full system state, including bus voltages, branch flows, nodal demands, and deferred load\-shifting buffers, and selects curtailment and load\-shifting fractions for each controllable load\. The reward is negative network loss, while voltage limit violations are recorded as the safety cost\. This is a single\-agent RL task for distribution\-level demand response and voltage\-aware load coordination\.

DERs \(Distributed Energy Resources\) — cooperative DER coordination\.Twelve DERs, including four batteries, four PV inverters333[https://bmrs\.elexon\.co\.uk/generation\-by\-fuel\-type](https://bmrs.elexon.co.uk/generation-by-fuel-type), Elexon\., and four flexible loads, are placed at fixed buses of the 141\-bus distribution networkcase141\. Each distributed energy resource agent observes its own bus voltage, the voltages of its electrically connected neighboring buses, a compact system\-wide voltage summary, and its local device state\. Each agent takes a 2\-dimensional device\-specific action, such as active/reactive power control for batteries, curtailment/reactive power control for PV inverters, and curtailment/load shifting for flexible loads\. All agents share a negative network loss reward, while voltage, thermal, and device feasibility violations are recorded as three separate safety costs\. This is a cooperative multi\-agent RL task with heterogeneous agents and local observations\.

DCMG \(Data Center Microgrid\) — long\-horizon multi\-objective scheduling\.A data center microgrid powers an AI data center using on\-site PV, battery storage, a diesel generator, and a connection to the main grid\. A microgrid agent controls the system over 288 five\-minute steps using real Google workload traces\[[30](https://arxiv.org/html/2609.36052#bib.bib16)\], jointly scheduling training and fine\-tuning jobs, cooling, battery operation, diesel generation, and grid interaction based on workload, thermal, and energy states\. Under normal conditions, the microgrid can operate in grid\-connected mode; during disturbances or outages, it can switch to islanded mode to maintain critical load service\. The reward combines energy use, operating cost, and carbon emissions, while power imbalance, server\-zone over\-temperature, and workload incompletion are recorded as three separate safety costs\. This is a single\-agent RL task for long\-horizon, multi\-objective scheduling of critical computing infrastructure\.

## 4Power System Problem, RL Challenge and Evaluation Scope

From the perspective of RL, the five tasks can be viewed as either single\-agent RL or multi\-agent RL problems\. This leads to one of our aims in proposing PowerZooJax: to reflect how RL techniques can address the emergent real\-world problems, while providing a clean, unified testbed for AI researchers to easily pinpoint the key challenges of RL problems\. Each task is grounded in a realistic operational setting, but is also designed to highlight a distinct RL challenge, such as partial observability, mixed action spaces, heterogeneous agents, long\-horizon scheduling, or multi\-objective control\.

Table[1](https://arxiv.org/html/2609.36052#S4.T1)summarizes the interrelationship between power system problems embedded in the tasks we defined and the corresponding RL challenges in the modern RL literature\. We categorize the tasks into centralized and decentralized settings according to their operational roles\. Centralized tasks correspond to decisions that are typically executed by system operators to maintain secure and reliable network operation, while decentralized tasks correspond to settings in which multiple users, assets, or market participants make individual decisions based on local information\.

From the perspective of RL evaluation, Table[1](https://arxiv.org/html/2609.36052#S4.T1)clearly captures the basic assumptions of popular RL problems for each task\. As a result, researchers can confidently get the information about what features of their algorithms \(RL challenges\) are expected to be meaningfully assessed\. More importantly, due to that our benchmark tasks follow the practical settings of power system problems, algorithms that work well on our benchmark can be explained as possessing the potential to handle real\-world problems\. This is what most of previous benchmarks that focused on either games or simplified grid\-world settings lack\.

Table 1:Interrelationship between power system problems and RL challenges\.Note:“dim” is short for dimensional\.

## 5How Does PowerZooJax Work?

Each of the five tasks in Section[3](https://arxiv.org/html/2609.36052#S3)runs a power flow \(PF\) or optimal power flow \(OPF\) computation at every environment step\. DSO and DERs use radial AC PF, TSO and GenCos rely on DC OPF clearing, and DCMG combines local power\-balance checks with device dynamics\. These computations involve fixed\-shape numerical operations, such as matrix updates, constraint evaluation, and batched linear\-algebra routines, which are well suited to compilation and fusion by accelerated linear algebra \(XLA\)\. Conventional pipelines typically write them as host\-side Python loops that traverse a feeder tree or call a CPU solver for a linear program, incurring a host round trip at every step\[[31](https://arxiv.org/html/2609.36052#bib.bib23)\]\. In contrast, our PowerZooJax rewrites each iteration as a fixed\-shape JAX computation graph and runs the full rollout, including batched parallel scenarios, inside a compiled accelerator program, as shown in Figure[1](https://arxiv.org/html/2609.36052#S0.F1)\.

This design enables PowerZooJax to keep the whole training loop on the GPU\. From episode reset, through transitions, rewards, safety costs, action sampling, and gradient updates, the main computation operates on GPU\-resident arrays\. Episode boundaries do not require returning to a Python control loop: when one episode terminates, the next reset can be executed as the next operation in the same compiled JAX program\. As a result, PowerZooJax avoids per\-step CPU\-GPU synchronization during rollout generation, reducing overhead and enabling efficient large\-scale parallel training\.

### 5\.1Rewriting Power Flow as a JAX Computation Graph

A traditional radial PF solver, e\.g\. as implemented in pandapower\[[32](https://arxiv.org/html/2609.36052#bib.bib26)\], is a host\-side fixed\-point iteration\. At each iterationkk, the solver walks the feeder tree backward to aggregate downstream loads into branch flows𝐏\(k\),𝐐\(k\)\\mathbf\{P\}^\{\(k\)\},\\mathbf\{Q\}^\{\(k\)\}, walks it forward to update bus voltages𝐯2,\(k\)\\mathbf\{v\}^\{2,\(k\)\}, and checks convergence conditionmaxn⁡\|𝐯n2,\(k\)−𝐯n2,\(k−1\)\|<ε\\max\_\{n\}\|\\mathbf\{v\}^\{2,\(k\)\}\_\{n\}\-\\mathbf\{v\}^\{2,\(k\-1\)\}\_\{n\}\|<\\varepsiloninside a Pythonwhileloop\. Since both sweeps and the convergence test are controlled on the host, each RL environment step incurs host\-side control overhead and potentially a device\-host synchronization\. PowerZooJax rewrites the same iterative procedure as a fixed\-shape JAX computation graph through three structural changes\. In each change below, the expression on the left of→\\rightarrowdenotes the traditional implementation pattern, while the expression on the right denotes the PowerZooJax replacement\.

Tree walking→\\rightarrowdense matrix\-vector products\.In a radial network, the backward and forward sweeps can be represented using fixed topology matrices\. PowerZooJax precomputes two such matrices at environment construction time: a downstream\-aggregation matrix𝐃∈ℝNℓ×N\\mathbf\{D\}\\in\\mathbb\{R\}^\{N\_\{\\ell\}\\times N\}, whereNNandNℓN\_\{\\ell\}are the numbers of buses and branches and𝐃ℓ,n=1\\mathbf\{D\}\_\{\\ell,n\}=1if busnnis downstream of lineℓ\\ell, and a path matrix𝚷∈ℝN×Nℓ\\boldsymbol\{\\Pi\}\\in\\mathbb\{R\}^\{N\\times N\_\{\\ell\}\}, where𝚷n,ℓ=1\\boldsymbol\{\\Pi\}\_\{n,\\ell\}=1if lineℓ\\elllies on the unique path from the slack bus to busnn\. The backward sweep then becomes a matrix\-vector multiplication that aggregates downstream active and reactive loads into branch flows, while the forward sweep becomes a matrix\-vector multiplication that accumulates voltage drops along each root\-to\-bus path\. One PF iteration can therefore be written as

𝐏\(k\)=𝐃⁡\(𝐩load\+𝐩loss,\(k−1\)\),\\displaystyle\\mathbf\{P\}^\{\(k\)\}=\\mathbf\{D\}\\,\\bigl\(\\mathbf\{p\}^\{\\mathrm\{load\}\}\+\\mathbf\{p\}^\{\\mathrm\{loss\},\(k\-1\)\}\\bigr\),\(1\)𝐯2,\(k\)=Vslack2​1−𝚷​𝜹2,\(k\),\\displaystyle\\mathbf\{v\}^\{2,\(k\)\}=V\_\{\\mathrm\{slack\}\}^\{2\}\\,\\mathbf\{1\}\-\\boldsymbol\{\\Pi\}\\,\\boldsymbol\{\\delta\}^\{2,\(k\)\},with𝐐\(k\)\\mathbf\{Q\}^\{\(k\)\}computed analogously\. Here,𝐩load\\mathbf\{p\}^\{\\mathrm\{load\}\}is the vector of nodal active loads,𝐩loss,\(k−1\)\\mathbf\{p\}^\{\\mathrm\{loss\},\(k\-1\)\}is the per\-bus loss correction recomputed from the previous iterate, and𝜹2,\(k\)\\boldsymbol\{\\delta\}^\{2,\(k\)\}denotes the per\-line squared\-voltage drop computed from\(𝐏\(k\),𝐐\(k\),𝐯2,\(k−1\)\)\(\\mathbf\{P\}^\{\(k\)\},\\mathbf\{Q\}^\{\(k\)\},\\mathbf\{v\}^\{2,\(k\-1\)\}\)using the standard DistFlow recurrence\[[33](https://arxiv.org/html/2609.36052#bib.bib25)\]\. Since𝐃\\mathbf\{D\}and𝚷\\boldsymbol\{\\Pi\}depend only on the network topology, they are precomputed once and reused throughout training\. This replaces explicit feeder\-tree traversal with fixed\-shape matrix\-vector products that are directly compatible with JAX compilation and batching\.

Pythonwhile→\\rightarrowjax\.lax\.while\_loop\.Equation \([1](https://arxiv.org/html/2609.36052#S5.E1)\) defines a fixed\-point iteration bodyf:\(𝐏,𝐐,𝐯2\)↦\(𝐏′,𝐐′,𝐯′2\)f:\(\\mathbf\{P\},\\mathbf\{Q\},\\mathbf\{v\}^\{2\}\)\\mapsto\(\\mathbf\{P\}^\{\\prime\},\\mathbf\{Q\}^\{\\prime\},\\mathbf\{v\}^\{\\prime 2\}\), whose output depends only on its input arrays\. In PowerZooJax, this body is implemented entirely withjnpandlaxoperations, without host\-side mutation or external state updates\. JAX therefore traces the iteration body once into a static XLA computation graph and runs the fixed\-point loop withjax\.lax\.while\_loop, replacing the host\-controlled Pythonwhileloop with a device\-side loop that can be compiled, fused, and batched across parallel environments\.

Per\-step solver call→\\rightarrowbatched PF rollout viajax\.vmap/jax\.lax\.scan\.Once the radial PF iteration is written as fixed\-shape matrix operations and a device\-sidelax\.while\_loop, PowerZooJax can evaluate many PF solves without launching a separate solver at every environment step\. The precomputed topology matrices𝐃\\mathbf\{D\}and𝚷\\boldsymbol\{\\Pi\}are shared across batched scenarios, so the same PF kernel is applied in parallel over a leading batch dimension withvmap\. Across time,lax\.scancomposes the per\-step PF solve with device updates, rewards, and safety\-cost calculations, producing a single compiled rollout for an entire RL episode\. In our tasks, this means that the 48 PF evaluations required by one DSO or DERs episode are executed as part of one accelerator\-resident computation graph, rather than as repeated Python calls to an external solver\.

### 5\.2Optimal Power Flow as a Computation Graph

The same computation\-graph principle applies to OPF kernels\. We illustrate this with DC OPF, which is used by the TSO and GenCos tasks\. DC OPF is a linear program that minimizes generation cost subject to system\-wide power balance, generator capacity bounds, and transmission line\-flow limits\. In a conventional pipeline, this linear program is passed to an external CPU solver once per environment step, making the OPF solve a host\-side bottleneck\. PowerZooJax instead implements OPF solvers directly inside the JAX graph: an ADMM\[[34](https://arxiv.org/html/2609.36052#bib.bib12)\]solver for TSO’s DC OPF, and a PDIPM\[[35](https://arxiv.org/html/2609.36052#bib.bib8)\]solver for GenCos market clearing\. For example, one iteration of the ADMM DC OPF solver can be written as

\[𝐩\(k\+1\)νeq\(k\+1\)\]=Aaug−1​\[𝐫\(k\)𝟏⊤​𝐩d\],\\displaystyle\\begin\{bmatrix\}\\mathbf\{p\}^\{\(k\+1\)\}\\\\ \\nu^\{\(k\+1\)\}\_\{\\mathrm\{eq\}\}\\end\{bmatrix\}=A\_\{\\mathrm\{aug\}\}^\{\-1\}\\begin\{bmatrix\}\\mathbf\{r\}^\{\(k\)\}\\\\ \\mathbf\{1\}^\{\\top\}\\mathbf\{p\}^\{d\}\\end\{bmatrix\},\(2\)𝐳1\(k\+1\)=clip⁡\(𝐩\(k\+1\)\+𝐲1\(k\)/ρ,𝐩¯,𝐩¯\),\\displaystyle\\mathbf\{z\}\_\{1\}^\{\(k\+1\)\}=\\mathrm\{clip\}\\bigl\(\\mathbf\{p\}^\{\(k\+1\)\}\+\\mathbf\{y\}\_\{1\}^\{\(k\)\}/\\rho,\\ \\underline\{\\mathbf\{p\}\},\\ \\overline\{\\mathbf\{p\}\}\\bigr\),where𝐩\\mathbf\{p\}is the generator dispatch vector with output bounds\(𝐩¯,𝐩¯\)\(\\underline\{\\mathbf\{p\}\},\\overline\{\\mathbf\{p\}\}\),𝐩d\\mathbf\{p\}^\{d\}is the nodal demand vector,νeq\\nu\_\{\\mathrm\{eq\}\}is the multiplier on the system\-wide power\-balance equality,𝐳1\\mathbf\{z\}\_\{1\}and𝐲1\\mathbf\{y\}\_\{1\}are the ADMM auxiliary and dual variables associated with the generator bounds,ρ\>0\\rho\>0is the ADMM penalty parameter, and𝐫\(k\)\\mathbf\{r\}^\{\(k\)\}is the residual vector assembled from the current auxiliary and dual iterates\. The projection step for the line\-flow auxiliary variable𝐳2\\mathbf\{z\}\_\{2\}and the dual updates for𝐲1,𝐲2\\mathbf\{y\}\_\{1\},\\mathbf\{y\}\_\{2\}complete one iteration \(Appendix[B\.2](https://arxiv.org/html/2609.36052#A2.SS2), Equation \([15](https://arxiv.org/html/2609.36052#A2.E15)\)\)\. The augmented KKT matrixAaugA\_\{\\mathrm\{aug\}\}depends only on grid parameters and does not change with the per\-step load, so its inverse is precomputed once at setup, similar to the topology matrices𝐃\\mathbf\{D\}and𝚷\\boldsymbol\{\\Pi\}used in the PF kernel\. Each ADMM iteration is then reduced to fixed\-shape matrix operations and elementwise projections, making the iteration body fully traceable by JAX\. The complete OPF solve is executed as a fixed\-lengthlax\.fori\_loop, so batched OPF solves can be compiled and run on GPU device without repeated calls to an external CPU solver\.

## 6Evaluation

We evaluate PowerZooJax on the five benchmark tasks introduced in Section[3](https://arxiv.org/html/2609.36052#S3)\. Section[6\.1](https://arxiv.org/html/2609.36052#S6.SS1)first evaluates the computational efficiency of the JAX\-based implementation described in Section[5](https://arxiv.org/html/2609.36052#S5), measuring wall\-clock speedups under matched training budgets\. Sections[6\.2](https://arxiv.org/html/2609.36052#S6.SS2)–[6\.4](https://arxiv.org/html/2609.36052#S6.SS4)then present representative task\-level results for GenCos, TSO, and DCMG, including training returns, safety violations, comparisons with task\-specific baselines, and performance on out\-of\-distribution \(OOD\) splits\. Analogous results for DSO and DERs follow the same protocol and are reported in Appendix[G\.3](https://arxiv.org/html/2609.36052#A7.SS3)–[G\.4](https://arxiv.org/html/2609.36052#A7.SS4)\.

Hardware\.All experiments are run on a single workstation equipped with an NVIDIA RTX 4500 Ada \(24 GB\) GPU and an AMD Ryzen Threadripper PRO 7985WX host with 256 GB of RAM\.

Evaluation Setup\.We evaluate each task with 5 random seeds\. Training budgets are task\-specific: 5M environment steps for GenCos, 20M for TSO, 3M for DSO, 10M for DERs, and 1M for DCMG\. We use PPO\[[36](https://arxiv.org/html/2609.36052#bib.bib30)\]for the single\-agent tasks \(TSO, DSO, and DCMG\), and IPPO\[[37](https://arxiv.org/html/2609.36052#bib.bib20)\]for the multi\-agent tasks \(GenCos and DERs\)\. DCMG is additionally trained with SAC\[[38](https://arxiv.org/html/2609.36052#bib.bib39)\]for the representative episode in Section[6\.4](https://arxiv.org/html/2609.36052#S6.SS4)\. After training, each seed is evaluated on a held\-out split, with task\-specific OOD stress splits when applicable\. For each seed, results are averaged over 10–50 evaluation episodes depending on the task, and final results are reported as mean values with 95% CIs across seeds\. Full hyperparameters and network architectures are in Appendix[F](https://arxiv.org/html/2609.36052#A6)\.

### 6\.1Execution Speed and Scaling Performance

We compare three training backends on the five benchmark tasks: \(i\) PowerZooJax, which runs both the environment simulation and policy training in JAX on the GPU; \(ii\) Stable\-Baselines3 \(SB3\)\[[39](https://arxiv.org/html/2609.36052#bib.bib40)\], which uses PowerZooPy444[https://github\.com/powerzoojax/PowerZooPy](https://github.com/powerzoojax/PowerZooPy), the companion object\-oriented Python implementation\., our CPU\-based Python implementation of the same tasks, for environment simulation and trains a PyTorch policy on the GPU; and \(iii\) SBX\[[40](https://arxiv.org/html/2609.36052#bib.bib3)\], the JAX port of SB3, which trains a JAX policy on the GPU while using the same CPU\-based PowerZooPy environment simulation\.

Figure 3:Cross\-backend speed for five benchmark tasks\. Top: wall\-clock training time under matched budgets\. Bottom: environment throughput with increasing parallelism, shown on log–log axes\.Table 2:Cross\-backend speed per benchmark task for three training backends\.Wall\-clock to matched budget \(s\)Throughput at2122^\{12\}envs \(steps/s\)TaskBudgetNenvsN\_\{\\text\{envs\}\}PowerZooJaxSBXSB3speedupPowerZooJaxSBXSB3speedupGenCos5M256679,60210,868161×\\times506,1344422442,077×\\timesTSO20M25651510,79910,80121×\\times53,7241,649411131×\\timesDSO3M128321,6673,801117×\\times134,6603,4352,20761×\\timesDERs10M1287961,88359,778780×\\times414,5062032332,040×\\timesDCMG1M64332212578×\\times25,1321,1671,29822×\\timesWall\-clock and Throughput\.Figure[3](https://arxiv.org/html/2609.36052#S6.F3)evaluates speed from two perspectives\. The wall\-clock measures how long each backend takes to finish the same training budget for each task, including episode resets, periodic evaluations, and gradient updates\. The throughput measures how many environment steps each backend can generate per second as the number of parallel environments increases from202^\{0\}to2122^\{12\}\. Table[2](https://arxiv.org/html/2609.36052#S6.T2)reports the final wall\-clock time and the throughput at2122^\{12\}parallel environments, and the corresponding speedup of PowerZooJax over the CPU\-based baselines\.

Speedup vs Per\-step Cost\.The largest gains occur when CPU baselines perform solver\-heavy work\. GenCos and DERs call CPU\-based OPF or PF routines in the CPU implementations, while PowerZooJax executes these kernels inside the compiled JAX graph, yielding throughput speedups above2,000×2\{,\}000\\times\. TSO and DSO implemented in JAX also benefit substantially from moving power system computation onto the GPU\. DCMG has the smallest speedup because its transitions are mostly arithmetic device updates without an external solver call, so the CPU baseline is already fast\.

### 6\.2GenCos: Bidding under Partial Observability

Figure[4](https://arxiv.org/html/2609.36052#S6.F4)evaluates GenCos under four bidding strategies and three evaluation splits\. The two extreme baselines define the strategic range: Truthful bidding at marginal cost achieves a daily profit of £6\.9k/day, while Max\-markup bidding at the maximum allowed price achieves £598k/day\. Uniform\-mid, which uses an intermediate markup between Truthful and Max\-markup bidding, achieves £303k/day\. IPPO learns a more profitable strategy than Uniform\-mid, reaching £395k/day on the in\-distribution split, while remaining below the Max\-markup upper baseline\. The same ordering is preserved under the two OOD stress splits: high demand, which increases load by 10%, and low renewable, which decreases wind generation by 5%\. The learned strategy also produces distinct locational marginal prices \(LMPs\)\. When system load exceeds 800 MW, Max\-markup clears at an LMP near £75/MWh, Uniform\-mid near £50/MWh, and IPPO around £45/MWh, indicating that IPPO learns to increase market profit while avoiding the most aggressive markup strategy\.

Figure 4:Generation company profit and market clearing price oncase5\. \(a\) Total daily profit for each strategy and evaluation split\. \(b\) Locational marginal price \(LMP\) vs system load on a representative evaluation\. Error bars in \(a\) show 95% CIs across 5 seeds\.
### 6\.3TSO: Economic–Safety Cost Trade\-off

Figure[5](https://arxiv.org/html/2609.36052#S6.F5)shows the economic–safety cost trade\-off for transmission operation\. The merit\-order baseline, which commits generators from cheapest to most expensive until capacity covers demand plus a 5% reserve margin, while ignoring thermal line limits, achieves an operating cost of £2\.89M/day\. The all\-on baseline keeps every generator online, improving safety but increasing cost to £4\.34M/day\. PPO, trained with negative operating cost as the reward, reduces to £1\.58M/day but violates thermal line limits in 36% and reserve margins in 7% over all evaluations\. PPO\-Lagrangian\[[41](https://arxiv.org/html/2609.36052#bib.bib33)\], which applies adaptive penalties to each violation type, eliminates reserve shortfalls and reduces the thermal overload rate to 6%, at a higher cost of £3\.50M/day\. The same pattern holds under the two OOD stress splits: line\-tightening, where line ratings are reduced to 85%, and load\-stress, where demand is scaled by 1\.15\. Under these stresses, PPO reaches thermal overload rates of 49% and 41%, respectively, while PPO\-Lagrangian reduces them to 16% and 9% and maintains zero reserve shortfall\. These results show that the Lagrangian formulation substantially improves safety, fully eliminating reserve violations on this network, but thermal overloads remain a harder constraint to satisfy under stressed conditions\.

### 6\.4DCMG: Long\-Horizon Scheduling for a Data Center

Figure[6](https://arxiv.org/html/2609.36052#S6.F6)shows a representative 24\-hour SAC policy trajectory over 288 five\-minute steps\. SAC charges the battery from the grid during the first two hours when electricity prices are low, discharges the battery through the morning price peak above £100/MWh, recharges during the midday when PV generation is available, and cycles again through the evening peak\. This repeated charge–discharge pattern indicates that SAC learns to coordinate storage schedules across the full 288\-step horizon\. Throughout the episode, the IT workload is served at every step, showing that SAC learns a price\-aware and constraint\-respecting scheduling policy directly from the multi\-objective reward\.

Figure 5:TSO economic–safety cost trade\-off oncase118\. \(a\) Operating cost vs thermal overload rate\. \(b\) Operating cost vs reserve shortfall rate\. Error bars show 95% CIs across 5 seeds\.Figure 6:DCMG daily scheduling on a representative day\. \(a\) Power supply to the IT workload\. \(b\) Daily battery state of charge \(SoC\) and grid electricity price\.

## 7Conclusion

PowerZooJax shows that power system RL benchmarks can simultaneously balance physical modeling and computational scalability\. By integrating training and evaluation into a GPU\-accelerated pipeline, the benchmark supports fast rollouts while retaining operational metrics that matter for real grid decision making\. The five tasks demonstrate how RL algorithms can be evaluated across diverse power system settings under a common constrained Markov decision process \(CMDP\) interface, with standardized reporting of returns, safety costs, and stress\-test performance\.

Limitations\.PowerZooJax is a benchmark suite, not a deployment\-ready power system simulator or controller\. Although it covers five representative tasks, the current version is still limited in grid scale, uncertainty modeling, operational constraints, and task diversity\. Finally, the current environments are built from public or synthetic grid models and public data, and therefore cannot fully capture the complexity of real power system with real\-time operational data\. Strong performance on PowerZooJax should therefore be interpreted as evidence of an algorithm’s ability to address representative power system RL challenges, not as a guarantee of safe real\-world power system operation\.

Future Work\.Future work will extend PowerZooJax in three directions\. First, we plan to broaden PowerZooJax with larger grid cases, richer uncertainty models, more realistic constraints, and additional tasks\. Second, we envision PowerZooJax evolving into a more intelligent programming tool for power system RL\. By integrating PowerZooJax with language models, users could specify power system requirements in natural language, and the tool could automatically generate the corresponding RL environment, task configuration, reward and safety cost definitions, and baseline algorithms\. Third, we aim to collaborate with the power industry to access grid models and operational data, build digital\-twin environments, and study how AI\-based control methods can be evaluated and transferred toward practical power system applications\.

## Acknowledgement

Dawei Qiu is supported by the Nanyang Technological University \(NTU\) Start\-Up Grant project \#026749\-00001 “Market Design for Low\-Carbon Power Systems: Towards a Reliable, Affordable, and Resilient Transition”\. Jianhong Wang is supported by the Engineering and Physical Sciences Research Council \(EPSRC\) \[Grant Ref: EP/Y028732/1\]\.

## References

- \[1\]\(2018\)How Decarbonization, Digitalization and Decentralization are changing key power infrastructures\.Renewable and Sustainable Energy Reviews93,pp\. 483–498\.External Links:ISSN 1364\-0321,LCCN 1Cited by:[§1](https://arxiv.org/html/2609.36052#S1.p1.1)\.
- \[2\]G\. Strbac, D\. Papadaskalopoulos, N\. Chrysanthopoulos, A\. Estanqueiro, H\. Algarvio, F\. Lopes, L\. de Vries, G\. Morales\-Espana, J\. Sijm, R\. Hernandez\-Serna, J\. Kiviluoma, and N\. Helisto\(2021\)Decarbonization of Electricity Systems in Europe: Market Design Challenges\.IEEE Power and Energy Magazine19\(1\),pp\. 53–63\.External Links:ISSN 1540\-7977Cited by:[§1](https://arxiv.org/html/2609.36052#S1.p1.1)\.
- \[3\]E\. Strubell, A\. Ganesh, and A\. McCallum\(2019\)Energy and Policy Considerations for Deep Learning in NLP\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,Florence, Italy,pp\. 3645–3650\.Cited by:[§1](https://arxiv.org/html/2609.36052#S1.p1.1)\.
- \[4\]T\. Xiao, F\. F\. Nerini, H\. D\. Matthews, M\. Tavoni, and F\. You\(2025\)Environmental impact and net\-zero pathways for sustainable artificial intelligence servers in the USA\.Nature Sustainability8\(12\),pp\. 1541–1553\.External Links:ISSN 2398\-9629,LCCN 1Cited by:[§1](https://arxiv.org/html/2609.36052#S1.p1.1)\.
- \[5\]W\. B\. Powell and S\. Meisel\(2016\)Tutorial on Stochastic Optimization in Energy—Part I: Modeling and Policies\.IEEE Transactions on Power Systems31\(2\),pp\. 1459–1467\.External Links:ISSN 0885\-8950,LCCN 1Cited by:[§1](https://arxiv.org/html/2609.36052#S1.p2.1)\.
- \[6\]Y\. Chen, F\. Pan, F\. Qiu, A\. S\. Xavier, T\. Zheng, M\. Marwali, B\. Knueven, Y\. Guan, P\. B\. Luh, L\. Wu, B\. Yan, M\. A\. Bragin, H\. Zhong, A\. Giacomoni, R\. Baldick, B\. Gisin, Q\. Gu, R\. Philbrick, and F\. Li\(2023\)Security\-Constrained Unit Commitment for Electricity Market: Modeling, Solution Methods, and Future Challenges\.IEEE Transactions on Power Systems38\(5\),pp\. 4668–4681\.External Links:ISSN 0885\-8950,LCCN 1Cited by:[§1](https://arxiv.org/html/2609.36052#S1.p2.1)\.
- \[7\]Y\. Yang, T\. Nishikawa, and A\. E\. Motter\(2017\)Small vulnerable sets determine large network cascades in power grids\.Science358\(6365\),pp\. eaan3184\.External Links:ISSN 0036\-8075,LCCN 1Cited by:[§1](https://arxiv.org/html/2609.36052#S1.p2.1)\.
- \[8\]H\. Wu, X\. Meng, M\. M\. Danziger, S\. P\. Cornelius, H\. Tian, and A\. Barabási\(2022\)Fragmentation of outage clusters during the recovery of power distribution grids\.Nature Communications13\(1\),pp\. 7372\.External Links:ISSN 2041\-1723,LCCN 1Cited by:[§1](https://arxiv.org/html/2609.36052#S1.p2.1)\.
- \[9\]B\. Kroposki, B\. Johnson, Y\. Zhang, V\. Gevorgian, P\. Denholm, B\. Hodge, and B\. Hannegan\(2017\)Achieving a 100% Renewable Grid: Operating Electric Power Systems with Extremely High Levels of Variable Renewable Energy\.IEEE Power and Energy Magazine15\(2\),pp\. 61–73\.External Links:ISSN 1540\-7977Cited by:[§1](https://arxiv.org/html/2609.36052#S1.p2.1)\.
- \[10\]R\. T\. Lange\(2022\)Gymnax: a JAX\-based reinforcement learning environment library\.External Links:[Link](http://github.com/RobertTLange/gymnax)Cited by:[§1](https://arxiv.org/html/2609.36052#S1.p4.1),[§2](https://arxiv.org/html/2609.36052#S2.p3.1)\.
- \[11\]S\. Bandyopadhyay, J\. Cook, C\. de Witt, B\. Ellis, J\. Foerster, M\. Gallici, R\. Hammond, N\. Hawes, G\. Ingvarsson, M\. Jiang, A\. Khan, B\. Lacerda, R\. Lange, C\. Lu, A\. Lupu, T\. Rocktäschel, A\. Rutherford, M\. Samvelyan, A\. Souly, S\. Whiteson, and T\. Willi\(2024\)JaxMARL: Multi\-Agent RL Environments and Algorithms in JAX\.InAdvances in Neural Information Processing Systems 37,pp\. 50925–50951\.External Links:ISSN 1558\-2914Cited by:[§1](https://arxiv.org/html/2609.36052#S1.p4.1),[§2](https://arxiv.org/html/2609.36052#S2.p4.1)\.
- \[12\]J\. Bradbury, R\. Frostig, P\. Hawkins, M\. J\. Johnson, Y\. Katariya, C\. Leary, D\. Maclaurin, G\. Necula, A\. Paszke, J\. VanderPlas, S\. Wanderman\-Milne, and Q\. Zhang\(2018\)JAX: composable transformations of Python\+NumPy programs\.External Links:[Link](http://github.com/jax-ml/jax)Cited by:[§1](https://arxiv.org/html/2609.36052#S1.p4.1)\.
- \[13\]T\. Fan, X\. Y\. Lee, and Y\. Wang\(2022\)PowerGym: A Reinforcement Learning Environment for Volt\-Var Control in Power Distribution Systems\.InProceedings of The 4th Annual Learning for Dynamics and Control Conference,pp\. 21–33\.External Links:ISSN 2640\-3498Cited by:[§1](https://arxiv.org/html/2609.36052#S1.p4.1)\.
- \[14\]C\. Yeh, V\. Li, R\. Datta, J\. Arroyo, N\. Christianson, C\. Zhang, Y\. Chen, M\. M\. Hosseini, A\. Golmohammadi, Y\. Shi, Y\. Yue, and A\. Wierman\(2023\)SustainGym: Reinforcement Learning Environments for Sustainable Energy Systems\.Advances in Neural Information Processing Systems36,pp\. 59464–59476\.Cited by:[§1](https://arxiv.org/html/2609.36052#S1.p4.1),[§2](https://arxiv.org/html/2609.36052#S2.p1.1)\.
- \[15\]E\. Marchesini, B\. Donnot, C\. Crozier, I\. Dytham, C\. Merz, L\. Schewe, N\. Westerbeck, C\. Wu, A\. Marot, and P\. L\. Donti\(2025\)RL2Grid: Benchmarking Reinforcement Learning in Power Grid Operations\.Preprint,arXiv\.External Links:2503\.23101Cited by:[§1](https://arxiv.org/html/2609.36052#S1.p4.1),[§2](https://arxiv.org/html/2609.36052#S2.p1.1)\.
- \[16\]J\. Wang, W\. Xu, Y\. Gu, W\. Song, and T\. C\. Green\(2021\)Multi\-agent reinforcement learning for active voltage control on power distribution networks\.Advances in neural information processing systems34,pp\. 3271–3284\.Cited by:[§2](https://arxiv.org/html/2609.36052#S2.p1.1)\.
- \[17\]H\. Cui and Y\. Zhang\(2022\)Andes\_gym: A Versatile Environment for Deep Reinforcement Learning in Power Systems\.In2022 IEEE Power & Energy Society General Meeting \(PESGM\),pp\. 01–05\.External Links:ISSN 1944\-9933Cited by:[§2](https://arxiv.org/html/2609.36052#S2.p1.1)\.
- \[18\]D\. Biagioni, X\. Zhang, D\. Wald, D\. Vaidhynathan, R\. Chintala, J\. King, and A\. S\. Zamzam\(2022\)PowerGridworld\.InProceedings of the Thirteenth ACM International Conference on Future Energy Systems,E\-Energy ’22,New York, NY, USA,pp\. 565–570\.External Links:ISBN 978\-1\-4503\-9397\-3Cited by:[§2](https://arxiv.org/html/2609.36052#S2.p1.1)\.
- \[19\]E\. Marchesini, E\. Boguslawski, A\. Leite, C\. Amato, M\. Dussartre, M\. Schoenauer, B\. Donnot, and P\. L\. Donti\(2026\)MARL2Grid\-TR: A Multi\-Agent RL Benchmark in Power Grid Operations\.InThe Fourteenth International Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2609.36052#S2.p1.1)\.
- \[20\]K\. Nweye, K\. Kaspar, G\. Buscemi, G\. Pinto, H\. Li, T\. Hong, M\. Ouf, A\. Capozzoli, and Z\. Nagy\(2023\)CityLearn v2: An OpenAI Gym environment for demand response control benchmarking in grid\-interactive communities\.InProceedings of the 10th ACM International Conference on Systems for Energy\-Efficient Buildings, Cities, and Transportation,BuildSys ’23,New York, NY, USA,pp\. 274–275\.External Links:ISBN 979\-8\-4007\-0230\-3Cited by:[§2](https://arxiv.org/html/2609.36052#S2.p1.1)\.
- \[21\]A\. Pigott, C\. Crozier, K\. Baker, and Z\. Nagy\(2022\)GridLearn: Multiagent reinforcement learning for grid\-aware building energy management\.Electric Power Systems Research213,pp\. 108521\.External Links:ISSN 0378\-7796,LCCN 3Cited by:[§2](https://arxiv.org/html/2609.36052#S2.p1.1)\.
- \[22\]G\. Henri, T\. Levent, A\. Halev, R\. Alami, and P\. Cordier\(2020\)Pymgrid: An Open\-Source Python Microgrid Simulator for Applied Artificial Intelligence Research\.arXiv\.External Links:2011\.08004Cited by:[§2](https://arxiv.org/html/2609.36052#S2.p1.1)\.
- \[23\]M\. Eichelbeck, H\. Markgraf, and M\. Althoff\(2025\)CommonPower: A Framework for Safe Data\-Driven Smart Grid Control\.IEEE Transactions on Smart Grid,pp\. 1–1\.External Links:ISSN 1949\-3053Cited by:[§2](https://arxiv.org/html/2609.36052#S2.p1.1)\.
- \[24\]C\. D\. Freeman, E\. Frey, A\. Raichuk, S\. Girgin, I\. Mordatch, and O\. Bachem\(2021\)Brax \- a differentiable physics engine for large scale rigid body simulation\.External Links:[Link](http://github.com/google/brax)Cited by:[§2](https://arxiv.org/html/2609.36052#S2.p3.1)\.
- \[25\]C\. Bonnet, D\. Luo, D\. Byrne, S\. Surana, S\. Abramowitz, P\. Duckworth, V\. Coyette, L\. I\. Midgley, E\. Tegegn, T\. Kalloniatis, O\. Mahjoub, M\. Macfarlane, A\. P\. Smit, N\. Grinsztajn, R\. Boige, C\. N\. Waters, M\. A\. Mimouni, U\. A\. M\. Sob, R\. de Kock, S\. Singh, D\. Furelos\-Blanco, V\. Le, A\. Pretorius, and A\. Laterre\(2024\)Jumanji: a diverse suite of scalable reinforcement learning environments in JAX\.External Links:2306\.09884Cited by:[§2](https://arxiv.org/html/2609.36052#S2.p3.1)\.
- \[26\]S\. Koyamada, S\. Okano, S\. Nishimori, Y\. Murata, K\. Habara, H\. Kita, and S\. Ishii\(2023\)Pgx: Hardware\-Accelerated Parallel Game Simulators for Reinforcement Learning\.InAdvances in Neural Information Processing Systems 36,Vol\.36,pp\. 45716–45743\.Cited by:[§2](https://arxiv.org/html/2609.36052#S2.p3.1)\.
- \[27\]M\. Matthews, M\. Beukman, B\. Ellis, M\. Samvelyan, M\. Jackson, S\. Coward, and J\. Foerster\(2024\)Craftax: A Lightning\-Fast Benchmark for Open\-Ended Reinforcement Learning\.InProceedings of the 41st International Conference on Machine Learning,Vol\.235\.Cited by:[§2](https://arxiv.org/html/2609.36052#S2.p4.1)\.
- \[28\]W\. Radji, T\. Michel, and H\. Piteau\(2025\)Octax: Accelerated CHIP\-8 arcade environments for reinforcement learning in JAX\.External Links:2510\.01764Cited by:[§2](https://arxiv.org/html/2609.36052#S2.p4.1)\.
- \[29\]X\. Chen, G\. Qu, Y\. Tang, S\. Low, and N\. Li\(2022\)Reinforcement Learning for Selective Key Applications in Power Systems: Recent Advances and Future Challenges\.IEEE Transactions on Smart Grid13\(4\),pp\. 2935–2958\.External Links:ISSN 1949\-3053,LCCN 1Cited by:[§A\.1](https://arxiv.org/html/2609.36052#A1.SS1.p1.1),[§3](https://arxiv.org/html/2609.36052#S3.p2.1)\.
- \[30\]\(2019\)Google/cluster\-data\.Note:GoogleExternal Links:[Link](https://github.com/google/cluster-data)Cited by:[§C\.5](https://arxiv.org/html/2609.36052#A3.SS5.p1.1),[§3](https://arxiv.org/html/2609.36052#S3.p7.1)\.
- \[31\]A\. Marot, B\. Donnot, G\. Dulac\-Arnold, A\. Kelly, A\. O’Sullivan, J\. Viebahn, M\. Awad, and I\. Guyon\(2021\)Learning to run a Power Network Challenge: a Retrospective Analysis\.Preprint,arXiv\.External Links:2103\.03104Cited by:[§5](https://arxiv.org/html/2609.36052#S5.p1.1)\.
- \[32\]L\. Thurner, A\. Scheidler, F\. Schafer, J\. Menke, J\. Dollichon, F\. Meier, S\. Meinecke, and M\. Braun\(2018\)Pandapower—An Open\-Source Python Tool for Convenient Modeling, Analysis, and Optimization of Electric Power Systems\.IEEE Transactions on Power Systems33\(6\),pp\. 6510–6521\.External Links:ISSN 0885\-8950,LCCN 1Cited by:[§5\.1](https://arxiv.org/html/2609.36052#S5.SS1.p1.1)\.
- \[33\]M\.E\. Baran and F\.F\. Wu\(1989\)Network reconfiguration in distribution systems for loss reduction and load balancing\.IEEE Transactions on Power Delivery4\(2\),pp\. 1401–1407\.External Links:ISSN 0885\-8977,LCCN 3Cited by:[§B\.1](https://arxiv.org/html/2609.36052#A2.SS1.p1.1),[§5\.1](https://arxiv.org/html/2609.36052#S5.SS1.p2.3)\.
- \[34\]T\. Erseghe\(2014\)Distributed Optimal Power Flow Using ADMM\.IEEE Transactions on Power Systems29\(5\),pp\. 2370–2380\.External Links:ISSN 0885\-8950,LCCN 1Cited by:[§B\.2](https://arxiv.org/html/2609.36052#A2.SS2.p2.2),[§5\.2](https://arxiv.org/html/2609.36052#S5.SS2.p1.2)\.
- \[35\]H\. Wang, C\. E\. Murillo\-Sanchez, R\. D\. Zimmerman, R\. J\. Thomas, and H\. Wang\(2007\)On Computational Issues of Market\-Based Optimal Power Flow\.IEEE Transactions on Power Systems22\(3\),pp\. 1185–1193\.External Links:ISSN 0885\-8950,LCCN 1Cited by:[§B\.3](https://arxiv.org/html/2609.36052#A2.SS3.p1.2),[§5\.2](https://arxiv.org/html/2609.36052#S5.SS2.p1.2)\.
- \[36\]J\. Schulman, F\. Wolski, P\. Dhariwal, A\. Radford, and O\. Klimov\(2017\)Proximal Policy Optimization Algorithms\.Preprint,arXiv\.External Links:1707\.06347Cited by:[1st item](https://arxiv.org/html/2609.36052#A6.I1.i1.p1.1),[§6](https://arxiv.org/html/2609.36052#S6.p3.1)\.
- \[37\]C\. S\. de Witt, T\. Gupta, D\. Makoviichuk, V\. Makoviychuk, P\. H\. S\. Torr, M\. Sun, and S\. Whiteson\(2020\)Is Independent Learning All You Need in the StarCraft Multi\-Agent Challenge?\.arXiv\.External Links:2011\.09533Cited by:[2nd item](https://arxiv.org/html/2609.36052#A6.I1.i2.p1.1),[§6](https://arxiv.org/html/2609.36052#S6.p3.1)\.
- \[38\]T\. Haarnoja, A\. Zhou, K\. Hartikainen, G\. Tucker, S\. Ha, J\. Tan, V\. Kumar, H\. Zhu, A\. Gupta, P\. Abbeel, and S\. Levine\(2018\)Soft Actor\-Critic Algorithms and Applications\.Preprint,arXiv\.External Links:1812\.05905Cited by:[1st item](https://arxiv.org/html/2609.36052#A6.I1.i1.p1.1),[§6](https://arxiv.org/html/2609.36052#S6.p3.1)\.
- \[39\]A\. Raffin, A\. Hill, A\. Gleave, A\. Kanervisto, M\. Ernestus, and N\. Dormann\(2021\)Stable\-baselines3: Reliable reinforcement learning implementations\.Journal of Machine Learning Research22\(268\),pp\. 1–8\.External Links:LCCN 3 \(CCF A\)Cited by:[§6\.1](https://arxiv.org/html/2609.36052#S6.SS1.p1.1)\.
- \[40\]A\. Raffin\(2026\)Stable Baselines Jax\.External Links:[Link](https://github.com/araffin/sbx)Cited by:[§6\.1](https://arxiv.org/html/2609.36052#S6.SS1.p1.1)\.
- \[41\]A\. Ray, J\. Achiam, and D\. Amodei\(2019\)Benchmarking safe exploration in deep reinforcement learning\.7\(1\),pp\. 2\.Cited by:[§6\.3](https://arxiv.org/html/2609.36052#S6.SS3.p1.1)\.
- \[42\]J\. Achiam, D\. Held, A\. Tamar, and P\. Abbeel\(2017\)Constrained Policy Optimization\.Preprint,arXiv\.External Links:1705\.10528Cited by:[§A\.1](https://arxiv.org/html/2609.36052#A1.SS1.p1.1)\.
- \[43\]A\. Sootla, A\. I\. Cowen\-Rivers, T\. Jafferjee, Z\. Wang, D\. Mguni, J\. Wang, and H\. Bou\-Ammar\(2022\)Saute RL: Almost Surely Safe Reinforcement Learning Using State Augmentation\.Preprint,arXiv\.External Links:2202\.06558Cited by:[§A\.3](https://arxiv.org/html/2609.36052#A1.SS3.p1.2),[1st item](https://arxiv.org/html/2609.36052#A6.I1.i1.p1.1)\.

## Appendix

This supplementary material is organized as follows\. Section[A](https://arxiv.org/html/2609.36052#A1)formalizes the CMDP tuple shared by all five tasks and fixes notation\. Section[B](https://arxiv.org/html/2609.36052#A2)gives the physical\-model derivations \(radial AC power flow, DC OPF with ADMM, bid\-based market clearing with PDIPM, and the device plus microgrid couplings\)\. Section[C](https://arxiv.org/html/2609.36052#A3)gives the per\-task observation, action, reward, and cost equations\. Section[D](https://arxiv.org/html/2609.36052#A4)documents external data sources and the split and stress\-test taxonomy\. Section[E](https://arxiv.org/html/2609.36052#A5)provides a quickstart with install, a minimal Python rollout, configuration excerpts, and CLI commands\. Section[F](https://arxiv.org/html/2609.36052#A6)lists algorithm choices, hyperparameters, network architectures, and evaluation protocol\. Section[G](https://arxiv.org/html/2609.36052#A7)reports the per\-task supplementary tables and figures\. Section[H](https://arxiv.org/html/2609.36052#A8)reports execution scaling and the backend comparison\.

## Appendix AConstrained Markov Decision Process Formulation

This section gives the formal constrained Markov decision process \(CMDP\) tuple shared by all five tasks, the two constraint forms used in the benchmark \(discounted expected cost and zero\-violation constraint\), and the Lagrangian formulation used by the safe\-RL baselines\. The notation used throughout the appendix is fixed in Table[3](https://arxiv.org/html/2609.36052#A1.T3)\(Section[A\.5](https://arxiv.org/html/2609.36052#A1.SS5)\)\.

### A\.1The CMDP tuple

Each task is an episodic CMDP\[[29](https://arxiv.org/html/2609.36052#bib.bib34),[42](https://arxiv.org/html/2609.36052#bib.bib9)\]

ℳ=\(𝒮,𝒜,P,r,𝐜,γ,𝐛,ρ0,T\),\\mathcal\{M\}=\\bigl\(\\mathcal\{S\},\\ \\mathcal\{A\},\\ P,\\ r,\\ \\mathbf\{c\},\\ \\gamma,\\ \\mathbf\{b\},\\ \\rho\_\{0\},\\ T\\bigr\),\(3\)where𝒮\\mathcal\{S\}and𝒜\\mathcal\{A\}are the state and action spaces,P:𝒮×𝒜→Δ⁡\(𝒮\)P:\\mathcal\{S\}\\times\\mathcal\{A\}\\to\\Delta\(\\mathcal\{S\}\)is the transition kernel,r:𝒮×𝒜→ℝr:\\mathcal\{S\}\\times\\mathcal\{A\}\\to\\mathbb\{R\}is the scalar reward,𝐜=\(c1,…,ck\):𝒮×𝒜→ℝk\\mathbf\{c\}=\(c\_\{1\},\\dots,c\_\{k\}\):\\mathcal\{S\}\\times\\mathcal\{A\}\\to\\mathbb\{R\}^\{k\}is the cost vector,γ∈\(0,1\]\\gamma\\in\(0,1\]is the discount factor,𝐛∈ℝ≥0k\\mathbf\{b\}\\in\\mathbb\{R\}\_\{\\geq 0\}^\{k\}is the constraint\-threshold vector,ρ0\\rho\_\{0\}is the initial\-state distribution, andTTis the episode horizon \(Table[5](https://arxiv.org/html/2609.36052#A3.T5)lists per\-taskTT\)\. Reward and cost are produced as separate outputs by every task environment so that safety violations cannot be silently absorbed into the scalar return; the cost\-vector order is fixed by the declared constraint names of each task\. In the multi\-agent tasks \(GenCos and DERs\), each agent observes only a projection of the state \(Section[C](https://arxiv.org/html/2609.36052#A3)\), so these tasks are constrained partially observable problems\.

For a stationary policyπ:𝒮→Δ⁡\(𝒜\)\\pi:\\mathcal\{S\}\\to\\Delta\(\\mathcal\{A\}\), define the reward and cost values

Jr​\(π\)\\displaystyle J\_\{r\}\(\\pi\)=𝔼π​\[∑t=0T−1γt​r​\(st,at\)\],\\displaystyle\\;=\\;\\mathbb\{E\}\_\{\\pi\}\\Bigl\[\\textstyle\\sum\_\{t=0\}^\{T\-1\}\\gamma^\{t\}\\,r\(s\_\{t\},a\_\{t\}\)\\Bigr\],\(4\)Jci​\(π\)\\displaystyle J\_\{c\_\{i\}\}\(\\pi\)=𝔼π\[∑t=0T−1γtci\(st,at\)\],i=1,…,k\.\\displaystyle\\;=\\;\\mathbb\{E\}\_\{\\pi\}\\Bigl\[\\textstyle\\sum\_\{t=0\}^\{T\-1\}\\gamma^\{t\}\\,c\_\{i\}\(s\_\{t\},a\_\{t\}\)\\Bigr\],\\quad i=1,\\dots,k\.\(5\)

### A\.2Constraint forms used in the suite

PowerZooJax tasks use one of two constraint forms\.Form 1 \(budgeted expected cost\)is applied to the DSO voltage channel, the three DERs channels \(voltage, thermal, resource\), and the three DCMG channels \(workload incompletion, server\-zone over\-temperature, power imbalance\), as well as the TSO Lagrangian baselines\.Form 2 \(zero\-violation constraint\)is applied to the TSO reserve\-shortfall and thermal\-overload channels reported in Table[11](https://arxiv.org/html/2609.36052#A7.T11)\. The single GenCos thermal\-overload channel is reported as a market\-side diagnostic only\.

Form 1 \(discounted expected\-cost budget\)\.

maxπ⁡Jr​\(π\)s\.t\.Jci​\(π\)≤bi,i=1,…,k\.\\max\_\{\\pi\}\\;J\_\{r\}\(\\pi\)\\quad\\text\{s\.t\.\}\\quad J\_\{c\_\{i\}\}\(\\pi\)\\leq b\_\{i\},\\ i=1,\\dots,k\.\(6\)Used by DSO, DERs, and TSO Lagrangian baselines, where thebib\_\{i\}are operational risk budgets per episode\.

Form 2 \(zero\-violation constraint\)\.

maxπJr\(π\)s\.t\.ℙπ\[ci\(st,at\)\>0\]=0,∀t,∀i∈Ihard\.\\max\_\{\\pi\}\\;J\_\{r\}\(\\pi\)\\quad\\text\{s\.t\.\}\\quad\\mathbb\{P\}\_\{\\pi\}\\\!\\bigl\[c\_\{i\}\(s\_\{t\},a\_\{t\}\)\>0\\bigr\]=0,\\ \\forall t,\\ \\forall i\\in I\_\{\\text\{hard\}\}\.\(7\)Used for the TSO reserve\-shortfall and thermal\-overload channels: a policy satisfies Form 2 only if both channels record zero cost at every step, so cost cannot be traded for safety\. The TSO results in Section[G\.2](https://arxiv.org/html/2609.36052#A7.SS2)are reported under this form\.

### A\.3Lagrangian formulation

For Form 1 with multipliers𝝀∈ℝ≥0k\\boldsymbol\{\\lambda\}\\in\\mathbb\{R\}\_\{\\geq 0\}^\{k\}, the Lagrangian

ℒ⁡\(π,𝝀\)=Jr​\(π\)−𝝀⊤​\(Jc​\(π\)−𝐛\)\\mathcal\{L\}\(\\pi,\\boldsymbol\{\\lambda\}\)\\;=\\;J\_\{r\}\(\\pi\)\\;\-\\;\\boldsymbol\{\\lambda\}^\{\\\!\\top\}\\bigl\(J\_\{c\}\(\\pi\)\-\\mathbf\{b\}\\bigr\)\(8\)gives the saddle\-point reformulationmaxπ⁡min𝝀≥0⁡ℒ⁡\(π,𝝀\)\\max\_\{\\pi\}\\min\_\{\\boldsymbol\{\\lambda\}\\geq 0\}\\mathcal\{L\}\(\\pi,\\boldsymbol\{\\lambda\}\)\. PPO\-Lagrangian and IPPO\-Lagrangian baselines update𝝀\\boldsymbol\{\\lambda\}by stochastic ascent onJc​\(π\)−𝐛J\_\{c\}\(\\pi\)\-\\mathbf\{b\}whileπ\\piis updated by PPO clipped\-objective ascent onℒ\\mathcal\{L\}\. Sauté PPO\[[43](https://arxiv.org/html/2609.36052#bib.bib36)\]instead augments the observation with the per\-channel remaining cost budgetbi−∑t′<tγt′​ci​\(st′,at′\)b\_\{i\}\-\\sum\_\{t^\{\\prime\}<t\}\\gamma^\{t^\{\\prime\}\}c\_\{i\}\(s\_\{t^\{\\prime\}\},a\_\{t^\{\\prime\}\}\)and trains plain PPO on the augmented state\.

### A\.4Per\-task instantiation

Each task instantiatesℳ\\mathcal\{M\}through aTaskSpec/ConstraintSpecpair \(theTaskSpecProtocol surface is shown in Section[E\.2](https://arxiv.org/html/2609.36052#A5.SS2)\)\. Concrete observation, action, reward, and cost equations are given in the task cards of Section[C](https://arxiv.org/html/2609.36052#A3); per\-task safety thresholds are part of each task’s configuration file, and split definitions are listed in Section[D](https://arxiv.org/html/2609.36052#A4)\.

### A\.5Notation

Table 3:Notation used across the main paper and the appendix\.SymbolMeaning*CMDP and policy*𝒮,𝒜\\mathcal\{S\},\\,\\mathcal\{A\}State and action spaces of one task\.PPTransition kernel;*deterministic given exogenous time\-series*\.r,𝐜∈ℝkr,\\,\\mathbf\{c\}\\in\\mathbb\{R\}^\{k\}Scalar reward and per\-step cost vector withkktask\-specific entries\.γ,T\\gamma,\\,TDiscount factor and episode horizon \(T=48T=48orT=288T=288\)\.𝐛∈ℝ≥0k\\mathbf\{b\}\\in\\mathbb\{R\}\_\{\\geq 0\}^\{k\}Cost thresholds \(Form 1\) or zero\-violation set \(Form 2\)\.π,ρ0\\pi,\\,\\rho\_\{0\}Policy and initial\-state distribution\.𝝀∈ℝ≥0k\\boldsymbol\{\\lambda\}\\in\\mathbb\{R\}\_\{\\geq 0\}^\{k\}Lagrange multiplier vector\.*Power system*𝐏,𝐐\\mathbf\{P\},\\,\\mathbf\{Q\}Branch active and reactive power flows\.𝐯2\\mathbf\{v\}^\{2\}Vector of squared bus voltage magnitudes \(DistFlow uses squared form\)\.𝐃∈ℝNℓ×N\\mathbf\{D\}\\in\\mathbb\{R\}^\{N\_\{\\ell\}\\times N\}Downstream\-aggregation matrix \(radial PF\)\.𝚷∈ℝN×Nℓ\\boldsymbol\{\\Pi\}\\in\\mathbb\{R\}^\{N\\times N\_\{\\ell\}\}Path matrix from slack bus to each bus \(radial PF\)\.𝐇,𝐇d,𝐇bus\\mathbf\{H\},\\,\\mathbf\{H\}\_\{d\},\\,\\mathbf\{H\}\_\{\\text\{bus\}\}Generator\-side, demand\-side, and full bus\-level PTDF \(DC OPF\)\.AaugA\_\{\\text\{aug\}\}Augmented KKT matrix in ADMM DC OPF\.𝐩,𝐩¯,𝐩¯\\mathbf\{p\},\\,\\underline\{\\mathbf\{p\}\},\\,\\overline\{\\mathbf\{p\}\}Generator dispatch vector and its lower and upper output bounds \(DC OPF\)\.𝐳1,𝐳2\\mathbf\{z\}\_\{1\},\\,\\mathbf\{z\}\_\{2\}ADMM auxiliary variables for the generator\-bound and line\-flow constraints\.𝐲1,𝐲2\\mathbf\{y\}\_\{1\},\\,\\mathbf\{y\}\_\{2\}ADMM dual variables for the generator\-bound and line\-flow constraints\.𝐫\(k\)\\mathbf\{r\}^\{\(k\)\}ADMM residual vector at iterationkk\(right\-hand side of the augmented KKT solve\)\.ρ,μ\\rho,\\,\\muADMM penalty parameter and PDIPM barrier parameter\.*JAX primitives used in PowerZooJax*jax\.jitCompiles the full rollout \(reset, step, reward, cost\) into a single GPU program\.jax\.vmapRuns the same environment in parallel across batched scenarios\.jax\.lax\.scanUnrolls the per\-episode loop \(48 or 288 steps\) on the device\.jax\.lax\.while\_loopDrives the radial PF fixed\-point iteration to convergence on the device\.jax\.lax\.fori\_loopRuns a fixed number of ADMM or PDIPM solver iterations on the device\.PRNG keyRandom seed passed toreset/stepfor reproducible scenario and policy sampling\.

## Appendix BPhysical Models

The five tasks share five physical kernels \(Table[4](https://arxiv.org/html/2609.36052#A2.T4)\)\. Main paper Sections 5\.1–5\.2 give one equation each for radial PF and DC OPF; this appendix gives the full derivations and shows how each kernel becomes a fixed\-shape JAX computation graph\. Section[B\.4](https://arxiv.org/html/2609.36052#A2.SS4)covers the device dynamics shared by DSO, DERs, and DCMG, and Section[B\.5](https://arxiv.org/html/2609.36052#A2.SS5)covers the DCMG couplings\.

Table 4:Physical kernels reused across tasks\.### B\.1Radial AC power flow as a JAX graph

Standard DistFlow recurrence\.Consider a radial feeder withNNbuses andNℓ=N−1N\_\{\\ell\}=N\-1branches, slack bus00, and per\-branch resistance and reactancerℓ,xℓr\_\{\\ell\},x\_\{\\ell\}\. Each branchℓ\\ellconnects parent busp⁡\(ℓ\)p\(\\ell\)to child busc⁡\(ℓ\)c\(\\ell\)\. Letpnload,qnloadp^\{\\mathrm\{load\}\}\_\{n\},q^\{\\mathrm\{load\}\}\_\{n\}be the active and reactive load at busnn, andPℓ,QℓP\_\{\\ell\},Q\_\{\\ell\}be the active and reactive flow on branchℓ\\ell\(positive in the parent→\\tochild direction\)\. The DistFlow recurrence\[[33](https://arxiv.org/html/2609.36052#bib.bib25)\]reads

Pℓ\\displaystyle P\_\{\\ell\}=pc⁡\(ℓ\)load\+∑ℓ′∈children​\(ℓ\)Pℓ′\+rℓ​Pℓ2\+Qℓ2vp⁡\(ℓ\)2,\\displaystyle=p^\{\\mathrm\{load\}\}\_\{c\(\\ell\)\}\+\\sum\_\{\\ell^\{\\prime\}\\in\\text\{children\}\(\\ell\)\}P\_\{\\ell^\{\\prime\}\}\+r\_\{\\ell\}\\,\\frac\{P\_\{\\ell\}^\{2\}\+Q\_\{\\ell\}^\{2\}\}\{v^\{2\}\_\{p\(\\ell\)\}\},\(9\)Qℓ\\displaystyle Q\_\{\\ell\}=qc⁡\(ℓ\)load\+∑ℓ′∈children​\(ℓ\)Qℓ′\+xℓ​Pℓ2\+Qℓ2vp⁡\(ℓ\)2,\\displaystyle=q^\{\\mathrm\{load\}\}\_\{c\(\\ell\)\}\+\\sum\_\{\\ell^\{\\prime\}\\in\\text\{children\}\(\\ell\)\}Q\_\{\\ell^\{\\prime\}\}\+x\_\{\\ell\}\\,\\frac\{P\_\{\\ell\}^\{2\}\+Q\_\{\\ell\}^\{2\}\}\{v^\{2\}\_\{p\(\\ell\)\}\},\(10\)vc⁡\(ℓ\)2\\displaystyle v^\{2\}\_\{c\(\\ell\)\}=vp⁡\(ℓ\)2−2​\(rℓ​Pℓ\+xℓ​Qℓ\)\+\(rℓ2\+xℓ2\)​Pℓ2\+Qℓ2vp⁡\(ℓ\)2,\\displaystyle=v^\{2\}\_\{p\(\\ell\)\}\-2\\bigl\(r\_\{\\ell\}P\_\{\\ell\}\+x\_\{\\ell\}Q\_\{\\ell\}\\bigr\)\+\(r\_\{\\ell\}^\{2\}\+x\_\{\\ell\}^\{2\}\)\\,\\frac\{P\_\{\\ell\}^\{2\}\+Q\_\{\\ell\}^\{2\}\}\{v^\{2\}\_\{p\(\\ell\)\}\},\(11\)wherechildren​\(ℓ\)\\text\{children\}\(\\ell\)denotes the set of branches whose parent bus isc⁡\(ℓ\)c\(\\ell\)\. Conventional pipelines walk the feeder tree backward \(children→\\toparent\) to evaluatePℓ,QℓP\_\{\\ell\},Q\_\{\\ell\}, then forward \(root→\\toleaves\) to evaluatevn2v^\{2\}\_\{n\}, iterating untilmaxn⁡\|vn2,\(k\)−vn2,\(k−1\)\|<ε\\max\_\{n\}\|v^\{2,\(k\)\}\_\{n\}\-v^\{2,\(k\-1\)\}\_\{n\}\|<\\varepsilon\(typicallyε=10−6\\varepsilon=10^\{\-6\}\)\. On a host CPU each pass is a Python tree walk and each iteration triggers a host\-device synchronization\.

Topology matrices\.For a radial network the backward and forward sweeps depend only on topology, so they can be precomputed once at environment construction time as two binary matrices:

𝐃\\displaystyle\\mathbf\{D\}∈\{0,1\}Nℓ×N,Dℓ,n=\{1if bus​n​is downstream of branch​ℓ,0else,\\displaystyle\\in\\\{0,1\\\}^\{N\_\{\\ell\}\\times N\},\\quad D\_\{\\ell,n\}=\\begin\{cases\}1&\\text\{if bus \}n\\text\{ is downstream of branch \}\\ell,\\\\ 0&\\text\{else,\}\\end\{cases\}\(12\)𝚷\\displaystyle\\boldsymbol\{\\Pi\}∈\{0,1\}N×Nℓ,Πn,ℓ=\{1if branch​ℓ​lies on the path from bus​0​to bus​n,0else\.\\displaystyle\\in\\\{0,1\\\}^\{N\\times N\_\{\\ell\}\},\\quad\\Pi\_\{n,\\ell\}=\\begin\{cases\}1&\\text\{if branch \}\\ell\\text\{ lies on the path from bus \}0\\text\{ to bus \}n,\\\\ 0&\\text\{else\.\}\\end\{cases\}The DistFlow iteration becomes

𝐏\(k\)\\displaystyle\\mathbf\{P\}^\{\(k\)\}=𝐃⁡\(𝐩load\+𝐩loss,\(k−1\)\),\\displaystyle=\\mathbf\{D\}\\,\\bigl\(\\mathbf\{p\}^\{\\mathrm\{load\}\}\+\\mathbf\{p\}^\{\\mathrm\{loss\},\(k\-1\)\}\\bigr\),\(13\)𝐐\(k\)\\displaystyle\\mathbf\{Q\}^\{\(k\)\}=𝐃⁡\(𝐪load\+𝐪loss,\(k−1\)\),\\displaystyle=\\mathbf\{D\}\\,\\bigl\(\\mathbf\{q\}^\{\\mathrm\{load\}\}\+\\mathbf\{q\}^\{\\mathrm\{loss\},\(k\-1\)\}\\bigr\),𝜹2,\(k\)\\displaystyle\\boldsymbol\{\\delta\}^\{2,\(k\)\}=2​\(𝐫⊙𝐏\(k\)\+𝐱⊙𝐐\(k\)\)−\(𝐫2\+𝐱2\)⊙\(𝐏\(k\)\)2\+\(𝐐\(k\)\)2𝐯p⁡\(⋅\)2,\(k−1\),\\displaystyle=2\\,\(\\mathbf\{r\}\\odot\\mathbf\{P\}^\{\(k\)\}\+\\mathbf\{x\}\\odot\\mathbf\{Q\}^\{\(k\)\}\)\-\(\\mathbf\{r\}^\{2\}\+\\mathbf\{x\}^\{2\}\)\\odot\\frac\{\(\\mathbf\{P\}^\{\(k\)\}\)^\{2\}\+\(\\mathbf\{Q\}^\{\(k\)\}\)^\{2\}\}\{\\mathbf\{v\}^\{2,\(k\-1\)\}\_\{p\(\\cdot\)\}\},𝐯2,\(k\)\\displaystyle\\mathbf\{v\}^\{2,\(k\)\}=Vslack2​1−𝚷​𝜹2,\(k\),\\displaystyle=V\_\{\\text\{slack\}\}^\{2\}\\,\\mathbf\{1\}\-\\boldsymbol\{\\Pi\}\\,\\boldsymbol\{\\delta\}^\{2,\(k\)\},with the per\-bus loss\-correction𝐩loss,𝐪loss\\mathbf\{p\}^\{\\mathrm\{loss\}\},\\mathbf\{q\}^\{\\mathrm\{loss\}\}recomputed from the previous iterate\. Equation \([13](https://arxiv.org/html/2609.36052#A2.E13)\) expresses one PF iteration as fixed\-shapejax\.numpymatrix\-vector products with no host\-side branching\.

Compiled fixed\-point loop and batching\.The whole iteration bodyf:\(𝐏,𝐐,𝐯2\)→\(𝐏′,𝐐′,𝐯′2\)f:\(\\mathbf\{P\},\\mathbf\{Q\},\\mathbf\{v\}^\{2\}\)\\to\(\\mathbf\{P\}^\{\\prime\},\\mathbf\{Q\}^\{\\prime\},\\mathbf\{v\}^\{\\prime 2\}\)is implemented withjnpandlaxops only, with no host\-side mutation\. PowerZooJax wraps it injax\.lax\.while\_loopwith the convergence predicate evaluated on traced arrays, then composes the per\-step PF solve with reward and cost computation throughjax\.lax\.scanfor time andjax\.vmapfor batched scenarios\. The 48 PF evaluations of one DSO or DERs episode are therefore one accelerator\-resident computation, not 48 host\-side solver calls\. The same kernel powers DSO and DERs\.

### B\.2DC optimal power flow with ADMM

LP formulation\.TSO uses a DC OPF with the bus\-level PTDF𝐇bus∈ℝNℓ×N\\mathbf\{H\}\_\{\\text\{bus\}\}\\in\\mathbb\{R\}^\{N\_\{\\ell\}\\times N\}from the prebuilt case file;𝐇∈ℝNℓ×Ng\\mathbf\{H\}\\in\\mathbb\{R\}^\{N\_\{\\ell\}\\times N\_\{g\}\}and𝐇d∈ℝNℓ×Nd\\mathbf\{H\}\_\{d\}\\in\\mathbb\{R\}^\{N\_\{\\ell\}\\times N\_\{d\}\}select the generator and demand bus columns of𝐇bus\\mathbf\{H\}\_\{\\text\{bus\}\}\. Let𝐩∈ℝNg\\mathbf\{p\}\\in\\mathbb\{R\}^\{N\_\{g\}\}be the generator dispatch vector,𝐜g\\mathbf\{c\}^\{g\}marginal cost,𝐩d∈ℝNd\\mathbf\{p\}^\{d\}\\in\\mathbb\{R\}^\{N\_\{d\}\}nodal demand,𝐟¯\\overline\{\\mathbf\{f\}\}line\-flow limits, and\(𝐩¯,𝐩¯\)\(\\underline\{\\mathbf\{p\}\},\\overline\{\\mathbf\{p\}\}\)generator output bounds\. The DC OPF is the LP

min𝐩⁡𝐜g⊤​𝐩s\.t\.𝟏⊤​𝐩=𝟏⊤​𝐩d,\|𝐇𝐩−𝐇d​𝐩d\|≤𝐟¯,𝐩¯≤𝐩≤𝐩¯\.\\min\_\{\\mathbf\{p\}\}\\;\\mathbf\{c\}^\{g\\,\\top\}\\mathbf\{p\}\\quad\\text\{s\.t\.\}\\quad\\mathbf\{1\}^\{\\top\}\\mathbf\{p\}=\\mathbf\{1\}^\{\\top\}\\mathbf\{p\}^\{d\},\\ \\ \|\\mathbf\{H\}\\mathbf\{p\}\-\\mathbf\{H\}\_\{d\}\\mathbf\{p\}^\{d\}\|\\leq\\overline\{\\mathbf\{f\}\},\\ \\ \\underline\{\\mathbf\{p\}\}\\leq\\mathbf\{p\}\\leq\\overline\{\\mathbf\{p\}\}\.\(14\)This DC OPF is the inner solver invoked at every step of the TSO task; commitment, ramp, and minimum up/down\-time constraints are enforced by the environment \(Section[C\.2](https://arxiv.org/html/2609.36052#A3.SS2)\)\.

ADMM iteration\.We split the box and line\-flow constraints into auxiliary variables𝐳1,𝐳2\\mathbf\{z\}\_\{1\},\\mathbf\{z\}\_\{2\}with dual variables𝐲1,𝐲2\\mathbf\{y\}\_\{1\},\\mathbf\{y\}\_\{2\}and penaltyρ\>0\\rho\>0\. The system\-wide power balance is kept as a hard equality and handled by an augmented KKT system\. After standard derivation\[[34](https://arxiv.org/html/2609.36052#bib.bib12)\]the ADMM update for iterationkkreads

\[𝐩\(k\+1\)νeq\(k\+1\)\]\\displaystyle\\begin\{bmatrix\}\\mathbf\{p\}^\{\(k\+1\)\}\\\\ \\nu^\{\(k\+1\)\}\_\{\\text\{eq\}\}\\end\{bmatrix\}=Aaug−1​\[𝐫\(k\)𝟏⊤​𝐩d\],\\displaystyle=A\_\{\\text\{aug\}\}^\{\-1\}\\begin\{bmatrix\}\\mathbf\{r\}^\{\(k\)\}\\\\ \\mathbf\{1\}^\{\\top\}\\mathbf\{p\}^\{d\}\\end\{bmatrix\},\(15\)𝐳1\(k\+1\)\\displaystyle\\mathbf\{z\}\_\{1\}^\{\(k\+1\)\}=clip⁡\(𝐩\(k\+1\)\+𝐲1\(k\)/ρ,𝐩¯,𝐩¯\),\\displaystyle=\\mathrm\{clip\}\\bigl\(\\mathbf\{p\}^\{\(k\+1\)\}\+\\mathbf\{y\}\_\{1\}^\{\(k\)\}/\\rho,\\ \\underline\{\\mathbf\{p\}\},\\ \\overline\{\\mathbf\{p\}\}\\bigr\),𝐳2\(k\+1\)\\displaystyle\\mathbf\{z\}\_\{2\}^\{\(k\+1\)\}=clip⁡\(𝐇𝐩\(k\+1\)−𝐇d​𝐩d\+𝐲2\(k\)/ρ,−𝐟¯,𝐟¯\),\\displaystyle=\\mathrm\{clip\}\\bigl\(\\mathbf\{H\}\\mathbf\{p\}^\{\(k\+1\)\}\-\\mathbf\{H\}\_\{d\}\\mathbf\{p\}^\{d\}\+\\mathbf\{y\}\_\{2\}^\{\(k\)\}/\\rho,\\ \-\\overline\{\\mathbf\{f\}\},\\ \\overline\{\\mathbf\{f\}\}\\bigr\),𝐲1\(k\+1\)\\displaystyle\\mathbf\{y\}\_\{1\}^\{\(k\+1\)\}=𝐲1\(k\)\+ρ⁡\(𝐩\(k\+1\)−𝐳1\(k\+1\)\),\\displaystyle=\\mathbf\{y\}\_\{1\}^\{\(k\)\}\+\\rho\\bigl\(\\mathbf\{p\}^\{\(k\+1\)\}\-\\mathbf\{z\}\_\{1\}^\{\(k\+1\)\}\\bigr\),𝐲2\(k\+1\)\\displaystyle\\mathbf\{y\}\_\{2\}^\{\(k\+1\)\}=𝐲2\(k\)\+ρ⁡\(𝐇𝐩\(k\+1\)−𝐇d​𝐩d−𝐳2\(k\+1\)\),\\displaystyle=\\mathbf\{y\}\_\{2\}^\{\(k\)\}\+\\rho\\bigl\(\\mathbf\{H\}\\mathbf\{p\}^\{\(k\+1\)\}\-\\mathbf\{H\}\_\{d\}\\mathbf\{p\}^\{d\}\-\\mathbf\{z\}\_\{2\}^\{\(k\+1\)\}\\bigr\),with the residual𝐫\(k\)=−𝐜g\+ρ​𝐳1\(k\)−𝐲1\(k\)\+ρ​𝐇⊤​\(𝐳2\(k\)\+𝐇d​𝐩d\)−𝐇⊤​𝐲2\(k\)\\mathbf\{r\}^\{\(k\)\}=\-\\mathbf\{c\}^\{g\}\+\\rho\\,\\mathbf\{z\}\_\{1\}^\{\(k\)\}\-\\mathbf\{y\}\_\{1\}^\{\(k\)\}\+\\rho\\mathbf\{H\}^\{\\top\}\(\\mathbf\{z\}\_\{2\}^\{\(k\)\}\+\\mathbf\{H\}\_\{d\}\\mathbf\{p\}^\{d\}\)\-\\mathbf\{H\}^\{\\top\}\\mathbf\{y\}\_\{2\}^\{\(k\)\}, the multiplierνeq\\nu\_\{\\text\{eq\}\}on𝟏⊤​𝐩=𝟏⊤​𝐩d\\mathbf\{1\}^\{\\top\}\\mathbf\{p\}=\\mathbf\{1\}^\{\\top\}\\mathbf\{p\}^\{d\}recovered alongside𝐩\(k\+1\)\\mathbf\{p\}^\{\(k\+1\)\}and discarded after extraction, and the augmented KKT matrix

Aaug=\[ρ⁡\(𝐈\+𝐇⊤​𝐇\)𝟏𝟏⊤0\]A\_\{\\text\{aug\}\}=\\begin\{bmatrix\}\\rho\\,\(\\mathbf\{I\}\+\\mathbf\{H\}^\{\\top\}\\mathbf\{H\}\)&\\mathbf\{1\}\\\\ \\mathbf\{1\}^\{\\top\}&0\\end\{bmatrix\}\(16\)depending only on the case parameters; PowerZooJax precomputesAaug−1A\_\{\\text\{aug\}\}^\{\-1\}at construction time, so each ADMM iteration becomes one matrix\-vector product plus elementwise clips\. The whole solve runs as a fixed\-lengthjax\.lax\.fori\_loop\(TSO uses 200 iterations and tolerance10−410^\{\-4\}\), which compiles to a single XLA program\. TSO uses this kernel\.

### B\.3Bid\-based market clearing with PDIPM

GenCos clears a security\-constrained economic dispatch \(SCED\) over piecewise\-linear bid curves\. Each agentiisubmits up toS=3S=3break\-points; the clearing problem with bid pricesci,sbc^\{b\}\_\{i,s\}and segment capsp¯i,s\\overline\{p\}\_\{i,s\}is

min⁡∑i,spi,s≥0⁡ci,sb​pi,ss\.t\.∑i,spi,s=Dt,\|𝐇𝐩−𝐇d​𝐩d\|≤𝐟¯,pi,s≤p¯i,s\.\\min\_\{p\_\{i,s\}\\geq 0\}\\;\\sum\_\{i,s\}c^\{b\}\_\{i,s\}\\,p\_\{i,s\}\\quad\\text\{s\.t\.\}\\quad\\sum\_\{i,s\}p\_\{i,s\}=D\_\{t\},\\ \\ \|\\mathbf\{H\}\\mathbf\{p\}\-\\mathbf\{H\}\_\{d\}\\mathbf\{p\}^\{d\}\|\\leq\\overline\{\\mathbf\{f\}\},\\ \\ p\_\{i,s\}\\leq\\overline\{p\}\_\{i,s\}\.\(17\)PowerZooJax solves it with a primal\-dual interior\-point method\[[35](https://arxiv.org/html/2609.36052#bib.bib8)\]\. With barrier parameterμ\\mu, the perturbed KKT system is

∇pℒ=0,sjξj=μ,∀j,sj≥0,ξj≥0,\\nabla\_\{p\}\\mathcal\{L\}=0,\\quad s\_\{j\}\\,\\xi\_\{j\}=\\mu,\\ \\forall j,\\quad s\_\{j\}\\geq 0,\\ \\xi\_\{j\}\\geq 0,\(18\)wheresj,ξjs\_\{j\},\\xi\_\{j\}are slacks and multipliers for the inequality constraints\. At each iteration the duality gapμ\(k\)=𝐬\(k\)⊤​𝝃\(k\)/m\\mu^\{\(k\)\}=\\mathbf\{s\}^\{\(k\)\\top\}\\boldsymbol\{\\xi\}^\{\(k\)\}/mis recomputed and shrunk via a centring term: the perturbed complementarity residual uses targetσ​μ\(k\)\\sigma\\,\\mu^\{\(k\)\}withσ=0\.2\\sigma=0\.2\. Each Newton step solves a reduced KKT system\. The loop runs as a fixed\-lengthjax\.lax\.fori\_loopwith early stopping whenμ\(k\)<10−5\\mu^\{\(k\)\}<10^\{\-5\}\. The dual variableν∗\\nu^\{\*\}on the power\-balance equality is recovered analytically from the primal solution; locational marginal prices are

𝝅∗=−ν∗​1−𝐇bus⊤​𝝃flow∗,\\boldsymbol\{\\pi\}^\{\*\}=\-\\nu^\{\*\}\\,\\mathbf\{1\}\-\\mathbf\{H\}\_\{\\text\{bus\}\}^\{\\top\}\\boldsymbol\{\\xi\}^\{\*\}\_\{\\text\{flow\}\},\(19\)where𝝃flow∗=𝝃upper∗−𝝃lower∗\\boldsymbol\{\\xi\}^\{\*\}\_\{\\text\{flow\}\}=\\boldsymbol\{\\xi\}^\{\*\}\_\{\\text\{upper\}\}\-\\boldsymbol\{\\xi\}^\{\*\}\_\{\\text\{lower\}\}is the net line\-constraint dual built from the duals of the upper and lower line\-flow inequalities in \([17](https://arxiv.org/html/2609.36052#A2.E17)\)\.

### B\.4Resource device dynamics

#### Battery storage\.

For batteryjjwith capacityEjmaxE^\{\\text\{max\}\}\_\{j\}\(MWh\) and rated powerPjmaxP^\{\\text\{max\}\}\_\{j\}\(MW\), the per\-step state\-of\-charge update with charge/discharge efficienciesηjc,ηjd\\eta^\{c\}\_\{j\},\\eta^\{d\}\_\{j\}and step lengthΔ​t\\Delta tis

SoCjt\+1=clip⁡\(SoCjt\+Δ​tEjmax​\(ηjc​pj,tc−1ηjd​pj,td\),SoCjmin,SoCjmax\),\\mathrm\{SoC\}\_\{j\}^\{t\+1\}=\\mathrm\{clip\}\\Bigl\(\\mathrm\{SoC\}\_\{j\}^\{t\}\+\\frac\{\\Delta t\}\{E^\{\\text\{max\}\}\_\{j\}\}\\bigl\(\\eta^\{c\}\_\{j\}p^\{c\}\_\{j,t\}\-\\tfrac\{1\}\{\\eta^\{d\}\_\{j\}\}\\,p^\{d\}\_\{j,t\}\\bigr\),\\ \\mathrm\{SoC\}^\{\\text\{min\}\}\_\{j\},\\ \\mathrm\{SoC\}^\{\\text\{max\}\}\_\{j\}\\Bigr\),\(20\)withpj,tc,pj,td∈\[0,Pjmax\]p^\{c\}\_\{j,t\},p^\{d\}\_\{j,t\}\\in\[0,P^\{\\text\{max\}\}\_\{j\}\]andpj,tc​pj,td=0p^\{c\}\_\{j,t\}\\,p^\{d\}\_\{j,t\}=0enforced by parameterizing the action aspj∈\[−Pjmax,Pjmax\]p\_\{j\}\\in\[\-P^\{\\text\{max\}\}\_\{j\},P^\{\\text\{max\}\}\_\{j\}\]and splitting at zero\. Tasks that include batteries \(DERs, DCMG\) terminate the episode by adding a quadratic penalty\(SoCjT−SoCjtarget\)2\(\\mathrm\{SoC\}\_\{j\}^\{T\}\-\\mathrm\{SoC\}^\{\\text\{target\}\}\_\{j\}\)^\{2\}\.

#### PV inverter\.

A PV inverterjjwith rated apparent powerSjmaxS^\{\\text\{max\}\}\_\{j\}, available DC powerpj,tavailp^\{\\text\{avail\}\}\_\{j,t\}, and active/reactive setpointspj,t,qj,tp\_\{j,t\},q\_\{j,t\}enforces the box\-projected PQ envelope

0≤pj,t≤pj,tavail,pj,t2\+qj,t2≤\(Sjmax\)2,0\\leq p\_\{j,t\}\\leq p^\{\\text\{avail\}\}\_\{j,t\},\\quad p\_\{j,t\}^\{2\}\+q\_\{j,t\}^\{2\}\\leq\(S^\{\\text\{max\}\}\_\{j\}\)^\{2\},\(21\)implemented as a closed\-form projection onto the inverter capability circle\.

#### Flexible load\.

Flexible loadjjhas nominal demanddj,tnomd^\{\\text\{nom\}\}\_\{j,t\}, curtailment capηjcur∈\[0,1\]\\eta^\{\\text\{cur\}\}\_\{j\}\\in\[0,1\], shifting capηjshift\\eta^\{\\text\{shift\}\}\_\{j\}, and a shift horizonHjH\_\{j\}\(steps\)\. Curtailed energy is permanently lost; shifted energy must be repaid withinHjH\_\{j\}steps\. The realized demand and the per\-load energy buffer evolve as

dj,treal\\displaystyle d^\{\\text\{real\}\}\_\{j,t\}=\(1−uj,tcur\)​dj,tnom−uj,tshift​dj,tnom\+uj,trecall,\\displaystyle=\(1\-u^\{\\text\{cur\}\}\_\{j,t\}\)\\,d^\{\\text\{nom\}\}\_\{j,t\}\-u^\{\\text\{shift\}\}\_\{j,t\}\\,d^\{\\text\{nom\}\}\_\{j,t\}\+u^\{\\text\{recall\}\}\_\{j,t\},\(22\)Ej,t\+1shift\\displaystyle E^\{\\text\{shift\}\}\_\{j,t\+1\}=Ej,tshift\+uj,tshift​dj,tnom−uj,trecall,\\displaystyle=E^\{\\text\{shift\}\}\_\{j,t\}\+u^\{\\text\{shift\}\}\_\{j,t\}\\,d^\{\\text\{nom\}\}\_\{j,t\}\-u^\{\\text\{recall\}\}\_\{j,t\},\(23\)withuj,tshift≤ηjshiftu^\{\\text\{shift\}\}\_\{j,t\}\\leq\\eta^\{\\text\{shift\}\}\_\{j\}anduj,tcur≤ηjcuru^\{\\text\{cur\}\}\_\{j,t\}\\leq\\eta^\{\\text\{cur\}\}\_\{j\}\. DSO and DERs use this model\.

#### Diesel generator\.

Diesel generatorjjhas rated powerPjmaxP^\{\\text\{max\}\}\_\{j\}and a linear fuel costCj​\(p\)=cjfuel​pC\_\{j\}\(p\)=c^\{\\text\{fuel\}\}\_\{j\}\\,pproportional to the delivered active power\. The continuous normalized commandaj,tdg∈\[0,1\]a^\{\\text\{dg\}\}\_\{j,t\}\\in\[0,1\]maps to

pj,t=clip⁡\(aj,tdg​Pjmax,0,Pjmax\),costj,t=Cj​\(pj,t\)​Δ​t,p\_\{j,t\}=\\mathrm\{clip\}\(a^\{\\text\{dg\}\}\_\{j,t\}\\,P^\{\\text\{max\}\}\_\{j\},\\ 0,\\ P^\{\\text\{max\}\}\_\{j\}\),\\qquad\\text\{cost\}\_\{j,t\}=C\_\{j\}\(p\_\{j,t\}\)\\,\\Delta t,\(24\)with an optional minimum\-loading threshold \(aj,tdg<Pjmin/\(2​Pjmax\)a^\{\\text\{dg\}\}\_\{j,t\}<P^\{\\text\{min\}\}\_\{j\}/\(2P^\{\\text\{max\}\}\_\{j\}\)snaps to00, in between snaps up toPjminP^\{\\text\{min\}\}\_\{j\}\) so the operating point respects diesel wet\-stacking guidance\. DCMG wraps the diesel as a same\-step residual\-slack actuator, so dispatch is realized after the workload, cooling, and battery decisions\.

### B\.5DCMG coupling

The DCMG task couples the device dynamics of Section[B\.4](https://arxiv.org/html/2609.36052#A2.SS4)with a lumped\-capacitance thermal model and a single\-bus power balance\. LetWtW\_\{t\}be the GPU\-equivalent compute load,θtIT\\theta\_\{t\}^\{\\text\{IT\}\}the IT\-zone temperature,θtout\\theta\_\{t\}^\{\\text\{out\}\}the outdoor temperature, andptcoolp^\{\\text\{cool\}\}\_\{t\}the cooling electrical power\. The IT\-zone thermal dynamics are

θt\+1IT=θtIT\+Δ​tCth​\(αW​Wt−βt​ptcool\+γout​\(θtout−θtIT\)\),\\theta\_\{t\+1\}^\{\\text\{IT\}\}=\\theta\_\{t\}^\{\\text\{IT\}\}\+\\frac\{\\Delta t\}\{C\_\{\\text\{th\}\}\}\\bigl\(\\alpha\_\{W\}\\,W\_\{t\}\-\\beta\_\{t\}\\,p^\{\\text\{cool\}\}\_\{t\}\+\\gamma\_\{\\text\{out\}\}\(\\theta\_\{t\}^\{\\text\{out\}\}\-\\theta\_\{t\}^\{\\text\{IT\}\}\)\\bigr\),\(25\)with thermal capacityCthC\_\{\\text\{th\}\}, IT\-to\-heat coefficientαW\\alpha\_\{W\}\(soαW​Wt\\alpha\_\{W\}W\_\{t\}is the IT electrical power at steptt, equal to its heat output\), wall\-to\-ambient coefficientγout\\gamma\_\{\\text\{out\}\}, and cooling COPβt=β⁡\(θtout\)\\beta\_\{t\}=\\beta\(\\theta\_\{t\}^\{\\text\{out\}\}\)that decreases with outdoor temperature\. The over\-temperature cost channelctth=\(θtIT−θmax\)\+c^\{\\text\{th\}\}\_\{t\}=\(\\theta\_\{t\}^\{\\text\{IT\}\}\-\\theta^\{\\text\{max\}\}\)\_\{\+\}is added to the cost vector at every step\. Cooling is implemented as a setpoint\-tracking controller: the RL actionatcoola^\{\\text\{cool\}\}\_\{t\}sets the IT\-zone setpoint, andptcoolp^\{\\text\{cool\}\}\_\{t\}equals the resulting heat\-removal rate divided byβt\\beta\_\{t\}\(Section[C\.5](https://arxiv.org/html/2609.36052#A3.SS5)\)\.

The local power balance at the point of common coupling is

ptDC=ptPV\+ptbatt\+ptDG\+ptgrid,p^\{\\text\{DC\}\}\_\{t\}\\;=\\;p^\{\\text\{PV\}\}\_\{t\}\+p^\{\\text\{batt\}\}\_\{t\}\+p^\{\\text\{DG\}\}\_\{t\}\+p^\{\\text\{grid\}\}\_\{t\},\(26\)whereptDC=ptIT\+ptcool\+ptauxp^\{\\text\{DC\}\}\_\{t\}=p^\{\\text\{IT\}\}\_\{t\}\+p^\{\\text\{cool\}\}\_\{t\}\+p^\{\\text\{aux\}\}\_\{t\}is the total data\-center load andptauxp^\{\\text\{aux\}\}\_\{t\}is the auxiliary load \(PSU losses and other site loads, modeled as a fixed fraction ofptITp^\{\\text\{IT\}\}\_\{t\}\)\. Grid importptgrid∈\[0,p¯grid\]p^\{\\text\{grid\}\}\_\{t\}\\in\[0,\\overline\{p\}^\{\\text\{grid\}\}\]in grid\-connected mode andptgrid=0p^\{\\text\{grid\}\}\_\{t\}=0in islanded mode\. SLA tracking compares served workload against committed jobs; the SLA\-deficit cost channel is the per\-step expired\-task count normalized by the GPU count\. The reward combines energy use, electricity cost \(GB MID market price\), and carbon emission \(grid\-import and DG carbon intensities\); the weights are given in Section[C\.5](https://arxiv.org/html/2609.36052#A3.SS5)\.

## Appendix CPer\-Task Cards

This section gives the formal CMDP instantiation, observation, action, reward, cost vector, splits, baselines, and safety budgets for each of the five tasks\. Concrete numerical thresholds, time\-step lengths, and split sizes are taken from the per\-task configuration files \(Section[E\.3](https://arxiv.org/html/2609.36052#A5.SS3)shows one such file\)\. Table[5](https://arxiv.org/html/2609.36052#A3.T5)gives a one\-line summary of the suite\.

#### Notation in the task cards\.

We use bold lowercase for vectors and bold uppercase for matrices; symbols without attsubscript are time\-invariant\. Per\-task dimensions follow each task’s case file and configuration; the explicit dimension counts under each observation/action equation in this section are evaluated at the canonical case \(case5,case118,case33bw,case141, and the data center microgrid configuration shipped with the benchmark\)\.

Table 5:One\-line task overview\.
### C\.1Generation\-company bidding task \(GenCos\)

Task\.Five generation companies \(GenCos\) compete in a repeated wholesale electricity market on the IEEE 5\-bus transmission grid \(case5, GB demand profiles555Great Britain demand traces from the National Energy System Operator \(NESO\),[https://www\.neso\.energy/data\-portal/historic\-demand\-data](https://www.neso.energy/data-portal/historic-demand-data)\.; see Figure[7](https://arxiv.org/html/2609.36052#A3.F7)\)\. The horizon is 24 hours at 30\-minute resolution, so an episode contains 48 sequential market clearing steps\. At each step the market operator clears all bids via a security\-constrained economic dispatch \(SCED, an optimization that picks the cheapest feasible bid mix subject to grid limits\), returning each unit’s dispatch levelPi,tgP^\{g\}\_\{i,t\}and the locational marginal prices \(LMPs, the per\-bus electricity prices\)\. Successive clearings are coupled by runtime ramp bounds \(per\-step limits on how fast a generator can change output\) enforced inside the SCED linear program: at steptteach unit’s box constraint is tightened topi,tmin=max⁡\(pimin,Pi,t−1g−Δ​Pidown\)p^\{\\min\}\_\{i,t\}=\\max\(p^\{\\min\}\_\{i\},P^\{g\}\_\{i,t\-1\}\-\\Delta P^\{\\mathrm\{down\}\}\_\{i\}\)andpi,tmax=min⁡\(pimax,Pi,t−1g\+Δ​Piup\)p^\{\\max\}\_\{i,t\}=\\min\(p^\{\\max\}\_\{i\},P^\{g\}\_\{i,t\-1\}\+\\Delta P^\{\\mathrm\{up\}\}\_\{i\}\), so the episode is a rolling sequential market rather than 48 independent auctions\. The market and the grid are shared across the five agents, which makes this a competitive multi\-agent task\.

Figure 7:IEEE 5\-bus market test system \(case5\) used by the GenCos task: 5 buses B1–B5 \(horizontal bars\), 6 transmission lines L1–L6, 5 thermal generators G1–G5 \(circled symbols, one per agent; G1 and G2 share bus B1, while G3, G4, G5 sit at buses B3, B4, B5 respectively\), and 3 load sites D2, D3, D4 \(outward arrows at buses B2, B3, B4\)\. Each step, the SCED clearing of Section[B\.3](https://arxiv.org/html/2609.36052#A2.SS3)returns the per\-bus LMPs and per\-unit dispatches that drive each agent’s profit\.Observation and Action\.Each agent owns one generator and observes only its own physical status plus a few public market signals\. For agentii,

𝐨i,t=\(cib,P~imax,Pi,t−1g,w~i,t−1,hi,tramp,Dt\+1fcst,𝝅¯thist,τt\)∈ℝ12\.\\mathbf\{o\}\_\{i,t\}=\\bigl\(c^\{b\}\_\{i\},\\ \\widetilde\{P\}^\{\\max\}\_\{i\},\\ P^\{g\}\_\{i,t\-1\},\\ \\widetilde\{w\}\_\{i,t\-1\},\\ h^\{\\mathrm\{ramp\}\}\_\{i,t\},\\ D^\{\\mathrm\{fcst\}\}\_\{t\+1\},\\ \\overline\{\\boldsymbol\{\\pi\}\}^\{\\mathrm\{hist\}\}\_\{t\},\\ \\tau\_\{t\}\\bigr\)\\in\\mathbb\{R\}^\{12\}\.\(27\)𝐚i,t∈\[−1,1\]Kseg\.\\mathbf\{a\}\_\{i,t\}\\in\[\-1,1\]^\{K\_\{\\mathrm\{seg\}\}\}\.\(28\)Each unit’s true cost isTCi​\(P\)=ai3​P3\+bi2​P2\+ci​P\\mathrm\{TC\}\_\{i\}\(P\)=\\tfrac\{a\_\{i\}\}\{3\}P^\{3\}\+\\tfrac\{b\_\{i\}\}\{2\}P^\{2\}\+c\_\{i\}P, the integral of a quadratic marginal cost \(the cost of producing one extra MW at outputPP\)MCi​\(P\)=ai​P2\+bi​P\+ci\\mathrm\{MC\}\_\{i\}\(P\)=a\_\{i\}P^\{2\}\+b\_\{i\}P\+c\_\{i\}; truthful prices are obtained by evaluatingMCi\\mathrm\{MC\}\_\{i\}at the midpoint of each ofKseg=3K\_\{\\mathrm\{seg\}\}\{=\}3equal\-width segments covering\[pimin,pimax\]\[p^\{\\min\}\_\{i\},p^\{\\max\}\_\{i\}\]\. The eight observation entries are:

- •cibc^\{b\}\_\{i\}: unit’s truthful first\-segment price \(lowest\-block marginal cost\), normalized by the configured price scale;
- •P~imax\\widetilde\{P\}^\{\\max\}\_\{i\}: capacity relative to the largest unit;
- •Pi,t−1gP^\{g\}\_\{i,t\-1\}: previous dispatch, normalized by the agent’s capacity;
- •w~i,t−1\\widetilde\{w\}\_\{i,t\-1\}: previous\-step profit, normalized by a fixed maximum\-revenue scale;
- •hi,tramp=min⁡\(pimax−Pi,t−1g,Δ​Piup\)/pimaxh^\{\\mathrm\{ramp\}\}\_\{i,t\}=\\min\(p^\{\\max\}\_\{i\}\-P^\{g\}\_\{i,t\-1\},\\Delta P^\{\\mathrm\{up\}\}\_\{i\}\)/p^\{\\max\}\_\{i\}: achievable ramp\-up next step \(the smaller of remaining capacity and the per\-step ramp\-up limit\);
- •Dt\+1fcstD^\{\\mathrm\{fcst\}\}\_\{t\+1\}: next\-step total\-demand forecast \(perfect one\-step look\-ahead, normalized by total capacity\);
- •𝝅¯thist∈ℝ4\\overline\{\\boldsymbol\{\\pi\}\}^\{\\mathrm\{hist\}\}\_\{t\}\\in\\mathbb\{R\}^\{4\}: circular buffer of the four most recent system\-mean LMPs \(oldest to newest\), giving every agent a coarse public price\-dynamics signal;
- •τt∈ℝ2\\tau\_\{t\}\\in\\mathbb\{R\}^\{2\}: sin/cos time\-of\-day encoding\.

The six task\-level scalars, the four\-element price history, and the two\-component time encoding sum to 12\.

The raw action𝐚i,t\\mathbf\{a\}\_\{i,t\}is mapped to a piecewise bid curve in three steps:

1. 1\.Shift to non\-negative magnitudesmi,t=\(𝐚i,t\+1\)/2∈\[0,1\]Ksegm\_\{i,t\}=\(\\mathbf\{a\}\_\{i,t\}\+1\)/2\\in\[0,1\]^\{K\_\{\\mathrm\{seg\}\}\}\.
2. 2\.Sort along the segment axis to enforce a monotone offer \(no later block cheaper than an earlier block\)\.
3. 3\.Apply a multiplicative markup capped by a configured maximum, so that the cleared offer for segmentkkisofferi,k,t=ci,kseg⋅\(1\+sort​\(mi,t\)k⋅m¯\)\\mathrm\{offer\}\_\{i,k,t\}=c^\{\\mathrm\{seg\}\}\_\{i,k\}\\cdot\\bigl\(1\+\\mathrm\{sort\}\(m\_\{i,t\}\)\_\{k\}\\cdot\\overline\{m\}\\bigr\)withci,ksegc^\{\\mathrm\{seg\}\}\_\{i,k\}the truthful segment price\.

Reward and Violation Cost\.Each agent’s reward is its realized profit at the cleared dispatch, equal to cleared revenue𝝅b⁡\(i\),t∗​Pi,tg​Δ​t\\boldsymbol\{\\pi\}^\{\*\}\_\{b\(i\),t\}\\,P^\{g\}\_\{i,t\}\\,\\Delta tminus true production costTCi​\(Pi,tg\)​Δ​t\\mathrm\{TC\}\_\{i\}\(P^\{g\}\_\{i,t\}\)\\,\\Delta t, whereb⁡\(i\)b\(i\)is the bus \(the network node where loads and generators connect\) of unitii:

ri,t=\(𝝅b⁡\(i\),t∗​Pi,tg−TCi​\(Pi,tg\)\)​Δ​t,ct=Ctthermal,r\_\{i,t\}=\\bigl\(\\boldsymbol\{\\pi\}^\{\*\}\_\{b\(i\),t\}\\,P^\{g\}\_\{i,t\}\-\\mathrm\{TC\}\_\{i\}\(P^\{g\}\_\{i,t\}\)\\bigr\)\\,\\Delta t,\\qquad c\_\{t\}=C^\{\\mathrm\{thermal\}\}\_\{t\},\(29\)with𝝅∗\\boldsymbol\{\\pi\}^\{\*\}the cleared LMP vector from Section[B\.3](https://arxiv.org/html/2609.36052#A2.SS3)\. The single violation channelCtthermalC^\{\\mathrm\{thermal\}\}\_\{t\}is the sum of post\-clearing line\-flow violations \(in MW\), shared across agents because line overloads are a property of the joint outcome rather than of any single bid\. This channel is reported as a market\-side diagnostic only\.

Baselines\.The benchmark releases three non\-learning strategic baselines \(Truthful, Uniform\-mid, and Max\-markup\) together with a learned IPPO baseline that uses one shared actor\-critic across the five GenCos\. The three heuristics bracket the strategic spectrum: Truthful \(bidding at marginal cost\) is the competitive\-equilibrium reference \(the outcome where every bidder offers at marginal cost\), Max\-markup \(bidding at the configured cap\) is the upper envelope of market power, and Uniform\-mid \(bidding at the midpoint markup\) sits between them as a neutral oligopoly reference\.

Challenge\.GenCos is a non\-stationary, partially observable, competitive multi\-agent task\. Each agent only sees its own dispatch, profit, and a small public LMP window, and the market reward depends on the joint action through the SCED clearing\. Policy improvement for one bidder shifts the optimization problem faced by every other bidder\. These factors together define the standard hard regime for current MARL\.

### C\.2Transmission dispatch task \(TSO\)

Task\.A single transmission system operator \(TSO\) runs the 54\-generator IEEE 118\-bus transmission grid \(case118, 186 lines; see Figure[8](https://arxiv.org/html/2609.36052#A3.F8)\) over a 24\-hour horizon at 30\-minute resolution, which gives 48 sequential decision steps per episode and is driven by GB net\-load profiles \(gross demand minus must\-take wind and solar\)666Demand from the National Energy System Operator \(NESO\),[https://www\.neso\.energy/data\-portal/historic\-demand\-data](https://www.neso.energy/data-portal/historic-demand-data); wind and solar from Elexon BMRS,[https://bmrs\.elexon\.co\.uk/generation\-by\-fuel\-type](https://bmrs.elexon.co.uk/generation-by-fuel-type)\.\. At each step the agent decides which generators are on \(*commitment*\) and how much each produces \(*dispatch*\); the environment then enforces ramp limits \(per\-step bounds on how fast a unit can change output\), minimum up/down times \(a unit must remain in its current state for at least this many consecutive steps after a transition\), and a DC optimal power flow solution \(a linear approximation that maps generator output to per\-line flows\)\. One centralized agent sees the whole grid, so this is a fully observable single\-agent task with a large mixed continuous/binary action space\.

![Refer to caption](https://arxiv.org/html/2609.36052v1/materials/figures/arxiv/fig_tso_topology.jpg)Figure 8:One\-line diagram of the IEEE 118\-bus transmission test system \(case118\) used by the TSO task: 118 buses \(horizontal bars\), 186 branches \(lines connecting buses\), 54 thermal generators \(circled G; arrows pointing into their bus\), and 91 load sides \(arrows pointing out of their bus\)\. The TSO task in Section[C\.2](https://arxiv.org/html/2609.36052#A3.SS2)commits and dispatches the 54 generators on a 30\-minute cadence under DC\-OPF redispatch\. Source: IIT Power Group, 2003\.Observation and Action\.The observation summarizes the whole grid; the action specifies a commitment intent and a dispatch preference per generator:

𝐨t=\(CLOSE\\displaystyle\\mathbf\{o\}\_\{t\}=\\bigl\(𝐮t,𝐭tstate,𝐏t−1g,𝐂g,𝐟t−1,\\displaystyle\\mathbf\{u\}\_\{t\},\\ \\mathbf\{t\}^\{\\mathrm\{state\}\}\_\{t\},\\ \\mathbf\{P\}^\{g\}\_\{t\-1\},\\ \\mathbf\{C\}^\{g\},\\ \\mathbf\{f\}\_\{t\-1\},OPENDt,ηtrsv,τt,𝐃tfcst\)∈ℝ410\.\\displaystyle D\_\{t\},\\ \\eta^\{\\mathrm\{rsv\}\}\_\{t\},\\ \\tau\_\{t\},\\ \\mathbf\{D\}^\{\\mathrm\{fcst\}\}\_\{t\}\\bigr\)\\in\\mathbb\{R\}^\{410\}\.\(30\)𝐚t=\(𝐮tcmd,𝐏tpref\)∈\[−1,1\]108\.\\mathbf\{a\}\_\{t\}=\\bigl\(\\mathbf\{u\}^\{\\mathrm\{cmd\}\}\_\{t\},\\ \\mathbf\{P\}^\{\\mathrm\{pref\}\}\_\{t\}\\bigr\)\\in\[\-1,1\]^\{108\}\.\(31\)The nine observation entries, withNg=54N\_\{g\}\{=\}54andNℓ=186N\_\{\\ell\}\{=\}186, are:

- •𝐮t∈\{0,1\}Ng\\mathbf\{u\}\_\{t\}\\in\\\{0,1\\\}^\{N\_\{g\}\}: on/off status of all generators, carried over from the previous step;
- •𝐭tstate∈ℝNg\\mathbf\{t\}^\{\\mathrm\{state\}\}\_\{t\}\\in\\mathbb\{R\}^\{N\_\{g\}\}: elapsed time each generator has been in its current on/off state \(normalized\), letting a learner respect minimum up/down constraints by inspection;
- •𝐏t−1g∈ℝNg\\mathbf\{P\}^\{g\}\_\{t\-1\}\\in\\mathbb\{R\}^\{N\_\{g\}\}: previous dispatch, normalized by per\-unit capacity;
- •𝐂g∈ℝNg\\mathbf\{C\}^\{g\}\\in\\mathbb\{R\}^\{N\_\{g\}\}: linear coefficientbib\_\{i\}of each unit’s quadratic marginal\-cost functionMCi​\(P\)=ai​P2\+bi​P\+ci\\mathrm\{MC\}\_\{i\}\(P\)=a\_\{i\}P^\{2\}\+b\_\{i\}P\+c\_\{i\}, normalized by its maximum across units;
- •𝐟t−1∈ℝNℓ\\mathbf\{f\}\_\{t\-1\}\\in\\mathbb\{R\}^\{N\_\{\\ell\}\}: active power flow on each transmission line, normalized by line capacity;
- •Dt∈ℝD\_\{t\}\\in\\mathbb\{R\}: aggregate system demand, normalized by total capacity;
- •ηtrsv∈ℝ\\eta^\{\\mathrm\{rsv\}\}\_\{t\}\\in\\mathbb\{R\}: committed\-headroom reserve ratio \(committed capacity minus current load, divided by load\);
- •τt∈ℝ2\\tau\_\{t\}\\in\\mathbb\{R\}^\{2\}: sin/cos time\-of\-day encoding;
- •𝐃tfcst∈ℝ4\\mathbf\{D\}^\{\\mathrm\{fcst\}\}\_\{t\}\\in\\mathbb\{R\}^\{4\}: four\-step \(two\-hour\) perfect look\-ahead total\-load forecast\.

The dimension count is4​Ng\+Nℓ\+4\+4=4104\\,N\_\{g\}\+N\_\{\\ell\}\+4\+4=410\.

The action is two\-headed:

- •𝐮tcmd∈\[−1,1\]Ng\\mathbf\{u\}^\{\\mathrm\{cmd\}\}\_\{t\}\\in\[\-1,1\]^\{N\_\{g\}\}: commitment intent applied by sign threshold \(entriesui,tcmd\>0u^\{\\mathrm\{cmd\}\}\_\{i,t\}\>0request unitiion for the next step\), with the active minimum up/down constraints overriding any request that would violate them;
- •𝐏tpref∈\[−1,1\]Ng\\mathbf\{P\}^\{\\mathrm\{pref\}\}\_\{t\}\\in\[\-1,1\]^\{N\_\{g\}\}: dispatch preference, denormalized linearly to each unit’s ramp\-adjusted feasible interval\[pi,tmin,pi,tmax\]\[p^\{\\min\}\_\{i,t\},p^\{\\max\}\_\{i,t\}\]and added as a linear bias to its marginal\-cost intercept inside the DC OPF, so the cleared dispatch follows the preference whenever it is feasible under all operational limits\.

The action dimension is2​Ng=1082N\_\{g\}=108\.

Reward and Violation Cost\.The reward is the negative total operating cost, summing fuel, startup, and no\-load components, rescaled by a fixed factor \(Section[F\.1](https://arxiv.org/html/2609.36052#A6.SS1)\)\. The two\-channel violation cost separately reports thermal line overloadsCtthermalC^\{\\mathrm\{thermal\}\}\_\{t\}\(the cumulative MW above each line’s capacity\) and reserve\-margin shortfallsCtreserveC^\{\\mathrm\{reserve\}\}\_\{t\}\(the MW deficit of committed headroom relative to the configured 5% margin\), so that a safe\-RL learner can trade them off explicitly rather than collapse them into a single penalty\. Both channels are reported under Form 2 \(zero\-violation constraint\):

rt=−Ctoperating,𝐜t=\(Ctthermal,Ctreserve\)\.r\_\{t\}=\-C^\{\\mathrm\{operating\}\}\_\{t\},\\qquad\\mathbf\{c\}\_\{t\}=\\bigl\(C^\{\\mathrm\{thermal\}\}\_\{t\},\\ C^\{\\mathrm\{reserve\}\}\_\{t\}\\bigr\)\.\(32\)
Baselines\.The benchmark provides two heuristic baselines \(All\-On, which leaves every unit committed, and Merit Order, which commits the cheapest unit first until total capacity covers demand plus a fixed 5% reserve margin\) and two learned baselines: PPO as the unconstrained reference and PPO\-Lagrangian with a primal\-dual update on the two cost channels\.

Challenge\.TSO mixes 54 binary commitments with 54 continuous setpoints, so the action space is high\-dimensional and combinatorial\. Each step also requires the policy to keep an OPF feasible against ramp, minimum up/down, thermal, and reserve constraints simultaneously\. Reserve violations are sparse and bursty, which hurts safe\-RL credit assignment\.

### C\.3Distribution demand response task \(DSO\)

Task\.A single distribution system operator \(DSO\) coordinates six flexible loads on the IEEE 33\-bus distribution network \(case33bw; see Figure[9](https://arxiv.org/html/2609.36052#A3.F9)\) over a 24\-hour episode at 30\-minute resolution, which gives 48 sequential decision steps per episode and is driven by real Ausgrid demand profiles777Ausgrid distribution demand traces,[https://www\.ausgrid\.com\.au/Industry/Our\-Research/Data\-to\-share/](https://www.ausgrid.com.au/Industry/Our-Research/Data-to-share/)\.\. The goal is to reduce total active power loss while keeping every bus voltage inside the safe band\[0\.94,1\.06\]\[0\.94,1\.06\]p\.u\. \(per unit, the normalized voltage scale used in power systems\)\. One centralized agent sees the whole network, so this is a fully observable single\-agent task with smooth continuous control\.

![Refer to caption](https://arxiv.org/html/2609.36052v1/fig_dso_topology.png)Figure 9:IEEE 33\-bus radial distribution network \(case33bw\) used by the DSO task\. The substation \(factory icon\) at bus 1 supplies the network through a main feeder \(buses 2–18\) and three lateral branches \(19–22, 23–25, 26–33\)\. Gray house icons mark passive loads; orange house icons mark the six controllable flexible loads at buses 6, 14, 18, 22, 28, and 33\. The DSO task in Section[C\.3](https://arxiv.org/html/2609.36052#A3.SS3)sets curtailment and load\-shift fractions for these six devices on a 30\-minute cadence under radial AC power flow\.Observation and Action\.

𝐨t=\(𝐕t,𝐏tbr,𝐐tbr,𝐏tload,𝐐tload,τt,𝐳tflex\)∈ℝ195\.\\mathbf\{o\}\_\{t\}=\\bigl\(\\mathbf\{V\}\_\{t\},\\ \\mathbf\{P\}^\{\\mathrm\{br\}\}\_\{t\},\\ \\mathbf\{Q\}^\{\\mathrm\{br\}\}\_\{t\},\\ \\mathbf\{P\}^\{\\mathrm\{load\}\}\_\{t\},\\ \\mathbf\{Q\}^\{\\mathrm\{load\}\}\_\{t\},\\ \\tau\_\{t\},\\ \\mathbf\{z\}^\{\\mathrm\{flex\}\}\_\{t\}\\bigr\)\\in\\mathbb\{R\}^\{195\}\.\(33\)𝐚t=\(𝐝tcurtail,𝐝tshift\)∈\[−1,1\]12\.\\mathbf\{a\}\_\{t\}=\\bigl\(\\mathbf\{d\}^\{\\mathrm\{curtail\}\}\_\{t\},\\ \\mathbf\{d\}^\{\\mathrm\{shift\}\}\_\{t\}\\bigr\)\\in\[\-1,1\]^\{12\}\.\(34\)The seven observation entries, withN=33N\{=\}33buses andNℓ=32N\_\{\\ell\}\{=\}32branches, are:

- •𝐕t∈ℝN\\mathbf\{V\}\_\{t\}\\in\\mathbb\{R\}^\{N\}: per\-bus voltage magnitude, centered and scaled as\(\|V\|−1\)/0\.1\(\|V\|\-1\)/0\.1so nominal voltages sit near 0 and the band edges0\.94/1\.060\.94/1\.06p\.u\. sit near±0\.6\\pm 0\.6;
- •𝐏tbr,𝐐tbr∈ℝNℓ\\mathbf\{P\}^\{\\mathrm\{br\}\}\_\{t\},\\mathbf\{Q\}^\{\\mathrm\{br\}\}\_\{t\}\\in\\mathbb\{R\}^\{N\_\{\\ell\}\}: active and reactive \(the imaginary part of AC power, in MVAr\) power flows on each branch, normalized by the reference load;
- •𝐏tload,𝐐tload∈ℝN\\mathbf\{P\}^\{\\mathrm\{load\}\}\_\{t\},\\mathbf\{Q\}^\{\\mathrm\{load\}\}\_\{t\}\\in\\mathbb\{R\}^\{N\}: active and reactive nodal demands, normalized by the reference load;
- •τt∈ℝ2\\tau\_\{t\}\\in\\mathbb\{R\}^\{2\}: sin/cos time\-of\-day encoding;
- •𝐳tflex∈ℝ30\\mathbf\{z\}^\{\\mathrm\{flex\}\}\_\{t\}\\in\\mathbb\{R\}^\{30\}: stacks five normalized variables per controllable load —\[ccur,sout,sin,bfill,beng\]\[c^\{\\mathrm\{cur\}\},s^\{\\mathrm\{out\}\},s^\{\\mathrm\{in\}\},b^\{\\mathrm\{fill\}\},b^\{\\mathrm\{eng\}\}\], namely current curtailment level, outgoing/incoming shift fractions, deferred\-buffer fill ratio, and deferred\-buffer energy — across the six flexible loads\.

The dimensions sum to3​N\+2​Nℓ\+2\+6×5=99\+64\+2\+30=1953N\+2N\_\{\\ell\}\+2\+6\{\\times\}5=99\+64\+2\+30=195\.

The action is two\-headed and per device:

- •𝐝tcurtail∈\[−1,1\]6\\mathbf\{d\}^\{\\mathrm\{curtail\}\}\_\{t\}\\in\[\-1,1\]^\{6\}: curtailment intent, clipped to\[0,1\]\[0,1\]inside each FlexLoad bundle and then multiplied by that device’s curtailment cap, applied at the current slot \(curtailed energy is permanently lost\);
- •𝐝tshift∈\[−1,1\]6\\mathbf\{d\}^\{\\mathrm\{shift\}\}\_\{t\}\\in\[\-1,1\]^\{6\}: load\-shift intent, clipped to\[0,1\]\[0,1\]and multiplied by that device’s shift cap, postponing demand to a later slot inside theH=4H\{=\}4\-step \(two\-hour\) shift horizon \(shifted energy must be repaid withinHHsteps\)\.

Reward and Violation Cost\.The reward is the negative total network lossPtlossP^\{\\mathrm\{loss\}\}\_\{t\}\(in MW summed over allNℓN\_\{\\ell\}branches\)\. The violation costCtvoltageC^\{\\mathrm\{voltage\}\}\_\{t\}is the count of buses whose voltage falls outside the safe band on the current half\-hour slot \(an integer in\[0,N\]\[0,N\], then averaged into the per\-step rate reported in Section[G\.3](https://arxiv.org/html/2609.36052#A7.SS3)\)\. This channel is reported under Form 1 \(budgeted expected cost\):

rt=−Ptloss,ct=Ctvoltage\.r\_\{t\}=\-P^\{\\mathrm\{loss\}\}\_\{t\},\\qquad c\_\{t\}=C^\{\\mathrm\{voltage\}\}\_\{t\}\.\(35\)
Baselines\.Three heuristic baselines are released:

- •*No Control*: zero curtailment, zero shift;
- •*TOU*\(time\-of\-use\): fixed peak\-window rule that curtails and shifts demand during 16:00–21:00 each day;
- •*Droop*: a local voltage droop response that reacts only to each controllable load’s own bus voltage\.

Four learned baselines are trained on the same observation and reward: PPO, SAC, Sauté PPO, and PPO\-Lagrangian\.

Challenge\.DSO has a smooth and continuous action space, but the reward \(loss\) and the cost \(voltage\) point in opposite directions: lowering loss often pushes voltages closer to the band edges, and the band must hold for every bus on every half\-hour slot\. Voltage couples the six controllers through the shared power flow equations, so learning has to discover a joint policy that respects a global constraint from local actions, which is the standard hard regime for safe\-RL methods even when control itself is smooth\.

### C\.4Distributed energy resource coordination task \(DERs\)

Task\.Twelve heterogeneous distributed energy resources — four batteries, four PV \(photovoltaic\) inverters, and four flexible loads — sit at fixed buses of the 141\-bus radial distribution feedercase141\(see Figure[10](https://arxiv.org/html/2609.36052#A3.F10)\)\. The horizon is 24 hours at 30\-minute resolution, which gives 48 sequential decision steps per episode, and is driven by real Ausgrid demand profiles888Ausgrid distribution demand traces,[https://www\.ausgrid\.com\.au/Industry/Our\-Research/Data\-to\-share/](https://www.ausgrid.com.au/Industry/Our-Research/Data-to-share/)\.for the bus\-level loads and GB solar availability traces999GB solar from Elexon BMRS,[https://bmrs\.elexon\.co\.uk/generation\-by\-fuel\-type](https://bmrs.elexon.co.uk/generation-by-fuel-type)\.for the PV bundles\. Each device is its own agent\. The agents cooperate to minimize total active power loss while keeping every bus voltage inside the safe band\[0\.94,1\.06\]\[0\.94,1\.06\]p\.u\. and respecting line thermal limits\. The actions are heterogeneous: the same 2\-dimensional action slot means different things for batteries, PV inverters, and flexible loads, so this is a partially observable cooperative MARL task with agents grouped by type\.

![Refer to caption](https://arxiv.org/html/2609.36052v1/fig_ders_topology.png)Figure 10:141\-bus radial distribution feeder \(case141\) used by the DERs task\. The substation \(factory icon\) at bus 1 supplies the network through deep radial branches\. Twelve DER agents are sited at fixed buses: 4 batteries \(BESS, teal vertical bars\) at buses 9, 17, 55, 122; 4 PV inverters \(sun\-panel icons\) at buses 6, 72, 73, 82; and 4 flexible loads \(orange houses\) at buses 24, 41, 70, 135\. Small black dots mark passive buses without an attached DER agent\. The 12 agents share a team reward \(negative network active power loss\) and cooperate under partial observability withK=4K\{=\}4BFS\-graph\-neighbor windows \(Section[C\.4](https://arxiv.org/html/2609.36052#A3.SS4)\)\.Observation and Action\.The 141\-bus network is too large for every agent to read every bus, so each agent samples the grid only through a small local window\. Agentiisits at busb⁡\(i\)b\(i\)and reads \(i\) its own bus voltage, \(ii\) the voltages at itsK=4K\{=\}4nearest BFS\-graph\-neighbor buses \(the part of the feeder it directly affects\), \(iii\) a three\-number summary of the system\-wide voltage profile \(its only global signal\), and \(iv\) its own device internal state:

𝐨i,t=\(Vb⁡\(i\),t,𝐕𝒩K​\(i\),t,𝐕~t,τt,𝐳i,tdevice\)∈ℝ15\.\\mathbf\{o\}\_\{i,t\}=\\bigl\(V\_\{b\(i\),t\},\\ \\mathbf\{V\}\_\{\\mathcal\{N\}\_\{K\}\(i\),t\},\\ \\widetilde\{\\mathbf\{V\}\}\_\{t\},\\ \\tau\_\{t\},\\ \\mathbf\{z\}^\{\\mathrm\{device\}\}\_\{i,t\}\\bigr\)\\in\\mathbb\{R\}^\{15\}\.\(36\)The action is per\-agent 2\-dimensional,𝐚i,t∈\[−1,1\]2\\mathbf\{a\}\_\{i,t\}\\in\[\-1,1\]^\{2\}, with type\-specific physical interpretation:

𝐚i,t=\{\(Pi,tbat,Qi,tbat\)battery,\(κi,tpv,Qi,tpv\)PV inverter,\(di,tcurtail,di,tshift\)flexible load\.\\mathbf\{a\}\_\{i,t\}=\\begin\{cases\}\\bigl\(P^\{\\mathrm\{bat\}\}\_\{i,t\},\\ Q^\{\\mathrm\{bat\}\}\_\{i,t\}\\bigr\)&\\text\{battery\},\\\\ \\bigl\(\\kappa^\{\\mathrm\{pv\}\}\_\{i,t\},\\ Q^\{\\mathrm\{pv\}\}\_\{i,t\}\\bigr\)&\\text\{PV inverter\},\\\\ \\bigl\(d^\{\\mathrm\{curtail\}\}\_\{i,t\},\\ d^\{\\mathrm\{shift\}\}\_\{i,t\}\\bigr\)&\\text\{flexible load\}\.\\end\{cases\}\(37\)The five observation entries are:

- •Vb⁡\(i\),t∈ℝV\_\{b\(i\),t\}\\in\\mathbb\{R\}: voltage at the agent’s own bus, centered and scaled the same way as in the DSO observation;
- •𝐕𝒩K​\(i\),t∈ℝK\\mathbf\{V\}\_\{\\mathcal\{N\}\_\{K\}\(i\),t\}\\in\\mathbb\{R\}^\{K\}: voltages at the agent’sK=4K\{=\}4nearest BFS\-graph\-neighbor buses, ordered by graph distance on the feeder topology;
- •𝐕~t=\(Vmin,Vmax,Vmean\)∈ℝ3\\widetilde\{\\mathbf\{V\}\}\_\{t\}=\(V\_\{\\min\},V\_\{\\max\},V\_\{\\mathrm\{mean\}\}\)\\in\\mathbb\{R\}^\{3\}: min, max, and mean of the full bus voltage vector — a coarse global signal that does not reveal per\-bus details;
- •τt∈ℝ2\\tau\_\{t\}\\in\\mathbb\{R\}^\{2\}: sin/cos time\-of\-day encoding;
- •𝐳i,tdevice∈ℝ5\\mathbf\{z\}^\{\\mathrm\{device\}\}\_\{i,t\}\\in\\mathbb\{R\}^\{5\}: the agent’s own device internal state, sized to the maximum per\-device observation length across DER types and zero\-padded for shorter types so all twelve agents share a fixed 15\-dimensional input shape under typed parameter sharing\. By type:33entries \[SoC, normalizedpp, normalizedqq\] for a battery;44entries \[capacity factor \(current solar output relative to nameplate\), dispatchedpp, dispatchedqq, curtailment level\] for a PV inverter;55entries \[current curtailment, outgoing shift fraction, incoming shift fraction, deferred\-buffer fill ratio, deferred\-buffer energy\] for a flexible load\.

The dimension count is1\+K\+3\+2\+5=151\+K\+3\+2\+5=15\.

The action is type\-specific \(the same 2\-dim slot but different physical meaning\):

- •*Battery*:\(Pi,tbat,Qi,tbat\)\(P^\{\\mathrm\{bat\}\}\_\{i,t\},Q^\{\\mathrm\{bat\}\}\_\{i,t\}\)active and reactive power, denormalized linearly from\[−1,1\]2\[\-1,1\]^\{2\}to the inverter capability envelope\{\(p,q\):p2\+q2≤\(Simax\)2\}\\\{\(p,q\):p^\{2\}\+q^\{2\}\\leq\(S^\{\\max\}\_\{i\}\)^\{2\}\\\}\(Section[B\.4](https://arxiv.org/html/2609.36052#A2.SS4)\) and clipped against the SoC\-feasible discharge / charge bounds;
- •*PV inverter*:\(κi,tpv,Qi,tpv\)\(\\kappa^\{\\mathrm\{pv\}\}\_\{i,t\},Q^\{\\mathrm\{pv\}\}\_\{i,t\}\)a curtailment factorκ∈\[0,1\]\\kappa\\in\[0,1\]\(fraction of available solar power to use\) and a reactive\-power setpoint, projected onto the same PQ envelope \(Section[B\.4](https://arxiv.org/html/2609.36052#A2.SS4)\) withpppriority;
- •*Flexible load*:\(di,tcurtail,di,tshift\)\(d^\{\\mathrm\{curtail\}\}\_\{i,t\},d^\{\\mathrm\{shift\}\}\_\{i,t\}\)curtailment and load\-shift fractions exactly as in the DSO task — clipped to\[0,1\]\[0,1\], scaled by per\-device caps, with a shift horizonH=4H\{=\}4steps\.

Reward and Violation Cost\.All twelve agents share a single team reward, the negative total active power lossPtlossP^\{\\mathrm\{loss\}\}\_\{t\}in MW summed over allNℓN\_\{\\ell\}branches:

rt=−Ptloss,𝐜t=\(Ctvoltage,Ctthermal,Ctresource\)\.r\_\{t\}=\-P^\{\\mathrm\{loss\}\}\_\{t\},\\qquad\\mathbf\{c\}\_\{t\}=\\bigl\(C^\{\\mathrm\{voltage\}\}\_\{t\},\\ C^\{\\mathrm\{thermal\}\}\_\{t\},\\ C^\{\\mathrm\{resource\}\}\_\{t\}\\bigr\)\.\(38\)The three\-channel violation cost separately reports:

- •CtvoltageC^\{\\mathrm\{voltage\}\}\_\{t\}: count of buses whose voltage falls outside the safe band\[0\.94,1\.06\]\[0\.94,1\.06\]p\.u\. on the current step;
- •CtthermalC^\{\\mathrm\{thermal\}\}\_\{t\}: count of branches whose apparent power flow\|S\|\|S\|exceeds the line capacity on the current step;
- •CtresourceC^\{\\mathrm\{resource\}\}\_\{t\}: per\-step bundle physical cost — e\.g\. an opportunity cost when a PV inverter curtails available solar power, or a clipping penalty when a battery dispatch hits its SoC bounds\.

All three channels are reported under Form 1 \(budgeted expected cost\)\.

Baselines\.Two heuristic baselines are released:

- •*No Control*: every device idle \(batteryP=Q=0P\{=\}Q\{=\}0, PV at MPPT with no curtailment, flexible loads serving nominal demand\);
- •*Voltage Droop*: a local reactive\-power droop rule keyed to each agent’s own bus voltage\.

Three learned baselines use type\-specific parameter sharing in IPPO, with one shared actor\-critic per agent type:

- •*IPPO*: the reference variant, trained with a voltage\-violation penalty of weight 4\.0;
- •*IPPO\-rs*: the same with the penalty doubled to 8\.0;
- •*IPPO\-Lagrangian*: separate cost critics for the three cost channels with a shared dual variable across types\.

Challenge\.DERs combines partial observability \(each agent only sees aK=4K\{=\}4local window plus a three\-number global summary\), heterogeneous agents \(batteries, PV, and flexible loads share a reward but not an action interface\), and a network\-wide constraint \(voltage band\) that couples all twelve agents through the power flow equations\. The reward is a single global signal \(loss in MW\), while the safety violations can fire from any one of the twelve agents, so credit assignment is hard\. These remain open problems for cooperative MARL, especially at feeder sizes where no agent can hold a global view\.

### C\.5Data Center Microgrid task \(DCMG\)

Task\.A microgrid powers an AI data center with on\-site PV, battery storage, a diesel generator, and a connection to the main grid\. A single operator schedules training and fine\-tuning workloads, cooling, battery dispatch, diesel generation, and grid import over a 24\-hour horizon at 5\-minute resolution \(288 steps per episode\), driven by real Google data center workload traces\[[30](https://arxiv.org/html/2609.36052#bib.bib16)\], GB solar availability traces101010GB solar from Elexon BMRS,[https://bmrs\.elexon\.co\.uk/generation\-by\-fuel\-type](https://bmrs.elexon.co.uk/generation-by-fuel-type)\.for the on\-site PV, and GB MID market grid\-import prices111111GB Market Index Data \(MID\) from Elexon BMRS,[https://data\.elexon\.co\.uk/bmrs/api/v1/datasets/MID](https://data.elexon.co.uk/bmrs/api/v1/datasets/MID)\.\. The microgrid runs grid\-connected by default; under simulated outage scenarios it switches to islanded mode where unmet demand counts as a hard violation\. The setting is a long\-horizon, multi\-objective single\-agent RL problem: trade off energy use, cost, and carbon while respecting service\-level, thermal, and power\-balance constraints\.

Observation and Action\.At each step the operator observes a joint workload, thermal, and energy snapshot:

𝐨t=\(CLOSE\\displaystyle\\mathbf\{o\}\_\{t\}=\\bigl\(utcpu,utmem,qttrain,qtft,ηt,Ttzone,Ttout,COPt,αtpv,SoCt,\\displaystyle u^\{\\mathrm\{cpu\}\}\_\{t\},\\ u^\{\\mathrm\{mem\}\}\_\{t\},\\ q^\{\\mathrm\{train\}\}\_\{t\},\\ q^\{\\mathrm\{ft\}\}\_\{t\},\\ \\eta\_\{t\},\\ T^\{\\mathrm\{zone\}\}\_\{t\},\\ T^\{\\mathrm\{out\}\}\_\{t\},\\ \\mathrm\{COP\}\_\{t\},\\ \\alpha^\{\\mathrm\{pv\}\}\_\{t\},\\ \\mathrm\{SoC\}\_\{t\},\(39\)OPENmtdg,ptload,ptnet,βtdis,βtchg,πtgrid,π¯tgrid,6​h,𝐚t−1,τt\)∈ℝ24\.\\displaystyle m^\{\\mathrm\{dg\}\}\_\{t\},\\ p^\{\\mathrm\{load\}\}\_\{t\},\\ p^\{\\mathrm\{net\}\}\_\{t\},\\ \\beta^\{\\mathrm\{dis\}\}\_\{t\},\\ \\beta^\{\\mathrm\{chg\}\}\_\{t\},\\ \\pi^\{\\mathrm\{grid\}\}\_\{t\},\\ \\overline\{\\pi\}^\{\\mathrm\{grid\},6h\}\_\{t\},\\ \\mathbf\{a\}\_\{t\-1\},\\ \\tau\_\{t\}\\bigr\)\\in\\mathbb\{R\}^\{24\}\.𝐚t=\(attrain,atft,atcool,atbatt,atdg\)∈\[0,1\]3×\[−1,1\]×\[0,1\]\.\\mathbf\{a\}\_\{t\}=\\bigl\(a^\{\\mathrm\{train\}\}\_\{t\},\\ a^\{\\mathrm\{ft\}\}\_\{t\},\\ a^\{\\mathrm\{cool\}\}\_\{t\},\\ a^\{\\mathrm\{batt\}\}\_\{t\},\\ a^\{\\mathrm\{dg\}\}\_\{t\}\\bigr\)\\in\[0,1\]^\{3\}\\times\[\-1,1\]\\times\[0,1\]\.\(40\)The 24 observation entries split into five thematic blocks plus the previous action and a time encoding:

- •*Workload state*\(5 dims\):utcpu,utmem∈\[0,1\]u^\{\\mathrm\{cpu\}\}\_\{t\},u^\{\\mathrm\{mem\}\}\_\{t\}\\in\[0,1\]are GPU compute and memory utilization;qttrain,qtft∈\[0,1\]q^\{\\mathrm\{train\}\}\_\{t\},q^\{\\mathrm\{ft\}\}\_\{t\}\\in\[0,1\]are the training and fine\-tuning queue fills \(waiting GPU demand normalized by total GPU count\);ηt∈\[−1,1\]\\eta\_\{t\}\\in\[\-1,1\]is the deadline urgency of the most pressing waiting job \(negative means already overdue\);
- •*Thermal state*\(3 dims\):Ttzone∈\[0,1\]T^\{\\mathrm\{zone\}\}\_\{t\}\\in\[0,1\]the server\-zone temperature,Ttout∈\[0,1\]T^\{\\mathrm\{out\}\}\_\{t\}\\in\[0,1\]the outdoor temperature \(both centered and scaled to the operating range\), andCOPt∈\[0,1\]\\mathrm\{COP\}\_\{t\}\\in\[0,1\]the chiller’s coefficient of performance \(higher means more cooling delivered per unit electricity\);
- •*Energy assets*\(3 dims\):αtpv∈\[0,1\]\\alpha^\{\\mathrm\{pv\}\}\_\{t\}\\in\[0,1\]the PV capacity factor \(current solar output relative to nameplate\),SoCt∈\[0,1\]\\mathrm\{SoC\}\_\{t\}\\in\[0,1\]the battery state of charge, andmtdg∈\[0,1\]m^\{\\mathrm\{dg\}\}\_\{t\}\\in\[0,1\]the remaining diesel headroom;
- •*Power balance and battery headroom*\(4 dims\):ptload∈\[0,1\]p^\{\\mathrm\{load\}\}\_\{t\}\\in\[0,1\]the total data\-center electrical load \(normalized by total supply capacity\),ptnet∈\[−1,1\]p^\{\\mathrm\{net\}\}\_\{t\}\\in\[\-1,1\]the residual net load after PV \(signed\), andβtdis,βtchg∈\[0,1\]\\beta^\{\\mathrm\{dis\}\}\_\{t\},\\beta^\{\\mathrm\{chg\}\}\_\{t\}\\in\[0,1\]the battery’s discharge and charge headroom as fractions of rated power;
- •*Grid price signal*\(2 dims\):πtgrid∈\[0,1\]\\pi^\{\\mathrm\{grid\}\}\_\{t\}\\in\[0,1\]the current GB MID grid\-import price \(normalized\) andπ¯tgrid,6​h∈\[0,1\]\\overline\{\\pi\}^\{\\mathrm\{grid\},6h\}\_\{t\}\\in\[0,1\]the maximum grid price expected over the next six hours, giving the policy a forward look\-ahead for arbitrage;
- •*Previous action and time encoding*\(5 \+ 2 dims\):𝐚t−1∈ℝ5\\mathbf\{a\}\_\{t\-1\}\\in\\mathbb\{R\}^\{5\}the previous\-step action andτt∈ℝ2\\tau\_\{t\}\\in\\mathbb\{R\}^\{2\}the sin/cos time\-of\-day encoding\.

The dimension count is11\+6\+5\+2=2411\+6\+5\+2=24\.

The 5\-dimensional action has mixed bounds:

- •attrain,atft∈\[0,1\]a^\{\\mathrm\{train\}\}\_\{t\},a^\{\\mathrm\{ft\}\}\_\{t\}\\in\[0,1\]: fractions of the training and fine\-tuning queues dispatched in the next 5\-minute slot;
- •atcool∈\[0,1\]a^\{\\mathrm\{cool\}\}\_\{t\}\\in\[0,1\]: cooling setpoint command \(lower means a colder zone setpoint, requiring more cooling power\);
- •atbatt∈\[−1,1\]a^\{\\mathrm\{batt\}\}\_\{t\}\\in\[\-1,1\]: battery dispatch signal \(negative charges, positive discharges; scaled by rated power and clipped against SoC\-feasible bounds inside the bundle\);
- •atdg∈\[0,1\]a^\{\\mathrm\{dg\}\}\_\{t\}\\in\[0,1\]: diesel generator preference — realized as a same\-step residual slack actuator after the workload, cooling, and battery decisions are applied, soadga^\{\\mathrm\{dg\}\}acts as a soft preference rather than a strict setpoint\.

Reward and Violation Cost\.The reward sums an energy\-use term, a grid\-cost term, and a carbon\-emissions term, with scalar weightswcost,wcarbonw\_\{\\mathrm\{cost\}\},w\_\{\\mathrm\{carbon\}\}that the operator can pick to encode a Pareto preference \(the canonical benchmark useswcost=0\.2w\_\{\\mathrm\{cost\}\}=0\.2,wcarbon=0\.1w\_\{\\mathrm\{carbon\}\}=0\.1\):

rt=rtenergy\+wcost​rtcost\+wcarbon​rtcarbon,𝐜t=\(CtSLA,Cttemp,Ctpwr\)\.r\_\{t\}=r^\{\\mathrm\{energy\}\}\_\{t\}\+w\_\{\\mathrm\{cost\}\}\\,r^\{\\mathrm\{cost\}\}\_\{t\}\+w\_\{\\mathrm\{carbon\}\}\\,r^\{\\mathrm\{carbon\}\}\_\{t\},\\qquad\\mathbf\{c\}\_\{t\}=\\bigl\(C^\{\\mathrm\{SLA\}\}\_\{t\},\\ C^\{\\mathrm\{temp\}\}\_\{t\},\\ C^\{\\mathrm\{pwr\}\}\_\{t\}\\bigr\)\.\(41\)The three\-channel violation cost separately reports:

- •CtSLAC^\{\\mathrm\{SLA\}\}\_\{t\}: workload incompletion — the per\-step count of expired tasks normalized by the GPU count;
- •CttempC^\{\\mathrm\{temp\}\}\_\{t\}: server\-zone over\-temperature,Cttemp=\(Ttzone−Tmax\)\+C^\{\\mathrm\{temp\}\}\_\{t\}=\(T^\{\\mathrm\{zone\}\}\_\{t\}\-T^\{\\max\}\)\_\{\+\};
- •CtpwrC^\{\\mathrm\{pwr\}\}\_\{t\}: per\-step power deficit — the shortfall when on\-site generation, battery, and grid import together cannot serve the data\-center load\. This channel is the hard constraint that becomes binding under simulated outage scenarios when the microgrid is forced into islanded mode\.

All three channels are reported under Form 1 \(budgeted expected cost\)\.

Baselines\.Three heuristic baselines are released:

- •*No Control*: always\-on policy with no scheduling;
- •*Max renewable*: serve workload from PV first, then battery, then diesel;
- •*Rule\-based*: a heuristic scheduler that respects SLA, temperature, and SoC bands\.

Two learned baselines are trained with fixed penalty weights on the SLA, over\-temperature, and deficit channels added to the reward:

- •*PPO*with a Beta distribution policy head for bounded actions;
- •*SAC*with a tanh\-squashed Gaussian\.

Challenge\.DCMG is a long\-horizon \(288 steps\), multi\-objective single\-agent task with three hard constraint channels\. The objectives \(energy, cost, carbon\) and the constraints \(SLA, over\-temperature, power deficit\) are cross\-coupled: deferring training jobs saves energy but risks SLA misses; running cooling harder protects servers but burns diesel; charging the battery during PV peaks improves later resilience but uses up current SLA headroom\. Credit must be propagated across hundreds of 5\-minute steps under stochastic workload, weather, and PV traces, which is precisely the regime that current model\-free RL handles poorly\.

## Appendix DData Sources and Splits

### D\.1Data and splits

Table[6](https://arxiv.org/html/2609.36052#A4.T6)summarizes the external time\-series traces used in the reported results, the network case for each task, the splits, and the seed and episode budget\. Processed copies of these traces are shipped with the PowerZooJax code release\.

Table 6:External time\-series traces, splits, and seed and episode budgets per task\. GB denotes Great Britain\. Bold marks the primary split used for the headline metric reported in Section[G](https://arxiv.org/html/2609.36052#A7); for DCMG the appendix\-only splits follow the semicolon\.TaskSystem / caseTime\-series source\(s\)SplitsSeed budgetGenCos5\-bus marketGB demandtrain,in\-distribution, demand shift, renewable shock5 seeds, 30 episodes/seedTSO118\-bus transmissionGB demand, wind, and solartrain,in\-distribution, load stress, line tightening5 seeds, 50 episodes/seedDSO33\-bus feederAusgrid distribution demandin\-distribution5 seeds, 50 episodes/seedDERs141\-bus feederAusgrid distribution demand and GB solartrain,in\-distribution, voltage tightening, PV shift, load stress5 seeds, 30 episodes/seedDCMGBehind\-the\-meter microgridGoogle workload \(Azure and Alibaba for workload\-OOD splits\); GB solar; GB MID pricetrain,in\-distribution, cooling stress, renewable drought; appendix: workload swap, workload shock, dg derating, SLA tightening5 seeds, 10 episodes/seed
### D\.2Split and stress\-test taxonomy

Each split asks a different physical question\. Table[7](https://arxiv.org/html/2609.36052#A4.T7)groups them by the mechanism that changes between train and evaluation or between routine and stressed evaluation, and states the role of each split\.

Table 7:Split and stress\-test taxonomy\. Each row states the physical mechanism and the role of the split\.Split typeTasks / split namesPhysical mechanismRoleRoutine in\-distributionTSO, DSO, DERs, GenCos, DCMG /in\-distributionHeld\-out episodes from the routine evaluation regimeMain task comparison for the primary metricLoad / demand shiftTSO / load stress; DERs / load stress; GenCos / demand shiftHigher or shifted demand changes reserve pressure, feeder loading, or market clearingWhether an in\-distribution policy remains useful when demand changesRenewable / PV shiftDERs / PV shift; GenCos / renewable shock; DCMG / renewable droughtRenewable generation changes the controllable/exogenous balanceSensitivity to renewable availabilityNetwork / security tighteningTSO / line tightening; DERs / voltage tighteningTighter physical operating limits \(line capacity or voltage band\)Whether lower cost or lower loss survives under tighter operating limitsCooling / workload stressDCMG / cooling stress \(main\); workload swap, workload shock, SLA tightening \(appendix\)Data center thermal load, workload mix, or service deadlines changeLong\-horizon resource scheduling under thermal, workload, or service stressResource\-availability stressDCMG / dg deratingBackup\-generation capacity is reducedSensitivity to reduced backup capacity

## Appendix EQuickstart and API

This section is a copy\-pasteable walk\-through of how to install PowerZooJax, run a minimal rollout, read a task config, read a training config, and trigger a benchmark run from the CLI\. Listings E\.1–E\.3 are minimal Python demos and theTaskSpecProtocol surface, Listings E\.4–E\.5 are abridged from the released config files, and the bash blocks in Sections[E\.1](https://arxiv.org/html/2609.36052#A5.SS1)and[E\.5](https://arxiv.org/html/2609.36052#A5.SS5)cover install and CLI usage\.

### E\.1Installation

PowerZooJax requires Python 3\.10–3\.12\. The base install ships a CPU\-only JAX; CUDA 12 support is opt\-in via thecuda12extra; the JAX single\-agent RL stack lives behindrl\(rejax,distrax\) and the multi\-agent stack behindmarl\(jaxmarl\); the benchmark workflow lives behindbenchmarks\(Stable\-Baselines3, SBX, PettingZoo\)\.

gitclonehttps://github\.com/powerzoojax/PowerZooJax&&cdPowerZooJax

pipinstall\-e"\.\[rl,marl,benchmarks\]"

pipinstall\-e"\.\[cuda12,rl,marl,benchmarks\]"

exportXLA\_PYTHON\_CLIENT\_PREALLOCATE=false

### E\.2Minimal rollout

The example below opens a transmission environment oncase5, resets it, then runs a 48\-step rollout underjax\.lax\.scanwith a fixed action schedule\. Every transition stays inside JIT\.

Listing E\.1: Minimal rollout on case5 \(Python\)[⬇](data:text/plain;base64,aW1wb3J0IGpheCwgamF4Lm51bXB5IGFzIGpucApmcm9tIHBvd2Vyem9vamF4LmNhc2UgaW1wb3J0IGNyZWF0ZV9jYXNlNQpmcm9tIHBvd2Vyem9vamF4LmVudnMgaW1wb3J0IFRyYW5zR3JpZEVudiwgbWFrZV90cmFuc19wYXJhbXMKZnJvbSBwb3dlcnpvb2pheC51dGlscyBpbXBvcnQgc2Nhbl9yb2xsb3V0CgplbnYgPSBUcmFuc0dyaWRFbnYoKQpwYXJhbXMgPSBtYWtlX3RyYW5zX3BhcmFtcyhjcmVhdGVfY2FzZTUoKSkKa2V5ID0gamF4LnJhbmRvbS5QUk5HS2V5KDApCnJlc2V0X2tleSwgcm9sbG91dF9rZXkgPSBqYXgucmFuZG9tLnNwbGl0KGtleSkKb2JzLCBzdGF0ZSA9IGVudi5yZXNldChyZXNldF9rZXksIHBhcmFtcykKYWN0aW9ucyA9IGpucC56ZXJvcygoNDgsICplbnYuYWN0aW9uX3NwYWNlKHBhcmFtcykuc2hhcGUpLCBkdHlwZT1qbnAuZmxvYXQzMikKZmluYWxfc3RhdGUsIG9ic190cmFqLCByZXdhcmRfdHJhaiwgY29zdF90cmFqLCBkb25lX3RyYWosIGluZm9fdHJhaiA9IHNjYW5fcm9sbG91dCgKICAgIGVudiwgcm9sbG91dF9rZXksIHN0YXRlLCBwYXJhbXMsIGFjdGlvbnMKKQ==)importjax,jax\.numpyasjnpfrompowerzoojax\.caseimportcreate\_case5frompowerzoojax\.envsimportTransGridEnv,make\_trans\_paramsfrompowerzoojax\.utilsimportscan\_rolloutenv=TransGridEnv\(\)params=make\_trans\_params\(create\_case5\(\)\)key=jax\.random\.PRNGKey\(0\)reset\_key,rollout\_key=jax\.random\.split\(key\)obs,state=env\.reset\(reset\_key,params\)actions=jnp\.zeros\(\(48,\*env\.action\_space\(params\)\.shape\),dtype=jnp\.float32\)final\_state,obs\_traj,reward\_traj,cost\_traj,done\_traj,info\_traj=scan\_rollout\(env,rollout\_key,state,params,actions\)

The same env contract works for the other physical models via the task\-level wrapper:

Listing E\.2: Task\-level interface \(Python\)[⬇](data:text/plain;base64,ZnJvbSBwb3dlcnpvb2pheC50YXNrcy50c28gaW1wb3J0IFRTT1Rhc2sKaW1wb3J0IGpheAp0YXNrID0gVFNPVGFzaygpCmVudiA9IHRhc2subWFrZV9lbnYoc3BsaXQ9ImlpZCIpCnBhcmFtcyA9IHRhc2suZXBpc29kZV9wYXJhbXMoCiAgICAiaWlkIiwgZXBpc29kZV9pZHg9MCwgbl9lcGlzb2Rlcz0xLCBtYXhfc3RlcHM9NDgsCiAgICBzdHJhdGVneT0ic2VlZGVkIiwgc2VlZD0wLAopCm9icywgc3RhdGUgPSBlbnYucmVzZXQoamF4LnJhbmRvbS5QUk5HS2V5KDApLCBwYXJhbXMp)frompowerzoojax\.tasks\.tsoimportTSOTaskimportjaxtask=TSOTask\(\)env=task\.make\_env\(split="iid"\)params=task\.episode\_params\("iid",episode\_idx=0,n\_episodes=1,max\_steps=48,strategy="seeded",seed=0,\)obs,state=env\.reset\(jax\.random\.PRNGKey\(0\),params\)

The five task classes live inpowerzoojax\.tasks\.\{tso,dso,ders,gencos,dc\_microgrid\}and all satisfy theTaskSpecProtocol \(Listing E\.3, abridged frompowerzoojax/tasks/base\.py\)\.

Listing E\.3:TaskSpecProtocol surface \(powerzoojax/tasks/base\.py, abridged\)[⬇](data:text/plain;base64,ZnJvbSB0eXBpbmcgaW1wb3J0IEFueSwgTGl0ZXJhbCwgUHJvdG9jb2wsIHJ1bnRpbWVfY2hlY2thYmxlCgpAcnVudGltZV9jaGVja2FibGUKY2xhc3MgVGFza1NwZWMoUHJvdG9jb2wpOgogICAgIiIiU3RydWN0dXJhbCBpbnRlcmZhY2UgZm9yIHRoZSBmaXZlIGJlbmNobWFyayB0YXNrcy4iIiIKCiAgICB0YXNrX25hbWU6IHN0ciAgICAgICAgICAgICAgICAgICAgICAgICAgIyAiZHNvIiB8ICJ0c28iIHwgImRlcnMiIHwgImdlbmNvcyIgfCAiZGNfbWljcm9ncmlkIgogICAgZGVmYXVsdF9zcGxpdHM6IHR1cGxlW3N0ciwgLi4uXSAgICAgICAgICMgZS5nLiAoInRyYWluIiwgImlpZCIsICJsb2FkX3N0cmVzcyIsIC4uLikKCiAgICBkZWYgbWFrZV9lbnYoc2VsZiwgc3BsaXQ6IHN0ciA9ICJ0cmFpbiIpIC0+IEFueTogLi4uCgogICAgZGVmIGVwaXNvZGVfcGFyYW1zKAogICAgICAgIHNlbGYsCiAgICAgICAgc3BsaXQ6IHN0ciwgZXBpc29kZV9pZHg6IGludCwgbl9lcGlzb2RlczogaW50LCBtYXhfc3RlcHM6IGludCwKICAgICAgICAqLCBzdHJhdGVneTogTGl0ZXJhbFsidW5pZm9ybSIsICJzZWVkZWQiXSA9ICJ1bmlmb3JtIiwgc2VlZDogaW50ID0gMCwKICAgICkgLT4gQW55OiAuLi4KCiAgICBkZWYgcm9sbG91dChzZWxmLCBlbnY6IEFueSwgcGFyYW1zOiBBbnksIGtleTogQW55LCBwb2xpY3lfZm46IEFueSkgLT4gQW55OiAuLi4KICAgIGRlZiBjb25zdHJhaW50X3NwZWMoc2VsZikgLT4gIkNvbnN0cmFpbnRTcGVjIjogLi4u)fromtypingimportAny,Literal,Protocol,runtime\_checkable@runtime\_checkableclassTaskSpec\(Protocol\):"""Structuralinterfaceforthefivebenchmarktasks\."""task\_name:strdefault\_splits:tuple\[str,\.\.\.\]defmake\_env\(self,split:str="train"\)\-\>Any:\.\.\.defepisode\_params\(self,split:str,episode\_idx:int,n\_episodes:int,max\_steps:int,\*,strategy:Literal\["uniform","seeded"\]="uniform",seed:int=0,\)\-\>Any:\.\.\.defrollout\(self,env:Any,params:Any,key:Any,policy\_fn:Any\)\-\>Any:\.\.\.defconstraint\_spec\(self\)\-\>"ConstraintSpec":\.\.\.

### E\.3Reading a task config

Every benchmark task is specified by a YAML file atbenchmarks/<task\>/configs/task\.yaml\. The TSO file is reproduced below in abridged form\.

Listing E\.4: TSO task config \(benchmarks/tso/configs/task\.yaml, abridged\)[⬇](data:text/plain;base64,dGFzazogdHNvCmNhc2U6IGNhc2UxMTggICAgICAgICAgICAgICAgICMgSUVFRSAxMTgtYnVzIC8gMTg2LWxpbmUgdHJhbnNtaXNzaW9uIGNhc2UKbl91bml0czogNTQgICAgICAgICAgICAgICAgICAgIyBUaGVybWFsIHVuaXRzIGNvbW1pdHRlZCBieSBTQ1VDCm1heF9zdGVwczogNDggICAgICAgICAgICAgICAgICMgNDggaGFsZi1ob3VyIGludGVydmFscyA9IDI0IGggaG9yaXpvbgpkdF9ob3VyczogMC41CgpkYXRhX3NvdXJjZTogZ2IgICAgICAgICAgICAgICAjIFJlYWwgR0IgdHJhbnNtaXNzaW9uIGRlbWFuZCAmIHdpbmQgdHJhY2UKZ2JfdHJhaW5fc3RhcnQ6ICIyMDI1LTA0LTAxIgpnYl90cmFpbl9lbmQ6ICAgIjIwMjUtMTItMzEiCmdiX2lpZF9zdGFydDogICAiMjAyNi0wMS0wMSIKZ2JfaWlkX2VuZDogICAgICIyMDI2LTAzLTMxIgoKcmVzZXJ2ZV9tYXJnaW5fZnJhYzogMC4wNSAgICAgIyBTQ1VDIHJlc2VydmUgbWFyZ2luCnJld2FyZF9zY2FsZTogMC4wMDAxICAgICAgICAgICMgcmV3YXJkIHNjYWxpbmcgZmFjdG9yCnNvbHZlcl9tb2RlOiAxICAgICAgICAgICAgICAgICMgREMtT1BGIHNvbHZlciBtb2RlCmRjb3BmX21heF9pdGVyOiAyMDAKZGNvcGZfdG9sOiAwLjAwMDEKZm9yZWNhc3RfaG9yaXpvbl9zdGVwczogNCAgICAgIyA0IHN0ZXBzID0gMiBoOyBzZXRzIFRTTyBvYnMgdG8gNDEwLWRpbSAoU2VjdGlvbiBDLjIpCgpiYXNlbGluZV9zZXQ6IFthbGxfb24sIG1lcml0X29yZGVyXQpldmFsX3NwbGl0czogIFt0cmFpbiwgaWlkLCBsb2FkX3N0cmVzcywgbGluZV90aWdodGVuaW5nXQpwcmltYXJ5X3NwbGl0OiBpaWQKc2VlZHM6IFswLCAxLCAyLCAzLCA0XQpldmFsX2VwaXNvZGVzOiA1MAoKc2FmZXR5X3RocmVzaG9sZHM6CiAgcmVzZXJ2ZV9zaG9ydGZhbGxfcmF0ZTogMC4wCiAgdGhlcm1hbF92aW9sYXRpb25fcmF0ZTogMC4w)task:tsocase:case118n\_units:54max\_steps:48dt\_hours:0\.5data\_source:gbgb\_train\_start:"2025\-04\-01"gb\_train\_end:"2025\-12\-31"gb\_iid\_start:"2026\-01\-01"gb\_iid\_end:"2026\-03\-31"reserve\_margin\_frac:0\.05reward\_scale:0\.0001solver\_mode:1dcopf\_max\_iter:200dcopf\_tol:0\.0001forecast\_horizon\_steps:4baseline\_set:\[all\_on,merit\_order\]eval\_splits:\[train,iid,load\_stress,line\_tightening\]primary\_split:iidseeds:\[0,1,2,3,4\]eval\_episodes:50safety\_thresholds:reserve\_shortfall\_rate:0\.0thermal\_violation\_rate:0\.0

### E\.4Reading a training config

Training\-side hyperparameters live next to the task config astrain\_<algo\>\.yaml\. The TSO PPO file is below; the master per\-algorithm\-per\-task hyperparameter table is in Section[F\.2](https://arxiv.org/html/2609.36052#A6.SS2)\.

Listing E\.5: TSO PPO config \(benchmarks/tso/configs/train\_ppo\.yaml\)[⬇](data:text/plain;base64,YWxnbzogcHBvCnRvdGFsX3RpbWVzdGVwczogMjAwMDAwMDAKbnVtX2VudnM6IDI1NiAgICAgICAgICAgICAgICAgIyBHUFUgcGFyYWxsZWxpc20gZm9yIGNhc2UxMTggKDU0IHVuaXRzLCA0MTAtZGltIG9icykKbl9zdGVwczogNDggICAgICAgICAgICAgICAgICAgIyBvbmUgZXBpc29kZSBwZXIgdXBkYXRlCmhpZGRlbl9kaW1zOiBbMjU2LCAyNTZdCmdhbW1hOiAwLjk5NQpnYWVfbGFtYmRhOiAwLjk1CmxyOiAwLjAwMDMKbm9ybWFsaXplX29ic2VydmF0aW9uczogdHJ1ZQp3cmFwcGVyOiBsb2cgICAgICAgICAgICAgICAgICAjIHVzZSAnc2FmZScgZm9yIENNRFAgY29zdCBjaGFubmVsCmV2YWxfZnJlcTogMTAwMDAwCmV2YWxfZXBpc29kZXM6IDggICAgICAgICAgICAgICMgaW4tdHJhaW5pbmcgbW9uaXRvciBvbmx5OyBiZW5jaG1hcmsgZXZhbCB1c2VzIHRhc2sueWFtbDo6ZXZhbF9lcGlzb2RlcwpyZWNvcmRfZXZhbF93YWxsX3RpbWU6IHRydWUgICAjIGxvZ3MgcGVyLWV2YWwgd2FsbHRpbWUgZm9yIGNyb3NzLWJhY2tlbmQgY3VydmVz)algo:ppototal\_timesteps:20000000num\_envs:256n\_steps:48hidden\_dims:\[256,256\]gamma:0\.995gae\_lambda:0\.95lr:0\.0003normalize\_observations:truewrapper:logeval\_freq:100000eval\_episodes:8record\_eval\_wall\_time:true

### E\.5CLI commands

The benchmark workflow exposes five entry points per task:baseline\(run non\-learning baselines\),train\(run a learning algorithm\),eval\(evaluate a trained run on a split\),summarize\(rebuild result summaries\), andplots\(regenerate paper figures\)\. The TSO end\-to\-end example is

pythonbenchmarks/tso/run\.pybaseline\-\-seeds0,1,2,3,4

pythonbenchmarks/tso/run\.pytrain\-\-algoppo\-\-seed0

pythonbenchmarks/tso/run\.pyeval\-\-run\-id<id\>\-\-splitiid

pythonbenchmarks/tso/run\.pyeval\-\-run\-id<id\>\-\-splitline\_tightening

pythonbenchmarks/tso/run\.pysummarize

pythonbenchmarks/tso/run\.pyplots

A lighter ad\-hoc CLI lives atpython \-m powerzoojaxfor preset and YAML\-driven runs:

python\-mpowerzoojax\-\-list\-presets

python\-mpowerzoojax\-\-presetcase5\-economic\-dispatch\-\-seed0

python\-mpowerzoojax\-\-presetcase5\-economic\-dispatch\-\-configexperiment\.yaml\-\-outputresult\.json

## Appendix FTraining Setup and Hyperparameters

This section gives the algorithm\-choice rationale, full per\-task per\-algorithm hyperparameters, network architectures, and evaluation protocol behind every reported number in the paper\.

### F\.1Algorithm choices

- •TSO, DSO, DCMG \(single\-agent\)\.PPO\[[36](https://arxiv.org/html/2609.36052#bib.bib30)\]as the unconstrained baseline; PPO\-Lagrangian as the safe\-RL baseline on TSO and DSO, with Sauté PPO\[[43](https://arxiv.org/html/2609.36052#bib.bib36)\]additionally reported on DSO\. DSO and DCMG additionally train SAC\[[38](https://arxiv.org/html/2609.36052#bib.bib39)\]; on DCMG, SAC supplies the appendix PPO\-vs\-SAC comparison in Section[G\.5](https://arxiv.org/html/2609.36052#A7.SS5)and the representative episode in main\-text Section[6\.4](https://arxiv.org/html/2609.36052#S6.SS4)\.
- •DERs \(cooperative MARL\)\.IPPO\[[37](https://arxiv.org/html/2609.36052#bib.bib20)\]with type\-specific parameter sharing: the 12 agents are partitioned into three types \(4 batteries, 4 PV inverters, 4 flexible loads\) and each type carries its own actor\-critic\. IPPO\-rs \(reward\-shaped IPPO with the voltage penalty doubled\) and IPPO\-Lagrangian \(separate cost critic per channel, one shared team\-level dual variable\) are the safety variants\.
- •GenCos \(competitive MARL\)\.IPPO with full parameter sharing: a single ActorCritic is shared across the five GenCos, where each agent is conditioned on its private 12\-dim observation and trained on its own profit reward\. Truthful, uniform\-mid, and max\-markup are reported as non\-learning strategic baselines\.

PPO\-Lagrangian on both TSO and DSO uses a zero cost\-budget threshold \(bi=0b\_\{i\}=0\)\. The Sauté PPO observation\-augmentation budget, per\-algorithm dual learning rates, and other safe\-RL specifics are listed with the master table in Section[F\.2](https://arxiv.org/html/2609.36052#A6.SS2)\.

#### TSO reward scaling\.

TSO trains with a reward\-scaling factorαr=10−4\\alpha\_\{r\}=10^\{\-4\}; the per\-step reward seen by the optimizer is

rt=−αr​\(Ctgen\+Ctstartup\+Ctnoload\)\.r\_\{t\}\\;=\\;\-\\alpha\_\{r\}\\,\\bigl\(C^\{\\mathrm\{gen\}\}\_\{t\}\+C^\{\\mathrm\{startup\}\}\_\{t\}\+C^\{\\mathrm\{noload\}\}\_\{t\}\\bigr\)\.\(42\)The operating\-cost numbers reported in main\-text Section[6\.3](https://arxiv.org/html/2609.36052#S6.SS3)and in Section[G\.2](https://arxiv.org/html/2609.36052#A7.SS2)are the unscaled per\-day operating cost\.

### F\.2Per\-task per\-algorithm hyperparameters

Table[8](https://arxiv.org/html/2609.36052#A6.T8)summarises the key training hyperparameters\.

Table 8:Master hyperparameter table\. Columns:TtotT\_\{\\text\{tot\}\}total environment steps \(millions\);NenvsN\_\{\\text\{envs\}\}parallel envs;nsn\_\{s\}steps\-per\-update;EEPPO epochs \(SAC update epochs\); LR learning rate; HD hidden dims; CL clip\-eps; EC entropy coef;γ\\gammadiscount;λ\\lambdaGAE lambda\. Optimizer is Adam throughout\.Safe\-RL specifics not in the table: the dual learning rate is5×10−35\\times 10^\{\-3\}for TSO PPO\-Lagrangian,2×10−52\\times 10^\{\-5\}for DSO PPO\-Lagrangian, and5×10−55\\times 10^\{\-5\}for DERs IPPO\-Lagrangian \(three independent cost channels, one shared team\-level dual variable\)\. The Sauté PPO observation\-augmentation budget on DSO is192192violation\-step\-equivalents\. IPPO\-rs on DERs sets the voltage penalty to8\.08\.0, twice the4\.04\.0of IPPO\.

### F\.3Network architectures

- •MLP backbone\.Actor and critic are independent MLPs with hidden sizes from Table[8](https://arxiv.org/html/2609.36052#A6.T8)\. IPPO \(GenCos, DERs\) and the CMDP wrapper used for PPO\-Lagrangian and Sauté PPO usetanhactivations; the rejax\-based PPO and SAC baselines useswishandrelu, respectively\.
- •Continuous policy heads\.Plain PPO via rejax \(TSO PPO, DSO PPO\) emits an unbounded diagonal Gaussian and clips the sampled action to the action box; the CMDP wrapper applies atanhsquash to the action mean before adding diagonal Gaussian noise; rejax SAC \(DSO SAC, DCMG SAC\) uses atanh\-squashed Gaussian whose support equals the action box; IPPO \(GenCos, DERs\) emits an unbounded diagonal Gaussian whose samples are clipped to\[−1,1\]\[\-1,1\]before the environment step\. DCMG PPO uses a bounded Beta policy over the\[0,1\]3×\[−1,1\]×\[0,1\]\[0,1\]^\{3\}\\times\[\-1,1\]\\times\[0,1\]action box \(Section[C\.5](https://arxiv.org/html/2609.36052#A3.SS5)\)\.
- •Mixed action space \(TSO\)\.The policy emits a single 108\-dim continuous head; the environment decomposes the output into 54 sign\-thresholded commitment intents and 54 dispatch preferences that bias the OPF cost intercept \(Section[C\.2](https://arxiv.org/html/2609.36052#A3.SS2)\)\.
- •MARL parameter sharing\.DERs partitions the 12 agents by type \(4 batteries, 4 PV inverters, 4 flexible loads\) and uses one shared actor\-critic per type; GenCos shares one actor\-critic across the five generation companies\.
- •Observation normalization\.The single\-agent baselines train with running\-mean / running\-std observation normalization; the MARL baselines \(GenCos and DERs IPPO variants\) train on raw observations\.

### F\.4Seeds and evaluation protocol

Each task uses 5 seeds \(0,1,2,3,4\) and an evaluation budget of 10–50 episodes per seed \(Table[6](https://arxiv.org/html/2609.36052#A4.T6)\)\. Episode\-level metrics — operating cost, network loss, profit, and episode return — are reported as across\-seed mean±\\pm95% bootstrap CI; per\-task comparisons on the primary metric use a paired sign\-flip permutation test on matched \(seed, episode\-index\) pairs\. Per\-step rate metrics \(thermal overload, reserve shortfall, voltage violation, SLA deficit\) are reported as across\-seed mean rates; the TSO reserve and thermal channels are reported under the zero\-violation constraint of Section[A\.2](https://arxiv.org/html/2609.36052#A1.SS2)\. Wall\-clock time is recorded at every evaluation point so that learning curves can be aligned across backends\.

## Appendix GAdditional Empirical Results

This section reports the per\-task supplementary tables and figures, one subsection per task \(Sections[G\.1](https://arxiv.org/html/2609.36052#A7.SS1)–[G\.5](https://arxiv.org/html/2609.36052#A7.SS5)\)\. Task definitions are in Section[C](https://arxiv.org/html/2609.36052#A3); split meanings are in Table[7](https://arxiv.org/html/2609.36052#A4.T7); per\-task seed and evaluation\-episode budgets are in Table[6](https://arxiv.org/html/2609.36052#A4.T6); the across\-seed mean±\\pm95% bootstrap CI and the paired sign\-flip permutation test are defined in Section[F\.4](https://arxiv.org/html/2609.36052#A6.SS4)\.

### G\.1GenCos

The task definition is given in Section[C\.1](https://arxiv.org/html/2609.36052#A3.SS1)\. We report the in\-distribution baseline comparison \(Table[9](https://arxiv.org/html/2609.36052#A7.T9)\), the cross\-split profit table \(Table[10](https://arxiv.org/html/2609.36052#A7.T10)\), and a representative\-episode behavior figure \(Figure[11](https://arxiv.org/html/2609.36052#A7.F11)\)\. Each method aggregates 5 seeds and 30 evaluation episodes per seed with 95% bootstrap CIs \(Section[F\.4](https://arxiv.org/html/2609.36052#A6.SS4)\)\.

The in\-distribution mean total profit places IPPO \(GBP​395,030\\mathrm\{GBP\}\\,395\{,\}030, 95% CI\[291,174,499,652\]\[291\{,\}174,\\ 499\{,\}652\]\) strictly between Uniform\-mid \(302,972302\{,\}972,\[297,376,309,015\]\[297\{,\}376,\\ 309\{,\}015\]\) and Max\-markup \(597,864597\{,\}864,\[587,181,609,873\]\[587\{,\}181,\\ 609\{,\}873\]\), and well above Truthful \(6,9346\{,\}934,\[6,412,7,432\]\[6\{,\}412,\\ 7\{,\}432\]\)\. Under the paired sign\-flip permutation test of Section[F\.4](https://arxiv.org/html/2609.36052#A6.SS4)on5​seeds×30​episodes=1505\\,\\text\{seeds\}\\times 30\\,\\text\{episodes\}=150matched \(seed, episode\-index\) pairs, the IPPO advantage over Truthful \(\+388,096\+388\{,\}096\) and Uniform\-mid \(\+92,058\+92\{,\}058\) and the IPPO shortfall against Max\-markup \(−202,834\-202\{,\}834\) are all significant atp<10−4p<10^\{\-4\}\. The cross\-split table preserves the sameTruthful<Uniform\-mid<IPPO<Max\-markup\\textrm\{Truthful\}<\\textrm\{Uniform\-mid\}<\\textrm\{IPPO\}<\\textrm\{Max\-markup\}ordering on the demand\-shift and renewable\-shock splits\.

Beyond the ranking, the in\-distribution results show two behavioral differences from the heuristics\. First, the three non\-learning baselines all bid the same shape and produce nearly identical Herfindahl–Hirschman Index values \(0\.56640\.5664,0\.56630\.5663,0\.56620\.5662for Truthful, Uniform\-mid, and Max\-markup\), while IPPO learns a less concentrated outcome \(HHI=0\.5384\\mathrm\{HHI\}=0\.5384\)\. Second, the ramp\-binding rate \(the fraction of clearing steps in which the SCED ramp constraint is active\) drops from0\.580\.58for the heuristics to0\.350\.35under IPPO\. Figure[11](https://arxiv.org/html/2609.36052#A7.F11)resolves these two signals over a single representative in\-distribution episode: the top\-left panel shows that IPPO clears at lower locational marginal prices than Max\-markup at high system load, the top\-right panel tracks the per\-step cumulative profit gap, the bottom\-left panel plots the HHI over the episode, and the bottom\-right panel plots the ramp\-binding indicator\.

Table 9:GenCos in\-distribution baseline comparison oncase5, 5 seeds, 30 evaluation episodes per seed\. Higher total profit is better\. Profit is the across\-seed mean with a 95% bootstrap CI \(Section[F\.4](https://arxiv.org/html/2609.36052#A6.SS4)\)\. HHI denotes the Herfindahl–Hirschman Index of cleared\-energy shares\.Table 10:GenCos total\-profit mean \(GBP\) across splits\. Each entry aggregates 5 seeds and 30 evaluation episodes per seed\.Figure 11:GenCos behavior on a representative in\-distribution episode under the same physical load\. Top\-left: locational marginal price as a function of system load\. Top\-right: cumulative profit per step\. Bottom\-left: Herfindahl–Hirschman Index of cleared\-energy shares\. Bottom\-right: per\-step ramp\-binding indicator\.
### G\.2TSO

Table[11](https://arxiv.org/html/2609.36052#A7.T11)reports operating cost and the two per\-step violation rates under the zero\-violation constraint of Section[A\.2](https://arxiv.org/html/2609.36052#A1.SS2)\. No method satisfies it on both channels: PPO\-Lagrangian and All\-On record zero reserve shortfall, while every method records a non\-zero thermal\-overload rate\. The paired sign\-flip permutation test on per\-episode operating cost \(Section[F\.4](https://arxiv.org/html/2609.36052#A6.SS4),n=250n=250pairs,p<10−4p<10^\{\-4\}for the reported comparisons\) confirms that the cost differences are significant\. The task card and 24\-hour SCUC formulation are described in Section[C\.2](https://arxiv.org/html/2609.36052#A3.SS2); Figure[12](https://arxiv.org/html/2609.36052#A7.F12)decomposes the cost\-versus\-violation pattern into worst\-rate and per\-channel views, separating the reserve channel \(PPO\-Lagrangian and All\-On at zero\) from the thermal channel \(every method above zero\)\.

Table 11:TSO cost and safety results\. Operating cost is reported per 24\-hour episode in £M; reserve and thermal columns are per\-step violation rates\.Figure 12:TSO cost–safety summary\. Left: total operating cost vs the worst safety\-violation rate \(the maximum of the reserve\-shortfall and thermal\-overload rates\)\. Right: the same methods separated by violation channel, showing that the reserve channel is satisfied for PPO\-Lagrangian and All\-On while the thermal channel is not satisfied for any method\.Figure[13](https://arxiv.org/html/2609.36052#A7.F13)expands the method\-level numbers in Table[11](https://arxiv.org/html/2609.36052#A7.T11)into per\-episode dispersion\. PPO has the lowest mean operating cost \(£1\.58M per 24\-hour episode\) but a wide spread along the thermal\-cost axis: its mean per\-episode total thermal cost is247\.67247\.67, and the worst\-decile episodes carry thermal cost more than an order of magnitude above this mean\. PPO\-Lagrangian shifts the cloud toward higher cost \(£3\.50M per episode\) while compressing mean thermal cost to1\.761\.76, roughly two orders of magnitude below PPO\. Merit Order and All\-On occupy the high\-cost low\-thermal corner\.

Figure 13:TSO per\-episode operating cost vs thermal\-overload cost on the in\-distribution split \(4 algorithms×\\times5 seeds×\\times50 episodes==1,000 points\)\. Marker area is proportional to the per\-episode thermal\-violation rate; the ellipse around each method covers its 95% covariance region\.Figure[14](https://arxiv.org/html/2609.36052#A7.F14)complements Table[11](https://arxiv.org/html/2609.36052#A7.T11)with train\-split checkpoint monitoring curves over the 20 M\-step training budget\. The curves aggregate the five seeds and show smoothed mean±\\pmone standard deviation over sparse greedy\-evaluation checkpoints\. PPO drives operating cost down toward its held\-out in\-distribution value of £1\.58M per episode \(Table[11](https://arxiv.org/html/2609.36052#A7.T11)\), but its safety panels move in the opposite direction, with both reserve\-shortfall and thermal\-overload incidence rising over training\. PPO\-Lagrangian holds reserve\-shortfall incidence at zero throughout and lowers final thermal\-overload incidence well below PPO, while its operating cost stays above both PPO and the merit\-order reference\. The learning dynamics therefore explain how the cost–safety frontier is reached\.

Figure 14:TSO train\-split checkpoint monitor over 20 M environment steps\. Panels show evaluation operating cost, reserve\-shortfall incidence, and thermal\-overload incidence\. Curves are smoothed means over five seeds; shaded bands show one standard deviation\.
### G\.3DSO

DSO reports raw in\-distribution physical metrics over 5 seeds×\\times50 evaluation episodes; the voltage channel is reported under Form 1 withbi=0b\_\{i\}\{=\}0\(Section[A\.2](https://arxiv.org/html/2609.36052#A1.SS2)\)\. Loss reduction in Table[12](https://arxiv.org/html/2609.36052#A7.T12)is computed relative to no control\.

Table 12:DSO in\-distribution physical metrics\. Lower total network loss is better; the voltage column is the per\-step count of buses outside the\[0\.94,1\.06\]\[0\.94,1\.06\]p\.u\. band\.Figure[15](https://arxiv.org/html/2609.36052#A7.F15)visualizes the three columns of Table[12](https://arxiv.org/html/2609.36052#A7.T12)as bar panels\.

Figure 15:DSO in\-distribution physical metrics\. Panels report total network loss, voltage\-violation count per step, and loss reduction relative to no control\.Figure 16:DSO learning curves over 3 M environment steps\. Left: episode total reward across the four learned algorithms\. Middle: per\-step voltage\-violation count\. Right: PPO\-Lagrangian Lagrange multiplier𝝀\\boldsymbol\{\\lambda\}\.Figure[17](https://arxiv.org/html/2609.36052#A7.F17)expands Table[12](https://arxiv.org/html/2609.36052#A7.T12)into per\-episode behavior distributions over the six methods, with each box overlaid by the 250 individual episodes so distributional tails remain visible\. PPO and Sauté PPO use roughly1717MWh per day of curtailment and1111MWh per day of shifting, achieving close to30%30\\%peak shaving with voltage violations driven to near zero on most episodes; tail steps are still visible above zero in the bottom\-right panel\. PPO\-Lagrangian curtails and shifts substantially less, yet has a heavier voltage\-violation tail than PPO\. SAC reaches lower total loss than PPO\-Lagrangian on every paired episode but carries the highest per\-step voltage\-violation rate of the four learned methods\. Droop’s curtailment and shifting stay well below11MWh per day, an order of magnitude under the four learned policies\.

Figure 17:DSO per\-episode in\-distribution behavior distributions over 6 algorithms×\\times5 seeds×\\times50 episodes\. Each panel reports a single behavior metric; boxes are overlaid by jittered individual episodes so distributional tails \(especially the voltage\-violation panel\) remain visible\.#### PPO\-Lagrangian training dynamics\.

Under no control the in\-distribution DSO carries roughly0\.210\.21voltage violations per step \(Table[12](https://arxiv.org/html/2609.36052#A7.T12)\)\. PPO and Sauté PPO drive the in\-distribution rate to near zero within the first half of the training budget; once the per\-step rate is small, the constraint signal that the dual update consumes becomes sparse, and PPO\-Lagrangian converges to a conservative region with low loss reduction and a heavier violation tail than PPO\. The corresponding learning curves are shown in Figure[16](https://arxiv.org/html/2609.36052#A7.F16)\.

### G\.4DERs

On the in\-distribution split, every learned and non\-learning method reports zero or near\-zero voltage violations \(Table[13](https://arxiv.org/html/2609.36052#A7.T13)\)\. Across all five algorithms and all five evaluation splits, the maximum bus voltage stays at the slack referenceVmax=1\.000V\_\{\\max\}=1\.000p\.u\., so the safe band is exercised only on its under\-voltage side and the operating regime never enters PV\-driven over\-voltage; this remains true under the PV\-shift split, where the PV bundle capacity is doubled\. The voltage\-tightening split \(Table[14](https://arxiv.org/html/2609.36052#A7.T14)\), which narrows the band from\[0\.94,1\.06\]\[0\.94,1\.06\]p\.u\. to\[0\.96,1\.04\]\[0\.96,1\.04\]p\.u\., is therefore the split on which the methods separate on safety\.

For the in\-distribution loss test, IPPO and IPPO\-rs reduce active\-power loss to0\.19740\.1974MW versus0\.20520\.2052MW for no control — a3\.8%3\.8\\%reduction — and beat both no control and voltage droop atp<10−4p<10^\{\-4\}under the paired sign\-flip test of Section[F\.4](https://arxiv.org/html/2609.36052#A6.SS4)\. IPPO\-Lagrangian has higher loss \(0\.20510\.2051MW\): it is statistically indistinguishable from no control on the loss test \(p=0\.87p\{=\}0\.87\) and is statistically worse than voltage droop \(mean difference\+0\.0020\+0\.0020MW,p<10−4p<10^\{\-4\}\)\.

Under tightening, IPPO and IPPO\-rs cut the violation count by44%44\\%relative to no control \(4\.914\.91vs8\.738\.73steps per episode\); voltage droop cuts it by32%32\\%\(5\.975\.97\); IPPO\-Lagrangian cuts it by9%9\\%\(7\.947\.94\), so the constrained variant trails IPPO and IPPO\-rs on both the in\-distribution loss test and the tightening split\. The load\-stress split shows the same ordering qualitatively: per\-episode violation steps are0\.230\.23for IPPO and IPPO\-rs,0\.330\.33for voltage droop,0\.830\.83for IPPO\-Lagrangian, and0\.930\.93for no control\. Figure[18](https://arxiv.org/html/2609.36052#A7.F18)traces battery state of charge, battery reactive\-power command, photovoltaic curtailment, photovoltaic reactive\-power command, flexible\-load action, and active\-power loss for each policy on the same in\-distribution episode, showing that IPPO and IPPO\-rs realize their loss reduction through coordinated reactive\-power dispatch rather than active\-power curtailment, while IPPO\-Lagrangian holds the cost critic near its budget and dispatches less aggressively on both channels\.

Table 13:DERs in\-distribution results\. Lower active\-power loss is better\. Loss values are mean across55seeds×\\times3030episodes\.Table 14:DERs voltage\-tightening stress split \(band tightened from\[0\.94,1\.06\]\[0\.94,1\.06\]to\[0\.96,1\.04\]\[0\.96,1\.04\]p\.u\.\)\. Violation steps are per episode; rates use the4848\-step horizon as denominator\. Active loss matches the in\-distribution column because tightening changes only the constraint band, not dispatch decisions\.Figure 18:DERs schedules on a single in\-distribution episode\. Panels compare battery state of charge, battery reactive power commands, PV curtailment, PV reactive\-power commands, flexible load actions, and active power loss across IPPO, IPPO\-rs, IPPO\-Lagrangian, voltage droop, and no control\.
### G\.5DCMG

DCMG \(Section[C\.5](https://arxiv.org/html/2609.36052#A3.SS5)\) is reported on the in\-distribution split with all three cost channels \(SLA, over\-temperature, power deficit\) at a zero threshold\. Episode returns are in Table[15](https://arxiv.org/html/2609.36052#A7.T15)\. Figure[19](https://arxiv.org/html/2609.36052#A7.F19)shows a representative SAC dispatch on one in\-distribution episode; the same SAC trajectory underlies main\-text Section[6\.4](https://arxiv.org/html/2609.36052#S6.SS4)\.

PPO records zero violations on all three channels; SAC and the three heuristic methods show SLA rates of10−710^\{\-7\}to10−610^\{\-6\}and zero on the other two channels\.

Table 15:DCMG in\-distribution results\. Higher episode return is better\. SLA rate is the per\-step rate of expired tasks; spill cost is the per\-step over\-generation penalty\. The power\-deficit and over\-temperature rates are00for every method and are not shown\.Figure 19:Representative DCMG dispatch from the SAC policy on one in\-distribution episode, displayed as 30\-minute means for readability\.Left:real GB solar capacity factor and the data center IT load\.Middle:net load after PV, grid import, diesel output, and battery charge/discharge\.Right:battery state of charge with SoC bounds, alongside the GB MID grid price\.

## Appendix HExecution Scaling and Backend Comparison

This section decomposes the suite\-level speed result of Section[6\.1](https://arxiv.org/html/2609.36052#S6.SS1)along three axes: backend architecture, parallelism range, and wall\-clock learning efficiency\.

### H\.1Backend decomposition atNenv=212N\_\{\\rm env\}=2^\{12\}

The three training backends differ in where the environment transition runs\. SB3 runs a PyTorch policy on the GPU and the PowerZooPy environment on the CPU\. SBX runs a JAX policy and optimizer on the GPU but keeps the same CPU\-hosted PowerZooPy environment\. PowerZooJax runs the policy, optimizer, and environment transition end\-to\-end inside a compiled JAX program on the GPU\. Table[16](https://arxiv.org/html/2609.36052#A8.T16)decomposes throughput atNenv=212N\_\{\\rm env\}=2^\{12\}across these three stacks: the SB3\-to\-SBX delta isolates the policy\-side acceleration \(PyTorch→\\toJAX\), and the SBX\-to\-PowerZooJax delta, reported in the rightmost column “×\\timesvs SBX”, isolates the environment\-side acceleration\. The environment\-side gain alone is1,145×1\{,\}145\\timeson GenCos and2,040×2\{,\}040\\timeson DERs, both of which call CPU\-hosted OPF or power\-flow solvers in PowerZooPy that PowerZooJax fuses into the compiled JAX graph\. The gain is smaller on DSO \(39×39\\times\) and TSO \(33×33\\times\), whose CPU steps already run JAX\-friendly kernels rather than an external solver\. DCMG \(22×22\\times\) is the smallest case: its step is dominated by small arithmetic device updates without an external solver call, so the CPU\-hosted environment is already fast\.

Table 16:Throughput atNenv=212N\_\{\\rm env\}=2^\{12\}parallel environments, with the environment\-side speedup isolated\. The headline speedup against the slower CPU baseline is in main\-text Table[2](https://arxiv.org/html/2609.36052#S6.T2)\.
### H\.2Per\-task throughput scaling

Figure[20](https://arxiv.org/html/2609.36052#A8.F20)reports throughput against parallelism for DCMG, DSO, and DERs\. PowerZooJax JAX/GPU is swept overnenv∈\{16,32,64,128,256\}\\text\{nenv\}\\in\\\{16,32,64,128,256\\\}on all three; SB3/CUDA is overlaid for DCMG and DSO under the same warm\-up and measurement protocol over the range where itsSubprocVecEnvbackend completes\. JAX/GPU throughput rises with parallelism on every task: DCMG from32\.132\.1k to307\.7307\.7k env steps/s, DSO from25\.025\.0k to388\.4388\.4k, and DERs from21\.621\.6k to283\.8283\.8k\. The matched\-range speedup over SB3/CUDA reaches7\.4×7\.4\\timesfor DCMG \(atnenv=64\\text\{nenv\}=64\) and39×39\\timesfor DSO \(atnenv=128\\text\{nenv\}=128\), the largest parallelism each SB3 path completes; SB3 throughput tops out at14\.114\.1k on DCMG and5\.35\.3k on DSO\. The high\-parallel endpoint in Table[16](https://arxiv.org/html/2609.36052#A8.T16)reports25\.125\.1k,134\.7134\.7k, and414\.5414\.5k atNenv=212N\_\{\\rm env\}=2^\{12\}for DCMG, DSO, and DERs respectively under a separate measurement protocol\.

Figure 20:Per\-task throughput scaling\. PowerZooJax JAX/GPU is plotted overnenv∈\{16,32,64,128,256\}\\text\{nenv\}\\in\\\{16,32,64,128,256\\\}for every task; SB3/CUDA is overlaid where itsSubprocVecEnvbackend completes \(throughnenv=64\\text\{nenv\}=64for DCMG, throughnenv=128\\text\{nenv\}=128for DSO; DERs has no paired SB3 sweep\)\. Endpoint labels mark the JAX/GPU and SB3/CUDA throughput at the right end of each curve; arrows mark the JAX\-vs\-SB3 ratio at the top of each matched range\.
### H\.3First\-compile cost

The scaling sweeps record the first JAX compilation separately from the steady\-state update loop\. Compile cost is single\-digit seconds for the three tasks in Figure[20](https://arxiv.org/html/2609.36052#A8.F20)\(Table[17](https://arxiv.org/html/2609.36052#A8.T17)\); subsequent updates run from the cached compiled program\.

Table 17:First JAX compilation time in task\-specific scaling sweeps\. Values are means over seeds\.
### H\.4Wall\-clock learning curves for DSO and TSO

For DSO and TSO, all three backends \(PowerZooJax, SBX/CUDA, SB3/CUDA\) ran to the same training budget and recorded per\-evaluation\-point wall\-clock timestamps, which allows learning curves to be plotted against time rather than environment steps\. Figure[21](https://arxiv.org/html/2609.36052#A8.F21)traces evaluation return against elapsed time for the three backends on these two tasks at the matched budgets of main\-text Table[2](https://arxiv.org/html/2609.36052#S6.T2)\. The PowerZooJax run reaches the same training budget in a fraction of the time required by either CPU\-hosted backend\.

Figure 21:Evaluation return against elapsed wall\-clock time for DSO and TSO across PowerZooJax, SBX/CUDA, and SB3/CUDA\.

相似文章

JAXBench:对自主TPU内核优化进行基准测试

arXiv cs.AI

JAXBench是一个新的基准测试套件,包含50个JAX工作负载,用于评估在谷歌云TPU上的AI生成内核优化,提供人工调优基线和智能体评估框架。论文发现,基于精选TPU文档进行条件化处理能显著提升正确性和加速比,其中Autocomp波束搜索在人工调优内核上相比XLA实现了高达1.6倍的几何平均加速。

PowerAtlas:面向电力系统的电算协同调度

arXiv cs.LG

PowerAtlas是一个LLM代理框架,用于在数据中心中联合调度电力和计算,确保电网可行性和任务SLA。通过与真实电力公司和包含2000个实例的新基准(ECBench)进行验证,它在多个开放权重LLM上显示出一致的性能提升。