Adaptive hybrid coupling with operator inference, the overlapping Schwarz alternating method and reinforcement learning

arXiv cs.LG Papers

Summary

The paper introduces a reinforcement learning-based approach for adaptively selecting full and reduced order models in hybrid domain decomposition simulations, using deep Q-networks to balance accuracy and cost in transient problems.

arXiv:2609.17837v1 Announce Type: new Abstract: Hybrid domain decomposition methods provide a flexible framework for coupling full order models (FOMs) and reduced order models (ROMs), but typically assume the model assigned to each subdomain is fixed throughout a simulation. This is limiting for transient problems in which localized features propagate through the domain and the regions requiring high-fidelity resolution change over time. We introduce a reinforcement learning (RL)-based approach for online adaptation of FOM-ROM models coupled via the overlapping Schwarz alternating method (O-SAM), an iterative domain decomposition method that solves subdomain-local problems while exchanging solution information through transmission boundary conditions on overlapping interfaces. Deep Q-networks (DQNs) are trained offline to select among subdomain-local FOMs and pre-trained Operator Inference (OpInf) ROMs using a reward balancing accuracy, cost, and model-switching frequency. Once trained, the policies are deployed predictively on problem instances not seen during training, without requiring a reference FOM solution. We demonstrate the approach on two examples: a 1D advection-diffusion problem with a moving front, and a 3D linear elastic wave propagation problem implemented in the Norma.jl solid mechanics code. For the advection-diffusion benchmark, the learned policy dynamically allocates high-fidelity resolution as the front propagates and outperforms static FOM/ROM assignments; letting the agent also adapt the domain decomposition provides no further benefit. For the elastic wave benchmark, learned policies for two and three subdomain decompositions track the propagating wave by assigning FOMs to subdomains containing the wave and ROMs elsewhere, as expected. Our results demonstrate the potential of RL to enable predictive online adaptation of model fidelity within Schwarz-based hybrid simulations.
Original Article
View Cached Full Text

Cached at: 09/17/26, 08:56 AM

# Adaptive hybrid coupling with operator inference, the overlapping Schwarz alternating method and reinforcement learning
Source: [https://arxiv.org/html/2609.17837](https://arxiv.org/html/2609.17837)
Irina Tezaur††thanks:Sandia National Laboratories, ikalash@sandia\.govAnthony Gruber††thanks:Sandia National Laboratories, adgrube@sandia\.gov

###### Abstract

Hybrid domain decomposition methods provide a flexible framework for coupling full order models \(FOMs\) and reduced order models \(ROMs\), but typically assume that the model assigned to each subdomain is selecteda prioriand remains fixed throughout a simulation\. This can be limiting for transient problems in which localized features propagate through the computational domain and the regions requiring high\-fidelity resolution change over time\. We introduce a reinforcement learning\- \(RL\)\-based approach for online adaptation of FOM\-ROM models coupled using the overlapping Schwarz alternating method \(O\-SAM\), an iterative domain decomposition method that solves subdomain\-local problems while exchanging solution information through transmission boundary conditions on the overlapping subdomain interfaces\. Deep Q\-networks \(DQNs\) are trained offline to select among subdomain\-local FOMs and pre\-trained Operator Inference \(OpInf\) ROMs using a reward that balances solution accuracy and computational cost, while simultaneously penalizing unnecessary model switching\. Once trained, the resulting policies are deployed predictively on problem instances not encountered during RL training without requiring a reference FOM solution\. We demonstrate the proposed approach on two numerical examples: a one\-dimensional advection\-diffusion problem with a moving front and a three\-dimensional linear elastic wave propagation problem implemented in theNorma\.jlsolid mechanics code\. For the advection\-diffusion benchmark, the learned policy dynamically allocates high\-fidelity resolution as the front propagates and provides a favorable accuracy\-cost tradeoff relative to static FOM/ROM assignments; allowing the agent to additionally adapt the domain decomposition provides no further benefit\. For the elastic wave benchmark, the learned policies for both two and three subdomain decompositions track the propagating wave by assigning FOMs to the subdomains containing the wave and ROMs elsewhere, as expected\. Importantly, this behavior is observed for problem instances not used during training\. Our results demonstrate the potential of RL to enable predictive online adaptation of model fidelity within Schwarz\-based hybrid simulations\.

## 1Introduction

Multi\-scale, multi\-physics modeling and simulation \(ModSim\) is essential for the design and qualification of engineered components and systems relevant to Sandia’s mission\. Analysts face significant delays due to both the mesh generation step of the ModSim workflow and the long run\-time requirements for simulations, which can preclude multi\-query analyses such as design optimization, control and uncertainty quantification \(UQ\)\. Emerging approaches for building data\-driven reduced order models \(ROMs\) have promised to reduce the runtime burden of ModSim and enable multi\-query analyses\. However, ROMs can suffer from their own deficiencies, including robustness, stability, and accuracy concerns, lengthy implementations, and a lack of systematic refinement mechanisms\.

In recent years, Sandia has expended a large effort to overcome both of the two aforementioned hurdles through the development of hybrid models based on optimal domain decomposition \(DD\)\[[26](https://arxiv.org/html/2609.17837#bib.bib8),[23](https://arxiv.org/html/2609.17837#bib.bib9),[19](https://arxiv.org/html/2609.17837#bib.bib10),[25](https://arxiv.org/html/2609.17837#bib.bib3)\], as illustrated in Figure[1](https://arxiv.org/html/2609.17837#S1.F1)\(a\)\. In our approach, the domain is decomposed into overlapping or non\-overlapping subdomains according to some physics\-based criteria, and different meshes and/or models are assigned to different subdomains according to these physics\. The subdomains are coupled via an iterative technique known as the Schwarz alternating method \(SAM\)\[[21](https://arxiv.org/html/2609.17837#bib.bib2)\]\. The key idea behind SAM is to replace a global monolithic problem in a domainΩ\\Omegaby a sequence of smaller problems in subdomainsΩi\\Omega\_\{i\}, whereΩ=∪iΩi\\Omega=\\cup\_\{i\}\\Omega\_\{i\}\. The coupling happens through carefully defined transmission boundary conditions prescribed on the interfaces between the subdomainsΩi\\Omega\_\{i\}\. In our past work\[[14](https://arxiv.org/html/2609.17837#bib.bib5),[15](https://arxiv.org/html/2609.17837#bib.bib7),[23](https://arxiv.org/html/2609.17837#bib.bib9),[26](https://arxiv.org/html/2609.17837#bib.bib8),[19](https://arxiv.org/html/2609.17837#bib.bib10),[25](https://arxiv.org/html/2609.17837#bib.bib3)\], we have shown that SAM is capable of seamlessly “gluing together” arbitrary combinations of subdomain\-local full order models \(FOMs\) and subdomain\-local ROMs in a plug\-and\-play fashion\. Since the method does not require conformal meshing of the subdomainsΩi\\Omega\_\{i\}, it is extremely effective at simplifying meshing workflows\. The introduction of ROMs into the workflow has the potential to improve the predictive viability of the reduced models by enabling their spatial localization via DD, as well as through the online integration of high\-fidelity information via FOM coupling\.

![Refer to caption](https://arxiv.org/html/2609.17837v1/figures/Schwarz_DD_FOM.png)\(a\)DD and model assignment
![Refer to caption](https://arxiv.org/html/2609.17837v1/figures/rom_switching.png)\(b\)Online ROM\-FOM\-ROM switching

Fig\. 1:Domain decomposition\-based hybrid coupling and online ROM\-FOM\-ROM switching within a SAM coupling framework\.Until now, our SAM\-based coupling workflow has assumed that the DD and the assignment of meshes and/or models to individual subdomains are performed once, prior to the start of a simulation, and remain fixed throughout the coupled simulation\. The primary limitation of this approach is that it cannot account for changing dynamics or traveling features that emerge and propagate during a simulation\. For example, consider a traveling shock that requires a FOM to be accurately resolved\. If the shock enters subdomains to which ROMs were assigneda priori, inaccuracies may be introduced and subsequently propagate through the coupled system, degrading the solution over time\.

The idea of dynamically switching between ROMs and FOMs online, as illustrated in Figure[1](https://arxiv.org/html/2609.17837#S1.F1)\(b\), is relatively new, but has been explored in several previous works\. In\[[18](https://arxiv.org/html/2609.17837#bib.bib12),[4](https://arxiv.org/html/2609.17837#bib.bib11)\], which consider solid mechanics problems, ROM\-FOM switching is determined by the onset of plasticity within a prescribed subdomain: a ROM is employed as long as the subdomain remains in the elastic regime, and the subdomain switches to a FOM once plasticity is detected\. Because plasticity is irreversible, switching is one\-way, from ROM to FOM\.

Several more recent works pursue closely related ideas, including adaptive enrichment, fidelity selection, and heterogeneous ROM\-FOM coupling, without performing dynamic online ROM\-FOM switching in the strict sense\. In\[[22](https://arxiv.org/html/2609.17837#bib.bib13)\], Smetana and Taddei propose an adaptive localized ROM in which residual\-based error indicators identify components requiring enrichment and high\-fidelity local correction problems are solved to augment the corresponding reduced spaces\. Thus, rather than switching a component from a ROM to a FOM, their approach selectively introduces high\-fidelity computations online when the ROM is deemed insufficient\. In\[[8](https://arxiv.org/html/2609.17837#bib.bib14)\], Huang and coauthors develop adaptive and component\-based ROM frameworks that incorporate online basis adaptation and permit heterogeneous coupling of ROM and FOM components; however, components are not dynamically switched between ROM and FOM representations during the simulation\. In\[[5](https://arxiv.org/html/2609.17837#bib.bib15)\], Ebrahimi and Yano develop a component\-based hyper\-reduced ROM that adaptively selects the hyper\-reduction fidelity of individual components online based on a system\-level error estimate\. This constitutes online adaptive fidelity selection rather than ROM\-FOM switching, as the components remain reduced order throughout the computation\. Finally, Hedayat et al\.\[[6](https://arxiv.org/html/2609.17837#bib.bib16)\]perform online temporal switching between ROM and FOM computations, periodically invoking the FOM to acquire high\-fidelity data used to adapt the ROM before returning to reduced order prediction\. In contrast to the present setting, their approach considers a single computational domain and therefore does not involve domain decomposition or coupling between distinct ROM and FOM subdomains\.

One fundamental challenge in enabling online ROM\-FOM switching is determining when and where a change in model fidelity should occur\. Fundamentally, this problem is related to extrapolation detection in machine learning and requires some measure of model adequacy\. While there has been some work on error estimation for certain types of ROMs, e\.g\., projection\-based ROMs\[[3](https://arxiv.org/html/2609.17837#bib.bib23)\], reliable error indicators are not readily available for many classes of data\-driven models\. Herein, we address this challenge by leveraging reinforcement learning \(RL\)\[[24](https://arxiv.org/html/2609.17837#bib.bib20)\]to learn an online model\-switching policy for subdomain\-local FOMs and Operator Inference \(OpInf\) ROMs\[[16](https://arxiv.org/html/2609.17837#bib.bib1)\]coupled via overlapping SAM \(O\-SAM\)\. We refer to the resulting RL\-based adaptive O\-SAM framework as “RL–O\-SAM”\. In our workflow, the RL agent is trained on a collection of problem instances to select the model employed in each subdomain over time, with a reward function that balances computational cost, solution error, and penalizes unnecessary model switching\. Once trained, the learned policy is deployed predictively on problem instances not encountered during RL training, including problems characterized by different parameters, boundary conditions, and/or initial conditions from those used to train the policy\. In this way, the proposed framework seeks to learn when and where high\-fidelity resolution is required while exploiting computationally inexpensive ROMs elsewhere, enabling the model assignment within the domain decomposition to adapt dynamically as the solution evolves\.

Toward this effect, the remainder of this paper is organized as follows\. Section[2](https://arxiv.org/html/2609.17837#S2)provides some preliminaries to keep this paper self\-contained, overviewing the OpInf approach to model order reduction and the O\-SAM algorithm applied to couple subdomain local OpInf ROMs with each other or with subdomain\-local FOMs\. Section[3](https://arxiv.org/html/2609.17837#S3)overviews RL and describes how it is utilized to perform online adaptation of O\-SAM\. Section[4](https://arxiv.org/html/2609.17837#S4)presents some numerical results illustrating the accuracy and efficiency of the proposed approach, applied to an advection\-diffusion benchmark with a moving shock, and to a linear elastic wave propagation problem in solid mechanics\. Finally, conclusions are offered in Section[5](https://arxiv.org/html/2609.17837#S5)\.

## 2Hybrid coupling with operator inference \(OpInf\) and the overlapping Schwarz alternating method \(O\-SAM\)

As discussed above, this paper presents an RL\-based approach for performing online switching between subdomain\-local FOMs and OpInf ROMs\[[16](https://arxiv.org/html/2609.17837#bib.bib1)\], coupled via O\-SAM\[[21](https://arxiv.org/html/2609.17837#bib.bib2)\]\. Before describing our RL\-based workflow \(Section[3\.3](https://arxiv.org/html/2609.17837#S3.SS3)\), we briefly describe OpInf model order reduction and its combination with O\-SAM\[[25](https://arxiv.org/html/2609.17837#bib.bib3)\]\.

### 2\.1Operator Inference \(OpInf\)

First proposed by Peherstorfer and Willcox in\[[16](https://arxiv.org/html/2609.17837#bib.bib1)\], OpInf is a projection\-inspired model reduction approach that is fully non\-intrusive, requiring no access to the underlying high\-fidelity simulation code\. The approach consists of three steps: \(i\) the creation of a reduced basis𝚽r\{\\bf\\Phi\}\_\{r\}from snapshot data collected during a high\-fidelity simulation, \(ii\) the solution of a regularized least\-squares optimization problem for the operators defining the ROM, and \(iii\) online deployment of the learned ROM\.

Following the approach in\[[25](https://arxiv.org/html/2609.17837#bib.bib3)\], we construct our reduced basis𝚽r\{\\bf\\Phi\}\_\{r\}using a common technique known as Proper Orthogonal Decomposition \(POD\)\[[7](https://arxiv.org/html/2609.17837#bib.bib4)\]\. Assume, without loss of generality, that we have collectedK\+1K\+1snapshots of a solutionu​\(t\)\\mathord\{\\btensor u\}\(t\)to a given transient partial differential equation \(PDE\) at timesT0,T1,…,TKT\_\{0\},T\_\{1\},\.\.\.,T\_\{K\}and placed them in the columns of a snapshot matrixU\\mathord\{\\btensor U\}, i\.e\.,

U:=\(u​\(T0\),u​\(T1\),⋯,u​\(TK\)\)∈ℝN×K,\\mathord\{\\btensor U\}:=\\left\(\\begin\{array\}\[\]\{cccc\}\\mathord\{\\btensor u\}\(T\_\{0\}\),&\\mathord\{\\btensor u\}\(T\_\{1\}\)&,\\cdots,&\\mathord\{\\btensor u\}\(\{T\_\{K\}\}\)\\end\{array\}\\right\)\\in\\mathbb\{R\}^\{N\\times K\},\(1\)whereNNis the size of each snapshot, corresponding to the number of degrees of freedom \(dofs\) in the PDE being solved\. POD is based on the assumption that the solution to a given PDE has a low rank relative to the FOM dof spaceℝN\\mathbb\{R\}^\{N\}, so that the snapshot matrixU\\mathord\{\\btensor U\}admits a low\-rank approximation\. POD begins by performing a thin singular value decomposition \(SVD\) ofU\\mathord\{\\btensor U\}, so thatU=��T\\mathord\{\\btensor U\}=\\bt@\\Phi\\bt@\\Sigma\{\}^\{T\}\. The firstr≤rank⁡\(U\)r\\leq\\operatorname\{rank\}\(\\mathord\{\\btensor U\}\)left singular vectors form the rank\-rrPOD basis�r∈ℝN×r\\bt@\\Phi\_\{r\}\\in\\mathbb\{R\}^\{N\\times r\}\. This basis provides the optimal rank\-rrapproximation of the snapshots in the Frobenius norm, so that

‖U−�r​�r⊤​U‖F2=∑i=r\+1rank⁡\(U\)σi2,\\\|\\mathord\{\\btensor U\}\-\\bt@\\Phi\_\{r\}\\bt@\\Phi\_\{r\}^\{\\top\}\\mathord\{\\btensor U\}\\\|\_\{F\}^\{2\}=\\sum\_\{i=r\+1\}^\{\\operatorname\{rank\}\(\\mathord\{\\btensor U\}\)\}\\sigma\_\{i\}^\{2\},\(2\)whereσi\\sigma\_\{i\}denotes theiith singular value ofU\\mathord\{\\btensor U\}\. The reduced dimensionrris commonly selected to retain a prescribed fractionξ∈\(0,1\)\\xi\\in\(0,1\)of the snapshot energy, i\.e\., as the smallest integer satisfying

∑i=1rσi2∑i=1rank⁡\(U\)σi2≥ξ\.\\frac\{\\sum\_\{i=1\}^\{r\}\\sigma\_\{i\}^\{2\}\}\{\\sum\_\{i=1\}^\{\\operatorname\{rank\}\(\\mathord\{\\btensor U\}\)\}\\sigma\_\{i\}^\{2\}\}\\geq\\xi\.\(3\)
OpInf is founded on the observation that projection\-based reduction of a polynomial FOM preserves its polynomial structure\. For example, consider a linear FOM of the form

u˙\+A​u=B​g\\dot\{\\mathord\{\\btensor u\}\}\+\\mathord\{\\btensor A\}\\mathord\{\\btensor u\}=\\mathord\{\\btensor B\}\\mathord\{\\btensor g\}\(4\)withu∈ℝN\\mathord\{\\btensor u\}\\in\\mathbb\{R\}^\{N\},A∈ℝN×N\\mathord\{\\btensor A\}\\in\\mathbb\{R\}^\{N\\times N\},B∈ℝN×n\\mathord\{\\btensor B\}\\in\\mathbb\{R\}^\{N\\times n\}andg∈ℝn\\mathord\{\\btensor g\}\\in\\mathbb\{R\}^\{n\}, forN,n∈ℕN,n\\in\\mathbb\{N\}\. A system of the form \([4](https://arxiv.org/html/2609.17837#S2.E4)\) is obtained by discretizing a linear PDE in space using the finite element method and enforcing Dirichlet boundary conditions using theB​g\\mathord\{\\btensor B\}\\mathord\{\\btensor g\}term, where theg\\mathord\{\\btensor g\}vector encodes the Dirichlet boundary data111For a derivation of how theB​g\\mathord\{\\btensor B\}\\mathord\{\\btensor g\}andB¯​g\\bar\{\\mathord\{\\btensor B\}\}\\mathord\{\\btensor g\}terms encode Dirichlet boundary condition in both the ROM and the FOM, the reader is referred to\[[25](https://arxiv.org/html/2609.17837#bib.bib3)\]\.\. Projecting \([4](https://arxiv.org/html/2609.17837#S2.E4)\) onto a reduced basis�r\\bt@\\Phi\_\{r\}yields the linear ROM

u^˙\+A^​u^=B¯​g,\\dot\{\\widehat\{\\mathord\{\\btensor u\}\}\}\+\\widehat\{\\mathord\{\\btensor A\}\}\\widehat\{\\mathord\{\\btensor u\}\}=\\bar\{\\mathord\{\\btensor B\}\}\\mathord\{\\btensor g\},\(5\)whereu^:=�rT​u\\widehat\{\\mathord\{\\btensor u\}\}:=\\bt@\\Phi\_\{r\}^\{T\}\\mathord\{\\btensor u\},A^:=�rT​A​�r\\widehat\{\\mathord\{\\btensor A\}\}:=\\bt@\\Phi\_\{r\}^\{T\}\\mathord\{\\btensor A\}\\bt@\\Phi\_\{r\}andB¯:=�rT​B\\bar\{\\mathord\{\\btensor B\}\}:=\\bt@\\Phi\_\{r\}^\{T\}\\mathord\{\\btensor B\}\. In traditional \(intrusive\) projection\-based model order reduction, computing theA^\\widehat\{\\mathord\{\\btensor A\}\}andB¯\\bar\{\\mathord\{\\btensor B\}\}operators in \([5](https://arxiv.org/html/2609.17837#S2.E5)\) requires access to the matrixA\\mathord\{\\btensor A\}and vectorf\\mathord\{\\btensor f\}, and hence to the code that was used to generate these matrices\. Non\-intrusive OpInf avoids this by learning the reduced operators in \([5](https://arxiv.org/html/2609.17837#S2.E5)\) directly from the snapshot data \([1](https://arxiv.org/html/2609.17837#S2.E1)\) by solving a convex least\-squares minimization problem\. Given data for the projected stateU^=�rT​U\\widehat\{\\mathord\{\\btensor U\}\}=\\bt@\\Phi\_\{r\}^\{T\}\\mathord\{\\btensor U\}and approximate derivative information,Dt​\(U^\)D\_\{t\}\(\\widehat\{\\mathord\{\\btensor U\}\}\)computed with, e\.g\., a finite difference operatorDtD\_\{t\}, the OpInf learning problem takes the form:

arg​minA^,B¯​‖Dt​\(U^\)−A^​U^−B¯​G‖F2\+η⁡\(‖A^‖F2\+‖B¯‖F2\)\.\\underset\{\\widehat\{\\mathord\{\\btensor A\}\},\\bar\{\\mathord\{\\btensor B\}\}\}\{\\operatorname\{arg\\,min\}\}\\left\\\|D\_\{t\}\(\\widehat\{\\mathord\{\\btensor U\}\}\)\-\\widehat\{\\mathord\{\\btensor A\}\}\\widehat\{\\mathord\{\\btensor U\}\}\-\\bar\{\\mathord\{\\btensor B\}\}\\mathord\{\\btensor G\}\\right\\\|\_\{F\}^\{2\}\+\\eta\\left\(\\\|\\hat\{\\mathord\{\\btensor A\}\}\\\|\_\{F\}^\{2\}\+\\\|\\bar\{\\mathord\{\\btensor B\}\}\\\|\_\{F\}^\{2\}\\right\)\.\(6\)Here,G\\mathord\{\\btensor G\}is a matrix whoseiith column isg​\(Ti\)\\mathord\{\\btensor g\}\(T\_\{i\}\), andη∈ℝ\+\\eta\\in\\mathbb\{R\}^\{\+\}is a \(scalar\-valued\) regularization parameter controlling the condition number of the resulting linear system\.

### 2\.2O\-SAM applied to Operator Inference

The second ingredient in our workflow is O\-SAM, which we use as a means to perform concurrent subdomain coupling\. Consider a physical domainΩ∈ℝd\\Omega\\in\\mathbb\{R\}^\{d\}, ford=1,2,3d=1,2,3, and assumeΩ\\Omegahas been decomposed into two overlapping subdomains,Ω1\\Omega\_\{1\}andΩ2\\Omega\_\{2\}, such thatΩ=Ω1∪Ω2\\Omega=\\Omega\_\{1\}\\cup\\Omega\_\{2\}andΩ1∩Ω2=∅\\Omega\_\{1\}\\cap\\Omega\_\{2\}=\\emptyset, as shown in Figure[2](https://arxiv.org/html/2609.17837#S2.F2)\. Let the interior interface boundaries be defined asΓ1:=∂Ω1∩Ω2\\Gamma\_\{1\}:=\\partial\\Omega\_\{1\}\\cap\\Omega\_\{2\}andΓ2:=∂Ω2∩Ω1\\Gamma\_\{2\}:=\\partial\\Omega\_\{2\}\\cap\\Omega\_\{1\}\. In the multiplicative O\-SAM algorithm, illustrated pictorally in Figure[2](https://arxiv.org/html/2609.17837#S2.F2), the governing PDE is solved by iterating sequentially between subdomain\-local problems, with boundary information exchanged to ensure compatibility across the “Schwarz boundaries”Γ1\\Gamma\_\{1\}andΓ2\\Gamma\_\{2\}, as shown in Figure[2](https://arxiv.org/html/2609.17837#S2.F2)\. It has been shown\[[11](https://arxiv.org/html/2609.17837#bib.bib6),[14](https://arxiv.org/html/2609.17837#bib.bib5)\]that, for an overlapping DD, the Schwarz iteration process is guaranteed to converge if Dirichlet transmission boundary conditions \(BCs\) are specified on the Schwarz boundaries, provided the underlying problem is well\-posed and the overlap region is non\-empty\.

![Refer to caption](https://arxiv.org/html/2609.17837v1/figures/o-sam_algorithm.png)Fig\. 2:The basic O\-SAM algorithm for the coupling of two overlapping subdomains,Ω1\\Omega\_\{1\}andΩ2\\Omega\_\{2\}\.There is a critical detail implied by Figures[2](https://arxiv.org/html/2609.17837#S2.F2)and[3](https://arxiv.org/html/2609.17837#S2.F3): the PDEs in each of the subdomains can be solved byanymethod onanymesh, meaning that the Schwarz algorithm allows for the plug\-and\-play coupling of a variety of models and meshes, including a combination of conventional and data\-driven models\. Indeed, it has been shown in previous work that O\-SAM is capable of coupling regions with different mesh resolutions, different element types, different time integration schemes \(e\.g\., implicit and explicit\), and even different models \(e\.g\., FOM and ROM\), all without introducing any artifacts exhibited by alternative coupling methods\[[14](https://arxiv.org/html/2609.17837#bib.bib5),[15](https://arxiv.org/html/2609.17837#bib.bib7),[23](https://arxiv.org/html/2609.17837#bib.bib9),[26](https://arxiv.org/html/2609.17837#bib.bib8),[19](https://arxiv.org/html/2609.17837#bib.bib10),[25](https://arxiv.org/html/2609.17837#bib.bib3)\]\. In addition, the method is minimally intrusive to implement in existing HPC software frameworks and possesses rigorous convergence properties/guarantees\[[14](https://arxiv.org/html/2609.17837#bib.bib5),[15](https://arxiv.org/html/2609.17837#bib.bib7)\]in the all\-FOM coupling setting\.

![Refer to caption](https://arxiv.org/html/2609.17837v1/figures/Schwarz_overview_figure_FOM-ROM.png)Fig\. 3:Graphical depiction of O\-SAM being used to couple a subdomain\-local FOM with a subdomain\-local ROM\.Herein, we adopt the O\-SAM OpInf\-OpInf and OpInf\-FOM coupling methodology developed in\[[25](https://arxiv.org/html/2609.17837#bib.bib3)\]and sketched succinctly below\. The approach is depicted graphically in Figure[3](https://arxiv.org/html/2609.17837#S2.F3)\. Consider the two subdomain decomposition shown in Figure[2](https://arxiv.org/html/2609.17837#S2.F2)and assume that an OpInf ROM is employed inΩ1\\Omega\_\{1\}while a FOM is used inΩ2\\Omega\_\{2\}\. Suppose, additionally, that we have pre\-learned the reduced basis and operators defining the former model by solving the minimization problem \([6](https://arxiv.org/html/2609.17837#S2.E6)\) restricted toΩ1\\Omega\_\{1\}\. For simplicity, we assume that there are no prescribed essential boundary conditions on the physical boundary, so that the only Dirichlet data imposed on each subdomain are the Schwarz transmission conditions\. At the semi\-discrete level, the corresponding subdomain models take the form:

u^˙1\+A^1​u^1=B¯1​g1,u˙2\+A2​u2=B2​g2,\\dot\{\\widehat\{\\mathord\{\\btensor u\}\}\}\_\{1\}\+\\widehat\{\\mathord\{\\btensor A\}\}\_\{1\}\\widehat\{\\mathord\{\\btensor u\}\}\_\{1\}=\\bar\{\\mathord\{\\btensor B\}\}\_\{1\}\\mathord\{\\btensor g\}\_\{1\},\\qquad\\dot\{\\mathord\{\\btensor u\}\}\_\{2\}\+\\mathord\{\\btensor A\}\_\{2\}\\mathord\{\\btensor u\}\_\{2\}=\\mathord\{\\btensor B\}\_\{2\}\\mathord\{\\btensor g\}\_\{2\},\(7\)wheregi\\mathord\{\\btensor g\}\_\{i\}contains the Dirichlet data imposed on the Schwarz boundaryΓi\\Gamma\_\{i\}\. More precisely, let

𝒫Ωj→Γi:ℝNj→ℝnΓi,i≠j,\\mathcal\{P\}\_\{\\Omega\_\{j\}\\to\\Gamma\_\{i\}\}:\\mathbb\{R\}^\{N\_\{j\}\}\\rightarrow\\mathbb\{R\}^\{n\_\{\\Gamma\_\{i\}\}\},\\qquad i\\neq j,\(8\)denote an operator that restricts \(or traces\) the solution onΩj\\Omega\_\{j\}to the Schwarz boundaryΓi\\Gamma\_\{i\}of the neighboring subdomain, whereNj∈ℕ\+N\_\{j\}\\in\\mathbb\{N\}^\{\+\}is the number of dofs inΩj\\Omega\_\{j\}andnΓi∈ℕ\+n\_\{\\Gamma\_\{i\}\}\\in\\mathbb\{N\}^\{\+\}is the number of dofs onΓi\\Gamma\_\{i\}\. Now, the Dirichlet data imposed onΓi\\Gamma\_\{i\}are given by

gi=𝒫Ωj→Γi​uj,i≠j\.\\mathord\{\\btensor g\}\_\{i\}=\\mathcal\{P\}\_\{\\Omega\_\{j\}\\to\\Gamma\_\{i\}\}\\mathord\{\\btensor u\}\_\{j\},\\qquad i\\neq j\.
Suppose now that𝚽r,1∈ℝN1×r1\\boldsymbol\{\\Phi\}\_\{r,1\}\\in\\mathbb\{R\}^\{N\_\{1\}\\times r\_\{1\}\}is the sizer1∈ℕ\+r\_\{1\}\\in\\mathbb\{N\}^\{\+\}POD basis used to construct the OpInf model onΩ1\\Omega\_\{1\}, so that

u1​\(t\)≈𝚽r,1​u^1​\(t\)\.\\mathord\{\\btensor u\}\_\{1\}\(t\)\\approx\\boldsymbol\{\\Phi\}\_\{r,1\}\\widehat\{\\mathord\{\\btensor u\}\}\_\{1\}\(t\)\.\(9\)Letn∈ℕn\\in\\mathbb\{N\}denote the Schwarz iteration index\. At Schwarz iterationnn, the coupled OpInf–FOM subdomain problems are given by

u^˙1\(n\)\+A^1​u^1\(n\)\\displaystyle\\dot\{\\widehat\{\\mathord\{\\btensor u\}\}\}\_\{1\}^\{\(n\)\}\+\\widehat\{\\mathord\{\\btensor A\}\}\_\{1\}\\widehat\{\\mathord\{\\btensor u\}\}\_\{1\}^\{\(n\)\}=B¯1​𝒫Ω2→Γ1​u2\(n−1\),\\displaystyle=\\bar\{\\mathord\{\\btensor B\}\}\_\{1\}\\mathcal\{P\}\_\{\\Omega\_\{2\}\\to\\Gamma\_\{1\}\}\\mathord\{\\btensor u\}\_\{2\}^\{\(n\-1\)\},\(10\)u˙2\(n\)\+A2​u2\(n\)\\displaystyle\\dot\{\\mathord\{\\btensor u\}\}\_\{2\}^\{\(n\)\}\+\\mathord\{\\btensor A\}\_\{2\}\\mathord\{\\btensor u\}\_\{2\}^\{\(n\)\}=B2​𝒫Ω1→Γ2​𝚽r,1​u^1\(n\)\.\\displaystyle=\\mathord\{\\btensor B\}\_\{2\}\\mathcal\{P\}\_\{\\Omega\_\{1\}\\to\\Gamma\_\{2\}\}\\boldsymbol\{\\Phi\}\_\{r,1\}\\widehat\{\\mathord\{\\btensor u\}\}\_\{1\}^\{\(n\)\}\.\(11\)forn=0,1,…n=0,1,\.\.\.\. The Schwarz iteration \([10](https://arxiv.org/html/2609.17837#S2.E10)\) continues until the differences in the solutionsui\(k\+1\)u\_\{i\}^\{\(k\+1\)\}andui\(k\)u\_\{i\}^\{\(k\)\}are sufficiently small, fori=1,…,ndi=1,\.\.\.,n\_\{d\}, withnd∈ℕ\+n\_\{d\}\\in\\mathbb\{N\}^\{\+\}denoting the number of subdomains \(in this case,nd=2n\_\{d\}=2\), or the number of Schwarz iterations performed hits a pre\-specified maximum Schwarz iteration value,maxit∈ℕ\+\\in\\mathbb\{N\}^\{\+\}\. Following the application of a time\-discretization scheme, at each time step, the subdomain\-local OpInf and FOM problems are solved alternately until the Schwarz convergence criterion is satisfied\. The converged solution is then advanced to the next time step, and this procedure is repeated until the final simulation time is reached\.

The above workflow describes the online Schwarz coupling phase of the O\-SAM algorithm, without considering the offline construction phase of the OpInf ROM inΩ1\\Omega\_\{1\}\. We construct our subdomain\-local ROMs following the methodology described in Algorithm 2 in\[[25](https://arxiv.org/html/2609.17837#bib.bib3)\], which is based on a top\-down training approach\. In this approach: \(i\) a set of O\-SAM\-based FOM\-FOM coupled simulations onΩ1\\Omega\_\{1\}andΩ2\\Omega\_\{2\}are performed, \(ii\) snapshots forU1\\mathord\{\\btensor U\}\_\{1\}anG1\\mathord\{\\btensor G\}\_\{1\}, the solution and boundary data fromΩ1\\Omega\_\{1\}are saved, \(iii\) a POD basis,𝚽r,1\{\\bf\\Phi\}\_\{r,1\}is computed, and \(iv\) the regularized minimization problem in \([6](https://arxiv.org/html/2609.17837#S2.E6)\), restricted toΩ1\\Omega\_\{1\}, is solved to obtain theA^1\\widehat\{\\mathord\{\\btensor A\}\}\_\{1\}andB¯1\\bar\{\\mathord\{\\btensor B\}\}\_\{1\}operators in \([10](https://arxiv.org/html/2609.17837#S2.E10)\)\.

Extending the above algorithm to the case of OpInf\-OpInf coupling and multiple subdomains is straightforward\. We note that various enhancements, such as the incorporation of essential boundary conditions on the outer boundaries ofΩ1\\Omega\_\{1\}andΩ2\\Omega\_\{2\}, and the reduction of the boundary dofs is also straightforward, but we omit these details herein for the sake of brevity\. The interested reader is referred to\[[25](https://arxiv.org/html/2609.17837#bib.bib3)\]for those and other details\.

## 3Reinforcement learning for online adaptation of O\-SAM

The objective of the present work is to use RL to enable online adaptation of the O\-SAM coupling framework described in Section[2\.2](https://arxiv.org/html/2609.17837#S2.SS2)\. In particular, we seek to learn a policy that dynamically selects between ROM and FOM representations within each subdomain as the solution evolves in time, thereby enabling both ROM\-to\-FOM and FOM\-to\-ROM switching during a simulation\. We additionally consider extending this framework to allow the underlying DD itself to vary in time, so that both the subdomain partition and the model assigned to each subdomain may be selected adaptively\.

The remainder of this section introduces the RL methodology employed to learn these adaptive policies and describes its integration within the O\-SAM iteration process\. In particular, Section[3\.1](https://arxiv.org/html/2609.17837#S3.SS1)provides an overview of reinforcement learning in a general setting, Section[3\.2](https://arxiv.org/html/2609.17837#S3.SS2)describes the specific RL algorithm used in this work, namely the deep Q\-network \(DQN\) approach, and Section[3\.3](https://arxiv.org/html/2609.17837#S3.SS3)describes the application of DQN to the online model\-switching and DD adaptation problems considered in this work\.

### 3\.1Reinforcement learning

RL\[[24](https://arxiv.org/html/2609.17837#bib.bib20),[17](https://arxiv.org/html/2609.17837#bib.bib24)\]is a machine learning paradigm in which an agent learns to make sequential decisions by interacting with an environment, in our case, a numerical simulation\. The RL decision\-making process begins with the agent observing the current statest∈𝒮s\_\{t\}\\in\\mathcal\{S\}\. Based on this state, the agent selects an actionat∈𝒜a\_\{t\}\\in\\mathcal\{A\}according to a policyπ⁡\(a∣s\)\\pi\(a\\mid s\)and receives a scalar rewardrt∈ℝr\_\{t\}\\in\\mathbb\{R\}quantifying the immediate benefit of that decision\. The state represents the information available to the agent at the time a decision is made, while the action represents one of the available decisions\. In the hybrid simulations considered herein, the state is constructed from the numerical solution at a given time, the actions correspond to different model selections, and the reward balances solution accuracy and computational cost, as detailed in Sections[3\.3](https://arxiv.org/html/2609.17837#S3.SS3)and[4](https://arxiv.org/html/2609.17837#S4)\.

Given a statests\_\{t\}an actionata\_\{t\}, the environment transitions to a new statest\+1s\_\{t\+1\}according to the transition dynamics

st\+1∼ℙ\(⋅∣st,at\),s\_\{t\+1\}\\sim\\mathbb\{P\}\(\\cdot\\mid s\_\{t\},a\_\{t\}\),\(12\)whereℙ\\mathbb\{P\}denotes the state transition probability distribution\. It is evident from \([12](https://arxiv.org/html/2609.17837#S3.E12)\) that an action affects not only the reward received at the current step, but also the state from which subsequent decisions are made\. This sequential dependence is central to RL: an action that appears favorable immediately may lead to less favorable outcomes later, and vice versa\. This sequential interaction between the agent and the environment is commonly modeled as a Markov Decision Process \(MDP\)\[[17](https://arxiv.org/html/2609.17837#bib.bib24)\], defined by the tuple

\(𝒮,𝒜,P,R,γ\),\(\\mathcal\{S\},\\mathcal\{A\},P,R,\\gamma\),\(13\)where𝒮\\mathcal\{S\}is the state space,𝒜\\mathcal\{A\}is the action space,PPis the transition model,γ∈\[0,1\)\\gamma\\in\[0,1\)is the discount factor, andR:𝒮×𝒜→ℝR:\\mathcal\{S\}\\times\\mathcal\{A\}\\rightarrow\\mathbb\{R\}is the*reward function*generating the scalar rewardrt=R⁡\(st,at,st\+1\)r\_\{t\}=R\(s\_\{t\},a\_\{t\},s\_\{t\+1\}\)received at every step\.

The objective of an RL agent is to learn a policy that maximizes the expected cumulative discounted reward,

J⁡\(π\)=𝔼π​\[∑t=0∞γt​rt\],J\(\\pi\)=\\mathbb\{E\}\_\{\\pi\}\\left\[\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}r\_\{t\}\\right\],\(14\)where the expectation is taken over trajectories induced by the policy, the transition dynamics, and the initial state distribution\. The discount factorγ\\gammadetermines the relative weighting of immediate and future rewards, so that values ofγ\\gammaclose to zero place greater emphasis on immediate rewards, whereas values closer to one place greater weight on the long\-term consequences of the agent’s decisions\. In \([14](https://arxiv.org/html/2609.17837#S3.E14)\),J⁡\(π\)J\(\\pi\)should be distinguished from the reward functionRR: whereasRRdefines the reward received at each step,J⁡\(π\)J\(\\pi\)measures the expected cumulative performance of a policyπ\\pi\.

In modern RL algorithms, NNs are often used to approximate the functions that define the RL agent’s decision\-making policy, providing a practical means of representing these functions over high\-dimensional state spaces\. In the present work, we employ such an NN\-based RL approach to enable adaptive decision\-making within an evolving numerical simulation\. Specifically, we augment the O\-SAM coupling algorithm described in Section[2](https://arxiv.org/html/2609.17837#S2)with an RL\-based model\-switching strategy that adaptively selects between FOM and ROM subdomain models as the simulation evolves\. The details of this approach are given in Section[3\.3](https://arxiv.org/html/2609.17837#S3.SS3), following the description of the specific value\-based RL algorithm employed in this work, the deep Q\-network \(DQN\)\[[12](https://arxiv.org/html/2609.17837#bib.bib25)\]\(Section[3\.2](https://arxiv.org/html/2609.17837#S3.SS2)\)\.

### 3\.2Deep Q\-networks \(DQNs\)

As stated above, the RL policy learned herein utilizes a DQN\[[12](https://arxiv.org/html/2609.17837#bib.bib25)\]\. DQNs are particularly well\-suited to problems with high\-dimensional state spaces and a discrete set of available actions\. This makes DQN a natural choice for the model selection problem considered here, in which the state contains high\-dimensional information describing the evolving solution, while the available actions correspond to a discrete and relatively small set of model assignments\.

As a value\-based RL method, DQN uses a NN to approximate an action value function that quantifies the expected long\-term return associated with taking a particular action in a given state\. Toward this effect, rather than representing the policyπ\\pidirectly, the RL agent learns the action value function

Qπ\(s,a\)=𝔼π\[∑t=0∞γtrt\|s0=s,a0=a\],Q^\{\\pi\}\(s,a\)=\\mathbb\{E\}\_\{\\pi\}\\left\[\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}r\_\{t\}\\,\\middle\|\\,s\_\{0\}=s,\\;a\_\{0\}=a\\right\],\(15\)the expected return from taking actionaain statessand followingπ\\pithereafter\. It can be shown\[[2](https://arxiv.org/html/2609.17837#bib.bib26)\]that the optimal action value function satisfies the Bellman optimality equation

Q∗\(s,a\)=𝔼\[r\+γmaxa′Q∗\(s′,a′\)\|s,a\],Q^\{\*\}\(s,a\)=\\mathbb\{E\}\\left\[r\+\\gamma\\max\_\{a^\{\\prime\}\}Q^\{\*\}\(s^\{\\prime\},a^\{\\prime\}\)\\,\\middle\|\\,s,a\\right\],\(16\)from which a greedy policyπ∗​\(s\)=arg⁡maxa​Q∗​\(s,a\)\\pi^\{\*\}\(s\)=\\arg\\max\_\{a\}Q^\{\*\}\(s,a\)is recovered\. Training follows the standard DQN procedure, in which the DQN is trained using the information collected as the agent interacts with the environment\. Each interaction produces a transition\(s,a,r,s′\)\(s,a,r,s^\{\\prime\}\), recording the current statess, the actionaataken by the agent, the resulting rewardrr, and the subsequent states′s^\{\\prime\}\. These transitions are stored in a replay buffer, which serves as a collection of the agent’s past experiences\. Rather than training the network using transitions in the order in which they occur, transitions are sampled randomly from this buffer in small groups, or minibatches\. This reduces correlations between successive training samples and improves the stability of the learning process\.

DQN employs two neural networks during training: an online networkQθQ\_\{\\theta\}, whose parameters are continually updated, and a target networkQθ¯Q\_\{\\bar\{\\theta\}\}, whose parameters are held fixed for a prescribed number of training steps\. For each sampled transition, the target network is used to construct the regression target

y=r\+γ⁡\(1−d\)​Qθ¯​\(s′,arg⁡maxa′​Qθ​\(s′,a′\)\),y=r\+\\gamma\\,\(1\-d\)\\,Q\_\{\\bar\{\\theta\}\}\\\!\\left\(s^\{\\prime\},\\arg\\max\_\{a^\{\\prime\}\}Q\_\{\\theta\}\(s^\{\\prime\},a^\{\\prime\}\)\\right\),\(17\)where the argmax is evaluated over the finite, discrete action set𝒜\\mathcal\{A\}: a single forward pass of the online network ons′s^\{\\prime\}produces one valueQθ¯​\(s′,a′\)Q\_\{\\bar\{\\theta\}\}\(s^\{\\prime\},a^\{\\prime\}\)for each candidate actiona′∈𝒜a^\{\\prime\}\\in\\mathcal\{A\}, and the action having the largest such value is selected\. The first term in \([17](https://arxiv.org/html/2609.17837#S3.E17)\) is the reward received for the action just taken, while the second estimates the future reward that can be obtained from the resulting states′s^\{\\prime\}\. Ifs′s^\{\\prime\}is the terminal state,d=1d=1and the future reward contribution vanishes\. In \([17](https://arxiv.org/html/2609.17837#S3.E17)\), the online networkQθQ\_\{\\theta\}is used to identify the action expected to be best in the new states′s^\{\\prime\}, while the target networkQθ¯Q\_\{\\bar\{\\theta\}\}evaluates the value of that action\. Separating these two operations reduces the systematic overestimation of action values that can arise when the same network is used for both\. The parametersθ\\thetaof the online network are updated by minimizing the Huber loss\[[9](https://arxiv.org/html/2609.17837#bib.bib27)\]between its current predictionQθ​\(s,a\)Q\_\{\\theta\}\(s,a\)and the target valueyy\. For a residualδ:=Qθ​\(s,a\)−y\\delta:=Q\_\{\\theta\}\(s,a\)\-y, the Huber loss is defined as

ℒκ​\(δ\)=\{12​δ2,\|δ\|≤κ,κ⁡\(\|δ\|−12​κ\),\|δ\|\>κ,\\mathcal\{L\}\_\{\\kappa\}\(\\delta\)=\\begin\{cases\}\\tfrac\{1\}\{2\}\\delta^\{2\},&\|\\delta\|\\leq\\kappa,\\\\\[4\.0pt\] \\kappa\\left\(\|\\delta\|\-\\tfrac\{1\}\{2\}\\kappa\\right\),&\|\\delta\|\>\\kappa,\\end\{cases\}\(18\)whereκ\>0\\kappa\>0is a threshold parameter\. The Huber loss is selected because it is less sensitive than a squared loss to occasional large prediction errors during training\. Through repeated updates using transitions sampled from the replay buffer, the network learns to estimate the expected long\-term return associated with each available action\. During deployment, the learned action value function is then used to select the action with the largest predicted return for the current state\. The specific application of this DQN framework to adaptive model selection and DD within O\-SAM is described next, in Section[3\.3](https://arxiv.org/html/2609.17837#S3.SS3)\.

### 3\.3RL\-based online adaptation of O\-SAM

We now specialize the RL framework described above to the adaptive O\-SAM problem\. Recall from Section[2\.2](https://arxiv.org/html/2609.17837#S2.SS2)that, for a given DD and model assignment, the O\-SAM algorithm advances the solution by solving the subdomain\-local problems and iteratively exchanging Dirichlet data until convergence is declared\. Once the Schwarz iteration has converged, the coupled solution is advanced forward in time\. In the present approach, RL does not modify the Schwarz iteration itself, but instead selects the subdomain models – and, possibly, the domain decomposition – used by O\-SAM in each decision window\.

Suppose that the time interval\[0,Tfinal\]\[0,T\_\{\\text\{final\}\}\]associated with onesimulated trajectory, i\.e\., one complete simulation of a particular problem instance, defined by a given choice of initial conditions, boundary conditions and/or physical parameters, is divided into a prescribed number ofdecision windows, as illustrated in Figure[5](https://arxiv.org/html/2609.17837#S3.F5)\. We refer to this simulated trajectory as anepisode\. The proposed RL–O\-SAM workflow consists of two stages: anoffline training stageand anonline deployment stage, as illustrated in Figure[4](https://arxiv.org/html/2609.17837#S3.F4)\.

![Refer to caption](https://arxiv.org/html/2609.17837v1/figures/train_vs_deploy.png)Fig\. 4:Offline training and online deployment of the RL–O\-SAM framework\. During training \(left\), a high\-fidelity reference solutionu∗​\(t\)u^\{\*\}\(t\)is computed for each problem instance and used to evaluate the error term in the reward \([20](https://arxiv.org/html/2609.17837#S3.E20)\) Each episode advances fromt=0t=0toTfinalT\_\{\\mathrm\{final\}\}, with the reward evaluated after each decision window and used to train the DQN\. During deployment \(right\), neither the reference solution nor the reward is required; the trained policy selects an action from the observed state alone\.In the offline training stage, we first define the discrete action space𝒜\\mathcal\{A\}available to the RL agent and construct the corresponding subdomain\-local ROMs\. We consider two approaches for defining𝒜\\mathcal\{A\}\.

![Refer to caption](https://arxiv.org/html/2609.17837v1/figures/rl-schwarz.png)Fig\. 5:RL–O\-SAM workflow for two subdomains, showing the first eight of the 32 decision windows of an episode\. The episode advances from left to right in time\. At the beginning of each window, the agent observes the current state and selects a FOM/ROM assignment, which is held fixed while O\-SAM advances the coupled solution over the window\. The agent therefore acts once per decision window, rather than at each Schwarz iteration or time step \(see Figure[4](https://arxiv.org/html/2609.17837#S3.F4)\); the O\-SAM coupling remains unchanged, while the models assigned to the subdomains may vary\.In the first approach, termed “fixed\-DD”, a fixed overlapping decomposition of the computational domain intondn\_\{d\}subdomains is prescribed\. To generate training data for the subdomain\-local OpInf ROMs, we perform all\-FOM simulations in which the subdomain\-local FOMs are coupled using O\-SAM, for a collection of problem instances characterized by different initial conditions, boundary conditions, and/or physical parameters\. The resulting subdomain\-local solution snapshots and Schwarz boundary data are then used to construct a pre\-trained OpInf ROM for each subdomain, following the procedure described in Section[2](https://arxiv.org/html/2609.17837#S2)\. Thus, both a FOM \(denoted by ‘F’\) and a pre\-trained OpInf ROM \(denoted by ‘R’\) are available in each subdomain\. The action space𝒜\\mathcal\{A\}now consists of all possible FOM/ROM assignments across thendn\_\{d\}subdomains, giving a total of\|𝒜\|=2nd\|\\mathcal\{A\}\|=2^\{n\_\{d\}\}possible actions\. For example, fornd=3n\_\{d\}=3,𝒜=\{FFF, FFR, FRF, FRR, RFF, RFR, RRF, RRR\}\\mathcal\{A\}=\\\{\\text\{FFF, FFR, FRF, FRR, RFF, RFR, RRF, RRR\}\\\}, where the position of each ‘F’ or ‘R’ identifies the model assigned to the corresponding subdomain\.

In the second approach, termed “adaptive\-DD”, both the DD and the FOM/ROM subdomain assignment are included in the action space𝒜\\mathcal\{A\}\. Rather than prescribing a single fixed decomposition, we define a discrete set ofNDDN\_\{\\mathrm\{DD\}\}candidate overlapping decompositions of the computational domain, each consisting ofndn\_\{d\}subdomains\. For example, for the one\-dimensional \(1D\) domainΩ=\[0,1\]\\Omega=\[0,1\]andnd=2n\_\{d\}=2, three candidate decompositions could be prescribed as

𝒫1\\displaystyle\\mathcal\{P\}\_\{1\}:\[0,0\.5\]∪\[0\.4,1\],\\displaystyle:\\quad\[0,0\.5\]\\cup\[0\.4,1\],\(19\)𝒫2\\displaystyle\\mathcal\{P\}\_\{2\}:\[0,0\.3\]∪\[0\.2,1\],\\displaystyle:\\quad\[0,0\.3\]\\cup\[0\.2,1\],𝒫3\\displaystyle\\mathcal\{P\}\_\{3\}:\[0,0\.8\]∪\[0\.5,1\]\.\\displaystyle:\\quad\[0,0\.8\]\\cup\[0\.5,1\]\.For each candidate decomposition, every subdomain may be assigned either a FOM \(‘F’\) or a pre\-trained OpInf ROM \(‘R’\)\. Thus, an action specifies both the candidate decomposition, and the FOM/ROM assignment on that decomposition\. The resulting action space will contain\|𝒜\|=NDD​2nd\|\\mathcal\{A\}\|=N\_\{\\mathrm\{DD\}\}2^\{n\_\{d\}\}possible actions\. For the example above,NDD=3N\_\{\\mathrm\{DD\}\}=3andnd=2n\_\{d\}=2, yielding3×22=123\\times 2^\{2\}=12possible actions\. The action\(𝒫2,FR\)\(\\mathcal\{P\}\_\{2\},\\text\{FR\}\), for instance, selects decomposition𝒫2\\mathcal\{P\}\_\{2\}, with a FOM employed on its first subdomain and a ROM on its second subdomain\.

Once the action space and all required subdomain\-local ROMs have been constructed, the RL agent is trained offline using a collection of problem instances characterized by different initial conditions, boundary conditions, and/or physical parameters\. Each problem instance defines an episode, comprised of prescribed decision windows, as illustrated in the left panel of Figure[4](https://arxiv.org/html/2609.17837#S3.F4)\. At the beginning of each decision window, the agent observes the current statesks\_\{k\}and selects an actionak∈𝒜a\_\{k\}\\in\\mathcal\{A\}\. The selected action determines the FOM/ROM assignment across the subdomains and, when adaptive\-DD is considered, the DD to be used\. This configuration is held fixed throughout the decision window while the coupled problem is advanced using O\-SAM\. At the end of each decision window, the agent receives a reward that balances computational efficiency, solution accuracy, and the cost of changing the selected configuration\. Specifically, we define the reward as

rk=−costk−α​errk−β​nflip,k,r\_\{k\}=\-\\mathrm\{cost\}\_\{k\}\-\\alpha\\,\\mathrm\{err\}\_\{k\}\-\\beta\\,n\_\{\\mathrm\{flip\},k\},\(20\)wherecostk\\mathrm\{cost\}\_\{k\}denotes the computational cost incurred during decision windowkk,errk\\mathrm\{err\}\_\{k\}denotes the error in the coupled solution relative to a reference solution, andnflip,kn\_\{\\mathrm\{flip\},k\}penalizes changes in the selected action between consecutive decision windows\. The precise definition oferrk\\mathrm\{err\}\_\{k\}is problem\-dependent and is given for the advection\-diffusion and clamped linear elastic wave propagation benchmarks in Sections[4\.1](https://arxiv.org/html/2609.17837#S4.SS1)and[4\.2](https://arxiv.org/html/2609.17837#S4.SS2), respectively\. The parametersα\\alphaandβ\\betain \([20](https://arxiv.org/html/2609.17837#S3.E20)\) control the relative importance of solution accuracy and switching, respectively\. During offline RL training, a high\-fidelity reference solutionu∗​\(t\)u^\{\*\}\(t\)is computed for each training problem instance and used to evaluateerrk\\mathrm\{err\}\_\{k\}\. The specific reference solution employed depends on the test case and is defined in Section[4](https://arxiv.org/html/2609.17837#S4)\. The resulting transition\(sk,ak,rk,sk\+1\)\(s\_\{k\},a\_\{k\},r\_\{k\},s\_\{k\+1\}\)is used to train the DQN according to the procedure described in Section[3\.2](https://arxiv.org/html/2609.17837#S3.SS2)\.

Having described the offline training phase of our RL algorithm, we now move on to the online deployment stage, depicted in Figure[5](https://arxiv.org/html/2609.17837#S3.F5)\. Once offline training is complete, the trained DQN policy and all pre\-trained subdomain\-local OpInf ROMs are held fixed and deployed for the online simulation of a new problem instance\. At the beginning of each decision window, the RL agent observes the current statesk∈𝒮s\_\{k\}\\in\\mathcal\{S\}, which consists of a fixed\-dimensional representation of the current coupled solution together with the actionak−1a\_\{k\-1\}selected in the preceding decision window\. Including the previous action provides the agent with information about the current FOM/ROM assignment and, in the adaptive\-DD setting, the current domain decomposition\. The specific representation of the state employed in the numerical examples is described in Section[4](https://arxiv.org/html/2609.17837#S4)\.

Given the statesks\_\{k\}, the trained DQN selects an actionak∈𝒜a\_\{k\}\\in\\mathcal\{A\}\. This action determines the FOM/ROM assignment on each subdomain and, in the adaptive\-DD setting, the DD to be employed during the current decision window\. The selected configuration is held fixed throughout the decision window while the coupled problem is advanced using O\-SAM\. In particular, the RL agent does not act at individual time steps or Schwarz iterations\. At the beginning of the subsequent decision window, a new state is constructed from the updated coupled solution and the preceding action, and the process is repeated until the final simulation timeTTis reached\.

In contrast to the offline training stage, online deployment requires neither a reference solution nor evaluation of the reward function\. The trained DQN selects actions based solely on the observed state, and neither the DQN nor the pre\-trained OpInf ROMs are updated during the online simulation\. Thus, the high\-fidelity reference solution used to evaluate the error contribution to the reward during training is not required for a new problem instance at deployment\. Importantly, the trained policy is deployed predictively on problem instances not encountered during training\. As demonstrated in Section[4](https://arxiv.org/html/2609.17837#S4), the numerical examples considered herein train the policy using one set of initial conditions and evaluate the resulting policy on initial conditions withheld from RL training, thereby assessing the ability of the learned policy to generalize beyond the problem instances used to train it\. Whether the underlying OpInf ROMs are themselves deployed predictively or reproductively depends on the benchmark and is specified in Section[4](https://arxiv.org/html/2609.17837#S4)\.

## 4Numerical results

We evaluate the proposed RL–O\-SAM framework on two time\-dependent benchmark problems: a 1D transient advection\-diffusion problem with a moving front \(Section[4\.1](https://arxiv.org/html/2609.17837#S4.SS1)\), and a linear elastic wave propagation problem in solid mechanics \(Section[4\.2](https://arxiv.org/html/2609.17837#S4.SS2)\)\. These examples are used to assess the ability of the learned policy to perform online FOM\-ROM switching as the solution evolves\. For the advection\-diffusion problem, we consider both a fixed\-DD case, in which the RL agent selects only the FOM/ROM assignment, as well as an adaptive\-DD case, in which the agent selects both the DD and the corresponding FOM/ROM assignment\.

For all numerical experiments reported below, the DQNQθQ\_\{\\theta\}\(see Section[3\.2](https://arxiv.org/html/2609.17837#S3.SS2)\) is represented by a multilayer perceptron \(MLP\) with two hidden layers of 256 units and ReLU activation functions, with one output for each available action\. Training is performed using the Adam optimizer with a learning rate of10−310^\{\-3\}, a replay buffer containing up to2×1042\\times 10^\{4\}transitions, minibatches of 128 transitions, and a discount factor ofγ=0\.99\\gamma=0\.99\. The target network is synchronized with the online network every 400 environment steps\. Problem\-specific definitions of the state, action space, reward terms, and associated parameters are provided within the corresponding numerical examples\.

The same Schwarz solver settings are used throughout the numerical experiments\. Schwarz convergence is declared when the change in the interface traces between successive iterations, measured in the discreteL2L^\{2\}norm over the decision window and accumulated over all interfaces, falls below an absolute tolerance of10−610^\{\-6\}\. A maximum of 25 Schwarz iterations is permitted per decision window\.

### 4\.1Advection\-diffusion benchmark

We first assess the learned policy on a 1D transient advection\-diffusion problem,

ut\+c​ux=ν​ux​x,x∈\(0,1\),t∈\[0,Tfinal\],u\_\{t\}\+c\\,u\_\{x\}=\\nu\\,u\_\{xx\},\\qquad x\\in\(0,1\),\\quad t\\in\[0,T\_\{\\text\{final\}\}\],\(21\)withc=1c=1,ν=10−2\\nu=10^\{\-2\}andTfinal=0\.6T\_\{\\text\{final\}\}=0\.6, which gives rise to a Péclet number ofc​L/ν=100cL/\\nu=100\. A Dirichlet conditionu⁡\(0,t\)=1u\(0,t\)=1is imposed at the inflow boundary, together with a homogeneous Neumann condition at the outflow\. The initial condition is given by the smoothed front

u⁡\(x,0\)=12​\[1−tanh⁡\(x−x0δ\)\],δ=0\.01,u\(x,0\)=\\frac\{1\}\{2\}\\left\[1\-\\tanh\\\!\\left\(\\frac\{x\-x\_\{0\}\}\{\\delta\}\\right\)\\right\],\\qquad\\delta=0\.01,\(22\)wherex0x\_\{0\}determines the initial location of the front\. During RL training,x0x\_\{0\}is sampled uniformly from\[0\.05,0\.5\]\[0\.05,0\.5\], so that each episode corresponds to a different problem instance\. As the sharp front propagates across the domain, different subdomains require higher\-fidelity resolution at different times, motivating adaptive FOM/ROM model selection\.

The domainΩ=\(0,1\)\\Omega=\(0,1\)is first discretized usingN=2048N=2048interior nodes\. Because conformal subdomain meshes are employed, the overlapping subdomain meshes are obtained by restricting this global discretization to the corresponding subdomain intervals\. The precise subdomain locations depend on whether a fixed\- or adaptive\-DD is employed and are specified below\. Second\-order central differences are used for diffusion and first\-order upwinding for advection\. The resulting semi\-discrete systems are integrated using an adaptive Runge\-Kutta method\. The subdomains are coupled using multiplicative O\-SAM with Dirichlet transmission conditions, as discussed in Section[2\.2](https://arxiv.org/html/2609.17837#S2.SS2)\.

The subdomain\-local OpInf ROMs utilized within our couplings are constructed offline using 10 FOM trajectories, each corresponding to a different initial condition \([22](https://arxiv.org/html/2609.17837#S4.E22)\) with

x0∈\{0\.280,0\.478,0\.115,0\.477,0\.190,0\.240,0\.423,0\.234,0\.297,0\.062\}\.x\_\{0\}\\in\\\{0\.280,\\,0\.478,\\,0\.115,\\,0\.477,\\,0\.190,\\,0\.240,\\,0\.423,\\,0\.234,\\,0\.297,\\,0\.062\\\}\.\(23\)Each trajectory contains 401 snapshots sampled uniformly over\[0,Tfinal\]\[0,T\_\{\\text\{final\}\}\]\. The same 10 trajectories are used to construct all subdomain\-local ROMs, with the snapshots restricted to the degrees of freedom associated with each subdomain\. Thus, the local bases are constructed from a common snapshot set rather than from separate FOM simulations\.

![Refer to caption](https://arxiv.org/html/2609.17837v1/advdiff_truth.png)Fig\. 6:Monolithic full order solution in the space\-time plane, used as the ground truth for both settings below\. The front enters at the inflow boundary and advects to the right at speedcc, remaining sharp over the whole time interval\.For the advection\-diffusion benchmark, the state supplied to the DQN consists of a fixed\-dimensional representation of the current solutionuutogether with the action selected in the preceding decision window\. TheN=2048N=2048\-dimensional solution field is max\-pooled to 200 components to reduce the dimension of the solution state provided to the DQN, and the previous action is represented using one\-hot encoding, i\.e\., a binary vector with a single entry equal to one at the index corresponding to the selected action and zeros elsewhere\. For the reward function \([20](https://arxiv.org/html/2609.17837#S3.E20)\),costk\\mathrm\{cost\}\_\{k\}is taken to be the wall time required to advance decision windowkk, whileerrk\\mathrm\{err\}\_\{k\}is the relative error with respect to the corresponding monolithic FOM solution\. To evaluate this error, the converged subdomain solutions are first assembled into a global solution field𝐮k\\mathbf\{u\}\_\{k\}on the global grid, with the solutions from neighboring subdomains averaged at grid points in the overlap regions\. The error is then computed as

errk=‖uk−uk∗‖2‖uk∗‖2,\\mathrm\{err\}\_\{k\}=\\frac\{\\\|\\mathord\{\\btensor u\}\_\{k\}\-\\mathord\{\\btensor u\}\_\{k\}^\{\*\}\\\|\_\{2\}\}\{\\\|\\mathord\{\\btensor u\}\_\{k\}^\{\*\}\\\|\_\{2\}\},\(24\)where𝐮k∗\\mathbf\{u\}\_\{k\}^\{\*\}denotes the corresponding monolithic FOM solution on the global grid and∥⋅∥2\\\|\\cdot\\\|\_\{2\}denotes the discreteℓ2\\ell\_\{2\}norm\. We setα=500\\alpha=500andβ=1\.5\\beta=1\.5, with both parameters selected through a coarse parameter sweep to balance computational cost and solution accuracy, as well as to avoid policies that collapse to either the all\-FOM or all\-ROM assignment\. The switching termnflip,kn\_\{\\mathrm\{flip\},k\}counts the number of subdomains whose FOM/ROM assignment changes relative to the preceding decision window; in the adaptive\-DD setting \(Section[4\.1\.2](https://arxiv.org/html/2609.17837#S4.SS1.SSS2)\), a change in the selected decomposition is also penalized\. The purpose of this term is to discourage unnecessary changes in the model assignment or DD between consecutive decision windows, which may incur additional computational overhead and lead to frequent oscillations between otherwise comparable actions\. No switching penalty is applied during the first decision window\. The DQN is trained for 500 episodes and subsequently evaluated predictively forx0=0\.22x\_\{0\}=0\.22\. Since this initial condition is not used during either DQN training or OpInf ROM construction, both the learned policy and the underlying OpInf ROMs are evaluated predictively for this test case\. The monolithic FOM solution forx0=0\.22x\_\{0\}=0\.22, shown in Figure[6](https://arxiv.org/html/2609.17837#S4.F6), is used solely as a reference for evaluating the accuracy of the deployed policy\.

We consider two configurations\. In the first fixed\-DD case, the DD is fixed and the agent selects only the FOM/ROM assignment among the three subdomains, yielding23=82^\{3\}=8possible actions \(Section[4\.1\.1](https://arxiv.org/html/2609.17837#S4.SS1.SSS1)\)\. In the second adaptive\-DD case, the agent selects both the domain decomposition and the FOM/ROM assignment, yielding 40 possible actions \(Section[4\.1\.2](https://arxiv.org/html/2609.17837#S4.SS1.SSS2)\)\.

#### 4\.1\.1Fixed\-DD case

First, we consider the first offline training approach described in Section[3\.3](https://arxiv.org/html/2609.17837#S3.SS3), in which the DD is fixed, and we are learning a policy that determines online whether each subdomain is assigned a FOM \(‘F’\) or a ROM \(’R’\)\. In this study, the domain is decomposed into three overlapping subdomains of equal length with 30% overlap,

Ω1=\[0,0\.4167\],Ω2=\[0\.2917,0\.7083\],Ω3=\[0\.5833,1\],\\Omega\_\{1\}=\[0,0\.4167\],\\qquad\\Omega\_\{2\}=\[0\.2917,0\.7083\],\\qquad\\Omega\_\{3\}=\[0\.5833,1\],\(25\)containing 853, 854, and 853 of the 2048 nodes, respectively\. Each subdomain is equipped with an OpInf ROM of dimensionr=12r=12\. The ROM dimension was selected by testing a range of candidate values, withr=12r=12being the smallest dimension that provided acceptable ROM accuracy for this configuration\. The OpInf least\-squares problem \([6](https://arxiv.org/html/2609.17837#S2.E6)\) uses a Tikhonov regularization withη=10−8\\eta=10^\{\-8\}\. For the RL training, the time interval\[0,Tfinal\]\[0,T\_\{\\text\{final\}\}\]is divided uniformly into 128 decision windows\. At the beginning of each window, the agent selects one of the23=82^\{3\}=8possible FOM/ROM assignments, which is held fixed while O\-SAM advances the coupled solution over the window\. After the Schwarz iteration converges, the RL agent observes the resulting state and selects the assignment for the next decision window\.

Table 1:Advection\-diffusion benchmark, fixed\-DD\. Results for the learned policy and the eight possible fixed FOM/ROM assignments\. ‘F’ and ‘R’ denote full order and reduced order subdomain models, respectively, ordered from left to right\. The relativeL2L^\{2\}error is computed with respect to the monolithic FOM reference solution\. ‘Schwarz iterations’ denotes the mean number of multiplicative Schwarz iterations per decision window, averaged over the 128 decision windows\. ‘Reward’ denotes the sum of the rewards over all decision windows of the episode\. The best value in each column is shown in bold\.Table[1](https://arxiv.org/html/2609.17837#S4.T1)compares the trained policy with each of the eight fixed FOM/ROM assignments\. The learned policy achieves a relativeL2L^\{2\}error of8\.02×10−48\.02\\times 10^\{\-4\}, approximately five times smaller than that of the most accurate fixed assignment containing at least one ROM \(FFR\), while requiring an average of 3\.00 Schwarz iterations per decision window, compared with 1\.61 for the all\-FOM assignment\. The increase in Schwarz iterations when ROMs are introduced is consistent with the observations in Wentland et al\.\[[27](https://arxiv.org/html/2609.17837#bib.bib17)\], where Schwarz couplings involving lower\-dimensional ROMs were found to generally require more iterations to converge than the corresponding all\-FOM coupling\. This behavior arises because inaccuracies in the ROM solution can lead to larger discrepancies between neighboring subdomain solutions, requiring the Schwarz iteration to work harder to reconcile the solutions across the overlap\. Despite the observed increase in Schwarz iterations, the learned policy reduces the wall time from3\.233\.23s for FFF to1\.911\.91s\. As expected, the all\-ROM \(RRR\) coupling yields the lowest wall time,0\.510\.51s, but at the expense of a substantially larger relativeL2L^\{2\}error of1\.18×10−21\.18\\times 10^\{\-2\}\. Thus, the learned policy provides a compromise between the high accuracy of the all\-FOM assignment and the low computational cost of the all\-ROM assignment\.

![Refer to caption](https://arxiv.org/html/2609.17837v1/advdiff_fixed_decision.png)Fig\. 7:Advection\-diffusion benchmark, fixed\-DD\. FOM/ROM model assignment selected by the learned policy for each subdomain as a function of time, with red denoting the FOM and blue denoting the ROM\. The FOM assignment shifts from left to right as the sharp front advects through the domain, preferentially providing full order resolution in the subdomain containing the moving front\.As shown in Figure[7](https://arxiv.org/html/2609.17837#S4.F7), the learned policy adapts the model assignment as the sharp front propagates through the domain, preferentially assigning the FOM to the subdomain containing the front while using ROMs elsewhere\. As the front advects from left to right, the FOM assignment correspondingly shifts from one subdomain to the next\. Figure[8](https://arxiv.org/html/2609.17837#S4.F8)shows the resulting hybrid solution and its pointwise error relative to the monolithic FOM reference solution\.

![Refer to caption](https://arxiv.org/html/2609.17837v1/advdiff_fixed_solerr.png)Fig\. 8:Advection\-diffusion benchmark, fixed\-DD\.Left:hybrid solution produced by the learned policy, with the two subdomain interfaces indicated in red\.Right:pointwise error relative to the monolithic FOM reference solution, shown on a logarithmic scale\. The error is largest in the reduced order subdomains and near the advecting front, and smallest where the policy selects the FOM\.Finally, the accuracy\-cost tradeoff is summarized by the Pareto plot in Figure[9](https://arxiv.org/html/2609.17837#S4.F9), which shows the relativeL2L^\{2\}error vs\. wall time for the learned policy and all fixed model assignments\. It is clear from this figure that the learned policy, denoted by ‘RL\(8\)’, is Pareto\-optimal\.

Fig\. 9:Advection\-diffusion benchmark, fixed\-DD\. Pareto plot showing the relativeL2L^\{2\}error vs\. measured wall time for the learned policy and all fixed FOM/ROM assignments\. The learned policy, denoted by ‘RL\(8\)’, is shown in red\. Policies farther toward the lower left provide a more favorable accuracy\-cost tradeoff, making the learned policy Pareto\-optimal\.4\.1\.1\.1\. Effect of the switching penalty\.

Fig\. 10:Advection–diffusion benchmark, fixed\-DD, effect of the switching penalty during training\. Number of FOM/ROM model switches per training episode forβ=0\\beta=0\(no switching penalty\) andβ=1\.5\\beta=1\.5\(including a switching penalty\)\. Faint lines show the number of switches in each episode, while solid lines show the corresponding moving averages\. Without the switching penalty, the number of switches levels off at approximately 33 per episode, whereas withβ=1\.5\\beta=1\.5, it decreases to approximately 13 per episode\.It is interesting to understand the effect of the switching penalty in the reward function \([20](https://arxiv.org/html/2609.17837#S3.E20)\), in particular, the extent to which it discourages unnecessary changes in the FOM/ROM assignment\. To isolate the effect of this term, we retrain the fixed\-DD policy withβ=0\\beta=0andβ=1\.5\\beta=1\.5in \([20](https://arxiv.org/html/2609.17837#S3.E20)\), while holding all other parameters fixed\. Figure[10](https://arxiv.org/html/2609.17837#S4.F10)shows the number of model switches per episode during training\. Both policies initially exhibit approximately 100 switches per episode, as the untrained agent explores the action space\. As training progresses, the number of switches decreases for both policies\. Without the switching penalty \(β=0\\beta=0\), however, the number eventually levels off at approximately 33 switches per episode, whereas withβ=1\.5\\beta=1\.5, it continues to decrease to approximately 13\.

![Refer to caption](https://arxiv.org/html/2609.17837v1/advdiff_mu_decision.png)Fig\. 11:Advection\-diffusion benchmark, fixed\-DD, effect of the switching penalty on the deployed policy\. FOM/ROM model assignments selected by policies trained without the switching penalty \(β=0\\beta=0, left\) and with the switching penalty \(β=1\.5\\beta=1\.5, right\), with red denoting the FOM and blue denoting the ROM\. The switching penalty suppresses the rapid FOM/ROM switching observed forβ=0\\beta=0, reducing the total number of model switches over the episode from 31 to 4\.The effect of the switching penalty is even more apparent when examining the deployed policies\. As shown in Figure[11](https://arxiv.org/html/2609.17837#S4.F11), the policy trained withβ=0\\beta=0\(no switching penalty\) switches models 31 times over the episode, with much of this switching occurring as rapid back\-and\-forth changes early in the simulation\. In contrast, the policy trained withβ=1\.5\\beta=1\.5switches only four times, producing a considerably more coherent model\-assignment pattern in which the FOM tracks the advecting front before transitioning to an all\-ROM assignment after the front leaves the domain\.

The above results reinforce the fact that the purpose of the switching penalty is not to improve solution accuracy, but rather to regularize the behavior of the learned policy by suppressing unnecessary FOM/ROM switching\. This is particularly important for deployment within a production code, where each model switch may require a state transfer between the FOM and ROM representations and therefore incur additional computational costs\.

#### 4\.1\.2Adaptive\-DD case

Next, we assess the second training strategy introduced in Section[3](https://arxiv.org/html/2609.17837#S3), in which the RL agent selects both the DD and the FOM/ROM assignment\. Five candidate DDs, denoted by𝒫i\\mathcal\{P\}\_\{i\},i=1,…,5i=1,\\ldots,5, are considered\. These decompositions are obtained by varying the locations of the two interior interfaces while maintaining a 25% overlap between adjacent subdomains\. The resulting subdomain extents and node counts are listed in Table[2](https://arxiv.org/html/2609.17837#S4.T2)\.

Table 2:Advection\-diffusion benchmark, adaptive\-DD\. Subdomain extents and node counts for the five candidate DDs𝒫i\\mathcal\{P\}\_\{i\},i=1,…,5i=1,\\ldots,5available to the RL agent\. Each decomposition consists of three overlapping subdomains with 25% overlap between adjacent subdomains\. A separate OpInf ROM is constructed offline for each of the 15 subdomains, so that all candidate ROMs are available prior to online deployment\.The set of candidate domain decompositions is prescribed offline, and a separate OpInf ROM is constructed offline for each subdomain of each𝒫i\\mathcal\{P\}\_\{i\}, resulting in a total of 15 subdomain\-local ROMs\. Thus, all ROMs that may be selected by the agent are available prior to deployment, and no ROM construction is performed online\. For this configuration, each ROM has dimensionr=10r=10, selected as the smallest dimension among the candidate values tested that provided acceptable ROM accuracy\. The OpInf least\-squares problem \([6](https://arxiv.org/html/2609.17837#S2.E6)\) utilizes a Tikhonov regularization parameter ofη=10−8\\eta=10^\{\-8\}\. At the beginning of each decision window, the agent selects one of the five candidate domain decompositions𝒫i\\mathcal\{P\}\_\{i\}together with one of the eight possible FOM/ROM assignments, yielding an action space of5×23=405\\times 2^\{3\}=40actions\. The selected decomposition and FOM/ROM assignment are held fixed throughout the decision window while O\-SAM advances the coupled solution\.

Table 3:Advection\-diffusion benchmark, adaptive\-DD\. Results for the learned policy with an action space comprising 40 actions, obtained by combining the five candidate domain decompositions𝒫i\\mathcal\{P\}\_\{i\},i=1,…,5i=1,\\ldots,5listed in Table[1](https://arxiv.org/html/2609.17837#S4.T1)with the eight possible FOM/ROM assignments\. ‘F’ and ‘R’ denote full order and reduced order subdomain models, respectively, ordered from left to right\. The relativeL2L^\{2\}error is computed with respect to the monolithic FOM reference solution\. ‘Schwarz iterations’ denotes the mean number of multiplicative Schwarz iterations per decision window, averaged over the 128 decision windows\. ‘Reward’ denotes the sum of the rewards over all decision windows of the episode\. The best value in each column is shown in bold\.Table[3](https://arxiv.org/html/2609.17837#S4.T3)summarizes the performance of the learned policy with an action space comprising 40 possible actions, obtained by combining the five candidate DDs𝒫i\\mathcal\{P\}\_\{i\},i=1,…,5i=1,\\ldots,5, with the eight possible FOM/ROM assignments\. The learned policy achieves a relativeL2L^\{2\}error of1\.03×10−31\.03\\times 10^\{\-3\}, making it approximately seven times more accurate than the FFR coupling, the most accurate baseline FOM/ROM assignment containing at least one ROM\. The learned policy requires an average of 2\.73 Schwarz iterations per decision window, compared with 1\.60 for the all\-FOM \(FFF\) assignment\. Although an individual ROM subdomain solve is approximately 34\-44 times less expensive than the corresponding FOM solve, it is interesting to observe that this reduction does not translate directly into an equivalent reduction in overall wall time\. The increased number of Schwarz iterations associated with FOM\-ROM coupling partially offsets the computational savings provided by the reduced order subdomain solves\. Despite requiring more Schwarz iterations, the learned policy has a lower wall time than FFF,2\.842\.84s compared with3\.643\.64s\. The all\-ROM \(RRR\) assignment has the lowest wall time,0\.390\.39s, but its relativeL2L^\{2\}error,1\.81×10−21\.81\\times 10^\{\-2\}, is approximately 18 times larger than that of the learned policy\.

![Refer to caption](https://arxiv.org/html/2609.17837v1/figures/adaptive_decision.png)Fig\. 12:Advection\-diffusion benchmark, adaptive\-DD\. Domain decomposition and FOM/ROM model assignment selected by the learned policy as functions of time, with red denoting the FOM and blue denoting the ROM\. Changes in the subdomain boundaries indicate changes in the selected candidate decomposition𝒫i\\mathcal\{P\}\_\{i\}, while changes in color indicate changes in the FOM/ROM assignment\.![Refer to caption](https://arxiv.org/html/2609.17837v1/figures/adaptive_sol.png)

![Refer to caption](https://arxiv.org/html/2609.17837v1/figures/adaptive_relerr.png)

Fig\. 13:Advection\-diffusion benchmark, adaptive\-DD\.Left:hybrid solution produced by the learned policy, with the selected subdomain interfaces shown in red and the three subdomains labeled\. The interface locations change between decision windows as the agent selects among the five candidate domain decompositions𝒫i\\mathcal\{P\}\_\{i\}\.Right:pointwise error relative to the monolithic FOM reference solution, shown on a logarithmic scale\. Changes in the selected domain decomposition are visible as discontinuities in the interface locations\.Figures[12](https://arxiv.org/html/2609.17837#S4.F12)and[13](https://arxiv.org/html/2609.17837#S4.F13)illustrate how the learned policy, ‘RL\(40\)’, jointly adapts the DD and FOM/ROM assignment as the solution evolves\. Figure[12](https://arxiv.org/html/2609.17837#S4.F12)shows the corresponding decomposition and model assignments selected by the agent over time, while Figure[13](https://arxiv.org/html/2609.17837#S4.F13)shows the resulting hybrid solution and pointwise error\. The reader can observe from these figures that the RL agent makes frequent changes to the selected DD, particularly during the early portion of the simulation; however, these changes do not translate into an improved accuracy\-cost tradeoff relative to the fixed\-DD strategy considered in Section[4\.1\.1](https://arxiv.org/html/2609.17837#S4.SS1.SSS1)\.

![Refer to caption](https://arxiv.org/html/2609.17837v1/figures/advdiff_compare.png)Fig\. 14:Advection\-diffusion benchmark, comparison of fixed and adaptive\-DD\. RelativeL2L^\{2\}error \(left\) and measured wall time \(right\) for the learned policy and the baseline FOM/ROM assignments\. The top row corresponds to the fixed\-DD with 8 actions, while the bottom row corresponds to the adaptive\-DD with 40 actions\. The learned policy is shown in red\.Finally, Figure[14](https://arxiv.org/html/2609.17837#S4.F14)provides a direct comparison of the accuracy and computational cost obtained with the fixed and adaptive\-DD strategies, complementing the quantitative results in Tables[1](https://arxiv.org/html/2609.17837#S4.T1)and[3](https://arxiv.org/html/2609.17837#S4.T3)\. In both cases, the learned policy achieves substantially lower error than the baseline FOM/ROM assignments containing at least one ROM\. However, expanding the action space from 8 to 40 actions by allowing the agent to select the DD does not improve the learned policy\. Compared with the fixed\-DD policy, the adaptive\-DD policy requires slightly fewer Schwarz iterations per decision window \(2\.73 vs\. 3\.00\), but has a higher relativeL2L^\{2\}error \(1\.03×10−31\.03\\times 10^\{\-3\}vs\.8\.02×10−48\.02\\times 10^\{\-4\}\) and a substantially higher wall time \(2\.84 s vs\. 1\.91 s\)\. We conclude that, for this benchmark, allowing the agent to select the DD in addition to the FOM/ROM assignment does not improve the overall accuracy\-cost tradeoff\. This result is favorable from an implementation standpoint, since dynamically changing the DD during a simulation is not feasible in current production implementations of O\-SAM\. The remainder of this work therefore focuses on online adaptation of the FOM/ROM assignment for a fixed\-DD\. Nevertheless, this experiment demonstrates that the proposed RL framework can accommodate a larger action space in which both the model assignment and domain decomposition are selected online\.

### 4\.2Clamped linear elastic wave propagation benchmark

The second example considered herein is a linear elastic wave propagation benchmark implemented within theNorma\.jlopen\-source three\-dimensional \(3D\) solid mechanics code\[[20](https://arxiv.org/html/2609.17837#bib.bib22)\]\. This test case is a variant of similar test cases considered in\[[15](https://arxiv.org/html/2609.17837#bib.bib7),[1](https://arxiv.org/html/2609.17837#bib.bib21),[25](https://arxiv.org/html/2609.17837#bib.bib3)\]

Consider the equations of dynamic linear elasticity on a domainΩ⊂ℝ3\\Omega\\subset\\mathbb\{R\}^\{3\}over the time intervalt∈\[0,Tfinal\]t\\in\[0,T\_\{\\text\{final\}\}\]\. Letu​\(x,t\):=\(ux​\(x,t\),uy​\(x,t\),uz​\(x,t\)\)∈ℝ3\\mathord\{\\btensor u\}\(\\mathord\{\\btensor x\},t\):=\\left\(\\begin\{array\}\[\]\{ccc\}u\_\{x\}\(\\mathord\{\\btensor x\},t\),&u\_\{y\}\(\\mathord\{\\btensor x\},t\),&u\_\{z\}\(\\mathord\{\\btensor x\},t\)\\end\{array\}\\right\)\\in\\mathbb\{R\}^\{3\}denote the displacement field\. In the absence of body forces, the governing equations are

ρ​u¨−∇⋅𝝈⁡\(u\)=𝟎in​Ω×\(0,T\],\\rho\\ddot\{\\mathord\{\\btensor u\}\}\-\\nabla\\cdot\\boldsymbol\{\\sigma\}\(\\mathord\{\\btensor u\}\)=\\mathbf\{0\}\\qquad\\text\{in \}\\Omega\\times\(0,T\],\(26\)whereρ\\rhois the material density and𝝈\\boldsymbol\{\\sigma\}is the Cauchy stress tensor\. For a homogeneous, isotropic, linear elastic material,

𝝈⁡\(u\)=λ​tr​\(𝜺⁡\(u\)\)​I\+2​μ​𝜺​\(u\),\\boldsymbol\{\\sigma\}\(\\mathord\{\\btensor u\}\)=\\lambda\\,\\mathrm\{tr\}\\\!\\left\(\\boldsymbol\{\\varepsilon\}\(\\mathord\{\\btensor u\}\)\\right\)\\mathord\{\\btensor I\}\+2\\mu\\,\\boldsymbol\{\\varepsilon\}\(\\mathord\{\\btensor u\}\),\(27\)with infinitesimal strain tensor

𝜺⁡\(u\)=12​\(∇u\+∇uT\)\.\\boldsymbol\{\\varepsilon\}\(\\mathord\{\\btensor u\}\)=\\frac\{1\}\{2\}\\left\(\\nabla\\mathord\{\\btensor u\}\+\\nabla\\mathord\{\\btensor u\}^\{T\}\\right\)\.\(28\)Here,λ\\lambdaandμ\\muare the Lamé parameters, which may be expressed in terms of the Young’s modulusEEand Poisson ratioν\\nuas

λ=E​ν\(1\+ν\)​\(1−2​ν\),μ=E2​\(1\+ν\)\.\\lambda=\\frac\{E\\nu\}\{\(1\+\\nu\)\(1\-2\\nu\)\},\\qquad\\mu=\\frac\{E\}\{2\(1\+\\nu\)\}\.\(29\)
For the clamped linear elastic wave propagation benchmark, the computational domain is a slender 3D beam geometry, so thatΩ=\(−5×10−4,5×10−4\)×\(−5×10−4,5×10−4\)×\(−0\.5,0\.5\)\\Omega=\\left\(\-5\\times 10^\{\-4\},5\\times 10^\{\-4\}\\right\)\\times\\left\(\-5\\times 10^\{\-4\},5\\times 10^\{\-4\}\\right\)\\times\(\-0\.5,0\.5\)m\. We utilize the following material properties:E=1E=1GPa,ν=0\\nu=0, andρ=1000\\rho=1000kg/m3\. The problem is evolved over0≤t≤Tfinal0\\leq t\\leq T\_\{\\text\{final\}\}, withTfinal=2×10−3T\_\{\\text\{final\}\}=2\\times 10^\{\-3\}s\. The spatial domain is discretized using eight\-node hexahedral elements with an axial mesh spacing ofΔ​z=0\.002\\Delta z=0\.002m, and withΔ​x=Δ​y=10−3\\Delta x=\\Delta y=10^\{\-3\}, making the mesh exactly one element thick in thexx\- andyy\-directions\. All overlapping subdomain meshes employed in the O\-SAM calculations are conformal\. The resulting semi\-discrete equations are integrated in time using the implicit Newmark\-β\\betatime integration scheme, withγ=12\\gamma=\\frac\{1\}\{2\}andβ=14\\beta=\\frac\{1\}\{4\}, with a time step ofΔ​t=3\.125×10−6\\Delta t=3\.125\\times 10^\{\-6\}s\.

Herein, for simplicity, we wish to restrict the dynamics to the axial \(zz\) direction in order to obtain a 1D problem solved using the 3D code,Norma\.jl\. In order to accomplish this, homogeneous Dirichlet boundary conditions are imposed on the transverse displacement components:

ux​\(x,t\)=0,atx=±5×10−4,uy​\(x,t\)=0,aty=±5×10−4,\\begin\{array\}\[\]\{cc\}u\_\{x\}\(\\mathord\{\\btensor x\},t\)=0,&\\text\{at \}x=\\pm 5\\times 10^\{\-4\},\\\\ u\_\{y\}\(\\mathord\{\\btensor x\},t\)=0,&\\text\{at \}y=\\pm 5\\times 10^\{\-4\},\\end\{array\}\(30\)The beam is additionally clamped in the axial direction at its two ends, giving

uz​\(x,t\)=0,at​z=±0\.5\.\\begin\{array\}\[\]\{cc\}u\_\{z\}\(\\mathord\{\\btensor x\},t\)=0,&\\text\{at \}z=\\pm 0\.5\.\\end\{array\}\(31\)The remaining unconstrained displacement components are subject to homogeneous Neumann boundary conditions\.

The initial conditions for our benchmark are:

u​\(x,0\)=\(00f⁡\(z\)\),u˙​\(x,0\)=\(00f′​\(z\)\)\\mathord\{\\btensor u\}\(\\mathord\{\\btensor x\},0\)=\\begin\{pmatrix\}0\\\\ 0\\\\ f\(z\)\\end\{pmatrix\},\\qquad\\dot\{\\mathord\{\\btensor u\}\}\(\\mathord\{\\btensor x\},0\)=\\begin\{pmatrix\}0\\\\ 0\\\\ f^\{\\prime\}\(z\)\\end\{pmatrix\}\(32\)wheref⁡\(z\)∈ℝf\(z\)\\in\\mathbb\{R\}specifies the initial axial displacement profile, andf′​\(z\):=d​fd​zf^\{\\prime\}\(z\):=\\frac\{df\}\{dz\}\. Here, we consider initial axial displacement profiles of the form

f⁡\(z\)=a​exp⁡\(−z22​s2\),f\(z\)=a\\exp\\left\(\-\\frac\{z^\{2\}\}\{2s^\{2\}\}\\right\),\(33\)fora,s∈ℝa,s\\in\\mathbb\{R\}, so that

f′​\(z\)=−a​cs2​z​exp⁡\(−z22​s2\),f^\{\\prime\}\(z\)=\-\\frac\{ac\}\{s^\{2\}\}z\\exp\\left\(\-\\frac\{z^\{2\}\}\{2s^\{2\}\}\\right\),\(34\)wherec=E/ρc=\\sqrt\{E/\\rho\}m/s is the speed of the traveling wave\.

To assess the ability of the learned policy to generalize across different initial conditions, we consider a family of problems in which both the initial pulse location and width are varied\. Specifically, we generalize the initial condition in \([33](https://arxiv.org/html/2609.17837#S4.E33)\) by defining

f⁡\(z\)=a​exp⁡\(−\(z−z0\)22​s2\)\.f\(z\)=a\\exp\\left\(\-\\frac\{\(z\-z\_\{0\}\)^\{2\}\}\{2s^\{2\}\}\\right\)\.\(35\)with the corresponding initial velocity obtained by differentiatingf⁡\(z\)f\(z\)with respect tozz, as in \([34](https://arxiv.org/html/2609.17837#S4.E34)\)\. We consider seven equally spaced values of the pulse locationz0∈\[−0\.05,0\.10\]z\_\{0\}\\in\[\-0\.05,0\.10\]m, with spacing0\.0250\.025m, and three pulse widthss∈\{0\.02,0\.03,0\.04\}s\\in\\\{0\.02,0\.03,0\.04\\\}m, yielding a total of 21 problem instances\. The pulse amplitude is fixed ata=1\.0×10−3a=1\.0\\times 10^\{\-3\}m\.

For all configurations considered below, each subdomain\-local OpInf ROM uses a reduced basis of dimensionr=10r=10\. As in Section[4\.1](https://arxiv.org/html/2609.17837#S4.SS1), the ROM dimension was selected by testing a range of candidate values, withr=10r=10being the smallest dimension that provided acceptable ROM accuracy\. The OpInf least\-squares problem \([6](https://arxiv.org/html/2609.17837#S2.E6)\) uses Tikhonov regularization with regularization parameterη=10−8\\eta=10^\{\-8\}\. The ROMs are constructed using FOM snapshot data from all 21 problem instances considered below, making the OpInf ROMs are reproductive with respect to the problem instances considered in this benchmark\. Snapshots are collected at fixed, uniformly spaced output times with a snapshot interval ofΔ​tsnap=6\.25×10−6\\Delta t\_\{\\mathrm\{snap\}\}=6\.25\\times 10^\{\-6\}s, corresponding to every two Newmark time steps\.

Based on the results of Section[4\.1\.2](https://arxiv.org/html/2609.17837#S4.SS1.SSS2), we restrict the present study to fixed\-DDs\. For the advection\-diffusion benchmark \(Section[4\.1](https://arxiv.org/html/2609.17837#S4.SS1)\), allowing the RL agent to additionally select among multiple candidate domain decompositions did not improve the overall accuracy\-cost tradeoff relative to adapting only the FOM/ROM assignment\. Moreover, dynamically changing the DD during a simulation is not currently supported within 3D code implementations such asNorma\.jl\.

For the subdomain\-local OpInf ROM training, each episode is divided uniformly into 32 decision windows, corresponding to 20 time steps per decision window\. The state supplied to the DQN consists of the displacement field together with the action selected in the preceding decision window\. For the reward function \([20](https://arxiv.org/html/2609.17837#S3.E20)\),costk\\mathrm\{cost\}\_\{k\}is the computational cost incurred during decision windowkk\. Since the linear elastic wave propagation benchmark is adynamicsproblem, there are three solution fields of interest: the displacement, the velocity and the acceleration, denoted byu\\mathord\{\\btensor u\},v\\mathord\{\\btensor v\}anda\\mathord\{\\btensor a\}, respectively; this means that all three fields should go into theerrk\\mathrm\{err\}\_\{k\}term in \([20](https://arxiv.org/html/2609.17837#S3.E20)\)\. Toward this effect, to evaluateerrk\\mathrm\{err\}\_\{k\}, the displacement, velocity, and acceleration solutions from the individual subdomains are first assembled into global fieldsuk\\mathord\{\\btensor u\}\_\{k\},vk\\mathord\{\\btensor v\}\_\{k\}, andak\\mathord\{\\btensor a\}\_\{k\}on the global grid spanning the full domainΩ\\Omega, with the solutions from neighboring subdomains averaged at grid points in the overlap regions\. The error is then defined as

𝐞𝐫𝐫k=‖u−u∗‖22\+‖v−v∗‖22\+‖a−a∗‖22‖u∗‖22\+‖v∗‖22\+‖a∗‖22,\\mathbf\{err\}\_\{k\}=\\sqrt\{\\frac\{\\\|\\mathord\{\\btensor u\}\-\\mathord\{\\btensor u\}^\{\*\}\\\|\_\{2\}^\{2\}\+\\\|\\mathord\{\\btensor v\}\-\\mathord\{\\btensor v\}^\{\*\}\\\|\_\{2\}^\{2\}\+\\\|\\mathord\{\\btensor a\}\-\\mathord\{\\btensor a\}^\{\*\}\\\|\_\{2\}^\{2\}\}\{\\\|\\mathord\{\\btensor u\}^\{\*\}\\\|\_\{2\}^\{2\}\+\\\|\\mathord\{\\btensor v\}^\{\*\}\\\|\_\{2\}^\{2\}\+\\\|\\mathord\{\\btensor a\}^\{\*\}\\\|\_\{2\}^\{2\}\}\},\(36\)where\(u∗,v∗,a∗\)\(\\mathord\{\\btensor u\}^\{\*\},\\mathord\{\\btensor v\}^\{\*\},\\mathord\{\\btensor a\}^\{\*\}\)denotes the corresponding all\-FOM O\-SAM reference solution and∥⋅∥2\\\|\\cdot\\\|\_\{2\}denotes the discreteℓ2\\ell\_\{2\}norm\. We setα=200\\alpha=200andβ=1\.5\\beta=1\.5in \([20](https://arxiv.org/html/2609.17837#S3.E20)\)\. As before, these parameters were fixed prior to training by sweeping each over several orders of magnitude and selecting values that kept the computational\-cost, accuracy, and switching contributions to the reward on comparable scales\.

In what follows, two DDs are considered: \(i\) a two subdomain decomposition \(Section[4\.3](https://arxiv.org/html/2609.17837#S4.SS3)\), and a \(ii\) three subdomain decomposition \(Section[4\.4](https://arxiv.org/html/2609.17837#S4.SS4)\)\. For each decomposition, the RL agent selects the FOM/ROM assignment among the subdomains, yielding22=42^\{2\}=4and23=82^\{3\}=8possible actions, respectively\. The high\-fidelity reference solution introduced in Section[2\.2](https://arxiv.org/html/2609.17837#S2.SS2)is the corresponding all\-FOM O\-SAM solution and includes the displacement, velocity and acceleration fields, as in \([36](https://arxiv.org/html/2609.17837#S4.E36)\)\. Thus, the reference is obtained using the FF assignment for the two subdomain decomposition and the FFF assignment for the three subdomain decomposition\. We adopt an all\-FOM O\-SAM reference because this problem is implemented in the production\-oriented𝙽𝚘𝚛𝚖𝚊\.𝚓𝚕\{\\tt Norma\.jl\}solid mechanics code, where, for problems of practical interest, construction of a single monolithic mesh may be difficult or impractical\. This reference is used to evaluate the error term \([36](https://arxiv.org/html/2609.17837#S4.E36)\) in the reward during training and to assess the accuracy of the learned policy and fixed FOM/ROM assignments during predictive testing\.

Separate DQN policies are trained for the two configurations\. Of the 21 problem instances, 11 are used for RL training and the remaining 10 are reserved for testing the learned policies\. The split is constructed so that the test cases represent combinations of pulse location and width not encountered during RL training, but within the parameter range spanned by the training set\. The two subdomain policy is trained for 1000 episodes, while the three subdomain policy is trained for 5000 episodes\. Thus, although the OpInf ROMs are reproductive for this benchmark, the RL policies are evaluated predictively on the held\-out test cases\.

### 4\.3Two subdomain study

We first consider a decomposition of the 3D beam into two overlapping subdomains,Ω1\\Omega\_\{1\}andΩ2\\Omega\_\{2\}\. Each subdomain spans the full cross\-section of the beam, with axial extentsz∈\[−0\.50,−0\.20\]z\\in\[\-0\.50,\-0\.20\]m forΩ1\\Omega\_\{1\}andz∈\[−0\.30,0\.50\]z\\in\[\-0\.30,0\.50\]m forΩ2\\Omega\_\{2\}\. Thus, the two subdomains overlap overz∈\[−0\.30,−0\.20\]z\\in\[\-0\.30,\-0\.20\]m\. The corresponding RL action space consists of the four possible FOM/ROM assignments to the two subdomains\.

![Refer to caption](https://arxiv.org/html/2609.17837v1/wave_2sd_frames.png)Fig\. 15:Linear elastic wave benchmark, two subdomain case\. Snapshots of the axial displacementuz​\(z,t\)u\_\{z\}\(z,t\)at six times over the episode, for the predictive test casez0=−0\.0375z\_\{0\}=\-0\.0375m\. Colors indicates the FOM/ROM assignment selected by the trained DQN policy \(red: full order; blue: reduced order\), and the vertical black line marks the divide betweenΩ1\\Omega\_\{1\}\(z∈\[−0\.50,−0\.20\]z\\in\[\-0\.50,\-0\.20\]m\) andΩ2\\Omega\_\{2\}\(z∈\[−0\.30,0\.50\]z\\in\[\-0\.30,0\.50\]m\)\.The DQN policy for this configuration is trained for 1000 episodes using the training problem instances described above\. We then evaluate the trained policy predictively on the held\-out test cases\. Figure[15](https://arxiv.org/html/2609.17837#S4.F15)shows the axial displacementuz​\(z,t\)u\_\{z\}\(z,t\)and the FOM/ROM assignments selected by the policy at six representative times for a held\-out test case withz0=−0\.0375z\_\{0\}=\-0\.0375m\. Att=0t=0, the pulse is located primarily inΩ2\\Omega\_\{2\}, and the policy assigns the FOM toΩ2\\Omega\_\{2\}and the ROM toΩ1\\Omega\_\{1\}\. As the pulse propagates to the left and entersΩ1\\Omega\_\{1\}, the policy assigns the FOM toΩ1\\Omega\_\{1\}and the ROM toΩ2\\Omega\_\{2\}\. Following reflection from the clamped boundary atz=−0\.5z=\-0\.5m, the pulse propagates back towardΩ2\\Omega\_\{2\}\. As the reflected wave moves fromΩ1\\Omega\_\{1\}towardΩ2\\Omega\_\{2\}, the policy temporarily assigns the FOM to both subdomains\. Once the wave has propagated farther intoΩ2\\Omega\_\{2\}, the policy returns to a ROM\-FOM assignment\. The policy therefore adapts the FOM/ROM assignment in response to the propagation of the localized wave between the two subdomains\. This behavior is expected: the FOM is preferentially assigned to the subdomain containing the propagating wave, while the ROM is used in regions where the solution is comparatively smooth\.

Table 4:Linear elastic wave benchmark, two subdomain case\. Results for the learned policy and the four possible fixed FOM/ROM assignments\. ‘F’ and ‘R’ denote full order and reduced order subdomain models, respectively, ordered from left to right\. The relativeL2L^\{2\}error is computed with respect to the FF assignment, which therefore has zero error by construction\.The results correspond to the predictive test casez0=−0\.0375z\_\{0\}=\-0\.0375m shown in Figure[15](https://arxiv.org/html/2609.17837#S4.F15)\.

Table[4](https://arxiv.org/html/2609.17837#S4.T4)compares the learned policy with each of the four fixed FOM/ROM assignments\. The learned policy achieves a relativeL2L^\{2\}error of1\.59×10−21\.59\\times 10^\{\-2\}, more than an order of magnitude smaller than that of the most accurate fixed assignment containing at least one ROM \(RF\), whose error is2\.43×10−12\.43\\times 10^\{\-1\}\. The learned policy also reduces the wall time from 52\.91 s for the all\-FOM \(FF\) assignment to 49\.77 s\. The all\-FOM relative error is identically zero, as the all\-FOM solution is utilized as the reference solution in theL2L^\{2\}error calculation\. As expected, the all\-ROM \(RR\) assignment has the lowest wall time, 33\.23 s, but its relativeL2L^\{2\}error of2\.69×10−12\.69\\times 10^\{\-1\}is approximately 17 times larger than that of the learned policy\. Thus, as in the advection\-diffusion benchmark \(Section[4\.1](https://arxiv.org/html/2609.17837#S4.SS1)\), the learned policy provides a compromise between the high accuracy of the all\-FOM assignment and the lower computational cost of assignments involving fixed ROMs\.

### 4\.4Three subdomain study

We next consider a decomposition of the 3D beam into three overlapping subdomains,Ω1\\Omega\_\{1\},Ω2\\Omega\_\{2\}, andΩ3\\Omega\_\{3\}\. Each subdomain spans the full cross\-section of the beam, with axial extentsz∈\[−0\.50,−0\.10\]z\\in\[\-0\.50,\-0\.10\]m,z∈\[−0\.20,0\.20\]z\\in\[\-0\.20,0\.20\]m, andz∈\[0\.10,0\.50\]z\\in\[0\.10,0\.50\]m, respectively\. Thus, adjacent subdomains overlap overz∈\[−0\.20,−0\.10\]z\\in\[\-0\.20,\-0\.10\]m andz∈\[0\.10,0\.20\]z\\in\[0\.10,0\.20\]m\. The corresponding RL action space consists of the23=82^\{3\}=8possible FOM/ROM assignments to the three subdomains\.

![Refer to caption](https://arxiv.org/html/2609.17837v1/wave_3sd_frames.png)Fig\. 16:Linear elastic wave benchmark, three subdomain case\. Snapshots of the axial displacementuz​\(z,t\)u\_\{z\}\(z,t\)at the same six times as Figure[15](https://arxiv.org/html/2609.17837#S4.F15), for the predictive test casez0=−0\.025z\_\{0\}=\-0\.025m,s=0\.02s=0\.02m\. Shading and vertical lines are as in Figure[15](https://arxiv.org/html/2609.17837#S4.F15), now marking the two divides betweenΩ1\\Omega\_\{1\}\(z∈\[−0\.50,−0\.10\]z\\in\[\-0\.50,\-0\.10\]m\),Ω2\\Omega\_\{2\}\(z∈\[−0\.20,0\.20\]z\\in\[\-0\.20,0\.20\]m\), andΩ3\\Omega\_\{3\}\(z∈\[0\.10,0\.50\]z\\in\[0\.10,0\.50\]m\)\.The DQN policy for the three subdomain configuration is trained for 5000 episodes using the training problem instances described above\. As for the two subdomain case, we evaluate the trained policy predictively on the held\-out test cases\. Figure[16](https://arxiv.org/html/2609.17837#S4.F16)shows the axial displacementuz​\(z,t\)u\_\{z\}\(z,t\)and the FOM/ROM assignments selected by the policy at six representative times for a held\-out test case withz0=−0\.025z\_\{0\}=\-0\.025m ands=0\.02s=0\.02m\. Att=0t=0, the pulse is located primarily inΩ2\\Omega\_\{2\}, and the policy assigns the FOM toΩ2\\Omega\_\{2\}and ROMs toΩ1\\Omega\_\{1\}andΩ3\\Omega\_\{3\}\. As the pulse propagates toward the left end of the beam, the policy assigns the FOM toΩ1\\Omega\_\{1\}and ROMs to the other two subdomains\. Following reflection from the left clamped boundary, the wave propagates toward the right end of the beam, with the FOM assignment shifting fromΩ1\\Omega\_\{1\}andΩ2\\Omega\_\{2\}, toΩ2\\Omega\_\{2\}andΩ3\\Omega\_\{3\}, and finally toΩ3\\Omega\_\{3\}as the wave approaches the right boundary\. After reflection, the FOM is again assigned toΩ2\\Omega\_\{2\}as the wave propagates back toward the center\. As in the two subdomain case \(Section[4\.3](https://arxiv.org/html/2609.17837#S4.SS3)\), the results described above exhibit the expected behavior: the policy preferentially assigns the FOM to the subdomain or subdomains containing the propagating wave and uses the ROM elsewhere\. The three subdomain example demonstrates that this behavior is retained as the number of subdomains and possible FOM/ROM assignments increases\.

Table 5:Linear elastic wave benchmark, three subdomain case\. Results for the learned policy and the eight possible fixed FOM/ROM assignments\. ‘F’ and ‘R’ denote full order and reduced order subdomain models, respectively, ordered from left to right\. The relativeL2L^\{2\}error is computed with respect to the FFF assignment, which therefore has zero error by construction\. The results correspond to the predictive test casez0=−0\.025z\_\{0\}=\-0\.025m ands=0\.02s=0\.02m shown in Figure[16](https://arxiv.org/html/2609.17837#S4.F16)\.Table[5](https://arxiv.org/html/2609.17837#S4.T5)compares the learned policy with each of the eight fixed FOM/ROM assignments\. The learned policy achieves a relativeL2L^\{2\}error of3\.90×10−33\.90\\times 10^\{\-3\}, approximately 40 times smaller than that of the most accurate fixed assignment containing at least one ROM \(FFR\), whose error is1\.54×10−11\.54\\times 10^\{\-1\}\. The learned policy also reduces the wall time from 62\.63 s for the all\-FOM \(FFF\) assignment to 54\.04 s\. As for the two subdomain version of this problem \(Section[4\.4](https://arxiv.org/html/2609.17837#S4.SS4)\), the all\-FOM relative error is identically zero, due to the fact that the all\-FOM solution is utilized as the reference solution in theL2L^\{2\}error calculation\. As expected, the all\-ROM \(RRR\) assignment has the lowest wall time, 42\.10 s, but its relativeL2L^\{2\}error of2\.94×10−12\.94\\times 10^\{\-1\}is approximately 75 times larger than that of the learned policy\. Thus, the three subdomain results exhibit the same overall behavior as the two subdomain results: the learned policy retains accuracy much closer to that of the all\-FOM solution while reducing its computational cost\.

## 5Conclusions

In this work, we have developed an RL\-based approach for online adaptation of hybrid FOM\-ROM models coupled through the overlapping Schwarz alternating method\. The objective is to address a limitation of conventional hybrid Schwarz coupling, in which the model assigned to each subdomain is typically selecteda prioriand remains fixed throughout the simulation\. Such a fixed assignment may be inadequate for transient problems involving moving or evolving features, since the regions requiring high\-fidelity resolution can change over time\. To address this challenge, we employed DQNs to learn policies that dynamically select between subdomain\-local FOMs and pre\-trained OpInf ROMs as the solution evolves\. The policies are trained offline using a reward function that balances solution accuracy and computational cost while penalizing unnecessary model switching, and are subsequently deployed online without requiring access to a reference FOM solution\.

We studied our proposed approach using two transient benchmark problems\. First, a 1D advection\-diffusion problem with a moving front was used to study online FOM/ROM switching for both fixed and adaptive domain decompositions and to examine the effect of the switching penalty in the reward function\. For this first benchmark, both the OpInf ROMs and the learned RL policy were evaluated predictively\. Second, a 3D linear elastic wave propagation problem implemented in theNorma\.jlcode\[[20](https://arxiv.org/html/2609.17837#bib.bib22)\]was used to evaluate the approach on a solid mechanics test case with two and three subdomain decompositions, in which the initial pulse location and width are varied and the learned policies are evaluated predictively on unseen problem instances\. For this second benchmark, the OpInf ROMs were reproductive, while the RL policies were evaluated predictively on problem instances withheld from RL training\.

For the advection\-diffusion benchmark, the learned policy successfully adapted the FOM/ROM assignment as the moving front propagated through the domain\. For the fixed domain decomposition, it achieved a relativeL2L^\{2\}error of8\.02×10−48\.02\\times 10^\{\-4\}, approximately five times smaller than that of the most accurate fixed FOM/ROM assignment containing at least one ROM, while also reducing the wall time relative to the all\-FOM coupled solution\. The switching penalty was found to play an important role in regularizing the learned policy, reducing unnecessary FOM/ROM switching without changing the underlying objective of balancing accuracy and computational cost\. Allowing the agent to adapt the DD in addition to the FOM/ROM assignment did not improve the overall accuracy\-cost tradeoff: although the adaptive\-DD policy required slightly fewer Schwarz iterations, it exhibited both higher error and higher wall time than the learned policy for the fixed domain decomposition\. These results suggest that online adaptation of the FOM/ROM assignment within a fixed\-DD setting provides the more effective and practically realizable strategy for the problems considered here\.

The 3D linear elastic wave propagation benchmark demonstrated that the same approach can be applied within theNorma\.jlsolid mechanics code\[[20](https://arxiv.org/html/2609.17837#bib.bib22)\]\. For both the two and three subdomain configurations, the learned policies exhibited the expected behavior, preferentially assigning the FOM to subdomains containing the propagating wave and using ROMs elsewhere\. As the wave propagated through the beam and reflected from the clamped boundaries, the policies adapted the FOM/ROM assignments accordingly\. The improvement relative to fixed FOM/ROM assignments was particularly pronounced for this benchmark, with the learned policies achieving errors more than an order of magnitude smaller than the most accurate fixed assignments containing at least one ROM, while also reducing the wall time relative to the all\-FOM assignments\. Importantly, these results were obtained for held\-out combinations of pulse location and width, indicating that the learned policies can generalize their model selection behavior to initial conditions not encountered during the training of the RL agent\.

Several directions for future work follow naturally from the present study\. First, the implementation developed for the linear elastic wave propagation benchmark provides a workflow for applying the proposed adaptive FOM/ROM coupling strategy to fully 3D problems inNorma\.jl\. A natural next step is therefore to apply this capability to more complex, nonlinear and truly 3D solid mechanics problems involving richer spatial and temporal dynamics, for which the regions requiring high\-fidelity resolution may evolve in less predictable ways\. Second, the OpInf ROMs considered in the present work are constructed offline and remain fixed throughout deployment\. An important direction for future work is to enable these ROMs to adapt online using newly available FOM information generated during the hybrid simulation\. Online adaptation of data\-driven ROMs remains an open research area; one promising approach to investigate is the streaming OpInf methodology of Koike et al\.\[[10](https://arxiv.org/html/2609.17837#bib.bib18)\], which enables reduced operators to be updated incrementally as new data become available\. Finally, the present work employs DQN to learn the adaptive model\-selection policy\. Future studies could investigate more advanced reinforcement learning approaches, including methods designed for larger or more structured action spaces, as the number of subdomains and available modeling choices increases\.

## Acknowledgements

Support for this work was received through Sandia National Laboratories’ Laboratory Directed Research and Development \(LDRD\) program and through the U\.S\. Department of Energy, Office of Science, Office of Advanced Scientific Computing Research, Mathematical Multifaceted Integrated Capability Centers \(MMICCs\) program, under Field Work Proposal 22025291 and the Multifaceted Mathematics for Predictive Digital Twins \(M2dt\) project\. Additionally, the writing of this manuscript was funded in part by Irina Tezaur’s Presidential Early Career Award for Scientists and Engineers \(PECASE\)\.

The authors wish to thank Chris Wentland for valuable discussions and insightful advice that helped shape the research presented in this paper\. The authors additionally wish to thank Alejandro Mota for his assistance in using and developing theNorma\.jlcode\.

Sandia National Laboratories is a multi\-mission laboratory managed and operated by National Technology and Engineering Solutions of Sandia, LLC\., a wholly owned subsidiary of Honeywell International, Inc\., for the U\.S\. Department of Energy’s National Nuclear Security Administration under contract DE\-NA0003525\.

## References

- \[1\]J\. Barnett, I\. Tezaur, and A\. Mota\(2022\)The Schwarz alternating method for the seamless coupling of nonlinear reduced order models and full order models\.External Links:2210\.12551,[Link](https://arxiv.org/abs/2210.12551)Cited by:[§4\.2](https://arxiv.org/html/2609.17837#S4.SS2.p1.1)\.
- \[2\]R\. Bellman\(1966\)Dynamic programming\.science153\(3731\),pp\. 34–37\.Cited by:[§3\.2](https://arxiv.org/html/2609.17837#S3.SS2.p2.2)\.
- \[3\]P\. J\. Blonigan and E\. J\. Parish\(2023\)Evaluation of dual\-weighted residual and machine learning error estimation for projection\-based reduced\-order models of steady partial differential equations\.Computer Methods in Applied Mechanics and Engineering409,pp\. 115988\.External Links:ISSN 0045\-7825,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.cma.2023.115988),[Link](https://www.sciencedirect.com/science/article/pii/S0045782523001111)Cited by:[§1](https://arxiv.org/html/2609.17837#S1.p6.1)\.
- \[4\]A\. Corigliano, M\. Dossi, and S\. Mariani\(2015\)Model order reduction and domain decomposition strategies for the solution of the dynamic elastic–plastic structural problem\.Computer Methods in Applied Mechanics and Engineering290\(C\),pp\. 127–155\.Cited by:[§1](https://arxiv.org/html/2609.17837#S1.p4.1)\.
- \[5\]M\. Ebrahimi and M\. Yano\(2024\)A hyperreduced reduced basis element method for reduced\-order modeling of component\-based nonlinear systems\.Computer Methods in Applied Mechanics and Engineering431,pp\. 117254\.External Links:ISSN 0045\-7825,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.cma.2024.117254),[Link](https://www.sciencedirect.com/science/article/pii/S0045782524005103)Cited by:[§1](https://arxiv.org/html/2609.17837#S1.p5.1)\.
- \[6\]A\. Hedayat, A\. Padovan, and K\. Duraisamy\(2026\)Toward adaptive non\-intrusive reduced\-order models: design and challenges\.Structural and Multidisciplinary Optimization69,pp\. 151\.External Links:[Document](https://dx.doi.org/10.1007/s00158-026-04355-1),ISSN 1615\-1488Cited by:[§1](https://arxiv.org/html/2609.17837#S1.p5.1)\.
- \[7\]P\. Holmes, J\. L\. Lumley, and G\. Berkooz\(1996\)Turbulence, coherent structures, dynamical systems and symmetry\.Cambridge University Press\.Cited by:[§2\.1](https://arxiv.org/html/2609.17837#S2.SS1.p2.1)\.
- \[8\]C\. Huang, K\. Duraisamy, and C\. Merkle\(2022\)Component\-based reduced order modeling of large\-scale complex systems\.Frontiers in Physics10\.External Links:[Link](https://www.frontiersin.org/journals/physics/articles/10.3389/fphy.2022.900064),[Document](https://dx.doi.org/10.3389/fphy.2022.900064),ISSN 2296\-424XCited by:[§1](https://arxiv.org/html/2609.17837#S1.p5.1)\.
- \[9\]P\. J\. Huber\(1992\)Robust estimation of a location parameter\.InBreakthroughs in statistics: Methodology and distribution,pp\. 492–518\.Cited by:[§3\.2](https://arxiv.org/html/2609.17837#S3.SS2.p3.2)\.
- \[10\]T\. Koike, P\. Mohan, M\. T\. H\. de Frahan, J\. Bessac, and E\. Qian\(2026\)Streaming operator inference for model reduction of large\-scale dynamical systems\.External Links:2601\.12161,[Link](https://arxiv.org/abs/2601.12161)Cited by:[§5](https://arxiv.org/html/2609.17837#S5.p5.1),[Remark 1](https://arxiv.org/html/2609.17837#Thmremark1.p1.1.1)\.
- \[11\]P\. Lions\(1988\)On the Schwarz alternating method I\.Note:in First International Symposium on Domain Decomposition Methods for Partial Differential EquationsCited by:[§2\.2](https://arxiv.org/html/2609.17837#S2.SS2.p1.1)\.
- \[12\]V\. Mnih, K\. Kavukcuoglu, D\. Silver, A\. Graves, I\. Antonoglou, D\. Wierstra, and M\. Riedmiller\(2013\)Playing atari with deep reinforcement learning\.arXiv preprint arXiv:1312\.5602\.Cited by:[§3\.1](https://arxiv.org/html/2609.17837#S3.SS1.p4.1),[§3\.2](https://arxiv.org/html/2609.17837#S3.SS2.p1.1)\.
- \[13\]A\. Mota, D\. Koliesnikova, I\. K\. Tezaur, and J\. Hoy\(2025\)A fundamentally new coupled approach to contact mechanics via the Dirichlet–Neumann Schwarz Alternating Method\.International Journal for Numerical Methods in Engineering126\(9\),pp\. e70039\.Cited by:[Remark 2](https://arxiv.org/html/2609.17837#Thmremark2.p1.1.1)\.
- \[14\]A\. Mota, I\. Tezaur, and C\. Alleman\(2017\)The Schwarz alternating method in solid mechanics\.Computer Methods in Applied Mechanics and Engineering319,pp\. 19–51\.External Links:[Document](https://dx.doi.org/10.1016/j.cma.2017.02.006)Cited by:[§1](https://arxiv.org/html/2609.17837#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.17837#S2.SS2.p1.1),[§2\.2](https://arxiv.org/html/2609.17837#S2.SS2.p2.1),[Remark 2](https://arxiv.org/html/2609.17837#Thmremark2.p1.1.1)\.
- \[15\]A\. Mota, I\. Tezaur, and G\. Phlipot\(2022\)The Schwarz alternating method for dynamic solid mechanics\.International Journal for Numerical Methods in Engineering,pp\. 1–36\.External Links:[Document](https://dx.doi.org/10.1002/nme.6982)Cited by:[§1](https://arxiv.org/html/2609.17837#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.17837#S2.SS2.p2.1),[§4\.2](https://arxiv.org/html/2609.17837#S4.SS2.p1.1),[Remark 2](https://arxiv.org/html/2609.17837#Thmremark2.p1.1.1)\.
- \[16\]B\. Peherstorfer and K\. Willcox\(2016\)Data\-driven operator inference for nonintrusive projection\-based model reduction\.Computer Methods in Applied Mechanics and Engineering306,pp\. 196–215\.Cited by:[§1](https://arxiv.org/html/2609.17837#S1.p6.1),[§2\.1](https://arxiv.org/html/2609.17837#S2.SS1.p1.1),[§2](https://arxiv.org/html/2609.17837#S2.p1.1)\.
- \[17\]M\. L\. Puterman\(2014\)Markov decision processes: discrete stochastic dynamic programming\.John Wiley & Sons\.Cited by:[§3\.1](https://arxiv.org/html/2609.17837#S3.SS1.p1.1),[§3\.1](https://arxiv.org/html/2609.17837#S3.SS1.p2.2)\.
- \[18\]A\. Radermacher and S\. Reese\(2014\)Model reduction in elastoplasticity: proper orthogonal decomposition combined with adaptive sub\-structuring\.Computational Mechanics54,pp\. 677–687\.Cited by:[§1](https://arxiv.org/html/2609.17837#S1.p4.1)\.
- \[19\]C\. Rodriguez, I\. Tezaur, A\. Mota, A\. Gruber, E\. Parish, and C\. Wentland\(2025\)Transmission conditions for the non\-overlapping schwarz coupling of full order and operator inference models\.External Links:2509\.12228,[Link](https://arxiv.org/abs/2509.12228)Cited by:[§1](https://arxiv.org/html/2609.17837#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.17837#S2.SS2.p2.1)\.
- \[20\]Sandia National Laboratories\(2026\)Norma\.jl: a julia testbed for solid mechanics, coupling, and multiphysics\.Note:https://github\.com/sandialabs/Norma\.jlGitHub repository, accessed August 18, 2026Cited by:[§4\.2](https://arxiv.org/html/2609.17837#S4.SS2.p1.1),[§5](https://arxiv.org/html/2609.17837#S5.p2.1),[§5](https://arxiv.org/html/2609.17837#S5.p4.1)\.
- \[21\]H\. A\. Schwarz\(1870\)Ueber einen Grenzübergang durch alternirendes Verfahren\.Zürcher u\. Furrer\.Cited by:[§1](https://arxiv.org/html/2609.17837#S1.p2.1),[§2](https://arxiv.org/html/2609.17837#S2.p1.1)\.
- \[22\]K\. Smetana and T\. Taddei\(2023\)Localized model reduction for nonlinear elliptic partial differential equations: localized training, partition of unity, and adaptive enrichment\.SIAM Journal on Scientific Computing45\(3\),pp\. A1300–A1331\.External Links:[Document](https://dx.doi.org/10.1137/22M148402X),[Link](https://doi.org/10.1137/22M148402X),https://doi\.org/10\.1137/22M148402XCited by:[§1](https://arxiv.org/html/2609.17837#S1.p5.1)\.
- \[23\]W\. Snyder, I\. Tezaur, and C\. Wentland\(2023\)Domain decomposition\-based coupling of physics\-informed neural networks via the Schwarz alternating method\.External Links:2311\.00224Cited by:[§1](https://arxiv.org/html/2609.17837#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.17837#S2.SS2.p2.1)\.
- \[24\]R\. S\. Sutton and A\. G\. Barto\(2018\)Reinforcement learning: an introduction\.2nd edition,The MIT Press,Cambridge, MA\.Cited by:[§1](https://arxiv.org/html/2609.17837#S1.p6.1),[§3\.1](https://arxiv.org/html/2609.17837#S3.SS1.p1.1),[Remark 3](https://arxiv.org/html/2609.17837#Thmremark3.p1.1.1)\.
- \[25\]I\. Tezaur, E\. Parish, A\. Gruber, I\. Moore, C\. Wentland, and A\. Mota\(2025\)Hybrid coupling with operator inference and the overlapping Schwarz alternating method\.arXiv preprint arXiv:2511\.20687\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2511.20687)Cited by:[§1](https://arxiv.org/html/2609.17837#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.17837#S2.SS1.p2.1),[§2\.2](https://arxiv.org/html/2609.17837#S2.SS2.p2.1),[§2\.2](https://arxiv.org/html/2609.17837#S2.SS2.p3.1),[§2\.2](https://arxiv.org/html/2609.17837#S2.SS2.p5.1),[§2\.2](https://arxiv.org/html/2609.17837#S2.SS2.p6.1),[§2](https://arxiv.org/html/2609.17837#S2.p1.1),[§4\.2](https://arxiv.org/html/2609.17837#S4.SS2.p1.1),[Remark 2](https://arxiv.org/html/2609.17837#Thmremark2.p1.1.1),[footnote 1](https://arxiv.org/html/2609.17837#footnote1)\.
- \[26\]C\. R\. Wentland, F\. Rizzi, J\. L\. Barnett, and I\. K\. Tezaur\(2025\)The role of interface boundary conditions and sampling strategies for schwarz\-based coupling of projection\-based reduced order models\.Journal of Computational and Applied Mathematics465,pp\. 116584\.External Links:[Document](https://dx.doi.org/10.1016/j.cam.2025.116584)Cited by:[§1](https://arxiv.org/html/2609.17837#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.17837#S2.SS2.p2.1),[Remark 2](https://arxiv.org/html/2609.17837#Thmremark2.p1.1.1)\.
- \[27\]C\. R\. Wentland, F\. Rizzi, J\. L\. Barnett, and I\. K\. Tezaur\(2025\)The role of interface boundary conditions and sampling strategies for Schwarz\-based coupling of projection\-based reduced order models\.Journal of Computational and Applied Mathematics465,pp\. 116584\.External Links:[Document](https://dx.doi.org/10.1016/j.cam.2025.116584)Cited by:[§4\.1\.1](https://arxiv.org/html/2609.17837#S4.SS1.SSS1.p2.1)\.

Similar Articles

Adaptive Multi-Horizon Reinforcement Learning

arXiv cs.LG

This paper proposes a multi-horizon reinforcement learning approach that adaptively selects and combines temporal horizons, enabling robust adaptation to changing reward structures without manual discount factor tuning, with empirical validation in MiniGrid environments.

Scalable Constrained Multi-Agent Reinforcement Learning via State Augmentation and Consensus for Separable Dynamics

arXiv cs.LG

This paper presents a distributed approach for constrained multi-agent reinforcement learning that uses state-augmented policy learning and neighbor-to-neighbor consensus over dual variables to satisfy global resource constraints while scaling linearly with the number of agents. Experiments on smart grid demand response demonstrate that consensus coordination is essential for feasibility, scaling to thousands of agents unlike centralized training approaches.