Inverse Critical Experiment Design via Gradient Optimization and a Multigroup Attention-Based Neural Network Architecture
Summary
Researchers from MIT present a methodology for inverse design of nuclear critical experiments using deep neural networks with a novel multigroup attention pooling architecture and gradient-based optimization to maximize neutronic similarity coefficients. The approach is applied to validate a HALEU fuel transportation cask, achieving high similarity scores for three configurations of interest.
View Cached Full Text
Cached at: 06/05/26, 02:18 AM
# Inverse Critical Experiment Design via Gradient Optimization and a Multigroup Attention-Based Neural Network Architecture
Source: [https://arxiv.org/html/2606.04033](https://arxiv.org/html/2606.04033)
###### Abstract
The validation of advanced nuclear reactor designs and fuel concepts requires critical experiments with high neutronic similarity to the target technology\. Neutronic similarity is quantified by the correlation coefficientckc\_\{k\}, which captures the shared bias inkeffk\_\{\\text\{eff\}\}induced by uncertainties in nuclear data\. Generally, ack≥0\.9c\_\{k\}\\geq 0\.9is needed for an experiment to be sufficiently similar to a target technology\. This work presents a methodology for the inverse design of critical experiments\. Deep neural network surrogate modeling and nonparametric gradient optimization are used to generate experiment geometries that maximizeckc\_\{k\}\.
A deep neural network is trained on OpenMC\-calculated sensitivity vectors for grid\-based critical experiment geometries\. The model architecture combines a U\-Net convolutional encoder\-decoder with a novel multigroup attention pooling layer, introduced to capture the differing spatial dependencies of sensitivities\. Multigroup attention pooling is shown to achieve better performance than traditional pooling, as well as interpretable internal behavior\. The differentiability of the surrogate enables gradient\-based optimization of the full combinatorial design space, allowingckc\_\{k\}to be maximized by directly changing the material assignment of each position in the geometry grid\.
The method is applied to the validation of the TN\-Americas TN\-LC transportation cask with HALEU fuel, for which existing critical experiment coverage is limited\. The optimization procedure is shown to produce experiment geometries achievingckc\_\{k\}scores of 0\.97757, 0\.81324, and 0\.93276 for three configurations of interest\. This approach demonstrates the potential of deep learning and gradient optimization to accelerate the development of advanced nuclear technology\.
###### keywords:
Criticality Experiments , Attention Mechanisms , Deep Learning , Differentiable Surrogates , HALEU , Inverse Design
††journal:Annals of Nuclear Energy\\affiliation
organization=Department of Nuclear Science and Engineering, Massachusetts Institute of Technology,addressline=60 Vassar St\., city=Cambridge, postcode=02139, state=MA, country=US
## 1Introduction
Driven by the rapid scaling of AI data centers, domestic reindustrialization, and a growing desire for energy independence, the public and private sectors alike are turning to nuclear power to meet the energy demands of the coming century\[[6](https://arxiv.org/html/2606.04033#bib.bib1)\]\[[29](https://arxiv.org/html/2606.04033#bib.bib2)\]\[[7](https://arxiv.org/html/2606.04033#bib.bib3)\]\. The revitalization of the nuclear industry has led to the introduction of a plethora of advanced reactor designs, fuels, and materials\. Critical experiments play a central role in supporting the validation and eventual deployment of these technologies\.
Critical \(or criticality\) experiments are low\-power assemblies of fissile material designed to precisely measure and validate the neutronic properties of a nuclear system without requiring an expensive and difficult full\-scale demonstration\[[28](https://arxiv.org/html/2606.04033#bib.bib4)\]\. The National Criticality Experiment Research Center \(NCERC\) supports criticality experiments for a variety of applications with multiple critical assemblies\. The Comet\[[27](https://arxiv.org/html/2606.04033#bib.bib5)\]and Planet\[[24](https://arxiv.org/html/2606.04033#bib.bib6)\]vertical\-lift assemblies at NCERC have been used to conduct a variety of criticality experiments\. Both assemblies are general\-purpose, heavy\-duty, vertical\-lift assemblies made of two vertically stacked platforms, each containing a subcritical fissile configuration\. Reactivity is controlled by moving the platforms towards or away from each other\. Historically, the Zero Power Reactor \(ZPR\) and Zero Power Plutonium/Physics Reactor \(ZPPR\) facilities were a series of HST critical experiment facilities operated at Argonne National Laboratory between 1955 and 1990\[[13](https://arxiv.org/html/2606.04033#bib.bib39)\]\. The ZPR/ZPPR approach to critical experiment design allowed for a range of experiments to be conducted within a single assembly by using a grid of interchangeable material tubes\. Interest in advanced reactor designs has demonstrated the need for a new critical experiment machine capable of supporting advanced fuels, geometries, and material compositions\. The System Physics Advanced Reactor Critical facility \(SPARC\) project at Idaho National Laboratory is a proposed critical experiment facility that will support this need\. Unlike the vertical\-lift mechanisms of Comet and Planet, SPARC will be a Horizontal Split\-Table \(HST\) machine, and will be capable of supporting a critical assembly with a mass of 24,000 kg and a length of around 2\.0 m in each dimension\. SPARC will be crucial to the validation and deployment of advanced reactor designs and fuel concepts\[[35](https://arxiv.org/html/2606.04033#bib.bib18)\]\.
To yield useful results, the neutronic properties of a criticality experiment must be similar to those of the full\-scale system\. Specifically, the uncertainties associated with each cross section involved in the full\-scale system should be targeted by the criticality experiment to encourage any calculation biases to manifest similarly in both systems\. This similarity can be measured by the correlation coefficientckc\_\{k\}described in Section[2](https://arxiv.org/html/2606.04033#S2)\. If a criticality experiment and full\-scale system have a highckc\_\{k\}, observed neutronic properties of the experiment can be safely transferred to the full\-scale system\. In the context of experiment design, Maldonado et al\.\[[16](https://arxiv.org/html/2606.04033#bib.bib8)\]rely onckc\_\{k\}values to guide the design of criticality experiments for the conceptual Snowflake microreactor\. As the correlation coefficient between an experiment design and Snowflake approaches unity, biases in calculatedkeffk\_\{\\text\{eff\}\}due to cross\-sectional uncertainty associated with the experiment approach those associated with the commercial technology\. This work demonstrates howckc\_\{k\}can be used as an objective similarity metric for experiment design\. Additionally, Walters & Roskoff\[[31](https://arxiv.org/html/2606.04033#bib.bib42)\]used both sensitivities ofkeffk\_\{\\text\{eff\}\}and reactivity to cross sections in order to evaluate the similarity between a 30 MWt graphite\-moderated test reactor and a criticality experiment\.
Many advanced reactor designs require high\-assay low\-enriched uranium \(HALEU\) as fuel, which will require safe transportation solutions\. The TN Americas TN\-LC cask\[[4](https://arxiv.org/html/2606.04033#bib.bib10)\]is a representative transportation package that could be plausibly adapted to support HALEU, avoiding the need to initiate a new design and regulatory approval process\. To validate TN\-LC for HALEU transportation, rigorous neutronic analysis must be carried out to ensure subcriticality in all conditions including situations such as water leakage\. To conduct neutronic analysis of the TN\-LC, Eidelpes et al\.\[[4](https://arxiv.org/html/2606.04033#bib.bib10)\]use correlation coefficients to identify benchmark critical experiments with highckc\_\{k\}\. Hall et al\.\[[5](https://arxiv.org/html/2606.04033#bib.bib9)\]consider two systems to be similar if they have ack≥0\.9c\_\{k\}\\geq 0\.9, while Eidelpes et al\. require that akeffk\_\{\\text\{eff\}\}penalty be applied to subcritical margins withoutck\>0\.8c\_\{k\}\>0\.8\. 130 experiments were found withck≥0\.8c\_\{k\}\\geq 0\.8for Case 1, but none withck≥0\.9c\_\{k\}\\geq 0\.9, and none withck≥0\.8c\_\{k\}\\geq 0\.8for Case 2\. This demonstrates the need for new experiments to be designed\.
Machine learning techniques have been explored to optimize criticality experiment design\. One such technique, Constrained Bayesian Optimization \(CBO\), strategically selects experimental configurations to evaluate while maintaining a probabilistic model of the objective function, minimizing expensive evaluations while enforcing the physical constraints of the design space\. Branco\-Ketcher et al\.\[[2](https://arxiv.org/html/2606.04033#bib.bib11)\]apply CBO to the design of criticality experiments for TerraPower’s Molten Chloride Reactor Experiment \(MCRE\) to simultaneously maximizeckc\_\{k\}and sensitivity to reactions of interest while remaining within the physical limits of the Comet assembly\. Targeting the reduction of intermediate\-energy Pu\-239 nuclear data, Kleedtke et al\.\[[9](https://arxiv.org/html/2606.04033#bib.bib26)\]used particle swarm optimization to adjust an experimental configuration such that the neutron flux and sensitivity ofkeffk\_\{\\text\{eff\}\}to the total fission cross section are both maximized in the 1 keV to 600 keV energy range\. Given the extreme potential computational cost of Monte Carlo based methods for evaluating each candidate model during the optimization procedure, a simplified model was used with limited sampling parameters\. An alternative approach to address this computational cost concern is to use a surrogate model capable of rapidly approximating the objective function\.
Surrogate models to replace the high computational cost of full\-order model evaluations within optimizers have seen frequent use in the nuclear engineering field\. Whyte and Parks\[[32](https://arxiv.org/html/2606.04033#bib.bib19)\]evaluated two surrogate models to be used in optimizing the design of a 36 assembly layout in a miniaturized pressurized water reactor\. These two deep learning based surrogate models were then embedded in a NSGA\-II\[[3](https://arxiv.org/html/2606.04033#bib.bib28)\]multiobjective optimizer\. Separately, a two part study\[[25](https://arxiv.org/html/2606.04033#bib.bib29),[26](https://arxiv.org/html/2606.04033#bib.bib30)\]optimized levelized cost of electricity for a heat pipe microreactor while maintaining constraints on safety\-relevant quantities\. Advanced reinforcement learning\-based multiobjective optimization methods were then applied to treat these constraints in a more comprehensive manner\. In both cases, neural networks were used as surrogate models\. One important practice used was to reintroduce the full\-order model at the conclusion of the optimization routine to evaluate the accuracy of the surrogate models specifically on the discovered optimal solution\. In general, this practice is fairly common\[[21](https://arxiv.org/html/2606.04033#bib.bib31),[23](https://arxiv.org/html/2606.04033#bib.bib32),[17](https://arxiv.org/html/2606.04033#bib.bib33)\]and acts to mitigate some of the drawbacks arising from inaccuracies of any surrogate model\.
In reference to the methods behind the surrogate models themselves, neutron transport calculations can be nonlinear and therefore difficult to accurately model with simple surrogate methods\. However, advances in deep learning\[[11](https://arxiv.org/html/2606.04033#bib.bib20)\]over recent years have produced architectures capable of approximating complex functions ranging from protein folding\[[36](https://arxiv.org/html/2606.04033#bib.bib34)\]to fluid dynamics\[[10](https://arxiv.org/html/2606.04033#bib.bib35)\], making deep neural networks a compelling candidate for neutronic surrogate modeling\. A number of deep learning techniques can be applied to build strong neutronic surrogate models\. Convolutional layers, for which Convolutional Neural Networks \(CNNs\) are named, are highly adept at extracting spatial and geometric information from multidimensional inputs, which has led to their ubiquitous presence in image\-processing models over the past few decades\[[12](https://arxiv.org/html/2606.04033#bib.bib22)\]\. Attention, the primary architectural advancement behind the emergence of Large Language Models, learns the relationships between all components of an input simultaneously to capture long\-ranged effects\[[30](https://arxiv.org/html/2606.04033#bib.bib21)\]\. When combined in a single neural network, these components present a powerful and promising method for capturing the neutronic behavior of complex geometries in a fraction of the time of a full Monte Carlo transport simulation\.
Deep neural networks are differentiable objects, meaning that they can be used with gradient\-based optimization methods\[[18](https://arxiv.org/html/2606.04033#bib.bib23)\]\. Gradient\-based optimization methods have empirically demonstrated strong optimization capabilities throughout machine learning, and scale well with the dimensionality of the optimization space\[[34](https://arxiv.org/html/2606.04033#bib.bib27)\]\. Backpropagation\[[8](https://arxiv.org/html/2606.04033#bib.bib24)\]can be used to compute the gradient of the objective with respect to each component of the input\. The gradient is then used to update the input in the direction that maximally increases or decreases \(depending on the formulation of the problem\) the objective\. In the context of nuclear system design, Pevey et al\.\[[20](https://arxiv.org/html/2606.04033#bib.bib25)\]use gradient methods to optimize the neutronic properties of a simplified core geometry\. If a differentiable surrogate for neutronic sensitivity profiles can be constructed, gradient\-based optimization can be applied to critical experiment design\.
This work exploits the differentiability of deep neural network surrogates to perform gradient\-based optimization ofckc\_\{k\}over the full combinatorial space of material assignments in a voxelized experimental geometry\. A model of the target technology is first built in OpenMC and used to calculate target sensitivity profiles for various scenarios of interest\. A dataset of procedurally generated material configurations is then used to train a surrogate neural network to predict sensitivities for arbitrary geometries\. The surrogate is comprised of a novel, heterogeneous neural network architecture designed to promote an inductive bias of the underlying transport physics during training\. Once trained, the surrogate is used with a gradient\-based optimization method to generate experiment geometries that maximizeckc\_\{k\}for any input sensitivity profile generated by OpenMC\.
## 2Methodology
### 2\.1Application Scenario
The TN Americas TN\-LC transportation cask\[[4](https://arxiv.org/html/2606.04033#bib.bib10)\]was originally developed for transporting moderate to high\-burnup spent nuclear fuel \(SNF\) assemblies, and is now being considered for future transportation of HALEU needed by advanced reactors\. If selected for this purpose, criticality safety evaluations would be required to show that it remains subcritical in all possible scenarios, including events of water intrusion into the cask and/or internalUO2\\text\{UO\}\_\{2\}canisters\. Data from criticality experiments with high neutronic similarity to the TN\-LC can be used to aid this analysis\. The neutronic similarity between an experiment and the target system can be quantified by the correlation coefficientckc\_\{k\}, which is the fraction of nuclear data\-induced uncertainty that the systems share\[[16](https://arxiv.org/html/2606.04033#bib.bib8)\]\. A sensitivity profile is constructed by combining the sensitivity vectorss∈ℝGs\\in\\mathbb\{R\}^\{G\}for the reactions of interest into a single vectorS∈ℝRGS\\in\\mathbb\{R\}^\{RG\}forRRreactions andGGdiscrete energy groups\. The correlation coefficient is defined as
ck=STARGET⊤CXXS\(STARGET⊤CXXSTARGET\)\(S⊤CXXS\)c\_\{k\}=\\frac\{S\_\{\\mathrm\{TARGET\}\}^\{\\top\}C\_\{XX\}S\}\{\\sqrt\{\\left\(S\_\{\\mathrm\{TARGET\}\}^\{\\top\}C\_\{XX\}S\_\{\\mathrm\{TARGET\}\}\\right\)\\left\(S^\{\\top\}C\_\{XX\}S\\right\)\}\}\(1\)whereSTARGETS\_\{TARGET\}is the sensitivity profile of the target,SSis the sensitivity profile of the experiment, andCXXC\_\{XX\}is a matrix containing the relative covariances of the nuclear data\.
In their analysis of a different candidate transportation cask, Hall et al\.\[[5](https://arxiv.org/html/2606.04033#bib.bib9)\]consider two systems with ack≥0\.9c\_\{k\}\\geq 0\.9to be “neutronically similar\.” In their evaluation of the TN\-LC, Eidelpes et al\.\[[4](https://arxiv.org/html/2606.04033#bib.bib10)\]use the TSUNAMI code to find existing critical experiment benchmarks with highckc\_\{k\}for two limiting configurations of the cask\. In both cases, the fuel canisters are flooded while the cask cavity remains dry\. In Case 1, the cask is modeled with a low\-density foam in the cask cavity and lowB10\{\}^\{10\}\\text\{B\}areal density in the canister poison coating\. In Case 2, the cask is modeled with a medium density foam and a highB10\{\}^\{10\}\\text\{B\}areal density\. While TSUNAMI identified 130 critical experiments withck≥0\.8c\_\{k\}\\geq 0\.8for Case 1, it identified no experiments withck≥0\.8c\_\{k\}\\geq 0\.8for Case 2, and none withck≥0\.9c\_\{k\}\\geq 0\.9for either Case 1 or Case 2\. The lack of highly correlated experiments for the TN\-LC highlights the need for new critical experiments to be designed, especially given the capability of the SPARC facility to accommodate experiments involving HALEU\.
ConfigurationDescriptionDry CaskNo water intrusion; no foam in cavity; low10B areal density\.Case 1Flooded canisters; low\-density foam in cavity; low10B areal density\.Case 2Flooded canisters; medium\-density foam in cavity; high10B areal density\.Table 1:Brief descriptions of the three TN\-LC configurations considered in this work\.Figure 1:Renderings of the simplified TN\-LC model in OpenMC showing the arrangement of HALEU canisters \(red\) within the cask\.A model of the TN\-LC was built in both OpenMC and Serpent\[[14](https://arxiv.org/html/2606.04033#bib.bib41)\]following a simplified design presented in Eidelpes et al\. and used to calculate sensitivity profiles for Case 1, Case 2, and a “dry” configuration of the cask without any water intrusion\. Table[1](https://arxiv.org/html/2606.04033#S2.T1)is provided as a concise summary of these cases\.
Serpent does not support reflective cylindrical boundaries, so a cuboid reflective boundary was placed around each cask model, unlike the cylindrical reflective boundary used by Eidelpes et al\[[4](https://arxiv.org/html/2606.04033#bib.bib10)\]\. As discussed in Section[3\.2](https://arxiv.org/html/2606.04033#S3.SS2), model agreement between the codes was necessary to benchmark the sensitivity calculation implementation in OpenMC against that of Serpent, which required the model to adhere to the limitations of the Serpent boundary conditions\.
The simplified model consists of an outer cylinder of 304 stainless steel, a layer of lead inside the outer steel cylinder, and an inner steel cylinder inside the lead layer\. There are 56 steel fuel canisters within the inner cylinder, each coated in a thin layer ofAlBneutron poison and containing 20% enrichedUO2\\text\{UO\}\_\{2\}\. The external neutron shield resin, aluminum shield boxes, and impact limiters are not modeled\. Cross\-sectional cutaways of the model are shown in Figure[1](https://arxiv.org/html/2606.04033#S2.F1)\.
Eidelpes et al\. present numerous dry cask configurations with different neutron poison arrangements\. The configuration modeled in this work had the highestkeffk\_\{\\text\{eff\}\}of all the options, and was thus the most relevant from a criticality safety perspective\. In Case 1 and Case 2, the void within the cask is filled with polyurethane foam, which is not present in the dry cask model\. Cask water intrusion is modeled via a homogeneous mixture ofUO2\\text\{UO\}\_\{2\}andH2O\\text\{H\}\_\{2\}\\text\{O\}determined by the desired volume ratio of each case\. The reactions chosen for the sensitivity profiles wereU235\{\}^\{235\}\\text\{U\}fission,C12\{\}^\{12\}\\text\{C\}elastic scattering,H1\{\}^\{1\}\\text\{H\}elastic scattering, andNi58\{\}^\{58\}\\text\{Ni\}absorption\. As shown in Section[3](https://arxiv.org/html/2606.04033#S3), sensitivity to theU235\{\}^\{235\}\\text\{U\}fission cross\-section is the primary contributor tokeffk\_\{\\text\{eff\}\}uncertainty\. TheH1\{\}^\{1\}\\text\{H\},C12\{\}^\{12\}\\text\{C\}, andNi58\{\}^\{58\}\\text\{Ni\}sensitivities are included as representative sensitivities for the other materials present in the model\. Modeling details and sensitivities for the different TN\-LC configurations are shown in Section[3\.1](https://arxiv.org/html/2606.04033#S3.SS1)\. To enable proper calculation of covariance\-weightedckc\_\{k\}scores, the uncertainty covariances and correlations for the four reactions were extracted from the SCALE\-2\.3\.2\[[33](https://arxiv.org/html/2606.04033#bib.bib40)\]code using the CADILLAC and COGNAC AMPX utility modules\. The correlation matrix for uncertainties in the235U fission cross section is shown in Figure[2](https://arxiv.org/html/2606.04033#S2.F2), using the SCALE\-252 energy group structure\.
Figure 2:Correlation matrix for uncertainties in theU235\{\}^\{235\}\\text\{U\}fission cross\-section using the SCALE\-252 energy group structure\.
### 2\.2Sensitivity Profile Calculation in OpenMC
An eigenvalue sensitivity coefficientSσ\(r,E\)S\_\{\\sigma\(r,E\)\}is a scalar quantity that describes the fractional change inkeffk\_\{\\text\{eff\}\}resulting from perturbations or uncertainties in nuclear dataσ\(r,E\)\\sigma\(r,E\)\. It is expressed as
Sσ\(r,E\)=δkeff/keffδσ\(r,E\)/σ\(r,E\)S\_\{\\sigma\(r,E\)\}=\\frac\{\\delta k\_\{\\text\{eff\}\}/k\_\{\\text\{eff\}\}\}\{\\delta\\sigma\(r,E\)/\\sigma\(r,E\)\}\(2\)wherekeffk\_\{\\text\{eff\}\}is the effective multiplication factor,σ\(r,E\)\\sigma\(r,E\)is the cross section for reactionrrat energyEE, andδ\\deltais the differential perturbation operation\. The IFP method tracks neutrons across successive generations to weight their importance by the fission contribution of their progeny, and is used in this work to calculate sensitivity vectors within the SCALE\-252 energy group structure\.
Sensitivities were calculated using the Iterated Fission Probability \(IFP\) method\[[19](https://arxiv.org/html/2606.04033#bib.bib14)\]on a non\-production branch of OpenMC\. The OpenMC sensitivity calculation implementation was validated against the Serpent\[[14](https://arxiv.org/html/2606.04033#bib.bib41)\]Monte Carlo code to ensure accuracy, as discussed in Section[3\.1](https://arxiv.org/html/2606.04033#S3.SS1)\. The codes were found to be highly consistent in their calculated sensitivities\.
### 2\.3Surrogate Modeling
A surrogate model is a fast approximation of an expensive simulation that comes with several advantages at the cost of accuracy\. In this work, a deep neural network was trained as a surrogate model of OpenMC’s sensitivity calculation capabilities\. Specifically, the neural network predicts the sensitivity profileS^\\hat\{S\}for a given experimental geometryxx, wherexxis encoded as an input tensorx∈ℝ4×32×32x\\in\\mathbb\{R\}^\{4\\times 32\\times 32\}with each entry being the one\-hot material assignment of that cell\. The predictedS^\\hat\{S\}is then used to computec^k\\hat\{c\}\_\{k\}for the criticality experiment represented byxxwith the target technology\. The model is trained to predictS^\\hat\{S\}rather thanc^k\\hat\{c\}\_\{k\}directly in order to finely evaluate its performance across different reaction types and energy groups as shown in Section[3\.2](https://arxiv.org/html/2606.04033#S3.SS2)\. This approach also limits the complexity of what the model must learn to the most computationally expensive aspect of the problem, since the step of calculatingc^k\\hat\{c\}\_\{k\}fromS^\\hat\{S\}is straightforward\. The^\\hat\{\}notation is used in this work to denote a quantity predicted by a surrogate model rather than calculated via Monte Carlo or other numerical or analytical method\.
Using a surrogate neural network for sensitivity calculation has several advantages\. It is extremely fast to evaluate, requiring only a single forward pass on a GPU to calculate the full sensitivity profile for a geometry\. It is also differentiable, meaning that the gradient of a predicted quantity with respect to the input can be obtained\. In the context of experiment design, this means that the gradient of the predictedc^k\\hat\{c\}\_\{k\}with respect to input geometryxxcan be obtained, andxxcan then be directly perturbed using gradient descent in order to increasec^k\\hat\{c\}\_\{k\}\. This approach is intractable using Monte Carlo sensitivity calculations\.
Several methods are used to evaluate the accuracy of the model\. For a single geometry, the Mean\-Squared Error \(MSE\) is calculated as the average squared difference between each scalar sensitivity coefficient in the predicted sensitivity profileS^\\hat\{S\}and the OpenMC\-calculated profileSS, of which there areRRreactions andGGenergy groups\. The MSE is extended for a set ofNNgeometries such that \(definingS^r,g\(i\)\\hat\{S\}^\{\(i\)\}\_\{r,g\}andSr,g\(i\)S^\{\(i\)\}\_\{r,g\}as the predicted and evaluated sensitivity coefficients of reactionrrat energy groupggfor geometryii\),
MSE=1N∑i=1N1RG∑r=1R∑g=1G\(S^r,g\(i\)−Sr,g\(i\)\)2\.\\text\{MSE\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\frac\{1\}\{RG\}\\sum\_\{r=1\}^\{R\}\\sum\_\{g=1\}^\{G\}\(\\hat\{S\}^\{\(i\)\}\_\{r,g\}\-S^\{\(i\)\}\_\{r,g\}\)^\{2\}\.\(3\)In the context of design optimization, it is also useful to evaluate the Mean Absolute Error \(MAE\), calculated forNNgeometries as
MAE=1N∑i=1N1RG∑r=1R∑g=1G\|S^r,g\(i\)−Sr,g\(i\)\|\.\\text\{MAE\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\frac\{1\}\{RG\}\\sum\_\{r=1\}^\{R\}\\sum\_\{g=1\}^\{G\}\|\\hat\{S\}^\{\(i\)\}\_\{r,g\}\-S^\{\(i\)\}\_\{r,g\}\|\.\(4\)
### 2\.4Data Structure & Generation
For this work, two\-dimensional experiment geometries were procedurally generated and used to build a dataset of OpenMC\-calculated sensitivity profiles\. Each geometry is a cube of material compatible with the rough 200 cm length scale proposed of the SPARC facility\[[35](https://arxiv.org/html/2606.04033#bib.bib18)\]\. The cube face is discretized into a 64x64 pixel grid \(exploiting quarter\-symmetry to reduce this to a 32x32 unique region\), where each3\.125cm23\.125\\,\\text\{cm\}^\{2\}pixel extends uniformly through the full depth of the cube to form a set of square prism voxels\. A palette of four materials was chosen representative of the target technology\. The palette included 20% enrichedUO2\\text\{UO\}\_\{2\}, stainless steel, borated aluminum, and polyethylene\.
Constraining critical experiment designs to a 2D grid of materials extruded through a cubic geometry has historical precedent in the field\. This approach was used by the ZPR\[[13](https://arxiv.org/html/2606.04033#bib.bib39)\]family of critical experiment facilities for 35 years to support highly flexible experiment designs\. Each facility contained an HST assembly comprised of a grid of nominally square0\.0400\.040in thick stainless steel tubes\. Grid dimensions ranged from 31x31 in the ZPR\-3 machine to 77x77 in the expanded ZPPR\-6 machine\. The tubes would be filled with easily interchangeable plates of various materials to represent the composition of materials in the target reactor design\.
Figure 3:An example critical experiment geometry\. The 64x64 material grid is extruded to form a cube with side lengths of200cm200\\,\\text\{cm\}\.To provide some structure to the spatial distribution of the explored experiment configurations, a cellular automaton algorithm was used to procedurally generate experiment geometries that include varying levels of structure without imposing human biases of what that structure should look like\. The cellular automaton process is shown for an example geometry in Figure[4](https://arxiv.org/html/2606.04033#S2.F4)\.
Figure 4:Procedural geometry generation over 40 cellular automaton steps\.In this context, a cellular automaton is a process that describes a discrete grid of cells that update over time according to the state of their neighbors\. The geometry starts as a grid of cells randomly initiated by a unique seed\. The state of each cell is its material assignment, which is updated at each step by sampling from a distribution of its neighboring material counts, weighted by a factorβ\\beta\. To encourage diverse results, the number of update stepstt, and the weighting of neighboring materialsβ\\betaare randomly chosen for each geometry that is generated\. Geometries generated with a higherttfeature larger, more homogeneous regions of material, since at each step it becomes more likely that a cell changes to the material of its neighbor\. Geometries generated with a higherβ\\betafeature sharper boundaries between regions, since the sampled probability distribution is much sharper\. By varying these quantities throughout the dataset a wide distribution of geometries can be generated\. The full cellular automaton algorithm is given in Algorithm[1](https://arxiv.org/html/2606.04033#alg1), and a comparison of geometries generated with differentttandβ\\betais shown in Figure[5](https://arxiv.org/html/2606.04033#S2.F5)\. It should be noted that the proposed experiment designs are not subject to this process, only the training data used to create the surrogate model\.
Algorithm 1Dataset Geometry Generation1:
x∼Uniform\{0,…,4\}32×32x\\sim\\text\{Uniform\}\\\{0,\\ldots,4\\\}^\{32\\times 32\}
2:for
1,…,t1,\\ldots,tdo
3:foreach cell
\(i,j\)\(i,j\)do
4:Update neighbor\_counts
5:
x′\[i,j\]∼softmax\(β⋅neighbor\_counts\)x^\{\\prime\}\[i,j\]\\sim\\text\{softmax\}\(\\beta\\cdot\\text\{neighbor\\\_counts\}\)
6:endfor
7:
x←x′x\\leftarrow x^\{\\prime\}
8:endfor
9:Reflect
xxto the full
64×6464\\times 64grid
10:return
xx
Figure 5:Geometries generated from the same initial seed with different combinations ofttandβ\\beta\.
### 2\.5Network Architecture
The deep neural network used as a surrogate model is not comprised of simple multilayer perceptrons\. Instead, it uses a heterogeneous architecture to more effectively encode the relationships between the experimental geometry and the sensitivities\. Model architectural design was guided by both the spatial nature of the input and certain aspects of the underlying transport physics\.
A geometry encodingxxis passed through a U\-Net\[[22](https://arxiv.org/html/2606.04033#bib.bib36)\]to produce a spatial feature map, which is then processed by a multigroup attention pooling layer \(introduced in this work\) that weights features by their spatial importance for each discretized energy group\. The resulting group descriptors are passed through a 1D convolutional regressor operating over the full energy group structure to produce a predicted sensitivity vectors^\\hat\{s\}for each reaction\. The sensitivity vectors are combined to create the final sensitivity predicted sensitivity profileS^\\hat\{S\}\.
Figure 6:Surrogate model architecture\. The experiment geometry is passed through a shared U\-Net to extract spatial features\. The features are weighted by isolated attention pools for each energy group and reaction, regressed into the scalar sensitivity coefficients, and combined to create the final predicted sensitivity profileS^∈ℝRG\\hat\{S\}\\in\\mathbb\{R\}^\{RG\}forRRunique reactions andGGdiscrete energy groups\.#### 2\.5\.1U\-Net Encoder
The U\-Net is an encoder\-decoder, fully convolutional network architecture first introduced in 2015 for biomedical image segmentation\[[22](https://arxiv.org/html/2606.04033#bib.bib36)\]\. The defining feature of the U\-Net is its use of skip\-connections, which transfer information from earlier, higher\-resolution feature maps directly to the decoder, allowing the network to recover fine spatial detail that would otherwise be lost during downsampling\. The encoder consists of four successive modules comprising a convolution, group normalization, ReLU activation, and residual block, followed by two max\-pooling layers between each stage\. The residual blocks consist of two group\-normalized convolutions with ReLU activations\. The decoder mirrors this structure with transposed convolutions for upsampling, concatenated with the corresponding encoder skip connection\. The output of the decoder is a feature mapΦ∈ℝC×H×W\\Phi\\in\\mathbb\{R\}^\{C\\times H\\times W\}whereCCis the number of channels in the first encoder stage, andH×WH\\times Wis the original input resolution\.
#### 2\.5\.2Multigroup Attention Pooling
Letϕr∈ℝC\\phi\_\{r\}\\in\\mathbb\{R\}^\{C\}denote the feature vector extracted by the U\-Net encoder at spatial positionr∈\{1,…,HW\}r\\in\\\{1,\\dots,HW\\\}\. For each energy groupg∈\{1,…,G\}g\\in\\\{1,\\dots,G\\\}, a learned attention vectorwg∈ℝCw\_\{g\}\\in\\mathbb\{R\}^\{C\}produces a scalar importance scorewg⊤ϕrw\_\{g\}^\{\\top\}\\phi\_\{r\}at every spatial location\. These scores are normalized via softmax to yieldag∈ℝHWa\_\{g\}\\in\\mathbb\{R\}^\{HW\}, a spatial attention map over the geometry,
ag,r=exp\(wg⊤ϕr\)∑r′=1HWexp\(wg⊤ϕr′\)\.a\_\{g,r\}=\\frac\{\\exp\(w\_\{g\}^\{\\top\}\\phi\_\{r\}\)\}\{\\sum\_\{r^\{\\prime\}=1\}^\{HW\}\\exp\(w\_\{g\}^\{\\top\}\\phi\_\{r^\{\\prime\}\}\)\}\.\(5\)Eachaga\_\{g\}records the regions of the geometry that are most informative for energy groupgg\. Because neutrons with different energies interact with geometry at different characteristic length scales, independently learned attention maps per group provide a physically motivated inductive bias for the model to learn\. The group descriptorzgz\_\{g\}is then taken as the attention\-weighted sum of the spatial feature vectors,
zg=∑r=1HWag,rϕr∈ℝC\.z\_\{g\}=\\sum\_\{r=1\}^\{HW\}a\_\{g,r\}\\,\\phi\_\{r\}\\;\\in\\;\\mathbb\{R\}^\{C\}\.\(6\)Stacking over allGGgroups gives the descriptor matrixZ∈ℝG×CZ\\in\\mathbb\{R\}^\{G\\times C\}\.
#### 2\.5\.3Group Regressor
The group descriptorsZ∈ℝG×CZ\\in\\mathbb\{R\}^\{G\\times C\}are treated as a 1D sequence of lengthGGand processed by a convolutional residual network that maps each group to a final scalar sensitivity coefficient\. The residual network is formed by six successive residual blocks, each comprised of a group normalization, ReLU activation, and 1D convolution\. The output of the residual network is projected through a final 1D convolution into the predicted sensitivity vectors^∈ℝG\\hat\{s\}\\in\\mathbb\{R\}^\{G\}for each reaction\.
When trained onR\>1R\>1reactions, the U\-Net encoder weights are shared across all reactions, but each reaction retains independent attention pool and regressor weights\. TheRRsensitivity vectors are flattened to form the full sensitivity profileS^∈ℝRG\\hat\{S\}\\in\\mathbb\{R\}^\{RG\}\.
#### 2\.5\.4Training
A training dataset was generated using the cellular automata algorithm described in Section[2\.4](https://arxiv.org/html/2606.04033#S2.SS4)with parameterst∈\[3,40\]t\\in\[3,40\]andβ∈\[0\.75,4\]\\beta\\in\[0\.75,4\]\. Each profile containedkeffk\_\{\\text\{eff\}\}sensitivities to the cross\-sectional uncertainties ofU235\{\}^\{235\}\\text\{U\}fission,H1\{\}^\{1\}\\text\{H\}elastic scattering,C12\{\}^\{12\}\\text\{C\}elastic scattering, andNi58\{\}^\{58\}\\text\{Ni\}absorption\. Sensitivities were calculated via the IFP method in OpenMC using 25000 particles, 20 inactive generations, 250 active generations, and 3 latent IFP generations\. Sensitivities to different reactions and energies span a wide range of magnitudes\. To encourage all sensitivity coefficients to be treated with equal importance during training, the model was trained on z\-scores of normalized sensitivities using a standard mean squared error \(MSE\) loss function, with the mean and standard deviation of each reaction saved in the model and recovered during inference\.
The AdamW\[[15](https://arxiv.org/html/2606.04033#bib.bib38)\]optimizer is used to train the model end\-to\-end by minimizing the mean squared error between the predicted and target sensitivity coefficients, normalized to𝒩\(0,1\)\\mathcal\{N\}\(0,1\), such that the training loss is
ℒ=1N∑i=1N1RG∑r=1R∑g=1G\(S^r,g\(i\)−Sr,g\(i\)\)2\\mathcal\{L\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\frac\{1\}\{RG\}\\sum\_\{r=1\}^\{R\}\\sum\_\{g=1\}^\{G\}\(\\hat\{S\}^\{\(i\)\}\_\{r,g\}\-S^\{\(i\)\}\_\{r,g\}\)^\{2\}\(7\)whereS^r,g\\hat\{S\}\_\{r,g\}andSr,gS\_\{r,g\}are the predicted and normalized target sensitivity coefficients respectively for reactionrrand energy groupgg\.
Since the focus of this work is the method and application of a novel model and task, rather than improvement upon an existing benchmark or neural network architecture, no structured hyperparameter sweep was run during training\. To validate the inclusion of the attention pooling layer, versions of the model were trained with different choices of pooling layer and compared on a set of benchmark geometries\. Results are shown in Table[3](https://arxiv.org/html/2606.04033#S3.T3)and Figure[6](https://arxiv.org/html/2606.04033#S2.F6)\. As discussed further in Section[3\.2](https://arxiv.org/html/2606.04033#S3.SS2), multigroup attention pooling demonstrated the lowest error of the benchmarked options under the chosen training hyperparameters, justifying its inclusion in the model\. This result indicates the potential value of custom neural network architectures for modeling neutronics, which is explored further in Section[3\.3](https://arxiv.org/html/2606.04033#S3.SS3)\.
The full architecture details and training hyperparameters for the surrogate model are given in[A](https://arxiv.org/html/2606.04033#A1)\.
### 2\.6Gradient Based Optimization
The design of an experiment geometry to maximizeckc\_\{k\}with a target technology can be mathematically framed as a single\-objective optimization problem\. In general, a single\-objective optimization problem can be expressed as
minx∈ℝn\\displaystyle\\min\_\{x\\in\\mathbb\{R\}^\{n\}\}ℒ\(x\)\\displaystyle\\mathcal\{L\}\(x\)\(8\)subject togi\(x\)≤0,i=1,2,…,ng,\\displaystyle g\_\{i\}\(x\)\\leq 0,\\quad i=1,2,\\dots,n\_\{g\},hj\(x\)=0,j=1,2,…,nh\.\\displaystyle h\_\{j\}\(x\)=0,\\quad j=1,2,\\dots,n\_\{h\}\.
Letfθf\_\{\\theta\}be the surrogate neural network, such thatS^=fθ\(x\)\\hat\{S\}=f\_\{\\theta\}\(x\)is the predicted sensitivity profile for input geometryxx\. For this application, the objective functionℒ\(x\)=−c^k\\mathcal\{L\}\(x\)=\-\\hat\{c\}\_\{k\}such that the optimization goal is to produce an experiment geometryxxthat maximizesckc\_\{k\}with the target technology of interest\. Thec^k\\hat\{c\}\_\{k\}notation is introduced to differentiate between the neural\-network predictedc^k\\hat\{c\}\_\{k\}and theckc\_\{k\}calculated via Monte Carlo\. The application is formulated such that there are no applicable constraints onxx, as any arrangement of materials in the grid produces a valid sensitivity profile\. It can therefore be thought of thatng=nh=0n\_\{g\}=n\_\{h\}=0\. However, an important shortcoming of this study in its current form is the lack of consideration for the actualkeffk\_\{\\text\{eff\}\}of the criticality experiment\. The constraint thatkeff\>1k\_\{\\text\{eff\}\}\>1for compatibility with the HST experiment approach should realistically be accounted for\.
The surrogatefθf\_\{\\theta\}is differentiable with respect toxx, meaning that the gradient ofc^k\\hat\{c\}\_\{k\}with respect to the input geometry can be computed by backpropagation through the model\. Since the candidate geometryxxis encoded as one\-hot material assignments over each spatial position in the grid, it is not immediately differentiable\. To allow for gradient\-based optimization, the problem is relaxed into optimization over continuous material logitsℓ∈ℝ4×32×32\\ell\\in\\mathbb\{R\}^\{4\\times 32\\times 32\}, where the four material channels correspond to the four possible material assignments and the32×3232\\times 32spatial dimensions correspond to the quarter\-grid design domain\. Before evaluating the surrogate, the quarter\-grid logits are reflected across both spatial axes to form full\-grid logits on a64×6464\\times 64grid\.
GivenMMpossible materials, the probabilityppof assigning materialmmto positionrris obtained by taking a temperature\-scaled softmax over the material logits
pm,r=exp\(ℓm,r/τ\)∑m′=1Mexp\(ℓm′,r/τ\)p\_\{m,r\}=\\frac\{\\exp\(\\ell\_\{m,r\}/\\tau\)\}\{\\sum\_\{m^\{\\prime\}=1\}^\{M\}\\exp\(\\ell\_\{m^\{\\prime\},r\}/\\tau\)\}\(9\)whereτ\\tauis the softmax temperature\. For optimization stepst=0,…,T−1t=0,\\ldots,T\-1, the temperature is held fixed atτ=1\\tau=1for the first third of the optimization to encourage exploration, and then cosine\-annealed to a minimum value of0\.20\.2\. The temperature schedule is then
τt=\{1,t<⌊T/3⌋0\.2\+0\.4\(1\+cos\(π⋅αt\)\)t≥⌊T/3⌋\\tau\_\{t\}=\\begin\{cases\}1,&t<\\lfloor T/3\\rfloor\\\\\[4\.0pt\] 0\.2\+0\.4\(1\+\\cos\(\\pi\\cdot\\alpha\_\{t\}\)\)&t\\geq\\lfloor T/3\\rfloor\\end\{cases\}\(10\)with
αt=t−⌊T/3⌋max\(1,T−⌊T/3⌋−1\)\\alpha\_\{t\}=\\frac\{t\-\\lfloor T/3\\rfloor\}\{\\max\(1,T\-\\lfloor T/3\\rfloor\-1\)\}\(11\)such that asτ\\taudecreases, the relaxed probabilities become increasingly concentrated on the material with the largest logit, so that in the limit
limτ→0pm,r=1\[m=argmaxm′ℓm′,r\]\\lim\_\{\\tau\\to 0\}p\_\{m,r\}=\{1\}\\\!\\left\[m=\\arg\\max\_\{m^\{\\prime\}\}\\ell\_\{m^\{\\prime\},r\}\\right\]\(12\)where1\[⋅\]1\[\\cdot\]denotes an indicator function, such that the expression approaches 1 for the selected material and 0 for all other materials\. This allows the softmax to remain differentiable for any finiteτ\>0\\tau\>0, while becoming increasingly faithful to the original discrete material\-assignment problem asτ\\taudecreases\.
At each optimization step, the surrogate must be evaluated on a discrete geometry, so the forward\-pass geometryxxis formed by taking the most probable material at each spatial position
xm,r=1\[m=argmaxm′pm′,r\]\.x\_\{m,r\}=1\\\!\\left\[m=\\arg\\max\_\{m^\{\\prime\}\}p\_\{m^\{\\prime\},r\}\\right\]\.\(13\)The tensor passed to the surrogate is implemented using a straight\-through estimator \(STE\)\[[1](https://arxiv.org/html/2606.04033#bib.bib37)\]
xSTE=x\+p−stopgrad\(p\)\.x\_\{\\mathrm\{STE\}\}=x\+p\-\\mathrm\{stopgrad\}\(p\)\.\(14\)In the forward pass,xSTE=xx\_\{\\mathrm\{STE\}\}=x, so the surrogate is evaluated on a hard one\-hot geometry\. In the backward pass, the hard assignment is treated as the identity map with respect topp, allowing gradients to flow through the relaxed probabilities and then through the softmax to the logits, so that
∂xSTE∂p≈I\.\\frac\{\\partial x\_\{\\mathrm\{STE\}\}\}\{\\partial p\}\\approx I\.\(15\)The surrogate prediction is denormalized before computing the correlation coefficient\. Iffθ\(xSTE\)f\_\{\\theta\}\(x\_\{\\mathrm\{STE\}\}\)denotes the normalized surrogate output, then the predicted sensitivity is
S^=σ⊙fθ\(xSTE\)\+μ\\hat\{S\}=\\sigma\\odot f\_\{\\theta\}\(x\_\{\\mathrm\{STE\}\}\)\+\\mu\(16\)whereμ\\muandσ\\sigmaare the normalization statistics stored with the surrogate checkpoint\. The predicted correlation coefficientc^k\\hat\{c\}\_\{k\}is then computed with Equation[1](https://arxiv.org/html/2606.04033#S2.E1)\.
The optimization objective is to maximizec^k\\hat\{c\}\_\{k\}, such that the minimized loss is
ℒ=−c^k\\mathcal\{L\}=\-\\hat\{c\}\_\{k\}\(17\)Using the straight\-through estimator, the gradient of the loss with respect to the material logits is
∇ℓℒ=∂ℒ∂ℓ=−∂c^k∂S^∂S^∂fθ∂fθ∂xSTE∂xSTE∂p∂p∂ℓ\\nabla\_\{\\ell\}\\mathcal\{L\}=\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\ell\}=\-\\frac\{\\partial\\hat\{c\}\_\{k\}\}\{\\partial\\hat\{S\}\}\\frac\{\\partial\\hat\{S\}\}\{\\partial f\_\{\\theta\}\}\\frac\{\\partial f\_\{\\theta\}\}\{\\partial x\_\{\\mathrm\{STE\}\}\}\\frac\{\\partial x\_\{\\mathrm\{STE\}\}\}\{\\partial p\}\\frac\{\\partial p\}\{\\partial\\ell\}\(18\)Since the STE replaces∂xSTE/∂p\\partial x\_\{\\mathrm\{STE\}\}/\\partial pby the identity as shown in Equation[15](https://arxiv.org/html/2606.04033#S2.E15), this becomes
∇ℓℒ=∂ℒ∂ℓ≈−∂c^k∂S^∂S^∂fθ∂fθ∂xSTE∂p∂ℓ\\nabla\_\{\\ell\}\\mathcal\{L\}=\\frac\{\\partial\\mathcal\{L\}\}\{\\partial\\ell\}\\approx\-\\frac\{\\partial\\hat\{c\}\_\{k\}\}\{\\partial\\hat\{S\}\}\\frac\{\\partial\\hat\{S\}\}\{\\partial f\_\{\\theta\}\}\\frac\{\\partial f\_\{\\theta\}\}\{\\partial x\_\{\\mathrm\{STE\}\}\}\\frac\{\\partial p\}\{\\partial\\ell\}\(19\)This formulation is fully differentiable with respect to the logits and can be treated with gradient descent\. In practice, the logits are updated with the AdamW\[[15](https://arxiv.org/html/2606.04033#bib.bib38)\]optimizer, such that to update from stepttto stept\+1t\+1is
ℓt\+1=AdamW\(ℓt,∇ℓtℒt\)\\ell\_\{t\+1\}=\\mathrm\{AdamW\}\\left\(\\ell\_\{t\},\\nabla\_\{\\ell\_\{t\}\}\\mathcal\{L\}\_\{t\}\\right\)\(20\)Minimizing the loss in this way is equivalent to maximizingc^k\\hat\{c\}\_\{k\}through gradient ascent\.
Upon convergence, a discrete geometry is recovered by assigning each positionrrthe material with the highest logit valueℓm,r\\ell\_\{m,r\}\. The resulting32×3232\\times 32quarter\-grid material assignment is reflected to form a64×6464\\times 64full\-grid geometry\. This full\-grid geometry is evaluated in OpenMC to produce a finalckc\_\{k\}\. The full procedure is given in Algorithm[2](https://arxiv.org/html/2606.04033#alg2)for a single initialization\.
In practice, this procedure is applied to many independently initialized candidate geometries in parallel to maximize coverage of the loss landscape from many different starting positions\. A batch ofNNgeometries are initialized and optimized in parallel\. Upon convergence, the geometries with the topkkpredicted correlations are evaluated in OpenMC, and the geometry with the highest trueckc\_\{k\}is returned\.
Algorithm 2Gradient\-Based Geometry Optimization1:Initialize logits
ℓ∈ℝ4×32×32,ℓr∼𝒩\(0,1\)∀r\\ell\\in\\mathbb\{R\}^\{4\\times 32\\times 32\},\\,\\,\\,\\ell\_\{r\}\\sim\\mathcal\{N\}\(0,1\)\\ \\forall r
2:for
t=0,…,T−1t=0,\\ldots,T\-1do
3:Compute
τt\\tau\_\{t\}using the annealing schedule
4:Reflect
ℓt\\ell\_\{t\}to the full
64×6464\\times 64grid
5:
pt←softmax\(ℓt/τt\)p\_\{t\}\\leftarrow\\mathrm\{softmax\}\(\\ell\_\{t\}/\\tau\_\{t\}\)
6:
xt←onehot\(argmaxmpt,m\)x\_\{t\}\\leftarrow\\mathrm\{onehot\}\(\\arg\\max\_\{m\}p\_\{t,m\}\)
7:
xSTE,t←xt\+pt−stopgrad\(pt\)x\_\{\\mathrm\{STE\},t\}\\leftarrow x\_\{t\}\+p\_\{t\}\-\\mathrm\{stopgrad\}\(p\_\{t\}\)
8:
S^t←σ⊙fθ\(xSTE,t\)\+μ\\hat\{S\}\_\{t\}\\leftarrow\\sigma\\odot f\_\{\\theta\}\(x\_\{\\mathrm\{STE\},t\}\)\+\\mu
9:
c^k,t←STARGET⊤CXXS^t\(STARGET⊤CXXSTARGET\)\(S^t⊤CXXS^t\)\\hat\{c\}\_\{k,t\}\\leftarrow\\dfrac\{S\_\{\\mathrm\{TARGET\}\}^\{\\top\}C\_\{XX\}\\hat\{S\}\_\{t\}\}\{\\sqrt\{\\left\(S\_\{\\mathrm\{TARGET\}\}^\{\\top\}C\_\{XX\}S\_\{\\mathrm\{TARGET\}\}\\right\)\\left\(\\hat\{S\}\_\{t\}^\{\\top\}C\_\{XX\}\\hat\{S\}\_\{t\}\\right\)\}\}
10:
ℒt←−c^k,t\\mathcal\{L\}\_\{t\}\\leftarrow\-\\hat\{c\}\_\{k,t\}
11:
ℓt\+1←AdamW\(ℓt,∇ℓtℒt\)\\ell\_\{t\+1\}\\leftarrow\\mathrm\{AdamW\}\(\\ell\_\{t\},\\nabla\_\{\\ell\_\{t\}\}\\mathcal\{L\}\_\{t\}\)
12:endfor
13:
x←onehot\(argmaxmℓm\)x\\leftarrow\\mathrm\{onehot\}\(\\arg\\max\_\{m\}\\ell\_\{m\}\)
14:Reflect
xxto the full
64×6464\\times 64grid
15:return
xxand OpenMC\-computed
ckc\_\{k\}
## 3Results and Discussion
### 3\.1TN\-LC Configurations and Sensitivities
In order to verify the models used for sensitivity calculation, OpenMC and Serpent were used to calculatekeffk\_\{\\text\{eff\}\}for the Dry Cask, Case 1 \(flooded cask, low areal poison density\), and Case 2 \(flooded cask, high areal poison density\) TN\-LC configurations\. Parameters and resultingkeffk\_\{\\text\{eff\}\}values are shown in Table[2](https://arxiv.org/html/2606.04033#S3.T2)\.
Dry CaskCase 1Case 2Foam DensityN/A0\.12g/cm30\.12\\text\{ g/cm\}^\{3\}0\.22g/cm30\.22\\text\{ g/cm\}^\{3\}10B Areal Density0\.0007g/cm20\.0007\\text\{ g/cm\}^\{2\}0\.0007g/cm20\.0007\\text\{ g/cm\}^\{2\}0\.0702g/cm20\.0702\\text\{ g/cm\}^\{2\}H2O VF0\.00\.00\.80\.80\.80\.8OpenMCkeffk\_\{\\text\{eff\}\}1\.08976±16pcm1\.08976\\pm 16\\text\{ pcm\}1\.19240±21pcm1\.19240\\pm 21\\text\{ pcm\}0\.92146±22pcm0\.92146\\pm 22\\text\{ pcm\}Serpentkeffk\_\{\\text\{eff\}\}1\.08832±3pcm1\.08832\\pm 3\\text\{ pcm\}1\.19275±5pcm1\.19275\\pm 5\\text\{ pcm\}0\.92061±7pcm0\.92061\\pm 7\\text\{pcm\}Table 2:Comparison of TN\-LC model configurations\.Full sensitivity profiles for the three TN\-LC configurations are shown with Monte Carlo uncertainties in[B](https://arxiv.org/html/2606.04033#A2)\. To benchmark the OpenMC sensitivity calculation implementation, the three TN\-LC configurations were also modeled in the Serpent\[[14](https://arxiv.org/html/2606.04033#bib.bib41)\]Monte Carlo code\.
Sensitivities in both codes were calculated using 500,000 particles over 500 active generations and 50 inactive generations\. Figure[7](https://arxiv.org/html/2606.04033#S3.F7)shows the close agreement235U fission sensitivity profiles calculated by OpenMC and Serpent\. This consistency between two independent Monte Carlo codes provides confidence in the sensitivity profiles used throughout the remainder of this analysis\.
Figure 7:235U sensitivity vectors calculated in both Serpent and OpenMC for the three TN\-LC configurations\.
### 3\.2Surrogate Dataset and Training
Sensitivities were calculated in OpenMC for a dataset of 5,000 experiment geometries generated using the cellular automata algorithm described in Algorithm[1](https://arxiv.org/html/2606.04033#alg1), with averagekeffk\_\{\\text\{eff\}\}uncertainty of about40pcm40\\,\\text\{pcm\}\.
The performance of the surrogate model is evaluated by comparing predicted sensitivity profiles to OpenMC\-evaluated ground truths\. As defined in Section[2\.3](https://arxiv.org/html/2606.04033#S2.SS3), mean squared error \(MSE\) and mean absolute error \(MAE\) are used as metrics for surrogate accuracy\. Test and validation sets of 1,000 geometries were held fixed, and the model was iteratively trained with increasingly large training sets until the test sensitivity error plateaued beneath 50pcm\. This point was found to occur once the size of the training set reached 3,000 geometries, at which point the training dataset was held constant for further evaluations\.
Figure 8:Test set sensitivity error as a function of training set size\.Figure 9:Distribution ofkeffk\_\{\\text\{eff\}\}across the generated dataset\.To benchmark the performance of the attention pooling layer against traditional alternatives, versions of the model were trained and evaluated with max and average pooling\. Each model configuration was evaluated across the training, validation, and test sets\. As shown in Table[3](https://arxiv.org/html/2606.04033#S3.T3)and Figure[10](https://arxiv.org/html/2606.04033#S3.F10), attention pooling resulted in the lowest error which supports its further use in the remainder of this work\. Final surrogate error metrics are shown in Table[4](https://arxiv.org/html/2606.04033#S3.T4)\. Average test error as a function of energy group is shown in Figure[11](https://arxiv.org/html/2606.04033#S3.F11), with per\-reaction error given in[C](https://arxiv.org/html/2606.04033#A3)\. While surrogate error may have some impact on the optimization process, the practice of evaluating multiple of the highestc^k\\hat\{c\}\_\{k\}experiment designs in OpenMC ensures that the final chosen geometry is that with the highest trueckc\_\{k\}of its peers, despite any errors inc^k\\hat\{c\}\_\{k\}prediction\.
Figure 10:Validation loss curves for different pooling configurations\.PoolingTrain MAEVal MAETest MAEAverage44\.6349\.4848\.27Max44\.3547\.7046\.13Attention40\.01\\bf 40\.0142\.73\\bf 42\.7341\.19\\bf 41\.19Table 3:Sensitivity MAE \(pcm\) for different pooling configurations with lowest values presented in bold\.SplitItemsMSEMAE \(pcm\)Train3,0007\.076⋅10−77\.076\\cdot 10^\{\-7\}40\.0140\.01Validation1,0008\.051⋅10−78\.051\\cdot 10^\{\-7\}42\.7342\.73Test1,0007\.620⋅10−77\.620\\cdot 10^\{\-7\}41\.1941\.19Table 4:Final surrogate error metrics across training, validation, and test sets\.Figure 11:Average test set sensitivity error across all reactions as a function of energy\.
### 3\.3Interpretability of Multigroup Attention Pools
The internal workings of deep learning models are generally considered “black box,” and difficult \(or impossible\) to interpret\. However, the spatial and physically motivated structure of the multigroup attention pooling layer introduced in Section[2\.5\.2](https://arxiv.org/html/2606.04033#S2.SS5.SSS2)means that it can be examined with a degree of human intuition\. To understand the behavior of the attention pooling layer, the model was run on several simple handcrafted geometries shown in Figure[12](https://arxiv.org/html/2606.04033#S3.F12)\.
For each geometry, the attention pool activations for the235U fission sensitivities were extracted and averaged for thermal \(00\\,eV to11\\,eV\) and fast \(10410^\{4\}\\,eV to10710^\{7\}\\,eV\) groups\. Since the attention scores do not have physically meaningful units, they were not flux\-weighted when rebinning\. While flux\-weighting would be necessary if converting energies or sensitivities to a coarser group structure, this interpretability study is nonphysical and qualitative in nature\. The goal is simply to gain a directional understanding of how the model internally represents the information it learns\.
Figure 12:Normalized attention pool activations for235U fission sensitivities\.A few observations can be made from the activation maps shown in Figure[12](https://arxiv.org/html/2606.04033#S3.F12)\. In both geometries, importance to the lower energy groups is almost entirely contained within the fuel regions themselves, while the higher energy group assigns more importance to the areas surrounding the fuel\. This is consistent with physical intuition: higher energy neutrons will travel further from the source before interacting, while thermal neutrons are more likely to stay contained within their source\.
Additionally, the low energy activations of the second geometry assign more importance to the internal fuel regions than those closer to the edges\. This is also consistent with intuition, as neutrons born in higher leakage regions are less likely to contribute to fission\.
While qualitative, these results indicate the potential of multigroup attention pooling as an interpretable neural network component for capturing the underlying physics of neutron transport\. Better understanding the internal behavior of surrogate models for neutron transport is of great interest, and is a target of future work\.
### 3\.4TN\-LC Experiment Design
For each target configuration of the TN\-LC, a batch of 500 randomly initialized grids of material assignments were created with an average initialckc\_\{k\}shown in Table[5](https://arxiv.org/html/2606.04033#S3.T5)\. The gradient optimization algorithm \(described in Section[2\.6](https://arxiv.org/html/2606.04033#S2.SS6)\) was then run for 2000 steps using the AdamW optimizer\[[15](https://arxiv.org/html/2606.04033#bib.bib38)\]\.
The batch range and highestc^k\\hat\{c\}\_\{k\}as a function of optimization timestep for the Case 2 configuration are shown in Figure[13](https://arxiv.org/html/2606.04033#S3.F13)\. The magnitude of difference between the mean and highestc^k\\hat\{c\}\_\{k\}within the batch of 500 initializations indicates the value of optimizing over a large set of geometries, rather than relying on a single initialization\.
Figure 13:Mean and maxc^k\\hat\{c\}\_\{k\}at each optimization step during the batch optimization for the Case 2 target configuration\.Figure 14:Snapshots of the geometry with the highestc^k\\hat\{c\}\_\{k\}at various optimization steps for the Case 2 target configuration\.Optimized experiment geometries and corresponding235U sensitivity vectors are shown for the Dry Cask, Case 1, and Case 2 TN\-LC configurations in Figures[15](https://arxiv.org/html/2606.04033#S3.F15),[16](https://arxiv.org/html/2606.04033#S3.F16), and[17](https://arxiv.org/html/2606.04033#S3.F17)respectively\. Full sensitivity profiles for each case are given in[B](https://arxiv.org/html/2606.04033#A2)\.
Dry CaskCase 1Case 2Mean Initialckc\_\{k\}0\.860830\.860830\.679630\.679630\.702780\.70278Finalc^k\\hat\{c\}\_\{k\}\(Predicted\)0\.982270\.982270\.845600\.845600\.926930\.92693Finalckc\_\{k\}\(OpenMC\)0\.977570\.977570\.813240\.813240\.932760\.93276Optimizationckc\_\{k\}Gain0\.116740\.133610\.22998Experimentkeffk\_\{\\text\{eff\}\}0\.971521\.065331\.07930Table 5:Optimization results for the Dry Cask, Case 1, and Case 2 TN\-LC configurations\.The optimization procedure was successful in all three cases, yielding experiments withck≥0\.9c\_\{k\}\\geq 0\.9for the Dry Cask and Case 2 TN\-LC configurations andck\>0\.8c\_\{k\}\>0\.8for the Case 1 configuration\. Notably, ackc\_\{k\}of 0\.93276 was achieved for the Case 2 configuration, for which no previous experiment was found with ack≥0\.8c\_\{k\}\\geq 0\.8\[[4](https://arxiv.org/html/2606.04033#bib.bib10)\]\.
There are several notable features present in the final geometries\. In each case, the material distribution is highly heterogeneous relative to the distribution of the surrogate training dataset\. Additionally, relatively little stainless steel or polyethylene are present in the final geometries; the optimizer seems to prefer to construct the bulk of the geometries fromUO2\\text\{UO\}\_\{2\}andAlB, as shown in Table[6](https://arxiv.org/html/2606.04033#S3.T6)\.
Dry CaskCase 1Case 2UO235\.74%38\.28%32\.03%SS30422\.66%12\.60%9\.57%AlB30\.96%38\.48%33\.30%Polyethylene10\.64%10\.64%25\.10%Table 6:Material composition percentages in the final optimized geometries for the Dry Cask, Case 1, and Case 2 TN\-LC configurations\.As shown in[B](https://arxiv.org/html/2606.04033#A2), across all cases the optimizer prioritizes matching the shape and magnitudes of the235U fission sensitivities, which is perhaps intuitive given the magnitude of sensitivities for that reaction\. It follows that the reaction with the greatest effect onkeffk\_\{\\text\{eff\}\}is fission itself\. While some effort is made to match theNi58\{\}^\{58\}\\text\{Ni\}thermal sensitivity magnitude, the12C and1H sensitivities seem to be lowest priority\.
Certain differences in final sensitivities may be an unavoidable result of the differences in scale and fidelity between the targets and the experiments\. The optimizer is unable to match the magnitudes of the resonance sensitivities for235U fission and1H elastic scattering, nor is it able to match the high energy sensitivities for58Ni absorption\.
It is unclear why the Case 1 configuration posed such difficulty for the optimizer, especially given the similarity in sensitivities to Case 2\. Nonetheless, the optimizer successfully increased
ckc\_\{k\}by 0\.13361 to achieve
ck\>0\.8c\_\{k\}\>0\.8, which Eidelpes et al\. set as the cutoff below which a
keffk\_\{\\text\{eff\}\}penalty must be applied when determining subcritical margins\[[4](https://arxiv.org/html/2606.04033#bib.bib10)\]\.
Further work is required to understand the degree to which the different reactions and energies contribute tockc\_\{k\}, especially given the effect of covariance weighting, which is not visible when overlaying raw sensitivity vectors\.
Figure 15:Optimized experiment geometry for the dry TN\-LC configuration\.Figure 16:Optimized experiment geometry for the TN\-LC Case 1 configuration \(lowB10\{\}^\{10\}\\text\{B\}areal density\)\.Figure 17:Optimized experiment geometry for the TN\-LC Case 2 configuration \(highB10\{\}^\{10\}\\text\{B\}areal density\)\.To explore the full capabilities of the method unencumbered by human biases, no constraints were imposed on the design space other than the material palette and grid dimensions\. As such, the optimized geometries are unintuitive and “alien” in design\. The optimization algorithm considers only the single\-objective problem of maximizingc^k\\hat\{c\}\_\{k\}without regard for factors such as cost, construction feasibility, leakage, orkeffk\_\{\\text\{eff\}\}itself\. However, these additional constraints can be enforced by the introduction of additional terms to the loss function of the optimizer\. In particular, optimization ofkeffk\_\{\\text\{eff\}\}would require training a separate surrogate model, which would be straightforward but is beyond the focus of this work\.
## 4Conclusion
This work presents a methodology for the inverse design of nuclear criticality experiments using deep neural network surrogate modeling and gradient\-based geometry optimization\. A dataset was created from 5,000 procedurally generated experiment geometries and corresponding sensitivity profiles, which was then used to train a surrogate model for sensitivity prediction\. The model combined a U\-Net\[[22](https://arxiv.org/html/2606.04033#bib.bib36)\]with a novel multigroup attention pooling layer, introduced to capture the importance of different regions in the geometry to the sensitivity of each energy group\. The attention pooling layer demonstrated superior predictive performance relative to standard pooling alternatives, as well as internal representations qualitatively consistent with the intuitively expected neutron flux distribution across each tested geometry\.
A gradient optimization algorithm was used to design criticality experiments given a target sensitivity profile with the maximization of the correlation coefficientckc\_\{k\}as the optimization objective\. Applied to the design of experiments for the validation of the TN\-LC HALEU transportation cask, the optimization procedure successfully produced experiment geometries withckc\_\{k\}greater than 0\.9 for two of three considered configurations, notably including a flooded cask scenario for which no previous experiment was identified withck\>0\.8c\_\{k\}\>0\.8\[[4](https://arxiv.org/html/2606.04033#bib.bib10)\]\. The correlation coefficients for the dry cask, flooded cask with low\-density poison, and flooded cask with high\-density poison were 0\.97757, 0\.81324, and 0\.93276 respectively\. The experimentkeffk\_\{\\text\{eff\}\}values were 0\.97152, 1\.06533, and 1\.07930\.
The primary shortcoming of the method is that it does not considerkeffk\_\{\\text\{eff\}\}during the optimization process\. The maximization ofckc\_\{k\}is the sole optimization objective, without constraints onkeffk\_\{\\text\{eff\}\}, construction feasibility, or material cost\. For realistic implementation,keffk\_\{\\text\{eff\}\}must be considered as part of the optimization process to ensure that all generated experiments reachkeff≥1k\_\{\\text\{eff\}\}\\geq 1\. The incorporation of these considerations through additional surrogate models or terms in the optimization objective is a natural and tractable extension of the method\. Extension to three\-dimensional geometries and the full material design space of a critical experiment facility such as SPARC\[[35](https://arxiv.org/html/2606.04033#bib.bib18)\]are also directions for future work\. Additionally, more rigorous interpretability analysis could be done to quantify the behavior of multigroup attention pooling\.
The ability to rapidly generate critical experiment designs with highckc\_\{k\}for arbitrary target systems has the potential to accelerate the validation of advanced nuclear technology for which historical experimental coverage is limited\. The performance and observed internal behavior of the surrogate model in this work indicates the potential of deep learning applied to neutronics\. There is potential for further work to explore the degree to which high\-fidelity surrogates can be trained to approximate more fundamental properties of neutron transport\. This would allow the methodology presented in this work to be applied to a wider range of nuclear system design tasks\.
## Acknowledgement
This research made use of the resources of the High Performance Computing Center at Idaho National Laboratory, which is supported by the Office of Nuclear Energy of the U\.S\. Department of Energy and the Nuclear Science User Facilities under Contract No\. DE\-AC07\-05ID14517\.
## References
- \[1\]Y\. Bengio, N\. Léonard, and A\. Courville\(2013\)Estimating or propagating gradients through stochastic neurons for conditional computation\.Note:arXiv preprintExternal Links:1308\.3432,[Document](https://dx.doi.org/10.48550/arXiv.1308.3432)Cited by:[§2\.6](https://arxiv.org/html/2606.04033#S2.SS6.p5.8)\.
- \[2\]M\. Branco\-Katcher, D\. Siefman, R\. Araj, T\. Cisernos, C\. Percher, and T\. S\. Palmer\(2025\)Design optimization of a criticality experiment for the Molten Chloride Reactor Experiment facility\.Nuclear Science and Engineering199\(11\),pp\. 1794–1815\.External Links:[Document](https://dx.doi.org/10.1080/00295639.2025.2464459)Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p5.2)\.
- \[3\]K\. Deb, A\. Pratap, S\. Agarwal, and T\. Meyarivan\(2002\)A fast and elitist multi\-objective genetic algorithm: NSGA\-II\.IEEE Transactions on Evolutionary Computation6\(2\),pp\. 182–197\.External Links:[Document](https://dx.doi.org/10.1109/4235.996017)Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p6.1)\.
- \[4\]E\. Eidelpes, J\. J\. Jarrell, H\. E\. Adkins, B\. M\. Hom, J\. M\. Scaglione, R\. A\. Hall, and B\. D\. Brickner\(2019\-11\)UO2 HALEU transportation package evaluation and recommendations\.Technical reportTechnical ReportINL/EXT\-19\-56333, Revision 0,Idaho National Laboratory,Idaho Falls, ID, USA\.Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p4.7),[§2\.1](https://arxiv.org/html/2606.04033#S2.SS1.p1.6),[§2\.1](https://arxiv.org/html/2606.04033#S2.SS1.p2.7),[§2\.1](https://arxiv.org/html/2606.04033#S2.SS1.p4.1),[§3\.4](https://arxiv.org/html/2606.04033#S3.SS4.p4.4),[§3\.4](https://arxiv.org/html/2606.04033#S3.SS4.p7.6),[§4](https://arxiv.org/html/2606.04033#S4.p2.4)\.
- \[5\]R\. Hall, W\. J\. Marshall, and W\. A\. Wieselquist\(2020\-10\)Assessment of existing transportation packages for use with HALEU\.Technical reportTechnical ReportORNL/TM\-2020/1725,Oak Ridge National Laboratory,Oak Ridge, TN, USA\.External Links:[Document](https://dx.doi.org/10.2172/1731046)Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p4.7),[§2\.1](https://arxiv.org/html/2606.04033#S2.SS1.p2.7)\.
- \[6\]International Energy Agency\(2025\)Energy and AI\.Technical reportIEA,Paris\.Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p1.1)\.
- \[7\]International Energy Agency\(2025\)The path to a new era for nuclear energy\.Technical reportIEA,Paris\.Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p1.1)\.
- \[8\]D\. P\. Kingma and J\. Ba\(2014\)Adam: a method for stochastic optimization\.Note:arXiv preprintExternal Links:1412\.6980,[Document](https://dx.doi.org/10.48550/arXiv.1412.6980)Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p8.1)\.
- \[9\]N\. Kleedtke, T\. Cutler, D\. Neudecker, J\. Hutchinson, P\. Brain, N\. Thompson, and R\. C\. Little\(2026\)Genetic algorithm optimization of nuclear criticality experiment for reduction of intermediate\-energy239Pu nuclear data uncertainties\.Annals of Nuclear Energy228,pp\. 112037\.External Links:[Document](https://dx.doi.org/10.1016/j.anucene.2025.112037)Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p5.2)\.
- \[10\]J\. N\. Kutz\(2017\)Deep learning in fluid dynamics\.Journal of Fluid Mechanics814,pp\. 1–4\.External Links:[Document](https://dx.doi.org/10.1017/jfm.2016.803)Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p7.1)\.
- \[11\]Y\. LeCun, Y\. Bengio, and G\. Hinton\(2015\)Deep learning\.Nature521,pp\. 436–444\.External Links:[Document](https://dx.doi.org/10.1038/nature14539)Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p7.1)\.
- \[12\]Y\. LeCun, L\. Bottou, Y\. Bengio, and P\. Haffner\(1998\)Gradient\-based learning applied to document recognition\.Proceedings of the IEEE86\(11\),pp\. 2278–2324\.External Links:[Document](https://dx.doi.org/10.1109/5.726791)Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p7.1)\.
- \[13\]R\. M\. Lell\(2024\-04\)ZPR/ZPPR critical experiment facilities and as\-built ZPR/ZPPR models\.Technical reportArgonne National Laboratory,Argonne, IL, USA\.External Links:[Document](https://dx.doi.org/10.2172/2506757)Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p2.1),[§2\.4](https://arxiv.org/html/2606.04033#S2.SS4.p2.1)\.
- \[14\]J\. Leppänen, V\. Valtavirta, A\. Rintala, and R\. Tuominen\(2025\)Status of the Serpent Monte Carlo code in 2024\.EPJ Nuclear Sciences & Technologies11,pp\. 3\.External Links:[Document](https://dx.doi.org/10.1051/epjn/2024031)Cited by:[§2\.1](https://arxiv.org/html/2606.04033#S2.SS1.p3.1),[§2\.2](https://arxiv.org/html/2606.04033#S2.SS2.p2.1),[§3\.1](https://arxiv.org/html/2606.04033#S3.SS1.p2.1)\.
- \[15\]I\. Loshchilov and F\. Hutter\(2019\)Decoupled weight decay regularization\.InProceedings of the 7th International Conference on Learning Representations \(ICLR\),New Orleans, LA, USA\.Cited by:[§2\.5\.4](https://arxiv.org/html/2606.04033#S2.SS5.SSS4.p2.1),[§2\.6](https://arxiv.org/html/2606.04033#S2.SS6.p6.4),[§3\.4](https://arxiv.org/html/2606.04033#S3.SS4.p1.1)\.
- \[16\]A\. Maldonado and C\. Perfetti\(2023\)Utilizing sensitivity and correlation coefficients from MCNP and WHISPER to guide microreactor experiment design\.Nuclear Science and Engineering197\(8\),pp\. 2086–2098\.External Links:[Document](https://dx.doi.org/10.1080/00295639.2022.2162782)Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p3.6),[§2\.1](https://arxiv.org/html/2606.04033#S2.SS1.p1.6)\.
- \[17\]E\. Nikolaou, S\. Kilimtzidis, and V\. Kostopoulos\(2025\)Multi\-fidelity surrogate\-assisted aerodynamic optimization of aircraft wings\.Aerospace12\(4\),pp\. 359\.External Links:[Document](https://dx.doi.org/10.3390/aerospace12040359)Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p6.1)\.
- \[18\]J\. Nocedal\(1980\)Updating quasi\-Newton matrices with limited storage\.Mathematics of Computation35\(151\),pp\. 773–782\.External Links:[Document](https://dx.doi.org/10.1090/S0025-5718-1980-0572855-7)Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p8.1)\.
- \[19\]X\. Peng, J\. Liang, A\. Alhajri, B\. Forget, and K\. Smith\(2017\)Development of continuous\-energy sensitivity analysis capability in OpenMC\.Annals of Nuclear Energy110,pp\. 362–383\.External Links:[Document](https://dx.doi.org/10.1016/j.anucene.2017.06.061)Cited by:[§2\.2](https://arxiv.org/html/2606.04033#S2.SS2.p2.1)\.
- \[20\]J\. Pevey, V\. Sobes, and W\. J\. Hines\(2022\)Neural network acceleration of genetic algorithms for the optimization of a coupled fast/thermal nuclear experiment\.Frontiers in Energy Research10,pp\. 874194\.External Links:[Document](https://dx.doi.org/10.3389/fenrg.2022.874194)Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p8.1)\.
- \[21\]D\. Price, M\. I\. Radaideh, and B\. Kochunas\(2022\)Multiobjective optimization of nuclear microreactor reactivity control system operation with swarm and evolutionary algorithms\.Nuclear Engineering and Design393,pp\. 111776\.External Links:[Document](https://dx.doi.org/10.1016/j.nucengdes.2022.111776)Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p6.1)\.
- \[22\]O\. Ronneberger, P\. Fischer, and T\. Brox\(2015\)U\-Net: convolutional networks for biomedical image segmentation\.Note:arXiv preprintExternal Links:1505\.04597,[Document](https://dx.doi.org/10.48550/arXiv.1505.04597)Cited by:[§2\.5\.1](https://arxiv.org/html/2606.04033#S2.SS5.SSS1.p1.3),[§2\.5](https://arxiv.org/html/2606.04033#S2.SS5.p2.3),[§4](https://arxiv.org/html/2606.04033#S4.p1.1)\.
- \[23\]C\. Sabater, P\. Bekemeyer, and S\. Görtz\(2020\)Efficient bilevel surrogate approach for optimization under uncertainty of shock control bumps\.AIAA Journal58\(12\),pp\. 5228–5242\.External Links:[Document](https://dx.doi.org/10.2514/1.J059480)Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p6.1)\.
- \[24\]R\. Sanchez, T\. Cutler, J\. Goda, T\. Grove, D\. Hayes, J\. Hutchinson, G\. McKenzie, A\. McSpaden, W\. Myers, R\. Rico, J\. Walker, and R\. Weldon\(2021\)A new era of nuclear criticality experiments: the first 10 years of PLANET operations at NCERC\.Nuclear Science and Engineering195\(sup\. 1\),pp\. S1–S16\.External Links:[Document](https://dx.doi.org/10.1080/00295639.2021.1951077)Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p2.1)\.
- \[25\]P\. Seurin, D\. Price, and L\. Nunez\(2025\)Techno\-economic optimization of a heat\-pipe microreactor, part I: theory and cost optimization\.Note:arXiv preprintExternal Links:2512\.16032,[Document](https://dx.doi.org/10.48550/arXiv.2512.16032)Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p6.1)\.
- \[26\]P\. Seurin and D\. Price\(2026\)Techno\-economic optimization of a heat\-pipe microreactor, part II: multi\-objective optimization analysis\.Note:arXiv preprintExternal Links:2601\.20079,[Document](https://dx.doi.org/10.48550/arXiv.2601.20079)Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p6.1)\.
- \[27\]N\. Thompson, R\. Sanchez, J\. Goda, K\. Amundson, T\. Cutler, T\. Grove, D\. Hayes, J\. Hutchinson, C\. Kostelac, G\. McKenzie, A\. McSpaden, W\. Myers, and J\. Walker\(2021\)A new era of nuclear criticality experiments: the first 10 years of COMET operations at NCERC\.Nuclear Science and Engineering195\(sup\. 1\),pp\. S17–S36\.External Links:[Document](https://dx.doi.org/10.1080/00295639.2021.1947105)Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p2.1)\.
- \[28\]N\. W\. Thompson, A\. Maldonado, T\. E\. Cutler, H\. R\. Trellue, K\. M\. Amundson, V\. Rao Dasari, J\. M\. Goda, T\. J\. Grove, D\. K\. Hayes, J\. D\. Hutchinson, H\. M\. Kistle, C\. M\. Kostelac, J\. R\. Lamproe, C\. Matthews, G\. McKenzie, G\. E\. McMath, A\. T\. McSpaden, R\. G\. Sanchez, K\. N\. Stolte, J\. L\. Walker, R\. Weldon, and N\. H\. Whitman\(2023\)The National Criticality Experiments Research Center and its role in support of advanced reactor design\.Frontiers in Energy Research10,pp\. 1082389\.External Links:[Document](https://dx.doi.org/10.3389/fenrg.2022.1082389)Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p2.1)\.
- \[29\]U\.S\. Department of Energy\(2024\-09\)Pathways to commercial liftoff: advanced nuclear\.Technical reportU\.S\. Department of Energy\.Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p1.1)\.
- \[30\]A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. Kaiser, and I\. Polosukhin\(2017\)Attention is all you need\.InAdvances in Neural Information Processing Systems,Vol\.30\.Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p7.1)\.
- \[31\]W\. Walters and N\. Roskoff\(2026\)Designing experiments for microreactor neutronic validation\.InPHYSOR 2026 – The International Conference on Physics of Reactors,Torino, Italy\.Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p3.6)\.
- \[32\]A\. Whyte and G\. Parks\(2021\)Surrogate model optimization of a ‘Micro Core’ PWR fuel assembly arrangement using deep learning models\.InProceedings of PHYSOR 2020,EPJ Web of Conferences, Vol\.247,pp\. 12003\.External Links:[Document](https://dx.doi.org/10.1051/epjconf/202124712003)Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p6.1)\.
- \[33\]W\. A\. Wieselquist and R\. A\. Lefebvre\(2024\-02\)SCALE 6\.3\.2 user manual\.Technical reportTechnical ReportORNL/TM\-2024/3386,Oak Ridge National Laboratory,Oak Ridge, TN, USA\.External Links:[Document](https://dx.doi.org/10.2172/2361197)Cited by:[§2\.1](https://arxiv.org/html/2606.04033#S2.SS1.p6.14)\.
- \[34\]A\. Williams, N\. Walton, A\. Maryanski, S\. Bogetic, W\. Hines, and V\. Sobes\(2023\)Stochastic gradient descent for optimization for nuclear systems\.Scientific Reports13,pp\. 8474\.External Links:[Document](https://dx.doi.org/10.1038/s41598-023-32112-7)Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p8.1)\.
- \[35\]N\. Woolstenhulme, M\. Cole, N\. Martin, T\. Reiss, J\. Johnson, M\. Lund, R\. Claghorn, K\. Gagnon, M\. DeHart, W\. Wieselquist, T\. Cutler, C\. Percher, A\. Barto, and D\. Algama\(2025\-07\)SPARC — plans for a new critical experiment facility with a horizontal split table\.Technical reportTechnical ReportINL/RPT\-25\-84855,Idaho National Laboratory\.Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p2.1),[§2\.4](https://arxiv.org/html/2606.04033#S2.SS4.p1.2),[§4](https://arxiv.org/html/2606.04033#S4.p3.5)\.
- \[36\]J\. Xu\(2019\)Distance\-based protein folding powered by deep learning\.Proceedings of the National Academy of Sciences116\(34\),pp\. 16856–16865\.External Links:[Document](https://dx.doi.org/10.1073/pnas.1821309116)Cited by:[§1](https://arxiv.org/html/2606.04033#S1.p7.1)\.
## Appendix ASurrogate Model Architecture and Training Details
Table 7:Surrogate architecture and training hyperparameters\.ParameterValueInputGeometry dimensions64×6464\\times 64One\-hot materials44U\-NetChannel widths32, 64, 128, 256Conv2d settingsk=3,p=1k=3,\\,\\,p=1GroupNorm settingsng=8n\_\{g\}=8Dropout0\.2Multigroup Attention PoolingSpatial dimensions64×6464\\times 64Energy group structureSCALE\-252Group RegressorHidden channels32Residual blocks6Conv1d settingsk=3,p=1k=3,\\,\\,p=1GroupNorm settingsng=8n\_\{g\}=8TrainingOptimizerAdamWLearning rate10−3→10−510^\{\-3\}\\to 10^\{\-5\}Epochs100100Batch size6464LossMSETraining items3000Validation/test items1000
## Appendix BTN\-LC Sensitivity Profiles
Figure 18:Full sensitivity profiles for the three optimized experiment geometries\.
## Appendix CPer\-Reaction Surrogate Error
Figure 19:Surrogate model test error \(MAE\) as a function of reaction and energy group\.Similar Articles
@BetaTomorrow: https://x.com/BetaTomorrow/status/2079015742738157750
An essay critiquing the common conflation of optimization and learning in neural network research, arguing that training should be understood as inverse reconstruction and studied through the evolving homology of weight-defined piecewise manifolds.
Steered Generation via Gradient-Based Optimization on Sparse Query Features
This paper introduces Prototype-Based Sparse Steering, a method that applies sparse autoencoders to attention query activations in LLMs, then uses gradient-based optimization during inference to steer generation toward target behaviors. The approach is validated in both a logical planning task and a stylistic educational domain, demonstrating interpretable and disentangled control.
A lift for input-convex neural network training
Proposes a 'lift' method for training input-convex neural networks (ICNNs) that uses an unconstrained hypernetwork to emit non-negative inter-layer weights, softening the loss landscape and escaping gradient attenuation, achieving lower test loss than projected gradient descent and softplus reparametrization.
In-Context Optimization for Retrieval-Augmented Generation: A Gradient-Descent Perspective
This paper studies retrieval-augmented generation as an in-context optimization process, showing that linear self-attention can implement gradient descent on a unified RAG objective. It proposes a lightweight method for frozen RAG LLMs that predicts context-conditioned updates, improving performance across multiple QA benchmarks.
From brute-force graph traversal to Cognitive Attention: an architectural redesign
The author redesigned the IONS protocol from brute-force graph traversal to a Cognitive Attention Architecture that progressively routes queries through relevant slices of the network, separating path confidence, relevance, and utility.