生成式反转用于竞争性地质解释的早期排名
摘要
本文提出了一种使用文本到图像基础模型和变分自编码器的工作流程,基于与水压头观测数据的一致性对竞争性地质解释进行排名,以改进早期地下决策制定。
arXiv:2609.20978v1 Announce Type: new
Abstract: High-consequence subsurface decisions are often made under severe data scarcity. Experts may arrive at competing interpretations of the same subsurface system, yet early in a project there is rarely a practical way to determine which one is most realistic. This uncertainty can persist until several wells are drilled, often costing millions of dollars. Existing approaches for evaluating geologic interpretations rely either on subjective judgment or on dense data that are rarely available in early-stage investigations. We present a workflow that addresses this challenge by translating competing geologic interpretations into alternative spatial priors and ranking them according to their consistency with hydraulic-head observations. For each interpretation, a text-to-image foundation model generates an ensemble of 1600 geologic images, and a separately trained variational autoencoder provides an interpretation-specific latent representation. A supervised inverse network maps the head observations into this latent space, and the frozen decoder produces an image that is mapped to a log-conductivity field. Steady-state flow simulation then provides predicted heads, and the resulting mismatch is converted into a Gaussian-form compatibility score. We evaluate the framework using a synthetic benchmark based on the Johansen Formation and three interpretations of decreasing consistency with the reference representation. Across 925 test cases, the mean head RMSE increases from 0.197 for the Precise \& Accurate interpretation to 0.227 for the Accurate interpretation and 0.280 for the Mismatched interpretation. We subsequently apply the workflow to two published conceptual models of the Culebra Dolomite Member at the Waste Isolation Pilot Plant. The revised model receives a compatibility weight of 0.991, compared with 0.009 for the original model, consistent with the independent evidence.
查看缓存全文
缓存时间: 2026/09/21 09:15
# Generative inversion for early ranking of competing geologic interpretations
Source: [https://arxiv.org/html/2609.20978](https://arxiv.org/html/2609.20978)
H\. U\. RashidAffiliation:Earth and Environmental Sciences Division, Los Alamos National Laboratory, Los Alamos, NM, USAD\. O’MalleyAffiliation:Earth and Environmental Sciences Division, Los Alamos National Laboratory, Los Alamos, NM, USA
###### Abstract
High\-consequence subsurface decisions are often made under severe data scarcity\. Experts may arrive at competing interpretations of the same subsurface system, yet early in a project there is rarely a practical way to determine which one is most realistic\. This uncertainty can persist until several wells are drilled, often costing millions of dollars\. Existing approaches for evaluating geologic interpretations rely either on subjective judgment or on dense data that are rarely available in early\-stage investigations\.
We present a workflow that addresses this challenge by translating competing geologic interpretations into alternative spatial priors and ranking them according to their consistency with hydraulic\-head observations\. For each interpretation, a text\-to\-image foundation model generates an ensemble of 1600 geologic images, and a separately trained variational autoencoder provides an interpretation\-specific latent representation\. A supervised inverse network maps the head observations into this latent space, and the frozen decoder produces an image that is mapped to a log\-conductivity field\. Steady\-state flow simulation then provides predicted heads, and the resulting mismatch is converted into a Gaussian\-form compatibility score\. Softmax normalization produces relative compatibility weights across the finite candidate set\. We evaluate the framework using a synthetic benchmark based on the Johansen Formation and three interpretations of decreasing consistency with the reference representation\. Across 925 test cases, the mean head RMSE increases from 0\.197 for the Precise & Accurate interpretation to 0\.227 for the Accurate interpretation and 0\.280 for the Mismatched interpretation\. We subsequently apply the workflow to two published conceptual models of the Culebra Dolomite Member at the Waste Isolation Pilot Plant\. The revised model receives a compatibility weight of 0\.991, compared with 0\.009 for the original model, consistent with the independent evidence that motivated the conceptual\-model revision\.
††journal:JGR Machine Learning and Computation††corresponding:H\. U\. Rashid, hrashid\.lsu@gmail\.com###### keypoints
We convert geologic interpretations into geologic maps that can be tested against flow observations\. Our framework ranks competing geologic interpretations based on their consistency with limited observations\. Our framework consistently ranks competing interpretations in real\-field case studies\.
## Plain Language Summary
Decisions about the deep subsurface often have to be made when very little is known about the site\. Geologists studying the same place frequently disagree about how it is put together, and each records their view as a written description\. Early in a project there is rarely a practical way to tell which description is closest to reality, and the disagreement can persist until several wells have been drilled at a cost of millions of dollars\. Existing ways of choosing between descriptions rely either on expert judgment or on far more data than an early\-stage project has\.
We built a method that tests written descriptions against measurements that already exist\. Artificial intelligence tools that generate images from text turn each description into a set of possible maps showing how easily water moves through the rock\. A computer model predicts the water pressures each map would produce, and those predictions are compared with pressures measured in wells\. The description whose predictions come closest is ranked highest, and the ranking carries a number showing by how much\.
In tests where the correct answer was known, the method found it\. Applied to the Waste Isolation Pilot Plant in New Mexico, where the operator revised its understanding of a key water\-bearing layer after new field data arrived, the method independently favored the revised description\.
## 1Introduction
Predictive subsurface models underpin decisions in groundwater management, hydrocarbon production, geothermal development, geologic carbon storage, and subsurface waste isolation\. A central component of these models is the geological conceptualization: an interpretation of which units and structures are present, how they are connected, and which spatial features control fluid flow and transport\. These interpretations are constructed from incomplete observations such as cores, well logs, seismic and other geophysical data, outcrop analogues, and geological reasoning\. Because the subsurface is only sparsely observed, the resulting geological interpretation is inherently nonunique[Caers \(2011\)](https://arxiv.org/html/2609.20978#bib.bib22);[Wellmann and Caumon \(2018\)](https://arxiv.org/html/2609.20978#bib.bib23)\. In a controlled experiment illustrating this ambiguity,[Bond et al\. \(2007\)](https://arxiv.org/html/2609.20978#bib.bib1)analyzed 412 interpretations of the same synthetic seismic section and found substantial variability among interpreters; only 21% identified the intended tectonic setting\. Thus, alternative geological interpretations can arise even when investigators are presented with the same underlying observations\.
The consequences of this interpretational uncertainty extend across subsurface applications\. In groundwater modeling, an incorrect conceptualization may reproduce available calibration data yet produce unreliable predictions[Bredehoeft \(2005\)](https://arxiv.org/html/2609.20978#bib.bib2), and alternative hydrogeological conceptual models can produce substantially different forecasts[Højberg and Refsgaard \(2005\)](https://arxiv.org/html/2609.20978#bib.bib3);[Refsgaard et al\. \(2006\)](https://arxiv.org/html/2609.20978#bib.bib4);[Refsgaard et al\. \(2012\)](https://arxiv.org/html/2609.20978#bib.bib5)\. Similar problems arise in petroleum reservoir modeling, where uncertainty in the geological scenario can strongly influence reservoir forecasts and development decisions[Demyanov et al\. \(2019\)](https://arxiv.org/html/2609.20978#bib.bib24)\. In geologic carbon storage, alternative representations of heterogeneity, stratigraphic architecture, and boundary conditions can alter predicted pressure evolution, plume migration, and storage performance[Li et al\. \(2011\)](https://arxiv.org/html/2609.20978#bib.bib25)\. Geological uncertainty is likewise central to geothermal reservoir characterization and long\-term assessments of subsurface waste isolation[Scheidt et al\. \(2018\)](https://arxiv.org/html/2609.20978#bib.bib26);[Wellmann and Caumon \(2018\)](https://arxiv.org/html/2609.20978#bib.bib23)\. Across these applications, parameter uncertainty within a single geological model therefore represents only part of the total predictive uncertainty\.
One strategy for addressing this problem is to construct multiple plausible geological conceptual models and propagate each through the forward model\. Their predictions can then be compared or combined using multimodel inference, Bayesian model averaging, or scenario\-based uncertainty quantification[Neuman \(2003\)](https://arxiv.org/html/2609.20978#bib.bib6);[Ye et al\. \(2004\)](https://arxiv.org/html/2609.20978#bib.bib7);[Demyanov et al\. \(2019\)](https://arxiv.org/html/2609.20978#bib.bib24)\. Such approaches account for alternative interpretations once numerical realizations of those interpretations are available\. A practical difficulty occurs earlier in the workflow: each geological interpretation must first be translated into a simulation\-ready representation\. Geological descriptions are commonly expressed through maps, cross sections, conceptual diagrams, and written interpretations, whereas numerical simulators require explicit spatial distributions of facies or physical properties\. Constructing these representations for many competing interpretations can require substantial manual effort and may introduce additional subjective choices during model construction[Wellmann and Caumon \(2018\)](https://arxiv.org/html/2609.20978#bib.bib23)\.
Dynamic observations such as pressure and hydraulic head provide an integrated test of geologic structure because they reflect the combined effects of unit geometry, connectivity, and property contrasts along flow paths\. However, these observations are commonly used to estimate parameters within a geological prior selected in advance rather than to evaluate the interpretation defining that prior\. Multimodel approaches can compare alternative interpretations, but only after each has been translated into a numerical representation[Refsgaard et al\. \(2012\)](https://arxiv.org/html/2609.20978#bib.bib5);[Demyanov et al\. \(2019\)](https://arxiv.org/html/2609.20978#bib.bib24)\. Establishing an efficient link between written interpretations, alternative spatial priors, and common dynamic observations would therefore enable earlier and more systematic evaluation of geological uncertainty\.
Generative methods provide a potential means of reducing this barrier\. Deep generative models can represent complex geological property fields using low\-dimensional latent variables while preserving spatial structures characteristic of the training data\.[Laloy et al\. \(2017\)](https://arxiv.org/html/2609.20978#bib.bib8)used a deep generative representation of complex binary geological media for inverse modeling, while[Laloy et al\. \(2018\)](https://arxiv.org/html/2609.20978#bib.bib9)developed a spatial generative adversarial network for training\-image\-based geostatistical inversion\.[Mo et al\. \(2020\)](https://arxiv.org/html/2609.20978#bib.bib10)combined an adversarial autoencoder with a convolutional surrogate for estimating non\-Gaussian conductivity fields\. More broadly, training\-image\-based geostatistical methods generate ensembles that reproduce geological patterns specified through representative training images[Strebelle \(2002\)](https://arxiv.org/html/2609.20978#bib.bib11);[Mariethoz and Caers \(2014\)](https://arxiv.org/html/2609.20978#bib.bib12)\. These approaches enable geologically realistic parameterizations, but the geological patterns to be represented must still be specified beforehand through training data or training images\. The choice of those patterns therefore encodes an important part of the assumed geological conceptualization\.
Text\-conditioned generative models offer a different route for connecting geological interpretation to spatial representation\. Latent diffusion models, for example, can generate images conditioned on natural\-language descriptions[Rombach et al\. \(2022\)](https://arxiv.org/html/2609.20978#bib.bib14)\. This capability creates the possibility of using written geological interpretations directly as conditioning information: alternative descriptions can generate alternative ensembles of geological maps, which can subsequently be tested against dynamic observations\. Recent work has also introduced large language models into process\-based subsurface modeling\.[Ma et al\. \(2026\)](https://arxiv.org/html/2609.20978#bib.bib15), for example, developed a multi\-agent framework that generates and iteratively modifies calibration code around simulators including MODFLOW\-2005 and TOUGHREACT\. Such approaches automate components of model implementation and calibration for a specified model structure\. Here, we address a complementary problem: whether alternative geological interpretations themselves can be translated into spatial models and evaluated against observations\.
This study introduces a framework that treats competing written geologic interpretations as testable hypotheses\. The framework converts each interpretation into an ensemble of spatial geologic realizations, conditions an interpretation\-specific latent model on sparse hydraulic\-head observations, and uses forward\-model likelihoods to quantify the relative support for each interpretation within the candidate set\. We first evaluate the framework using a controlled benchmark based on the Johansen Formation, comprising three interpretations with progressively decreasing fidelity to the reference geology\. We then apply it to two published conceptual models of the Culebra Dolomite Member at the Waste Isolation Pilot Plant \(WIPP\), where a documented conceptual\-model revision provides an independent basis for evaluating the ranking\. The central objective is to determine whether limited hydraulic observations contain sufficient information to distinguish among competing geologic interpretations\.
The remainder of this paper is organized as follows\. Section 2 presents the proposed framework, including the generation of interpretation\-specific image ensembles, latent\-space representation and inversion, forward flow simulation, and likelihood\-based ranking\. Section 3 evaluates the framework using a controlled synthetic benchmark comprising three geologic interpretation classes\. Section 4 applies the framework to two competing conceptual models of the Culebra Dolomite at WIPP\. Section 5 discusses the interpretation of the rankings, the information provided by existing observations, connections to related machine\-learning approaches, and the limitations of the framework\. Section 6 summarizes the principal findings and identifies priorities for future research\.
## 2Method
### 2\.1Problem formulation and workflow
Let\{C1,…,CK\}\\\{C\_\{1\},\\ldots,C\_\{K\}\\\}denote a finite set of candidate geologic interpretations, each expressed as a written description\. For observation caseii, hydraulic heads𝐝i∈ℝNobs\\mathbf\{d\}\_\{i\}\\in\\mathbb\{R\}^\{N\_\{\\mathrm\{obs\}\}\}are available at a fixed set of monitoring locations\. The objective is to evaluate how consistently each candidate interpretation can reproduce these observations after being converted into an explicit spatial conductivity model\.
The workflow contains four stages \(Figure[1](https://arxiv.org/html/2609.20978#S2.F1)\)\. First, a text\-conditioned foundation model converts each written interpretation into an ensemble of geologic images\. Second, a variational autoencoder \(VAE\) learns an interpretation\-specific low\-dimensional representation of that ensemble\. Third, a supervised inverse network maps the head observations to the latent space of each interpretation\. Finally, the frozen decoder converts the inferred latent vector to a conductivity field, a forward flow simulation predicts heads at the monitoring locations, and the resulting residual is used to rank the candidate interpretations\.
For each candidateCkC\_\{k\},fk:ℝNobs→ℝdzf\_\{k\}:\\mathbb\{R\}^\{N\_\{\\mathrm\{obs\}\}\}\\rightarrow\\mathbb\{R\}^\{d\_\{z\}\}denotes the interpretation\-specific inverse network\. The vector𝐳^ik∈ℝdz\\hat\{\\mathbf\{z\}\}\_\{ik\}\\in\\mathbb\{R\}^\{d\_\{z\}\}, wheredz=128d\_\{z\}=128, is the latent representation predicted byfkf\_\{k\}from the observation vector𝐝i\\mathbf\{d\}\_\{i\}\. The decoderDkD\_\{k\}maps𝐳^ik\\hat\{\\mathbf\{z\}\}\_\{ik\}to a normalized geologic image associated with candidateCkC\_\{k\}, andℳ\\mathcal\{M\}maps the image intensities to the inferred cellwise natural log\-conductivity field𝐘^ik\\hat\{\\mathbf\{Y\}\}\_\{ik\}, where𝐘=ln𝐊\\mathbf\{Y\}=\\ln\\mathbf\{K\}\. The forward operatorℱ\\mathcal\{F\}solves the flow equation and produces the complete simulated head field𝐡^ik∈ℝNgrid\\hat\{\\mathbf\{h\}\}\_\{ik\}\\in\\mathbb\{R\}^\{N\_\{\\mathrm\{grid\}\}\}, whereNgrid=16,384N\_\{\\mathrm\{grid\}\}=16\{,\}384\. Finally, the fixed observation operatorℋ:ℝNgrid→ℝNobs\\mathcal\{H\}:\\mathbb\{R\}^\{N\_\{\\mathrm\{grid\}\}\}\\rightarrow\\mathbb\{R\}^\{N\_\{\\mathrm\{obs\}\}\}extracts the simulated heads at the monitoring locations, producing the predicted observation vector𝐝^ik\\hat\{\\mathbf\{d\}\}\_\{ik\}\. The complete prediction sequence is
𝐳^ik\\displaystyle\\hat\{\\mathbf\{z\}\}\_\{ik\}=fk\(𝐝i\),\\displaystyle=f\_\{k\}\(\\mathbf\{d\}\_\{i\}\),\(1\)𝐘^ik\\displaystyle\\hat\{\\mathbf\{Y\}\}\_\{ik\}=ℳ\[Dk\(𝐳^ik\)\],\\displaystyle=\\mathcal\{M\}\\\!\\left\[D\_\{k\}\(\\hat\{\\mathbf\{z\}\}\_\{ik\}\)\\right\],𝐡^ik\\displaystyle\\hat\{\\mathbf\{h\}\}\_\{ik\}=ℱ\(𝐘^ik\),\\displaystyle=\\mathcal\{F\}\(\\hat\{\\mathbf\{Y\}\}\_\{ik\}\),𝐝^ik\\displaystyle\\hat\{\\mathbf\{d\}\}\_\{ik\}=ℋ\(𝐡^ik\)\.\\displaystyle=\\mathcal\{H\}\(\\hat\{\\mathbf\{h\}\}\_\{ik\}\)\.
Here, a hat denotes an inferred or predicted quantity\. Becausefkf\_\{k\}andDkD\_\{k\}are trained separately for each candidate interpretation,𝐘^ik\\hat\{\\mathbf\{Y\}\}\_\{ik\}is restricted to the spatial patterns represented by the image ensemble associated withCkC\_\{k\}\. The same observation vector𝐝i\\mathbf\{d\}\_\{i\}and observation operatorℋ\\mathcal\{H\}are used for all candidate interpretations, ensuring that differences among the predicted observations𝐝^ik\\hat\{\\mathbf\{d\}\}\_\{ik\}reflect differences among the interpretation\-specific models\.
Figure 1:Workflow for comparing competing geologic interpretations\. A written interpretation is expanded into an image ensemble by a text\-conditioned foundation model \(red\)\. An interpretation\-specific variational autoencoder learns a low\-dimensional representation of that ensemble \(green\)\. A supervised inverse network maps hydraulic\-head observations to a latent vector, and the frozen decoder maps that vector to a conductivity field \(blue\)\. Forward simulation produces predicted heads whose mismatch with the observations is converted into a relative compatibility weight across the candidate set \(black\)\.
### 2\.2From written interpretation to image ensemble
Each written interpretation is used as conditioning text for a text\-to\-image foundation model\. For interpretationCkC\_\{k\}, the model produces an ensemble𝒳k=\{𝐱kn\}n=1Nk\\mathcal\{X\}\_\{k\}=\\\{\\mathbf\{x\}\_\{kn\}\\\}\_\{n=1\}^\{N\_\{k\}\}of grayscale images with a resolution of256×256256\\times 256pixels\. Pixel intensities are represented on\[0,1\]\[0,1\]and are treated as normalized indicators of spatial variation in log conductivity: larger values correspond to more conductive regions\. The generated images are not calibrated property fields\. Rather, they represent spatial characteristics conveyed by the interpretation, including layering, contact geometry, continuity, and the presence or absence of connected high\-conductivity features\. The conversion to numerical conductivity values is imposed separately throughℳ\\mathcal\{M\}\.
Written descriptions convey depositional setting and qualitative spatial relations more readily than detailed geometry\. For applications with an irregular external boundary or a geometrically complex internal contact, an outline image can therefore be supplied as an additional conditioning input\. The text then specifies the interpreted internal architecture, whereas the outline constrains its location and geometry\. This auxiliary conditioning is used for the WIPP application described in Section 4\.2\.
### 2\.3Interpretation\-specific latent representation
One convolutional VAE[Kingma and Welling \(2014\)](https://arxiv.org/html/2609.20978#bib.bib13)is trained independently for each interpretation using only the corresponding image ensemble\. The encoder contains four convolutional layers with4×44\\times 4kernels, stride 2, unit padding, and 32, 64, 128, and 256 output channels, respectively\. Each convolution is followed by a ReLU activation\. For a256×256256\\times 256input image, the resulting16×16×25616\\times 16\\times 256feature array is flattened and passed to two linear layers that return the mean𝝁k\(𝐱\)\\boldsymbol\{\\mu\}\_\{k\}\(\\mathbf\{x\}\)and log varianceℓk\(𝐱\)\\boldsymbol\{\\ell\}\_\{k\}\(\\mathbf\{x\}\)of adz=128d\_\{z\}=128dimensional Gaussian latent distribution\. During training, a latent sample is obtained using the reparameterization
𝐳=𝝁k\+exp\(12ℓk\)⊙ϵ,ϵ∼𝒩\(𝟎,𝐈\)\.\\mathbf\{z\}=\\boldsymbol\{\\mu\}\_\{k\}\+\\exp\\\!\\left\(\\tfrac\{1\}\{2\}\\boldsymbol\{\\ell\}\_\{k\}\\right\)\\odot\\boldsymbol\{\\epsilon\},\\qquad\\boldsymbol\{\\epsilon\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)\.\(2\)The decoder applies a linear layer followed by four transposed\-convolution layers that reverse the encoder dimensions; the final sigmoid activation constrains reconstructed intensities to\[0,1\]\[0,1\]\.
The VAE objective is
ℒVAE=MAE\(𝐱^,𝐱\)\+λeE\(𝐱^,𝐱\)\+β\(t\)DKL\[qk\(𝐳∣𝐱\)∥𝒩\(𝟎,𝐈\)\],\\mathcal\{L\}\_\{\\mathrm\{VAE\}\}=\\operatorname\{MAE\}\(\\hat\{\\mathbf\{x\}\},\\mathbf\{x\}\)\+\\lambda\_\{e\}E\(\\hat\{\\mathbf\{x\}\},\\mathbf\{x\}\)\+\\beta\(t\)D\_\{\\mathrm\{KL\}\}\\\!\\left\[q\_\{k\}\(\\mathbf\{z\}\\mid\\mathbf\{x\}\)\\,\\\|\\,\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)\\right\],\(3\)whereMAE\\operatorname\{MAE\}denotes the mean absolute pixel error and
E\(𝐱^,𝐱\)=MAE\(Δ1𝐱^,Δ1𝐱\)\+MAE\(Δ2𝐱^,Δ2𝐱\)E\(\\hat\{\\mathbf\{x\}\},\\mathbf\{x\}\)=\\operatorname\{MAE\}\(\\Delta\_\{1\}\\hat\{\\mathbf\{x\}\},\\Delta\_\{1\}\\mathbf\{x\}\)\+\\operatorname\{MAE\}\(\\Delta\_\{2\}\\hat\{\\mathbf\{x\}\},\\Delta\_\{2\}\\mathbf\{x\}\)\(4\)penalizes errors in first differences along the two image axes\. This edge term reduces smoothing of contacts and other sharp features that distinguish the candidate interpretations\. We setλe=0\.1\\lambda\_\{e\}=0\.1\. The Kullback–Leibler weightβ\(t\)\\beta\(t\)increases linearly from 0 to 0\.01 over the first 100 epochs and remains at 0\.01 thereafter\.
The VAEs are trained for 3000 epochs with Adam, a learning rate of10−410^\{\-4\}, and a batch size of 32\. Each image ensemble is divided randomly into 80% training and 20% validation subsets\. Training uses stochastic latent samples, whereas validation decodes the posterior mean𝝁k\\boldsymbol\{\\mu\}\_\{k\}to remove sampling variability\. Gradients are averaged across MPI ranks\. The retained checkpoint is the one with the smallest deterministic validation reconstruction loss, defined as the sum of the pixel MAE and the weighted edge loss; the Kullback–Leibler term is not used in checkpoint selection\.
### 2\.4Flow model and observation operator
Forward simulations are performed with DPFEHM[O’Malley et al\. \(2023\)](https://arxiv.org/html/2609.20978#bib.bib16), a differentiable subsurface\-flow package implemented in Julia[Bezanson et al\. \(2017\)](https://arxiv.org/html/2609.20978#bib.bib19)and used here with Flux[Innes \(2018\)](https://arxiv.org/html/2609.20978#bib.bib18)\. The synthetic benchmark solves two\-dimensional, steady\-state, single\-phase flow,
∇⋅\[K\(𝐬\)∇h\(𝐬\)\]\+q\(𝐬\)=0,\\nabla\\\!\\cdot\\\!\\left\[K\(\\mathbf\{s\}\)\\nabla h\(\\mathbf\{s\}\)\\right\]\+q\(\\mathbf\{s\}\)=0,\(5\)on a regular128×128128\\times 128grid spanning\[−100,100\]m\[\-100,100\]\\,\\mathrm\{m\}in both horizontal directions and having unit thickness\. No internal source or sink is applied, soq=0q=0\.
A decoded imageg∈\[0,1\]g\\in\[0,1\]is sampled at the flow\-grid locations and mapped affinely to natural log conductivity,
Y=lnK=Ymin\+g\(Ymax−Ymin\),Ymin=−30,Ymax=−25\.Y=\\ln K=Y\_\{\\min\}\+g\(Y\_\{\\max\}\-Y\_\{\\min\}\),\\qquad Y\_\{\\min\}=\-30,\\qquad Y\_\{\\max\}=\-25\.\(6\)Thus, the image range spans five natural\-log units, corresponding to a conductivity ratio ofexp\(5\)≈148\\exp\(5\)\\approx 148between the endpoints\. Conductivity on the face between adjacent cellsaaandbbis evaluated using the geometric mean,
Kab=exp\(Ya\+Yb2\)\.K\_\{ab\}=\\exp\\\!\\left\(\\frac\{Y\_\{a\}\+Y\_\{b\}\}\{2\}\\right\)\.\(7\)
Prescribed heads impose a lateral gradient:h=10mh=10\\,\\mathrm\{m\}atx=−100mx=\-100\\,\\mathrm\{m\}andh=0mh=0\\,\\mathrm\{m\}atx=100mx=100\\,\\mathrm\{m\}\. The upper and lower boundaries are assignedh=0\.5mh=0\.5\\,\\mathrm\{m\}andh=0mh=0\\,\\mathrm\{m\}, respectively\. In the current implementation, the upper or lower value takes precedence at corner nodes\. A reproducible set ofNobs=250N\_\{\\mathrm\{obs\}\}=250monitoring nodes is sampled without replacement from the 16,384 grid nodes and is held fixed for inverse\-model training and comparison\.
For a synthetic field with clean monitored\-head vector𝐝i0\\mathbf\{d\}^\{0\}\_\{i\}, the comparison script generates observations as
𝐝i=𝐝i0\+𝜼i,𝜼i∼𝒩\(𝟎,ση,i2𝐈\),ση,i=0\.01\|𝐝i0\|¯,\\mathbf\{d\}\_\{i\}=\\mathbf\{d\}^\{0\}\_\{i\}\+\\boldsymbol\{\\eta\}\_\{i\},\\qquad\\boldsymbol\{\\eta\}\_\{i\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\sigma\_\{\\eta,i\}^\{2\}\\mathbf\{I\}\),\\qquad\\sigma\_\{\\eta,i\}=0\.01\\,\\overline\{\|\\mathbf\{d\}^\{0\}\_\{i\}\|\},\(8\)where\|𝐝i0\|¯\\overline\{\|\\mathbf\{d\}^\{0\}\_\{i\}\|\}is the mean absolute clean head\. The same noisy observation vector is supplied to every candidate interpretation for a given test field\. The synthetic comparison uses the noisy observations as the reference for ranking; residuals relative to the clean heads are retained only as diagnostics\.
### 2\.5Supervised latent inverse operator
An inverse networkfkf\_\{k\}is trained separately for each interpretation\. Because the monitoring locations are fixed, their heads are represented as an ordered vector\. The network contains three fully connected hidden layers with 256 units and ReLU activations, followed by a linear layer that mapsℝ250\\mathbb\{R\}^\{250\}to the 128\-dimensional latent space\.
For every training image𝐱\\mathbf\{x\}, the frozen VAE encoder provides the target latent mean𝝁k\(𝐱\)\\boldsymbol\{\\mu\}\_\{k\}\(\\mathbf\{x\}\)\. The original image, rather than its VAE reconstruction, is mapped through Equation[6](https://arxiv.org/html/2609.20978#S2.E6)to obtain the target field𝐘\\mathbf\{Y\}\. A forward simulation of that field supplies the clean head vector𝐝0\\mathbf\{d\}^\{0\}, and a noisy version is used as the inverse\-network input\. During inverse training and validation, independent Gaussian noise is regenerated for each mini\-batch with a standard deviation equal to 1% of the mean absolute head over that batch\.
For a mini\-batch of sizeBB, the inverse\-network loss is
ℒinv=λYMSE\(𝐘^,𝐘\)\+λzMSE\(𝐳^,𝝁k\)\+λr1B∑b=1B∥𝐳^b∥22,\\mathcal\{L\}\_\{\\mathrm\{inv\}\}=\\lambda\_\{Y\}\\operatorname\{MSE\}\(\\hat\{\\mathbf\{Y\}\},\\mathbf\{Y\}\)\+\\lambda\_\{z\}\\operatorname\{MSE\}\(\\hat\{\\mathbf\{z\}\},\\boldsymbol\{\\mu\}\_\{k\}\)\+\\lambda\_\{r\}\\frac\{1\}\{B\}\\sum\_\{b=1\}^\{B\}\\lVert\\hat\{\\mathbf\{z\}\}\_\{b\}\\rVert\_\{2\}^\{2\},\(9\)where𝐳^=fk\(𝐝\)\\hat\{\\mathbf\{z\}\}=f\_\{k\}\(\\mathbf\{d\}\),𝐘^=ℳ\[Dk\(𝐳^\)\]\\hat\{\\mathbf\{Y\}\}=\\mathcal\{M\}\[D\_\{k\}\(\\hat\{\\mathbf\{z\}\}\)\], andMSE\\operatorname\{MSE\}denotes the mean squared error over all elements of its arguments\. The first term constrains the decoded conductivity field, the second constrains the latent estimate to the encoder mean, and the third discourages latent vectors with large norms\. We useλY=1\\lambda\_\{Y\}=1,λz=1\\lambda\_\{z\}=1, andλr=10−2\\lambda\_\{r\}=10^\{\-2\}\. Only the inverse\-network parameters are updated; gradients pass through the frozen decoder to the latent estimate but do not update the VAE\.
The reported inverse networks do not include a head\-misfit term in Equation[9](https://arxiv.org/html/2609.20978#S2.E9)\. The forward solver is used to generate the head inputs before training and to evaluate decoded predictions during comparison, but it is not differentiated through during inverse\-network optimization\. The method is therefore described as supervised latent inversion rather than end\-to\-end physics\-informed training\.
For each interpretation, 1550 images are used for training and 50 for validation\. At each of 2000 epochs, 500 training images are sampled without replacement and divided into mini\-batches of 32\. Optimization uses Adam with a learning rate of10−410^\{\-4\}\. The checkpoint with the smallest total validation loss in Equation[9](https://arxiv.org/html/2609.20978#S2.E9)is retained\.
### 2\.6Compatibility scoring, ranking, and aggregation
For observation caseiiand candidate interpretationCkC\_\{k\}, the head RMSE is
rik=\[1Nobs∑m=1Nobs\(dim−d^ikm\)2\]1/2\.r\_\{ik\}=\\left\[\\frac\{1\}\{N\_\{\\mathrm\{obs\}\}\}\\sum\_\{m=1\}^\{N\_\{\\mathrm\{obs\}\}\}\\left\(d\_\{im\}\-\\hat\{d\}\_\{ikm\}\\right\)^\{2\}\\right\]^\{1/2\}\.\(10\)The comparison uses an effective residual scale
σi=\(σobs,i2\+σmod,i2\)1/2,σobs,i=0\.01\|𝐝i\|¯,σmod,i=0\.02\|𝐝i\|¯,\\sigma\_\{i\}=\\left\(\\sigma\_\{\\mathrm\{obs\},i\}^\{2\}\+\\sigma\_\{\\mathrm\{mod\},i\}^\{2\}\\right\)^\{1/2\},\\qquad\\sigma\_\{\\mathrm\{obs\},i\}=0\.01\\,\\overline\{\|\\mathbf\{d\}\_\{i\}\|\},\\qquad\\sigma\_\{\\mathrm\{mod\},i\}=0\.02\\,\\overline\{\|\\mathbf\{d\}\_\{i\}\|\},\(11\)where the 2% term is a prescribed model\-error floor intended to account approximately for inverse\-network, decoder, and forward\-model discrepancy\. It is a sensitivity parameter rather than an independently estimated error variance\.
The normalized error iseik=rik/σie\_\{ik\}=r\_\{ik\}/\\sigma\_\{i\}\. We convert it to a Gaussian\-form compatibility score
sik=−Nobs2eik2\.s\_\{ik\}=\-\\frac\{N\_\{\\mathrm\{obs\}\}\}\{2\}e\_\{ik\}^\{2\}\.\(12\)Given candidate prior weightsπk\\pi\_\{k\}\(uniform in the reported comparisons\), the per\-case relative compatibility weight is
wik=exp\(sik\+logπk\)∑j=1Kexp\(sij\+logπj\),w\_\{ik\}=\\frac\{\\exp\(s\_\{ik\}\+\\log\\pi\_\{k\}\)\}\{\\sum\_\{j=1\}^\{K\}\\exp\(s\_\{ij\}\+\\log\\pi\_\{j\}\)\},\(13\)evaluated using a numerically stable softmax\.
ForNNtest cases, the implementation averages scores before normalization,
s¯k=1N∑i=1Nsik,Wk=exp\(s¯k\+logπk\)∑j=1Kexp\(s¯j\+logπj\)\.\\bar\{s\}\_\{k\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}s\_\{ik\},\\qquad W\_\{k\}=\\frac\{\\exp\(\\bar\{s\}\_\{k\}\+\\log\\pi\_\{k\}\)\}\{\\sum\_\{j=1\}^\{K\}\\exp\(\\bar\{s\}\_\{j\}\+\\log\\pi\_\{j\}\)\}\.\(14\)The mean\-score aggregation is equivalent to using a geometric\-mean likelihood and prevents a collection of correlated or overlapping test cases from being treated as independent replications\. It is not a joint likelihood\. The aggregate rank is determined by increasing root\-mean\-square normalized error,
Rk=\(1N∑i=1Neik2\)1/2\.R\_\{k\}=\\left\(\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}e\_\{ik\}^\{2\}\\right\)^\{1/2\}\.\(15\)With uniformπk\\pi\_\{k\}, ranking byRkR\_\{k\}is equivalent to ranking bys¯k\\bar\{s\}\_\{k\}\. The synthetic comparison evaluatesN=595N=595test fields\.
Equations[12](https://arxiv.org/html/2609.20978#S2.E12)–[14](https://arxiv.org/html/2609.20978#S2.E14)have the algebraic form of likelihood\-based model weights, but the implementation evaluates one inverse\-network point estimate per candidate and uses the same observations both to infer that estimate and to calculate its residual\. It does not integrate over interpretation\-specific field uncertainty to evaluate the marginal likelihoodp\(𝐝i∣Ck\)p\(\\mathbf\{d\}\_\{i\}\\mid C\_\{k\}\)\. We therefore reportwikw\_\{ik\}andWkW\_\{k\}as relative compatibility or support weights, not as Bayesian posterior model probabilities\.
For the synthetic benchmark, where the reference conductivity field is known, we additionally calculate
EikK=∥𝐊^ik−𝐊i∥2∥𝐊i∥2,𝐊^ik=exp\(𝐘^ik\),𝐊i=exp\(𝐘i\)\.E^\{K\}\_\{ik\}=\\frac\{\\lVert\\hat\{\\mathbf\{K\}\}\_\{ik\}\-\\mathbf\{K\}\_\{i\}\\rVert\_\{2\}\}\{\\lVert\\mathbf\{K\}\_\{i\}\\rVert\_\{2\}\},\\qquad\\hat\{\\mathbf\{K\}\}\_\{ik\}=\\exp\(\\hat\{\\mathbf\{Y\}\}\_\{ik\}\),\\qquad\\mathbf\{K\}\_\{i\}=\\exp\(\\mathbf\{Y\}\_\{i\}\)\.\(16\)This relativeL2L\_\{2\}error is computed on conductivity, not log conductivity\. It is reported only as a synthetic diagnostic and does not enter the compatibility scores, aggregate weights, or rankings\.
## 3Synthetic benchmark
### 3\.1Interpretation classes and generated image ensembles
We constructed a controlled benchmark based on the Johansen Formation, a candidate geological CO2\-storage unit offshore Norway for which geological and flow models have been published[Eigestad et al\. \(2009\)](https://arxiv.org/html/2609.20978#bib.bib20)\. The benchmark uses three written geologic interpretations that differ in their consistency with the reference representation:
1. 1\.*Precise & Accurate*\(C1C\_\{1\}\), which closely represents the reference unit ordering, structural geometry, continuity, and contacts;
2. 2\.*Accurate*\(C2C\_\{2\}\), which preserves the general geologic setting and principal structure but omits or simplifies features that influence connectivity; and
3. 3\.*Mismatched*\(C3C\_\{3\}\), which describes a substantially different, predominantly horizontal architecture that is inconsistent with the reference structure\.
For each interpretation, the text\-conditioned foundation model generated an ensemble of 1600 grayscale images\. A separate VAE was trained for each 1600\-image ensemble, and a separate inverse network was then trained in the corresponding latent space, following Sections 2\.2–2\.5\. Consequently, each candidate pipeline can produce only conductivity structures represented within its own generated ensemble\. The three pipelines were evaluated using the same 925 synthetic reference fields and the same set of 250 monitoring locations\.
Figure[2](https://arxiv.org/html/2609.20978#S3.F2)presents one generated image from each interpretation\-specific ensemble\. The*Precise & Accurate*and*Accurate*ensembles both contain an inclined or folded conductive unit, although the latter represents its internal and bounding geometry less completely\. In contrast, the*Mismatched*ensemble is characterized primarily by subhorizontal layering\. These differences establish three distinct geologic priors before any head observations are introduced\.



Figure 2:Representative images from the three interpretation\-specific ensembles\. Left:*Precise & Accurate*\. Center:*Accurate*\. Right:*Mismatched*\. Each ensemble contains 1600 generated images\. Brighter intensities are mapped to larger natural log\-conductivity values during flow\-model construction\.
### 3\.2Representative recovered conductivity fields and head responses
Figures[3](https://arxiv.org/html/2609.20978#S3.F3)–[6](https://arxiv.org/html/2609.20978#S3.F6)illustrate the inferred fields and hydraulic responses for two representative test cases\. For each case, the same noisy head\-observation vector is supplied to all three interpretation\-specific inverse networks\. The displayed fields are natural log conductivity,Y=lnKY=\\ln K, whereas the reported relativeL2L\_\{2\}errors are calculated on conductivity,K=exp\(Y\)K=\\exp\(Y\), using Equation[16](https://arxiv.org/html/2609.20978#S2.E16)\. These field errors are available because the reference fields are known in the synthetic benchmark; they are diagnostic only and do not contribute to the interpretation ranking\.
In the first test case \(Figure[3](https://arxiv.org/html/2609.20978#S3.F3)\), the reference field contains an inclined high\-conductivity unit bounded by lower\-conductivity material\. The*Precise & Accurate*pipeline reproduces the position and overall inclination of this unit most closely, with a relative conductivity error of 0\.331\. The*Accurate*result retains the principal inclined geometry but smooths and displaces parts of the contacts, giving an error of 0\.438\. The*Mismatched*result imposes a horizontally layered structure and has an error of 0\.486\.
Figure 3:Recovered conductivity structure for representative test case 1\. Fields are displayed as natural log conductivity,Y=lnKY=\\ln K\. \(a\) Reference field\. \(b\)*Precise & Accurate*result, with relative conductivityL2L\_\{2\}error 0\.331\. \(c\)*Accurate*result, with error 0\.438\. \(d\)*Mismatched*result, with error 0\.486\.The corresponding head predictions are shown in Figure[4](https://arxiv.org/html/2609.20978#S3.F4)\. The*Precise & Accurate*prediction remains closest to the 1:1 line and gives a head RMSE of 0\.0708\. The*Accurate*and*Mismatched*predictions yield similar RMSE values of 0\.154 and 0\.161, respectively\. Thus, although the*Mismatched*field does not reproduce the reference geometry, the monitoring configuration only weakly distinguishes it from the*Accurate*interpretation for this particular case\.
Figure 4:Predicted versus reference hydraulic heads at the 250 monitoring locations for test case 1\. The dashed line denotes exact agreement\. The head RMSE values are 0\.0708 for*Precise & Accurate*, 0\.154 for*Accurate*, and 0\.161 for*Mismatched*\.The second test case contains a more pronounced lateral change in the reference\-unit geometry \(Figure[5](https://arxiv.org/html/2609.20978#S3.F5)\)\. The*Precise & Accurate*and*Accurate*pipelines recover the overall inclined conductive unit, with relative conductivity errors of 0\.345 and 0\.418, respectively, although both smooth finer geometric details\. The*Mismatched*pipeline again produces subhorizontal layering and has a substantially larger error of 0\.873\.
Figure 5:Recovered conductivity structure for representative test case 2\. Fields are displayed as natural log conductivity,Y=lnKY=\\ln K\. \(a\) Reference field\. \(b\)*Precise & Accurate*result, with relative conductivityL2L\_\{2\}error 0\.345\. \(c\)*Accurate*result, with error 0\.418\. \(d\)*Mismatched*result, with error 0\.873\. The relative errors are evaluated onK=exp\(Y\)K=\\exp\(Y\)and are not used for ranking\.For the second case, the head responses distinguish the three interpretations more clearly \(Figure[6](https://arxiv.org/html/2609.20978#S3.F6)\)\. The*Precise & Accurate*and*Accurate*predictions remain near the 1:1 line, with RMSE values of 0\.092 and 0\.112, respectively\. The*Mismatched*prediction deviates systematically over the middle and upper portions of the head range and yields an RMSE of 0\.321\. The contrast between the two representative cases demonstrates that the ability to distinguish interpretations depends on how their structural differences affect heads at the selected monitoring locations\.
Figure 6:Predicted versus reference hydraulic heads at the 250 monitoring locations for test case 2\. The dashed line denotes exact agreement\. The head RMSE values are 0\.092 for*Precise & Accurate*, 0\.112 for*Accurate*, and 0\.321 for*Mismatched*\.
### 3\.3Head\-misfit distributions and interpretation ranking
The interpretation ranking is based only on hydraulic\-head mismatch; the conductivity errors shown in Figures[3](https://arxiv.org/html/2609.20978#S3.F3)and[5](https://arxiv.org/html/2609.20978#S3.F5)are excluded from the scoring calculation\. For each test caseiiand interpretationkk, the raw head RMSErikr\_\{ik\}is normalized by the effective residual scaleσi\\sigma\_\{i\}to obtaineik=rik/σie\_\{ik\}=r\_\{ik\}/\\sigma\_\{i\}, as described in Section 2\.6\. The aggregate rank is determined from the root\-mean\-square normalized errorRkR\_\{k\}in Equation[15](https://arxiv.org/html/2609.20978#S2.E15)\.
Figure[7](https://arxiv.org/html/2609.20978#S3.F7)shows the distributions of raw head RMSE across the 925 test cases\. The mean RMSE is 0\.197 for*Precise & Accurate*, 0\.227 for*Accurate*, and 0\.280 for*Mismatched*\. The distributions therefore shift progressively toward larger errors as the candidate interpretation becomes less consistent with the reference representation\. However, all three distributions overlap substantially\. In particular, the*Accurate*and*Mismatched*interpretations cannot be distinguished reliably from every individual observation case, as also demonstrated by Figure[4](https://arxiv.org/html/2609.20978#S3.F4)\. The aggregate evidence arises from a systematic difference across the complete test set rather than complete separation in any single case\.
Figure 7:Density distributions of raw hydraulic\-head RMSE across the 925 synthetic test cases\. Mean RMSE values are 0\.197 for*Precise & Accurate*, 0\.227 for*Accurate*, and 0\.280 for*Mismatched*\. Lower RMSE indicates greater consistency with the reference heads\. The distributions overlap, but their systematic displacement produces an aggregate ordering across the complete test set\.Figure[8](https://arxiv.org/html/2609.20978#S3.F8)presents the distributions of the per\-case compatibility weightswikw\_\{ik\}defined by Equation[13](https://arxiv.org/html/2609.20978#S2.E13)\. The*Precise & Accurate*interpretation has the greatest density near a weight of one, whereas the*Mismatched*interpretation is concentrated most strongly near zero and has the smallest density near one\. The*Accurate*interpretation exhibits intermediate behavior\. Nevertheless, each distribution contains mass near both limits, confirming that individual test cases can favor different candidates\. These case\-level variations motivate the mean\-score aggregation in Equation[14](https://arxiv.org/html/2609.20978#S2.E14)\. The aggregate comparison ranks*Precise & Accurate*first,*Accurate*second, and*Mismatched*third\.
Figure 8:Density distributions of the per\-case compatibility weights for the three candidate interpretations across the 925 synthetic test cases\. The*Precise & Accurate*interpretation has the strongest concentration near one, the*Mismatched*interpretation has the strongest concentration near zero, and the*Accurate*interpretation lies between them\. The weights are relative to the three\-candidate set and should not be interpreted as Bayesian posterior model probabilities\.Taken together, the field reconstructions, head\-response comparisons, and test\-set distributions show that the framework recovers the expected aggregate ordering while retaining case\-level ambiguity\. This behavior is desirable: an interpretation is strongly disfavored only when its restricted latent representation produces systematically poorer head predictions, whereas interpretations with hydraulically similar realizations remain difficult to distinguish from a single observation case\.
## 4Application to the Culebra Dolomite at WIPP
### 4\.1Site and competing conceptual models
The Waste Isolation Pilot Plant is the U\.S\. Department of Energy’s geologic repository for transuranic waste, constructed 655 m below ground in bedded halite of the Permian Salado Formation in southeastern New Mexico\. The Culebra Dolomite Member of the overlying Rustler Formation is the most transmissive saturated unit above the repository horizon and the most likely groundwater pathway for radionuclides released by inadvertent human intrusion, which makes its conceptual model consequential for the compliance case[Beauheim \(2008\)](https://arxiv.org/html/2609.20978#bib.bib21)\.
WIPP is useful here because the conceptual model was revised, publicly, for stated reasons\. The model supporting the original 1996 license application treated all units below the water\-table aquifer as effectively isolated from surface hydrologic processes, with heads slowly declining since the end of the last glacial pluvial period roughly 14,000 years ago, and the Culebra as a fully confined unit whose heads would appear steady over the operational period\. Subsequent monitoring contradicted this\. Culebra heads were found to be rising and to respond to discrete present\-day events including major rainstorms, and recalibration failed to reproduce the high\-transmissivity offsite pathway the original model placed in the southeastern part of the site\. The revised model holds that strata overlying the Culebra in Nash Draw have lost their effectiveness as confining beds, so that heads there respond to rainfall and those changes propagate east into the confined site area; it also treats the region east of the site, where halite occupies Culebra pore space, as having very low transmissivity and very high head[Beauheim \(2008\)](https://arxiv.org/html/2609.20978#bib.bib21)\.
These two accounts are exactly the kind of input the workflow takes\. They are prose, they disagree about connectivity and confinement, and there is independent evidence about which is better supported\.
### 4\.2Conditioning on a boundary outline
The Culebra application exposed the geometric limitation noted in Section 2\.2\. Written descriptions of the two conceptual models were not sufficient on their own: the geologic boundary relevant to the comparison is too intricate to specify in words, and images generated from text alone placed structure in the wrong locations\. We therefore supplied a contour outline of the domain boundary as an additional conditioning input, so that the foundation model controls the internal architecture while the outline fixes the geometry\. Geologic contour maps are routinely available at characterized sites, which makes this a mild requirement in practice, but it does mean the workflow is not text\-only for structurally complex settings\.
Figure[9](https://arxiv.org/html/2609.20978#S4.F9)shows the outline and one generated realization for each conceptual model\. The field generated from the original description is smooth and broadly zoned, consistent with a confined unit whose properties vary gradually\. The field generated from the revised description contains a branching high\-conductivity network, consistent with the fracture connectivity and recharge pathways the revision introduced\.



Figure 9:Conditioning and generated fields for the Culebra comparison\. Left: domain boundary outline supplied to the foundation model alongside the written description\. Center: realization generated from the original conceptual model\. Right: realization generated from the revised conceptual model\.
### 4\.3Ranking result
Each conceptual model received its own autoencoder and inverse network, trained exactly as in the synthetic benchmark, and both were scored against the same observed head data\. Figure[10](https://arxiv.org/html/2609.20978#S4.F10)shows the resulting misfit distributions\. They are separated: the revised model concentrates between roughly 3\.7 and 3\.9 m while the original lies between 4\.2 and 4\.45 m, with only a thin tail of overlap\.
Applying Equations[12](https://arxiv.org/html/2609.20978#S2.E12)through[13](https://arxiv.org/html/2609.20978#S2.E13)gives a relative probability of 0\.991 for the revised conceptual model against 0\.009 for the original\. The ranking agrees with the direction of the documented revision, and it was obtained without supplying any information about which model was current or why the revision was made\. We emphasize that 0\.991 is a probability relative to a two\-member candidate set, not a statement that the revised model is correct in absolute terms\.
Figure 10:Head misfit distributions for the two Culebra conceptual models scored against observed heads\. The revised model yields systematically lower misfit, giving relative probabilities of 0\.991 and 0\.009\.
## 5Discussion
### 5\.1What the ranking establishes
The result is a comparison, not a verdict\. Equation[13](https://arxiv.org/html/2609.20978#S2.E13)normalizes over the candidate set, so the probabilities describe how the data distribute support among the interpretations that were offered\. An interpretation nobody wrote down cannot be discovered, and if every candidate is wrong the method will still return a winner\. This is the familiar limitation of Bayesian model selection over a finite set[Neuman \(2003\)](https://arxiv.org/html/2609.20978#bib.bib6), and it applies here with the additional wrinkle that the candidate set is generated from prose, so its coverage depends on how the descriptions were written\.
Within that scope, two things are gained over current practice\. The comparison happens before new characterization data are collected, which is when the choice actually has to be made and when the cost of getting it wrong is highest\. And the output carries a magnitude\. Knowing that the data favor one interpretation by 0\.991 to 0\.009 supports a different decision than knowing they favor it by 0\.55 to 0\.45, and conventional practice, which selects a conceptual model by judgment, produces no such number\.
### 5\.2Existing data are more informative than they are usually allowed to be
The synthetic benchmark and the WIPP application both discriminate using head observations that already exist\. No new wells were required\. This runs against the common assumption that conceptual\-model uncertainty can only be reduced by characterization, and it suggests that hydraulic monitoring records at characterized sites contain unexploited information about which conceptual model is right\. The reason the information is normally unavailable is not that it is absent but that there is no mechanism for confronting a written interpretation with a head record\. Supplying that mechanism is the contribution here\.
The counterpart is that discrimination is bounded by what the observation network can see\. Our two closest interpretation classes differed by 0\.213 against 0\.273 in mean misfit and overlapped heavily case by case\. Where two interpretations imply similar flow behavior at the monitoring points, no amount of processing will separate them, and the correct output is a weak ranking\.
### 5\.3Relation to other machine learning approaches
Latent\-space inversion with generative priors is established[Laloy et al\. \(2017\)](https://arxiv.org/html/2609.20978#bib.bib8);[Laloy et al\. \(2018\)](https://arxiv.org/html/2609.20978#bib.bib9);[Mo et al\. \(2020\)](https://arxiv.org/html/2609.20978#bib.bib10), and our inversion machinery is conventional\. What differs is where the prior comes from\. In existing work the training image is chosen by the modeler and fixed; here the prior is generated from the interpretation, one prior per interpretation, and comparing priors is the point of the exercise rather than a preliminary to it\.
The relationship to agent\-based approaches is complementary\.[Ma et al\. \(2026\)](https://arxiv.org/html/2609.20978#bib.bib15)used language model agents to generate and debug calibration scripts for process\-based simulators, automating implementation while operating within a predefined conceptual model; their own analysis notes that conceptual model development remains dependent on expert geological judgment\. Our workflow addresses that layer and leaves implementation alone\. The two could be composed, with agents handling model construction for whichever conceptualization ranks highest\.
### 5\.4Limitations
Several caveats bear on how far these results should be pushed\.
The images are not calibrated property fields\. Equation[6](https://arxiv.org/html/2609.20978#S2.E6)imposes a fixed six\-order\-of\-magnitude range on image intensity, and the foundation model has no knowledge of absolute permeability\. What the method compares is spatial organization, and the absolute conductivity scale must come from elsewhere\.
The model\-error floor in Equation[11](https://arxiv.org/html/2609.20978#S2.E11)affects reported confidence\. WithNobs=250N\_\{\\mathrm\{obs\}\}=250, the log\-likelihood scales with the square of the normalized misfit multiplied by 125, so probabilities are sensitive toσi\\sigma\_\{i\}\. We set the floor at 2% of the head scale, chosen so that residuals attributable to inverse\-operator and decoder error are not read as evidence against an interpretation\. A smaller floor sharpens every ranking, including rankings that should be weak\. We regard this as the most consequential tuning choice in the workflow and recommend reporting results across a range of floors\.
The forward model is steady\-state, single\-phase, and two\-dimensional\. Interpretations that differ mainly in transient or multiphase behavior would not be separated by it, though the ranking construction carries over to richer physics without modification[Pachalieva et al\. \(2022\)](https://arxiv.org/html/2609.20978#bib.bib17)\.
Finally, the priors are only as good as the descriptions\. Two geologists writing about the same conceptual model will produce different text and therefore different ensembles, and we have not characterized that sensitivity\. Establishing how much of the ranking depends on prompt wording rather than on geological content is the most important open question, and it is a prerequisite for using the method in a regulatory setting, where conceptual models are subject to independent peer review[Beauheim \(2008\)](https://arxiv.org/html/2609.20978#bib.bib21)\.
## 6Conclusions
We have described a workflow that takes competing written geologic interpretations, converts each into an ensemble of permeability fields using a text\-to\-image foundation model, learns an interpretation\-specific latent representation, inverts sparse hydraulic head observations into that latent space, and scores the resulting forward prediction against the observations to produce relative probabilities across the candidate set\.
On a synthetic benchmark with three interpretation classes, mean head misfit separated the classes at 0\.213, 0\.273, and 0\.793, and the aggregate probability assigned to the most accurate class was 1\.000\. Applied to two published conceptual models of the Culebra Dolomite at WIPP, the workflow assigned 0\.991 to the revised model and 0\.009 to the original, matching the direction of a revision that was made for independent field reasons\.
Three conclusions follow\. Written conceptual models can be treated as testable hypotheses rather than as fixed inputs\. Existing hydraulic observations carry enough information to discriminate among them, so the test can be run before new characterization is commissioned\. And text alone is not a sufficient interface to geology at structurally complex sites, where a boundary outline must be supplied alongside the description\. The main open question is the sensitivity of the ranking to how the interpretations are worded, which we regard as the necessary next step before the method is used to support a decision\.
## Open Research Section
The DPFEHM differentiable flow simulator used for all forward solves is openly available[O’Malley et al\. \(2023\)](https://arxiv.org/html/2609.20978#bib.bib16)\. Source code for autoencoder training, latent inversion, and interpretation ranking, together with the generated image ensembles and the scripts that produce every figure in this paper, are archived at https://github\.com/hrashid10/GeolEarlyRank \. Observed head data for the Culebra are available through the WIPP records described by[Beauheim \(2008\)](https://arxiv.org/html/2609.20978#bib.bib21)\.
## Conflict of Interest disclosure
The authors declare there are no conflicts of interest for this manuscript\.
###### Acknowledgements\.
This work was supported by the Laboratory Directed Research and Development program at Los Alamos National Laboratory as part of the project ”Physics\-informed Generative Inversion for Early Ranking of Competing Geologic Interpretations” under award number: 20261470IR\. Los Alamos National Laboratory is operated by Triad National Security, LLC, for the National Nuclear Security Administration of the U\.S\. Department of Energy\.
## References
- R\. L\. BeauheimCollection and integration of geoscience information to revise the WIPP hydrology conceptual model\.Technical reportTechnical ReportSAND2008\-3257P,Sandia National Laboratories,Albuquerque, NM\.Cited by:[§4\.1](https://arxiv.org/html/2609.20978#S4.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.20978#S4.SS1.p2.1),[§5\.4](https://arxiv.org/html/2609.20978#S5.SS4.p5.1),[Open Research Section](https://arxiv.org/html/2609.20978#Sx2.p1.1)\.
- Bezansonet al\.\(2017\)J\. Bezanson, A\. Edelman, S\. Karpinski, and V\. B\. ShahJulia: a fresh approach to numerical computing\.SIAM Review59\(1\),pp\. 65–98\.External Links:[Document](https://dx.doi.org/10.1137/141000671)Cited by:[§2\.4](https://arxiv.org/html/2609.20978#S2.SS4.p1.1)\.
- Bondet al\.\(2007\)C\. E\. Bond, A\. D\. Gibbs, Z\. K\. Shipton, and S\. JonesWhat do you think this is? “conceptual uncertainty” in geoscience interpretation\.GSA Today17\(11\),pp\. 4–10\.External Links:[Document](https://dx.doi.org/10.1130/GSAT01711A.1)Cited by:[§1](https://arxiv.org/html/2609.20978#S1.p1.1)\.
- Bredehoeft \(2005\)J\. BredehoeftThe conceptualization model problem—surprise\.Hydrogeology Journal13\(1\),pp\. 37–46\.External Links:[Document](https://dx.doi.org/10.1007/s10040-004-0430-5)Cited by:[§1](https://arxiv.org/html/2609.20978#S1.p2.1)\.
- Caers \(2011\)J\. CaersModeling uncertainty in the earth sciences\.John Wiley & Sons\.External Links:[Document](https://dx.doi.org/10.1002/9781119995920),ISBN 9781119992622Cited by:[§1](https://arxiv.org/html/2609.20978#S1.p1.1)\.
- Demyanovet al\.\(2019\)V\. Demyanov, D\. Arnold, T\. Rojas, and M\. ChristieUncertainty quantification in reservoir prediction: part 2—handling uncertainty in the geological scenario\.Mathematical Geosciences51,pp\. 241–264\.External Links:[Document](https://dx.doi.org/10.1007/s11004-018-9755-9)Cited by:[§1](https://arxiv.org/html/2609.20978#S1.p2.1),[§1](https://arxiv.org/html/2609.20978#S1.p3.1),[§1](https://arxiv.org/html/2609.20978#S1.p4.1)\.
- Eigestadet al\.\(2009\)G\. T\. Eigestad, H\. K\. Dahle, B\. Hellevang, F\. Riis, W\. T\. Johansen, and E\. ØianGeological modeling and simulation of CO2\{\}\_\{2\}injection in the Johansen formation\.Computational Geosciences13\(4\),pp\. 435–450\.External Links:[Document](https://dx.doi.org/10.1007/s10596-009-9153-y)Cited by:[§3\.1](https://arxiv.org/html/2609.20978#S3.SS1.p1.1)\.
- Højberg and Refsgaard \(2005\)A\. L\. Højberg and J\. C\. RefsgaardModel uncertainty—parameter uncertainty versus conceptual models\.Water Science and Technology52\(6\),pp\. 177–186\.External Links:[Document](https://dx.doi.org/10.2166/wst.2005.0166)Cited by:[§1](https://arxiv.org/html/2609.20978#S1.p2.1)\.
- Innes \(2018\)M\. InnesFlux: elegant machine learning with Julia\.Journal of Open Source Software3\(25\),pp\. 602\.External Links:[Document](https://dx.doi.org/10.21105/joss.00602)Cited by:[§2\.4](https://arxiv.org/html/2609.20978#S2.SS4.p1.1)\.
- Kingma and Welling \(2014\)D\. P\. Kingma and M\. WellingAuto\-encoding variational Bayes\.InProceedings of the 2nd International Conference on Learning Representations \(ICLR\),Note:arXiv:1312\.6114Cited by:[§2\.3](https://arxiv.org/html/2609.20978#S2.SS3.p1.1)\.
- Laloyet al\.\(2018\)E\. Laloy, R\. Hérault, D\. Jacques, and N\. LindeTraining\-image based geostatistical inversion using a spatial generative adversarial neural network\.Water Resources Research54\(1\),pp\. 381–406\.External Links:[Document](https://dx.doi.org/10.1002/2017WR022148)Cited by:[§1](https://arxiv.org/html/2609.20978#S1.p5.1),[§5\.3](https://arxiv.org/html/2609.20978#S5.SS3.p1.1)\.
- Laloyet al\.\(2017\)E\. Laloy, R\. Hérault, J\. Lee, D\. Jacques, and N\. LindeInversion using a new low\-dimensional representation of complex binary geological media based on a deep neural network\.Advances in Water Resources110,pp\. 387–405\.External Links:[Document](https://dx.doi.org/10.1016/j.advwatres.2017.09.029)Cited by:[§1](https://arxiv.org/html/2609.20978#S1.p5.1),[§5\.3](https://arxiv.org/html/2609.20978#S5.SS3.p1.1)\.
- Liet al\.\(2011\)S\. Li, Y\. Zhang, and X\. ZhangA study of conceptual model uncertainty in large\-scale CO2\{\}\_\{2\}storage simulation\.Water Resources Research47\(5\),pp\. W05534\.External Links:[Document](https://dx.doi.org/10.1029/2010WR009707)Cited by:[§1](https://arxiv.org/html/2609.20978#S1.p2.1)\.
- Maet al\.\(2026\)F\. Ma, J\. Chen, Z\. Dai, F\. Cai, and Y\. HuAutonomous inverse modeling of complex groundwater systems via a physics\-integrated large language model multi\-agent framework\.Water Research299,pp\. 125886\.External Links:[Document](https://dx.doi.org/10.1016/j.watres.2026.125886)Cited by:[§1](https://arxiv.org/html/2609.20978#S1.p6.1),[§5\.3](https://arxiv.org/html/2609.20978#S5.SS3.p2.1)\.
- Mariethoz and Caers \(2014\)G\. Mariethoz and J\. CaersMultiple\-point geostatistics: stochastic modeling with training images\.John Wiley & Sons,Chichester, UK\.External Links:[Document](https://dx.doi.org/10.1002/9781118662953)Cited by:[§1](https://arxiv.org/html/2609.20978#S1.p5.1)\.
- Moet al\.\(2020\)S\. Mo, N\. Zabaras, X\. Shi, and J\. WuIntegration of adversarial autoencoders with residual dense convolutional networks for estimation of non\-Gaussian hydraulic conductivities\.Water Resources Research56\(2\),pp\. e2019WR026082\.External Links:[Document](https://dx.doi.org/10.1029/2019WR026082)Cited by:[§1](https://arxiv.org/html/2609.20978#S1.p5.1),[§5\.3](https://arxiv.org/html/2609.20978#S5.SS3.p1.1)\.
- Neuman \(2003\)S\. P\. NeumanMaximum likelihood Bayesian averaging of uncertain model predictions\.Stochastic Environmental Research and Risk Assessment17\(5\),pp\. 291–305\.External Links:[Document](https://dx.doi.org/10.1007/s00477-003-0151-7)Cited by:[§1](https://arxiv.org/html/2609.20978#S1.p3.1),[§5\.1](https://arxiv.org/html/2609.20978#S5.SS1.p1.1)\.
- O’Malleyet al\.\(2023\)D\. O’Malley, S\. Y\. Greer, A\. Pachalieva, W\. Hao, D\. Harp, and V\. V\. VesselinovDPFEHM: a differentiable subsurface physics simulator\.Journal of Open Source Software8\(90\),pp\. 4560\.External Links:[Document](https://dx.doi.org/10.21105/joss.04560)Cited by:[§2\.4](https://arxiv.org/html/2609.20978#S2.SS4.p1.1),[Open Research Section](https://arxiv.org/html/2609.20978#Sx2.p1.1)\.
- Pachalievaet al\.\(2022\)A\. Pachalieva, D\. O’Malley, D\. R\. Harp, and H\. ViswanathanPhysics\-informed machine learning with differentiable programming for heterogeneous underground reservoir pressure management\.Scientific Reports12,pp\. 18734\.External Links:[Document](https://dx.doi.org/10.1038/s41598-022-22832-7)Cited by:[§5\.4](https://arxiv.org/html/2609.20978#S5.SS4.p4.1)\.
- Refsgaardet al\.\(2012\)J\. C\. Refsgaard, S\. Christensen, T\. O\. Sonnenborg, D\. Seifert, A\. L\. Højberg, and L\. TroldborgReview of strategies for handling geological uncertainty in groundwater flow and transport modeling\.Advances in Water Resources36,pp\. 36–50\.External Links:[Document](https://dx.doi.org/10.1016/j.advwatres.2011.04.006)Cited by:[§1](https://arxiv.org/html/2609.20978#S1.p2.1),[§1](https://arxiv.org/html/2609.20978#S1.p4.1)\.
- Refsgaardet al\.\(2006\)J\. C\. Refsgaard, J\. P\. van der Sluijs, J\. Brown, and P\. van der KeurA framework for dealing with uncertainty due to model structure error\.Advances in Water Resources29\(11\),pp\. 1586–1597\.External Links:[Document](https://dx.doi.org/10.1016/j.advwatres.2005.11.013)Cited by:[§1](https://arxiv.org/html/2609.20978#S1.p2.1)\.
- Rombachet al\.\(2022\)R\. Rombach, A\. Blattmann, D\. Lorenz, P\. Esser, and B\. OmmerHigh\-resolution image synthesis with latent diffusion models\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 10684–10695\.External Links:[Document](https://dx.doi.org/10.1109/CVPR52688.2022.01042)Cited by:[§1](https://arxiv.org/html/2609.20978#S1.p6.1)\.
- C\. Scheidt, L\. Li, and J\. Caers \(Eds\.\) \(2018\)C\. Scheidt, L\. Li, and J\. Caers \(Eds\.\)Quantifying uncertainty in subsurface systems\.Geophysical Monograph Series, Vol\.236,American Geophysical Union and John Wiley & Sons\.External Links:[Document](https://dx.doi.org/10.1002/9781119325888),ISBN 9781119325833Cited by:[§1](https://arxiv.org/html/2609.20978#S1.p2.1)\.
- Strebelle \(2002\)S\. StrebelleConditional simulation of complex geological structures using multiple\-point statistics\.Mathematical Geology34\(1\),pp\. 1–21\.External Links:[Document](https://dx.doi.org/10.1023/A%3A1014009426274)Cited by:[§1](https://arxiv.org/html/2609.20978#S1.p5.1)\.
- Wellmann and Caumon \(2018\)F\. Wellmann and G\. Caumon3\-D structural geological models: concepts, methods, and uncertainties\.InAdvances in Geophysics,C\. Schmelzbach \(Ed\.\),Vol\.59,pp\. 1–121\.External Links:[Document](https://dx.doi.org/10.1016/bs.agph.2018.09.001)Cited by:[§1](https://arxiv.org/html/2609.20978#S1.p1.1),[§1](https://arxiv.org/html/2609.20978#S1.p2.1),[§1](https://arxiv.org/html/2609.20978#S1.p3.1)\.
- Yeet al\.\(2004\)M\. Ye, S\. P\. Neuman, and P\. D\. MeyerMaximum likelihood Bayesian averaging of spatial variability models in unsaturated fractured tuff\.Water Resources Research40\(5\),pp\. W05113\.External Links:[Document](https://dx.doi.org/10.1029/2003WR002557)Cited by:[§1](https://arxiv.org/html/2609.20978#S1.p3.1)\.相似文章
语言先验对达西流反演有何贡献?一项机制性审计
本文研究了句子嵌入能否作为推理时接口,将地质知识注入到学习型达西流反演求解器中,发现文本条件化相比无文本反事实降低了81%的重构误差,且大部分增益来自类别级约束。
LithoFormer:基于Transformers的稳健地层推断框架
LithoFormer是一个基于Seq2Seq Transformer的框架,用于从测井数据中进行地层推断。它采用PatchTST主干网络,结合旋转位置编码和多任务头,共同预测地质分带和边界,实现了边界误差降低90%,并消除了地层顺序违反。
基于Flow Matching的概率反演
本文应用Flow Matching(一种生成式人工智能技术)进行地震全波形反演中的概率反演,并在合成数据集上证明了其有效性。
RankE:面向离散文本到图像生成的端到端后训练与解码器协同进化
RankE 提出了一种用于离散文本到图像生成的端到端后训练框架,通过联合优化生成器和解码器来解决潜在协变量偏移问题,同时提升对齐度与保真度。
用于地质碳封存可变操作建模与不确定性量化的多模态自回归Transformer代理模型
本文提出了一种多模态自回归Transformer代理模型,用于对地质碳封存中的可变井操作和地质不确定性进行建模,实现了准确的预测,并通过MCMC数据同化实现不确定性量化。