HarmoCore: Functional Latent Diffusion for Sparse Reconstruction of Oscillatory Wave Fields

arXiv cs.LG Papers

Summary

HarmoCore introduces a functional latent diffusion method for reconstructing oscillatory wave fields from sparse sensor data, achieving significant improvements in efficiency and accuracy with minimal sensing in 2D and 3D scenarios.

arXiv:2609.00679v1 Announce Type: new Abstract: Reconstructing oscillatory wave fields from scattered sensors is a severely underdetermined inverse problem. Beyond the challenges of general physical-field reconstruction, wave responses are complex-valued, frequency-sensitive, and highly oscillatory, while costly simulation and sensing often leave only extreme-sparse observations. Existing low-rank, operator, and diffusion approaches are largely designed for real-valued, smoother fields; dense pixel-space diffusion is particularly inefficient for oscillatory complex fields and difficult to scale to 3D. We propose HarmoCore, which places a generative prior in a compact, continuous, and structured wave-field latent. HarmoCore represents joint real--imaginary channels with Functional Tucker cores over shared continuous spatial bases, learns a frequency-conditioned core diffusion prior, and performs Diffusion Posterior Sampling directly in core space. At fixed sensor coordinates, the multilinear decoder induces an explicit likelihood guidance operator, avoiding dense pixel-space correction. Optional target-equation residual guidance further promotes physical consistency. Experiments on 2D Helmholtz, 2D synthetic wave fields, and 3D Helmholtz show substantial gains under 1%--2% sensing while remaining practical in three dimensions.
Original Article
View Cached Full Text

Cached at: 09/02/26, 06:19 AM

# HarmoCore: Functional Latent Diffusion for Sparse Reconstruction of Oscillatory Wave Fields
Source: [https://arxiv.org/html/2609.00679](https://arxiv.org/html/2609.00679)
Xinyu ZhangPanqi ChenLei ChengTing ZhangJianlong LiShikai Fang††thanks:Corresponding author: Shikai Fang $¡$fsk@zju\.edu\.cn$¿$

###### Abstract

Reconstructing oscillatory wave fields from scattered sensors is a severely underdetermined inverse problem\. Beyond the challenges of general physical\-field reconstruction, wave responses are complex\-valued, frequency\-sensitive, and highly oscillatory, while costly simulation and sensing often leave only extreme\-sparse observations\. Existing low\-rank, operator, and diffusion approaches are largely designed for real\-valued, smoother fields; dense pixel\-space diffusion is particularly inefficient for oscillatory complex fields and difficult to scale to 3D\. We proposeHarmoCore, which places a generative prior in a compact, continuous, and structured wave\-field latent\. HarmoCore represents joint real–imaginary channels with Functional Tucker cores over shared continuous spatial bases, learns a frequency\-conditioned core diffusion prior, and performs Diffusion Posterior Sampling directly in core space\. At fixed sensor coordinates, the multilinear decoder induces an explicit likelihood guidance operator, avoiding dense pixel\-space correction\. Optional target\-equation residual guidance further promotes physical consistency\. Experiments on 2D Helmholtz, 2D synthetic wave fields, and 3D Helmholtz show substantial gains under1%1\\%–2%2\\%sensing while remaining practical in three dimensions\.

1College of Information Science and Electronic Engineering, Zhejiang University

## 1Introduction

Reconstructing physical fields from sparse measurements is a fundamental inverse problem across science and engineering\([Arridge et al\. 2019](https://arxiv.org/html/2609.00679#bib.bib1);[Manohar et al\. 2018](https://arxiv.org/html/2609.00679#bib.bib20)\)\. Wave\-field reconstruction is particularly important in electromagnetic simulation\([Colton and Kress 2013](https://arxiv.org/html/2609.00679#bib.bib4)\), ocean acoustics\([Jensen et al\. 2011](https://arxiv.org/html/2609.00679#bib.bib11)\), geophysical imaging\([Virieux and Operto 2009](https://arxiv.org/html/2609.00679#bib.bib32)\), and many other sensing and modeling problems\. We focus on time\-harmonic wave fields, which describe the steady\-state response to excitation at a fixed frequency\. In these applications, high\-fidelity simulation or dense sensing can be prohibitively expensive, especially when responses must be acquired across many frequencies\. The available training fields and sensor measurements are therefore often severely limited\. Recovering a complete wave response from scattered sensors is thus both practically important and profoundly underdetermined\.

Existing sparse\-field reconstruction methods broadly include low\-rank fitting, deep sparse\-to\-dense regression, neural operators, and diffusion\-based generative reconstruction\. Low\-rank methods exploit compact spatial structure\([Dolgov, Kressner, and Strössner 2021](https://arxiv.org/html/2609.00679#bib.bib5);[Luo et al\. 2024](https://arxiv.org/html/2609.00679#bib.bib19)\); neural networks and operators learn direct field mappings\([Fukami et al\. 2021](https://arxiv.org/html/2609.00679#bib.bib7);[Li et al\. 2021](https://arxiv.org/html/2609.00679#bib.bib16);[Tran et al\. 2023](https://arxiv.org/html/2609.00679#bib.bib31);[Lu et al\. 2021](https://arxiv.org/html/2609.00679#bib.bib18);[Li et al\. 2024](https://arxiv.org/html/2609.00679#bib.bib17);[Li et al\. 2023](https://arxiv.org/html/2609.00679#bib.bib15);[Kovachki et al\. 2023](https://arxiv.org/html/2609.00679#bib.bib14)\); and diffusion priors can be combined with partial observations through posterior sampling\([Chung et al\. 2023](https://arxiv.org/html/2609.00679#bib.bib3);[Huang et al\. 2024](https://arxiv.org/html/2609.00679#bib.bib10)\)\. These paradigms have shown strong results, but most have been developed for real\-valued, relatively smooth spatial or spatiotemporal fields\. Sparse reconstruction of frequency\-sensitive, highly oscillatory complex\-valued wave fields remains largely underexplored\.

Oscillatory wave fields, however, differ sharply from the smoother data targeted by most existing methods\. In the time\-harmonic setting, the Helmholtz, frequency\-domain Maxwell, and elastodynamic equations define a complex spatial response at each frequency\. Its real and imaginary components jointly encode amplitude and phase, while sources, materials, and boundaries create nonlocal interference\. Short wavelengths and sensitivity to frequency or medium parameters produce rapid spatial variation and can shift nodes and antinodes throughout the domain, leaving local sensors weakly informative about unobserved regions\. Reconstruction is therefore not merely local interpolation: the model must infer a globally coherent phase pattern from sparse evidence, and small phase errors can alter interference across the domain\. This structure mismatches common representations\. Fixed low\-rank models can suppress frequency\-dependent modes; learned regressors and operators require broad training coverage to distinguish phase\-sensitive responses; and pixel\-space diffusion must model every rapidly varying complex value, making guidance costly and 3D scaling difficult\. Generic visual latents remain grid\-bound and are not designed for continuous\-coordinate spatial queries or coupled complex channels\. The central challenge is therefore to build a generative prior aligned with both the oscillatory field structure and its sparse observations\.

To address these challenges, we proposeHarmoCore, a functional latent diffusion framework for sparse complex wave\-field reconstruction\. We represent the real and imaginary components of each field as joint channels of a compact Functional Tucker core over shared continuous spatial bases\. The bases capture continuous coordinate dependence, while the core retains compact, sample\- and frequency\-specific coefficients\. We train a frequency\-conditioned diffusion model on these cores and perform Diffusion Posterior Sampling directly in core space\. We evaluate the shared bases at the sensor coordinates, so we can use scattered observations without rasterization\. We exploit the decoder’s multilinearity to write the sensor measurements as a linear operator on the core, which gives a closed\-form observation\-likelihood gradient\. For optional governing\-equation guidance, the same decoder provides a fixed core\-to\-field Jacobian that efficiently propagates the gradient of a possibly nonlinear residual\. We thus keep posterior correction in the compact core space instead of the dense pixel space\. Experiments on 2D Helmholtz, 2D synthetic wave fields, and 3D Helmholtz demonstrate substantial gains under1%1\\%–2%2\\%sensing\.

We summarize our contributions as follows:

- •We propose a frequency\-aware, compact representation of complex oscillatory wave fields based on Functional Tucker models, with joint real–imaginary channels and continuous\-coordinate decoding\.
- •We train a frequency\-conditioned diffusion prior in the core space and exploit the multilinear decoder to obtain a closed\-form observation\-likelihood gradient and efficient governing\-equation residual guidance\.
- •We demonstrate substantial improvements under extreme sparsity across 2D Helmholtz, 2D synthetic wave fields, and 3D Helmholtz reconstruction settings\.

## 2Preliminaries and Problem Setup

### 2\.1Sparse Reconstruction of Time\-Harmonic Wave Fields

A time\-harmonic wave field describes the steady\-state complex spatial response to excitation at angular frequencyω\\omega\. The physical fieldu~s​\(𝐫,ω\)∈ℂK\\tilde\{u\}\_\{s\}\(\\mathbf\{r\},\\omega\)\\in\\mathbb\{C\}^\{K\}encodes amplitude and phase jointly; real and imaginary parts together determine interference, node locations, and energy distribution throughout the domain\. We represent the field through stacked real and imaginary channels and formalize the sparse observation model as

us​\(𝐫,ω\)\\displaystyle u\_\{s\}\(\\mathbf\{r\},\\omega\)=\[Re\(u~s\),Im\(u~s\)\]∈ℝC,C=2K,\\displaystyle=\\bigl\[\\operatorname\{Re\}\(\\tilde\{u\}\_\{s\}\),\\,\\operatorname\{Im\}\(\\tilde\{u\}\_\{s\}\)\\bigr\]\\in\\mathbb\{R\}^\{C\},\\quad C=2K,\(1\)𝐲m\\displaystyle\\mathbf\{y\}\_\{m\}=us\(𝐫m,ω\)\+εm,εm∼𝒩\(𝟎,σobs2I\),\\displaystyle=u\_\{s\}\(\\mathbf\{r\}\_\{m\},\\omega\)\+\\varepsilon\_\{m\},\\quad\\varepsilon\_\{m\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\sigma\_\{\\mathrm\{obs\}\}^\{2\}I\),wheressindexes the field instance \(source location, material parameters, boundary configuration\),𝐫m∈Ω\\mathbf\{r\}\_\{m\}\\in\\Omegais themm\-th sensor coordinate, and𝐲m∈ℝC\\mathbf\{y\}\_\{m\}\\in\\mathbb\{R\}^\{C\}is the measured channel vector\. For scalar wave fieldsK=1K=1\(C=2C=2, one real and one imaginary channel\); in the experiments reported here all three benchmarks useK=1K=1\. The sparse observation set\{\(𝐫m,𝐲m\)\}m=1M\\\{\(\\mathbf\{r\}\_\{m\},\\mathbf\{y\}\_\{m\}\)\\\}\_\{m=1\}^\{M\}may cover only a small fraction of the evaluation grid; in our extreme\-sparse experiments, for example, the sensing ratio is as low as11%–22%\. The primary reconstruction target is the full continuous fieldus​\(⋅,ω\)u\_\{s\}\(\\cdot,\\omega\)overΩ\\Omega, evaluated on both observed and unobserved regions, with the unobserved region being the main indicator of reconstruction quality\.

The fields satisfy a governing equationℒξs,ω​u~s=fξs,ω\\mathcal\{L\}\_\{\\xi\_\{s\},\\omega\}\\,\\tilde\{u\}\_\{s\}=f\_\{\\xi\_\{s\},\\omega\}, whereℒ\\mathcal\{L\}is the differential operator \(e\.g\., the Helmholtz operator−Δ−ω2/c​\(𝐫\)2\-\\Delta\-\\omega^\{2\}/c\(\\mathbf\{r\}\)^\{2\}\),ξs\\xi\_\{s\}collects the instance\-specific medium parameters, andfξs,ωf\_\{\\xi\_\{s\},\\omega\}is the source term\. Operator and source metadata available at test time are dataset\-specific and used for optional equation guidance in Section[3\.3](https://arxiv.org/html/2609.00679#S3.SS3)\.

### 2\.2Functional Tucker Representations

Tucker decomposition approximates a tensor𝒳∈ℝn1×⋯×nd\\mathcal\{X\}\\in\\mathbb\{R\}^\{n\_\{1\}\\times\\cdots\\times n\_\{d\}\}with a compact coreG∈ℝr1×⋯×rdG\\in\\mathbb\{R\}^\{r\_\{1\}\\times\\cdots\\times r\_\{d\}\}and factor matricesUk∈ℝnk×rkU\_\{k\}\\in\\mathbb\{R\}^\{n\_\{k\}\\times r\_\{k\}\},rk≪nkr\_\{k\}\\ll n\_\{k\}, reducing the parameter count from∏nk\\prod n\_\{k\}to∏rk\+∑nk​rk\\prod r\_\{k\}\+\\sum n\_\{k\}r\_\{k\}\. Functional Tucker\([Dolgov, Kressner, and Strössner 2021](https://arxiv.org/html/2609.00679#bib.bib5);[Fang et al\. 2024](https://arxiv.org/html/2609.00679#bib.bib6)\)replaces the discrete row\-lookupUk​\[𝐢\]U\_\{k\}\[\\mathbf\{i\}\]with a continuous coordinate\-evaluable basis functionϕk:ℝ→ℝrk\\phi\_\{k\}:\\mathbb\{R\}\\to\\mathbb\{R\}^\{r\_\{k\}\}:

u\(𝐫;G\)=G×1ϕ1\(r1\)×2⋯×dϕd\(rd\),u\(\\mathbf\{r\};\\,G\)=G\\times\_\{1\}\\phi\_\{1\}\(r\_\{1\}\)\\times\_\{2\}\\cdots\\times\_\{d\}\\phi\_\{d\}\(r\_\{d\}\),\(2\)where×k\\times\_\{k\}denotes mode\-kkcontraction\. At any fixed coordinate𝐫\\mathbf\{r\}, the mappingG↦u⁡\(𝐫,G\)G\\mapsto u\(\\mathbf\{r\};\\,G\)is multilinear—linear in each mode separately\. Becauseϕk\\phi\_\{k\}can be evaluated at arbitrary real\-valued inputs, the decoder supports queries at scattered off\-grid sensor locations without rasterization\. When a single set of basis functions is shared across a field family and only the coreGGvaries per sample, the model separates common spatial structure \(captured in the bases\) from sample\-specific coefficients \(captured in the core\)\.

### 2\.3Diffusion Priors and Posterior Sampling

A diffusion model places a generative priorpθ​\(x0\)p\_\{\\theta\}\(x\_\{0\}\)over an unknown variablex0x\_\{0\}by training a denoiserϵθ​\(xt,t\)\\epsilon\_\{\\theta\}\(x\_\{t\},t\)to reverse a forward noising processq⁡\(xt∣x0\)=𝒩⁡\(α¯t​x0,\(1−α¯t\)​I\)q\(x\_\{t\}\\mid x\_\{0\}\)=\\mathcal\{N\}\(\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,x\_\{0\},\(1\{\-\}\\bar\{\\alpha\}\_\{t\}\)I\)\. Given a noisy statextx\_\{t\}, the denoiser produces a clean estimatex^0,t=\(xt−1−α¯t​ϵθ​\(xt,t\)\)/α¯t\\hat\{x\}\_\{0,t\}=\(x\_\{t\}\-\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\,\\epsilon\_\{\\theta\}\(x\_\{t\},t\)\)/\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\. For an inverse problem with partial observations𝐲\\mathbf\{y\}related tox0x\_\{0\}through a forward measurement operator𝒜\\mathcal\{A\}, the inference target is the posteriorp⁡\(x0∣𝐲\)∝pθ​\(x0\)​p​\(𝐲∣x0\)p\(x\_\{0\}\\mid\\mathbf\{y\}\)\\propto p\_\{\\theta\}\(x\_\{0\}\)\\,p\(\\mathbf\{y\}\\mid x\_\{0\}\)\. Diffusion Posterior Sampling \(DPS\)\([Chung et al\. 2023](https://arxiv.org/html/2609.00679#bib.bib3)\)approximates sampling from this posterior by correcting each unconditional reverse step with a measurement\-consistency gradient evaluated on the clean estimate:

xt−1=xt−1′−ζt​∇xt‖𝐲−𝒜⁡\(x^0,t\)‖22,x\_\{t\-1\}=x^\{\\prime\}\_\{t\-1\}\-\\zeta\_\{t\}\\,\\nabla\_\{x\_\{t\}\}\\bigl\\\|\\mathbf\{y\}\-\\mathcal\{A\}\(\\hat\{x\}\_\{0,t\}\)\\bigr\\\|\_\{2\}^\{2\},\(3\)wherext−1′x^\{\\prime\}\_\{t\-1\}is the unconditional reverse sample obtained fromx^0,t\\hat\{x\}\_\{0,t\}andϵθ​\(xt,t\)\\epsilon\_\{\\theta\}\(x\_\{t\},t\), andζt\\zeta\_\{t\}is a step\-size schedule\. The efficiency and accuracy of this guidance depend critically on how cheaply the measurement operator𝒜\\mathcal\{A\}and its gradient can be evaluated on the clean estimatex^0,t\\hat\{x\}\_\{0,t\}\.

## 3Method

HarmoCore reconstructs sparse complex wave fields through two training stages and a guided inference stage\.Training stage 1:shared continuous spatial basis networks\(ϕx,ϕy\)\(\\phi\_\{x\},\\phi\_\{y\}\)and per\-field compact coresGs,ωG\_\{s,\\omega\}are learned jointly from sparse training observations, compressing the field family into a structured latent\.Training stage 2:a frequency\-conditioned diffusion model is trained on the normalized cores, capturing the distribution of valid latent coefficients\.Inference:the frozen continuous basis networks are evaluated at the test\-case sensor coordinates, yielding a precomputed observation operatorH𝒪H\_\{\\mathcal\{O\}\}; core\-space posterior sampling guided byH𝒪H\_\{\\mathcal\{O\}\}and an optional governing\-equation residual reconstructs the core, which is decoded to the continuous wave field\. Figure[1](https://arxiv.org/html/2609.00679#S3.F1)illustrates the overall pipeline\.

![Refer to caption](https://arxiv.org/html/2609.00679v1/figures/method_overview.png)Figure 1:HarmoCore pipeline overview\.\(2d field as example\)*Training*\(left and center\): sparse training observations are compressed into per\-field joint real–imaginary channel cores via shared continuous basis networks; normalized cores are used to train a frequency\-conditioned diffusion prior\.*Inference*\(right\): basis evaluation at sensor coordinates yields the precomputed observation operatorH𝒪H\_\{\\mathcal\{O\}\}; core\-space posterior sampling guided byH𝒪H\_\{\\mathcal\{O\}\}and an optional equation residual produces the reconstructed complex wave field\.### 3\.1Functional Latent Modeling of Complex Wave Fields

HarmoCore instantiates the Functional Tucker representation of Eq\. \([2](https://arxiv.org/html/2609.00679#S2.E2)\) with a*joint real–imaginary channel core*\. For each \(sample, frequency\) pair\(s,ω\)\(s,\\omega\), the field is encoded by a single core tensor and decoded as

Gs,ω\\displaystyle G\_\{s,\\omega\}∈ℝRx×Ry×C,C=2K,\\displaystyle\\in\\mathbb\{R\}^\{R\_\{x\}\\times R\_\{y\}\\times C\},\\quad C=2K,\(4\)us​\(𝐫,Gs,ω\)\\displaystyle u\_\{s\}\(\\mathbf\{r\};\\,G\_\{s,\\omega\}\)=Gs,ω×1ϕx\(rx\)×2ϕy\(ry\),\\displaystyle=G\_\{s,\\omega\}\\times\_\{1\}\\phi\_\{x\}\(r\_\{x\}\)\\times\_\{2\}\\phi\_\{y\}\(r\_\{y\}\),where channels1,…,K1,\\ldots,Kare the real parts of theKKfield components and channelsK\+1,…,2​KK\{\+\}1,\\ldots,2Kare the imaginary parts\. For scalar wave fieldsK=1K=1\(C=2C=2\)\. In 3D a third basis networkϕz\\phi\_\{z\}is added and the core lives inℝRx×Ry×Rz×C\\mathbb\{R\}^\{R\_\{x\}\\times R\_\{y\}\\times R\_\{z\}\\times C\}\. A single pair of sine\-activated MLP networks\(ϕx,ϕy\)\(\\phi\_\{x\},\\phi\_\{y\}\)\([Sitzmann et al\. 2020](https://arxiv.org/html/2609.00679#bib.bib25)\)is*shared*across all samples, all frequencies, and all channels; onlyGs,ωG\_\{s,\\omega\}is sample\- and frequency\-specific\. Each \(sample, frequency\) pair therefore yields aC×Rx×RyC\{\\times\}R\_\{x\}\{\\times\}R\_\{y\}core \(R=Rx​RyR=R\_\{x\}R\_\{y\}coefficients per channel\), which for the ranks used in our experiments \(Section[5\.1](https://arxiv.org/html/2609.00679#S5.SS1)\) is far smaller than a full\-resolution field of sizeC×H×WC\{\\times\}H\{\\times\}W\.

#### Joint optimization of bases and cores\.

We learn\(ϕx,ϕy\)\(\\phi\_\{x\},\\phi\_\{y\}\)and all cores simultaneously by Adam optimization\. Letgc=vec⁡\(G⁡\[⋅,⋅,c\]\)∈ℝRg\_\{c\}=\\operatorname\{vec\}\(G\[\\cdot,\\cdot,c\]\)\\in\\mathbb\{R\}^\{R\}denote the per\-channel vectorized core andΦ=\[ϕx​\(xi\)⊗ϕy​\(yj\)\]i​j∈ℝP×R\\Phi=\[\\phi\_\{x\}\(x\_\{i\}\)\\otimes\\phi\_\{y\}\(y\_\{j\}\)\]\_\{ij\}\\in\\mathbb\{R\}^\{P\\times R\}\(P=H×WP=H\{\\times\}W\) the full\-grid basis matrix\. Let𝒪\\mathcal\{O\}index the observation coordinates used in the training objective below; restrictingΦ\\Phito the rows indexed by𝒪\\mathcal\{O\}gives the observation operatorH𝒪=Φ⁡\[𝒪\]∈ℝM×RH\_\{\\mathcal\{O\}\}=\\Phi\[\\mathcal\{O\}\]\\in\\mathbb\{R\}^\{M\\times R\}, so thatH𝒪​gcH\_\{\\mathcal\{O\}\}g\_\{c\}evaluates the decoded field at those positions against the corresponding measurements𝐲c\\mathbf\{y\}\_\{c\}\. The training objective is

ℒFTM=12​∑c=1C𝔼s,ω​\[‖H𝒪​gc−𝐲c‖‖𝐲c‖\]\+λ​𝒮​\(G\),\\mathcal\{L\}\_\{\\mathrm\{FTM\}\}=\\frac\{1\}\{2\}\\sum\_\{c=1\}^\{C\}\\mathbb\{E\}\_\{s,\\omega\}\\\!\\left\[\\frac\{\\\|H\_\{\\mathcal\{O\}\}g\_\{c\}\-\\mathbf\{y\}\_\{c\}\\\|\}\{\\\|\\mathbf\{y\}\_\{c\}\\\|\}\\right\]\+\\lambda\\,\\mathcal\{S\}\(G\),\(5\)where𝒪\\mathcal\{O\}indexes the observed positions and𝒮⁡\(G\)\\mathcal\{S\}\(G\)is a frequency\-weighted spatial\-smoothness regularizer on the core matrices\. After training, each core is normalized per channel using global mean and standard deviation computed over all \(sample, frequency\) pairs, yielding zero\-mean, unit\-variance inputs for the diffusion model\.

### 3\.2Frequency\-Conditioned Core Diffusion

The normalized cores are used to train a frequency\-conditioned diffusion model\. A conditional UNetϵθ\\epsilon\_\{\\theta\}takes a noisy core imagegt∈ℝC×Rx×Ryg\_\{t\}\\in\\mathbb\{R\}^\{C\\times R\_\{x\}\\times R\_\{y\}\}, diffusion timesteptt, and normalized frequencyωnorm=\(ω−ωmin\)/\(ωmax−ωmin\)∈\[0,1\]\\omega\_\{\\mathrm\{norm\}\}=\(\\omega\-\\omega\_\{\\min\}\)/\(\\omega\_\{\\max\}\-\\omega\_\{\\min\}\)\\in\[0,1\]injected via FiLM conditioning at each ResBlock\. The model is trained with the standard noise\-prediction objective\([Ho, Jain, and Abbeel 2020](https://arxiv.org/html/2609.00679#bib.bib8)\)

ℒdiff\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{diff\}\}=𝔼g0,ϵ,t​\[‖ϵ−ϵθ​\(gt,t,ω\)‖22\],\\displaystyle=\\mathbb\{E\}\_\{g\_\{0\},\\epsilon,t\}\\bigl\[\\\|\\epsilon\-\\epsilon\_\{\\theta\}\(g\_\{t\},t,\\omega\)\\\|\_\{2\}^\{2\}\\bigr\],\(6\)gt\\displaystyle g\_\{t\}=α¯t​g0\+1−α¯t​ϵ,\\displaystyle=\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,g\_\{0\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\,\\epsilon,using a variance\-preserving linear schedule \(T=500T=500steps,β1=10−4\\beta\_\{1\}=10^\{\-4\},βT=2×10−2\\beta\_\{T\}=2\{\\times\}10^\{\-2\}\)\. This yields the conditional priorpθ​\(g0∣ω\)p\_\{\\theta\}\(g\_\{0\}\\mid\\omega\)over normalized joint\-channel cores\.

### 3\.3Core\-Space Posterior Reconstruction

#### Explicit observation operator\.

At test time, the frozen basis networks are evaluated on the full evaluation grid to buildΦ∈ℝP×R\\Phi\\in\\mathbb\{R\}^\{P\\times R\}once, and the observation operatorH𝒪=Φ⁡\[𝒪\]H\_\{\\mathcal\{O\}\}=\\Phi\[\\mathcal\{O\}\]is re\-instantiated for the test\-case sensor coordinates𝒪\\mathcal\{O\}\. The resulting observation loss and its core\-space gradient are

ℒobs​\(gc\)\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{obs\}\}\(g\_\{c\}\)=‖H𝒪​gc−𝐲c‖2,\\displaystyle=\\\|H\_\{\\mathcal\{O\}\}g\_\{c\}\-\\mathbf\{y\}\_\{c\}\\\|^\{2\},\(7\)∇gcℒobs\\displaystyle\\nabla\_\{g\_\{c\}\}\\mathcal\{L\}\_\{\\mathrm\{obs\}\}=H𝒪⊤​\(H𝒪​gc−𝐲c\)\.\\displaystyle=H\_\{\\mathcal\{O\}\}^\{\\top\}\(H\_\{\\mathcal\{O\}\}g\_\{c\}\-\\mathbf\{y\}\_\{c\}\)\.BothH𝒪H\_\{\\mathcal\{O\}\}and the gradient require only matrix–vector products inℝR\\mathbb\{R\}^\{R\}, andH𝒪H\_\{\\mathcal\{O\}\}is reused across all reverse steps\. The same spatial mask applies to all channels \(real and imaginary\), so a singleH𝒪H\_\{\\mathcal\{O\}\}serves both\.

#### Core\-space sampling update\.

We run a DDPM\-style reverse process with guidance applied directly to the clean estimate at each step:

g^0,t\\displaystyle\\hat\{g\}\_\{0,t\}=gt−1−α¯t​ϵ^θα¯t,ϵ^θ=ϵθ\(gt,t,ω\),\\displaystyle=\\frac\{g\_\{t\}\-\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\,\\hat\{\\epsilon\}\_\{\\theta\}\}\{\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\},\\quad\\hat\{\\epsilon\}\_\{\\theta\}=\\epsilon\_\{\\theta\}\(g\_\{t\},t,\\omega\),\(8\)g^0,t\\displaystyle\\hat\{g\}\_\{0,t\}←g^0,t−αt​\(λobs​∇g^0,tℒobs\+λeq​∇g^0,tℒeq\),\\displaystyle\\leftarrow\\hat\{g\}\_\{0,t\}\-\\alpha\_\{t\}\\\!\\left\(\\lambda\_\{\\mathrm\{obs\}\}\\nabla\_\{\\hat\{g\}\_\{0,t\}\}\\mathcal\{L\}\_\{\\mathrm\{obs\}\}\+\\lambda\_\{\\mathrm\{eq\}\}\\nabla\_\{\\hat\{g\}\_\{0,t\}\}\\mathcal\{L\}\_\{\\mathrm\{eq\}\}\\right\),gt−1\\displaystyle g\_\{t\-1\}=α¯t−1​g^0,t\+1−α¯t−1​ϵ^θ,\\displaystyle=\\sqrt\{\\bar\{\\alpha\}\_\{t\-1\}\}\\,\\hat\{g\}\_\{0,t\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\-1\}\}\\,\\hat\{\\epsilon\}\_\{\\theta\},whereαt\\alpha\_\{t\}is a step\-specific guidance weight and the noise estimateϵ^θ\\hat\{\\epsilon\}\_\{\\theta\}is not recomputed after the guidance correction\.

#### Optional governing\-equation guidance\.

For the Helmholtz benchmarks, the discretized operatorAωA\_\{\\omega\}\(sparse finite\-difference matrix\) and source fieldffare available from dataset metadata at test time\. The full\-grid decoded field per channel isu^c=Φ​gc\\hat\{u\}\_\{c\}=\\Phi g\_\{c\}; the equation residual\([Raissi, Perdikaris, and Karniadakis 2019](https://arxiv.org/html/2609.00679#bib.bib23)\)is evaluated on the decoded dense grid via autograd\. BecauseΦ\\Phiis frozen after FTM training, the decoder Jacobian is constant across all reverse steps and is reused in each guidance evaluation without re\-evaluation through the basis networks\. For the linear Helmholtz residualAω​u^\+fA\_\{\\omega\}\\hat\{u\}\+f, the core\-space gradient is∇gcℒeq=2​Φ⊤​Aω⊤​\(Aω​Φ​gc\+fc\)\\nabla\_\{g\_\{c\}\}\\mathcal\{L\}\_\{\\mathrm\{eq\}\}=2\\,\\Phi^\{\\top\}A\_\{\\omega\}^\{\\top\}\(A\_\{\\omega\}\\Phi g\_\{c\}\+f\_\{c\}\); in practice we compute this via autograd on the decoded output\. The equation term is optional and supportive: ablations in Section[5\.3](https://arxiv.org/html/2609.00679#S5.SS3)confirm that removing observation\-guided posterior sampling causes a far larger accuracy collapse than removing the equation term alone\.

## 4Related Work

Reconstructing dense fields from sparse sensors has been studied through sensor\-to\-dense networks, which pool irregular observations onto a surrogate regular input before prediction\([Fukami et al\. 2021](https://arxiv.org/html/2609.00679#bib.bib7)\), and neural operators, which learn direct field mappings that generalize across PDE parameters\([Li et al\. 2021](https://arxiv.org/html/2609.00679#bib.bib16);[Tran et al\. 2023](https://arxiv.org/html/2609.00679#bib.bib31);[Lu et al\. 2021](https://arxiv.org/html/2609.00679#bib.bib18);[Li et al\. 2024](https://arxiv.org/html/2609.00679#bib.bib17);[Li et al\. 2023](https://arxiv.org/html/2609.00679#bib.bib15);[Kovachki et al\. 2023](https://arxiv.org/html/2609.00679#bib.bib14)\)\. Both perform well when training coverage captures the relevant response patterns, but under extreme sparsity \(11%–22% sensing\) the oscillatory structure and frequency sensitivity of time\-harmonic wave fields make a single deterministic estimate depend heavily on whether training instances cover the specific frequency and boundary configuration\. A generative model that captures the distribution of valid field configurations can supply the missing structural constraint when local sensors cannot resolve the global wave pattern alone\.

Diffusion models\([Ho, Jain, and Abbeel 2020](https://arxiv.org/html/2609.00679#bib.bib8)\)supply such priors and can be combined with partial observations through posterior sampling\([Song et al\. 2021](https://arxiv.org/html/2609.00679#bib.bib29);[Chung et al\. 2023](https://arxiv.org/html/2609.00679#bib.bib3);[Kawar et al\. 2022](https://arxiv.org/html/2609.00679#bib.bib12);[Song et al\. 2022](https://arxiv.org/html/2609.00679#bib.bib28);[Song et al\. 2023](https://arxiv.org/html/2609.00679#bib.bib27)\)\. building on efficient and controllable sampling techniques\([Song, Meng, and Ermon 2021](https://arxiv.org/html/2609.00679#bib.bib26);[Ho and Salimans 2022](https://arxiv.org/html/2609.00679#bib.bib9)\); Diffusion Posterior Sampling\([Chung et al\. 2023](https://arxiv.org/html/2609.00679#bib.bib3)\)steers a learned denoiser toward measurement\-consistent states during the reverse process, and related work applies this directly in pixel space to PDE\-governed field completion\([Huang et al\. 2024](https://arxiv.org/html/2609.00679#bib.bib10)\)\. Modeling and guiding every value of the dense spatial field is costly for rapidly varying, channel\-coupled complex fields, and grows more expensive as resolution or dimensionality increases—motivating a compact, continuous representation aligned with the real/imaginary channel coupling, in the spirit of latent\-space diffusion\([Rombach et al\. 2022](https://arxiv.org/html/2609.00679#bib.bib24)\)\.

Tucker and tensor\-train methods offer compact multi\-mode decompositions of structured fields\([Kolda and Bader 2009](https://arxiv.org/html/2609.00679#bib.bib13);[Oseledets 2011](https://arxiv.org/html/2609.00679#bib.bib22);[Dolgov, Kressner, and Strössner 2021](https://arxiv.org/html/2609.00679#bib.bib5)\), and Functional Tucker models\([Fang et al\. 2024](https://arxiv.org/html/2609.00679#bib.bib6)\)replace discrete factor matrices with continuous\-coordinate basis functions\([Tancik et al\. 2020](https://arxiv.org/html/2609.00679#bib.bib30);[Mildenhall et al\. 2020](https://arxiv.org/html/2609.00679#bib.bib21)\), separating shared spatial variation \(the bases\) from sample\-specific coefficients \(the core\)\. Closest to our work,\([Chen, Sun et al\. 2025](https://arxiv.org/html/2609.00679#bib.bib2)\)combines learned Functional Tucker cores with a diffusion prior for spatiotemporal field reconstruction\. HarmoCore instead targets time\-harmonic complex wave fields under extreme\-sparse scattered sensing, conditioning the diffusion prior on frequency, using a joint real–imaginary channel core, and exploiting the multilinear decoder at fixed sensor coordinates to form an explicit observation operator for efficient core\-space posterior sampling—keeping guidance in a compact, structured latent rather than the dense pixel space\.

## 5Experiments

![Refer to caption](https://arxiv.org/html/2609.00679v1/figures/helmholtz_2d_qualitative.png)Figure 2:Qualitative comparison on 2D Helmholtz at extreme sparsity \(1%1\\%sensing\)\. HarmoCore recovers the global oscillatory structure more faithfully than competing baselines at this sensing density\.### 5\.1Experimental Setup

We evaluate on three benchmarks, each isolating a different experimental question\.2D Helmholtzis the primary benchmark and drives the main sparsity comparison across sensor densities\.2D Synthetic Wave Fieldstest whether the same core\-space formulation transfers to a physically different wave\-field family generated from a closed\-form ray model rather than a PDE solve: each field is a direct\-plus\-reflected superposition of per\-source ray terms,

u~s\(𝐫,ω\)=∑k=1Ksws,k\[a1\(𝐫,𝐫s,k\)ei​2​π​f​τ1​\(𝐫,𝐫s,k\)\\displaystyle\\tilde\{u\}\_\{s\}\(\\mathbf\{r\},\\omega\)=\\sum\_\{k=1\}^\{K\_\{s\}\}w\_\{s,k\}\\Bigl\[a\_\{1\}\(\\mathbf\{r\},\\mathbf\{r\}\_\{s,k\}\)\\,e^\{\\,i2\\pi f\\,\\tau\_\{1\}\(\\mathbf\{r\},\\mathbf\{r\}\_\{s,k\}\)\}\(9\)\+βa2\(𝐫,𝐫s,k′\)ei​2​π​f​τ2​\(𝐫,𝐫s,k′\)\],\\displaystyle\+\\beta\\,a\_\{2\}\(\\mathbf\{r\},\\mathbf\{r\}\_\{s,k\}^\{\\prime\}\)\\,e^\{\\,i2\\pi f\\,\\tau\_\{2\}\(\\mathbf\{r\},\\mathbf\{r\}\_\{s,k\}^\{\\prime\}\)\}\\Bigr\],withω=2​π​f\\omega=2\\pi f,a1,a2a\_\{1\},a\_\{2\}distance\-dependent amplitude decays,τ1,τ2\\tau\_\{1\},\\tau\_\{2\}the direct/reflected travel times, and𝐫s,k′\\mathbf\{r\}\_\{s,k\}^\{\\prime\}each source’s mirror\-image reflector \(full parameterization in Appendix[A\.1](https://arxiv.org/html/2609.00679#A1.SS1)\); having no governing equation, this benchmark receives no physics\-residual metric or equation guidance \(Section[3\.3](https://arxiv.org/html/2609.00679#S3.SS3)\)\.3D Helmholtzreuses the same governing equation on a higher\-dimensional domain to test whether the framework extends to 3D via the natural addition of a third shared basis networkϕz\\phi\_\{z\}\. All methods on a given benchmark are evaluated on the same held\-out test set\.

At each sensing ratior∈\{1%,2%,5%,10%\}r\\in\\\{1\\%,2\\%,5\\%,10\\%\\\}, the observation mask is an i\.i\.d\. Bernoulli\(rr\) mask over grid points—each grid location is retained as an observed sensor independently with probabilityrr—generated once per ratio and shared by every method on a given test case\. We report1%1\\%,2%2\\%, and5%5\\%, the extreme\-to\-moderate sparsity range that is the focus of our method\.

We compare against three categories of baselines: a continuous low\-rank tensor\-function representation, LRTFR\([Luo et al\. 2024](https://arxiv.org/html/2609.00679#bib.bib19)\); deterministic operator\-regression networks trained on dense observations, FNO\([Li et al\. 2021](https://arxiv.org/html/2609.00679#bib.bib16)\), F\-FNO\([Tran et al\. 2023](https://arxiv.org/html/2609.00679#bib.bib31)\), and VoronoiCNN\([Fukami et al\. 2021](https://arxiv.org/html/2609.00679#bib.bib7)\); and a generative diffusion\-based baseline, DiffusionPDE\([Huang et al\. 2024](https://arxiv.org/html/2609.00679#bib.bib10)\)\.

The primary metric is the relative reconstruction error \(relativeℓ2\\ell\_\{2\}norm, mean±\\pmstd over the test set\), which we refer to throughout as*Relative L2 Error*\(*Rel\. L2*in tables\); it is computed over the*full*evaluation grid including sensor locations, and the primary comparison focuses on1%1\\%–2%2\\%sensing, where the inverse problem is most underdetermined\. For the two Helmholtz benchmarks we additionally report a physics\-residual metric under the discretized Helmholtz operator; the 2D Synthetic benchmark has no governing PDE, so this column is marked “—” for all methods there\. Lower values are better for all metrics\. Unless otherwise noted, the Functional Tucker core uses ranksRx=Ry=24R\_\{x\}=R\_\{y\}=24\(Rx=Ry=Rz=24R\_\{x\}=R\_\{y\}=R\_\{z\}=24in 3D\), yieldingR=576R=576\(2D\) orR=13,824R=13\{,\}824\(3D\) coefficients per channel per \(sample, frequency\) pair \(Section[3\.1](https://arxiv.org/html/2609.00679#S3.SS1)\)\. Details of the experimental setup is given in Appendix[A](https://arxiv.org/html/2609.00679#A1)\.

Table 1:Reconstruction Relative L2 Error and physics residual across three benchmarks \(mean±\\pmstd\)\. Physics residual is a PDE residual under the Helmholtz operator for 2D/3D Helmholtz; the 2D Synthetic benchmark has no governing PDE \(fields are generated by a closed\-form ray model\), so this column is “—” for all methods there\.
### 5\.2Main Results across Benchmarks

Table[1](https://arxiv.org/html/2609.00679#S5.T1)reports reconstruction Rel\. L2 error and physics residual across all three benchmarks\. HarmoCore achieves the lowest Rel\. L2 error and best physics consistency on every benchmark at all three sensing ratios \(1%1\\%,2%2\\%, and5%5\\%\), with the largest margin over the best baseline at1%1\\%–2%2\\%sensing, where the inverse problem is least constrained\. Figure[2](https://arxiv.org/html/2609.00679#S5.F2)and Figure[4](https://arxiv.org/html/2609.00679#S5.F4)show representative qualitative examples\. Full qualitative examples are shown in Appendix[B\.4](https://arxiv.org/html/2609.00679#A2.SS4)\.

![Refer to caption](https://arxiv.org/html/2609.00679v1/figures/helmholtz_2d_rmse_vs_frequency.png)Figure 3:2D Helmholtz mean reconstruction Relative L2 Error vs\. frequencyω\\omegaat2%2\\%observation rate\. HarmoCore \(Ours\) stays low and comparatively flat across the frequency range, while baseline error fluctuates irregularly withω\\omega\.Figure[3](https://arxiv.org/html/2609.00679#S5.F3)further breaks this down by frequency at2%2\\%sensing: HarmoCore’s error stays low and comparatively flat across the tested frequency range, while the baselines fluctuate irregularly withω\\omegarather than following a stable trend, underscoring HarmoCore’s comparative robustness across the frequency range\.

![Refer to caption](https://arxiv.org/html/2609.00679v1/figures/helmholtz_3d_qualitative.png)Figure 4:Qualitative 3D Helmholtz reconstructions at1%1\\%sensing , comparing ground truth, HarmoCore, and the strongest baselines\. HarmoCore recovers the correct interference and phase structure throughout the volume, while baselines flatten detail or drift in phase away from observed sensors\.A consistent pattern across all three benchmarks is that HarmoCore is most valuable in the truly underdetermined regime: when the sensor ratio is extremely low, sparse observations alone do not constrain the full field, and the learned core\-space prior resolves this ambiguity, which is exactly where HarmoCore’s margin over every baseline is largest\.

### 5\.3Mechanism Analysis and Ablation

Table[2](https://arxiv.org/html/2609.00679#S5.T2)reports ablations on 2D Helmholtz\. Removing DPS guidance entirely—keeping only the PDE\-residual term—causes a large accuracy collapse at both1%1\\%and2%2\\%, confirming that the learned prior and observation\-guided posterior sampling are the primary source of reconstruction accuracy\. Removing only the equation term produces a smaller but notable degradation, particularly at1%1\\%sensing, which supports the interpretation that equation guidance acts as a supplementary physical regularizer rather than the primary reconstruction mechanism\.

The “w/o DPS” variant achieves a lower PDE residual than the full method despite substantially worse reconstruction accuracy\. This is expected: that variant optimizes directly toward equation consistency while receiving no constraint from observations\. A low PDE residual alone is not sufficient for correct field recovery under extreme sparsity; physical consistency must be interpreted jointly with reconstruction error\.

Table 2:Ablation on 2D Helmholtz \(lower is better\)\.Together, these ablations indicate that the dominant ingredient is the combination of a compact Functional Tucker core with observation\-guided DPS: the shared continuous basis reduces the dimensionality of the inverse problem while the diffusion prior and DPS inject the actual observations into the posterior reconstruction, whereas the PDE\-residual term acts mainly as a supporting constraint that improves physical consistency without being the main reason the method works\.

### 5\.4Additional Analyses

#### Representation capacity\.

![Refer to caption](https://arxiv.org/html/2609.00679v1/figures/rank_sweep_2d_helmholtz.png)Figure 5:Reconstruction Relative L2 Error \(mean±\\pmstd\) vs\. Functional Tucker rankRx=RyR\_\{x\}=R\_\{y\}on 2D Helmholtz at2%2\\%sensing \(102102test cases per rank\), with the diffusion\-prior architecture and training configuration held fixed across ranks\. The default rankR=24R=24\(dashed line\) attains the lowest error under this fixed training budget\.Figure[5](https://arxiv.org/html/2609.00679#S5.F5)sweeps the Functional Tucker rankRx=Ry∈\{4,8,16,24,32,48,64\}R\_\{x\}=R\_\{y\}\\in\\\{4,8,16,24,32,48,64\\\}on 2D Helmholtz at2%2\\%sensing, retraining the diffusion prior at each rank with the same architecture and training budget\. Reconstruction error is non\-monotonic in rank: it falls sharply fromR=4R\{=\}4\(0\.891±0\.4000\.891\\pm 0\.400\) toR=16R\{=\}16\(0\.509±0\.3190\.509\\pm 0\.319\), reaches its minimum at the defaultR=24R\{=\}24\(0\.068±0\.0390\.068\\pm 0\.039, matching Table[1](https://arxiv.org/html/2609.00679#S5.T1)\), stays close atR=32R\{=\}32\(0\.075±0\.0520\.075\\pm 0\.052\), and rises again atR=64R\{=\}64\(0\.246±0\.1060\.246\\pm 0\.106\)\. Since a larger core is at least as expressive as a smaller one, this reflects a fixed\-budget diffusion prior becoming increasingly under\-provisioned as the core grows, rather than an intrinsic capacity ceiling—we readR=24R\{=\}24as the best operating point under the current, rank\-independent training budget\.

#### Robustness to distribution shift\.

To test robustness to a shift in the underlying generative distribution at test time, we evaluate all methods, without retraining, on an out\-of\-distribution \(OOD\) variant of the 2D Synthetic benchmark that redraws the per\-sample wave speed from a substantially wider range \(full construction in Appendix[B\.2](https://arxiv.org/html/2609.00679#A2.SS2)\)\. Table[3](https://arxiv.org/html/2609.00679#S5.T3)reports Rel\. L2 error under this shift: HarmoCore’s error is essentially unchanged from its in\-distribution values and LRTFR is similarly robust, while the remaining baselines degrade substantially, most sharply DiffusionPDE\.

Table 3:Reconstruction Rel\. L2 error on the 2D Synthetic OOD test set \(shifted wave\-speed range, no retraining\)\.

## 6Conclusion

We presented HarmoCore, a latent generative framework for sparse wave\-field reconstruction representing complex fields as a compact Functional Tucker core over shared continuous spatial bases and performs diffusion posterior sampling directly in this core space rather than dense pixel space\. Across 2D Helmholtz, 2D synthetic wave fields, and 3D Helmholtz, this formulation is most effective in the most underdetermined regimes, where deterministic reconstruction and prior\-free low\-rank fitting fail, while maintaining better physical consistency than dense operator and pixel\-space generative baselines\. The method relies on globally learned spatial parameterization and on first obtaining a sufficiently expressive FTM representation, and the paper offers stronger evidence for sparse reconstruction than for frequency extrapolation or uncertainty calibration; extending the core\-space prior along these directions is left to future work\.

## References

- Arridge et al\. \(2019\)Arridge, S\.; Maass, P\.; Öktem, O\.; and Schönlieb, C\.\-B\. 2019\.Solving Inverse Problems Using Data\-Driven Models\.*Acta Numerica*, 28: 1–174\.
- Chen, Sun et al\. \(2025\)Chen, P\.; Sun, Y\.; et al\. 2025\.Generating Full\-field Evolution of Physical Dynamics from Irregular Sparse Observations\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*\.ArXiv:2505\.09284\.
- Chung et al\. \(2023\)Chung, H\.; Kim, J\.; Mccann, M\. T\.; Klasky, M\. L\.; and Ye, J\. C\. 2023\.Diffusion Posterior Sampling for General Noisy Inverse Problems\.In*International Conference on Learning Representations*\.
- Colton and Kress \(2013\)Colton, D\.; and Kress, R\. 2013\.*Inverse Acoustic and Electromagnetic Scattering Theory*\.Springer, 3rd edition\.
- Dolgov, Kressner, and Strössner \(2021\)Dolgov, S\.; Kressner, D\.; and Strössner, C\. 2021\.Functional Tucker Approximation Using Chebyshev Interpolation\.*SIAM Journal on Scientific Computing*, 43\(3\): A2190–A2210\.ArXiv:2007\.16126\. \.
- Fang et al\. \(2024\)Fang, S\.; Yu, X\.; Wang, Z\.; Li, S\.; Kirby, R\. M\.; and Zhe, S\. 2024\.Functional Bayesian Tucker Decomposition for Continuous\-indexed Tensor Data\.In*International Conference on Learning Representations \(ICLR\)*\.ArXiv:2311\.04829\.
- Fukami et al\. \(2021\)Fukami, K\.; Maulik, R\.; Ramachandra, N\.; Fukagata, K\.; and Taira, K\. 2021\.Global Field Reconstruction from Sparse Sensors with Voronoi Tessellation\-Assisted Deep Learning\.*Nature Machine Intelligence*, 3\(11\): 945–951\.ArXiv:2101\.00554; code: github\.com/kfukami/Voronoi\-CNN\.
- Ho, Jain, and Abbeel \(2020\)Ho, J\.; Jain, A\.; and Abbeel, P\. 2020\.Denoising Diffusion Probabilistic Models\.In*Advances in Neural Information Processing Systems*\.
- Ho and Salimans \(2022\)Ho, J\.; and Salimans, T\. 2022\.Classifier\-Free Diffusion Guidance\.*arXiv preprint*\.ArXiv:2207\.12598\.
- Huang et al\. \(2024\)Huang, J\.; Yang, G\.; Wang, Z\.; and Park, J\. J\. 2024\.DiffusionPDE: Generative PDE\-Solving Under Partial Observation\.In*Advances in Neural Information Processing Systems*\.ArXiv:2406\.17763\.
- Jensen et al\. \(2011\)Jensen, F\. B\.; Kuperman, W\. A\.; Porter, M\. B\.; and Schmidt, H\. 2011\.*Computational Ocean Acoustics*\.Springer, 2nd edition\.
- Kawar et al\. \(2022\)Kawar, B\.; Elad, M\.; Ermon, S\.; and Song, J\. 2022\.Denoising Diffusion Restoration Models\.In*Advances in Neural Information Processing Systems*\.ArXiv:2201\.11793\.
- Kolda and Bader \(2009\)Kolda, T\. G\.; and Bader, B\. W\. 2009\.Tensor Decompositions and Applications\.*SIAM Review*, 51\(3\): 455–500\.
- Kovachki et al\. \(2023\)Kovachki, N\.; Li, Z\.; Liu, B\.; Azizzadenesheli, K\.; Bhattacharya, K\.; Stuart, A\.; and Anandkumar, A\. 2023\.Neural Operator: Learning Maps Between Function Spaces\.*Journal of Machine Learning Research*, 24\.ArXiv:2108\.08481\.
- Li et al\. \(2023\)Li, Z\.; Huang, D\. Z\.; Liu, B\.; and Anandkumar, A\. 2023\.Fourier Neural Operator with Learned Deformations for PDEs on General Geometries\.*Journal of Machine Learning Research*, 24\.ArXiv:2207\.05209\.
- Li et al\. \(2021\)Li, Z\.; Kovachki, N\.; Azizzadenesheli, K\.; Liu, B\.; Bhattacharya, K\.; Stuart, A\.; and Anandkumar, A\. 2021\.Fourier Neural Operator for Parametric Partial Differential Equations\.In*International Conference on Learning Representations*\.
- Li et al\. \(2024\)Li, Z\.; Zheng, H\.; Kovachki, N\.; Jin, D\.; Chen, H\.; Liu, B\.; Azizzadenesheli, K\.; and Anandkumar, A\. 2024\.Physics\-Informed Neural Operator for Learning Partial Differential Equations\.*ACM/JMS Journal of Data Science*\.ArXiv:2111\.03794\.
- Lu et al\. \(2021\)Lu, L\.; Jin, P\.; Pang, G\.; Zhang, Z\.; and Karniadakis, G\. E\. 2021\.Learning Nonlinear Operators via DeepONet Based on the Universal Approximation Theorem of Operators\.*Nature Machine Intelligence*, 3\(3\): 218–229\.
- Luo et al\. \(2024\)Luo, Y\.; Zhao, X\.; Li, Z\.; Ng, M\. K\.; and Meng, D\. 2024\.Low\-Rank Tensor Function Representation for Multi\-Dimensional Data Recovery\.*IEEE Transactions on Pattern Analysis and Machine Intelligence*, 46\(5\): 3351–3369\.
- Manohar et al\. \(2018\)Manohar, K\.; Brunton, B\. W\.; Kutz, J\. N\.; and Brunton, S\. L\. 2018\.Data\-Driven Sparse Sensor Placement for Reconstruction: Demonstrating the Benefits of Exploiting Known Patterns\.*IEEE Control Systems Magazine*\.ArXiv:1701\.07569\.
- Mildenhall et al\. \(2020\)Mildenhall, B\.; Srinivasan, P\. P\.; Tancik, M\.; Barron, J\. T\.; Ramamoorthi, R\.; and Ng, R\. 2020\.NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis\.In*European Conference on Computer Vision*\.
- Oseledets \(2011\)Oseledets, I\. V\. 2011\.Tensor\-Train Decomposition\.*SIAM Journal on Scientific Computing*, 33\(5\): 2295–2317\.
- Raissi, Perdikaris, and Karniadakis \(2019\)Raissi, M\.; Perdikaris, P\.; and Karniadakis, G\. E\. 2019\.Physics\-Informed Neural Networks: A Deep Learning Framework for Solving Forward and Inverse Problems Involving Nonlinear Partial Differential Equations\.*Journal of Computational Physics*, 378: 686–707\.
- Rombach et al\. \(2022\)Rombach, R\.; Blattmann, A\.; Lorenz, D\.; Esser, P\.; and Ommer, B\. 2022\.High\-Resolution Image Synthesis with Latent Diffusion Models\.In*IEEE/CVF Conference on Computer Vision and Pattern Recognition*\.
- Sitzmann et al\. \(2020\)Sitzmann, V\.; Martel, J\. N\. P\.; Bergman, A\. W\.; Lindell, D\. B\.; and Wetzstein, G\. 2020\.Implicit Neural Representations with Periodic Activation Functions\.In*Advances in Neural Information Processing Systems*\.
- Song, Meng, and Ermon \(2021\)Song, J\.; Meng, C\.; and Ermon, S\. 2021\.Denoising Diffusion Implicit Models\.In*International Conference on Learning Representations*\.
- Song et al\. \(2023\)Song, J\.; Vahdat, A\.; Mardani, M\.; and Kautz, J\. 2023\.Pseudoinverse\-Guided Diffusion Models for Inverse Problems\.In*International Conference on Learning Representations*\.
- Song et al\. \(2022\)Song, Y\.; Shen, L\.; Xing, L\.; and Ermon, S\. 2022\.Solving Inverse Problems in Medical Imaging with Score\-Based Generative Models\.In*International Conference on Learning Representations*\.ArXiv:2111\.08005\.
- Song et al\. \(2021\)Song, Y\.; Sohl\-Dickstein, J\.; Kingma, D\. P\.; Kumar, A\.; Ermon, S\.; and Poole, B\. 2021\.Score\-Based Generative Modeling through Stochastic Differential Equations\.In*International Conference on Learning Representations*\.
- Tancik et al\. \(2020\)Tancik, M\.; Srinivasan, P\. P\.; Mildenhall, B\.; Fridovich\-Keil, S\.; Raghavan, N\.; Singhal, U\.; Ramamoorthi, R\.; Barron, J\. T\.; and Ng, R\. 2020\.Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains\.In*Advances in Neural Information Processing Systems*\.
- Tran et al\. \(2023\)Tran, A\.; Mathews, A\. G\. d\. G\.; Xie, L\.; and Ong, C\. S\. 2023\.Factorized Fourier Neural Operators\.In*International Conference on Learning Representations*\.
- Virieux and Operto \(2009\)Virieux, J\.; and Operto, S\. 2009\.An Overview of Full\-Waveform Inversion in Exploration Geophysics\.*Geophysics*, 74\(6\): WCC1–WCC26\.

## Appendix AImplementation Details

This appendix collects implementation details that are omitted from the main paper for space reasons\.

### A\.1Datasets and Preprocessing

All three benchmarks \(2D Helmholtz, 2D synthetic wave fields, 3D Helmholtz; Section[5\.1](https://arxiv.org/html/2609.00679#S5.SS1)\) use scalar complex fields \(K=1K=1,C=2C=2: one real and one imaginary channel per \(sample, frequency\) pair; Eq\. \([1](https://arxiv.org/html/2609.00679#S2.E1)\)\)\.

All methods in Table[1](https://arxiv.org/html/2609.00679#S5.T1)are evaluated on the same held\-out test set per benchmark:153153held\-out fields for 2D Helmholtz,170170held\-out cases \(1010samples×\\times1717frequencies\) for 2D Synthetic, and410410held\-out cases \(1010samples×\\times4141frequencies\) for 3D Helmholtz\.

Sensor ratios of1%1\\%,2%2\\%, and5%5\\%of the evaluation grid are used in the main paper; the10%10\\%ratio is reported in Appendix[B\.1](https://arxiv.org/html/2609.00679#A2.SS1)\.

After FTM training, each core is normalized per channel using the global mean and standard deviation computed over all \(sample, frequency\) pairs \(Section[3\.1](https://arxiv.org/html/2609.00679#S3.SS1)\), producing zero\-mean, unit\-variance inputs for diffusion training\.

#### 2D and 3D Helmholtz generation\.

Both Helmholtz benchmarks numerically solve the PML\-damped Helmholtz equation

Δ​u~s​\(𝐫,ω\)\+\(ω/c\)2​u~s​\(𝐫,ω\)=−fs​\(𝐫,ω\),\\Delta\\tilde\{u\}\_\{s\}\(\\mathbf\{r\},\\omega\)\+\(\\omega/c\)^\{2\}\\,\\tilde\{u\}\_\{s\}\(\\mathbf\{r\},\\omega\)=\-f\_\{s\}\(\\mathbf\{r\},\\omega\),\(10\)withL=1L=1,c=1c=1,𝐫∈Ω=\[0,L\]d\\mathbf\{r\}\\in\\Omega=\[0,L\]^\{d\},d∈\{2,3\}d\\in\\\{2,3\\\}, discretized by second\-order finite differences on a uniform1282128^\{2\}grid \(2D\) or32332^\{3\}grid \(3D\)\.Δ\\Deltais realized as a complex\-coordinate\-stretched Laplacian∑k=1d∂rk\[sk\(rk\)∂rk\]\\sum\_\{k=1\}^\{d\}\\partial\_\{r\_\{k\}\}\[\\,s\_\{k\}\(r\_\{k\}\)\\,\\partial\_\{r\_\{k\}\}\\,\]that implements a quadratic\-profile PML absorbing layer of widthη​L\\eta Lon every face \(η=0\.12\\eta=0\.12in 2D,0\.150\.15in 3D\):

σ⁡\(rk\)=σmax​\(η​L−dist⁡\(rk,∂Ω\)η​L\)2,\\sigma\(r\_\{k\}\)=\\sigma\_\{\\max\}\\Bigl\(\\tfrac\{\\eta L\-\\operatorname\{dist\}\(r\_\{k\},\\partial\\Omega\)\}\{\\eta L\}\\Bigr\)^\{2\},\(11\)withσmax=50\\sigma\_\{\\max\}=50\(2D\) or4040\(3D\);u~s=0\\tilde\{u\}\_\{s\}=0is imposed on the outermost grid rows \(Dirichlet\)\. The source term is a sum ofKsK\_\{s\}random Gaussian point sources with unit\-magnitude, random\-phase complex amplitudes,

fs\(𝐫,ω\)=∑k=1Ksei​ϕs,kexp\(−∥𝐫−𝐫s,k∥2/2σsrc2\),f\_\{s\}\(\\mathbf\{r\},\\omega\)=\\sum\_\{k=1\}^\{K\_\{s\}\}e^\{i\\phi\_\{s,k\}\}\\,\\exp\\\!\\bigl\(\-\\\|\\mathbf\{r\}\-\\mathbf\{r\}\_\{s,k\}\\\|^\{2\}/2\\sigma\_\{\\mathrm\{src\}\}^\{2\}\\bigr\),\\\\\(12\)withϕs,k∼𝒰⁡\(0,2​π\)\\phi\_\{s,k\}\\sim\\mathcal\{U\}\(0,2\\pi\),Ks∈\{1,…,4\}K\_\{s\}\\in\\\{1,\\dots,4\\\}\(2D,σsrc=0\.025\\sigma\_\{\\mathrm\{src\}\}=0\.025\) orKs∈\{1,2,3\}K\_\{s\}\\in\\\{1,2,3\\\}\(3D\), source positions𝐫s,k\\mathbf\{r\}\_\{s,k\}drawn uniformly insideΩ\\Omegaaway from the PML layer, and held fixed across every frequency of sampless\(onlyω\\omegachanges the linear system, so\(𝐫s,1:Ks,ϕs,1:Ks\)\(\\mathbf\{r\}\_\{s,1\{:\}K\_\{s\}\},\\phi\_\{s,1\{:\}K\_\{s\}\}\)is the per\-sample instance descriptorξs\\xi\_\{s\}inℒξs,ω\\mathcal\{L\}\_\{\\xi\_\{s\},\\omega\}, Section[2\.1](https://arxiv.org/html/2609.00679#S2.SS1)\)\. Frequencies are sampled on a linear grid: 2D Helmholtz usesω∈\[2,52\]\\omega\\in\[2,52\]and 3D Helmholtz usesω∈\[2,22\]\\omega\\in\[2,22\]\. Each complex solution is split into real/imaginary channels and the whole dataset is divided by a single global scale \(the maximum absolute value over all samples, frequencies, and grid points\) before FTM fitting\.

#### 2D synthetic wave\-field generation\.

The synthetic benchmark replaces the PDE solve with a closed\-form direct\-plus\-reflected ray superposition on the same128×128128\\times 128grid overΩ=\[0,1\]2\\Omega=\[0,1\]^\{2\}\. Samplessdraws a fixed wave speedvs∼𝒰⁡\(0\.8,1\.2\)v\_\{s\}\\sim\\mathcal\{U\}\(0\.8,1\.2\)andKs∈\{1,2,3\}K\_\{s\}\\in\\\{1,2,3\\\}source positions𝐫s,k\\mathbf\{r\}\_\{s,k\}\(uniform inΩ\\Omega\), each held fixed across that sample’s frequency sweep; every source has a mirror\-image reflector𝐫s,k′\\mathbf\{r\}\_\{s,k\}^\{\\prime\}across the boundaryr2=0r\_\{2\}=0and a random weightws,k∼𝒰⁡\(0\.8,1\.2\)w\_\{s,k\}\\sim\\mathcal\{U\}\(0\.8,1\.2\)\. The field at frequencyω=2​π​f\\omega=2\\pi fis given by Eq\. \([9](https://arxiv.org/html/2609.00679#S5.E9)\) \(Section[5\.1](https://arxiv.org/html/2609.00679#S5.SS1)\), whereβ=0\.18\\beta=0\.18scales the reflected term; the direct/reflected travel times areτ1=\(r1/vs\)​\(1\+ε​Φ1​\(𝐫\)\)\\tau\_\{1\}=\(r\_\{1\}/v\_\{s\}\)\\bigl\(1\+\\varepsilon\\,\\Phi\_\{1\}\(\\mathbf\{r\}\)\\bigr\)andτ2=\(\(r2\+δ\)/vs\)​\(1\+ε​Φ2​\(𝐫\)\)\\tau\_\{2\}=\\bigl\(\(r\_\{2\}\+\\delta\)/v\_\{s\}\\bigr\)\\bigl\(1\+\\varepsilon\\,\\Phi\_\{2\}\(\\mathbf\{r\}\)\\bigr\), withr1,r2r\_\{1\},r\_\{2\}the distances from𝐫\\mathbf\{r\}to the source and its reflector,δ=0\.35\\delta=0\.35a fixed delay bias,ε=0\.08\\varepsilon=0\.08a phase\-perturbation strength, andΦ1,Φ2\\Phi\_\{1\},\\Phi\_\{2\}fixed smooth sinusoidal fields \(insin,cos\\sin,\\cosofr1,r2r\_\{1\},r\_\{2\}over the domain extent\) shared by every source in a sample; the amplitude decays area1=e−α1​r1/\(r1\+r0\)pa\_\{1\}=e^\{\-\\alpha\_\{1\}r\_\{1\}\}/\(r\_\{1\}\+r\_\{0\}\)^\{p\}anda2=e−α2​r2/r2a\_\{2\}=e^\{\-\\alpha\_\{2\}r\_\{2\}\}/\\sqrt\{r\_\{2\}\}, withα1=0\\alpha\_\{1\}=0,α2=0\.12\\alpha\_\{2\}=0\.12,r0=0\.4r\_\{0\}=0\.4,p=0\.3p=0\.3\. In the unperturbed limit \(ε→0\\varepsilon\\to 0\) the direct\-wave phase satisfies the eikonal relation‖∇θ‖=2​π​f/vs\\\|\\nabla\\theta\\\|=2\\pi f/v\_\{s\}; because the field is produced by this closed\-form summation rather than a governing\-equation solve, no PDE residual is available for this benchmark \(Section[3\.3](https://arxiv.org/html/2609.00679#S3.SS3)\)\. The evaluation frequency grid is1717linearly spaced points in\[1,5\]​Hz\[1,5\]\\,\\mathrm\{Hz\}; data are globally rescaled by the maximum absolute value, as in the Helmholtz benchmarks\.

### A\.2Functional Tucker Representation

The shared spatial bases\(ϕx,ϕy\)\(\\phi\_\{x\},\\phi\_\{y\}\)\(andϕz\\phi\_\{z\}in 3D\) are sine\-activated SIREN MLPs\([Sitzmann et al\. 2020](https://arxiv.org/html/2609.00679#bib.bib25)\), one pair \(triple in 3D\) shared across all samples, frequencies, and channels \(Section[3\.1](https://arxiv.org/html/2609.00679#S3.SS1)\)\. Default ranks areRx=Ry=24R\_\{x\}=R\_\{y\}=24, givingR=Rx​Ry=576R=R\_\{x\}R\_\{y\}=576coefficients per channel per \(sample, frequency\) pair, versus a full\-resolution field of sizeC×H×WC\\times H\\times W\(Eq\. \([4](https://arxiv.org/html/2609.00679#S3.E4)\)\)\. Bases and all per\-\(sample, frequency\) cores are optimized jointly with Adam, using the objective in Eq\. \([5](https://arxiv.org/html/2609.00679#S3.E5)\): a relative reconstruction loss on the observed positions plus a frequency\-weighted spatial\-smoothness regularizer𝒮⁡\(G\)\\mathcal\{S\}\(G\)\(a spatial\-gradient penalty weighted by normalized frequency\)\. Both the basis networks and the cores use learning rate10−410^\{\-4\}, with a batch size of6464samples over25,00025\{,\}000training iterations; the smoothness regularizer weight isλ=105\\lambda=10^\{5\}in Eq\. \([5](https://arxiv.org/html/2609.00679#S3.E5)\)\. Each SIREN basis network has44hidden layers of width512512\.

### A\.3Latent Diffusion Model

The diffusion prior is a conditional UNetϵθ\\epsilon\_\{\\theta\}operating on the normalized joint\-channel core imagegt∈ℝC×Rx×Ryg\_\{t\}\\in\\mathbb\{R\}^\{C\\times R\_\{x\}\\times R\_\{y\}\}\(Section[3\.2](https://arxiv.org/html/2609.00679#S3.SS2)\)\. Frequency conditioning uses the normalized scalarωnorm∈\[0,1\]\\omega\_\{\\mathrm\{norm\}\}\\in\[0,1\]injected via FiLM at each ResBlock\. Training uses the noise\-prediction objective \(Eq\. \([6](https://arxiv.org/html/2609.00679#S3.E6)\)\) under a variance\-preserving linear schedule withT=500T=500steps,β1=10−4\\beta\_\{1\}=10^\{\-4\},βT=2×10−2\\beta\_\{T\}=2\\times 10^\{\-2\}\. Optimization uses AdamW with learning rate10−410^\{\-4\}, weight decay10−610^\{\-6\}, and a batch size of3232, trained for500500epochs; no EMA of model weights is used\. We additionally tried richer frequency encodings \(Fourier features\) during development; these did not yield consistent gains over the scalar\-FiLM default and are treated as a negative result \(Appendix[B\.3](https://arxiv.org/html/2609.00679#A2.SS3)\)\.

### A\.4Posterior Sampling and Guidance

At inference, the frozen bases are evaluated once at the test\-case sensor coordinates to build the observation operatorH𝒪H\_\{\\mathcal\{O\}\}\(Eq\. \([7](https://arxiv.org/html/2609.00679#S3.E7)\)\), which is reused across all reverse steps\. Guidance is applied directly to the clean estimateg^0,t\\hat\{g\}\_\{0,t\}at each reverse step \(Eq\. \([8](https://arxiv.org/html/2609.00679#S3.E8)\)\), combining the observation\-likelihood gradient \(weightλobs\\lambda\_\{\\mathrm\{obs\}\}\) and, optionally, a governing\-equation residual gradient \(weightλeq\\lambda\_\{\\mathrm\{eq\}\}\); the step\-specific scaleαt\\alpha\_\{t\}multiplies the combined correction\. Observations are treated as noiseless \(σobs=0\\sigma\_\{\\mathrm\{obs\}\}=0\) in all experiments, soλobs\\lambda\_\{\\mathrm\{obs\}\}absorbs the likelihood scaling\. For the Helmholtz benchmarks, the equation term uses the closed\-form residual gradient given in Section[3\.3](https://arxiv.org/html/2609.00679#S3.SS3)\. The 2D Synthetic benchmark has no governing PDE, so no equation\-guidance term is applied there \(λeq=0\\lambda\_\{\\mathrm\{eq\}\}=0\); reconstruction on that benchmark uses the observation term only\.

### A\.5Evaluation Metrics

For a predicted fieldu^\\hat\{u\}and ground truthuu\(channel\-stacked real/imaginary parts, sizeC×H×W\[×D\]C\\times H\\times W\[\\times D\]\), the reconstruction error reported throughout the main tables is the relativeℓ2\\ell\_\{2\}error over the full evaluation grid, which we refer to as*Relative L2 Error*\(*Rel\. L2*\),

Rel\.L2=‖u^−u‖2‖u‖2,\\mathrm\{Rel\.\\,L2\}=\\frac\{\\\|\\hat\{u\}\-u\\\|\_\{2\}\}\{\\\|u\\\|\_\{2\}\},\(13\)computed jointly over all channels and grid points—including sensor locations—then aggregated as mean±\\pmstd over the test set; this full\-field definition is used consistently for HarmoCore and every baseline\.

For the two Helmholtz benchmarks, the physics\-residual metric evaluates the discretized governing operatorAωA\_\{\\omega\}\(Eq\. \([10](https://arxiv.org/html/2609.00679#A1.E10)\)\) against the predicted field and known sourceffon interior grid pointsℐ\\mathcal\{I\}\(the domain with the outermost boundary row/column excluded\):

PhysRes=meanℐ⁡\(\|Aω​u^\+f\|2\)meanℐ⁡\(\|f\|2\)\.\\mathrm\{PhysRes\}=\\sqrt\{\\frac\{\\operatorname\{mean\}\_\{\\mathcal\{I\}\}\\bigl\(\|A\_\{\\omega\}\\hat\{u\}\+f\|^\{2\}\\bigr\)\}\{\\operatorname\{mean\}\_\{\\mathcal\{I\}\}\\bigl\(\|f\|^\{2\}\\bigr\)\}\}\.\(14\)

### A\.6Baseline Settings

All learned baselines are trained on the same80%/20%80\\%/20\\%train/validation split of the dense training set for each benchmark \(Appendix[A\.1](https://arxiv.org/html/2609.00679#A1.SS1)\) and evaluated on the same held\-out test set as HarmoCore\.

FNO\([Li et al\. 2021](https://arxiv.org/html/2609.00679#bib.bib16)\):44Fourier layers, width6464,1212Fourier modes per spatial dimension, trained for200200epochs \(Adam, learning rate10−310^\{\-3\}, batch size6464\)\.

F\-FNO\([Tran et al\. 2023](https://arxiv.org/html/2609.00679#bib.bib31)\):44factorized Fourier layers, width6464,1212modes per spatial dimension, trained for200200epochs \(Adam, learning rate10−310^\{\-3\}, batch size6464\)\.

VoronoiCNN\([Fukami et al\. 2021](https://arxiv.org/html/2609.00679#bib.bib7)\): convolutional encoder–decoder, base width6464, trained for200200epochs \(Adam, learning rate10−310^\{\-3\}, batch size3232\)\.

DiffusionPDE\([Huang et al\. 2024](https://arxiv.org/html/2609.00679#bib.bib10)\): conditional UNet \(base width6464, conditioning dimension256256\) trained with the sameT=500T=500\-step variance\-preserving schedule as HarmoCore’s diffusion prior \(Appendix[A\.3](https://arxiv.org/html/2609.00679#A1.SS3)\), for200200epochs \(Adam, learning rate10−410^\{\-4\}, batch size3232\); inference uses DPS\-style guidance with weightζ=0\.3\\zeta=0\.3\.

LRTFR\([Luo et al\. 2024](https://arxiv.org/html/2609.00679#bib.bib19)\): at test time, per\-case core coefficients are recovered by a linear least\-squares fit of the observed positions against a fixed spatial basis, then decoded to the full grid\.

## Appendix BAdditional Results and Ablations

This appendix collects supporting quantitative and qualitative results that complement the main paper\.

### B\.1Full Quantitative Tables

Table[4](https://arxiv.org/html/2609.00679#A2.T4)extends Table[1](https://arxiv.org/html/2609.00679#S5.T1)to10%10\\%sensing, the ratio dropped from the main\-paper table for width \(Section[5\.1](https://arxiv.org/html/2609.00679#S5.SS1)\)\. The 2D Synthetic benchmark has no governing PDE at any sensing ratio \(Section[3\.3](https://arxiv.org/html/2609.00679#S3.SS3)\), so its Phys\. Res\. column is fixed at “—” for all methods, matching Table[1](https://arxiv.org/html/2609.00679#S5.T1)\.

Table 4:Reconstruction Relative L2 Error and physics residual at10%10\\%sensing \(mean±\\pmstd\), complementing Table[1](https://arxiv.org/html/2609.00679#S5.T1)\.Table 5:Reconstruction Rel\. L2 error \(mean±\\pmstd\) on the 2D Synthetic OOD test set \(shifted wave\-speed range, no retraining\), complementing the condensed Table[3](https://arxiv.org/html/2609.00679#S5.T3)\.
### B\.2Out\-of\-Distribution Robustness \(2D Synthetic\)

To probe sensitivity to a shift in the underlying generative distribution at test time, we build a second 2D Synthetic test set in which the per\-sample wave speedvsv\_\{s\}is redrawn from a substantially wider range than the in\-distribution setting used everywhere else in the paper \(vs∼𝒰⁡\(0\.8,1\.2\)v\_\{s\}\\sim\\mathcal\{U\}\(0\.8,1\.2\), empiricallyvs∈\[0\.84,1\.19\]v\_\{s\}\\in\[0\.84,1\.19\]across its held\-out samples; Appendix[A\.1](https://arxiv.org/html/2609.00679#A1.SS1)\): the out\-of\-distribution \(OOD\) test set instead drawsvsv\_\{s\}from a shifted, wider range \(empiricallyvs∈\[0\.44,1\.76\]v\_\{s\}\\in\[0\.44,1\.76\]across its1010samples\), with every other generative factor—source count, source/reflector positions and weights, the reflection coefficientβ\\beta, and the evaluation frequency grid—held fixed via the same random seed\. No model is retrained: every method uses the same checkpoint evaluated in Table[1](https://arxiv.org/html/2609.00679#S5.T1), applied unchanged to this shifted test set\. Results are reported in Table[3](https://arxiv.org/html/2609.00679#S5.T3)and discussed in Section[5\.4](https://arxiv.org/html/2609.00679#S5.SS4)\.

Table[5](https://arxiv.org/html/2609.00679#A2.T5)reports the same comparison in full \(mean±\\pmstd, all four sensing ratios\), complementing the condensed version in Table[3](https://arxiv.org/html/2609.00679#S5.T3), which omits standard deviations and10%10\\%sensing for space\. Every baseline with a dedicated OOD evaluation shows increased error under the shift, most sharply for DiffusionPDE \(e\.g\.0\.038→0\.1400\.038\\to 0\.140at5%5\\%sensing, more than a3×3\\timesincrease\) and, to a lesser degree, VoronoiCNN, FNO, and F\-FNO—consistent with these methods relying on a mapping fit to the in\-distribution wave\-speed range\. Ours is flat to marginally lower than its in\-distribution values at every sensing ratio \(e\.g\.0\.035→0\.0330\.035\\to 0\.033at5%5\\%sensing\)\.

### B\.3Frequency\-Conditioning Encoding for the Diffusion Prior: A Negative Result

The diffusion prior conditions on frequency through the normalized scalarωnorm∈\[0,1\]\\omega\_\{\\mathrm\{norm\}\}\\in\[0,1\]injected via FiLM at each ResBlock \(Appendix[A\.3](https://arxiv.org/html/2609.00679#A1.SS3)\)\. During development we also tried replacing this scalar with a richer Fourier\-style encoding ofωnorm\\omega\_\{\\mathrm\{norm\}\},

γ⁡\(ωnorm\)=\\displaystyle\\gamma\(\\omega\_\{\\mathrm\{norm\}\}\)=\{sin⁡\(k​π​ωnorm\),cos⁡\(k​π​ωnorm\)\}k=0K−1,\\displaystyle\\bigl\\\{\\sin\(k\\pi\\,\\omega\_\{\\mathrm\{norm\}\}\),\\ \\cos\(k\\pi\\,\\omega\_\{\\mathrm\{norm\}\}\)\\bigr\\\}\_\{k=0\}^\{K\-1\},\(15\)withK=8K=8frequency bands, concatenated and fed through the same FiLM conditioning path in place of the scalar default\.

Table[6](https://arxiv.org/html/2609.00679#A2.T6)compares this Fourier encoding against several variants—augmenting it with explicit low\-order polynomial terms, dropping the Fourier features entirely in favor of the polynomial terms alone, further adding a linear spectral positional term, and the scalar\-only default—on 2D Helmholtz at10%10\\%sensing\. None of the richer encodings improves over the plain scalar conditioning: the scalar\-only variant attains the lowest error \(0\.030\), while every richer variant is worse, with no consistent ordering among them—adding a linear spectral term on top of the polynomial encoding gives the worst result tested \(0\.052\), Fourier encoding alone is only slightly better \(0\.049\), and augmenting Fourier features with explicit polynomial terms or using the polynomial terms alone both land at an intermediate 0\.044\. We treat this as a negative result: for this benchmark, more elaborate frequency encodings for the diffusion prior’s conditioning do not translate into improved reconstruction accuracy, consistent with the summary in Appendix[A\.3](https://arxiv.org/html/2609.00679#A1.SS3)\.

Table 6:Reconstruction Rel\. L2 error on 2D Helmholtz at10%10\\%sensing under different frequency\-conditioning encodings for the diffusion prior\.
### B\.4Additional Qualitative Comparisons

Qualitative 3D Helmholtz reconstructions at1%1\\%sensing are shown in Figure[4](https://arxiv.org/html/2609.00679#S5.F4)in Section[5\.2](https://arxiv.org/html/2609.00679#S5.SS2); Figure[6](https://arxiv.org/html/2609.00679#A2.F6)and Figure[7](https://arxiv.org/html/2609.00679#A2.F7)below extend both Helmholtz benchmarks to additional test cases\. Figure[8](https://arxiv.org/html/2609.00679#A2.F8)shows qualitative 2D Synthetic wave\-field reconstructions, complementing the quantitative results in Section[5\.2](https://arxiv.org/html/2609.00679#S5.SS2); the main paper itself has no qualitative figure for this benchmark\.

![Refer to caption](https://arxiv.org/html/2609.00679v1/figures/2D.png)Figure 6:Additional qualitative 2D Helmholtz reconstructions across multiple test cases, with real and imaginary channels shown as separate rows per case, complementing the single\-case comparison in Figure[2](https://arxiv.org/html/2609.00679#S5.F2)\.![Refer to caption](https://arxiv.org/html/2609.00679v1/figures/3D.png)Figure 7:Additional qualitative 3D Helmholtz reconstructions across multiple test cases, with real and imaginary channels shown as separate rows per case, complementing the single\-case comparison in Figure[4](https://arxiv.org/html/2609.00679#S5.F4)\.![Refer to caption](https://arxiv.org/html/2609.00679v1/figures/synthetic_2d_qualitative.png)Figure 8:Qualitative 2D Synthetic wave\-field reconstructions across three test cases, with real and imaginary channels shown as separate rows per case\.

Similar Articles

Low-Cost High-Order Singular Value Decomposition for Tensor-Based Reconstruction from Sparse Sensor Measurements: Urban Flow and Air-Quality Applications

arXiv cs.LG

This paper introduces low-cost High-Order Singular Value Decomposition (lcHOSVD), a tensor-based method for reconstructing high-dimensional environmental fields from sparse sensor measurements. Applied to urban flow and air-quality datasets, it achieves lower reconstruction errors and greater robustness to uneven sensor distributions compared to matrix-based approaches.

Learning to Discretize: Diffusion-Based Adaptive Mesh with Spectral Guidance

arXiv cs.LG

This paper proposes a diffusion-based framework for learning adaptive mesh discretization conditioned on observed PDE dynamics, using spectral guidance and physics constraints to allocate resolution where needed. The method achieves competitive or superior performance across five PDE regimes.