From Hybrid Mechanistic--Data-Driven Modeling Toward Neuro-Symbolic AI: What, Why, and How

arXiv cs.LG Papers

Summary

This paper introduces the Hybrid-to-NeSy (H2N) framework, which systematically translates hybrid mechanistic-data-driven models into neuro-symbolic AI designs, enabling the derivation of metrics for structural violation and belief dispersion as measures of epistemic uncertainty in the mechanistic part.

arXiv:2607.22811v1 Announce Type: new Abstract: Hybrid mechanistic/data-driven models, which combine first-principles with learned components, are increasingly used in process engineering and scientific machine learning. Common hybrid modeling designs are specified primarily through their architectures and training losses, which offers a limited basis for a shared semantic interface to compare or verify them across domains, with comparatively little attention paid to epistemic uncertainty in the mechanistic part. We bridge hybrid modeling and neuro-symbolic (NeSy) AI by reconstructing these designs as instances of NeSy interface. The resulting translation, Hybrid-to-NeSy (H2N), places mechanistic knowledge on the language side, learned modules on the belief side, and validity domains together with constraints on the logic side. For each design, H2N then yields an explicit NeSy inference functional and a logic-belief decomposition. From this decomposition we derive two metrics: structural violation rate (SVR), measuring whether the learned belief respects the mechanistic structure; and belief dispersion (BD), measuring how concentrated the learned plausibility is, serving as a hybrid model's epistemic uncertainty in its mechanistic part. We instantiate H2N on a case study of a structured hybrid model for binary classification under label noise and show that models with higher SVR and BD exhibit greater variability in held-out accuracy. Under structural distribution shift, H2N further quantifies a model's uncertainty during extrapolations, whereas test accuracy reveals the same shift only post hoc.
Original Article
View Cached Full Text

Cached at: 07/28/26, 06:22 AM

# From Hybrid Mechanistic–Data-Driven Modeling Toward Neuro-Symbolic AI: What, Why, and How
Source: [https://arxiv.org/html/2607.22811](https://arxiv.org/html/2607.22811)
Moein E\. Samadi1,2, Andreas Schuppert∗,1,2

1Institute for Computational Biomedicine, RWTH Aachen University, Aachen, Germany\. 2Center for Computational Life Sciences, RWTH Aachen University, Aachen, Germany\. ∗Correspondence:aschuppert@ukaachen\.de

###### Abstract

Hybrid mechanistic–data\-driven models, which combine first\-principles with learned components, are increasingly used in process engineering and scientific machine learning\. Common hybrid modeling designs are specified primarily through their architectures and training losses, which offers a limited basis for a shared semantic interface to compare or verify them across domains, with comparatively little attention paid to epistemic uncertainty in the mechanistic part\.

We bridge hybrid modeling and neuro\-symbolic \(NeSy\) AI by reconstructing these designs as instances of NeSy interface\. The resulting translation, Hybrid\-to\-NeSy \(H2N\), places mechanistic knowledge on the language side, learned modules on the belief side, and validity domains together with constraints on the logic side\. For each design, H2N then yields an explicit NeSy inference functional and a logic–belief decomposition\.

From this decomposition we derive two metrics: structural violation rate \(SVR\), measuring whether the learned belief respects the mechanistic structure; and belief dispersion \(BD\), measuring how concentrated the learned plausibility is, serving as a hybrid model’s epistemic uncertainty in its mechanistic part\. We instantiate H2N on a case study of a structured hybrid model for binary classification under label noise and show that models with higher SVR and BD exhibit greater variability in held\-out accuracy\. Under structural distribution shift, H2N further quantifies a model’s uncertainty during extrapolations, whereas test accuracy reveals the same shift only post hoc\.

## 1Introduction

In many science and engineering domains, the system of interest is partially understood: mass and energy balances, reaction stoichiometry, structural topology, or known kinetic forms specify what must be true, while the remaining components \(constitutive closure, an unobserved rate, an environment\-dependent residua\) resist first\-principles treatment and must be learned from data\.*Hybrid mechanistic–data\-driven models*, also called semi\-parametric\[[25](https://arxiv.org/html/2607.22811#bib.bib29)\]or grey\-box\[[27](https://arxiv.org/html/2607.22811#bib.bib28)\]models, embed mechanistic knowledge while learning the rest\[[18](https://arxiv.org/html/2607.22811#bib.bib7),[26](https://arxiv.org/html/2607.22811#bib.bib4),[29](https://arxiv.org/html/2607.22811#bib.bib5),[1](https://arxiv.org/html/2607.22811#bib.bib25),[22](https://arxiv.org/html/2607.22811#bib.bib6)\]\. A canonical example: in a bioreactor model, mass balances and stoichiometry are fixed equations, while the cell\-specific growth rateμ​\(c,T\)\\mu\(c,T\)is learned from data and inserted into the otherwise\-mechanistic dynamics\. The hybrid model thus inherits the interpretability\[[5](https://arxiv.org/html/2607.22811#bib.bib2)\], data efficiency\[[7](https://arxiv.org/html/2607.22811#bib.bib3)\], and extrapolation properties\[[30](https://arxiv.org/html/2607.22811#bib.bib9)\]of mechanistic models, while using data\-driven approximation where mechanism is incomplete\.

Across application areas, hybrid modeling has converged on a handful of reusable design patterns\[[33](https://arxiv.org/html/2607.22811#bib.bib10)\]: \(i\)*serial*and*parallel*arrangements of mechanistic and learned sub\-models\[[13](https://arxiv.org/html/2607.22811#bib.bib15)\], \(ii\) residual correction by output superposition\[[4](https://arxiv.org/html/2607.22811#bib.bib14)\], \(iii\) mixture\-of\-experts with gating\[[16](https://arxiv.org/html/2607.22811#bib.bib13),[17](https://arxiv.org/html/2607.22811#bib.bib12)\]and fuzzy\-rule analogues\[[24](https://arxiv.org/html/2607.22811#bib.bib16),[31](https://arxiv.org/html/2607.22811#bib.bib18)\], and \(iv\) validity\-domain rules that modulate the learned component when inputs leave the training support\[[12](https://arxiv.org/html/2607.22811#bib.bib11),[23](https://arxiv.org/html/2607.22811#bib.bib20),[19](https://arxiv.org/html/2607.22811#bib.bib22)\]\.

Similar patterns appear throughout scientific machine learning\[[10](https://arxiv.org/html/2607.22811#bib.bib26)\], but they are described at the level of architectures and training losses\. Two consequences follow\. First, there is no widely adopted shared semantic account of what a hybrid modelmeansas an inference object, which can make cross\-domain comparison and verification difficult\. Second, hybrid models often lack a principled way to represent and propagateepistemicuncertainty in mechanistic assumptions

Neuro\-symbolic \(NeSy\) AI has long studied how to combine logical structure with learned components\[[2](https://arxiv.org/html/2607.22811#bib.bib31),[11](https://arxiv.org/html/2607.22811#bib.bib32),[15](https://arxiv.org/html/2607.22811#bib.bib38),[8](https://arxiv.org/html/2607.22811#bib.bib33),[9](https://arxiv.org/html/2607.22811#bib.bib34),[28](https://arxiv.org/html/2607.22811#bib.bib36),[3](https://arxiv.org/html/2607.22811#bib.bib35),[14](https://arxiv.org/html/2607.22811#bib.bib37)\]\. De Smet and De Raedt\[[6](https://arxiv.org/html/2607.22811#bib.bib1)\]recently consolidated this literature into a general definition: a NeSy model is a tuple\(L,μ,Ω,b𝜽\)\(L,\\mu,\\Omega,b\_\{\\boldsymbol\{\\theta\}\}\)whereLLis a language with semanticsμ\\muover an interpretation spaceΩ\\Omegaandb𝜽b\_\{\\boldsymbol\{\\theta\}\}is a belief weight, with inference defined by an integral functional

F𝜽,x​\(φ\)=∫Ω′l​\(φ,ω\)​b𝜽,x​\(ω\)​dm​\(ω\)F\_\{\\boldsymbol\{\\theta\},x\}\(\\varphi\)=\\int\_\{\\Omega^\{\\prime\}\}l\(\\varphi,\\omega\)\\,b\_\{\\boldsymbol\{\\theta\},x\}\(\\omega\)\\,\\mathrm\{d\}m\(\\omega\)\(1\)that combines a logic functionll\(evaluating queriesφ\\varphiin interpretationsω\\omega\) with the beliefb𝜽,xb\_\{\\boldsymbol\{\\theta\},x\}\. Crucially, this interface separates the*logic side*\(what isadmissible, given byLL,μ\\mu,ll, and the integration domainΩ′⊆Ω\\Omega^\{\\prime\}\\subseteq\\Omega\) from the*belief side*\(what isplausible, given byb𝜽,xb\_\{\\boldsymbol\{\\theta\},x\}\)\.

Hybrid modeling and NeSy AI address closely related integration problems but from opposite directions: hybrid modeling has rich architectural patterns but limited basis for a shared semantic interface; NeSy has the semantics but emphasises logical languages rather than mechanistic equation systems and engineering patterns\. Our work establishes a connection between these perspectives\.

#### Thesis\.

Our central claim is that hybrid mechanistic–data\-driven models can be reconstructed, systematically, as NeSy models in the sense of De Smet and De Raedt\[[6](https://arxiv.org/html/2607.22811#bib.bib1)\]\. Under this correspondence the mechanistic equations and structural constraints supply the languageLLand its semanticsμ\\mu; the learned components induce a belief functionb𝜽b\_\{\\boldsymbol\{\\theta\}\}over the unknown quantities; and validity rules and constraints are expressed either as logic functionsllor as restrictions of the integration domainΩ′\\Omega^\{\\prime\}\. Every hybrid architecture then induces an explicit inference functional of the form \([1](https://arxiv.org/html/2607.22811#S1.E1)\)\.

#### Consequences of placement\.

Where each assumption is placed within the NeSy tuple\(L,μ,Ω,b𝜽\)\(L,\\mu,\\Omega,b\_\{\\boldsymbol\{\\theta\}\}\)determines which inference functionals are well\-defined\. In particular, for structured hybrid models of Ref\.\[[7](https://arxiv.org/html/2607.22811#bib.bib3)\], keeping the structural partition on the logic side \(encoding it in the integration domainΩ′\\Omega^\{\\prime\}rather than in the beliefb𝜽b\_\{\\boldsymbol\{\\theta\}\}\) yields an additive decompositionBD=BDseen\+BDunseen\\mathrm\{BD\}=\\mathrm\{BD\}\_\{\\text\{seen\}\}\+\\mathrm\{BD\}\_\{\\text\{unseen\}\}\(Section[5](https://arxiv.org/html/2607.22811#S5)\), in whichBDunseen\\mathrm\{BD\}\_\{\\text\{unseen\}\}is computable from the coverage ofΩ′\\Omega^\{\\prime\}at deployment time, before any out\-of\-distribution \(OOD\) sample is observed\. The conventional architecture\-and\-loss descriptions of hybrid models express no such quantity, because they do not separate the structural partition from the learned predictor\.

#### Contributions\.

\(i\) We introduce a principled translation*procedure*that maps any hybrid model description into a NeSy tuple\(L,μ,Ω,b𝜽\)\(L,\\mu,\\Omega,b\_\{\\boldsymbol\{\\theta\}\}\)with an explicit inference functional \(Sections[2](https://arxiv.org/html/2607.22811#S2)–[3](https://arxiv.org/html/2607.22811#S3)\)\. \(ii\) Applying the procedure to canonical hybrid design patterns yields a compact*mapping table*\(Table[1](https://arxiv.org/html/2607.22811#S3.T1)\) that records, for each pattern, what changes in\(L,μ,Ω,b𝜽\)\(L,\\mu,\\Omega,b\_\{\\boldsymbol\{\\theta\}\}\)and in the induced functional, together with representative NeSy architectures that realize it\. \(iii\) We derive an evaluation protocol comprising two metrics that measure violations of logical structure and the learned belief’s concentration \(Section[4](https://arxiv.org/html/2607.22811#S4)\)\. \(iv\) We instantiate the translation on a structured hybrid model for binary classification under label noise and show that the resulting metrics quantify the trained model’s epistemic uncertainty in the mechanistic component at deployment time, as well as uncertainty during extrapolation, failure modes that test accuracy alone does not reveal \(Section[5](https://arxiv.org/html/2607.22811#S5)\)\.

## 2Problem Setting and Definitions

#### Hybrid mechanistic–data\-driven models\.

Letx∈𝒳x\\in\\mathcal\{X\}denote inputs,z∈𝒵z\\in\\mathcal\{Z\}a latent state, andy∈𝒴y\\in\\mathcal\{Y\}an output\. Letα\\alphadenote mechanistic parameters, and letψ\\psicollect unknown terms \(residuals, gates, noise terms\) required to complete the mechanistic description\. Here, a*closure*is any constitutive specification of such unknown terms that renders the mechanistic constraints solvable oncexxandα\\alphaare fixed\. A hybrid model is specified by

\(Mechanistic constraints\)𝒢​\(x,z;α,ψ\)=0,\\displaystyle\\quad\\mathcal\{G\}\(x,z;\\alpha,\\psi\)=0,\(2\)\(Observation model\)y=ℋ​\(x,z;α,ψ\)∈𝒴,\\displaystyle\\quad y=\\mathcal\{H\}\(x,z;\\alpha,\\psi\)\\in\\mathcal\{Y\},\(3\)\(Learned closure\)ψ=gβ​\(q​\(x,z\),ξ\),ξ∼p​\(ξ\),\\displaystyle\\quad\\psi=g\_\{\\beta\}\\\!\\big\(q\(x,z\),\\xi\\big\),\\qquad\\xi\\sim p\(\\xi\),\(4\)whereqqis a feature map andgβg\_\{\\beta\}is a learned module \(deterministic or stochastic\), andξ\\xiis a noise variable with a fixed base distributionp​\(ξ\)p\(\\xi\)\. Givenxx, the predictiony^α,β​\(x\)\\hat\{y\}\_\{\\alpha,\\beta\}\(x\)is obtained by solving \([2](https://arxiv.org/html/2607.22811#S2.E2)\)–\([4](https://arxiv.org/html/2607.22811#S2.E4)\) for\(z,y\)\(z,y\); if multiple solutions exist, the hybrid specification is assumed to include a fixed solver\. Ifgβg\_\{\\beta\}is stochastic, the specification induces a predictive distribution; a point prediction can be taken asy^α,β​\(x\):=𝔼​\[y∣x\]\\hat\{y\}\_\{\\alpha,\\beta\}\(x\):=\\mathbb\{E\}\[y\\mid x\]\.

###### Definition 1\(Hybrid model\)\.

A*hybrid model*is the specificationℳα,β=\(𝒢,ℋ,gβ\)\\mathcal\{M\}\_\{\\alpha,\\beta\}=\(\\mathcal\{G\},\\mathcal\{H\},g\_\{\\beta\}\)together with the induced prediction mapx↦y^α,β​\(x\)x\\mapsto\\hat\{y\}\_\{\\alpha,\\beta\}\(x\)\.

Many hybrid modeling designs can also be summarized at the architecture level as

y=Comp⁡\(Mα​\(x\),Nβ​\(x,Mα​\(x\)\)\),y=\\operatorname\{Comp\}\\\!\\left\(M\_\{\\alpha\}\(x\),\\,N\_\{\\beta\}\\big\(x,M\_\{\\alpha\}\(x\)\\big\)\\right\),\(5\)whereMαM\_\{\\alpha\}denotes the mechanistic operator,NβN\_\{\\beta\}collects learned components, andComp\\operatorname\{Comp\}is the composition rule that combines the mechanistic output with the learned component\.

#### NeSy models and the inference functional\.

We adopt the interface of De Smet and De Raedt\[[6](https://arxiv.org/html/2607.22811#bib.bib1)\]\. A NeSy model specifies a languageLL, a semanticsμ\\muover an interpretation spaceΩ\\Omega, and a belief functionb𝜽b\_\{\\boldsymbol\{\\theta\}\}that weights interpretations; inference aggregates a logic functionllagainst belief as in Equation \([1](https://arxiv.org/html/2607.22811#S1.E1)\)\. We distinguish the full spaceΩ\\Omegafrom the restricted spaceΩ′⊆Ω\\Omega^\{\\prime\}\\subseteq\\Omegarelevant to a given translation \(unknown parameters, closures, latent states\)\. Under mild measurability conditions on\(Ω,l,B𝜽,x\)\(\\Omega,l,B\_\{\\boldsymbol\{\\theta\},x\}\), stated in Appendix[A](https://arxiv.org/html/2607.22811#A1), the functional \([1](https://arxiv.org/html/2607.22811#S1.E1)\) is well\-defined; Dirac beliefs correspond toB𝜽,x=δω∗​\(x\)B\_\{\\boldsymbol\{\\theta\},x\}=\\delta\_\{\\omega^\{\\ast\}\(x\)\}\.

###### Definition 2\(Hybrid\-to\-NeSy decomposition\)\.

A hybrid modelℳα,β\\mathcal\{M\}\_\{\\alpha,\\beta\}induces a NeSy tuple\(L,μ,Ω,b𝛉\)\(L,\\mu,\\Omega,b\_\{\\boldsymbol\{\\theta\}\}\)and an induced inference functionalFF, where

1. 1\.LLandμ\\mucapture the*structural core*given by Equations \([2](https://arxiv.org/html/2607.22811#S2.E2)\)–\([3](https://arxiv.org/html/2607.22811#S2.E3)\) and the composition pattern in Equation \([5](https://arxiv.org/html/2607.22811#S2.E5)\);
2. 2\.Ω\\Omegais the*interpretation space*of assignments to unknown quantities \(e\.g\.α,ψ,z\\alpha,\\psi,z, gates, noise\);
3. 3\.b𝜽b\_\{\\boldsymbol\{\\theta\}\}is a*belief function*overΩ\\Omegainduced by learned part\(s\); \(Dirac in deterministic settings; non\-degenerate in stochastic/Bayesian settings\);
4. 4\.FFis the induced inference functional of the form Equation \([1](https://arxiv.org/html/2607.22811#S1.E1)\)\.

###### Proposition 3\(Logic–belief separation\)\.

Let\(L,μ,Ω,b𝛉\)\(L,\\mu,\\Omega,b\_\{\\boldsymbol\{\\theta\}\}\)be obtained via Definition[2](https://arxiv.org/html/2607.22811#Thmtheorem2)\. Fixxx,φ∈L\\varphi\\in L, and measurableΩ′⊆Ω\\Omega^\{\\prime\}\\subseteq\\Omega\. LetB𝛉,xB\_\{\\boldsymbol\{\\theta\},x\}be the induced \(finite\) belief measure onΩ′\\Omega^\{\\prime\}and assumel​\(φ,⋅\)l\(\\varphi,\\cdot\)is measurable and non\-negative\. Forτ≥0\\tau\\geq 0define

Ωviolτ​\(φ\):=\{ω∈Ω′:l​\(φ,ω\)≤τ\},Ωadmτ​\(φ\):=Ω′∖Ωviolτ​\(φ\)\.\\Omega\_\{\\mathrm\{viol\}\}^\{\\tau\}\(\\varphi\):=\\\{\\omega\\in\\Omega^\{\\prime\}:l\(\\varphi,\\omega\)\\leq\\tau\\\},\\qquad\\Omega\_\{\\mathrm\{adm\}\}^\{\\tau\}\(\\varphi\):=\\Omega^\{\\prime\}\\setminus\\Omega\_\{\\mathrm\{viol\}\}^\{\\tau\}\(\\varphi\)\.Then:

1. 1\.*Logic side \(admissibility\)\.*The violation/admissibility partitionΩviolτ​\(φ\)\\Omega\_\{\\mathrm\{viol\}\}^\{\\tau\}\(\\varphi\)vs\.Ωadmτ​\(φ\)\\Omega\_\{\\mathrm\{adm\}\}^\{\\tau\}\(\\varphi\)is determined entirely byll\(hence by\(L,μ\)\(L,\\mu\)and the chosen feasibility/scoring rule\)\. In the hard caseτ=0\\tau=0,F𝜽,x​\(φ\)=∫Ω′l​\(φ,ω\)​dB𝜽,x​\(ω\)=∫Ωadm0​\(φ\)l​\(φ,ω\)​dB𝜽,x​\(ω\)\.F\_\{\\boldsymbol\{\\theta\},x\}\(\\varphi\)=\\int\_\{\\Omega^\{\\prime\}\}l\(\\varphi,\\omega\)\\,\\mathrm\{d\}B\_\{\\boldsymbol\{\\theta\},x\}\(\\omega\)=\\int\_\{\\Omega\_\{\\mathrm\{adm\}\}^\{0\}\(\\varphi\)\}l\(\\varphi,\\omega\)\\,\\mathrm\{d\}B\_\{\\boldsymbol\{\\theta\},x\}\(\\omega\)\.
2. 2\.*Belief side \(plausibility\)\.*Conditional on admissibility, inference is governed by howB𝜽,xB\_\{\\boldsymbol\{\\theta\},x\}allocates*probability mass*withinΩadmτ​\(φ\)\\Omega\_\{\\mathrm\{adm\}\}^\{\\tau\}\(\\varphi\)\(concentrated vs\. dispersed\)\. In particular, ifB𝜽,x=δω∗​\(x\)B\_\{\\boldsymbol\{\\theta\},x\}=\\delta\_\{\\omega^\{\\ast\}\(x\)\}, thenF𝜽,x​\(φ\)=l​\(φ,ω∗​\(x\)\)F\_\{\\boldsymbol\{\\theta\},x\}\(\\varphi\)=l\(\\varphi,\\omega^\{\\ast\}\(x\)\)and the induced behavior is admissible iffω∗​\(x\)∈Ωadmτ​\(φ\)\\omega^\{\\ast\}\(x\)\\in\\Omega\_\{\\mathrm\{adm\}\}^\{\\tau\}\(\\varphi\); non\-degenerate beliefs distribute probability mass across multiple admissible interpretations\.

###### Proof sketch\.

The admissibility partition uses onlyl​\(φ,⋅\)l\(\\varphi,\\cdot\)andτ\\tau, hence only\(L,μ,l\)\(L,\\mu,l\)and notB𝜽,xB\_\{\\boldsymbol\{\\theta\},x\}\. In the hard caseτ=0\\tau=0,l​\(φ,⋅\)=0l\(\\varphi,\\cdot\)=0onΩviol0​\(φ\)\\Omega\_\{\\mathrm\{viol\}\}^\{0\}\(\\varphi\)by non\-negativity together with the definition ofΩviol0​\(φ\)\\Omega\_\{\\mathrm\{viol\}\}^\{0\}\(\\varphi\), so the violating set contributes nothing; a Dirac belief collapsesF𝜽,x​\(φ\)F\_\{\\boldsymbol\{\\theta\},x\}\(\\varphi\)tol​\(φ,ω∗​\(x\)\)l\(\\varphi,\\omega^\{\\ast\}\(x\)\)\. The full argument is in Appendix[B](https://arxiv.org/html/2607.22811#A2)\. ∎

Problem\.Given a hybrid model \(Equation \([5](https://arxiv.org/html/2607.22811#S2.E5)\)\), construct\(L,μ,Ω,b𝜽\)\(L,\\mu,\\Omega,b\_\{\\boldsymbol\{\\theta\}\}\)andllsuch that structures are evaluated viallunder\(L,μ\)\(L,\\mu\), unknowns live inΩ\\Omega, learning inducesb𝜽b\_\{\\boldsymbol\{\\theta\}\}, and the induced functionalFFreproduces the hybrid’s predictive behaviour \(up to selection rules\)\.

## 3Method: A translation from hybrid modeling to NeSy

This section gives a constructive solution toH2N: given a hybrid design, we build\(L,μ,Ω,b𝜽\)\(L,\\mu,\\Omega,b\_\{\\boldsymbol\{\\theta\}\}\)andllso that the induced functional in \([1](https://arxiv.org/html/2607.22811#S1.E1)\) becomes an explicit, comparable interface\. The inputxxis external toω\\omega; the learned part inducesb𝜽,x​\(ω\)b\_\{\\boldsymbol\{\\theta\},x\}\(\\omega\)conditional onxx, while query dependence enters throughl​\(φ,ω\)l\(\\varphi,\\omega\)\. If a design requires query\-dependent beliefs, this can be represented by extendingω\\omegaor lettingb𝜽,xb\_\{\\boldsymbol\{\\theta\},x\}depend onφ\\varphi; our default conditions onxxonly\.

### 3\.1Translation principles

H2Nis governed by five principles, stated in full in Appendix[C](https://arxiv.org/html/2607.22811#A3)and summarised here: \(P1\)Mechanistic structure determines the languageLL, with unknown closures and latent states introduced as explicit symbols\. \(P2\)Semanticsμ\\muspecify evaluation; the logic functionllis the operational object inside the integral, implementing admissibility or scoring\. \(P3\)Learned modules induce beliefb𝜽,x​\(ω\)b\_\{\\boldsymbol\{\\theta\},x\}\(\\omega\)over interpretations; deterministic learners are the Dirac case\. \(P4\)Validity domains and hard constraints live on the logic side: by restricting the integration domainΩ′⊆Ω\\Omega^\{\\prime\}\\subseteq\\Omegaand/or by selection insidell, separating*admissibility*\(logic\) from*plausibility*\(belief\)\. Hard physical laws are encoded on the logic side rather than inb𝜽b\_\{\\boldsymbol\{\\theta\}\}: such laws express data\-independent admissibility, whereasb𝜽b\_\{\\boldsymbol\{\\theta\}\}is by construction data\-dependent\. Encoding them inb𝜽b\_\{\\boldsymbol\{\\theta\}\}would couple their enforcement to sample size, which is inconsistent with their meaning\. Soft or empirical constraints, whose strength legitimately depends on data, may enterb𝜽b\_\{\\boldsymbol\{\\theta\}\}\. \(P5\)Architectural factorisations \(serial, parallel, mixture\) are mirrored by factorisations oflland/orb𝜽,xb\_\{\\boldsymbol\{\\theta\},x\}\. As summarised in Algorithm[1](https://arxiv.org/html/2607.22811#algorithm1)in Appendix[C](https://arxiv.org/html/2607.22811#A3), the result is a NeSy inference object whose structural core\(L,μ,l\)\(L,\\mu,l\)and uncertainty modelb𝜽b\_\{\\boldsymbol\{\\theta\}\}make comparison of hybrid designs systematic\.

### 3\.2Mapping of hybrid modeling designs to NeSy elements

Table[1](https://arxiv.org/html/2607.22811#S3.T1)summarizes how canonical hybrid modeling design patterns map to NeSy elements and how each pattern changes the inference functionalF𝜽,xF\_\{\\boldsymbol\{\\theta\},x\}in Equation \([1](https://arxiv.org/html/2607.22811#S1.E1)\)\. Across the patterns, the translation separates structure from learned uncertainty\. The main differences are whether a pattern changes \(i\) the structural languageLLthat determines what is expressible, \(ii\) operational admissibility or scoring throughllthat determines how structure is enforced, \(iii\) which unknowns are integrated over through the choice ofΩ′\\Omega^\{\\prime\}withinΩ\\Omega, or \(iv\) how uncertainty is allocated through belief factorization and dispersion inb𝜽,xb\_\{\\boldsymbol\{\\theta\},x\}\.

Table 1:Canonical hybrid modeling design patterns mapped to NeSy elements and their primary impact on the inference functionalF𝜽,xF\_\{\\boldsymbol\{\\theta\},x\}in Equation \([1](https://arxiv.org/html/2607.22811#S1.E1)\)\. The final column lists representative NeSy architectures that realise each pattern\.

## 4Evaluation Protocol

Let𝒟test=\{\(xi,yi\)\}i=1n\\mathcal\{D\}\_\{\\mathrm\{test\}\}=\\\{\(x\_\{i\},y\_\{i\}\)\\\}\_\{i=1\}^\{n\}be a test set, and letF𝜽,xF\_\{\\boldsymbol\{\\theta\},x\}be the induced inference functional in Equation \([1](https://arxiv.org/html/2607.22811#S1.E1)\)\. We evaluate the H2N translation with two kinds of queries, each evaluable throughl​\(φ,ω\)l\(\\varphi,\\omega\)and integrable underb𝜽,xb\_\{\\boldsymbol\{\\theta\},x\}\. Aconstraint queryφxcon\\varphi^\{\\mathrm\{con\}\}\_\{x\}expresses structural admissibility; apredictive queryφx,ypred\\varphi^\{\\mathrm\{pred\}\}\_\{x,y\}expresses agreement with data \(e\.g\.‖y^​\(x,ω\)−y‖≤ε\\\|\\hat\{y\}\(x,\\omega\)\-y\\\|\\leq\\varepsilonfor a toleranceε\\varepsilon\), of which ordinary test accuracy is the point\-estimate reading\. The protocol is otherwise agnostic to the domain\.

For a constraint queryφxcon\\varphi^\{\\mathrm\{con\}\}\_\{x\}and a toleranceτ≥0\\tau\\geq 0, the logic–belief separation partitionsΩ′\\Omega^\{\\prime\}into an admissible regionΩadmτ​\(φ\)\\Omega\_\{\\mathrm\{adm\}\}^\{\\tau\}\(\\varphi\)and a violating regionΩviolτ​\(φ\)\\Omega\_\{\\mathrm\{viol\}\}^\{\\tau\}\(\\varphi\)\. This partition is fixed entirely byll\(hence by\(L,μ\)\(L,\\mu\)and the chosen feasibility rule\), independently of the belief \(Proposition[3](https://arxiv.org/html/2607.22811#Thmtheorem3), Remark[4](https://arxiv.org/html/2607.22811#Thmtheorem4)\)\. The belief\-weighted violation mass atxx,

V𝜽​\(x;φ\):=∫Ω′𝕀​\[l​\(φ,ω\)≤τ\]​dB𝜽,x​\(ω\),d​B𝜽,x=b𝜽,x​d​m,V\_\{\\boldsymbol\{\\theta\}\}\(x;\\varphi\):=\\int\_\{\\Omega^\{\\prime\}\}\\mathbb\{I\}\\\!\\left\[l\(\\varphi,\\omega\)\\leq\\tau\\right\]\\,\\mathrm\{d\}B\_\{\\boldsymbol\{\\theta\},x\}\(\\omega\),\\qquad\\mathrm\{d\}B\_\{\\boldsymbol\{\\theta\},x\}=b\_\{\\boldsymbol\{\\theta\},x\}\\,\\mathrm\{d\}m,is the \(normalised\) belief mass falling in the violating region; it couples the logic\-side partition to the belief and reduces to∫Ω′𝕀​\[l​\(φ,ω\)=0\]​dB𝜽,x\\int\_\{\\Omega^\{\\prime\}\}\\mathbb\{I\}\\\!\\left\[l\(\\varphi,\\omega\)=0\\right\]\\,\\mathrm\{d\}B\_\{\\boldsymbol\{\\theta\},x\}in the hard Boolean caseτ=0\\tau=0\.

### 4\.1Metrics

#### Structural violation rate \(SVR\)\.

Given a constraint queryφxcon\\varphi^\{\\mathrm\{con\}\}\_\{x\},

SVR​\(𝜽\):=1n​∑i=1nV𝜽​\(xi;φxicon\),\\mathrm\{SVR\}\(\\boldsymbol\{\\theta\}\):=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}V\_\{\\boldsymbol\{\\theta\}\}\(x\_\{i\};\\varphi^\{\\mathrm\{con\}\}\_\{x\_\{i\}\}\),\(6\)the average belief mass placed on interpretations that violate the structural constraints\. SVR is near0when the logic side reliably admits interpretations that satisfy the constraints under the learned belief, and increases as belief mass concentrates on invalid interpretations\. Three properties characterise SVR\.*\(i\) It lives on the logic side:*its value is governed by the admissibility partitionΩadmτ/Ωviolτ\\Omega\_\{\\mathrm\{adm\}\}^\{\\tau\}/\\Omega\_\{\\mathrm\{viol\}\}^\{\\tau\}, a property of\(L,μ,l\)\(L,\\mu,l\), weighted by the belief\.*\(ii\) It is tolerance\-dependent:*the strictness of the constraint is set byτ\\tau\(the violation budgetκ\\kappain the case study of Section[5](https://arxiv.org/html/2607.22811#S5)\); tighteningτ\\tauenlarges the violating region and raises SVR, loosening it lowers SVR\.*\(iii\) It is a feasibility measure, not a dispersion measure:*for a Dirac belief,V𝜽​\(x;⋅\)∈\{0,1\}V\_\{\\boldsymbol\{\\theta\}\}\(x;\\cdot\)\\in\\\{0,1\\\}records only whether the point estimate is admissible, and SVR says nothing about how concentrated the belief is\.

#### Belief dispersion \(BD\)\.

To assess epistemic uncertainty, we choose an uncertain projectionωU=πU​\(ω\)\\omega\_\{U\}=\\pi\_\{U\}\(\\omega\)\(e\.g\. closure values, parameters\) and report

BD​\(𝜽\):=1n​∑i=1ntr​\(Covω∼b𝜽,xi​\[ωU\]\),\\mathrm\{BD\}\(\\boldsymbol\{\\theta\}\):=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\mathrm\{tr\}\\\!\\left\(\\mathrm\{Cov\}\_\{\\omega\\sim b\_\{\\boldsymbol\{\\theta\},x\_\{i\}\}\}\\\!\[\\omega\_\{U\}\]\\right\),\(7\)where samplingω∼b𝜽,x\\omega\\sim b\_\{\\boldsymbol\{\\theta\},x\}is from the normalised belief onΩ′\\Omega^\{\\prime\}\(proportional tob𝜽,x​\(ω\)​d​m​\(ω\)b\_\{\\boldsymbol\{\\theta\},x\}\(\\omega\)\\,dm\(\\omega\)\)\. BD has the dual set of properties\.*\(i\) It lives on the belief side:*it is a functional of the second moment ofb𝜽,xb\_\{\\boldsymbol\{\\theta\},x\}alone, and the logic functionlldoes not enter\.*\(ii\) It is tolerance\-independent:*BD never referencesτ\\tauor the admissibility partition, so changing the constraint’s strictness leaves it unchanged\.*\(iii\) It is a dispersion measure:*BD is zero for Dirac beliefs and grows as the belief spreads; it is invariant under orthonormal reparameterisations ofωU\\omega\_\{U\}but not scale\-invariant, soωU\\omega\_\{U\}should be standardised when coordinates differ in scale\.

## 5Case Study: Structured networks for noisy binary classification

We instantiateH2Non structured Boolean networks for binary classification with prior feature grouping\. The mechanistic part is a partition of theNNbinary\-represented features intoMMdisjoint groups routed to first\-layer Boolean modulesFm:\{0,1\}nm→\{0,1\}F\_\{m\}:\\\{0,1\\\}^\{n\_\{m\}\}\\to\\\{0,1\\\}, combined by an output moduleFO:\{0,1\}M→\{0,1\}F\_\{O\}:\\\{0,1\\\}^\{M\}\\to\\\{0,1\\\}\. The learned components are the module truth\-tables, identified by the learning strategy of NoiseCut\[[21](https://arxiv.org/html/2607.22811#bib.bib23)\]\.

#### H2Ntranslation\.

The unknowns are the module truth tables, so we introduce one Boolean atom per table entry:fm,kf\_\{m,k\}encodesFm​\(k\)F\_\{m\}\(k\)at local inputk∈\{0,1\}nmk\\in\\\{0,1\\\}^\{n\_\{m\}\}, andoKo\_\{K\}encodesFO​\(K\)F\_\{O\}\(K\)at second\-layerrowK∈\{0,1\}MK\\in\\\{0,1\\\}^\{M\}\. An interpretationω∈Ω=Ω′=𝔹𝒜\\omega\\in\\Omega=\\Omega^\{\\prime\}=\\mathbb\{B\}^\{\\mathcal\{A\}\}\(𝒜\\mathcal\{A\}the set of all atoms\) is then a complete assignment of module outputs, and induces a predictor: routingxxthrough the first layer selects the rowKω​\(x\)=\(ω​\(f1,k1​\(x\)\),…,ω​\(fM,kM​\(x\)\)\)K^\{\\omega\}\(x\)=\\big\(\\omega\(f\_\{1,k\_\{1\}\(x\)\}\),\\dots,\\omega\(f\_\{M,k\_\{M\}\(x\)\}\)\\big\), with predicted labelOω​\(x\)=ω​\(oKω​\(x\)\)O^\{\\omega\}\(x\)=\\omega\(o\_\{K^\{\\omega\}\(x\)\}\)\. Under Boolean semanticsμ=μB\\mu=\\mu\_\{B\}, per\-sample consistency isφs=\(Oω\(x\(s\)\)↔y\(s\)\)\\varphi\_\{s\}=\\big\(O^\{\\omega\}\(x^\{\(s\)\}\)\\\!\\leftrightarrow\\\!y^\{\(s\)\}\\big\)and the constraint query isφcon=⋀s=1Sφs\\varphi^\{\\mathrm\{con\}\}=\\bigwedge\_\{s=1\}^\{S\}\\varphi\_\{s\}over theSStraining samples\. Writingviol​\(ω\)=∑s𝕀​\[μB​\(φs,ω\)=0\]\\mathrm\{viol\}\(\\omega\)=\\sum\_\{s\}\\mathbb\{I\}\\\!\\left\[\\mu\_\{B\}\(\\varphi\_\{s\},\\omega\)=0\\right\]for the number of training samples misclassified by the interpretationω\\omega, the noise robustness of the logical side is determined by the choice of logical evaluation function:

lhard=μB​\(φcon,ω\),ltol=𝕀​\[viol​\(ω\)≤κ\]\.l\_\{\\mathrm\{hard\}\}=\\mu\_\{B\}\(\\varphi^\{\\mathrm\{con\}\},\\omega\),\\qquad l\_\{\\mathrm\{tol\}\}=\\mathbb\{I\}\\\!\\left\[\\mathrm\{viol\}\(\\omega\)\\leq\\kappa\\right\]\.These correspond respectively to exact consistency \(κ=0\\kappa=0\) and consistency subject to a violation budgetκ\\kappa\. Consequently, noise robustness is a design choice of the logical evaluation function rather than a property of the learned belief; SVR is evaluated underltoll\_\{\\mathrm\{tol\}\}, with the budgetκ\\kappaplaying the role of the toleranceτ\\tauof Section[4](https://arxiv.org/html/2607.22811#S4)\.

#### Belief\.

NoiseCut identifies the truth tables of the first\-layer modules, yielding point estimatesF^m\\hat\{F\}\_\{m\}, which are incorporated as Dirac factors\. The truth table of the output module is estimated via label counting: for each rowKK, letn0​\(K\)n\_\{0\}\(K\)andn1​\(K\)n\_\{1\}\(K\)denote the numbers of training samples labeled0and11, respectively\. This defines a Bernoulli parameterpK=n1​\(K\)/\(n0​\(K\)\+n1​\(K\)\)p\_\{K\}=n\_\{1\}\(K\)/\(n\_\{0\}\(K\)\+n\_\{1\}\(K\)\)withpK=12p\_\{K\}=\\tfrac\{1\}\{2\}assigned to rows receiving no training support \(i\.e\.,n0​\(K\)\+n1​\(K\)=0n\_\{0\}\(K\)\+n\_\{1\}\(K\)=0\)\. The belief factorises as

b𝜽​\(ω\)=∏m𝕀​\[ω\|Fm=F^m\]⏟Dirac \(first\-layer modules\)​∏KpKω​\(oK\)​\(1−pK\)1−ω​\(oK\)⏟Bernoulli \(output module\)\.b\_\{\\boldsymbol\{\\theta\}\}\(\\omega\)=\\underbrace\{\\prod\_\{m\}\\mathbb\{I\}\\\!\\left\[\\omega\|\_\{F\_\{m\}\}=\\hat\{F\}\_\{m\}\\right\]\}\_\{\\text\{Dirac \(first\-layer modules\)\}\}\\;\\underbrace\{\\prod\_\{K\}p\_\{K\}^\{\\omega\(o\_\{K\}\)\}\(1\-p\_\{K\}\)^\{1\-\\omega\(o\_\{K\}\)\}\}\_\{\\text\{Bernoulli \(output module\)\}\}\.The non\-degenerate Bernoulli factors capture the epistemic uncertainty, while unobserved rows are assigned the maximum\-entropy valuepK=12p\_\{K\}=\\tfrac\{1\}\{2\}\. Under the counting measure, the functional in \([1](https://arxiv.org/html/2607.22811#S1.E1)\) corresponds to a weighted model count\. The corresponding estimators are described in Appendix[D](https://arxiv.org/html/2607.22811#A4)\.

#### Experimental setup\.

The data are generated from random functionality\-preserving modules withnm=4n\_\{m\}=4inputs per first\-layer module\. All experiments useM=4M=4first\-layer modules \(N=16N=16,2162^\{16\}samples\) and a50%/50%50\\%/50\\%train/test split\. The noise sweep study uses a fixed toleranceκ=⌊0\.5​S⌋\\kappa=\\lfloor 0\.5\\,S\\rfloor, and5050seeds at each of ten noise levels from0%0\\%to45%45\\%\. The tolerance sweep uses1010seeds and fixes10%10\\%label noise and sweepsκ/S\\kappa/Sfrom0to0\.50\.5\(Appendix[E](https://arxiv.org/html/2607.22811#A5)\); the OOD study uses1010seeds and fixesκ=⌊0\.10​S⌋\\kappa=\\lfloor 0\.10\\,S\\rfloorand10%10\\%label noise and holds outNood∈\{1,2,4,6\}N\_\{\\mathrm\{ood\}\}\\in\\\{1,2,4,6\\\}of the1616reachable second\-layer rows \(Appendix[F](https://arxiv.org/html/2607.22811#A6)\)\. Label noise flips training labels only, keeping testing labels clean\. All H2N metrics use20,00020\{,\}000Monte CarloFOF\_\{O\}draws, with BD additionally available in closed form \(Appendix[D](https://arxiv.org/html/2607.22811#A4)\)\.

#### Result 1: quantifying uncertainty in the mechanistic part\.

Table[2](https://arxiv.org/html/2607.22811#S5.T2)and Figure[1](https://arxiv.org/html/2607.22811#S5.F1)report the H2N metrics across label\-noise levels\. As the noise increases, mean test accuracy falls from1\.001\.00to0\.630\.63while mean BD measured at the deployment time rises from0to near its maximum2M/4=42^\{M\}/4=4, the value attained when every second\-layer row is maximally uncertain,pK=12p\_\{K\}=\\tfrac\{1\}\{2\}\. Test accuracy and BD are strongly anti\-correlated across seeds \(Spearmanρ=−0\.94\\rho=\-0\.94\), and therefore track related but distinct properties: the former is a point estimate of predictive error, the latter the dispersion of the row\-Bernoulli belief\.

As shown in Table[2](https://arxiv.org/html/2607.22811#S5.T2)\(as well as in Figure[G1](https://arxiv.org/html/2607.22811#A7.F1)\), SVR, under the toleranceκ=⌊0\.5​S⌋\\kappa=\\lfloor 0\.5\\,S\\rfloor, remains near zero while the realized violation count stays within the budget and rises to0\.4280\.428only at at45%45\\%noise levels, where violations routinely exceedκ\\kappa\.

![Refer to caption](https://arxiv.org/html/2607.22811v1/x1.png)Figure 1:Noise sweep,5050seeds per level \(struct\[4,4,4,4\]\[4,4,4,4\],κ=⌊0\.5​S⌋\\kappa=\\lfloor 0\.5\\,S\\rfloor\)\. \(a\) Per\-seed scatter of test accuracy against BD, coloured by noise level: the spread of accuracy widens as BD grows\. \(b\) Standard deviation of test accuracy across seeds, computed within bins of BD; the accuracy uncertainty rises monotonically with BD\.Table 2:Noise sweep with H2N metrics; struct\[4,4,4,4\]\[4,4,4,4\], train fraction50%50\\%,κ=⌊0\.5​S⌋\\kappa=\\lfloor 0\.5\\,S\\rfloor\. Each entry reports the mean over5050seeds with the95%95\\%confidence interval\. SVR and BD are measured at the deployment time using training data with label noise\. Acc\. and F1 score are measured using test data with clean labels\.The central observation concerns not the means but the*spread*of accuracy\. When the500500trained models are stratified by their BD value, the seed\-to\-seed standard deviation of test accuracy increases monotonically with BD, from0atBD≈0\\mathrm\{BD\}\\approx 0to about0\.0730\.073asBD→4\\mathrm\{BD\}\\to 4\(Figure[1](https://arxiv.org/html/2607.22811#S5.F1)\(b\)\); the per\-seed scatter in Figure[1](https://arxiv.org/html/2607.22811#S5.F1)\(a\) exhibits the same effect\. A high BD thus signals a wider, less predictable accuracy distribution\. Because BD is a function only of the training\-label countsn0​\(K\),n1​\(K\)n\_\{0\}\(K\),n\_\{1\}\(K\)and the mechanistic structure of first\-layer partitions, it provides ana prioriestimate of how uncertain a trained model’s accuracy will be, a quantity that test accuracy can report only after labeled test data have been collected\.

#### Result 2: quantifying uncertainty during extrapolation\.

We holdNood∈\{1,2,4,6\}N\_\{\\mathrm\{ood\}\}\\in\\\{1,2,4,6\\\}reachable second\-layer rows entirely out of training \(with10%10\\%noise\) and stratify the test set into covered and held\-out rows, for which predictions amounts to extrapolation\. Because the row partition is logical \(fixed byΩ′\\Omega^\{\\prime\}rather than by the learned predictor\) BD decomposes exactly asBD=BDseen\+BDunseen\\mathrm\{BD\}=\\mathrm\{BD\}\_\{\\text\{seen\}\}\+\\mathrm\{BD\}\_\{\\text\{unseen\}\}, whereBDunseen=14​\|\{K:n0​\(K\)\+n1​\(K\)=0\}\|\\mathrm\{BD\}\_\{\\text\{unseen\}\}=\\tfrac\{1\}\{4\}\\,\|\\\{K:n\_\{0\}\(K\)\+n\_\{1\}\(K\)=0\\\}\|sums the maximum\-entropy variance of the unobserved rows and is read off the*learned*first\-layer functions\. Crucially,BDunseen\\mathrm\{BD\}\_\{\\text\{unseen\}\}is computable from row coverage alone, before any OOD sample is seen \(Remark[4](https://arxiv.org/html/2607.22811#Thmtheorem4)\)\. Across1010seeds, mean in\-distribution accuracy stays at0\.990\.99–1\.001\.00while mean OOD or extrapolation accuracy collapses to0\.35,0\.35,0\.56,0\.490\.35,0\.35,0\.56,0\.49and meanBDunseen\\mathrm\{BD\}\_\{\\text\{unseen\}\}rises monotonically as0\.25,0\.50,0\.68,0\.950\.25,0\.50,0\.68,0\.95; the full numbers are in Appendix[F](https://arxiv.org/html/2607.22811#A6)\. The decomposition thus reports the model’s epistemic state about coverage fromΩ′\\Omega^\{\\prime\}directly, whereas accuracy reveals the same shift only after the OOD labels are observed\.

## 6Discussion

Hybrid models are often presented as architectures and losses\.H2Nreconstructs each design as a NeSy inference object\(L,μ,Ω,b𝜽\)\(L,\\mu,\\Omega,b\_\{\\boldsymbol\{\\theta\}\}\)with an explicit functionalFF, making it clear which assumptions act as*structural admissibility*\(language/semantics/logic function\) versus*learned plausibility*\(belief over unknown parameters, functions, or states\)\. This does not remove modeling choices \(hard vs\. soft constraints, choice ofμ\\mu, belief parameterisation\), but makes them comparable across canonical hybrid design patterns\. SVR quantifies how much belief mass violates the structural constraints \(logic side\), while BD summarises epistemic dispersion of the learned belief \(belief side\); the two are complementary and decoupled, since the tolerance controls SVR without affecting BD\. Robustness under structured shift becomes a targeted intervention: modifyll\(changed regimes/constraints\) versus modifyb𝜽b\_\{\\boldsymbol\{\\theta\}\}\.

#### Measurable consequence of placement\.

Sec\.[5](https://arxiv.org/html/2607.22811#S5)instantiates the placement argument empirically: BD quantifies the model’s uncertainty on its structure and on uncovered regions ofΩ′\\Omega^\{\\prime\}at deployment time, while test accuracy reports the same shift only post hoc\. The diagnostic follows from logic–belief separation \(Prop\.[3](https://arxiv.org/html/2607.22811#Thmtheorem3), Rem\.[4](https://arxiv.org/html/2607.22811#Thmtheorem4)\) and is well\-defined under H2N but undefined under the architecture\-and\-loss description of the same model\.

#### What NeSy gains from hybrid modeling\.

Hybrid modeling contributes assets that NeSy has lacked an interface to: an established library of*compositional patterns*\(serial closure, parallel residual, mixture\-of\-experts, modular networks\) with documented identifiability and validity\-domain analyses;*principled constraint families*\(conservation, balance, monotonicity\) that translate naturally intollrather than into ad\-hoc penalty terms; and*diagnostic practices*from process engineering \(validity envelopes, residual checks\) that the H2N metrics generalise\. The flow is therefore bidirectional:H2Ngives hybrid modelers a semantic interface, and gives NeSy researchers a pre\-processed catalogue of structured\-uncertainty designs to import\.

#### Limitations\.

The translation is not unique: redistributing assumptions betweenΩ′\\Omega^\{\\prime\}andllcan alterFFunless normalisation is controlled\. Principles P1–P5 fix a canonical placement per pattern \(Table[1](https://arxiv.org/html/2607.22811#S3.T1)\); within this discipline, alternative placements correspond to alternative modeling commitments rather than equivalent notations\. For large or continuousΩ\\Omega, diagnostics rely on approximate inference and inherit its calibration error\.

## References

- \[1\]\(1997\)Combining neural and conventional paradigms for modelling, prediction and control\.International Journal of Systems Science28\(1\),pp\. 65–81\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p1.1)\.
- \[2\]S\. Bader and P\. Hitzler\(2005\)Dimensions of neural\-symbolic integration\-a structured survey\.arXiv preprint cs/0511042\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p4.5)\.
- \[3\]V\. Belle, A\. Passerini, and G\. Van den Broeck\(2015\)Probabilistic inference in hybrid domains by weighted model integration\.InProceedings of the Twenty\-Fourth International Joint Conference on Artificial Intelligence, IJCAI 2015, Buenos Aires, Argentina, July 25\-31, 2015,pp\. 2770–2776\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p4.5)\.
- \[4\]N\. Bhutani, G\. Rangaiah, and A\. Ray\(2006\)First\-principles, data\-based, and hybrid modeling and optimization of an industrial hydrocracking unit\.Industrial & engineering chemistry research45\(23\),pp\. 7807–7816\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p2.1),[Table 1](https://arxiv.org/html/2607.22811#S3.T1.17.15.6.1.1)\.
- \[5\]I\. T\. Cameron and K\. Hangos\(2001\)Process modelling and model analysis\.Vol\.4,Elsevier\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p1.1)\.
- \[6\]L\. De Smet and L\. De Raedt\(2025\)Defining neurosymbolic ai\.arXiv preprint arXiv:2507\.11127\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.SS0.SSS0.Px1.p1.5),[§1](https://arxiv.org/html/2607.22811#S1.p4.5),[§2](https://arxiv.org/html/2607.22811#S2.SS0.SSS0.Px2.p1.9)\.
- \[7\]B\. Fiedler and A\. Schuppert\(2008\)Local identification of scalar hybrid models with tree structure\.IMA Journal of Applied Mathematics73\(3\),pp\. 449–476\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.SS0.SSS0.Px2.p1.6),[§1](https://arxiv.org/html/2607.22811#S1.p1.1),[Table 1](https://arxiv.org/html/2607.22811#S3.T1.26.24.4.1.1)\.
- \[8\]A\. d\. Garcez, M\. Gori, L\. C\. Lamb, L\. Serafini, M\. Spranger, and S\. N\. Tran\(2019\)Neural\-symbolic computing: an effective methodology for principled integration of machine learning and reasoning\.arXiv preprint arXiv:1905\.06088\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p4.5)\.
- \[9\]A\. d\. Garcez and L\. C\. Lamb\(2023\)Neurosymbolic ai: the 3 rd wave\.Artificial Intelligence Review56\(11\),pp\. 12387–12406\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p4.5)\.
- \[10\]J\. Glassey and M\. Von Stosch\(2018\)Hybrid modeling in process industries\.CRC Press\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p3.1)\.
- \[11\]P\. Hitzler, M\. Sarker, T\. Besold, A\. Garcez, S\. Bader, H\. Bowman, P\. Domingos, P\. Hitzler, K\. Kühnberger, L\. Lamb,et al\.\(2022\)Neural\-symbolic learning and reasoning: a survey and interpretation\.Frontiers in artificial intelligence and applications342,pp\. 1–51\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p4.5)\.
- \[12\]O\. Kahrs and W\. Marquardt\(2007\)The validity domain of hybrid models and its application in process optimization\.Chemical Engineering and Processing: Process Intensification46\(11\),pp\. 1054–1066\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p2.1)\.
- \[13\]D\. S\. Lee, C\. O\. Jeon, J\. M\. Park, and K\. S\. Chang\(2002\)Hybrid neural network modeling of a full\-scale industrial wastewater treatment process\.Biotechnology and bioengineering78\(6\),pp\. 670–682\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p2.1),[Table 1](https://arxiv.org/html/2607.22811#S3.T1.17.15.6.1.1)\.
- \[14\]E\. Marconato, S\. Teso, A\. Vergari, and A\. Passerini\(2023\)Not all neuro\-symbolic concepts are created equal: analysis and mitigation of reasoning shortcuts\.Advances in Neural Information Processing Systems36,pp\. 72507–72539\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p4.5)\.
- \[15\]G\. Marra, S\. Dumančić, R\. Manhaeve, and L\. De Raedt\(2024\)From statistical relational to neurosymbolic artificial intelligence: a survey\.Artificial Intelligence328,pp\. 104062\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p4.5)\.
- \[16\]J\. Peres, R\. Oliveira, and S\. F\. De Azevedo\(2001\)Knowledge based modular networks for process modelling and control\.Computers & Chemical Engineering25\(4\-6\),pp\. 783–791\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p2.1),[Table 1](https://arxiv.org/html/2607.22811#S3.T1.20.18.4.1.1)\.
- \[17\]J\. Peres, R\. Oliveira, and S\. F\. de Azevedo\(2008\)Bioprocess hybrid parametric/nonparametric modelling based on the concept of mixture of experts\.Biochemical Engineering Journal39\(1\),pp\. 190–206\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p2.1),[Table 1](https://arxiv.org/html/2607.22811#S3.T1.20.18.4.1.1)\.
- \[18\]D\. C\. Psichogios and L\. H\. Ungar\(1992\)A hybrid neural network\-first principles approach to process modeling\.AIChE Journal38\(10\),pp\. 1499–1511\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p1.1),[Table 1](https://arxiv.org/html/2607.22811#S3.T1.12.10.3.1.1)\.
- \[19\]M\. E\. Samadi, J\. Guzman\-Maldonado, K\. Nikulina, H\. Mirzaieazar, K\. Sharafutdinov, S\. J\. Fritsch, and A\. Schuppert\(2024\)A hybrid modeling framework for generalizable and interpretable predictions of icu mortality across multiple hospitals\.Scientific reports14\(1\),pp\. 5725\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p2.1)\.
- \[20\]M\. E\. Samadi, S\. Kiefer, S\. J\. Fritsch, J\. Bickenbach, and A\. Schuppert\(2022\)A training strategy for hybrid models to break the curse of dimensionality\.Plos one17\(9\),pp\. e0274569\.Cited by:[Table 1](https://arxiv.org/html/2607.22811#S3.T1.26.24.4.1.1)\.
- \[21\]M\. E\. Samadi, H\. Mirzaieazar, A\. Mitsos, and A\. Schuppert\(2024\)Noisecut: a python package for noise\-tolerant classification of binary data using prior knowledge integration and max\-cut solutions\.BMC bioinformatics25\(1\),pp\. 155\.Cited by:[§5](https://arxiv.org/html/2607.22811#S5.p1.4)\.
- \[22\]A\. A\. Schuppert\(2000\)Extrapolability of structured hybrid models: a key to optimization of complex processes\.InEquadiff 99: \(In 2 Volumes\),pp\. 1135–1151\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p1.1)\.
- \[23\]A\. A\. Schuppert\(2011\)Efficient reengineering of meso\-scale topologies for functional networks in biomedical applications\.Journal of Mathematics in Industry1\(1\),pp\. 6\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p2.1)\.
- \[24\]T\. Takagi and M\. Sugeno\(1985\)Fuzzy identification of systems and its applications to modeling and control\.IEEE transactions on systems, man, and cybernetics\(1\),pp\. 116–132\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p2.1)\.
- \[25\]A\. P\. Teixeira, N\. Carinhas, J\. M\. Dias, P\. Cruz, P\. M\. Alves, M\. J\. Carrondo, and R\. Oliveira\(2007\)Hybrid semi\-parametric mathematical systems: bridging the gap between systems biology and process engineering\.Journal of biotechnology132\(4\),pp\. 418–425\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p1.1)\.
- \[26\]M\. L\. Thompson and M\. A\. Kramer\(1994\)Modeling chemical processes using prior knowledge and neural networks\.AIChE Journal40\(8\),pp\. 1328–1340\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p1.1),[Table 1](https://arxiv.org/html/2607.22811#S3.T1.12.10.3.1.1)\.
- \[27\]H\. J\. Tulleken\(1993\)Grey\-box modelling and identification using physical knowledge and bayesian techniques\.Automatica29\(2\),pp\. 285–308\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p1.1)\.
- \[28\]M\. van Bekkum, M\. de Boer, F\. van Harmelen, A\. Meyer\-Vitali, and A\. t\. Teije\(2021\)Modular design patterns for hybrid learning and reasoning systems: a taxonomy, patterns and use cases\.Applied Intelligence51\(9\),pp\. 6528–6546\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p4.5)\.
- \[29\]H\. J\. Van Can, C\. Hellinga, K\. C\. A\. Luyben, J\. J\. Heijnen, and H\. A\. Te Braake\(1996\)Strategy for dynamic process modeling based on neural networks in macroscopic balances\.AIChE Journal42\(12\),pp\. 3403–3418\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p1.1)\.
- \[30\]H\. J\. Van Can, H\. A\. Te Braake, S\. Dubbelman, C\. Hellinga, K\. C\. A\. Luyben, and J\. J\. Heijnen\(1998\)Understanding and applying the extrapolation properties of serial gray\-box models\.AIChE journal44\(5\),pp\. 1071–1089\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p1.1)\.
- \[31\]P\. F\. van Lith, B\. H\. Betlem, and B\. Roffel\(2002\)A structured modeling approach for dynamic hybrid fuzzy\-first principles models\.Journal of Process Control12\(5\),pp\. 605–615\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p2.1),[Table 1](https://arxiv.org/html/2607.22811#S3.T1.23.21.4.1.1)\.
- \[32\]P\. F\. van Lith, B\. H\. Betlem, and B\. Roffel\(2003\)Combining prior knowledge with data driven modeling of a batch distillation column including start\-up\.Computers & chemical engineering27\(7\),pp\. 1021–1030\.Cited by:[Table 1](https://arxiv.org/html/2607.22811#S3.T1.23.21.4.1.1)\.
- \[33\]M\. Von Stosch, R\. Oliveira, J\. Peres, and S\. F\. De Azevedo\(2014\)Hybrid semi\-parametric modeling in process systems engineering: past, present and future\.Computers & Chemical Engineering60,pp\. 86–101\.Cited by:[§1](https://arxiv.org/html/2607.22811#S1.p2.1)\.

## Appendix AMeasurability assumptions for the inference functional

The functional \([1](https://arxiv.org/html/2607.22811#S1.E1)\) is well\-defined under the following standard conditions\. We assume\(Ω,ΣΩ\)\(\\Omega,\\Sigma\_\{\\Omega\}\)is measurable withΩ′∈ΣΩ\\Omega^\{\\prime\}\\in\\Sigma\_\{\\Omega\}; thatω↦l​\(φ,ω\)\\omega\\mapsto l\(\\varphi,\\omega\)is measurable for each queryφ\\varphi; that the belief induces a finite measureB𝜽,xB\_\{\\boldsymbol\{\\theta\},x\}onΩ′\\Omega^\{\\prime\}; and thatl​\(φ,⋅\)∈L1​\(Ω′,B𝜽,x\)l\(\\varphi,\\cdot\)\\in L^\{1\}\(\\Omega^\{\\prime\},B\_\{\\boldsymbol\{\\theta\},x\}\)\. Dirac beliefs correspond toB𝜽,x=δω∗​\(x\)B\_\{\\boldsymbol\{\\theta\},x\}=\\delta\_\{\\omega^\{\\ast\}\(x\)\}, recovering pointwise evaluationF𝜽,x​\(φ\)=l​\(φ,ω∗​\(x\)\)F\_\{\\boldsymbol\{\\theta\},x\}\(\\varphi\)=l\(\\varphi,\\omega^\{\\ast\}\(x\)\)\.

## Appendix BProof of Proposition[3](https://arxiv.org/html/2607.22811#Thmtheorem3)

###### Proof\.

For \(1\), the partitionΩviolτ​\(φ\)\\Omega\_\{\\mathrm\{viol\}\}^\{\\tau\}\(\\varphi\)versusΩadmτ​\(φ\)\\Omega\_\{\\mathrm\{adm\}\}^\{\\tau\}\(\\varphi\)is defined directly froml​\(φ,⋅\)l\(\\varphi,\\cdot\)andτ\\tau, hence depends only on\(L,μ,l\)\(L,\\mu,l\)and not onB𝜽,xB\_\{\\boldsymbol\{\\theta\},x\}\. For the integral identity in the hard caseτ=0\\tau=0, note that under the hypothesisl​\(φ,⋅\)≥0l\(\\varphi,\\cdot\)\\geq 0the setΩviol0​\(φ\)=\{ω∈Ω′:l​\(φ,ω\)≤0\}\\Omega\_\{\\mathrm\{viol\}\}^\{0\}\(\\varphi\)=\\\{\\omega\\in\\Omega^\{\\prime\}:l\(\\varphi,\\omega\)\\leq 0\\\}coincides with\{ω∈Ω′:l​\(φ,ω\)=0\}\\\{\\omega\\in\\Omega^\{\\prime\}:l\(\\varphi,\\omega\)=0\\\}\. Hence the integrand vanishes onΩviol0​\(φ\)\\Omega\_\{\\mathrm\{viol\}\}^\{0\}\(\\varphi\), that subset contributes zero, and

F𝜽,x​\(φ\)=∫Ω′l​\(φ,ω\)​dB𝜽,x​\(ω\)=∫Ωadm0​\(φ\)l​\(φ,ω\)​dB𝜽,x​\(ω\)\.F\_\{\\boldsymbol\{\\theta\},x\}\(\\varphi\)=\\int\_\{\\Omega^\{\\prime\}\}l\(\\varphi,\\omega\)\\,\\mathrm\{d\}B\_\{\\boldsymbol\{\\theta\},x\}\(\\omega\)=\\int\_\{\\Omega\_\{\\mathrm\{adm\}\}^\{0\}\(\\varphi\)\}l\(\\varphi,\\omega\)\\,\\mathrm\{d\}B\_\{\\boldsymbol\{\\theta\},x\}\(\\omega\)\.For \(2\), sincel​\(φ,⋅\)l\(\\varphi,\\cdot\)is measurable, substitutingB𝜽,x=δω∗​\(x\)B\_\{\\boldsymbol\{\\theta\},x\}=\\delta\_\{\\omega^\{\\ast\}\(x\)\}intoF𝜽,x​\(φ\)=∫Ω′l​\(φ,ω\)​dB𝜽,x​\(ω\)F\_\{\\boldsymbol\{\\theta\},x\}\(\\varphi\)=\\int\_\{\\Omega^\{\\prime\}\}l\(\\varphi,\\omega\)\\,\\mathrm\{d\}B\_\{\\boldsymbol\{\\theta\},x\}\(\\omega\)collapses the integral tol​\(φ,ω∗​\(x\)\)l\(\\varphi,\\omega^\{\\ast\}\(x\)\)\. By the definition ofΩadmτ​\(φ\)\\Omega\_\{\\mathrm\{adm\}\}^\{\\tau\}\(\\varphi\), this value exceedsτ\\tauiffω∗​\(x\)∈Ωadmτ​\(φ\)\\omega^\{\\ast\}\(x\)\\in\\Omega\_\{\\mathrm\{adm\}\}^\{\\tau\}\(\\varphi\), which is the stated admissibility criterion\. For non\-Dirac beliefs,F𝜽,x​\(φ\)F\_\{\\boldsymbol\{\\theta\},x\}\(\\varphi\)is theB𝜽,xB\_\{\\boldsymbol\{\\theta\},x\}\-weighted average ofl​\(φ,⋅\)l\(\\varphi,\\cdot\), so \(in the hard case\) its value depends only on howB𝜽,xB\_\{\\boldsymbol\{\\theta\},x\}allocates mass withinΩadm0​\(φ\)\\Omega\_\{\\mathrm\{adm\}\}^\{0\}\(\\varphi\)\. ∎

## Appendix CH2N principles \(full statements with algorithm\)

P1 \(Structure as language\)\.Mechanistic equations and modular composition graphs determine the*symbols*and well\-formed expressions ofLL\. Unknown closures and latent states are introduced as explicit symbols inLLso that they can be quantified over by interpretations\.

P2 \(Semantics as evaluation\)\.The semanticsμ\\muspecifies how formulas are evaluated in an interpretation \(Boolean truth, fuzzy degree, etc\.\)\. The logic functionllis the*operational*object used inside the integral: it implements admissibility and scoring by selecting or reweighting semantic values, e\.g\. hard satisfaction indicators or residual\-based penalties\.

P3 \(Learning as belief conditioned on input\)\.Learned modules induce a belief componentb𝜽b\_\{\\boldsymbol\{\\theta\}\}by producing a non\-negative density/weightb𝜽,x​\(ω\)b\_\{\\boldsymbol\{\\theta\},x\}\(\\omega\)over interpretationsω∈Ω′\\omega\\in\\Omega^\{\\prime\}given inputxx\. Deterministic learners correspond to Dirac beliefs concentrated atω∗​\(x\)\\omega^\{\\ast\}\(x\); Bayesian/ensemble/energy\-based learners induce non\-degenerate beliefs\.

P4 \(Constraints and structures live on the logic side\)\.Validity domains, trust regions, hard constraints \(e\.g\. conservation, positivity\), and regime declarations are encoded as*structural admissibility*rather than plausibility: \(i\) by restricting the integration domainΩ′\\Omega^\{\\prime\}to an admissible subset ofΩ\\Omega, and/or \(ii\) by enforcing selection/penalties insidell\. This separates*what is admissible*\(logic/structure\) from*what is plausible*\(belief/learning\)\.

P5 \(Factorization mirrors architectural composition\)\.When the hybrid design factorizes \(serial, parallel, mixture\), the NeSy construction should make this explicit by factorizinglland/orb𝜽,xb\_\{\\boldsymbol\{\\theta\},x\}over corresponding sub\-interpretations and sub\-formulas\. This yields comparable semantics across patterns and clarifies where uncertainty and structures enter\.

Hence, comparing hybrid modeling designs reduces to comparing how they change \(i\) the language/semantics, \(ii\) admissibility/scoring viall, \(iii\) the unknowns integrated over inΩ′\\Omega^\{\\prime\}, and \(iv\) the belief factorisation and dispersion\.

Input:Hybrid modeling design with mechanistic partMαM\_\{\\alpha\}, learned part\(s\)gβg\_\{\\beta\}, compositionComp\\operatorname\{Comp\}, and structures𝒞\\mathcal\{C\}

Output:NeSy tuple

\(L,μ,Ω,b𝜽\)\(L,\\mu,\\Omega,b\_\{\\boldsymbol\{\\theta\}\}\), logic function

ll, and induced functional

F𝜽,xF\_\{\\boldsymbol\{\\theta\},x\}
1

2Build the languageLL\.Encode mechanistic equations, wiring/composition predicates, and rule symbols\. Introduce explicit symbols for unknown closures/residuals/gates/slacks and latent states\.

3Define the interpretation spaceΩ\\Omega\.Let

ω∈Ω\\omega\\in\\Omegaassign values to the unknown entities introduced in

LL\(e\.g\.

α,ψ,z\\alpha,\\psi,z, gating variables, mixture weights, memberships, noise, slack variables\)\.

4Choose semanticsμ\\muand define the logic functionll\.Fix an evaluation scheme \(Boolean, fuzzy, residual\-/trajectory\-based\) and define

l​\(φ,ω\)l\(\\varphi,\\omega\)accordingly:*hard*admissibility \(indicator selectors\),*tolerant*admissibility \(thresholded residuals\), or*graded*scoring \(penalties/likelihood\-like scores\)\.

5Encode structures𝒞\\mathcal\{C\}on the logic side\.Implement hard feasibility by restricting the measurable integration domain to

Ω′⊆Ω\\Omega^\{\\prime\}\\subseteq\\Omega, and/or implement soft feasibility by penalties/selection inside

ll\. \(These choices affect

FFunless one renormalizes/conditions the belief\.\)

6Define beliefbθb\_\{\\boldsymbol\{\\theta\}\}\(conditional on inputsxx\)\.Specify a non\-negative density function

b𝜽,x​\(ω\)b\_\{\\boldsymbol\{\\theta\},x\}\(\\omega\)w\.r\.t\.

mminduced by the learned components \(posterior, likelihood\-weighted prior, ensemble mixture\)\. Dirac beliefs arise as degenerate cases; weights can be normalized when probabilities are required\.

Assemble inference\.Instantiate

F𝜽,x​\(φ\)F\_\{\\boldsymbol\{\\theta\},x\}\(\\varphi\)by Equation \([1](https://arxiv.org/html/2607.22811#S1.E1)\) over

Ω′\\Omega^\{\\prime\}\. If the pattern factorizes \(serial, parallel, mixture\), make the factorization of

lland/or

b𝜽,xb\_\{\\boldsymbol\{\\theta\},x\}explicit\.

Algorithm 1H2Ntranslation
## Appendix DMonte Carlo procedure for H2N metrics in the case study

WithΩ=𝔹𝒜\\Omega=\\mathbb\{B\}^\{\\mathcal\{A\}\}and counting measure, the inference functional reduces to a sum overω\\omega:F𝜽,x​\(φ\)=∑ωl​\(φ,ω\)​b𝜽,x​\(ω\)F\_\{\\boldsymbol\{\\theta\},x\}\(\\varphi\)=\\sum\_\{\\omega\}l\(\\varphi,\\omega\)b\_\{\\boldsymbol\{\\theta\},x\}\(\\omega\)\. NoiseCut yields Dirac belief on the first\-layer functionsF^m\\hat\{F\}\_\{m\}and an independent Bernoulli belief on each second\-layer atomoKo\_\{K\}with parameterpK=n1​\(K\)/\(n0​\(K\)\+n1​\(K\)\)p\_\{K\}=n\_\{1\}\(K\)/\(n\_\{0\}\(K\)\+n\_\{1\}\(K\)\)\(andpK=12p\_\{K\}=\\tfrac\{1\}\{2\}whenn0​\(K\)\+n1​\(K\)=0n\_\{0\}\(K\)\+n\_\{1\}\(K\)=0\)\. Under this factorisation:

#### BD\.

Closed formBD=∑KpK​\(1−pK\)\\mathrm\{BD\}=\\sum\_\{K\}p\_\{K\}\(1\-p\_\{K\}\)because the first\-layer is Dirac \(zero variance\) and the2M2^\{M\}second\-layer atoms are independent\.

#### SVR\.

We drawT=20,000T=20\{,\}000Monte Carlo samplesFO\(t\)∈\{0,1\}2MF\_\{O\}^\{\(t\)\}\\in\\\{0,1\\\}^\{2^\{M\}\}by independent Bernoulli sampling per row\. With Dirac first\-layer, the number of violated training samples decomposes per row:

viol​\(ω\(t\)\)=∑K\[n1​\(K\)​𝕀​\[FO\(t\)​\(K\)=0\]\+n0​\(K\)​𝕀​\[FO\(t\)​\(K\)=1\]\]\.\\mathrm\{viol\}\(\\omega^\{\(t\)\}\)=\\sum\_\{K\}\\bigl\[n\_\{1\}\(K\)\\mathbb\{I\}\\\!\\left\[F\_\{O\}^\{\(t\)\}\(K\)=0\\right\]\+n\_\{0\}\(K\)\\mathbb\{I\}\\\!\\left\[F\_\{O\}^\{\(t\)\}\(K\)=1\\right\]\\bigr\]\.The tolerance enters as the violation budgetκ\\kappainsideltol=𝕀​\[viol≤κ\]l\_\{\\mathrm\{tol\}\}=\\mathbb\{I\}\\\!\\left\[\\mathrm\{viol\}\\leq\\kappa\\right\], so an interpretation is admissible whenviol≤κ\\mathrm\{viol\}\\leq\\kappaand violating otherwise\. HenceSVR=1−1T​∑t𝕀​\[viol​\(ω\(t\)\)≤κ\]\\mathrm\{SVR\}=1\-\\tfrac\{1\}\{T\}\\sum\_\{t\}\\mathbb\{I\}\\\!\\left\[\\mathrm\{viol\}\(\\omega^\{\(t\)\}\)\\leq\\kappa\\right\]is the Monte Carlo estimate of the belief mass on violating interpretations\.T=20,000T=20\{,\}000yields standard errors below0\.0050\.005; we verified stability by replicate runs with different MC seeds\. \(BD requires no sampling: it is the closed form above\.\)

## Appendix ETolerance sweep

This appendix isolates the role of the toleranceτ\\tau\(instantiated here and in the case study of Section[5](https://arxiv.org/html/2607.22811#S5)by the violation budgetκ\\kappa\) as a purely*logic\-side*design choice, and turns the two contrasting tolerance properties of the metrics established in Section[4](https://arxiv.org/html/2607.22811#S4)into a measurement\. SVR is tolerance\-dependent \(property \(ii\) of the SVR\): it reads the admissibility partitionΩadmκ/Ωviolκ\\Omega\_\{\\mathrm\{adm\}\}^\{\\kappa\}/\\Omega\_\{\\mathrm\{viol\}\}^\{\\kappa\}, whichκ\\kapparesizes\. BD is tolerance\-independent \(property \(ii\) of the BD\): as a second moment of the beliefb𝜽b\_\{\\boldsymbol\{\\theta\}\}it never referencesκ\\kappaor the admissibility partition\. Sweepingκ\\kappaat a fixed belief should therefore move SVR across its full range\[0,1\]\[0,1\]while leaving BD \(and the deployed predictor’s accuracy\) exactly fixed\.

We fix one learned model per seed \(struct\[4,4,4,4\]\[4,4,4,4\],10%10\\%label noise,50%50\\%training fraction,1010seeds\) and re\-evaluate the H2N metrics under eleven budgets, withκ/S\\kappa/Sranging from0to0\.50\.5\. The fit is performed once: neither the first\-layer Dirac factorsF^m\\hat\{F\}\_\{m\}nor the row\-Bernoulli parameterspKp\_\{K\}are refitted between budgets\. Only the logic functionltol=𝕀​\[viol​\(ω\)≤κ\]l\_\{\\mathrm\{tol\}\}=\\mathbb\{I\}\\\!\\left\[\\mathrm\{viol\}\(\\omega\)\\leq\\kappa\\right\]changes, so any variation across the columns of Table[E1](https://arxiv.org/html/2607.22811#A5.T1)is attributable to the logic side alone\.

Two columns are flat by construction\. Accuracy is constant at0\.9980\.998because the deployed predictor is the per\-row threshold of the Bernoulli belief, which does not referenceκ\\kappa\. BD is constant at0\.7830\.783because it is a functional ofb𝜽b\_\{\\boldsymbol\{\\theta\}\}alone; its within\-seed spread across the eleven budgets is exactly0, so the column mean and every per\-seed value coincide\. The two logic\-side columns instead move monotonically and are mirror images, sincePr⁡\[viol≤κ\]=1−SVR\\Pr\[\\mathrm\{viol\}\\leq\\kappa\]=1\-\\mathrm\{SVR\}\. At budgets below the noise floor \(κ/S∈\{0,0\.05\}\\kappa/S\\in\\\{0,0\.05\\\}\) essentially every sampled interpretation exceeds the budget \(the cleanest belief consistent with the mechanistic structure still misclassifies the∼10%\\sim\\\!10\\%of flipped training labels\) so SVR saturates at1\.0001\.000and the admissible mass is0\. As the budget crosses the noise level nearκ/S=0\.10\\kappa/S=0\.10, SVR collapses \(0\.4030\.403atκ/S=0\.10\\kappa/S=0\.10,0\.1600\.160at0\.150\.15,0\.0550\.055at0\.200\.20\) and reaches0byκ/S=0\.40\\kappa/S=0\.40, where the budget comfortably absorbs the label noise and every sampled interpretation is admissible\.

This makes the SVR/BD distinction of Section[4](https://arxiv.org/html/2607.22811#S4)concrete and complements Appendix[G](https://arxiv.org/html/2607.22811#A7): there, fixingκ\\kappaand varying the noise moves BD while leaving SVR near zero; here, fixing the noise and varyingκ\\kappamoves SVR across its entire range while leaving BD untouched\.

Table E1:Tolerance dependence of the H2N metrics \(struct\[4,4,4,4\]\[4,4,4,4\],10%10\\%noise,50%50\\%train\)\. Mean over1010seeds\. Accuracy and BD are independent ofκ\\kappa; SVR and the admissible massPr⁡\[viol≤κ\]=1−SVR\\Pr\[\\mathrm\{viol\}\\leq\\kappa\]=1\-\\mathrm\{SVR\}vary monotonically\.
## Appendix FOOD experiment: full numbers

This appendix gives the full numbers behind Result 2 of Section[5](https://arxiv.org/html/2607.22811#S5)and discusses the structural out\-of\-distribution \(OOD\) protocol\. The mechanistic partition routes each input throughM=4M=4first\-layer modules whose joint output indexes one of2M=162^\{M\}=16second\-layer rows; for the randomly drawn ground\-truth modules all1616rows are reachable\. We holdNood∈\{1,2,4,6\}N\_\{\\mathrm\{ood\}\}\\in\\\{1,2,4,6\\\}of these rows out of training entirely \(with10%10\\%label noise and budgetκ=⌊0\.10​S⌋\\kappa=\\lfloor 0\.10\\,S\\rfloor\), so that any test input landing on a held\-out row forces a prediction on a region ofΩ′\\Omega^\{\\prime\}that received no training support: structural extrapolation rather than interpolation\. Because the row partition is*logical*\(fixed byΩ′\\Omega^\{\\prime\}rather than by the learned predictor\) belief dispersion decomposes exactly asBD=BDseen\+BDunseen\\mathrm\{BD\}=\\mathrm\{BD\}\_\{\\text\{seen\}\}\+\\mathrm\{BD\}\_\{\\text\{unseen\}\}, withBDunseen=14​\|\{K:n0​\(K\)\+n1​\(K\)=0\}\|\\mathrm\{BD\}\_\{\\text\{unseen\}\}=\\tfrac\{1\}\{4\}\\,\|\\\{K:n\_\{0\}\(K\)\+n\_\{1\}\(K\)=0\\\}\|collecting the maximum\-entropy variance \(pK=12p\_\{K\}=\\tfrac\{1\}\{2\}\) of the uncovered rows\. By Remark[4](https://arxiv.org/html/2607.22811#Thmtheorem4)this term is read off the*learned*first\-layer factorisation alone, before any OOD label is observed, which is what makes it a deployment\-time quantity rather than a post\-hoc one\.

![Refer to caption](https://arxiv.org/html/2607.22811v1/x2.png)Figure F1:Structural OOD shift \(struct\[4,4,4,4\]\[4,4,4,4\],10%10\\%noise,1010seeds\)\. The additive H2N decompositionBD=BDseen\+BDunseen\\mathrm\{BD\}=\\mathrm\{BD\}\_\{\\text\{seen\}\}\+\\mathrm\{BD\}\_\{\\text\{unseen\}\}:BDunseen\\mathrm\{BD\}\_\{\\text\{unseen\}\}, computed from row coverage alone, grows monotonically withNoodN\_\{\\mathrm\{ood\}\}and identifies the belief mass on the held\-out region without observing OOD test data \(whiskers:±1\\pm 1std ofBDunseen\\mathrm\{BD\}\_\{\\text\{unseen\}\}\)\.Table F1:Structural OOD shift on second\-layer rows \(struct\[4,4,4,4\]\[4,4,4,4\],10%10\\%label noise,κ=⌊0\.10​S⌋\\kappa=\\lfloor 0\.10\\,S\\rfloor\)\. Mean \(std\) over1010seeds\. All2M=162^\{M\}=16second\-layer rows are reachable for the randomly drawn ground\-truth modules; the experiment holds outNoodN\_\{\\mathrm\{ood\}\}of them entirely from training\. “unseen frac\.” is the mean fraction of rows left uncovered under the*learned*factorisation\.As shown in Figure[F1](https://arxiv.org/html/2607.22811#A6.F1)and Table[F1](https://arxiv.org/html/2607.22811#A6.T1), meanBDunseen\\mathrm\{BD\}\_\{\\text\{unseen\}\}rises monotonically with the hold\-out size \(0\.25,0\.50,0\.68,0\.950\.25,0\.50,0\.68,0\.95\), and at the smaller hold\-outs \(Nood=1,2N\_\{\\mathrm\{ood\}\}=1,2\) it is*deterministic*across seeds \(std0\): the learned first\-layer partition reliably keeps every held\-out row as a distinct uncovered row, soBDunseen=14​Nood\\mathrm\{BD\}\_\{\\text\{unseen\}\}=\\tfrac\{1\}\{4\}N\_\{\\mathrm\{ood\}\}exactly\. At the larger hold\-outs \(Nood=4,6N\_\{\\mathrm\{ood\}\}=4,6\) a seed\-to\-seed variance appears \(std0\.470\.47and0\.730\.73\): in33of the1010seeds at each level the learner recovers a*label\-equivalent*factorisation that maps held\-out ground\-truth rows into already\-observed learned equivalence classes, givingBDunseen=0\\mathrm\{BD\}\_\{\\text\{unseen\}\}=0\. This is exactly what the decomposition is meant to report:BDunseen\\mathrm\{BD\}\_\{\\text\{unseen\}\}measures the belief mass on rows uncovered*under the actually\-learned factorisation*, the only one accessible at deployment time, not under a hypothetical ground\-truth one\.

The OOD\-accuracy variance has a separate origin\. A held\-out row receives the maximum\-entropy beliefpK=12p\_\{K\}=\\tfrac\{1\}\{2\}, whose thresholded prediction is either right or wrong for that entire row; when few rows are held out \(especiallyNood=1N\_\{\\mathrm\{ood\}\}=1\) the per\-seed OOD accuracy is therefore near\-bimodal \(std0\.470\.47\) and concentrates as more rows are averaged\. Despite this variance the qualitative pattern is robust: in\-distribution accuracy is essentially unchanged acrossNoodN\_\{\\mathrm\{ood\}\}while OOD accuracy is far lower, andBDunseen\\mathrm\{BD\}\_\{\\text\{unseen\}\}\(computed before any OOD label is seen\) grows with the hold\-out\. Finally,BDunseen\\mathrm\{BD\}\_\{\\text\{unseen\}\}is informative about*severity*, not only presence: pooling all seeds, those withBDunseen\>0\\mathrm\{BD\}\_\{\\text\{unseen\}\}\>0average0\.390\.39OOD accuracy against0\.690\.69for the seeds where the learner absorbed the held\-out rows \(BDunseen=0\\mathrm\{BD\}\_\{\\text\{unseen\}\}=0\)\. A zeroBDunseen\\mathrm\{BD\}\_\{\\text\{unseen\}\}thus does not certify OOD robustness \(those models still fall well below their in\-distribution accuracy\) but signals that the model*believes*it has coverage, which is the deployment\-time statement the decomposition is designed to make\.

## Appendix GSVR and BD are complementary

SVR and BD probe opposite sides of the logic–belief separation \(Proposition[3](https://arxiv.org/html/2607.22811#Thmtheorem3)\) and are, by construction, free to vary independently\. SVR is a functional of the admissibility partitionΩadmτ/Ωviolτ\\Omega\_\{\\mathrm\{adm\}\}^\{\\tau\}/\\Omega\_\{\\mathrm\{viol\}\}^\{\\tau\}, which is a property of\(L,μ,l\)\(L,\\mu,l\), weighted by the belief, and it is governed by the toleranceτ\\tau\. BD is the trace of the belief covariance and depends onb𝜽,xb\_\{\\boldsymbol\{\\theta\},x\}alone, never referencingllorτ\\tau\. The two therefore read out distinct degrees of freedom: moving the tolerance at a fixed belief slides SVR while leaving BD unchanged \(Appendix[E](https://arxiv.org/html/2607.22811#A5)\), and concentrating or dispersing the belief at a fixed admissibility moves BD while leaving SVR fixed\.

All four combinations are realisable, so neither metric can be recovered from the other\. A model may disperse its belief widely \(large BD\) yet rarely violate its constraints \(small SVR\) when the admissible region is large or the tolerance is loose; conversely, a model with a sharp belief \(small BD\) can still place that belief largely outside the admissible region \(large SVR\)\. SVR thus answers whether the learned belief respects the mechanistic structure, while BD answers how concentrated the learned plausibility is; neither subsumes the other, and reporting both separates structural violations from epistemic uncertainty rather than collapsing them into a single accuracy figure\.

The case study \(Section[5](https://arxiv.org/html/2607.22811#S5)\) exhibits this decoupling directly\. Under a fixed toleranceκ=⌊0\.5​S⌋\\kappa=\\lfloor 0\.5\\,S\\rfloor, Table[2](https://arxiv.org/html/2607.22811#S5.T2)and Figure[G1](https://arxiv.org/html/2607.22811#A7.F1)show SVR pinned at or near zero across a wide noise range \(≤0\.011\\leq 0\.011through20%20\\%noise, and still only0\.0830\.083at30%30\\%\): the generous budget keeps almost all sampled interpretations admissible\. BD, by contrast, climbs monotonically over the same range, from0at0%0\\%noise to3\.5393\.539at30%30\\%and toward its ceiling2M/4=42^\{M\}/4=4as the row beliefs approach maximum entropy\. SVR departs from zero only once the noise pushes the realized violation count past the fixed budget, rising to0\.4600\.460at50%50\\%, by which point BD has already saturated\. The two metrics therefore move on different schedules: BD tracks the growth of epistemic uncertainty in the learned rows from the very first noise increment, whereas SVR is a logic\-side feasibility statement that fires only when the chosen tolerance is exceeded\.

![Refer to caption](https://arxiv.org/html/2607.22811v1/x3.png)Figure G1:Noise sweep,5050seeds per level \(struct\[4,4,4,4\]\[4,4,4,4\],κ=⌊0\.5​S⌋\\kappa=\\lfloor 0\.5\\,S\\rfloor\)\. Mean test accuracy, mean belief dispersion BD, and mean SVR \(shaded band:±1\\pm 1standard deviation across seeds\) as functions of the label\-noise level\. The y\-axis for the two metrics has been normalized for better visualizations\.

Similar Articles

Neuro-Symbolic AI for LEED compliance: Document-Centric Benchmarking, Deterministic Numeric Checking, and When Multimodal Hurts

arXiv cs.AI

This paper introduces a neuro-symbolic pipeline for automating LEED v4.1 BD+C compliance verification using small locally deployed language models and deterministic numeric checking. Experiments on four university buildings show that a 4B model outperforms an 8B model, and the deterministic checker corrects arithmetic errors on key credits, though multimodal inputs reduce accuracy.

The Dynamic Concept Graph: Toward Persistent Multimodal World Models for Artificial Intelligence

Reddit r/ArtificialInteligence

This proposal introduces the Dynamic Concept Graph (DCG), a hybrid cognitive architecture that combines neural representation learning, symbolic knowledge structures, multimodal perception, and analogical reasoning to provide persistent, evolving world models for AI, addressing limitations such as inconsistent reasoning and lack of causal understanding in large language models.