Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning

arXiv cs.AI Papers

Summary

The paper distinguishes generated context from governed state in clinical AI, arguing that accountable longitudinal reasoning requires a governed patient state representation and proposes a tiered governance standard and maturity framework.

arXiv:2608.14804v1 Announce Type: new Abstract: Large language models (LLMs) have become the dominant interface of clinical artificial intelligence, yet the interface they expose (text in, text out, one context window at a time) maintains no explicit, persistent, governed representation of what is currently true about a patient. This paper argues that longitudinal clinical reasoning is a state-estimation problem under partial observability, and that the axis on which clinical AI succeeds or fails is not the fluency of the model reading the record but the governance of the patient state it reasons over. We distinguish generated context from governed state; separate five objects that clinical AI habitually conflates (true state, observations, evidence, belief, and simulated state); define a tiered governance standard against which any clinical AI system can be audited; and show that an operational definition of accountability decomposes into four information requirements: an immutable evidence ledger with awareness-time versioning, a belief state distinct from accumulated evidence, an observation-process model, and claim-level causal typing. We are explicit that this decomposition is analytic rather than a necessity theorem, and that its value is conceptual hygiene: it converts "accountable clinical AI" from a slogan into an audit instrument. A six-level maturity framework separates what a system makes governable from what it can compute, locating current LLM-centric practice at high capability but low maturity. The paper is fully self-contained: the four research questions the framework poses are stated in the introduction, and the conclusion records what the paper establishes toward each; future work develops the buildable core of the architecture and the research program toward full Clinical World Models. No empirical result is claimed here.
Original Article
View Cached Full Text

Cached at: 08/18/26, 10:09 AM

# Generated Context versus Governed State:Functional Conditions for Accountable Longitudinal ClinicalReasoning
Source: [https://arxiv.org/html/2608.14804](https://arxiv.org/html/2608.14804)
## Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical ReasoningThanks:MyndwareMed builds governed clinical\-state infrastructure for healthcare organizations\. The evidence layer specified in this line of work \(maturity Level 3\) operates in production; the belief layer \(Level 4\) is in engineering and the full Clinical World Model \(Level 5\) in research\.[https://www\.myndwaremed\.com](https://www.myndwaremed.com/)

Victor Lorena de Farias SouzaThanks:victor\.lorena@myndware\.comAffiliation:MyndwareMed

August 2026

###### Abstract

Large language models \(LLMs\) have become the dominant interface of clinical artificial intelligence, yet the interface they expose — text in, text out, one context window at a time — maintains no explicit, persistent, governed representation of what is currently true about a patient\. This paper argues that longitudinal clinical reasoning is a state\-estimation problem under partial observability, and that the axis on which clinical AI succeeds or fails is not the fluency of the model reading the record but the*governance*of the patient state it reasons over\. We distinguish generated context from governed state; separate five objects that clinical AI habitually conflates \(true state, observations, evidence, belief, and simulated state\); define a tiered governance standard against which any clinical AI system can be audited; and show that an operational definition of accountability decomposes into four information requirements — an immutable evidence ledger with awareness\-time versioning, a belief state distinct from accumulated evidence, an observation\-process model, and claim\-level causal typing\. We are explicit that this decomposition is analytic rather than a necessity theorem, and that its value is conceptual hygiene: it converts “accountable clinical AI” from a slogan into an audit instrument\. A six\-level maturity framework separates what a system makes governable from what it can compute, locating current LLM\-centric practice at high capability but low maturity\. The paper is fully self\-contained: the four research questions the framework poses are stated in the introduction, and the conclusion records what the paper establishes toward each; future work develops the buildable core of the architecture and the research program toward full Clinical World Models\. No empirical result is claimed here\.

## 1Introduction

For sixty years, clinical information systems have stored documents about patients rather than the patient’s state itself\. Paper charts became scanned charts; scanned charts became structured documents; departmental records became Clinical Document Architecture \(CDA\) exchanges and then Fast Healthcare Interoperability Resources \(FHIR\) bundles\. Each generation made clinical documents easier to create, exchange, and display — while the work of reconstructing “what is true of this patient right now” stayed where it had always been: in the head of whoever reads the chart\. Large language models are the first technology fluent enough to perform that reconstruction on demand, and their fluency makes it easy to mistake reconstruction\-on\-demand for something standard LLM\-centric architectures do not inherently maintain: an explicit, accountable, continuously governed representation of the patient\. That mistake, and its architectural remedy, are the subject of this paper\.

In the last three years, clinical artificial intelligence has become nearly synonymous with LLMs\[[1](https://arxiv.org/html/2608.14804#bib.bib1),[3](https://arxiv.org/html/2608.14804#bib.bib3)\]\. Hospitals, payers, and health\-technology vendors now route documentation, summarization, coding assistance, and parts of decision support through general\-purpose models, most commonly over retrieval\-augmented substrates\[[4](https://arxiv.org/html/2608.14804#bib.bib4)\]\. The productivity gains are real, and this paper does not dispute them\. But the interface a language model exposes carries no guarantee of a persistent notion of what is currently true about a person’s physiology, no versioned record of how that truth changed, and no native mechanism for distinguishing an observed fact from an inferred one, or an absence of evidence from evidence of absence\. Whatever latent structure such a model may learn internally, none of it is exposed as an object that a hospital can query, audit, correct, or govern\.

We should be precise about how contrarian our thesis is, and is not\. It is not a claim that language models should be abandoned, nor that neural representation is the problem — we concede below that a language model can be wrapped in governed external memory, and that the real contrast is generated context versus governed state, not neural versus symbolic\. The operative claim is narrower and, we think, sharper: governance of patient state — not the fluency of the model reading it — is the axis on which accountable longitudinal clinical reasoning succeeds or fails\.

One reframing organizes everything that follows, and we state it here rather than let it emerge by page eight: clinical reasoning over a longitudinal record is state estimation under partial observability\[[14](https://arxiv.org/html/2608.14804#bib.bib14)\]— closer in kind to robotics, econometrics, or epidemiological modeling than to summarization — that happens to be documented primarily in natural language\. The patient’s true state is never possessed; a care process observes it selectively; a system forms and revises beliefs from the evidence that survives\. Once the problem is seen this way, the question is not which model reads the record best but what artifact plays the role of the estimator’s state — and the deficit of current practice is that no governed artifact plays it\. Put as one sentence: current clinical AI research is increasingly learning patient dynamics; this paper asks a different question — what state must a clinical AI system maintain so that its longitudinal reasoning can be reconstructed, corrected, audited, and safely separated from simulation? Everything in this paper unpacks that question\.

#### Research questions\.

Fluency is now abundant; governed clinical state is not — and the gap between the two is what makes the framework empirical rather than rhetorical\. Four research questions follow from it, stated here so the reader knows from the outset what the paper is trying to establish; the conclusion records what this paper achieves toward each, and what remains to be answered next\.RQ1— does an architecture that maintains persistent, governed patient state answer questions about a patient’s current and historical status more consistently, and with better\-supported provenance, than long\-context or retrieval baselines?RQ2— do dynamics over a governed belief state derived from governed temporal evidence predict multi\-horizon clinical outcomes better than next\-event sequence modeling of the raw record?RQ3— does a deterministic verifier materially reduce clinically impossible outputs without materially reducing sensitivity?RQ4— does modeling the process that generates clinical observations recover a material fraction of the performance lost when models transport across institutions?

#### Contributions\.

This paper makes three contributions\. First, a precise statement of the deficit of LLM\-centric clinical systems: not representation in general, but concrete governance properties of state, organized into a tiered audit standard \(Section[3](https://arxiv.org/html/2608.14804#S3)\)\. Second, a vocabulary that keeps separate the five objects clinical AI habitually conflates, together with an epistemic\-type discipline for assertions and absences \(Section[4](https://arxiv.org/html/2608.14804#S4)\)\. Third, an operational definition of accountability and its decomposition into four information requirements, with an honest account of the decomposition’s logical character, plus a six\-level maturity framework that separates representation maturity from computational capability \(Sections[6](https://arxiv.org/html/2608.14804#S6)–[7](https://arxiv.org/html/2608.14804#S7)\)\. This paper is self\-contained: it is a conceptual and architectural contribution, complete on its own terms, and no companion reading is required to evaluate it\. Future work develops the buildable core — the full specification of the evidence ledger, the operational belief state, and the governed update operator, with its implementability at population scale — and the research program toward full Clinical World Models: the observation\-process model, causal qualification, action\-conditioned dynamics, and quarantined simulation\. Everything in this paper is specification and argument; no empirical or clinical result is claimed\.

## 2Background: where clinical AI stands

General\-purpose LLMs encode substantial medical knowledge — instruction\-tuned models approach expert\-level performance on licensing\-exam\-style question answering\[[1](https://arxiv.org/html/2608.14804#bib.bib1),[2](https://arxiv.org/html/2608.14804#bib.bib2)\], with applications spanning documentation, triage, and decision support\[[3](https://arxiv.org/html/2608.14804#bib.bib3)\]— and models trained on next\-token prediction demonstrably learn internal representations beyond surface statistics\. Nothing in this paper depends on denying that\. What benchmark performance does not establish is that a model answering medical questions well thereby maintains a consistent, auditable representation of a specific patient over years of documentation: a high exam score is a statement about a model’s marginal distribution over medical text, not about its ability to track, in a form anyone can inspect, which of two conflicting medication lists is current and who documented each\.

Retrieval\-augmented generation \(RAG\)\[[4](https://arxiv.org/html/2608.14804#bib.bib4)\]grounds answers in a specific chart rather than the training distribution, and it is the substrate many clinical AI products are built on; agent frameworks chain calls and add tools and memory\. But retrieval is a search operation over a document store, not state estimation over a model of the patient, and each agent in a chain still reasons over a transient context rather than a shared, governed state object — chaining language models does not, by itself, produce a state machine, although one can build a governed state machine and let language models operate against it\. The agentic frontier is, in fact, already moving in this direction at the knowledge layer: Agents\-K1\[[21](https://arxiv.org/html/2608.14804#bib.bib21)\]replaces flat retrieval with an agent\-native knowledge substrate — scientific corpora transformed into structured graphs that preserve entities, claims, evidence, and provenance for multi\-hop reasoning\. That movement supports this paper’s thesis while marking exactly where it stops short: a claims\-and\-evidence graph over documents governs what the literature*says*; it is not a belief state about an evolving individual patient, and it carries none of the multitemporal, correction, or simulation discipline that longitudinal clinical accountability requires\. Clinical decision support, meanwhile, sets an instructive credibility bar from an earlier era: a rule\-based alert can be traced, line by line, to the guideline that produced it, and across regulatory regimes the ability to explain a recommendation materially eases deployment\. Purely generative systems do not clear this bar by default — a point this paper develops as a governance property, not a capability claim\.

An autoregressive language model is trained to estimate

P⁡\(ti∣t1,t2,…,ti−1\)\.P\(t\_\{i\}\\mid t\_\{1\},t\_\{2\},\\ldots,t\_\{i\-1\}\)\.\(1\)Two inferences are commonly drawn from equation \([1](https://arxiv.org/html/2608.14804#S3.E1)\), and only one is valid\. The invalid inference is that such a model contains no representation of state, transition, or persistence — the training objective constrains the interface, not the internal solution\. The valid inference concerns the interface itself: everything the model knows about a specific patient at inference time enters through a transient context window and leaves as generated text\. Whatever latent patient representation forms inside the forward pass has no identity that survives the call, no address at which a second system could query it, no version history, no semantics an auditor could check, and no update discipline distinguishing a correction from a contradiction\.*The representation may exist; the object does not\.*Accountable longitudinal clinical reasoning requires the object\.

The claim in its strongest defensible form: a standard autoregressive language\-model architecture does not provide an externally addressable, persistent, versioned, and clinically governed patient\-state object whose invariants can be independently verified\. Figure[1](https://arxiv.org/html/2608.14804#S3.F1)shows the two regimes; the same neural machinery may participate in both, which is exactly why the difference must be located in the object, not the network\. The properties the governed regime guarantees are concrete engineering requirements:*identity*\(one canonical state object per patient, stably addressable\);*persistence*\(survival across calls, sessions, encounters, years, by construction\);*explicit semantics*\(ontology bindings, units, value sets\[[6](https://arxiv.org/html/2608.14804#bib.bib6),[7](https://arxiv.org/html/2608.14804#bib.bib7),[8](https://arxiv.org/html/2608.14804#bib.bib8)\]\);*update governance*\(writes through a defined operator, with rules for who may assert, correct, or supersede\);*provenance*\(every element traceable to the evidence that produced it\);*inspectability*;*deterministic constraints*;*calibration*;*rollback and correction*; and*historical query*\(“what did we believe at awareness timeτ\\tau, about clinical timett, and why?”\)\.

![Refer to caption](https://arxiv.org/html/2608.14804v1/figures/fig1-generated-context-vs-governed-state.png)Figure 1:Generated context versus governed state\. Left: knowledge of the patient is reconstituted into a transient buffer on every call and vanishes with it\. Right: one canonical, versioned belief state — fed by an immutable evidence ledger — is the object every component reads from and writes through, and the object an auditor queries\.We treat these properties as a normative engineering standard, not as necessary and sufficient conditions, and we organize them — together with the additional requirements each higher tier introduces — into tiers matched to stakes; Definition[1](https://arxiv.org/html/2608.14804#Thmdefinition1)states the tiers, and Table[1](https://arxiv.org/html/2608.14804#S3.T1)normalizes their requirements\.

> ###### Definition 1\(Governed Clinical State, tiered\)\. *Core*governed state requires identity, persistence, explicit semantics, provenance, update governance, and historical query — the minimum for any representation other components treat as authoritative about a patient\.*Safety\-critical*governed state additionally requires deterministic constraint validation, correction propagation, and explicit uncertainty representation — required wherever state feeds clinical decisions\.*Predictive*governed state additionally requires outcome calibration, model\-version binding, and monitored validity under distribution shift — required wherever state includes forecasts or simulations\. The definition is architecture\-neutral: it does not require that the representation be*symbolic*\(structure held in explicit, human\-readable units with declared semantics — graphs, rules, ontology\-bound assertions\),*neural*\(structure held in learned continuous parameters and activations\), or*hybrid*\(a composition of both, with defined interfaces between them\) — only that the tier’s properties hold and can be independently audited\. We use these three terms in exactly this sense throughout\.

Table 1:The tiers of Definition[1](https://arxiv.org/html/2608.14804#Thmdefinition1)\(Governed Clinical State\), normalized\. Each tier inherits the tier below it; a system is audited at the tier matching what its state is used for\.Note what this definition does not say\. It does not say a language model cannot participate in maintaining governed state — on the contrary, language models are essential at the perception boundary, reading and drafting the documents the ledger ingests\. It does not say symbolic systems automatically qualify: a knowledge graph with silent overwrites and no provenance fails the core tier as surely as a context window does\. Persistence and auditability are systems properties, not gifts of symbolic representation — a neural system can be wrapped in event\-sourced external memory; symbolic structure earns its place for specific reasons \(checkable semantics, human inspectability, ontology binding\), not because symbols are where persistence lives\. And the definition converts this paper’s thesis into an audit instrument: for any proposed clinical AI system, ask which tier its patient representation reaches, and demand evidence property by property\. Table[2](https://arxiv.org/html/2608.14804#S3.T2)previews the audit against architectural classes rather than products; the properties it interrogates are exactly the ones this section defined\.

Table 2:Architectural classes, not products, against the governance properties of Definition[1](https://arxiv.org/html/2608.14804#Thmdefinition1)\(Governed Clinical State\)\. The five columns compress the audit for readability rather than replacing it: Persistent summarizes identity and persistence; Governed summarizes explicit semantics, update governance, provenance, deterministic constraints, and calibration; Replay summarizes historical query and rollback/correction; the last two columns audit the object separations of Section[4](https://arxiv.org/html/2608.14804#S4)\(belief distinct from evidence; quarantined simulation\), which the tier requirements presuppose\. Entries indicate properties guaranteed by the architectural class as such, not properties a particular implementation could add; “partial” marks properties implementations may add without a class\-level guarantee; the final row states what this paper requires, not what any deployed system has demonstrated — the audit questions, not the grades, are the contribution\.
## 4Five objects that must not be conflated

Before the formalism, the picture\. A patient exists; a care process looks at the patient, selectively; what it records becomes evidence; from evidence the system forms a belief about the patient; from belief it predicts, and in quarantined branches it simulates; a human decides\. Every arrow in that chain loses, shapes, or creates information \(Figure[2](https://arxiv.org/html/2608.14804#S4.F2)\)\. The vocabulary of clinical AI habitually collapses this entire chain into the single word “state,” and that collapse is not a terminological quibble: each of the stages is a different mathematical object, with different update rules and different failure modes, and every pathology this paper diagnoses is a conflation of two*neighboring*stages of this chain: treating observations as the truth \(oto\_\{t\}mistaken forxtx\_\{t\}\), treating extracted assertions as established facts \(eie\_\{i\}mistaken foroto\_\{t\}\), treating the accumulated record as what the system believes \(E≤τawareE^\{\\mathrm\{aware\}\}\_\{\\leq\\tau\}mistaken forbτb\_\{\\tau\}— the most common and most damaging\), or letting simulated values leak back into the record \(x~tbranch\\tilde\{x\}^\{\\,\\mathrm\{branch\}\}\_\{t\}contaminating ledger or belief\)\.

![Refer to caption](https://arxiv.org/html/2608.14804v1/figures/fig2-five-objects.png)Figure 2:The five objects and the epistemic boundary\. The true statextx\_\{t\}sits above a boundary no system crosses; observationsoto\_\{t\}, evidence unitseie\_\{i\}, and beliefbτ​\(xt\)b\_\{\\tau\}\(x\_\{t\}\)are successive constructions below it, and simulated statex~tbranch\\tilde\{x\}^\{\\,\\mathrm\{branch\}\}\_\{t\}lives in a quarantined branch whose values never re\-enter evidence or belief\. The pathologies of LLM\-centric practice are conflations of adjacent stages — most commonly, treating evidence as if it were belief\.This paper therefore assigns each stage its own symbol and keeps all five separate:

The true statextx\_\{t\}is the patient’s actual physiological and clinical condition; no system — and no clinician — possesses it\. Observationsoto\_\{t\}are what the care process happened to measure and record — always the product of decisions, never a neutral sample\. Evidence comes in immutable unitsei=\(oi,tievent,tidoc,tiingest,τiaware,pi\)e\_\{i\}=\(o\_\{i\},t^\{\\mathrm\{event\}\}\_\{i\},t^\{\\mathrm\{doc\}\}\_\{i\},t^\{\\mathrm\{ingest\}\}\_\{i\},\\tau^\{\\mathrm\{aware\}\}\_\{i\},p\_\{i\}\): an observation wrapped with provenancepip\_\{i\}and four time stamps — event, documentation, ingestion, and awareness time, the last being when the system’s belief first reflected it\. Corrections create new evidence linked to old, never edits in place\. The evidence available at awareness timeτ\\tauis the setE≤τaware=\{ei:τiaware≤τ\}E^\{\\mathrm\{aware\}\}\_\{\\leq\\tau\}=\\\{\\,e\_\{i\}:\\tau^\{\\mathrm\{aware\}\}\_\{i\}\\leq\\tau\\,\\\}— exactly the conditioning set of equation \([2](https://arxiv.org/html/2608.14804#S4.E2)\) — so the chain runsxt→oi→ei→E≤τaware→bτ​\(xt\)x\_\{t\}\\to o\_\{i\}\\to e\_\{i\}\\to E^\{\\mathrm\{aware\}\}\_\{\\leq\\tau\}\\to b\_\{\\tau\}\(x\_\{t\}\)\. The belief state requires two clocks, and conflating them would undo the very multitemporal distinction this paper claims as a contribution\. Letttdenote*clinical time*— the time at which the patient’s latent state held — andτ\\taudenote*awareness time*— the time at which the system’s knowledge is being evaluated\. The belief state is then the system’s probability distribution over the true state at clinical timett, given everything the system was aware of byτ\\tau:

bτ​\(xt\)=P⁡\(Xt=x∣E≤τaware,A≤τaware\),b\_\{\\tau\}\(x\_\{t\}\)=P\\\!\\left\(X\_\{t\}=x\\mid E^\{\\mathrm\{aware\}\}\_\{\\leq\\tau\},\\ A^\{\\mathrm\{aware\}\}\_\{\\leq\\tau\}\\right\),\(2\)whereE≤τawareE^\{\\mathrm\{aware\}\}\_\{\\leq\\tau\}is all evidence with awareness time at mostτ\\tauandA≤τawareA^\{\\mathrm\{aware\}\}\_\{\\leq\\tau\}the care actions — prescriptions, procedures, orders — known to the system byτ\\tau\. The two indices are what make replay a mathematical statement rather than a slogan:bτ1​\(xt\)b\_\{\\tau\_\{1\}\}\(x\_\{t\}\)versusbτ2​\(xt\)b\_\{\\tau\_\{2\}\}\(x\_\{t\}\)asks what did the system believe at awareness timeτ1\\tau\_\{1\}about the patient’s condition at clinical timett, compared with what it believed later — exactly the comparison a delayed discharge summary forces, where an event of January \(tt\) enters awareness only in February \(τ\\tau\), and exactly the query an auditor asks after a correction\. Where the two clocks coincide we writebtb\_\{t\}as shorthand, but the object is alwaysbτ​\(xt\)b\_\{\\tau\}\(x\_\{t\}\)\. We reserveBτB\_\{\\tau\}for the complete stored belief\-state object — the software artifact, one per patient — of which eachbτ​\(xt\)b\_\{\\tau\}\(x\_\{t\}\)is a component posterior about clinical state at one time; the composition claim of Section[8](https://arxiv.org/html/2608.14804#S8)is stated overBτB\_\{\\tau\}\. Howbτ​\(xt\)b\_\{\\tau\}\(x\_\{t\}\)is parameterized computationally — factored parametric distributions, probabilistic\-logic assertions over ontology\-bound variables, neural latent state with declared decoders, or a hybrid — is deliberately left open at this level: the requirements below constrain any choice, and an operational minimum profile is future work\. The belief state is required to be versioned, calibrated where its components admit calibration, attributed, and honest about ignorance — unmeasured represented as unmeasured, not imputed to normal — and it may become multimodal when evidence genuinely supports incompatible alternatives\. Simulated statex~tbranch\\tilde\{x\}^\{\\,\\mathrm\{branch\}\}\_\{t\}, finally, is belief evolved under hypothetical actions, and its role is prospective: it is the object that answers “what if?” — comparing candidate treatment plans, projecting a trajectory under an alternative medication, stress\-testing a discharge decision — before any action is taken on the real patient\. That usefulness is exactly why its quarantine matters: branches can be created, compared, and discarded freely only because they live where they cannot contaminate the record of what was actually observed, and simulated values never re\-enter the ledger or the belief\. Systems that collapse the first four objects \(xtx\_\{t\},oto\_\{t\},eie\_\{i\},bτb\_\{\\tau\}\) into one — most commonly by defining state as the accumulated record — become simulators of documentation rather than estimators of the patient, however well they predict\.

A worked contrast makes the belief object concrete\. Two active medication lists survive ingestion: the hospital discharge list includes apixaban; the primary\-care list, updated later by a clinic that may not have known about the admission, does not\. A record\-as\-state system answers “is this patient anticoagulated?” with whichever fragment retrieval happens to surface\. The belief state instead answers with an explicitly uncertain — possibly multimodal — representation, preserving the evidence chains that support each alternative; how probability mass is assigned across alternatives is a calibration question this paper deliberately leaves to future work\. The uncertainty is the answer; a system that hides it has not resolved the conflict, only concealed it\.

### 4\.1Epistemic types, and the taxonomy of absence

Governed state requires that every assertion carry its epistemic type — the answer to “how do we know this?” — as a machine\-checkable attribute rather than a nuance of phrasing\. Two orthogonal attributes must not be merged\. The*epistemic origin*answers how the assertion was produced: observed, reported, inferred, predicted, simulated, guideline\-derived, or unknown\. The*assertion status*records its standing relative to other evidence: supported, contradicted, superseded, unresolved\-conflict, or retracted — the attribute through which the managed consistency of the next subsection operates\. Two rules make the taxonomy load\-bearing\. First, type is inherited and never laundered: an inference built on a reported value is at most as strong as the report, and a prediction can never be silently re\-typed as an observation\. Second, absence is typed with the same care as presence:

unknown≠negative≠not measured≠not documented≠unavailable\.\\text\{unknown\}\\neq\\text\{negative\}\\neq\\text\{not measured\}\\neq\\text\{not documented\}\\neq\\text\{unavailable\.\}The canonical illustration is the allergy question\. “No allergies documented” \(nothing was asked\), “patient reports no allergies” \(a reported negative\), “allergy testing negative” \(an observed negative\), and “allergy list not accessible from this source” \(an availability fact — which is why unavailable is a first\-class member of the taxonomy: absence at this system is not absence in the world\) authorize four different clinical actions — and a system that renders all four as the reassuring phrase “no known allergies” has erased a patient\-safety distinction\. In text, these distinctions survive only if a writer chose careful words and a reader parses them; in governed state, they are types, and a verifier can refuse an action whose preconditions demand an observed negative where only an undocumented absence exists\.

### 4\.2Managed consistency, not guaranteed consistency

It is tempting to demand that a clinical representation be guaranteed free of contradictions\. The demand is misguided, because clinical evidence is legitimately contradictory: clinicians disagree; assays disagree; a diagnosis is provisional and then revised; documentation lags reality\. A system that guarantees a contradiction\-free representation can do so only by forcing conflicting evidence into a single truth — precisely the silent merging that destroys auditability\. What we require instead is*managed consistency*: contradictions are detected mechanically; contradictory evidence is preserved, never overwritten; incompatible assertions are never silently merged; belief estimation consumes the disagreement as evidence structure — widening or splittingbτ​\(xt\)b\_\{\\tau\}\(x\_\{t\}\)rather than pretending to a resolution the evidence does not license; and conflicts remain visibly open until evidence, or a human, closes them\. A diagnosis copied forward for two years after the infiltrate resolved is not deleted but superseded — and an auditor can still see exactly how long the stale assertion survived and which report ended it\.

## 5One question, asked of both architectures

“Can this patient safely receive iodinated contrast today?” looks like a retrieval problem\. It is, in miniature, everything this paper formalizes: it requires current renal function \(not the most recently documented creatinine\); the status of metformin therapy, where relevant to medication management around contrast administration \(an order is care\-process state, not a sentence in a note\); prior contrast reactions \(where the difference between “no reaction documented” and “no reaction” can be the entire answer\); and knowledge of what the record is missing \(prior imaging happened at another institution\)\. Table[3](https://arxiv.org/html/2608.14804#S5.T3)shows how the two regimes of Figure[1](https://arxiv.org/html/2608.14804#S3.F1)fare\.

Table 3:One clinical question, component by component\. The failure on the left is not that the language model reads badly — it reads perfectly\. The failure is that four of the five objects of Section[4](https://arxiv.org/html/2608.14804#S4)do not exist anywhere in that architecture, so no amount of reading can consult them\.Notice the output the governed architecture can represent: not “yes” or “no,” but a set of typed gaps and the actions that would close them — a language model may of course recommend such a measurement; the architectural distinction is that a governed observation model can represent the missing measurement itself as a typed information gap and justify the measurement as an uncertainty\-reducing action \(the observation\-process model, developed in future work\)\.

### 5\.1The contrast in full: a longitudinal case

This subsection develops what the section has just presented in more detailed and practical form: the same contrast, now with a concrete timeline — dates, values, documents, and gaps — rather than isolated components\. The component\-level table above is deliberately synchronic; the failure modes that motivate this paper are diachronic — they accumulate over years — so we close the argument with a case\-based illustration: one realistic six\-year record, read by both architectures at the same query point \(Figure[3](https://arxiv.org/html/2608.14804#S5.F3)\)\.

The record: a creatinine of 0\.9 mg/dL in 2019 \(normal\); chronic kidney disease \(CKD\) stage 2 added to the problem list in 2020; an acute kidney injury \(AKI\) with creatinine peaking at 2\.4 during a 2021 admission; recovery to 1\.1 in January 2022 — documented in a discharge summary written thirty days after the events it describes; laboratory work drawn at an outside institution in 2023 that never reaches this system; and then nothing — no measurement for fourteen months — until the 2024 query: can this patient safely receive iodinated contrast today?

Each event exercises a different object from Section[4](https://arxiv.org/html/2608.14804#S4)\. The 2022 discharge summary makes documentation time diverge from event time: a system that conditions on it as of January has silently read the future\. The 2023 outside labs are an availability failure: the patient’s true state was measured, and the measurement exists — in a silo this system cannot see — so the absence in the record is not absence in the world\. And the fourteen\-month gap is itself potentially informative about the observation process — but it does not identify its cause: follow\-up may not have occurred, may have occurred outside the observable system, or may have generated evidence that never reached this one\. What governed state owes the clinician is exactly that typed ambiguity, not a story; converting a missing observation into one specific inferred explanation would repeat, in reverse, the overconfidence this paper diagnoses\.

![Refer to caption](https://arxiv.org/html/2608.14804v1/figures/fig3-one-patient-six-years.png)Figure 3:One patient, six years, one query\. Top: the timeline, with a documentation lag \(2022\), an availability exit \(2023\), and an informative measurement gap \(2023–24\)\. Bottom: the same record read at the same moment by the two architectures of Figure[1](https://arxiv.org/html/2608.14804#S3.F1)\. The governed answer is not produced by a better reader but by objects the generated\-context architecture does not possess\. The case is an illustrative construction, not an executed comparison: the claim is about mechanisms each architecture can and cannot express, not measured performance\.A record\-as\-state, generated\-context architecture can produce a fluent but unjustifiably confident answer\. Retrieval surfaces the most recent creatinine — 1\.1, from the 2022 summary — and treats it as the patient’s current renal status; the value’s staleness is invisible in prose, the 2023 outside measurements are simply unavailable to this system \(retrieval cannot return documents that never arrived\), and the measurement gap is not represented anywhere, because a record cannot list what it lacks\. The answer is “last creatinine 1\.1 — proceed,” delivered with the confidence of a system that has read everything it was given\. The governed\-state architecture instead represents the unresolved uncertainty explicitly — and not because its reader is smarter: the same language model may sit at its perception boundary\. The belief over renal function has widened with every unmeasured month since the last awareness\-timed observation; the 2021 AKI persists in belief, as a risk modifier, rather than only in a note; the 2023 gap is typed as an availability failure rather than rendered as reassurance; and the absence of an expected measurement has itself been digested as evidence\. The answer is “current renal function unknown; last known 1\.1 in January 2022; a point\-of\-care creatinine would resolve it” — a set of typed gaps and the action that closes them\. Every difference between the two columns of Figure[3](https://arxiv.org/html/2608.14804#S5.F3)traces to exactly one cause: the objects of Section[4](https://arxiv.org/html/2608.14804#S4)exist in one architecture and not the other\.

## 6An accountability decomposition

### 6\.1The demand

We call a system’s longitudinal clinical reasoning*accountable*if, for any patient\-specific temporal conclusion it produces, a third party can: \(R1\) reconstruct the evidence the system was aware of when it produced the conclusion; \(R2\) distinguish what the system observed from what it inferred, predicted, or simulated — including distinguishing observed absence from unmeasured absence; and \(R3\) determine what kind of claim, associational or causal, the conclusion makes\. Nothing in this definition mentions architecture; it is a demand a hospital, a court, or a regulator can state without knowing how the system works\.

### 6\.2The decomposition

> ###### Operational Decomposition of Accountability\. Given requirements R1–R3, the required information content decomposes into four functional artifacts: any system whose longitudinal clinical reasoning is accountable in the sense above maintains, explicitly or in an informationally equivalent form, four artifacts: \(A1\) an immutable evidence ledger with awareness\-time versioning; \(A2\) a belief state distinct from accumulated evidence, with epistemic typing of its contents; \(A3\) an observation model — an account of how observations came to exist; and \(A4\) claim\-level causal typing of every predictive output\. We say*artifacts*to keep them distinct from the five epistemic objects of Section[4](https://arxiv.org/html/2608.14804#S4), to which they map directly: A1 materializes the evidence unitseie\_\{i\}and their ledger, A2 materializes the belief objectbτ​\(xt\)b\_\{\\tau\}\(x\_\{t\}\), A3 models the passage fromxtx\_\{t\}tooto\_\{t\}, and A4 types the predictions derived frombτb\_\{\\tau\}\. The decomposition asserts functional necessity only: an accountable system must maintain the information content and update discipline of A1–A4, but need not expose them as four physically separate services or data structures — belief and evidence may share storage, an observation model may be folded into a generative data model, and causal qualification may be enforced at query time, provided the functional properties hold and can be audited\.

The argument is elimination, one requirement at a time\. R1 forces A1: to reconstruct “what the system was aware of at awareness timeτ\\tau,” the information distinguishing awareness time from event time must have been preserved at ingestion — it cannot be recovered afterward from content alone, since a discharge summary describing day 2 reads identically whether it arrived on day 3 or day 30\. R2 forces A2: if no functional distinction between evidence and belief is maintained, observed and inferred are distinguished nowhere, and the distinction cannot be reconstructed after the fact\. R2’s second clause forces A3: the difference between “tested negative” and “never tested” cannot in general be recovered reliably from the record’s content alone; it lives in the process that generated the record\. R3 forces A4: a conclusion’s causal status is a property of how it was derived, and unless attached at derivation time, an associational estimate and an interventional one are indistinguishable downstream\.

### 6\.3What kind of result this is — and is not

We should be candid about the logical character of the argument\. Because we define accountability as R1–R3 and those requirements essentially name the information the four objects carry, the “forcing” is analytic — a clarification of what our definition already entails — rather than a synthetic discovery about the world\. That is a genuine and, we think, useful contribution: it pins down what accountability must mean operationally and shows the definition is coherent and decomposable\. But it is not a theorem in the strong sense, and it should not be read as one; its force is conceptual hygiene\. A reviewer who exhibits an accountable system that factors the objects differently confirms the decomposition rather than refuting it, so long as the functional properties survive\. Its empirical counterpart — whether maintaining the four artifacts actually improves clinical reasoning — is a separate question, staked as the research questions summarized in the conclusion\.

Two consequences follow\.Context\-window insufficiency:an architecture in which patient information exists at reasoning time only as context assembled for the call — whatever the scale or quality of the model that reads it — maintains none of the four artifacts, and its longitudinal reasoning cannot be made accountable solely by prompting, retrieval, fine\-tuning, or increased context length: accountability requires persistent governed information outside the transient context itself — changing the architecture, not the model\.The evaluation shift:if the decomposition holds, evaluation of clinical AI intended for longitudinal use must weigh the governance properties of the patient state it maintains at least as heavily as the quality of the text it generates\. Output\-only benchmarks cannot distinguish a system that maintains the four artifacts from one that merely sounds as if it does\. The audit questions of Definition[1](https://arxiv.org/html/2608.14804#Thmdefinition1)\(Governed Clinical State\) can\.

## 7A maturity framework for clinical representations

Clinical knowledge is, overwhelmingly, encoded in text, and good clinical prose carries temporality, uncertainty, negation, causal language, hypotheses, and disagreement\. The defensible claim is operational: the structure text carries is implicit, inconsistently expressed, entangled with authorship, and therefore not queryable, not validatable, not updatable, and not governable at the standard of Definition[1](https://arxiv.org/html/2608.14804#Thmdefinition1)\(Governed Clinical State\)\. Take one sentence of ordinary clinical prose: “Given her worsening renal function, we held the lisinopril and will recheck creatinine on Monday; if stable, resume at half dose\.” In one line: an observation trend, a causal attribution, an executed action with a reason, a scheduled observation, a conditional plan, and an implicit baseline\. A language model, prompted well, extracts most of it, most of the time\. But the sentence’s structure becomes usable infrastructure only when each element lands in governed state with its epistemic type — and left as prose, every distinction survives only until the next summarization, the next copy\-forward, the next context\-window eviction\. The failure mode of text is not that it cannot express structure; it is that nothing enforces the structure it expresses\.

Before walking any levels, the distinction that makes a maturity framework defensible at all: two properties of a representation must be kept apart —*representation maturity*, what the system makes explicit and governable, and*computational capability*, what it can infer internally\. A multimodal time\-series model can estimate a latent physiological state without ever building an entity graph; a mechanistic digital twin can host world\-model machinery over tabular variables with no NLP anywhere\. The framework’s ordering claim is therefore about governance, not capability, and it does not assert that systems must be built level by level: a system can compute at levelkkwhile only being accountable at levelj<kj<k, and clinical deployment is constrained by the accountable level, not the computed one\.

With that guard in place, the framework of Figure[4](https://arxiv.org/html/2608.14804#S7.F4)orders representations by how much clinical structure is explicit, persistent, and governable: Level 0, raw text; Level 1, canonical entities; Level 2, typed relations — the classical clinical knowledge graph\[[5](https://arxiv.org/html/2608.14804#bib.bib5)\]; Level 3,*Governed Temporal Evidence*— time\-qualified, provenance\-tagged assertions, the temporal graph of what has been*asserted*; Level 4, a governed belief state over the patient’s condition \(thebτ​\(xt\)b\_\{\\tau\}\(x\_\{t\}\)of equation \([2](https://arxiv.org/html/2608.14804#S4.E2)\)\); Level 5, a Clinical World Model — belief plus governed dynamics, an observation model, and causal qualification \(future work\)\. In compact form: L0 Text→\\toL1 Entities→\\toL2 Relations→\\toL3 Evidence→\\toL4 Belief→\\toL5 World Model\. The count of levels is a convention, not a claim — a finer or coarser partition would serve — and what the framework actually asserts is one ordering of governance commitments with a single qualitative jump: Level 3 to Level 4, from what has been*asserted*to what the system*believes*\. Levels 0–2 are a standard natural\-language\-processing \(NLP\) and knowledge\-graph pipeline; the jump to belief is the step that record\-as\-state systems, however sophisticated their predictors, do not expose\. LLM\-centric systems have high capability at low maturity; that combination is exactly what Section[3](https://arxiv.org/html/2608.14804#S3)diagnosed, and “raising maturity to meet capability” is a fair one\-sentence summary of this paper’s program\.

![Refer to caption](https://arxiv.org/html/2608.14804v1/figures/fig4-maturity-levels.png)Figure 4:The six\-level framework: increasing degrees of explicit, persistent, governable clinical structure\. An architectural maturity framework, not a claimed law of intelligence\.
## 8Positioning, novelty, and what this paper does not claim

Several communities have built parts of what this paper assembles\. Digital twins in healthcare\[[9](https://arxiv.org/html/2608.14804#bib.bib9)\]pursue persistent computational models of individual patients, most maturely in mechanistic, organ\- or disease\-specific form — and we should not undersell them: a good mechanistic twin already delivers individualized dynamics at higher fidelity than any general belief state for the specific decision it was built for; the honest claim is that our architecture is governance\-general where the twin is decision\-specific, with the twin a natural high\-fidelity dynamics component to host inside the belief layer\. Electronic\-health\-record \(EHR\) representation learning \(BEHRT and successors\[[10](https://arxiv.org/html/2608.14804#bib.bib10)\]\) learns predictive latent representations over longitudinal coded records\. The world\-model literature outside medicine — latent dynamics models\[[11](https://arxiv.org/html/2608.14804#bib.bib11)\], joint\-embedding predictive architectures\[[12](https://arxiv.org/html/2608.14804#bib.bib12)\], continuous\-time models\[[13](https://arxiv.org/html/2608.14804#bib.bib13)\], and the belief\-state formalism of partially observable Markov decision processes \(POMDPs\)\[[14](https://arxiv.org/html/2608.14804#bib.bib14)\]— supplies the mathematical frame\. Within medicine, EHRWorld\[[22](https://arxiv.org/html/2608.14804#bib.bib22)\]demonstrates that explicit temporally evolving state improves rollout stability over frontier LLMs, and SMB\-Structure \(Standard Model Biomedicine\)\[[23](https://arxiv.org/html/2608.14804#bib.bib23)\]— whose title, “The Patient is not a Moving Document,” shares our exact rhetorical move — shows on large oncology and pulmonary\-embolism cohorts \(23,319 and 19,402 patients respectively\) that joint\-embedding predictive architecture \(JEPA\)\-style embeddings capture disease dynamics autoregressive baselines miss\. A 2026 roadmap organizes biomedical world models around data engines, simulators, and planning substrates\[[24](https://arxiv.org/html/2608.14804#bib.bib24)\]; and the term “Clinical World Model” itself is used by Safavi\-Naini et al\. for a different purpose — a competency\-evaluation framework\[[25](https://arxiv.org/html/2608.14804#bib.bib25)\]— so we make no priority claim over the term\. Provenance and auditability are likewise not novel in isolation: auditable, source\-verified clinical\-AI frameworks combining RAG with provenance and tamper\-evident logging exist\[[26](https://arxiv.org/html/2608.14804#bib.bib26)\], and neurosymbolic clinical decision support with ontology grounding is an actively worked area\[[27](https://arxiv.org/html/2608.14804#bib.bib27),[28](https://arxiv.org/html/2608.14804#bib.bib28)\]\.

Three works from mid\-2026 sharpen the neighborhood further\. MedBeads\[[15](https://arxiv.org/html/2608.14804#bib.bib15)\]independently develops an immutable, provenance\-bearing clinical data substrate — records as tamper\-evident, causally linked objects with amendments and retractions, supplied to agents by deterministic traversal rather than semantic retrieval\. It is closest to this paper’s Level\-3 evidence layer, and it reinforces the claim that retrieval alone is not a governance mechanism; our framework begins where such context governance ends — the governed evidence remains distinct from the canonical belief state, the observation process, causal typing, and quarantined simulation\. Chen et al\.’s review of medical world models\[[16](https://arxiv.org/html/2608.14804#bib.bib16)\]screened 1,455 records to find fourteen empirical systems, organizing the field around state representation, dynamics, intervention\-conditioned simulation, and planning while naming causal identifiability, calibration, uncertainty, and external validation as the field’s open limitations — a capability\-side map whose named gaps are, almost item for item, the governance\-side requirements this paper specifies\. And CalTwin\[[17](https://arxiv.org/html/2608.14804#bib.bib17)\]directly targets calibration and distribution\-shift robustness in latent medical\-world\-model trajectories — an example of the machinery that could one day satisfy part of Definition[1](https://arxiv.org/html/2608.14804#Thmdefinition1)’s predictive tier, which is precisely how the tiered standard is meant to be used\.

Closest of all at the claim\-typing layer is Qazi et al\.’s “Beyond Generative AI”\[[18](https://arxiv.org/html/2608.14804#bib.bib18)\], which grades clinical world models on a four\-level capability hierarchy — temporal prediction, action\-conditioned prediction, counterfactual rollouts, and planning — a rubric adjacent to the causal typing this paper requires as artifact A4, and one any reader of this paper should study alongside it\. The distinction to keep sharp: theirs is an evaluative hierarchy of what a model*can do*; ours is enforced output metadata — every prediction carries its causal level, and promotion between levels passes a per\-query identification gate, not a capability judgment\. The “patient as evolving state, not a document” thesis itself is likewise now shared ground across the 2026 systems above, and we claim no priority over it\.

Our novelty therefore cannot rest on “neurosymbolic plus provenance plus ontology,” which is occupied, nor on the patient\-as\-state thesis, which is shared\. It rests on four commitments that remain under\-occupied across all of the works surveyed: \(a\) awareness time as a first\-class fourth temporal axis, distinct from ingestion; \(b\) the replay invariant — belief as a deterministic replay of the ledger prefix, making historical belief a computation rather than an archaeology; \(c\) quarantined simulation as a formal isolation construct rather than a usage convention; and \(d\) causal\-level typing as enforced output metadata with a per\-query identification gate\. These four, combined with the belief\-versus\-evidence separation and the observation\-process model into one accountable operator, are the claim — and the claim is a composition claim, not a component claim\. We do not assert novelty for the individual ingredients; we propose that accountable longitudinal clinical reasoning requires their governed composition around a replayable distinction between evidence and belief\. WithBτB\_\{\\tau\}the canonical stored belief\-state object introduced in Section[4](https://arxiv.org/html/2608.14804#S4), the composition is

E≤τaware→𝒰Bτand\(Bτ,a\)→simulationX~branch,E^\{\\mathrm\{aware\}\}\_\{\\leq\\tau\}\\;\\xrightarrow\{\\ \\mathscr\{U\}\\ \}\\;B\_\{\\tau\}\\qquad\\text\{and\}\\qquad\(B\_\{\\tau\},a\)\\;\\xrightarrow\{\\ \\text\{simulation\}\\ \}\\;\\widetilde\{X\}^\{\\,\\mathrm\{branch\}\},with a governed, replayable update operator𝒰\\mathscr\{U\}and no reverse write path from simulation to evidence or belief\. The architectural invariant, not any component, is the contribution\. One qualification keeps the replay commitment honest: replay is deterministic with respect to the ledger prefix together with the pinned update\-operator version, model versions, ontology and rule versions, configuration, and stored stochastic state where inference is sampled — given the same, replay reproduces the same governed belief state\. This is precisely why the predictive tier of Definition[1](https://arxiv.org/html/2608.14804#Thmdefinition1)requires model\-version binding; without it, the same ledger could yield a different posterior after a model update, and the invariant would silently fail\. That is a synthesis claim, not a claim to have invented a category; and it is an architectural claim whose empirical standing depends entirely on the research program sketched in the conclusion\.

## 9Conclusion

This paper set out to relocate the debate about clinical AI: from what language models can generate to what clinical systems must maintain\. Its argument, compressed: clinical reasoning over a patient record is state estimation under partial observability; the interface of an autoregressive language model, whatever its internal representations, exposes no governed state to estimate with; and the remedy is architectural\. The accountability demand decomposes into four information requirements — an evidence ledger with awareness\-time versioning, a belief state distinct from the record, an observation model, and causal typing of predictive claims — and the maturity framework gives the field a way to say precisely what a system makes governable, as opposed to what it can compute\. Fluency is now abundant; governed clinical state is not — and it is the difference between a system that talks about patients and a system that can be held to account for what it believes about them\.

#### The research questions, answered as far as this paper can\.

The introduction posed four questions; here is the state of each\.

- •RQ1 \(temporal consistency\)\.Answer at this stage: not yet empirically answered — predicted by the architecture\. The architectural analysis predicts improved temporal consistency and provenance, because governed state exposes information that retrieval\-only architectures do not guarantee; the contrast of Section[5](https://arxiv.org/html/2608.14804#S5)and the constructed longitudinal case of Section[5\.1](https://arxiv.org/html/2608.14804#S5.SS1)demonstrate the mechanism, not the margin\. The quantitative comparison against strong long\-context and retrieval baselines remains to be measured, and is the first measurement to be taken up in future work\.
- •RQ2 \(multi\-horizon dynamics\)\.Answer at this stage: not yet answerable — and this paper says precisely why\. The maturity framework locates what the question tests \(the Level\-3 to Level\-4 jump\), and answering it requires an operational belief state first; that operationalization is the critical path of the program\.
- •RQ3 \(verification\)\.Answer at this stage: the question is now well\-posed\. Definition[1](https://arxiv.org/html/2608.14804#Thmdefinition1)\(Governed Clinical State\) supplies the constraint classes a verifier enforces; the experiment is cheap to run and informative either way, including against us\.
- •RQ4 \(transportability\)\.Answer at this stage: predicted, not demonstrated\. The observation\-model artifact \(A3\) of the accountability decomposition predicts the transport benefit; before any measurement can be trusted, future work must state the estimation procedure and its identifiability assumptions\.

Each question will be taken up in future work with quantitative thresholds preregistered before results are reported — and each can go against us: a reader who believes that scale, context length, or retrieval make governed state unnecessary will find in these four questions the concrete places to show it\.

#### Future work\.

Two fronts, in order\.*Empirical*: answer RQ1–RQ4, beginning with RQ1 and RQ3 \(which need only the evidence layer already in production plus a minimal belief layer\), then RQ2 and RQ4 \(which require the operational belief state and observation model first\)\.*Specification*: publish the full engineering treatment of the evidence ledger and governed update operator; an operational schema for the belief state; the observation model’s estimation procedure with its identifiability assumptions stated \(fitting a measurement policy conditioned on latent state requires the very state being estimated\); and the per\-query causal identification contracts that govern promotion between prediction levels\.

#### About MyndwareMed\.

This paper is the conceptual foundation of a system being built\. MyndwareMed’s platform implements the governed evidence layer specified here — ontology\-bound, provenance\-tagged, multitemporal — in production today \(maturity Level 3\), is engineering the governed belief layer \(Level 4\), and is researching the full Clinical World Model \(Level 5\) against the research questions above\. The architecture in this paper is, deliberately, the standard that program is accountable to\.

#### Competing interests and scope statement\.

Both authors are employees of MyndwareMed, which develops a commercial platform implementing the architecture this paper specifies; the reader should weigh the positioning claims of Section[8](https://arxiv.org/html/2608.14804#S8)accordingly\. This paper is a position\-and\-specification paper: it uses no patient data, reports no experiments, and claims no empirical or clinical result\. Statements about MyndwareMed’s implementation are claims about engineering artifacts, not about demonstrated clinical performance, and the commitments outlined in the conclusion will be preregistered with quantitative thresholds before any results are reported\.

## References

- \[1\]Singhal, K\. et al\. Large language models encode clinical knowledge\.*Nature*620, 172–180 \(2023\)\.
- \[2\]Nori, H\. et al\. Capabilities of GPT\-4 on medical challenge problems\. arXiv:2303\.13375 \(2023\)\.
- \[3\]Thirunavukarasu, A\. J\. et al\. Large language models in medicine\.*Nature Medicine*29, 1930–1940 \(2023\)\.
- \[4\]Lewis, P\. et al\. Retrieval\-augmented generation for knowledge\-intensive NLP tasks\.*NeurIPS*\(2020\)\.
- \[5\]Hogan, A\. et al\. Knowledge graphs\.*ACM Computing Surveys*54\(4\) \(2021\)\.
- \[6\]SNOMED International\. SNOMED CT\.[https://www\.snomed\.org/](https://www.snomed.org/)
- \[7\]Regenstrief Institute\. LOINC\.[https://loinc\.org/](https://loinc.org/)
- \[8\]W3C\. OWL 2 Web Ontology Language\.[https://www\.w3\.org/TR/owl2\-overview/](https://www.w3.org/TR/owl2-overview/)
- \[9\]Katsoulakis, E\. et al\. Digital twins for health: a scoping review\.*npj Digital Medicine*7\(2024\)\.
- \[10\]Li, Y\. et al\. BEHRT: transformer for electronic health records\.*Scientific Reports*10\(2020\)\.
- \[11\]Hafner, D\. et al\. Mastering diverse domains through world models \(DreamerV3\)\. arXiv:2301\.04104 \(2023\)\.
- \[12\]Assran, M\. et al\. Self\-supervised learning from images with a joint\-embedding predictive architecture \(I\-JEPA\)\.*CVPR*\(2023\)\.
- \[13\]Chen, R\. T\. Q\. et al\. Neural ordinary differential equations\.*NeurIPS*\(2018\)\.
- \[14\]Kaelbling, L\. P\., Littman, M\. L\. & Cassandra, A\. R\. Planning and acting in partially observable stochastic domains\.*Artificial Intelligence*101, 99–134 \(1998\)\.
- \[15\]Nakajima, T\. et al\. MedBeads: an agent\-native, immutable data substrate for trustworthy medical AI\. arXiv:2602\.01086 \(2026\)\.
- \[16\]Chen, Z\. et al\. Medical world models in healthcare: foundations, applications, and challenges for trustworthy clinical translation\. arXiv:2607\.25242 \(2026\)\.
- \[17\]Khan, B\. et al\. CalTwin: towards calibrated, shift\-robust medical world models via Fisher\-information regularisation\. arXiv:2607\.26752 \(2026\)\.
- \[18\]Qazi, M\. A\., Nadeem, M\. & Yaqub, M\. Beyond generative AI: world models for clinical prediction, counterfactuals, and planning\. arXiv:2511\.16333 \(2025\)\.
- \[19\]HL7 International\. Fast Healthcare Interoperability Resources \(FHIR\)\.[https://www\.hl7\.org/fhir/](https://www.hl7.org/fhir/)
- \[20\]Wang, L\. et al\. A survey on large language model based autonomous agents\.*Frontiers of Computer Science*18\(2024\)\.
- \[21\]Cao, Z\., Zhan, B\., Shi, J\. et al\. Agents\-K1: towards agent\-native knowledge orchestration\. arXiv:2606\.13669 \(2026\)\.
- \[22\]Mu, L\. et al\. EHRWorld: a patient\-centric medical world model for long\-horizon clinical trajectories\. arXiv:2602\.03569 \(2026\)\.
- \[23\]Adam, I\. et al\. The patient is not a moving document: a world model training paradigm for longitudinal EHR\. arXiv:2601\.22128 \(2026\)\.
- \[24\]Wang, G\. et al\. Towards world models in biomedical research\. arXiv:2606\.05925 \(2026\)\.
- \[25\]Safavi\-Naini, S\. A\. A\. et al\. Grounding clinical AI competency in human cognition through the clinical world model and skill\-mix framework\. arXiv:2604\.08226 \(2026\)\.
- \[26\]Alu, F\. F\. & Oluwadare, S\. An auditable and source\-verified framework for clinical AI decision support\.*Frontiers in Artificial Intelligence*9\(2026\)\.
- \[27\]Garcez, A\. S\. d’Avila, Lamb, L\. C\. & Gabbay, D\. M\.*Neural\-Symbolic Cognitive Reasoning*\. Springer \(2009\)\.
- \[28\]Besold, T\. R\. et al\. Neural\-symbolic learning and reasoning: a survey and interpretation\. arXiv:1711\.03902 \(2017\)\.

Similar Articles