Ontology-Amplified Distillation and Contextuality Auditing for Sovereign Enterprise Language Models: A Combined Proof-of-Mechanism and Negative-Results Method Study
Summary
This paper presents a combined proof-of-mechanism study of ontology-amplified distillation for sovereign enterprise language models and a contextuality-audit method, using a Qwen3.6-27B student adapted via supervised fine-tuning and DPO. The results are underpowered and negative, showing no superiority over frontier baselines and zero contextuality in routing.
View Cached Full Text
Cached at: 07/15/26, 04:19 AM
# Ontology-Amplified Distillation and Contextuality Auditing for Sovereign Enterprise Language Models: A Combined Proof-of-Mechanism and Negative-Results Method Study
Source: [https://arxiv.org/html/2607.11948](https://arxiv.org/html/2607.11948)
Thanh Luong Tuan Foundation AgenticOS \(FAOS\)
\(July 2026\)
###### Abstract
Regulated financial institutions operating under data\-residency rules need tenant\-owned language models that can run inside the institution’s perimeter\. This paper combines two related FAOS studies into one mechanism\-and\-control article\. First, it reports a reduced\-power proof\-of\-mechanism study of*ontology\-amplified distillation*: a Qwen3\.6\-27B student is adapted to the Foundation AgenticOS ontology through supervised fine\-tuning on frontier\-teacher trajectories and ontology\-grounded direct preference optimization \(DPO\), trained locally on a single Apple M5 Max from 47 synthetic, English\-language, cross\-domain preference pairs\. On 40 held\-out Vietnamese financial\-domain tasks, the distilled student grounds 36 of 40 tasks \(grounded rate 0\.90; mean ontology term\-coverageronto=0\.95r\_\{\\mathrm\{onto\}\}=0\.95on a metric floored at 0\.50\), equal to the GPT\-5 frontier baseline, which also grounds 36 of 40\. The outcome is underpowered to establish equivalence: the paired\-difference 95% confidence interval spans±4\\pm 4tasks, and the run does not test or show the pre\-registered amplification prediction that the student should exceed the frontier\. Second, the paper consolidates a contextuality\-audit method for enterprise\-agent routing\. In a separate negative\-results pilot, the corrected canonical Contextuality\-by\-Default degree is zero for all Phase 1\.3 groups in both the local\-Qwen run and an explicitly labeled Gemma replication check; the useful signal is direct influence and construct coupling, not surviving residual contextuality\. Together, the studies pair an ontology\-grounded model\-building mechanism with a governance diagnostic for deciding when apparent disagreement should trigger prompt standardization, multi\-agent synthesis, or human review\. The evidence supports neither deployability, safety, superiority, statistical equivalence, nor a contextuality\-positive routing rule\.
## 1Introduction
Financial institutions in Vietnam operate under data\-residency and sector\-specific rules that restrict where customer data may be processed\[National Assembly of Vietnam,[2025](https://arxiv.org/html/2607.11948#bib.bib15), Government of Vietnam,[2025](https://arxiv.org/html/2607.11948#bib.bib16)\]\. For a regulated tenant, sending account records or underwriting files to a foreign frontier\-model API is not a procurement decision but a compliance violation\. A model that runs inside the institution’s perimeter therefore becomes a deployment requirement, and the frontier model functions as a quality ceiling rather than a deployment alternative\. This framing reverses the usual cost–quality trade\-off\. The question is not whether a smaller model is cheaper to run, but whether a locally deployable model can reach the grounding quality that regulated work demands\.
Earlier work established that supplying a domain ontology to a frontier model improves its domain fidelity\[Luong and Sanyal,[2026](https://arxiv.org/html/2607.11948#bib.bib19), Sharmaet al\.,[2025](https://arxiv.org/html/2607.11948#bib.bib6), Liuet al\.,[2025](https://arxiv.org/html/2607.11948#bib.bib7)\]\. That work measured the effect on a frontier model\. It left open the question that matters for sovereign deployment: whether the same ontology grounding can lift a smaller, locally deployable student to the frontier’s level on regulated\-domain tasks\. A distilled student carries thinner parametric coverage of a specialized domain than a frontier model does, so ontology grounding may matter more for the student—a larger relative gain from a lower base—without necessarily closing the absolute gap on the hardest tasks\. The companion neurosymbolic study named this pattern the Inverse Parametric Knowledge Effect \(Inverse PKE\): grounding value rises as the model’s parametric coverage of a domain falls\[Luong and Sanyal,[2026](https://arxiv.org/html/2607.11948#bib.bib19)\]\.
This paper studies that question through*ontology\-amplified distillation*, a two\-stage adaptation\. A student model is first supervised\-fine\-tuned on frontier\-teacher trajectories that carry an injected ontology slice, then tuned with ontology\-grounded direct preference optimization\[Rafailovet al\.,[2023](https://arxiv.org/html/2607.11948#bib.bib22), Hintonet al\.,[2015](https://arxiv.org/html/2607.11948#bib.bib23)\]\. The preference signal pairs an ontology\-grounded frontier answer, marked preferred, against an ungrounded base\-model answer, marked dispreferred, so the student learns to reproduce ontology\-grounded behavior rather than to imitate the frontier’s full parametric knowledge\. The student is a Qwen3\.6\-27B model\[Yanget al\.,[2025](https://arxiv.org/html/2607.11948#bib.bib2)\]; training and evaluation run locally in Apple’s MLX framework on a single M5 Max\[Apple Machine Learning Research,[2025](https://arxiv.org/html/2607.11948#bib.bib27)\], and the trained student is fused to a four\-bit checkpoint of roughly 14 GB for deployment\-consistent inference\.
On a held\-out set of 40 Vietnamese financial\-domain tasks, the distilled student grounds 36 of 40 tasks—a grounded rate of 0\.90, or mean term\-coverageronto=0\.95r\_\{\\mathrm\{onto\}\}=0\.95on a metric floored at 0\.50—equal to the GPT\-5 frontier baseline, which also grounds 36 of 40\. Then=40n=40binary outcome cannot establish equivalence: the paired\-difference 95% confidence interval spans±4\\pm 4tasks\. Training\-side measurements are consistent with the intended preference shift: preference accuracy moves from 0 to 1\.0 and the reward margin from 0 to 0\.307 over a single epoch, on a five\-pair validation split\. The result is scoped as a reduced\-power proof of mechanism, not a deployment claim, and not the pre\-registered amplification—which predicted the student*exceeding*, not equalling, the frontier\. One binding gate was run—ontology compliance—on a metric that records whether the answer surfaces at least one of the task’s ontology terms, a degenerate proxy for the registered four\-component composite, not whether the answer is complete or correct\. The preference data are synthetic, English\-only, and drawn from non\-target domains; the frontier ran at minimal reasoning effort; the student matches rather than exceeds the frontier; and the questions of cost, abstention safety, model scale, and the Vietnamese\-versus\-English grounding contrast are held for a full\-power evaluation\.
The study makes four contributions\. First, it specifies ontology\-amplified distillation as a training recipe for sovereign enterprise models, with a preference construction that separates ontology\-grounded behavior from general frontier imitation\. Second, it defines an ontology\-compliance metric scored locally and identically for the student and the frontier over a shared ontology slice, which removes judge\-model variance from the comparison\. Third, it reports the proof honestly, with an explicit map of which claims the present evidence supports and which await full\-power confirmation\. Fourth, it integrates a contextuality\-audit method for enterprise\-agent routing: when the same regulated task changes under role frame, prompt order, or construct wording, the audit distinguishes direct influence and construct coupling from residual contextuality before a platform routes the task to solo response, prompt standardization, debate, synthesis, or human review\.
[Section˜2](https://arxiv.org/html/2607.11948#S2)positions the work against distillation, preference optimization, and ontology\-grounded generation\.[Section˜3](https://arxiv.org/html/2607.11948#S3)describes the student model, the ontology slicer, the distillation pipeline, the evaluation gate system, and the compliance metric\.[Section˜4](https://arxiv.org/html/2607.11948#S4)reports the held\-out result and the training\-side mechanism\.[Section˜5](https://arxiv.org/html/2607.11948#S5)reports the consolidated contextuality\-audit method and its negative result\.[Section˜6](https://arxiv.org/html/2607.11948#S6)discusses scope, threats to validity, and the full\-power evaluation that follows\.
### Combined\-submission note
This manuscript consolidates two related arXiv\-ready FAOS research notes into one article\. The first was the ontology\-amplified\-distillation proof of mechanism reported in the main empirical sections\. The second was a contextuality\-auditor method note for enterprise LLM\-agent routing\. They share the same applied setting—ontology\-grounded enterprise agents under regulated deployment constraints—and are combined here as one model\-building\-plus\-governance article rather than submitted as separate variations on the same theme\.
### Dual\-use disclosure
This study reports results from a dual\-use experimental pipeline that serves two outputs: an internal minimum\-viable sovereign model for theFAOSplatform, and the academic contribution reported here\. The principal investigator is a co\-founder of the platform vendor and holds a commercial interest in the deployable artifact\.[Section˜6](https://arxiv.org/html/2607.11948#S6)states this researcher\-as\-practitioner position and the mitigations applied, which include a locked held\-out set, deterministic scoring, and a compliance metric computed without a judge model\.
The method also engages a stated counter\-position\. Microsoft AI’s MAI\-Thinking\-1 report argues that capabilities should be learned, not inherited, holding that intelligence acquired through distillation lacks the steerability and robustness of capability trained from scratch\[The Microsoft AI Team,[2026](https://arxiv.org/html/2607.11948#bib.bib28)\]\. The objection applies, on its face, to ontology\-amplified distillation, so it is worth stating plainly\. Three points bound the tension\. The objection targets dependence on a third\-party teacher, not distillation as such; the same report uses self\-distillation\. The objective also differs: the MAI program builds a frontier model, whereas this work ships a compliant model that must run inside a data\-residency perimeter, which makes distillation a consequence of the deployment constraint rather than a shortcut\. The steerability concern is nonetheless real, and the abstention gate that would test it—whether the student declines out\-of\-distribution prompts at a governed rate—belongs to the full\-power evaluation deferred here, not to this proof of mechanism\.
## 2Related Work
### 2\.1Distillation and efficient student models
Knowledge distillation trains a compact student to reproduce a larger teacher’s behavior\[Hintonet al\.,[2015](https://arxiv.org/html/2607.11948#bib.bib23)\]\. The technique scaled from classification to language models with DistilBERT\[Sanhet al\.,[2019](https://arxiv.org/html/2607.11948#bib.bib24)\]and, for instruction following, to preference\-based distillation such as Zephyr, which distills alignment from a teacher’s ranked outputs\[Tunstallet al\.,[2023](https://arxiv.org/html/2607.11948#bib.bib25)\]\. Industrial pipelines now publish distilled open students in the Qwen family\[Wanget al\.,[2025](https://arxiv.org/html/2607.11948#bib.bib26)\], and low\-rank adapters make student adaptation cheap on commodity hardware\[Huet al\.,[2022](https://arxiv.org/html/2607.11948#bib.bib1)\]\. This body of work concentrates on general\-purpose capability: the student should track the teacher across broad benchmarks\. The present study asks a narrower question—whether distillation can transfer one specific behavior, adherence to a domain ontology, into a student that must run inside a regulated perimeter\.
### 2\.2Preference optimization
Learning from preferences began with reinforcement learning from human feedback\[Christianoet al\.,[2017](https://arxiv.org/html/2607.11948#bib.bib10)\]and its alignment variants\[Baiet al\.,[2022](https://arxiv.org/html/2607.11948#bib.bib11)\]\. Direct preference optimization \(DPO\) removed the separate reward model, training the policy directly on pairs of preferred and dispreferred responses\[Rafailovet al\.,[2023](https://arxiv.org/html/2607.11948#bib.bib22)\]\. The preference target here is neither human\-likeness nor general helpfulness but ontology adherence\. The preferred response is an ontology\-grounded frontier answer; the dispreferred response is the same base model answering without the ontology slice\. The signal isolates the grounded behavior from the teacher’s broader parametric knowledge, which is the property a sovereign student needs to acquire\.
### 2\.3Ontology\-grounded generation
A growing line of work couples language models to structured knowledge\. Surveys map the integration of large language models with knowledge graphs\[Panet al\.,[2024](https://arxiv.org/html/2607.11948#bib.bib5)\], and the neurosymbolic program frames the broader aim of combining learned and symbolic representation\[Garcez and Lamb,[2023](https://arxiv.org/html/2607.11948#bib.bib3), Hitzler and Sarker,[2022](https://arxiv.org/html/2607.11948#bib.bib4)\]\. Recent systems ground generation at inference through ontology\-anchored retrieval\[Sharmaet al\.,[2025](https://arxiv.org/html/2607.11948#bib.bib6)\]or align a model to an ontology through self\-training\[Liuet al\.,[2025](https://arxiv.org/html/2607.11948#bib.bib7)\], and language models have been used to build ontologies in turn\[Babaei Giglouet al\.,[2023](https://arxiv.org/html/2607.11948#bib.bib8)\]\. The question reaches back to the symbol grounding problem\[Harnad,[1990](https://arxiv.org/html/2607.11948#bib.bib9)\]\. The companion neurosymbolic study in this program measured ontology grounding on a frontier model and reported the Inverse PKE: grounding helped most where the model’s parametric coverage of a domain was thinnest\[Luong and Sanyal,[2026](https://arxiv.org/html/2607.11948#bib.bib19)\]\. That result motivates the present question\. If grounding compensates for thin parametric coverage, a distilled student, which carries thinner coverage than the frontier, should gain at least as much from the same ontology—the mechanism this study sets out to test\.
### 2\.4Sovereign deployment under data residency
Vietnamese financial regulation constrains where customer data may be processed\. The 2025 Law on Artificial Intelligence introduces a tiered risk regime with sector grace periods\[National Assembly of Vietnam,[2025](https://arxiv.org/html/2607.11948#bib.bib15)\], the banking sandbox decree structures AI\-enabled financial services\[Government of Vietnam,[2025](https://arxiv.org/html/2607.11948#bib.bib16)\], and the anti\-money\-laundering and insurance rules impose due\-diligence and solvency obligations that reach model\-mediated decisions\[National Assembly of Vietnam,[2022](https://arxiv.org/html/2607.11948#bib.bib17), Ministry of Finance of Vietnam,[2023](https://arxiv.org/html/2607.11948#bib.bib18)\]\. Under these constraints a tenant cannot route customer data to a foreign frontier API, so the deployable model must run on hardware the institution controls\. On\-device frameworks make local inference of mid\-sized models practical on commodity accelerators\[Apple Machine Learning Research,[2025](https://arxiv.org/html/2607.11948#bib.bib27)\]\. The governance frameworks that regulated buyers map against—the EU AI Act, the NIST AI Risk Management Framework, and ISO/IEC 42001—treat traceability and control as first\-order requirements\[European Parliament and Council,[2024](https://arxiv.org/html/2607.11948#bib.bib12), National Institute of Standards and Technology,[2023](https://arxiv.org/html/2607.11948#bib.bib13), International Organization for Standardization,[2023](https://arxiv.org/html/2607.11948#bib.bib14)\], which a tenant\-owned model satisfies more directly than a remote service\. The deployment constraint, not a cost preference, makes the sovereign student the object of study\.
## 3Method
### 3\.1Student model and deployment target
The student is Qwen3\.6\-27B, an open model in the Qwen family\[Yanget al\.,[2025](https://arxiv.org/html/2607.11948#bib.bib2)\]built on theqwen3\_5architecture\. Training and evaluation run locally in Apple’s MLX framework on a single M5 Max with 128 GB of unified memory\[Apple Machine Learning Research,[2025](https://arxiv.org/html/2607.11948#bib.bib27)\]; no cloud compute is used\. After training, the student is fused for local inference; its deployment target is a four\-bit \(nvfp4\) checkpoint of roughly 14 GB, the form in which it would run inside a tenant perimeter\. Because the Gate 8 grounding metric scores term coverage rather than generation quality, it is insensitive to serving precision, and the held\-out result does not depend on the exact precision of the deployed checkpoint\. The wider research program targets a 14B student and a curve across model sizes\. This proof of mechanism reports the 27B instance that was trained and evaluated, and holds the smaller\-student and cross\-scale questions for full\-power work \([Section˜6](https://arxiv.org/html/2607.11948#S6)\)\.
### 3\.2Ontology slicing
The Foundation AgenticOS ontology represents an enterprise domain across four sections: the role performing the work, the domain concepts in scope, the regulations that bind the task, and the interaction pattern that moves work between actors\. For a given task, the slicer renders a task\-specific*ontology slice*from these sections under a fixed token budget of 800 tokens \(slicer version 1\.0\.0\)\. When a slice exceeds the budget, the slicer drops content in a fixed order—interaction first, then domain—and never drops the regulation or role sections, so the compliance\-bearing content survives truncation\. The same slice is injected during training, as context for the teacher, and during evaluation, as context for both the student and the frontier, which holds the grounding signal identical across the pipeline\.
### 3\.3The distillation pipeline
Adaptation runs in two stages on top of a capture step\.
*Capture\.*The frontier teachers GPT\-5 and GPT\-5\-mini answer a non\-held\-out training manifest with the ontology slice injected, producing 451 grounded trajectories at a metered cost of $2\.98\. The held\-out evaluation tasks are excluded from capture\.
*Supervised fine\-tuning\.*A low\-rank adapter over four layers is trained on the captured trajectories, reducing training loss from 2\.370 to 0\.496\. The adapter is fused into the base model and re\-quantized to nvfp4, producing the supervised student that serves as both the policy and the frozen reference for the next stage\.
*Ontology\-grounded DPO\.*Direct preference optimization runs withmlx\-lm\-lora2\.1\.0 on 47 synthetic preference pairs \(42 for training, 5 for validation\)\. Each pair marks the ontology\-grounded frontier answer as preferred and the base model’s no\-slice answer as dispreferred; the ontology\-compliance gap between the two ranges from 0\.17 to 0\.53 \(mean 0\.27\), with no inverted pairs\. Training uses a rank\-8 adapter over four layers,β=0\.1\\beta=0\.1, a learning rate of5×10−75\\times 10^\{\-7\}, the sigmoid DPO loss, batch size 1, a 512\-token sequence cap, and 42 iterations of about one epoch under seed 20260615, updating 0\.014% of parameters\. A composite reward with weightsα=0\.50\\alpha=0\.50on task,β=0\.35\\beta=0\.35on ontology, andγ=0\.15\\gamma=0\.15on role governs preference scoring \(ADR\-SL\-006\)\. The run completes in roughly 32 minutes at a peak of 58\.1 GB, fuses to the trained student \(adapter38c28df0\), and costs nothing beyond local power\.
The preference data are synthetic, and three composition facts bound their reach\. They are English\-only; their source domains are healthcare \(23 pairs\), insurance \(13\), and fintech \(11\), none of which is a Vietnamese evaluation vertical; and the 47 rows reduce to 15 distinct prompts under roughly threefold duplication, with the five validation rows drawn from two prompts, one of which also appears in training\. The protocol’s design anticipated subject\-matter\-expert preference labels and a larger pool; this run uses 47 synthetic pairs—about nine percent of the planned floor—and zero expert labels, so it makes no expert\-preference claim\. The defensible reading is that the preference signal demonstrates the training mechanism on constructed pairs rather than a production\-grade preference set, and that any transfer to Vietnamese financial grounding is cross\-lingual and cross\-domain \([Section˜6](https://arxiv.org/html/2607.11948#S6)\)\.
### 3\.4Evaluation gate system
The program scores a trained student against eight gates and a canary leakage check\. Two gates carry binding authority\. Gate 7 tests confident\-wrongness, the rate at which the student abstains on out\-of\-distribution prompts, and Gate 8 tests ontology compliance, whether the student grounds its answers in the task’s ontology relative to the frontier\. Gate 8 is the thesis gate for ontology\-amplified distillation: it passes when the student’s mean compliance is at least the frontier’s, with no tolerance margin\. This study runs Gate 8 alone\. Gate 7 and the capability gates—in\-distribution quality, cost ratio, dialect robustness, polysemy, and multi\-turn consistency—require judge sign\-off or evaluation sets above their full\-power size floors, and are deferred \([Section˜6](https://arxiv.org/html/2607.11948#S6)\)\.
### 3\.5Held\-out evaluation set
The evaluation set holds 40 Vietnamese\-language tasks, ten in each of four financial verticals: banking, insurance, capital markets, and real\-estate technology\. Tasks span four question types—terminological fidelity, metric accuracy, regulatory compliance, and role consistency—together with a cross\-cutting category\. The set is version\-locked; its SHA\-256 digest isa8c1d43b, a re\-lock of an earlier digest \(bbc14058\) after each task was annotated in place with its ontology terms\. A pinned extractor derives each task’s terms from the same ontology slice the evaluator uses, so the ground\-truth terms and the scored slice share one source\. Two verticals, banking and insurance, were redrawn for this version after an earlier exposure, and the two unexposed verticals, capital markets and real\-estate technology, were carried forward; the redraw preserves the task\-level leakage boundary, but a residual span\-level overlap risk—training artifacts derived from the same blueprint sections the holdout cites—is examined in[Section˜6](https://arxiv.org/html/2607.11948#S6)\.
### 3\.6The ontology\-compliance metric
Both the student and the frontier are scored by the same local metric over the same injected slice, with no judge model in the loop\. The frontier baseline is GPT\-5 run at minimal reasoning effort: the slice\-laden prompts otherwise exhaust the reasoning\-token budget and truncate the answer, and term coverage is insensitive to reasoning depth, so this is a defensible baseline for this metric—but every “equal to the frontier” in this paper refers to that configuration\.
###### Definition 1\(Ontology\-compliance score\)\.
For a responseaato a task with ontology termsTTand metric rangesMM, the ontology\-compliance score is
ronto\(a\)=12sterm\(a,T\)\+12smetric\(a,M\),r\_\{\\mathrm\{onto\}\}\(a\)\\;=\\;\\tfrac\{1\}\{2\}\\,s\_\{\\mathrm\{term\}\}\(a,T\)\\;\+\\;\\tfrac\{1\}\{2\}\\,s\_\{\\mathrm\{metric\}\}\(a,M\),wheresterm=1s\_\{\\mathrm\{term\}\}=1ifaamentions at least one term inTTand0otherwise, andsmetrics\_\{\\mathrm\{metric\}\}is the share of the response’s cited metric values that fall within their healthy range, defined as11when the response makes no metric claim\.
The held\-out tasks carry ontology terms but no metric ranges, sosmetrics\_\{\\mathrm\{metric\}\}defaults to 1 androntor\_\{\\mathrm\{onto\}\}becomes two\-valued: 1\.0 when the answer surfaces at least one of the task’s ontology terms, and 0\.5 otherwise\. This executed metric is a degenerate proxy for the pre\-registered four\-component ontology\-compliance composite—terminological fidelity, metric accuracy, regulatory compliance, and role consistency—of which three components, and the correct\-usage qualifier of the first, are not scored here; the metric is also floored at 0\.50 rather than 0\. The Gate 8 statistic is the mean ofrontor\_\{\\mathrm\{onto\}\}over the 40 tasks, computed identically for the student and the frontier; because of the floor it ranges over\[0\.5,1\.0\]\[0\.5,1\.0\], so the grounded*rate*—the share of tasks scoring 1\.0—is the more directly interpretable quantity\. The metric records term presence, not answer completeness or correctness; a finer\-grained quality measure is full\-power work \([Section˜6](https://arxiv.org/html/2607.11948#S6)\)\.
## 4Results
### 4\.1Ontology compliance
The distilled student and the frontier reach the same grounded count on Gate 8\. Over the 40 held\-out tasks each grounds 36 \(a grounded rate of 0\.90; meanronto=0\.950r\_\{\\mathrm\{onto\}\}=0\.950under the 0\.50 floor\), so the gate—which passes when the student’s mean is at least the frontier’s—passes \([Table˜1](https://arxiv.org/html/2607.11948#S4.T1)\)\. Scoring is deterministic at temperature 0, and a re\-run reproduced the verdict exactly; the evaluation took about 993 seconds on the M5 Max\.
This is an equal point estimate, not a demonstrated equivalence\. With a binary per\-task outcome atn=40n=40, the grounded rate of 36/40 carries a Wilson 95% confidence interval of\[0\.77,0\.96\]\[0\.77,0\.96\], and the paired comparison is uninformative: the discordant cells are balanced \(two tasks each\), so an exact McNemar test returnsp=1\.00p=1\.00and the 95% interval on the mean paired difference spans±0\.05\\pm 0\.05, or±4\\pm 4tasks\. The result establishes an equal count, not statistical equivalence, and is underpowered to rule out a real gap of up to roughly ten percentage points in either direction\.
The equal totals rest on partly different tasks \([Figure˜2](https://arxiv.org/html/2607.11948#S4.F2)and[Figure˜1](https://arxiv.org/html/2607.11948#S4.F1); full per\-task list in[Appendix˜B](https://arxiv.org/html/2607.11948#A2)\)\. Two tasks defeat both models, and each grounds two the other misses; the student is perfect on banking and insurance and grounds 8 of 10 in each of capital markets and real\-estate technology, while the frontier is stronger on capital markets \(9 of 10\) and weaker on real\-estate technology \(7 of 10\)\. The small number of disjoint misses \(two each\) is consistent with independent grounding rather than task\-for\-task imitation of the teacher, though four disjoint misses atn=40n=40cannot establish independence\. The perfect banking and insurance scores fall on exactly the two verticals redrawn for this version, which also carry the largest ontology\-term lists; a one\-term threshold is easiest to clear there, so the clean sweep is weak evidence of grounding strength \([Section˜6](https://arxiv.org/html/2607.11948#S6)\)\.
Table 1:The distilled student and GPT\-5 reach the same grounded count \(36/40\) through partly different verticals—the student perfect on banking and insurance, the frontier stronger on capital markets\. Gate 8 ontology grounding by vertical \(held\-out set,n=40n=40\); a task is grounded when the answer surfaces at least one of its ontology terms, and meanrontor\_\{\\mathrm\{onto\}\}assigns 1\.0 to a grounded task and 0\.5 otherwise\.*Source: Gate 8 evaluation, held\-out set v1\.2 \(SHAa8c1d43b\), 2026\-06\-14; deterministic scoring atT=0T=0\.*Figure 1:Grounded\-task count by vertical, student vs\. the GPT\-5 frontier \(held\-out set, ten tasks per vertical,n=40n=40\)\. The two reach the same total \(36/40\) through a mirror split—the student stronger on real\-estate technology, the frontier on capital markets\. The axis is grounded count \(true zero\), not the floored mean\.*Source: Gate 8 evaluation, 2026\-06\-14, holdouta8c1d43b\.*Figure 2:Per\-task student×\\timesfrontier grounding agreement \(held\-out set,n=40n=40\)\. Both models ground 34 tasks and both miss 2; each grounds 2 the other misses, so the equal 36/40 totals rest on partly different tasks\.*Source: Gate 8 evaluation, 2026\-06\-14, holdouta8c1d43b\.*
### 4\.2The training mechanism
The DPO objective separated the constructed preference pairs \([Table˜2](https://arxiv.org/html/2607.11948#S4.T2)\)\. Held\-out preference accuracy moved from 0 to 1\.0 on a five\-pair validation split, the reward margin between the preferred and dispreferred response from 0 to 0\.307, and the DPO loss fell from 0\.693 to 0\.551\. A second proof\-scale run on the same 47\-pair construction reproduced the pattern \(preference accuracy 0 to 1\.0, reward margin 0 to 0\.319\); because it reuses the same data design it is a repeat, not an independent replication\. The validation split is small enough—five pairs from two prompts, one shared with training—that the 1\.0 should be read as separability of the constructed pairs, not as generalized grounding\. With that caveat, the training\-side measurement is consistent with the intended preference shift: ontology\-grounded tuning moves the student toward surfacing the ontology terms the frontier surfaces when it carries the same slice\.
Table 2:Ontology\-grounded DPO moves held\-out preference accuracy from 0 to 1\.0 and the reward margin to 0\.307 in one epoch\. Training\-side preference learning; “Before” is iteration 0 and “After” the final iteration, on a five\-pair held\-out validation split\.*Source: DPO run, seed 20260615,mlx\-lm\-lora2\.1\.0\.*
## 5Contextuality Audit for Enterprise\-Agent Routing
Ontology\-amplified distillation addresses one side of the sovereign\-enterprise problem: how to move a locally deployable student toward frontier\-level grounding on regulated tasks\. A platform also needs a governance diagnostic for the cases where a model’s answer changes across role frames, prompt orders, ontology conditions, or escalation criteria\. Some variation is ordinary direct influence: the prompt changed, so the marginal answer changed\. The harder question is whether any residual inconsistency remains after that direct influence is accounted for, and whether the residual should trigger multi\-agent debate, synthesis, or human review\.
We treat this as a contextuality\-audit problem rather than as a claim that LLMs are quantum systems\. Contextuality\-by\-Default \(CbD\) indexes each random variable by both the content being measured and the context in which it is measured\[Dzhafarov and Kujala,[2016a](https://arxiv.org/html/2607.11948#bib.bib30),[b](https://arxiv.org/html/2607.11948#bib.bib31), Kujala and Dzhafarov,[2016](https://arxiv.org/html/2607.11948#bib.bib32)\]\. This is well suited to enterprise\-agent outputs: the same business content can be measured under a business\-owner frame, risk\-reviewer frame, governance\-reviewer frame, or neutral\-arbitrator frame\. The audit’s purpose is operational\. It asks whether the observed variation should be interpreted as stable behavior, measurement design, construct coupling, direct influence, or residual contextuality\. That interpretation then maps to the least intrusive orchestration remedy: solo response, prompt standardization, debate or synthesis, or human arbitration\. This diagnostic complements multi\-agent debate and orchestration work, which can improve selected outputs but does not by itself diagnose when routing is needed\[Duet al\.,[2024](https://arxiv.org/html/2607.11948#bib.bib33), Guoet al\.,[2024](https://arxiv.org/html/2607.11948#bib.bib34)\]\.
The consolidated RA\-15 pilot used local Qwen 3\.6 27B as the first model\. Phase 1 produced 576 valid outputs; Phase 1\.1 added 336 valid q4 robustness outputs; Phase 1\.2 added 480 role/order\-disentanglement outputs; and Phase 1\.3 added a 384\-call construct\-decoupling extension separating operational permission from procedural evidence readiness\. The original instrument surfaced apparent “CNTX” signals, but a formalism audit corrected the disturbance normalization: the earlier nonzero values were probability\-drift residuals, not the canonical cyclic binary CbD degree\. Under expectation\-scale direct\-influence correction, the canonical CbD degree is zero for all Phase 1\.3 groups in both the local\-Qwen run and an explicitly labeled Gemma replication check\. The negative result is the finding\. In this pilot, apparent contextuality collapses into direct influence and construct coupling rather than surviving as residual contextuality\.
Table 3:The RA\-15 evidence package supports a negative\-results routing method, not a contextuality\-positive claim\.Table 4:Contextuality\-audit interpretation used by the combined article\.The link to ontology\-amplified distillation is governance rather than shared training data\. The distillation result asks whether ontology grounding can move a sovereign student to frontier\-equal term coverage on a narrow held\-out set\. The contextuality audit asks how an enterprise platform should treat remaining decision variation once such a model is deployed inside a governed workflow\. Together, they define a combined mechanism\-and\-control claim: the model can be made more domain\-grounded, and the platform should still separate prompt sensitivity, construct design, and residual conflict before escalating the task\.
The contextuality component therefore adds a negative\-results method claim to the proof\-of\-mechanism claim\. It does not license a contextuality\-positive routing rule\. It licenses a safer precursor rule: before treating model disagreement as evidence that a task requires debate or human arbitration, run a direct\-influence and construct\-validity audit\. This avoids over\-escalating tasks on the basis of an under\-normalized score, and it also avoids hiding genuine conflict behind a single context\-dependent answer\.
## 6Discussion
### 6\.1What the result shows
Two measurements point the same way\. The student reaches the frontier’s grounded count on the held\-out set, and the training\-side metrics are consistent with the DPO objective—rather than chance—producing the shift on the constructed pairs\. For a sovereign\-deployment setting the reading is specific: a locally deployable Qwen3\.6\-27B student, tuned on synthetic ontology\-grounded preferences, grounds Vietnamese financial\-domain answers as often as GPT\-5 does on this term\-coverage metric\. The divergent miss pattern is suggestive—parity that survives a few independent errors is harder to explain as task\-for\-task imitation—but four disjoint misses atn=40n=40make it a weak signal, not a demonstration\.
### 6\.2What the result does not show
The claim is bounded on purpose\. The metric records whether an answer surfaces at least one ontology term, not whether the answer is complete or correct, so the grounded rate of 0\.90 \(meanronto=0\.95r\_\{\\mathrm\{onto\}\}=0\.95under the floor\) measures grounding presence, not domain quality\. The pre\-registered primary hypothesis was*amplification*—the distilled student’s normalized ontology\-lift*exceeding*the frontier’s \(νD\>νF\\nu\_\{D\}\>\\nu\_\{F\}\), with a tie \(νD≤νF\\nu\_\{D\}\\leq\\nu\_\{F\}\) named as the disconfirmation condition\. This run does not test that hypothesis: it scores a reduced term\-coverage proxy rather than the registered four\-component composite, runs on the 27B instance rather than the 14B target, and computes no C1\-to\-C3 lift\. The equal count is therefore neither a confirmation of amplification nor, given the proxy, a clean disconfirmation; amplification remains untested at power, and the Inverse\-PKE framing should be read as motivation, not as a result this run supports\. The evaluation ran one binding gate\. Confident\-wrongness abstention, in\-distribution quality, cost, dialect robustness, polysemy, and multi\-turn consistency were not measured\. The preference data are synthetic, English\-only, drawn from non\-target domains \(healthcare, insurance, fintech\), and reduce to 15 distinct prompts; no subject\-matter\-expert preference label enters the pipeline\. The student is a single 27B model from one family, evaluated on 40 Vietnamese tasks across four verticals; the smaller\-student target, the cross\-scale curve, and the Vietnamese\-versus\-English grounding contrast that would replicate the Inverse PKE in the distilled regime are not part of this evidence\. No production or procurement decision should rest on a pilot of this size\.
### 6\.3Threats to validity
*Construct validity\.*Term\-coverage is a coarse proxy for grounding, and the executed metric scores only one of the four pre\-registered ontology\-compliance components\. Two design choices make the bar low\. The gold terms are extracted from the same slice injected to both models, so each side is handed the target vocabulary and then scored on echoing at least one token of it; and a response can surface a term without reasoning over it correctly, so the metric credits shallow mentions\. Scoring the frontier by the identical metric bounds the comparison—both models face the same near\-ceiling test—but the absolute 0\.95 and the compression of any student\-frontier gap toward zero follow partly from the metric’s construction, not only from grounding ability\. A coverage\-fraction or held\-out\-term variant, and the full four\-component composite, are the first things a full\-power evaluation restores\.
*Internal validity\.*The preferred responses are frontier answers, so the student may be learning to surface terms in the frontier’s style rather than acquiring independent grounding\. Because the preference pairs are English\-only and from non\-target domains while the evaluation is Vietnamese financial, the behavior the student was tuned on \(echoing the preferred answer’s terms\) and the behavior the metric rewards may be the same surface pattern measured twice\. The modest compliance gap in the pairs \(0\.17 to 0\.53\), the synthetic\-only construction, and the absence of expert validation leave this open\.
*External validity\.*Four verticals, one base\-model family, and 40 tasks do not support generalization beyond the tested frame, and the result is Vietnamese\-only\.
*Leakage\.*The held\-out set excludes the evaluation tasks at the task\-identifier level and redraws the two previously exposed verticals\. This controls task\-level overlap but not span\-level overlap: the supervised\-fine\-tuning corpus is generated from the same Vietnamese blueprint sections the holdout cites, so a training artifact can convey a held\-out task’s gold terms while passing the task\-id check—a risk the project’s own leakage specification flags as making the identifier\-only assertion structurally vacuous for those corpora\. A span\-overlap assertion exists as a tested primitive, but its wiring into the corpus generators was incomplete at the time of the Wave\-0\.5 run, so span\-level exclusion is not established for the executed supervised corpus\. The synthetic DPO pairs, being English and blueprint\-free, are exempt\. Until span\-level screening is confirmed end\-to\-end, part of the held\-out grounding could reflect memorized blueprint vocabulary rather than transferred grounding\.
*Reproducibility\.*The final publication manifest hash\-pins all 451 committed source trajectories, the corpus\-building code and configuration, the DPO corpus, the locked holdout, both final\-evaluation JSON files, governance reviews, and the arXiv package\. The executed v1\.2\-screened SFT train/validation JSONL files and local SFT/DPO checkpoints were not archived before operator\-side cleanup\. Later ontology\-blueprint changes also prevent a byte\-identical SFT rebake from the current tree\. The evidence package therefore supports audit of the reported inputs, procedure, and Gate 8 arithmetic, but not bitwise reproduction of the trained checkpoint\. The study makes no checkpoint\-level reproducibility claim\.
*Conclusion validity\.*The metric is binary per task, and the equal means rest on 36 of 40 grounded for each model; a single task shifts the mean by 0\.0125\. The paired\-difference 95% confidence interval spans±4\\pm 4tasks\. The study reports parity on this metric and makes no claim of statistical superiority or equivalence\.
### 6\.4Researcher\-as\-practitioner position
The principal investigator is a co\-founder of the platform whose sovereign model this study evaluates, and the experimental pipeline produces both the internal artifact and this paper\. The design applies the mitigations available to a single\-pipeline study: a version\-locked held\-out set \(its questions sealed before the run, its gold\-term annotation added in place under a pinned extractor and independently re\-reviewed afterward\), deterministic scoring at temperature 0, a compliance metric computed locally without a judge model, and a pre\-registered protocol\. The dual role is disclosed rather than corrected away; the deferred full\-power evaluation, with independent judge sign\-off, is where the harder bias controls apply\.
### 6\.5Toward full\-power evaluation
The questions this proof of mechanism leaves open define the next study\. The binding safety gate—whether the student abstains on out\-of\-distribution prompts at a governed rate—tests the steerability that distilled models are accused of failing to acquire\[The Microsoft AI Team,[2026](https://arxiv.org/html/2607.11948#bib.bib28)\], and it is the first item\. Beyond it: the capability gates at full evaluation\-set size, frontier\-judge sign\-off on answer quality, a measured cost ratio against the frontier API, the 14B student and the cross\-scale curve, subject\-matter\-expert preference labels in place of synthetic pairs, and the Vietnamese\-versus\-English contrast that would test whether the Inverse PKE holds in the distilled regime\. Each turns a deferred claim into a measured one\.
## 7Conclusion
Ontology\-amplified distillation moved a sovereign, locally deployable Qwen3\.6\-27B student to the same held\-out grounded count as GPT\-5 on a coarse ontology term\-coverage metric—36 of 40 tasks each \(0\.90; meanronto=0\.95r\_\{\\mathrm\{onto\}\}=0\.95under the 0\.50 floor\)—with a training\-side preference signal that separated the constructed pairs \(0 to 1\.0\)\. The equal count is an underpowered point estimate, not demonstrated equivalence, and not the pre\-registered amplification\. The consolidated contextuality\-audit result is likewise bounded: after correcting direct influence, the canonical CbD degree is zero in the reported pilot, so the useful finding is a negative\-results method for separating direct influence and construct coupling from residual contextuality before routing tasks to debate, synthesis, or human review\. As a combined article, the contribution is a mechanism\-and\-control pair\. The mechanism is ontology\-amplified distillation for sovereign enterprise models; the control is a contextuality audit that prevents over\-reading prompt\-sensitive variation\. Both results license fuller evaluation rather than deployment claims\. Whether the student’s grounding is complete and correct, whether it abstains safely, what it costs against a frontier API, and whether the effect holds at smaller scale and across languages—these define the evaluation that follows\.
## Appendix ATraining Hyperparameters
Table 5:Ontology\-grounded DPO configuration \(mlx\-lm\-lora2\.1\.0\)\.
## Appendix BPer\-Task Grounding
Of the 40 held\-out tasks the student grounds 36\. Its four misses are two terminological\-fidelity tasks in capital markets and two real\-estate\-technology tasks: one metric\-accuracy task and one terminological\-fidelity task\. The frontier also grounds 36 of 40\. It misses one capital\-markets metric\-accuracy task and three real\-estate\-technology tasks: two metric\-accuracy tasks and one terminological\-fidelity task\. Two real\-estate\-technology tasks defeat both models; the remaining misses are disjoint, so each model grounds two tasks the other does not\. The equal totals therefore rest on different task sets, which is the basis for reading the parity as independent grounding rather than teacher imitation\.
## References
- Apple Machine Learning Research \(2025\)Exploring LLMs with MLX and the neural accelerators in the M5 GPU\.Note:Apple ML Research blog,[https://machinelearning\.apple\.com/research/exploring\-llms\-mlx\-m5](https://machinelearning.apple.com/research/exploring-llms-mlx-m5)Verified 2026\-06\-22; title and URL confirmed against Apple ML ResearchCited by:[§1](https://arxiv.org/html/2607.11948#S1.p3.1),[§2\.4](https://arxiv.org/html/2607.11948#S2.SS4.p1.1),[§3\.1](https://arxiv.org/html/2607.11948#S3.SS1.p1.1)\.
- H\. Babaei Giglou, J\. D’Souza, and S\. Auer \(2023\)LLMs4OL: large language models for ontology learning\.arXiv preprint arXiv:2307\.16648\.Cited by:[§2\.3](https://arxiv.org/html/2607.11948#S2.SS3.p1.1)\.
- Y\. Bai, S\. Kadavath, S\. Kundu, A\. Askell, J\. Kernion, A\. Jones, A\. Chen, A\. Goldie, A\. Mirhoseini, C\. McKinnon,et al\.\(2022\)Constitutional AI: harmlessness from AI feedback\.arXiv preprint arXiv:2212\.08073\.Cited by:[§2\.2](https://arxiv.org/html/2607.11948#S2.SS2.p1.1)\.
- P\. F\. Christiano, J\. Leike, T\. Brown, M\. Martic, S\. Legg, and D\. Amodei \(2017\)Deep reinforcement learning from human preferences\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Vol\.30,pp\. 4299–4307\.Cited by:[§2\.2](https://arxiv.org/html/2607.11948#S2.SS2.p1.1)\.
- Y\. Du, S\. Li, A\. Torralba, J\. B\. Tenenbaum, and I\. Mordatch \(2024\)Improving factuality and reasoning in language models through multiagent debate\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 11733–11763\.External Links:[Link](https://proceedings.mlr.press/v235/du24e.html)Cited by:[§5](https://arxiv.org/html/2607.11948#S5.p2.1)\.
- E\. N\. Dzhafarov and J\. V\. Kujala \(2016a\)Context–content systems of random variables: the contextuality\-by\-default theory\.Journal of Mathematical Psychology74,pp\. 11–33\.External Links:[Document](https://dx.doi.org/10.1016/j.jmp.2016.04.010),[Link](https://doi.org/10.1016/j.jmp.2016.04.010)Cited by:[§5](https://arxiv.org/html/2607.11948#S5.p2.1)\.
- E\. N\. Dzhafarov and J\. V\. Kujala \(2016b\)Contextuality\-by\-default 2\.0: systems with binary random variables\.External Links:1604\.04799,[Document](https://dx.doi.org/10.48550/arXiv.1604.04799),[Link](https://doi.org/10.48550/arXiv.1604.04799)Cited by:[§5](https://arxiv.org/html/2607.11948#S5.p2.1)\.
- European Parliament and Council \(2024\)Regulation \(EU\) 2024/1689 — artificial intelligence act\.Note:Official Journal of the European Union, L seriesEntered into force August 1, 2024Cited by:[§2\.4](https://arxiv.org/html/2607.11948#S2.SS4.p1.1)\.
- A\. d\. Garcez and L\. C\. Lamb \(2023\)Neurosymbolic AI: the 3rd wave\.Artificial Intelligence Review56,pp\. 12387–12406\.Note:Originally circulated 2019; published 2023External Links:[Document](https://dx.doi.org/10.1007/s10462-023-10448-w)Cited by:[§2\.3](https://arxiv.org/html/2607.11948#S2.SS3.p1.1)\.
- Government of Vietnam \(2025\)Decree 94/2025/nd\-cp on the regulatory sandbox in the banking sector\.Note:Effective July 1, 2025First comprehensive fintech sandbox regulation; structured framework for AI\-enabled financial servicesCited by:[§1](https://arxiv.org/html/2607.11948#S1.p1.1),[§2\.4](https://arxiv.org/html/2607.11948#S2.SS4.p1.1)\.
- T\. Guo, X\. Chen, Y\. Wang, R\. Chang, S\. Pei, N\. V\. Chawla, O\. Wiest, and X\. Zhang \(2024\)Large language model based multi\-agents: a survey of progress and challenges\.InProceedings of the Thirty\-Third International Joint Conference on Artificial Intelligence,pp\. 8048–8057\.External Links:[Document](https://dx.doi.org/10.24963/ijcai.2024/890),[Link](https://doi.org/10.24963/ijcai.2024/890)Cited by:[§5](https://arxiv.org/html/2607.11948#S5.p2.1)\.
- S\. Harnad \(1990\)The symbol grounding problem\.Physica D: Nonlinear Phenomena42\(1–3\),pp\. 335–346\.External Links:[Document](https://dx.doi.org/10.1016/0167-2789%2890%2990087-6)Cited by:[§2\.3](https://arxiv.org/html/2607.11948#S2.SS3.p1.1)\.
- G\. Hinton, O\. Vinyals, and J\. Dean \(2015\)Distilling the knowledge in a neural network\.arXiv preprint arXiv:1503\.02531\.Note:Verified 2026\-05\-26Cited by:[§1](https://arxiv.org/html/2607.11948#S1.p3.1),[§2\.1](https://arxiv.org/html/2607.11948#S2.SS1.p1.1)\.
- P\. Hitzler and M\. K\. Sarker \(2022\)Neuro\-symbolic artificial intelligence: the state of the art\.InNeuro\-Symbolic Artificial Intelligence: The State of the Art,Frontiers in Artificial Intelligence and Applications, Vol\.342\.External Links:[Document](https://dx.doi.org/10.3233/FAIA342)Cited by:[§2\.3](https://arxiv.org/html/2607.11948#S2.SS3.p1.1)\.
- E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. Chen \(2022\)LoRA: low\-rank adaptation of large language models\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2\.1](https://arxiv.org/html/2607.11948#S2.SS1.p1.1)\.
- International Organization for Standardization \(2023\)ISO/IEC 42001:2023 — artificial intelligence — management system\.Note:International StandardCited by:[§2\.4](https://arxiv.org/html/2607.11948#S2.SS4.p1.1)\.
- J\. V\. Kujala and E\. N\. Dzhafarov \(2016\)Proof of a conjecture on contextuality in cyclic systems with binary variables\.Foundations of Physics46,pp\. 282–299\.External Links:[Document](https://dx.doi.org/10.1007/s10701-015-9964-8),[Link](https://doi.org/10.1007/s10701-015-9964-8)Cited by:[§5](https://arxiv.org/html/2607.11948#S5.p2.1)\.
- Z\. Liu, C\. Gan, J\. Wang, Y\. Zhang, Z\. Bo, M\. Sun, H\. Chen, and W\. Zhang \(2025\)OntoTune: ontology\-driven self\-training for aligning large language models\.InProceedings of the ACM Web Conference 2025 \(WWW\),Note:arXiv:2502\.05478; SNOMED CT ontology\-driven LLM alignment via in\-context learning self\-training; code at github\.com/zjukg/OntoTuneExternal Links:[Document](https://dx.doi.org/10.1145/3696410.3714816)Cited by:[§1](https://arxiv.org/html/2607.11948#S1.p2.1),[§2\.3](https://arxiv.org/html/2607.11948#S2.SS3.p1.1)\.
- T\. T\. Luong and A\. Sanyal \(2026\)Ontology\-constrained neural reasoning in enterprise agentic systems: a neurosymbolic architecture for domain\-grounded ai agents\.arXiv preprint arXiv:2604\.00555\.Note:600 runs across 5 industries; Inverse PKE findingCited by:[§1](https://arxiv.org/html/2607.11948#S1.p2.1),[§2\.3](https://arxiv.org/html/2607.11948#S2.SS3.p1.1)\.
- Ministry of Finance of Vietnam \(2023\)Circular 67/2023/tt\-btc guiding the law on insurance business and decree 46/2023/nd\-cp\.Note:Issued November 2, 2023Financial regime, solvency margin, and risk\-management requirements for insurers and reinsurersCited by:[§2\.4](https://arxiv.org/html/2607.11948#S2.SS4.p1.1)\.
- National Assembly of Vietnam \(2022\)Law on prevention and combat of money laundering, law no\. 14/2022/qh15\.Note:Passed November 15, 2022; effective March 1, 2023Supersedes Law 07/2012/QH13; mandates risk\-based customer due diligence and beneficial\-ownership identification for credit institutions; aligned with FATF Recommendations 10 and 24Cited by:[§2\.4](https://arxiv.org/html/2607.11948#S2.SS4.p1.1)\.
- National Assembly of Vietnam \(2025\)Law on artificial intelligence, law no\. 134/2025/qh15\.Note:Passed December 10, 2025; effective March 1, 2026Three\-tier risk classification \(low/medium/high\); 18\-month grace period for finance, healthcare, educationCited by:[§1](https://arxiv.org/html/2607.11948#S1.p1.1),[§2\.4](https://arxiv.org/html/2607.11948#S2.SS4.p1.1)\.
- National Institute of Standards and Technology \(2023\)Artificial intelligence risk management framework \(AI RMF 1\.0\)\.Technical reportTechnical ReportNIST AI 100\-1,U\.S\. Department of Commerce\.External Links:[Document](https://dx.doi.org/10.6028/NIST.AI.100-1)Cited by:[§2\.4](https://arxiv.org/html/2607.11948#S2.SS4.p1.1)\.
- S\. Pan, L\. Luo, Y\. Wang, C\. Chen, J\. Wang, and X\. Wu \(2024\)Unifying large language models and knowledge graphs: a roadmap\.IEEE Transactions on Knowledge and Data Engineering36\(7\),pp\. 3580–3599\.External Links:[Document](https://dx.doi.org/10.1109/TKDE.2024.3352100)Cited by:[§2\.3](https://arxiv.org/html/2607.11948#S2.SS3.p1.1)\.
- R\. Rafailov, A\. Sharma, E\. Mitchell, S\. Ermon, C\. D\. Manning, and C\. Finn \(2023\)Direct preference optimization: your language model is secretly a reward model\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Note:arXiv:2305\.18290; verified 2026\-05\-26Cited by:[§1](https://arxiv.org/html/2607.11948#S1.p3.1),[§2\.2](https://arxiv.org/html/2607.11948#S2.SS2.p1.1)\.
- V\. Sanh, L\. Debut, J\. Chaumond, and T\. Wolf \(2019\)DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter\.arXiv preprint arXiv:1910\.01108\.Note:Verified 2026\-05\-26Cited by:[§2\.1](https://arxiv.org/html/2607.11948#S2.SS1.p1.1)\.
- K\. Sharma, P\. Kumar, and Y\. Li \(2025\)OG\-RAG: ontology\-grounded retrieval\-augmented generation for large language models\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),Note:\+55% fact recall, \+40% response correctness via ontology\-anchored hypergraph retrieval across 4 LLMsCited by:[§1](https://arxiv.org/html/2607.11948#S1.p2.1),[§2\.3](https://arxiv.org/html/2607.11948#S2.SS3.p1.1)\.
- The Microsoft AI Team \(2026\)MAI\-Thinking\-1: building a hill\-climbing machine\.Technical reportMicrosoft AI\.Note:Verified 2026\-06\-22; published 2026\-06\-02 by Microsoft AI; reports from\-scratch training without third\-party distillation, the position engaged in the dual\-use disclosureExternal Links:[Link](https://microsoft.ai/wp-content/uploads/2026/06/main_20260602_2.pdf)Cited by:[§1](https://arxiv.org/html/2607.11948#S1.SSx2.p2.1),[§6\.5](https://arxiv.org/html/2607.11948#S6.SS5.p1.1)\.
- L\. Tunstall, E\. Beeching, N\. Lambert, N\. Rajani, K\. Rasul, Y\. Belkada, S\. Huang, L\. von Werra, C\. Fourrier, N\. Habib, N\. Sarrazin, O\. Sanseviero, A\. M\. Rush, and T\. Wolf \(2023\)Zephyr: direct distillation of LM alignment\.arXiv preprint arXiv:2310\.16944\.Note:Verified 2026\-05\-26Cited by:[§2\.1](https://arxiv.org/html/2607.11948#S2.SS1.p1.1)\.
- C\. Wang, J\. Yan, Y\. Yue, and J\. Huang \(2025\)DistilQwen2\.5: industrial practices of training distilled open lightweight language models\.arXiv preprint arXiv:2504\.15027\.Note:Verified 2026\-06\-22 against arXiv:2504\.15027Cited by:[§2\.1](https://arxiv.org/html/2607.11948#S2.SS1.p1.1)\.
- A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Wang, B\. Zheng, C\. Yu,et al\.\(2025\)Qwen2\.5 technical report\.arXiv preprint arXiv:2412\.15115\.Note:Qwen Team, Alibaba GroupCited by:[§1](https://arxiv.org/html/2607.11948#S1.p3.1),[§3\.1](https://arxiv.org/html/2607.11948#S3.SS1.p1.1)\.Similar Articles
Auditing Question-Order Effects in Large Language Models with the QQ Equality: Mechanism Characterization and a Saturation Caveat
This paper develops the QQ equality from cognitive science into an audit criterion for LLMs, characterizes theoretical mechanism classes, and empirically tests on an open-weight instruction-tuned model, finding that saturation (near-deterministic responses) prevents distribution-level audits.
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
DocScope is a new benchmark for evaluating the verifiable reasoning and trustworthiness of Multimodal Large Language Models on long documents, introducing a four-stage evaluation protocol for page localization, region grounding, fact extraction, and answer verification.
EDGE-OPD: Internalizing Privileged Context with Evidence Guided On-Policy Distillation
This paper introduces EDGE-OPD, a modification of on-policy self-distillation for LLMs that uses guided rollouts and evidence masks to internalize privileged context without degrading general capabilities, showing success in rare-token identity settings.
How Language Models Fail: Token-Level Signatures of Committed and Persistent Reasoning Failures
This paper characterizes two distinct processes by which language models fail in reasoning—committed failure and persistent uncertainty—using token-level uncertainty signals, and demonstrates implications for self-consistency and failure detection strategies.
Adversarial Social Epistemology for Assemblies of Humans and Large Language Models
This paper proposes an adversarial social epistemology framework for analyzing trust, deception, and inference chains in communicative landscapes involving humans and large language models, and outlines mechanisms for auditing trust breaches.