From Token Probabilities to Semantic Constraints: Towards Declarative Probabilistic Evaluation of Language Models

arXiv cs.CL Papers

Summary

The paper proposes ModelLog, a declarative probabilistic framework for evaluating language models by defining semantic constraints over token predictions, linking evaluation to learning through shared semantics.

arXiv:2609.13520v1 Announce Type: new Abstract: While Large Language Models have improved rapidly, many fundamental questions remain about how to evaluate the knowledge and reasoning abilities they acquire, and how such evaluations relate to the learning signals used in pre-training. In this paper, we propose ModelLog, a declarative probabilistic framework for pre-training evaluation that makes the semantic structure of model behavior explicit and provides new formal tools for relating evaluation to learning. ModelLog specifies evaluation targets as symbolic constraints over token-level predictions and measures how strongly a model's distribution satisfies those constraints. We explore the framework through a new suite of tasks targeting negation, mutual exclusivity, and consistency, finding systematic failures that are difficult to characterize through token likelihood or answer accuracy alone. We further show that these evaluation scores can also be interpreted as losses, whose gradients reflect logical strength, informativeness, and variable-level sensitivity. This links evaluation and learning through a shared semantics, suggesting evaluation methods that diagnose model behavior while also helping to clarify the semantic structure of learning.
Original Article
View Cached Full Text

Cached at: 09/15/26, 08:34 AM

# From Token Probabilities to Semantic Constraints: Towards Declarative Probabilistic Evaluation of Language Models
Source: [https://arxiv.org/html/2609.13520](https://arxiv.org/html/2609.13520)
Kyle Richardson Cullen Anderson Pranav Balakrishnan Takuto BanDaksha LadiaAffiliation:Allen Institute for AIAffiliation:University of Massachusetts Amherstkyler@allenai\.org,\{cyanderson,pranavbalakr,tban,dladia\}@umass\.eduAnkita GuptaAffiliation:University of Massachusetts Amherstkyler@allenai\.org,\{cyanderson,pranavbalakr,tban,dladia\}@umass\.eduMarisa HudspethAffiliation:University of Massachusetts Amherstkyler@allenai\.org,\{cyanderson,pranavbalakr,tban,dladia\}@umass\.edu

###### Abstract

While Large Language Models have improved rapidly, many fundamental questions remain about how to evaluate the knowledge and reasoning abilities they acquire, and how such evaluations relate to the learning signals used in pre\-training\. In this paper, we proposeModelLog, a declarative probabilistic framework for pre\-training evaluation that makes the semantic structure of model behavior explicit and provides new formal tools for relating evaluation to learning\.ModelLogspecifies evaluation targets as symbolic constraints over token\-level predictions and measures how strongly a model’s distribution satisfies those constraints\. We explore the framework through a new suite of tasks targeting negation, mutual exclusivity, and consistency, finding systematic failures that are difficult to characterize through token likelihood or answer accuracy alone\. We further show that these evaluation scores can also be interpreted as losses, whose gradients reflect logical strength, informativeness, and variable\-level sensitivity\. This links evaluation and learning through a shared semantics, suggesting evaluation methods that diagnose model behavior while also helping to clarify the semantic structure of learning\.

## 1Introduction

Modern pre\-training relies on a remarkably simple learning signal: cross\-entropy loss for next\-token prediction\. This objective has given rise to models with broad linguistic, factual, and reasoning capabilities, but at first glance it seems limited\. While it provides positive, local supervision over observed continuations, many other richer semantic relations remain implicit\. For example in Figure[1](https://arxiv.org/html/2609.13520#S1.F1), next\-token supervision might tell us that*alcohol*is a sensible completion to the generic statement*When driving, it isnotsafe to drink \_\_*, which expresses a broadly applicable factual or normative commitment\. However, the objective does not directly state that*alcohol*should be unlikely in the contradictory context*When driving, it is safe to drink \_\_*\. Nor does it explicitly encode relations among multiple token predictions: paraphrases should preserve the same commitments, while contexts that commit to incompatible states of affairs should induce divergent predictions\. This raises the natural question:*why is cross\-entropy such an effective signal for learning?*Relatedly,*can richer learning signals be developed that improve the robustness of cross\-entropy pre\-training?*

\(A\)LM Token PredictionsExample local next\-tokendecisionsfor a language modelt∼ℙℳθ​\(⋅\)t\\sim\\mathbb\{P\}\_\{\\mathcal\{M\}\_\{\\theta\}\}\(\\cdot\)\.When driving, it is not safe to drinkalcoholaIt is not true that while driving it is safe to drinkalcoholbWhen driving, it is safe to drinkalcoholcIt is true that while driving it is safe to drinkalcohold\(C\)Queries over formulas𝒯\\mathcal\{T\}
Evaluation as structured probabilistic queries over constraints𝒯i\\mathcal\{T\}\_\{i\}weighted byℙℳθ\\mathbb\{P\}\_\{\\mathcal\{M\}\_\{\\theta\}\}Soft satisfaction:ℙ⁡\(𝒯i,ℳθ\)\\mathbb\{P\}\(\\mathcal\{T\}\_\{i\};\\mathcal\{M\}\_\{\\theta\}\)Marginals:ℙ⁡\(CLOSE\\mathbb\{P\}\(vv∣𝒯i;ℳθ\)\\,\\mid\\mathcal\{T\}\_\{i\};\\mathcal\{M\}\_\{\\theta\}\)Gradients :∇θℙ​\(𝒯i,ℳθ\)\\nabla\_\{\\theta\}\\,\\mathbb\{P\}\(\\mathcal\{T\}\_\{i\};\\mathcal\{M\}\_\{\\theta\}\)\(B\)Declarative Constraints𝒯\\mathcal\{T\}Boolean formulas specifying ideal token decisions for any modelℳ⁡\(⋅\)\\mathcal\{M\}\(\\cdot\)\.Likely tokens𝒯p\\mathcal\{T\}\_\{p\}
Modelsℳ\\mathcal\{M\}should prefer local factual token decisions\.ℳ⁡\(CLOSE\\mathcal\{M\}\(a\)\\scriptstyle\)∧\\landℳ⁡\(CLOSE\\mathcal\{M\}\(b\)\\scriptstyle\)Unlikely tokens𝒯n\\mathcal\{T\}\_\{n\}
Modelsℳ\\mathcal\{M\}should disprefer non\-factual local tokens decisions\.¬ℳ⁡\(CLOSE\\neg\\mathcal\{M\}\(c\)\\scriptstyle\)∧\\land¬ℳ⁡\(CLOSE\\neg\\mathcal\{M\}\(d\)\\scriptstyle\)Paraphrase Symmetries𝒯s\\mathcal\{T\}\_\{s\}
Token decisions involving paraphrases should be compatible
ℳ⁡\(CLOSE\\mathcal\{M\}\(aOPEN\)↔ℳ⁡\(CLOSE\)\\text\{\{\\color\[rgb\]\{0\.66,0\.13,0\.24\}$\\leftrightarrow$\}\}\\mathcal\{M\}\(bOPEN\)∧\)\\,\\text\{\{\\color\[rgb\]\{0\.66,0\.13,0\.24\}$\\land$\}\}\\,ℳ⁡\(CLOSE\\quad\\mathcal\{M\}\(cOPEN\)↔ℳ⁡\(CLOSE\)\\text\{\{\\color\[rgb\]\{0\.66,0\.13,0\.24\}$\\leftrightarrow$\}\}\\mathcal\{M\}\(d\)\)Incompatibilities𝒯m\\mathcal\{T\}\_\{m\}
Contradictory token predictions should be distinct and mutually exclusive\.\(ℳ⁡\(a\)∧¬ℳ⁡\(b\)\)∨\(\\mathcal\{M\}\(\\hbox\{\\pagecolor\{red\!50\}\\color\[rgb\]\{1,1,1\}a\}\)\\,\\text\{\{\\color\[rgb\]\{0\.66,0\.13,0\.24\}$\\land$\}\}\\,\\neg\\mathcal\{M\}\(\\hbox\{\\pagecolor\{blue\!50\}\\color\[rgb\]\{1,1,1\}b\}\)\)\\,\\text\{\{\\color\[rgb\]\{0\.66,0\.13,0\.24\}$\\lor$\}\}\\,\(¬ℳ⁡\(a\)∧ℳ⁡\(b\)\)\\quad\(\\neg\\mathcal\{M\}\(\\hbox\{\\pagecolor\{red\!50\}\\color\[rgb\]\{1,1,1\}a\}\)\\,\\text\{\{\\color\[rgb\]\{0\.66,0\.13,0\.24\}$\\land$\}\}\\,\\mathcal\{M\}\(\\hbox\{\\pagecolor\{blue\!50\}\\color\[rgb\]\{1,1,1\}b\}\)\)ℙ⁡\(ℳ⁡\(⋅\)\)∼\\mathbb\{P\}\(\\mathcal\{M\}\(\\cdot\)\)\\simℙℳθ​\(⋅\)\\mathbb\{P\}\_\{\\mathcal\{M\}\_\{\\theta\}\}\(\\cdot\)Figure 1:An illustration of our evaluation frameworkModelLogthat maps local next\-token predictions from a pre\-trained language model \(A\) into propositional variablesℳ⁡\(⋅\)\\mathcal\{M\}\(\\cdot\)denoting prediction events\. Symbolic constraints over these variables \(B\) specify the local semantic relations that should hold among them\. Evaluation then becomes probabilistic inference over constraints \(C\): model token probabilitiesℙℳθ​\(⋅\)\\mathbb\{P\}\_\{\\mathcal\{M\}\_\{\\theta\}\}\(\\cdot\)provide weights over assignmentsℙ⁡\(ℳ⁡\(⋅\)\)\\mathbb\{P\}\(\\mathcal\{M\}\(\\cdot\)\),satisfactionormarginalprobability measures how strongly the model supports the desired semantic structure, and thegradientsof such inferences provide a direct link to learning\.In this paper, we approach these broad questions*solely*through the lens of pre\-training evaluation, focusing in particular on how toformalizeand precisely measure the degree to whichpre\-trainedmodels satisfy semantic constraints\. As illustrated in Figure[1](https://arxiv.org/html/2609.13520#S1.F1), we develop a framework calledModelLogwhere we express the target of an evaluation as an explicitdeclarative theory: a set of logical formulas specifying the relations that should ideally hold among a model’s local token\-level predictions\. For example, to express that*alcohol*is compatible with both the prefix*When driving, it is not safe to drink \_\_*and its paraphrased form*It is not true that when driving it is safe to drink \_\_*, we uselogical variablessuch asℳ⁡\(a\)\\mathcal\{M\}\(\\hbox\{\\pagecolor\{red\!50\}\{\\color\[rgb\]\{1,1,1\}a\}\}\)andℳ⁡\(b\)\\mathcal\{M\}\(\\hbox\{\\pagecolor\{red\!50\}\{\\color\[rgb\]\{1,1,1\}b\}\}\)to denote the corresponding prediction events\. The paraphrase relation between these contexts can then be encoded by the constraintℳ⁡\(a\)↔ℳ⁡\(b\)\\mathcal\{M\}\(\\hbox\{\\pagecolor\{red\!50\}\{\\color\[rgb\]\{1,1,1\}a\}\}\)\\text\{\{\\color\[rgb\]\{0\.66,0\.13,0\.24\}$\\leftrightarrow$\}\}\\mathcal\{M\}\(\\hbox\{\\pagecolor\{red\!50\}\{\\color\[rgb\]\{1,1,1\}b\}\}\), requiring that the two prediction events agree\. In this way, declarative theories let us move from isolated token likelihoods to explicit statements about the semantic relations that should hold among them\.

Questions about model behavior can then be formulated as structured queries over the theories themselves\. We focus onprobabilistic queries: given a model’s token probabilities, we ask how much probability mass the model assigns to interpretations thatsatisfy the constraints of a given theory\.This builds on work in statistical relational learning\([Getoor and Taskar, 2007](https://arxiv.org/html/2609.13520#bib.bib2);[De Raedt et al\., 2016](https://arxiv.org/html/2609.13520#bib.bib3)\)and neuro\-symbolic modeling\([Marra et al\., 2024](https://arxiv.org/html/2609.13520#bib.bib45);[Feldstein et al\., 2024](https://arxiv.org/html/2609.13520#bib.bib44)\), where logicalconstraintsare combined with probabilistic weights to reason about structured events\. In this way, probabilistic inference moves evaluation beyond binary correctness, measuring how strongly a model’s distribution supportsa targetsemantic structure\.

The main goal of this paper is to give a technical outline ofModelLogand show how such a framework can both clarify the semantics of evaluation and connect such semantics directly to learning\. As such, the paper is primarily methodological in nature and offers a mix of formal and empirical results\. On the formal side, we identify two interpretability challenges that arise when doing semantic reasoning over token probabilities\. First, raw next\-token probabilities measure how probability mass is divided among competing continuations, making them difficult to interpret directly as probabilities of semantic validity\. We therefore rescale them using information about the local prediction distribution, treating semantic event probability as membership in a model’s effective “live” region of plausible continuations\([Holtzman et al\., 2020](https://arxiv.org/html/2609.13520#bib.bib43);[Hewitt et al\., 2022](https://arxiv.org/html/2609.13520#bib.bib41)\)\. Second, raw constraint probabilities can overstate model competence, since some formulas are easy to satisfy by chance\. We therefore introduce a notion ofconstraint informativeness, which accounts for chance satisfaction and is incorporated into our main probabilistic evaluation metric\.

Since semantic satisfaction is differentiable in model probabilities, our evaluation scores can also be interpreted as losses\. Surprisingly, we show how constraint informativeness, introduced as a measurement for evaluation, appears directly in the corresponding gradients for learning\. This offers a formal perspective on why likelihood\-style next\-token objectives – which we show inModelLogcan be represented as maximally informative conjunctive formulas – can potentially induce strong learning signals and provides a basis for systematically comparing them to other candidate objectives\.

On the empirical side, we introduceProbCT\(ProbabilisticConsistencyTests\) a set of diagnostic tests inspired by prior work on consistency probing\([Elazar et al\., 2021](https://arxiv.org/html/2609.13520#bib.bib37);[Kassner and Schütze, 2020](https://arxiv.org/html/2609.13520#bib.bib36)\)\. We find that current pre\-trained models perform poorly across these tasks, suggesting that pre\-training does not reliably induce the structured semantic relations captured by our constraints\. More importantly,ProbCTillustrates how declarative theories and probabilistic queries can expose failures that are difficult to see from token likelihoods or accuracy alone\. This includes an intriguing inverse\-scaling\([McKenzie et al\., 2023](https://arxiv.org/html/2609.13520#bib.bib8)\)pattern we observe, where larger models systematically satisfy certain constraints less often than smaller ones\.

#### Contributions

In line with the special track on*New Missions for NLP Research*, we proposeModelLog, a new evaluation framework that connects language model evaluation with techniques from statistical relational learning and probabilistic logic\. Our main contribution is methodological: we show how declarative probabilistic evaluation can help to bring more semantic clarity to evaluation by treating evaluation data as compositional semantic objects\. We illustrate this methodology through a mixture of formal results and empirical case studies on a new diagnostic benchmark calledProbCT\.

## 2Related Work

#### Language model evaluation

Our work connects to loss\-based language model evaluation, which measures models using quantities such as perplexity\([Jelinek, 1980](https://arxiv.org/html/2609.13520#bib.bib19);[Goodman, 2001](https://arxiv.org/html/2609.13520#bib.bib26);[Bengio et al\., 2003](https://arxiv.org/html/2609.13520#bib.bib40);[Brown et al\., 2020](https://arxiv.org/html/2609.13520#bib.bib38);[Lozhkov et al\., 2025](https://arxiv.org/html/2609.13520#bib.bib30);[Groeneveld et al\., 2024](https://arxiv.org/html/2609.13520#bib.bib29);[Magnusson et al\., 2024](https://arxiv.org/html/2609.13520#bib.bib39)\), and post\-hoc behavioral evaluation, which tests models on diagnostic tasks and benchmarks\([Srivastava et al\., 2023](https://arxiv.org/html/2609.13520#bib.bib24);[Wang et al\., 2018](https://arxiv.org/html/2609.13520#bib.bib22);[Petroni et al\., 2019](https://arxiv.org/html/2609.13520#bib.bib9);[Richardson et al\., 2020](https://arxiv.org/html/2609.13520#bib.bib13);[Ribeiro et al\., 2020](https://arxiv.org/html/2609.13520#bib.bib25);[Liang et al\., 2022](https://arxiv.org/html/2609.13520#bib.bib23),*inter alia*\)\.We specifically take inspiration from behavioral studies that use consistency as a core diagnostic of model behavior\([Elazar et al\., 2021](https://arxiv.org/html/2609.13520#bib.bib37);[Kassner and Schütze, 2020](https://arxiv.org/html/2609.13520#bib.bib36);[Liu et al\., 2025](https://arxiv.org/html/2609.13520#bib.bib31);[Jang et al\., 2022](https://arxiv.org/html/2609.13520#bib.bib33);[Novikova et al\., 2025](https://arxiv.org/html/2609.13520#bib.bib32)\)\.ModelLogsits between these traditions: it uses model probabilities while targeting interpretable semantic capabilities\. Unlike either, it evaluates rich declarative constraints over multiple predictions, yielding probabilistic, differentiable measurements that directly connect the semantics of evaluation with learning\.

\(A\) Token predictionsWhen driving, it is not safe to drinkalcoholaWhen driving, it is safe to drinkalcoholc\(B\) Constraint𝖥\\mathsf\{F\}‘Alcohol’ is a valid token completion in one sentence, but not in both sentences\.𝖥=\(ℳ⁡\(CLOSECLOSE\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}=\\big\(\\mathcal\{M\}\(aOPEN\)∧¬ℳ⁡\(CLOSE\)\\,\\text\{\{\\color\[rgb\]\{0\.66,0\.13,0\.24\}$\\land$\}\}\\,\\neg\\mathcal\{M\}\(cOPENOPEN\)\)\)\\big\)∨\(¬ℳ⁡\(CLOSECLOSE\\;\\text\{\{\\color\[rgb\]\{0\.66,0\.13,0\.24\}$\\lor$\}\}\\;\\big\(\\neg\\mathcal\{M\}\(aOPEN\)∧ℳ⁡\(CLOSE\)\\,\\text\{\{\\color\[rgb\]\{0\.66,0\.13,0\.24\}$\\land$\}\}\\,\\mathcal\{M\}\(cOPENOPEN\)\)\)\\big\)\(C\) Boolean semantics\(D\) Language model\(E\) Uninformed Priorℳθ\\mathcal\{M\}\_\{\\theta\}ℳ0\\mathcal\{M\}\_\{0\}ℳ⁡\(a\)\\mathcal\{M\}\(\\hbox\{\\pagecolor\{red\!50\}\\color\[rgb\]\{1,1,1\}a\}\)ℳ⁡\(c\)\\mathcal\{M\}\(\\hbox\{\\pagecolor\{blue\!50\}\\color\[rgb\]\{1,1,1\}c\}\)I∈𝖨⁡\(𝖥\)I\\in\\mathsf\{I\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)S⁡\(Ij,θ\)S\(I\_\{j\};\\theta\)uniformI1I\_\{1\}TTℙθ​\(ℳ⁡\(a\)\)⋅ℙθ​\(ℳ⁡\(c\)\)\\mathbb\{P\}\_\{\\theta\}\(\\mathcal\{M\}\(\\hbox\{\\pagecolor\{red\!50\}\\color\[rgb\]\{1,1,1\}a\}\)\)\\cdot\\mathbb\{P\}\_\{\\theta\}\(\\mathcal\{M\}\(\\hbox\{\\pagecolor\{blue\!50\}\\color\[rgb\]\{1,1,1\}c\}\)\)0\.25I2I\_\{2\}TF✓ℙθ​\(ℳ⁡\(a\)\)⋅\(1−ℙθ​\(ℳ⁡\(c\)\)\)\\mathbb\{P\}\_\{\\theta\}\(\\mathcal\{M\}\(\\hbox\{\\pagecolor\{red\!50\}\\color\[rgb\]\{1,1,1\}a\}\)\)\\cdot\(1\-\\mathbb\{P\}\_\{\\theta\}\(\\mathcal\{M\}\(\\hbox\{\\pagecolor\{blue\!50\}\\color\[rgb\]\{1,1,1\}c\}\)\)\)0\.25I3I\_\{3\}FT✓\(1−ℙθ​\(ℳ⁡\(a\)\)\)⋅ℙθ​\(ℳ⁡\(c\)\)\(1\-\\mathbb\{P\}\_\{\\theta\}\(\\mathcal\{M\}\(\\hbox\{\\pagecolor\{red\!50\}\\color\[rgb\]\{1,1,1\}a\}\)\)\)\\cdot\\mathbb\{P\}\_\{\\theta\}\(\\mathcal\{M\}\(\\hbox\{\\pagecolor\{blue\!50\}\\color\[rgb\]\{1,1,1\}c\}\)\)0\.25I4I\_\{4\}FF\(1−ℙθ​\(ℳ⁡\(a\)\)\)⋅\(1−ℙθ​\(ℳ⁡\(c\)\)\)\(1\-\\mathbb\{P\}\_\{\\theta\}\(\\mathcal\{M\}\(\\hbox\{\\pagecolor\{red\!50\}\\color\[rgb\]\{1,1,1\}a\}\)\)\)\\cdot\(1\-\\mathbb\{P\}\_\{\\theta\}\(\\mathcal\{M\}\(\\hbox\{\\pagecolor\{blue\!50\}\\color\[rgb\]\{1,1,1\}c\}\)\)\)0\.25WMC​\(𝖥,θ\)=S⁡\(I2\)\+S⁡\(I3\)\\qquad\\,\\,\\text\{WMC\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)=S\(I\_\{2\}\)\+S\(I\_\{3\}\)WMC0​\(𝖥\)=0\.5\\text\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)=0\.5ℙ^θ​\(ℳ⁡\(x,wj\)\)∼ℙℳθ​\(wj∣x<j\)\\hat\{\\mathbb\{P\}\}\_\{\\theta\}\(\\mathcal\{M\}\(x,\\hbox\{\\pagecolor\{red\!50\}\{\\color\[rgb\]\{1,1,1\}$w\_\{j\}$\}\}\)\)\\sim\\mathbb\{P\}\_\{\\mathcal\{M\}\_\{\\theta\}\}\(\\hbox\{\\pagecolor\{red\!50\}\{\\color\[rgb\]\{1,1,1\}$w\_\{j\}$\}\}\\mid x\_\{<\\,j\}\)\\qquad\\qquad\\,\\,ℙ0​\(ℳ⁡\(x,wj\)\)=0\.5\\mathbb\{P\}\_\{0\}\(\\mathcal\{M\}\(x,\\hbox\{\\pagecolor\{red\!50\}\{\\color\[rgb\]\{1,1,1\}$w\_\{j\}$\}\}\)\)=0\.5\\hskip 79\.6678ptFigure 2:Giventoken predictions\(A\) and aconstraint𝖥\\mathsf\{F\}over those predictions \(B\), evaluation reduces to weighted model counting \(WMC\) over interpretationsIiI\_\{i\}ofFF\(C\), weighted by model\-induced probabilities fromℳθ\\mathcal\{M\}\_\{\\theta\}\(D\)\. To make satisfaction scores interpretable, we also compare against anuninformed prior\(E\), which performs weighted model counting under uniform weights, yieldingWMC0​\(𝖥\)\\textbf\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\.
#### Neuro\-symbolic modeling

We employ techniques from statistical relational learning that combine logical constraints with probabilistic weights\([De Raedt and Kimmig, 2015](https://arxiv.org/html/2609.13520#bib.bib14);[Manhaeve et al\., 2018](https://arxiv.org/html/2609.13520#bib.bib16);[Li et al\., 2019](https://arxiv.org/html/2609.13520#bib.bib15);[Li et al\., 2023](https://arxiv.org/html/2609.13520#bib.bib6)\)and exact inference techniques\([Chavira and Darwiche, 2008](https://arxiv.org/html/2609.13520#bib.bib46);[Fierens et al\., 2015](https://arxiv.org/html/2609.13520#bib.bib42)\)\. Similar techniques are also used in constraint\-based neuro\-symbolic learning methods, notably semantic loss\([Xu et al\., 2018](https://arxiv.org/html/2609.13520#bib.bib47);[Ahmed et al\., 2023a](https://arxiv.org/html/2609.13520#bib.bib20);[Ahmed et al\., 2023b](https://arxiv.org/html/2609.13520#bib.bib4);[Richardson et al\., 2025](https://arxiv.org/html/2609.13520#bib.bib18);[Calanzone et al\., 2025](https://arxiv.org/html/2609.13520#bib.bib17)\)and semantic probabilistic layers\([Ahmed et al\., 2022](https://arxiv.org/html/2609.13520#bib.bib21)\),as well as other approaches that couple structured probabilistic inference with LLMs\([Dohan et al\., 2022](https://arxiv.org/html/2609.13520#bib.bib35);[Kassner et al\., 2023](https://arxiv.org/html/2609.13520#bib.bib34);[Lew et al\., 2023](https://arxiv.org/html/2609.13520#bib.bib28);[Cheng et al\., 2026](https://arxiv.org/html/2609.13520#bib.bib27);[Garg et al\., 2026](https://arxiv.org/html/2609.13520#bib.bib7);[Richardson et al\., 2026](https://arxiv.org/html/2609.13520#bib.bib1)\)\.Our focus differs: rather than using constraints primarily as training objectives or output\-layer structure for inference, we use them for LM evaluation and formal analysis\.

## 3TheModelLogFramework

In this section, we define the core concepts and notation underlying the evaluation frameworkModelLogillustrated in Figure[2](https://arxiv.org/html/2609.13520#S2.F2), starting with the five principles outlined below\. In §[3\.1](https://arxiv.org/html/2609.13520#S3.SS1), we then turn to the core technical obstacles that arise when applying our approach to pre\-trained language model evaluation:*token probability calibration*,*constraint probability rescaling*and*selection*\.

#### 1\. Token prediction events are symbolic objects\.

Letx=w1,…,wnx=w\_\{1\},\\ldots,w\_\{n\}be a token sequence andℙℳθ​\(wj∣x<j\)\\mathbb\{P\}\_\{\\mathcal\{M\}\_\{\\theta\}\}\(\\hbox\{\\pagecolor\{red\!50\}\{\\color\[rgb\]\{1,1,1\}$w\_\{j\}$\}\}\\mid x\_\{<\\,j\}\)be the next token probability assigned by an autoregressive modelℳθ\\mathcal\{M\}\_\{\\theta\}\. InModelLog, local model predictions are treated as symbolic propositions\. We writeℳ⁡\(x,wj\)\\mathcal\{M\}\(x,\\hbox\{\\pagecolor\{red\!50\}\{\\color\[rgb\]\{1,1,1\}$w\_\{j\}$\}\}\)to denote the proposition that for modelℳ\\mathcal\{M\}, the tokenwjw\_\{j\}is a valid completion at positionjjin the contextx<jx\_\{<j\}\. Intuitively, this proposition abstracts away from the raw token probability and asks if the model supportswjw\_\{j\}as a plausible continuation\.

#### 2\. Relations between predictions are formulas\.

Semantic evaluation requires reasoning not only about individual predictions, but about relations among predictions\. InModelLog, these relations are expressed as Boolean formulas𝖥\\mathsf\{F\}over prediction variablesℳ⁡\(⋅\)\\mathcal\{M\}\(\\cdot\)\(denoted below in short form as𝖯\\mathsf\{P\}\), using standard logical operators such as∧\\land,∨\\lor,→\\to,↔\\leftrightarrowand Booleans⊤/⊥\\top/\\bot\(true/false\)\. Thus, a formula𝖥\\mathsf\{F\}specifies the semantic structure that should hold among a set of local token predictions\.

#### 3\. Formulas are weighted by model probabilities\.

To evaluate a formula probabilistically, we assign each prediction variable a weight derived from the language model\. For a prediction eventℳ⁡\(x,wj\)\\mathcal\{M\}\(x,\\hbox\{\\pagecolor\{red\!50\}\{\\color\[rgb\]\{1,1,1\}$w\_\{j\}$\}\}\), this weight is based on the language model probabilityℙℳθ​\(wj∣x<j\)\\mathbb\{P\}\_\{\\mathcal\{M\}\_\{\\theta\}\}\(\\hbox\{\\pagecolor\{red\!50\}\{\\color\[rgb\]\{1,1,1\}$w\_\{j\}$\}\}\\mid x\_\{<j\}\)\. SinceModelLogrepresents prediction events as logical atoms, their semantics are naturally Bernoulli: in any interpretation, the event either holds or does not hold \(we examine this closely in §[3\.1](https://arxiv.org/html/2609.13520#S3.SS1)\)\. We write the resulting event probability asℙθ​\(ℳ​\(x,wj\)\)\\mathbb\{P\}\_\{\\theta\}\(\\mathcal\{M\}\(x,\\hbox\{\\pagecolor\{red\!50\}\{\\color\[rgb\]\{1,1,1\}$w\_\{j\}$\}\}\)\)\. In this way, a language model induces a probability distribution over truth assignments\.

#### 4\. Evaluation is probabilistic satisfaction\.

Given a formula𝖥\\mathsf\{F\}and probabilities for its variables,ModelLogevaluates the degree to which the model satisfies the formula by computingℙθ​\(𝖥\)\\mathbb\{P\}\_\{\\theta\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\), the probability that𝖥\\mathsf\{F\}holds under the model\-induced distribution over truth assignments\. Equivalently, this is the total probability mass assigned to interpretationsIIthat satisfy the declarative theory𝖥\\mathsf\{F\}\. We compute this quantity usingweighted model counting\([Chavira and Darwiche, 2008](https://arxiv.org/html/2609.13520#bib.bib46)\):

ℙθ​\(𝖥\)\\displaystyle\\mathbb\{P\}\_\{\\theta\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)=WMC​\(𝖥,θ\)\\displaystyle=\\text\{WMC\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\(1\):=∑I∈𝖨⁡\(𝖥\)∏𝖯:I⁡\(𝖯\)=Tℙθ\(𝖯\)⋅∏𝖯:I⁡\(𝖯\)=F\(1−ℙθ\(𝖯\)\)⏟probability​S​\(I,θ\)​of interpretation​I\\displaystyle:=\\sum\_\{I\\in\\mathsf\{I\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\}\\underbrace\{\\prod\_\{\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}:I\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\)=\\text\{T\}\}\\hskip\-8\.5359pt\\mathbb\{P\}\_\{\\theta\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\)\\cdot\\prod\_\{\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}:I\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\)=\\text\{F\}\}\\hskip\-8\.5359pt\\left\(1\-\\mathbb\{P\}\_\{\\theta\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\)\\right\)\}\_\{\\text\{probability \}S\(I;\\theta\)\\text\{ of interpretation \}I\}ℙθ​\(𝖥\)\\mathbb\{P\}\_\{\\theta\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)is then our main evaluation score: a graded measure of how strongly the model’s probabilities support the desired semantic constraint\.

#### 5\. Learning is tied to satisfaction gradients\.

Since satisfaction probabilities are differentiable functions of the model\-induced variable probabilities,ModelLogexposes a natural connectiontolearning\. For a formula𝖥\\mathsf\{F\}, we can define the correspondingsemantic lossfrom[Xu et al\. \(2018\)](https://arxiv.org/html/2609.13520#bib.bib47):

ℓsl​\(𝖥,θ\):=−log⁡ℙθ​\(𝖥\)\.\\displaystyle\\ell\_\{\\mathrm\{sl\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\):=\-\\log\\mathbb\{P\}\_\{\\theta\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\.\(2\)This loss penalizes the model when it assigns low probability mass to interpretations satisfying the declarative theory\. Thus, the same quantity used for evaluation can also be differentiated to ask how changes in local prediction probabilities would affect semantic satisfaction\. Later, we use this connection to show that semantic properties of the constraints that are relevant for evaluation appear in the gradients of the semantic loss∇ℓsl​\(𝖥,θ\)\\nabla\\ell\_\{\\text\{sl\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\.

###### Example 1\(Likelihood formula\)\.

An important special case is standard language\-model likelihood\. Given a token sequencex=w1,…,wnx=w\_\{1\},\\ldots,w\_\{n\}, define thelikelihood formulabelow as the conjunction of the observed token\-prediction events𝖥ℓ​\(x\)\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{\\ell\}\(x\)\.

‘$\\formula\_\{\\ell\}\(x\):=\\bigwedge\\limits\_\{j=1\}^\{n\}\\mathcal\{M\}\\big\(x,\\colorbox\{red\!50\}\{\\textcolor\{white\}\{$w\_\{j\}$\}\}\\big\)$''

Formula 1:A symbolic formula for likelihood\.Because𝖥ℓ​\(x\)\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{\\ell\}\(x\)is a pure conjunction, its satisfaction probability factors into the product of the probabilities of the observed token events\. This gives the standard likelihood objective as a special case \(see[Richardson et al\. \(2026\)](https://arxiv.org/html/2609.13520#bib.bib1)for a similar result\)\.

###### Proposition 1\(Likelihood special case\)\.

If each event probability is identified with the model’s next\-token probability,ℙθ​\(ℳ⁡\(x,wj\)\)=ℙℳθ​\(wj∣x<j\)\\mathbb\{P\}\_\{\\theta\}\(\\mathcal\{M\}\(x,\\hbox\{\\pagecolor\{red\!50\}\{\\color\[rgb\]\{1,1,1\}$w\_\{j\}$\}\}\)\)=\\mathbb\{P\}\_\{\\mathcal\{M\}\_\{\\theta\}\}\(\\hbox\{\\pagecolor\{red\!50\}\{\\color\[rgb\]\{1,1,1\}$w\_\{j\}$\}\}\\mid x\_\{<j\}\), then the satisfaction probability of𝖥ℓ​\(x\)\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{\\ell\}\(x\)is equal to the language\-model likelihood ofxx:

ℙθ​\(𝖥ℓ​\(x\)\)=∏j=1nℙℳθ​\(wj∣x<j\)\.\\mathbb\{P\}\_\{\\theta\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{\\ell\}\(x\)\)=\\prod\_\{j=1\}^\{n\}\\mathbb\{P\}\_\{\\mathcal\{M\}\_\{\\theta\}\}\(\\hbox\{\\pagecolor\{red\!50\}\{\\color\[rgb\]\{1,1,1\}$w\_\{j\}$\}\}\\mid x\_\{<j\}\)\.Consequently, the semantic lossℓsl​\(𝖥ℓ​\(x\),θ\)\\ell\_\{\\mathrm\{sl\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{\\ell\}\(x\);\\theta\)is the negative log\-likelihood, i\.e\., the \(unnormalized\)*cross\-entropy loss*ℓce​\(𝖥ℓ​\(x\),θ\)=−log⁡Pℳθ​\(x\)\\ell\_\{\\text\{ce\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{\\ell\}\(x\),\\theta\)=\-\\log P\_\{\\mathcal\{M\}\_\{\\theta\}\}\(x\)\.

Thus, ordinary next\-token training is recovered inModelLogin the case where thedeclarative theory is a conjunction of observed token events\.We note that other likelihood\-like losses, such as unlikelihood loss from[Welleck et al\. \(2019\)](https://arxiv.org/html/2609.13520#bib.bib48), can be expressed as a similar conjunctive formula extended with logical negation\.This restates thecentral puzzle of cross\-entropyin semantic terms:*why should optimizing satisfaction of this very particular kind of formula –a simple conjunction of prediction events– produce models that satisfy much richer semantic constraints?*Later, we argue that part of the answermay liein the inherent informativeness of likelihood formulas\. To make this idea precise, we will compare a model’s satisfaction probability against the probability of satisfying the same formula by chance\. This motivates anuninformed prior: the weighted model count of a formula𝖥\\mathsf\{F\}under uniform weights, where each prediction variable has probability0\.5:

WMC0​\(𝖥\):=∣𝖨⁡\(𝖥\)∣2∣vars⁡\(𝖥\)∣\.\\displaystyle\\text\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\):=\\frac\{\\mid\\mathsf\{I\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\\mid\}\{2^\{\\mid\\mathrm\{vars\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\\mid\}\}\.\(3\)Here,vars⁡\(𝖥\)\\mathrm\{vars\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)is the set ofatomicvariables appearing in𝖥\\mathsf\{F\}\. As illustrated in Figure[2](https://arxiv.org/html/2609.13520#S2.F2)\(E\), this quantity captures the baseline satisfiability of the formula independent of a model, and it becomes central to the issues we discuss next and our evaluation\.

### 3\.1Subtleties in pre\-training evaluation

The semantics above is deliberately close to standard neuro\-symbolic formulations\([Manhaeve et al\., 2018](https://arxiv.org/html/2609.13520#bib.bib16);[Xu et al\., 2018](https://arxiv.org/html/2609.13520#bib.bib47)\)that have been applied to language model training\([Ahmed et al\., 2023a](https://arxiv.org/html/2609.13520#bib.bib20);[Richardson et al\., 2025](https://arxiv.org/html/2609.13520#bib.bib18);[Calanzone et al\., 2025](https://arxiv.org/html/2609.13520#bib.bib17)\)\. Inthese settings, the main role of the probabilistic semantics is to provide a differentiable signal for learning;probabilities, such as the raw token likelihoods considered above, therefore need only define a useful training objective, not an interpretable or calibrated evaluation score\. InModelLog, our initial goal is measurement:ℙθ​\(𝖥\)\\mathbb\{P\}\_\{\\theta\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)is meant to quantify how strongly a pre\-trained model satisfies a semantic constraint\. This evaluation setting makes interpretability of the probabilities themselves essential, leading to the following issues involving calibration and constraint selection\.

#### What do token variables mean semantically?

Treating local token decisions as logical variables makes them Bernoulli\-like, i\.e\., either true/false in a given context\. Following[Richardson et al\. \(2025\)](https://arxiv.org/html/2609.13520#bib.bib18), we interpret truth as local validity, or whether a token is an acceptable continuation and its probability exceeds some acceptability thresholdϵ\\epsilon\. The difficulty is that language models output categorical distributions over the vocabulary, not Bernoulli probabilities over validity events\. Thus, raw token probabilities must be calibrated into event probabilities that reflect membership in the model’s locally valid region of continuations\.

Wecalibrate token event probabilitiesby using structural information about the local distribution to offset the odds of the raw token probability:

ℙ^θC​\(ℳ​\(x,wj\)\)\\displaystyle\\hat\{\\mathbb\{P\}\}\_\{\\theta\_\{C\}\}\\big\(\{\\mathcal\{M\}\(x,\\hbox\{\\pagecolor\{red\!50\}\{\\color\[rgb\]\{1,1,1\}$w\_\{j\}$\}\}\)\}\\big\)\\hskip\-2\.84544pt:=σ⁡\(logit​\(ℙℳθ​\(wj∣x\)\)\+C\)\\displaystyle:=\\sigma\\bigg\(\\text\{logit\}\\big\(\\mathbb\{P\}\_\{\\mathcal\{M\}\_\{\\theta\}\}\(\{\\hbox\{\\pagecolor\{red\!50\}\{\\color\[rgb\]\{1,1,1\}$w\_\{j\}$\}\}\\mid x\}\)\\big\)\+C\\bigg\)∝eC​ℙℳθ​\(wj∣x\)\.\\displaystyle\\propto e^\{C\}\\mathbb\{P\}\_\{\\mathcal\{M\}\_\{\\theta\}\}\(\{\\hbox\{\\pagecolor\{red\!50\}\{\\color\[rgb\]\{1,1,1\}$w\_\{j\}$\}\}\\mid x\}\)\.\(4\)Here,logit​\(⋅\)\\text\{logit\}\(\\cdot\)corresponds to the log odds of the raw token probability andCCis an offset term representing the size of the valid region of continuations\. Intuitively, larger values ofCCincrease the odds of a token being treated as valid, correcting for the fact that raw probability mass may be spread across many plausible continuations\. The calibrated value therefore estimates live\-region membership rather than raw next\-token likelihood\.

While different choices can be made forCC, we use the entropy of the top\-pptokens ornucleusof the local distribution\([Holtzman et al\., 2020](https://arxiv.org/html/2609.13520#bib.bib43)\), which we later refer to asnucleus entropy calibrationand discuss below through an example\.

###### Example 2\(Token calibration via nucleus entropy\)\.

For the prefixes*A computeris not/isan \_\_*, the completion*airplane*is semantically valid only in the negative context\. Yet underSmolLM2\-135M\([Lozhkov et al\., 2025](https://arxiv.org/html/2609.13520#bib.bib30)\), its raw probabilities are small in both contexts:0\.0055after*is not*and3\.2678e\-05after*is*\. Nucleus entropy calibration rescales these probabilities by the effective size of the local live region: entropies of3\.97and4\.11for the top\-pptokens \(with nucleusp=0\.95p=0\.95\) correspond to perplexities,eHe^\{H\}, of roughly53and61, estimating the number of plausible alternatives in the nucleus\. This yields calibrated event probabilities of0\.58and0\.001, recovering the intended contrast:*airplane*is plausible in the negative context but not the affirmative one\.

#### What do constraint probabilities tell us?

A second issue is that raw constraint probabilities conflate model behavior with formula structure:𝖯1∨𝖯2\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{1\}\\,\\text\{\{\\color\[rgb\]\{0\.66,0\.13,0\.24\}$\\lor$\}\}\\,\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{2\}can receive higher raw probability than𝖯1∧𝖯2\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{1\}\\,\\text\{\{\\color\[rgb\]\{0\.66,0\.13,0\.24\}$\\land$\}\}\\,\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{2\}because it admits more satisfying assignments\. In learning,analogous effects appear as*reasoning shortcuts*\([Van Krieken et al\., 2024](https://arxiv.org/html/2609.13520#bib.bib10);[Marconato et al\., 2023](https://arxiv.org/html/2609.13520#bib.bib12);[Marconato et al\., 2026](https://arxiv.org/html/2609.13520#bib.bib11)\),whichcreate spurious learning patterns; such issues arise in evaluation too\. We therefore compare model\-weighted satisfaction to the uninformed baselineWMC0​\(𝖥\)\\mathrm\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)from Eq\.[3](https://arxiv.org/html/2609.13520#S3.E3)and define our mainprobabilistic consistency metricρ\\rho:

ρ⁡\(𝖥,θC\)=WMC⁡\(𝖥,θC\)−WMC0​\(𝖥\)1−WMC0​\(𝖥\)\.\\displaystyle\\rho\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\_\{C\}\)=\\frac\{\\mathrm\{WMC\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\_\{C\}\)\-\\mathrm\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\}\{1\-\\mathrm\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\}\.\(5\)This measures thefraction ofpossible improvement over chance satisfactionachieved by the model:ρ=0\\rho=0matches the uninformed baseline,ρ=1\\rho=1is perfect satisfaction\. We also use a hard,accuracy\-like consistency metric,

CAcc\(𝖥;θC\)=\[WMC\(𝖥;θC\)\>WMC0\(𝖥\)\],\\displaystyle\\mathrm\{CAcc\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\_\{C\}\)=\\mathbbm\{1\}\\\!\\left\[\\mathrm\{WMC\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\_\{C\}\)\>\\mathrm\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\\right\],\(6\)which records whether the model satisfies the constraint better than chance\. Averaging this over examples gives a dataset\-level score, which we report later in §[5](https://arxiv.org/html/2609.13520#S5)and Table[2](https://arxiv.org/html/2609.13520#S3.T2)\.

#### How can we determine if a constraint is useful?

InModelLog, experiment designers must come up with specific constraints to test\. A natural question then is:*how can we know if a constraint is meaningful to use for testing?*Intuitively, one should prioritize constraints that are inherently informative and hence not subject to spurious satisfaction\. Using the uninformed baseline, we define the notion ofconstraint informativeness:I⁡\(𝖥\):=−log2⁡WMC0​\(𝖥\)I\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\):=\-\\log\_\{2\}\\mathrm\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\.

HereI⁡\(𝖥\)I\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)measures how many bits of information are gained by knowing that𝖥\\mathsf\{F\}holds under an otherwise uniform assignment\. As we discuss in the example below, this gives us a useful tool for reasoning about what to test and how it might connect to learning\.

###### Example 3\(Constraint informativeness\)\.

For a sequencex=w1,…,wnx=w\_\{1\},\\ldots,w\_\{n\}, the likelihood formula𝖥ℓ​\(x\)\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{\\ell\}\(x\)from Formula[1](https://arxiv.org/html/2609.13520#LST1)requires every observed token\-prediction eventℳ⁡\(x,wj\)\\mathcal\{M\}\(x,\{\\hbox\{\\pagecolor\{red\!50\}\{\\color\[rgb\]\{1,1,1\}$w\_\{j\}$\}\}\}\)to hold\. It is therefore a complete assignment, or minterm, and is maximally informative among satisfiable constraints over these variables \(see §[A](https://arxiv.org/html/2609.13520#A1)\)\.

###### Proposition 2\(Likelihood is a maximally informative minterm\)\.

Let𝖥ℓ​\(x\)\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{\\ell\}\(x\)be the likelihood formula for a text sequencexxovernnobserved token\-prediction events\. For any satisfiable formula𝖥⁡\(x\)\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\(x\)over the same variables:

I⁡\(𝖥⁡\(x\)\)≤I⁡\(𝖥ℓ​\(x\)\)=n\.I\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\(x\)\)\\leq I\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{\\ell\}\(x\)\)=n\.Moreover, equality holds iff𝖥⁡\(x\)\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\(x\)has exactly one satisfying interpretation\.

Thus, likelihood is maximally informativebecause it corresponds to a complete assignment of the observed token events\. Richer constraints, such as implications or biconditionals, may test different structure but typically admit more satisfying interpretations and are therefore less informative in this formal sense\. This reframes likelihood\-style next token prediction as optimizing satisfaction of a maximally informative positive constraint from the text\.

More generally, informativeness provides a way to reason about the logical strength of evaluation constraints\. If one formula*semantically entails*\(⊧\\models\) another, then satisfying the stronger formula also guarantees satisfaction of the weaker one\. This gives the following monotonicity property:

###### Proposition 3\(Monotonicity of satisfaction and informativeness\)\.

For any two formulas𝖥1\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{1\}and𝖥2\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{2\}over the same variables, if𝖥1⊧𝖥2\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{1\}\\models\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{2\}, then

ℙθ​\(𝖥1\)≤ℙθ​\(𝖥2\)andI⁡\(𝖥1\)≥I⁡\(𝖥2\)\.\\mathbb\{P\}\_\{\\theta\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{1\}\)\\leq\\mathbb\{P\}\_\{\\theta\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{2\}\)\\quad\\text\{and\}\\quad I\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{1\}\)\\geq I\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{2\}\)\.

This highlights an important difference between learning and evaluation\. For learning, highly informative constraints may be desirable because they can induce stronger losses,a topic we revisit in Section[6](https://arxiv.org/html/2609.13520#S6)\. For evaluation, however, the most informative constraint is not always the most useful diagnostic\. A very strong formula may simply fail, while weaker entailed constraints can reveal partial semantic knowledge and provide a more interpretable picture of model behavior\.As a simple but useful consequence, monotonicity can be used constructively: by weakening the likelihood conjunction, we obtain likelihood\-style upper bounds that provide principled new evaluation targets\.

###### Example 4\(Deriving novel likelihood bounds\)\.

Let⋁≥k\\bigvee\_\{\\geq k\}denote a threshold disjunction, true when at leastkkof its arguments are true:

‘$\\formula\_\{\\geqk\}\(x\):=\\sideset\{\}\{\_\{\\geqk\}\}\\bigvee\\limits\_\{j=1\}^\{n\}\\mathcal\{M\}\(x,\\colorbox\{red\!50\}\{\\textcolor\{white\}\{$w\_j$\}\}\)$''

Formula 2:A relaxed form of Formula[1](https://arxiv.org/html/2609.13520#LST1)\.Herek=1k=1gives ordinary disjunction \(∨\\lor\) andk=nk=ngives𝖥ℓ​\(x\)\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{\\ell\}\(x\), which is ordinary sequence likelihood under the raw weighting from Prop\.[1](https://arxiv.org/html/2609.13520#Thmprop1)\. Since𝖥≥n​\(x\)⊧⋯⊧𝖥≥1​\(x\)\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{\\geq n\}\(x\)\\models\\cdots\\models\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{\\geq 1\}\(x\), monotonicity gives a hierarchy of increasingly tight upper bounds:

ℙθ​\(𝖥≥1​\(x\)\)⏟∨≥ℙθ​\(𝖥≥2​\(x\)\)≥…≥ℙθ​\(𝖥≥n​\(x\)\)⏟∧\.\\underbrace\{\\mathbb\{P\}\_\{\\theta\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{\\geq 1\}\(x\)\)\}\_\{\\text\{\{\\color\[rgb\]\{0\.66,0\.13,0\.24\}$\\lor$\}\}\}\\geq\\mathbb\{P\}\_\{\\theta\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{\\geq 2\}\(x\)\)\\geq\.\.\.\\geq\\underbrace\{\\mathbb\{P\}\_\{\\theta\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{\\geq n\}\(x\)\)\}\_\{\\text\{\{\\color\[rgb\]\{0\.66,0\.13,0\.24\}$\\land$\}\}\}\.While simple, this shows how bounds that are not obvious from the usual product form of likelihood become immediate once likelihood is represented semantically as a logical conjunction\. This relates in spirit to selective\-modeling objectives that train on subsets of tokens\([Lin et al\., 2024](https://arxiv.org/html/2609.13520#bib.bib5)\), and shows how such relaxations can be derived semantically\.

Task SubsetConstraint𝖥\\mathsf\{F\}I⁡\(𝖥\)≈I\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\\approxExampleTransitivity\(ℳ⁡\(x1,y1\)∧ℳ⁡\(x2,y2\)\)\(\\mathcal\{M\}\(x\_\{1\},\\hbox\{\\pagecolor\{red\!50\}\\color\[rgb\]\{1,1,1\}$y\_\{1\}$\}\)\\land\\mathcal\{M\}\(x\_\{2\},\\hbox\{\\pagecolor\{blue\!50\}\\color\[rgb\]\{1,1,1\}$y\_\{2\}$\}\)\)→ℳ⁡\(x1,y2\)\\text\{\{\\color\[rgb\]\{0\.66,0\.13,0\.24\}$\\to$\}\}\\mathcal\{M\}\(x\_\{1\},\\hbox\{\\pagecolor\{blue\!50\}\\color\[rgb\]\{1,1,1\}$y\_\{2\}$\}\)0\.193A Golden Retriever is adog∧\\landEvery dog is amammal→\\toA Golden Retriever is amammalForward Implicationℳ⁡\(x,y1\)→ℳ⁡\(x,y2\)\\mathcal\{M\}\(x,\\hbox\{\\pagecolor\{red\!50\}\\color\[rgb\]\{1,1,1\}$y\_\{1\}$\}\)\\text\{\{\\color\[rgb\]\{0\.66,0\.13,0\.24\}$\\to$\}\}\\mathcal\{M\}\(x,\\hbox\{\\pagecolor\{blue\!50\}\\color\[rgb\]\{1,1,1\}$y\_\{2\}$\}\)0\.415A Golden Retriever is adog→\\toA Golden Retriever is amammalNegation Consistencyℳ⁡\(x,y\)⊕ℳ⁡\(¬x,y\)\\mathcal\{M\}\(x,\\hbox\{\\pagecolor\{red\!50\}\\color\[rgb\]\{1,1,1\}$y$\}\)\\oplus\\mathcal\{M\}\(\\neg x,\\hbox\{\\pagecolor\{red\!50\}\\color\[rgb\]\{1,1,1\}$y$\}\)1\.0A Golden Retriever is adog⊕\\oplusA Golden Retriever is not adogEntityMutual Exclusivityℳ⁡\(x,y1\)⊕ℳ⁡\(x,y2\)\\mathcal\{M\}\(x,\\hbox\{\\pagecolor\{red\!50\}\\color\[rgb\]\{1,1,1\}$y\_\{1\}$\}\)\\oplus\\mathcal\{M\}\(x,\\hbox\{\\pagecolor\{blue\!50\}\\color\[rgb\]\{1,1,1\}$y\_\{2\}$\}\)1\.0A Golden Retriever is amammal⊕\\oplusA Golden Retriever is areptileSpatialMutual ExclusivityParis is inGermany⊕\\oplusParis is inFranceTable 1:A description of theProbCTdiagnostic benchmark in terms of its fivetask subsets, the structure of theconstraints𝖥\\mathsf\{F\}being tested with their informativenessI⁡\(𝖥\)I\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)and anexampleset of prefixes\.Table 2:Main results \(CAcc%\) across the different reasoning tasks inProbCTand model families, averaged over different nucleus valuesp∈\{0\.8,0\.85,0\.9,0\.95\}p\\in\\\{0\.8,0\.85,0\.9,0\.95\\\}used for calibration\.

## 4Testing the framework

To evaluateModelLogempirically, we introduceProbCT\(ProbabilisticConsistencyTests\), a diagnostic suite described in §[4\.1](https://arxiv.org/html/2609.13520#S4.SS1)\. We then describe our experimental setup in §[4\.2](https://arxiv.org/html/2609.13520#S4.SS2)111Our code and data are available at[https://github\.com/cullena20/ModelLog](https://github.com/cullena20/ModelLog)\.\.

### 4\.1A case study onProbCT

Following prior work on knowledge probing\([Petroni et al\., 2019](https://arxiv.org/html/2609.13520#bib.bib9);[Kassner and Schütze, 2020](https://arxiv.org/html/2609.13520#bib.bib36);[Elazar et al\., 2021](https://arxiv.org/html/2609.13520#bib.bib37)\), we automatically generate tasks from factual knowledge graphs \(§[B\.1](https://arxiv.org/html/2609.13520#A2.SS1)\)\. As shown in Table[1](https://arxiv.org/html/2609.13520#S3.T1),ProbCTcontains five task subsets, inspired by prior work, that cover constraints with different logical structures and levels of informativeness:transitivity\([Li et al\., 2019](https://arxiv.org/html/2609.13520#bib.bib15)\),forward implication\([Richardson et al\., 2020](https://arxiv.org/html/2609.13520#bib.bib13)\),negation consistency\([Kassner and Schütze, 2020](https://arxiv.org/html/2609.13520#bib.bib36)\), and two forms ofmutual exclusivityinvolving exclusive\-or \(⊕\\oplus\) reasoning\([Xu et al\., 2018](https://arxiv.org/html/2609.13520#bib.bib47)\)\.

We applyLLM\-as\-a\-judge stylefiltering withGPT\-5o\-minito improve correctness and naturalness of the generated data \(§[B\.2](https://arxiv.org/html/2609.13520#A2.SS2)\),yielding around500\-1000examples per subset\.As a check on quality, a subset of the authors validated 50 randomly sampled examples, finding overallhigh quality and acceptance\(§[B\.4](https://arxiv.org/html/2609.13520#A2.SS4)\)\. Finally, we restrict examples to single\-token substitutions, producing the model\-specificsplits further detailed in §[B\.3](https://arxiv.org/html/2609.13520#A2.SS3)\.

### 4\.2Models and Experimental Setup

We evaluate the following pre\-trained models:Gemma\-3\([Team et al\., 2025](https://arxiv.org/html/2609.13520#bib.bib51)\),Llama\-3\([Grattafiori et al\., 2024](https://arxiv.org/html/2609.13520#bib.bib49)\), andQwen3\([Yang et al\., 2025](https://arxiv.org/html/2609.13520#bib.bib50)\),with parameter sizes from270Mto14B\. For each model andProbCTsubset, we report average constraint accuracy CAcc \(Eq\.[6](https://arxiv.org/html/2609.13520#S3.E6)\) and per\-example probabilistic consistencyρ\\rho\(Eq\.[5](https://arxiv.org/html/2609.13520#S3.E5)\)\. Unless noted, all results use nucleus\-entropy calibration \(§[3\.1](https://arxiv.org/html/2609.13520#S3.SS1)\); for CAcc, we report mean performance across different nucleus thresholdsppwith standard deviations\.

## 5Results and Findings

Our main results onProbCTare reported in Table[2](https://arxiv.org/html/2609.13520#S3.T2)and Figure[3](https://arxiv.org/html/2609.13520#S5.F3)\(see also §[F](https://arxiv.org/html/2609.13520#A6)\)\.

Gemma\-3\-12B\-pt Llama\-3\.1\-8B Qwen3\-14BTransitivityForward ImplicationNegation ConsistencyMutual ExclusivitySpatial Exclusivity

Figure 3:Probabilistic Consistency score histogramsρ⁡\(⋅,θC\)\\rho\(\\cdot,\\theta\_\{C\}\)for our largest models in each model family, using nucleus valuep=0\.95p=0\.95\. Points highlighted in red are inconsistent, while points highlighted in blue are consistent\. The solid vertical black line corresponds to the average probabilistic consistency \(see also §[F](https://arxiv.org/html/2609.13520#A6)\)\.#### Models often fail declarative consistency tests\.

Across model families, performance is low on severalProbCTsubsets, especially onnegation consistencyandspatial exclusivity\. For example, in the latter case, the largestGemma\-3andQwen3models improve over the uninformed baseline on only23\.8% and12\.6% of examples, respectively\. Figure[3](https://arxiv.org/html/2609.13520#S5.F3)shows that these failures are not only aggregate effects: many examples have negative probabilistic consistency scoresρ\\rho, meaning that the model supports the constraint less than the uninformed baseline\. Thus,ModelLogreveals failures not simply in individual token predictions, but in the semantic relations among predictions\.

Table 3:Inverse\-scaling\([McKenzie et al\., 2023](https://arxiv.org/html/2609.13520#bib.bib8)\)behaviorfor the forward implication:𝖯1​→𝖯2\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{1\}\\text\{\{\\color\[rgb\]\{0\.66,0\.13,0\.24\}$\\to$\}\}\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{2\}forGemma\-3\.ℙθC​\(𝖯2\)−ℙθC​\(𝖯1\)\\mathbb\{P\}\_\{\\theta\_\{C\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{2\}\)\{\-\}\\mathbb\{P\}\_\{\\theta\_\{C\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{1\}\)is the mean difference between conclusion and premise probabilities\.ℙθC​\(𝖯2\>𝖯1\)\\mathbb\{P\}\_\{\\theta\_\{C\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{2\}\{\>\}\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{1\}\)is the percentage of points whereℙθC​\(𝖯2\)\>ℙθC​\(𝖯1\)\\mathbb\{P\}\_\{\\theta\_\{C\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{2\}\)\>\\mathbb\{P\}\_\{\\theta\_\{C\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{1\}\)\.See §[C](https://arxiv.org/html/2609.13520#A3)for similar patterns across our full set of models\.
#### Consistency can inverse\-scale\.

A more surprising pattern is that larger models are not always more consistent\. Ontransitivityandforward implication, CAcc often decreases with model scale across all three families\. Further analysis suggests that this reflects systematic changes in how models distribute probability across premise and conclusion facts \(see an example in Table[8](https://arxiv.org/html/2609.13520#A3.T8)\)\. For implication constraints𝖯1​→𝖯2\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{1\}\\text\{\{\\color\[rgb\]\{0\.66,0\.13,0\.24\}$\\to$\}\}\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{2\}, where𝖯1\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{1\}is a more specific fact and𝖯2\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{2\}a broader consequence, smaller models are more likely to assign higher probability to𝖯2\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{2\}than to𝖯1\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{1\}\. As shown in Table[3](https://arxiv.org/html/2609.13520#S5.T3), this trend reverses with scale: for example,ℙθC​\(𝖯2\)−ℙθC​\(𝖯1\)\\mathbb\{P\}\_\{\\theta\_\{C\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{2\}\)\-\\mathbb\{P\}\_\{\\theta\_\{C\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{1\}\)drops from0\.17to\-0\.24acrossGemma\-3\. Similar patterns are found for our other models \(Table[6](https://arxiv.org/html/2609.13520#A3.T6)\) and for thetransitivityrule \(Table[7](https://arxiv.org/html/2609.13520#A3.T7)\)\.

Importantly, this should not be read as showing that larger models are worse reasoners overall\. Larger models may distribute probability differently across logically related predictions, sometimes reducing consistency under one constraint while improving other aspects of robustness\.ModelLogmakes such shifts measurable, providing a framework for testing how model scale changes the semantic structure of probabilistic predictions\.

#### Findings are stable across calibration choices\.

Results are stable across nucleus thresholds, suggesting that the observed patterns are not artifacts of calibration \(for results across different calibration methods, see §[E](https://arxiv.org/html/2609.13520#A5)\)\. Taken together, these case studies show thatModelLogcan reveal systematic failures, inverse\-scaling patterns, and qualitative changes in model behavior that are difficult to capture with standard evaluation metrics\.

## 6Discussion and Future Directions

We began with the puzzle of cross\-entropy:*why does a local next\-token objective provide such an effective learning signal, and can richer semantic losses improve pre\-training?*Although we focused on evaluation and the semantics of model constraints,ModelLogdefines evaluation scores as differentiable satisfaction probabilities, connected through the semantic lossℓsl​\(𝖥,θ\)\\ell\_\{\\text\{sl\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\(Eq\.[2](https://arxiv.org/html/2609.13520#S3.E2)\)\. We end by returning to this learning question, showing that concepts introduced for evaluation, such as the uninformed baselineWMC0​\(𝖥\)\\mathrm\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)and constraint informativenessI⁡\(𝖥\)I\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\), reappear directly in the gradients of semantic loss∇ℓsl​\(𝖥,θ\)\\nabla\\ell\_\{\\text\{sl\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\. This suggests that they are not only diagnostic quantities, but also potential tools for studying how the evaluation failures observed above might be turned into more effective training strategies\.

#### Informativeness shapes semantic\-loss gradients\.

The first connection is that constraint informativeness contributes to the potential scale of the semantic\-loss gradient\. Since gradients are vector\-valued, we measure this scale standardly using a norm\. Highly informative constraints are hard to satisfy by chance, so failures on such constraints can receive larger loss amplification\. More precisely, we have the following bound \(see proof in §[A](https://arxiv.org/html/2609.13520#A1)\):

###### Proposition 4\(Gradient scaling by informativeness\)\.

For any formula𝖥\\mathsf\{F\}such thatWMC⁡\(𝖥,θ\)≥WMC0​\(𝖥\)\\mathrm\{WMC\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\\geq\\mathrm\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\),

‖∇θℓsl​\(𝖥,θ\)‖≤2I⁡\(𝖥\)​‖∇θWMC​\(𝖥,θ\)‖\.\\left\\\|\\nabla\_\{\\theta\}\\ell\_\{\\mathrm\{sl\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\\right\\\|\\leq 2^\{I\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\}\\left\\\|\\nabla\_\{\\theta\}\\mathrm\{WMC\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\\right\\\|\.

This bound shows that informativeness appears as an exponential amplification term in the possible learning signal induced by a constraint\. It does not imply that more informative constraints always produce larger or better updates: the actual gradient can be reweighted externally and also depends on the local derivative of the satisfaction probability, which may point in different directions across noisy examples\. Still, the result suggests that highly informative constraints, such as the likelihood constraint \(Formula[1](https://arxiv.org/html/2609.13520#LST1)\), can provide strong learning signals when their gradients are coherent\. This may therefore help explain part of their success\.

#### Informativeness appears directly in the gradient\.

The previous bound controls gradient scale; the next identity decomposes the gradient into interpretable factors\. First, we define the following:

A⁡\(𝖥,θ\)=WMC0​\(𝖥\)WMC⁡\(𝖥,θ\)\.A\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)=\\frac\{\\mathrm\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\}\{\\mathrm\{WMC\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\}\.This compares the model’s satisfaction probability to the uninformed baseline:A⁡\(𝖥,θ\)<1A\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)<1means that the model satisfies the constraint better than chance, whileA⁡\(𝖥,θ\)\>1A\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\>1means that it does worse\. Based on this, the semantic\-loss gradient then factors as follows:

###### Proposition 5\(Semantic\-loss gradient identity\)\.

For any formula𝖥\\mathsf\{F\},

∇θℓsl​\(𝖥,θ\)=−A⁡\(𝖥,θ\)​2I⁡\(𝖥\)​∇θWMC​\(𝖥,θ\)\.\\nabla\_\{\\theta\}\\ell\_\{\\mathrm\{sl\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)=\-A\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\\,2^\{I\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\}\\,\\nabla\_\{\\theta\}\\mathrm\{WMC\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\.

Thus, the gradient decomposes into three pieces: a fit termA⁡\(𝖥,θ\)A\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\), an intrinsic informativeness term2I⁡\(𝖥\)2^\{I\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\}, and a local direction term∇θWMC​\(𝖥,θ\)\\nabla\_\{\\theta\}\\mathrm\{WMC\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\. This gives a more nuanced view of semantic learning signals: informativeness appears as a structural factor, whileA⁡\(𝖥,θ\)A\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)reflects how much the model currently under\- or over\-satisfies the constraint relative to the uninformed baseline\. The actual update direction then remains determined by the local derivative∇θWMC​\(𝖥,θ\)\\nabla\_\{\\theta\}\\mathrm\{WMC\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\.

The informativeness term again helps explain why likelihood\-style formulas can induce strong learning signals\. More importantly, by separating constraint strength from model fit and local update direction, the identity makes the latter terms diagnostic of whether a constraint provides useful learning evidence\.

#### Putting the pieces together\.

These results give a partial answer to the cross\-entropy puzzle\. Likelihood is recovered inModelLogas a conjunction of observed token\-prediction events whose satisfaction probability is sequence likelihood \(Prop\.[1](https://arxiv.org/html/2609.13520#Thmprop1)\)\. This formula is maximally informative \(Prop\.[2](https://arxiv.org/html/2609.13520#Thmprop2)\), and the gradient results show that informativeness can amplify the semantic\-loss signal\. Thus, likelihood\-style objectives may be effective because they optimize highly informative constraints supplied by text, even though they do not directly encode richer semantic relations\.

At the same time, these results complicate a simple training story in which semantic failures are fixed by replacing likelihood with weaker constraint\-based losses: since likelihood\-style constraints are already maximally informative, such replacements may not provide stronger pre\-training signals\. Richer constraints may instead be most useful for identifying which examples or constraint types provide useful learning evidence\. Future work can use the gradient decomposition above to study when such constraints provide coherent learning signals, connecting evaluation failures to questions of constraint informativeness, data coverage, selection, and weighting\.

## 7Conclusion

We introducedModelLog, a declarative probabilistic framework for evaluating pre\-trained language models through logical constraints over token\-level predictions\. Using exact probabilistic inference,ModelLogmeasures graded satisfaction of semantic properties that are difficult to capture with likelihood or accuracy alone\. ThroughProbCT, we showed how this framework can reveal consistency failures in current models, and our formal analysis characterized the structure and informativeness of evaluation constraints, including their connection to semantic\-loss gradients\. This suggests a broader role for evaluation as a methodology for making model behavior semantically measurable while also informing more interpretable learning objectives\.

## Limitations

This work has three main limitations\. First,ModelLogcurrently focuses on single\-token prediction events and evaluates constraints under a factorized product model over local prediction events\. This is a deliberate semantic abstraction: the events correspond to separate local queries to the language model, while dependencies among them are introduced declaratively by the constraints rather than assumed to be part of the model’s autoregressive distribution\. This keeps inference tractable and enables clean formal analysis, but the current framework does not yet address multi\-token spans, variable\-length paraphrases, or full sequence\-level constraints\. Second, our empirical results are intended as small case studies of the framework and are based on synthetic diagnostic tasks generated from knowledge\-graph relations, which, despite our efforts at filtering, may still contain noise\. These tasks provide controlled tests of declarative constraints, but they do not capture the full complexity of natural language evaluation\. Third, while our formal analysis connects probabilistic evaluation to semantic\-loss gradients, we do not train models with these losses at scale\. The learning discussion should therefore be read as a formal and diagnostic contribution, with actual training left for future work\.

## Acknowledgements

We thank Kareem Ahmed, Gregor Betz, Yanai Elazar, Poorva Garg, Ronan Le Bras, William Merrill, Jackson Petty and Sahil Sidheekh for useful feedback at different stages of this work\. We also thank Andrew McCallum and the Manning College of Information & Computer Sciences at the University of Massachusetts Amherst for their support of this project, as well as the UMass Unity Cluster \([www\.umass\.edu/research\-computing/unity\-research\-computing\-platform](https://www.umass.edu/research-computing/unity-research-computing-platform)\) for providing GPU compute\. We acknowledge that bothChatGPTandClaudewere used to improve some of the writing and overall presentation, and to provide feedback on some of the technical results\. All mistakes remain our own\.

## References

- K\. Ahmed, K\. Chang, and G\. Van den BroeckA pseudo\-semantic loss for autoregressive models with logical constraints\.Proceeding of NeurIPS\.External Links:[Link](https://neurips.cc/virtual/2023/poster/70782)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px2.p1.1),[§3\.1](https://arxiv.org/html/2609.13520#S3.SS1.p1.1)\.
- Ahmedet al\.\(2022\)K\. Ahmed, S\. Teso, K\. Chang, G\. Van den Broeck, and A\. VergariSemantic probabilistic layers for neuro\-symbolic learning\.Proceedings of NeurIPS\.External Links:[Link](https://arxiv.org/abs/2206.00426)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px2.p1.1)\.
- Ahmedet al\.\(2023b\)K\. Ahmed, S\. Teso, P\. Morettin, L\. Di Liello, P\. Ardino, J\. Gobbi, Y\. Liang, E\. Wang, K\. Chang, A\. Passerini,et al\.Semantic loss functions for neuro\-symbolic structured prediction\.Proceedings of ICLR\.External Links:[Link](https://arxiv.org/abs/2210.01941)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px2.p1.1)\.
- Bengioet al\.\(2003\)Y\. Bengio, R\. Ducharme, P\. Vincent, and C\. JauvinA neural probabilistic language model\.Journal of machine learning research3\(Feb\),pp\. 1137–1155\.External Links:[Link](https://www.jmlr.org/papers/v3/bengio03a.html)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px1.p1.1)\.
- Brownet al\.\(2020\)T\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. D\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell,et al\.Language models are few\-shot learners\.Proceedings of NeurIPS\.External Links:[Link](https://papers.nips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px1.p1.1)\.
- Calanzoneet al\.\(2025\)D\. Calanzone, S\. Teso, and A\. VergariLogically consistent language models via neuro\-symbolic integration\.InProceedings of ICLR,External Links:[Link](https://arxiv.org/abs/2409.13724)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px2.p1.1),[§3\.1](https://arxiv.org/html/2609.13520#S3.SS1.p1.1)\.
- Chavira and Darwiche \(2008\)M\. Chavira and A\. DarwicheOn probabilistic inference by weighted model counting\.Artificial Intelligence172\(6\-7\),pp\. 772–799\.External Links:[Link](https://www.sciencedirect.com/science/article/pii/S0004370207001889)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px2.p1.1),[§3](https://arxiv.org/html/2609.13520#S3.SS0.SSS0.Px4.p1.1)\.
- Chenget al\.\(2026\)J\. Cheng, K\. Richardson, and P\. ChinAnalytica: soft propositional reasoning for robust and scalable llm\-driven analysis\.Proceedings of ICLR\.External Links:[Link](https://iclr.cc/virtual/2026/poster/10011098)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px2.p1.1.1)\.
- De Raedtet al\.\(2016\)L\. De Raedt, K\. Kersting, S\. Natarajan, and D\. PooleStatistical relational artificial intelligence: logic, probability, and computation\.Morgan & Claypool Publishers\.External Links:[Link](https://link.springer.com/book/10.1007/978-3-031-01574-8)Cited by:[§1](https://arxiv.org/html/2609.13520#S1.p3.1.3)\.
- De Raedt and Kimmig \(2015\)L\. De Raedt and A\. KimmigProbabilistic \(logic\) programming concepts\.Machine Learning100\(1\),pp\. 5–47\.External Links:[Link](https://link.springer.com/article/10.1007/s10994-015-5494-z)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px2.p1.1)\.
- Dohanet al\.\(2022\)D\. Dohan, W\. Xu, A\. Lewkowycz, J\. Austin, D\. Bieber, R\. G\. Lopes, Y\. Wu, H\. Michalewski, R\. A\. Saurous, J\. Sohl\-Dickstein,et al\.Language model cascades\.arXiv preprint arXiv:2207\.10342\.External Links:[Link](https://arxiv.org/abs/2207.10342)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px2.p1.1.1)\.
- Elazaret al\.\(2021\)Y\. Elazar, N\. Kassner, S\. Ravfogel, A\. Ravichander, E\. Hovy, H\. Schütze, and Y\. GoldbergMeasuring and improving consistency in pretrained language models\.Transactions of the Association for Computational Linguistics9,pp\. 1012–1031\.External Links:[Link](https://aclanthology.org/2021.tacl-1.60/)Cited by:[§1](https://arxiv.org/html/2609.13520#S1.p6.1.1),[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px1.p1.1.2),[§4\.1](https://arxiv.org/html/2609.13520#S4.SS1.p1.1)\.
- Feldsteinet al\.\(2024\)J\. Feldstein, P\. Dilkas, V\. Belle, and E\. TsamouraMapping the neuro\-symbolic ai landscape by architectures: a handbook on augmenting deep learning through symbolic reasoning\.arXiv preprint arXiv:2410\.22077\.External Links:[Link](https://arxiv.org/abs/2410.22077)Cited by:[§1](https://arxiv.org/html/2609.13520#S1.p3.1.3)\.
- Fierenset al\.\(2015\)D\. Fierens, G\. Van den Broeck, J\. Renkens, D\. Shterionov, B\. Gutmann, I\. Thon, G\. Janssens, and L\. De RaedtInference and learning in probabilistic logic programs using weighted boolean formulas\.Theory and Practice of Logic Programming15\(3\),pp\. 358–401\.External Links:[Link](https://www.cambridge.org/core/journals/theory-and-practice-of-logic-programming/article/inference-and-learning-in-probabilistic-logic-programs-using-weighted-boolean-formulas/9455B07774BA31AA4F6AB81FB0A6B013)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px2.p1.1)\.
- Garget al\.\(2026\)P\. Garg, R\. L\. Geh, D\. Israel, T\. Millstein, K\. Richardson, and G\. V\. d\. BroeckProbabilistic programs of thought\.arXiv preprint arXiv:2604\.17290\.External Links:[Link](https://arxiv.org/abs/2604.17290)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px2.p1.1.1)\.
- Getoor and Taskar \(2007\)L\. Getoor and B\. TaskarIntroduction to statistical relational learning\.MIT press\.Cited by:[§1](https://arxiv.org/html/2609.13520#S1.p3.1.3)\.
- Goodman \(2001\)J\. T\. GoodmanA bit of progress in language modeling\.Computer Speech & Language15\(4\),pp\. 403–434\.External Links:[Link](https://arxiv.org/abs/cs/0108005)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px1.p1.1)\.
- Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan,et al\.The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.External Links:[Link](https://arxiv.org/abs/2407.21783)Cited by:[§4\.2](https://arxiv.org/html/2609.13520#S4.SS2.p1.1)\.
- Groeneveldet al\.\(2024\)D\. Groeneveld, I\. Beltagy, E\. Walsh, A\. Bhagia, R\. Kinney, O\. Tafjord, A\. Jha, H\. Ivison, I\. Magnusson, Y\. Wang,et al\.OLMo: accelerating the science of language models\.InProceedings of ACL,External Links:[Link](https://aclanthology.org/2024.acl-long.841/)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px1.p1.1)\.
- Hewittet al\.\(2022\)J\. Hewitt, C\. D\. Manning, and P\. LiangTruncation sampling as language model desmoothing\.InFindings of EMNLP,External Links:[Link](https://arxiv.org/abs/2210.15191)Cited by:[§1](https://arxiv.org/html/2609.13520#S1.p4.1.1)\.
- Holtzmanet al\.\(2020\)A\. Holtzman, J\. Buys, L\. Du, M\. Forbes, and Y\. ChoiThe curious case of neural text degeneration\.Proceedings of ICLR\.External Links:[Link](https://iclr.cc/virtual_2020/poster_rygGQyrFvH.html)Cited by:[§1](https://arxiv.org/html/2609.13520#S1.p4.1.1),[§3\.1](https://arxiv.org/html/2609.13520#S3.SS1.SSS0.Px1.p3.1)\.
- Janget al\.\(2022\)M\. E\. Jang, D\. Kwon, and T\. LukasiewiczBECEL: benchmark for consistency evaluation of language models\.InProceedings of COLING,External Links:[Link](https://aclanthology.org/2022.coling-1.324/)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px1.p1.1.2)\.
- Jelinek \(1980\)F\. JelinekInterpolated estimation of markov source parameters from sparse data\.InProc\. Workshop on Pattern Recognition in Practice, 1980,Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px1.p1.1)\.
- Kassner and Schütze \(2020\)N\. Kassner and H\. SchützeNegated and misprimed probes for pretrained language models: birds can talk, but cannot fly\.InProceedings of ACL,External Links:[Link](https://arxiv.org/abs/1911.03343)Cited by:[§1](https://arxiv.org/html/2609.13520#S1.p6.1.1),[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px1.p1.1.2),[§4\.1](https://arxiv.org/html/2609.13520#S4.SS1.p1.1)\.
- Kassneret al\.\(2023\)N\. Kassner, O\. Tafjord, A\. Sabharwal, K\. Richardson, H\. Schuetze, and P\. ClarkLanguage models with rationality\.InProceedings of EMNLP,External Links:[Link](https://aclanthology.org/2023.emnlp-main.877.pdf)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px2.p1.1.1)\.
- Lewet al\.\(2023\)A\. K\. Lew, T\. Zhi\-Xuan, G\. Grand, and V\. K\. MansinghkaSequential monte carlo steering of large language models using probabilistic programs\.arXiv preprint arXiv:2306\.03081\.External Links:[Link](https://arxiv.org/abs/2306.03081)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px2.p1.1.1)\.
- Liet al\.\(2019\)T\. Li, V\. Gupta, M\. Mehta, and V\. SrikumarA logic\-driven framework for consistency of neural models\.InProceedings of EMNLP,External Links:[Link](https://aclanthology.org/D19-1405/)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2609.13520#S4.SS1.p1.1)\.
- Liet al\.\(2023\)Z\. Li, J\. Huang, and M\. NaikScallop: a language for neurosymbolic programming\.Proceedings of the ACM on Programming Languages7\(PLDI\),pp\. 1463–1487\.External Links:[Link](https://dl.acm.org/doi/10.1145/3591280)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px2.p1.1)\.
- Lianget al\.\(2022\)P\. Liang, R\. Bommasani, T\. Lee, D\. Tsipras, D\. Soylu, M\. Yasunaga, Y\. Zhang, D\. Narayanan, Y\. Wu, A\. Kumar,et al\.Holistic evaluation of language models\.arXiv preprint arXiv:2211\.09110\.External Links:[Link](https://arxiv.org/abs/2211.09110)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px1.p1.1)\.
- Linet al\.\(2024\)Z\. Lin, Z\. Gou, Y\. Gong, X\. Liu, Y\. Shen, R\. Xu, C\. Lin, Y\. Yang, J\. Jiao, N\. Duan,et al\.Rho\-1: not all tokens are what you need\.arXiv preprint arXiv:2404\.07965\.External Links:[Link](https://arxiv.org/abs/2404.07965)Cited by:[Example 4](https://arxiv.org/html/2609.13520#Thmexample4.p2.2.1)\.
- Liuet al\.\(2025\)Y\. Liu, Z\. Guo, T\. Liang, E\. Shareghi, I\. Vulić, and N\. CollierAligning with logic: measuring, evaluating and improving logical preference consistency in large language models\.Proceedings of ICML\.External Links:[Link](https://arxiv.org/abs/2410.02205)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px1.p1.1.2)\.
- Lozhkovet al\.\(2025\)A\. Lozhkov, E\. Bakouch, G\. M\. Blazquez, G\. Penedo, L\. Tunstall, A\. Marafioti, A\. P\. Lajarín, H\. Kydlíček, V\. Srivastav, J\. Lochner,et al\.Smollm2: when smol goes big—data\-centric training of a fully open small language model\.InProceedings of COLM,External Links:[Link](https://arxiv.org/abs/2502.02737)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px1.p1.1),[Example 2](https://arxiv.org/html/2609.13520#Thmexample2.p1.1.1)\.
- Magnussonet al\.\(2024\)I\. Magnusson, A\. Bhagia, V\. Hofmann, L\. Soldaini, A\. H\. Jha, O\. Tafjord, D\. Schwenk, E\. Walsh, Y\. Elazar, K\. Lo,et al\.Paloma: a benchmark for evaluating language model fit\.Proceedings of NeurIPS\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/760b2d94398aa61468aa3bc11506d9ea-Paper-Datasets_and_Benchmarks_Track.pdf)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px1.p1.1)\.
- Manhaeveet al\.\(2018\)R\. Manhaeve, S\. Dumancic, A\. Kimmig, T\. Demeester, and L\. De RaedtDeepproblog: neural probabilistic logic programming\.Proceedings of NeurIPS\.External Links:[Link](https://proceedings.neurips.cc/paper/2018/hash/dc5d637ed5e62c36ecb73b654b05ba2a-Abstract.html)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px2.p1.1),[§3\.1](https://arxiv.org/html/2609.13520#S3.SS1.p1.1)\.
- Marconatoet al\.\(2026\)E\. Marconato, S\. Bortolotti, E\. van Krieken, P\. Morettin, E\. Umili, A\. Vergari, E\. Tsamoura, A\. Passerini, and S\. TesoSymbol grounding in neuro\-symbolic ai: a gentle introduction to reasoning shortcuts\.Journal of Artificial Intelligence Research86\.External Links:[Link](https://www.jair.org/index.php/jair/article/view/22389)Cited by:[§3\.1](https://arxiv.org/html/2609.13520#S3.SS1.SSS0.Px2.p1.1)\.
- Marconatoet al\.\(2023\)E\. Marconato, S\. Teso, A\. Vergari, and A\. PasseriniNot all neuro\-symbolic concepts are created equal: analysis and mitigation of reasoning shortcuts\.Proceedings of NeurIPS\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/hash/e560202b6e779a82478edb46c6f8f4dd-Abstract-Conference.html)Cited by:[§3\.1](https://arxiv.org/html/2609.13520#S3.SS1.SSS0.Px2.p1.1)\.
- Marraet al\.\(2024\)G\. Marra, S\. Dumančić, R\. Manhaeve, and L\. De RaedtFrom statistical relational to neurosymbolic artificial intelligence: a survey\.Artificial Intelligence328,pp\. 104062\.External Links:ISSN 0004\-3702,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.artint.2023.104062),[Link](https://www.sciencedirect.com/science/article/pii/S0004370223002084)Cited by:[§1](https://arxiv.org/html/2609.13520#S1.p3.1.3)\.
- McCraeet al\.\(2019\)J\. P\. McCrae, A\. Rademaker, F\. Bond, E\. Rudnicka, and C\. FellbaumEnglish WordNet 2019 – an open\-source WordNet for English\.InProceedings of the 10th Global Wordnet Conference,External Links:[Link](https://aclanthology.org/2019.gwc-1.31/),[Document](https://dx.doi.org/10.18653/v1/2019.gwc-1.31)Cited by:[§B\.1](https://arxiv.org/html/2609.13520#A2.SS1.p1.1)\.
- McKenzieet al\.\(2023\)I\. R\. McKenzie, A\. Lyzhov, M\. Pieler, A\. Parrish, A\. Mueller, A\. Prabhu, E\. McLean, A\. Kirtland, A\. Ross, A\. Liu,et al\.Inverse scaling: when bigger isn’t better\.arXiv preprint arXiv:2306\.09479\.External Links:[Link](https://arxiv.org/abs/2306.09479)Cited by:[§1](https://arxiv.org/html/2609.13520#S1.p6.1.1),[Table 3](https://arxiv.org/html/2609.13520#S5.T3)\.
- Novikovaet al\.\(2025\)J\. Novikova, C\. Anderson, B\. Blili\-Hamelin, D\. Rosati, and S\. MajumdarConsistency in language models: current landscape, challenges, and future directions\.arXiv preprint arXiv:2505\.00268\.External Links:[Link](https://arxiv.org/abs/2505.00268)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px1.p1.1.2)\.
- Petroniet al\.\(2019\)F\. Petroni, T\. Rocktäschel, S\. Riedel, P\. Lewis, A\. Bakhtin, Y\. Wu, and A\. MillerLanguage models as knowledge bases?\.InProceedings of EMNLP,External Links:[Link](https://arxiv.org/abs/1909.01066)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2609.13520#S4.SS1.p1.1)\.
- Ribeiroet al\.\(2020\)M\. T\. Ribeiro, T\. Wu, C\. Guestrin, and S\. SinghBeyond accuracy: behavioral testing of nlp models with checklist\.InProceedings of ACL,External Links:[Link](https://aclanthology.org/2020.acl-main.442/)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px1.p1.1)\.
- Richardsonet al\.\(2026\)K\. Richardson, Y\. Feng, P\. Garg, J\. Cheng, G\. Van den Broeck, and D\. RothCoTs as tractable probabilistic programs\.InProceedings of the The Ninth Workshop on Tractable Probabilistic Models,External Links:[Link](https://openreview.net/pdf?id=j1wLo06bmx)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px2.p1.1.1),[Example 1](https://arxiv.org/html/2609.13520#Thmexample1.p2.1.1)\.
- Richardsonet al\.\(2020\)K\. Richardson, H\. Hu, L\. Moss, and A\. SabharwalProbing natural language inference models through semantic fragments\.InProceedings of the AAAI,External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/6397/6253)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2609.13520#S4.SS1.p1.1)\.
- Richardsonet al\.\(2025\)K\. Richardson, V\. Srikumar, and A\. SabharwalUnderstanding the logic of direct preference alignment through logic\.Proceedings of ICML\.External Links:[Link](https://icml.cc/virtual/202https://icml.cc/virtual/2025/poster/464815/poster/46481)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px2.p1.1),[§3\.1](https://arxiv.org/html/2609.13520#S3.SS1.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2609.13520#S3.SS1.p1.1)\.
- Srivastavaet al\.\(2023\)A\. Srivastava, A\. Rastogi, A\. Rao, A\. A\. M\. Shoeb, A\. Abid, A\. Fisch, A\. R\. Brown, A\. Santoro, A\. Gupta, A\. Garriga\-Alonso,et al\.Beyond the imitation game: quantifying and extrapolating the capabilities of language models\.Transactions on machine learning research\.External Links:[Link](https://openreview.net/forum?id=uyTL5Bvosj)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px1.p1.1)\.
- Suchaneket al\.\(2007\)F\. M\. Suchanek, G\. Kasneci, and G\. WeikumYAGO: a core of semantic knowledge\.InProceedings of the 16th International Conference on World Wide Web,External Links:[Link](https://10.0.4.121/1242572.1242667)Cited by:[§B\.1](https://arxiv.org/html/2609.13520#A2.SS1.p1.1)\.
- Teamet al\.\(2025\)G\. Team, A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Rivière, L\. Rouillard, T\. Mesnard, G\. Cideron, J\. Grill, S\. Ramos, E\. Yvinec, M\. Casbon, E\. Pot, I\. Penchev, G\. Liu, F\. Visin, K\. Kenealy, L\. Beyer, X\. Zhai, A\. Tsitsulin, R\. Busa\-Fekete, A\. Feng, N\. Sachdeva, B\. Coleman, Y\. Gao, B\. Mustafa, I\. Barr, E\. Parisotto, D\. Tian, M\. Eyal, C\. Cherry, J\. Peter, D\. Sinopalnikov, S\. Bhupatiraju, R\. Agarwal, M\. Kazemi, D\. Malkin, R\. Kumar, D\. Vilar, I\. Brusilovsky, J\. Luo, A\. Steiner, A\. Friesen, A\. Sharma, A\. Sharma, A\. M\. Gilady, A\. Goedeckemeyer, A\. Saade, A\. Feng, A\. Kolesnikov, A\. Bendebury, A\. Abdagic, A\. Vadi, A\. György, A\. S\. Pinto, A\. Das, A\. Bapna, A\. Miech, A\. Yang, A\. Paterson, A\. Shenoy, A\. Chakrabarti, B\. Piot, B\. Wu, B\. Shahriari, B\. Petrini, C\. Chen, C\. L\. Lan, C\. A\. Choquette\-Choo, C\. Carey, C\. Brick, D\. Deutsch, D\. Eisenbud, D\. Cattle, D\. Cheng, D\. Paparas, D\. S\. Sreepathihalli, D\. Reid, D\. Tran, D\. Zelle, E\. Noland, E\. Huizenga, E\. Kharitonov, F\. Liu, G\. Amirkhanyan, G\. Cameron, H\. Hashemi, H\. Klimczak\-Plucińska, H\. Singh, H\. Mehta, H\. T\. Lehri, H\. Hazimeh, I\. Ballantyne, I\. Szpektor, I\. Nardini, J\. Pouget\-Abadie, J\. Chan, J\. Stanton, J\. Wieting, J\. Lai, J\. Orbay, J\. Fernandez, J\. Newlan, J\. Ji, J\. Singh, K\. Black, K\. Yu, K\. Hui, K\. Vodrahalli, K\. Greff, L\. Qiu, M\. Valentine, M\. Coelho, M\. Ritter, M\. Hoffman, M\. Watson, M\. Chaturvedi, M\. Moynihan, M\. Ma, N\. Babar, N\. Noy, N\. Byrd, N\. Roy, N\. Momchev, N\. Chauhan, N\. Sachdeva, O\. Bunyan, P\. Botarda, P\. Caron, P\. K\. Rubenstein, P\. Culliton, P\. Schmid, P\. G\. Sessa, P\. Xu, P\. Stanczyk, P\. Tafti, R\. Shivanna, R\. Wu, R\. Pan, R\. Rokni, R\. Willoughby, R\. Vallu, R\. Mullins, S\. Jerome, S\. Smoot, S\. Girgin, S\. Iqbal, S\. Reddy, S\. Sheth, S\. Põder, S\. Bhatnagar, S\. R\. Panyam, S\. Eiger, S\. Zhang, T\. Liu, T\. Yacovone, T\. Liechty, U\. Kalra, U\. Evci, V\. Misra, V\. Roseberry, V\. Feinberg, V\. Kolesnikov, W\. Han, W\. Kwon, X\. Chen, Y\. Chow, Y\. Zhu, Z\. Wei, Z\. Egyed, V\. Cotruta, M\. Giang, P\. Kirk, A\. Rao, K\. Black, N\. Babar, J\. Lo, E\. Moreira, L\. G\. Martins, O\. Sanseviero, L\. Gonzalez, Z\. Gleicher, T\. Warkentin, V\. Mirrokni, E\. Senter, E\. Collins, J\. Barral, Z\. Ghahramani, R\. Hadsell, Y\. Matias, D\. Sculley, S\. Petrov, N\. Fiedel, N\. Shazeer, O\. Vinyals, J\. Dean, D\. Hassabis, K\. Kavukcuoglu, C\. Farabet, E\. Buchatskaya, J\. Alayrac, R\. Anil, Dmitry, Lepikhin, S\. Borgeaud, O\. Bachem, A\. Joulin, A\. Andreev, C\. Hardin, R\. Dadashi, and L\. HussenotGemma 3 technical report\.External Links:2503\.19786,[Link](https://arxiv.org/abs/2503.19786)Cited by:[§4\.2](https://arxiv.org/html/2609.13520#S4.SS2.p1.1)\.
- Van Kriekenet al\.\(2024\)E\. Van Krieken, P\. Minervini, E\. M\. Ponti, and A\. VergariOn the independence assumption in neurosymbolic learning\.Proceedings of ICML\.External Links:[Link](https://arxiv.org/abs/2404.08458)Cited by:[§3\.1](https://arxiv.org/html/2609.13520#S3.SS1.SSS0.Px2.p1.1)\.
- Wanget al\.\(2018\)A\. Wang, A\. Singh, J\. Michael, F\. Hill, O\. Levy, and S\. BowmanGLUE: a multi\-task benchmark and analysis platform for natural language understanding\.InProceedings of BlackboxNLP,External Links:[Link](https://aclanthology.org/W18-5446/)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px1.p1.1)\.
- Wellecket al\.\(2019\)S\. Welleck, I\. Kulikov, S\. Roller, E\. Dinan, K\. Cho, and J\. WestonNeural text generation with unlikelihood training\.arXiv preprint arXiv:1908\.04319\.External Links:[Link](https://arxiv.org/abs/1908.04319)Cited by:[§3](https://arxiv.org/html/2609.13520#S3.SS0.SSS0.Px5.p2.1.3)\.
- Xuet al\.\(2018\)J\. Xu, Z\. Zhang, T\. Friedman, Y\. Liang, and G\. BroeckA semantic loss function for deep learning with symbolic knowledge\.InProceedings of ICML,External Links:[Link](https://proceedings.mlr.press/v80/xu18h.html)Cited by:[§2](https://arxiv.org/html/2609.13520#S2.SS0.SSS0.Px2.p1.1),[§3](https://arxiv.org/html/2609.13520#S3.SS0.SSS0.Px5.p1.1),[§3\.1](https://arxiv.org/html/2609.13520#S3.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.13520#S4.SS1.p1.1)\.
- Yanget al\.\(2025\)A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.External Links:[Link](https://arxiv.org/abs/2505.09388)Cited by:[§4\.2](https://arxiv.org/html/2609.13520#S4.SS2.p1.1)\.

## Appendix AProofs

See[2](https://arxiv.org/html/2609.13520#Thmprop2)

###### Proof\.

Let𝖥\\mathsf\{F\}be any satisfiable formula overnnvariables\. Since𝖥\\mathsf\{F\}, in virtue of being satisfiable, has at least one satisfying interpretation, it follows that:

\|𝖨⁡\(𝖥\)\|≥1\.\|\\mathsf\{I\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\|\\geq 1\.Under uniform weights,

WMC0​\(𝖥\)=\|𝖨⁡\(𝖥\)\|2n≥2−n\.\\mathrm\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)=\\frac\{\|\\mathsf\{I\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\|\}\{2^\{n\}\}\\geq 2^\{\-n\}\.Applying−log2\-\\log\_\{2\}gives

I⁡\(𝖥\)=−log2⁡WMC0​\(𝖥\)≤n\.I\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)=\-\\log\_\{2\}\\mathrm\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\\leq n\.The likelihood formula𝖥ℓ\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{\\ell\}has exactly one satisfying interpretation, namely the all\-true assignment, soWMC0​\(𝖥ℓ\)=2−n\\mathrm\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{\\ell\}\)=2^\{\-n\}andI⁡\(𝖥ℓ\)=nI\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{\\ell\}\)=n\. Equality holds exactly when\|𝖨⁡\(𝖥\)\|=1\|\\mathsf\{I\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\|=1\. ∎

See[3](https://arxiv.org/html/2609.13520#Thmprop3)

###### Proof\.

If𝖥1⊧𝖥2\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{1\}\\models\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{2\}, then every interpretation satisfying𝖥1\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{1\}also satisfies𝖥2\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{2\}\. By definition then𝖨⁡\(𝖥1\)⊆𝖨⁡\(𝖥2\)\.\\mathsf\{I\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{1\}\)\\subseteq\\mathsf\{I\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{2\}\)\.Sinceℙθ​\(𝖥\)\\mathbb\{P\}\_\{\\theta\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)is computed by summing nonnegative interpretation weights over𝖨⁡\(𝖥\)\\mathsf\{I\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\), it follows that:

ℙθ​\(𝖥1\)≤ℙθ​\(𝖥2\)\.\\mathbb\{P\}\_\{\\theta\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{1\}\)\\leq\\mathbb\{P\}\_\{\\theta\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{2\}\)\.The same subset relation holds under the uniform weighting used byWMC0\\mathrm\{WMC\}\_\{0\}, giving

WMC0​\(𝖥1\)≤WMC0​\(𝖥2\)\.\\mathrm\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{1\}\)\\leq\\mathrm\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{2\}\)\.Applying−log2\-\\log\_\{2\}reverses the inequality, so that:

I⁡\(𝖥1\)\\displaystyle I\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{1\}\)=−log2⁡WMC0​\(𝖥1\)\\displaystyle=\-\\log\_\{2\}\\mathrm\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{1\}\)≥−log2⁡WMC0​\(𝖥2\)=I⁡\(𝖥2\)\.\\displaystyle\\geq\-\\log\_\{2\}\\mathrm\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{2\}\)=I\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\_\{2\}\)\.∎

See[4](https://arxiv.org/html/2609.13520#Thmprop4)

###### Proof\.

By applying the chain rule to the semantic loss, we have

∇θℓsl​\(𝖥,θ\)=−1WMC⁡\(𝖥,θ\)​∇θWMC​\(𝖥,θ\)\.\\displaystyle\\nabla\_\{\\theta\}\\ell\_\{\\mathrm\{sl\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)=\-\\frac\{1\}\{\\mathrm\{WMC\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\}\\nabla\_\{\\theta\}\\mathrm\{WMC\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\.Taking norms gives

‖∇θℓsl​\(𝖥,θ\)‖=1WMC⁡\(𝖥,θ\)​‖∇θWMC​\(𝖥,θ\)‖\.\\displaystyle\\left\\\|\\nabla\_\{\\theta\}\\ell\_\{\\mathrm\{sl\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\\right\\\|=\\frac\{1\}\{\\mathrm\{WMC\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\}\\left\\\|\\nabla\_\{\\theta\}\\mathrm\{WMC\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\\right\\\|\.By assuming then thatWMC⁡\(𝖥,θ\)≥WMC0​\(𝖥\)\\mathrm\{WMC\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\\geq\\mathrm\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\), it follows that

1WMC⁡\(𝖥,θ\)≤1WMC0​\(𝖥\)\.\\frac\{1\}\{\\mathrm\{WMC\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\}\\leq\\frac\{1\}\{\\mathrm\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\}\.SinceI⁡\(𝖥\)=−log2⁡WMC0​\(𝖥\)I\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)=\-\\log\_\{2\}\\mathrm\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\), we have the following equivalence:

1WMC0​\(𝖥\)=2I⁡\(𝖥\)\.\\displaystyle\\frac\{1\}\{\\mathrm\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\}=2^\{I\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\}\.Combining these \(in\)equalities then yields:

‖∇θℓsl​\(𝖥,θ\)‖≤2I⁡\(𝖥\)​‖∇θWMC​\(𝖥,θ\)‖\.\\left\\\|\\nabla\_\{\\theta\}\\ell\_\{\\mathrm\{sl\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\\right\\\|\\leq 2^\{I\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\}\\left\\\|\\nabla\_\{\\theta\}\\mathrm\{WMC\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\\right\\\|\.∎

See[5](https://arxiv.org/html/2609.13520#Thmprop5)

###### Proof\.

This again relies on the chain rule from above for the semantic loss:

∇θℓsl​\(𝖥,θ\)=−1WMC⁡\(𝖥,θ\)​∇θWMC​\(𝖥,θ\)\.\\nabla\_\{\\theta\}\\ell\_\{\\mathrm\{sl\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)=\-\\frac\{1\}\{\\mathrm\{WMC\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\}\\nabla\_\{\\theta\}\\mathrm\{WMC\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\.By definition,

A⁡\(𝖥,θ\)=WMC0​\(𝖥\)WMC⁡\(𝖥,θ\)\.A\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)=\\frac\{\\mathrm\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\}\{\\mathrm\{WMC\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\}\.Importantly, the following holds:

1WMC⁡\(𝖥,θ\)=A⁡\(𝖥,θ\)WMC0​\(𝖥\)\.\\frac\{1\}\{\\mathrm\{WMC\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\}=\\frac\{A\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\}\{\\mathrm\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\}\.Since again we have:

I⁡\(𝖥\)=−log2⁡WMC0​\(𝖥\),I\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)=\-\\log\_\{2\}\\mathrm\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\),and

1WMC0​\(𝖥\)=2I⁡\(𝖥\)\.\\frac\{1\}\{\\mathrm\{WMC\}\_\{0\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\}=2^\{I\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\}\.Substituting these identities into the semantic\-loss gradient gives the desired output:

∇θℓsl​\(𝖥,θ\)=−A⁡\(𝖥,θ\)​2I⁡\(𝖥\)​∇θWMC​\(𝖥,θ\)\.\\nabla\_\{\\theta\}\\ell\_\{\\mathrm\{sl\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)=\-A\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\\,2^\{I\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\}\)\}\\nabla\_\{\\theta\}\\mathrm\{WMC\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{F\}$\}\};\\theta\)\.∎

## Appendix BDataset Generation

### B\.1Large\-Scale Dataset Generation via Knowledge Graphs

We construct evaluation data for each logical constraint from two complementary knowledge bases \(KBs\): Open English WordNet222[https://github\.com/globalwordnet/english\-wordnet](https://github.com/globalwordnet/english-wordnet)\(OEWN\)\([McCrae et al\., 2019](https://arxiv.org/html/2609.13520#bib.bib52)\), which provides a large\-scale lexical hierarchy, and YAGO\([Suchanek et al\., 2007](https://arxiv.org/html/2609.13520#bib.bib53)\), which contributes physical and historical facts about named entities\.

For each constraint type, we extract relational tuples from these KBs and instantiate them into predefined natural language templates; for example, a tuple relating an entitye1e\_\{1\}to its categorye2e\_\{2\}is rendered as “Ae1e\_\{1\}is ae2e\_\{2\}\.” The KB’s relational structure then determines how individual facts are combined into constraint instances, following the schemas in Table[1](https://arxiv.org/html/2609.13520#S3.T1)\.Forward Implicationpairs two facts sharing an entity such that the KB hierarchy entails one from the other\.Transitivitychains two such hierarchy edges to yield a three\-fact example\.Mutual ExclusivityandSpatial Mutual Exclusivitypair sibling nodes in the KB taxonomy that cannot jointly apply to a single entity\.Negation Consistencypairs a fact with its negated form under the same templates\. This procedure yields a large candidate pool for each constraint, which is subsequently filtered through an LLM\-As\-Judge pipeline\.

### B\.2LLM\-As\-Judge Filtering

Automatically constructing logical constraint datasets from a knowledge base may introduce numerous artifacts\. To address this, we employ an LLM\-As\-Judge to filter the generated evaluation sets according to a strict criteria\. Each extracted datapoint, which consists of a series of automatically extracted facts placed into predefined templates and a constraint between them, is scored according to the following binary criteria\.

Table 4:Dataset sizes at each filtering stage\.Raw: initial generated examples\.Filtered: after LLM\-as\-judge filtering\.Gemma\-3,Llama\-3,Qwen3: remaining examples after removing multi\-token prediction events per model family\.Table 5:Human validation of the KB to LLM filtering pipeline onn=50n=50sampled examples per dataset\.Human Acceptis the fraction accepted by both annotators\.LLM Acceptis the fraction accepted byGPT\-5o\-mini\.Ann1 PrecisionandAnn2 Precisiongive the fraction of LLM\-accepted items each annotator independently judged correct i\.e\., the LLM filter’s precision as measured by each annotator\.1. 1\.Grammatical Coherence:All facts must be grammatically natural\. Because facts are extracted from knowledge graphs and inserted into predefined templates, grammatical errors and unnatural phrasings can emerge\.
2. 2\.Monosemous:All facts must concern entities with a single, unambiguous meaning\. Knowledge graphs contain polysemous nodes whose extracted facts may be interpretable under multiple senses \(e\.g\., “bat” may refer to the animal or the sports equipment\), undermining the logical validity of the constraint\.
3. 3\.Entity Salience:All facts must concern well\-known, real\-world entities\. Knowledge graphs sometimes contain overly specific, archaic, or obscure entries that, while technically valid graph nodes, produce evaluation examples that are unreasonable to expect a language model to have encountered during pre\-training\.
4. 4\.Factual Accuracy:All facts must indeed be true\. This controls for any errors that may exist in the knowledge graph\.

We useGPT\-5o\-minias the LLM judge, and only keep datapoints passing each criteria\.

### B\.3Dataset Sizes

To limit our evaluation to single\-token substitutions, we additionally apply a final model\-specific filtering step to remove any data points whose trigger word is split into multiple tokens\. We provide the size of each dataset inProbCTat each filtering stage in Table[4](https://arxiv.org/html/2609.13520#A2.T4)\. This consists of the raw dataset sizes extracted from the knowledge bases, the sizes after LLM\-as\-judge filtering, and the sizes after removing multi\-token prediction events for each model family\.

### B\.4Human Validation of LLM Filter

Table[5](https://arxiv.org/html/2609.13520#A2.T5)reports a human validation study of the knowledge base to final LLM\-filtered dataset\. For each dataset, we samplen=50n=50examples from the knowledge base and collect independent judgments from two annotators alongside the LLM filter’s decision\. We report acceptance rates from human annotators through majority vote and from the LLM filter, along with the percentage of LLM\-accepted items also accepted by each annotator\.

Both annotators estimate LLM precision at0\.880\.88or higher on every dataset\. Therefore items retained by the filter are reliably correct under independent human review, and the underlying judgments appear consistent across annotators rather than annotator\-specific\. Moreover, the LLM is generally the stricter filter, retaining a similar or smaller fraction of examples than human annotators\.

## Appendix CInverse Scaling Investigation

Table 6:Inverse scaling analysis for the forward implication𝖯1​→𝖯2\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{1\}\\text\{\{\\color\[rgb\]\{0\.66,0\.13,0\.24\}$\\to$\}\}\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{2\}across all model families\.Δ\\Deltais shorthand forℙθC​\(𝖯2\)−ℙθC​\(𝖯1\)\\mathbb\{P\}\_\{\\theta\_\{C\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{2\}\)\{\-\}\\mathbb\{P\}\_\{\\theta\_\{C\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{1\}\), the mean difference between conclusion and premise probabilities\.ℙθC​\(𝖯2\>𝖯1\)\\mathbb\{P\}\_\{\\theta\_\{C\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{2\}\{\>\}\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{1\}\)is the percentage of points whereℙθC​\(𝖯2\)\>ℙθC​\(𝖯1\)\\mathbb\{P\}\_\{\\theta\_\{C\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{2\}\)\>\\mathbb\{P\}\_\{\\theta\_\{C\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{1\}\)\. Calibration is done usingnucleus entropyaveraged over nucelus valuesp∈\{0\.80,0\.85,0\.90,0\.95\}p\\in\\\{0\.80,0\.85,0\.90,0\.95\\\}\.Table 7:Inverse scaling analysis for transitivity\(𝖯1​∧𝖯2\)​→𝖯3\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{1\}\\text\{\{\\color\[rgb\]\{0\.66,0\.13,0\.24\}$\\land$\}\}\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{2\}\)\\text\{\{\\color\[rgb\]\{0\.66,0\.13,0\.24\}$\\to$\}\}\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{3\}, across all model families\.Δ\\Deltais shorthand forℙθC​\(𝖯3\)−ℙθC​\(𝖯1\)​ℙθC​\(𝖯2\)\\mathbb\{P\}\_\{\\theta\_\{C\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{3\}\)\{\-\}\\mathbb\{P\}\_\{\\theta\_\{C\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{1\}\)\\mathbb\{P\}\_\{\\theta\_\{C\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{2\}\), the mean gap between the conclusion probability and the product of premise probabilities\.ℙθC​\(𝖯3\>𝖯1​𝖯2\)\\mathbb\{P\}\_\{\\theta\_\{C\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{3\}\{\>\}\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{1\}\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{2\}\)is the percentage of points whereℙθC​\(𝖯3\)\>ℙθC​\(𝖯1\)​ℙθC​\(𝖯2\)\\mathbb\{P\}\_\{\\theta\_\{C\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{3\}\)\>\\mathbb\{P\}\_\{\\theta\_\{C\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{1\}\)\\mathbb\{P\}\_\{\\theta\_\{C\}\}\(\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{2\}\)\. Calibration is done usingnucleus entropyaveraged over nucleus valuesp∈\{0\.80,0\.85,0\.90,0\.95\}p\\in\\\{0\.80,0\.85,0\.90,0\.95\\\}\.Table 8:Examples of inverse scaling for forward implication, comparingGemma\-3\-270MandGemma\-3\-12Busingnucleus entropywith nucleus valuep=0\.95p=0\.95\. Each row defines𝖯1=ℳ⁡\(Prefix,a\)\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{1\}=\\mathcal\{M\}\(\\text\{Prefix\},a\)and𝖯2=ℳ⁡\(Prefix,b\)\\text\{\{\\color\[rgb\]\{0,0,1\}$\\mathsf\{P\}$\}\}\_\{2\}=\\mathcal\{M\}\(\\text\{Prefix\},b\), whereaaandbbare the premise and conclusion completions respectively\.We provide additional evidence explaining the inverse scaling phenomenon observed in the Forward Implication and Transitivity datasets in Tables[6](https://arxiv.org/html/2609.13520#A3.T6)and[7](https://arxiv.org/html/2609.13520#A3.T7)\. Across all model families, the smallest model is the most consistent and the largest model is the least consistent, where consistency is measured by aggregate CAcc scores \(Qwen3is a minor exception, with the second largest model being slightly more inconsistent than the largest model\)\. In all cases the probabilities assigned to the conclusions, which are always broader consequences of the premises inProbCT, decrease with model size on average\. The probabilities assigned to the more narrower premises increases with size on average for theGemma\-3andQwen3model families, and remain similar across model sizes in theLlama\-3family\. Specific examples exhibiting this phenomenon betweenGemma\-3\-270MandGemma\-3\-12Bare shown in Table[8](https://arxiv.org/html/2609.13520#A3.T8)\.

## Appendix DEnforcing Factuality

Table 9:Enforcing Factuality: CAcc \(%\) over conjunction of true statements usingnucleus entropycalibration averaged overp∈\{0\.8,0\.85,0\.9,0\.95\}p\\in\\\{0\.8,0\.85,0\.9,0\.95\\\}\.Our main results utilize general logical constraints without imposing factuality\. However,ProbCTis generated from knowledge bases and therefore comes with ground truth factual labels\. This allows us to evaluate each dataset under a modified logical constraint that consists of the conjunction of all true statements in a data point\. These results are presented in Figure[9](https://arxiv.org/html/2609.13520#A4.T9)\.

## Appendix EAdditional Calibration Methods

Here we consider alternative choices for the calibration method; that is we consider different choices of the valueCCin Equation[4](https://arxiv.org/html/2609.13520#S3.E4)\. As a reminder, our default method uses the entropy of the top\-ppnucleus of the local distribution, and results are reproduced in Table[10](https://arxiv.org/html/2609.13520#A5.T10)for ease of comparison\.Entropy calibration, shown in Table[11](https://arxiv.org/html/2609.13520#A5.T11), setsCCto be the entropy of the full next\-token distribution, dropping the nucleus restriction\.Nucleus calibration, shown in Table[12](https://arxiv.org/html/2609.13520#A5.T12), setsCCto be the log count of tokens in the top\-ppnucleus, replacing entropy with a uniform count over the same set\.

Both alternatives generally preserve the qualitative findings of Table[2](https://arxiv.org/html/2609.13520#S3.T2): models generally perform poorly, inverse scaling on Transitivity and Forward Implication is present across model families, and the task difficulty ordering is unchanged\. Spatial Exclusivity is the most sensitive to calibration method, with meaningfully improved scores here compared to nucleus entropy calibration across model families and sizes\. For example,Gemma\-3\-12Bhas an average score of18\.518\.5for nucleus entropy calibration, while it has scores of59\.459\.4and44\.444\.4for nucleus calibration and entropy calibration respectively\. Scores on Negation Consistency remain similarly poor across model families and calibration methods\.

Table 10:CAcc \(%\) usingnucleus entropycalibration averaged over nucleus valuesp∈\{0\.8,0\.85,0\.9,0\.95\}p\\in\\\{0\.8,0\.85,0\.9,0\.95\\\}\. Reproduced from Table[2](https://arxiv.org/html/2609.13520#S3.T2)for ease of comparison with the alternative calibrations in Tables[11](https://arxiv.org/html/2609.13520#A5.T11)and[12](https://arxiv.org/html/2609.13520#A5.T12)Table 11:CAcc \(%\) usingentropycalibration\.Table 12:CAcc \(%\) usingnucleuscalibration averaged over nucleus valuesp∈\{0\.8,0\.85,0\.9,0\.95\}p\\in\\\{0\.8,0\.85,0\.9,0\.95\\\}\.
## Appendix FAdditional Probabilistic Consistency Histograms

Gemma\-3\-270M

Gemma\-3\-1B\-pt

Gemma\-3\-4B\-pt

Gemma\-3\-12B\-ptTransitivityForward ImplicationNegation ConsistencyMutual ExclusivitySpatial Exclusivity

Figure 4:WMC score distributions forGemma\-3base models usingnucleus entropycalibration with nucleus valuep=0\.95p=0\.95\. Points highlighted in red are inconsistent, while points highlighted in blue are consistent\. The solid vertical black line corresponds to the average probabilistic consistency\. Each figure is labeled with the corresponding aggregate CAcc score\.
Llama\-3\.2\-1B

Llama\-3\.2\-3B

Llama\-3\.1\-8BTransitivityForward ImplicationNegation ConsistencyMutual ExclusivitySpatial Exclusivity

Figure 5:Probabilistic Consistency score distributions forLlama\-3pretrained models usingnucleus entropycalibration with nucleus valuep=0\.95p=0\.95\. Points highlighted in red are inconsistent, while points highlighted in blue are consistent\. The solid vertical black line corresponds to the average probabilistic consistency\. Each figure is labeled with the corresponding aggregate CAcc score\.
Qwen3\-0\.6B

Qwen3\-1\.7B

Qwen3\-4B

Qwen3\-8B

Qwen3\-14BTransitivityForward ImplicationNegation ConsistencyMutual ExclusivitySpatial Exclusivity

Figure 6:Probabilistic Consistency score distributions forQwen3pretrained models usingnucleus entropycalibration with nucleus valuep=0\.95p=0\.95\. Points highlighted in red are inconsistent, while points highlighted in blue are consistent\. The solid vertical black line corresponds to the average probabilistic consistency\. Each figure is labeled with the corresponding aggregate CAcc score\.

Similar Articles

Probabilistic Attribution For Large Language Models

arXiv cs.CL

This paper proposes a model-agnostic probabilistic token attribution measure for LLMs using Bayes' rule to invert next-token log probabilities, capturing the model's internal representation of token sequences and improving interpretability through entropy analysis.