Selection Shapes the Boundary: A Preregistered Replication of Monotonicity and Label Agreement in Unselected NLI Populations

arXiv cs.CL Papers

Summary

This preregistered replication tests whether the monotonicity effect on label agreement in NLI generalizes from selected low-agreement items to unselected populations, finding that the effect reverses and is small, suggesting the earlier finding was conditional on selection.

arXiv:2607.19231v1 Announce Type: new Abstract: Prior work on human label variation (HLV) in natural language inference (NLI) has often relied on re-annotation resources that select items by disagreement level. An earlier study (arXiv:2607.15870) found that hypotheses containing non-upward monotonicity operators showed lower label agreement in ChaosNLI (Cliff's delta = -0.284), which is restricted to items whose majority label carries exactly three of five votes. We preregistered a replication of this boundary in the unselected populations that ChaosNLI was drawn from: the SNLI and MultiNLI development sets, using the same operator tagger and a four-level ordinal agreement outcome. The registered prediction fails. All seven contrasts return a positive Cliff's delta (non-upward items agree slightly more, not less), the only significant confirmatory contrast has the opposite sign to the registration, and every effect is far below our smallest effect size of interest (0.10). Robustness checks support the measurement: simulated tagger misclassification shrinks the effects rather than manufacturing them, and a manual re-tagging audit reaches four-class agreement of 0.875 on a fresh 200-item sample. We conclude that the earlier negative boundary is plausibly a structure conditional on low-agreement selection rather than a population-level property, and that HLV structure claims built on selected re-annotation resources should state their selection conditional explicitly.
Original Article
View Cached Full Text

Cached at: 07/22/26, 08:25 AM

# A Preregistered Replication of Monotonicity and Label Agreement in Unselected NLI Populations
Source: [https://arxiv.org/html/2607.19231](https://arxiv.org/html/2607.19231)
###### Abstract

Prior work on human label variation \(HLV\) in natural language inference \(NLI\) has often relied on re\-annotation resources that select items by disagreement level\. An earlier study\(Choi,[2026](https://arxiv.org/html/2607.19231#bib.bib3)\)found that hypotheses containing non\-upward monotonicity operators showed lower label agreement in ChaosNLI \(Cliff’sδ=−0\.284\\delta=\-0\.284\), which is restricted to items whose majority label carries exactly three of five votes\. We preregistered a replication of this boundary in the unselected populations that ChaosNLI was drawn from: the SNLI and MultiNLI development sets, using the same operator tagger and a four\-level ordinal agreement outcome\. The registered prediction fails\. All seven contrasts return a positive Cliff’sδ\\delta\(non\-upward items agree slightly more, not less\), the only significant confirmatory contrast has the opposite sign to the registration, and every effect is far below our smallest effect size of interest \(0\.10\)\. Robustness checks support the measurement: simulated tagger misclassification shrinks the effects rather than manufacturing them, and a manual re\-tagging audit reaches four\-class agreement of 0\.875 on a fresh 200\-item sample\. We conclude that the earlier negative boundary is plausibly a structure conditional on low\-agreement selection rather than a population\-level property, and that HLV structure claims built on selected re\-annotation resources should state their selection conditional explicitly\.

Selection Shapes the Boundary: A Preregistered Replication of Monotonicity and Label Agreement in Unselected NLI Populations

Haram ChoiUniversity of Bremenhachoi@uni\-bremen\.de

## 1Introduction

Human label variation in NLI is increasingly treated as signal rather than annotation noise: multiple labels per item reflect genuine interpretive variation that models and evaluations should represent\(Plank,[2022](https://arxiv.org/html/2607.19231#bib.bib9); Pavlick and Kwiatkowski,[2019](https://arxiv.org/html/2607.19231#bib.bib8)\)\. Under this perspectivist frame, disagreement is a measurement target, and a natural research question is whether linguistic properties of an item predict how much annotators will disagree on it\.

One candidate predictor is monotonicity\.Choi \([2026](https://arxiv.org/html/2607.19231#bib.bib3)\)tagged hypothesis sentences for monotonicity operators and found that in ChaosNLI, items whose hypotheses contained non\-upward operators \(downward\-entailing or non\-monotone triggers\) had lower label agreement than purely upward items, with a Cliff’sδ\\deltaof−0\.284\-0\.284\(95% CI\[−0\.332,−0\.235\]\[\-0\.332,\-0\.235\]\) on a 100\-label entropy outcome\.

That estimate, however, was conditional on a strong selection rule\. ChaosNLI re\-annotated only items whose majority label carried exactly three of the five original SNLI/MNLI annotator labels\. Whether the monotonicity boundary exists in the unselected populations those items came from was never measured\. This paper measures it, under preregistration\. We apply the identical tagger, whose rules were frozen before analysis, to the SNLI and MNLI development sets and test the registered prediction that non\-upward items sit lower on a four\-level agreement ordinal built from the original five labels\.

The registered prediction is rejected\. The sign reverses in all seven contrasts, and every effect is small: the largestδ\\deltais\+0\.059\+0\.059, far below our preregistered smallest effect size of interest \(SESOI\) of 0\.10, and no confidence interval reaches it\. The contribution of this paper is the measured selection dependence itself: a boundary that is present inside a disagreement\-selected resource can vanish and even faintly reverse in the unselected population\.

## 2Related Work

### NLI disagreement resources\.

SNLI and MNLI development items carry five labels each, the original writer’s label plus four validation labels\(Bowman et al\.,[2015](https://arxiv.org/html/2607.19231#bib.bib1); Williams et al\.,[2018](https://arxiv.org/html/2607.19231#bib.bib12)\); ChaosNLI re\-annotates a selected subset with 100 labels each\(Nie et al\.,[2020](https://arxiv.org/html/2607.19231#bib.bib6), Section 3\.1\)\. The perspectivist literature argues for modeling the resulting label distributions directly\(Plank,[2022](https://arxiv.org/html/2607.19231#bib.bib9), Section 3\); see alsoPavlick and Kwiatkowski \([2019](https://arxiv.org/html/2607.19231#bib.bib8)\)\.

### Monotonicity and NLI\.

Monotonicity reasoning is a classic locus of difficulty in NLI\.MacCartney \([2009](https://arxiv.org/html/2607.19231#bib.bib5), Section 1\.4\.1\)notes that downward\-entailing behavior is not confined to obvious negation: “constructions which are not upward monotone are surprisingly widespread”, spanning quantifiers, conditional antecedents, superlatives, and open\-class triggers across several parts of speech\. The MED dataset targets monotonicity reasoning directly\(Yanaka et al\.,[2019](https://arxiv.org/html/2607.19231#bib.bib13)\); we use it to anchor tagger error rates\.

### Operator semantics is itself contested\.

For several operators the monotonicity classification is a live semantic debate rather than a lookup:*only*has been analyzed as Strawson downward\-entailing rather than simply downward\-entailing\(von Fintel,[1999](https://arxiv.org/html/2607.19231#bib.bib11)\), and*many*has received both cardinal and proportional readings\(Rett,[2018](https://arxiv.org/html/2607.19231#bib.bib10), Section 2\.1\.2\); for that ambiguity, Rett citesPartee \([1989](https://arxiv.org/html/2607.19231#bib.bib7)\)among others\. This is relevant to the audit in Section[5](https://arxiv.org/html/2607.19231#S5): part of the tagger\-human disagreement we observe is irreducible for exactly this reason\.

## 3Method

### Data\.

SNLI dev \(9,986 items after excluding 14 items with four validation labels\) and MNLI matched dev \(10,000 items\)\. MNLI mismatched dev \(9,946 items after excluding 54 rows from 27 duplicated pair IDs\) is analyzed as a preregistered secondary corpus\. All corpora carry five labels per item\. MED \(5,382 premise\-hypothesis pairs\) is used only for tagger validation, and the ChaosNLI overlap \(1,514 SNLI rows, of which 1,507 remain after the same four\-label exclusion, and 1,599 MNLI matched items\) only for the bridge analysis in Section[4](https://arxiv.org/html/2607.19231#S4)\.

### Predictor\.

A rule\-based monotonicity operator tagger \(v0\.3, frozen before this study under preregistration; no rule changes\) labels each hypothesis upward, downward, non\-monotone, or mixed by a surface trigger scan\. The analysis contrast is binary: purely upward versus non\-upward\.

### Outcome\.

A four\-level agreement ordinal from the five validation labels: no majority<<3/5<<4/5<<5/5\. The primary analysis includes the no\-majority level; excluding it is sensitivity analysis A\.

### Statistics\.

Tie\-corrected Cliff’sδ\\deltawith two\-sided Mann\-Whitney U tests, percentile bootstrap CIs \(10,000 resamples, seed 20260717\), Holm correction over the two confirmatory contrasts \(SNLI primary, MNLI matched primary\), and a preregistered SESOI of\|δ\|=0\.10\|\\delta\|=0\.10\. The registered prediction wasδ<0\\delta<0\(non\-upward items lower in agreement\)\. The registration is a plan document written into our working repository on 2026\-07\-17, before the confirmatory analysis was run\. Its timestamp is a version\-control record under our own control rather than a third\-party registry entry, so we report the date as a statement about our procedure and not as independently attested evidence;AUDIT\.mdin the reproduction repository sets out the registration timeline in those terms, and the full registration\-versus\-result concordance is in Appendix[A](https://arxiv.org/html/2607.19231#A1)\.

### Bridge to the earlier estimate\.

The originally planned correlation of the ordinal outcome with ChaosNLI entropy inside the overlap is undefined by arithmetic: ChaosNLI selected only items with a 3/5 modal label, so the agreement ordinal is constant \(variance zero\) on the entire overlap\. We state this openly and replace the bridge with level\-wise descriptive statistics \(Section[4](https://arxiv.org/html/2607.19231#S4)\)\. Scale caveat: the earlierδ\\deltawas computed on a 100\-label entropy outcome and is not numerically commensurable withδ\\deltavalues on the four\-level ordinal; we compare sign and magnitude class only\.

## 4Results

Table 1:All contrasts\. Registered prediction:δ<0\\delta<0\. SESOI:\|δ\|=0\.10\|\\delta\|=0\.10\. No CI reaches the SESOI boundary\. The earlier interval\[−0\.332,−0\.235\]\[\-0\.332,\-0\.235\]sits on a non\-commensurable outcome scale, so comparison with it is one of sign and magnitude class only\.### The registered prediction is rejected\.

Table[1](https://arxiv.org/html/2607.19231#S4.T1)reports all seven contrasts\. Everyδ\\deltais positive: non\-upward items sit very slightly higher, not lower, on the agreement ordinal\. The SNLI primary contrast is not significant \(δ=\+0\.032\\delta=\+0\.032, 95% CI\[−0\.011,\+0\.075\]\[\-0\.011,\+0\.075\], Holmp=0\.162p=0\.162\)\. The MNLI matched primary contrast is significant but with the sign opposite to the registration \(δ=\+0\.045\\delta=\+0\.045, 95% CI\[\+0\.024,\+0\.066\]\[\+0\.024,\+0\.066\], Holmp=7\.1×10−5p=7\.1\\times 10^\{\-5\}\)\. The secondary mismatched corpus behaves like matched \(δ=\+0\.059\\delta=\+0\.059\)\. Dropping no\-majority items \(sensitivity A\) and merging SNLI with matched \(sensitivity B\) change nothing material\.

![Refer to caption](https://arxiv.org/html/2607.19231v1/x1.png)Figure 1:The seven contrasts of Table[1](https://arxiv.org/html/2607.19231#S4.T1)\. Points are tie\-corrected Cliff’sδ\\deltawith percentile bootstrap 95% confidence intervals; the shaded band marks the preregistered SESOI at\|δ\|=0\.10\|\\delta\|=0\.10\. The gray reference row is labeled*Sprint 1 \(entropy scale\)*after the earlier study’s internal name and reports that study’s estimate\. It is computed on a 100\-label entropy outcome that is not numerically commensurable with the four\-level ordinal used here and is shown for sign and magnitude comparison only\. Generated byanalysis/fig\_forest\.py\.
### Effects are below the SESOI everywhere\.

All\|δ\|<0\.10\|\\delta\|<0\.10, and no bootstrap CI touches the boundary \(Figure[1](https://arxiv.org/html/2607.19231#S4.F1)\)\. Item\-level discriminability is essentially absent: ordered logistic pseudo\-R2R^\{2\}is 0\.000093 \(SNLI\) and 0\.00084 \(matched\), and the AUC for separating high\-agreement \(4/5, 5/5\) from low\-agreement items is at chance in both corpora \(0\.492 SNLI, 0\.484 matched\)\.

### Outcome resolution does not explain the reversal\.

The earlier study read disagreement off 100 labels per item, while the ordinal used here has five levels, so one might ask whether the boundary failed to replicate only because the outcome lost resolution\. Coarsening an outcome generally attenuates an association toward zero rather than reversing its sign, and everyδ\\deltawe observe is positive, so lost resolution cannot manufacture the direction we report\. What resolution does bound is how precisely the two studies can be compared, which is why we restrict that comparison to sign and magnitude class\. Subsampling experiments that vary annotation count from one to one hundred labels per item on ChaosNLI make that dependence explicit\(Kadasi and Singh,[2023](https://arxiv.org/html/2607.19231#bib.bib4)\)\.

### Complexity confounders do not explain the reversal\.

As an auxiliary, non\-confirmatory defense we fit proportional\-odds models of the agreement ordinal on the binary monotonicity predictor, without \(M0\) and with \(M1\) length, parse depth, and genre covariates\. In matched and mismatched the monotonicity coefficient keeps its sign and stays away from zero after adjustment \(matched M1:−0\.202\-0\.202,p=2\.4×10−6p=2\.4\\times 10^\{\-6\}; mismatched M1:−0\.248\-0\.248,p=7\.3×10−9p=7\.3\\times 10^\{\-9\}; negative coefficients here correspond to the positiveδ\\deltavalues above, because upward items sit lower in agreement\)\. In SNLI the proportional\-odds assumption fails a Brant test\(Brant,[1990](https://arxiv.org/html/2607.19231#bib.bib2)\), and the multinomial substitute shows no significant category\-wise monotonicity effect, which is consistent with the non\-significant SNLIδ\\delta\.

![Refer to caption](https://arxiv.org/html/2607.19231v1/x2.png)Figure 2:Share of non\-upward hypotheses at each agreement level, with Wilson 95% confidence intervals, for SNLI dev and MNLI matched dev\. The share does not decrease as agreement rises\. Generated byanalysis/fig\_bridge\_shares\.py\.
### Level\-wise structure runs against the registered direction\.

The share of non\-upward hypotheses does not decrease with agreement level\. In matched it is highest among unanimous items: 0\.303 \(no majority\), 0\.308 \(3/5\), 0\.312 \(4/5\), 0\.350 \(5/5\)\. SNLI shows a flat low profile \(0\.045, 0\.041, 0\.060, 0\.057\)\. Figure[2](https://arxiv.org/html/2607.19231#S4.F2)plots both corpora; Appendix[B](https://arxiv.org/html/2607.19231#A2)gives the arithmetic\.

## 5Validity of the Measurement

### Tagger misclassification \(Tier 2\)\.

We anchored flip rates in MED measurements of the tagger \(symmetric error542/5,382=0\.101542/5\{,\}382=0\.101; upward to non\-upward229/1,818=0\.126229/1\{,\}818=0\.126; non\-upward to upward313/3,564=0\.088313/3\{,\}564=0\.088; denominators recomputed from the released MED file, which differs by two items from the class counts reported inYanaka et al\.,[2019](https://arxiv.org/html/2607.19231#bib.bib13); MED domain, not a dev\-domain error claim\) and simulated label flips over a 36\-cell grid \(three corpora, three scenarios, anchored plus fixed rates, 1,000 replications each\)\. Misclassification shrinksδ\\deltatoward zero rather than manufacturing it: at the MED\-anchored symmetric rate the matchedδ\\deltaaverages 0\.0345 with significance retained in 94\.9% of replications, and mismatched retains it in every replication\. Retention drops only under the extreme fixed rates, most steeply for matched at symmetric 0\.20 \(0\.594\), where mismatched still holds at 0\.933\. Across all 36 cells the 97\.5th percentile of simulatedδ\\deltanever exceeds 0\.064, so the below\-SESOI conclusion never flips \(Appendix[C](https://arxiv.org/html/2607.19231#A3)\)\.

### Manual audit \(Tier 3\)\.

We hand\-tagged a fresh blind sample of 200 hypotheses \(50 each from SNLI dev, MNLI matched dev, and the two ChaosNLI overlap strata\) against a written codebook that fixes the judgment target: does the hypothesis contain a monotonicity trigger, judged by semantic definition rather than a word list\. Four\-class agreement with the tagger was 0\.875 \(Wilson 95% CI\[0\.822,0\.914\]\[0\.822,0\.914\]\), binary agreement 0\.905, Cohen’sκ\\kappa0\.607 \(four\-class; both marginals are heavily skewed toward upward, soκ\\kappaunderstates agreement and is read alongside the raw rates\)\. An earlier round without a codebook produced agreement of 0\.135\. The gap is an instrument\-specification result rather than an error: without a fixed judgment target, the human rater measured a different construct \(premise\-relative entailment direction rather than trigger presence\)\. We report both rounds\.

### Three layers of residual disagreement\.

We read the 25 disagreements by hand \(Appendix[D](https://arxiv.org/html/2607.19231#A4)\) and assigned each to an error layer\. This layer coding is a qualitative judgment, not a script output; the counts below are approximate and are the one set of numbers in this paper that no script regenerates\. They decompose into \(a\) about 10 tagger lexicon coverage gaps, \(b\) about 5 polysemy and construction failures, and \(c\) about 10 items where the operator’s monotonicity classification is itself semantically contested, including*only*\(rated downward by the human 4 times and left upward 3 times within this sample;von Fintel,[1999](https://arxiv.org/html/2607.19231#bib.bib11)\) and*many*\(Rett,[2018](https://arxiv.org/html/2607.19231#bib.bib10)\)\. Layers \(a\) and \(b\) are engineering\-reducible; we judge layer \(c\) not to be, and it is itself direct evidence for treating disagreement as a measurement target\. We deliberately did not patch the tagger in response to this audit, because reacting to the audit sample would circularize the validation\.

## 6Discussion

The leading interpretation is that the earlier boundary is a structure conditional on low\-agreement selection\. Inside a resource restricted to 3/5\-split items, non\-upward operators separated items by residual disagreement; in the unselected population, they do not, and the residual association even runs slightly the other way\. Two implications follow\. First, HLV structure claims estimated on selected re\-annotation resources should carry their selection conditional explicitly; the ChaosNLI selection rule is strong enough to zero out the variance of the four\-level agreement ordinal on the overlap, which is why the natural bridge analysis is undefined by arithmetic\. Second, the disagreement decomposition shows that part of the tagger\-human gap is the contested semantics of the operators themselves, which lexicon engineering is unlikely to resolve given the live debates documented invon Fintel \([1999](https://arxiv.org/html/2607.19231#bib.bib11)\)andRett \([2018](https://arxiv.org/html/2607.19231#bib.bib10)\); measurement frames that treat disagreement as noise to be eliminated would misread exactly this layer\. The preregistration, SESOI, and freeze discipline are what give the negative result its evidential value; we report it as registered\.

## 7Reproducibility

Every quantity this paper reports is printed in the paper itself, including the full appendix tables, so no claim requires external material\. The reproduction repository \([https://github\.com/oudeis01/nli\-hlv\-selection](https://github.com/oudeis01/nli-hlv-selection)\) supplies the code that regenerates those quantities\. File paths named here are relative to that repository’s root\.

All reported numbers regenerate from the analysis scripts, except the manual error layer counts in Section[5](https://arxiv.org/html/2607.19231#S5), which come from manual inspection of the items in Table[9](https://arxiv.org/html/2607.19231#A4.T9)\. Every run that involves randomness fixes a seed and records it in that run’s output file\. The repository README maps each table, figure, and statistic to the step that produces it, and gives the run order\.

AUDIT\.mdis the audit trail\. It records the registration timeline, what had already been observed at each registration point, every deviation from the registered plan, and every negative result\. The repository is a curated snapshot, not a full development history, andAUDIT\.mdstates what that costs\.

## Limitations

Manual tagging was performed by a single author, blind to tagger output but without an independent second rater; the reportedκ\\kappais tagger\-versus\-human, not inter\-human\. The tagger is a surface trigger scanner with measured coverage gaps \(layers \(a\) and \(b\) of Section[5](https://arxiv.org/html/2607.19231#S5)\)\. MED anchor rates are measured on MED sentences, not on the SNLI/MNLI dev domain\. The four\-level agreement ordinal is coarse, and all data are English\. The earlier study and this replication use non\-commensurable outcome scales, so the replication verdict rests on sign and magnitude class, not on interval overlap\.

## References

- Bowman et al\. \(2015\)Samuel R\. Bowman, Gabor Angeli, Christopher Potts, and Christopher D\. Manning\. 2015\.[A large annotated corpus for learning natural language inference](https://doi.org/10.18653/v1/D15-1075)\.In*Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing*, pages 632–642, Lisbon, Portugal\. Association for Computational Linguistics\.
- Brant \(1990\)Rollin Brant\. 1990\.[Assessing proportionality in the proportional odds model for ordinal logistic regression](https://doi.org/10.2307/2532457)\.*Biometrics*, 46\(4\):1171–1178\.
- Choi \(2026\)Haram Choi\. 2026\.[How much human label variation does formal semantic structure explain?: Group\-level effects and item\-level ceilings in NLI](https://arxiv.org/abs/2607.15870)\.*Preprint*, arXiv:2607\.15870\.ArXiv preprint, v1, submitted 2026\-07\-17\.
- Kadasi and Singh \(2023\)Pritam Kadasi and Mayank Singh\. 2023\.[Unveiling the multi\-annotation process: Examining the influence of annotation quantity and instance difficulty on model performance](https://doi.org/10.18653/v1/2023.findings-emnlp.96)\.In*Findings of the Association for Computational Linguistics: EMNLP 2023*, pages 1371–1388, Singapore\. Association for Computational Linguistics\.
- MacCartney \(2009\)Bill MacCartney\. 2009\.[*Natural Language Inference*](https://nlp.stanford.edu/~wcmac/papers/nli-diss.pdf)\.Ph\.D\. thesis, Stanford University\.
- Nie et al\. \(2020\)Yixin Nie, Xiang Zhou, and Mohit Bansal\. 2020\.[What can we learn from collective human opinions on natural language inference data?](https://doi.org/10.18653/v1/2020.emnlp-main.734)In*Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\)*, pages 9131–9143, Online\. Association for Computational Linguistics\.
- Partee \(1989\)Barbara H\. Partee\. 1989\.Many quantifiers\.In*Proceedings of the 5th Eastern States Conference on Linguistics*, pages 383–402, Columbus, OH\. Ohio State University\.Reprinted in Compositionality in Formal Semantics: Selected Papers by Barbara H\. Partee, pages 241–258, Blackwell, Oxford, 2004\.
- Pavlick and Kwiatkowski \(2019\)Ellie Pavlick and Tom Kwiatkowski\. 2019\.[Inherent disagreements in human textual inferences](https://doi.org/10.1162/tacl_a_00293)\.*Transactions of the Association for Computational Linguistics*, 7:677–694\.
- Plank \(2022\)Barbara Plank\. 2022\.[The “problem” of human label variation: On ground truth in data, modeling and evaluation](https://doi.org/10.18653/v1/2022.emnlp-main.731)\.In*Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing*, pages 10671–10682, Abu Dhabi, United Arab Emirates\. Association for Computational Linguistics\.
- Rett \(2018\)Jessica Rett\. 2018\.[The semantics of many, much, few, and little](https://doi.org/10.1111/lnc3.12269)\.*Language and Linguistics Compass*, 12\(1\):e12269\.
- von Fintel \(1999\)Kai von Fintel\. 1999\.[NPI licensing, Strawson entailment, and context dependency](https://doi.org/10.1093/jos/16.2.97)\.*Journal of Semantics*, 16\(2\):97–148\.
- Williams et al\. \(2018\)Adina Williams, Nikita Nangia, and Samuel R\. Bowman\. 2018\.[A broad\-coverage challenge corpus for sentence understanding through inference](https://doi.org/10.18653/v1/N18-1101)\.In*Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long Papers\)*, pages 1112–1122, New Orleans, Louisiana\. Association for Computational Linguistics\.
- Yanaka et al\. \(2019\)Hitomi Yanaka, Koji Mineshima, Daisuke Bekki, Kentaro Inui, Satoshi Sekine, Lasha Abzianidze, and Johan Bos\. 2019\.[Can neural networks understand monotonicity reasoning?](https://doi.org/10.18653/v1/W19-4804)In*Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP*, pages 31–40, Florence, Italy\. Association for Computational Linguistics\.

## Appendix ARegistration\-versus\-Result Concordance

The preregistration is section 1 of the plan document written into the working repository and frozen on 2026\-07\-17; the item numbers below refer to it, and the Registration column reproduces what each item registered, so the concordance can be read without consulting that document\. Status values: “as registered” \(followed without change\), “prediction rejected” \(procedure followed, registered direction failed\), “documented deviation” \(change made during the study, recorded in the analysis log before or at the point of change\)\. Tables[2](https://arxiv.org/html/2607.19231#A1.T2)and[3](https://arxiv.org/html/2607.19231#A1.T3)give the full concordance\.

Table 2:Registration\-versus\-result concordance, registered items 1\.1 to 1\.4\.Table 3:Registration\-versus\-result concordance, registered items 1\.5 to 1\.10\.
## Appendix BBridge Details and Overlap Arithmetic

Table 4:ChaosNLI overlap arithmetic\. Every matched item falls at a single agreement level, which is why the registered bridge correlation is undefined: the agreement ordinal has zero variance on the overlap\.Table 5:ChaosNLI 100\-label entropy inside the constant\-ordinal overlap\. Substantial graded variation survives within a single agreement level, which is the roughness the ordinal cannot express\.Table 6:Share of non\-upward hypotheses at each agreement level, with Wilson 95% confidence intervals\. The share does not decrease as agreement rises\. Plotted as Figure[2](https://arxiv.org/html/2607.19231#S4.F2)\.The original bridge design \(correlating five\-label agreement with ChaosNLI 100\-label entropy on the overlap\) is undefined by arithmetic: ChaosNLI selected items whose majority label carries exactly three votes, so the agreement ordinal has zero variance on the overlap\. The overlap arithmetic confirms this exactly\. All 1,599 ChaosNLI\-M rows match MNLI matched dev items at level 3/5, and the 1,514 ChaosNLI\-S rows decompose into 1,507 five\-label items at 3/5 plus 7 four\-label items at 3/4; the latter fall under the registered four\-label exclusion, leaving 1,507 SNLI bridge items \(Table[4](https://arxiv.org/html/2607.19231#A2.T4)\)\.

The registered replacement has three components\. Component 1 reports the spread of ChaosNLI 100\-label entropy inside the constant\-ordinal overlap, measuring the coarseness of the ordinal \(Table[5](https://arxiv.org/html/2607.19231#A2.T5)\): the SNLI overlap has mean entropy 0\.800 \(sd 0\.332\) and the matched overlap 1\.072 \(sd 0\.267\), so substantial graded variation survives within a single ordinal level\. Component 2 reports level\-wise non\-upward shares with Wilson CIs \(Table[6](https://arxiv.org/html/2607.19231#A2.T6), plotted as Figure[2](https://arxiv.org/html/2607.19231#S4.F2)\)\. Component 3 is this transparency statement itself\.

## Appendix CFull Tier 2 Grid

ScenarioRatennvalidnndegen\.Meanδ\\delta2\.5th pct97\.5th pctSig\. ret\.MNLI matcheddirectional\_non\_to\_up0\.0500100000\.04390\.03810\.04921\.000directional\_non\_to\_up0\.0878 \(anchor\)100000\.04300\.03540\.05101\.000directional\_non\_to\_up0\.1000100000\.04300\.03520\.05141\.000directional\_non\_to\_up0\.2000100000\.04090\.02870\.05290\.990directional\_up\_to\_non0\.0500100000\.04070\.03340\.04831\.000directional\_up\_to\_non0\.1000100000\.03730\.02710\.04710\.995directional\_up\_to\_non0\.1260 \(anchor\)100000\.03620\.02580\.04630\.992directional\_up\_to\_non0\.2000100000\.03200\.01880\.04540\.909symmetric0\.0500100000\.03960\.03060\.05011\.000symmetric0\.1000100000\.03450\.02180\.04700\.946symmetric0\.1007 \(anchor\)100000\.03450\.02200\.04650\.949symmetric0\.2000100000\.02490\.00740\.03980\.594MNLI mismatcheddirectional\_non\_to\_up0\.0500100000\.05730\.05140\.06311\.000directional\_non\_to\_up0\.0878 \(anchor\)100000\.05620\.04860\.06351\.000directional\_non\_to\_up0\.1000100000\.05570\.04770\.06331\.000directional\_non\_to\_up0\.2000100000\.05310\.04130\.06391\.000directional\_up\_to\_non0\.0500100000\.05370\.04590\.06071\.000directional\_up\_to\_non0\.1000100000\.04900\.03910\.05881\.000directional\_up\_to\_non0\.1260 \(anchor\)100000\.04730\.03550\.05811\.000directional\_up\_to\_non0\.2000100000\.04230\.02840\.05500\.998symmetric0\.0500100000\.05180\.04210\.06071\.000symmetric0\.1000100000\.04540\.03200\.05811\.000symmetric0\.1007 \(anchor\)100000\.04500\.03220\.05771\.000symmetric0\.2000100000\.03300\.01640\.04990\.933SNLIdirectional\_non\_to\_up0\.0500100000\.03170\.02160\.04130\.000directional\_non\_to\_up0\.0878 \(anchor\)100000\.03170\.01900\.04540\.001directional\_non\_to\_up0\.1000100000\.03190\.01660\.04650\.000directional\_non\_to\_up0\.2000100000\.03170\.00910\.05470\.016directional\_up\_to\_non0\.0500100000\.0172\-0\.00650\.04130\.038directional\_up\_to\_non0\.1000100000\.0120\-0\.01080\.03420\.037directional\_up\_to\_non0\.1260 \(anchor\)100000\.0103\-0\.01200\.03400\.044directional\_up\_to\_non0\.2000100000\.0070\-0\.01350\.02700\.025symmetric0\.0500100000\.0161\-0\.00800\.03980\.030symmetric0\.1000100000\.0108\-0\.01310\.03530\.041symmetric0\.1007 \(anchor\)100000\.0108\-0\.01470\.03650\.045symmetric0\.2000100000\.0050\-0\.01810\.02700\.026Table 7:Tier 2 misclassification grid, all 36 cells\. Seed 20260717, 1000 replications per cell\. Rates marked \(anchor\) are the MED\-anchored rate for that scenario\. The maximum 97\.5th percentile across all cells is 0\.064, below the registered SESOI of\|δ\|=0\.10\|\\delta\|=0\.10\.Table[7](https://arxiv.org/html/2607.19231#A3.T7)lists all 36 cells: three corpora, three flip scenarios \(symmetric, upward\-to\-non\-upward only, reverse\), and four rates per scenario arm \(fixed 0\.05, 0\.10, 0\.20 plus the MED\-anchored rate for that scenario\), with 1,000 replications per cell at seed 20260717\. Reported per cell: mean simulatedδ\\delta, the 2\.5th and 97\.5th percentiles, and significance retention\. The maximum 97\.5th percentile across all cells is 0\.064, which is the basis for the claim in Section[5](https://arxiv.org/html/2607.19231#S5)that the below\-SESOI conclusion does not flip under simulated misclassification\.

## Appendix DTier 3 Codebook Summary and Disagreements

Table 8:Tier 3 agreement between the manual codebook labels and the tagger, with per\-type disagreement counts\. Directions read manual to tagger\. Agreement is tagger versus human, not inter\-human\. The rows run down the left panel and continue in the right\.ItemCellManualTaggerTriggersHypothesisdownward \-\> upward \(7 items\)7chaosnli\_sdownwardupward\(none detected\)50 people walked to the wrong subway hall\.34chaosnli\_mdownwardupward\(none detected\)It is better to plant when it is colder\.38mnli\_matcheddownwardupward\(none detected\)It’s impossible to have a plate hand\-painted to your own design in Hong Kong\.47mnli\_matcheddownwardupward\(none detected\)European members of NATO might consider the US’s efforts to be less credible\.143chaosnli\_mdownwardupward\(none detected\)Was it Jane Eyre or not?146chaosnli\_mdownwardupward\(none detected\)They would get upset whenever anyone would speak to them\.173mnli\_matcheddownwardupward\(none detected\)We went to the office to see if there was anything we could rent\.upward \-\> non\_monotone \(6 items\)21mnli\_matchedupwardnon\_monotoneonly \(non\_monotone\)Zelon is the only student\-run chapter of the American Civil Liberties Union in New York state\.58chaosnli\_mupwardnon\_monotoneonly \(non\_monotone\)Lister and Simpson were the only ones to use carbolic acid and chloroform for this purpose\.121mnli\_matchedupwardnon\_monotonelast \(non\_monotone\)Your mistress wrote letters last night\.128chaosnli\_mupwardnon\_monotonemany \(non\_monotone\)Our higher\-end stores have been suffering due to the recession, and many have shut down for lack of revenue\.184chaosnli\_supwardnon\_monotoneonly \(non\_monotone\)There are only two people in the field\.188mnli\_matchedupwardnon\_monotonelast \(non\_monotone\)It was 37 degrees last night\.downward \-\> non\_monotone \(3 items\)24chaosnli\_mdownwardnon\_monotoneonly \(non\_monotone\)There were only a few villas the whole way along, until we reached a small village that seemed to be the end\.95mnli\_matcheddownwardnon\_monotoneonly \(non\_monotone\)Price hikes are only possible if there is more money in circulation\.119mnli\_matcheddownwardnon\_monotonesmallest \(non\_monotone\), only \(non\_monotone\)The library is the smallest estate in Jamaica, with only three books\.upward \-\> downward \(3 items\)3chaosnli\_supwarddownwardeach \(downward\)Two dogs chase each other in the high grass\.66chaosnli\_mupwarddownwardn’t \(downward\)I wish you hadn’t revealed your identity, that was a mistake\.181snli\_devupwarddownwardlittle \(downward\)A little girl is blowing the petals\.mixed \-\> downward \(2 items\)82mnli\_matchedmixeddownwardall \(downward\)All of the homes in the hillside have been converted into art galleries and shops selling collectibles\.112mnli\_matchedmixeddownwardn’t \(downward\)I chose to become an actor, but I wasn’t very good at it\.mixed \-\> upward \(2 items\)65mnli\_matchedmixedupward\(none detected\)Everything can be found inside a shopping mall\.111chaosnli\_smixedupward\(none detected\)The men are higher than the wall\.downward \-\> mixed \(1 items\)133chaosnli\_mdownwardmixedonly \(non\_monotone\), barely \(downward\)The tunnel of Eupalinos is only one foot in diameter, barely large enough for a child to squeeze through\.non\_monotone \-\> upward \(1 items\)166chaosnli\_mnon\_monotoneupward\(none detected\)Although it was unnecessary, some of the equipment was adjacent\.Table 9:All 25 Tier 3 disagreement items, grouped by disagreement type\. Group headings read manual label to tagger label\. These items are the basis for the three\-layer decomposition in Section[5](https://arxiv.org/html/2607.19231#S5)\.The codebook \(v1, 2026\-07\-19\) fixes the judgment target: whether the hypothesis sentence contains a monotonicity trigger, judged by semantic definition \(does the expression license or block upward or downward substitution in its scope\) rather than by membership in a word list\. It defines the four\-class label set \(upward, downward, non\-monotone, mixed\), an existence\-based decision rule, boundary\-case guidance, and a circularity note: the codebook was written from semantic definitions, not from the tagger’s rule inventory, so agreement with the tagger is not built in\.

Table[8](https://arxiv.org/html/2607.19231#A4.T8)reports the full agreement metrics for the codebook round, including per\-type disagreement counts\. Table[9](https://arxiv.org/html/2607.19231#A4.T9)lists all 25 disagreement items verbatim \(cell, manual label, tagger label, detected triggers, hypothesis text\), grouped by disagreement type\. These 25 items are the basis for the three\-layer decomposition in Section[5](https://arxiv.org/html/2607.19231#S5)\.

Similar Articles

Demographic Prompting at Scale: When More Attributes Hurt LLM--Human Agreement

arXiv cs.CL

This paper investigates how adding demographic attributes in prompts affects LLM-human agreement across tasks, finding that while a few high-signal attributes improve alignment, over-specification degrades it. The study uses five open-source LLMs and neuron probing to show that attribute signal quality and coherence matter more than quantity.

Sample-Size Scaling of the African Languages NLI Evaluation

arXiv cs.CL

This paper examines the effect of labeled data size on natural language inference performance for 16 African languages using the AfriXNLI benchmark. The results show that scaling behavior is language-sensitive and often non-monotonic, challenging the common assumption of monotonic improvement, and emphasizing the need for language-specific dataset creation and stronger multilingual strategies.

Hidden Consensus:Preference-Validity Compression in Human Feedback

arXiv cs.CL

This paper argues that standard RLHF's scalarization of human preferences collapses multiple valid interpretations into a single target, mis-measuring alignment in culturally plural societies. Analyzing a Malaysian dataset, they find 79% of prompts have multiple majority-supported responses that single-winner aggregation discards.