Structured Output Collapses Answer Diversity Across 44 Language Models

arXiv cs.CL Papers

Summary

A study shows that when LLMs are asked to output in JSON format, their answer diversity collapses significantly compared to plain chat, with modal answers becoming more common and distinctive models losing half their uniqueness.

arXiv:2607.18476v1 Announce Type: new Abstract: When a language model must choose one answer from a large space of equally valid options, a format clause -- "Reply with JSON only" -- changes which answer it chooses. We re-run the One-Word Census (arXiv:2607.12796): 31 wide-answer-space category prompts asked of 44 models, now with the reply requested in JSON -- no schema enforcement, no constrained decoding, only the request. Convergence deepens sharply: on the unconstrained "Pick a word" prompt the modal answer rises from 41% to 64% of the pool and distinct answers fall from 52 to 36; mean answer-choice surprisal drops from 1.80 to 1.58 bits. The tax is progressive: six of 44 models move individually (BH-FDR q=.10), all toward the mode, led by the most distinctive models, while the conformist floor is immobile. It is a sharpener, not a re-indexer -- the plain-chat modal answer survives in 28 of 31 categories. Defaults are register-indexed: a within-run re-sample (n=20) finds JSON shifts 53% of a model's stable chat defaults, mostly back to the crowd, and installs defaults absent from chat (Claude Fable 5 answers "cerulean" for colour 0% of the time in chat, 100% in JSON). Full-battery controls reveal a register gradient: compression is significant and specific to the answer-delivery formats models are trained to speak (JSON -0.22 bits, p=.0002; XML -0.19, p=.002), absent for YAML and CSV, and reversed for an arbitrary bracket wrapper (+0.13, p=.009) -- weighing the mechanism toward tool-use post-training. Enforcing the schema at the decoder (response_format) compresses no further than the request (-0.03 bits): the collapse lives in the model's response to the register, not the decoder. Structured output is how software consumes language models, and that surface is served by a measurably more homogeneous model than the chat surface on which models are evaluated, compared, and chosen.
Original Article
View Cached Full Text

Cached at: 07/22/26, 08:23 AM

# Structured Output Collapses Answer Diversity Across 44 Language Models
Source: [https://arxiv.org/html/2607.18476](https://arxiv.org/html/2607.18476)
\(July 2026\)

###### Abstract

When a language model must choose one answer from a large space of equally valid options, a format clause —*“Reply with JSON only”*— changes which answer it chooses\. We re\-run the One\-Word Census\[[7](https://arxiv.org/html/2607.18476#bib.bib39)\], 31 category prompts with wide answer spaces asked of 44 models, with the reply requested in JSON: no schema enforcement, no constrained decoding, only the request\. The field’s convergence deepens sharply: on the unconstrained prompt \(“Pick a word\.”\), the modal answer’s share rises from 41% to 64% and distinct answers fall from 52 to 36; mean answer\-choice surprisal drops 1\.80 to 1\.58 bits\. The tax is*progressive*and respects the instrument’s resolution: six of 44 models move individually \(BH\-FDRq=\.10q=\.10\), all toward the mode, led by the panel’s most distinctive frontier model \(the strongest explorer loses 1\.31 bits — half its measured distinctiveness\)\. The format is a sharpener, not a re\-indexer: the plain\-chat modal answer survives in 28 of 31 categories, the three flips landing where the plain mode was weakest\.*Defaults are register\-indexed*: a within\-run re\-sample \(n=20n\{=\}20\) finds JSON shifts 53% of a model’s stable chat defaults, mostly back to the crowd, and*installs*defaults absent from chat \(Claude Fable 5 answers*cerulean*for colour 0% of the time in chat, 100% in JSON\)\. Full\-battery controls reveal a register gradient: compression is significant and specific to the answer\-delivery formats models are trained to speak \(JSON−0\.22\-0\.22bits,p=\.0002p=\.0002; XML−0\.19\-0\.19,p=\.002p=\.002; permutation over models and categories\), absent for YAML and CSV, and*reversed*for an arbitrary bracket wrapper \(\+0\.13\+0\.13,p=\.009p=\.009\) — weighing the mechanism toward tool\-use post\-training, though every serialization concentrates the unconstrained pool, so a corpus\-register component remains\. Enforcing the schema at the decoder \(response\_format\) compresses no further than the request \(−0\.03\-0\.03bits\): the collapse lives in the model’s response to the register, not the decoder\. Structured output is how software consumes language models; that surface is served by a measurably more homogeneous model than the chat surface on which models are evaluated, compared, and chosen\.

## 1Introduction

Software does not consume language models in prose\. It asks for JSON\. Every agent that calls a tool, every pipeline that extracts a field, classifies a record, fills a schema, or routes a request receives the model’s answer inside a data structure — and that traffic increasingly dwarfs the chat window in which models are benchmarked, ranked, and chosen\. This paper asks a question about that surface: when the same model answers the same question in a serialization format instead of prose, does it give the same answer?

The companion One\-Word Census\[[7](https://arxiv.org/html/2607.18476#bib.bib39)\]built a deliberately narrow instrument\. When a model must choose one answer from a large space of equally valid options —*“Name a tree\. Reply with one word only\.”*— which does it choose, and how often is it the same answer the rest of the field chooses? The metric is*answer\-choice surprisal*: a leave\-one\-out measure, in bits, of how unlikely a model’s answers are under the pooled answers of every other model, computed by exact match on normalized one\-word replies — no embeddings, no LLM judge, no human annotation\. Across 44 models it found a monoculture \(*oak*takes 94% of tree answers;*serendipity*41% of answers to an unconstrained “Pick a word\.”\) structured by a wide per\-model conformity range and, for a handful of models, stable off\-modal defaults that constitute a measurable character\.

We re\-run that instrument unchanged in every respect but one: the reply is requested in a serialization format\.*“Name a tree\. Reply with JSON only, in the form\{"word": "<your answer\>"\}\.”*No schema enforcement, no constrained decoding, no token masking — only the request\. The 31 prompts, the 44\-model panel, and the conditions \(no system prompt, requested temperature 1\.0, four samples\) are identical to the census; the single manipulated variable is the register in which the answer is delivered\. Because surprisal is computed*within*each format column — JSON answers scored against the JSON field — a column’s convergence is an internal property of that column, not an artifact of comparing one register against another\.

Our contributions:

1. 1\.A progressive format tax\(§[4\.1](https://arxiv.org/html/2607.18476#S4.SS1)\)\. Field\-mean surprisal falls from 1\.80 to 1\.58 bits \(−0\.22\-0\.22per model,p=\.0002p=\.0002\), but the drop is not uniform: six of 44 models move individually \(BH\-FDRq=\.10q=\.10\), all toward the mode, led by the panel’s most distinctive models \(the strongest explorer loses 1\.31 bits\), while the conformist floor is immobile and part of the positive\-Δ\\Deltatail is register\-invariant stranding rather than divergence — a panel\-free check separates the two\. The remaining deltas form a noise plateau we decline to rank\.
2. 2\.A sharpener, not a re\-indexer\(§[4\.2](https://arxiv.org/html/2607.18476#S4.SS2)\)\. The plain\-chat modal answer survives in 28 of 31 categories; the three flips occur exactly where the plain mode was weakest\.
3. 3\.Register\-indexed defaults\(§[4\.4](https://arxiv.org/html/2607.18476#S4.SS4)\)\. A within\-run re\-sample \(n=20n\{=\}20\) shows the register significantly shifts 53% of a model’s stable chat defaults \(mostly back to the crowd\) and*installs*register\-only defaults absent from chat — read from discrete answer behavior, not noisy score deltas\.
4. 4\.A register gradient\(§[4\.5](https://arxiv.org/html/2607.18476#S4.SS5)\)\. Compression is significant and specific to the answer\-delivery formats models are trained to speak \(JSON, XML\), absent for YAML and CSV, and*reversed*for an arbitrary non\-data wrapper — weighing the mechanism toward tool\-use post\-training while leaving a corpus\-register component visible\.
5. 5\.A cheap, reusable method\. The manipulation is one extra column on a frozen instrument, and format compliance falls out of every run as a standing capability metric\.

## 2Related work

#### Format restrictions and accuracy\.

The closest prior work asks whether format*restrictions*degrade task performance\.Tamet al\.\[[10](https://arxiv.org/html/2607.18476#bib.bib32)\]report that requiring structured output lowers reasoning accuracy;Kurt \[[3](https://arxiv.org/html/2607.18476#bib.bib33)\]contest the result, attributing it to prompt design and parsing rather than the format itself\. Our question is orthogonal to this debate: our prompts have no wrong answers, so there is no accuracy to degrade\. We measure which valid answer is chosen, not whether a correct one is reached\.

#### Decoder\-level enforcement\.

A 2026 line names the cost of*enforcing*structure at the decoder — the “constraint tax”\[[8](https://arxiv.org/html/2607.18476#bib.bib34),[5](https://arxiv.org/html/2607.18476#bib.bib35)\]and “format tax”\[[4](https://arxiv.org/html/2607.18476#bib.bib36)\]: grammar\-constrained decoding raises validity while lowering correctness and suppressing behaviors such as tool calls\. But they constrain the sampler with token masking; we mask nothing and only ask\. The effect we measure therefore lives in the model’s response to a register, not in a decoding constraint\.

#### Diversity within grammars\.

Luanet al\.\[[6](https://arxiv.org/html/2607.18476#bib.bib37)\]engineer diversity back into constrained generation with automata\-based steering\. That work presupposes exactly the collapse we characterize — it is an engineering fix for a phenomenon the present paper measures at the level of a mere request, with no grammar in the loop\.

#### Surface form and register\.

Sclaret al\.\[[9](https://arxiv.org/html/2607.18476#bib.bib38)\]show model accuracy is sensitive to spurious formatting choices; the census showed answer*choice*is sensitive to wording \(*gemstone*vs\.*precious stone*\)\. This paper adds a third axis: not how the question is worded but the register in which the answer is requested\. Underlying all of this is the mode\-collapse literature\[[1](https://arxiv.org/html/2607.18476#bib.bib1),[11](https://arxiv.org/html/2607.18476#bib.bib3),[2](https://arxiv.org/html/2607.18476#bib.bib5)\], which documents the convergence we exploit and notes that LLM judges systematically prefer modal outputs — one reason our instrument stays judge\-free\.

## 3Method

### 3\.1Instrument

The instrument is the*one\-word census*unchanged\[[7](https://arxiv.org/html/2607.18476#bib.bib39)\]: 31 single\-turn prompts \(30 category prompts of the form*“Name a\[n\]XX”*plus one unconstrained*“Pick a word”*\), no system prompt, requested temperature 1\.0, four samples per cell\. Each answer is scored by its add\-one\-smoothed leave\-one\-out answer\-choice surprisal against the pooled answers of the other 43 models, and a model’s headline score is the mean over its valid answers, in bits\. Replies are reduced to a single normalized token by exact\-match rules and a mechanical junk guard \(§[3\.3](https://arxiv.org/html/2607.18476#S3.SS3)\)\.

Surprisal is computed*within*each format column\. JSON answers are scored against the JSON field, plain answers against the plain field\. A column’s convergence is therefore an internal property of that column — it is never a comparison of JSON text against a chat reference, and so cannot be an artifact of the register shift itself\. The per\-model*Δ\\Delta\-surprisal*\(JSON minus plain\) is the headline quantity: how many bits of a model’s distinctiveness the register removes\.

### 3\.2Format columns

The plain\-chat column is the census transcripts\. Each of five format columns re\-runs the full 31\-category battery for all 44 models at four samples per cell \(5,456 calls per column\): JSON, XML, YAML, CSV, and a non\-data “brackets” wrapper, the last included to separate serialization from mere structure\. The clauses are appended verbatim to the census prompt:

- •json:Reply with JSON only, in the form \{"word": "<your answer\>"\}\.
- •xml:Reply with XML only, in the form <word\>your answer</word\>\.
- •yaml:Reply with YAML only, in the form ‘word: <your answer\>‘\.
- •csv:Reply with CSV only: a header row ‘word‘, then one data row with your answer\.
- •brackets:Reply with your answer inside square brackets only, like \[answer\]\.

Noresponse\_formatparameter, no tool schema, and no constrained decoding is used anywhere: every column is an ordinary chat completion whose only difference from the census is the appended clause\. This is a deliberate scope choice: these columns measure the effect of the*request*\. An engine\-enforcedresponse\_formatcolumn is analyzed separately in §[4\.6](https://arxiv.org/html/2607.18476#S4.SS6)\.

### 3\.3Parsing and hygiene

Each reply is reduced to a single answer token in two steps\. First the format wrapper is stripped by a per\-format regular expression — the contents of the JSON"word"field, the text between<word\>tags, and so on — and the extracted string is then passed through the census normalization and junk guard unchanged\. A reply that does not match its format’s wrapper falls back to the census rule applied to the raw text, so that a malformed reply cannot smuggle in a spurious “novel” answer: the last word of a broken\-JSON essay is scored exactly as the census would score it\. One guard is added beyond the census rule, because a fill\-in wrapper invites an artifact the census prompt cannot: a reply whose token is the category noun or the template placeholder — a literal\[city\]for “name a city”, oranswer/wordlifted from the slot — is dropped as a failed cell\. Parse\-plus\-junk survival is at least 99% per model per column, and we audited per\-model survival for every model whose surprisal*rose*, since non\-compliance is the one artifact that manufactures apparent divergence \(§[4\.5](https://arxiv.org/html/2607.18476#S4.SS5)\)\.

Of the 27,280 format cells, 91 \(0\.3%\) are unrecoverable nulls, concentrated in YAML \(56\) and negligible in JSON \(1\); all distributional claims are computed over the surviving cells, with the per\-column parse survival audited above\.

## 4Results

### 4\.1The format tax is progressive

Field\-mean answer\-choice surprisal falls from 1\.80 bits in plain chat to 1\.58 in JSON\. Averaged over the 44 models the per\-modelΔ\\Deltais−0\.22\-0\.22bits, significant on both exchangeable units \(§[4\.5](https://arxiv.org/html/2607.18476#S4.SS5)\), but the effect is concentrated rather than spread\. The shape is a divergent tail collapsing onto a fixed floor \(Figure[1](https://arxiv.org/html/2607.18476#S4.F1)\)\.

The models that lose the most are the ones that had the most to lose\.deepseek\-v3\.2, the census’s one genuine explorer among frontier\-lab models, falls from 2\.63 to 1\.32 bits \(−1\.31\-1\.31\) — half its measured distinctiveness — the moment the answer is requested as JSON\.hermes\-4loses 0\.91,gpt\-4o\-mini0\.92,gpt\-4\-turbo0\.87,gpt\-5\.6\-sol0\.65\. The conformist floor barely moves:claude\-opus\-4\.8−0\.16\-0\.16,grok\-4\.5−0\.07\-0\.07,claude\-sonnet\-5−0\.05\-0\.05\. There is nowhere below the floor to go, and the tail falls toward it, though the raw range is nearly unchanged \(2\.2 to 2\.0 bits\): a register\-invariant minority holds its ground while the field collapses around it \(§[4\.5](https://arxiv.org/html/2607.18476#S4.SS5)\), so the compression is in the mean and the collapsing tail, not the extremes\. On the unconstrained “Pick a word” prompt the same collapse appears with no category to anchor it:*serendipity*rises from 41% of the pool to 64% and the number of distinct words falls from 52 to 36\.

Which individual movements are real we take up in §[4\.5](https://arxiv.org/html/2607.18476#S4.SS5); the progressive*shape*— tail down, floor fixed — is a property of the aggregate and does not depend on ranking the middle of the table\. Nor is the shape regression to the mean, the classic alternative for “the ones that lost most had the most to lose”: split\-half test\-retest reliability of the census score isr=0\.94r\{=\}0\.94, so the shrinkage a pure re\-measurement predicts for the most distinctive model is0\.050\.05bits against the1\.311\.31observed, and every individually significant compressor \(§[4\.5](https://arxiv.org/html/2607.18476#S4.SS5)\) clears the re\-measurement\-noise null by3\.73\.7to10​σ10\\sigma\.

![Refer to caption](https://arxiv.org/html/2607.18476v1/x1.png)Figure 1:Answer\-choice surprisal for all 44 models, open chat \(gray dot\) versus JSON \(colored dot\), sorted by the open\-chat score\. Blue marks a model that becomes*more generic*under JSON \(31 of 44\); amber marks one whose relative score*rises*\(13 — mostly register\-invariant models the collapsing field strands, §[4\.5](https://arxiv.org/html/2607.18476#S4.SS5)\)\. The distinctive models at the top slide farthest toward the conformist floor, which is itself immobile, and the field mean falls from 1\.80 to 1\.58 bits\. The raw range is nearly unchanged \(2\.2→2\.02\.2\\to 2\.0\) because the amber models hold up the top: the compression is in the mean and the tail, not the extremes\.
### 4\.2A sharpener, not a re\-indexer

If the register merely added noise, or merely imposed structure, it could move probability anywhere\. Instead it sharpens the distribution the model already had\. In 28 of 31 categories the JSON modal answer is the same word as the plain\-chat mode, and less than 4% of JSON answer\-mass lands on words the plain census never produced\. The register moves mass*onto the existing mode*, not onto new answers\.

The three exceptions are diagnostic: each is a category whose plain\-chat mode was weak\. Insect flips from*ant*\(43%\) to*butterfly*\(60%\), board game from*chess*to*monopoly*, dance from*salsa*to*tango*\. The register amplifies its own conditional prior, and that prior wins only where the chat prototype was too weak to hold\.

### 4\.3The register moves the mode, not the temperature

A natural worry is that we are measuring a sampling change rather than a change in preference — that some providers quietly cool the sampler when a request looks like structured output\. The census’s self\-distinctness statistic \(distinct answers divided by samples, within model and category\) is the effective\-temperature proxy that settles it\. At the field level it is nearly flat across all six columns — plain 0\.42, JSON 0\.39, XML 0\.40, YAML 0\.42, CSV 0\.43, brackets 0\.43 — while surprisal drops 0\.22 bits under JSON\. A serving\-layer temperature cut would crater self\-distinctness wherever it cut; it does not\. The collapse is*positional*: mass relocates onto the field’s mode without the model sampling any less around it\.

The within\-model narrowing that does exist is progressive, on the same models as the surprisal effect:hermes0\.71→0\.520\.71\\to 0\.52,deepseek\-v3\.20\.56→0\.420\.56\\to 0\.42,sonar0\.50→0\.400\.50\\to 0\.40,haiku\-4\.50\.42→0\.330\.42\\to 0\.33, while the conformist floor is flat \(sonnet\-50\.30→0\.270\.30\\to 0\.27\) andclaude\-fable\-5holds at 0\.32 in both registers — stable self\-distinctness rather than a cooled sampler\. The tax is thus progressive on both of the census’s axes, and the dissociation between a large surprisal move and a small self\-distinctness move is itself the signature of a reshaped conditional distribution rather than a rescaled one\.

### 4\.4Defaults are register\-indexed

The sharpest evidence that personality is register\-dependent comes not from surprisal scores, which are noisy for individual models \(§[4\.5](https://arxiv.org/html/2607.18476#S4.SS5)\), but from discrete answer behavior\. Call a model’s*stable default*in a category an off\-modal answer it gives repeatedly — a standing refusal of the field’s first choice, likefable’s*gouda*for cheese\. The four\-sample census flags 144 such defaults across the panel\. Four samples, however, cannot tell a genuine default from a lucky streak, nor a register effect from ordinary run\-to\-run drift, so we re\-sampled every default cell atn=20n=20in*both*registers within a single run\. This settles three questions the four\-sample data could not\.

#### The defaults are real, and the register erodes them\.

Atn=20n=20the flagged defaults are genuine \(median per\-sample probability0\.900\.90\), not sampling artifacts\. Requesting JSON then*significantly*shifts the answer distribution for 76 of 144 \(53%\) of them \(Fisher exact per cell, Benjamini–Hochbergq=\.10q=\.10\), against a false\-positive floor of∼10%\{\\sim\}10\\%, and the shift is directional: 29% revert outright to the field’s modal answer\. The register pulls a model’s idiosyncratic default back toward the crowd\.

#### The register also installs defaults\.

The move is not only subtractive\. Of the JSON answers a model gives four\-of\-four but*never*produces in chat, 81% survive then=20n=20re\-sample as genuine register\-only defaults:fablesays*cerulean*for colour 0% of the time in chat and 100% in JSON \(p≈10−11p\\approx 10^\{\-11\}\), and*carpenter*for occupation0%→90%0\\%\\to 90\\%\(p≈10−9p\\approx 10^\{\-9\}\);opus\-4\.8acquires*phoenix*,sonnet\-4\.6*gold*\. Some of the field\-level mode flips of §[4\.2](https://arxiv.org/html/2607.18476#S4.SS2)are these acquisitions in aggregate — one model’s JSON\-only*butterfly*\(15%→100%15\\%\\to 100\\%\) or*monopoly*\(5%→100%5\\%\\to 100\\%\)\.

#### Same default, opposite response\.

Because the register both erases and installs, models that look identical in chat diverge under it\.gpt\-5\.6\-terraandfableboth answer*mango*for fruit, twenty runs of twenty in chat; under JSONterraflips to*apple*\(20/20,p≈10−11p\\approx 10^\{\-11\}\) whilefableholds*mango*\(19/20\)\. The right object is therefore not one personality the register reveals or hides but a*per\-register defaults profile*: a model’s character is indexed to the channel it is asked through, the same way the census found answers indexed to the wording of the question\.

### 4\.5Serialization vs\. structure: the register gradient

#### The gradient\.

Extending the same battery to four further formats separates*serialization*from mere*structure*, and the result is a gradient rather than a uniform effect\. Conditioning on compliance — restricting each format to the models that produce its wrapper at least 90% of the time, so incapacity stays out of the distributional signal — the field\-meanΔ\\Delta\-surprisal is−0\.22\-0\.22bits for JSON,−0\.23\-0\.23for XML,−0\.09\-0\.09for YAML,−0\.09\-0\.09for CSV, and\+0\.13\+0\.13for brackets\. The two formats that compress the field are the two models are trained to*answer*in: JSON and XML, the registers of function calls, tool use, and structured\-output modes\. The data formats models mostly only*read*— YAML and CSV — do not reliably compress it, and an arbitrary non\-data wrapper \(square brackets\)*loosens*it\.

#### Significance\.

We test both exchangeable units with a sign\-flip permutation \(20,000 draws\): field entropy paired over the 31 categories, and within\-column surprisal paired over the 44 models\. JSON and XML compress on both units \(entropy−0\.20\-0\.20,p=\.0006p=\.0006and−0\.20\-0\.20,p=\.0004p=\.0004; per model−0\.22\-0\.22,p=\.0002p=\.0002and−0\.19\-0\.19,p=\.002p=\.002\)\. YAML and CSV are not significant on either unit \(p=\.75p=\.75and\.46\.46by category;\.80\.80and\.35\.35by model\)\. Brackets significantly*loosens*on both \(\+0\.12\+0\.12entropy,p=\.014p=\.014;\+0\.13\+0\.13per model,p=\.009p=\.009\), and the loosening survives every echo\-guard variant —\+0\.16\+0\.16\(no guard\) through\+0\.13\+0\.13\(full guard\) per model,p<\.01p<\.01throughout — so it is a real reversal, not a hygiene artifact\. Why an arbitrary wrapper should*widen*the field we can only speculate: a bare\[answer\]slot reads less like a data record than a fill\-in\-the\-blank, a framing that may invite a more playful completion than prose does; we report the reversal as a settled effect without a settled mechanism\. The JSON\-versus\-YAML difference, paired by category, is itself significant \(p<10−4p<10^\{\-4\}\), anddeepseek\-v3\.2’s individual collapse clearsp<10−4p<10^\{\-4\}alone\. The compression is thus specific to the trained answer\-delivery formats, which moves the weight of the mechanism toward tool\-use post\-training rather than a generic property of serialized text\. The gradient is not an artifact of the differing per\-format compliance subsets \(compliantnn: JSON 43, XML 41, YAML 37, CSV 39, brackets 43 of 44\): on the 34 models compliant in*all*five formats it reproduces and sharpens \(JSON−0\.27\-0\.27, XML−0\.26\-0\.26, brackets\+0\.13\+0\.13\), while YAML and CSV stay weakly negative and non\-significant\.

#### But not post\-training alone\.

The corpus\-register account is not dead, because on the unconstrained “Pick a word” prompt*every*serialization concentrates the pool, including the two with no net battery effect:*serendipity*rises from 41% in plain chat to 52% \(CSV\), 59% \(XML\), 64% \(JSON\), and 65% \(YAML\), while brackets holds at 41%\. YAML’s single\-category concentration is real; it simply washes out across the full battery, which is why its net effect is null\. A black\-box study cannot fully separate what the training corpus taught — that text inside a data structure is the canonical value — from what tool\-use tuning reflexively reinforced; our reading is that both are present and the post\-training component dominates the battery\-wide gradient\.

#### Two companion phenomena\.

Compliance is not a nuisance to discard but a control that must be applied\.*Format incompetence*masquerades as divergence:graniteandmythomaxemit a valid CSV wrapper only about 1% of the time, and their unwrapped replies, scored naively, read as high surprisal —graniteshows a spurious\+1\.49\+1\.49bits in CSV before conditioning\. This is why every distributional figure above is compliance\-conditioned\. Distinct from incompetence, a fully\-compliant model can still carry a positiveΔ\\Deltathat is real signal rather than a parse artifact:llama\-4\-maverickis 100% compliant in JSON and CSV yet sits\+0\.4\+0\.4to\+0\.7\+0\.7bits above the field in every format — not format failure but the*stranding*described below, the field moving while the model holds still\. Register reaction is heterogeneous, not uniformly compressive\.

#### Per\-model deltas: three tiers\.

Applying the census’s tiers\-not\-ranks discipline to the deltas: \(a\) exactly six of 44 models have an individually significantΔ\\Deltaat BH\-FDRq=\.10q=\.10, and all six are compressions —deepseek\-v3\.2−1\.31\-1\.31,gpt\-4o\-mini−0\.92\-0\.92,gpt\-4\-turbo−0\.87\-0\.87,gpt\-5\.6\-sol−0\.65\-0\.65,qwen3−0\.45\-0\.45,gemini\-3\.1\-pro−0\.31\-0\.31\(hermes,−0\.91\-0\.91, falls just short — its within\-model dispersion is the panel’s largest \(self\-distinctness0\.710\.71, §[4\.3](https://arxiv.org/html/2607.18476#S4.SS3)\), which widens its interval enough that a near\-identical delta misses the threshold\)\. \(b\) The remaining∼38\{\\sim\}38models — including every model with a positiveΔ\\Delta— are individually indistinguishable from zero, and their mid\-table ordering carries no information; we do not interpret it\.

#### Divergence is often stranding, not motion\.

Because surprisal is scored within\-column against the pool, a model that keeps its answers while the field converges around it earns a risingΔ\\Deltawith no change in its own behavior — passive*stranding*, not active divergence\. A panel\-free check separates the two: for each model we compare its own chat answer distribution to its own JSON distribution \(mean Jensen–Shannon over the 31 categories, never touching the pool\)\. The panel’s largest positive delta,llama\-4\-maverick\(\+0\.56\+0\.56, 100% JSON\-compliant\), has a*below*\-median self\-divergence \(JSD0\.200\.20vs the field’s0\.270\.27; 7 of 31 modal answers change\): it barely moves its own answers, and its rising score is the collapsing field stranding a fixed, distinctive point\. Register\-invariance is thus a model trait — some models code\-switch under serialization, some hold their answers regardless — and it is what most of the amber tail is\. The trait is not lineage\-clean \(mythomax,\+0\.43\+0\.43,*does*move its own answers\), so we read the positive tail as a mix of stranding and genuine divergence the panel\-relativeΔ\\Deltacannot separate, and decline to rank it\.

### 4\.6Enforcement adds little the request did not

A natural objection is that the collapse is a decoding artifact\. It is not: enforcing the schema at the decoder compresses no further than asking for it\. We ran one more column — the JSON clause plus a strictresponse\_formatjson\_schema — on the 36 of 44 models whose providers support it \(the other eight return no valid output, a feature\-acquisition signal in itself\)\. Field\-mean surprisal — plain, request, and enforcement all recomputed on the 36\-model supported subset and scored within that subset’s own pool, so the three are comparable — is 1\.56 bits under the request and 1\.53 under enforcement, against 1\.79 in open chat: the request does the−0\.22\-0\.22\-bit work and enforcement adds−0\.03\-0\.03more\. The effect lives in the model’s response to the register, not in the sampler\. Per\-model reactions are heterogeneous and confounded —response\_formatis a native decoder constraint for some providers and a gateway coercion for others — so we read only the battery\-wide mean and release the data for a native\-only replication\.

## 5Limitations

Several limitations bound these claims\. The data are a single snapshot served through one channel \(OpenRouter\), and requested temperature 1\.0 is not honored uniformly across providers — which is why we report self\-distinctness alongside surprisal and rest the personality claims on discrete four\-of\-four behavior\. Surprisal is panel\-relative: a model’s bits depend on the field it is scored against, so the absolute numbers are not portable across panels, although the within\-column andΔ\\Deltacomparisons are internal to this one\. Each format is probed with a*single*clause wording, so we cannot separate the register from the particular phrasing that invokes it; a clause\-paraphrase column is the natural control\. The gradient does, however, already rule out the most obvious phrasing confounds: all four serialization clauses embed the samewordslot name and a fill\-in template, yet only JSON and XML compress, so neither key\-priming nor the fill\-in framing can be doing the work; and the one residual axis, clause length, runs the wrong way for a length account — CSV carries the wordiest clause and is null\. Finally, we vary the register through prompt\-level*requests*and one enforcedresponse\_formatcounterpart \(§[4\.6](https://arxiv.org/html/2607.18476#S4.SS6)\); tool\-call framing and other structured\-output pathways remain unmeasured, and enforcement is itself served heterogeneously across providers, so a provider\-controlled native\-enforcement replication would sharpen the per\-model picture\.

## 6Discussion

The result has one blunt practical consequence\. Models are evaluated, compared, ranked, and felt out in chat — leaderboards, vibe checks, and human preference are all collected on the prose surface — but software consumes them through structured output, and that surface is measurably more collapsed\. Every diversity number the census reported is, for the deployment tier that increasingly matters, an overstatement: the model an agent consults has narrower tastes than the one a user chatted with, and the gap is a fixed cost of the register rather than a tail risk\. Synthetic\-survey, LLM\-judge, and tool\-selecting pipelines all read the JSON persona\. This bites where the task admits many acceptable answers — surveys, recommendations, brainstorming, judgement, tool choice — not where structured output carries a single correct value \(extraction, classification, routing\), for which reduced diversity is no cost\.

The register is also a cheap instrument in its own right, and format compliance falls out for free as a generational capability track — older models cannot speak some of these registers at all, and that acquisition has a date a re\-run will record\. If the battery\-wide compression we attribute mostly to tool\-use post\-training is being trained in per release, a fixed public instrument run on every model is what would detect it — the same argument the census makes for measuring conformity over time, now pointed at the register\. The immediate next columns are a provider\-controlled native\-enforcement replication and tool\-call framing \(§[4\.6](https://arxiv.org/html/2607.18476#S4.SS6)\) and clause paraphrases\.

## Data and code availability

All code and data are in thestudies/structured/directory of the modelun repository:run\_formats\.py\(the battery runner\),analyze\.py\(the metric, junk guard, andΔ\\Delta\-surprisal\),probe\_significance\.py\(the permutation tests\),analyze\_rtm\_splithalf\.py\(split\-half reliability and the regression\-to\-the\-mean null\),analyze\_register\_invariance\.py\(the panel\-free self\-divergence check\), andanalyze\_brackets\_robustness\.py\(the echo\-guard sweep\), againstprobes/format\_register\.json\(the raw replies for all six columns\)\. The plain\-chat baseline is the census transcripts; this study scores against the 44\-model panel\. An interactive explorer — scorecard, per\-category and per\-format tables, per\-model drill\-downs, and a compliance tab — is published alongside\.

## Note on AI usage

This work was done in collaboration with Claude \(Opus 4\.8 and Fable 5\), which helped run the battery, build the analysis, and draft the text; the research questions and interpretation are the author’s\. Both models are also subjects of the study\.

## References

- \[1\]L\. Jiang, Y\. Chai, M\. Li, M\. Liu, R\. Fok, N\. Dziri, Y\. Tsvetkov, M\. Sap, A\. Albalak, and Y\. Choi\(2025\)Artificial Hivemind: the open\-ended homogeneity of language models \(and beyond\)\.Advances in Neural Information Processing Systems \(NeurIPS\)\.Note:arXiv:2510\.22954Cited by:[§2](https://arxiv.org/html/2607.18476#S2.SS0.SSS0.Px4.p1.1)\.
- \[2\]R\. Kirk, I\. Mediratta, C\. Nalmpantis,et al\.\(2024\)Understanding the effects of RLHF on LLM generalisation and diversity\.InInternational Conference on Learning Representations \(ICLR\),Note:arXiv:2310\.06452Cited by:[§2](https://arxiv.org/html/2607.18476#S2.SS0.SSS0.Px4.p1.1)\.
- \[3\]W\. Kurt\(2024\)Say what you mean: a response to “let me speak freely”\.Note:[https://blog\.dottxt\.ai/say\-what\-you\-mean\.html](https://blog.dottxt.ai/say-what-you-mean.html)Cited by:[§2](https://arxiv.org/html/2607.18476#S2.SS0.SSS0.Px1.p1.1)\.
- \[4\]I\. Y\. Lee, L\. D’Antoni, and T\. Berg\-Kirkpatrick\(2026\)The format tax\.Note:arXiv:2604\.03616Cited by:[§2](https://arxiv.org/html/2607.18476#S2.SS0.SSS0.Px2.p1.1)\.
- \[5\]F\. Li, A\. Zhang, and C\. Lv\(2026\)Constraint tax in open\-weight llms: an empirical study of tool calling suppression under structured output constraints\.Note:arXiv:2606\.25605Cited by:[§2](https://arxiv.org/html/2607.18476#S2.SS0.SSS0.Px2.p1.1)\.
- \[6\]X\. Luan, Z\. Wei, Y\. Zhang, and M\. Sun\(2025\)Automata\-based steering of large language models for diverse structured generation\.Note:arXiv:2511\.11018Cited by:[§2](https://arxiv.org/html/2607.18476#S2.SS0.SSS0.Px3.p1.1)\.
- \[7\]T\. Parikh\(2026\)The one\-word census: answer\-choice conformity across 44 language models\.Note:arXiv:2607\.12796; data and explorer at[https://github\.com/tap2k/modelun](https://github.com/tap2k/modelun)Cited by:[§1](https://arxiv.org/html/2607.18476#S1.p2.1),[§3\.1](https://arxiv.org/html/2607.18476#S3.SS1.p1.1)\.
- \[8\]J\. Ray\(2026\)The constraint tax: measuring validity\-correctness tradeoffs in structured outputs for small language models\.Note:arXiv:2605\.26128Cited by:[§2](https://arxiv.org/html/2607.18476#S2.SS0.SSS0.Px2.p1.1)\.
- \[9\]M\. Sclar, Y\. Choi, Y\. Tsvetkov, and A\. Suhr\(2024\)Quantifying language models’ sensitivity to spurious features in prompt design, or: how i learned to start worrying about prompt formatting\.InICLR,Note:arXiv:2310\.11324Cited by:[§2](https://arxiv.org/html/2607.18476#S2.SS0.SSS0.Px4.p1.1)\.
- \[10\]Z\. R\. Tam, C\. Wu, Y\. Tsai, C\. Lin, H\. Lee, and Y\. Chen\(2024\)Let me speak freely? a study on the impact of format restrictions on performance of large language models\.InProceedings of EMNLP 2024: Industry Track,Note:arXiv:2408\.02442Cited by:[§2](https://arxiv.org/html/2607.18476#S2.SS0.SSS0.Px1.p1.1)\.
- \[11\]Y\. Zhang, H\. Diddee, S\. Holm, H\. Liu, X\. Liu, V\. Samuel, B\. Wang, and D\. Ippolito\(2025\)NoveltyBench: evaluating language models for humanlike diversity\.InConference on Language Modeling \(COLM\),Note:arXiv:2504\.05228Cited by:[§2](https://arxiv.org/html/2607.18476#S2.SS0.SSS0.Px4.p1.1)\.

## Appendix APrompts

The battery is 31 single\-turn prompts, frozen before data collection — the census stimulus unchanged\. Thirty have the form*“Name a\[n\]XX\.”*and the thirty\-first removes the category constraint \(*“Pick a word\.”*\)\. Each is followed by*“Reply with one word only\.”*in plain chat, which in a format column is replaced by that format’s clause \(§[3](https://arxiv.org/html/2607.18476#S3)\), with no other text changing\. There is no system prompt, and each prompt is a fresh single\-turn conversation at requested temperature 1\.0\.

- •Name a color\.
- •Name an animal\.
- •Name a fruit\.
- •Name a vegetable\.
- •Name a city\.
- •Name a country\.
- •Name a flower\.
- •Name a sport\.
- •Name a musical instrument\.
- •Name a bird\.
- •Name a gemstone\.
- •Name a tree\.
- •Name a beverage\.
- •Name an insect\.
- •Name an occupation\.
- •Name a language\.
- •Name a dessert\.
- •Name a tool\.
- •Name a fish\.
- •Name a metal\.
- •Name a fabric\.
- •Name an herb\.
- •Name a dance\.
- •Name a hobby\.
- •Name a condiment\.
- •Name a cheese\.
- •Name a dinosaur\.
- •Name a mythical creature\.
- •Name a board game\.
- •Name an emotion\.
- •Pick a word\.

## Appendix BNormalization and junk guard

Each reply is reduced to one answer token, following the census\. For a format column the wrapper is first stripped by the per\-format regular expression \(§[3\.3](https://arxiv.org/html/2607.18476#S3.SS3)\); the extracted string — or the raw reply if no wrapper matches — is then normalized: lowercased, stripped of surrounding punctuation and emoji, and reduced to its final alphabetic word \(so*“A common color is blue\.”*→\\to*blue*\)\. A mechanical junk guard treats the following as*failed cells*rather than answers: chat\-template artifacts \(e\.g\.,\[/INST\]or markup tags\), replies longer than 15 words \(a truncated chain of thought would otherwise contribute a spurious novel final word\), bare acknowledgements \(*“Okay\.”*\), and single\-letter tokens\. Within a category, bare plurals are merged with their singulars when both occur\. This is the identical rule set used by the census, imported rather than re\-implemented so the two studies cannot drift\. The format columns add one guard the plain census does not need: a reply whose normalized token is the category noun or the fill\-in placeholder \(\[city\],answer,word\) is an*echo*of the wrapper’s own slot rather than an answer, and is treated as a failed cell\. The guard runs on every column alike — it removes nothing from plain chat, which has no slot to echo — and its only material effect is to remove the echo mass a fill\-in wrapper elicits \(most of it in CSV and brackets\), which trims the apparent brackets loosening from\+0\.16\+0\.16to\+0\.13\+0\.13bits and is why the raw and compliance\-conditioned brackets deltas agree after it is applied\.

Similar Articles

Where does output diversity collapse in post-training?

arXiv cs.CL

This paper investigates where and why output diversity collapses during post-training of language models, analyzing three OLMo 3 lineages (Think, Instruct, RL-Zero) across multiple tasks and metrics. The authors find that diversity collapse is primarily determined by training data composition and embedded in model weights during training, not addressable at inference time alone.

Do Large Language Models Always Tell The Same Stories?

arXiv cs.CL

This paper investigates whether large language models generate diverse stories. Using narrative similarity analysis, the authors find that LLM-generated narratives are consistently more similar to each other than human-written stories, and that common mitigation strategies like negative prompting and temperature scaling fail to address this homogeneity.

Response drift across frontier large language models

arXiv cs.CL

A large-scale human evaluation of 10 frontier LLMs across 62 questions finds that all models exhibit response drift, with most converging to a 78-81% deviation ceiling, while two achieve lower deviation. Drift varies by domain and question, and automated metrics explain little of human judgments, highlighting the need for human evaluation.