通信瓶颈:语言模型中树结构表达序列化的往返研究

arXiv cs.AI 论文

摘要

本文介绍了一种往返协议,用于评估树结构表达在语言模型序列化为自然语言时的完整性,揭示了该过程是有损的、非对称的,并且可以通过微调进行训练。

arXiv:2609.21509v1 Announce Type: new Abstract: When language models reason in chain-of-thought or exchange free-text intermediates, they serialize structured information into natural language. How much tree-structured compositional content survives this bottleneck? We propose a round-trip protocol that answers this question empirically for tree-structured expressions. A generator converts a procedurally generated arithmetic expression into a word problem, a separate extractor recovers the expression from the word problem alone, and symbolic equivalence provides an exact oracle. Evaluating all pairwise combinations of sixteen models yields a communication matrix whose marginals separate generation quality from extraction quality. Three main findings emerge. First, the channel is lossy and asymmetric: swapping which model generates and which extracts shifts accuracy by up to 60.4 points, and the best pair reaches 92.9% by combining different models on each end rather than the same model on both. Second, at least 73.6% of round-trip failures originate at generation, and difficulty is driven by tree structure (operator count, depth, right-branching) rather than model family. Third, the channel is trainable: ~3600 fine-tuning examples that share the evaluation's operators and tree shapes lift every open-weight model above untrained Gemini-3.1-Pro, an upper bound under matched semantics. A disjoint-domain regime with new operators and vocabulary also raises every open-weight model, confirming the gain is not an artifact of matched semantics, though a gap to the frontier remains. Together these results identify tree-structured expression serialization as a primary limiting factor when models communicate hierarchical structure through natural language.
查看原文
查看缓存全文

缓存时间: 2026/09/21 09:26

# The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models
Source: [https://arxiv.org/html/2609.21509](https://arxiv.org/html/2609.21509)
Alex FerrandoAffiliation:AppleLuca ZappellaAffiliation:AppleSamy BengioAffiliation:Apple

###### Abstract

When language models reason in chain\-of\-thought or exchange free\-text intermediates, they serialize structured information into natural language\. How much tree\-structured compositional content survives this bottleneck? We propose a round\-trip protocol that answers this question empirically for tree\-structured expressions\. A generator converts a procedurally generated arithmetic expression into a word problem, a separate extractor recovers the expression from the word problem alone, and symbolic equivalence provides an exact oracle\. Evaluating all pairwise combinations of sixteen models yields a communication matrix whose marginals separate generation quality from extraction quality\. Three main findings emerge\. First, the channel is lossy and asymmetric: swapping which model generates and which extracts shifts accuracy by up to 60\.4 points, and the best pair reaches 92\.9% by combining different models on each end rather than the same model on both\. Second, at least 73\.6% of round\-trip failures originate at generation, and difficulty is driven by tree structure \(operator count, depth, right\-branching\) rather than model family\. Third, the channel is trainable:∼3600\\sim\\\!3600fine\-tuning examples that share the evaluation’s operators and tree shapes lift every open\-weight model above untrained Gemini\-3\.1\-Pro, an upper bound under matched semantics\. A disjoint\-domain regime with new operators and vocabulary also raises every open\-weight model, confirming the gain is not an artifact of matched semantics, though a gap to the frontier remains\. Together these results identify tree\-structured expression serialization as a primary limiting factor when models communicate hierarchical structure through natural language\.

![Refer to caption](https://arxiv.org/html/2609.21509v1/fig1_v2.png)Figure 1:The round\-trip protocol\.A generator𝖦\\mathsf\{G\}serializes an arithmetic expressione\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\(a tree\) into a word problemw\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}that passes through anatural\-language bottleneck\. A separate extractor𝖤\\mathsf\{E\}recoverse^\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}\\hat\{e\}\}fromw\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}alone, and SymPy serves as an exact oracle for symbolic equivalencee≡e^\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\\equiv\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}\\hat\{e\}\}\. Any compositional detail the generator fails to encode is irreversibly lost, regardless of𝖤\\mathsf\{E\}\.## 1Introduction

Language models routinely serialize structured information into natural language \(NL\): chain\-of\-thought reasoning linearizes intermediate computations into prose\([Wei et al\., 2022](https://arxiv.org/html/2609.21509#bib.bib19)\), model outputs are consumed by downstream models in pipelines\([Wu et al\., 2024](https://arxiv.org/html/2609.21509#bib.bib37);[Hong et al\., 2024](https://arxiv.org/html/2609.21509#bib.bib36)\), and tool\-use descriptions encode structured calls as text\([Schick et al\., 2023](https://arxiv.org/html/2609.21509#bib.bib17)\)\. If this serialization is lossy\([Herbst et al\., 2025](https://arxiv.org/html/2609.21509#bib.bib34)\), any structure that the encoding step fails to express unambiguously is irreversibly lost, regardless of the downstream consumer’s capability\. How much compositional structure actually survives the natural\-language bottleneck? We study this for hierarchical structure using tree\-structured expressions built from arithmetic and logical operators, which admit an exact correctness oracle via symbolic equivalence\.

Answering this question requires measuring communication*between*models, but existing benchmarks are not designed for it\. Whether evaluating math\([Cobbe et al\., 2021](https://arxiv.org/html/2609.21509#bib.bib6);[Hendrycks et al\., 2021b](https://arxiv.org/html/2609.21509#bib.bib7);[Glazer et al\., 2025](https://arxiv.org/html/2609.21509#bib.bib18)\), reasoning\([Rein et al\., 2023](https://arxiv.org/html/2609.21509#bib.bib20);[Phan et al\., 2026](https://arxiv.org/html/2609.21509#bib.bib27)\), or code\([Jimenez et al\., 2024](https://arxiv.org/html/2609.21509#bib.bib28);[Jain et al\., 2024](https://arxiv.org/html/2609.21509#bib.bib29)\), they test single models on fixed problem sets\. The closest prior art is[Allamanis et al\. \(2024\)](https://arxiv.org/html/2609.21509#bib.bib63), who round\-trip code through natural language but deliberately keep encoder and decoder fixed to the same model to sidestep what they term the ‘communication chasm’ between distinct models\. We instead treat that chasm as the object of study: separating generator from extractor across anN×NN\{\\times\}Ngrid of models yields a measurement instrument that exposes role asymmetry, fault attribution, and structural predictors of failure, all invisible under same\-model analysis\.

One might ask why study natural language at all\. While structured formats like JSON, code or S\-expressions natively preserve compositional hierarchy and should be utilized whenever feasible \([Section6](https://arxiv.org/html/2609.21509#S6)\), they do not reflect how models typically operate\. Natural language remains the*de facto*channel in all the settings listed above, despite evidence that alternative formats can improve both reasoning and communication\([Chen et al\., 2024b](https://arxiv.org/html/2609.21509#bib.bib53)\)\. Measuring NL compositional fidelity requires a domain with an exact correctness oracle, independent control over structural complexity, and freedom from contamination\([Zhang et al\., 2024](https://arxiv.org/html/2609.21509#bib.bib8)\)\. Arithmetic satisfies all three: \(i\) symbolic equivalence gives a ground\-truth oracle, \(ii\) expression trees expose operator count, nesting depth, and branching topology as controllable parameters, and \(iii\) instances can be procedurally generated\.

To quantify the channel’s capacity for compositional structure, we propose around\-trip protocol\([Figure1](https://arxiv.org/html/2609.21509#S0.F1)\)\. A generator𝖦\\mathsf\{G\}encodes an arithmetic expressione\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}, sampled from an enumerated space of tree skeletons, into a word problemw\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}\. A separate extractor𝖤\\mathsf\{E\}recoverse^\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}\\hat\{e\}\}fromw\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}alone, verified by symbolic equivalence\. A rule\-based guard ensuresw\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}contains no mathematical notation, forcing structure through a linguistic channel \([Section2\.1](https://arxiv.org/html/2609.21509#S2.SS1)\)\. EvaluatingNNmodels in both roles yields anN×NN\{\\times\}N*communication matrix*whose marginals decompose performance into generation and extraction quality \([Section2\.3](https://arxiv.org/html/2609.21509#S2.SS3)\)\. We replicate the protocol on propositional logic as a robustness check \([AppendixR](https://arxiv.org/html/2609.21509#A18)\)\.

We evaluate 16 models: 12 open\-weight \(0\.6B–32B\) across five families and 4 frontier, on 2450 expressions with operator counts 2–8 and depths 2–6\. Our contributions are:

- •A round\-trip protocol for communication fidelity\.A generator encodes a procedurally generated expression as a word problem, an extractor recovers it, and symbolic equivalence serves as an exact oracle\. EvaluatingNNmodels in both roles yields anN×NN\{\\times\}Ncommunication matrix that decomposes pairwise performance into generation and extraction quality \([Section2](https://arxiv.org/html/2609.21509#S2)\)\. Code will be released upon acceptance\.
- •Generation is the bottleneck, not extraction\.At least 73\.6% of round\-trip failures originate at the generation step\. Role assignment alone shifts accuracy by up to 60\.4 pp, and the highest fidelity is achieved by a cross\-model pair, not by self\-communication \([Sections3\.2](https://arxiv.org/html/2609.21509#S3.SS2)and[3\.3](https://arxiv.org/html/2609.21509#S3.SS3)\)\.
- •Failure is governed by tree structure, not model family\.Operator count, nesting depth, and a linearization complexity measureℓ⁡\(T\)\\ell\(T\)dominate failure prediction\. Right\-branching trees yield lower accuracy than left\-branching ones of identical size, andℓ⁡\(T\)\\ell\(T\)captures this asymmetry directly \([Section3\.3](https://arxiv.org/html/2609.21509#S3.SS3)\)\. A propositional\-logic replication preserves the main rankings and fault pattern \([AppendixR](https://arxiv.org/html/2609.21509#A18)\)\.
- •The bottleneck is compositional serialization, not domain knowledge\.Fine\-tuning on a novel operator vocabulary with no arithmetic overlap lifts every open\-weight generator, and in\-domain fine\-tuning on approximately 3600 examples pushes all of them above the strongest untrained frontier generator\. The bottleneck is trainable under matched semantics and partially trainable across domains \([Section4](https://arxiv.org/html/2609.21509#S4)\)\.

## 2Protocol Design

### 2\.1Expression Suite

Standard math\-word\-problem datasets such as GSM8K offer no control over structural complexity: expressions are hand\-authored with uncontrolled nesting depth, operator count, and tree shape\. We instead generate expressions procedurally from abstract syntax tree*skeletons*\. Given a target operator countkkand depthdd, we recursively enumerate every binary\-tree skeleton matching the pair\(k,d\)\(k,d\)\. Each skeleton is then instantiated by assigning operators uniformly from\{\+,−,×,÷\}\\\{\+,\-,\\times,\\div\\\}to internal nodes and filling leaves with named variables from\{A,B,…,H\}\\\{A,B,\\ldots,H\\\}\(probability 0\.7\) or integer constants in\[1,10\]\[1,10\]\(probability 0\.3\)\. Locally degenerate subexpressions are rejected at instantiation: same\-leaf patterns under a non\-commutative operator \(X−XX\-X,X/XX/X\) and multiplicative identities \(X×1X\\times 1,1×X1\\times X,X/1X/1\)\. The operator set, variable alphabet, and leaf\-type probabilities are free parameters of the procedure\. The values above are reasonable defaults, not tuned choices\. Since the skeleton fixes the topology before instantiation, every resulting expression has the targetkkandddby construction \([Figure2](https://arxiv.org/html/2609.21509#S2.F2)\), enabling the factorial analyses of depth, operator count, and tree shape reported in[Section3](https://arxiv.org/html/2609.21509#S3)\.

k=2,d=2k\{=\}2,d\{=\}2\(A\+B\)×C\(A\{\+\}B\)\{\\times\}C\(D/3\)−E\(D\{/\}3\)\{\-\}E\(F×G\)\+7\(F\{\\times\}G\)\{\+\}7k=2,d=2k\{=\}2,d\{=\}2A−\(B×C\)A\{\-\}\(B\{\\times\}C\)F/\(D\+2\)F\{/\}\(D\{\+\}2\)H\+\(E×G\)H\{\+\}\(E\{\\times\}G\)k=3,d=3k\{=\}3,d\{=\}3\(\(A\+B\)×C\)−D\(\(A\{\+\}B\)\{\\times\}C\)\{\-\}D\(\(F/3\)\+G\)×H\(\(F\{/\}3\)\{\+\}G\)\{\\times\}H\(\(E−4\)×A\)\+B\(\(E\{\-\}4\)\{\\times\}A\)\{\+\}BFigure 2:Expression generation via skeleton enumeration\.Circles denote operator nodes, squares denote leaf positions\. We fill each operator node with a random element of\{\+,−,×,÷\}\\\{\+,\-,\\times,\\div\\\}and each leaf with a variable or constant, producing distinct expressions that share the same structure\. The first two skeletons share\(k=2,d=2\)\(k\{=\}2,d\{=\}2\)but differ in branching direction; the third has\(k=3,d=3\)\(k\{=\}3,d\{=\}3\)\.
### 2\.2Round\-Trip Protocol

The round\-tripe→𝖦w→𝖤e^\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\\\!\\xrightarrow\{\\mathsf\{G\}\}\\\!\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}\\\!\\xrightarrow\{\\mathsf\{E\}\}\\\!\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}\\hat\{e\}\}\([Figure1](https://arxiv.org/html/2609.21509#S0.F1)\) proceeds as follows\. Agenerator𝖦\\mathsf\{G\}, prompted withπ𝖦γ\{\\color\[rgb\]\{0,0,0\}\{\\color\[rgb\]\{0,0,0\}\\pi\}\_\{\\mathsf\{G\}\}^\{\\gamma\}\}containingγ\\gammain\-context\(e,w\)\(\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\},\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}\)exemplars, produces a short word problemw=𝖦⁡\(e,π𝖦γ\)\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}=\\mathsf\{G\}\(\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\};\{\\color\[rgb\]\{0,0,0\}\{\\color\[rgb\]\{0,0,0\}\\pi\}\_\{\\mathsf\{G\}\}^\{\\gamma\}\}\)using variable placeholders and no mathematical symbols\. A rule\-based*leakage guard*\(Leak:w↦\{0,1\}\\mathrm\{Leak\}:\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}\\mapsto\\\{0,1\\\}\) rejects outputs that leak mathematical syntax, ensuring structural information passes through a genuinely linguistic channel\. Outputs missing required placeholders or exhibiting degenerate repetition are counted separately as generator faults \([AppendixC](https://arxiv.org/html/2609.21509#A3)\)\. Guard false negatives bias round\-trip accuracy upward, so the reported figures are conservative estimates of channel fidelity\. Anextractor𝖤\\mathsf\{E\}, prompted withπ𝖤ε\{\\color\[rgb\]\{0,0,0\}\{\\color\[rgb\]\{0,0,0\}\\pi\}\_\{\\mathsf\{E\}\}^\{\\varepsilon\}\}containingε\\varepsilonexemplars, recoverse^=𝖤⁡\(w,π𝖤ε\)\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}\\hat\{e\}\}=\\mathsf\{E\}\(\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\};\{\\color\[rgb\]\{0,0,0\}\{\\color\[rgb\]\{0,0,0\}\\pi\}\_\{\\mathsf\{E\}\}^\{\\varepsilon\}\}\), verified by symbolic equivalencesympy\.simplify​\(e^−e\)=0\\texttt\{sympy\.simplify\}\(\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}\\hat\{e\}\}\-\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\)=0\([Meurer et al\., 2017](https://arxiv.org/html/2609.21509#bib.bib12)\), which accepts valid rearrangements \(*e\.g\.,*A\+BA\+Bvs\.B\+AB\+A\) but catches all semantic errors\. Any detail the generator fails to express unambiguously is irreversibly lost, regardless of the extractor\. Prompts are detailed in[AppendixA](https://arxiv.org/html/2609.21509#A1), worked examples in[AppendixK](https://arxiv.org/html/2609.21509#A11)\. Shot exemplars are drawn from a pool disjoint from the evaluation set, and the shot countsγ,ε\\gamma,\\varepsilonare fixed in[Section3\.1](https://arxiv.org/html/2609.21509#S3.SS1)\. Because our contribution is the protocol itself,π𝖦γ\{\\color\[rgb\]\{0,0,0\}\{\\color\[rgb\]\{0,0,0\}\\pi\}\_\{\\mathsf\{G\}\}^\{\\gamma\}\}andπ𝖤ε\{\\color\[rgb\]\{0,0,0\}\{\\color\[rgb\]\{0,0,0\}\\pi\}\_\{\\mathsf\{E\}\}^\{\\varepsilon\}\}are parameters rather than fixed choices, and prompt tuning is orthogonal to the measurements we report\.

### 2\.3Communication Matrix

GivenNNmodels, each serving as both𝖦\\mathsf\{G\}\{\}and𝖤\\mathsf\{E\}\{\}, the protocol produces anN×NN\{\\times\}Ncommunication matrix𝑸\{\\bm\{Q\}\}whose entryqi​jq\_\{ij\}is the round\-trip accuracy when extractoriireads generatorjj’s word problems\. For eache∈D\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\\in D, generatorjjproducesw=𝖦j​\(e\)\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}=\\mathsf\{G\}\_\{j\}\(\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\)\. Ifw\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}fails the leakage guard \(Leak⁡\(w\)=0\\mathrm\{Leak\}\(\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}\)=0\), the round trip scores zero, otherwise extractoriirecoverse^=𝖤i​\(w\)=𝖤i​\(𝖦j​\(e\)\)\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}\\hat\{e\}\}=\\mathsf\{E\}\_\{i\}\(\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}\)=\\mathsf\{E\}\_\{i\}\\big\(\\mathsf\{G\}\_\{j\}\(\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\)\\big\):

qi​j=1\|D\|∑e∈DLeak\(𝖦j\(e\)\)⋅\[𝖤i\(𝖦j\(e\)\)≡e\]\.q\_\{ij\}\\;=\\;\\frac\{1\}\{\|D\|\}\\sum\_\{\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\\in D\}\\mathrm\{Leak\}\\big\(\\mathsf\{G\}\_\{j\}\(\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\)\\big\)\\;\\cdot\\;\\mathds\{1\}\\\!\\bigl\[\\,\\mathsf\{E\}\_\{i\}\\big\(\\mathsf\{G\}\_\{j\}\(\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\)\\big\)\\equiv\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\\,\\bigr\]\.\(1\)The column averageqj𝖦=1N​∑iqi​jq^\{\\mathsf\{G\}\}\_\{j\}=\\frac\{1\}\{N\}\\sum\_\{i\}q\_\{ij\}

measures thegeneration qualityof modeljj: how decodable its word problems are across extractors\. The row averageqi𝖤=1N​∑jqi​jq^\{\\mathsf\{E\}\}\_\{i\}=\\frac\{1\}\{N\}\\sum\_\{j\}q\_\{ij\}measures theextraction qualityof modelii: how reliably it recovers expressions across generators\. The diagonal entryqi​iq\_\{ii\}is theself\-communicationscore\. The denominator is the full suite\|D\|\|D\|, not each generator’s guard\-passing subset, so guard failures count as communication failures, eliminating the selection bias from evaluating generators only on expressions they can encode\. A consequence is thatqi𝖤q^\{\\mathsf\{E\}\}\_\{i\}is panel\-relative: guard failures enter as zeros in every extractor row, so absolute values reflect the generator mix\. This balanced design, in which every𝖦/𝖤\\mathsf\{G\}/\\mathsf\{E\}pair is evaluated on the same expression set, ensures that marginal averages are fair comparisons\. A latent\-ability analysis confirming that raw marginals track maximum\-likelihood skill estimates is reported in[AppendixN](https://arxiv.org/html/2609.21509#A14)\.

### 2\.4Fault Attribution

When a round trip fails \(e^≢e\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}\\hat\{e\}\}\\not\\equiv\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\), the error may originate at the generation or extraction step\. Because the expression actually encoded byw\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}is unobserved, we estimate it via multi\-extractor consensus\. For a word problemw=𝖦j​\(e\)\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}=\\mathsf\{G\}\_\{j\}\(\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\), we define the*extractor vote*for any candidate expressione′e^\{\\prime\}as

P\(e′∣w\)=1N∑i=1N\[𝖤i\(w\)≡e′\]\.P\(e^\{\\prime\}\\mid\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}\)\\;=\\;\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\mathds\{1\}\\\!\\bigl\[\\mathsf\{E\}\_\{i\}\(\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}\)\\equiv e^\{\\prime\}\\bigr\]\.\(2\)IfP⁡\(e∣w\)\>0P\(\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\\mid\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}\)\>0, at least one extractor recovers the target, proving the word problem faithful, and any remaining failures are consideredextractor faults\. IfP⁡\(e∣w\)=0P\(\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\\mid\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}\)=0, no extractor succeeds and we classify the failure as agenerator fault\. The severity of the generator fault is gauged by theextraction agreementp∗​\(w\)=maxe′⁡P⁡\(e′∣w\)p^\{\*\}\(\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}\)=\\max\_\{e^\{\\prime\}\}P\(e^\{\\prime\}\\mid\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}\),*i\.e\.,*the fraction of extractors that converge on the most common output\. Highp∗p^\{\*\}means most extractors recover the same wrong expression, indicating that the generator produced a coherent word problem that systematically describes a different expression thane\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\. Lowp∗p^\{\*\}means extractors disagree, indicating that the word problem is garbled or ambiguous\. Attribution results are reported in[Section3\.3](https://arxiv.org/html/2609.21509#S3.SS3)\.

A generation counts as faithful if*any*of theN=16N\{=\}16extractors succeeds, so enlarging the panel can only shift failures from the generator\-fault bucket to the extractor\-fault bucket\. The reported generator\-fault rate is therefore conservative with respect to the bottleneck claim: every smaller sub\-panel we tested reports a higher rate, up to91\.1%91\.1\\%\. Misattribution is negligible at thisNN\([SectionM\.1](https://arxiv.org/html/2609.21509#A13.SS1)\)\.

Table 1:The𝟏𝟔×𝟏𝟔\\bm\{16\\times 16\}communication matrix on all generated expressions \(guard failures counted as errors\)\.Each cell shows round\-trip accuracy \(%\)\. Rows are extractors, columns are generators\. The rightmost column and bottom row show extraction quality \(row average\) and generation quality \(column average\)\. Generator and extractor are run with promptsπ𝖦0\{\\color\[rgb\]\{0,0,0\}\{\\color\[rgb\]\{0,0,0\}\\pi\}\_\{\\mathsf\{G\}\}^\{0\}\}andπ𝖤ε\{\\color\[rgb\]\{0,0,0\}\{\\color\[rgb\]\{0,0,0\}\\pi\}\_\{\\mathsf\{E\}\}^\{\\varepsilon\}\}respectively \(configurationOPENγ=0,ε=3\)\\gamma\{=\}0,\\varepsilon\{=\}3\); see[AppendixP](https://arxiv.org/html/2609.21509#A16)for the full ablation\. The±\\pmvalues denote standard deviation across the 16 pairings in each row or column\. TheLeak\\mathrm\{Leak\}% row reports each model’s leakage\-guard pass rate as a generator\. Thin rules separate open\-weight auto\-regressive, diffusion, and frontier models\.GeneratorsExtractorsQwen3\-0\.6B Qwen3\-1\.7B Qwen3\-4B Qwen3\-8B Qwen3\-14B Qwen3\-32B Gemma\-3\-4B Gemma\-3\-12B Gemma\-3\-27B Phi\-4 Dream\-7B LLaDA\-8B C\-Haiku\-4\.5 GPT\-5 G\-3\-Flash G\-3\.1\-Pro Extr\. quality Leak\\mathrm\{Leak\}pass %206094989999989696976495100100100100Qwen3\-0\.6B0\.01\.23\.24\.53\.42\.20\.10\.40\.34\.70\.54\.07\.21\.91\.91\.62\.3±\\pm2Qwen3\-1\.7B0\.01\.87\.711\.29\.97\.90\.52\.01\.314\.41\.09\.224\.911\.311\.813\.68\.0±\\pm7Qwen3\-4B0\.02\.813\.620\.321\.319\.00\.73\.43\.126\.32\.115\.849\.139\.034\.646\.018\.6±\\pm16Qwen3\-8B0\.12\.714\.219\.821\.720\.70\.83\.13\.827\.62\.016\.549\.044\.239\.152\.419\.9±\\pm17Qwen3\-14B0\.12\.916\.123\.324\.923\.60\.84\.04\.332\.02\.419\.157\.157\.150\.665\.124\.0±\\pm22Qwen3\-32B0\.13\.116\.223\.325\.724\.80\.93\.94\.431\.32\.919\.357\.264\.151\.770\.225\.0±\\pm23Gemma\-3\-4B0\.12\.110\.213\.513\.310\.80\.62\.91\.817\.81\.411\.729\.517\.417\.819\.810\.7±\\pm8Gemma\-3\-12B0\.02\.413\.318\.419\.418\.00\.73\.23\.124\.52\.315\.545\.338\.032\.744\.317\.6±\\pm15Gemma\-3\-27B0\.12\.815\.921\.524\.123\.11\.13\.84\.329\.82\.417\.654\.556\.346\.364\.823\.0±\\pm21Phi\-40\.02\.917\.123\.326\.724\.50\.94\.24\.532\.82\.619\.958\.365\.052\.372\.825\.5±\\pm24Dream\-7B0\.01\.87\.511\.910\.610\.40\.42\.21\.813\.11\.19\.122\.718\.219\.431\.310\.1±\\pm9LLaDA\-8B0\.02\.312\.417\.117\.615\.90\.53\.22\.821\.52\.415\.138\.937\.732\.239\.416\.2±\\pm14C\-Haiku\-4\.50\.03\.116\.523\.326\.725\.70\.93\.84\.633\.03\.019\.059\.679\.156\.382\.527\.3±\\pm27GPT\-50\.13\.317\.624\.928\.124\.00\.94\.64\.434\.72\.921\.964\.791\.364\.588\.529\.8±\\pm30G\-3\-Flash0\.13\.316\.824\.427\.225\.61\.03\.94\.434\.22\.920\.462\.589\.360\.688\.929\.1±\\pm29G\-3\.1\-Pro0\.02\.816\.724\.127\.125\.60\.84\.14\.435\.53\.020\.862\.992\.962\.488\.729\.5±\\pm30Gen\. quality0\.1±\\pm02\.6±\\pm113\.4±\\pm419\.0±\\pm620\.5±\\pm718\.8±\\pm70\.7±\\pm03\.3±\\pm13\.3±\\pm125\.8±\\pm92\.2±\\pm115\.9±\\pm546\.5±\\pm1750\.2±\\pm2839\.6±\\pm1954\.4±\\pm27

## 3Experimental Results

### 3\.1Setup

We evaluateN=16N=16models across three groups\. Ten open\-weight auto\-regressive: Qwen3\([Yang and others, 2025](https://arxiv.org/html/2609.21509#bib.bib3)\)\(0\.6B, 1\.7B, 4B, 8B, 14B, 32B\), Gemma\-3\([Gemma Team, 2025](https://arxiv.org/html/2609.21509#bib.bib4)\)\(4B, 12B, 27B\), and Phi\-4\([Abdin et al\., 2024](https://arxiv.org/html/2609.21509#bib.bib5)\)\. Two diffusion LMs, Dream\-v0\-7B\([Ye et al\., 2025a](https://arxiv.org/html/2609.21509#bib.bib9)\)and LLaDA\-8B\([Nie and others, 2025](https://arxiv.org/html/2609.21509#bib.bib10)\), to test whether parallel denoising yields a different serialization profile\. Four frontier models via API: Claude Haiku 4\.5\([Anthropic, 2025b](https://arxiv.org/html/2609.21509#bib.bib52)\), GPT\-5\([Singh et al\., 2025](https://arxiv.org/html/2609.21509#bib.bib26)\), Gemini\-3\-Flash\([Google DeepMind, 2025a](https://arxiv.org/html/2609.21509#bib.bib76)\), and Gemini\-3\.1\-Pro\([Google DeepMind, 2025b](https://arxiv.org/html/2609.21509#bib.bib77)\)\. All use greedy decoding \(temperature 0, thinking disabled\),[AppendixE](https://arxiv.org/html/2609.21509#A5)confirmstemp=0\\mathrm\{temp\}=0yields the best accuracy\. The expression suite contains\|D\|=2450\|D\|=2450expressions, 5 instantiations per skeleton with up to 50 skeletons per\(k,d\)\(k,d\)cell, fork∈\[2,8\],d∈\[2,6\]k\\in\[2,8\],\\;d\\in\[2,6\], with uniform random subsample with fixed seed when a cell exceeds 50 \(not all pairs are feasible, see[AppendixB](https://arxiv.org/html/2609.21509#A2)\)\. Guard pass rates are reported alongside all results in[Table1](https://arxiv.org/html/2609.21509#S2.T1)\.

Shot configuration\.We ablate\(γ,ε\)∈\{0,3\}2\(\\gamma,\\varepsilon\)\\in\\\{0,3\\\}^\{2\}and report[Table1](https://arxiv.org/html/2609.21509#S2.T1)at the trade\-off optimum\(0,3\)\(0,3\): generator shots \(γ=3\\gamma\{=\}3\) do not help the strongest auto\-regressive models \(Qwen3≥\\geq4B, Phi\-4\), while extractor shots uniformly improve extraction overε=0\\varepsilon\{=\}0\. Frontier models were evaluated only atγ=0\\gamma\{=\}0for tractability\. Full per\-condition grids are in[AppendixP](https://arxiv.org/html/2609.21509#A16)\.

### 3\.2Characterizing the Natural\-Language Channel

The full communication matrix \([Table1](https://arxiv.org/html/2609.21509#S2.T1); per\-depth breakdowns in[AppendixF](https://arxiv.org/html/2609.21509#A6)\) reveals a channel whose capacity rises with model capability and depends on how models are paired across roles\. Across the 12 open\-weight models, mean round\-trip accuracy \(qi​jq\_\{ij\},[Equation1](https://arxiv.org/html/2609.21509#S2.E1)\) is9\.4%9\.4\\%, rising to17\.0%17\.0\\%at≥\\geq8B\. Frontier models raise the ceiling: the4×44\\times 4frontier submatrix averages74\.7%74\.7\\%, and the best cell reaches92\.9%92\.9\\%\(Gemini\-3\.1\-Pro reading GPT\-5\), above the best self\-communication score \(max91\.3%91\.3\\%, GPT\-5\) and4\.24\.2pp above Gemini\-3\.1\-Pro reading from itself\. The best cell is itself a role\-asymmetry result: the strongest channel pairs the strongest generator with the strongest extractor, not the same model on both sides\. Accuracy still degrades with complexity at every tier\. Frontier generators drop from88%88\\%atk=3k\{=\}3to71%71\\%atk=8k\{=\}8, and open\-weight≥\\geq8B generators from64%64\\%atk=2k\{=\}2to13%13\\%atk=8k\{=\}8\. Diffusion parallel denoising does not escape the bottleneck\. LLaDA\-8B \(gen/extr15\.9/16\.215\.9/16\.2\) and Dream\-v0\-7B \(2\.2/10\.12\.2/10\.1\) both sit within the autoregressive envelope\.

Generation and extraction are distinct skills\.q𝖦q^\{\\mathsf\{G\}\}andq𝖤q^\{\\mathsf\{E\}\}are not directly comparable in absolute terms because each marginal is upper\-bounded by the panel’s quality on the opposite role, but within\-cohort ranks are comparable\.[Figure3](https://arxiv.org/html/2609.21509#S3.F3)\-left plots each model’s rank as𝖦\\mathsf\{G\}against its rank as𝖤\\mathsf\{E\}\. Had both roles reflected one ability, models would sit on the diagonal, but they disperse instead\. GPT\-5 is the top extractor \(q𝖤=29\.8%q^\{\\mathsf\{E\}\}=29\.8\\%, rank 1\) yet ranks second as generator \(q𝖦=50\.2%q^\{\\mathsf\{G\}\}=50\.2\\%\); Gemini\-3\.1\-Pro leads generation \(q𝖦=54\.4%q^\{\\mathsf\{G\}\}=54\.4\\%, rank 1\) but is only the second\-best extractor \(q𝖤=29\.5%q^\{\\mathsf\{E\}\}=29\.5\\%\)\. The dispersion is largest for Gemma\-3\-27B \(rank 11 generator at3\.3%3\.3\\%, rank 8 extractor at23\.0%23\.0\\%\): a model whose narration loses structure more than its parsing does\.

Role assignment matters\.Swapping which model generates and which extracts shifts accuracy by up to60\.460\.4pp \(Gemma\-3\-27B with Gemini\-3\.1\-Pro, larger than any row or column spread in[Table1](https://arxiv.org/html/2609.21509#S2.T1)\), where Gemma’s stronger extraction \(rank 8\) and weak generation \(rank 11\) make role assignment decisive\. The effect also appears within the frontier panel: GPT\-5 generating with Gemini\-3\.1\-Pro extracting yields92\.9%92\.9\\%, the reverse88\.5%88\.5\\%\(4\.44\.4pp\)\. Pairing the strongest model on both sides is suboptimal when the panel has a model with complementary strengths\.

Figure 3:\(Left\) Role asymmetry\.Generator rank vs\. extractor rank \(1 = best\)\. Gemini\-3\.1\-Pro leads generation but is second as extractor, GPT\-5 is the top extractor but second generator\. Swapping roles shifts round\-trip accuracy by up to 60\.4 pp\.\(Right\) Channel asymmetry\.Qwen3 extraction on JSON/S\-expressions \(structured, unambiguous\) vs\. Frontier\-generated word problems\. The shaded area is the*NL tax*: capacity large extractors have but no NL generator in our panel reaches\.The NL channel is quantifiably lossy\.To separate channel loss from extractor weakness, we need a reference point where the encoding is lossless\. We implement it by replacing the generator with deterministic S\-expression and JSON\-AST serializations and measure theNL tax=max⁡\(qS\-expr𝖤,qJSON𝖤\)−qNL𝖤\\text\{NL tax\}=\\max\(q^\{\\mathsf\{E\}\}\_\{\\text\{S\-expr\}\},q^\{\\mathsf\{E\}\}\_\{\\text\{JSON\}\}\)\-q^\{\\mathsf\{E\}\}\_\{\\text\{NL\}\}, the extraction gap between a lossless encoding and natural language\.[Figure3](https://arxiv.org/html/2609.21509#S3.F3)\-right shows this for the Qwen3 family: at 14B–32B, lossless extraction reaches8585–88%88\\%while NL extraction from Frontier models plateaus at5151–70%70\\%, an1818–2020pp tax that grows with complexity \([AppendixQ](https://arxiv.org/html/2609.21509#A17)\)\. The gap is relative to the panel’s best generator and would narrow with a stronger one\. Below 4B, models lack fluency in structured notation and extract similarly from prose, so the structured\-format advantage requires extractor familiarity\.

### 3\.3Why Communication Fails

What predicts failure?We train a gradient\-boosted classifier on all\(e,𝖦,𝖤\)\(\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\},\\mathsf\{G\},\\mathsf\{E\}\)triples from the 12 open\-weight models and compute SHAP attributions\([Lundberg and Lee, 2017](https://arxiv.org/html/2609.21509#bib.bib61);[Lundberg et al\., 2019](https://arxiv.org/html/2609.21509#bib.bib1)\)\([Figure4](https://arxiv.org/html/2609.21509#S3.F4), left; details in[AppendixG](https://arxiv.org/html/2609.21509#A7)\)\. Frontier models are excluded because their unknown parameter counts would confound the model\-size features\. Model size dominates: generator and extractor sizes together account for57\.0%57\.0\\%of the attribution\.RmathR\_\{\\text\{math\}\}\(19\.5%19\.5\\%,↑\\uparrow\), the fraction of words in the word problem belonging to a curated arithmetic lexicon \([AppendixD](https://arxiv.org/html/2609.21509#A4)\), ranks second overall, ahead of extractor size, and is comparable in magnitude to the combined structural features \(21\.6%21\.6\\%\), indicating that pseudo\-mathematical phrasing produces more decodable problems\. Among structural features, depth \(9\.2%9\.2\\%\), operator count \(7\.8%7\.8\\%\), and*linearization complexity*ℓ⁡\(T\)\\ell\(T\)\(4\.6%4\.6\\%\) contribute in that order\. We defineℓ⁡\(T\)\\ell\(T\)as the number of internal nodes ofTTwhose right child is itself internal, counted tree\-wide \(ℓ=0\\ell\{=\}0for a left\-caterpillar of any depth,ℓ=d−1\\ell\{=\}d\{\-\}1for a right\-caterpillar of depthdd\), reminiscent of the minimum register count in expression evaluation\([Sethi and Ullman, 1970](https://arxiv.org/html/2609.21509#bib.bib15)\)\. Controlling for depth,ℓ⁡\(T\)\\ell\(T\)predicts difficulty monotonically, with a mean within\-depth Spearman correlation ofρ\|d¯=−0\.48\\overline\{\\rho\_\{\|d\}\}=\-0\.48\([Figure4](https://arxiv.org/html/2609.21509#S3.F4), center\)\. Right\-branching skeletons are therefore harder than left\-branching ones of matched size\. Per\-cell breakdowns appear in[AppendixH](https://arxiv.org/html/2609.21509#A8)\. Family match contributes1\.9%1\.9\\%, providing no evidence of architecture\-specific dialects\.

Where does failure occur?Applying the consensus\-based attribution \([Section2\.4](https://arxiv.org/html/2609.21509#S2.SS4), with the full breakdown in[AppendixM](https://arxiv.org/html/2609.21509#A13)\) to all failed round trips, we find that at least73\.6%73\.6\\%are generator faults \(no extractor recovers the target\) and at most26\.4%26\.4\\%are extractor faults\. Wrong operator count is the dominant error type, accounting for71\.8%71\.8\\%of all errors \([AppendixI](https://arxiv.org/html/2609.21509#A9), with qualitative failure traces in[AppendixL](https://arxiv.org/html/2609.21509#A12)\)\. The extractor\-fault rate scales steeply with generator capability, from under4%4\\%for sub\-2B open\-weight generators to6060–97%97\\%for frontier models \([Figure4](https://arxiv.org/html/2609.21509#S3.F4), right\)\. The73\.6%73\.6\\%figure is robust to panel composition: across nine extractor\-panel subsets the rate ranges73\.673\.6–91\.1%91\.1\\%, with the full 16\-model panel yielding the lowest observation \([SectionM\.1](https://arxiv.org/html/2609.21509#A13.SS1)\)\. No panel we tested reports a generator\-fault share below73\.6%73\.6\\%\.

FeatureAttributionGenerator size39\.1%↑\\uparrowRmathR\_\{\\text\{math\}\}19\.5%↑\\uparrowExtractor size17\.9%↑\\uparrowDepth9\.2%↓\\downarrowOpCount7\.8%↓\\downarrowℓ⁡\(T\)\\ell\(T\)4\.6%↓\\downarrowFamily match1\.9%↑\\uparrow

Figure 4:Why communication fails\.Over depths22–66\.\(Left\)Global SHAP attributions, computed on the 12 open\-weight models only \(frontier parameter counts are unknown and would confound the size features\): generator size39\.1%39\.1\\%,RmathR\_\{\\text\{math\}\}19\.5%19\.5\\%, and extractor size17\.9%17\.9\\%dominate\.\(Center\)Mean within\-depth Spearman correlation between linearization complexityℓ⁡\(T\)\\ell\(T\)and accuracy isρ\|d¯=−0\.48\\overline\{\\rho\_\{\|d\}\}=\-0\.48: right\-branching skeletons are harder at matched depth\.\(Right\)Extractor\-fault rate rises with generator size, computed on all 16 models \(12 open\-weight plus Haiku, GPT\-5 and the two Geminis\), from<4%\{<\}4\\%for the smallest open\-weight generators to6060–97%97\\%for the frontier models\.
### 3\.4External Validity

Do round\-trip scores recapitulate existing benchmarks or capture a distinct capability? We correlate per\-model round\-trip scores with published results on 9 benchmarks across four categories \(at least seven verified scores each\)\. Extraction correlates strongly with general knowledge \(MMLUρ=0\.98\\rho=0\.98\), code \(MBPPρ=0\.93\\rho=0\.93, LiveCodeBenchρ=0\.93\\rho=0\.93\), and reasoning \(GPQA\-Diamondρ=0\.95\\rho=0\.95\), but generation correlates only moderately \(ρ=0\.47\\rho=0\.47–0\.890\.89\)\. This gap is consistent across all four categories, confirming that serializing compositional structure into prose is partially distinct from the comprehension skills standard benchmarks test\. Details in[AppendixJ](https://arxiv.org/html/2609.21509#A10)\.

Cross\-domain replication\.We re\-run the protocol on propositional logic on a reduced14×1414\\times 14roster \(five connectives, depths 2–5, logical equivalence as oracle; frontier Gemini omitted to limit compute;[AppendixR](https://arxiv.org/html/2609.21509#A18)\)\. Absolute accuracy drops \(mean cell8\.7%8\.7\\%vs\.14\.9%14\.9\\%on the shared 14\-model block\), but model rankings mostly carry over: Spearmanρ=0\.85\\rho=0\.85for generation and0\.690\.69for extraction, with GPT\-5 and Claude Haiku 4\.5 again on top\. The main exception is Phi\-4, which falls from extraction rank 3 to 14, suggesting its arithmetic extraction relies on format\-specific parsing\. Rankings are stable overall but not uniform, reinforcing a core takeaway: the best model and role assignment depend on the domain, not just on overall capability\.

## 4Improving Communication Through Fine\-Tuning

Generation is the bottleneck: at least 73\.6% of round\-trip failures originate there \([Section3\.3](https://arxiv.org/html/2609.21509#S3.SS3)\)\. If the bottleneck is trainable rather than architectural, what kind of training lifts it and how far does it transfer? Two fine\-tuning regimes test this, both sharing the round\-trip format\.

Arithmetic fine\-tuningprobes how far structural fluency carries across unseen variables and constants\. Training pairs \(e→w\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\\\!\\to\\\!\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}andw→e\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}\\\!\\to\\\!\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\) use the\{\+,−,×,÷\}\\\{\+,\-,\\times,\\div\\\}operators and the evaluation tree topology, with disjoint letters\{Q​…​Z\}\\\{Q\\ldots Z\\\}and constants\[11​…​20\]\[11\\ldots 20\]\.

Assembly fine\-tuningprobes whether the linearization skill transfers across domains\. Assembly has mixed\-arity operators \(three unary, six binary, four ternary\), breaking the binary\-only isomorphism, and coloured\-object leaves \(13 colours×\\times13 objects\) disjoint from any arithmetic vocabulary\. Each operator maps to multiple paraphrase templates in a NL procedural style\. Every \(expression, NL\) pair yields two training examples \(e→w\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\\\!\\to\\\!\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}andw→e\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}\\\!\\to\\\!\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\)\. Full specification in[AppendixO](https://arxiv.org/html/2609.21509#A15)\.

Training and evaluation\.Both regimes use QLoRA\([Dettmers et al\., 2023](https://arxiv.org/html/2609.21509#bib.bib13)\)\(rank 32,α=64\\alpha\{=\}64, 4\-bit NF4, 3 epochs, AdamW\([Loshchilov and Hutter, 2019](https://arxiv.org/html/2609.21509#bib.bib14)\), lr2×10−42\{\\times\}10^\{\-4\}, cosine schedule\) on∼3600\\sim\\\!3600training messages across the 10 open\-weight autoregressive models\. Adapters are merged before evaluation\.111A single hyperparameter setting is used across models\. Per\-model tuning is orthogonal to the structural\-transfer question\.We report each condition at its strongest shot configuration on the full grid \([AppendixP](https://arxiv.org/html/2609.21509#A16)\): base and arithmetic FT useπ𝖦0,π𝖤3\{\\color\[rgb\]\{0,0,0\}\{\\color\[rgb\]\{0,0,0\}\\pi\}\_\{\\mathsf\{G\}\}^\{0\}\},\{\\color\[rgb\]\{0,0,0\}\{\\color\[rgb\]\{0,0,0\}\\pi\}\_\{\\mathsf\{E\}\}^\{3\}\}, assembly FT usesπ𝖦3,π𝖤3\{\\color\[rgb\]\{0,0,0\}\{\\color\[rgb\]\{0,0,0\}\\pi\}\_\{\\mathsf\{G\}\}^\{3\}\},\{\\color\[rgb\]\{0,0,0\}\{\\color\[rgb\]\{0,0,0\}\\pi\}\_\{\\mathsf\{E\}\}^\{3\}\}because assembly\-FT generators benefit from three exemplars but base and arithmetic\-FT generators do not\.

### 4\.1Fine\-tuning Results

Table 2:Fine\-tuning comparison across 10 open\-weight models\.Three conditions are evaluated at their strongest shot configuration \([AppendixP](https://arxiv.org/html/2609.21509#A16)\): Base\(γ=0,ε=3\)\(\\gamma\{=\}0,\\varepsilon\{=\}3\), assembly FT\(γ=3,ε=3\)\(\\gamma\{=\}3,\\varepsilon\{=\}3\), arithmetic FT\(γ=0,ε=3\)\(\\gamma\{=\}0,\\varepsilon\{=\}3\)\.Leak\\mathrm\{Leak\}failures count as errors; all marginals are averaged over the 10\-model extractor panel\.Bold= best per row;underline= second\-best\.Leak\\mathrm\{Leak\}pass %q𝖦q^\{\\mathsf\{G\}\}q𝖤q^\{\\mathsf\{E\}\}SelfModelBaseFTasm\{\}\_\{\\text\{asm\}\}FTarith\{\}\_\{\\text\{arith\}\}BaseFTasm\{\}\_\{\\text\{asm\}\}FTarith\{\}\_\{\\text\{arith\}\}BaseFTasm\{\}\_\{\\text\{asm\}\}FTarith\{\}\_\{\\text\{arith\}\}BaseFTasm\{\}\_\{\\text\{asm\}\}FTarith\{\}\_\{\\text\{arith\}\}Qwen3\-0\.6B20521000\.10\.151\.93\.03\.714\.10\.10\.061\.8Qwen3\-1\.7B59661002\.66\.557\.26\.68\.434\.82\.13\.870\.4Qwen3\-4B948210014\.017\.355\.312\.420\.964\.415\.321\.579\.4Qwen3\-8B989410020\.532\.659\.512\.921\.166\.222\.737\.383\.0Qwen3\-14B999810021\.426\.255\.914\.524\.272\.727\.336\.782\.1Qwen3\-32B989810019\.427\.658\.014\.725\.373\.626\.541\.685\.6Gemma\-3\-4B98751000\.85\.555\.48\.411\.937\.10\.64\.648\.4Gemma\-3\-12B96791003\.916\.757\.811\.719\.259\.84\.221\.181\.2Gemma\-3\-27B96931003\.820\.458\.414\.224\.270\.65\.127\.284\.6Phi\-4979710026\.932\.758\.215\.126\.674\.435\.759\.384\.0

Arithmetic fine\-tuning closes the gap to the frontier\.Arithmetic FT raises generation quality from a baseline range of0\.10\.1–26\.9%26\.9\\%to51\.951\.9–59\.5%59\.5\\%across all 10 open\-weight models \([Table2](https://arxiv.org/html/2609.21509#S4.T2)\), and guard pass rates reach100%100\\%, eliminating leakage entirely\. Self\-communication reaches7979–86%86\\%above 4B and62%62\\%even for Qwen3\-0\.6B\. The bottleneck quantified in[Section3](https://arxiv.org/html/2609.21509#S3)is therefore not architectural: training on the same operator set with disjoint variables and constants is enough to push every panel model above the strongest untrained frontier generator \(Gemini\-3\.1\-Pro,q𝖦=45\.1%q^\{\\mathsf\{G\}\}=45\.1\\%\)\. Because training and evaluation expressions share operator semantics and tree topology, this is best read as an upper bound on what the channel admits under matched semantics, not a generalization claim\.

Structural fine\-tuning lifts all generators, weakest most\.Assembly FT raises generation quality for every open\-weight model, by\+3\.3\+3\.3to\+16\.6\+16\.6pp, with the largest gains on the weakest generators \(Gemma\-3\-12B\+12\.8\+12\.8, Gemma\-3\-27B\+16\.6\+16\.6\) and smaller but consistently positive gains on the strongest \(Qwen3\-8B\+12\.1\+12\.1, Phi\-4\+5\.8\+5\.8\)\. Extraction quality improves uniformly,\+0\.7\+0\.7to\+11\.5\+11\.5pp, with the strongest extractors gaining most\. Self\-communication moves with both, reaching59\.3%59\.3\\%for Phi\-4 and41\.6%41\.6\\%for Qwen3\-32B\. Guard pass rate falls on average from85\.5%85\.5\\%\(base\) to83\.4%83\.4\\%\. The drop is concentrated in Gemma\-3\-4B \(−22\.9\-22\.9pp\) and Gemma\-3\-12B \(−17\.5\-17\.5\), while Phi\-4 and the larger Qwen3 models are essentially unchanged \([AppendixO](https://arxiv.org/html/2609.21509#A15)\)\.

Cross\-domain gains are real but a gap to the frontier persists\.Because the 10 base extractors never see the FT data, assembly\-FT gains reflect more decodable problems rather than a shared artifact\. Even so, the strongest assembly\-FT generator \(Phi\-4,32\.7%32\.7\\%\) stays12\.412\.4pp below untrained Gemini\-3\.1\-Pro \(45\.1%45\.1\\%\)\. Arithmetic FT closes that gap: every open\-weight model exceeds Gemini\-3\.1\-Pro, but only by sharing operator semantics and tree topology with evaluation, so it is best read as an upper bound\.

## 5Related Work

Round\-trip and consistency evaluation\.Round\-trip protocols are established quality proxies, going back to back\-translation\([Sennrich et al\., 2016](https://arxiv.org/html/2609.21509#bib.bib45)\)\. Closest to us,[Allamanis et al\. \(2024\)](https://arxiv.org/html/2609.21509#bib.bib63)formalize round\-trip correctness for code LLMs, using the*same*model on both ends to avoid a “communication chasm”\.[Maveli et al\. \(2026\)](https://arxiv.org/html/2609.21509#bib.bib64)round\-trip code through compression–decompression bijections, finding persistent gaps across prompting, fine\-tuning, and self\-reflection, and[Kong et al\. \(2025\)](https://arxiv.org/html/2609.21509#bib.bib65)extend the framing as a training reward for molecule–text LLMs\. Both use a same\-model round\-trip\.[Wigler et al\. \(2026\)](https://arxiv.org/html/2609.21509#bib.bib49)use a cross\-model NL round\-trip like ours to recover personality profiles from life stories, but transmit a trait vector at fixed complexity\. We study Allamanis’ chasm directly for compositional content: anN×NN\{\\times\}Ngenerator–extractor matrix, symbolic equivalence as the oracle, and controlled complexity, with failures attributed via row and column marginals\. Unlike self\-consistency\([Wang et al\., 2023](https://arxiv.org/html/2609.21509#bib.bib46)\), which aggregates reasoning paths from one model, we test whether a second reader recovers the sender’s structure through the NL channel, complementing work on internal faithfulness\([Lanham et al\., 2023](https://arxiv.org/html/2609.21509#bib.bib47);[Turpin et al\., 2023](https://arxiv.org/html/2609.21509#bib.bib48)\)with an externally verifiable signal\.

Multi\-agent LLM systems\.MetaGPT\([Hong et al\., 2024](https://arxiv.org/html/2609.21509#bib.bib36)\), CAMEL\([Li et al\., 2023](https://arxiv.org/html/2609.21509#bib.bib35)\), AutoGen\([Wu et al\., 2024](https://arxiv.org/html/2609.21509#bib.bib37)\), and Mixture\-of\-Agents\([Wang et al\., 2025](https://arxiv.org/html/2609.21509#bib.bib22)\)route coordination through NL but do not measure information loss at that interface, even when debate improves factuality\([Du et al\., 2024](https://arxiv.org/html/2609.21509#bib.bib21)\)\.[Zhou et al\. \(2025\)](https://arxiv.org/html/2609.21509#bib.bib24)argue NL is fundamentally misaligned with LLM representations\.[He et al\. \(2026\)](https://arxiv.org/html/2609.21509#bib.bib59)formalize this as an information bottleneck, and[Tran and Kiela \(2026\)](https://arxiv.org/html/2609.21509#bib.bib60)show that multi\-agent decomposition inherently loses information through the NL channel\. Our communication matrix provides direct empirical support, and the 60\.4 pp role\-swap effect has immediate implications for pipeline design\. Cross\-model\([Yin et al\., 2023](https://arxiv.org/html/2609.21509#bib.bib70)\)and heterogeneous\-LLM\([Ye et al\., 2025b](https://arxiv.org/html/2609.21509#bib.bib71)\)pipelines improve over single\-model baselines but do not measure text\-interface loss, while LatentMAS\([Zou et al\., 2025](https://arxiv.org/html/2609.21509#bib.bib72)\)sidesteps it by collaborating in a shared latent space\.

Structured channels for LLMs\.Early frameworks bridged LLMs with structured interfaces via SQL\-like operations\([Jiang et al\., 2023](https://arxiv.org/html/2609.21509#bib.bib31)\), but recent work identifies a persistent serialization gap: linearizing structured data into prose breaks order invariance and makes reasoning brittle to arbitrary formatting choices\([Herbst et al\., 2025](https://arxiv.org/html/2609.21509#bib.bib34)\)\. Remediation moves beyond text\-wrapping through hypergraph encodings\([Huang et al\., 2025](https://arxiv.org/html/2609.21509#bib.bib30)\)and table\-native architectures\([Li et al\., 2025a](https://arxiv.org/html/2609.21509#bib.bib33)\), yet frontier models still show format\-dependent accuracy on structured outputs\([Elnashar et al\., 2025](https://arxiv.org/html/2609.21509#bib.bib32)\)\. Our NL tax quantifies this loss in a controlled symbolic setting\. Recent proposals replace NL with dense vectors\([Wu and Wang, 2025](https://arxiv.org/html/2609.21509#bib.bib74)\)or a layered telecom\-style protocol\([Li et al\., 2025b](https://arxiv.org/html/2609.21509#bib.bib75)\), and our NL tax is the floor such alternatives would have to beat\.

Compositional generalization\.SCAN\([Lake and Baroni, 2018](https://arxiv.org/html/2609.21509#bib.bib38)\)and CFQ\([Keysers et al\., 2020](https://arxiv.org/html/2609.21509#bib.bib39)\)evaluate compositional*parsing*\(text\-to\-structure\)\. Transformer performance decays with depth\([Dziri et al\., 2023](https://arxiv.org/html/2609.21509#bib.bib40)\), shows a “compositionality gap”\([Press et al\., 2023](https://arxiv.org/html/2609.21509#bib.bib41)\), and collapses at high complexity\([Thomm et al\., 2024](https://arxiv.org/html/2609.21509#bib.bib42);[Mirzadeh et al\., 2024](https://arxiv.org/html/2609.21509#bib.bib69);[Shojaee et al\., 2025](https://arxiv.org/html/2609.21509#bib.bib68)\)\. The same gradients appear in math reasoning\([Stolfo et al\., 2023](https://arxiv.org/html/2609.21509#bib.bib23);[Opedal et al\., 2025](https://arxiv.org/html/2609.21509#bib.bib25)\), evaluated as solving rather than transmission\. We invert the direction to*generation*\(structure\-to\-text\), connecting these failures to surface realization\([Reiter and Dale, 2000](https://arxiv.org/html/2609.21509#bib.bib43);[Zhang et al\., 2006](https://arxiv.org/html/2609.21509#bib.bib44)\)via linearization complexityℓ⁡\(T\)\\ell\(T\)\. Recent work uses equation\-to\-word\-problem generation for data augmentation\([Chen et al\., 2024a](https://arxiv.org/html/2609.21509#bib.bib62)\), but does not ask whether the text preserves enough structure for a second model to recover the original\.

## 6Discussion

The bottleneck is compositional serialization, not domain knowledge\.Swapping arithmetic for logical operators changes absolute difficulty but preserves rankings and the generator\-dominant fault pattern \([AppendixR](https://arxiv.org/html/2609.21509#A18)\)\. Fine\-tuning on a novel operator vocabulary transfers broadly, lifting every open\-weight generator but still leaving a gap to the frontier \([Section4\.1](https://arxiv.org/html/2609.21509#S4.SS1)\)\. Under matched operator semantics,∼3600\\sim\\\!3600examples lift every open\-weight model above untrained Gemini\-3\.1\-Pro, so the limit is trainable rather than architectural\. What stays hard is flattening a hierarchical expression into words another model can re\-parse, a skill chain\-of\-thought and multi\-agent pipelines typically rely on\.

Practical recommendations\.Role assignment matters: swapping generator and extractor shifts accuracy by up to60\.460\.4pp, and the best pair \(92\.9%92\.9\\%\) mixes two different models, so match models to roles by their per\-role rank rather than overall strength\. Prefer structured intermediates when the consumer can parse them\([Gao et al\., 2023](https://arxiv.org/html/2609.21509#bib.bib50);[Chen et al\., 2023](https://arxiv.org/html/2609.21509#bib.bib51)\)or when an interoperability standard such as MCP\([Anthropic, 2025c](https://arxiv.org/html/2609.21509#bib.bib78)\)or A2A\([Google LLC and Linux Foundation Contributors, 2025](https://arxiv.org/html/2609.21509#bib.bib79)\)already provides the schema, the NL tax reaches 18–20 pp at 14B–32B \([Figure3](https://arxiv.org/html/2609.21509#S3.F3)\-right\)\.

Limitations\.The protocol covers arithmetic with 2–8 binary operators and a logic replication \([AppendixR](https://arxiv.org/html/2609.21509#A18)\) which both admit a canonical form, so any two correct answers are provably equivalent\. Our protocol requires this kind of oracle, so it does not directly transfer to domains where correctness admits many answers with no canonical form to check against, only validators specific to each instance, like planning, spatial reasoning, and program synthesis\. Evaluation is monolingual \(English\) with one prompt pair and no tuning, so scores are a reproducible lower bound\. Thinking modes are disabled to isolate serialization from test\-time reasoning and keep the panel comparable\.

## 7Conclusion

We introduced a round\-trip protocol that measures how faithfully language models transmit compositional structure through natural language\. What fails is predictable: operator count and tree shape dominate, while model family does not\. A logic replication keeps most of the ranking structure, so the failure pattern is not unique to arithmetic even though absolute difficulty changes\. What fails is also partially fixable: arithmetic fine\-tuning saturates the channel for every tested model, and structural gains transfer to non\-fine\-tuned extractors\. Natural language is a lossy channel for compositional structure, one whose capacity scales with model capability but remains structurally bounded\. As language models increasingly communicate with each other, understanding the capacity and failure modes of this channel becomes a first\-order concern for system design\.

## References

- Abdinet al\.\(2024\)M\. Abdin, J\. Aneja, H\. Behl, S\. Bubeck, R\. Eldan, S\. Gunasekar,et al\.Phi\-4 technical report\.arXiv preprint arXiv:2412\.08905\.Cited by:[Table 11](https://arxiv.org/html/2609.21509#A10.T11.6.4.2),[Appendix J](https://arxiv.org/html/2609.21509#A10.p2.1),[§3\.1](https://arxiv.org/html/2609.21509#S3.SS1.p1.1)\.
- Allamaniset al\.\(2024\)M\. Allamanis, S\. Panthaplackel, and P\. YinUnsupervised evaluation of code LLMs with round\-trip correctness\.InProceedings of the 41st International Conference on Machine Learning,Vol\.235,pp\. 1050–1066\.External Links:[Link](https://proceedings.mlr.press/v235/allamanis24a.html)Cited by:[§1](https://arxiv.org/html/2609.21509#S1.p2.1),[§5](https://arxiv.org/html/2609.21509#S5.p1.1)\.
- Anthropic \(2025a\)AnthropicClaude 3\.7 sonnet and claude code\.External Links:[Link](https://www.anthropic.com/news/claude-3-7-sonnet)Cited by:[Appendix J](https://arxiv.org/html/2609.21509#A10.p2.1)\.
- Anthropic \(2025b\)AnthropicClaude haiku 4\.5\.External Links:[Link](https://www.anthropic.com/claude/haiku)Cited by:[Table 11](https://arxiv.org/html/2609.21509#A10.T11.6.5.2),[Appendix J](https://arxiv.org/html/2609.21509#A10.p2.1),[§3\.1](https://arxiv.org/html/2609.21509#S3.SS1.p1.1)\.
- Anthropic \(2025c\)AnthropicModel Context Protocol Specification\.Note:[https://modelcontextprotocol\.io](https://modelcontextprotocol.io/)Accessed: 2026\-07\-26Cited by:[§6](https://arxiv.org/html/2609.21509#S6.p2.1)\.
- Anthropic \(2026a\)AnthropicIntroducing claude opus 4\.6\.Note:[https://www\.anthropic\.com/news/claude\-opus\-4\-6](https://www.anthropic.com/news/claude-opus-4-6)Cited by:[§T\.2](https://arxiv.org/html/2609.21509#A20.SS2.p1.1)\.
- Anthropic \(2026b\)AnthropicIntroducing claude opus 4\.7\.Note:[https://www\.anthropic\.com/news/claude\-opus\-4\-7](https://www.anthropic.com/news/claude-opus-4-7)Cited by:[§T\.2](https://arxiv.org/html/2609.21509#A20.SS2.p1.1)\.
- Austinet al\.\(2021\)J\. Austin, A\. Odena, M\. Nye, M\. Bosma, H\. Michalewski, D\. Dohan, E\. Jiang, C\. Cai, M\. Terry, Q\. Le, and C\. SuttonProgram synthesis with large language models\.arXiv preprint arXiv:2108\.07732\.External Links:[Link](https://arxiv.org/abs/2108.07732)Cited by:[Appendix J](https://arxiv.org/html/2609.21509#A10.p2.1)\.
- Chenet al\.\(2024a\)N\. Chen, N\. Wu, J\. Chang, and J\. LiControlMath: controllable data generation promotes math generalist models\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,External Links:[Link](https://arxiv.org/abs/2409.15376)Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p4.1)\.
- Chenet al\.\(2024b\)W\. Chen, C\. Yuan, J\. Yuan, Y\. Su, C\. Qian, C\. Yang, R\. Xie, Z\. Liu, and M\. SunBeyond natural language: llms leveraging alternative formats for enhanced reasoning and communication\.InFindings of the Association for Computational Linguistics: EMNLP 2024,pp\. 10626–10641\.Cited by:[§1](https://arxiv.org/html/2609.21509#S1.p3.1)\.
- Chenet al\.\(2023\)W\. Chen, X\. Ma, X\. Wang, and W\. W\. CohenProgram of thoughts prompting: disentangling computation from reasoning for numerical reasoning tasks\.TMLR\.Cited by:[§6](https://arxiv.org/html/2609.21509#S6.p2.1)\.
- Cobbeet al\.\(2021\)K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano, C\. Hesse, and J\. SchulmanTraining verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168\.Cited by:[Appendix J](https://arxiv.org/html/2609.21509#A10.p2.1),[§1](https://arxiv.org/html/2609.21509#S1.p2.1)\.
- Dettmerset al\.\(2023\)T\. Dettmers, A\. Pagnoni, A\. Holtzman, and L\. ZettlemoyerQLORA: efficient finetuning of quantized llms\.InProceedings of the 37th International Conference on Neural Information Processing Systems,NIPS ’23,Red Hook, NY, USA\.Cited by:[§4](https://arxiv.org/html/2609.21509#S4.p4.1)\.
- Duet al\.\(2024\)Y\. Du, S\. Li, A\. Torralba, J\. B\. Tenenbaum, and I\. MordatchImproving factuality and reasoning in language models through multiagent debate\.InForty\-first international conference on machine learning,Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p2.1)\.
- Dziriet al\.\(2023\)N\. Dziri, X\. Lu, M\. Sclar, X\. L\. Li, L\. Jiang, B\. Y\. Lin, S\. Welleck, P\. West, C\. Bhagavatula, R\. Le Bras,et al\.Faith and fate: limits of transformers on compositionality\.Advances in neural information processing systems36,pp\. 70293–70332\.Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p4.1)\.
- Elnasharet al\.\(2025\)A\. Elnashar, J\. White, and D\. C\. SchmidtEnhancing structured data generation with gpt\-4o evaluating prompt efficiency across prompt styles\.Frontiers in Artificial IntelligenceVolume 8 \- 2025\.External Links:[Link](https://www.frontiersin.org/journals/artificial-intelligence/articles/10.3389/frai.2025.1558938),[Document](https://dx.doi.org/10.3389/frai.2025.1558938),ISSN 2624\-8212Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p3.1)\.
- Gaoet al\.\(2023\)L\. Gao, A\. Madaan, S\. Zhou, U\. Alon, P\. Liu, Y\. Yang, J\. Callan, and G\. NeubigPal: program\-aided language models\.InInternational conference on machine learning,pp\. 10764–10799\.Cited by:[§6](https://arxiv.org/html/2609.21509#S6.p2.1)\.
- Gemma Team \(2025\)Gemma TeamGemma 3 technical report\.arXiv preprint arXiv:2503\.19786\.Cited by:[Table 11](https://arxiv.org/html/2609.21509#A10.T11.6.3.2),[Appendix J](https://arxiv.org/html/2609.21509#A10.p2.1),[§3\.1](https://arxiv.org/html/2609.21509#S3.SS1.p1.1)\.
- Glazeret al\.\(2025\)E\. Glazer, E\. Erdil, T\. Besiroglu, D\. Chicharro, E\. Chen, A\. Gunning, C\. F\. Olsson, J\. Denain, A\. Ho, E\. de Oliveira Santos, O\. Järviniemi, M\. Barnett, R\. Sandler, M\. Vrzala, J\. Sevilla, Q\. Ren, E\. Pratt, L\. Levine, G\. Barkley, N\. Stewart, B\. Grechuk, T\. Grechuk, S\. V\. Enugandla, and M\. WildonFrontierMath: a benchmark for evaluating advanced mathematical reasoning in ai\.External Links:2411\.04872,[Link](https://arxiv.org/abs/2411.04872)Cited by:[§1](https://arxiv.org/html/2609.21509#S1.p2.1)\.
- Google DeepMind \(2025a\)Google DeepMindGemini 3 flash model card\.External Links:[Link](https://deepmind.google/models/model-cards/)Cited by:[§3\.1](https://arxiv.org/html/2609.21509#S3.SS1.p1.1)\.
- Google DeepMind \(2025b\)Google DeepMindGemini 3 pro model card\.External Links:[Link](https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf)Cited by:[§3\.1](https://arxiv.org/html/2609.21509#S3.SS1.p1.1)\.
- Google LLC and Linux Foundation Contributors \(2025\)Google LLC and Linux Foundation ContributorsAgent2Agent \(A2A\) Protocol Specification\.Note:[https://github\.com/google\-a2a/A2A](https://github.com/google-a2a/A2A)Apache 2\.0 License\. Accessed: 2026\-07\-03Cited by:[§6](https://arxiv.org/html/2609.21509#S6.p2.1)\.
- Heet al\.\(2026\)S\. He, A\. Narayan, I\. S\. Khare, S\. Linderman, C\. Ré, and D\. BidermanAn information theoretic perspective on agentic system design\.InInternational Conference on Learning Representations,External Links:[Link](https://arxiv.org/abs/2512.21720)Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p2.1)\.
- Hendryckset al\.\(2021a\)D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. SteinhardtMeasuring massive multitask language understanding\.InInternational Conference on Learning Representations,External Links:[Link](https://arxiv.org/abs/2009.03300)Cited by:[Appendix J](https://arxiv.org/html/2609.21509#A10.p2.1)\.
- Hendryckset al\.\(2021b\)D\. Hendrycks, C\. Burns, S\. Kadavath, A\. Arora, S\. Basart, E\. Tang, D\. Song, and J\. SteinhardtMeasuring mathematical problem solving with the math dataset\.NeurIPS\.Cited by:[Appendix J](https://arxiv.org/html/2609.21509#A10.p2.1),[§1](https://arxiv.org/html/2609.21509#S1.p2.1)\.
- Herbstet al\.\(2025\)D\. Herbst, L\. Karbevska, D\. Kumar, A\. Ahuja, F\. G\. Nasrabadi, and F\. FrascaLost in serialization: invariance and generalization of llm graph reasoners\.AAAI Workshop on Graphs and more Complex Structures For Learning and Reasoning \(GCLR\)\.External Links:[Link](https://arxiv.org/abs/2511.10234)Cited by:[§1](https://arxiv.org/html/2609.21509#S1.p1.1),[§5](https://arxiv.org/html/2609.21509#S5.p3.1)\.
- Honget al\.\(2024\)S\. Hong, M\. Zhuge, J\. Chen, X\. Zheng, Y\. Cheng, J\. Wang, C\. Zhang, Z\. Wang, S\. K\. S\. Yau, Z\. Lin,et al\.MetaGPT: meta programming for a multi\-agent collaborative framework\.InThe Twelfth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.21509#S1.p1.1),[§5](https://arxiv.org/html/2609.21509#S5.p2.1)\.
- Huanget al\.\(2025\)S\. Huang, H\. Li, Y\. Gu, X\. Hu, Q\. Li, and G\. XuHyperg: hypergraph\-enhanced llms for structured knowledge\.InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 1218–1228\.Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p3.1)\.
- Jainet al\.\(2024\)N\. Jain, K\. Han, A\. Gu, W\. Li, F\. Yan, T\. Zhang, S\. Wang, A\. Solar\-Lezama, K\. Sen, and I\. StoicaLiveCodeBench: holistic and contamination free evaluation of large language models for code\.arXiv preprint arXiv:2403\.07974\.Cited by:[Appendix J](https://arxiv.org/html/2609.21509#A10.p2.1),[§1](https://arxiv.org/html/2609.21509#S1.p2.1)\.
- Jianget al\.\(2023\)J\. Jiang, K\. Zhou, Z\. Dong, K\. Ye, W\. X\. Zhao, and J\. WenStructgpt: a general framework for large language model to reason over structured data\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 9237–9251\.Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p3.1)\.
- Jimenezet al\.\(2024\)C\. E\. Jimenez, J\. Yang, A\. Wettig, S\. Yao, K\. Pei, O\. Press, and K\. R\. NarasimhanSWE\-bench: can language models resolve real\-world github issues?\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=VTF8yNQM66)Cited by:[§1](https://arxiv.org/html/2609.21509#S1.p2.1)\.
- Keyserset al\.\(2020\)D\. Keysers, N\. Schärli, N\. Scales, H\. Buisman, D\. Furrer, S\. Kashubin, N\. Momchev, D\. Sinopalnikov, L\. Stafiniak, T\. Tihon, D\. Tsarkov, X\. Wang, M\. van Zee, and O\. BousquetMeasuring compositional generalization: a comprehensive method on realistic data\.InInternational Conference on Learning Representations,Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p4.1)\.
- Konget al\.\(2025\)L\. Kong, X\. Wang, Y\. Chen, and M\. ZhangRound\-trip reinforcement learning: self\-consistent training for better chemical LLMs\.arXiv preprint arXiv:2510\.01527\.External Links:[Link](https://arxiv.org/abs/2510.01527)Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p1.1)\.
- Lake and Baroni \(2018\)B\. Lake and M\. BaroniGeneralization without systematicity: on the compositional skills of sequence\-to\-sequence recurrent networks\.InInternational conference on machine learning,pp\. 2873–2882\.Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p4.1)\.
- Lanhamet al\.\(2023\)T\. Lanham, A\. Chen, A\. Radhakrishnan, B\. Steiner, C\. Denison, D\. Hernandez, D\. Li, E\. Durmus, E\. Hubinger, J\. Kernion,et al\.Measuring faithfulness in chain\-of\-thought reasoning\.arXiv preprint arXiv:2307\.13702\.External Links:[Link](https://arxiv.org/abs/2307.13702)Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p1.1)\.
- Liet al\.\(2023\)G\. Li, H\. Hammoud, H\. Itani, D\. Khizbullin, and B\. GhanemCamel: communicative agents for" mind" exploration of large language model society\.Advances in neural information processing systems36,pp\. 51991–52008\.Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p2.1)\.
- Liet al\.\(2025a\)L\. Li, C\. Ye, W\. Ye, Y\. Sun, Z\. Jiang, H\. Wang, J\. Tian, Y\. Zhang, N\. Wang, X\. Fu,et al\.Table as a modality for large language models\.Advances in Neural Information Processing Systems\.Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p3.1)\.
- Liet al\.\(2025b\)X\. Li, M\. Liu, and C\. YuenLLM agent communication protocol \(LACP\) requires urgent standardization: a telecom\-inspired protocol is necessary\.arXiv preprint arXiv:2510\.13821\.Note:NeurIPS 2025 AI4NextG WorkshopExternal Links:[Link](https://arxiv.org/abs/2510.13821)Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p3.1)\.
- Loshchilov and Hutter \(2019\)I\. Loshchilov and F\. HutterDecoupled weight decay regularization\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=Bkg6RiCqY7)Cited by:[§4](https://arxiv.org/html/2609.21509#S4.p4.1)\.
- Lundberget al\.\(2019\)S\. M\. Lundberg, G\. Erion, H\. Chen, A\. DeGrave, J\. M\. Prutkin, B\. Nair, R\. Katz, J\. Himmelfarb, N\. Bansal, and S\. LeeExplainable ai for trees: from local explanations to global understanding\.arXiv preprint arXiv:1905\.04610\.Cited by:[Appendix G](https://arxiv.org/html/2609.21509#A7.p3.1),[§3\.3](https://arxiv.org/html/2609.21509#S3.SS3.p1.1)\.
- Lundberg and Lee \(2017\)S\. M\. Lundberg and S\. LeeA unified approach to interpreting model predictions\.InAdvances in Neural Information Processing Systems,External Links:[Link](https://proceedings.neurips.cc/paper/7062-a-unified-approach-to-interpreting-model-predictions)Cited by:[§3\.3](https://arxiv.org/html/2609.21509#S3.SS3.p1.1)\.
- Maveliet al\.\(2026\)N\. Maveli, A\. Vergari, and S\. B\. CohenCan LLMs compress \(and decompress\)? evaluating code understanding and execution via invertibility\.arXiv preprint arXiv:2601\.13398\.External Links:[Link](https://arxiv.org/abs/2601.13398)Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p1.1)\.
- Meureret al\.\(2017\)A\. Meurer, C\. P\. Smith, M\. Paprocki, O\. Čertík, S\. B\. Kirpichev, M\. Rocklin, A\. Kumar, S\. Ivanov, J\. K\. Moore, S\. Singh, T\. Rathnayake, S\. Vig, B\. E\. Granger, R\. P\. Muller, F\. Bonazzi, H\. Gupta, S\. Vats, F\. Johansson, F\. Pedregosa, M\. J\. Curry, A\. R\. Terrel, Š\. Roučka, A\. Saboo, I\. Fernando, S\. Kulal, R\. Cimrman, and A\. ScopatzSymPy: symbolic computing in python\.PeerJ Computer Science3,pp\. e103\.External Links:ISSN 2376\-5992,[Link](https://doi.org/10.7717/peerj-cs.103),[Document](https://dx.doi.org/10.7717/peerj-cs.103)Cited by:[§2\.2](https://arxiv.org/html/2609.21509#S2.SS2.p1.1)\.
- Mirzadehet al\.\(2024\)I\. Mirzadeh, K\. Alizadeh, H\. Shahrokhi, O\. Tuzel, S\. Bengio, and M\. FarajtabarGSM\-symbolic: understanding the limitations of mathematical reasoning in large language models\.External Links:[Link](https://arxiv.org/abs/2410.05229)Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p4.1)\.
- Nieet al\.\(2025\)S\. Nieet al\.LLaDA: large language diffusion models\.39th Conference on Neural Information Processing Systems\.Cited by:[§3\.1](https://arxiv.org/html/2609.21509#S3.SS1.p1.1)\.
- Opedalet al\.\(2025\)A\. Opedal, H\. Shirakami, B\. Schölkopf, A\. Saparov, and M\. SachanMathGAP: out\-of\-distribution evaluation on problems with arbitrarily complex proofs\.International Conference on Learning Representations\.Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p4.1)\.
- Paperno \(2022\)D\. PapernoOn learning interpreted languages with recurrent models\.Computational Linguistics48\(2\),pp\. 471–482\.External Links:[Link](https://aclanthology.org/2022.cl-2.7/)Cited by:[Appendix H](https://arxiv.org/html/2609.21509#A8.p4.1)\.
- Pedregosaet al\.\(2011\)F\. Pedregosa, G\. Varoquaux, A\. Gramfort, V\. Michel, B\. Thirion, O\. Grisel, M\. Blondel, P\. Prettenhofer, R\. Weiss, V\. Dubourg, J\. Vanderplas, A\. Passos, D\. Cournapeau, M\. Brucher, M\. Perrot, and É\. DuchesnayScikit\-learn: machine learning in Python\.Journal of Machine Learning Research12,pp\. 2825–2830\.Cited by:[Appendix G](https://arxiv.org/html/2609.21509#A7.p2.1)\.
- Phanet al\.\(2026\)L\. Phan, A\. Gatti, N\. Li, A\. Khoja, R\. Kim, R\. Ren, J\. Hausenloy, O\. Zhang, M\. Mazeika, D\. Hendrycks,et al\.A benchmark of expert\-level academic questions to assess ai capabilities\.Nature649\(8099\),pp\. 1139–1146\.Cited by:[§1](https://arxiv.org/html/2609.21509#S1.p2.1)\.
- Presset al\.\(2023\)O\. Press, M\. Zhang, S\. Min, L\. Schmidt, N\. A\. Smith, and M\. LewisMeasuring and narrowing the compositionality gap in language models\.InFindings of the Association for Computational Linguistics: EMNLP 2023,pp\. 5687–5711\.Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p4.1)\.
- Rasch \(1980\)G\. RaschProbabilistic models for some intelligence and attainment tests\.The SAGE Encyclopedia of Research Design\.External Links:[Link](https://api.semanticscholar.org/CorpusID:61203382)Cited by:[Appendix N](https://arxiv.org/html/2609.21509#A14.p1.1)\.
- Reinet al\.\(2023\)D\. Rein, B\. L\. Hou, A\. C\. Stickland, J\. Petty, R\. Y\. Pang, J\. Dirani, J\. Michael, and S\. R\. BowmanGPQA: a graduate\-level google\-proof q&a benchmark\.External Links:2311\.12022,[Link](https://arxiv.org/abs/2311.12022)Cited by:[Appendix J](https://arxiv.org/html/2609.21509#A10.p2.1),[§1](https://arxiv.org/html/2609.21509#S1.p2.1)\.
- Reiter and Dale \(2000\)E\. Reiter and R\. DaleBuilding natural language generation systems\.Studies in Natural Language Processing,Cambridge University Press\.Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p4.1)\.
- Schicket al\.\(2023\)T\. Schick, J\. Dwivedi\-Yu, R\. Dessì, R\. Raileanu, M\. Lomeli, E\. Hambro, L\. Zettlemoyer, N\. Cancedda, and T\. ScialomToolformer: language models can teach themselves to use tools\.Advances in neural information processing systems36,pp\. 68539–68551\.Cited by:[§1](https://arxiv.org/html/2609.21509#S1.p1.1)\.
- Sennrichet al\.\(2016\)R\. Sennrich, B\. Haddow, and A\. BirchImproving neural machine translation models with monolingual data\.InProceedings of the 54th annual meeting of the association for computational linguistics \(volume 1: long papers\),pp\. 86–96\.Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p1.1)\.
- Sethi and Ullman \(1970\)R\. Sethi and J\. D\. UllmanThe generation of optimal code for arithmetic expressions\.J\. ACM17,pp\. 715–728\.External Links:[Link](https://api.semanticscholar.org/CorpusID:10845848)Cited by:[§3\.3](https://arxiv.org/html/2609.21509#S3.SS3.p1.1)\.
- Shiet al\.\(2023\)F\. Shi, M\. Suzgun, M\. Freitag, X\. Wang, S\. Srivats, S\. Vosoughi, H\. W\. Chung, Y\. Tay, S\. Ruder, D\. Zhou, D\. Das, and J\. WeiLanguage models are multilingual chain\-of\-thought reasoners\.InInternational Conference on Learning Representations,External Links:[Link](https://arxiv.org/abs/2210.03057)Cited by:[Appendix J](https://arxiv.org/html/2609.21509#A10.p2.1)\.
- Shojaeeet al\.\(2025\)P\. Shojaee, I\. Mirzadeh, K\. Alizadeh, M\. Horton, S\. Bengio, and M\. FarajtabarThe illusion of thinking: understanding the strengths and limitations of reasoning models via the lens of problem complexity\.Advances in Neural Information Processing Systems\.External Links:[Link](https://arxiv.org/abs/2506.06941)Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p4.1)\.
- Singhet al\.\(2025\)A\. Singh, A\. Fry, A\. Perelman, A\. Tart, A\. Ganesh, A\. El\-Kishky, A\. McLaughlin, A\. Low, A\. Ostrow, A\. Ananthram,et al\.Openai gpt\-5 system card\.arXiv preprint arXiv:2601\.03267\.External Links:[Link](https://arxiv.org/abs/2601.03267)Cited by:[Table 11](https://arxiv.org/html/2609.21509#A10.T11.6.6.2),[Appendix J](https://arxiv.org/html/2609.21509#A10.p2.1),[§3\.1](https://arxiv.org/html/2609.21509#S3.SS1.p1.1)\.
- Stolfoet al\.\(2023\)A\. Stolfo, Z\. Jin, K\. Shridhar, B\. Schölkopf, and M\. SachanA causal framework to quantify the robustness of mathematical reasoning with language models\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 545–561\.Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p4.1)\.
- Suzgunet al\.\(2023\)M\. Suzgun, N\. Scales, N\. Schärli, S\. Gehrmann, Y\. Tay, H\. W\. Chung, A\. Chowdhery, Q\. V\. Le, E\. H\. Chi, D\. Zhou, and J\. WeiChallenging BIG\-Bench tasks and whether chain\-of\-thought can solve them\.InFindings of the Association for Computational Linguistics: ACL 2023,pp\. 13003–13051\.External Links:[Link](https://arxiv.org/abs/2210.09261)Cited by:[Appendix J](https://arxiv.org/html/2609.21509#A10.p2.1)\.
- Thommet al\.\(2024\)J\. Thomm, G\. Camposampiero, A\. Terzic, M\. Hersche, B\. Schölkopf, and A\. RahimiLimits of transformer language models on learning to compose algorithms\.Advances in Neural Information Processing Systems37,pp\. 7631–7674\.Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p4.1)\.
- Tran and Kiela \(2026\)D\. Tran and D\. KielaSingle\-agent LLMs outperform multi\-agent systems on multi\-hop reasoning under equal thinking token budgets\.arXiv preprint arXiv:2604\.02460\.External Links:[Link](https://arxiv.org/abs/2604.02460)Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p2.1)\.
- Turpinet al\.\(2023\)M\. Turpin, J\. Michael, E\. Perez, and S\. BowmanLanguage models don’t always say what they think: unfaithful explanations in chain\-of\-thought prompting\.Advances in Neural Information Processing Systems36,pp\. 74952–74965\.Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p1.1)\.
- Wanget al\.\(2025\)J\. Wang, J\. Wang, B\. Athiwaratkun, C\. Zhang, and J\. ZouMixture\-of\-agents enhances large language model capabilities\.International Conference on Learning Representations\.Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p2.1)\.
- Wanget al\.\(2023\)X\. Wang, J\. Wei, D\. Schuurmans, Q\. Le, E\. Chi, S\. Narang, A\. Chowdhery, and D\. ZhouSelf\-consistency improves chain of thought reasoning in language models\.International Conference on Learning Representations\.External Links:[Link](https://arxiv.org/abs/2203.11171)Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p1.1)\.
- Wanget al\.\(2024\)Y\. Wang, X\. Ma, G\. Zhang, Y\. Ni, A\. Chandra, S\. Guo, W\. Ren, A\. Arulraj, X\. He, Z\. Jiang, T\. Li, M\. Ku, K\. Wang, A\. Zhuang, R\. Fan, X\. Yue, and W\. ChenMMLU\-Pro: a more robust and challenging multi\-task language understanding benchmark\.InAdvances in Neural Information Processing Systems,External Links:[Link](https://arxiv.org/abs/2406.01574)Cited by:[Appendix J](https://arxiv.org/html/2609.21509#A10.p2.1)\.
- Weiet al\.\(2022\)J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, F\. Xia, E\. Chi, Q\. V\. Le, D\. Zhou,et al\.Chain\-of\-thought prompting elicits reasoning in large language models\.Advances in neural information processing systems35,pp\. 24824–24837\.Cited by:[§1](https://arxiv.org/html/2609.21509#S1.p1.1)\.
- Wigleret al\.\(2026\)B\. Wigler, M\. Tsfasman, and T\. M\. HrkalovicStories of your life as others: a round\-trip evaluation of llm\-generated life stories conditioned on rich psychometric profiles\.arXiv preprint arXiv:2604\.06071\.Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p1.1)\.
- Wuet al\.\(2024\)Q\. Wu, G\. Bansal, J\. Zhang, Y\. Wu, B\. Li, E\. Zhu, L\. Jiang, X\. Zhang, S\. Zhang, J\. Liu,et al\.Autogen: enabling next\-gen llm applications via multi\-agent conversations\.InFirst conference on language modeling,Cited by:[§1](https://arxiv.org/html/2609.21509#S1.p1.1),[§5](https://arxiv.org/html/2609.21509#S5.p2.1)\.
- Wu and Wang \(2025\)S\. Wu and Y\. WangDense communication between language models\.arXiv preprint arXiv:2505\.12741\.External Links:[Link](https://arxiv.org/abs/2505.12741)Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p3.1)\.
- Yanget al\.\(2025\)A\. Yanget al\.Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[Table 11](https://arxiv.org/html/2609.21509#A10.T11.6.2.2),[Appendix J](https://arxiv.org/html/2609.21509#A10.p2.1),[§3\.1](https://arxiv.org/html/2609.21509#S3.SS1.p1.1)\.
- Yeet al\.\(2025a\)J\. Ye, Z\. Xie, L\. Zheng, J\. Gao, Z\. Wu, X\. Jiang, Z\. Li, and L\. KongDream 7b: diffusion large language models\.arXiv preprint arXiv:2508\.15487\.Cited by:[§3\.1](https://arxiv.org/html/2609.21509#S3.SS1.p1.1)\.
- Yeet al\.\(2025b\)R\. Ye, X\. Liu, Q\. Wu, X\. Pang, Z\. Yin, L\. Bai, and S\. ChenX\-mas: towards building multi\-agent systems with heterogeneous llms\.arXiv preprint arXiv:2505\.16997\.External Links:[Link](https://arxiv.org/abs/2505.16997)Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p2.1)\.
- Yinet al\.\(2023\)Z\. Yin, Q\. Sun, C\. Chang, Q\. Guo, J\. Dai, X\. Huang, and X\. QiuExchange\-of\-thought: enhancing large language model capabilities through cross\-model communication\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),External Links:[Link](https://arxiv.org/abs/2312.01823)Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p2.1)\.
- Zhanget al\.\(2006\)H\. Zhang, L\. Huang, D\. Gildea, and K\. KnightSynchronous binarization for machine translation\.InProceedings of the Human Language Technology Conference of the NAACL, Main Conference,R\. C\. Moore, J\. Bilmes, J\. Chu\-Carroll, and M\. Sanderson \(Eds\.\),New York City, USA,pp\. 256–263\.External Links:[Link](https://aclanthology.org/N06-1033/)Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p4.1)\.
- Zhanget al\.\(2024\)H\. Zhang, J\. Da, D\. Lee, V\. Robinson, C\. Wu, W\. Song, T\. Zhao, P\. Raja, C\. Zhuang, D\. Slack,et al\.A careful examination of large language model performance on grade school arithmetic\.Advances in Neural Information Processing Systems37,pp\. 46819–46836\.Cited by:[§1](https://arxiv.org/html/2609.21509#S1.p3.1)\.
- Zhouet al\.\(2025\)P\. Zhou, Y\. Feng, H\. Julaiti, and Z\. YangWhy do ai agents communicate in human language?\.arXiv preprint arXiv:2506\.02739\.Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p2.1)\.
- Zouet al\.\(2025\)J\. Zou, X\. Yang, R\. Qiu, G\. Li, K\. Tieu, P\. Lu, K\. Shen, H\. Tong, Y\. Choi, J\. He, J\. Zou, M\. Wang, and L\. YangLatent collaboration in multi\-agent systems\.arXiv preprint arXiv:2511\.20639\.External Links:[Link](https://arxiv.org/abs/2511.20639)Cited by:[§5](https://arxiv.org/html/2609.21509#S5.p2.1)\.

## Appendix APrompt Templates

Generator system prompt\.The following system prompt is used for the generation phase, with\{domain\}and\{variables\}filled per expression:

> ``` Convert a mathematical expression into a short word problem. Rules: 1. Set the word problem in the domain of: {domain}. 2. Use EXACTLY these variable placeholders in curly braces: {variables}. 3. If the expression contains integer constants, use them as literal numbers in the word problem (e.g. the constant 8 becomes "8" in the text). 4. Each operation in the expression must correspond to a clear action in the story. 5. Do NOT include any mathematical symbols (+, -, *, /, parentheses) in your response. 6. Keep it to 2-3 sentences. 7. Output ONLY the word problem - no explanation, no title. ```

Generator in\-context exemplars\.When the generator runs atγ=3\\gamma\{=\}3, three exemplars are appended to its prompt to formπ𝖦3\{\\color\[rgb\]\{0,0,0\}\{\\color\[rgb\]\{0,0,0\}\\pi\}\_\{\\mathsf\{G\}\}^\{3\}\}\. The exemplars are drawn from a fixed stratified pool \(five entries spanningk∈\{2,3,4,6,8\}k\\in\\\{2,3,4,6,8\\\}andd∈\{2,3,4,5\}d\\in\\\{2,3,4,5\\\}, disjoint from the evaluation set\); the first three entries are used atγ=3\\gamma\{=\}3:

> ``` Expression: ( A + B ) * C [domain: baking] Word problem: A bakery made {A} loaves in the morning and {B} in the afternoon. Each loaf was packed into a box with {C} rolls. How many rolls were packed in total? Expression: ( C - ( A + B + 2 ) ) [domain: real estate] Word problem: A developer had {C} plots of land. They sold {A} plots in the spring, {B} plots in the summer, and set aside 2 for a park. How many plots remained? Expression: ( ( E + D ) - A ) - ( A / C ) [domain: mountaineering] Word problem: A team of climbers reached a peak with {E} supplies and found {D} more along the way. They used {A} for shelter and rationed {A} units of fuel over {C} days, consuming one day’s ration. How many supplies remained? ```

Extractor system prompt\.The following system prompt is used for the extraction phase:

> ``` You are given a word problem where some numbers are replaced by variable placeholders and others appear as literal integers. Identify the mathematical expression that represents the computation described. Rules: 1. Use ONLY the variable names listed and any integer constants from the text. 2. Use parentheses to make the order of operations explicit. 3. Use plain ASCII operators: +, -, *, /. Do NOT use LaTeX, \frac, \times, or any markup. 4. Output ONLY the expression - no explanation, no prose. ```

Extractor in\-context exemplars\.When the extractor runs atε=3\\varepsilon\{=\}3, three exemplars are appended to formπ𝖤3\{\\color\[rgb\]\{0,0,0\}\{\\color\[rgb\]\{0,0,0\}\\pi\}\_\{\\mathsf\{E\}\}^\{3\}\}, drawn from a stratified pool paired with the generator exemplars above but reformatted asVariables / Problem / Expressiontriples:

> ``` Variables: A, B, C Problem: A bakery made {A} loaves in the morning and {B} in the afternoon. Each loaf was packed into a box with {C} rolls. How many rolls were packed in total? Expression: ( A + B ) * C Variables: A, B, C Problem: A developer had {C} plots of land. They sold {A} plots in the spring, {B} plots in the summer, and set aside 2 for a park. How many plots remained? Expression: C - ( A + B + 2 ) Variables: A, C, D, E Problem: A team of climbers reached a peak with {E} supplies and found {D} more along the way. They used {A} for shelter and rationed {A} units of fuel over {C} days, consuming one day’s ration. How many supplies remained? Expression: ( ( E + D ) - A ) - ( A / C ) ```

The main matrix evaluates every pairing at\(γ,ε\)=\(0,3\)\(\\gamma,\\varepsilon\)=\(0,3\); see[AppendixP](https://arxiv.org/html/2609.21509#A16)for the full ablation over\(γ,ε\)∈\{0,3\}2\(\\gamma,\\varepsilon\)\\in\\\{0,3\\\}^\{2\}\.

## Appendix BExpression Suite Details

[Table3](https://arxiv.org/html/2609.21509#A2.T3)shows the number of expressions generated for each feasible combination of operator countkkand depthdd\. Each cell is populated by enumerating all skeletons for that\(k,d\)\(k,d\)pair \(up to 50\) and sampling 5 instantiations per skeleton, following the skeleton enumeration procedure described in[Section2\.1](https://arxiv.org/html/2609.21509#S2.SS1)\. Not all\(k,d\)\(k,d\)pairs are feasible:kkoperators require at least⌈log2⁡\(k\+1\)⌉\\lceil\\log\_\{2\}\(k\+1\)\\rceillevels of nesting, so shallow depths cannot accommodate high operator counts\. All 2450 expressions are used in the main evaluation; each column of the communication matrix is evaluated on the subset for which that generator passed the leakage guard\.

Table 3:Number of expressions per operator countkkand depthdd\. Cells marked “—” correspond to infeasible or empty\(k,d\)\(k,d\)pairs\.d=2d\{=\}2d=3d\{=\}3d=4d\{=\}4d=5d\{=\}5d=6d\{=\}6Totalk=2k\{=\}210————10k=3k\{=\}3520———25k=4k\{=\}4—3040——70k=5k\{=\}5—3010080—210k=6k\{=\}6—20200250160630k=7k\{=\}7—5250250250755k=8k\{=\}8——250250250750Total151058408306602450
## Appendix CLeakage Guard and Generator Faults

The*leakage guard*is a rule\-based filter that rejects any generated word problem in which mathematical syntax has leaked into the prose\. Two additional checks catch further generator faults \(missing placeholders and degenerate repetition\); these are counted as generator errors in the reported accuracies but are not reported as leakage\.

1\. Leakage check: operator adjacency rejection\.Four regular\-expression patterns reject any word problem in which a mathematical operator appears adjacent to a variable letter or digit:

- •\[A\-Z0\-9\]\\s\*\[\+\*/\]\(variable or digit followed by operator\)
- •\[\+\*/\]\\s\*\[A\-Z0\-9\]\(operator followed by variable or digit\)
- •\\\(\\s\*\[A\-Z\]\(open parenthesis followed by variable\)
- •\[A\-Z\]\\s\*\\\)\(variable followed by close parenthesis\)

Before pattern matching, all placeholder tokens\{X\}are replaced with the letterX, so the check operates on the surface form that a human reader would see\. These patterns catch expressions like “A \+ B” or “\(A \* C\)” that leak the formal structure into the word problem\.

2\. Placeholder completeness \(generator fault\)\.Every variable in the original expression must appear as a curly\-brace placeholder \(e\.g\.\{A\},\{B\}\) in the generated text\. A missing placeholder means the model dropped a variable, making faithful extraction impossible\. This is a generator fault rather than a leakage failure\.

3\. Repetition loop detection \(generator fault\)\.The guard checks for degenerate outputs containing three or more consecutive repetitions of any word\-levelnn\-gram \(forn∈\{2,3,4,5\}n\\in\\\{2,3,4,5\\\}\)\. Such repetition indicates the model has entered a generation loop rather than producing a coherent word problem\. This is also a generator fault rather than a leakage failure\.

A word problem must pass all three checks to enter the extraction phase\. TheLeak\\mathrm\{Leak\}% rows in the communication matrices report pass rates on the leakage check only; the other two fault types are counted as errors in the reported accuracies but are not shown as leakage errors\.

#### A tunable guard\.

The guard is a deliberately simple, rule\-based filter that downstream users can tighten or relax for their own setting\. Practitioners building round\-trip evaluations on new domains can reuse the three checks directly, swap in stricter syntactic patterns \(e\.g\., additional operator glyphs or domain\-specific markers\), or replace the rule\-based check with a learned classifier\. The bias\-direction argument below holds under any stricter replacement that still counts guard rejections as round\-trip failures\.

#### Known regex gaps\.

Two specifications in the leakage check are deliberately permissive and admit narrow leak forms that a stricter filter would catch\. The operator class\[\+\*/\]omits the minus sign because\-also serves as an English hyphen, and including it would reject compound adjectives and hyphenated numerals in otherwise leak\-free prose\. Leak strings of the form “A \- B” or “5 \- 3” therefore pass\. The parenthesis patterns match only letter\-adjacent parentheses, so numeric expressions such as “\(5 \- 3\)” also pass\. Both are instances of guard false negatives: they admit word problems into the extraction phase that a stricter filter would reject\. The bias\-direction argument below covers this case\. A stricter guard \(adding\-contextually, and digit\-adjacent parentheses\) would catch more leaks at the cost of some false positives on natural prose\. Because guard rejections remain charged as round\-trip failures, the lower\-bound direction on the NL tax is preserved under any such tightening\.

#### Bias direction under guard false negatives\.

The leakage guard is conservative with respect to the paper’s central claims\. If the guard admits a word problem that contains implicit structural hints \(a guard false negative\), extractors have an easier time recovering the expression, inflating round\-trip accuracy\. Because guard rejections are already charged as round\-trip failures in our protocol \([Section2\.3](https://arxiv.org/html/2609.21509#S2.SS3)\), a stricter guard can only move additional word problems into the failure bucket\. Consequently, the NL tax \([Section3](https://arxiv.org/html/2609.21509#S3)\) and the generator\-fault rate \([Section3\.3](https://arxiv.org/html/2609.21509#S3.SS3)\) are lower bounds on their true values, and the direction of bias strengthens rather than undermines the conclusion that natural language is a lossy channel for compositional structure\. The reported role\-swap asymmetry is a difference of two accuracies and its sign is not pinned down by this argument, so we do not claim it as a lower bound\.

## Appendix DRmathR\_\{\\text\{math\}\}Definition and Lexicon

RmathR\_\{\\text\{math\}\}measures the density of mathematical vocabulary in a generated word problem\. Given a word problemw=\(w1,w2,…,wn\)w=\(w\_\{1\},w\_\{2\},\\ldots,w\_\{n\}\)and a curated lexiconℒ\\mathcal\{L\}, we define

Rmath​\(w\)=∑i=1n𝟙\[wi∈ℒ\]n,R\_\{\\text\{math\}\}\(w\)\\;=\\;\\frac\{\\sum\_\{i=1\}^\{n\}\\mathds\{1\}\[w\_\{i\}\\in\\mathcal\{L\}\]\}\{n\},\(3\)where eachwiw\_\{i\}is lowercased and stripped of punctuation before lookup\. HigherRmathR\_\{\\text\{math\}\}indicates the model relied more heavily on arithmetic\-signalling vocabulary \(e\.g\. “sum,” “divided,” “remaining”\) rather than domain\-specific narrative\.

The lexiconℒ\\mathcal\{L\}contains 109 terms curated from 200 sampled word problems from the full expression set, organised into seven semantic categories\. The complete lexicon is listed below\.

Addition\(11\):*added, adding, adds, combine, combined, plus, sum, total, altogether, collected, gathered\.*

Subtraction\(28\):*subtracted, subtracting, subtracts, minus, remaining, remain, remained, remains, left, leftover, removed, removing, remove, deducted, deducting, lost, loses, fewer, less, spent, spend, spending, spends, reduced, reducing, discarded, unused, withdrawn\.*

Multiplication\(9\):*multiplied, multiplies, times, doubled, tripled, twice, per, each, every\.*

Division\(20\):*divided, divides, dividing, division, split, shared, share, distributed, distributing, distribution, allocated, evenly, equally, equal, among, portion, portions, fraction, half, halved\.*

Result / query\(19\):*calculate, calculating, determine, determined, find, finds, estimate, estimated, how, many, much, what, number, amount, result, average, count, counted, counts\.*

Change\(18\):*increase, increased, decrease, decreased, additional, extra, more, gain, gained, gains, earn, earned, earnings, save, saved, cost, costing, costs\.*

Scale\(4\):*scaled, scaling, factor, rate\.*

## Appendix ETemperature Sensitivity

All main experiments use greedy decoding \(temperatureτ=0\\tau=0\)\. To verify that this choice does not critically bias results, we sweepτ∈\{0\.0,0\.1,…,1\.0\}\\tau\\in\\\{0\.0,0\.1,\\ldots,1\.0\\\}on Phi\-4 with 5 independent seeds per temperature point\.

[Figure5](https://arxiv.org/html/2609.21509#A5.F5)summarises the findings\. Guard pass rate \(left panel\) declines gradually from89\.6%89\.6\\%\(τ=0\\tau=0\) to70\.5%70\.5\\%\(τ=1\\tau=1\), confirming that the generation prompt remains usable across the sampling range\. Round\-trip accuracy \(centre panel\) is relatively stable forτ≤0\.4\\tau\\leq 0\.4\(35\.7%35\.7\\%→\\to30\.7%30\.7\\%,−5\-5pp\), then drops more steeply to13\.9%13\.9\\%atτ=1\.0\\tau=1\.0\. The per\-depth breakdown \(right panel\) reveals that deeper expressions are disproportionately affected: depth\-5 accuracy falls from34\.6%34\.6\\%to12\.7%12\.7\\%\(−21\.9\-21\.9pp\), whereas depth\-2 accuracy stays in a narrow range of5353–63%63\\%across the entire sweep\. Across all temperature points the inter\-seed standard deviation stays below1\.51\.5pp for the overall round\-trip metric, indicating low run\-to\-run variance\.

These results support the use of greedy decoding in the main experiments:τ=0\\tau=0maximises round\-trip accuracy while eliminating seed variance, and the relative model rankings are unlikely to change at moderate temperatures given the flat region belowτ≈0\.4\\tau\\approx 0\.4\.

Figure 5:Phi\-4 temperature sweep \(5 seeds per point\)\. Left: guard pass rate\. Centre: overall round\-trip accuracy\. Right: round\-trip accuracy by expression depth\. Shaded bands show±\\pm1 standard deviation across seeds\.
## Appendix FPer\-Depth Communication Matrices

[Table1](https://arxiv.org/html/2609.21509#S2.T1)in the main text aggregates across all depths\.[Tables4](https://arxiv.org/html/2609.21509#A6.T4),[5](https://arxiv.org/html/2609.21509#A6.T5),[6](https://arxiv.org/html/2609.21509#A6.T6),[7](https://arxiv.org/html/2609.21509#A6.T7)and[8](https://arxiv.org/html/2609.21509#A6.T8)show the communication matrix separately for each depth level, revealing how pairwise accuracy degrades with increasing nesting\.

Table 4:Communication matrix at depthd=2d=2\(15 expressions\)\. Each cell shows round\-trip accuracy \(%\)\. Rows are extractors, columns are generators\. Dotted lines separate open\-weight AR, diffusion, and frontier models\.GeneratorsExtractorsQwen3\-0\.6B Qwen3\-1\.7B Qwen3\-4B Qwen3\-8B Qwen3\-14B Qwen3\-32B Gemma\-3\-4B Gemma\-3\-12B Gemma\-3\-27B Phi\-4 Dream\-7B LLaDA\-8B C\-Haiku\-4\.5 GPT\-5 G\-3\-Flash G\-3\.1\-Pro Extr\. quality Leak\\mathrm\{Leak\}pass %33739310010010010010010010053100100100100100Qwen3\-0\.6B0\.06\.733\.353\.360\.026\.70\.06\.76\.733\.36\.726\.726\.740\.026\.733\.324\.2±\\pm18Qwen3\-1\.7B0\.020\.020\.060\.033\.320\.06\.720\.020\.053\.30\.013\.346\.760\.046\.760\.030\.0±\\pm21Qwen3\-4B0\.026\.733\.366\.773\.353\.36\.720\.033\.346\.713\.340\.066\.766\.773\.393\.344\.6±\\pm26Qwen3\-8B6\.720\.033\.353\.380\.060\.06\.726\.726\.760\.020\.046\.780\.080\.086\.793\.348\.8±\\pm29Qwen3\-14B6\.726\.733\.360\.073\.360\.013\.320\.033\.360\.020\.040\.080\.080\.093\.3100\.050\.0±\\pm29Qwen3\-32B6\.726\.733\.360\.073\.353\.313\.320\.033\.360\.026\.746\.786\.780\.093\.3100\.050\.8±\\pm29Gemma\-3\-4B6\.720\.020\.053\.360\.046\.76\.733\.326\.740\.020\.020\.053\.366\.773\.393\.340\.0±\\pm24Gemma\-3\-12B0\.033\.333\.360\.073\.360\.013\.326\.726\.753\.313\.346\.780\.066\.7100\.093\.348\.7±\\pm29Gemma\-3\-27B6\.726\.733\.360\.073\.353\.36\.726\.733\.360\.013\.346\.780\.093\.386\.7100\.050\.0±\\pm30Phi\-40\.026\.726\.746\.773\.360\.06\.720\.033\.360\.06\.746\.780\.080\.0100\.0100\.047\.9±\\pm32Dream\-7B0\.020\.013\.366\.753\.333\.36\.720\.026\.760\.00\.026\.753\.380\.066\.766\.737\.1±\\pm26LLaDA\-8B0\.020\.026\.760\.086\.760\.00\.033\.340\.060\.026\.753\.373\.393\.386\.793\.350\.8±\\pm30C\-Haiku\-4\.50\.020\.046\.753\.380\.053\.313\.326\.733\.360\.026\.746\.786\.7100\.086\.7100\.052\.1±\\pm30GPT\-56\.726\.726\.753\.380\.060\.06\.726\.746\.760\.013\.353\.373\.3100\.093\.3100\.051\.7±\\pm31G\-3\-Flash6\.720\.040\.053\.380\.060\.06\.726\.740\.060\.026\.746\.786\.7100\.086\.7100\.052\.5±\\pm30G\-3\.1\-Pro0\.013\.340\.060\.073\.366\.76\.726\.733\.360\.026\.746\.773\.3100\.0100\.0100\.051\.7±\\pm32Gen\. quality2\.9±\\pm322\.1±\\pm630\.8±\\pm857\.5±\\pm570\.4±\\pm1351\.7±\\pm137\.5±\\pm423\.8±\\pm630\.8±\\pm955\.4±\\pm816\.2±\\pm940\.4±\\pm1270\.4±\\pm1680\.4±\\pm1781\.2±\\pm2089\.2±\\pm19

Table 5:Communication matrix at depthd=3d=3\(105 expressions\)\. Each cell shows round\-trip accuracy \(%\)\. Rows are extractors, columns are generators\. Dotted lines separate open\-weight AR, diffusion, and frontier models\.GeneratorsExtractorsQwen3\-0\.6B Qwen3\-1\.7B Qwen3\-4B Qwen3\-8B Qwen3\-14B Qwen3\-32B Gemma\-3\-4B Gemma\-3\-12B Gemma\-3\-27B Phi\-4 Dream\-7B LLaDA\-8B C\-Haiku\-4\.5 GPT\-5 G\-3\-Flash G\-3\.1\-Pro Extr\. quality Leak\\mathrm\{Leak\}pass %1270999910099100999795659899100100100Qwen3\-0\.6B1\.06\.713\.319\.013\.313\.30\.02\.91\.015\.24\.816\.219\.07\.612\.47\.69\.6±\\pm6Qwen3\-1\.7B1\.09\.521\.025\.731\.421\.93\.812\.47\.629\.52\.932\.446\.735\.231\.436\.221\.8±\\pm14Qwen3\-4B1\.014\.326\.730\.540\.040\.04\.815\.213\.341\.98\.640\.066\.772\.456\.269\.533\.8±\\pm23Qwen3\-8B1\.013\.328\.633\.339\.041\.95\.713\.311\.440\.010\.543\.874\.373\.361\.981\.035\.8±\\pm25Qwen3\-14B1\.09\.530\.535\.239\.043\.82\.913\.316\.242\.911\.445\.777\.187\.673\.387\.638\.6±\\pm29Qwen3\-32B1\.013\.331\.437\.139\.048\.66\.713\.313\.347\.611\.442\.977\.189\.573\.393\.339\.9±\\pm29Gemma\-3\-4B1\.010\.526\.727\.630\.528\.65\.712\.48\.639\.08\.635\.254\.347\.636\.248\.626\.3±\\pm16Gemma\-3\-12B1\.012\.427\.630\.538\.136\.26\.713\.310\.541\.08\.645\.771\.464\.856\.269\.533\.3±\\pm23Gemma\-3\-27B1\.011\.431\.436\.240\.042\.96\.713\.316\.243\.89\.542\.977\.184\.867\.684\.838\.1±\\pm27Phi\-41\.013\.331\.433\.339\.042\.95\.714\.314\.343\.810\.549\.575\.285\.766\.794\.338\.8±\\pm28Dream\-7B1\.09\.521\.021\.921\.921\.93\.810\.55\.726\.74\.831\.438\.135\.242\.949\.521\.6±\\pm14LLaDA\-8B1\.011\.427\.628\.623\.833\.31\.914\.37\.633\.311\.439\.061\.069\.560\.052\.429\.8±\\pm21C\-Haiku\-4\.51\.014\.332\.437\.140\.047\.67\.612\.416\.243\.811\.445\.776\.295\.275\.291\.440\.5±\\pm29GPT\-51\.014\.328\.637\.139\.041\.05\.716\.210\.542\.911\.446\.774\.395\.273\.394\.339\.5±\\pm30G\-3\-Flash1\.014\.329\.539\.039\.042\.97\.613\.310\.545\.710\.548\.670\.595\.270\.597\.139\.7±\\pm30G\-3\.1\-Pro1\.014\.328\.635\.235\.233\.36\.717\.112\.447\.612\.450\.574\.3100\.077\.198\.140\.2±\\pm31Gen\. quality1\.0±\\pm012\.0±\\pm227\.3±\\pm531\.7±\\pm634\.3±\\pm836\.2±\\pm105\.1±\\pm213\.0±\\pm311\.0±\\pm439\.0±\\pm89\.3±\\pm341\.0±\\pm864\.6±\\pm1671\.2±\\pm2658\.4±\\pm1872\.2±\\pm26

Table 6:Communication matrix at depthd=4d=4\(840 expressions\)\. Each cell shows round\-trip accuracy \(%\)\. Rows are extractors, columns are generators\. Dotted lines separate open\-weight AR, diffusion, and frontier models\.GeneratorsExtractorsQwen3\-0\.6B Qwen3\-1\.7B Qwen3\-4B Qwen3\-8B Qwen3\-14B Qwen3\-32B Gemma\-3\-4B Gemma\-3\-12B Gemma\-3\-27B Phi\-4 Dream\-7B LLaDA\-8B C\-Haiku\-4\.5 GPT\-5 G\-3\-Flash G\-3\.1\-Pro Extr\. quality Leak\\mathrm\{Leak\}pass %195794989998989695966495100100100100Qwen3\-0\.6B0\.01\.54\.65\.54\.22\.10\.10\.50\.45\.70\.85\.68\.82\.42\.52\.02\.9±\\pm2Qwen3\-1\.7B0\.01\.59\.813\.812\.68\.60\.42\.01\.016\.71\.411\.227\.613\.514\.916\.19\.4±\\pm8Qwen3\-4B0\.02\.115\.626\.023\.620\.50\.63\.52\.729\.02\.518\.256\.443\.039\.055\.821\.2±\\pm19Qwen3\-8B0\.02\.315\.726\.024\.222\.30\.73\.53\.832\.32\.718\.355\.549\.845\.464\.922\.9±\\pm21Qwen3\-14B0\.02\.918\.529\.628\.625\.40\.63\.84\.235\.73\.022\.563\.665\.054\.678\.027\.2±\\pm25Qwen3\-32B0\.03\.018\.329\.230\.026\.10\.74\.64\.034\.33\.622\.961\.371\.955\.880\.727\.9±\\pm26Gemma\-3\-4B0\.01\.912\.717\.316\.112\.40\.63\.31\.721\.81\.514\.434\.919\.222\.023\.312\.7±\\pm10Gemma\-3\-12B0\.02\.315\.525\.023\.320\.40\.54\.02\.628\.23\.018\.251\.446\.138\.255\.120\.9±\\pm18Gemma\-3\-27B0\.02\.418\.227\.527\.925\.21\.04\.03\.833\.93\.320\.259\.863\.548\.177\.526\.0±\\pm24Phi\-40\.02\.319\.228\.729\.526\.30\.84\.34\.835\.43\.222\.062\.571\.554\.683\.828\.1±\\pm26Dream\-7B0\.01\.38\.914\.513\.511\.40\.22\.02\.015\.51\.29\.925\.622\.323\.041\.012\.0±\\pm11LLaDA\-8B0\.01\.913\.121\.818\.115\.50\.53\.32\.524\.33\.016\.941\.544\.236\.546\.818\.1±\\pm16C\-Haiku\-4\.50\.02\.618\.128\.729\.527\.40\.74\.04\.236\.03\.622\.163\.181\.960\.691\.929\.7±\\pm29GPT\-50\.02\.618\.630\.430\.724\.40\.54\.54\.236\.73\.224\.565\.491\.366\.895\.731\.2±\\pm31G\-3\-Flash0\.02\.918\.630\.131\.326\.50\.74\.33\.936\.53\.522\.665\.690\.763\.196\.331\.0±\\pm31G\-3\.1\-Pro0\.02\.617\.930\.430\.826\.40\.64\.24\.237\.53\.523\.065\.293\.765\.896\.131\.4±\\pm31Gen\. quality0\.0±\\pm02\.3±\\pm015\.2±\\pm424\.0±\\pm723\.4±\\pm820\.1±\\pm70\.6±\\pm03\.5±\\pm13\.1±\\pm128\.7±\\pm92\.7±\\pm118\.3±\\pm550\.5±\\pm1754\.4±\\pm2843\.2±\\pm1962\.8±\\pm29

Table 7:Communication matrix at depthd=5d=5\(830 expressions\)\. Each cell shows round\-trip accuracy \(%\)\. Rows are extractors, columns are generators\. Dotted lines separate open\-weight AR, diffusion, and frontier models\.GeneratorsExtractorsQwen3\-0\.6B Qwen3\-1\.7B Qwen3\-4B Qwen3\-8B Qwen3\-14B Qwen3\-32B Gemma\-3\-4B Gemma\-3\-12B Gemma\-3\-27B Phi\-4 Dream\-7B LLaDA\-8B C\-Haiku\-4\.5 GPT\-5 G\-3\-Flash G\-3\.1\-Pro Extr\. quality Leak\\mathrm\{Leak\}pass %215993979999989697976495100100100100Qwen3\-0\.6B0\.00\.51\.72\.92\.31\.80\.10\.10\.24\.30\.03\.16\.61\.00\.60\.71\.6±\\pm2Qwen3\-1\.7B0\.01\.35\.710\.07\.77\.60\.21\.61\.115\.11\.08\.125\.28\.99\.212\.47\.2±\\pm6Qwen3\-4B0\.01\.712\.018\.221\.318\.10\.53\.53\.327\.31\.715\.147\.738\.233\.745\.818\.0±\\pm16Qwen3\-8B0\.01\.613\.717\.821\.920\.40\.52\.94\.028\.41\.116\.350\.043\.139\.351\.419\.5±\\pm18Qwen3\-14B0\.01\.714\.821\.624\.223\.30\.64\.74\.332\.21\.818\.258\.855\.351\.863\.923\.6±\\pm22Qwen3\-32B0\.01\.714\.921\.725\.124\.60\.53\.54\.831\.42\.218\.858\.762\.251\.270\.824\.5±\\pm23Gemma\-3\-4B0\.01\.18\.412\.512\.010\.40\.22\.71\.716\.71\.010\.828\.417\.116\.917\.29\.8±\\pm8Gemma\-3\-12B0\.01\.311\.715\.417\.617\.20\.42\.53\.425\.41\.714\.643\.736\.932\.441\.816\.6±\\pm15Gemma\-3\-27B0\.01\.814\.619\.523\.422\.40\.74\.14\.830\.41\.816\.754\.256\.148\.963\.722\.7±\\pm21Phi\-40\.01\.716\.022\.726\.524\.10\.74\.74\.134\.12\.219\.459\.562\.855\.272\.525\.4±\\pm24Dream\-7B0\.01\.26\.011\.88\.710\.60\.42\.21\.812\.71\.28\.624\.217\.118\.131\.89\.8±\\pm9LLaDA\-8B0\.01\.411\.916\.318\.217\.50\.52\.82\.822\.81\.614\.940\.635\.932\.937\.316\.1±\\pm14C\-Haiku\-4\.50\.02\.015\.821\.927\.025\.30\.53\.64\.834\.02\.519\.261\.978\.757\.186\.727\.6±\\pm28GPT\-50\.01\.917\.823\.428\.719\.50\.74\.84\.536\.12\.722\.266\.491\.267\.594\.930\.1±\\pm31G\-3\-Flash0\.02\.316\.123\.026\.925\.70\.73\.74\.835\.52\.420\.864\.789\.262\.395\.329\.6±\\pm31G\-3\.1\-Pro0\.01\.815\.822\.027\.626\.30\.43\.94\.536\.62\.420\.763\.593\.364\.696\.330\.0±\\pm31Gen\. quality0\.0±\\pm01\.6±\\pm012\.3±\\pm417\.5±\\pm619\.9±\\pm818\.4±\\pm70\.5±\\pm03\.2±\\pm13\.4±\\pm126\.4±\\pm91\.7±\\pm115\.5±\\pm547\.1±\\pm1749\.2±\\pm2940\.1±\\pm2055\.2±\\pm30

Table 8:Communication matrix at depthd=6d=6\(660 expressions\)\. Each cell shows round\-trip accuracy \(%\)\. Rows are extractors, columns are generators\. Dotted lines separate open\-weight AR, diffusion, and frontier models\.GeneratorsExtractorsQwen3\-0\.6B Qwen3\-1\.7B Qwen3\-4B Qwen3\-8B Qwen3\-14B Qwen3\-32B Gemma\-3\-4B Gemma\-3\-12B Gemma\-3\-27B Phi\-4 Dream\-7B LLaDA\-8B C\-Haiku\-4\.5 GPT\-5 G\-3\-Flash G\-3\.1\-Pro Extr\. quality Leak\\mathrm\{Leak\}pass %216293971009999959597639599100100100Qwen3\-0\.6B0\.00\.61\.12\.01\.10\.30\.20\.00\.01\.70\.00\.83\.50\.60\.60\.60\.8±\\pm1Qwen3\-1\.7B0\.01\.25\.26\.15\.34\.80\.30\.50\.67\.40\.34\.417\.16\.87\.47\.44\.7±\\pm4Qwen3\-4B0\.02\.610\.613\.014\.114\.20\.31\.11\.218\.60\.89\.438\.328\.925\.628\.913\.0±\\pm12Qwen3\-8B0\.02\.410\.011\.514\.114\.70\.30\.91\.817\.90\.69\.434\.833\.226\.232\.313\.1±\\pm12Qwen3\-14B0\.02\.912\.014\.817\.917\.70\.61\.72\.024\.50\.811\.242\.943\.839\.245\.917\.4±\\pm16Qwen3\-32B0\.03\.012\.414\.817\.919\.10\.51\.72\.424\.11\.210\.946\.452\.342\.651\.718\.8±\\pm19Gemma\-3\-4B0\.01\.86\.26\.77\.65\.80\.20\.60\.310\.30\.25\.519\.59\.59\.212\.16\.0±\\pm5Gemma\-3\-12B0\.01\.59\.810\.812\.612\.10\.30\.81\.815\.50\.97\.734\.724\.120\.828\.611\.4±\\pm11Gemma\-3\-27B0\.02\.611\.713\.216\.517\.40\.61\.21\.820\.90\.810\.844\.142\.136\.446\.116\.6±\\pm16Phi\-40\.03\.013\.315\.020\.218\.90\.51\.52\.425\.61\.112\.648\.255\.942\.355\.219\.7±\\pm19Dream\-7B0\.01\.45\.35\.86\.56\.40\.20\.80\.57\.60\.34\.714\.110\.511\.814\.55\.6±\\pm5LLaDA\-8B0\.02\.09\.29\.513\.610\.60\.31\.11\.713\.60\.98\.329\.225\.320\.329\.410\.9±\\pm10C\-Haiku\-4\.50\.02\.712\.115\.319\.520\.00\.31\.72\.325\.61\.110\.048\.972\.946\.163\.521\.4±\\pm23GPT\-50\.03\.814\.117\.121\.125\.50\.82\.12\.728\.61\.213\.660\.090\.855\.870\.025\.4±\\pm28G\-3\-Flash0\.03\.212\.715\.819\.420\.60\.51\.52\.627\.11\.212\.154\.186\.553\.269\.723\.8±\\pm26G\-3\.1\-Pro0\.02\.313\.816\.219\.221\.40\.61\.72\.728\.91\.212\.957\.090\.352\.068\.024\.3±\\pm27Gen\. quality0\.0±\\pm02\.3±\\pm110\.0±\\pm411\.7±\\pm414\.2±\\pm614\.3±\\pm70\.4±\\pm01\.2±\\pm11\.7±\\pm118\.6±\\pm80\.8±\\pm09\.0±\\pm337\.1±\\pm1642\.1±\\pm2930\.6±\\pm1739\.0±\\pm22

## Appendix GSHAP Attribution Methodology and Additional Experiments

To quantify which features drive round\-trip success, we fit a gradient\-boosted decision tree \(GBT\) classifier to the binary outcome \(correct/incorrect\) of each*triple*\(expression, generator, extractor\)\.

Model\.We use scikit\-learn’s\[[Pedregosa et al\., 2011](https://arxiv.org/html/2609.21509#bib.bib2)\]GradientBoostingClassifierwith 200 trees, maximum depth 4, learning rate 0\.1, and row subsampling rate 0\.8\. The classifier is trained on the full set of triples\. A held\-out 80/20 split yields a sanity\-check accuracy of 88\.1%, confirming the features carry substantial predictive signal\. For context, the majority\-class baseline \(always predicting failure\) achieves 87\.8%, so the GBT’s improvement is modest in absolute terms, and the value of the classifier lies in the SHAP decomposition rather than raw accuracy\. The corresponding AUC\-ROC scores are0\.8120\.812for the high\-level classifier and0\.8220\.822for the AST classifier, consistent with a model that captures substantial but not saturating signal on the round\-trip outcome\.

Attribution\.We compute exact SHAP values viaTreeExplainer\[[Lundberg et al\., 2019](https://arxiv.org/html/2609.21509#bib.bib1)\], which exploits the tree structure to compute Shapley values in polynomial time without sampling\. For each featureff, we report*mean absolute SHAP*normalised to a percentage of the total attribution:

Attribution⁡\(f\)=1n​∑t=1n\|ϕf\(t\)\|∑f′1n​∑t=1n\|ϕf′\(t\)\|×100%,\\mathrm\{Attribution\}\(f\)\\;=\\;\\frac\{\\frac\{1\}\{n\}\\sum\_\{t=1\}^\{n\}\|\\phi\_\{f\}^\{\(t\)\}\|\}\{\\sum\_\{f^\{\\prime\}\}\\frac\{1\}\{n\}\\sum\_\{t=1\}^\{n\}\|\\phi\_\{f^\{\\prime\}\}^\{\(t\)\}\|\}\\;\\times\\;100\\%,\(4\)whereϕf\(t\)\\phi\_\{f\}^\{\(t\)\}is the SHAP value of featurefffor triplettand the sum in the denominator runs over all features\.

Direction\.We additionally report the Pearson correlation between each feature’s raw values and its SHAP values across triples\. A positive correlation means higher feature values push the prediction toward success; a negative correlation means higher values push toward failure\.

### G\.1Complete Structural SHAP Attribution

The main text \([Section3\.3](https://arxiv.org/html/2609.21509#S3.SS3)\) reports the high\-level SHAP analysis with six summary features and summarizes the key findings of the structural analysis\.[Table9](https://arxiv.org/html/2609.21509#A7.T9)lists all 16 features from the AST\-level gradient\-boosted classifier, sorted by mean absolute SHAP value\.

After controlling for model size \(50\.8%50\.8\\%combined\), four structural findings stand out:

- •Division is the hardest operator\.Division \(n÷n\_\{\\div\}: 6\.4%\) is the single largest structural predictor, followed by subtraction \(n−n\_\{\-\}: 4\.1%\) and multiplication \(n×n\_\{\\times\}: 3\.7%\); addition \(n\+n\_\{\+\}: 1\.4%\) matters least\. The two non\-commutative operators thus rank first and second among the four\. Non\-commutativity, not multiplicativity, drives operator difficulty: argument order \(a÷b≠b÷aa\\div b\\neq b\\div a,a−b≠b−aa\-b\\neq b\-a\) must be conveyed unambiguously in prose, whereas the commutative operators tolerate reordering\. Consistent with this, the deepest tree level containing a non\-commutative operator \(5\.5%,↓\\downarrow\) is the largest structural predictor after depth itself, confirming that ordering ambiguity compounds with nesting\.
- •Tree shape matters beyond size and depth\.Skeleton identity \(5\.5%\) is a strong predictor alongside depth \(7\.4%\), indicating that expressions with different branching topologies but identical\(k,d\)\(k,d\)can differ substantially in round\-trip accuracy\. Center of mass \(3\.0%,↓\\downarrow\) reveals the direction of this effect: right\-branching trees, which require the word problem to hold an incomplete computation open while describing a nested subexpression, tend to be harder than left\-branching ones\.
- •Variable reuse introduces ambiguity\.Expressions where a variable appears in multiple leaves \(3\.5%,↓\\downarrow\) require coreference in the word problem, forcing the generator to signal that two quantities refer to the same variable\. More unique variables \(3\.0%,↑\\uparrow\) have the opposite effect, reducing coreference demands and improving round\-trip fidelity\.
- •Root operator sets the frame\.Addition as root \(2\.2%,↓\\downarrow\) and subtraction as root \(2\.0%,↓\\downarrow\) are the most penalising root choices, whereas multiplicative roots \(÷\\div: 0\.9%,×\\times: 0\.7%, both↑\\uparrow\) provide more flexibility at the top of the tree\.

Table 9:Complete SHAP feature attribution for the structural analysis\. All features from the AST\-level gradient\-boosted classifier are shown, sorted by mean\|SHAP\|\|\\text\{SHAP\}\|as a percentage of total attribution\. Green \(↑\\uparrow\) indicates higher feature values increase round\-trip success; red \(↓\\downarrow\) indicates higher values decrease success\.FeatureAttributionGenerator size34\.3%↑\\uparrowExtractor size16\.5%↑\\uparrowDepth7\.4%↓\\downarrown÷n\_\{\\div\}6\.4%↓\\downarrowNon\-comm\. depth5\.5%↓\\downarrowSkeleton ID5\.5%↓\\downarrown−n\_\{\-\}4\.1%↓\\downarrown×n\_\{\\times\}3\.7%↓\\downarrowVariable reuse3\.5%↓\\downarrowCenter of mass3\.0%↓\\downarrowUnique variables3\.0%↑\\uparrowRoot==\+2\.2%↓\\downarrowRoot==−\-2\.0%↓\\downarrown\+n\_\{\+\}1\.4%↑\\uparrowRoot==÷\\div0\.9%↑\\uparrowRoot==×\\times0\.7%↑\\uparrow

## Appendix HSkeleton Difficulty and Center of Mass

The SHAP analysis identifies skeleton identity \(5\.5%\) and center of mass \(3\.0%\) as significant predictors of difficulty\. We now examine these effects directly by comparing skeletons that share the same operator count and depth\.

[Figure6](https://arxiv.org/html/2609.21509#A8.F6)plots the round\-trip accuracy of each skeleton, grouped by\(k,d\)\(k,d\)\. The spread within groups is substantial: at\(k=4,d=4\)\(k\{=\}4,d\{=\}4\), the seven observed skeletons span 22\.1 percentage points, from 52\.7% for the easiest to 30\.6% for the hardest\. At\(k=2,d=2\)\(k\{=\}2,d\{=\}2\), the only two skeletons differ by 1\.1 percentage points \(left\-branching: 60\.8% vs\. right\-branching: 59\.7%\)\. These differences are controlled for operator count and depth by construction; they reflect the effect of tree shape alone\.

Center of mass as predictor\.Points in[Figure6](https://arxiv.org/html/2609.21509#A8.F6)are coloured by their normalised center of mass \(CoM\), ranging from blue \(left\-branching, CoM=0=0\) to red \(right\-branching, CoM=1=1\)\. A visual pattern emerges: within each group, bluer points tend to appear to the right \(higher accuracy\) and redder points to the left \(lower accuracy\)\. The within\-group weighted correlation between CoM and accuracy isr=−0\.148r=\-0\.148\(p=0\.007p=0\.007, one\-tailed permutation test with 10000 shuffles,N=455N=455skeletons\), indicating that right\-branching trees tend to be harder after controlling for\(k,d\)\(k,d\)\.

Why right\-branching is harder\.Left\-branching trees \(e\.g\.\(\(A\+B\)×C\)−D\(\(A\+B\)\\times C\)\-D\) correspond to a sequence of operations applied to a running result, which maps naturally onto left\-to\-right prose: “start withAAplusBB, multiply byCC, subtractDD\.” Right\-branching trees \(e\.g\.A−\(B×\(C\+D\)\)A\-\(B\\times\(C\+D\)\)\) require the generator to introduce a nested subcomputation before the main operation, demanding forward references or subordinate clauses that are harder to produce and parse unambiguously\. An analogous right\-branching deficit has been reported for LSTM/GRU learners on synthetic interpreted languages\[[Paperno, 2022](https://arxiv.org/html/2609.21509#bib.bib73)\], suggesting the bias is not specific to transformer LLMs\.

![Refer to caption](https://arxiv.org/html/2609.21509v1/skeleton_difficulty.png)Figure 6:Skeleton difficulty within\(k,d\)\(k,d\)groups\. Each point is one skeleton; thexx\-axis shows its round\-trip accuracy averaged across all models and expressions sharing that skeleton\. Colour encodes center of mass \(blue = left\-branching, red = right\-branching\)\. The spread in parentheses on each row label shows the gap between the easiest and hardest skeleton in that group\. Within\-group weightedr=−0\.148r=\-0\.148\(CoM vs\. accuracy,p=0\.007p=0\.007, permutation test\)\.
## Appendix IError Taxonomy Details

[Table10](https://arxiv.org/html/2609.21509#A9.T10)shows the per\-extractor error profile across all generators and depths\. The classification algorithm processes each incorrect extraction in the following order: \(1\) attempt to parsee^\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}\\hat\{e\}\}into a symbolic expression tree; if parsing fails, labelErrparse\\mathrm\{Err\}\_\{\\mathrm\{parse\}\}; \(2\) count the operators in bothe\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}ande^\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}\\hat\{e\}\}; if they differ, labelErrarity\\mathrm\{Err\}\_\{\\mathrm\{arity\}\}; \(3\) compare the unlabelled tree skeletons; if they differ, labelErrstruct\\mathrm\{Err\}\_\{\\mathrm\{struct\}\}; \(4\) otherwise labelErrcontent\\mathrm\{Err\}\_\{\\mathrm\{content\}\}\(the skeleton matches but leaf values or operator labels differ\)\.

Table 10:Per\-extractor error profile\. Each row shows the percentage of that extractor’s errors falling into each category, aggregated across all generators and depths\. Models are grouped by family and sorted by size\.ExtractorErrparse\\mathrm\{Err\}\_\{\\mathrm\{parse\}\}Errarity\\mathrm\{Err\}\_\{\\mathrm\{arity\}\}Errstruct\\mathrm\{Err\}\_\{\\mathrm\{struct\}\}Errcontent\\mathrm\{Err\}\_\{\\mathrm\{content\}\}Qwen3\-0\.6B0\.1%89\.1%10\.3%0\.5%Qwen3\-1\.7B0\.4%74\.4%23\.4%1\.8%Qwen3\-4B1\.2%67\.8%29\.3%1\.7%Qwen3\-8B1\.6%66\.2%30\.7%1\.6%Qwen3\-14B1\.8%65\.7%31\.3%1\.2%Qwen3\-32B1\.6%64\.7%32\.4%1\.2%Gemma\-3\-4B5\.4%67\.9%25\.5%1\.1%Gemma\-3\-12B0\.3%70\.6%27\.9%1\.3%Gemma\-3\-27B0\.6%68\.0%29\.8%1\.6%Phi\-42\.3%67\.0%28\.8%1\.9%Dream\-7B10\.5%69\.6%18\.6%1\.3%LLaDA\-8B2\.6%77\.2%17\.6%2\.5%C\-Haiku\-4\.51\.9%67\.0%29\.7%1\.5%GPT\-52\.1%70\.3%26\.2%1\.4%Gemini\-3\-Flash0\.6%73\.3%24\.9%1\.2%Gemini\-3\.1\-Pro1\.4%71\.0%26\.3%1\.2%[Figure7](https://arxiv.org/html/2609.21509#A9.F7)breaks the same categories down by expression depth, showing that arity errors grow with depth while content errors shrink\.

Figure 7:Distribution of error categories by expression depth\.Errarity\\mathrm\{Err\}\_\{\\mathrm\{arity\}\}\(lost or extra operators\) dominates at every depth and grows from 46% at depth 2 to 74% at depth 6\.Errcontent\\mathrm\{Err\}\_\{\\mathrm\{content\}\}shrinks correspondingly, indicating that models that preserve the correct operator count and tree shape rarely misidentify leaves\.Errparse\\mathrm\{Err\}\_\{\\mathrm\{parse\}\}errors are small \(2\.0% overall\)\.
## Appendix JCorrelation with External Benchmarks

A compact summary appears in[Section3\.4](https://arxiv.org/html/2609.21509#S3.SS4); this appendix provides the full per\-benchmark breakdown\.

We compute Spearman rank correlations between per\-model round\-trip scores and published results on 9 external benchmarks spanning four categories: math \(GSM8K\[[Cobbe et al\., 2021](https://arxiv.org/html/2609.21509#bib.bib6)\], MATH\[[Hendrycks et al\., 2021b](https://arxiv.org/html/2609.21509#bib.bib7)\], MGSM\[[Shi et al\., 2023](https://arxiv.org/html/2609.21509#bib.bib58)\]\), code \(MBPP\[[Austin et al\., 2021](https://arxiv.org/html/2609.21509#bib.bib57)\], LiveCodeBench\[[Jain et al\., 2024](https://arxiv.org/html/2609.21509#bib.bib29)\]\), general knowledge \(MMLU\[[Hendrycks et al\., 2021a](https://arxiv.org/html/2609.21509#bib.bib54)\], MMLU\-Pro\[[Wang et al\., 2024](https://arxiv.org/html/2609.21509#bib.bib55)\], BBH\[[Suzgun et al\., 2023](https://arxiv.org/html/2609.21509#bib.bib56)\]\), and reasoning \(GPQA\[[Rein et al\., 2023](https://arxiv.org/html/2609.21509#bib.bib20)\]\)\. Benchmark scores are drawn exclusively from official technical reports and vendor announcements\[[Yang and others, 2025](https://arxiv.org/html/2609.21509#bib.bib3),[Gemma Team, 2025](https://arxiv.org/html/2609.21509#bib.bib4),[Abdin et al\., 2024](https://arxiv.org/html/2609.21509#bib.bib5),[Anthropic, 2025a](https://arxiv.org/html/2609.21509#bib.bib16),[Anthropic, 2025b](https://arxiv.org/html/2609.21509#bib.bib52),[Singh et al\., 2025](https://arxiv.org/html/2609.21509#bib.bib26)\]; benchmarks with verified scores for fewer than seven models are excluded\. For each benchmark we correlate its model rankings with three round\-trip scores \([Section2\.3](https://arxiv.org/html/2609.21509#S2.SS3)\):*generation quality*\(column\-average of the communication matrix\),*extraction quality*\(row\-average\), and*self\-communication*\(diagonal cell\)\. Because frontier vendors did not publish standard\-mode scores on most benchmarks, the analysis is dominated by the 10 open\-weight models; each benchmark uses pairwise deletion, so the effectivennvaries \(7–10; see[Table13](https://arxiv.org/html/2609.21509#A10.T13)\)\.[Table11](https://arxiv.org/html/2609.21509#A10.T11)lists the source document, evaluation protocol, and coverage for each model family\. Because evaluation protocols differ across families \(e\.g\. base\-model few\-shot for Qwen3 vs\. 0\-shot instruct for Gemma 3\), the correlations should be read as rough indicators rather than precise estimates\.

Table 11:Provenance of external benchmark scores\. Each score traces to the listed source;nullentries are omitted from the correlation analysis via pairwise deletion\.Model familySourceEval modeBenchmarks \(of 9\)Qwen3 \(0\.6–32B\)[Yang and others \[2025\]](https://arxiv.org/html/2609.21509#bib.bib3)base few\-shot†9/9Gemma 3 IT \(4–27B\)[Gemma Team \[2025\]](https://arxiv.org/html/2609.21509#bib.bib4)0\-shot instruct7/9Phi\-4 \(14B\)[Abdin et al\. \[2024\]](https://arxiv.org/html/2609.21509#bib.bib5)simple\-evals5/9Claude Haiku 4\.5[Anthropic \[2025b\]](https://arxiv.org/html/2609.21509#bib.bib52)no standard\-mode scores publishedGPT\-5[Singh et al\. \[2025\]](https://arxiv.org/html/2609.21509#bib.bib26)non\-reasoning1/9
†MATH\-500, LiveCodeBench, and GPQA\-Diamond use instruct non\-thinking scores where base\-model scores are unavailable\.

[Table13](https://arxiv.org/html/2609.21509#A10.T13)reports the per\-benchmark correlations sorted by averageρ\\rhoacross the three round\-trip phases\. Several patterns emerge:

Extraction quality is consistently well\-predicted\.For 7 of 9 benchmarks, extractionρ\\rhoreaches0\.830\.83or higher, peaking at MMLU \(ρ=0\.98\\rho=0\.98\) and MGSM \(ρ=0\.96\\rho=0\.96\)\. The models that extract expressions most accurately are, broadly, those that score highest on standard evaluations\.

Generation quality is harder to predict\.Generationρ\\rhoranges from0\.470\.47\(MMLU\) to0\.890\.89\(MGSM\) and fails to reach significance for MMLU, MMLU\-Pro, MATH, and GPQA\-Diamond individually\. The gap between extraction and generationρ\\rhois largest for MMLU \(0\.980\.98vs\.0\.470\.47\) and GPQA\-Diamond \(0\.950\.95vs\.0\.500\.50\), confirming that faithfully*generating*a word problem from an expression is a skill not fully predicted by standard evaluations\.

Code benchmarks correlate strongly\.MBPP and LiveCodeBench both reachρ=0\.93\\rho=0\.93for extraction and0\.770\.77for generation, placing code among the best\-predicted categories rather than the weakest\.

Math benchmarks are the weakest category\.MATH shows uniformly moderate correlations \(ρ=0\.52\\rho=0\.52/0\.500\.50/0\.300\.30for generation/extraction/self\-communication\), and GSM8K follows a similar pattern \(0\.680\.68/0\.720\.72/0\.500\.50\)\. MGSM is an outlier \(0\.890\.89/0\.960\.96/0\.890\.89,n=7n\{=\}7\)\. Standard math\-solving ability overlaps only partially with the structural linearization our protocol measures\.

[Table12](https://arxiv.org/html/2609.21509#A10.T12)aggregates these findings at the category level via Fisher\-zzaveraging of per\-benchmarkρ\\rhovalues\. Extraction correlates strongly across all four categories \(ρ=0\.82\\rho=0\.82–0\.950\.95\)\. Generation is highest for code \(ρ=0\.77\\rho=0\.77\) and math \(ρ=0\.74\\rho=0\.74\) and lowest for reasoning and general knowledge \(ρ=0\.50\\rho=0\.50–0\.550\.55\)\.

Table 12:Fisher\-zzaveraged Spearmanρ\\rhobetween round\-trip scores and external benchmarks, grouped by category:Math,Code,General,Reasoning\.CodeReasoningGeneralMathbenchmarks2133Generation Acc\.0\.770\.500\.550\.74Extraction Acc\.0\.930\.950\.940\.82Self\-Consistency0\.770\.900\.860\.64Table 13:Spearman rank correlation \(ρ\\rho\) between round\-trip scores and external benchmarks, sorted by averageρ\\rho\. Columns are colored by category:Math,Code,General,Reasoning\.MGSM

LiveCodeBench

MBPP

MMLU

GPQA\-Diamond

MMLU\-Pro

BBH

GSM8K

MATH

nn7991010109910Generation Acc\.0\.890\.770\.770\.470\.500\.500\.670\.680\.52Extraction Acc\.0\.960\.930\.930\.980\.950\.930\.830\.720\.50Self\-Consistency0\.890\.770\.770\.970\.900\.790\.630\.500\.30
## Appendix KRound\-Trip Examples

[Table14](https://arxiv.org/html/2609.21509#A11.T14)presents example round\-trips generated by Qwen3\-32B and frontier models \(as both𝖦\\mathsf\{G\}and𝖤\\mathsf\{E\}\) at increasing complexity\. Each row shows the skeleton \(tree shape\), the source expression, the generated word problem, and the expression recovered by the extractor\. All five examples were successfully verified by SymPy, despite the recovered expressions differing syntactically from the originals \(e\.g\. reordering commutative operands or distributing subtraction\)\.

Table 14:Round\-trip examples at increasing operator countkkand depthdd\. Skeletons use the notation from[Figure2](https://arxiv.org/html/2609.21509#S2.F2):□\\square= leaf,∘\\circ= operator\. The “Recovered” column shows the extractor’s output, which SymPy confirms as symbolically equivalent to the original expression\. The first four rows use Qwen3\-32B as both𝖦\\mathsf\{G\}and𝖤\\mathsf\{E\}; the last two use frontier models\.\(k,d\)\(k,d\)SkeletonExpressionWord ProblemRecovered\(2,2\)\(2,2\)\(□∘□\)∘□\(\\square\\circ\\square\)\\circ\\square\(E−H\)−H\(E\-H\)\-HA farmer had \{E\} acres of land\. He used \{H\} acres for planting corn and another \{H\} acres for growing soybeans\. How many acres were left unplanted?E−\(H\+H\)E\-\(H\+H\)\(3,3\)\(3,3\)□∘\(□∘\(□∘□\)\)\\square\\circ\(\\square\\circ\(\\square\\circ\\square\)\)D−\(1−\(D\+C\)\)D\-\(1\-\(D\+C\)\)A construction site has \{D\} tons of steel\. Workers use 1 ton for a beam, then recover \{D\} tons from an old structure and add \{C\} tons from storage\. How much steel is available now?D−1\+D\+CD\-1\+D\+C\(5,4\)\(5,4\)\(\(\(□∘□\)∘□\)∘□\)∘\(□∘□\)\(\(\(\\square\\circ\\square\)\\circ\\square\)\\circ\\square\)\\circ\(\\square\\circ\\square\)\(\(\(F/3\)\+1\)−C\)−\(A\+9\)\(\(\(F/3\)\+1\)\-C\)\-\(A\+9\)A traveler drove \{F\} miles and then took a break after every 3 miles\. They added 1 rest stop but skipped \{C\} due to road closures\. How many rest stops were left after subtracting the \{A\} planned stops plus 9 extra ones?\(F/3\+1−C\)−\(A\+9\)\(F/3\{\+\}1\{\-\}C\)\{\-\}\(A\{\+\}9\)\(6,5\)\(6,5\)\(\(□∘\(\(□∘□\)∘□\)\)∘□\)∘\(□∘□\)\(\(\\square\\circ\(\(\\square\\circ\\square\)\\circ\\square\)\)\\circ\\square\)\\circ\(\\square\\circ\\square\)\(\(E−\(\(C×6\)×B\)\)−9\)\+\(G/8\)\(\(E\-\(\(C\{\\times\}6\)\{\\times\}B\)\)\-9\)\+\(G/8\)A research team counted \{E\} fish in a coral reef\. They observed that \{C\} divers each removed 6 fish per hour for \{B\} hours, and 9 fish swam away\. Later, \{G\} baby fish were evenly distributed among 8 areas\. How many fish remained in the reef?E−\(C×6×B\+9\)\+G/8E\{\-\}\(C\{\\times\}6\{\\times\}B\{\+\}9\)\{\+\}G/8Frontier models\(5,5\)\(5,5\)G×\(H\+G/\(C\+H\+C\)\)G\{\\times\}\(H\{\+\}G/\(C\{\+\}H\{\+\}C\)\)𝖦\\mathsf\{G\}/𝖤\\mathsf\{E\}: Gemini\-3\.1\-ProA dispatcher calculates the total fleet size by combining \{C\} cargo jets, \{H\} transport helicopters, and another \{C\} light aircraft\. A stockpile of \{G\} crates of aviation fuel is divided evenly among this entire fleet\. Each aircraft takes its share of crates, adds \{H\} emergency supply boxes to its payload, and multiplies this new load by \{G\} scheduled flights to find the total cargo transported\.\(H\+\(G/\(C\+H\+C\)\)\)×G\(H\{\+\}\(G/\(C\{\+\}H\{\+\}C\)\)\)\{\\times\}G\(5,5\)\(5,5\)C−F×\(F−\(7\+F\+D\)\)C\{\-\}F\{\\times\}\(F\{\-\}\(7\{\+\}F\{\+\}D\)\)𝖦\\mathsf\{G\}/𝖤\\mathsf\{E\}: GPT\-5In a basketball game, start with 7 bonus points, then add \{F\} for free throws and \{D\} for assists, and take this total away from \{F\}\. Multiply that result by \{F\}, then take that amount away from \{C\} to find the team’s final tally\.C−\(\(F−\(7\+F\+D\)\)×F\)C\{\-\}\(\(F\{\-\}\(7\{\+\}F\{\+\}D\)\)\{\\times\}F\)
## Appendix LRound\-Trip Failure Examples

[Table14](https://arxiv.org/html/2609.21509#A11.T14)shows successful round\-trips\.[Table15](https://arxiv.org/html/2609.21509#A12.T15)complements it with representative failures covering all four error categories defined in[AppendixI](https://arxiv.org/html/2609.21509#A9), including both extractor and generator faults\. All examples are drawn from strong model pairs \(≥\\geq14B\) and frontier models to illustrate that these failures are not trivial\.

Table 15:Representative round\-trip failures covering all four error categories, with a mix of extractor and generator faults\. For each row: the source expression, the word problem written by𝖦\\mathsf\{G\}, the expression recovered by𝖤\\mathsf\{E\}, and a diagnosis identifying whether the fault lies in generation or extraction, based on the consensus vote across the 16 extractors\. Underlining highlights the locus of the error in the recovered expression\.CategoryExpression & Word ProblemRecoveredDiagnosisErrparse\\mathrm\{Err\}\_\{\\mathrm\{parse\}\}
𝖦\\mathsf\{G\}: GPT\-5
𝖤\\mathsf\{E\}: GPT\-5\(\(F\+H\)×\(10\+F/E\)\)\(\(F\+H\)\\times\(10\+F/E\)\)
“You bake \{F\} trays of cookies and \{H\} trays of brownies, then count all the trays together\. For each tray, you set the glaze amount by splitting \{F\} cups evenly among \{E\} bowls and then adding 10 more cups, and you make that amount for every tray\.”\(F \+ H\)\(F/E \+ 10\)
Juxtaposition without an explicit operator, so the expression does not parse\.Extractor fault\.The expected expression is the product of “trays” and “glaze per tray,” but the extractor drops the explicit×\\timesand emits two parenthesised groups side by side, which the parser rejects\.Errarity\\mathrm\{Err\}\_\{\\mathrm\{arity\}\}
𝖦\\mathsf\{G\}: Qwen3\-14B
𝖤\\mathsf\{E\}: GPT\-5G\+3\+7−7G\+3\+7\-7
“A soccer team scored \{G\} goals in the first half, then scored 3 more goals, and later scored 7 additional goals before giving up 7 in the second half\. How many goals did they have at the end of the game?”\(G\+3\)\+7\(G\+3\)\+7*\[drops−7\-7\]*
2 operators instead of 3\.Extractor fault\.The word problem clearly narrates four events, including “giving up 7,” but the extractor stops after the third and omits the final subtraction, effectively simplifying\+7−7\+7\-7away\.Errarity\\mathrm\{Err\}\_\{\\mathrm\{arity\}\}
𝖦\\mathsf\{G\}: Qwen3\-14B
𝖤\\mathsf\{E\}: Phi\-4D\+F×9\+8D\+F\\times 9\+8
“A designer creates \{D\} dresses and \{F\} fashion sets, each set requiring 9 dresses\. After making 8 additional dresses, how many total dresses has the designer created?”\(C/F\)¯\\underline\{\(C/F\)\}
1 operator instead of 3, and with wrong variable \(CCdoes not appear in the problem\)\.Generator fault\.The narration flattens the two additions and the multiplication into a single sentence and loses the link between “sets,” “each requiring 9 dresses,” and the final “\+8\+8” term\. The extractors, including the frontier models, fail to recover the three\-operation expression, and several emit a single division using a variable \(CC\) that never appears in the problem\.Errstruct\\mathrm\{Err\}\_\{\\mathrm\{struct\}\}
𝖦\\mathsf\{G\}: GPT\-5
𝖤\\mathsf\{E\}: Qwen3\-14BE\+B−\(1−9\)E\+B\-\(1\-9\)
“During a forest survey, a ranger counts \{E\} birch saplings and then adds \{B\} pine saplings to the tally\. From this total, they subtract the result of taking 9 away from 1 to correct an earlier note\. How many saplings are recorded now?”\(E\+B\)−\(9−1\)¯\(E\+B\)\-\\underline\{\(9\-1\)\}
Same three operators \(\+\+,−\-,−\-\), with operands of the inner subtraction swapped\.Extractor fault\.The word problem states “taking 9 away from 1,” i\.e\.1−91\-9, but the extractor flips the operands to the more natural9−19\-1, preserving the outer structure while inverting the inner subtraction\.Errcontent\\mathrm\{Err\}\_\{\\mathrm\{content\}\}
𝖦\\mathsf\{G\}: Qwen3\-14B
𝖤\\mathsf\{E\}: GPT\-5B×F\+E−DB\\times F\+E\-D
“During a cycling race, \{B\} cyclists each rode \{F\} laps around the track\. They then received an extra \{E\} bonus points for completing the race, but \{D\} points were deducted for penalties\. What is the total points earned by all cyclists?”B×\(F\+E−D\)¯B\\times\\underline\{\(F\+E\-D\)\}
Same skeleton after canonicalisation, withBBdistributed over a parenthesised sum rather than multiplied only withFF\.Extractor fault\.The word problem distinguishes “laps per cyclist” \(B×FB\\times F\) from the additive bonus\+E\+Eand penalty−D\-D, but the extractor binds all three addends under the multiplication, changing the value at every operand assignment\. Most extractors recover the intended grouping despite the mild unit slip \(laps vs\. points\), indicating the English is clear enough for most readers\.Errcontent\\mathrm\{Err\}\_\{\\mathrm\{content\}\}
𝖦\\mathsf\{G\}: Qwen3\-32B
𝖤\\mathsf\{E\}: Claude Haiku 4\.5F\+8\+EF\+8\+E
“\{F\} flower bulbs were planted in the spring, and by summer there were 8 more blooming flowers\. Then, \{E\} extra flowers were added to the garden bed\.”D¯\+8\+E\\underline\{D\}\+8\+E
Same skeleton, with leafFFreplaced byDD\.Generator fault\.The word problem starts with “\{F\} flower bulbs,” but the surface form places the variable in sentence\-initial position with no anchoring cue, and extractors consistently read the leading token asDD\. The extractors exhibit the sameF→DF\\\!\\to\\\!Dconfusion: they recover the tree skeleton correctly but mis\-bind the leading variable, a universal failure that points to the surface form rather than any individual extractor\.
## Appendix MFault Attribution Details

This appendix provides implementation details and per\-model breakdowns for the fault attribution analysis of[Section3\.3](https://arxiv.org/html/2609.21509#S3.SS3), based on the extractor voteP⁡\(e′∣w\)P\(e^\{\\prime\}\\mid\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}\)\([Equation2](https://arxiv.org/html/2609.21509#S2.E2)\) and extraction agreementp∗p^\{\*\}\([Section2\.4](https://arxiv.org/html/2609.21509#S2.SS4)\)\.

### Computing extractor votes

The symbolic equivalence≡\\equivin[Equation2](https://arxiv.org/html/2609.21509#S2.E2)is evaluated via SymPy: each predicted string is parsed into an expression and simplified; two predictions are equivalent if their simplified forms match\. Predictions that fail to parse are treated as distinct\. This is applied to allNNextractor outputs per word problem to computeP⁡\(e′∣w\)P\(e^\{\\prime\}\\mid\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}\)and, for non\-decodable word problems, the extraction agreementp∗p^\{\*\}\([Section2\.4](https://arxiv.org/html/2609.21509#S2.SS4)\)\. Simplification is parallelized across expressions usingProcessPoolExecutor\.

### Aggregate results

[Table16](https://arxiv.org/html/2609.21509#A13.T16)reports the aggregate breakdown\. Of 392276 errors, 73\.6% arise from non\-decodable word problems \(P⁡\(e∣w\)=0P\(\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\\mid\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}\)=0\), matching the headline rate reported in[Section3\.3](https://arxiv.org/html/2609.21509#S3.SS3)\.

Table 16:Fault attribution summary for all incorrect round\-trips\.CategoryCount% of errorsExtractor fault \(P⁡\(e∣w\)\>0P\(\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\\mid\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}\)\>0\)10366826\.4Non\-decodable \(P⁡\(e∣w\)=0P\(\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\\mid\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}\)=0\)28860873\.6
### Per\-depth breakdown

[Table17](https://arxiv.org/html/2609.21509#A13.T17)shows the per\-depth breakdown\. At depth 2, 37\.9% of errors are extractor faults \(decodable word problems\); as depth increases, the extractor\-fault rate falls to 24\.7% at depth 6\. Deeper word problems are longer and more ambiguous, causing extractor outputs to diverge\.

Table 17:Fault attribution by expression depth\. Ext\-fault % is the share of errors from decodable word problems\.DepthErrorsExt\-fault %2162237\.931392629\.2412874926\.8513351627\.1611446324\.7
### Per\-generator breakdown

[Table18](https://arxiv.org/html/2609.21509#A13.T18)breaks down the two fault\-attribution metrics by generator model, ordered by parameter count\.

Table 18:Fault attribution by generator model\. Larger generators produce more decodable word problems \(higher ext\-fault %\)\.GeneratorErrorsExt\-fault %Qwen3\-0\.6B68890\.1Qwen3\-1\.7B196303\.8Qwen3\-4B2824014\.3Gemma\-3\-4B352351\.3Dream\-7B187535\.5LLaDA\-8B2391019\.6Qwen3\-8B2781317\.4Gemma\-3\-12B344195\.2Qwen3\-14B2867621\.0Phi\-42481729\.5Gemma\-3\-27B339396\.0Qwen3\-32B2952426\.6Haiku\-4\.51962459\.6Gemini\-3\-Flash2345668\.6Gemini\-3\.1\-Pro1783890\.2GPT\-51951396\.9Small models produce systematic errors\.Qwen3\-0\.6B and Gemma\-3\-4B have the lowest extractor\-fault rates \(<<2%\)\. Their word problems are wrong in a consistent, predictable way: nearly all extractors decode the same incorrect expression, and the word problems are almost never decodable\. In contrast, GPT\-5 shows the opposite pattern: 97% of its errors are extractor faults, meaning the word problem is faithful but extractors fail to recover the target expression\.

Large models shift errors to the extractor\.The frontier generators \(Haiku\-4\.5, Gemini\-3\-Flash, Gemini\-3\.1\-Pro, GPT\-5\) have extractor\-fault rates of 60–97%: their word problems are far more often decodable, so when failures occur they are predominantly attributable to extractor limitations\.

### Takeaway

Fault attribution reinforces a central finding of this paper: generation is harder than extraction\. Over 73% of errors arise from non\-decodable word problems, with the rate reaching essentially 100% for the smallest generator and dropping to 3% for GPT\-5\. This asymmetry is consistent with the generation\-extraction gap observed in[Section3\.2](https://arxiv.org/html/2609.21509#S3.SS2)and the IRT generation\-clarity coefficients in[AppendixN](https://arxiv.org/html/2609.21509#A14)\. Improving round\-trip fidelity will require advances primarily on the encoding side, particularly for deep expressions where generators most frequently distort the intended arithmetic structure\.

### M\.1Misattribution Sensitivity

The consensus\-based attribution \([Section2\.4](https://arxiv.org/html/2609.21509#S2.SS4)\) classifies a word problem as a generator fault when no extractor in the panel recovers the target expression\. A faithful word problem will be misclassified whenever every extractor independently fails on it\. We bound this probability as follows\.

Letfif\_\{i\}denote the success rate of extractoriion confirmed\-faithful word problems \(those where at least one extractor succeeds\)\. Assuming conditional independence across extractors, the probability that a faithful word problem is misclassified as a generator fault is

P⁡\(misattribution\)=∏i=1N\(1−fi\)\.P\(\\text\{misattribution\}\)\\;=\\;\\prod\_\{i=1\}^\{N\}\(1\-f\_\{i\}\)\.\(5\)
[Table19](https://arxiv.org/html/2609.21509#A13.T19)reportsfif\_\{i\}for all 16 extractors, estimated from 14232 confirmed\-faithful word problems\. Values range from 0\.064 \(Qwen3\-0\.6B\) to 0\.820 \(GPT\-5\), with a mean off¯=0\.55\\bar\{f\}=0\.55\. Using the per\-extractor values in[Equation5](https://arxiv.org/html/2609.21509#A13.E5)givesP⁡\(misattribution\)=4\.5×10−7P\(\\text\{misattribution\}\)=4\.5\\times 10^\{\-7\}, yielding fewer than 1 expected false attribution among the 18038 word problems classified as generator faults by the consensus rule\.

The independence assumption is optimistic: extractors may share failure modes on the same hard word problems\. However, the panel spans six architectural families and four frontier APIs, limiting correlated failures\. Even under moderate positive correlation, the false\-attribution rate remains negligible given the diversity of the panel\.

Table 19:Per\-extractor success ratefif\_\{i\}on 14232 confirmed\-faithful word problems \(those where at least one extractor recovers the target\)\. Models sorted byfif\_\{i\}\.ExtractorCorrect / 14232fif\_\{i\}Qwen3\-0\.6B9130\.064Qwen3\-1\.7B31540\.222Dream\-7B39590\.278Gemma\-3\-4B41790\.294LLaDA\-8B63490\.446Gemma\-3\-12B68900\.484Qwen3\-4B72810\.512Qwen3\-8B77820\.547Gemma\-3\-27B90280\.634Qwen3\-14B93930\.660Qwen3\-32B97830\.687Phi\-499920\.702Haiku\-4\.5107080\.752Gemini\-3\-Flash114040\.801Gemini\-3\.1\-Pro115570\.812GPT\-5116720\.820Mean0\.545
### Panel\-composition sensitivity

The consensus\-based attribution depends on the extractor panel: a generation is classified as faithful whenever any single extractor recovers the target, so the generator\-fault rate is a function of panel composition\.[Table20](https://arxiv.org/html/2609.21509#A13.T20)reports the rate under nine panel subsets, all reusing the same extraction data\. Across these subsets the rate ranges 73\.6–91\.1%, with the full 16\-model panel yielding the lowest value: every restricted panel classifies more failures as generator faults, because smaller panels afford fewer chances for “at least one” extractor to succeed\. Within same\-size panels \(\|P\|=3\|P\|\{=\}3\), the Gemma\-3\-only panel gives 85\.8% and the other\-open\-weight panel gives 82\.6%, indicating that stronger extractors are if anything more likely to collectively fail on genuinely bad word problems\. The diversity\-controlled “one per family” panel \(77\.3%\) is close to Qwen3\-only \(79\.3%\) and above the full open\-weight panel \(74\.4%\), offering no evidence that correlated failures within a single family inflate the headline figure\.

Table 20:Generator\-fault rate under different extractor panel subsets \(incl\-guard convention\)\. The full\-panel rate of 73\.6% is the lowest observation; every restricted panel yields a higher rate\.Panel\|P\|\|P\|ErrorsGenF %Full \(all models\)1639227673\.6Open\-weight only1230853774\.4Qwen3 only615531479\.3Gemma\-3 only37671385\.8Other \(Phi/Dream/LLaDA\)37651082\.6Frontier only48373991\.1Strong \(≥\\geq14B \+ frontier\)817462384\.8Weak \(≤\\leq4B\)411355383\.7One per family \(median\)614697677\.3

## Appendix NLatent\-Ability Model \(IRT\)

As a validation of the additive decomposition in[Section2\.3](https://arxiv.org/html/2609.21509#S2.SS3), we fit a Rasch model\[[Rasch, 1980](https://arxiv.org/html/2609.21509#bib.bib11)\]to the binary per\-example outcomes in the communication matrix\. The fit uses all 16 models \(12 open\-weight and 4 frontier\)\. For each triple \(extractorii, generatorjj, expressionkk\), we model the probability of a correct round\-trip as

log⁡P⁡\(correct\)1−P⁡\(correct\)=μ\+θi\+βj−δk,\\log\\frac\{P\(\\text\{correct\}\)\}\{1\-P\(\\text\{correct\}\)\}\\;=\\;\\mu\+\\theta\_\{i\}\+\\beta\_\{j\}\-\\delta\_\{k\},\(6\)whereθi\\theta\_\{i\}is the extraction skill of modelii,βj\\beta\_\{j\}is the generation clarity of modeljj, andδk\\delta\_\{k\}is the difficulty of expressionkk\. Equivalently, the predicted success probability is

P⁡\(correct\)=11\+exp⁡\[−\(μ\+θi\+βj−δk\)\]\.P\(\\text\{correct\}\)\\;=\\;\\frac\{1\}\{1\+\\exp\\\!\\bigl\[\-\(\\mu\+\\theta\_\{i\}\+\\beta\_\{j\}\-\\delta\_\{k\}\)\\bigr\]\}\.\(7\)For example, given the estimated global baselineμ=−1\.90\\mu=\-1\.90, pairing Phi\-4 \(θ=\+0\.53\\theta=\+0\.53,β=\+1\.07\\beta=\+1\.07\) on a median\-difficulty expression \(δ=0\.1\\delta=0\.1\) givesP=σ⁡\(−1\.90\+0\.53\+1\.07−0\.1\)≈40\.1%P=\\sigma\(\-1\.90\+0\.53\+1\.07\-0\.1\)\\approx 40\.1\\%, while a hard expression \(δ=1\.2\\delta=1\.2,∼\\sim90th percentile\) yieldsP=σ⁡\(−1\.90\+0\.53\+1\.07−1\.2\)≈18\.2%P=\\sigma\(\-1\.90\+0\.53\+1\.07\-1\.2\)\\approx 18\.2\\%\. Parameters are estimated byL2L\_\{2\}\-regularised logistic regression \(regularisation strengthC=10C=10\)\.

[Table21](https://arxiv.org/html/2609.21509#A14.T21)reports the estimated coefficients for each model alongside its raw marginal accuracy from the communication matrix\. Spearman rank correlations between the IRT coefficients and the matrix marginals in[Table1](https://arxiv.org/html/2609.21509#S2.T1)areρ=1\.00\\rho=1\.00for extraction andρ=0\.96\\rho=0\.96for generation, confirming that the simple averages faithfully summarise latent ability\. Among the three parameter groups, generation clarity has the largest variance \(σβ=1\.93\\sigma\_\{\\beta\}=1\.93, compared withσδ=1\.00\\sigma\_\{\\delta\}=1\.00for expression difficulty andσθ=0\.98\\sigma\_\{\\theta\}=0\.98for extraction\), indicating that variation in how well models encode structure exceeds variation in expression difficulty or extraction skill\. Including all three parameter groups reduces the log\-loss from 0\.44 \(model\-only,θ\+β\\theta\+\\beta\) to 0\.38 \(full model withδ\\delta\), confirming that expression difficulty adds predictive signal beyond the model coefficients\.

Table 21:Rasch model coefficients and matrix marginal accuracies\.θ\\theta: extraction skill;β\\beta: generation clarity; Raw Ext\. / Gen\.: row / column averages of the16×1616\\times 16communication matrix in[Table1](https://arxiv.org/html/2609.21509#S2.T1)\(guard failures counted as errors\)\. Models are sorted by family and size\.Modelθ\\thetaRaw Ext\. \(%\)β\\betaRaw Gen\. \(%\)Qwen3\-0\.6B\-2\.892\.3\-4\.230\.1Qwen3\-1\.7B\-1\.438\.0\-1\.612\.6Qwen3\-4B\-0\.0818\.6\+0\.0213\.4Qwen3\-8B\+0\.1119\.9\+0\.5319\.0Qwen3\-14B\+0\.4224\.0\+0\.5520\.5Qwen3\-32B\+0\.4925\.0\+0\.4418\.8Gemma\-3\-4B\-0\.9210\.7\-3\.200\.7Gemma\-3\-12B\-0\.1417\.6\-1\.423\.3Gemma\-3\-27B\+0\.3623\.0\-1\.743\.3Phi\-4\+0\.5325\.5\+1\.0725\.8Dream\-7B\-1\.0710\.1\-1\.602\.2LLaDA\-8B\-0\.3416\.2\+0\.6615\.9Haiku\-4\.5\+0\.6727\.3\+2\.1046\.5GPT\-5\+0\.8129\.8\+2\.2250\.2Gemini\-3\-Flash\+0\.7729\.1\+1\.7439\.6Gemini\-3\.1\-Pro\+0\.8229\.5\+2\.5954\.4Spearmanρ\\rho1\.000\.96
## Appendix OFine\-Tuning: Assembly Domain Details

This appendix provides the full domain specification for both the nested assembly fine\-tuning data, along with training examples and qualitative analysis supporting the fine\-tuning experiment in[Section4](https://arxiv.org/html/2609.21509#S4)\.

### O\.1Assembly Domain Specification

Expressions are random trees over eleven assembly operators of mixed arity: three unary \(label,inspect,polish\), six binary \(three commutative:weld,glue,rivet; three non\-commutative:mount,pour,load\), and four ternary \(two commutative:mix,assort; two non\-commutative:layer,thread\)\. Operators are applied to coloured\-object leaves drawn from 13 colours×\\times13 objects\. Tree depth is sampled uniformly from 2 to 5, matching the benchmark range\. Each operator maps to multiple paraphrase templates describing physical assembly steps in a*procedural*style \(ordinal step markers\)\. Every \(expression, NL\) pair yields two training examples \(e→w\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\\\!\\to\\\!\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}andw→e\{\\color\[rgb\]\{0\.2344,0\.457,0\.6875\}w\}\\\!\\to\\\!\{\\color\[rgb\]\{0\.8555,0\.2188,0\.1953\}e\}\)\. We generate 500 pairs per depth with a 90/10 train/test split, producing approximately 3600 training messages\.

### O\.2Training Examples

[Table22](https://arxiv.org/html/2609.21509#A15.T22)shows three nested assembly expressions rendered in procedural style, illustrating the mixed arity \(unarypolish, binarymount, ternaryassort/load/mix\) and non\-commutative argument ordering\.

Table 22:Sample assembly expressions\(depth 3\) with procedural NL renderings\.ExpressionNatural\-language renderingassort\(label\(copper gem\), gold ring, yellow card\)First, mark the copper gem with a tag to get the first component\. Second, assort the first component, the gold ring, and the yellow card into one set as the final product\.mount\(rivet\(bronze ball, copper coin\), black cube\)First, clinch the bronze ball to the copper coin with a pop rivet to get the first component\. Second, position the first component on the face of the black cube as the final product\.polish\(layer\(silver ring, gold rod, blue cube\)\)First, deposit the silver ring at the bottom, the gold rod in the centre, and the blue cube at the top to get the first component\. Second, buff the first component to a shine as the final product\.
### O\.3Procedural assembly FT: failure taxonomy

The2\.12\.1pp guard\-pass drop reported for assembly FT \([Table2](https://arxiv.org/html/2609.21509#S4.T2)\) is concentrated in two models, not shared across the panel\.[Table23](https://arxiv.org/html/2609.21509#A15.T23)gives the per\-modelLeak\\mathrm\{Leak\}pass rate \(operator adjacency, the only category counted in the guard column\) for the base and assembly FT conditions\. Gemma\-3\-4B and Gemma\-3\-12B account for nearly all of the aggregate drop \(−22\.9\-22\.9and−17\.5\-17\.5pp\)\. Phi\-4 and Qwen3\-14B/32B stay within±0\.7\\pm 0\.7pp of their base, and Qwen3\-0\.6B/1\.7B improve \(\+32\.6\+32\.6and\+7\.0\+7\.0pp\) because their base outputs leak heavily and FT reduces that tendency\. We do not have a mechanistic account for the Gemma\-specific drop\. The*missing\-variable*fault \(a generator error under our reporting rule, not a guard failure\) rises by\+7\.0\+7\.0pp on average under FT, dominated again by Qwen3\-0\.6B \(\+42\.2\+42\.2\) and Qwen3\-1\.7B \(\+17\.8\+17\.8\), where the smallest generators drop placeholders more often after training\.

Table 23:Leak\\mathrm\{Leak\}pass rate under base and procedural assembly FT\(operator adjacency, the only failure category counted in theLeak\\mathrm\{Leak\}column of[Table2](https://arxiv.org/html/2609.21509#S4.T2)\)\. Values are the share of 1790 problems that clear the operator\-leak check\. Two Gemma models concentrate the aggregate drop\.ModelBase %FTasm\{\}\_\{\\text\{asm\}\}%Δ\\DeltappQwen3\-0\.6B19\.652\.2\+32\.6\+32\.6Qwen3\-1\.7B58\.965\.9\+7\.0\+7\.0Qwen3\-4B93\.982\.2−11\.7\-11\.7Qwen3\-8B97\.594\.0−3\.5\-3\.5Qwen3\-14B99\.298\.5−0\.7\-0\.7Qwen3\-32B98\.497\.7−0\.7\-0\.7Gemma\-3\-4B98\.075\.1−22\.9\\mathbf\{\-22\.9\}Gemma\-3\-12B96\.278\.7−17\.5\\mathbf\{\-17\.5\}Gemma\-3\-27B96\.192\.7−3\.4\-3\.4Phi\-496\.696\.5−0\.1\-0\.1Panel mean85\.483\.3−2\.1\-2\.1
### O\.4Qualitative Analysis

[Table24](https://arxiv.org/html/2609.21509#A15.T24)illustrates how assembly fine\-tuning change the style of generated word problems for Phi\-4\.

Table 24:Word problems generated by Phi\-4 \(base and assembly fine\-tuned FT\) for three expressions\. Assembly fine\-tuning shifts output toward explicit arithmetic descriptions\.Expression:\( C \- \( A \+ B \+ 2 \) \)\[depth 3\]BaseA real estate agent is calculating the remaining budget for a property renovation\. The total budget is \{C\}, and the costs include \{A\} for materials, \{B\} for labor, and an additional 2 for miscellaneous expenses\. The agent needs to determine how much budget is left after these expenses\.FTA developer had \{C\} plots of land\. They sold \{A\} plots in the spring, \{B\} plots in the summer, and set aside 2 for a park\. How many plots remained?Expression:\( B \- C \* \( A \+ G \) \)\[depth 3\]BaseA delivery truck starts with \{B\} gallons of fuel\. It uses \{C\} gallons for each mile driven, and it travels a total of \{A\} miles plus an additional \{G\} miles for detours\. How many gallons of fuel does the truck have left after completing the trip?FTA bus company had \{B\} buses in service\. They assigned \{A\} buses to the north route and \{G\} buses to the south route, then multiplied the total by \{C\} to get the number of buses used for express service\. How many buses remained in service?Expression:\( G \* \( D \+ C \) / H \)\[depth 3\]BaseA musician is composing a piece and wants to determine the total number of beats in a section\. They have \{G\} measures, each containing \{D\} beats plus an additional \{C\} beats for a special rhythm\. To find the average number of beats per measure, they divide the total beats by \{H\}\.FTFirst, add \{D\} and \{C\} together to get the first component\. Second, multiply \{G\} by the first component to get the second component\. Third, divide the second component by \{H\} to get the third component\. Fourth, the third component is the final product\.The base model produces fluent narratives whose structure often diverges from the target expression: in the first example,C \- \(A \+ B \+ 2\)becomes “remaining budget” after listed expenses, flattening the subtraction\-of\-a\-sum into a single subtractive relation\. The assembly\-fine\-tuned model makes operations explicit and sequential: narrative outputs preserve operator order \(“sold … in the spring, \{B\} plots in the summer, and set aside 2 for a park”\), and the third example collapses to a pure step\-marker form \(“First, add … Second, multiply … Third, divide …”\)\. This structural fidelity comes at a stylistic cost: two Gemma models contribute most of the guard\-rate drop reported in[Table23](https://arxiv.org/html/2609.21509#A15.T23)\.

## Appendix PShot\-Configuration Grid

[Figures8](https://arxiv.org/html/2609.21509#A16.F8),[9](https://arxiv.org/html/2609.21509#A16.F9)and[10](https://arxiv.org/html/2609.21509#A16.F10)report generation quality \(q𝖦q^\{\\mathsf\{G\}\}\) and extraction quality \(q𝖤q^\{\\mathsf\{E\}\}\) across all evaluated\(γ,ε\)\(\\gamma,\\varepsilon\)shot configurations for each training condition\. The base figure justifies the\(γ=0,ε=3\)\(\\gamma\{=\}0,\\varepsilon\{=\}3\)configuration used in[Table1](https://arxiv.org/html/2609.21509#S2.T1); the FT figures justify the per\-condition selections used in[Table2](https://arxiv.org/html/2609.21509#S4.T2)\. Bold marks the per\-row maximum in each panel\. All values \(%\) are averaged over the 10\-model open\-weight extractor panel with guard failures counted as errors\.

\(a\) Generation quality \(q𝖦q^\{\\mathsf\{G\}\}\)

γ=0\\gamma=0γ=3\\gamma=3Modelε=0\\varepsilon\{=\}0ε=3\\varepsilon\{=\}3ε=0\\varepsilon\{=\}0ε=3\\varepsilon\{=\}3Qwen3\-0\.6B0\.10\.10\.30\.3Qwen3\-1\.7B2\.62\.60\.30\.3Qwen3\-4B13\.714\.02\.93\.1Qwen3\-8B19\.720\.55\.35\.6Qwen3\-14B20\.721\.410\.611\.3Qwen3\-32B18\.319\.410\.010\.7Gemma\-3\-4B0\.80\.81\.01\.1Gemma\-3\-12B3\.83\.92\.62\.5Gemma\-3\-27B3\.63\.84\.14\.1Phi\-425\.926\.910\.410\.9
\(b\) Extraction quality \(q𝖤q^\{\\mathsf\{E\}\}\)

γ=0\\gamma=0γ=3\\gamma=3Modelε=0\\varepsilon\{=\}0ε=3\\varepsilon\{=\}3ε=0\\varepsilon\{=\}0ε=3\\varepsilon\{=\}3Qwen3\-0\.6B2\.63\.00\.70\.7Qwen3\-1\.7B3\.16\.61\.22\.1Qwen3\-4B12\.512\.45\.45\.4Qwen3\-8B12\.812\.95\.56\.0Qwen3\-14B14\.314\.56\.57\.1Qwen3\-32B14\.014\.76\.77\.0Gemma\-3\-4B8\.68\.43\.13\.3Gemma\-3\-12B11\.911\.75\.25\.3Gemma\-3\-27B14\.014\.26\.46\.5Phi\-415\.315\.16\.76\.7

Figure 8:Base condition: shot\-configuration grid\.\(a\) Generation quality \(q𝖦q^\{\\mathsf\{G\}\}\)

γ=0\\gamma=0γ=3\\gamma=3Modelε=0\\varepsilon\{=\}0ε=3\\varepsilon\{=\}3ε=0\\varepsilon\{=\}0ε=3\\varepsilon\{=\}3Qwen3\-0\.6B0\.00\.00\.30\.1Qwen3\-1\.7B0\.20\.12\.66\.5Qwen3\-4B20\.811\.212\.117\.3Qwen3\-8B22\.322\.823\.332\.6Qwen3\-14B20\.324\.827\.626\.2Qwen3\-32B23\.721\.320\.527\.6Gemma\-3\-4B6\.07\.43\.45\.5Gemma\-3\-12B7\.911\.512\.016\.7Gemma\-3\-27B23\.223\.125\.820\.4Phi\-436\.935\.234\.132\.7
\(b\) Extraction quality \(q𝖤q^\{\\mathsf\{E\}\}\)

γ=0\\gamma=0γ=3\\gamma=3Modelε=0\\varepsilon\{=\}0ε=3\\varepsilon\{=\}3ε=0\\varepsilon\{=\}0ε=3\\varepsilon\{=\}3Qwen3\-0\.6B3\.33\.92\.53\.7Qwen3\-1\.7B4\.17\.33\.48\.4Qwen3\-4B18\.217\.718\.320\.9Qwen3\-8B19\.618\.317\.621\.1Qwen3\-14B21\.920\.721\.824\.2Qwen3\-32B22\.822\.622\.525\.3Gemma\-3\-4B9\.19\.510\.011\.9Gemma\-3\-12B17\.314\.517\.319\.2Gemma\-3\-27B20\.020\.522\.824\.2Phi\-424\.922\.625\.626\.6

Figure 9:Assembly FT \(procedural\) condition: shot\-configuration grid\.Panel means are16\.13/15\.75/16\.17/18\.5416\.13/15\.75/16\.17/18\.54for\(γ,ε\)∈\{\(0,0\),\(0,3\),\(3,0\),\(3,3\)\}\(\\gamma,\\varepsilon\)\\in\\\{\(0,0\),\(0,3\),\(3,0\),\(3,3\)\\\}, so the panel optimum is\(γ=3,ε=3\)\(\\gamma\{=\}3,\\varepsilon\{=\}3\)\. Per\-row optima vary: Phi\-4 and Qwen3\-4B peak at\(0,0\)\(0,0\), Qwen3\-14B and Gemma\-3\-27B at\(3,0\)\(3,0\)\.\(a\) Generation quality \(q𝖦q^\{\\mathsf\{G\}\}\)

γ=0\\gamma=0Modelε=0\\varepsilon\{=\}0ε=3\\varepsilon\{=\}3Qwen3\-0\.6B47\.751\.9Qwen3\-1\.7B57\.057\.2Qwen3\-4B53\.255\.3Qwen3\-8B56\.459\.5Qwen3\-14B54\.855\.9Qwen3\-32B56\.058\.0Gemma\-3\-4B53\.955\.4Gemma\-3\-12B56\.557\.8Gemma\-3\-27B55\.058\.4Phi\-456\.658\.2
\(b\) Extraction quality \(q𝖤q^\{\\mathsf\{E\}\}\)

γ=0\\gamma=0Modelε=0\\varepsilon\{=\}0ε=3\\varepsilon\{=\}3Qwen3\-0\.6B13\.714\.1Qwen3\-1\.7B17\.234\.8Qwen3\-4B62\.864\.4Qwen3\-8B64\.766\.2Qwen3\-14B71\.072\.7Qwen3\-32B70\.073\.6Gemma\-3\-4B39\.237\.1Gemma\-3\-12B62\.259\.8Gemma\-3\-27B72\.270\.6Phi\-474\.074\.4

Figure 10:Arithmetic FT condition: shot\-configuration grid\.Arithmetic FT was evaluated withγ=0\\gamma\{=\}0only \(in\-domain generators do not benefit from exemplars\)\.
## Appendix QLossless\-Channel Extraction Baselines

To quantify the cost of using natural language as the communication channel, we replace the NL word problem with an unambiguous structured representation and measure how accurately the extractor recovers the original infix expression\. We evaluate two formats:

S\-expression \(prefix notation\)\.Each expression is deterministically converted to prefix notation,*e\.g\.,*\(A\+B\)×C\(A\+B\)\\times Cbecomes\(\* \(\+ A B\) C\)\.

JSON abstract syntax tree\.Each expression is converted to a nested JSON dictionary,*e\.g\.,*\(A\+B\)×C\(A\+B\)\\times Cbecomes\{"op": "\*", "left": \{"op": "\+", "left": "A", "right": "B"\}, "right": "C"\}\.

Both formats encode tree structure without ambiguity, so extraction accuracy on either task upper\-bounds what any faithful NL encoding could achieve for a given extractor\.

The S\-expression vs\. NL comparison scales with operator count for the Qwen3 family: for the strongest extractors \(14B–32B\), S\-expression accuracy stays high even atk=8k\{=\}8\(78–82%\), whereas NL extraction falls off sharply as expressions grow more complex\. The NL tax is thus negligible at low complexity and widens at high complexity, confirming that the natural\-language bottleneck specifically targets structurally complex expressions;[Figure3](https://arxiv.org/html/2609.21509#S3.F3)\-right quantifies this gap against the best available lossless format across extractor sizes\.

Format dependence at mid\-scale\.Comparing the two lossless channels across model sizes reveals that the structured\-format ceiling is not format\-invariant\. JSON extraction substantially exceeds S\-expression extraction for mid\-range models \(Qwen3\-4B: 63\.2% vs\. 40\.8%; Qwen3\-8B: 75\.9% vs\. 37\.5%\), while both converge at 32B \(≈88%\{\\approx\}\\,88\\%\)\. This gap likely reflects pretraining exposure: models encounter far more JSON than S\-expressions in their training corpora\. The NL tax should therefore be measured against the best available lossless format \(max⁡\(S\-expr,JSON\)\\max\(\\text\{S\-expr\},\\text\{JSON\}\)\), as reported in[Figure3](https://arxiv.org/html/2609.21509#S3.F3)\-right\.

## Appendix RPropositional\-Logic Replication

To test whether the communication bottleneck is specific to arithmetic or reflects a more general structural limitation, we replicate the protocol on propositional logic on a reduced14×1414\\times 14roster \(the arithmetic frontier additions Gemini\-3\-Flash and Gemini\-3\.1\-Pro are omitted to limit compute\)\. Formulas are built from five connectives \(AND, OR, NOT, IMPLIES, XOR\) over propositional variablesPP–WW, at depths 2–5\. Correctness is judged by*logical equivalence*via SymPy, so surface\-form variation \(e\.g\. commutativity of AND\) does not penalise the extractor\. A domain\-adapted leakage guard forbids logical operator words and symbols in the generated scenario, mirroring the arithmetic guard\.

Protocol\.The generator receives a formula and a randomly sampled real\-world domain \(access control, safety regulations, game rules, etc\.\) and must produce a 2–3 sentence scenario using proposition placeholders\{P\},\{Q\},…\\\{P\\\},\\\{Q\\\},\\ldotswithout any logical vocabulary\. The extractor receives the scenario and must recover the original formula using only the connectives AND, OR, NOT, IMPLIES, XOR\. The same 14 models serve as both generators and extractors\.

Example \(depth 3\)\.

> ``` Formula: ( ( P AND Q ) IMPLIES R ) Domain: safety regulations Scenario: When equipment {P} is running and the temperature reading {Q} exceeds the threshold, the automatic shutdown protocol {R} activates. ```

Results\.[Table25](https://arxiv.org/html/2609.21509#A18.T25)shows the full logic communication matrix\. Overall accuracy is lower than in arithmetic \(mean cell accuracy8\.7%8\.7\\%vs\.14\.9%14\.9\\%on the shared 14\-model block\), consistent with the additional ambiguity introduced by logical connectives that lack the concrete semantics of arithmetic operators\. Despite this difficulty shift, the relative ordering of models is largely preserved\.

[Figure11](https://arxiv.org/html/2609.21509#A18.F11)compares arithmetic and logic ranks for both generation and extraction quality\. Generation ranks are well preserved \(Spearmanρ=0\.85\\rho=0\.85, Kendallτb=0\.69\\tau\_\{b\}=0\.69\): GPT\-5 and Claude Haiku 4\.5 lead in both domains\. Extraction ranks show comparable pairwise agreement \(τb=0\.67\\tau\_\{b\}=0\.67\), though the Spearman correlation is lower \(ρ=0\.69\\rho=0\.69\) due to a single large displacement: Phi\-4 drops from extraction rank 3 in arithmetic to rank 14 in logic, showing a strong imbalance in terms of extraction abilities \(ranked 3rd on arithmetic\)\. This suggests that its strong arithmetic extraction reflects format\-specific parsing fluency rather than general structural competence\.

Excluding Phi\-4, extractionρ\\rhorises above generationρ\\rho, indicating that for all other models the extraction ranking transfers at least as reliably as the generation ranking\.

Table 25:The𝟏𝟒×𝟏𝟒\\bm\{14\\times 14\}communication matrix on all generated expressions \(guard failures counted as errors\)\.Each cell shows round\-trip accuracy \(%\)\. Rows are extractors, columns are generators\. The rightmost column and bottom row show extraction quality \(row average\) and generation quality \(column average\)\. Generator and extractor are run with promptsπ𝖦0\{\\color\[rgb\]\{0,0,0\}\{\\color\[rgb\]\{0,0,0\}\\pi\}\_\{\\mathsf\{G\}\}^\{0\}\}andπ𝖤ε\{\\color\[rgb\]\{0,0,0\}\{\\color\[rgb\]\{0,0,0\}\\pi\}\_\{\\mathsf\{E\}\}^\{\\varepsilon\}\}respectively \(configurationOPENγ=0,ε=3\)\\gamma\{=\}0,\\varepsilon\{=\}3\); see[AppendixP](https://arxiv.org/html/2609.21509#A16)for the full ablation\. The±\\pmvalues denote standard deviation across the 14 pairings in each row or column\. TheLeak\\mathrm\{Leak\}% row reports each model’s leakage\-guard pass rate as a generator\. Thin rules separate open\-weight auto\-regressive, diffusion, and frontier models\.GeneratorsExtractorsQwen3\-0\.6B Qwen3\-1\.7B Qwen3\-4B Qwen3\-8B Qwen3\-14B Qwen3\-32B Gemma\-3\-4B Gemma\-3\-12B Gemma\-3\-27B Phi\-4 Dream\-7B LLaDA\-8B C\-Haiku\-4\.5 GPT\-5 Extr\. quality Leak\\mathrm\{Leak\}pass %92969597989991959899878798100Qwen3\-0\.6B1\.42\.73\.51\.93\.83\.91\.52\.23\.25\.72\.72\.05\.65\.83\.3±\\pm1Qwen3\-1\.7B2\.74\.35\.95\.58\.47\.42\.13\.45\.212\.04\.44\.014\.219\.67\.1±\\pm5Qwen3\-4B3\.04\.37\.95\.59\.58\.92\.24\.06\.214\.65\.63\.718\.834\.59\.2±\\pm8Qwen3\-8B2\.64\.88\.26\.210\.510\.92\.14\.27\.715\.75\.04\.221\.540\.810\.3±\\pm10Qwen3\-14B2\.64\.88\.86\.911\.311\.62\.24\.77\.817\.14\.63\.622\.837\.810\.5±\\pm9Qwen3\-32B2\.64\.89\.48\.011\.713\.42\.05\.68\.119\.05\.44\.526\.254\.812\.5±\\pm13Gemma\-3\-4B2\.23\.86\.24\.47\.46\.82\.13\.04\.611\.64\.03\.214\.315\.96\.4±\\pm4Gemma\-3\-12B2\.54\.78\.26\.310\.710\.92\.14\.77\.815\.65\.34\.420\.135\.29\.9±\\pm9Gemma\-3\-27B2\.74\.58\.77\.411\.411\.22\.14\.57\.317\.44\.94\.522\.038\.910\.5±\\pm10Phi\-41\.72\.41\.10\.81\.12\.10\.60\.40\.52\.10\.70\.53\.64\.11\.5±\\pm1Dream\-7B1\.83\.02\.92\.54\.23\.41\.51\.62\.36\.22\.22\.17\.49\.13\.6±\\pm2LLaDA\-8B2\.43\.76\.66\.59\.09\.12\.34\.16\.914\.44\.34\.215\.627\.68\.3±\\pm7C\-Haiku\-4\.52\.65\.19\.78\.312\.813\.81\.94\.69\.819\.25\.85\.228\.057\.313\.1±\\pm14GPT\-52\.84\.99\.98\.514\.014\.52\.34\.68\.221\.15\.55\.028\.380\.415\.0±\\pm19Gen\. quality2\.4±\\pm04\.1±\\pm16\.9±\\pm35\.6±\\pm29\.0±\\pm49\.1±\\pm41\.9±\\pm03\.7±\\pm16\.1±\\pm313\.7±\\pm54\.3±\\pm13\.6±\\pm117\.8±\\pm833\.0±\\pm21

Figure 11:Arithmetic vs\. logic rank correlation for generation quality \(left\) and extraction quality \(right\)\. Each point is one model; the dashed diagonal marks perfect agreement\. Spearmanρ\\rhoand Kendallτb\\tau\_\{b\}are shown in each panel\. Phi\-4 \(bottom\-right in panel b\) is the sole large outlier, dropping from extraction rank 3 to rank 14\.
## Appendix SCompute Resources

All GPU experiments were run on NVIDIA A100 80GB nodes\. Frontier models were accessed over API, we report call counts \(one request per generated or extracted example\)\.

Table 26:Compute consumed by experiments reported in this paper\. GPU hours are wall\-clock on NVIDIA A100 80GB nodes, summed across generator and extractor roles\. API calls count requests to frontier models \(one request per generated or extracted example\) and are listed separately because they do not consume our GPU budget\.ExperimentA100 80GB \(h\)API callsSectionCommunication matrix \(arithmetic\)182\.585,644[Section3\.2](https://arxiv.org/html/2609.21509#S3.SS2)Communication matrix \(logic\)119\.272,714[Section3\.4](https://arxiv.org/html/2609.21509#S3.SS4)Lossless baseline \(S\-expression\)0\.6–[Section3\.2](https://arxiv.org/html/2609.21509#S3.SS2)Lossless baseline \(JSON AST\)0\.5–[Section3\.2](https://arxiv.org/html/2609.21509#S3.SS2)Fine\-tuning \(arithmetic\)68\.5–[Section4\.1](https://arxiv.org/html/2609.21509#S4.SS1)Fine\-tuning \(assembly\)57\.3–[Section4\.1](https://arxiv.org/html/2609.21509#S4.SS1)Temperature sweep \(Phi\-4\)13\.3–[AppendixE](https://arxiv.org/html/2609.21509#A5)Total441\.9158,358
## Appendix TAuthor Notes

### T\.1Broader Impact

This is a measurement study with no new model, human\-subjects data, or deployed system, so direct societal impact is limited\. Positively, the round\-trip protocol gives practitioners a principled way to audit model\-to\-model and chain\-of\-thought pipelines for silent structural errors before deployment\. Negatively, our fine\-tuning results show that modest data suffices to improve structured\-to\-text serialization, which could marginally aid production of persuasive structured content, a risk we judge small relative to existing generation capabilities and which is mitigated by open release of our code and data\.

### T\.2Use of LLMs

Claude Opus 4\.6\[[Anthropic, 2026a](https://arxiv.org/html/2609.21509#bib.bib66)\]and 4\.7\[[Anthropic, 2026b](https://arxiv.org/html/2609.21509#bib.bib67)\]were used under human supervision to assist with code, drafting, and discussion of research directions\. All scientific claims, experimental designs, and final phrasings were reviewed and validated by the authors, who take full responsibility for the content\.

## NeurIPS Paper Checklist

1. 1\.Claims
2. Question: Do the main claims made in the abstract and introduction accurately reflect the paper’s contributions and scope?
3. Answer:\[Yes\]
4. Justification: The abstract’s three findings \(lossy asymmetric channel with up to 60\.4 pp role\-swap gap and 92\.9% best pair,≥\\geq73\.6% of failures at generation with structure\-driven difficulty, and∼\\sim3600 fine\-tuning examples lifting open\-weight models above the strongest non\-fine\-tuned frontier generator\) are supported by[Sections3\.2](https://arxiv.org/html/2609.21509#S3.SS2),[3\.3](https://arxiv.org/html/2609.21509#S3.SS3)and[4](https://arxiv.org/html/2609.21509#S4)respectively\. Scope \(16 models, arithmetic plus logic replication\) is stated in the introduction\.
5. Guidelines: - •The answer\[N/A\]means that the abstract and introduction do not include the claims made in the paper\. - •The abstract and/or introduction should clearly state the claims made, including the contributions made in the paper and important assumptions and limitations\. A\[No\]or\[N/A\]answer to this question will not be perceived well by the reviewers\. - •The claims made should match theoretical and experimental results, and reflect how much the results can be expected to generalize to other settings\. - •It is fine to include aspirational goals as motivation as long as it is clear that these goals are not attained by the paper\.
6. 2\.Limitations
7. Question: Does the paper discuss the limitations of the work performed by the authors?
8. Answer:\[Yes\]
9. Justification: The limitations are openly discussed in[Section6](https://arxiv.org/html/2609.21509#S6)\.
10. Guidelines: - •The answer\[N/A\]means that the paper has no limitation while the answer\[No\]means that the paper has limitations, but those are not discussed in the paper\. - •The authors are encouraged to create a separate “Limitations” section in their paper\. - •The paper should point out any strong assumptions and how robust the results are to violations of these assumptions \(e\.g\., independence assumptions, noiseless settings, model well\-specification, asymptotic approximations only holding locally\)\. The authors should reflect on how these assumptions might be violated in practice and what the implications would be\. - •The authors should reflect on the scope of the claims made, e\.g\., if the approach was only tested on a few datasets or with a few runs\. In general, empirical results often depend on implicit assumptions, which should be articulated\. - •The authors should reflect on the factors that influence the performance of the approach\. For example, a facial recognition algorithm may perform poorly when image resolution is low or images are taken in low lighting\. Or a speech\-to\-text system might not be used reliably to provide closed captions for online lectures because it fails to handle technical jargon\. - •The authors should discuss the computational efficiency of the proposed algorithms and how they scale with dataset size\. - •If applicable, the authors should discuss possible limitations of their approach to address problems of privacy and fairness\. - •While the authors might fear that complete honesty about limitations might be used by reviewers as grounds for rejection, a worse outcome might be that reviewers discover limitations that aren’t acknowledged in the paper\. The authors should use their best judgment and recognize that individual actions in favor of transparency play an important role in developing norms that preserve the integrity of the community\. Reviewers will be specifically instructed to not penalize honesty concerning limitations\.
11. 3\.Theory assumptions and proofs
12. Question: For each theoretical result, does the paper provide the full set of assumptions and a complete \(and correct\) proof?
13. Answer:\[N/A\]
14. Justification: The paper does not contain formal theorems or theoretical results\. The main numbered equations in[Sections2\.3](https://arxiv.org/html/2609.21509#S2.SS3)and[2\.4](https://arxiv.org/html/2609.21509#S2.SS4)are derived from basic algebra and do not require formal proofs\.
15. Guidelines: - •The answer\[N/A\]means that the paper does not include theoretical results\. - •All the theorems, formulas, and proofs in the paper should be numbered and cross\-referenced\. - •All assumptions should be clearly stated or referenced in the statement of any theorems\. - •The proofs can either appear in the main paper or the supplemental material, but if they appear in the supplemental material, the authors are encouraged to provide a short proof sketch to provide intuition\. - •Inversely, any informal proof provided in the core of the paper should be complemented by formal proofs provided in appendix or supplemental material\. - •Theorems and Lemmas that the proof relies upon should be properly referenced\.
16. 4\.Experimental result reproducibility
17. Question: Does the paper fully disclose all the information needed to reproduce the main experimental results of the paper to the extent that it affects the main claims and/or conclusions of the paper \(regardless of whether the code and data are provided or not\)?
18. Answer:\[Yes\]
19. Justification: We provide all details about inference and prompts in[AppendixA](https://arxiv.org/html/2609.21509#A1), as well as detailed information about the fine\-tuning procedure in[AppendixO](https://arxiv.org/html/2609.21509#A15)and the S\-expr/JSON lossless experiments in[AppendixQ](https://arxiv.org/html/2609.21509#A17)\. Details for other side experiments are provided throughout the appendix\.
20. Guidelines: - •The answer\[N/A\]means that the paper does not include experiments\. - •If the paper includes experiments, a\[No\]answer to this question will not be perceived well by the reviewers: Making the paper reproducible is important, regardless of whether the code and data are provided or not\. - •If the contribution is a dataset and/or model, the authors should describe the steps taken to make their results reproducible or verifiable\. - •Depending on the contribution, reproducibility can be accomplished in various ways\. For example, if the contribution is a novel architecture, describing the architecture fully might suffice, or if the contribution is a specific model and empirical evaluation, it may be necessary to either make it possible for others to replicate the model with the same dataset, or provide access to the model\. In general\. releasing code and data is often one good way to accomplish this, but reproducibility can also be provided via detailed instructions for how to replicate the results, access to a hosted model \(e\.g\., in the case of a large language model\), releasing of a model checkpoint, or other means that are appropriate to the research performed\. - •While NeurIPS does not require releasing code, the conference does require all submissions to provide some reasonable avenue for reproducibility, which may depend on the nature of the contribution\. For example 1. \(a\)If the contribution is primarily a new algorithm, the paper should make it clear how to reproduce that algorithm\. 2. \(b\)If the contribution is primarily a new model architecture, the paper should describe the architecture clearly and fully\. 3. \(c\)If the contribution is a new model \(e\.g\., a large language model\), then there should either be a way to access this model for reproducing the results or a way to reproduce the model \(e\.g\., with an open\-source dataset or instructions for how to construct the dataset\)\. 4. \(d\)We recognize that reproducibility may be tricky in some cases, in which case authors are welcome to describe the particular way they provide for reproducibility\. In the case of closed\-source models, it may be that access to the model is limited in some way \(e\.g\., to registered users\), but it should be possible for other researchers to have some path to reproducing or verifying the results\.
21. 5\.Open access to data and code
22. Question: Does the paper provide open access to the data and code, with sufficient instructions to faithfully reproduce the main experimental results, as described in supplemental material?
23. Answer:\[No\]
24. Justification: We will release the code upon acceptance\. The data is procedurally generated with the code\.
25. Guidelines: - •The answer\[N/A\]means that paper does not include experiments requiring code\. - • - •While we encourage the release of code and data, we understand that this might not be possible, so\[No\]is an acceptable answer\. Papers cannot be rejected simply for not including code, unless this is central to the contribution \(e\.g\., for a new open\-source benchmark\)\. - •The instructions should contain the exact command and environment needed to run to reproduce the results\. See the NeurIPS code and data submission guidelines \([https://neurips\.cc/public/guides/CodeSubmissionPolicy](https://neurips.cc/public/guides/CodeSubmissionPolicy)\) for more details\. - •The authors should provide instructions on data access and preparation, including how to access the raw data, preprocessed data, intermediate data, and generated data, etc\. - •The authors should provide scripts to reproduce all experimental results for the new proposed method and baselines\. If only a subset of experiments are reproducible, they should state which ones are omitted from the script and why\. - •At submission time, to preserve anonymity, the authors should release anonymized versions \(if applicable\)\. - •Providing as much information as possible in supplemental material \(appended to the paper\) is recommended, but including URLs to data and code is permitted\.
26. 6\.Experimental setting/details
27. Question: Does the paper specify all the training and test details \(e\.g\., data splits, hyperparameters, how they were chosen, type of optimizer\) necessary to understand the results?
28. Answer:\[Yes\]
29. Justification: Details are provided in[Section3\.1](https://arxiv.org/html/2609.21509#S3.SS1), and fine\-tuning params in[Section4](https://arxiv.org/html/2609.21509#S4)\.
30. Guidelines: - •The answer\[N/A\]means that the paper does not include experiments\. - •The experimental setting should be presented in the core of the paper to a level of detail that is necessary to appreciate the results and make sense of them\. - •The full details can be provided either with the code, in appendix, or as supplemental material\.
31. 7\.Experiment statistical significance
32. Question: Does the paper report error bars suitably and correctly defined or other appropriate information about the statistical significance of the experiments?
33. Answer:\[Yes\]
34. Justification: Yes, std intervals are provided whenever available\.
35. Guidelines: - •The answer\[N/A\]means that the paper does not include experiments\. - •The authors should answer\[Yes\]if the results are accompanied by error bars, confidence intervals, or statistical significance tests, at least for the experiments that support the main claims of the paper\. - •The factors of variability that the error bars are capturing should be clearly stated \(for example, train/test split, initialization, random drawing of some parameter, or overall run with given experimental conditions\)\. - •The method for calculating the error bars should be explained \(closed form formula, call to a library function, bootstrap, etc\.\) - •The assumptions made should be given \(e\.g\., Normally distributed errors\)\. - •It should be clear whether the error bar is the standard deviation or the standard error of the mean\. - •It is OK to report 1\-sigma error bars, but one should state it\. The authors should preferably report a 2\-sigma error bar than state that they have a 96% CI, if the hypothesis of Normality of errors is not verified\. - •For asymmetric distributions, the authors should be careful not to show in tables or figures symmetric error bars that would yield results that are out of range \(e\.g\., negative error rates\)\. - •If error bars are reported in tables or plots, the authors should explain in the text how they were calculated and reference the corresponding figures or tables in the text\.
36. 8\.Experiments compute resources
37. Question: For each experiment, does the paper provide sufficient information on the computer resources \(type of compute workers, memory, time of execution\) needed to reproduce the experiments?
38. Answer:\[Yes\]
39. Justification: Compute resources \(NVIDIA A100 80GB GPU hours and API hours\) are reported per experiment in[Table26](https://arxiv.org/html/2609.21509#A19.T26)\([AppendixS](https://arxiv.org/html/2609.21509#A19)\)\. Preliminary and failed runs are not included in those totals, so the full project consumed additional compute not reported here\.
40. Guidelines: - •The answer\[N/A\]means that the paper does not include experiments\. - •The paper should indicate the type of compute workers CPU or GPU, internal cluster, or cloud provider, including relevant memory and storage\. - •The paper should provide the amount of compute required for each of the individual experimental runs as well as estimate the total compute\. - •The paper should disclose whether the full research project required more compute than the experiments reported in the paper \(e\.g\., preliminary or failed experiments that didn’t make it into the paper\)\.
41. 9\.Code of ethics
43. Answer:\[Yes\]
44. Justification: This work has no conflict with the given code of ethics\.
45. Guidelines: - •The answer\[N/A\]means that the authors have not reviewed the NeurIPS Code of Ethics\. - •If the authors answer\[No\], they should explain the special circumstances that require a deviation from the Code of Ethics\. - •The authors should make sure to preserve anonymity \(e\.g\., if there is a special consideration due to laws or regulations in their jurisdiction\)\.
46. 10\.Broader impacts
47. Question: Does the paper discuss both potential positive societal impacts and negative societal impacts of the work performed?
48. Answer:\[Yes\]
49. Justification: We include a Broader Impact section in appendix \([SectionT\.1](https://arxiv.org/html/2609.21509#A20.SS1)\), discussing the very limited negative impact of our work\.
50. Guidelines: - •The answer\[N/A\]means that there is no societal impact of the work performed\. - •If the authors answer\[N/A\]or\[No\], they should explain why their work has no societal impact or why the paper does not address societal impact\. - •Examples of negative societal impacts include potential malicious or unintended uses \(e\.g\., disinformation, generating fake profiles, surveillance\), fairness considerations \(e\.g\., deployment of technologies that could make decisions that unfairly impact specific groups\), privacy considerations, and security considerations\. - •The conference expects that many papers will be foundational research and not tied to particular applications, let alone deployments\. However, if there is a direct path to any negative applications, the authors should point it out\. For example, it is legitimate to point out that an improvement in the quality of generative models could be used to generate Deepfakes for disinformation\. On the other hand, it is not needed to point out that a generic algorithm for optimizing neural networks could enable people to train models that generate Deepfakes faster\. - •The authors should consider possible harms that could arise when the technology is being used as intended and functioning correctly, harms that could arise when the technology is being used as intended but gives incorrect results, and harms following from \(intentional or unintentional\) misuse of the technology\. - •If there are negative societal impacts, the authors could also discuss possible mitigation strategies \(e\.g\., gated release of models, providing defenses in addition to attacks, mechanisms for monitoring misuse, mechanisms to monitor how a system learns from feedback over time, improving the efficiency and accessibility of ML\)\.
51. 11\.Safeguards
52. Question: Does the paper describe safeguards that have been put in place for responsible release of data or models that have a high risk for misuse \(e\.g\., pre\-trained language models, image generators, or scraped datasets\)?
53. Answer:\[N/A\]
54. Justification: This work poses no risk related to data or model misuse\.
55. Guidelines: - •The answer\[N/A\]means that the paper poses no such risks\. - •Released models that have a high risk for misuse or dual\-use should be released with necessary safeguards to allow for controlled use of the model, for example by requiring that users adhere to usage guidelines or restrictions to access the model or implementing safety filters\. - •Datasets that have been scraped from the Internet could pose safety risks\. The authors should describe how they avoided releasing unsafe images\. - •We recognize that providing effective safeguards is challenging, and many papers do not require this, but we encourage authors to take this into account and make a best faith effort\.
56. 12\.Licenses for existing assets
57. Question: Are the creators or original owners of assets \(e\.g\., code, data, models\), used in the paper, properly credited and are the license and terms of use explicitly mentioned and properly respected?
58. Answer:\[Yes\]
59. Justification: We cite all models used in this work\. All open\-source models have Apache2\.0 or MIT licenses\. We additionally have direct permission from Anthropic and OpenAI to use their respective models\.
60. Guidelines: - •The answer\[N/A\]means that the paper does not use existing assets\. - •The authors should cite the original paper that produced the code package or dataset\. - •The authors should state which version of the asset is used and, if possible, include a URL\. - •The name of the license \(e\.g\., CC\-BY 4\.0\) should be included for each asset\. - •For scraped data from a particular source \(e\.g\., website\), the copyright and terms of service of that source should be provided\. - •If assets are released, the license, copyright information, and terms of use in the package should be provided\. For popular datasets,[paperswithcode\.com/datasets](https://paperswithcode.com/datasets)has curated licenses for some datasets\. Their licensing guide can help determine the license of a dataset\. - •For existing datasets that are re\-packaged, both the original license and the license of the derived asset \(if it has changed\) should be provided\. - •If this information is not available online, the authors are encouraged to reach out to the asset’s creators\.
61. 13\.New assets
62. Question: Are new assets introduced in the paper well documented and is the documentation provided alongside the assets?
63. Answer:\[N/A\]
64. Justification: We do not release new assets\. Upon acceptance, code will be released, which will allow to procedurally generate new arithmetic expressions\.
65. Guidelines: - •The answer\[N/A\]means that the paper does not release new assets\. - •Researchers should communicate the details of the dataset/code/model as part of their submissions via structured templates\. This includes details about training, license, limitations, etc\. - •The paper should discuss whether and how consent was obtained from people whose asset is used\. - •At submission time, remember to anonymize your assets \(if applicable\)\. You can either create an anonymized URL or include an anonymized zip file\.
66. 14\.Crowdsourcing and research with human subjects
67. Question: For crowdsourcing experiments and research with human subjects, does the paper include the full text of instructions given to participants and screenshots, if applicable, as well as details about compensation \(if any\)?
68. Answer:\[N/A\]
69. Justification: No human subjects are used in this work\.
70. Guidelines: - •The answer\[N/A\]means that the paper does not involve crowdsourcing nor research with human subjects\. - •Including this information in the supplemental material is fine, but if the main contribution of the paper involves human subjects, then as much detail as possible should be included in the main paper\. - •According to the NeurIPS Code of Ethics, workers involved in data collection, curation, or other labor should be paid at least the minimum wage in the country of the data collector\.
71. 15\.Institutional review board \(IRB\) approvals or equivalent for research with human subjects
72. Question: Does the paper describe potential risks incurred by study participants, whether such risks were disclosed to the subjects, and whether Institutional Review Board \(IRB\) approvals \(or an equivalent approval/review based on the requirements of your country or institution\) were obtained?
73. Answer:\[N/A\]
74. Justification: Our work does not involve crowdsourcing nor research with human subjects\.
75. Guidelines: - •The answer\[N/A\]means that the paper does not involve crowdsourcing nor research with human subjects\. - •Depending on the country in which research is conducted, IRB approval \(or equivalent\) may be required for any human subjects research\. If you obtained IRB approval, you should clearly state this in the paper\. - •We recognize that the procedures for this may vary significantly between institutions and locations, and we expect authors to adhere to the NeurIPS Code of Ethics and the guidelines for their institution\. - •For initial submissions, do not include any information that would break anonymity \(if applicable\), such as the institution conducting the review\.
76. 16\.Declaration of LLM usage
77. Question: Does the paper describe the usage of LLMs if it is an important, original, or non\-standard component of the core methods in this research? Note that if the LLM is used only for writing, editing, or formatting purposes and does*not*impact the core methodology, scientific rigor, or originality of the research, declaration is not required\.
78. Answer:\[Yes\]
79. Justification: We describe the use of LLMs in[SectionT\.2](https://arxiv.org/html/2609.21509#A20.SS2)\.
80. Guidelines: - •The answer\[N/A\]means that the core method development in this research does not involve LLMs as any important, original, or non\-standard components\. - •Please refer to our LLM policy in the NeurIPS handbook for what should or should not be described\.

相似文章