LexLattice: Multilingual Extractive Summarization via Neural Cellular Automata on Document Hierarchies
Summary
LexLattice introduces a multilingual extractive summarization method using neural cellular automata on document hierarchies, achieving state-of-the-art ROUGE scores on the EUR-Lex-Sum dataset with a compact model surpassing large instruction-tuned baselines.
View Cached Full Text
Cached at: 09/24/26, 09:13 AM
# LexLattice: Multilingual Extractive Summarization via Neural Cellular Automata on Document Hierarchies
Source: [https://arxiv.org/html/2609.27032](https://arxiv.org/html/2609.27032)
Sujay Uday RittikarAffiliation:Applied Computer ScienceAffiliation:The University of WinnipegAffiliation:Winnipeg, MBEmail:[rittikar\-s@webmail\.uwinnipeg\.ca](mailto:)Sheela RamannaAffiliation:Applied Computer ScienceAffiliation:The University of WinnipegAffiliation:Winnipeg, MBEmail:[s\.ramanna@uwinnipeg\.ca](mailto:)
###### Abstract
Faithfulness is a central concern in legal text summarization, which motivates extractive approaches that select verbatim content traceable to its source\. Such methods typically rank paragraphs or other structural units in isolation, yet give little attention to*consolidating*evidence that is distributed across, and shares salience between, distant parts of a document\. We introduce LexLattice, an extractive summarizer that reifies a legal act’s hierarchy as a two\-dimensional semantic lattice and consolidates over it with a masked 2D neural cellular automata before selection\. LexLattice attains state\-of\-the art ROUGE across all 24 languages of EUR\-Lex\-Sum in both multilingual and cross\-lingual settings, surpassing instruction\-tuned baselines with billions of parameters, despite concentrating all trainable capacity in a 1\.8M parameter consolidator over a frozen multilingual encoder\. A consolidator trained only on high\-resource languages further transfers to unseen languages with near\-lossless retention \(0\.99\), indicating that the model operates on language\-agnostic semantic geometry rather than surface form\. Our results position explicit consolidation over document structure as a compact and traceable alternative to scale for multilingual legal summarization\.
*K*eywordsNeural Cellular Automata⋅\\cdotReinforcement Learning⋅\\cdotMultilingual Summarization⋅\\cdotExtractive Summarization
## 1Introduction
The European Union \(EU\) enacts its legislation in 24 official languages, and a single legal act often runs to tens of thousands of tokens of dense, formulaic prose that binds not only legal professionals but the citizens and businesses subject to it\. Automatic summarization is therefore a natural instrument for access to law\. It is an important tool to turn lengthy legislative instruments into concise, faithful accounts of what they require to present\. However, this is a demanding setting: Documents run well beyond standard encoder limits, and a system must work uniformly across languages that differ sharply in available resources\. Faithfulness is the sharper of these pressures\[[1](https://arxiv.org/html/2609.27032#bib.bib14)\], and it is what makes extractive approaches especially attractive here: a summary assembled from verbatim provisions can be traced back to the exact article it came from, and it must not invent legal content that the act does not contain\. Yet the standard neural extractive recipe encodes the document and then scores each sentence in reading order, as one long flat sequence\[[2](https://arxiv.org/html/2609.27032#bib.bib9),[3](https://arxiv.org/html/2609.27032#bib.bib6)\], discarding structure that legal drafting marks out explicitly\. An EU act is not a stream of sentences but a hierarchy of recitals, articles, paragraphs, and annexes, and a passage’s position within it says a great deal about whether the passage belongs in a summary: the first paragraph of an operative article typically states an obligation, a recital supplies background reasoning, an annex lists technical detail\. A model must infer these roles from wording alone, even though the document already declares them\[[4](https://arxiv.org/html/2609.27032#bib.bib3),[5](https://arxiv.org/html/2609.27032#bib.bib4),[6](https://arxiv.org/html/2609.27032#bib.bib11)\]\.
EU legal act \(plain text\)Hierarchy parsermT5\-base encoder∗\\astfrozenTile to lattice48×3248\{\\times\}32, maskMM2D\-NCATTshared stepsReadoutMLP→\\toscoresis\_\{i\}Budgeted selection\(s,p\)\(s,p\)N×768N\{\\times\}768768×48×32768\{\\times\}48\{\\times\}32\[h0∥hT\]\[h^\{0\}\\\|h^\{T\}\]×T\\times Thth\_\{t\}3×33\{\\times\}3DW perception\[ht∥DWConvht\]\[\\,h\_\{t\}\\\|\\mathrm\{DWConv\}\\,h\_\{t\}\\,\]1×11\{\\times\}1conv→\\toReLU1536→2561536\\to 2561×11\{\\times\}1conv \(0\-init\)256→768256\\to 768\+\+⊙\\odotht\+1h\_\{t\+1\}residualmaskMMone NCA stepFigure 1:The LexLattice architecture\. A rule\-based parser \(Section[3\.1](https://arxiv.org/html/2609.27032#S3.SS1)\) recovers the act’s hierarchy; a frozen mT5\-base encoder embeds each paragraph, and embeddings are tiled onto the semantic lattice \(cf\. Fig\.[2](https://arxiv.org/html/2609.27032#S4.F2)\)\. The masked 2D\-NCA applies one shared local updateT=8T\{=\}8times \(right: one step\), and a small readout scores each un\-tiled paragraph for budgeted selection\. Blue blocks are the only trainable components \(≈\\approx1\.8M parameters\); the maskMMkeeps padding out of every update\.In this work, we proposeLexLattice, an extractive summarizer that accounts for document structure by reifying it as a two\-dimensionalsemantic lattice: rows index the sections of an act, columns index paragraph position within a section, and each cell carries the frozen multilingual encoder state\[[7](https://arxiv.org/html/2609.27032#bib.bib13)\]of the corresponding text unit\. Over this lattice we run a masked two\-dimensional Neural Cellular Automata \(NCA\)\[[8](https://arxiv.org/html/2609.27032#bib.bib8)\], a parameter\-light convolutional update rule applied iteratively, so that salience evidence propagates locallyalongsections \(between adjacent paragraphs\) andacrossthem \(between structurally parallel positions\) before any selection decision is made\. Our hypothesis is that document structure, once made geometric, allows a small recurrent local model perform the global evidence consolidation that sequence models otherwise, obtain through long\-range self\-attention and far heavier parameterization\. We evaluate on EUR\-Lex\-Sum across all 24 languages in multilingual and all\-pairs cross\-lingual settings, comparing a capacity\-matched 1D\-NCA control, the 2D LexLattice, and its reinforcement\-tuned LexLattice \+ RLOO\. With only 1\.8M trainable parameters over a frozen mT5\-base encoder, LexLattice \+ RLOO achieves the best ROUGE among all systems, including recent LLM and SLM baselines, while remaining competitive on BERTScore\-F1\.
## 2Related Works
The EUR\-Lex\-Sum benchmark\[[9](https://arxiv.org/html/2609.27032#bib.bib2)\]provides reference baselines in all languages from a zero\-shot multilingual LexRank extractor\[[10](https://arxiv.org/html/2609.27032#bib.bib23)\], and includes a cross\-lingual English\-to\-Spanish track evaluated with translate\-then\-summarize pipelines over OPUS\-MT\[[11](https://arxiv.org/html/2609.27032#bib.bib25)\]and a Longformer Encoder\-Decoder\[[12](https://arxiv.org/html/2609.27032#bib.bib24)\]\.[Sie et al\. \[13\]](https://arxiv.org/html/2609.27032#bib.bib19)propose a multi\-step extractive\-then\-abstractive pipeline for long regulatory documents, coupling a LexLM extractor to LongT5\[[14](https://arxiv.org/html/2609.27032#bib.bib26)\], Pegasus\[[15](https://arxiv.org/html/2609.27032#bib.bib27)\], or a QLoRA\-tuned Llama\-3 decoder\[[16](https://arxiv.org/html/2609.27032#bib.bib28)\]\. Their two\-stage design first extracts salient content and then rewrites it abstractively, improving fluency and coherence over the extracted intermediate summary\. The extractive stage that precedes it remains essential for keeping the generator’s input faithful to the source, a concern that becomes central in the legal domain\. While the above mentioned systems focus on the summarization pipeline,[T\.Y\.S\.S et al\. \[17\]](https://arxiv.org/html/2609.27032#bib.bib22)instead emphasize the model: they introduce LexT5, a sequence\-to\-sequence model pre\-trained, fine\-tuned, and probed on legal\-domain knowledge, and benchmark it against Longformer and PRIMERA\[[18](https://arxiv.org/html/2609.27032#bib.bib29)\]variants\. Complementing these,[Bendahman et al\. \[19\]](https://arxiv.org/html/2609.27032#bib.bib21)presented an analysis of entity\-level faithfulness on the same corpus, illustrating distinction between expert\-introduced abstractive entities and spurious hallucinations\. Recently, the trend of Small Language Models \(SLMs\) has encouraged various works on text summarization\[[20](https://arxiv.org/html/2609.27032#bib.bib31),[21](https://arxiv.org/html/2609.27032#bib.bib30)\]\. A study by[Medina\-Ramírez et al\. \[22\]](https://arxiv.org/html/2609.27032#bib.bib20)benchmarked the small instruction\-tuned language models against the large language models \(LLMs\) on six languages of EUR\-Lex\-Sum across an extensive list of models\. Their results indicate that instruction tuning alone outperforms task fine\-tuning under the LLM\-as\-judge evaluations\.
## 3Dataset
Table 1:EUR\-Lex\-Sum statistics per language after structural pre\-processing\. Validation and test splits contain 187 and 188 \(act, summary\) pairs for every language\. Word, section, and paragraph counts are means over the training split\.*Tier*marks the cross\-lingual transfer split: we train on the 12 high\-resource \(H\) languages and evaluate zero\-shot on the 12 low\-resource \(L\) languages\. The bottom row gives the total training documents and macro\-averages over languages\.†Irish: 16 training documents, and the longest source documents in the corpus\.We use EUR\-Lex\-Sum\[[9](https://arxiv.org/html/2609.27032#bib.bib2)\], a manually curated multilingual and cross\-lingual dataset of legal acts and their summaries, drawn from the EUR\-Lex law platform and covering all 24 official EU languages\. We use the released train/validation/test partitions in full, without language or document sampling\. Table[1](https://arxiv.org/html/2609.27032#S3.T1)reports per\-language statistics after pre\-processing\. Each record consists of a CELEX identifier\[[23](https://arxiv.org/html/2609.27032#bib.bib16)\], the full text of the act, and the reference summary, all as plain newline\-delimited text\. The release carries no structural markup, thus, the document structure must be recovered from the text itself\.
### 3\.1Rule\-based hierarchy parsing
We recover each act’s hierarchy with a deterministic, rule\-based parser built on EU legal drafting conventions, using a manually verified keyword table for all languages covering the head\-words forarticle,chapter,section,title, andannex\. The article patterns accommodate bothword\-first layouts\("Article 1", "Artigo 1\.º"\) andnumber\-first layouts\("1 artikla", "1\. cikk"\); recitals are detected by the pan\-European "\(N\)" line convention and numbered paragraphs by leading "N\." markers\. A line is considered as a potential header only if it is shorter than 60 characters, guarding against body text that merely contains a keyword\. The parser emits a list of sections, each holding an ordered list of paragraphs\. Only four section kinds become structural units downstream:preamble, therecitals block,articles, andannexes, while chapter/section/title headings act purely as paragraph\-flow separators\. Sections with no paragraphs are discarded\.
#### Text units\.
The atomic unit for selection is theparagraph, addressed by its coordinate \(section index, paragraph index\)\. For encoding only, paragraphs longer than 256 multilingual T5 \(mT5\)\[[7](https://arxiv.org/html/2609.27032#bib.bib13)\]tokens are split into consecutive 256\-token chunks whose embeddings are later mean\-pooled back to a single paragraph vector\. Chunking therefore stays internal to encoding: the model still selects, and is supervised on, whole paragraphs rather than chunks, so the unit of prediction matches the unit of the oracle target\.
## 4Preliminaries
### 4\.1The Semantic Lattice
PreambleRecitalsArticle 1Article 2Article 3Article 4Article 5Article 6Article 7Article 8Article 9Annex IAnnex IIAnnex III123456789101112paragraph position within section→\\rightarrowsection \(lattice row\)
PreambleRecitalsArticleAnnexPadding \(masked\)
Figure 2:Semantic\-lattice tiling of a EUR\-Lex act \(32007R0458\)\. Rows index sections recovered by the hierarchy parser \(Section[3\.1](https://arxiv.org/html/2609.27032#S3.SS1)\), columns index paragraph position within a section, and each occupied cell holds the frozen mT5 embedding of its paragraph\. Grey cells are masked padding\. Only the occupied14×1214\{\\times\}12region of the full48×3248\{\\times\}32lattice is shown\.We represent the parsed hierarchy as a fixed two\-dimensional grid of sizeH×W=48×32H\\times W=48\\times 32, where rows index sections and columns index paragraph positions within a section \(Figure[2](https://arxiv.org/html/2609.27032#S4.F2)\)\. This grid is the substrate of an NCA, a dynamical computational model that learns a single shared update rule through which cells self\-organize over a lattice: each cell holds a 768\-dimensional state, initialized to its paragraph’s frozen embedding, and is updated only from its immediate neighbourhood\. We set the grid dimensions from a scan of the training corpus, choosing a grid that accommodates the large majority of documents without inflating the number of empty cells\. Because a fixed grid cannot cover every document, we adopt aclamp\-and\-mergepolicy rather than truncating: a paragraph whose section or position index falls outside the grid is clamped into the last row or column, and a cell receiving several paragraphs stores their mean vector\.Tilingscatters each paragraph vector into its cell and records a binary occupancy mask, and masked cells never contribute to updates;un\-tilingbroadcasts the updated cell state back to every paragraph mapped to that cell\. Both operations are differentiable, so the consolidator trains end to end and every paragraph receives a score\. Because the axes are semantic, a cell’s neighbors are its adjacent paragraphs within a section and the paragraphs at the same relative position in neighboring sections\.
### 4\.2Supervision: a Greedy Extractive Oracle
Following standard practice in extractive summarization\[[2](https://arxiv.org/html/2609.27032#bib.bib9),[3](https://arxiv.org/html/2609.27032#bib.bib6)\], we construct a greedy oracle per document to serve as the training target\. LetD=\{p1,…,pN\}D=\\\{p\_\{1\},\\dots,p\_\{N\}\\\}be theNNparagraphs of a documentDD, andRRits reference summary\. The quality of a selectionS⊆DS\\subseteq Dagainst the referenceRRis
g\(S\)=12\(R\-1F1\(S,R\)\+R\-2F1\(S,R\)\),g\(S\)=\\frac\{1\}\{2\}\\Bigl\(\\text\{R\-1\}\_\{F\_\{1\}\}\(S,R\)\+\\text\{R\-2\}\_\{F\_\{1\}\}\(S,R\)\\Bigr\),\(1\)whereR\-1F1\\text\{R\-1\}\_\{F\_\{1\}\}andR\-2F1\\text\{R\-2\}\_\{F\_\{1\}\}denote unigram and bigramRougeF1F\_\{1\}respectively\. Starting from an empty selectionS0=∅S\_\{0\}=\\varnothing, we writeℛt=D∖St−1\\mathcal\{R\}\_\{t\}=D\\setminus S\_\{t\-1\}for the paragraphs still available at steptt, and add the one whose marginal gain inggis largest,
πt=argmaxp∈ℛt\[g\(St−1∪\{p\}\)−g\(St−1\)\],\\pi\_\{t\}=\\operatorname\*\{arg\\,max\}\_\{p\\,\\in\\,\\mathcal\{R\}\_\{t\}\}\\Bigl\[\\,g\\bigl\(S\_\{t\-1\}\\cup\\\{p\\\}\\bigr\)\-g\(S\_\{t\-1\}\)\\,\\Bigr\],\(2\)settingSt=St−1∪\{πt\}S\_\{t\}=S\_\{t\-1\}\\cup\\\{\\pi\_\{t\}\\\}and terminating when no remaining paragraph yields a positive gain or when\|St\|=40\|S\_\{t\}\|=40\. This cap is a safeguard: the zero\-gain criterion terminates first for the large majority of training documents, and where the cap does bind, it removes only the lowest\-gain tail of the supervision sequence\. The procedure yields both a selection setSSand the order in which its members were accepted\. It is the ordered sequenceπ=\(π1,…,πK\)\\pi=\(\\pi\_\{1\},\\dots,\\pi\_\{K\}\), withK=\|S\|≤NK=\|S\|\\leq N, that provides the supervision signal for the warm\-start stage \(Section[5\.1](https://arxiv.org/html/2609.27032#S5.SS1)\)\.
### 4\.3Model: LexLattice
Figure[1](https://arxiv.org/html/2609.27032#S1.F1)shows the full pipeline; we describe each component in turn, using the notation of the figure throughout\.
#### Frozen multilingual encoder\.
Every paragraph is embedded with a frozen mT5\-base encoder as the attention\-masked mean of its final\-layer token states \(768\-d;≤\\leq256 tokens per chunk, with longer paragraphs mean\-pooling their chunk embeddings\), yielding theN×768N\\times 768paragraph states that are tiled onto the lattice\. Frozen encoder serves three purposes: \(i\) it holds the shared multilingual representation space fixed as an experimental constant, \(ii\) it concentrates all trainable capacity in the consolidator and readout, and \(iii\) it allows embeddings to be computed once and cached\.
#### Masked 2D NCA\.
The consolidator applies the same local update to the lattice statehth^\{t\}forT=8T\{=\}8steps \(Figure[1](https://arxiv.org/html/2609.27032#S1.F1), right\)\. A perception stage concatenates each cell with a learned depthwise \(DW\)3×33\{\\times\}3convolution of its neighbourhood,\[ht∥DWConvht\]\[\\,h^\{t\}\\\|\\mathrm\{DWConv\}\\,h^\{t\}\\,\]\(768→1,536768\\to 1\{,\}536channels\), and an update MLP of two1×11\{\\times\}1convolutions \(1,536→256→7681\{,\}536\\to 256\\to 768, ReLU\) produces a residual state change\. The final projection is zero\-initialized, so the untrained automata is the identity map\. The update is deterministic and is multiplied by the occupancy maskMMat every step, so padding cells never contribute to, or drift from, the zero state\. Salience evidence thus propagates locally bothalongsections \(adjacent paragraphs\) andacrossstructurally parallel positions in neighboring sections\. Optionally, a learned language embedding \(24×76824\\times 768\) is added to every occupied cell before the first step\. The consolidator and readout together hold≈1\.80\{\\approx\}1\.80M trainable parameters: the frozen encoder does all representation work, and the automata only consolidates\.
#### Capacity\-matched 1D control\.
To isolate the contribution of the second, structural axis we train an otherwise identical 1D\-NCA: the same perception and update architecture with width\-3 1D convolutions over paragraphs in reading order, and the same step count, hidden width, initialization, readout, loss, and training loop\. Any difference between the two systems is attributable to lattice geometry alone\.
#### Readout and selection\.
A residual readout scores each paragraph from the concatenation of its initial and consolidated states,\[h0∥hT\]\[\\,h^\{0\}\\\|h^\{T\}\\,\]:Linear\(1,536→768\)→GELU→Linear\(768→1\)\\mathrm\{Linear\}\(1\{,\}536\\to 768\)\\to\\mathrm\{GELU\}\\to\\mathrm\{Linear\}\(768\\to 1\), yielding a scalar salience scoresis\_\{i\}\. At inference, paragraphs are ranked bysis\_\{i\}and greedily accepted until the cumulative word count reaches the reference summary’s length; accepted paragraphs are emitted in document order\.
## 5Methodology
Stage 1 trains the model that we evaluate asLexLattice; Stage 2 then asks whether reinforcement learning provides any additional value\.
### 5\.1Stage 1: Listwise Warm\-Start
The readout assigns a score to every paragraph of the document, givings1,…,sNs\_\{1\},\\dots,s\_\{N\}\. These scores define a Plackett–Luce distribution over paragraphs\[[24](https://arxiv.org/html/2609.27032#bib.bib7),[25](https://arxiv.org/html/2609.27032#bib.bib10)\]: a paragraph is drawn without replacement with probability proportional toexp\(si\)\\exp\(s\_\{i\}\), so an ordered selection ofKKparagraphs has likelihood
P\(π\)=∏t=1Kexp\(sπt\)∑j∈ℛtexp\(sj\),P\(\\pi\)=\\prod\_\{t=1\}^\{K\}\\frac\{\\exp\\\!\\bigl\(s\_\{\\pi\_\{t\}\}\\bigr\)\}\{\\sum\_\{j\\in\\mathcal\{R\}\_\{t\}\}\\exp\\\!\\bigl\(s\_\{j\}\\bigr\)\},\(3\)withℛt=D∖St−1\\mathcal\{R\}\_\{t\}=D\\setminus S\_\{t\-1\}as above, so thatℛ1=D\\mathcal\{R\}\_\{1\}=Dand each factor renormalizes over the paragraphs not yet taken\. Taking the oracle sequence of Section[4\.2](https://arxiv.org/html/2609.27032#S4.SS2)as the target, we minimize the Plackett\-Luce negative log\-likelihood of that ordering \(a listwise ranking loss, equivalent to ListMLE\[[26](https://arxiv.org/html/2609.27032#bib.bib32)\],
ℒPL=−∑t=1K\[sπt−log∑j∈ℛtexp\(sj\)\],\\mathcal\{L\}\_\{\\mathrm\{PL\}\}=\-\\sum\_\{t=1\}^\{K\}\\Bigl\[\\,s\_\{\\pi\_\{t\}\}\-\\log\\\!\\\!\\sum\_\{j\\in\\mathcal\{R\}\_\{t\}\}\\exp\(s\_\{j\}\)\\Bigr\],\(4\)which we evaluate with a vectorized reverse log\-cumulative\-sum\-exponential \(log\-cumsum\-exp\) rather than a per\-step loop\. The model therefore scores allNNparagraphs while the target constrains the firstKK, leaving the ordering of the unselected remainder free\.
Because the listwise objective is invariant to a constant shift of all scores, it fixes only their relative order\. We therefore add a binary cross\-entropy termℒBCE\\mathcal\{L\}\_\{\\mathrm\{BCE\}\}over allNNparagraphs, with oracle membership as the label, which gives the readout an absolute reference point:
ℒ=ℒPL\+λℒBCE,λ=0\.3\.\\mathcal\{L\}=\\mathcal\{L\}\_\{\\mathrm\{PL\}\}\+\\lambda\\,\\mathcal\{L\}\_\{\\mathrm\{BCE\}\},\\qquad\\lambda=0\.3\.\(5\)
We setλ=0\.3\\lambda=0\.3without tuning\. The role ofℒBCE\\mathcal\{L\}\_\{\\mathrm\{BCE\}\}is only to resolve the shift\-invariance ofℒPL\\mathcal\{L\}\_\{\\mathrm\{PL\}\}by anchoring the absolute scale of the scores, which any moderate positive weight accomplishes; a value well below one keeps the listwise ranking term dominant, so membership classification refines rather than overrides the learned ordering\. Because every model variant is trained with the sameλ\\lambda, our comparisons are unaffected by its precise value\.
### 5\.2Stage 2: Reinforcement Fine\-Tuning
Stage 1 optimizes agreement with an oracle sequence, a proxy for the quantity we actually care about: the quality of the extract the selector finally emits\. Stage 2 optimizes that quantity directly with REINFORCE Leave\-One\-Out\[[27](https://arxiv.org/html/2609.27032#bib.bib5),[28](https://arxiv.org/html/2609.27032#bib.bib1)\], testing whether the proxy leaves headroom rather than serving as a component the model depends on\.
Letπ∼Pθ\\pi\\sim P\_\{\\theta\}denote an extract sampled from the Plackett–Luce policy of Eq\.[3](https://arxiv.org/html/2609.27032#S5.E3): paragraphs are drawn sequentially without replacement until the word budget is met, then assembled in document order and truncated to the budget, exactly matching the evaluation\-time selector\. We maximize the expected reward of the sampled extract,
𝒥\(θ\)=𝔼π∼Pθ\[r\(π\)\],\\mathcal\{J\}\(\\theta\)=\\mathbb\{E\}\_\{\\pi\\sim P\_\{\\theta\}\}\\\!\\bigl\[\\,r\(\\pi\)\\,\\bigr\],\(6\)wherer\(π\)∈\[0,1\]r\(\\pi\)\\in\[0,1\]scores the extract against the referenceRRas the mean of threeRougeF1F\_\{1\}variants,
r\(π\)=13∑v∈\{1,2,Lsum\}R\-vF1\(π,R\)\.r\(\\pi\)=\\tfrac\{1\}\{3\}\\\!\\\!\\sum\_\{v\\,\\in\\,\\\{1,\\,2,\\,\\mathrm\{Lsum\}\\\}\}\\\!\\\!\\text\{R\-\}v\_\{F\_\{1\}\}\(\\pi,R\)\.\(7\)
#### Language\-normalized reward\.
Raw reward magnitudes differ systematically across languages, so before forming the gradient we standardizerrper languageℓ\\ellusing running estimatesμℓ,σℓ\\mu\_\{\\ell\},\\sigma\_\{\\ell\}maintained online,
r~\(π\)=r\(π\)−μℓσℓ,\\tilde\{r\}\(\\pi\)=\\frac\{r\(\\pi\)\-\\mu\_\{\\ell\}\}\{\\sigma\_\{\\ell\}\},\(8\)which prevents high\-scoring languages from dominating the update\.
#### Leave\-one\-out advantage\.
For each document we drawG=4G\{=\}4rolloutsπ1,…,πG\\pi\_\{1\},\\dots,\\pi\_\{G\}and use the mean normalized reward of the remainingG−1G\{\-\}1as each rollout’s baseline, giving a critic\-free, unbiased advantageAgA\_\{g\}where,
Ag=r~\(πg\)−1G−1∑g′≠gr~\(πg′\)\.A\_\{g\}=\\tilde\{r\}\(\\pi\_\{g\}\)\-\\frac\{1\}\{G\-1\}\\sum\_\{g^\{\\prime\}\\neq g\}\\tilde\{r\}\(\\pi\_\{g^\{\\prime\}\}\)\.\(9\)The resulting policy\-gradient estimate is
∇θ𝒥≈∑g=1GAg∇θlogPθ\(πg\),\\nabla\_\{\\theta\}\\mathcal\{J\}\\;\\approx\\;\\sum\_\{g=1\}^\{G\}A\_\{g\}\\,\\nabla\_\{\\theta\}\\log P\_\{\\theta\}\(\\pi\_\{g\}\),\(10\)wherelogPθ\(πg\)\\log P\_\{\\theta\}\(\\pi\_\{g\}\)is the Plackett\-Luce log\-likelihood of Eq\. \(1\), differentiated through the readout scores\.
## 6Experiments
Experiments run on a single workstation with an AMD Ryzen Threadripper 5975WX CPU and three NVIDIA RTX A6000 GPUs \(49 GB VRAM each\), with all models implemented in PyTorch\[[29](https://arxiv.org/html/2609.27032#bib.bib15)\]\. Frozen encoder outputs are cached once, after which the warm\-start, RL, transfer, and ablation runs proceed without the encoder in the loop\. RLOO is the most expensive stage at≈\\approx8\.4 GPU\-hours; warm\-start converges within roughly one epoch\. All runs use seed 0\.
## 7Results and Discussion
### 7\.1Multilingual Analysis
Table 2:Per\-language results on EUR\-Lex\-Sum \(R1: ROUGE\-1, R2: ROUGE\-2, RL: ROUGE\-L, BS: BERTScore\-F1\)\. Best per metric per language inbold\. LexLattice is our 2D\-NCA over the semantic lattice; 1D\-NCA is the capacity\-matched control \(Section[4\.3](https://arxiv.org/html/2609.27032#S4.SS3.SSS0.Px3)\)\.†[Aumiller et al\. \[9\]](https://arxiv.org/html/2609.27032#bib.bib2);‡LexT5\[[17](https://arxiv.org/html/2609.27032#bib.bib22)\];§SLM benchmark\[[22](https://arxiv.org/html/2609.27032#bib.bib20)\]\.Table[2](https://arxiv.org/html/2609.27032#S7.T2)reports per\-language results across six languages, and three patterns hold consistently\.First, the structural axis pays off: LexLattice improves over the 1D\-NCA on every metric in every language, by1\.71\.7\-2\.02\.0ROUGE\-1 and up to2\.72\.7ROUGE\-2, confirming that the section×\\timesparagraph geometry drives the gain\.Second, LexLattice\+RLOO attains the best ROUGE\-1, ROUGE\-2, and ROUGE\-L in all six languages, ahead of instruction\-tuned baselines as large as 8B parameters despite carrying only1\.81\.8M trainable parameters over a frozen encoder\. The margin over the strongest prior system reaches7\.77\.7ROUGE\-1 in English\.Third, RLOO adds marginal gains to ROUGE\-1 or ROUGE\-2 metrics\. The warm\-start already captures most of that quality, but yields a consistent ROUGE\-L gain that secures the ROUGE\-L lead: in Spanish, for instance, only after RLOO does our model surpass the strongest baseline on this metric\. The BERTScore picture is more balanced: LexLattice is strongest in English, French, and Italian, while a larger Qwen baseline edges it in German, Spanish, and Portuguese, in most cases by under a point\.
### 7\.2Cross\-Lingual Transfer
If the consolidator operates on the encoder’s shared semantic geometry rather than on language\-specific surface cues, a model trained on one set of languages should transfer to unseen ones with little loss\. We test this in two settings\. In both, the language embedding is omitted, since it is undefined for a language never seen in training, and we reportretention: the ratio of a transferred model’s ROUGE\-2 to that of a model jointly supervised on the target language, where a value of11denotes lossless transfer\.
#### Held\-out languages\.
We split the languages at the median training\-document count into HIGH \(12 languages, 12,473 documents\) and LOW \(12 languages, including Irish with only 16\), train the 2D\-NCA on HIGH alone, and evaluate zero\-shot on LOW\. Table[3](https://arxiv.org/html/2609.27032#S7.T3)reports the result\. Transfer is near\-lossless throughout: retention averages0\.9930\.993and never falls below0\.9740\.974, and for five of the twelve languages the zero\-shot model matches or exceeds its jointly\-supervised counterpart\. A model which has never seen a language can perform nearly equal to the one trained on it indicates that the consolidation rule the automata learns is largely language\-independent\.
Table 3:Zero\-shot transfer on the 12 held\-out \(LOW\) languages, ROUGE\-2\.Transfer: 2D\-NCA trained on HIGH only;Joint: trained on all languages\.Retention= Transfer / Joint;≥1\{\\geq\}1indicates no loss\. Sorted by retention\.
#### English→\{\\rightarrow\}Spanish\.
As a stricter case mirroring the cross\-lingual baseline of the original EUR\-Lex\-Sum study, we train LexLattice on English alone, select on English validation, and evaluate zero\-shot on Spanish, so the target language is never seen before test time \(Table[4](https://arxiv.org/html/2609.27032#S7.T4)\)\. Among methods with access only to the English source, zero\-shot LexLattice improves over both the translate\-then\-summarize LED pipeline and English LexRank by a wide margin on every metric\. It also comes within1\.81\.8ROUGE\-1 of the Spanish\-supervised model, reinforcing the finding above: transferring from English costs little even when no target\-language signal is available at any stage\.
Table 4:Cross\-lingual English→\{\\rightarrow\}Spanish results\. Upper block: English\-source\-only methods \(bold= best\)\. Lower block: Spanish\-side references \(italic\)\.†[Aumiller et al\. \[9\]](https://arxiv.org/html/2609.27032#bib.bib2)\.
### 7\.3LLM\-as\-a\-Judge Evaluation
Table 5:LLM\-as\-a\-judge evaluation on the full test set \(all 24 languages, candidates blinded\), macro\-averaged over languages\.*Coverage*: fraction of reference LDPs recovered;*relevance*: fraction of candidate content judged pertinent; LDP\-F1F\_\{1\}: their harmonic mean\.†/‡: improvement over 1D\-NCA / LexLattice \(paired Wilcoxon, per\-document,p<0\.05p<0\.05\)\.ROUGE under\-credits heavily inflected languages and behaves inconsistently across scripts\. We therefore complement it with a reference\-based adaptation of LeMAJ\[[30](https://arxiv.org/html/2609.27032#bib.bib17)\], which decomposes each reference intoLegal Data Points\(LDPs\): atomic, self\-contained statements of legal information\. A single fixed judge \(Claude Sonnet 4\.5,[31](https://arxiv.org/html/2609.27032#bib.bib18)\) scores each candidate pointwise and blind, withcoveragemarking every reference LDP as covered, partial, or missing, andrelevancemarking every candidate unit as relevant, marginal, or irrelevant \(partial credit0\.50\.5\); LDP\-F1F\_\{1\}is their harmonic mean\. Because the systems are extractive and therefore near\-perfectly grounded in the source, the protocol isolates*selection*quality rather than factual consistency\. The judge\-based results in Table[5](https://arxiv.org/html/2609.27032#S7.T5)demonstrate the ROUGE findings under an evaluation largely insensitive to inflection and script\. The structural axis accounts for most of the improvement: LexLattice improves LDP\-F1F\_\{1\}over the 1D control by7\.87\.8points \(\+17%\+17\\%\), the majority of it attributable to coverage \(\+7\.1\+7\.1\), which suggests that structural adjacency helps the model recover salient content that a flat pass over paragraphs overlooks\. RLOO contributes a further1\.31\.3points, concentrated in relevance, and all improvements are significant under a per\-document paired Wilcoxon test\.
### 7\.4Coverage Across the Remaining Languages
The six languages of Section[7\.1](https://arxiv.org/html/2609.27032#S7.SS1)are the only ones with instruction\-tuned LLM or SLM variants; for the remaining 18, the zero\-shot multilingual LexRank of[Aumiller et al\. \[9\]](https://arxiv.org/html/2609.27032#bib.bib2)is the sole published baseline\. Figure[3](https://arxiv.org/html/2609.27032#S7.F3)shows ROUGE\-1 for these languages, and the gains are uniform: LexLattice model with RLOO improves on LexRank in every one, by19\.019\.0points on average and by at least11\.811\.8, spanning scripts from Latin to Cyrillic \(Bulgarian\) and Greek\. The 1D→\\rightarrow2D and RLOO trends observed on the high\-resource languages persist here as well, with the semantic lattice providing the bulk of the improvement over the 1D control; full per\-system, per\-metric numbers appear in Section[7\.5](https://arxiv.org/html/2609.27032#S7.SS5)\.
30405060ROUGE\-1HungarianIrishCzechMalteseSwedishDutchLatvianSlovakRomanianDanishPolishSloveneCroatianLithuanianEstonianFinnishGreekBulgarianLexRankLexLattice\+RLOO \(ours\)Figure 3:ROUGE\-1 on the remaining 18 EUR\-Lex\-Sum languages \(jointly trained, sorted by our score\)\. LexRank is the only external baseline available for these languages; our final model improves on it by19\.019\.0points on average \(11\.811\.8–24\.624\.6\), in every language\. Full per\-system, per\-metric results are in Appendix[6](https://arxiv.org/html/2609.27032#S7.T6)\.
### 7\.5Complete Results Across All Languages
Table[6](https://arxiv.org/html/2609.27032#S7.T6)reports the complete per\-system, per\-metric scores for the 18 EUR\-Lex\-Sum languages not shown in Table[2](https://arxiv.org/html/2609.27032#S7.T2), jointly trained on all 24 languages\. LexLattice or its RLOO variant attains the best ROUGE\-1 and ROUGE\-2 in every language, with the 1D control competitive only in Bulgarian\.
Table 6:Full per\-language, per\-system results on the 18 EUR\-Lex\-Sum languages not shown in Table[2](https://arxiv.org/html/2609.27032#S7.T2)\(jointly trained on all 24 languages\)\. Best value per metric per language inbold\. LexLattice is our 2D\-NCA;\+RLOOadds the reinforcement stage; 1D\-NCA is the capacity\-matched control\.
### 7\.6Visualizing NCA Salience Evolution
Figure[4](https://arxiv.org/html/2609.27032#S7.F4)illustrates the consolidation dynamics of LexLattice on a representative English document\. Each panel shows the semantic lattice \(rows: preamble, recitals, articles, annexes; columns: paragraph position within the section; grey cells are padding\), coloured by the readout salience of each paragraph, min–max normalized over the whole trajectory\. The leftmost panel is the encoder input \(t=0t\{=\}0\); subsequent panels show the state after each of theT=8T\{=\}8local NCA update steps\. Salience is diffuse att=0t\{=\}0and progressively sharpens as information propagates through the3×33\{\\times\}3neighbourhood, concentrating on a sparse set of cells\. Stars in the final panel mark the paragraphs ranked into the summary under the reference\-length word budget; these cluster in contiguous lattice regions rather than isolated cells, evidence that the learned dynamics exploit the document’s hierarchical structure rather than scoring paragraphs independently\.
Figure 4:Per\-cell readout salience on the semantic lattice att=0t\{=\}0and after each of the 8 NCA update steps \(single sequential colour ramp; grey = padding\)\.⋆\\starmarks paragraphs selected into the final extract\.
## 8Conclusion and Future Work
We introduced LexLattice, an extractive summarizer that casts a legal act’s hierarchy as a 2D semantic lattice and consolidates over it with a masked 2D NCA before selection\. Concentrating all trainable capacity in a1\.81\.8M\-parameter consolidator over a frozen multilingual encoder, it attains state\-of\-the\-art ROUGE scores across all 24 languages of EUR\-Lex\-Sum, surpassing far larger instruction\-tuned baselines, with a uniform per\-language evaluation and a broad cross\-lingual transfer study\. The learned dynamics sharpen diffuse salience onto contiguous, structurally coherent regions of the lattice \(Section[7\.6](https://arxiv.org/html/2609.27032#S7.SS6)\), indicating that explicit consolidation over document structure is a compact substitute for scale\. Proposed extensions of our work include: applying the lattice to other long, hierarchically organized corpora, both legal\[[32](https://arxiv.org/html/2609.27032#bib.bib12)\]and beyond\[[33](https://arxiv.org/html/2609.27032#bib.bib33)\]; and coupling the traceable extract with an abstractive generator to add fluency while preserving auditability\.
## 9Limitations
LexLattice relies on a rule\-based parser to recover document hierarchy, so it presupposes structurally marked text and would degrade on documents without explicit sectioning\. Being extractive, it is bounded by the source and a fixed length budget, which caps coverage and forgoes the fluency of abstraction\. Finally, for most of the languages the only published baseline is extractive LexRank, so our comparisons rest on limited external baselines\.
## Ethical Statement
This work uses EUR\-Lex\-Sum, a publicly released corpus of EU legal acts and their official summaries; it contains no personal or private data, and we use it in accordance with its license\. Our aim is to broaden access to law across the official EU languages, which is of particular benefit to speakers of lower\-resourced languages underserved by English\-centric systems\. We caution, however, that automatic summaries of legal text are not legal advice and must not substitute for the authoritative instruments they condense\. Although our extractive design keeps summaries traceable to the source and limits fabricated content, it guarantees neither completeness nor faithfulness, and quality is weaker for the most resource\-scarce languages, where reliance should be correspondingly cautious\. Because the method trains only a small consolidator over a frozen encoder, it also carries a modest computational and energy cost relative to large generative summarizers\.
## Acknowledgments
This research is supported by NSERC Discovery grant \#194376\.
## References
- \[1\]\(2023\)Improving abstractive summarization of legal rulings through textual entailment\.Artificial Intelligence and Law31\(1\),pp\. 91–113\.External Links:ISSN 1572\-8382,[Document](https://dx.doi.org/10.1007/s10506-021-09305-4),[Link](https://doi.org/10.1007/s10506-021-09305-4)Cited by:[§1](https://arxiv.org/html/2609.27032#S1.p1.1)\.
- \[2\]R\. Nallapati, F\. Zhai, and B\. Zhou\(2017\)Summarunner: a recurrent neural network based sequence model for extractive summarization of documents\.InProceedings of the AAAI conference on artificial intelligence,Vol\.31\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/10958)Cited by:[§1](https://arxiv.org/html/2609.27032#S1.p1.1),[§4\.2](https://arxiv.org/html/2609.27032#S4.SS2.p1.1)\.
- \[3\]Y\. Liu and M\. Lapata\(2019\)Text summarization with pretrained encoders\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),K\. Inui, J\. Jiang, V\. Ng, and X\. Wan \(Eds\.\),Hong Kong, China,pp\. 3730–3740\.External Links:[Link](https://aclanthology.org/D19-1387/),[Document](https://dx.doi.org/10.18653/v1/D19-1387)Cited by:[§1](https://arxiv.org/html/2609.27032#S1.p1.1),[§4\.2](https://arxiv.org/html/2609.27032#S4.SS2.p1.1)\.
- \[4\]A\. Cohan, F\. Dernoncourt, D\. S\. Kim, T\. Bui, S\. Kim, W\. Chang, and N\. Goharian\(2018\)A discourse\-aware attention model for abstractive summarization of long documents\.InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 \(Short Papers\),M\. Walker, H\. Ji, and A\. Stent \(Eds\.\),New Orleans, Louisiana,pp\. 615–621\.External Links:[Link](https://aclanthology.org/N18-2097/),[Document](https://dx.doi.org/10.18653/v1/N18-2097)Cited by:[§1](https://arxiv.org/html/2609.27032#S1.p1.1)\.
- \[5\]H\. Y\. Koh, J\. Ju, M\. Liu, and S\. Pan\(2022\)An empirical survey on long document summarization: datasets, models, and metrics\.ACM Comput\. Surv\.55\(8\)\.External Links:ISSN 0360\-0300,[Link](https://doi.org/10.1145/3545176),[Document](https://dx.doi.org/10.1145/3545176)Cited by:[§1](https://arxiv.org/html/2609.27032#S1.p1.1)\.
- \[6\]Q\. Ruan, M\. Ostendorff, and G\. Rehm\(2022\)HiStruct\+: improving extractive text summarization with hierarchical structure information\.InFindings of the Association for Computational Linguistics: ACL 2022,S\. Muresan, P\. Nakov, and A\. Villavicencio \(Eds\.\),Dublin, Ireland,pp\. 1292–1308\.External Links:[Link](https://aclanthology.org/2022.findings-acl.102/),[Document](https://dx.doi.org/10.18653/v1/2022.findings-acl.102)Cited by:[§1](https://arxiv.org/html/2609.27032#S1.p1.1)\.
- \[7\]L\. Xue, N\. Constant, A\. Roberts, M\. Kale, R\. Al\-Rfou, A\. Siddhant, A\. Barua, and C\. Raffel\(2021\)MT5: a massively multilingual pre\-trained text\-to\-text transformer\.InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,K\. Toutanova, A\. Rumshisky, L\. Zettlemoyer, D\. Hakkani\-Tur, I\. Beltagy, S\. Bethard, R\. Cotterell, T\. Chakraborty, and Y\. Zhou \(Eds\.\),Online,pp\. 483–498\.External Links:[Link](https://aclanthology.org/2021.naacl-main.41/),[Document](https://dx.doi.org/10.18653/v1/2021.naacl-main.41)Cited by:[§1](https://arxiv.org/html/2609.27032#S1.p2.1),[§3\.1](https://arxiv.org/html/2609.27032#S3.SS1.SSS0.Px1.p1.1)\.
- \[8\]A\. Mordvintsev, E\. Randazzo, E\. Niklasson, and M\. Levin\(2020\)Growing neural cellular automata\.Distill\.Note:https://distill\.pub/2020/growing\-caExternal Links:[Document](https://dx.doi.org/10.23915/distill.00023)Cited by:[§1](https://arxiv.org/html/2609.27032#S1.p2.1)\.
- \[9\]D\. Aumiller, A\. Chouhan, and M\. Gertz\(2022\)EUR\-lex\-sum: a multi\- and cross\-lingual dataset for long\-form summarization in the legal domain\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,Y\. Goldberg, Z\. Kozareva, and Y\. Zhang \(Eds\.\),Abu Dhabi, United Arab Emirates,pp\. 7626–7639\.External Links:[Link](https://aclanthology.org/2022.emnlp-main.519/),[Document](https://dx.doi.org/10.18653/v1/2022.emnlp-main.519)Cited by:[§2](https://arxiv.org/html/2609.27032#S2.p1.1),[§3](https://arxiv.org/html/2609.27032#S3.p1.1),[§7\.4](https://arxiv.org/html/2609.27032#S7.SS4.p1.1),[Table 2](https://arxiv.org/html/2609.27032#S7.T2),[Table 4](https://arxiv.org/html/2609.27032#S7.T4)\.
- \[10\]G\. Erkan and D\. R\. Radev\(2004\)LexRank: graph\-based lexical centrality as salience in text summarization\.Journal of Artificial Intelligence Research22,pp\. 457–479\.External Links:ISSN 1076\-9757,[Link](http://dx.doi.org/10.1613/jair.1523),[Document](https://dx.doi.org/10.1613/jair.1523)Cited by:[§2](https://arxiv.org/html/2609.27032#S2.p1.1)\.
- \[11\]J\. Tiedemann and S\. Thottingal\(2020\)OPUS\-MT – building open translation services for the world\.InProceedings of the 22nd Annual Conference of the European Association for Machine Translation,A\. Martins, H\. Moniz, S\. Fumega, B\. Martins, F\. Batista, L\. Coheur, C\. Parra, I\. Trancoso, M\. Turchi, A\. Bisazza, J\. Moorkens, A\. Guerberof, M\. Nurminen, L\. Marg, and M\. L\. Forcada \(Eds\.\),Lisboa, Portugal,pp\. 479–480\.External Links:[Link](https://aclanthology.org/2020.eamt-1.61/)Cited by:[§2](https://arxiv.org/html/2609.27032#S2.p1.1)\.
- \[12\]I\. Beltagy, M\. E\. Peters, and A\. Cohan\(2020\)Longformer: the long\-document transformer\.External Links:2004\.05150,[Link](https://arxiv.org/abs/2004.05150)Cited by:[§2](https://arxiv.org/html/2609.27032#S2.p1.1)\.
- \[13\]M\. Sie, R\. Beek, M\. Bots, S\. Brinkkemper, and A\. Gatt\(2024\)Summarizing long regulatory documents with a multi\-step pipeline\.InProceedings of the Natural Legal Language Processing Workshop 2024,N\. Aletras, I\. Chalkidis, L\. Barrett, C\. Goanță, D\. Preoțiuc\-Pietro, and G\. Spanakis \(Eds\.\),Miami, FL, USA,pp\. 18–32\.External Links:[Link](https://aclanthology.org/2024.nllp-1.2/),[Document](https://dx.doi.org/10.18653/v1/2024.nllp-1.2)Cited by:[§2](https://arxiv.org/html/2609.27032#S2.p1.1)\.
- \[14\]M\. Guo, J\. Ainslie, D\. Uthus, S\. Ontañón, J\. Ni, Y\. Sung, and Y\. Yang\(2022\)LongT5: Efficient text\-to\-text transformer for long sequences\.InFindings of the Association for Computational Linguistics: NAACL 2022,M\. Carpuat, M\. de Marneffe, and I\. V\. Meza Ruiz \(Eds\.\),Seattle, United States,pp\. 724–736\.External Links:[Link](https://aclanthology.org/2022.findings-naacl.55/),[Document](https://dx.doi.org/10.18653/v1/2022.findings-naacl.55)Cited by:[§2](https://arxiv.org/html/2609.27032#S2.p1.1)\.
- \[15\]J\. Zhang, Y\. Zhao, M\. Saleh, and P\. Liu\(2020\)PEGASUS: pre\-training with extracted gap\-sentences for abstractive summarization\.InProceedings of the 37th International Conference on Machine Learning,H\. D\. III and A\. Singh \(Eds\.\),Proceedings of Machine Learning Research, Vol\.119,pp\. 11328–11339\.External Links:[Link](https://proceedings.mlr.press/v119/zhang20ae.html)Cited by:[§2](https://arxiv.org/html/2609.27032#S2.p1.1)\.
- \[16\]A\. a\. M\. Llama Team\(2024\)The llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[§2](https://arxiv.org/html/2609.27032#S2.p1.1)\.
- \[17\]S\. T\.Y\.S\.S, C\. Weiss, and M\. Grabmair\(2024\)LexSumm and LexT5: benchmarking and modeling legal summarization tasks in English\.InProceedings of the Natural Legal Language Processing Workshop 2024,N\. Aletras, I\. Chalkidis, L\. Barrett, C\. Goanță, D\. Preoțiuc\-Pietro, and G\. Spanakis \(Eds\.\),Miami, FL, USA,pp\. 381–403\.External Links:[Link](https://aclanthology.org/2024.nllp-1.35/),[Document](https://dx.doi.org/10.18653/v1/2024.nllp-1.35)Cited by:[§2](https://arxiv.org/html/2609.27032#S2.p1.1),[Table 2](https://arxiv.org/html/2609.27032#S7.T2)\.
- \[18\]W\. Xiao, I\. Beltagy, G\. Carenini, and A\. Cohan\(2022\)PRIMERA: pyramid\-based masked sentence pre\-training for multi\-document summarization\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),S\. Muresan, P\. Nakov, and A\. Villavicencio \(Eds\.\),Dublin, Ireland,pp\. 5245–5263\.External Links:[Link](https://aclanthology.org/2022.acl-long.360/),[Document](https://dx.doi.org/10.18653/v1/2022.acl-long.360)Cited by:[§2](https://arxiv.org/html/2609.27032#S2.p1.1)\.
- \[19\]N\. Bendahman, K\. Pinel\-Sauvagnat, G\. Hubert, and M\. B\. Billami\(2025\)Not all hallucinations are good to throw away when it comes to legal abstractive summarization\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),L\. Chiruzzo, A\. Ritter, and L\. Wang \(Eds\.\),Albuquerque, New Mexico,pp\. 5331–5344\.External Links:[Link](https://aclanthology.org/2025.naacl-long.275/),[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.275),ISBN 979\-8\-89176\-189\-6Cited by:[§2](https://arxiv.org/html/2609.27032#S2.p1.1)\.
- \[20\]V\. Chheda, A\. Ghaisas, A\. Sankhe, and N\. Shekokar\(2025\)Extract\-explain\-abstract: a rhetorical role\-driven domain\-specific summarisation framework for Indian legal documents\.InProceedings of the Natural Legal Language Processing Workshop 2025,N\. Aletras, I\. Chalkidis, L\. Barrett, C\. Goanță, D\. Preoțiuc\-Pietro, and G\. Spanakis \(Eds\.\),Suzhou, China,pp\. 439–455\.External Links:[Link](https://aclanthology.org/2025.nllp-1.32/),[Document](https://dx.doi.org/10.18653/v1/2025.nllp-1.32),ISBN 979\-8\-89176\-338\-8Cited by:[§2](https://arxiv.org/html/2609.27032#S2.p1.1)\.
- \[21\]A\. Bailly, A\. Saubin, G\. Kocevar, and J\. Bodin\(2025\)Divide and summarize: improve slm text summarization\.Frontiers in Artificial IntelligenceVolume 8 \- 2025\.External Links:[Link](https://www.frontiersin.org/journals/artificial-intelligence/articles/10.3389/frai.2025.1604034),[Document](https://dx.doi.org/10.3389/frai.2025.1604034),ISSN 2624\-8212Cited by:[§2](https://arxiv.org/html/2609.27032#S2.p1.1)\.
- \[22\]M\. Medina\-Ramírez, C\. Estupiñán\-Ojeda, V\. Torres\-Rodríguez, E\. Sánchez\-Nielsen, C\. Guerra\-Artal, and M\. Hernández\-Tejera\(2026\)Small language models for legislative summarization: an empirical evaluation of performance and suitability\.IEEE Access14\(\),pp\. 55918–55942\.External Links:[Document](https://dx.doi.org/10.1109/ACCESS.2026.3679718)Cited by:[§2](https://arxiv.org/html/2609.27032#S2.p1.1),[Table 2](https://arxiv.org/html/2609.27032#S7.T2)\.
- \[23\]T\. König, B\. Luetgert, and T\. Dannwolf\(2006\)Quantifying european legislative research: using celex and prelex in eu legislative studies\.European Union Politics7\(4\),pp\. 553–574\.External Links:[Document](https://dx.doi.org/10.1177/1465116506069444),[Link](https://doi.org/10.1177/1465116506069444),https://doi\.org/10\.1177/1465116506069444Cited by:[§3](https://arxiv.org/html/2609.27032#S3.p1.1)\.
- \[24\]R\.D\. Luce\(1959\)Individual choice behavior: a theoretical analysis\.Wiley\.External Links:LCCN 59009346,[Link](https://books.google.ca/books?id=a80DAQAAIAAJ)Cited by:[§5\.1](https://arxiv.org/html/2609.27032#S5.SS1.p1.1)\.
- \[25\]R\. L\. Plackett\(1975\)The analysis of permutations\.Journal of the Royal Statistical Society Series C: Applied Statistics24\(2\),pp\. 193–202\.External Links:ISSN 0035\-9254,[Document](https://dx.doi.org/10.2307/2346567),[Link](https://doi.org/10.2307/2346567),https://academic\.oup\.com/jrsssc/article\-pdf/24/2/193/48619663/jrsssc\_24\_2\_193\.pdfCited by:[§5\.1](https://arxiv.org/html/2609.27032#S5.SS1.p1.1)\.
- \[26\]F\. Xia, T\. Liu, J\. Wang, W\. Zhang, and H\. Li\(2008\)Listwise approach to learning to rank: theory and algorithm\.InProceedings of the 25th International Conference on Machine Learning,ICML ’08,New York, NY, USA,pp\. 1192–1199\.External Links:ISBN 9781605582054,[Link](https://doi.org/10.1145/1390156.1390306),[Document](https://dx.doi.org/10.1145/1390156.1390306)Cited by:[§5\.1](https://arxiv.org/html/2609.27032#S5.SS1.p1.2)\.
- \[27\]W\. Kool, H\. van Hoof, and M\. Welling\(2019\)Buy 4 REINFORCE samples, get a baseline for free\!\.External Links:[Link](https://openreview.net/forum?id=r1lgTGL5DE)Cited by:[§5\.2](https://arxiv.org/html/2609.27032#S5.SS2.p1.1)\.
- \[28\]A\. Ahmadian, C\. Cremer, M\. Gallé, M\. Fadaee, J\. Kreutzer, O\. Pietquin, A\. Üstün, and S\. Hooker\(2024\)Back to basics: revisiting REINFORCE\-style optimization for learning from human feedback in LLMs\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 12248–12267\.External Links:[Link](https://aclanthology.org/2024.acl-long.662/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.662)Cited by:[§5\.2](https://arxiv.org/html/2609.27032#S5.SS2.p1.1)\.
- \[29\]A\. Paszke, S\. Gross, F\. Massa, A\. Lerer, J\. Bradbury, G\. Chanan, T\. Killeen, Z\. Lin, N\. Gimelshein, L\. Antiga, A\. Desmaison, A\. Kopf, E\. Yang, Z\. DeVito, M\. Raison, A\. Tejani, S\. Chilamkurthy, B\. Steiner, L\. Fang, J\. Bai, and S\. Chintala\(2019\)PyTorch: an imperative style, high\-performance deep learning library\.InAdvances in Neural Information Processing Systems,H\. Wallach, H\. Larochelle, A\. Beygelzimer, F\. d'Alché\-Buc, E\. Fox, and R\. Garnett \(Eds\.\),Vol\.32,pp\.\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2019/file/bdbca288fee7f92f2bfa9f7012727740-Paper.pdf)Cited by:[§6](https://arxiv.org/html/2609.27032#S6.p1.1)\.
- \[30\]J\. Enguehard, M\. Van Ermengem, K\. Atkinson, S\. Cha, A\. G\. Chowdhury, P\. K\. Ramaswamy, J\. Roghair, H\. R\. Marlowe, C\. S\. Negreanu, K\. Boxall, and D\. Mincu\(2025\)LeMAJ \(legal LLM\-as\-a\-judge\): bridging legal reasoning and LLM evaluation\.InProceedings of the Natural Legal Language Processing Workshop 2025,N\. Aletras, I\. Chalkidis, L\. Barrett, C\. Goanță, D\. Preoțiuc\-Pietro, and G\. Spanakis \(Eds\.\),Suzhou, China,pp\. 318–337\.External Links:[Link](https://aclanthology.org/2025.nllp-1.23/),[Document](https://dx.doi.org/10.18653/v1/2025.nllp-1.23),ISBN 979\-8\-89176\-338\-8Cited by:[§7\.3](https://arxiv.org/html/2609.27032#S7.SS3.p1.1)\.
- \[31\]Anthropic\(2025\)Claude sonnet 4\.5\.Note:[https://www\.anthropic\.com/claude\-sonnet\-4\-5\-system\-card](https://www.anthropic.com/claude-sonnet-4-5-system-card)Model card and system card; accessed 23 July 2026Cited by:[§7\.3](https://arxiv.org/html/2609.27032#S7.SS3.p1.1)\.
- \[32\]Z\. Shen, K\. Lo, L\. Yu, N\. Dahlberg, M\. Schlanger, and D\. Downey\(2022\)Multi\-lexsum: real\-world summaries of civil rights lawsuits at multiple granularities\.InAdvances in Neural Information Processing Systems,S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(Eds\.\),Vol\.35,pp\. 13158–13173\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/552ef803bef9368c29e53c167de34b55-Paper-Datasets_and_Benchmarks.pdf)Cited by:[§8](https://arxiv.org/html/2609.27032#S8.p1.1)\.
- \[33\]W\. Kryscinski, N\. Rajani, D\. Agarwal, C\. Xiong, and D\. Radev\(2022\)BOOKSUM: a collection of datasets for long\-form narrative summarization\.InFindings of the Association for Computational Linguistics: EMNLP 2022,Y\. Goldberg, Z\. Kozareva, and Y\. Zhang \(Eds\.\),Abu Dhabi, United Arab Emirates,pp\. 6536–6558\.External Links:[Link](https://aclanthology.org/2022.findings-emnlp.488/),[Document](https://dx.doi.org/10.18653/v1/2022.findings-emnlp.488)Cited by:[§8](https://arxiv.org/html/2609.27032#S8.p1.1)\.Similar Articles
A Tree-of-Thoughts Inspired Hybrid Approach for Legal Case Judgement Summarization using LLMs
Proposes a tree-of-thoughts inspired extractive-abstractive approach for legal case judgement summarization using LLMs, with experiments on DeepSeek and LLama showing improved summaries over extractive or abstractive methods alone.
Large Language Models for Token-Efficient and Semantic-Preserving Opinion Summarization
This paper presents a framework for opinion summarization using LLMs that combines multidimensional classification and stratified sampling to reduce token usage while preserving semantic diversity and balance across viewpoints.
Good Summarization SLMs for < 2000 tokens
A novice asks for recommendations on small language models and prompting strategies to build an employee note summarization engine under 2000 tokens, after experiencing hallucinations with Qwen2.5-7B-Instruct.
A Hybrid Hierarchical 1D-CNN-BiLSTM Framework for Extractive Summarization of Biomedical and Clinical Text
The paper introduces a hybrid hierarchical 1D-CNN-BiLSTM framework for extractive summarization of biomedical and clinical text, designed to preserve factuality by selecting sentences directly from source documents rather than generating new text.
MIDAS: Multi-LLM Iterative Data-Adaptive Summarization
This paper proposes MIDAS, a multi-LLM framework for data-adaptive summarization that automates prompt optimization for domain-specific enterprise use cases, achieving strong improvements over prior methods on customer ticket summarization benchmarks.