Improving Cross-Lingual Transfer for Sequential Sentence Classification in Research Papers via Structural Similarity

arXiv cs.CL Papers

Summary

This paper explores improving cross-lingual transfer for sequential sentence classification in research papers by leveraging structural similarity, proposing methods that enhance performance in multilingual settings based on experiments with encoder-based and generative models.

arXiv:2609.19650v1 Announce Type: new Abstract: Sequential sentence classification (SSC) is an essential task for structuring scientific publications, and extending SSC research to languages other than English can improve accessibility to scientific knowledge in multilingual digital libraries. Cross-lingual transfer is a promising approach to address the scarcity of training data in non-English languages. Prior work on other natural language processing tasks has shown the benefits of capturing linguistic similarity between source and target languages. However, SSC inherently depends on patterns at the discourse level, such as label sequences and positional regularities, which appear consistently across languages regardless of linguistic differences. To examine the factors that determine transfer success in SSC, we constructed a multilingual SSC dataset covering 13 non-English languages collected from five academic databases. Our cross-lingual transfer experiments, using both encoder-based and generative models, show that linguistic proximity has no consistent predictive power for transfer performance, whereas structural similarity in rhetorical organization shows a weak but consistent positive correlation across models. After controlling for source-language performance, the similarity of label distributions is the most consistent predictor. Building on this finding, we propose a set of three methods that explicitly leverage structural information using generative models. In the in-domain evaluation, the best combination reaches parity with the strongest encoder baselines, and in transfer to languages unseen during training, it outperforms the strongest encoder baseline.
Original Article
View Cached Full Text

Cached at: 09/18/26, 08:59 AM

# Improving Cross-Lingual Transfer for Sequential Sentence Classification in Research Papers via Structural Similarity
Source: [https://arxiv.org/html/2609.19650](https://arxiv.org/html/2609.19650)
Conference:The 2026 ACM/IEEE Joint Conference on Digital Libraries; October 13–16, 2026; Frisco, TX, USAThe 2026 ACM/IEEE Joint Conference on Digital Libraries \(JCDL ’26\), October 13–16, 2026, Frisco, TX, USADOI:[10\.1145/3805696\.3846040](https://doi.org/10.1145/3805696.3846040)ISBN:979\-8\-4007\-2597\-5/2026/10CCS:Information systems Document representationCCS:Information systems Information extractionCCS:Information systems Digital libraries and archives© cc

###### Abstract\.

Sequential sentence classification \(SSC\) is an essential task for structuring scientific publications, and extending SSC research to languages other than English can improve accessibility to scientific knowledge in multilingual digital libraries\. Cross\-lingual transfer is a promising approach to address the scarcity of training data in non\-English languages\. Prior work on other natural language processing tasks has shown the benefits of capturing linguistic similarity between source and target languages\. However, SSC inherently depends on patterns at the discourse level, such as label sequences and positional regularities, which appear consistently across languages regardless of linguistic differences\. To examine the factors that determine transfer success in SSC, we constructed a multilingual SSC dataset covering 13 non\-English languages collected from five academic databases\. Our cross\-lingual transfer experiments, using both encoder\-based and generative models, show that linguistic proximity has no consistent predictive power for transfer performance, whereas structural similarity in rhetorical organization shows a weak but consistent positive correlation across models\. After controlling for source\-language performance, the similarity of label distributions is the most consistent predictor\. Building on this finding, we propose a set of three methods that explicitly leverage structural information using generative models\. In the in\-domain evaluation, the best combination reaches parity with the strongest encoder baselines, and in transfer to languages unseen during training, it outperforms the strongest encoder baseline\.

###### Keywords:

sequential sentence classification, cross\-lingual transfer, multilingual dataset, structural similarity, digital libraries, scholarly document processing

††cc\-license:by## 1\.Introduction

Sequential sentence classification \(SSC\) is the task of assigning each sentence in a scientific abstract to a rhetorical role such as background, objective, method, result, or conclusion\([Dernoncourt and Lee, 2017](https://arxiv.org/html/2609.19650#bib.bib1)\)\. Decomposing an abstract into these roles enables retrieval based on content that goes beyond keyword matching, for example finding papers with a similar method but a different background\. SSC thus serves as a foundational technology for downstream applications such as literature search, automatic summarization, and paper recommendation\. Recent advancements in SSC have leveraged hierarchical architectures built on Transformer models\([Devlin et al\., 2019](https://arxiv.org/html/2609.19650#bib.bib20);[Jin and Szolovits, 2018](https://arxiv.org/html/2609.19650#bib.bib2);[Cohan et al\., 2019](https://arxiv.org/html/2609.19650#bib.bib10);[Brack et al\., 2024](https://arxiv.org/html/2609.19650#bib.bib3)\)and large language models \(LLMs\)\([Lan et al\., 2024](https://arxiv.org/html/2609.19650#bib.bib4)\), reaching unprecedented performance levels\. However, these developments have been predominantly centered on English, creating a gap in the accessibility of non\-English academic research\.

Cross\-lingual transfer learning, facilitated by multilingual pre\-trained language models \(mPLMs\)\([Devlin et al\., 2019](https://arxiv.org/html/2609.19650#bib.bib20)\), has emerged as a key strategy to bridge this gap\. By transferring knowledge from a language with abundant labeled data such as English to a target language with few or no labeled abstracts, it removes the need to build a labeled dataset for each target language\. The success of this transfer has long been attributed to the linguistic proximity between the source and target languages\([Lin et al\., 2019](https://arxiv.org/html/2609.19650#bib.bib6);[Philippy et al\., 2023](https://arxiv.org/html/2609.19650#bib.bib5)\)\. However, recent studies show this proximity to be an inconsistent predictor\([Blaschke et al\., 2025](https://arxiv.org/html/2609.19650#bib.bib7);[De Souza et al\., 2024](https://arxiv.org/html/2609.19650#bib.bib34)\)and point instead to factors specific to the task\([Lin et al\., 2024](https://arxiv.org/html/2609.19650#bib.bib8);[Yun et al\., 2023](https://arxiv.org/html/2609.19650#bib.bib9)\)\. This motivates identifying, for SSC, what governs cross\-lingual transfer beyond linguistic distance\.

In this study, we focus on the intrinsic structural properties of SSC\. SSC assigns a label to each sentence, and these labels form sequences whose ordering and positions follow recurring patterns across abstracts; we refer to these patterns as an abstract’s rhetorical structure\. Part of this structure is shared across languages because academic abstracts follow standardized conventions typified by IMRaD \(Introduction, Methods, Results, and Discussion\)\([Sollaci and Pereira, 2004](https://arxiv.org/html/2609.19650#bib.bib41)\), so that Background tends to appear at the beginning and Result tends to follow Method\. Other parts vary by language: in our data, for instance, Russian and Chinese omit Background almost entirely and begin directly with Objective\. We quantify the degree to which two languages share this structure as their structural similarity, a measure that captures both the shared conventions and the departures specific to individual languages\. To test whether this property predicts transfer, we constructed a multilingual SSC dataset covering 13 non\-English languages with approximately 32,000 abstracts in these languages \(52,487 including the English portion of PubMed\-RCT\)\. Our empirical analysis revealed that structural similarity, defined by the alignment of label distributions, label transitions, and positional regularities, is a more consistent predictor of transfer performance than linguistic proximity\.

Building on these insights, we propose a set of three methods that explicitly leverage structural information to enhance cross\-lingual SSC: injecting structural knowledge into prompts, reranking candidate sequences with a trained verifier, and enforcing prediction consistency in zero\-shot settings\. With these methods, we show that structural information can be used not only to predict transfer but also to improve SSC directly\.

The main contributions of this study are as follows:

- •We introduce a multilingual SSC dataset covering 13 languages and demonstrate, through empirical analysis, that structural similarity correlates more consistently with cross\-lingual transfer performance than linguistic proximity\.
- •We propose a set of three methods that leverage structural information and improve macro F1 scores over existing multilingual baselines\.

## 2\.Related Work

### 2\.1\.Sequential Sentence Classification

Since the release of the PubMed\-RCT benchmark\([Dernoncourt and Lee, 2017](https://arxiv.org/html/2609.19650#bib.bib1)\), SSC models have evolved from hierarchical neural architectures\([Jin and Szolovits, 2018](https://arxiv.org/html/2609.19650#bib.bib2)\)to approaches based on Transformers\([Cohan et al\., 2019](https://arxiv.org/html/2609.19650#bib.bib10)\)\. To effectively capture sequential dependencies and rhetorical flow, modern architectures leverage hierarchical modeling to integrate both representations at the sentence level and context at the document level\([Brack et al\., 2024](https://arxiv.org/html/2609.19650#bib.bib3)\)\. The hierarchical sequential labeling network \(HSLN\)\([Jin and Szolovits, 2018](https://arxiv.org/html/2609.19650#bib.bib2)\)is representative: Bi\-LSTM encodes the tokens of each sentence and attention pooling aggregates them into a sentence vector, a second Bi\-LSTM enriches each sentence vector with context from its neighbors, a linear layer maps the result to label scores, and a conditional random field \(CRF\)\([Lafferty et al\., 2001](https://arxiv.org/html/2609.19650#bib.bib22)\)predicts the label sequence with the highest joint probability rather than maximizing each label independently\.[Brack et al\. \(2024\)](https://arxiv.org/html/2609.19650#bib.bib3)replaced the word embedding layer of HSLN with SciBERT\([Beltagy et al\., 2019](https://arxiv.org/html/2609.19650#bib.bib40)\), an encoder pretrained on scientific text, and evaluated the resulting model across four scientific domains\. More recently, approaches based on LLMs have shown competitive performance\([Lan et al\., 2024](https://arxiv.org/html/2609.19650#bib.bib4)\)\. However, existing SSC research remains predominantly focused on English\.

The CRF learns a transition matrix over labels, so that regularities such as Method preceding Result are encoded in the model rather than left to the classifier for each sentence\.[Brack et al\. \(2024\)](https://arxiv.org/html/2609.19650#bib.bib3)report that when tasks from different domains are given a common label set, sharing this output layer across them degrades performance, and attribute this to each domain having its own transition distribution over labels\. The difference between domains thus lies in the label sequence rather than in the sentences\. We compare languages on the same terms: Section[4\.3](https://arxiv.org/html/2609.19650#S4.SS3)measures how two languages differ in the label distributions, transitions, and positions that CRF features encode\.

### 2\.2\.Cross\-Lingual Transfer Learning

Zero\-shot cross\-lingual transfer, in which a model is trained on a source language and evaluated on a target language with no labeled examples, became widely studied with the release of multilingual BERT \(mBERT\)\([Devlin et al\., 2019](https://arxiv.org/html/2609.19650#bib.bib20)\), an mPLM pretrained on 104 languages\. Transfer performance has been reported to correlate with typological similarity\([Lauscher et al\., 2020](https://arxiv.org/html/2609.19650#bib.bib11)\)and alignment in word order\([Deshpande et al\., 2022](https://arxiv.org/html/2609.19650#bib.bib21)\)\. However, linguistic proximity is an inconsistent predictor: its effect varies across tasks\([Blaschke et al\., 2025](https://arxiv.org/html/2609.19650#bib.bib7)\), and competitive transfer occurs even without shared script or family\([De Souza et al\., 2024](https://arxiv.org/html/2609.19650#bib.bib34)\)\. A separate line of work finds that similarity in the internal representations of mPLMs tracks transfer more closely than surface features\([Lin et al\., 2024](https://arxiv.org/html/2609.19650#bib.bib8);[Yun et al\., 2023](https://arxiv.org/html/2609.19650#bib.bib9)\)\. These findings point to factors specific to the task, which motivates a structural view\.

Discourse\-level work provides a starting point\.[Zeyrek et al\. \(2020\)](https://arxiv.org/html/2609.19650#bib.bib13)annotate parallel transcripts in six languages and report that the distribution of discourse relations differs across them even when the content is held constant, and[Braud et al\. \(2017\)](https://arxiv.org/html/2609.19650#bib.bib12)report that in cross\-lingual discourse parsing, the segmentation of a document into related spans transfers reasonably well, whereas the relations themselves degrade sharply, which they attribute in part to differences in the relation distributions of the source and target corpora\. Comparative studies of research article abstracts report variation of the same kind in the genre we target: English abstracts tend to argue for the significance of the work, whereas abstracts addressed to local scientific communities tend to report what was done, with corresponding differences in which rhetorical roles appear and how often\([Martín\-Martín, 2003](https://arxiv.org/html/2609.19650#bib.bib37);[Van Bonn and Swales, 2007](https://arxiv.org/html/2609.19650#bib.bib38);[Yakhontova, 2002](https://arxiv.org/html/2609.19650#bib.bib39)\)\. These are manual analyses of small corpora, each covering a single language pair, and they do not connect the variation they document to automatic classification\. Since SSC operates on sequences of rhetorical roles, this variation is a candidate predictor of transfer that is independent of linguistic proximity, and it is the one we test\.

### 2\.3\.Structure\-Aware Methods

Recent structure\-aware approaches have improved document modeling\([Buchmann et al\., 2024](https://arxiv.org/html/2609.19650#bib.bib15)\)and argumentation mining\([Sun et al\., 2024](https://arxiv.org/html/2609.19650#bib.bib16)\)\. The research has employed grammar constrained decoding\([Geng et al\., 2023](https://arxiv.org/html/2609.19650#bib.bib17);[Park et al\., 2024](https://arxiv.org/html/2609.19650#bib.bib18)\)and discriminative reranking\([Wang et al\., 2024](https://arxiv.org/html/2609.19650#bib.bib19)\)to ensure output consistency\. In light of these developments, our study adapts the verifier\-reranking paradigm\([Cobbe et al\., 2021](https://arxiv.org/html/2609.19650#bib.bib14)\)to SSC, using structural features to estimate the candidate quality without requiring gold labels\. In this paradigm, a model samples several candidate outputs and a separately trained verifier scores each one, replacing the single greedy decode with a selection among candidates; the verifier is trained on candidates paired with their observed correctness, so it needs no gold labels at inference\. We score whole candidate label sequences rather than individual labels\. We do not adopt grammar\-constrained decoding, since the constraint we need is distributional rather than categorical, as Section[4\.4](https://arxiv.org/html/2609.19650#S4.SS4)shows\.

## 3\.Multilingual SSC Dataset

### 3\.1\.Data Sources and Collection

The dataset was constructed from structured scientific abstracts\. Section headers supply sentence labels without manual annotation; the resulting data are intended for training models that are applied to abstracts without headers, where rhetorical roles are not directly available\. Following the labeling scheme of PubMed\-RCT 20k\([Dernoncourt and Lee, 2017](https://arxiv.org/html/2609.19650#bib.bib1)\), we built a multilingual dataset from non\-English abstracts whose section headers provide the rhetorical labels\. PubMed\-RCT is derived from the MEDLINE/PubMed Baseline Database and is not released under an explicit license; we use it strictly for non\-commercial research, consistent with the National Library of Medicine’s terms\.

We selected five academic databases based on three criteria: \(1\) provision of API access to abstracts, \(2\) substantial non\-English content, and \(3\) inclusion of structured abstracts with section headers:

- •DOAJ111[https://doaj\.org/](https://doaj.org/): Open\-access journals across multiple languages, accessed via REST API\.
- •HAL\(Hyper Articles en Ligne\)222[https://hal\.science/](https://hal.science/): the French national open archive, accessed via OAI\-PMH interface\.
- •Dialnet333[https://dialnet\.unirioja\.es/](https://dialnet.unirioja.es/): Spanish bibliographic database containing Spanish and Portuguese publications\.
- •
- •CiNii Research555[https://cir\.nii\.ac\.jp/](https://cir.nii.ac.jp/): Japanese academic database, accessed via REST API\.

We targeted 13 non\-English languages: French, Japanese, Spanish, Chinese, Russian, Portuguese, Italian, Indonesian, Turkish, Korean, Polish, Dutch, and Estonian\. Since structured abstracts typically contain Method and Result sections, we translated these two terms into each target language and used them as search queries\. We used only these two terms because nearly all structured abstracts contain them and their wording varies little across journals, whereas headers for Background, Objective, and Conclusion vary widely\. The queries thus condition on the presence of Method and Result headers and place no condition on the other sections\. We then verified the presence of section headers through pattern matching for three formats: the XML format<sec\> <title\> Methods </title\> \.\.\. </sec\>, the bracket format\[Methods\], and the colon formatMethods:\. Data collection was conducted between February and May 2025\.

Headers were mapped to the five labels using keyword lists compiled for each language \(for example, Aim, Purpose, and Objectives to Objective; Introduction to Background; Discussion to Conclusion\)\. A header not in the lists was not treated as a header, and the text following it was kept in the preceding section; text preceding the first header was discarded\. The lists are part of the released code\.

### 3\.2\.Preprocessing and Quality Assurance

We applied the following preprocessing steps: \(1\) HTML entity conversion, \(2\) NFKC Unicode normalization, \(3\) language detection using langdetect\([Shuyo, 2010](https://arxiv.org/html/2609.19650#bib.bib30)\)to discard mismatched abstracts, and \(4\) duplicate removal based on title matching that ignores case\. We retained only abstracts with two or more sections and used NLTK\([Bird et al\., 2009](https://arxiv.org/html/2609.19650#bib.bib28)\)for sentence segmentation in European languages and spaCy\([Honnibal et al\., 2020](https://arxiv.org/html/2609.19650#bib.bib29)\)for Asian languages\. The headers in all three formats were removed before sentence segmentation; in the colon format, the header string was stripped from the first sentence of the section\. The inputs to all models in Sections[4](https://arxiv.org/html/2609.19650#S4)and[6](https://arxiv.org/html/2609.19650#S6)therefore contain none of the headers from which the labels were derived\. PubMed\-RCT likewise provides sentences without headers\.

To verify dataset quality, we randomly sampled 50 abstracts from each of the nine languages with the largest amounts of data and manually inspected the labels of individual sentences\. On average, 2\.71 erroneous sentences were identified per 50 abstracts, which we consider negligible\.

### 3\.3\.Dataset Statistics

The dataset contained 52,487 abstracts and 504,416 sentences across 14 languages, comprising the 13 collected languages plus English from PubMed\-RCT 20k\. Table[1](https://arxiv.org/html/2609.19650#S3.T1)shows statistics by language\.

Because structured abstracts are most common in medicine and life sciences, the dataset is drawn predominantly from these fields, as Table[2](https://arxiv.org/html/2609.19650#S3.T2)shows\. Domain information was unavailable for Chinese, Turkish, Korean, Polish, Dutch, and Estonian from the DOAJ/TRdizin metadata; for these languages we expect a similar skew toward medicine and life sciences, since the convention of structured abstracts is largely confined to these fields\. Whether the structural differences observed in Section[4](https://arxiv.org/html/2609.19650#S4)partly reflect domain similarity is examined in Section[7](https://arxiv.org/html/2609.19650#S7)\.

The dataset and code are available on GitHub666[https://github\.com/mm\-doshisha/multilingual\-SSC](https://github.com/mm-doshisha/multilingual-SSC); parts of the code were written with the assistance of Anthropic’s Claude and were reviewed by the authors\. For data from CiNii Research, Dialnet, and TRdizin, we provide document IDs instead of full abstracts to comply with their data usage policies\.

Table 1\.Dataset statistics\.Table 2\.Approximate domain distribution per language; the primary domain is shown\.

## 4\.Cross\-Lingual Transfer Analysis

For SSC, which assigns rhetorical roles at the sentence level, we hypothesize that similarity in rhetorical structure across languages predicts cross\-lingual transfer more consistently than linguistic proximity\. To investigate this hypothesis, we selected nine languages with 200\+ abstracts from the dataset: Chinese, Spanish, English, French, Indonesian, Italian, Japanese, Portuguese, and Russian\. We conducted comprehensive cross\-lingual transfer experiments across9×99\\times 9language pairs, including pairs of the same language, using both BERT and LLM\-based models, examining whether the findings generalize across different model architectures\. Our experimental setting was zero\-shot cross\-lingual transfer: For each of the 81 language pairs, models were trained on the source language and evaluated on the target language without any training examples in the target language\.

### 4\.1\.Models

We employed mBERT with the hierarchical sequence labeling network \(HSLN\) architecture\([Brack et al\., 2024](https://arxiv.org/html/2609.19650#bib.bib3)\), hereafter referred to as mBERT\-HSLN\. HSLN uses an architecture with two layers: representations at the sentence level are first obtained by aggregating token embeddings via a Bi\-LSTM, and then a second Bi\-LSTM performs contextualized sequence labeling over the sentence representations\. We followed the hyperparameters from[Brack et al\. \(2024\)](https://arxiv.org/html/2609.19650#bib.bib3)\.

For LLM\-based models, we selected three models fine\-tuned to follow natural language instructions: Gemma2\-2B\-it\([Gemma Team, 2024](https://arxiv.org/html/2609.19650#bib.bib31)\), Qwen2\.5\-3B\-Instruct\([Qwen Team, 2024](https://arxiv.org/html/2609.19650#bib.bib32)\), and Llama\-3\.2\-3B\-Instruct\([Llama Team, AI @ Meta, 2024](https://arxiv.org/html/2609.19650#bib.bib33)\)\. The prompt used was a simplified version of that proposed by[Lan et al\. \(2024\)](https://arxiv.org/html/2609.19650#bib.bib4), from which we removed the demonstration examples \(to focus on zero\-shot transfer\) and the abstract context \(for the reason given below\)\. We unified prompts in English for all experiments in this section\. Each prompt contains only the sentence to be classified, without the surrounding sentences; each sentence is classified in a separate call, and the label sequence of an abstract is the concatenation of these outputs\. We chose this setting for the transfer analysis so that the LLMs cannot exploit the order and position of sentences, for which they could rely on knowledge acquired in pretraining rather than on the source corpus, whereas HSLN models the sequence only through components trained from scratch on that corpus\. The template is shown in Figure[1](https://arxiv.org/html/2609.19650#S4.F1)\.

Instruction:’You must categorize the given sentence into one of these five labels: Background, Objective, Method, Result, Conclusion\. Respond with ONLY the label name\.’Input:’Target Sentence: Individuals who received CM targeting psycho\-stimulants were 79% more likely to submit a smoking\-negative breath\-sample relative to controls\.’Output:’Question: What is the rhetorical role of the Target Sentence? Answer with one word from the labels list\.’

Figure 1\.Prompt template used in the transfer analysis\. The example sentence is taken from the training split of PubMed\-RCT 20k\.We applied low\-rank adaptation \(LoRA;\([Hu et al\., 2022](https://arxiv.org/html/2609.19650#bib.bib26)\)\) fine\-tuning withr=32r=32andα=64\\alpha=64for all attention and feedforward network layers, using a learning rate of2×10−42\\times 10^\{\-4\}, batch size of 4, and training for three epochs\.

For each language, we split the data as follows: 70% training, 15% validation, and 15% testing\. The test set for each language comprisedmin⁡\(200,⌊0\.15×N⌋\)\\min\(200,\\ \\lfloor 0\.15\\times N\\rfloor\)abstracts, whereNNis the total number of abstracts for that language\. For each language pair, we conducted three runs with different random seeds and recorded the average Macro F1 scores\.

### 4\.2\.Results of Zero\-Shot Cross\-Lingual Transfer

Figure[2](https://arxiv.org/html/2609.19650#S4.F2)shows the transfer results for Qwen2\.5\-3B\-Instruct as a representative example\. We observed consistent transfer patterns across all four models, namely mBERT\-HSLN and the three LLMs, which suggests that the findings are not specific to particular model architectures\. Pairs of the same language reached an average Macro F1 of 0\.611 for this model, while cross\-lingual transfer averaged 0\.525\. Transfer performance varied considerably: Japanese to Chinese and Chinese to Indonesian reached 0\.642 and 0\.655, whereas French to English and Spanish to Japanese fell to 0\.355 and 0\.391\. For mBERT\-HSLN, the corresponding averages were 0\.488 for pairs of the same language and 0\.394 for cross\-lingual pairs, and the strongest and weakest pairs were largely the same across architectures\.

Common patterns held across all models, including stable performance with English as source\. Transfer was also asymmetric; for example, Qwen transferred Japanese to Chinese at 0\.642 but Chinese to Japanese at only 0\.451, likely reflecting model capability differences for each language\. Notably, Japanese and English consistently achieved high performance as source languages across all four models, while French and Russian were among the weakest sources\. This asymmetry is consistent with the source language effect discussed in Section[4\.5](https://arxiv.org/html/2609.19650#S4.SS5), where transfer performance correlates with the source language’s in\-domain performance\.

![Refer to caption](https://arxiv.org/html/2609.19650v1/fig_transfer_heatmap.png)Figure 2\.Cross\-lingual transfer performance for Qwen2\.5\-3B\-Instruct \(Macro F1\)\. Rows: training languages; Columns: evaluation languages\.
### 4\.3\.Linguistic and Structural Similarity Measures

To determine which factors predict cross\-lingual transfer performance, we compared two similarity measures between language pairs: linguistic proximity and structural similarity\. For linguistic proximity, we used lang2vec\([Littell et al\., 2017](https://arxiv.org/html/2609.19650#bib.bib23)\), which provides typological feature vectors for languages\. We concatenated feature vectors from five categories: syntax, phonology, inventory, geography, and family, and calculated the cosine similarity between each language pair\.

For structural similarity, we designed measures based on the rhetorical structure patterns found in the abstracts\. These measures are computed on the corpus collected for each language, so the similarity between two languages is measured on these corpora and may partly reflect their source databases and disciplines rather than the languages alone \(Section[7](https://arxiv.org/html/2609.19650#S7)\)\. Drawing on feature representations used in conditional random fields \(CRF;\([Lafferty et al\., 2001](https://arxiv.org/html/2609.19650#bib.bib22)\)\) and research on abstract composition patterns\([Martín\-Martín, 2003](https://arxiv.org/html/2609.19650#bib.bib37)\), we defined six distance measures between language pairs, all normalized to the\[0,1\]\[0,1\]range\. Label distribution distance uses the Jensen\-Shannon divergence \(JSD\) between the frequency distributions of the five labels:

\(1\)dlabel​\(L1,L2\)=JSD​\(PL1,PL2\)d\_\{\\text\{label\}\}\(L\_\{1\},L\_\{2\}\)=\\text\{JSD\}\(P\_\{L\_\{1\}\},P\_\{L\_\{2\}\}\)wherePL​\(l\)P\_\{L\}\(l\)is the frequency of labelllin languageLL, computed as the number of sentences labeledlldivided by the total number of sentences inLL\. Section length distance averages the JSD of section length distributions across labels, where a section length is the number of consecutive sentences with the same label:

\(2\)dsection​\(L1,L2\)=15​∑lJSD​\(PL1,l,PL2,l\)d\_\{\\text\{section\}\}\(L\_\{1\},L\_\{2\}\)=\\frac\{1\}\{5\}\\sum\_\{l\}\\text\{JSD\}\(P\_\{L\_\{1\},l\},P\_\{L\_\{2\},l\}\)Continuation probability distance uses the normalized Euclidean distance between vectors of label continuation probabilitiespL​\(l\)=P⁡\(yi\+1=l∣yi=l\)p\_\{L\}\(l\)=P\(y\_\{i\+1\}=l\\mid y\_\{i\}=l\):

\(3\)dcont​\(L1,L2\)=15​∑l\(pL1​\(l\)−pL2​\(l\)\)2d\_\{\\text\{cont\}\}\(L\_\{1\},L\_\{2\}\)=\\frac\{1\}\{\\sqrt\{5\}\}\\sqrt\{\\sum\_\{l\}\(p\_\{L\_\{1\}\}\(l\)\-p\_\{L\_\{2\}\}\(l\)\)^\{2\}\}Boundary position distance computes the weighted difference in average relative positions of four major transitions such as Method→\\toResult, with weights proportional to the minimum transition frequency\. Block count distance applies JSD to the distribution of label spans that are not contiguous, analogously to section length distance\. Transition entropy distance normalizes the absolute difference in Shannon entropy of label transition distributions bylog2⁡25\\log\_\{2\}25, the maximum entropy for 25 transition patterns\.

The overall structural distance is the simple average of all six normalized distances,d⁡\(L1,L2\)=16​∑i=16di​\(L1,L2\)d\(L\_\{1\},L\_\{2\}\)=\\frac\{1\}\{6\}\\sum\_\{i=1\}^\{6\}d\_\{i\}\(L\_\{1\},L\_\{2\}\), and structural similarity is defined as1−d⁡\(L1,L2\)1\-d\(L\_\{1\},L\_\{2\}\)\. Equal weighting was adopted as a neutral baseline to avoid overfitting given the limited sample size of 72 language pairs; weighting learned or adapted to the task was left for future work\.

Figure[3](https://arxiv.org/html/2609.19650#S4.F3)shows the two similarity matrices\. While linguistic proximity largely reflects language family membership, structural similarity reveals different patterns: Japanese–Spanish and Japanese–Portuguese both reach 0\.93 despite belonging to different language families\. Chinese–Russian also shows high structural similarity at 0\.89, as both tend to omit Background and start with Objective\.

![Refer to caption](https://arxiv.org/html/2609.19650v1/fig_similarity_matrices.png)Figure 3\.Linguistic \(left\) and structural \(right\) similarity matrices\.
### 4\.4\.Structural Patterns Across Languages

To understand why structural similarity diverges from linguistic proximity, we analyzed the rhetorical patterns of each language using label transition probabilities and positional distributions\. Two primary patterns emerged, with additional variation in positional distributions\.

Background omission\.Russian abstracts contained no Background section, and Chinese abstracts contained only four Background sentences out of 24,649\. Both languages show strong Objective→\\toMethod transitions, at 0\.96 for Russian and 0\.95 for Chinese, indicating that abstracts begin directly with Objectives\. This shared convention explains the high structural similarity of 0\.89 and the above\-average transfer from Chinese to Russian \(0\.605 for Qwen\) despite their typological distance\. Among the nine analyzed languages, only Chinese and Russian exhibit this pattern, forming a distinct structural cluster\.

Canonical IMRaD adherence\.Indonesian and Italian show strong sequential transitions consistent with the canonical IMRaD structure; Indonesian reaches 0\.73 for Method→\\toResult and 0\.63 for Result→\\toConclusion, and Italian reaches 0\.72 and 0\.79 for these same transitions, with Background appearing in the first 20% of abstracts, followed by Objective, Method, Result, and Conclusion in sequence\. Similarly, Spanish, Portuguese, and Japanese abstracts maintain clear positional separation of sections, which accounts for the high structural similarity among these five languages, all pairwise similarities being≥0\.91\\geq 0\.91\. This cluster also includes English, whose pairwise similarity with all five is≥0\.86\\geq 0\.86, reflecting the influence of international publication norms in biomedicine\.

Positional variation\.Beyond these two patterns, languages differ in where sections are positioned within the abstract\. Because Chinese and Russian omit Background, they concentrate Objective and Method in the first 30% of the abstract, whereas Indonesian and Italian show a more evenly distributed positional profile\. French occupies an intermediate position: it retains Background sections, but the positional distributions are more compressed than those of Spanish or Portuguese, potentially reflecting differences in editorial conventions across journals in French\. These positional differences are captured by the boundary position and label distribution metrics defined in Section[4\.3](https://arxiv.org/html/2609.19650#S4.SS3)\.

These patterns connect to cross\-lingual transfer\. Pairs that share rhetorical conventions despite typological distance, such as Chinese and Russian, which both omit Background, or Japanese together with Spanish and Portuguese, which all maintain clear positional separation, show high structural similarity, and across the 72 language pairs structural similarity correlates with transfer performance, as shown in Section[4\.5](https://arxiv.org/html/2609.19650#S4.SS5)\. At the level of individual pairs, transfer is also shaped by source language capability and is often asymmetric, so structural similarity alone does not determine any single outcome; structurally divergent pairs such as French→\\toEnglish at 0\.355 transfer poorly\. The structural conventions themselves likely reflect a combination of disciplinary norms, such as medical journals requiring structured abstracts, and editorial practices specific to each language\.

### 4\.5\.Correlation Analysis

To determine which similarity measure better predicts transfer performance, we analyzed the 72 pairs, excluding pairs of the same language, through correlation analysis\. Figure[4](https://arxiv.org/html/2609.19650#S4.F4)shows the results\. Statistically significant positive Pearson correlations between structural similarity and transfer performance were found for all four models: mBERT\-HSLN \(r=0\.372r=0\.372,p<0\.001p<0\.001\), Qwen2\.5\-3B \(r=0\.343r=0\.343,p<0\.001p<0\.001\), Gemma2\-2B \(r=0\.308r=0\.308,p<0\.01p<0\.01\), and Llama\-3\.2\-3B \(r=0\.312r=0\.312,p<0\.01p<0\.01\)\. Spearman correlations showed similar tendencies\. The p\-values here and in Tables[3](https://arxiv.org/html/2609.19650#S4.T3)and[4](https://arxiv.org/html/2609.19650#S4.T4)treat the 72 pairs as independent; pairs that share a language are not, so these values are nominal and likely optimistic \(Section[7](https://arxiv.org/html/2609.19650#S7)\)\.

For linguistic proximity, no model presented consistent significant correlations\. In mBERT\-HSLN, neither the Pearson nor the Spearman correlation was significant \(r=0\.087r=0\.087,p=0\.438p=0\.438;ρ=0\.024\\rho=0\.024,p=0\.831p=0\.831\)\. Among the LLMs, only Gemma2\-2B showed a significant Pearson correlation \(r=0\.278r=0\.278,p=0\.012p=0\.012\), but the Spearman correlation was not significant\. In Qwen2\.5\-3B and Llama\-3\.2\-3B, no significant correlations were found \(allp\>0\.05p\>0\.05\)\. Overall, linguistic proximity showed weak and inconsistent predictive power\.

However, transfer performance is also strongly correlated with the source language’s in\-domain performance \(mBERT\-HSLN:r=0\.86r=0\.86; Qwen:r=0\.49r=0\.49; Gemma:r=0\.71r=0\.71; Llama:r=0\.72r=0\.72; allp<0\.001p<0\.001\), which is a potential confound for the correlations reported above\.

Figure 4\.Correlation between similarity measures and transfer performance for Qwen2\.5\-3B \(N=72N=72\)\. Left: structural similarity \(Pearsonr=0\.343r=0\.343,p<0\.001p<0\.001\)\. Right: linguistic proximity \(Pearsonr=−0\.042r=\-0\.042,p=0\.726p=0\.726\)\.Partial correlation analysis\.To assess the independent contribution of structural similarity, we computed partial correlations controlling for the source language’s in\-domain performance\. The combined structural similarity metric retained a significant partial correlation for Qwen \(r=0\.274r=0\.274,p=0\.021p=0\.021\) and Llama \(r=0\.269r=0\.269,p=0\.024p=0\.024\), but not for mBERT\-HSLN \(r=0\.153r=0\.153,p=0\.204p=0\.204\) or Gemma \(r=0\.080r=0\.080,p=0\.506p=0\.506\)\. Linguistic proximity showed no significant partial correlation for any model \(allp\>0\.45p\>0\.45\)\. This indicates that while the overall structural similarity signal is partially attributable to source language performance, it retains independent predictive power in some architectures, whereas linguistic proximity does not\.

Individual metric contributions\.Before controlling for the confound, label distribution was significant across all four architectures, withrrranging from 0\.288 to 0\.432\. The full Pearson correlations for each metric across the four models are reported in Table[3](https://arxiv.org/html/2609.19650#S4.T3)\.

Table 3\.Pearsonrrbetween individual structural metrics and transfer performance across four models\.∗∗∗p<0\.001\{\}^\{\*\*\*\}p<0\.001,p∗⁣∗<0\.01\{\}^\{\*\*\}p<0\.01,∗p<0\.05\{\}^\{\*\}p<0\.05, n\.s\. = not significant\.Partial correlations of individual metrics\.To identify which structural aspects retain predictive power after controlling for source language performance, we computed partial correlations for each individual metric, reported in Table[4](https://arxiv.org/html/2609.19650#S4.T4)\. Label distribution was the only individual metric to remain significant in three of four models \(mBERT:r=0\.347r=0\.347,p<0\.01p<0\.01; Qwen:r=0\.251r=0\.251,p<0\.05p<0\.05; Llama:r=0\.406r=0\.406,p<0\.001p<0\.001\)\. Block count and transition entropy each reached significance in two models, whereas boundary position lost all significance, indicating that its apparent predictive power was largely attributable to source language performance\. The variation across models, with Gemma showing no significant partial correlations for any metric, likely reflects differences in multilingual capability and pretraining data composition\. We also note that for Qwen, the strongest model in the in\-domain evaluation reported in Table[6](https://arxiv.org/html/2609.19650#S6.T6), the highest individual partial correlation was transition entropy at 0\.337 rather than label distribution; we treat this as a pattern specific to the model and do not read it as an explanation of Qwen’s overall performance\. The prominence of label distribution does not stem from the sentence\-level prompt of the LLMs \(Section[4\.1](https://arxiv.org/html/2609.19650#S4.SS1)\): mBERT\-HSLN, which uses the sequence, shows it as well\.

Table 4\.Partial Pearsonrrbetween individual structural metrics and transfer performance, controlling for source language in\-domain performance\.∗∗∗p<0\.001\{\}^\{\*\*\*\}p<0\.001,p∗⁣∗<0\.01\{\}^\{\*\*\}p<0\.01,∗p<0\.05\{\}^\{\*\}p<0\.05\.Together, these results indicate that structural similarity is a more relevant predictor of cross\-lingual SSC transfer than linguistic proximity\. The combined metric retains an independent contribution in two of four models; label distribution similarity is the individual metric that most consistently retains significance, whereas boundary position retains none\.

## 5\.Leveraging Structural Information

Section[4](https://arxiv.org/html/2609.19650#S4)shows that the rhetorical structure of abstracts is associated with cross\-lingual transfer, which motivates giving models explicit structural guidance\. We therefore developed a set of three methods for explicitly leveraging structural information: structure\-informed prompting \(SIP\), structure\-guided verifier reranking \(SGVR\), and structure\-adaptive verifier \(SAV\)\.

Unlike the structure\-aware approaches in Section[2\.3](https://arxiv.org/html/2609.19650#S2.SS3), our methods use the structural features of the abstract itself, not the similarity between languages measured in Section[4](https://arxiv.org/html/2609.19650#S4), as cues at the input stage and during candidate selection, and require no gold labels at inference time\.

### 5\.1\.SIP

SIP guides LLM predictions by explicitly incorporating structural constraints into prompts\. Unlike the prompt in Section[4](https://arxiv.org/html/2609.19650#S4), SIP provides the full abstract together with the target sentence, and it specifies the typical label ordering of Background, Objective, Method, Result, and Conclusion so that the model can relate the position of the target sentence within the abstract to this ordering\. The prompt consisted of: \(1\) a task description with label definitions, \(2\) structural constraints indicating typical ordering, \(3\) one demonstration example randomly selected from the training data, independently for each test abstract, and \(4\) the full abstract and the target sentence\. The template is shown in Figure[5](https://arxiv.org/html/2609.19650#S5.F5)\.

Instruction:You must categorize the given sentence into one of these five labels: Background, Objective, Method, Result, Conclusion\. Respond with ONLY the label name\.Structural Constraints:Academic abstracts typically follow this order: Background \(introducing the topic\)→\\toObjective \(stating the research goal\)→\\toMethod \(describing the approach\)→\\toResult \(presenting findings\)→\\toConclusion \(summarizing implications\)\.Example:One abstract from the training data with each sentence followed by its label, selected independently for each test abstract\.Input:’Abstract: Treatments for drug addiction and smoking in severely mentally ill \(SMI\) adults are needed\. To investigate the effect of a contingency management \(CM\) intervention targeting psycho\-stimulant on cigarette smoking\. 126 stimulant dependent SMI smokers were assigned to CM or a non\-contingent control condition\. Rates of smoking\-negative \(<<3 ppm\) carbon monoxide breath\-samples were compared\. Individuals who received CM targeting psycho\-stimulants were 79% more likely to submit a smoking\-negative breath\-sample relative to controls\. This study provides initial evidence that a behavioral treatment for drug use results in reductions in cigarette smoking in SMI adults\.’ ’Target Sentence: Individuals who received CM targeting psycho\-stimulants were 79% more likely to submit a smoking\-negative breath\-sample relative to controls\.’Output:Question: What is the rhetorical role of the Target Sentence? Answer with one word from the labels list\.

Figure 5\.Prompt template for SIP\. The abstract in the Input field is taken from the training split of PubMed\-RCT 20k\.
### 5\.2\.SGVR

SGVR reranks multiple candidates generated by an LLM via temperature sampling, using structural features\. We adapt the verifier\-reranking paradigm\([Cobbe et al\., 2021](https://arxiv.org/html/2609.19650#bib.bib14)\)to SSC: following quality estimation in machine translation\([Specia and Shah, 2018](https://arxiv.org/html/2609.19650#bib.bib24)\), the verifier predicts output quality without reference labels and selects the most valid candidate\.

Specifically, for each training abstract, the LLM generatedKKcandidate label sequences𝐲=\(y1,…,yn\)\\mathbf\{y\}=\(y\_\{1\},\\ldots,y\_\{n\}\), wherennis the number of sentences in the abstract andyiy\_\{i\}is the predicted label for theii\-th sentence\. Each candidate was obtained by sampling one label per sentence with temperature sampling and concatenating the samples in sentence order\. We computed the accuracy at the sentence level for each candidate using the formulaq=1n∑i=1n𝟏\[yi=yi∗\]q=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\mathbf\{1\}\[y\_\{i\}=y\_\{i\}^\{\*\}\], whereyi∗y\_\{i\}^\{\*\}denotes the ground truth label\. From each candidate, we also extracted a feature vector of 44 dimensions𝐟⁡\(𝐲\)\\mathbf\{f\}\(\\mathbf\{y\}\), listed in Table[5](https://arxiv.org/html/2609.19650#S5.T5), capturing \(1\) label distribution features, namely occurrence counts and cosine similarity with the training data distribution; \(2\) transition features, namely the mean log probability of adjacent label transitions, binary indicators of major transitions, and entropy; and \(3\) position features, namely scores evaluating whether label positions fall within expected ranges and one\-hot representations of start and end labels\. We trained a regression modelVϕV\_\{\\phi\}that predictsqqfrom structural features𝐟⁡\(𝐲\)\\mathbf\{f\}\(\\mathbf\{y\}\)\.

Table 5\.Structural features for SGVR, totaling 44 dimensions\.At inference, the verifier predicts accuracyq^k=Vϕ​\(𝐟⁡\(𝐲k\)\)\\hat\{q\}\_\{k\}=V\_\{\\phi\}\(\\mathbf\{f\}\(\\mathbf\{y\}\_\{k\}\)\)for each candidate and selects𝐲pred=arg⁡maxk⁡q^k\\mathbf\{y\}^\{\\text\{pred\}\}=\\arg\\max\_\{k\}\\hat\{q\}\_\{k\}\.

### 5\.3\.SAV

SAV is an unsupervised method for zero\-shot settings, where no validation data are available for verifier training\. Instead, it dynamically estimates structural properties from the candidate set and favors labels that are consistent across candidates\. Simple majority voting, such as self\-consistency\([Wang et al\., 2023](https://arxiv.org/html/2609.19650#bib.bib25)\), is ineffective for sequence labeling if exact matches across full sequences are rare\. SAV instead evaluates agreement at each position and transition consistency separately to select the most structurally valid candidate\.

For each abstract, as in SGVR, an LLM generatedKKcandidate predictions\. We denote thekk\-th candidate as𝐲\(k\)=\(y1\(k\),…,yn\(k\)\)\\mathbf\{y\}^\{\(k\)\}=\(y\_\{1\}^\{\(k\)\},\\ldots,y\_\{n\}^\{\(k\)\}\)\. Each candidate was evaluated using three scores:

Theconsistency scoreScons​\(𝐲\(k\)\)S\_\{\\text\{cons\}\}\(\\mathbf\{y\}^\{\(k\)\}\)measures the average agreement rate at each position between candidate𝐲\(k\)\\mathbf\{y\}^\{\(k\)\}and all other candidates\. Theconfidence scoreSconf​\(𝐲\(k\)\)S\_\{\\text\{conf\}\}\(\\mathbf\{y\}^\{\(k\)\}\)first identifies the most frequent labelli∗l\_\{i\}^\{\*\}at each positioniiacross theKKcandidates and its agreement rateri=1K∑j=1K𝟏\[yi\(j\)=li∗\]r\_\{i\}=\\frac\{1\}\{K\}\\sum\_\{j=1\}^\{K\}\\mathbf\{1\}\[y\_\{i\}^\{\(j\)\}=l\_\{i\}^\{\*\}\]; the score rewards candidates that agree with these majority labels at each position, weighted by the agreement rates\. Thetransition scoreStrans​\(𝐲\(k\)\)S\_\{\\text\{trans\}\}\(\\mathbf\{y\}^\{\(k\)\}\)applies the same principle to transitions between adjacent labels rather than individual labels\.

The final score isS⁡\(𝐲\(k\)\)=13​\(Scons\+Sconf\+Strans\)S\(\\mathbf\{y\}^\{\(k\)\}\)=\\frac\{1\}\{3\}\(S\_\{\\text\{cons\}\}\+S\_\{\\text\{conf\}\}\+S\_\{\\text\{trans\}\}\), and the prediction is𝐲pred=arg⁡maxk⁡S⁡\(𝐲\(k\)\)\\mathbf\{y\}^\{\\text\{pred\}\}=\\arg\\max\_\{k\}S\(\\mathbf\{y\}^\{\(k\)\}\)\. Unlike simple majority voting at each position, SAV always outputs a sequence that exists in the candidate set, avoiding inconsistent reconstructions\.

Together, SIP, SGVR, and SAV incorporate structural constraints at the input stage as well as during candidate selection\. SIP\+SGVR was used when the target language is included in the training data, enabling verifier training, and SIP\+SAV was used for zero\-shot settings\.

## 6\.Experiments

### 6\.1\.Experimental Settings

Unlike the analysis of language pairs in Section[4](https://arxiv.org/html/2609.19650#S4), we trained models on mixed multilingual data and evaluated them in two settings: \(1\) in\-domain evaluation, testing on trained languages, and \(2\) zero\-shot evaluation, testing on languages not included during training\.

For in\-domain evaluation, we used nine languages: English, French, Japanese, Spanish, Chinese, Indonesian, Portuguese, Italian, and Russian, splitting each into a 70:15:15 ratio for training, validation, and testing, respectively\. The model was trained on combined training data from all nine languages\. For zero\-shot evaluation, we used five languages not included during training: Estonian, Korean, Dutch, Polish, and Turkish\. Since these languages were not used for training, their entire sets served as test sets\.

We clarify how the single demonstration in SIP relates to our zero\-shot claim, since the two operate at different levels\.*Zero\-shot cross\-lingual transfer*refers to the absence of any training data in the target language, whereas the single demonstration in SIP is a form of*1\-shot prompting*that conditions the model with one labeled example\. In the zero\-shot evaluation, this demonstration is drawn from the in\-domain training languages, so the target language remains unseen during both fine\-tuning and prompting\.

We applied LoRA withr=32r=32andα=64\\alpha=64for all attention and feedforward network layers to Qwen2\.5\-3B\-Instruct, Llama\-3\.2\-3B\-Instruct, and Gemma2\-2B\-it, training with a learning rate of2×10−42\\times 10^\{\-4\}and batch size 4, for three epochs\. For SGVR, we setK=3K=3; for SAV in zero\-shot evaluation, we setK=5K=5to balance candidate diversity against computational cost\. In preliminary experiments, performance gains plateaued beyondK=5K=5, while inference time increased roughly linearly withKK\. The SGVR verifierVϕV\_\{\\phi\}is a LightGBM\([Ke et al\., 2017](https://arxiv.org/html/2609.19650#bib.bib27)\)regressor, chosen for its ability to capture non\-linear feature interactions\. SGVR withK=3K=3and SAV withK=5K=5increase inference time by factors of approximately three and five relative to decoding a single candidate, and the SIP prompt is longer than the baseline prompt of Table[8](https://arxiv.org/html/2609.19650#S6.T8)by the ordering constraint and the demonstration\. On an H100 GPU, SGVR withK=3K=3raised inference time from 1\.08 to 3\.35 seconds per abstract, and SAV withK=5K=5from 1\.33 to 6\.55 seconds\. For applications sensitive to latency, SIP alone offers a competitive alternative \(Table[6](https://arxiv.org/html/2609.19650#S6.T6)\)\. All experiments were run with three inference iterations, and we report mean values with standard deviations\.

We used the following baselines: \(1\)multilingual LLM\-SSC, a multilingual extension of[Lan et al\. \(2024\)](https://arxiv.org/html/2609.19650#bib.bib4)with their original prompt, which includes the full abstract and demonstration examples; \(2\)mB ERT\-HSLN, the HSLN used by[Brack et al\. \(2024\)](https://arxiv.org/html/2609.19650#bib.bib3)with mBERT encoder; and additionally \(3\)mmBERT\-HSLN,XLM\-R\-HSLN\([Conneau et al\., 2020](https://arxiv.org/html/2609.19650#bib.bib35)\), andLaBSE\-HSLN\([Feng et al\., 2022](https://arxiv.org/html/2609.19650#bib.bib36)\), which replace the encoder component with larger or more recent multilingual encoders\. We report accuracy at the sentence level and Macro F1 across five classes\. Overall scores in Tables[6](https://arxiv.org/html/2609.19650#S6.T6),[7](https://arxiv.org/html/2609.19650#S6.T7), and[9](https://arxiv.org/html/2609.19650#S6.T9)are computed over the pooled test sentences of all languages\.

### 6\.2\.In\-Domain Evaluation

Table 6\.Overall in\-domain evaluation results on 9 languages\.Table[6](https://arxiv.org/html/2609.19650#S6.T6)shows the results aggregated across test sets\. SIP\+SGVR with Qwen achieves an accuracy of 92\.9 and Macro F1 of 84\.9, outperforming mBERT\-HSLN by \+2\.3 in Macro F1 and matching or exceeding the strongest encoder baselines, mmBERT\-HSLN and XLM\-R\-HSLN, and clearly surpassing LaBSE\-HSLN\. The use of SIP and SGVR substantially improved performance over LLM\-SSC across all three LLMs\.

Table[7](https://arxiv.org/html/2609.19650#S6.T7)shows the results for the four major languages\. Our method consistently outperformed baselines across all four languages, with the largest gains in Spanish and Japanese at \+9\.3 and \+9\.2 Macro F1, a moderate gain of \+5\.8 in English, and a smaller gain of \+0\.9 in Chinese\.

Table 7\.Macro F1 results for individual major languages\. The Overall row reports the score over the pooled test sentences of all nine in\-domain languages\. SIP\+SGVR uses Qwen2\.5\-3B\-Instruct\.We also analyzed component contributions, reported in Table[8](https://arxiv.org/html/2609.19650#S6.T8)\. SGVR with a baseline prompt, namely the SIP prompt without the structural constraints and the demonstration but with the full abstract, achieved an accuracy of 91\.5 and Macro F1 of 81\.6, compared to 92\.9 and 84\.9 for SIP\+SGVR with Qwen\. The−3\.3\-3\.3decrease in Macro F1 when removing SIP indicates that the SIP prompt, which adds the ordering constraint and one demonstration to this baseline, contributes substantially to accurate classification \(see Section[7](https://arxiv.org/html/2609.19650#S7)for what this comparison does not separate\), and the larger drop in Macro F1 than in accuracy suggests an effect on minority classes\. SIP\+SGVR with Qwen shows improvements of \+1\.0 in accuracy and \+0\.8 in Macro F1 over SIP alone, confirming SGVR’s contribution\.

Table 8\.Ablation results for Qwen2\.5\-3B\-Instruct, overall\.
### 6\.3\.Zero\-Shot Cross\-Lingual Evaluation

Table 9\.Overall zero\-shot evaluation results for five unseen languages\.We evaluated models trained on the nine in\-domain languages on five unseen languages\. Since Qwen2\.5\-3B\-Instruct was the strongest in\-domain model in Table[6](https://arxiv.org/html/2609.19650#S6.T6), we conducted the zero\-shot evaluation with Qwen alone\. In this setting, validation data for the target languages were unavailable, making verifier training impossible; we therefore used SAV, which selects predictions based on candidate consistency\. SIP\+SAV with Qwen reaches 84\.9 accuracy and 78\.0 Macro F1, outperforming the strongest encoder baseline XLM\-R\-HSLN at 72\.3 Macro F1 by \+5\.7, and mBERT\-HSLN by \+10\.3, as reported in Table[9](https://arxiv.org/html/2609.19650#S6.T9)\. This indicates that encoder strength alone does not close the cross\-lingual generalization gap\. The larger improvement in Macro F1 suggests more stable performance for minority classes\.

Macro F1 varied across the five unseen languages: 88\.3 for Polish \(30 abstracts\), 87\.4 for Dutch \(14\), 78\.5 for Korean \(48\), 75\.4 for Turkish \(179\), and 70\.9 for Estonian \(7\)\. Because the test sets are this small, we do not interpret the differences among languages; in particular, the structural similarity of the unseen languages to the training languages cannot be estimated reliably from sets of this size and was not computed\.

## 7\.Limitations

Corpus\-level confounding\.Each language is represented by a single collected corpus, and the corpora differ in disciplinary composition and editorial conventions and, across databases, in retrieval procedure\. The structural differences reported in Section[4](https://arxiv.org/html/2609.19650#S4)may therefore partly reflect these factors rather than conventions of the languages themselves; since differences also appear among languages collected from the same database \(Chinese and Russian versus Italian and Indonesian, all from DOAJ\), the database alone does not explain them\.

Subject domain bias\.Our dataset primarily comprises abstracts with explicit section headers, which are prevalent in medicine and life sciences\. Its effectiveness in subject domains with different rhetorical organizations, such as humanities and social sciences, remains to be validated\. A related question is whether structural similarity merely reflects domain similarity\. The domain composition in Table[2](https://arxiv.org/html/2609.19650#S3.T2)differs across languages: Japanese from CiNii is almost entirely medical and life sciences, whereas Spanish and Portuguese from Dialnet include a larger share of health and social sciences\. These languages nonetheless show high structural similarity and good transfer, which suggests that structural similarity is not reducible to broad domain composition alone\. This is a single illustrative comparison rather than a controlled test, and the domain figures rest on partial metadata coverage, so they are approximate\. The more relevant factor may be the prevalence of the IMRaD convention for structured abstracts rather than broad subject domain, and separating the two is left for future work\.

Language coverage\.Although the dataset covers 13 languages, it lacks representation for several major linguistic regions and families, including Sub\-Saharan African \(Swahili\), Indic \(Hindi\), and Southeast Asian \(Thai\) languages\. The observed structural patterns may differ for these underrepresented groups\.

Limited hyperparameter exploration\.We setK=3K=3for SGVR andK=5K=5for SAV based on preliminary experiments\. The optimal value ofKKlikely varies by language, domain, and compute budget\.

Zero\-shot evaluation scope\.The zero\-shot evaluation was limited to five languages with small test sets of 7 to 179 abstracts, and we do not report confidence intervals for them\. Scores for individual languages should therefore be read as indicative, and those from the smallest sets, Dutch and Estonian, warrant particular caution\. The manual label check in Section[3\.2](https://arxiv.org/html/2609.19650#S3.SS2)did not include these languages, so their label quality is unverified\. A larger, more diverse evaluation would further strengthen our findings\.

Attribution of gains\.Table[8](https://arxiv.org/html/2609.19650#S6.T8)compares prompts that differ in both the ordering constraint and the demonstration, so the contribution of the constraint alone is not isolated\. Likewise, the gain of SIP\+SAV over LLM\-SSC in Table[9](https://arxiv.org/html/2609.19650#S6.T9)combines the SIP prompt, sampling of five candidates, and consistency\-based selection; we did not separate these effects or compare SAV with position\-wise majority voting\.

Lack of independence between language pairs\.The correlation and partial correlation analyses in Section[4](https://arxiv.org/html/2609.19650#S4)treat the 72 language pairs as independent observations, but pairs that share a language are not statistically independent\. Conventional p\-values may therefore be optimistic\. We did not apply a permutation or Mantel test to account for this dependence, so the reported significance levels should be interpreted with this caveat in mind\. A permutation test that permutes language identities, or a regression with random effects for source and target language, would account for this dependence; we leave this to future work\.

Source language confounding\.Transfer performance correlates with source language in\-domain performance, which reflects both measurable fine\-tuning performance and pretraining coverage that we did not directly measure\. While this confounding factor cannot be fully controlled, and is addressed in part by the partial correlations in Section[4\.5](https://arxiv.org/html/2609.19650#S4.SS5), our comparative analysis between structural similarity and linguistic proximity remains less affected, as both predictors are evaluated using the same set of language pairs under identical conditions\.

## 8\.Conclusion

We constructed a multilingual SSC dataset covering 13 non\-English languages from five major academic databases and analyzed the factors that determine cross\-lingual transfer success\. Our key finding is that linguistic proximity has no consistent predictive power for transfer, whereas structural similarity in rhetorical organization, and the distribution of rhetorical roles in particular, shows a weak but consistent positive correlation\. While the strength of this relationship varies across model architectures and is partially attributable to source language performance, label distribution similarity retains independent predictive power in three of four tested models\.

We proposed SIP, SGVR, and SAV to leverage this structural information\. In the in\-domain evaluation, SIP\+SGVR reaches parity with the strongest encoder baselines, mmBERT\-HSLN and XLM\-R\-HSLN, while clearly outperforming the mBERT\-HSLN and LLM\-SSC baselines, and in zero\-shot transfer to unseen languages SIP\+SAV outperforms the strongest encoder baseline XLM\-R\-HSLN by \+5\.7 Macro F1\.

These results have practical implications for multilingual digital library systems\. First, our dataset provides rhetorical annotations at the sentence level, released as full abstracts for DOAJ and HAL and as document IDs for CiNii Research, Dialnet, and TRdizin, enabling structured metadata enrichment for these collections\. These annotations can support faceted search interfaces that allow users to retrieve papers by specific rhetorical components, such as finding papers with similar methods but different conclusions\. Second, the correlational finding that structural similarity is associated with transfer success suggests a hypothesis for selecting training data when deploying SSC to new languages: source languages with similar rhetorical conventions may be preferable to typologically related ones\. Our analysis is correlational and does not directly evaluate this selection strategy, which we leave for future work\. Third, the zero\-shot capability of SIP\+SAV makes it feasible to extend rhetorical structure analysis to new languages in digital libraries without collecting labeled training data in each target language, lowering the barrier for multilingual deployment\.

Future work includes analyzing internal representations to understand how structural patterns are encoded, extending our approach to other tasks at the discourse level such as citation function classification, verifying the structural similarity hypothesis on subsets with matched label distributions, and exploring weighting of the structural similarity components adapted to the task\.

###### Acknowledgements\.

This work was supported by JSPS KAKENHI Grant Number JP25K03419\. The computations were carried out on the TSUBAME4\.0 supercomputer at Institute of Science Tokyo\.

## References

- Beltagyet al\.\(2019\)I\. Beltagy, K\. Lo, and A\. CohanSciBERT: a pretrained language model for scientific text\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),K\. Inui, J\. Jiang, V\. Ng, and X\. Wan \(Eds\.\),Hong Kong, China,pp\. 3615–3620\.External Links:[Link](https://aclanthology.org/D19-1371/),[Document](https://dx.doi.org/10.18653/v1/D19-1371)Cited by:[§2\.1](https://arxiv.org/html/2609.19650#S2.SS1.p1.1)\.
- Birdet al\.\(2009\)S\. Bird, E\. Klein, and E\. LoperNatural language processing with Python: analyzing text with the natural language toolkit\.O’Reilly Media, Inc\.\.Cited by:[§3\.2](https://arxiv.org/html/2609.19650#S3.SS2.p1.1)\.
- Blaschkeet al\.\(2025\)V\. Blaschke, M\. Fedzechkina, and M\. Ter HoeveAnalyzing the effect of linguistic similarity on cross\-lingual transfer: tasks and experimental setups matter\.InFindings of the Association for Computational Linguistics: ACL 2025,Vienna, Austria,pp\. 8653–8684\.External Links:[Link](https://aclanthology.org/2025.findings-acl.454/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.454)Cited by:[§1](https://arxiv.org/html/2609.19650#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.19650#S2.SS2.p1.1)\.
- Bracket al\.\(2024\)A\. Brack, E\. Entrup, M\. Stamatakis, P\. Buschermöhle, A\. Hoppe, and R\. EwerthSequential sentence classification in research papers using cross\-domain multi\-task learning\.International Journal on Digital Libraries25,pp\. 377–400\.External Links:[Document](https://dx.doi.org/10.1007/s00799-023-00392-z)Cited by:[§1](https://arxiv.org/html/2609.19650#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.19650#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2609.19650#S2.SS1.p2.1),[§4\.1](https://arxiv.org/html/2609.19650#S4.SS1.p1.1),[§6\.1](https://arxiv.org/html/2609.19650#S6.SS1.p5.1)\.
- Braudet al\.\(2017\)C\. Braud, M\. Coavoux, and A\. SøgaardCross\-lingual RST discourse parsing\.InProceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers,Valencia, Spain,pp\. 292–304\.External Links:[Link](https://aclanthology.org/E17-1028/)Cited by:[§2\.2](https://arxiv.org/html/2609.19650#S2.SS2.p2.1)\.
- Buchmannet al\.\(2024\)J\. Buchmann, M\. Eichler, J\. Bodensohn, I\. Kuznetsov, and I\. GurevychDocument structure in long document transformers\.InProceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),St\. Julian’s, Malta,pp\. 1056–1073\.External Links:[Link](https://aclanthology.org/2024.eacl-long.64/),[Document](https://dx.doi.org/10.18653/v1/2024.eacl-long.64)Cited by:[§2\.3](https://arxiv.org/html/2609.19650#S2.SS3.p1.1)\.
- Cobbeet al\.\(2021\)K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano, C\. Hesse, and J\. SchulmanTraining verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168\.Cited by:[§2\.3](https://arxiv.org/html/2609.19650#S2.SS3.p1.1),[§5\.2](https://arxiv.org/html/2609.19650#S5.SS2.p1.1)\.
- Cohanet al\.\(2019\)A\. Cohan, I\. Beltagy, D\. King, B\. Dalvi, and D\. WeldPretrained language models for sequential sentence classification\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),Hong Kong, China,pp\. 3693–3699\.External Links:[Link](https://aclanthology.org/D19-1383/),[Document](https://dx.doi.org/10.18653/v1/D19-1383)Cited by:[§1](https://arxiv.org/html/2609.19650#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.19650#S2.SS1.p1.1)\.
- Conneauet al\.\(2020\)A\. Conneau, K\. Khandelwal, N\. Goyal, V\. Chaudhary, G\. Wenzek, F\. Guzmán, E\. Grave, M\. Ott, L\. Zettlemoyer, and V\. StoyanovUnsupervised cross\-lingual representation learning at scale\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),Online,pp\. 8440–8451\.External Links:[Link](https://aclanthology.org/2020.acl-main.747/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.747)Cited by:[§6\.1](https://arxiv.org/html/2609.19650#S6.SS1.p5.1)\.
- De Souzaet al\.\(2024\)L\. De Souza, T\. Almeida, R\. Lotufo, and R\. Frassetto NogueiraMeasuring cross\-lingual transfer in bytes\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),K\. Duh, H\. Gomez, and S\. Bethard \(Eds\.\),Mexico City, Mexico,pp\. 7526–7537\.External Links:[Link](https://aclanthology.org/2024.naacl-long.418/),[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.418)Cited by:[§1](https://arxiv.org/html/2609.19650#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.19650#S2.SS2.p1.1)\.
- Dernoncourt and Lee \(2017\)F\. Dernoncourt and J\. Y\. LeePubMed 200k RCT: a dataset for sequential sentence classification in medical abstracts\.InProceedings of the Eighth International Joint Conference on Natural Language Processing \(Volume 2: Short Papers\),Taipei, Taiwan,pp\. 308–313\.Cited by:[§1](https://arxiv.org/html/2609.19650#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.19650#S2.SS1.p1.1),[§3\.1](https://arxiv.org/html/2609.19650#S3.SS1.p1.1)\.
- Deshpandeet al\.\(2022\)A\. Deshpande, P\. Talukdar, and K\. NarasimhanWhen is BERT multilingual? isolating crucial ingredients for cross\-lingual transfer\.InProceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,M\. Carpuat, M\. de Marneffe, and I\. V\. Meza Ruiz \(Eds\.\),Seattle, United States,pp\. 3610–3623\.External Links:[Link](https://aclanthology.org/2022.naacl-main.264),[Document](https://dx.doi.org/10.18653/v1/2022.naacl-main.264)Cited by:[§2\.2](https://arxiv.org/html/2609.19650#S2.SS2.p1.1)\.
- Devlinet al\.\(2019\)J\. Devlin, M\. Chang, K\. Lee, and K\. ToutanovaBERT: pre\-training of deep bidirectional transformers for language understanding\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,pp\. 4171–4186\.External Links:[Link](https://aclanthology.org/N19-1423)Cited by:[§1](https://arxiv.org/html/2609.19650#S1.p1.1),[§1](https://arxiv.org/html/2609.19650#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.19650#S2.SS2.p1.1)\.
- Fenget al\.\(2022\)F\. Feng, Y\. Yang, D\. Cer, N\. Arivazhagan, and W\. WangLanguage\-agnostic BERT sentence embedding\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),S\. Muresan, P\. Nakov, and A\. Villavicencio \(Eds\.\),Dublin, Ireland,pp\. 878–891\.External Links:[Link](https://aclanthology.org/2022.acl-long.62/),[Document](https://dx.doi.org/10.18653/v1/2022.acl-long.62)Cited by:[§6\.1](https://arxiv.org/html/2609.19650#S6.SS1.p5.1)\.
- Gemma Team \(2024\)Gemma TeamGemma 2: improving open language models at a practical size\.arXiv preprint arXiv:2408\.00118\.Cited by:[§4\.1](https://arxiv.org/html/2609.19650#S4.SS1.p2.1)\.
- Genget al\.\(2023\)S\. Geng, M\. Josifoski, M\. Peyrard, and R\. WestGrammar\-constrained decoding for structured NLP tasks without finetuning\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,Singapore,pp\. 10932–10952\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.674/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.674)Cited by:[§2\.3](https://arxiv.org/html/2609.19650#S2.SS3.p1.1)\.
- Honnibalet al\.\(2020\)spaCy: industrial\-strength natural language processing in PythonExternal Links:[Document](https://dx.doi.org/10.5281/zenodo.1212303)Cited by:[§3\.2](https://arxiv.org/html/2609.19650#S3.SS2.p1.1)\.
- Huet al\.\(2022\)E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. ChenLoRA: low\-rank adaptation of large language models\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=nZeVKeeFYf9)Cited by:[§4\.1](https://arxiv.org/html/2609.19650#S4.SS1.p3.1)\.
- Jin and Szolovits \(2018\)D\. Jin and P\. SzolovitsHierarchical neural networks for sequential sentence classification in medical scientific abstracts\.InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,Brussels, Belgium,pp\. 3100–3109\.External Links:[Document](https://dx.doi.org/10.18653/v1/D18-1349)Cited by:[§1](https://arxiv.org/html/2609.19650#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.19650#S2.SS1.p1.1)\.
- Keet al\.\(2017\)G\. Ke, Q\. Meng, T\. Finley, T\. Wang, W\. Chen, W\. Ma, Q\. Ye, and T\. LiuLightGBM: a highly efficient gradient boosting decision tree\.InAdvances in Neural Information Processing Systems,Vol\.30,pp\. 3146–3154\.Cited by:[§6\.1](https://arxiv.org/html/2609.19650#S6.SS1.p4.1)\.
- Laffertyet al\.\(2001\)J\. Lafferty, A\. McCallum, and F\. C\. PereiraConditional random fields: probabilistic models for segmenting and labeling sequence data\.InProceedings of the 18th International Conference on Machine Learning,pp\. 282–289\.Cited by:[§2\.1](https://arxiv.org/html/2609.19650#S2.SS1.p1.1),[§4\.3](https://arxiv.org/html/2609.19650#S4.SS3.p2.1)\.
- Lanet al\.\(2024\)M\. Lan, L\. Zheng, S\. Ming, and H\. KilicogluMulti\-label sequential sentence classification via large language model\.InFindings of the Association for Computational Linguistics: EMNLP 2024,Miami, Florida, USA,pp\. 16086–16104\.External Links:[Link](https://aclanthology.org/2024.findings-emnlp.944/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.944)Cited by:[§1](https://arxiv.org/html/2609.19650#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.19650#S2.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.19650#S4.SS1.p2.1),[§6\.1](https://arxiv.org/html/2609.19650#S6.SS1.p5.1)\.
- Lauscheret al\.\(2020\)A\. Lauscher, V\. Ravishankar, I\. Vulić, and G\. GlavašFrom zero to hero: on the limitations of zero\-shot language transfer with multilingual transformers\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),Online,pp\. 4483–4499\.External Links:[Link](https://aclanthology.org/2020.emnlp-main.363/),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.363)Cited by:[§2\.2](https://arxiv.org/html/2609.19650#S2.SS2.p1.1)\.
- Linet al\.\(2024\)P\. Lin, C\. Hu, Z\. Zhang, A\. Martins, and H\. SchuetzeMPLM\-sim: better cross\-lingual similarity and transfer in multilingual pretrained language models\.InFindings of the Association for Computational Linguistics: EACL 2024,St\. Julian’s, Malta,pp\. 276–310\.Cited by:[§1](https://arxiv.org/html/2609.19650#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.19650#S2.SS2.p1.1)\.
- Linet al\.\(2019\)Y\. Lin, C\. Chen, J\. Lee, Z\. Li, Y\. Zhang, M\. Xia, S\. Rijhwani, J\. He, Z\. Zhang, X\. Ma, A\. Anastasopoulos, P\. Littell, and G\. NeubigChoosing transfer languages for cross\-lingual learning\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,Florence, Italy,pp\. 3125–3135\.External Links:[Document](https://dx.doi.org/10.18653/v1/P19-1301)Cited by:[§1](https://arxiv.org/html/2609.19650#S1.p2.1)\.
- Littellet al\.\(2017\)P\. Littell, D\. R\. Mortensen, K\. Lin, K\. Kairis, C\. Turner, and L\. LevinURIEL and lang2vec: representing languages as typological, geographical, and phylogenetic vectors\.InProceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers,pp\. 8–14\.Cited by:[§4\.3](https://arxiv.org/html/2609.19650#S4.SS3.p1.1)\.
- Llama Team, AI @ Meta \(2024\)Llama Team, AI @ MetaThe Llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.Cited by:[§4\.1](https://arxiv.org/html/2609.19650#S4.SS1.p2.1)\.
- Martín\-Martín \(2003\)P\. Martín\-MartínA genre analysis of english and spanish research paper abstracts in experimental social sciences\.English for Specific Purposes22\(1\),pp\. 25–43\.External Links:[Document](https://dx.doi.org/10.1016/S0889-4906%2801%2900033-3)Cited by:[§2\.2](https://arxiv.org/html/2609.19650#S2.SS2.p2.1),[§4\.3](https://arxiv.org/html/2609.19650#S4.SS3.p2.1)\.
- Parket al\.\(2024\)K\. Park, J\. Wang, T\. Berg\-Kirkpatrick, N\. Polikarpova, and L\. D’AntoniGrammar\-aligned decoding\.InAdvances in Neural Information Processing Systems 37 \(NeurIPS 2024\),Cited by:[§2\.3](https://arxiv.org/html/2609.19650#S2.SS3.p1.1)\.
- Philippyet al\.\(2023\)F\. Philippy, S\. Guo, and S\. HaddadanTowards a common understanding of contributing factors for cross\-lingual transfer in multilingual language models: a review\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Toronto, Canada,pp\. 5877–5891\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.323)Cited by:[§1](https://arxiv.org/html/2609.19650#S1.p2.1)\.
- Qwen Team \(2024\)Qwen TeamQwen2\.5 technical report\.arXiv preprint arXiv:2412\.15115\.Cited by:[§4\.1](https://arxiv.org/html/2609.19650#S4.SS1.p2.1)\.
- Shuyo \(2010\)N\. ShuyoLanguage detection library for Java\.External Links:[Link](https://github.com/shuyo/language-detection)Cited by:[§3\.2](https://arxiv.org/html/2609.19650#S3.SS2.p1.1)\.
- Sollaci and Pereira \(2004\)L\. B\. Sollaci and M\. G\. PereiraThe introduction, methods, results, and discussion \(IMRAD\) structure: a fifty\-year survey\.Journal of the Medical Library Association92\(3\),pp\. 364–367\.Note:PMID 15243643, PMCID PMC442179Cited by:[§1](https://arxiv.org/html/2609.19650#S1.p3.1)\.
- Specia and Shah \(2018\)L\. Specia and K\. ShahMachine translation quality estimation: applications and future perspectives\.InTranslation Quality Assessment: From Principles to Practice,J\. Moorkens, S\. Castilho, F\. Gaspari, and S\. Doherty \(Eds\.\),pp\. 201–235\.External Links:ISBN 978\-3\-030\-08206\-2,[Document](https://dx.doi.org/10.1007/978-3-319-91241-7%5F10)Cited by:[§5\.2](https://arxiv.org/html/2609.19650#S5.SS2.p1.1)\.
- Sunet al\.\(2024\)Y\. Sun, G\. Chen, C\. Yang, J\. Bao, B\. Liang, X\. Zeng, M\. Yang, and R\. XuDiscourse structure\-aware prefix for generation\-based end\-to\-end argumentation mining\.InFindings of the Association for Computational Linguistics: ACL 2024,Bangkok, Thailand,pp\. 11597–11613\.External Links:[Link](https://aclanthology.org/2024.findings-acl.689/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.689)Cited by:[§2\.3](https://arxiv.org/html/2609.19650#S2.SS3.p1.1)\.
- Van Bonn and Swales \(2007\)S\. Van Bonn and J\. M\. SwalesEnglish and french journal abstracts in the language sciences: three exploratory studies\.Journal of English for Academic Purposes6\(2\),pp\. 93–108\.External Links:[Document](https://dx.doi.org/10.1016/j.jeap.2007.04.001)Cited by:[§2\.2](https://arxiv.org/html/2609.19650#S2.SS2.p2.1)\.
- Wanget al\.\(2023\)X\. Wang, J\. Wei, D\. Schuurmans, Q\. Le, E\. H\. Chi, S\. Narang, A\. Chowdhery, and D\. ZhouSelf\-consistency improves chain of thought reasoning in language models\.InProceedings of the Eleventh International Conference on Learning Representations \(ICLR\),External Links:[Link](https://openreview.net/forum?id=1PL1NIMMrw)Cited by:[§5\.3](https://arxiv.org/html/2609.19650#S5.SS3.p1.1)\.
- Wanget al\.\(2024\)Z\. Wang, L\. F\. R\. Ribeiro, A\. Papangelis, R\. Mukherjee, T\. Wang, X\. Zhao, A\. Biswas, J\. Caverlee, and A\. MetallinouFANTAstic SEquences and where to find them: faithful and efficient API call generation through state\-tracked constrained decoding and reranking\.InFindings of the Association for Computational Linguistics: EMNLP 2024,Miami, Florida, USA,pp\. 6179–6191\.External Links:[Link](https://aclanthology.org/2024.findings-emnlp.359/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.359)Cited by:[§2\.3](https://arxiv.org/html/2609.19650#S2.SS3.p1.1)\.
- Yakhontova \(2002\)T\. Yakhontova“Selling” or “telling”? the issue of cultural variation in research genres\.InAcademic Discourse,J\. Flowerdew \(Ed\.\),pp\. 216–232\.Cited by:[§2\.2](https://arxiv.org/html/2609.19650#S2.SS2.p2.1)\.
- Yunet al\.\(2023\)T\. Yun, J\. Kim, D\. Kang, S\. Lim, J\. Kim, and T\. KimX\-SNS: cross\-lingual transfer prediction through sub\-network similarity\.InFindings of the Association for Computational Linguistics: EMNLP 2023,Singapore,pp\. 13131–13144\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.findings-emnlp.875)Cited by:[§1](https://arxiv.org/html/2609.19650#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.19650#S2.SS2.p1.1)\.
- Zeyreket al\.\(2020\)D\. Zeyrek, A\. Mendes, Y\. Grishina, M\. Kurfalı, S\. Gibbon, and M\. OgrodniczukTED multilingual discourse bank \(ted\-mdb\): a parallel corpus annotated in the pdtb style\.Language Resources and Evaluation54\(2\),pp\. 587–613\.External Links:[Document](https://dx.doi.org/10.1007/s10579-019-09445-9)Cited by:[§2\.2](https://arxiv.org/html/2609.19650#S2.SS2.p2.1)\.

Similar Articles

An In-Vitro Study on Cross-Lingual Generalization in Language Models

arXiv cs.CL

This paper introduces an in-vitro framework with two procedurally generated languages to study cross-lingual generalization in language models, finding that tokenization's preservation of reusable substructure is more critical than lexical similarity or data balance for transferring capabilities across languages.

Structural priors for data-efficient language learning

arXiv cs.CL

This paper investigates structural transfer from non-language data to improve data efficiency in language modeling, finding that while it reduces next-token prediction loss, it does not reliably enhance downstream linguistic performance.

Improving Cross-Lingual Token Representations by Adding a Pinch of SALT

arXiv cs.CL

The paper proposes SALT, a lightweight post-training method that injects span-level supervision into cross-lingual sentence encoders to improve token representations, achieving top results on multilingual token-level benchmarks and enhancing sentence-level performance.