Rate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic Content
Summary
This paper compares subword tokens, raw bytes, and rendered pixels as text encodings for language models under controlled linguistic content across 13 languages. It traces rate–utility frontiers and finds that no encoding dominates across tasks, with pixels preserving surface form best, bytes preserving cross-lingual alignment best, and tokens supporting topic prediction best.
View Cached Full Text
Cached at: 07/20/26, 09:37 AM
# Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic Content
Source: [https://arxiv.org/html/2607.16117](https://arxiv.org/html/2607.16117)
## Rate–Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic Content
Ingo Ziegler Martin Krebs Desmond Elliott Department of Computer Science, University of Copenhagen inzi@di\.ku\.dk,lsw275@alumni\.ku\.dk,de@di\.ku\.dk
###### Abstract
Language models encode text as subword tokens, raw bytes, or rendered pixels, but these encodings are usually compared under modeling constraints that expose different amounts of linguistic content to models across different languages\. We instead ask what each encoding preserves when both the content and the downstream capacity are controlled\. Using verified parallel sentences across thirteen languages and five scripts, we compare tokens, bytes, and pixels through a shared bottleneck whose width is swept to trace rate–utility frontiers\. This separates three quantities that are often conflated: the number of input positions an encoding creates, the latent capacity available after encoding, and the task\-relevant information that survives compression\. We evaluate three utilities: surface form preservation, cross\-lingual sentence alignment, and topic classification\. No encoding dominates across tasks or capacity regimes\. Pixels preserve surface form best, bytes preserve cross\-lingual alignment best, especially in same\-script multilingual settings, and tokens support topic prediction best\. These performances are not explained by sequence length alone\. Short inputs can discard useful meaning, while long inputs can preserve information that compresses well\. Choosing an encoding is therefore not a fixed preference for tokens, bytes, or pixels, but a rate–utility tradeoff that depends on the task, language mix, capacity regime, and compute budget\.
## 1Introduction
Before a language model can process text, the text must be encoded as a sequence of units\. Most models use subword tokens\(Radford et al\.,[2019](https://arxiv.org/html/2607.16117#bib.bib35); Grattafiori et al\.,[2024](https://arxiv.org/html/2607.16117#bib.bib12); Yang et al\.,[2025](https://arxiv.org/html/2607.16117#bib.bib53); GLM\-5\-Team et al\.,[2026](https://arxiv.org/html/2607.16117#bib.bib11)\), some read raw bytes\(Xue et al\.,[2022](https://arxiv.org/html/2607.16117#bib.bib52); Yu et al\.,[2023](https://arxiv.org/html/2607.16117#bib.bib54); Pagnoni et al\.,[2025](https://arxiv.org/html/2607.16117#bib.bib33)\), and a few render text as images and process pixels\(Salesky et al\.,[2021](https://arxiv.org/html/2607.16117#bib.bib41); Rust et al\.,[2023](https://arxiv.org/html/2607.16117#bib.bib40); Kesen et al\.,[2025](https://arxiv.org/html/2607.16117#bib.bib17)\)\. Although these encodings preserve linguistic content, they change the representation the model sees\. Tokens form short sequences of meaning\-bearing units from a learned but fixed vocabulary\(Sennrich et al\.,[2016](https://arxiv.org/html/2607.16117#bib.bib42)\)\. Bytes form longer sequences from a small, universal alphabet\(Xue et al\.,[2022](https://arxiv.org/html/2607.16117#bib.bib52)\)\. Pixels form a two\-dimensional image, with spatial layout and visual detail but no discrete symbols\(Rust et al\.,[2023](https://arxiv.org/html/2607.16117#bib.bib40)\)\.
These differences widen across languages and scripts, which makes encodings hard to compare\. Prior work has shown that pre\-trained tokenizers expose unequal amounts of linguistic content across languages\(Petrov et al\.,[2023](https://arxiv.org/html/2607.16117#bib.bib34); Ahia et al\.,[2023](https://arxiv.org/html/2607.16117#bib.bib2)\)\. Figure[1](https://arxiv.org/html/2607.16117#S1.F1)shows that the problem generalizes beyond tokenization\. A fixed sequence length, batch size, or compute budget can hold different amounts of content depending on both the encoding and the language\(Petrov et al\.,[2023](https://arxiv.org/html/2607.16117#bib.bib34)\)\. A difference measured between two encodings may therefore reflect how much content the model was shown, or how much capacity it was given, rather than what the encoding preserves\.
A controlled comparison needs two quantities held constant\. First, the content must be identical, so that nothing but the encoding varies\. Second, the downstream capacity must be controlled, so that an encoding is judged by how well it uses a given budget, not by how large a budget it takes\. This changes the central question from which encoding is shorter to which encoding preserves useful task\-relevant information under compression\.
We meet both conditions directly\. We work with SIB\-200\(Adelani et al\.,[2024](https://arxiv.org/html/2607.16117#bib.bib1)\), a human\-translated parallel dataset, where the same meaning appears in every language, and we pass every encoding through a shared bottleneck of fixed width\. Complementary toAhia et al\. \([2023](https://arxiv.org/html/2607.16117#bib.bib2)\)andPetrov et al\. \([2023](https://arxiv.org/html/2607.16117#bib.bib34)\), we re\-train tokenizers on matched content for each language regime, so the token baseline is adapted to the same languages as the byte and pixel encodings\. This lets us separate three quantities that are often collapsed: how many input positions an encoding creates, how much latent capacity we allow after encoding, and which task\-relevant information survives compression\. We call the first the*source rate*, sweep the second by varying the bottleneck width, and measure the third through three utilities that probe complementary uses of text representations: preserving form, aligning translations, and supporting downstream semantic prediction\. We evaluate these quantities across language regimes that range from a single language, to five languages in one script, to five scripts at once\.
We find that the utility frontiers do not follow source rate\. Instead, encoding comparisons depend strongly on how content and capacity are controlled\. A short encoding can misallocate capacity, and a long encoding can preserve structure that compresses well\. The preferred encoding changes with the task and the capacity regime: pixels preserve surface form best, bytes preserve cross\-lingual alignment most reliably, and tokens support topic prediction best\. The same script can also change roles across utilities\. Chinese is difficult to align across scripts, especially for pixels, yet its logographic form narrows the gap between encodings for topic prediction\.
Figure 1:The same linguistic content produces different*source rates*across encodings and languages\. Each column shows a translation of the same SIB\-200 sentence, with token, byte, and pixel patch lengths annotated\. A fixed input length does not expose equal content across representations to a model\. For example, Chinese is longer than English in tokens, even with tokenizers re\-trained on matched\-content and regime specific languages, but shorter in bytes and patches\. Our experiments examine how source rate differences interact with bottleneck capacity across tasks that require different kinds of information\.Our results demonstrate that a cheaper encoding is not necessarily a better one, and what counts as useful depends on what one wants to model\. Our contributions are as follows:
- •We introduce a controlled rate–utility comparison for text encodings\. The comparison holds linguistic content fixed with parallel sentences, fixes downstream capacity with a shared bottleneck, and traces utility as capacity changes, together with the training compute each frontier requires\.
- •We show that encoding preferences are utility\-dependent\. Pixels preserve surface form, bytes preserve cross\-lingual alignment, and tokens support topic prediction, with script mix and bottleneck capacity changing the margins\.
- •We measure source rate for tokens, bytes, and pixels on identical linguistic content across thirteen languages and five scripts, confirming and extending prior findings on length disparities to rendered text and matched\-content tokenizer regimes\.
## 2Related Work
#### Tokenization and multilingual imbalance\.
Prior work has shown that multilingual tokenizers differ in fertility\(Nayeem et al\.,[2025](https://arxiv.org/html/2607.16117#bib.bib30)\), vocabulary coverage\(Limisiewicz et al\.,[2023](https://arxiv.org/html/2607.16117#bib.bib26)\), and downstream quality across languages\(Rust et al\.,[2021](https://arxiv.org/html/2607.16117#bib.bib39)\), especially for low\-resource\(Raj S et al\.,[2025](https://arxiv.org/html/2607.16117#bib.bib37)\)and morphologically rich languages\(Asgari et al\.,[2025](https://arxiv.org/html/2607.16117#bib.bib4)\)and for scripts underrepresented in the tokenizer training data\(Limisiewicz et al\.,[2023](https://arxiv.org/html/2607.16117#bib.bib26)\)\. Closest to our source\-rate analysis,Petrov et al\. \([2023](https://arxiv.org/html/2607.16117#bib.bib34)\)andAhia et al\. \([2023](https://arxiv.org/html/2607.16117#bib.bib2)\)measure length disparities on parallel text from FLORES\-200\(NLLB Team et al\.,[2022](https://arxiv.org/html/2607.16117#bib.bib32)\)\. They show that the same content can cost more tokens in one language than another, withPetrov et al\. \([2023](https://arxiv.org/html/2607.16117#bib.bib34)\)also reporting similar gaps at the byte and character level\. These studies focus on pre\-trained tokenizers and frame length disparity as a question of fairness, access, and commercial cost\. We re\-train a tokenizer per language regime on matched parallel text, which isolates source rate also from the tokenizer’s own training data, and we add rendered pixels as a third encoding\. Centrally, we move past length itself by asking which task\-relevant information each encoding preserves once every encoding passes through the same capacity bottleneck\.
#### Vocabulary\-free encoders\.
These replace learned subword vocabularies with smaller atomic units such as UTF\-8 bytes or Unicode characters\. ByT5\(Xue et al\.,[2022](https://arxiv.org/html/2607.16117#bib.bib52)\)uses uncompressed bytes, while CANINE\(Clark et al\.,[2022](https://arxiv.org/html/2607.16117#bib.bib7)\)downsamples characters before its deeper layers\. Pixel models\(Salesky et al\.,[2021](https://arxiv.org/html/2607.16117#bib.bib41); Rust et al\.,[2023](https://arxiv.org/html/2607.16117#bib.bib40)\)drop symbols entirely by rendering text to images and processing pixel patches, and Pix2Struct\(Lee et al\.,[2023](https://arxiv.org/html/2607.16117#bib.bib24)\)reads language directly from rendered screenshots\. Both pay a structural cost: bytes in sequence length and pixels in a two\-dimensional image\. These works ask whether such models can match a token model after changing the architecture, objective, or scale\. We ask instead what each encoding preserves when architecture, content, and capacity are held fixed\.
## 3Experimental Setup
### 3\.1Controlled content and language regimes
Our comparison controls for content, so that any differences we measure come from the encoding, not from the content shown to the model\. Therefore, all experiments use SIB\-200\(Adelani et al\.,[2024](https://arxiv.org/html/2607.16117#bib.bib1)\), a topic\-labeled, human\-translated parallel dataset\. Each sentence appears in every language with the same topic label, so content and labels are aligned across languages by design\. We use the standard splits: 701 training, 99 validation, and 204 test sentences per language\.
We study five language regimes that isolate one factor at a time\. Two are monolingual:English, written in the alphabetic Latin script, andChinese, written in the logographic Han script\. Two are multilingual within a single script:Latin\-5\(English, Spanish, Turkish, Indonesian, and Swahili\) andCyrillic\-5\(Russian, Bulgarian, Ukrainian, Serbian, and Kazakh\)\. The fifth,Multiscript\-5, pairs five languages with five scripts: English in Latin, Russian in Cyrillic, Arabic in Arabic, Hindi in Devanagari, and Chinese in Han\. Each multilingual regime holds the same sentences as a monolingual one, repeated once per language, so it adds languages but no new content\. The monolingual regimes remove multilinguality, the same\-script regimes add languages while holding the script fixed, and the multiscript regime adds scripts\. Considering each independently separates the effect of language count from the effect of script diversity\.
We use three probes, each corresponding to a different use of text representations, to measure which information each encoding preserves under compression\. Form preservation asks whether the written sentence itself can be recovered\. Cross\-lingual retrieval asks whether translations of the same sentence stay close in representation space\. Topic classification asks whether the compressed representation retains enough information to support a downstream prediction head\. Each task is defined in full where it is used in Section[4](https://arxiv.org/html/2607.16117#S4)\.
These probes matter for different applications\. Form matters when the surface itself is part of the task, as in OCR\(Kim et al\.,[2022](https://arxiv.org/html/2607.16117#bib.bib19); Li et al\.,[2023](https://arxiv.org/html/2607.16117#bib.bib25)\), scene text\(Jang et al\.,[2026](https://arxiv.org/html/2607.16117#bib.bib16)\), document understanding\(Lee et al\.,[2023](https://arxiv.org/html/2607.16117#bib.bib24); Kim et al\.,[2022](https://arxiv.org/html/2607.16117#bib.bib19)\), spelling\-sensitive processing\(Salesky et al\.,[2021](https://arxiv.org/html/2607.16117#bib.bib41); Vesalainen et al\.,[2026](https://arxiv.org/html/2607.16117#bib.bib48)\), and visual text generation\(Liu et al\.,[2023](https://arxiv.org/html/2607.16117#bib.bib27);[2024](https://arxiv.org/html/2607.16117#bib.bib28)\)\. Cross\-lingual meaning matters when representations require semantic equivalence across scripts and languages, such as bitext mining\(Heffernan et al\.,[2022](https://arxiv.org/html/2607.16117#bib.bib13); Artetxe & Schwenk,[2019](https://arxiv.org/html/2607.16117#bib.bib3)\), multilingual sentence embedding\(Reimers & Gurevych,[2020](https://arxiv.org/html/2607.16117#bib.bib38); Enevoldsen et al\.,[2025](https://arxiv.org/html/2607.16117#bib.bib10)\), or cross\-lingual retrieval\(Zhang et al\.,[2023](https://arxiv.org/html/2607.16117#bib.bib57)\)\. Topic classification captures the setting closest to encoder\-style inference\(Wang et al\.,[2018](https://arxiv.org/html/2607.16117#bib.bib49); Devlin et al\.,[2019](https://arxiv.org/html/2607.16117#bib.bib8)\), where a fixed representation must support a classification head\.
### 3\.2Encodings
We compare three encodings of the same content\. Tokens expose a learned segmentation: frequent subwords become single units, which compresses common text but fixes a vocabulary\. Bytes expose a fixed universal alphabet: every string is a sequence of UTF\-8 bytes, at the cost of more positions\. Pixels expose visual form: text is rendered to an image, which removes symbolic units but adds surface detail such as spacing and glyph shape\.
Tokens\.We use SentencePiece\(Kudo & Richardson,[2018](https://arxiv.org/html/2607.16117#bib.bib23)\)with a Unigram model\(Kudo,[2018](https://arxiv.org/html/2607.16117#bib.bib22)\)and a vocabulary of 32,000\. We train one tokenizer per language regime on the union of Bible translations\(Christodouloupoulos & Steedman,[2015](https://arxiv.org/html/2607.16117#bib.bib6); Wordproject,[2026](https://arxiv.org/html/2607.16117#bib.bib50); HTML Bible,[2026](https://arxiv.org/html/2607.16117#bib.bib14)\)for that regime’s languages, then apply it to SIB\-200\. As the Bible translations are parallel, tokenizer training is matched across languages within each regime\. The shift from Bible text to SIB\-200 leaves some vocabulary slots unused due to the small dataset size\.
Bytes\.We encode text as raw UTF\-8 bytes, a fixed alphabet of 256 values\.
Pixels\.We render text in grayscale with the size 16 Noto Sans fonts onto a grid with maximum height of 36 pixels\. We slice each rendered line into non\-overlapping patches of36×3236\\times 32pixels, so each patch is 32 pixels wide, and treat a sentence as a one\-dimensional sequence of patches\(Rust et al\.,[2023](https://arxiv.org/html/2607.16117#bib.bib40)\)\. We render right\-to\-left scripts in their natural order and never truncate across any modality\. Shorter lines are padded to the regime’s maximum sequence length\. Additional details on encodings can be found in Appendix[C\.1](https://arxiv.org/html/2607.16117#A3.SS1)\.
### 3\.3Encoder model and capacity bottleneck
All experiments share the same Transformer\-based\(Vaswani et al\.,[2017](https://arxiv.org/html/2607.16117#bib.bib47)\)encoder architecture, so that a difference between encodings reflects the encoding and not the model\. Only the part attached after the bottleneck changes across experiments, and we describe those parts where they are used\.
Each encoding is first mapped to its embedding\. Tokens and bytes use an embedding table\. Pixels use a linear projection of each36×3236\\times 32patch, equivalent to a non\-overlapping convolution\(Rust et al\.,[2023](https://arxiv.org/html/2607.16117#bib.bib40)\)\. The resulting sequence passes through a standard pre\-norm Transformer\(Xiong et al\.,[2020](https://arxiv.org/html/2607.16117#bib.bib51)\)encoder\. Across all experiments and encodings, total model sizes vary between 5\.6M and 13\.8M parameters\. Full hyperparameters and additional implementation details are provided in Appendices[C\.1](https://arxiv.org/html/2607.16117#A3.SS1)–[C\.3](https://arxiv.org/html/2607.16117#A3.SS3), and we provide code for all experiments at[github\.com/ziegler\-ingo/rate\-utility\-frontiers](https://github.com/ziegler-ingo/rate-utility-frontiers)\.
We then compress the encoder output through a bottleneck\. The bottleneck is a single learned query that attends to the encoder output and returns one vector, in the style of a Perceiver\(Jaegle et al\.,[2021](https://arxiv.org/html/2607.16117#bib.bib15)\)\. This holds downstream representational capacity fixed across encodings\. Without it, an encoding that exposes more positions could carry more total information, breaking the controlled comparison\. The width of this vector,DD, is the capacity available to every downstream task: all information must pass through it\. We treatDDas the controlled variable and sweep it across fifteen values, from 256 down to 1:\{256,128,64,48,32,24,20,16,12,10,8,6,4,2,1\}\\\{256,128,64,48,32,24,20,16,12,10,8,6,4,2,1\\\}\. A smallerDDforces more compression\. Plotting utility againstDDgives a frontier: one encoding is more compact than another for a given utility if it reaches the same score at a smallerDD, or a higher score at the sameDD\.
This sweep is important because source length and preserved information need not move together\. An encoding can expose many positions but store the information needed for a task in a narrow latent\. Another encoding can expose fewer positions but require more latent capacity to preserve the same utility\. A single bottleneck width would hide this distinction\. The frontier shows how each encoding degrades as capacity is removed, and therefore measures not only whether an encoding works, but how compactly it stores the information a task needs\.
## 4Results
We report results in three parts\. We first measure the source rate of each encoding, a fixed property of the data\. We then train models and trace how much utility each encoding preserves as we vary the bottleneck widthDD, for form \(Section[4\.2](https://arxiv.org/html/2607.16117#S4.SS2)\), alignment \(Section[4\.3](https://arxiv.org/html/2607.16117#S4.SS3)\), and prediction \(Section[4\.4](https://arxiv.org/html/2607.16117#S4.SS4)\), and then measure what each frontier costs to train \(Section[4\.6](https://arxiv.org/html/2607.16117#S4.SS6)\)\. We discuss limitations and future work in Appendix[A](https://arxiv.org/html/2607.16117#A1)\.
Figure 2:Source rate relative to monolingual English, by language and encoding\. Each cell is a language’s mean source rate divided by English’s mean in the same encoding\. Values above1×1\\timesare more expensive than English, below1×1\\timesare cheaper\. The cheapest encoding depends on the language: tokens are cheapest for English yet most expensive for Chinese, while pixel patches are cheapest for Chinese, Arabic, and Hindi\.### 4\.1Source rates
Table 1:Mean units per sentence\.The source rate is the number of positions an encoding produces for a piece of content: tokens after segmentation, bytes, or image patches after rendering\. It varies across encodings and languages even when the content is held constant\(Petrov et al\.,[2023](https://arxiv.org/html/2607.16117#bib.bib34); Ahia et al\.,[2023](https://arxiv.org/html/2607.16117#bib.bib2)\)\. We confirm their findings in our controlled setup, extend them to rendered pixels and to tokenizers trained per regime\. We then ask whether these source\-rate differences translate into utility differences in the following experiments\. We measure on all 1004 SIB\-200\(Adelani et al\.,[2024](https://arxiv.org/html/2607.16117#bib.bib1)\)sentences per language, averaged within each regime\. In raw sequence length the three encodings are ordered consistently: patches are the fewest positions, tokens the middle, and bytes the most \(Table[1](https://arxiv.org/html/2607.16117#S4.T1)\)\.
The cheapest encoding depends on the language\. We make this visible by normalizing each language to monolingual English within each encoding \(Figure[2](https://arxiv.org/html/2607.16117#S4.F2)\)\. Chinese is the clearest case\. A Chinese sentence costs1\.71×1\.71\\timesas many tokens as English, but only0\.94×0\.94\\timesas many bytes and0\.72×0\.72\\timesas many patches\. The ranking inverts: tokens are the most expensive encoding for Chinese and the cheapest for English\. The tokenizer, even when trained only on matched\-content Chinese, segments a Chinese sentence into more units than an English one\. A Han character is three bytes but a whole morpheme, so it needs fewer positions and less visual space than meaning\-aligned alphabetic versions\. The pattern holds across scripts: dense non\-Latin scripts are costly in bytes, with Hindi highest at2\.59×2\.59\\timesEnglish, yet the same scripts are the cheapest in pixels, with Chinese, Arabic, and Hindi all below English \(0\.720\.72–0\.83×0\.83\\timesthe patches\)\.
These source rates show that the cheapest encoding depends on the language and the script mix, so the standard way of comparing encodings is confounded\. A fixed sequence length or compute budget therefore exposes each encoding to a different amount of content\(Petrov et al\.,[2023](https://arxiv.org/html/2607.16117#bib.bib34); Ahia et al\.,[2023](https://arxiv.org/html/2607.16117#bib.bib2)\), and which language is disadvantaged changes with the encoding\. The rest of the paper removes this confound by holding content fixed and varying only the capacity each encoding is allowed\.
Figure 3:Form preservation\. Test Recall@1 for self\-retrieval through a width\-DDbottleneck, by regime and encoding, averaged over five seeds \(shaded band:±1\\pm\{\}1standard deviation\)\. Pixels lead wherever capacity is ample\. In the same\-script multilingual regimes \(Latin\-5, Cyrillic\-5\) bytes overtake below aboutD=10D\{=\}10\. In Multiscript\-5, pixels stay ahead to the smallest latent\.
### 4\.2Form preservation
The form experiment asks how much of a sentence’s surface a single latent vector can hold\. The model encodes one input, compresses it to a latent of widthDD, and is trained to recover its own embedded input from that latent via InfoNCE\(van den Oord et al\.,[2019](https://arxiv.org/html/2607.16117#bib.bib46)\)\. We score recovery by retrieval: a sentence counts as correct if its compressed encoding matches its own input more closely than any other test sentence\. Appendix[C\.3](https://arxiv.org/html/2607.16117#A3.SS3)provides more details on how retrieval is implemented\. We report test Recall@1 across the sweep \(Figure[3](https://arxiv.org/html/2607.16117#S4.F3)\) but select models on the validation set by mean reciprocal rank \(MRR\)\.
Pixels lead wherever capacity is ample\. In the three multilingual regimes a pixel model retrieves almost perfectly down to a width of 48 \(Recall@10\.970\.97–0\.990\.99\), and it stays ahead through the mid\-range: atD=32D\{=\}32pixels reach0\.940\.94–0\.970\.97, while tokens sit at0\.610\.61–0\.710\.71and bytes at0\.500\.50–0\.690\.69\. English and Chinese are closer at the top, but pixels remain the most robust asDDfalls\.
Tokens are competitive only at the widest latents, and they collapse fastest\. In Latin\-5, tokens match pixels atD=256D\{=\}256\(0\.980\.98against0\.990\.99\) but drop from0\.710\.71atD=32D\{=\}32to0\.260\.26atD=16D\{=\}16\. The other multilingual regimes fall just as steeply\. However, bytes behave oppositely\. They are the weakest encoding at the top: atD=256D\{=\}256byte Recall@1 is0\.830\.83in Latin\-5 and about0\.620\.62in Cyrillic\-5 and Multiscript\-5, well below pixels, but they degrade slowly\.
These opposite slopes cross\. In the two same\-script multilingual regimes, bytes overtake both other encodings once the latent is small: below aboutD=10D\{=\}10in Latin\-5 and Cyrillic\-5, bytes lead, and atD=8D\{=\}8byte Recall@1 is0\.340\.34and0\.330\.33against0\.300\.30for pixels and0\.100\.10and0\.060\.06for tokens\. Multiscript\-5 is the exception\. There, pixels stay on top to the bottom of the sweep: atD=8D\{=\}8pixels reach0\.540\.54against0\.330\.33for bytes, and bytes overtake only tokens, never pixels\.
Pixels carry the most distinctive surface, an image of the text, and they are also the shortest sequence to reconstruct \(Section[4\.1](https://arxiv.org/html/2607.16117#S4.SS1)\), so a generous latent recovers them faithfully\. Compression erodes that surface from the fine detail upward: it keeps coarse, script\-level visual features and loses the word\-level detail that separates two sentences written in the same script\. In the same\-script regimes every distractor shares the script, so the only distinctions left are word\-level, which are exactly the details compression destroys\. Therefore, the pixel representations become mutually confusable\. Bytes keep the exact UTF\-8 sequence in a 256\-symbol alphabet and degrade slowly, so they take the lead once the latent is narrow\. Multiscript\-5 differs only in the makeup of the pool and not in the capacity available\. There, four\-fifths of the distractors are written in another script, representing a difference coarse enough to survive compression, so they fall away as easy negatives\. Pixel models then only have to separate the same\-script sentences\.
Sequence length alone does not explain the ordering\. The crossings show that capacity interacts with the structure of each encoding, not only its length\. Tokens make the point from the other side: they are shorter than bytes, yet they collapse first, because each position is part of a large vocabulary and a narrow latent cannot preserve which one\.
Figure 4:Cross\-lingual retrieval\. Test cross\-lingual Recall@1 through a width\-DDbottleneck, by regime and encoding, averaged over five seeds and over the twenty ordered language pairs \(shaded band:±1\\pm\{\}1standard deviation across seeds\)\. Bytes lead across widths and regimes\. Pixels are weakest and nearly flat, while tokens are in\-between\. All encodings drop sharply in the multiscript regime\.
### 4\.3Cross\-lingual retrieval
This experiment asks whether translated sentences can still be matched after compression\. It requires no task\-specific output head, so the bottleneck latent is the sentence representation\. We train it with a contrastive objective that pulls the five parallel versions of a sentence together and pushes other sentences apart\(Khosla et al\.,[2020](https://arxiv.org/html/2607.16117#bib.bib18); van den Oord et al\.,[2019](https://arxiv.org/html/2607.16117#bib.bib46)\)\. For evaluation, each sentence retrieves its translation across every ordered language pair, and we report test Recall@1 averaged over all twenty pairs\. Appendix[C\.3](https://arxiv.org/html/2607.16117#A3.SS3)provides more details on how retrieval is implemented, and we select models by validation MRR\. Only the three multilingual regimes appear, since cross\-lingual retrieval needs at least two languages\. We note that chance Recall@1 is1/204=0\.00491/204\{=\}0\.0049\.
The ranking from the form experiment inverts\. Bytes preserve cross\-lingual alignment best, by a wide margin\. AtD=256D\{=\}256a byte model retrieves at0\.450\.45in Latin\-5 and0\.470\.47in Cyrillic\-5, against0\.270\.27and0\.340\.34for tokens and0\.080\.08and0\.100\.10for pixels \(Figure[4](https://arxiv.org/html/2607.16117#S4.F4)\)\. Bytes also gain the most from capacity: their Recall@1 falls from0\.450\.45to0\.230\.23betweenD=256D\{=\}256andD=8D\{=\}8in Latin\-5, so most of their advantage is information that a wider latent carries\.
Pixels are the weakest encoding at every width, and their curve is nearly flat\. In Latin\-5, pixel Recall@1 stays near0\.070\.07fromD=256D\{=\}256all the way toD=8D\{=\}8, while bytes more than halve and tokens collapse from0\.270\.27to0\.040\.04\. One hypothesis is that pixels carry little semantic signal to begin with, so the bottleneck has little to keep or lose\. Pixels do better within a script than across one:0\.080\.08in Latin\-5 and0\.100\.10in Cyrillic\-5, but only0\.030\.03in Multiscript\-5 at a third of the same\-script level\. Same\-script languages share glyphs and cognate spellings, so two aligned sentences look somewhat alike as images\. However, across five unrelated scripts that surface overlap is gone, and pixel retrieval falls with it\.
Bytes show the mirror image of the pixel pattern\. Their lead is largest where languages share characters, reaching0\.450\.45–0\.470\.47against tokens’0\.270\.27–0\.340\.34\. In Multiscript\-5, however, where no two languages share a script, the byte lead over tokens nearly closes \(0\.150\.15against0\.140\.14\)\. Byte\-level matching appears to exploit character overlap between languages that share a script\. Over five unrelated scripts, that overlap disappears, so the task is hard for every encoding, with even bytes reaching only0\.150\.15, a third of their same\-script performance\.
Figure 5:Topic classification\. Test macro\-F1 for SIB\-200 topic classification through a width\-DDbottleneck, by regime and encoding, averaged over five seeds \(shaded band:±1\\pm\{\}1standard deviation across seeds\)\. Tokens dominate every regime and resist compression best\. Pixels are weakest and nearly flat, except in Chinese, where a meaning\-bearing script lifts them\.
### 4\.4Topic classification
The topic classification experiment asks whether the compressed representation preserves information useful for a downstream semantic decision\. We attach a linear classifier to the bottleneck output and classify each sentence into one of SIB\-200’s seven topics\. We select models based on highest validation macro\-F1\-score, and report test set macro\-F1\. All five language regimes are considered\. Labels are content\-level, so a sentence carries the same topic in every language\. A uniform guess over the seven topics is0\.140\.14macro\-F1\.
Tokens dominate, by the widest margin of any experiment\. AtD=256D\{=\}256, tokens reach0\.460\.46–0\.500\.50macro\-F1 in four regimes and0\.630\.63in Chinese, against0\.260\.26–0\.300\.30for bytes and0\.160\.16–0\.200\.20for pixels \(Figure[5](https://arxiv.org/html/2607.16117#S4.F5)\)\. Tokens are also the most robust to compression as they hold most of their performance down to aboutD=8D\{=\}8\. In Latin\-5, macro\-F1 slips only from0\.470\.47to0\.420\.42and collapses only at the smallest latents\.
Bytes sit in the middle\. Pixels are weakest, and in every regime except Chinese their macro\-F1 stays between0\.160\.16and0\.200\.20across almost the whole sweep, close to the uniform guess floor and nearly flat, mirroring the shape they showed for meaning \(Section[4\.3](https://arxiv.org/html/2607.16117#S4.SS3)\)\. We hypothesize that there is little topic signal in the pixel representation to expose\.
Chinese is the exception, and it is the best performing regime for pixels\. While tokens remain best at 0\.63 macro\-F1, followed by bytes at 0\.46, pixels double to0\.360\.36, marking their largest gain across regimes and closing much of the gap to tokens\. The script is the likely reason\. A Chinese character is a meaning\-bearing unit, so an image of Chinese text shows semantic units directly, whereas an image of alphabetic text shows letter shapes that carry meaning only in combination\. Where surface form coincides with meaning, pixels can recover it\.
The ordering reflects how directly each encoding exposes lexical content for a coarse semantic decision\. Tokens are dense meaning\-bearing units, so a few of them signal a topic and a narrow latent can still preserve which\. Bytes spread the same words over many character positions, so the latent has to reassemble them\. Pixels show surface that, outside a logographic script, is only loosely tied to topic\. This also reverses the ordering from the meaning experiment, where bytes beat tokens\. Character\-level overlap helped match translations across languages, but it does not help a topic decision, whereas lexical density does\.
Figure 6:Per\-pair cross\-lingual Recall@1 atD=256D\{=\}256, with source language as rows and target language as columns, for each encoding in Multiscript\-5 \(averaged over five seeds\)\. Darker cells mark harder retrieval pairs\. Chinese rows and columns are hardest for every encoding\. Pixels are near the1/204=0\.00491/204\{=\}0\.0049floor throughout\. The same logographic structure that isolates Chinese here is what lets pixels classify it in Section[4\.4](https://arxiv.org/html/2607.16117#S4.SS4)\.
### 4\.5Per\-pair alignment, and the Chinese case
Looking at the pair\-level breakdown of cross\-lingual retrieval shows that the averaged rankings are broad\. Figure[6](https://arxiv.org/html/2607.16117#S4.F6)presents the Multiscript\-5 case atD=256D\{=\}256, while Figure[8](https://arxiv.org/html/2607.16117#A2.F8)in Appendix[B](https://arxiv.org/html/2607.16117#A2)provides the full breakdown for all nine multilingual regimes and encodings atD=256D\{=\}256\. Within a regime, the spread across language pairs is smaller than the gap between encodings, with Latin\-5 byte Recall@1 running from 0\.36 to 0\.57 across pairs while pixels do not reach 0\.10\. Where languages are close, every encoding does better, especially the surface\-bound ones: in Cyrillic\-5 the Russian–Ukrainian pair, which is nearly identical in spelling, is the easiest for bytes \(0\.56\) and for pixels \(0\.14 against a 0\.10 average\)\. Chinese is the outlier\. In Multiscript\-5 it is the hardest language to align for every encoding\. Byte Recall@1 on Chinese pairs falls to 0\.09–0\.13, against 0\.16–0\.20 among the other four, and pixels sit between 0\.01 and 0\.04\. A logographic script shares no characters, no subwords, and no visual form with the alphabetic languages, so every encoding struggles to align it\.
Yet, the same property reverses sign for topic classification\. Monolingual Chinese is the one regime where pixels classify well \(0\.36, Section[4\.4](https://arxiv.org/html/2607.16117#S4.SS4)\)\. Chinese characters often correspond to morphemes, so an image of Chinese text exposes units that are more directly tied to lexical meaning than letter shapes in alphabetic scripts\. The same visible structure that isolates Chinese in cross\-lingual retrieval can therefore help pixels preserve topic information\.
### 4\.6Accounting for computation
The frontiers so far hold content and capacity fixed, but they do not show what each encoding costs to train\. We therefore measure the total training FLOPs of each run: the forward and backward cost of one training epoch, including padding, multiplied by the number of epochs until the best validation score was reached\. Removing all padding from this accounting lowers every total but does not change ordering \(Appendix[D](https://arxiv.org/html/2607.16117#A4)\)\.
Table 2:Training FLOPs per input, relative to tokens, atD=256D\{=\}256\.Total FLOPs are nearly independent of the bottleneck width: sweepingDDfrom 1 to 256 changes the cost of a run by about 4%, because the shared encoder dominates and the bottleneck is small\. The cost of a run is therefore set by two factors that are independent of capacity: how many positions the input has, and how many epochs the model trains\. The first factor follows the source rates of Section[4\.1](https://arxiv.org/html/2607.16117#S4.SS1)\. Table[2](https://arxiv.org/html/2607.16117#S4.T2)shows the cost of one input relative to tokens: bytes are the most expensive encoding in every regime except Chinese, and patches the cheapest\. The second factor depends on the task\. For form preservation, token models converge in roughly half the epochs of byte and pixel models\. For cross\-lingual retrieval the order reverses: byte models converge fastest, on average 40 epochs against 57 for tokens, which offsets part of their higher cost per input\. For topic classification all three encodings converge at similar speed\.
Figure[7](https://arxiv.org/html/2607.16117#S4.F7)reports test utility against total training FLOPs atD=256D\{=\}256\. For form preservation, accounting for compute strengthens the pixel result\. Pixels reach the best Recall@1 in every multilingual regime and are among the cheapest models to train, while bytes need 11–22×\\timesthe FLOPs of pixels and still score lower \(0\.62–0\.83 against 0\.99\)\. The byte overtake at narrow latents \(Section[4\.2](https://arxiv.org/html/2607.16117#S4.SS2)\) is similarly expensive: atD=8D\{=\}8in Latin\-5, bytes lead pixels by 0\.34 to 0\.30, but need 4\.7×\\timesthe FLOPs\. For topic classification, tokens are the most accurate and the cheapest at the top of the utility range\. Bytes reach a lower macro\-F1 at a higher cost in every regime, in Latin\-5 0\.26 against tokens’ 0\.47 at 1\.4×\\timestheir FLOPs, so bytes are never the preferred encoding for this task at any budget\. Pixels stay cheapest, and in Chinese they reach over half of token utility \(0\.36 against 0\.63\) at less than half of token compute\.
Cross\-lingual retrieval is the one task where the additional byte compute yields utility that no other encoding reaches\. In the same\-script regimes, bytes reach a Recall@1 of 0\.45–0\.47 against at most 0\.27–0\.34 for tokens, and the faster convergence of bytes keeps the additional cost at only 1\.4–1\.6×\\timesthe token FLOPs\. In Multiscript\-5, the regime without character overlap between languages, this advantage disappears\. Bytes cost 4\.2×\\timesthe FLOPs of tokens for a gain of only 0\.01 Recall@1 \(0\.15 against 0\.14\)\.
Figure 7:Test utility against total training FLOPs atD=256D\{=\}256, by task and encoding, averaged over five seeds\. Each point is one language regime \(En, Zh, La5, Cy5, Ms5\)\. Pixels preserve form best and are among the cheapest to train\. Tokens classify best at fewer FLOPs than bytes, and the cross\-lingual retrieval advantage of bytes costs 1\.4–1\.6×\\timesthe token FLOPs within a script but4\.2×4\.2\\timesacross scripts\.
## 5Conclusion
We compared subword tokens, raw bytes, and rendered pixels under matched linguistic content and matched latent capacity\. This control separates three quantities that standard comparisons often collapse: how many input positions an encoding creates, how much information the model is allowed to keep, and which task\-relevant information survives\. Across thirteen languages and five scripts, these quantities behave differently\. Source rate varies sharply with language, script, and tokenizer composition, so fixed\-length multilingual training does not expose equal content to all languages\.
The rate\-utility frontiers show that no encoding is best in general\. Pixels preserve surface form most compactly\. Bytes preserve cross\-lingual alignment most reliably\. Tokens preserve topic information best\. These differences are not explained by source rate alone\. Short inputs can lose useful meaning, and long inputs can contain structure that compresses well\.
The practical takeaway is therefore conditional\. For visual text, OCR\-like processing, and other settings where the surface matters, pixels are a strong and computationally efficient interface\. For multilingual retrieval and sentence alignment, bytes deserve stronger consideration, especially when languages share a script\. For encoder\-style prediction, tokens remain the strongest default in our experiments\. Input encoding should therefore be chosen alongside the model architecture and training objective rather than treated as a fixed preprocessing step\. Its effectiveness depends on the utility being optimized, the languages being represented, and the capacity available to the system\.
## Acknowledgments
IZ and DE have been supported by the European Union’s Horizon 2020 research and innovation program under grant agreement No\. 101135671 \(TrustLLM\)\. This work was supported by a research grant \(VIL53122\) from VILLUM FONDEN\.
## References
- Adelani et al\. \(2024\)David Ifeoluwa Adelani, Hannah Liu, Xiaoyu Shen, Nikita Vassilyev, Jesujoba O\. Alabi, Yanke Mao, Haonan Gao, and En\-Shiun Annie Lee\.SIB\-200: A simple, inclusive, and big evaluation dataset for topic classification in 200\+ languages and dialects\.In Yvette Graham and Matthew Purver \(eds\.\),*Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pp\. 226–245, St\. Julian’s, Malta, March 2024\. Association for Computational Linguistics\.doi:10\.18653/v1/2024\.eacl\-long\.14\.URL[https://aclanthology\.org/2024\.eacl\-long\.14/](https://aclanthology.org/2024.eacl-long.14/)\.
- Ahia et al\. \(2023\)Orevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai, David Mortensen, Noah Smith, and Yulia Tsvetkov\.Do all languages cost the same? tokenization in the era of commercial language models\.In Houda Bouamor, Juan Pino, and Kalika Bali \(eds\.\),*Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing*, pp\. 9904–9923, Singapore, December 2023\. Association for Computational Linguistics\.doi:10\.18653/v1/2023\.emnlp\-main\.614\.URL[https://aclanthology\.org/2023\.emnlp\-main\.614/](https://aclanthology.org/2023.emnlp-main.614/)\.
- Artetxe & Schwenk \(2019\)Mikel Artetxe and Holger Schwenk\.Massively multilingual sentence embeddings for zero\-shot cross\-lingual transfer and beyond\.*Transactions of the Association for Computational Linguistics*, 7:597–610, 2019\.doi:10\.1162/tacl˙a˙00288\.URL[https://aclanthology\.org/Q19\-1038/](https://aclanthology.org/Q19-1038/)\.
- Asgari et al\. \(2025\)Ehsaneddin Asgari, Yassine El Kheir, and Mohammad Ali Sadraei Javaheri\.Morphbpe: A morpho\-aware tokenizer bridging linguistic complexity for efficient llm training across morphologies, 2025\.URL[https://arxiv\.org/abs/2502\.00894](https://arxiv.org/abs/2502.00894)\.
- Beltagy et al\. \(2020\)Iz Beltagy, Matthew E\. Peters, and Arman Cohan\.Longformer: The long\-document transformer, 2020\.URL[https://arxiv\.org/abs/2004\.05150](https://arxiv.org/abs/2004.05150)\.
- Christodouloupoulos & Steedman \(2015\)Christos Christodouloupoulos and Mark Steedman\.A massively parallel corpus: the bible in 100 languages\.*Lang\. Resour\. Eval\.*, 49\(2\):375–395, June 2015\.ISSN 1574\-020X\.doi:10\.1007/s10579\-014\-9287\-y\.URL[https://doi\.org/10\.1007/s10579\-014\-9287\-y](https://doi.org/10.1007/s10579-014-9287-y)\.
- Clark et al\. \(2022\)Jonathan H\. Clark, Dan Garrette, Iulia Turc, and John Wieting\.Canine: Pre\-training an efficient tokenization\-free encoder for language representation\.*Transactions of the Association for Computational Linguistics*, 10:73–91, 2022\.doi:10\.1162/tacl˙a˙00448\.URL[https://aclanthology\.org/2022\.tacl\-1\.5/](https://aclanthology.org/2022.tacl-1.5/)\.
- Devlin et al\. \(2019\)Jacob Devlin, Ming\-Wei Chang, Kenton Lee, and Kristina Toutanova\.BERT: Pre\-training of deep bidirectional transformers for language understanding\.In Jill Burstein, Christy Doran, and Thamar Solorio \(eds\.\),*Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\)*, pp\. 4171–4186, Minneapolis, Minnesota, June 2019\. Association for Computational Linguistics\.doi:10\.18653/v1/N19\-1423\.URL[https://aclanthology\.org/N19\-1423/](https://aclanthology.org/N19-1423/)\.
- Egli et al\. \(2025\)Eric Egli, Matteo Manica, and Jannis Born\.Multiscale byte language models – a hierarchical architecture for causal million\-length sequence modeling\.In*ICML 2025 Workshop on Long\-Context Foundation Models*, 2025\.URL[https://openreview\.net/forum?id=2r7YTSdWYD](https://openreview.net/forum?id=2r7YTSdWYD)\.
- Enevoldsen et al\. \(2025\)Kenneth Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos, Ashwin Mathur, David Stap, Jay Gala, Wissam Siblini, Dominik Krzemiński, Genta Indra Winata, Saba Sturua, Saiteja Utpala, Mathieu Ciancone, Marion Schaeffer, Diganta Misra, Shreeya Dhakal, Jonathan Rystrøm, Roman Solomatin, Ömer Veysel Çağatan, Akash Kundu, Martin Bernstorff, Shitao Xiao, Akshita Sukhlecha, Bhavish Pahwa, Rafał Poświata, Kranthi Kiran GV, Shawon Ashraf, Daniel Auras, Björn Plüster, Jan Philipp Harries, Loïc Magne, Isabelle Mohr, Dawei Zhu, Hippolyte Gisserot\-Boukhlef, Tom Aarsen, Jan Kostkan, Konrad Wojtasik, Taemin Lee, Marek Suppa, Crystina Zhang, Roberta Rocca, Mohammed Hamdy, Andrianos Michail, John Yang, Manuel Faysse, Aleksei Vatolin, Nandan Thakur, Manan Dey, Dipam Vasani, Pranjal A Chitale, Simone Tedeschi, Nguyen Tai, Artem Snegirev, Mariya Hendriksen, Michael Günther, Mengzhou Xia, Weijia Shi, Xing Han Lù, Jordan Clive, Gayatri K, Maksimova Anna, Silvan Wehrli, Maria Tikhonova, Henil Shalin Panchal, Aleksandr Abramov, Malte Ostendorff, Zheng Liu, Simon Clematide, Lester James Validad Miranda, Alena Fenogenova, Guangyu Song, Ruqiya Bin Safi, Wen\-Ding Li, Alessia Borghini, Federico Cassano, Lasse Hansen, Sara Hooker, Chenghao Xiao, Vaibhav Adlakha, Orion Weller, Siva Reddy, and Niklas Muennighoff\.MMTEB: Massive multilingual text embedding benchmark\.In*The Thirteenth International Conference on Learning Representations*, 2025\.URL[https://openreview\.net/forum?id=zl3pfz4VCV](https://openreview.net/forum?id=zl3pfz4VCV)\.
- GLM\-5\-Team et al\. \(2026\)GLM\-5\-Team, :, Aohan Zeng, Xin Lv, Zhenyu Hou, Zhengxiao Du, Qinkai Zheng, Bin Chen, Da Yin, Chendi Ge, Chenghua Huang, Chengxing Xie, Chenzheng Zhu, Congfeng Yin, Cunxiang Wang, Gengzheng Pan, Hao Zeng, Haoke Zhang, Haoran Wang, Huilong Chen, Jiajie Zhang, Jian Jiao, Jiaqi Guo, Jingsen Wang, Jingzhao Du, Jinzhu Wu, Kedong Wang, Lei Li, Lin Fan, Lucen Zhong, Mingdao Liu, Mingming Zhao, Pengfan Du, Qian Dong, Rui Lu, Shuang\-Li, Shulin Cao, Song Liu, Ting Jiang, Xiaodong Chen, Xiaohan Zhang, Xuancheng Huang, Xuezhen Dong, Yabo Xu, Yao Wei, Yifan An, Yilin Niu, Yitong Zhu, Yuanhao Wen, Yukuo Cen, Yushi Bai, Zhongpei Qiao, Zihan Wang, Zikang Wang, Zilin Zhu, Ziqiang Liu, Zixuan Li, Bojie Wang, Bosi Wen, Can Huang, Changpeng Cai, Chao Yu, Chen Li, Chengwei Hu, Chenhui Zhang, Dan Zhang, Daoyan Lin, Dayong Yang, Di Wang, Ding Ai, Erle Zhu, Fangzhou Yi, Feiyu Chen, Guohong Wen, Hailong Sun, Haisha Zhao, Haiyi Hu, Hanchen Zhang, Hanrui Liu, Hanyu Zhang, Hao Peng, Hao Tai, Haobo Zhang, He Liu, Hongwei Wang, Hongxi Yan, Hongyu Ge, Huan Liu, Huanpeng Chu, Jia’ni Zhao, Jiachen Wang, Jiajing Zhao, Jiamin Ren, Jiapeng Wang, Jiaxin Zhang, Jiayi Gui, Jiayue Zhao, Jijie Li, Jing An, Jing Li, Jingwei Yuan, Jinhua Du, Jinxin Liu, Junkai Zhi, Junwen Duan, Kaiyue Zhou, Kangjian Wei, Ke Wang, Keyun Luo, Laiqiang Zhang, Leigang Sha, Liang Xu, Lindong Wu, Lintao Ding, Lu Chen, Minghao Li, Nianyi Lin, Pan Ta, Qiang Zou, Rongjun Song, Ruiqi Yang, Shangqing Tu, Shangtong Yang, Shaoxiang Wu, Shengyan Zhang, Shijie Li, Shuang Li, Shuyi Fan, Wei Qin, Wei Tian, Weining Zhang, Wenbo Yu, Wenjie Liang, Xiang Kuang, Xiangmeng Cheng, Xiangyang Li, Xiaoquan Yan, Xiaowei Hu, Xiaoying Ling, Xing Fan, Xingye Xia, Xinyuan Zhang, Xinze Zhang, Xirui Pan, Xu Zou, Xunkai Zhang, Yadi Liu, Yandong Wu, Yanfu Li, Yidong Wang, Yifan Zhu, Yijun Tan, Yilin Zhou, Yiming Pan, Ying Zhang, Yinpei Su, Yipeng Geng, Yong Yan, Yonglin Tan, Yuean Bi, Yuhan Shen, Yuhao Yang, Yujiang Li, Yunan Liu, Yunqing Wang, Yuntao Li, Yurong Wu, Yutao Zhang, Yuxi Duan, Yuxuan Zhang, Zezhen Liu, Zhengtao Jiang, Zhenhe Yan, Zheyu Zhang, Zhixiang Wei, Zhuo Chen, Zhuoer Feng, Zijun Yao, Ziwei Chai, Ziyuan Wang, Zuzhou Zhang, Bin Xu, Minlie Huang, Hongning Wang, Juanzi Li, Yuxiao Dong, and Jie Tang\.Glm\-5: from vibe coding to agentic engineering, 2026\.URL[https://arxiv\.org/abs/2602\.15763](https://arxiv.org/abs/2602.15763)\.
- Grattafiori et al\. \(2024\)Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al\-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aurelien Rodriguez, Austen Gregerson, Ava Spataru, Baptiste Roziere, Bethany Biron, Binh Tang, Bobbie Chern, Charlotte Caucheteux, Chaya Nayak, Chloe Bi, Chris Marra, Chris McConnell, Christian Keller, Christophe Touret, Chunyang Wu, Corinne Wong, Cristian Canton Ferrer, Cyrus Nikolaidis, Damien Allonsius, Daniel Song, Danielle Pintz, Danny Livshits, Danny Wyatt, David Esiobu, Dhruv Choudhary, Dhruv Mahajan, Diego Garcia\-Olano, Diego Perino, Dieuwke Hupkes, Egor Lakomkin, Ehab AlBadawy, Elina Lobanova, Emily Dinan, Eric Michael Smith, Filip Radenovic, Francisco Guzmán, Frank Zhang, Gabriel Synnaeve, Gabrielle Lee, Georgia Lewis Anderson, Govind Thattai, Graeme Nail, Gregoire Mialon, Guan Pang, Guillem Cucurell, Hailey Nguyen, Hannah Korevaar, Hu Xu, Hugo Touvron, Iliyan Zarov, Imanol Arrieta Ibarra, Isabel Kloumann, Ishan Misra, Ivan Evtimov, Jack Zhang, Jade Copet, Jaewon Lee, Jan Geffert, Jana Vranes, Jason Park, Jay Mahadeokar, Jeet Shah, Jelmer van der Linde, Jennifer Billock, Jenny Hong, Jenya Lee, Jeremy Fu, Jianfeng Chi, Jianyu Huang, Jiawen Liu, Jie Wang, Jiecao Yu, Joanna Bitton, Joe Spisak, Jongsoo Park, Joseph Rocca, Joshua Johnstun, Joshua Saxe, Junteng Jia, Kalyan Vasuden Alwala, Karthik Prasad, Kartikeya Upasani, Kate Plawiak, Ke Li, Kenneth Heafield, Kevin Stone, Khalid El\-Arini, Krithika Iyer, Kshitiz Malik, Kuenley Chiu, Kunal Bhalla, Kushal Lakhotia, Lauren Rantala\-Yeary, Laurens van der Maaten, Lawrence Chen, Liang Tan, Liz Jenkins, Louis Martin, Lovish Madaan, Lubo Malo, Lukas Blecher, Lukas Landzaat, Luke de Oliveira, Madeline Muzzi, Mahesh Pasupuleti, Mannat Singh, Manohar Paluri, Marcin Kardas, Maria Tsimpoukelli, Mathew Oldham, Mathieu Rita, Maya Pavlova, Melanie Kambadur, Mike Lewis, Min Si, Mitesh Kumar Singh, Mona Hassan, Naman Goyal, Narjes Torabi, Nikolay Bashlykov, Nikolay Bogoychev, Niladri Chatterji, Ning Zhang, Olivier Duchenne, Onur Çelebi, Patrick Alrassy, Pengchuan Zhang, Pengwei Li, Petar Vasic, Peter Weng, Prajjwal Bhargava, Pratik Dubal, Praveen Krishnan, Punit Singh Koura, Puxin Xu, Qing He, Qingxiao Dong, Ragavan Srinivasan, Raj Ganapathy, Ramon Calderer, Ricardo Silveira Cabral, Robert Stojnic, Roberta Raileanu, Rohan Maheswari, Rohit Girdhar, Rohit Patel, Romain Sauvestre, Ronnie Polidoro, Roshan Sumbaly, Ross Taylor, Ruan Silva, Rui Hou, Rui Wang, Saghar Hosseini, Sahana Chennabasappa, Sanjay Singh, Sean Bell, Seohyun Sonia Kim, Sergey Edunov, Shaoliang Nie, Sharan Narang, Sharath Raparthy, Sheng Shen, Shengye Wan, Shruti Bhosale, Shun Zhang, Simon Vandenhende, Soumya Batra, Spencer Whitman, Sten Sootla, Stephane Collot, Suchin Gururangan, Sydney Borodinsky, Tamar Herman, Tara Fowler, Tarek Sheasha, Thomas Georgiou, Thomas Scialom, Tobias Speckbacher, Todor Mihaylov, Tong Xiao, Ujjwal Karn, Vedanuj Goswami, Vibhor Gupta, Vignesh Ramanathan, Viktor Kerkez, Vincent Gonguet, Virginie Do, Vish Vogeti, Vítor Albiero, Vladan Petrovic, Weiwei Chu, Wenhan Xiong, Wenyin Fu, Whitney Meers, Xavier Martinet, Xiaodong Wang, Xiaofang Wang, Xiaoqing Ellen Tan, Xide Xia, Xinfeng Xie, Xuchao Jia, Xuewei Wang, Yaelle Goldschlag, Yashesh Gaur, Yasmine Babaei, Yi Wen, Yiwen Song, Yuchen Zhang, Yue Li, Yuning Mao, Zacharie Delpierre Coudert, Zheng Yan, Zhengxing Chen, Zoe Papakipos, Aaditya Singh, Aayushi Srivastava, Abha Jain, Adam Kelsey, Adam Shajnfeld, Adithya Gangidi, Adolfo Victoria, Ahuva Goldstand, Ajay Menon, Ajay Sharma, Alex Boesenberg, Alexei Baevski, Allie Feinstein, Amanda Kallet, Amit Sangani, Amos Teo, Anam Yunus, Andrei Lupu, Andres Alvarado, Andrew Caples, Andrew Gu, Andrew Ho, Andrew Poulton, Andrew Ryan, Ankit Ramchandani, Annie Dong, Annie Franco, Anuj Goyal, Aparajita Saraf, Arkabandhu Chowdhury, Ashley Gabriel, Ashwin Bharambe, Assaf Eisenman, Azadeh Yazdan, Beau James, Ben Maurer, Benjamin Leonhardi, Bernie Huang, Beth Loyd, Beto De Paola, Bhargavi Paranjape, Bing Liu, Bo Wu, Boyu Ni, Braden Hancock, Bram Wasti, Brandon Spence, Brani Stojkovic, Brian Gamido, Britt Montalvo, Carl Parker, Carly Burton, Catalina Mejia, Ce Liu, Changhan Wang, Changkyu Kim, Chao Zhou, Chester Hu, Ching\-Hsiang Chu, Chris Cai, Chris Tindal, Christoph Feichtenhofer, Cynthia Gao, Damon Civin, Dana Beaty, Daniel Kreymer, Daniel Li, David Adkins, David Xu, Davide Testuggine, Delia David, Devi Parikh, Diana Liskovich, Didem Foss, Dingkang Wang, Duc Le, Dustin Holland, Edward Dowling, Eissa Jamil, Elaine Montgomery, Eleonora Presani, Emily Hahn, Emily Wood, Eric\-Tuan Le, Erik Brinkman, Esteban Arcaute, Evan Dunbar, Evan Smothers, Fei Sun, Felix Kreuk, Feng Tian, Filippos Kokkinos, Firat Ozgenel, Francesco Caggioni, Frank Kanayet, Frank Seide, Gabriela Medina Florez, Gabriella Schwarz, Gada Badeer, Georgia Swee, Gil Halpern, Grant Herman, Grigory Sizov, Guangyi, Zhang, Guna Lakshminarayanan, Hakan Inan, Hamid Shojanazeri, Han Zou, Hannah Wang, Hanwen Zha, Haroun Habeeb, Harrison Rudolph, Helen Suk, Henry Aspegren, Hunter Goldman, Hongyuan Zhan, Ibrahim Damlaj, Igor Molybog, Igor Tufanov, Ilias Leontiadis, Irina\-Elena Veliche, Itai Gat, Jake Weissman, James Geboski, James Kohli, Janice Lam, Japhet Asher, Jean\-Baptiste Gaya, Jeff Marcus, Jeff Tang, Jennifer Chan, Jenny Zhen, Jeremy Reizenstein, Jeremy Teboul, Jessica Zhong, Jian Jin, Jingyi Yang, Joe Cummings, Jon Carvill, Jon Shepard, Jonathan McPhie, Jonathan Torres, Josh Ginsburg, Junjie Wang, Kai Wu, Kam Hou U, Karan Saxena, Kartikay Khandelwal, Katayoun Zand, Kathy Matosich, Kaushik Veeraraghavan, Kelly Michelena, Keqian Li, Kiran Jagadeesh, Kun Huang, Kunal Chawla, Kyle Huang, Lailin Chen, Lakshya Garg, Lavender A, Leandro Silva, Lee Bell, Lei Zhang, Liangpeng Guo, Licheng Yu, Liron Moshkovich, Luca Wehrstedt, Madian Khabsa, Manav Avalani, Manish Bhatt, Martynas Mankus, Matan Hasson, Matthew Lennie, Matthias Reso, Maxim Groshev, Maxim Naumov, Maya Lathi, Meghan Keneally, Miao Liu, Michael L\. Seltzer, Michal Valko, Michelle Restrepo, Mihir Patel, Mik Vyatskov, Mikayel Samvelyan, Mike Clark, Mike Macey, Mike Wang, Miquel Jubert Hermoso, Mo Metanat, Mohammad Rastegari, Munish Bansal, Nandhini Santhanam, Natascha Parks, Natasha White, Navyata Bawa, Nayan Singhal, Nick Egebo, Nicolas Usunier, Nikhil Mehta, Nikolay Pavlovich Laptev, Ning Dong, Norman Cheng, Oleg Chernoguz, Olivia Hart, Omkar Salpekar, Ozlem Kalinli, Parkin Kent, Parth Parekh, Paul Saab, Pavan Balaji, Pedro Rittner, Philip Bontrager, Pierre Roux, Piotr Dollar, Polina Zvyagina, Prashant Ratanchandani, Pritish Yuvraj, Qian Liang, Rachad Alao, Rachel Rodriguez, Rafi Ayub, Raghotham Murthy, Raghu Nayani, Rahul Mitra, Rangaprabhu Parthasarathy, Raymond Li, Rebekkah Hogan, Robin Battey, Rocky Wang, Russ Howes, Ruty Rinott, Sachin Mehta, Sachin Siby, Sai Jayesh Bondu, Samyak Datta, Sara Chugh, Sara Hunt, Sargun Dhillon, Sasha Sidorov, Satadru Pan, Saurabh Mahajan, Saurabh Verma, Seiji Yamamoto, Sharadh Ramaswamy, Shaun Lindsay, Shaun Lindsay, Sheng Feng, Shenghao Lin, Shengxin Cindy Zha, Shishir Patil, Shiva Shankar, Shuqiang Zhang, Shuqiang Zhang, Sinong Wang, Sneha Agarwal, Soji Sajuyigbe, Soumith Chintala, Stephanie Max, Stephen Chen, Steve Kehoe, Steve Satterfield, Sudarshan Govindaprasad, Sumit Gupta, Summer Deng, Sungmin Cho, Sunny Virk, Suraj Subramanian, Sy Choudhury, Sydney Goldman, Tal Remez, Tamar Glaser, Tamara Best, Thilo Koehler, Thomas Robinson, Tianhe Li, Tianjun Zhang, Tim Matthews, Timothy Chou, Tzook Shaked, Varun Vontimitta, Victoria Ajayi, Victoria Montanez, Vijai Mohan, Vinay Satish Kumar, Vishal Mangla, Vlad Ionescu, Vlad Poenaru, Vlad Tiberiu Mihailescu, Vladimir Ivanov, Wei Li, Wenchen Wang, Wenwen Jiang, Wes Bouaziz, Will Constable, Xiaocheng Tang, Xiaojian Wu, Xiaolan Wang, Xilun Wu, Xinbo Gao, Yaniv Kleinman, Yanjun Chen, Ye Hu, Ye Jia, Ye Qi, Yenda Li, Yilin Zhang, Ying Zhang, Yossi Adi, Youngjin Nam, Yu, Wang, Yu Zhao, Yuchen Hao, Yundi Qian, Yunlu Li, Yuzi He, Zach Rait, Zachary DeVito, Zef Rosnbrick, Zhaoduo Wen, Zhenyu Yang, Zhiwei Zhao, and Zhiyu Ma\.The llama 3 herd of models, 2024\.URL[https://arxiv\.org/abs/2407\.21783](https://arxiv.org/abs/2407.21783)\.
- Heffernan et al\. \(2022\)Kevin Heffernan, Onur Çelebi, and Holger Schwenk\.Bitext mining using distilled sentence representations for low\-resource languages\.In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang \(eds\.\),*Findings of the Association for Computational Linguistics: EMNLP 2022*, pp\. 2101–2112, Abu Dhabi, United Arab Emirates, December 2022\. Association for Computational Linguistics\.doi:10\.18653/v1/2022\.findings\-emnlp\.154\.URL[https://aclanthology\.org/2022\.findings\-emnlp\.154/](https://aclanthology.org/2022.findings-emnlp.154/)\.
- HTML Bible \(2026\)HTML Bible\.HTML Bible Index – Ukrainian\.[https://www\.htmlbible\.com/sacrednamebiblecom/ukrainian/index\.htm](https://www.htmlbible.com/sacrednamebiblecom/ukrainian/index.htm), 2026\.Accessed: 2026\-05\.
- Jaegle et al\. \(2021\)Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joao Carreira\.Perceiver: General perception with iterative attention\.In Marina Meila and Tong Zhang \(eds\.\),*Proceedings of the 38th International Conference on Machine Learning*, volume 139 of*Proceedings of Machine Learning Research*, pp\. 4651–4664\. PMLR, 18–24 Jul 2021\.URL[https://proceedings\.mlr\.press/v139/jaegle21a\.html](https://proceedings.mlr.press/v139/jaegle21a.html)\.
- Jang et al\. \(2026\)Leeje Jang, Yijun Lin, Yao\-Yi Chiang, and Jerod Weinman\.Ticls: Tightly coupled language text spotter\.In*Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision \(WACV\)*, pp\. 3730–3740, March 2026\.
- Kesen et al\. \(2025\)Ilker Kesen, Jonas F\. Lotz, Ingo Ziegler, Phillip Rust, and Desmond Elliott\.Multilingual pretraining for pixel language models\.In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng \(eds\.\),*Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing*, pp\. 29594–29611, Suzhou, China, November 2025\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-332\-6\.doi:10\.18653/v1/2025\.emnlp\-main\.1504\.URL[https://aclanthology\.org/2025\.emnlp\-main\.1504/](https://aclanthology.org/2025.emnlp-main.1504/)\.
- Khosla et al\. \(2020\)Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan\.Supervised contrastive learning\.In H\. Larochelle, M\. Ranzato, R\. Hadsell, M\.F\. Balcan, and H\. Lin \(eds\.\),*Advances in Neural Information Processing Systems*, volume 33, pp\. 18661–18673\. Curran Associates, Inc\., 2020\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2020/file/d89a66c7c80a29b1bdbab0f2a1a94af8\-Paper\.pdf](https://proceedings.neurips.cc/paper_files/paper/2020/file/d89a66c7c80a29b1bdbab0f2a1a94af8-Paper.pdf)\.
- Kim et al\. \(2022\)Geewook Kim, Teakgyu Hong, Moonbin Yim, JeongYeon Nam, Jinyoung Park, Jinyeong Yim, Wonseok Hwang, Sangdoo Yun, Dongyoon Han, and Seunghyun Park\.Ocr\-free document understanding transformer\.In*Computer Vision – ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXVIII*, pp\. 498–517, Berlin, Heidelberg, 2022\. Springer\-Verlag\.ISBN 978\-3\-031\-19814\-4\.doi:10\.1007/978\-3\-031\-19815\-1˙29\.URL[https://doi\.org/10\.1007/978\-3\-031\-19815\-1\_29](https://doi.org/10.1007/978-3-031-19815-1_29)\.
- Kingma & Ba \(2015\)Diederik P Kingma and Jimmy Ba\.Adam: A method for stochastic optimization\.In*International Conference on Learning Representations \(ICLR\)*, 2015\.
- Koehn \(2005\)Philipp Koehn\.Europarl: A parallel corpus for statistical machine translation\.In*Proceedings of Machine Translation Summit X: Papers*, pp\. 79–86, Phuket, Thailand, September 13\-15 2005\.URL[https://aclanthology\.org/2005\.mtsummit\-papers\.11/](https://aclanthology.org/2005.mtsummit-papers.11/)\.
- Kudo \(2018\)Taku Kudo\.Subword regularization: Improving neural network translation models with multiple subword candidates\.In Iryna Gurevych and Yusuke Miyao \(eds\.\),*Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pp\. 66–75, Melbourne, Australia, July 2018\. Association for Computational Linguistics\.doi:10\.18653/v1/P18\-1007\.URL[https://aclanthology\.org/P18\-1007/](https://aclanthology.org/P18-1007/)\.
- Kudo & Richardson \(2018\)Taku Kudo and John Richardson\.SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing\.In Eduardo Blanco and Wei Lu \(eds\.\),*Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations*, pp\. 66–71, Brussels, Belgium, November 2018\. Association for Computational Linguistics\.doi:10\.18653/v1/D18\-2012\.URL[https://aclanthology\.org/D18\-2012/](https://aclanthology.org/D18-2012/)\.
- Lee et al\. \(2023\)Kenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu, Fangyu Liu, Julian Martin Eisenschlos, Urvashi Khandelwal, Peter Shaw, Ming\-Wei Chang, and Kristina Toutanova\.Pix2Struct: Screenshot parsing as pretraining for visual language understanding\.In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett \(eds\.\),*Proceedings of the 40th International Conference on Machine Learning*, volume 202 of*Proceedings of Machine Learning Research*, pp\. 18893–18912\. PMLR, 23–29 Jul 2023\.URL[https://proceedings\.mlr\.press/v202/lee23g\.html](https://proceedings.mlr.press/v202/lee23g.html)\.
- Li et al\. \(2023\)Minghao Li, Tengchao Lv, Jingye Chen, Lei Cui, Yijuan Lu, Dinei Florencio, Cha Zhang, Zhoujun Li, and Furu Wei\.Trocr: transformer\-based optical character recognition with pre\-trained models\.In*Proceedings of the Thirty\-Seventh AAAI Conference on Artificial Intelligence and Thirty\-Fifth Conference on Innovative Applications of Artificial Intelligence and Thirteenth Symposium on Educational Advances in Artificial Intelligence*, AAAI’23/IAAI’23/EAAI’23\. AAAI Press, 2023\.ISBN 978\-1\-57735\-880\-0\.doi:10\.1609/aaai\.v37i11\.26538\.URL[https://doi\.org/10\.1609/aaai\.v37i11\.26538](https://doi.org/10.1609/aaai.v37i11.26538)\.
- Limisiewicz et al\. \(2023\)Tomasz Limisiewicz, Jiří Balhar, and David Mareček\.Tokenization impacts multilingual language modeling: Assessing vocabulary allocation and overlap across languages\.In Anna Rogers, Jordan Boyd\-Graber, and Naoaki Okazaki \(eds\.\),*Findings of the Association for Computational Linguistics: ACL 2023*, pp\. 5661–5681, Toronto, Canada, July 2023\. Association for Computational Linguistics\.doi:10\.18653/v1/2023\.findings\-acl\.350\.URL[https://aclanthology\.org/2023\.findings\-acl\.350/](https://aclanthology.org/2023.findings-acl.350/)\.
- Liu et al\. \(2023\)Rosanne Liu, Dan Garrette, Chitwan Saharia, William Chan, Adam Roberts, Sharan Narang, Irina Blok, Rj Mical, Mohammad Norouzi, and Noah Constant\.Character\-aware models improve visual text rendering\.In Anna Rogers, Jordan Boyd\-Graber, and Naoaki Okazaki \(eds\.\),*Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pp\. 16270–16297, Toronto, Canada, July 2023\. Association for Computational Linguistics\.doi:10\.18653/v1/2023\.acl\-long\.900\.URL[https://aclanthology\.org/2023\.acl\-long\.900/](https://aclanthology.org/2023.acl-long.900/)\.
- Liu et al\. \(2024\)Zeyu Liu, Weicong Liang, Zhanhao Liang, Chong Luo, Ji Li, Gao Huang, and Yuhui Yuan\.Glyph\-byt5: A customized text encoder for accurate visual text rendering\.In*Computer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Part LXXV*, pp\. 361–377, Berlin, Heidelberg, 2024\. Springer\-Verlag\.ISBN 978\-3\-031\-73225\-6\.doi:10\.1007/978\-3\-031\-73226\-3˙21\.URL[https://doi\.org/10\.1007/978\-3\-031\-73226\-3\_21](https://doi.org/10.1007/978-3-031-73226-3_21)\.
- Loshchilov & Hutter \(2019\)Ilya Loshchilov and Frank Hutter\.Decoupled weight decay regularization\.In*International Conference on Learning Representations*, 2019\.URL[https://openreview\.net/forum?id=Bkg6RiCqY7](https://openreview.net/forum?id=Bkg6RiCqY7)\.
- Nayeem et al\. \(2025\)Mir Tafseer Nayeem, Sawsan Alqahtani, Md Tahmid Rahman Laskar, Tasnim Mohiuddin, and M Saiful Bari\.Beyond fertility: STRR as a metric for multilingual tokenization evaluation\.In*NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling*, 2025\.URL[https://openreview\.net/forum?id=QphCVdMCyG](https://openreview.net/forum?id=QphCVdMCyG)\.
- Neitemeier et al\. \(2025\)Pit Neitemeier, Björn Deiseroth, Constantin Eichenberg, and Lukas Balles\.Hierarchical autoregressive transformers: Combining byte\- and word\-level processing for robust, adaptable language models\.In*The Thirteenth International Conference on Learning Representations*, 2025\.URL[https://openreview\.net/forum?id=tU074jg2vS](https://openreview.net/forum?id=tU074jg2vS)\.
- NLLB Team et al\. \(2022\)NLLB Team, Marta R\. Costa\-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang\.No language left behind: Scaling human\-centered machine translation, 2022\.URL[https://arxiv\.org/abs/2207\.04672](https://arxiv.org/abs/2207.04672)\.
- Pagnoni et al\. \(2025\)Artidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodriguez, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason E Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, and Srini Iyer\.Byte latent transformer: Patches scale better than tokens\.In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar \(eds\.\),*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pp\. 9238–9258, Vienna, Austria, July 2025\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-251\-0\.doi:10\.18653/v1/2025\.acl\-long\.453\.URL[https://aclanthology\.org/2025\.acl\-long\.453/](https://aclanthology.org/2025.acl-long.453/)\.
- Petrov et al\. \(2023\)Aleksandar Petrov, Emanuele La Malfa, Philip Torr, and Adel Bibi\.Language model tokenizers introduce unfairness between languages\.In A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(eds\.\),*Advances in Neural Information Processing Systems*, volume 36, pp\. 36963–36990\. Curran Associates, Inc\., 2023\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2023/file/74bb24dca8334adce292883b4b651eda\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2023/file/74bb24dca8334adce292883b4b651eda-Paper-Conference.pdf)\.
- Radford et al\. \(2019\)Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever\.Language Models are Unsupervised Multitask Learners, 2019\.URL[https://cdn\.openai\.com/better\-language\-models/language\_models\_are\_unsupervised\_multitask\_learners\.pdf](https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf)\.
- Radford et al\. \(2021\)Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever\.Learning transferable visual models from natural language supervision\.In Marina Meila and Tong Zhang \(eds\.\),*Proceedings of the 38th International Conference on Machine Learning*, volume 139 of*Proceedings of Machine Learning Research*, pp\. 8748–8763\. PMLR, 18–24 Jul 2021\.URL[https://proceedings\.mlr\.press/v139/radford21a\.html](https://proceedings.mlr.press/v139/radford21a.html)\.
- Raj S et al\. \(2025\)Bharath Raj S, Garvit Suri, Vikrant Dewangan, and Raghav Sonavane\.When every token counts: Optimal segmentation for low\-resource language models\.In Hansi Hettiarachchi, Tharindu Ranasinghe, Paul Rayson, Ruslan Mitkov, Mohamed Gaber, Damith Premasiri, Fiona Anting Tan, and Lasitha Uyangodage \(eds\.\),*Proceedings of the First Workshop on Language Models for Low\-Resource Languages*, pp\. 294–308, Abu Dhabi, United Arab Emirates, January 2025\. Association for Computational Linguistics\.URL[https://aclanthology\.org/2025\.loreslm\-1\.24/](https://aclanthology.org/2025.loreslm-1.24/)\.
- Reimers & Gurevych \(2020\)Nils Reimers and Iryna Gurevych\.Making monolingual sentence embeddings multilingual using knowledge distillation\.In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu \(eds\.\),*Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\)*, pp\. 4512–4525, Online, November 2020\. Association for Computational Linguistics\.doi:10\.18653/v1/2020\.emnlp\-main\.365\.URL[https://aclanthology\.org/2020\.emnlp\-main\.365/](https://aclanthology.org/2020.emnlp-main.365/)\.
- Rust et al\. \(2021\)Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, and Iryna Gurevych\.How good is your tokenizer? on the monolingual performance of multilingual language models\.In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli \(eds\.\),*Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\)*, pp\. 3118–3135, Online, August 2021\. Association for Computational Linguistics\.doi:10\.18653/v1/2021\.acl\-long\.243\.URL[https://aclanthology\.org/2021\.acl\-long\.243/](https://aclanthology.org/2021.acl-long.243/)\.
- Rust et al\. \(2023\)Phillip Rust, Jonas F\. Lotz, Emanuele Bugliarello, Elizabeth Salesky, Miryam de Lhoneux, and Desmond Elliott\.Language modelling with pixels\.In*The Eleventh International Conference on Learning Representations*, 2023\.URL[https://openreview\.net/forum?id=FkSp8VW8RjH](https://openreview.net/forum?id=FkSp8VW8RjH)\.
- Salesky et al\. \(2021\)Elizabeth Salesky, David Etter, and Matt Post\.Robust open\-vocabulary translation from visual text representations\.In Marie\-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen\-tau Yih \(eds\.\),*Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing*, pp\. 7235–7252, Online and Punta Cana, Dominican Republic, November 2021\. Association for Computational Linguistics\.doi:10\.18653/v1/2021\.emnlp\-main\.576\.URL[https://aclanthology\.org/2021\.emnlp\-main\.576/](https://aclanthology.org/2021.emnlp-main.576/)\.
- Sennrich et al\. \(2016\)Rico Sennrich, Barry Haddow, and Alexandra Birch\.Neural machine translation of rare words with subword units\.In Katrin Erk and Noah A\. Smith \(eds\.\),*Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pp\. 1715–1725, Berlin, Germany, August 2016\. Association for Computational Linguistics\.doi:10\.18653/v1/P16\-1162\.URL[https://aclanthology\.org/P16\-1162/](https://aclanthology.org/P16-1162/)\.
- Srivastava et al\. \(2014\)Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov\.Dropout: A simple way to prevent neural networks from overfitting\.*Journal of Machine Learning Research*, 15\(56\):1929–1958, 2014\.URL[http://jmlr\.org/papers/v15/srivastava14a\.html](http://jmlr.org/papers/v15/srivastava14a.html)\.
- Su et al\. \(2024\)Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu\.Roformer: Enhanced transformer with rotary position embedding\.*Neurocomput\.*, 568\(C\), February 2024\.ISSN 0925\-2312\.doi:10\.1016/j\.neucom\.2023\.127063\.URL[https://doi\.org/10\.1016/j\.neucom\.2023\.127063](https://doi.org/10.1016/j.neucom.2023.127063)\.
- Tay et al\. \(2022\)Yi Tay, Vinh Q\. Tran, Sebastian Ruder, Jai Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu, and Donald Metzler\.Charformer: Fast character transformers via gradient\-based subword tokenization\.In*International Conference on Learning Representations*, 2022\.URL[https://openreview\.net/forum?id=JtBRnrlOEFN](https://openreview.net/forum?id=JtBRnrlOEFN)\.
- van den Oord et al\. \(2019\)Aaron van den Oord, Yazhe Li, and Oriol Vinyals\.Representation learning with contrastive predictive coding, 2019\.URL[https://arxiv\.org/abs/1807\.03748](https://arxiv.org/abs/1807.03748)\.
- Vaswani et al\. \(2017\)Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin\.Attention is all you need\.In I\. Guyon, U\. Von Luxburg, S\. Bengio, H\. Wallach, R\. Fergus, S\. Vishwanathan, and R\. Garnett \(eds\.\),*Advances in Neural Information Processing Systems*, volume 30\. Curran Associates, Inc\., 2017\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa\-Paper\.pdf](https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf)\.
- Vesalainen et al\. \(2026\)Ari Vesalainen, Eetu Mäkelä, Laura Ruotsalainen, and Mikko Tolonen\.Error patterns in historical ocr: A comparative analysis of trocr and a vision\-language model, 2026\.URL[https://arxiv\.org/abs/2602\.14524](https://arxiv.org/abs/2602.14524)\.
- Wang et al\. \(2018\)Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R\. Bowman\.GLUE: A multi\-task benchmark and analysis platform for natural language understanding\.In Tal Linzen, Grzegorz Chrupała, and Afra Alishahi \(eds\.\),*Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP*, pp\. 353–355, Brussels, Belgium, November 2018\. Association for Computational Linguistics\.doi:10\.18653/v1/W18\-5446\.URL[https://aclanthology\.org/W18\-5446/](https://aclanthology.org/W18-5446/)\.
- Wordproject \(2026\)Wordproject\.The Holy Bible international: Text and audio bibles\.[https://www\.wordproject\.org/](https://www.wordproject.org/), 2026\.Accessed: 2026\-05\.
- Xiong et al\. \(2020\)Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu\.On layer normalization in the transformer architecture\.In Hal Daumé III and Aarti Singh \(eds\.\),*Proceedings of the 37th International Conference on Machine Learning*, volume 119 of*Proceedings of Machine Learning Research*, pp\. 10524–10533\. PMLR, 13–18 Jul 2020\.URL[https://proceedings\.mlr\.press/v119/xiong20b\.html](https://proceedings.mlr.press/v119/xiong20b.html)\.
- Xue et al\. \(2022\)Linting Xue, Aditya Barua, Noah Constant, Rami Al\-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel\.ByT5: Towards a token\-free future with pre\-trained byte\-to\-byte models\.*Transactions of the Association for Computational Linguistics*, 10:291–306, 2022\.doi:10\.1162/tacl˙a˙00461\.URL[https://aclanthology\.org/2022\.tacl\-1\.17/](https://aclanthology.org/2022.tacl-1.17/)\.
- Yang et al\. \(2025\)An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, Le Yu, Lianghao Deng, Mei Li, Mingfeng Xue, Mingze Li, Pei Zhang, Peng Wang, Qin Zhu, Rui Men, Ruize Gao, Shixuan Liu, Shuang Luo, Tianhao Li, Tianyi Tang, Wenbiao Yin, Xingzhang Ren, Xinyu Wang, Xinyu Zhang, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yinger Zhang, Yu Wan, Yuqiong Liu, Zekun Wang, Zeyu Cui, Zhenru Zhang, Zhipeng Zhou, and Zihan Qiu\.Qwen3 technical report, 2025\.URL[https://arxiv\.org/abs/2505\.09388](https://arxiv.org/abs/2505.09388)\.
- Yu et al\. \(2023\)LILI Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, and Mike Lewis\.Megabyte: Predicting million\-byte sequences with multiscale transformers\.In A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(eds\.\),*Advances in Neural Information Processing Systems*, volume 36, pp\. 78808–78823\. Curran Associates, Inc\., 2023\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2023/file/f8f78f8043f35890181a824e53a57134\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2023/file/f8f78f8043f35890181a824e53a57134-Paper-Conference.pdf)\.
- Zaheer et al\. \(2020\)Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed\.Big bird: Transformers for longer sequences\.In H\. Larochelle, M\. Ranzato, R\. Hadsell, M\.F\. Balcan, and H\. Lin \(eds\.\),*Advances in Neural Information Processing Systems*, volume 33, pp\. 17283–17297\. Curran Associates, Inc\., 2020\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2020/file/c8512d142a2d849725f31a9a7a361ab9\-Paper\.pdf](https://proceedings.neurips.cc/paper_files/paper/2020/file/c8512d142a2d849725f31a9a7a361ab9-Paper.pdf)\.
- Zhang & Sennrich \(2019\)Biao Zhang and Rico Sennrich\.Root mean square layer normalization\.In H\. Wallach, H\. Larochelle, A\. Beygelzimer, F\. d'Alché\-Buc, E\. Fox, and R\. Garnett \(eds\.\),*Advances in Neural Information Processing Systems*, volume 32\. Curran Associates, Inc\., 2019\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2019/file/1e8a19426224ca89e83cef47f1e7f53b\-Paper\.pdf](https://proceedings.neurips.cc/paper_files/paper/2019/file/1e8a19426224ca89e83cef47f1e7f53b-Paper.pdf)\.
- Zhang et al\. \(2023\)Xinyu Zhang, Nandan Thakur, Odunayo Ogundepo, Ehsan Kamalloo, David Alfonso\-Hermelo, Xiaoguang Li, Qun Liu, Mehdi Rezagholizadeh, and Jimmy Lin\.MIRACL: A multilingual retrieval dataset covering 18 diverse languages\.*Transactions of the Association for Computational Linguistics*, 11:1114–1131, 2023\.doi:10\.1162/tacl˙a˙00595\.URL[https://aclanthology\.org/2023\.tacl\-1\.63/](https://aclanthology.org/2023.tacl-1.63/)\.
- Ziemski et al\. \(2016\)Michał Ziemski, Marcin Junczys\-Dowmunt, and Bruno Pouliquen\.The United Nations parallel corpus v1\.0\.In Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Sara Goggi, Marko Grobelnik, Bente Maegaard, Joseph Mariani, Helene Mazo, Asuncion Moreno, Jan Odijk, and Stelios Piperidis \(eds\.\),*Proceedings of the Tenth International Conference on Language Resources and Evaluation \(LREC’16\)*, pp\. 3530–3534, Portorož, Slovenia, May 2016\. European Language Resources Association \(ELRA\)\.URL[https://aclanthology\.org/L16\-1561/](https://aclanthology.org/L16-1561/)\.
## Appendix ALimitations and future work
Our comparison is controlled by design\. This control also limits what the experiments can show\. SIB\-200\(Adelani et al\.,[2024](https://arxiv.org/html/2607.16117#bib.bib1)\)gives us parallel sentences across languages, which is essential for matching content, but it is small compared to open\-domain pretraining corpora\. As a result, the token models are not strongly punished for large vocabularies\. The dataset does not fully expose long\-tail morphology, domain terms, spelling noise, code switching, or new scripts\. These are settings where bytes and pixels may have a larger advantage than we observe here\. Future work could repeat the comparison on larger controlled parallel corpora\. Resources such as Europarl\(Koehn,[2005](https://arxiv.org/html/2607.16117#bib.bib21)\)or the United Nations Parallel Corpus\(Ziemski et al\.,[2016](https://arxiv.org/html/2607.16117#bib.bib58)\)provide many more aligned sentences than SIB\-200, which would make the token setting more realistic and better expose morphology, rare words, and domain\-specific vocabulary\.
We measure training FLOPs, but not memory or wall\-clock time\. Byte activation memory grows with sequence length, and our sentence\-length inputs keep the quadratic part of attention\(Vaswani et al\.,[2017](https://arxiv.org/html/2607.16117#bib.bib47)\)small\. At document\-length contexts, the byte costs of Section[4\.6](https://arxiv.org/html/2607.16117#S4.SS6)would therefore grow faster than linearly, so our totals understate the byte disadvantage at scale\.
Our byte encoder follows the direct byte\-processing setup used by ByT5\-style models\(Xue et al\.,[2022](https://arxiv.org/html/2607.16117#bib.bib52)\)\. It attends over the full byte sequence\. This is a simple and useful baseline, but it is not the only possible byte interface\. Hierarchical pooling\(Yu et al\.,[2023](https://arxiv.org/html/2607.16117#bib.bib54); Egli et al\.,[2025](https://arxiv.org/html/2607.16117#bib.bib9); Neitemeier et al\.,[2025](https://arxiv.org/html/2607.16117#bib.bib31)\), local attention\(Beltagy et al\.,[2020](https://arxiv.org/html/2607.16117#bib.bib5); Zaheer et al\.,[2020](https://arxiv.org/html/2607.16117#bib.bib55)\), learned downsampling\(Clark et al\.,[2022](https://arxiv.org/html/2607.16117#bib.bib7); Tay et al\.,[2022](https://arxiv.org/html/2607.16117#bib.bib45)\), or byte\-level compression\(Pagnoni et al\.,[2025](https://arxiv.org/html/2607.16117#bib.bib33)\)could reduce the cost before the bottleneck\. Such designs may change the tradeoff between source rate and utility\.
Our models are encoder\-based and are trained for fixed\-representation utilities\. They do not evaluate autoregressive generation\. The topic classification task \(Section[4\.4](https://arxiv.org/html/2607.16117#S4.SS4)\) is related to masked or encoder\-style inference, where a representation supports a prediction head\. It is only suggestive for decoder\-only language modeling, where the output interface, decoding process, and long\-context behavior introduce additional constraints\.
The pixel experiments also depend on a particular rendering and patching pipeline\. Pixels preserve form well, but they preserve cross\-lingual meaning and topic information poorly in our setup \(Section[4](https://arxiv.org/html/2607.16117#S4)\)\. It points to a useful open problem: pixel representations have attractive source\-rate properties \(Section[4\.1](https://arxiv.org/html/2607.16117#S4.SS1)\), especially for dense scripts and multiscript settings, but they currently spend much of their capacity on surface detail \(Section[4\.2](https://arxiv.org/html/2607.16117#S4.SS2)\)\. Making them more semantic without losing their compact input interface could improve the compute and multilingual tradeoffs of visual text encoders\.
Finally, our experiments use controlled models trained from scratch\. This makes the comparison clean, but it does not replace large\-scale pretraining\. Scale, data mixture, tokenizer training, and hardware optimality of certain architectures and objectives may shift the frontiers\.
## Appendix BPer\-pair cross\-lingual retrieval
Section[4\.3](https://arxiv.org/html/2607.16117#S4.SS3)reports cross\-lingual retrieval averaged over all ordered language pairs\. Figure[8](https://arxiv.org/html/2607.16117#A2.F8)shows the full per\-pair breakdown behind that average\. For every multilingual regime and encoding, it shows test Recall@1 atD=256D\{=\}256for each ordered source\-target pair, averaged over five seeds, with darker cells marking harder pairs and the diagonal left empty because a language is never retrieved against itself\. The per\-pair patterns summarized in the main text are visible directly across the grid\.
Figure 8:Per\-pair cross\-lingual Recall@1 atD=256D\{=\}256for every multilingual regime and encoding, averaged over five seeds\. Rows are regimes \(Latin\-5, Cyrillic\-5, Multiscript\-5\) and columns are encodings \(token, byte, pixel\)\. Within each panel, cell\(A,B\)\(A,B\)is Recall@1 for retrieving the language\-BBtranslation of a language\-AAsentence, and darker cells are harder pairs\.
## Appendix CExperimental details
### C\.1Input encodings and padded source rates
All experiments use fixed padded input lengths for each regime and modality\. For tokens and bytes, the padded length is the maximum number of token or byte positions observed in that regime\. For pixels, the padded length is the maximum number of image patches, rounded up to the next multiple of 32 pixels before patching\.
Table 3:Maximum source lengths per language, grouped by regime\.Boldvalues are the regime\-level maxima\. Token and byte rates are sequence lengths\. Pixel rates are counts of non\-overlapping36×3236\\times 32image patches\.Token inputs are encoded with the SentencePiece\(Kudo & Richardson,[2018](https://arxiv.org/html/2607.16117#bib.bib23); Kudo,[2018](https://arxiv.org/html/2607.16117#bib.bib22)\)model associated with the regime\. Byte inputs are UTF\-8 byte sequences with vocabulary size 256\. Pixel inputs are rendered as grayscale images with height 36\. We render text with Noto Sans fonts at font size 16, using a custom lightweight rendering engine based on HarfBuzz111[https://github\.com/harfbuzz/harfbuzz](https://github.com/harfbuzz/harfbuzz)and FreeType222[https://github\.com/freetype/freetype](https://github.com/freetype/freetype)\. We crop each rendered sentence to its non\-empty horizontal content region, align the crop to patch boundaries, normalize pixel values with mean 0\.09 and standard deviation 0\.25, and pad with the normalized background value\. Pixels are then projected into patch embeddings with a non\-overlapping convolution\(Rust et al\.,[2023](https://arxiv.org/html/2607.16117#bib.bib40)\)over36×3236\\times 32patches\.
### C\.2Model and training details
All experiments use the same encoder architecture\. Each input first passes through a modality\-specific adapter\. Tokens and bytes use learned embedding tables\. Pixels use a non\-overlapping convolutional patch projection\(Rust et al\.,[2023](https://arxiv.org/html/2607.16117#bib.bib40)\)\. The adapted sequence is processed by a 6\-layer pre\-norm Transformer\(Xiong et al\.,[2020](https://arxiv.org/html/2607.16117#bib.bib51)\)encoder with hidden size 256, 4 attention heads, RoPE positional embeddings\(Su et al\.,[2024](https://arxiv.org/html/2607.16117#bib.bib44)\), RMSNorm\(Zhang & Sennrich,[2019](https://arxiv.org/html/2607.16117#bib.bib56)\), and dropout\(Srivastava et al\.,[2014](https://arxiv.org/html/2607.16117#bib.bib43)\)0\.1\. Attention\(Vaswani et al\.,[2017](https://arxiv.org/html/2607.16117#bib.bib47)\)uses a padding mask, so padded positions are ignored\. The encoder output is compressed by a learned\-query bottleneck\(Jaegle et al\.,[2021](https://arxiv.org/html/2607.16117#bib.bib15)\)withL=1L=1latent and bottleneck widthDD\. We sweepD∈256,128,64,48,32,24,20,16,12,10,8,6,4,2,1D\\in\{256,128,64,48,32,24,20,16,12,10,8,6,4,2,1\}, so the bottleneck rate isR=DR=Din all reported runs\. The bottleneck attention usesmax\(1,⌊D/16⌋\)\\max\(1,\\lfloor D/16\\rfloor\)heads\.
Training hyperparameters are shared across experiments unless noted otherwise\. All models are trained for at most 200 epochs with validation after every epoch\. We use AdamW\(Kingma & Ba,[2015](https://arxiv.org/html/2607.16117#bib.bib20); Loshchilov & Hutter,[2019](https://arxiv.org/html/2607.16117#bib.bib29)\)with learning rate10−410^\{\-4\}, weight decay 0\.01, batch size 32, gradient clipping at 1\.0, and a learning\-rate schedule with 15 warmup epochs followed by cosine decay to10−610^\{\-6\}\. Runs use five seeds: 9109, 2943, 3127, 7081, and 4517\. We select checkpoints by validation MRR for form and semantic utility, and by validation macro\-F1 for predictive utility\. Training stops early when the selection metric does not improve for 30 epochs after a minimum training duration of 30 epochs\. For form preservation and cross\-lingual retrieval, we also stop when validation Recall@1 reaches 1\.0\. For predictive utility, we also stop when validation macro\-F1 reaches 1\.0\.
Table[4](https://arxiv.org/html/2607.16117#A3.T4)reports parameter counts for the largest bottleneck setting,D=256D=256andL=1L=1\. Counts include the modality adapter, encoder, bottleneck, and task\-specific head or readout\. They change withDD, but we report only the largest setting for compactness\. The adapter size is shown separately because the encodings differ substantially at the input interface\. Token adapter size depends on the tokenizer vocabulary\. The byte adapter has256×256=65,536256\\times 256=65\{,\}536parameters\. The pixel adapter has256×36×32=294,912256\\times 36\\times 32=294\{,\}912parameters\.
Table 4:Parameter counts for theD=256D=256,L=1L=1configuration\. Values are in millions \(M\) and include adapters, rounded to one decimal\. Token adapters are embedding tables and scale with vocabulary size: 11,653→\\rightarrow3\.0M, 4,265→\\rightarrow1\.1M, 32,000→\\rightarrow8\.2M\. Byte adapter is fixed at 0\.07M \(65,536 params\) and pixel adapter at 0\.29M \(294,912 params\) across all regimes\.
### C\.3Task objectives and evaluation
#### Form preservation\.
The form task measures whether the bottleneck preserves the input representation itself\. For a sentencexx, the model first computes the modality adapter sequence and uses this detached sequence as the target\. The encoder \(Appendix[C\.2](https://arxiv.org/html/2607.16117#A3.SS2)\) then processes the same input, compresses it through the bottleneck\(Jaegle et al\.,[2021](https://arxiv.org/html/2607.16117#bib.bib15)\), and reads the latent back into a sequence\-shaped representation with learned output queries\. We train with a symmetric InfoNCE\(van den Oord et al\.,[2019](https://arxiv.org/html/2607.16117#bib.bib46)\)loss between the readout and the detached target representation\. At evaluation time, every example in the split is encoded once, and each readout retrieves its matching target from the full split\. This task is therefore a contrastive form\-preservation probe\. Models are selected by highest validation MRR because MRR is smoother than Recall@1\. We train with a trainable contrastive temperature initialized at0\.070\.07\(Radford et al\.,[2021](https://arxiv.org/html/2607.16117#bib.bib36)\)\.
#### Cross\-lingual retrieval\.
The cross\-lingual retrieval task measures cross\-lingual alignment under compression\. Each batch containsBBparallel content units andKKlanguages\. We flatten the batch intoBKBKencoded sentences and train with a multi\-positive\(Khosla et al\.,[2020](https://arxiv.org/html/2607.16117#bib.bib18)\)InfoNCE\(van den Oord et al\.,[2019](https://arxiv.org/html/2607.16117#bib.bib46)\)loss\. For each anchor, the positives are all other languages of the same content unit, and the negatives are all sentences from other content units in the batch\. At evaluation time, we build one embedding bank per language\. For every ordered language pairA→BA\\rightarrow B, each sentence in languageAAretrieves its translation from all sentences in languageBB\. We average Recall@1 over allK\(K−1\)K\(K\-1\)ordered language pairs\. Models are selected by highest validation MRR because MRR is smoother than Recall@1\. We train with a trainable contrastive temperature initialized at0\.070\.07\(Radford et al\.,[2021](https://arxiv.org/html/2607.16117#bib.bib36)\)\.
#### Topic classification\.
The classification task measures whether the bottleneck retains information useful for topic classification\. The model compresses the input through the shared bottleneck and uses the single learned query vector to predict the topic with a linear classifier head\. We train with standard cross\-entropy\. We report macro\-F1 and select models by highest validation macro\-F1\.
## Appendix DTraining FLOP accounting
Section[4\.6](https://arxiv.org/html/2607.16117#S4.SS6)reports the total training FLOPs of each run as the cost of one training epoch multiplied by the number of epochs until the best validation score\. This appendix describes how both factors are obtained and shows that the comparison does not depend on padding\.
#### Measurement\.
We measure the cost of one epoch with PyTorch’sFlopCounterMode: we build each model exactly as trained, run one forward and backward pass including the task loss, and count every operation, with a matrix product of shapes\(m,k\)\(m,k\)and\(k,n\)\(k,n\)counted as2mnk2mnkFLOPs\. Every input is padded to its regime’s maximum length \(Appendix[C\.1](https://arxiv.org/html/2607.16117#A3.SS1)\), and attention over padded positions is masked but still computed, so the cost of an epoch is the per\-input cost times the number of training sentences\. The embedding lookup of token and byte models is a gather and costs no FLOPs\. However, its backward pass is a scatter\-add, which we count\. The number of epochs is the epoch of the checkpoint that early stopping selected \(Appendix[C\.2](https://arxiv.org/html/2607.16117#A3.SS2)\)\. Validation and test passes are excluded\.
#### Removing padding\.
The padded totals reflect the computation our runs performed, but they count every sentence at the length of the longest sentence in its regime\. This could distort the comparison: in Multiscript\-5 every byte sequence is padded to Hindi’s maximum of 983 bytes,4\.6×4\.6\\timesthe mean byte length\. We therefore recompute every total with padding removed per\-sequence\. Every counted operation is a matrix product whose dimensions are affine in the sequence length, so the per\-input cost is a degree\-two polynomial in the sequence length, with the quadratic term contributed by attention\. We recover the coefficients of this polynomial from three measured probe lengths and verify them against a fourth measured length, which they match exactly\. Summing the polynomial over the true lengths of all training sentences then gives the cost of an epoch without padding\.
Figure 9:Recalculation of Figure[7](https://arxiv.org/html/2607.16117#S4.F7)under the zero\-padding accounting: test utility against total training FLOPs atD=256D\{=\}256, with every training sentence charged at its own length\. Totals are2\.42\.4–4\.7×4\.7\\timeslower than in Figure[7](https://arxiv.org/html/2607.16117#S4.F7), but conclusions remain unchanged\.Removing padding lowers every total by 2\.4–4\.7×\\timesbut affects the encodings unevenly\. The per\-input cost of bytes relative to tokens falls from 4\.86×\\timesto 3\.55×\\timesin Multiscript\-5, where padding inflated bytes the most, but rises from 3\.78×\\timesto 4\.66×\\timesin Cyrillic\-5, where padding inflated tokens more: the regime’s token maximum of 187 is 3\.5×\\timesits token mean, driven by a single long Serbian tokenization\. Padding therefore does not systematically favor one encoding\. Figure[9](https://arxiv.org/html/2607.16117#A4.F9)repeats Figure[7](https://arxiv.org/html/2607.16117#S4.F7)under the zero\-padding convention\. The conclusions in Section[4\.6](https://arxiv.org/html/2607.16117#S4.SS6)remain unchanged\.
#### Cost decomposition\.
Figure[10](https://arxiv.org/html/2607.16117#A4.F10)separates the two factors behind the totals\. Each point places one regime and encoding by its cost per epoch and its number of epochs to the selected checkpoint, averaged over all bottleneck widths and seeds, with diagonal lines marking equal totals\. The horizontal spread repeats the source rates of Section[4\.1](https://arxiv.org/html/2607.16117#S4.SS1): within each regime, bytes sit rightmost and patches leftmost\. The vertical spread shows the convergence behavior described in Section[4\.6](https://arxiv.org/html/2607.16117#S4.SS6): token models converge fastest for form preservation, byte models converge fastest for cross\-lingual retrieval, and all three encodings converge at similar speed for topic classification\. Under the zero\-padding accounting, the points shift left by the per\-input factors above while the epochs stay the same\.
Figure 10:Decomposition of the training cost into FLOPs per epoch and epochs to the selected checkpoint, by task, regime, and encoding, averaged over all bottleneck widths and five seeds\. Dashed diagonals mark equal total FLOPs\. Bytes are the most expensive encoding per epoch in every multilingual regime, but their position on the vertical axis depends on the task: they are slowest to converge for form preservation but fastest for cross\-lingual retrieval\.Similar Articles
Byte-level models
Discusses whether byte-level tokenizers outperform subword tokenizers for precise tasks like distinguishing similar names, counting characters, and case sensitivity, and asks for current recommendations.
Equity with Efficiency: An Empirical Study of Tokenizers for Multilingual Large Language Models
This paper systematically compares equitable tokenizers for multilingual LLMs across 11 Southeast Asian languages, finding that Parity-aware BPE achieves the best efficiency-equity trade-off and that cross-lingual fairness and tokenization efficiency are not fundamentally at odds.
Compute Optimal Tokenization (2 minute read)
This paper systematically derives compression-aware neural scaling laws by training nearly 1,300 models, demonstrating that the widely used heuristic of 20 tokens per parameter is an artifact of specific tokenizers. The authors propose a tokenizer-agnostic scaling law based on bytes, offering a new framework for compute-efficient training across diverse languages and modalities.
The Tokenizer Tax Across 25 European Languages: Domain Invariance, Cross-Lingual Few-Shot Effects, and the Ukrainian Penalty
This paper measures tokenizer fertility across 25 European languages on parallel text, revealing a 2.5x spread from English to Greek/Maltese, with Ukrainian paying a 15-18% penalty. It demonstrates domain invariance of fertility rankings, analyzes subword fragmentation, and evaluates cross-lingual few-shot effects.
The African Language Tax: Quantifying the Cost, Latency, and Context Penalty of Tokenizing African Languages in Frontier LLMs
This paper systematically quantifies the tokenization penalty for 20 African languages across 11 frontier and open tokenizers, finding up to 8.9× inference cost and latency multipliers and as little as 11% effective context window compared to English, highlighting a structural digital divide encoded in subword vocabularies.