The Functionalizer: Lossless Functional Decomposition for Subword Tokenization
Summary
The Functionalizer is a lossless pre-tokenizer framework that factors orthographic and structural variations into opcodes and operands, reducing vocabulary requirements by up to 16% and improving code tokenization while maintaining text coherence.
View Cached Full Text
Cached at: 09/16/26, 08:34 AM
# The Functionalizer: Lossless Functional Decomposition for Subword Tokenization
Source: [https://arxiv.org/html/2609.15991](https://arxiv.org/html/2609.15991)
Connor Makowski Center for Transportation & Logistics Massachusetts Institute of Technology Cambridge, MA, USA conmak@mit\.eduWillem Guter Center for Transportation & Logistics Massachusetts Institute of Technology Cambridge, MA, USA wjguter@mit\.edu
###### Abstract
Standard subword tokenizers either treat every orthographic variation of a word \(such ashello,Hello,HELLO, andHéllo\) as unrelated vocabulary entries, which fragments the embedding space, or discard this variation through lossy normalization\. We present theFunctionalizer, a lossless pre\-tokenizer framework that factors orthographic and structural variations into a compositionalopcode/operandprefix stream before tokenization: a canonical base token \(operand\) prefixed by parametric transformation operators \(opcodes\) encoded in the Unicode Private Use Area\. We introduce operators covering casing \(CAPITALIZE\), diacritics \(13 dedicated opcodes\), and character repetition \(REPEAT,MULTIREPEAT\), which are fully reversible\. Across six natural language and code corpora, the Functionalizer enables complete corpus coverage with significantly smaller vocabularies under unconstrained conditions, reducing actual vocabulary slot requirements by up to16%\. When looking at sequence lengths, we observe a sharp domain\-dependent tradeoff: it compresses indentation\-heavy code sequences but inflates natural\-language prose sequences\. Preliminary downstream evaluations on 25M parameter GPT\-2 scale models show that at this scale, the Functionalizer drastically improves code syntax validity and improves code character perplexity while maintaining similar text coherence on prose\. These findings demonstrate that functional decomposition can be an effective mechanism for vocabulary\-efficient, structurally aware language modeling, and motivate further validation at production scale\.
## 1Introduction
Subword tokenizers face a dilemma\. Treatinghello,Hello, andHélloas independent tokens causesvocabulary expansion: redundant surface forms consume embedding slots, and a gradient update toHellonever benefitshello\. The alternative, aggressive lowercasing and accent\-stripping, islossy\. It compresses the vocabulary but permanently discards information the downstream model can never recover\.
The Functionalizer takes a third path:lossless functional decomposition\. Rather than memorizing surface forms or destroying them, it factors variation out into reusable*operators*applied to a single*canonical base*\. The design is borrowed directly from Instruction Set Architecture: a CPU does not implement a distinct instruction for every constant \(ADD\_1,ADD\_2, …\); it separates the operation \(opcode\) from its data \(operand\), as inADD A, \#1\. The Functionalizer applies the same factoring; the base token is the operand, the transformation prefix is the opcode\.
##### Contributions\.
1. 1\.A unified opcode/operand framework that handles casing, diacritics, and character repetition under one compositional, parametric, lossless scheme\. Prior work addresses these orthographic variations in isolation, either with lossy normalization or with a single operator that does not necessarily reduce tokenization overhead\.
2. 2\.A concrete Private Use Area \(PUA\) encoding that is usable in most tokenizers \(such as Hugging Face’s BPE\) and fully reversible\.
3. 3\.Empirical validation across six natural language and code corpora, demonstrating that when unconstrained by target vocabulary limits, the Functionalizer reduces the total vocabulary slots required to fully cover a corpus by up to16%through the collapse of formatting variations\. Regarding sequences, we show a clear domain split: it significantly compresses indentation\-heavy code sequences \(yielding inference throughput speedups of up to22\.8%\) while introducing a measured sequence length cost \(6–9%\) on natural\-language prose\.
4. 4\.Downstream language modeling training and evaluations at the∼\\sim25M parameter scale showing that the Functionalizer significantly improves next\-token code predictability \(a12\.6%relative improvement in character perplexity for python\) and drastically improves basic syntax validity \(achieving up to9\.20%syntax success compared to0\.00%–2\.20%for standard baselines\) while matching coherence on natural language prose\.
## 2Related Work
##### Subword tokenization\.
BPE\(Sennrich et al\.,[2016](https://arxiv.org/html/2609.15991#bib.bib1)\), WordPiece\(Schuster and Nakajima,[2012](https://arxiv.org/html/2609.15991#bib.bib3)\), and Unigram\(Kudo,[2018](https://arxiv.org/html/2609.15991#bib.bib2)\)build vocabularies bottom\-up from frequent merges; byte\-level BPE\(Radford et al\.,[2019](https://arxiv.org/html/2609.15991#bib.bib4)\)avoids out\-of\-vocabulary failures by operating over 256 byte values\. Existing pre\-tokenizers that address orthographic variation are either lossy \(lowercasing, accent\-stripping\) or limited to a single operator\. For example,Bayram et al\. \([2025](https://arxiv.org/html/2609.15991#bib.bib8)\)introduce a Turkish tokenizer that employs a single<uppercase\>token to fold casing into a shared root embedding\. The Functionalizer generalizes these approaches into a fully parametric, multi\-operator instruction set covering position, diacritic kind, and repeat count among other potential future extensions\.
##### Morphology\-aware tokenization\.
Morfessor\(Creutz and Lagus,[2002](https://arxiv.org/html/2609.15991#bib.bib5),[2007](https://arxiv.org/html/2609.15991#bib.bib6)\)performs unsupervised morpheme segmentation; MorphBPE\(Asgari et al\.,[2025](https://arxiv.org/html/2609.15991#bib.bib7)\)constrains BPE merges to morpheme boundaries\. The Functionalizer is complementary: where morphology\-aware methods target linguistic structure, the Functionalizer targets orthographic surface variation, and the two could be composed\.
##### Tokenization\-free models\.
ByT5\(Xue et al\.,[2022](https://arxiv.org/html/2609.15991#bib.bib9)\), MrT5\(Kallini et al\.,[2024](https://arxiv.org/html/2609.15991#bib.bib10)\), and MegaByte\(Yu et al\.,[2023](https://arxiv.org/html/2609.15991#bib.bib11)\)pursue orthographic robustness by operating on raw bytes or characters, at the cost of much longer sequences\. The Functionalizer moves towards similar invariance while retaining subword granularity and avoiding the sequence\-length penalty on code\.
##### Structured Unicode encoding\.
SCRIPT\-BPE\(Land and Arnett,[2025](https://arxiv.org/html/2609.15991#bib.bib12)\)re\-encodes characters by Unicode script and category to remove cross\-lingual bias\. The normalization literature \(e\.g\.,Gorman and Pinter,[2024](https://arxiv.org/html/2609.15991#bib.bib13)\) documents the downstream cost of inconsistent Unicode handling\. The Functionalizer is complementary to both: a structured prefix scheme that can be layered on top of any existing pipeline\.
## 3The Functionalizer Framework
The Functionalizer establishes a parametric, lossless, prefix\-based pre\-tokenization framework\. Instead of tokenizing raw surface forms directly, the framework decomposes orthographic and structural variations into a compositional sequence of non\-destructive operators \(opcodes\) prepended to a canonical base token \(operand\)\.
By design, any transformation that is fully reversible \(bijective\) and can be mapped to character indices can be integrated as an operator within this framework\. This includes not only capitalization and combining diacritics, but also structural transformations \(e\.g\., character repetition\)\. Potential future extensions include reversible morphological lemma folding, number/date normalization, prepending of whitespace, and machine learning preprocessing \(e\.g\., correcting likely spelling errors\)\. By isolating the parametric transformation from the semantic root, the framework enables downstream language models to process clean, shared canonical bases while preserving all orthographic details for lossless reconstruction\.
### 3\.1PUA Instruction Layout
Instructions are prepended to the base token as a sequence of PUA codepoints, composed of an operator followed by numeric parameters:
```
[operator] [param_1] [param_2] ... -> [base_token]
```
- •Numeric parameters\(U\+E000–U\+E0FF\): encode integer values 0–255 \(value = codepoint \- 0xE000\)\.
- •Operators\(U\+E100–U\+EFFF\): opcodes consuming a fixed number of parameters\. Decoupling opcode from arguments leaves the full plane available for future operators\.
### 3\.2Encoding and Reversibility
The transformation pipeline is designed to be fully bijective, ensuring lossless recovery of the original string\. For example:
##### Encode:
1. 1\.Extract diacritic marks \(generating serialization operators\) and uppercase positions \(generatingCAPITALIZE\)\.
2. 2\.Strip combining marks and lowercase the remaining characters\.
3. 3\.Prepend the operator prefix\.
##### Decode:
1. 1\.Apply operators in reverse order to the base token, restoring diacritics and capitalization\.
2. 2\.Remove the operator prefix, yielding the original string\.
Repeating text can be detected and encoded in the same way\. In the decode order, repetition operators execute last: per\-piece transforms \(CAPITALIZE, diacritics\) are applied first\. All position parameters \(includingposforCAPITALIZE/REPEAT, andstart/endforMULTIREPEAT\) refer to character positionswithin their individual piece\(after applying any prior operations, but before expansion\)\. The transform is bijective such that the original text is recovered exactly \(see Table[6](https://arxiv.org/html/2609.15991#A1.T6)in the Appendix for worked encoding examples\)\. This fixed decoding sequence introduces a representation constraint: since per\-piece transforms are applied first and then repeated, any capitalization or formatting is duplicated across all expanded units \(for example, applyingCAPITALIZE\(0\)to baseabcand repeating it 3 times yieldsAbcAbcAbc, rather than heterogeneously cased units likeAbcabcabc\)\. In such heterogeneous cases, the encoder must fall back to an uncollapsed representation\. While this limitation is minor in standard corpora, the framework could in theory be extended to support arbitrary decoding orders \(e\.g\., left\-to\-right or right\-to\-left evaluation\) to handle complex formatting combinations\.
### 3\.3Pipeline Integration
We recommend that the Functionalizer should run on pre\-split pieces, not raw text\. With the current design, this is because operators address only the first 256 characters \(pos≤255\\text\{pos\}\\leq 255\)\. This can be extended to cover a larger index space in future iterations\. From a more strategic level, this recommendation holds as it can allow a prepended operator/opcode sequence like “capitalize the next character” to be learned as one token during tokenization\.
In this implementation’s tests, the custom Split regex isolates each individual space character as a standalone space piece\. The Functionalizer’sREPEATthen aggregates consecutive identical space pieces back into a single opcode and one base space piece \(e\.g\., 12 isolated space pieces→\\rightarrow\[REPEAT\(0,12\)\]\+ space\)\. Because each space is its own piece, the repetition collapse does not affect character offsets within adjacent word pieces\.
By default \(split\_operators = true\), operators and parameters are emitted as separate tokens from the canonical base such that each operator and its parameters are tokenized independently\. This allows the model to learn operator semantics and parameter distributions separately from the base token embeddings\. Other options can be used to emit the entire operator sequence as a single token, which may be beneficial for certain downstream tasks\.
During vocabulary training, the BPE merge process is allowed to naturally merge these separate operator and parameter pieces back into combined tokens \(such as merging\[CAPITALIZE\]and\[0\]into a single\[CAPITALIZE\(0\)\]token\)\. In our experiments, BPE training naturally merges high\-frequency operator\-parameter pairs, which consumes extra vocabulary slots\.
## 4Current Operators and Expected Performance Profiles
### 4\.1Current Operators Specification
The current implementation of the Functionalizer includes operators for capitalization, combining diacritics \(detailed in Table[1](https://arxiv.org/html/2609.15991#S4.T1)\), and character repetition \(detailed in Table[2](https://arxiv.org/html/2609.15991#S4.T2)\)\.
Each capitalization and diacritic variant is mapped to a dedicated 1\-parameter operator in theU\+E100–U\+E10Fspace \(last two reserved\), consuming exactly one parameter \(pos\):
Table 1:Capitalization and Diacritic Operators SpecificationRepetition operators start atU\+E200:
Table 2:Repetition Operators Specification
### 4\.2Expected Performance Profiles
We expect performance improvements and training dynamics to manifest in different ways depending on the function of the operator applied\. While these effects can be greatly different depending on the operator, we present here two examples we expect to be representative of the current operator categories:
1. 1\.Semantic Invariance and Case Sharing \(CAPITALIZE\):Casing variations typically do not alter the core semantic identity of a word \(e\.g\.,hivs\.Hiorhellovs\.Hello\)\. By extracting the capitalization operation to a parametric prefix, casing variants are collapsed into the same base token embedding\. The downstream model shares parameter updates across all occurrences, boosting representation sharing and accelerating gradient updates\.
2. 2\.Structural Predictability and Compression \(REPEAT\):In contrast, repetition operators target structural and syntactic formatting, most notably in programming languages\. Repeated spaces or characters \(e\.g\., indentation levels\) are represented compactly by a single space character and a parameter value\. This avoids fragmenting the sequence or vocabulary with many variants of whitespace blocks \(e\.g\.,,,\)\. Consequently, repeated whitespace is rendered much more understandable and predictable for models, reducing context sequence lengths and enabling cleaner spatial representation of code block structures\.
## 5Experimental Setup
### 5\.1Pipelines and Configurations
To clarify the mechanics of the test configurations, we define the primary components and flags used across our experimental pipelines:
- •Unicode Normalization \(NFC\): All tokenization configurations apply Unicode NFC \(Normalization Form Canonical Composition\) normalization to the raw input text prior to pre\-tokenization\. NFC standardizes all visual variations of pre\-composed characters and decomposed character\-combining mark sequences \(such as different forms of accented characters\) into a single, unified canonical composed character representation \(e\.g\., composing decomposede \+ Combining Acute Accentinto the single pre\-composedé\)\. While not necessary, this can help reduce text variation and improve the efficiency of the Functionalizer’s decomposition\.
- •Split Operators \(split\_operators\): A configuration flag that determines whether prepended PUA operators \(opcodes\) and their numeric arguments \(parameters\) are split into independent string pieces before BPE subword training\. For example, when enabled \(split\_operators = true\), BPE receives\[CAPITALIZE\(0\)\], andhelloas separate tokens rather than the fused sequence\[CAPITALIZE\(0\)\]hello\. This prevents the vocabulary from memorizing fused operator\-operand structures and allows the model to learn reusable, generalizable embeddings for the operators and arguments\. When disabled, the prefix remains attached to the base word\.
- •Regex Splitting \(Split\): Segments raw text into localized character runs \(words, numbers, symbols, spaces, or newlines\) prior to subword training\. This pre\-segmentation isolates indentation sequences and keeps character offset counts compact\. We compare a custom regex splitter \(`\\n\+\|\\t\+\|\[ \]\| \|\[^\\p\{L\}\\p\{N\}\\s \_\]\+`\) and the standard Llama 3 pattern \(LLAMA\_SPLIT\_PAT\)\.
- •Functionalizer Decompositions: Governed by three key parameters in the pre\-tokenizer: - –capitalize: Identifies uppercase characters, generates aCAPITALIZE\(pos\)opcode, and lowercases the base character\. - –serialize: Detects combining diacritical marks, generates diacritic\-specific opcodes \(e\.g\.,ACUTE\(pos\)\), and strips the diacritics from the base character\. - –repeat: Identifies runs of 3\+ consecutive identical pre\-tokenized pieces \(or 3\+ identical characters within a single piece\), collapses them to a single instance, and generatesREPEATorMULTIREPEATopcodes\. When upstream pre\-tokenizers \(e\.g\., the Split regex\) isolate each space as its own piece,repeataggregates those split pieces back into one opcode plus one base piece\.
For all Functionalizer configurations, the corresponding decoder is configured to mirror the pre\-tokenizer flags to ensure losslessness\. For downstream language model training \(Section[6\.2](https://arxiv.org/html/2609.15991#S6.SS2)\) and downstream inference task evaluation \(Section[6\.3](https://arxiv.org/html/2609.15991#S6.SS3)\), we evaluateSplit Only,Split \+ Functionalizer, andSplit \+ Functionalizer \(Repeat\)configurations\.
### 5\.2Datasets, Vocabulary, and Tokenizer Metrics
We conduct experiments across six distinct natural language and source code corpora:
- •Prose: Wikitext \(Salesforce/wikitext\-2\-raw\-v1\), TinyStories \(roneneldan/TinyStories\)\.
- •Source Code: Python\-Codes \(flytech/python\-codes\-25k\) and CodeSearchNet \(CSN\-Python,CSN\-Java,CSN\-Go\)\.
We train tokenizers with target vocabulary sizes of128k\. All of our chosen corpora are exhausted before reaching the 128k target\. This means that training continues until no further BPE merges can be performed, representing corpus exhaustion\.
To evaluate vocabulary efficiency and compression directly at the tokenizer level, we track two metrics:
1. 1\.Token Compression \(Chars/Token\): The average number of visual characters represented per token\. A higher ratio indicates stronger text compression and shorter sequence lengths\.
2. 2\.Vocab Diff \(%\): The percentage reduction in actual vocabulary size achieved by the Functionalizer under unconstrained/exhaustion conditions \(actual\_vocab\_size < target\_vocab\_size\), which measures the framework’s ability to cover the same corpus with a smaller vocabulary\.
### 5\.3Downstream Training and Evaluation
To assess downstream performance, we train a GPT\-2 model \(6 layers, 512 embedding dim, 8 attention heads, 16k target vocabulary,∼\\sim25M parameters\) on each configuration\.
- •Training Details: Trained with a batch size of 16, a learning rate of 5e\-4, an AdamW optimizer \(weight decay 0\.01\), a linear learning rate schedule with 200 warmup steps decaying to zero, gradient clipping at 1\.0, and a context length of 256 tokens\. Training runs for 5,000 steps\. Training is repeated across five distinct seeds:10,42,100,2026, and7\.
- •Inference Downstream Prompts: Evaluated on 100 validation prompts per dataset\. Generation uses greedy decoding up to 50 new tokens \(for natural language\) or 64 new tokens \(for source code\)\.
- •Downstream Evaluation Metrics: - –Per\-Character Perplexity \(Char PPL\): A normalized measure of language modeling loss to enable direct comparison across tokenizers with differing token lengths\. Because tokenizers partition the same text into different numbers of tokens, comparing token\-level perplexity is unfair \(tokenizers with shorter tokens will show artificially lower perplexity\)\. To normalize, we convert the average token\-level cross\-entropy loss \(ℒtoken\\mathcal\{L\}\_\{\\text\{token\}\}\) to character\-level cross\-entropy loss \(ℒchar\\mathcal\{L\}\_\{\\text\{char\}\}\) using the corpus\-level average character\-to\-token ratio \(Rchar/token=NcharsNtokensR\_\{\\text\{char/token\}\}=\\frac\{N\_\{\\text\{chars\}\}\}\{N\_\{\\text\{tokens\}\}\}\): ℒchar=ℒtokenRchar/token\\mathcal\{L\}\_\{\\text\{char\}\}=\\frac\{\\mathcal\{L\}\_\{\\text\{token\}\}\}\{R\_\{\\text\{char/token\}\}\}\(1\)The per\-character perplexity is then computed as: PPLchar=exp\(ℒchar\)=exp\(ℒtokenRchar/token\)\\text\{PPL\}\_\{\\text\{char\}\}=\\exp\\left\(\\mathcal\{L\}\_\{\\text\{char\}\}\\right\)=\\exp\\left\(\\frac\{\\mathcal\{L\}\_\{\\text\{token\}\}\}\{R\_\{\\text\{char/token\}\}\}\\right\)\(2\)This ensures that character perplexity represents a fair, tokenizer\-independent metric of information density\. - –Throughput \(Chars/Sec\): During training, this is computed as the total tokens processed multiplied by the dataset’s average character\-to\-token ratio \(based on raw text length including all spaces and formatting\) divided by total training time\. During inference, this text speed is calculated on the reconstructed decoded string length \(including spaces and newlines, with PUA opcodes decoded back to standard characters\) divided by the coherent generation time\. To prevent collapsed models from scoring artificially high throughput, this speed is computed exclusively over the sequence prefix generated prior to any repetition collapse \(i\.e\., before any single token repeats\>3\>3times consecutively\)\. - –Average Tokens Pre\-Collapse: The average number of tokens generated before a single token repeats consecutively more than three times \(the threshold for pathological loop collapse\)\. This measures the length of coherent sequence generation before degeneration\. - –Coherence \(1 \- Repetition\): For prose, the percentage of generated sequences free of pathological n\-gram loops\. If the model produces an empty or whitespace\-only sequence, it is tracked as 0% coherent\. - –Syntax Success Rate: For source code, the percentage of generations that parse cleanly under Python’sast\.parseor Go’sgofmt\. Empty generations \(or whitespace only\) are counted as failures\. - –% Empty: The percentage of generated sequences that are empty or contain only whitespace, metaspace, or raw PUA control characters\. While empty sequences are classified as failures for quality metrics \(coherence and syntax success\), their tokens and generation time are still fully accounted for in the throughput and speed calculations\.
## 6Results
### 6\.1Tokenizer Metrics
To analyze vocabulary usage, token compression, and unconstrained vocabulary requirements \(corpus exhaustion\), we evaluate configurations across all six datasets\. Table[3](https://arxiv.org/html/2609.15991#S6.T3)presents consolidated metrics at the128k target vocabulary size\(which represents the unconstrained/corpus exhaustion regime where the tokenizer’s bounds are determined by corpus entropy rather than a hard limit\)\. Each configuration consists of aBaseversion \(without decomposition\) and aFunctionalizerversion \(with decomposition\)\. The configurations labeled as\(Repeat\)\(e\.g\. Split Only \(Repeat\) or Split \+ Functionalizer \(Repeat\)\) represent baseline and Functionalizer configurations where only the character repetition operators \(REPEAT,MULTIREPEAT\) are active, while capitalization and diacritic serialization operators are disabled\.
Table 3:Tokenizer Metrics at 128k Target Vocabulary \(Corpus Exhaustion\)#### Key Insights and Analysis
Distinct vocabulary and corpus exhaustion\.At the 128k target scale, all of our datasets saturate before reaching the maximum vocabulary size, leading to corpus exhaustion\. Table[3](https://arxiv.org/html/2609.15991#S6.T3)demonstrates that when unconstrained by a target vocabulary limit, the collapse of formatting variants allows the Functionalizer to achieve complete corpus coverage with a significantly smaller total vocabulary size\. The necessary vocabulary size drops by up to16%on standard configurations:
- •Wikitext: Reduces by15\.60%\(43,700 vs\. 51,780\)\.
- •Python\-Codes: Reduces by16\.11%\(28,017 vs\. 33,399\)\.
- •TinyStories: Reduces by9\.99%\(12,853 vs\. 14,280\)\.
- •CSN\-Python: Reduces by12\.72%\(82,575 vs\. 94,605\)\.
- •CSN\-Java: Reduces by12\.92%\(54,371 vs\. 62,436\)\.
OnCSN\-Go, the unconstrained vocabulary reduction is minor \(2\.89%\), matching the minimal casing and repetition structure targeted by the operators\.
Token compression: the domain split\.The character\-to\-token ratio \(Chars/Token\) in Table[3](https://arxiv.org/html/2609.15991#S6.T3)reveals a sharp division between prose and code\.
- •Prose: On Wikitext and TinyStories, the Functionalizer reduces the Chars/Token ratio \(resulting in a\+6\.42% to \+16\.50%token inflation across configurations, or\+6\.42% to \+9\.02%on standard splits\)\. This sequence length cost occurs because natural language prose lacks the long repeated sequences thatREPEATcollapses, meaning the casing operators represent a net sequence length overhead\.
- •Source Code: On CSN\-Python and CSN\-Java, the Functionalizer substantially increases the Chars/Token ratio, compressing sequence lengths\. This compression saves up to26\.77%in tokens on CSN\-Python and8\.44%on CSN\-Java \(increasing up to30\.35%and19\.99%respectively when considering repetition collapse only\)\. This is driven by indentation collapse, where long space runs are collapsed into a single character plus aREPEAToperator\.
The Llama Split Code Inflation\.While custom Split configurations compress source code, theLlama Splitconfiguration exhibits the opposite behavior, showing sequence length inflation on both CSN\-Python \(\+17\.67%token inflation, with Chars/Token dropping from 4\.0416 to 3\.4345\) and CSN\-Java \(\+33\.22%token inflation, with Chars/Token dropping from 4\.2262 to 3\.1723\)\. This sequence inflation occurs because the Llama split regex does not isolate spaces from adjacent word pieces, which prevents the Functionalizer’sREPEAToperator from collapsing indentation blocks\. Consequently, the casing operators introduce a net token overhead without any offsetting spatial compression\. This sensitivity highlights that the Functionalizer’s compression effectiveness is tightly coupled to the split regex of the base tokenizer, which we flag as a key area for follow\-up research\.
The Go Indentation and CamelCase Exception\.CSN\-Goacts as the clearest exception to the source code compression rule\. As shown in Table[3](https://arxiv.org/html/2609.15991#S6.T3), the Functionalizer inflates Go sequence lengths by\+6\.97%\(reduced Chars/Token\)\. This occurs because:
1. 1\.Go’s standard formatter \(gofmt\) enforces tab\-based indentation instead of spaces\. Tabs are already single\-character tokens in BPE, leaving no long space runs forREPEATto compress\.
2. 2\.Go’s strict CamelCase conventions for exported functions and identifiers lead to a high frequency of uppercase transitions, which translates to a high frequency ofCAPITALIZEoperators without corresponding compression\.
This exception demonstrates that the Functionalizer’s sequence efficiency is highly predictable and depends entirely on the presence of the orthographic structures its operators are designed to target\.
### 6\.2Training Dynamics and Language Modeling Performance
We analyze the training behavior and next\-token prediction performance of the∼\\sim25M parameter GPT\-2 model under each tokenizer configuration\. Table[4](https://arxiv.org/html/2609.15991#S6.T4)summarizes the training metrics across TinyStories, CSN\-Python, and CSN\-Go, averaged over five random seeds\.
Table 4:Training Dynamics and Language Modeling Metrics \(5 Seeds\)\. Inflation metrics are calculated from unrounded raw token counts\.Tokenizer TypeVocab SizeChars/TokenInflationTokens/SecChars/SecFinal LossToken PPLChar PPLDataset: TinyStoriesSplit Only14,2802\.2950\.00%0\.00\\%55498\.7±661\.855498\.7\\pm 661\.8127354\.4±1518\.5127354\.4\\pm 1518\.51\.4980±0\.00481\.4980\\pm 0\.00484\.47±0\.024\.47\\pm 0\.021\.9209±0\.00401\.9209\\pm 0\.0040Split \+ Functionalizer12,8532\.156\+6\.45%\+6\.45\\%53675\.9±1568\.653675\.9\\pm 1568\.6115705\.6±3381\.3115705\.6\\pm 3381\.31\.4402±0\.00111\.4402\\pm 0\.00114\.22±0\.004\.22\\pm 0\.001\.9506±0\.00101\.9506\\pm 0\.0010Split \+ Functionalizer \(Repeat\)14,6652\.294\+0\.03%\+0\.03\\%54220\.7±2187\.054220\.7\\pm 2187\.0124387\.2±5017\.1124387\.2\\pm 5017\.11\.5125±0\.00231\.5125\\pm 0\.00234\.54±0\.014\.54\\pm 0\.011\.9335±0\.00201\.9335\\pm 0\.0020Dataset: CSN\-PythonSplit Only16,0001\.9010\.00%0\.00\\%52237\.9±17\.352237\.9\\pm 17\.399306\.9±33\.099306\.9\\pm 33\.02\.5551±0\.00412\.5551\\pm 0\.004112\.87±0\.0512\.87\\pm 0\.053\.8345±0\.00823\.8345\\pm 0\.0082Split \+ Functionalizer16,0002\.620−27\.44%\-27\.44\\%54584\.6±19\.954584\.6\\pm 19\.9143018\.9±52\.3143018\.9\\pm 52\.33\.1686±0\.00403\.1686\\pm 0\.004023\.77±0\.1023\.77\\pm 0\.103\.3512±0\.00523\.3512\\pm 0\.0052Split \+ Functionalizer \(Repeat\)16,0002\.771−31\.39%\-31\.39\\%55470\.3±25\.355470\.3\\pm 25\.3153699\.3±70\.0153699\.3\\pm 70\.03\.3496±0\.00493\.3496\\pm 0\.004928\.49±0\.1428\.49\\pm 0\.143\.3498±0\.00593\.3498\\pm 0\.0059Dataset: CSN\-GoSplit Only15,7352\.2260\.00%0\.00\\%60367\.7±23\.360367\.7\\pm 23\.3134352\.3±51\.8134352\.3\\pm 51\.82\.8526±0\.00932\.8526\\pm 0\.009317\.33±0\.1617\.33\\pm 0\.163\.6029±0\.01513\.6029\\pm 0\.0151Split \+ Functionalizer15,2811\.959\+13\.63%\+13\.63\\%60554\.2±36\.460554\.2\\pm 36\.4118598\.1±71\.3118598\.1\\pm 71\.32\.6738±0\.00662\.6738\\pm 0\.006614\.50±0\.1014\.50\\pm 0\.103\.9166±0\.01333\.9166\\pm 0\.0133Split \+ Functionalizer \(Repeat\)15,6722\.224\+0\.07%\+0\.07\\%60627\.7±30\.960627\.7\\pm 30\.9134833\.6±68\.8134833\.6\\pm 68\.82\.8677±0\.00852\.8677\\pm 0\.008517\.60±0\.1517\.60\\pm 0\.153\.6308±0\.01393\.6308\\pm 0\.0139
#### Key Insights and Analysis
Analyzing Per\-Character Perplexity \(Char PPL\)\.Because sequence lengths differ under different tokenizers, Token Perplexity \(Token PPL\) is not directly comparable\. We therefore compute the Per\-Character Perplexity \(Char PPL\) as a normalized measure of language modeling performance\.
- •CSN\-Python \(Source Code\): The Functionalizer configurations consistently and significantly outperform their baseline counterparts\. For instance, theSplit \+ Functionalizerachieves a Char PPL of3\.3512compared to theSplit Onlybaseline at3\.8345\(a relative improvement of12\.6%\)\. This indicates that collapsing space runs into simple structural parameters makes code indentation structures highly predictable, boosting the model’s overall representational capability\.
- •CSN\-Go \(Source Code\): Go represents the counter\-case\. Due to tab indentation \(which already behaves as clean single\-character tokens in BPE\) and capital\-dense CamelCase exports, the Functionalizer inflates Go sequence lengths without substantial offsetting compression \(causing a \+13\.63% sequence inflation\)\. Consequently, Char PPL degrades under the Go configuration \(e\.g\.,3\.9166for Split \+ Functionalizer vs\.3\.6029for Split Only\)\.
- •TinyStories \(Prose\): Consistent with the sequence inflation on prose \(\+6\.45%\), the Split \+ Functionalizer experiences a minor Char PPL degradation \(1\.9506 vs\. 1\.9209\)\.
Computational Throughput\.The Functionalizer can improve character level throughput for certain corpora\.
- •Throughput \(Chars/Sec\): The Functionalizer significantly increases content throughput \(Chars/Sec\) on corpora where it reduces sequence length\. For example, on CSN\-Python, the Split \+ Functionalizer generates143,018\.9 chars/seccompared to99,306\.9 chars/secfor Split Only\. This represents a44\.0%increase in content throughput\. This gain is driven primarily by the higher chars/token ratio achieved through sequence compression: raw token generation rates \(Tokens/Sec\) remain comparable across all configurations \(∼\\sim52,000–62,000 tokens/sec\), so more characters are delivered per generated token rather than more tokens per second\. On prose \(TinyStories\), character throughput slightly decreases \(115,705\.6vs\.127,354\.4 chars/sec\), in line with the sequence length cost\.
### 6\.3Downstream Inference and Downstream Tasks
We perform greedy decoding evaluations on validation prompts to assess inference latency and text generation success\. Table[5](https://arxiv.org/html/2609.15991#S6.T5)lists the results \(Part a for prose, Part b for code\)\.
Table 5:Downstream Inference Metrics \(100 Prompts, 5 Seeds\)\(a\) Prose Generation Metrics
\(b\) Source Code Generation Metrics
Tokenizer TypeAvg TokensAvg Tokens Pre\-CollapseAvg CharsTokens / SecChars / Sec \(Text Speed\)% EmptySyntax Success Rate \(%\)Dataset: CSN\-Python \(Code Generation\)Split Only6\.0±1\.36\.0\\pm 1\.32\.0±1\.32\.0\\pm 1\.37\.9±2\.47\.9\\pm 2\.4313\.3±6\.0313\.3\\pm 6\.0626\.7±79\.1626\.7\\pm 79\.152\.0±25\.7%52\.0\\pm 25\.7\\%0\.00±0\.00%0\.00\\pm 0\.00\\%Split \+ Functionalizer9\.1±1\.19\.1\\pm 1\.15\.1±1\.15\.1\\pm 1\.115\.6±3\.115\.6\\pm 3\.1316\.8±2\.9316\.8\\pm 2\.9507\.6±200\.1507\.6\\pm 200\.177\.6±11\.6%77\.6\\pm 11\.6\\%9\.20±\\pm9\.00%Split \+ Functionalizer \(Repeat\)7\.3±0\.27\.3\\pm 0\.23\.3±0\.23\.3\\pm 0\.212\.1±0\.612\.1\\pm 0\.6317\.5±4\.1317\.5\\pm 4\.1769\.8±25\.2769\.8\\pm 25\.292\.6±2\.9%92\.6\\pm 2\.9\\%0\.00±0\.00%0\.00\\pm 0\.00\\%Dataset: CSN\-Go \(Code Generation\)Split Only8\.5±3\.78\.5\\pm 3\.74\.6±3\.84\.6\\pm 3\.810\.6±6\.110\.6\\pm 6\.1314\.6±3\.3314\.6\\pm 3\.3401\.1±74\.1401\.1\\pm 74\.173\.2±23\.7%73\.2\\pm 23\.7\\%2\.20±3\.12%2\.20\\pm 3\.12\\%Split \+ Functionalizer10\.7±7\.610\.7\\pm 7\.67\.0±8\.27\.0\\pm 8\.213\.6±11\.313\.6\\pm 11\.3314\.6±1\.9314\.6\\pm 1\.9431\.5±28\.4431\.5\\pm 28\.471\.8±26\.3%71\.8\\pm 26\.3\\%4\.00±\\pm4\.65%Split \+ Functionalizer \(Repeat\)9\.9±3\.99\.9\\pm 3\.96\.1±4\.06\.1\\pm 4\.013\.4±7\.413\.4\\pm 7\.4312\.2±4\.7312\.2\\pm 4\.7436\.9±81\.8436\.9\\pm 81\.867\.6±21\.7%67\.6\\pm 21\.7\\%2\.80±4\.17%2\.80\\pm 4\.17\\%
#### Key Insights and Analysis
Downstream Code Generation Syntax Success Rate\.One of our most striking results is the impact of decomposition on source code syntax validity under highly limited model size\. At the 25M parameter GPT\-2 scale, models struggle to learn basic syntax and indentation, as evidenced by the zero success rates of standard tokenizers on CSN\-Python \(e\.g\.,0\.00%for Split Only\)\.
By contrast, integrating the Functionalizer pre\-tokenizer drastically improves the model’s capacity to output valid code syntax at this scale\. TheSplit \+ Functionalizerconfiguration reaches a syntax success rate of9\.20%\(compared to0\.00%for the Split Only baseline\)\. The customSplitregex preserves newlines \(\\n\) and tabs \(\\t\) losslessly, allowing the model to process multi\-line sequences\. In this setup, the standard Split pre\-tokenizer splits the space character from adjacent words, which allows theREPEAToperator to fully aggregate the run of consecutive spaces \(e\.g\., 4 spaces collapse to\\ue200\\ue000\\ue004and a space\)\. We note that these syntax success rates carry substantial variance across seeds \(e\.g\.,9\.20±\\pm9\.00%\), reflecting sensitivity to weight initialization at this model scale; the consistent direction of improvement of the Functionalizer configuration on CSN\-Python, however, supports the pattern as a real effect rather than noise\. This indicates that by reducing the cognitive load of structural whitespace into deterministic, parameter\-based instructions, small language models can generate syntactically sound structures that compile/parse\. This does not mean that the results are production\-ready, but it does suggest that the Functionalizer can help small models learn to represent and generate structured content more effectively\.
Coherent Generation Length \(Avg Tokens Pre\-Collapse\)\.A critical challenge in low\-parameter language modeling is the tendency to fall into early pathological looping or generation collapse \(e\.g\., repeating spaces, newlines, or specific phrases infinitely\)\. In Table[5](https://arxiv.org/html/2609.15991#S6.T5), we track the average number of tokens generated before a single token repeats consecutively more than three times \(Avg Tokens Pre\-Collapse\)\. We find that models trained with the Functionalizer generate substantially longer sequences before collapsing\. For example, on theTinyStoriesprose dataset, the Split \+ Functionalizer model generates an average of11\.7tokens pre\-collapse compared to only3\.6tokens for the Split Only baseline\. Similarly, onCSN\-Python, the Functionalizer increases pre\-collapse tokens from2\.0to5\.1, and onCSN\-Gofrom4\.6to7\.0\. When paired with similar coherence levels on prose \(e\.g\., 97\.80% vs\. 99\.00%\) and significantly higher syntax success rates on code, these metrics indicate that the Functionalizer allows the model to produce longer content without degenerating\. While we avoid making definitive causal claims, this behavior suggests that by offloading surface variations, the model can maintain a more stable context representation over longer generation sequences, which may correlate with better representation learning or training dynamics\.
Downstream Latency and Generation Speed\.During inference generation \(which is autoregressive and token\-by\-token\), we observe mixed results in generation speed in terms of actual content \(characters generated per second\)\. On CSN\-Go, the Split \+ Functionalizer configuration achieves a text speed of431\.5 chars/seccompared to401\.1 chars/secfor Split Only \(a7\.6%speedup\)\. For configurations targeting repetition only \(Split \+ Functionalizer \(Repeat\)\), we observe more consistent improvements: on CSN\-Python, it generates769\.8 chars/seccompared to626\.7 chars/secfor Split Only \(a22\.8%speedup\), and on TinyStories it generates1134\.5 chars/seccompared to1117\.9 chars/sec\(a1\.5%speedup\)\. However, the full Split \+ Functionalizer configuration is slower on TinyStories \(954\.5 chars/sec\) and CSN\-Python \(507\.6 chars/sec\) compared to the Split Only baseline, reflecting the sequence length overhead of casing operators when they are not offset by structural compression\. We note that the raw generation speeds during inference are orders of magnitude lower than the training throughput metrics reported in Table[4](https://arxiv.org/html/2609.15991#S6.T4)\. This discrepancy is expected and stems from the difference in computation modes\.
## 7Discussion
The tokenizer\-level results support a specific, bounded claim:the Functionalizer is a vocabulary\-efficiency and code\-compression tool, not a universal improvement\.It eliminates casing\-induced vocabulary fragmentation and shortens code sequences substantially \(token usage\), at the price of longer prose sequences\. For a code\-domain tokenizer, this tradeoff is clearly favorable; for prose\-dominant workloads, the value depends on downstream gains not yet demonstrated at production scale\.
The Serving Cost and Sequence Length Trade\-off\.A critical open question is the trade\-off between vocabulary size reduction and sequence length inflation\. Under modern LLM deployment paradigms, inference is frequently bottlenecked by memory bandwidth, where serving costs and latencies scale with context sequence length \(driven by KV cache size and attention computation overhead\)\. In this light, a6–9%sequence length inflation on prose represents a non\-trivial penalty\. Conversely, a9–16%\(excepting Go\) reduction in vocabulary size reduces the model’s embedding parameters, which can decrease peak memory footprint and training requirements\. Characterizing the exact boundary where vocabulary slot savings outweigh sequence length inflation, under varying hardware constraints \(compute\-bound training vs\. memory\-bandwidth\-bound inference\), remains an important open question for future research\.
The preliminary inference results in Section[6\.3](https://arxiv.org/html/2609.15991#S6.SS3)are consistent with the vocabulary compaction story\. The Functionalizer matches Split Only coherence on prose while achieving higher vocabulary efficiency\. With that said, a 25M parameter model at short generation lengths is insufficient to validate or invalidate any production scale hypothesis\. The open questions are squarely enumerated in Subsection[7\.1](https://arxiv.org/html/2609.15991#S7.SS1)\.
It is also worth noting that at this small model scale \(25M parameters\) and short context length \(maximum 50–64 generated tokens\), the computation time is not the primary bottleneck during inference\. Host\-device synchronization \(e\.g\., retrieving the token ID via blocking CPU\-GPU calls\) and CUDA kernel launch overhead dominate the latency\. Consequently, hardware\-level optimizations like Key\-Value \(KV\) caching yield negligible differences in generation throughput at this scale\. The reported downstream speeds reflect these systemic overhead constraints, and we expect the relative generation speedups of Functionalizer configurations to widen as model parameters and generation sequence lengths scale up\.
Additionally, we expect that as more operators are implemented within the Functionalizer framework, they could further improve performance, especially when training models on specific tasks\.
### 7\.1Limitations
1. 1\.Downstream evaluation is preliminary\.Inference results are at∼\\sim25M parameters \(Section[6\.3](https://arxiv.org/html/2609.15991#S6.SS3)\)\. The vocabulary compression→\\rightarrowperformance link is not validated at production scale, where we expect gains may shift to training and inference efficiency rather than test performance\.
2. 2\.Prose costs sequence length\.Up to \+9% more tokens on Wikitext at 128k\. For prose\-dominant workloads this is a net cost unless offset by downstream gains not yet shown\.
3. 3\.256\-character position limit\.Transforms beyond index 255 are silently dropped; runs of≥\\geq256 identical characters are not collapsed\. Rare in practice, but a hard boundary under adversarial or unusual inputs\. Bumping the starting codepoints of the operators to higher PUA sub\-blocks \(e\.g\., shifts starting atU\+EA00andU\+EB00\) would expand the numeric parameter range and mitigate this constraint, though we do not focus on this extension in the current work\.
4. 4\.Fixed diacritic coverage\.The serialization framework supports 13 combining marks\. Scripts and marks outside this set fall back to non\-functionalizer behavior\.
5. 5\.Pareto frontier uncharacterised\.We measure vocabulary size and token usage independently; we do not characterise the joint tradeoff or its interaction with model size\.
6. 6\.Incomplete operator ablations\.Beyond some tests with repetition only, we did not isolate the individual contribution of each operator \(e\.g\., casing vs\. repetition\) in our downstream evaluations\.
### 7\.2Future Work
While our preliminary evaluations establish the feasibility and efficiency of functional decomposition, several promising directions remain to expand the framework:
- •Production\-Scale Validation:The most critical next step is extending downstream training and inference validation to standard LLM scales \(1B\+ parameters\)\. This will evaluate whether the vocabulary compression gains translate directly to training throughput, data efficiency, and downstream task accuracy \(especially on casing\-sensitive tasks like Named Entity Recognition\)\.
- •Richer Operator Sets:Expanding and validating advanced bijective operators, including: - –*Morphological Lemma Folding:*Factoring morphological inflection markers \(e\.g\.,\[PLURAL\],\[PAST\_TENSE\]\) from canonical base lemmas to compress morphologically rich languages\. - –*Structured Pattern Folding:*Factoring structured data runs \(e\.g\., dates, IP addresses, hashes\) into parameterized instructions over base operand tokens\. - –*Lossless Spelling Normalization:*Factoring minor typographical errors into parameterized edit operators applied to shared base roots\.
## 8Conclusion
The Functionalizer factors orthographic and structural variation out of the subword tokenizer vocabulary into a compositional, parametric opcode/operand prefix stream encoded in the Unicode Private Use Area\. By isolating surface variations from semantic roots, the framework achieves complete corpus coverage with much smaller vocabularies under unconstrained conditions, reducing actual vocabulary slot requirements by up to16%\(averaging11\.7%fewer slots overall across split only vs split \+ functionalizer configurations and datasets\)\.
Downstream training and evaluation on 25M parameter language models reveal a domain\-dependent tradeoff: on indentation\-dense code, the Functionalizer delivers substantial sequence\-length compression, improves next\-token predictability, and drastically improves code generation syntax validity \(from near\-zero for standard baselines up to9\.20%\)\. On natural\-language prose, the casing operators introduce a minor sequence length cost \(\+6% to \+9%\) which degrades generation speed \(showing a 14\.6% slowdown\), though the repetition\-only configuration \(Split \+ Functionalizer \(Repeat\)\) maintains a minor 1\.5% speedup on TinyStories \(while matching baseline coherence\)\. On CSN\-Python code generation, the repetition\-only configuration yields a substantial speedup of 22\.8% \(at the cost of a higher empty generation rate\)\.
While production\-scale validation remains an open question, these preliminary results demonstrate that factoring orthographic complexity out of tokenizers is a viable path toward vocabulary\-efficient, structurally aware language modeling\.
## Appendix AWorked Encoding Examples
Table 6:Worked Encoding Examples
## Data and Code Availability
## References
- Sennrich et al\. \(2016\)Sennrich, R\., Haddow, B\., & Birch, A\. \(2016\)\. Neural Machine Translation of Rare Words with Subword Units\.*arXiv preprint arXiv:1508\.07909*\.
- Kudo \(2018\)Kudo, T\. \(2018\)\. Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates\.*arXiv preprint arXiv:1804\.10959*\.
- Schuster and Nakajima \(2012\)Schuster, M\., & Nakajima, K\. \(2012\)\. Japanese and Korean voice search\. In*2012 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\)*\(pp\. 5149–5152\)\. IEEE\.
- Radford et al\. \(2019\)Radford, A\., Wu, J\., Child, R\., Luan, D\., Amodei, D\., & Sutskever, I\. \(2019\)\. Language Models are Unsupervised Multitask Learners\.*OpenAI Blog*\.
- Creutz and Lagus \(2002\)Creutz, M\., & Lagus, K\. \(2002\)\. Unsupervised Discovery of Morphemes\. In*Proceedings of the ACL\-02 Workshop on Morphological and Phonological Learning*\.
- Creutz and Lagus \(2007\)Creutz, M\., & Lagus, K\. \(2007\)\. Unsupervised Models for Morpheme Segmentation and Morphology Learning\.*ACM Transactions on Speech and Language Processing \(TSLP\)*, 4\(1\), 3\-es\.
- Asgari et al\. \(2025\)Asgari, E\., El Kheir, M\., & Sadraei Javaheri, A\. \(2025\)\. MorphBPE: A Morpho\-Aware Tokenizer Bridging Linguistic Complexity for Efficient LLM Training Across Morphologies\.*arXiv preprint arXiv:2502\.00894*\.
- Bayram et al\. \(2025\)Bayram, M\. A\., Fincan, A\. A\., Gümüş, A\. S\., Karakaş, S\., Diri, B\., Yıldırım, S\., & Çelik, D\. \(2025\)\. Tokens with Meaning: A Hybrid Tokenization Approach for Turkish\.*arXiv preprint arXiv:2508\.14292*\.
- Xue et al\. \(2022\)Xue, L\., Barua, A\., Constant, N\., Al\-Rfou, R\., Narang, S\., Kale, M\., Roberts, A\., & Raffel, C\. \(2022\)\. ByT5: Towards a Token\-Free Future with Pre\-trained Byte\-to\-Byte Models\.*arXiv preprint arXiv:2105\.13626*\.
- Kallini et al\. \(2024\)Kallini, J\., Murty, S\., Manning, C\. D\., Potts, C\., & Csordás, R\. \(2024\)\. MrT5: Dynamic Token Merging for Efficient Byte\-level Language Models\.*arXiv preprint arXiv:2410\.20771*\.
- Yu et al\. \(2023\)Yu, L\., Simig, D\., Flaherty, C\., Aghajanyan, A\., Zettlemoyer, L\., & Lewis, M\. \(2023\)\. MEGABYTE: Predicting Million\-byte Sequences with Multiscale Transformers\.*arXiv preprint arXiv:2305\.07185*\.
- Land and Arnett \(2025\)Land, S\., & Arnett, C\. \(2025\)\. BPE Stays on SCRIPT: Structured Encoding for Robust Multilingual Pretokenization\.*arXiv preprint arXiv:2505\.24689*\.
- Gorman and Pinter \(2024\)Gorman, K\., & Pinter, Y\. \(2024\)\. Don’t Touch My Diacritics\.*arXiv preprint arXiv:2410\.24140*\.Similar Articles
Balancing Image Compression and Generation with Bootstrapped Tokenization
Introduces SelfBootTok, a self-bootstrapped tokenization method that separates global and local information, reducing generator computation by ~40% and achieving a new state-of-the-art gFID of 1.56 with only 64 tokens.
Joint Optimization for Greedy Longest-match Tokenization
This paper introduces JOLT, an integer programming approach to optimize subword tokenization for greedy left-to-right longest-match decoding (WordPiece). JOLT achieves near-optimal compression, closing most of the gap between BPE and the theoretical lower bound, reducing token count by up to 0.78% over BPE.
Beyond Atomic Tokens: Factorizing Syllables for Language Model Pretraining
This paper introduces a linguistically motivated phonemic tokenizer for Vietnamese and Chinese that factorizes syllables into onset, rime, and tone components, reducing vocabulary size and improving efficiency. The proposed PhonemicBERT models demonstrate competitive or superior performance in language understanding tasks compared to existing tokenizers and pretrained models.
Compute Optimal Tokenization (2 minute read)
This paper systematically derives compression-aware neural scaling laws by training nearly 1,300 models, demonstrating that the widely used heuristic of 20 tokens per parameter is an artifact of specific tokenizers. The authors propose a tokenizer-agnostic scaling law based on bytes, offering a new framework for compute-efficient training across diverse languages and modalities.
Removable and Irreducible: A Token-Cost Ledger for the Multilingual Tokenization Tax
This paper introduces a token-cost ledger to analyze the multilingual tokenization tax, decomposing it into removable and intrinsic components, showing that much excess token cost for non-English text is removable through improved tokenization codes.