Compute Optimal Tokenization (2 minute read)
Summary
This paper systematically derives compression-aware neural scaling laws by training nearly 1,300 models, demonstrating that the widely used heuristic of 20 tokens per parameter is an artifact of specific tokenizers. The authors propose a tokenizer-agnostic scaling law based on bytes, offering a new framework for compute-efficient training across diverse languages and modalities.
View Cached Full Text
Cached at: 05/14/26, 12:10 AM
Similar Articles
Finding Optimal Tokenizers
This blog post presents an algorithm using integer linear programming to compute optimal tokenizers for language models, drawing parallels to solving the Traveling Salesman Problem. It notes that while the result is theoretically interesting, practical tokenizers are already near-optimal and the method may not generalize well.
Lifecycle-Optimal Tokenization: Vocabulary Size as a Deployment-Regime-Dependent Infrastructure Parameter
This paper argues that the optimal tokenizer vocabulary size is not fixed but depends on the deployment regime, such as batch size and inference volume. Through roofline analysis and experiments on A10G and A100 GPUs, it shows the lifecycle-optimal vocabulary can shift by up to 16x between on-device and datacenter serving, with minimal quality impact.
A Pilot Study of Autocompleting Tokenizers
This paper proposes a compression scheme for byte-level tokenization using an autocomplete model to remove predictable bytes from input sequences, reducing sequence length while maintaining machine translation performance across diverse languages.
Beyond Selection: Token Parameterization for Extreme Visual Token Compression
The paper introduces Braco, a lightweight token parameterization method for visual token compression in vision-language models, achieving significant efficiency gains and accuracy under extreme compression budgets.
What Tokens are Learned when Tokenization is Optimized Jointly with Language Modeling?
This paper investigates the tokens learned when tokenization is optimized jointly with language modeling, comparing tokenizer-free methods across multiple languages and finding that they produce distinct, efficient vocabularies for NLP.