Tag
The paper proposes a densing law for user representation learning that quantifies the relationship between data scale and tokenization capacity, and introduces an adaptive tokenization method ALGN to improve efficiency in billion-scale scenarios.
A tweet recommends a Stanford class that provides a clear and complete explanation of how AI models like ChatGPT and Claude are built, covering tokenization, Transformer architecture, and training processes.
A new open-source project called FreeToken has been released, featuring a research paper and GitHub repository, with tests showing efficient token processing speeds of around 100 tokens per second on consumer hardware.
CZ advocates for tokenization on all blockchains to attract FDI. BNB Chain reports a 370% increase in RWA holders in 30 days.
SuTRA is a morphology-aware tokenization algorithm that preserves akshara indivisibility for Indic languages, reducing morphological shattering and achieving improvements in machine translation metrics over standard BPE methods.
This paper investigates the tokens learned when tokenization is optimized jointly with language modeling, comparing tokenizer-free methods across multiple languages and finding that they produce distinct, efficient vocabularies for NLP.
This paper challenges the standard assumptions about the linguistic units of the Voynich manuscript, showing through quantitative analysis that glyphs, tokens, and spaces do not correspond to letters, words, and word spaces as commonly believed.
This paper proposes a compression scheme for byte-level tokenization using an autocomplete model to remove predictable bytes from input sequences, reducing sequence length while maintaining machine translation performance across diverse languages.
ReconSpan introduces an adaptive latent tokenization method that divides text into reconstructible chunks using a backward decoder, enabling variable-length latent tokens and post-training control of granularity.
This paper argues that the optimal tokenizer vocabulary size is not fixed but depends on the deployment regime, such as batch size and inference volume. Through roofline analysis and experiments on A10G and A100 GPUs, it shows the lifecycle-optimal vocabulary can shift by up to 16x between on-device and datacenter serving, with minimal quality impact.
A user shares an observation that Qwen and Gemma tokenize code very differently, with Qwen using far fewer tokens for the same HTML/JS input, which may explain differences in coding and language performance. They also note a potential retraining project by LiquidAI using a more efficient tokenizer.
This arXiv paper evaluates federated training of tokenized generative event models (GEMs) on ICU EHR data from three health systems, showing that federated learning preserves most centralized performance and improves cross-site transportability compared to conventional supervised models.
This paper proposes Pruned BPE, a post-training method that prunes low-exposure tokens from a BPE vocabulary and reallocates slots to better-exposed candidates, reducing encoded length without increasing model-visible vocabulary size. Experiments on English and Chinese corpora show approximately 0.27–0.36% encoded length reduction over standard BPE.
TokTier is a stateful tokenization service that cuts tokenization overhead in agentic LLM serving by reusing stable token boundaries and running exact CPU/GPU tokenization, reducing time-to-first-token by 16–34% under vLLM.
This paper proposes GenCDSR, a generative framework for cross-domain sequential recommendation with hybrid tokenization and serial-parallel decoding, achieving improved accuracy and significantly reduced inference latency compared to state-of-the-art baselines.
This paper investigates whether language models remain robust to alternative (non-canonical) tokenizations across 27 languages, finding that invariance observed in English does not generalize and that languages with higher token fragmentation show greater sensitivity. The authors demonstrate that LoRA fine-tuning with multi-tokenization data can mitigate this sensitivity.
This paper introduces JOLT, an integer programming approach to optimize subword tokenization for greedy left-to-right longest-match decoding (WordPiece). JOLT achieves near-optimal compression, closing most of the gap between BPE and the theoretical lower bound, reducing token count by up to 0.78% over BPE.
This paper introduces BHARATI, a set of SentencePiece BPE tokenizers trained on a balanced multilingual corpus for classical Indian languages. Subword fertility analysis shows significant improvements in tokenization efficiency, reducing sequence length by up to 90% compared to baseline tokenizers, thereby enhancing effective context length for downstream language models.
Brazilian farmers tokenized 10 dairy cows on the B3 stock exchange, raising nearly $20,000 in credit backed by livestock, using AI-powered tracking collars to prevent fraud and enable movable collateral.
GigaToken is an ultra-fast tokenizer library that claims up to 1000x speedup over HuggingFace tokenizers, supporting most common LLM tokenizers and providing drop-in compatibility.