tokenization

Tag

Cards List
#tokenization

Towards a Densing Law for User Representation Learning at Billion-Scale Capacity

Hugging Face Daily Papers ↗ · 2026-08-24 Cached

The paper proposes a densing law for user representation learning that quantifies the relationship between data scale and tokenization capacity, and introduces an adaptive tokenization method ALGN to improve efficiency in billion-scale scenarios.

0 favorites 0 likes
#tokenization

@dkare1009: Leave Netflix tonight. Watch this 2 h 34 min Stanford class. It's the clearest, most complete, and brutally honest expl…

X AI KOLs Timeline ↗ · 2026-08-23 Cached

A tweet recommends a Stanford class that provides a clear and complete explanation of how AI models like ChatGPT and Claude are built, covering tokenization, Transformer architecture, and training processes.

0 favorites 0 likes
#tokenization

Freetokens project is impressive

Reddit r/LocalLLaMA ↗ · 2026-08-22

A new open-source project called FreeToken has been released, featuring a research paper and GitHub repository, with tests showing efficient token processing speeds of around 100 tokens per second on consumer hardware.

0 favorites 0 likes
#tokenization

@cz_binance: Let's tokenize everything. Tokenization is one of the best ways for countries to "raise money", or attract FDI (Foreign…

X AI KOLs Following ↗ · 2026-08-21 Cached

CZ advocates for tokenization on all blockchains to attract FDI. BNB Chain reports a 370% increase in RWA holders in 30 days.

0 favorites 0 likes
#tokenization

SuTRA : Structurally-Unified Tokenization with Root Awareness

arXiv cs.CL ↗ · 2026-08-20 Cached

SuTRA is a morphology-aware tokenization algorithm that preserves akshara indivisibility for Indic languages, reducing morphological shattering and achieving improvements in machine translation metrics over standard BPE methods.

0 favorites 0 likes
#tokenization

What Tokens are Learned when Tokenization is Optimized Jointly with Language Modeling?

arXiv cs.CL ↗ · 2026-08-19 Cached

This paper investigates the tokens learned when tokenization is optimized jointly with language modeling, comparing tokenizer-free methods across multiple languages and finding that they produce distinct, efficient vocabularies for NLP.

0 favorites 0 likes
#tokenization

A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space: What the Units of Voynichese Are Not

arXiv cs.CL ↗ · 2026-08-19 Cached

This paper challenges the standard assumptions about the linguistic units of the Voynich manuscript, showing through quantitative analysis that glyphs, tokens, and spaces do not correspond to letters, words, and word spaces as commonly believed.

0 favorites 0 likes
#tokenization

A Pilot Study of Autocompleting Tokenizers

arXiv cs.CL ↗ · 2026-08-18 Cached

This paper proposes a compression scheme for byte-level tokenization using an autocomplete model to remove predictable bytes from input sequences, reducing sequence length while maintaining machine translation performance across diverse languages.

0 favorites 0 likes
#tokenization

ReconSpan: Reconstruction-Guided Adaptive Latent Tokenization

arXiv cs.CL ↗ · 2026-08-14 Cached

ReconSpan introduces an adaptive latent tokenization method that divides text into reconstructible chunks using a backward decoder, enabling variable-length latent tokens and post-training control of granularity.

0 favorites 0 likes
#tokenization

Lifecycle-Optimal Tokenization: Vocabulary Size as a Deployment-Regime-Dependent Infrastructure Parameter

arXiv cs.LG ↗ · 2026-08-13 Cached

This paper argues that the optimal tokenizer vocabulary size is not fixed but depends on the deployment regime, such as batch size and inference volume. Through roofline analysis and experiments on A10G and A100 GPUs, it shows the lifecycle-optimal vocabulary can shift by up to 16x between on-device and datacenter serving, with minimal quality impact.

0 favorites 0 likes
#tokenization

No wonder Qwen and Gemma are so different

Reddit r/LocalLLaMA ↗ · 2026-08-09

A user shares an observation that Qwen and Gemma tokenize code very differently, with Qwen using far fewer tokens for the same HTML/JS input, which may explain differences in coding and language performance. They also note a potential retraining project by LiquidAI using a more efficient tokenizer.

0 favorites 0 likes
#tokenization

Federated generative event models for tokenized electronic health records

arXiv cs.LG ↗ · 2026-08-05 Cached

This arXiv paper evaluates federated training of tokenized generative event models (GEMs) on ICU EHR data from three health systems, showing that federated learning preserves most centralized performance and improves cross-site transportability compared to conventional supervised models.

0 favorites 0 likes
#tokenization

Pruned BPE: Post-training Visibility Pruning and Token Reallocation for Byte Pair Encoding

arXiv cs.CL ↗ · 2026-08-04 Cached

This paper proposes Pruned BPE, a post-training method that prunes low-exposure tokens from a BPE vocabulary and reallocates slots to better-exposed candidates, reducing encoded length without increasing model-visible vocabulary size. Experiments on English and Chinese corpora show approximately 0.27–0.36% encoded length reduction over standard BPE.

0 favorites 0 likes
#tokenization

@omarsar0: TokTier makes tokenization stateful for agentic serving. Across 153,951 real agent calls with a 94.1% prompt-cache hit …

X AI KOLs Following ↗ · 2026-08-03 Cached

TokTier is a stateful tokenization service that cuts tokenization overhead in agentic LLM serving by reusing stable token boundaries and running exact CPU/GPU tokenization, reducing time-to-first-token by 16–34% under vLLM.

0 favorites 0 likes
#tokenization

Empowering Cross-Domain Sequential Recommendation with Hybrid Tokenization and Serial-Parallel Decoding

arXiv cs.AI ↗ · 2026-08-03 Cached

This paper proposes GenCDSR, a generative framework for cross-domain sequential recommendation with hybrid tokenization and serial-parallel decoding, achieving improved accuracy and significantly reduced inference latency compared to state-of-the-art baselines.

0 favorites 0 likes
#tokenization

Language Models are not Equally Robust to Non-Canonical Tokenization across Languages

arXiv cs.CL ↗ · 2026-07-30 Cached

This paper investigates whether language models remain robust to alternative (non-canonical) tokenizations across 27 languages, finding that invariance observed in English does not generalize and that languages with higher token fragmentation show greater sensitivity. The authors demonstrate that LoRA fine-tuning with multi-tokenization data can mitigate this sensitivity.

0 favorites 0 likes
#tokenization

Joint Optimization for Greedy Longest-match Tokenization

arXiv cs.CL ↗ · 2026-07-28 Cached

This paper introduces JOLT, an integer programming approach to optimize subword tokenization for greedy left-to-right longest-match decoding (WordPiece). JOLT achieves near-optimal compression, closing most of the gap between BPE and the theoretical lower bound, reducing token count by up to 0.78% over BPE.

0 favorites 0 likes
#tokenization

BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis

arXiv cs.CL ↗ · 2026-07-28 Cached

This paper introduces BHARATI, a set of SentencePiece BPE tokenizers trained on a balanced multilingual corpus for classical Indian languages. Subword fertility analysis shows significant improvements in tokenization efficiency, reducing sequence length by up to 90% compared to baseline tokenizers, thereby enhancing effective context length for downstream language models.

0 favorites 0 likes
#tokenization

Brazilian farmers tokenized dairy cows to get loans, bypassing bank limits

Hacker News Top ↗ · 2026-07-25 Cached

Brazilian farmers tokenized 10 dairy cows on the B3 stock exchange, raising nearly $20,000 in credit backed by livestock, using AI-powered tracking collars to prevent fraud and enable movable collateral.

0 favorites 0 likes
#tokenization

GigaToken: ~1000x faster Language model tokenization

Hacker News Top ↗ · 2026-07-22 Cached

GigaToken is an ultra-fast tokenizer library that claims up to 1000x speedup over HuggingFace tokenizers, supporting most common LLM tokenizers and providing drop-in compatibility.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback