tokenization

Tag

Cards List
#tokenization

An In-Vitro Study on Cross-Lingual Generalization in Language Models

arXiv cs.CL ↗ · 2026-05-27 Cached

This paper introduces an in-vitro framework with two procedurally generated languages to study cross-lingual generalization in language models, finding that tokenization's preservation of reusable substructure is more critical than lexical similarity or data balance for transferring capabilities across languages.

0 favorites 0 likes
#tokenization

BrickAnything: Geometry-Conditioned Buildable Brick Generation with Structure-Aware Tokenization

arXiv cs.AI ↗ · 2026-05-27 Cached

BrickAnything is an autoregressive framework that generates physically buildable brick structures from diverse 3D representations using point clouds and structure-aware tree tokenization, ensuring geometric fidelity and structural stability.

0 favorites 0 likes
#tokenization

Do machines think or tokenize?

Reddit r/artificial ↗ · 2026-05-26

This paper introduces the SAPS (Synthetic Algorithmic Predictive Systems) framework, arguing that modern AI systems do not think but tokenize and compute statistical patterns, and clarifies the critical distinction between artificial and synthetic systems.

0 favorites 0 likes
#tokenization

@gordic_aleksa: new in-depth blog post time: Inside the Transformer: The Life of a Token a deep dive into a modern dense transformer, i…

X AI KOLs Timeline ↗ · 2026-05-26 Cached

An in-depth blog post exploring the inner workings of modern dense transformers, covering topics such as YaRN for positional information, hybrid attention for long context lengths, soft capping, QK normalization, and transformer math including FLOPs/token formulas and cluster sizing.

0 favorites 0 likes
#tokenization

Lost in Translation? Exploring the Shift in Grammatical Gender from Latin to Occitan

Hugging Face Daily Papers ↗ · 2026-05-26 Cached

A deep learning framework is developed to analyze grammatical gender evolution from Latin to Romance languages, focusing on low-resource historical settings using lexical and contextual analysis.

0 favorites 0 likes
#tokenization

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Hugging Face Daily Papers ↗ · 2026-05-25 Cached

LLaVA-OneVision-2 introduces codec-stream tokenization and windowed attention for efficient video understanding, achieving state-of-the-art performance across multiple multimodal benchmarks including video, spatial, and tracking tasks.

0 favorites 0 likes
#tokenization

@shabnam_774: https://x.com/shabnam_774/status/2058517919760355729

X AI KOLs Timeline ↗ · 2026-05-24 Cached

This article provides a comprehensive step-by-step breakdown of how modern Large Language Models like ChatGPT and Claude are built from scratch, covering data collection, tokenization, transformer architectures, training, alignment, and deployment.

0 favorites 0 likes
#tokenization

@Tabbu_ai: https://x.com/Tabbu_ai/status/2058145123444347339

X AI KOLs Timeline ↗ · 2026-05-23 Cached

An educational thread explaining 11 key lessons for understanding and building LLM architectures from scratch, covering tokens, embeddings, attention, positional encoding, data quality, and common misconceptions.

0 favorites 0 likes
#tokenization

AI agents are making tokenization platforms far more usable than I expected

Reddit r/AI_Agents ↗ · 2026-05-20

A developer shares how AI agents are improving tokenization platforms through intelligent orchestration of humans and systems, rather than full autonomy.

0 favorites 0 likes
#tokenization

Learning Faster with Better Tokens: Parameter-Efficient Vocabulary Adaptation for Specialized Text Summarization

arXiv cs.CL ↗ · 2026-05-19 Cached

This paper proposes a parameter-efficient vocabulary adaptation method for LLM-based text summarization in specialized domains, augmenting pretrained tokenizers with domain-specific tokens and selectively replacing under-trained ones to reduce training time by 35-55% and parameter counts by up to 37%.

0 favorites 0 likes
#tokenization

@nemild: Y Combinator is hosting fintech happy hour on Thursday in New York City. Thinking of a startup in stablecoins, tokeniza…

X AI KOLs Timeline ↗ · 2026-05-18 Cached

Y Combinator is hosting a fintech happy hour on Thursday in New York City, inviting startups working on stablecoins, tokenization, financial AI, agentic commerce, and prediction markets.

0 favorites 0 likes
#tokenization

Dywave: Event-Aligned Dynamic Tokenization for Heterogeneous IoT Sensing Signal

arXiv cs.LG ↗ · 2026-05-15 Cached

Dywave is a dynamic tokenization framework for IoT sensing signals that uses wavelet-based hierarchical decomposition to align tokens with semantic events, achieving up to 12% higher accuracy and 75% reduction in input token length on five real-world datasets.

0 favorites 0 likes
#tokenization

A Calculus-Based Framework for Determining Vocabulary Size in End-to-End ASR

arXiv cs.CL ↗ · 2026-05-15 Cached

This paper presents a calculus-based framework that uses first and second derivative tests to estimate the optimal vocabulary size hyper-parameter for end-to-end ASR systems, improving performance on the Librispeech corpus.

0 favorites 0 likes
#tokenization

From Token to Token Pair: Efficient Prompt Compression for Large Language Models in Clinical Prediction

arXiv cs.CL ↗ · 2026-05-13 Cached

This paper introduces MedTPE, a method for efficient, lossless prompt compression of electronic health records for large language models, significantly reducing token length and inference latency in clinical prediction tasks.

0 favorites 0 likes
#tokenization

Compute Optimal Tokenization (2 minute read)

TLDR AI ↗ · 2026-05-13 Cached

This paper systematically derives compression-aware neural scaling laws by training nearly 1,300 models, demonstrating that the widely used heuristic of 20 tokens per parameter is an artifact of specific tokenizers. The authors propose a tokenizer-agnostic scaling law based on bytes, offering a new framework for compute-efficient training across diverse languages and modalities.

0 favorites 0 likes
#tokenization

The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval

arXiv cs.CL ↗ · 2026-05-11 Cached

This research paper investigates the 'Text Uncanny Valley,' a phenomenon where LLM performance in information retrieval tasks degrades non-monotonically as word-boundary corruption increases. The authors propose a mode transition hypothesis to explain this U-shaped performance curve and demonstrate its relevance to real-world noisy text inputs.

0 favorites 0 likes
#tokenization

@0xLogicrw: MiniMax published a technical blog post detailing the root cause analysis for its M2 series large models' inability to output the person's name "Ma Jiaqi". Starting from a single case study, the investigation ultimately revealed a systematic degradation issue affecting nearly 5% of the entire vocabulary. The root cause was a severe disconnect in data coverage between the two training stages of the large model. In the first stage (pre-training), massive amounts of internet text were used to cre…

X AI KOLs Timeline ↗ · 2026-05-10

MiniMax published a technical blog post providing an in-depth analysis of the systematic vocabulary degradation issue behind its M2 series large models' inability to output specific personal names. It reveals parameter shifts caused by a disconnect in data coverage between pre-training and post-training stages, and proposes an effective solution involving full-scale synthetic data for remediation.

0 favorites 0 likes
#tokenization

Human typing habits and token counts

Hacker News Top ↗ · 2026-05-08 Cached

A blog post exploring how human typing habits like typos, shorthand, filler words, and whitespace affect token counts in OpenAI and Claude tokenizers, noting that common misspellings can inflate token usage and costs without changing meaning.

0 favorites 0 likes
#tokenization

When Informal Text Breaks NLI: Tokenization Failure, Distribution Shift, and Targeted Mitigations

arXiv cs.CL ↗ · 2026-04-21 Cached

This paper investigates how informal text (slang, emoji, Gen-Z filler tokens) degrades NLI accuracy in ELECTRA-small and RoBERTa-large models, identifying two distinct failure mechanisms—tokenization failure (emoji mapped to [UNK]) and distribution shift (out-of-domain noise tokens)—and proposes targeted mitigations that recover accuracy without harming clean-text performance.

0 favorites 0 likes
#tokenization

Defragmenting Language Models: An Interpretability-based Approach for Vocabulary Expansion

arXiv cs.CL ↗ · 2026-04-21 Cached

Researchers from University of Utah and CMU propose FragMend, an interpretability-based approach for vocabulary expansion in LLMs that addresses token over-fragmentation in non-Latin script languages. Their method outperforms frequency-based vocabulary selection and baseline embedding initialization by ~20 points for several underrepresented languages.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback