tokenization

Tag

Cards List
#tokenization

@maximelabonne: How do you add new languages to a model? Efficient tokenization plays a major role here. In this new blog post, we talk…

X AI KOLs Following ↗ · 2026-07-21 Cached

Liquid AI shares a recipe for upgrading a pretrained model's tokenizer in place, expanding LFM2.5's tokenizer from 65K to 128K to improve efficiency for languages like Thai, Vietnamese, and Hindi, resulting in up to 4x fewer tokens and 2.2-3.7x faster generation.

0 favorites 0 likes
#tokenization

Tokenizing Crosslingual Homographs

arXiv cs.CL ↗ · 2026-07-21 Cached

This paper investigates how multilingual tokenizers handle cross-lingual homographs (identical surface forms with different meanings across languages) and proposes a lightweight language-cue intervention that introduces language-specific characters to reduce token sharing. Experiments show modest improvements in machine translation, particularly with BPE tokenization.

0 favorites 0 likes
#tokenization

Discrete Diffusion Models: A Unified Framework from Tokenization to Generation

arXiv cs.LG ↗ · 2026-07-16 Cached

This paper introduces a unified conceptual framework for discrete diffusion models, analyzing their design space through tokenization, state space construction, and highlighting trade-offs in training, inference, and scaling.

0 favorites 0 likes
#tokenization

Real Up Cannes Summit Unveils "Anvita" to Power the Agent-to-Agent Economy

Reddit r/ArtificialInteligence ↗ · 2026-07-11 Cached

Ant Digital Technologies launched Anvita, a new Web3+AI brand, at the Real Up Cannes Summit to build infrastructure for the agent-to-agent (A2A) digital economy, featuring Anvita Flow for autonomous agent collaboration and Anvita TaaS for tokenization-as-a-service.

0 favorites 0 likes
#tokenization

Let's build a simple interpreter for APL

Lobsters Hottest ↗ · 2026-07-10 Cached

This blog post starts a series on building an APL interpreter in Python, covering tokenization and parsing of basic APL expressions with numbers, functions, and operators.

0 favorites 0 likes
#tokenization

Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES

arXiv cs.CL ↗ · 2026-07-08 Cached

This paper systematically compares BPE and Unigram-LM tokenization methods for SMILES strings in chemistry, finding they produce near-disjoint vocabularies and differ significantly in segmentation granularity. The results show that the choice of subword algorithm is a critical modeling decision, not a free default.

0 favorites 0 likes
#tokenization

Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization

Hugging Face Daily Papers ↗ · 2026-07-05 Cached

This paper proposes a speaker-disentangled syllabic tokenizer that uses regression of perturbed student representations toward clean teacher targets, achieving state-of-the-art syllable boundary detection and a 7% relative improvement in speech language model understanding over SpiRit-LM.

0 favorites 0 likes
#tokenization

@VictoriaLinML: The video of my Stanford CS25 guest lecture, From Language Models to Native Multimodal Intelligence, is now online. I d…

X AI KOLs Timeline ↗ · 2026-07-03 Cached

Victoria Lin's Stanford CS25 lecture discusses the paradigm shift from language models to native multimodal intelligence, covering tokenization approaches and comparing models like Chameleon and Transfusion.

0 favorites 0 likes
#tokenization

AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation

Hugging Face Daily Papers ↗ · 2026-06-29 Cached

AVTok proposes a unified 1D tokenizer for audio-video generation using a dual-stream transformer with shared encoder-decoder and modal-specific queries, achieving compact latent representations and excelling in reconstruction and downstream generation tasks.

0 favorites 0 likes
#tokenization

BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language

Hugging Face Daily Papers ↗ · 2026-06-29 Cached

BrainJanus is the first unified brain model that integrates brain, vision, and language via a shared Omni space, enabling bidirectional mapping between neural activity and sensory stimuli through tokenized representation and autoregressive next-token prediction.

0 favorites 0 likes
#tokenization

Scenes as Objects, Not Primitives: Instance-Structured 3D Tokenization from Unposed Views

Hugging Face Daily Papers ↗ · 2026-06-28 Cached

This paper proposes a feed-forward framework that decomposes 3D scenes into instance-structured token groups from unposed multi-view images, enabling direct object-level reconstruction, segmentation, and manipulation without 3D annotations.

0 favorites 0 likes
#tokenization

The African Language Tax: Quantifying the Cost, Latency, and Context Penalty of Tokenizing African Languages in Frontier LLMs

arXiv cs.CL ↗ · 2026-06-24 Cached

This paper systematically quantifies the tokenization penalty for 20 African languages across 11 frontier and open tokenizers, finding up to 8.9× inference cost and latency multipliers and as little as 11% effective context window compared to English, highlighting a structural digital divide encoded in subword vocabularies.

0 favorites 0 likes
#tokenization

Best Preprocessing Techniques for Sentiment Analysis

arXiv cs.CL ↗ · 2026-06-24 Cached

This paper systematically investigates the optimal order of preprocessing techniques for sentiment analysis on Twitter data, finding that tokenisation is most impactful and spelling correction least, with the best order being tokenisation, cleaning, stemming, then stopword removal.

0 favorites 0 likes
#tokenization

QuechuaTok: Morphological Boundary Accuracy as a Necessary Metric for Tokenizer Evaluation in Agglutinative Low-Resource Languages

arXiv cs.CL ↗ · 2026-06-24 Cached

This paper presents QuechuaTok, a benchmark for evaluating tokenization strategies for Southern Quechua, and introduces Morphological Boundary Accuracy (MorphAcc) as a necessary metric. It shows that BPE achieves low fertility but poor morphological accuracy, while a morphology-aware PRPE tokenizer achieves 83% MorphAcc, demonstrating that fertility rate alone is insufficient for agglutinative languages.

0 favorites 0 likes
#tokenization

The Galaxy's Guide to the Tokenizer: A Benchmark for Scientific Foundation Models

Hugging Face Daily Papers ↗ · 2026-06-24 Cached

This paper compares four tokenization methods (Affine, AIM, JetFormer, VQ-VAE) for astronomical images within a unified transformer framework, using 640,000 galaxy images to evaluate reconstruction quality, physical property prediction, and morphological preservation. It finds that no single method excels across all tasks, highlighting trade-offs in representation learning.

0 favorites 0 likes
#tokenization

Toten: Knowledge-Based Ontological Tokenization Of Physical Quantities And Technical Notation In Brazilian Portuguese

arXiv cs.AI ↗ · 2026-06-20 Cached

TOTEN is a knowledge-based ontological tokenization framework that replaces statistical tokenization with declarative classification grounded in a formal ontology of engineering entities, achieving high ontological atomicity and numerical reconstruction for physical quantities and technical notation in Brazilian Portuguese.

0 favorites 0 likes
#tokenization

BioMatrix: Towards a Comprehensive Biological Foundation Model Spanning the Modality Matrix of Sequences, Structures, and Language

Hugging Face Daily Papers ↗ · 2026-06-20 Cached

BioMatrix is a multimodal foundation model that unifies molecular sequences, structures, and natural language in a single decoder-only architecture, achieving state-of-the-art performance on 77 out of 80 biological tasks.

0 favorites 0 likes
#tokenization

Hallucinations = Imagination

Reddit r/ArtificialInteligence ↗ · 2026-06-18

A developer working on an AI agent wrapper observes that the agent's hallucinations of user responses can actually aid problem-solving, and proposes treating such hallucinations as imagined events rather than errors.

0 favorites 0 likes
#tokenization

Beyond Tokenization: Direct Timestep Embedding and Contrastive Alignment for Time-Series Question Answering

arXiv cs.CL ↗ · 2026-06-18 Cached

This paper introduces CADE, a framework for time-series question answering that maps each timestep directly into the LLM embedding space and uses a one-directional supervised contrastive loss to align time-series representations with frozen text anchors, outperforming existing baselines on the Time-MQA benchmark.

0 favorites 0 likes
#tokenization

Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish

arXiv cs.CL ↗ · 2026-06-18 Cached

This paper presents Morpheus, a neural tokenizer and word embedder for Turkish that learns morpheme boundaries without string normalization, achieving lossless tokenization and competitive embeddings for lexical retrieval, while using less GPU memory than subword tokenizers.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback