Tag
A comprehensive survey on tokenization in modern NLP, compiled by 32 tokenizer researchers, covering algorithms, evaluations, multilinguality, encodings, theory, and alternatives like latent or visual tokenization, plus adjacent topics such as constrained generation and tokenizer security.
This paper investigates the curse of multilinguality in lexical normalization, finding that training a single model on multiple languages leads to decreased per-language accuracy, with optimal performance when languages are trained in small groups.
This paper introduces the 'culture funnel' concept, demonstrating that cultural signals in LLM training data sharply decline during post-training stages. The authors release a 5.6M-sample tagged dataset to help preserve cultural grounding in model alignment.