Tag
This paper compares subword tokens, raw bytes, and rendered pixels as text encodings for language models under controlled linguistic content across 13 languages. It traces rate–utility frontiers and finds that no encoding dominates across tasks, with pixels preserving surface form best, bytes preserving cross-lingual alignment best, and tokens supporting topic prediction best.