byte-level-language-model

Tag

Cards List
#byte-level-language-model

Kathleen Writes: Autoregressive Generation and Data Scaling Without Attention

arXiv cs.CL · 2026-08-06 Cached

This paper from the Kathleen series shows that an attention-free, byte-level model with ~0.5M parameters can beat a parameter-matched transformer on WikiText-103 language modeling and generation, introduces a non-parametric 'Form Distance' metric for evaluating text realism, and demonstrates that retrieval-augmented decoding from the model's own training corpus improves generation quality.

0 favorites 0 likes
← Back to home

Submit Feedback