Tag
A research paper from Meta demonstrates that byte-level language models, initially inferior to token-based models, can outperform them as computational resources increase, shown through distilled 1B models trained on up to 1 trillion bytes.
A paper from Meta shows that byte-level models start behind token models but surpass them with increasing compute, demonstrated with distilled 1B models trained on up to 1 trillion bytes.