Tag
This paper identifies a curse of ambiguity in language models, where more ambiguous next-token distributions are harder to learn, tracing this to architectural and learning roots and validating on synthetic and real data.