Tag
New paper finds that emergent capabilities in LLMs arise randomly from learning sparse attention patterns; the bottleneck is learning which tokens to attend to, which is slow and unpredictable.
XLGoBench introduces a synthetic benchmark of algorithmic tasks to detect cross-lingual skill gaps in LLMs, demonstrating persistent gaps across multiple state-of-the-art models.