Tag
Google DeepMind introduces a routing method framed as a Pandora's Box problem to efficiently allocate compute by deciding when to invest in better model selection estimates, demonstrating improved performance on benchmarks like MATH, RAG, and EmbedLLM.
The paper introduces the Skaling law, a generalized neural scaling law that couples model capacity and data through an interaction exponent, reducing prediction error by 1.5-3x and enabling full-grid extrapolation using roughly 10x less compute.
This paper demonstrates that scaling laws fit on small transformer models can accurately predict the loss of much larger models trained on particle physics jet data, enabling compute budgets to be translated into expected physics performance before large training runs. They release five pretrained models and the full training recipe.