Tag
This paper introduces a composite metric for quantization in small large language models that balances speed and quality by combining information retention and throughput gains, evaluated on models like Gemma 3 1B.
This paper argues that the optimal tokenizer vocabulary size is not fixed but depends on the deployment regime, such as batch size and inference volume. Through roofline analysis and experiments on A10G and A100 GPUs, it shows the lifecycle-optimal vocabulary can shift by up to 16x between on-device and datacenter serving, with minimal quality impact.
A roofline model from the LLM Engineer's Almanac estimates speedups from speculative decoding for different draft lengths across models and hardware, with a note that it may underestimate benefits when overhead is significant.
On Friday, we released six new state-of-the-art drafters for accelerated inference, along with a blog post on speculative decoding and a roofline model tool to estimate speedups.