roofline-model

Tag

Cards List
#roofline-model

A Layer Importance Metric for Quantization Accounting for the Speed-Quality Trade-off in Autoregressive Models

arXiv cs.LG ↗ · 2026-08-28 Cached

This paper introduces a composite metric for quantization in small large language models that balances speed and quality by combining information retention and throughput gains, evaluated on models like Gemma 3 1B.

0 favorites 0 likes
#roofline-model

Lifecycle-Optimal Tokenization: Vocabulary Size as a Deployment-Regime-Dependent Infrastructure Parameter

arXiv cs.LG ↗ · 2026-08-13 Cached

This paper argues that the optimal tokenizer vocabulary size is not fixed but depends on the deployment regime, such as batch size and inference volume. Through roofline analysis and experiments on A10G and A100 GPUs, it shows the lifecycle-optimal vocabulary can shift by up to 16x between on-device and datacenter serving, with minimal quality impact.

0 favorites 0 likes
#roofline-model

@charles_irl: If you're interested in speculative decoding, take some time to grok this chart! And read the article from @haoailab.ht…

X AI KOLs Timeline ↗ · 2026-07-07 Cached

A roofline model from the LLM Engineer's Almanac estimates speedups from speculative decoding for different draft lengths across models and hardware, with a note that it may underestimate benefits when overhead is significant.

0 favorites 0 likes
#roofline-model

@charles_irl: On Friday, we released six new state-of-the-art drafters for accelerated inference. We also put out a blog post on why …

X AI KOLs Following ↗ · 2026-06-22 Cached

On Friday, we released six new state-of-the-art drafters for accelerated inference, along with a blog post on speculative decoding and a roofline model tool to estimate speedups.

0 favorites 0 likes
← Back to home

Submit Feedback