Tag
This article details a series of low-level optimizations for binary search in Rust, achieving a 6x speedup by leveraging CPU architecture features like branch prediction and SIMD, applied to a scikit-learn gradient boosting use case.
A blog post demonstrates how adding a conditional check that appears useless can dramatically improve loop performance by allowing the CPU's branch predictor to eliminate data dependencies, achieving up to 4x speedup in a specific compression algorithm.
Speculative decoding, inspired by 1990s CPU branch prediction, is now used by Anthropic, Google, and Meta to speed up LLM inference 2-3x. It uses a small model to guess future tokens and a large model to verify them in parallel, avoiding idle GPU time during decoding.