Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration
Summary
This paper explores approximating softmax in pretrained LLMs for kernel acceleration, demonstrating performance gains like up to 25.8% speedup on Blackwell B200 with minimal perplexity impact.
View Cached Full Text
Cached at: 09/30/26, 04:23 AM
Paper page - Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration
Source: https://huggingface.co/papers/2609.33586
On Blackwell B200, tensor cores outrun the exponential unit by ~500× (8192 vs 16 ops/clk/SM), soexpbecomes exposed in fused attention. We ask what afrozenpretrained LLM actually needs from softmax — and use the answer to replaceexp2inside FlashAttention-4.
What frozen models need(10 models, 0.5B–72B, no retraining):
- Attention support and within-row resolution can be cut substantially — but uniform weighting on the same support hurts at every tested layer.
- Whereresolution goes matters as much as how much: finer intervals near the row maximum lower NLL in all 10 models, even when overall approximation error goes up.
- A scalar distortion budget is not enough: flattening vs. sharpening at equal attention JSD gives opposite-sign losses depending on the model.
**Rowmax-H15 in FA4:**weights snap to {1, 1.5}×2^k anchored at the kernel’s running row max — no calibration, one FP32 add + one bit shift per element.
- FP8 attention forward on B200:+12.4%(causal 8K),+25.8%(non-causal 8K); 9.3% faster than the best stock FA4 emulation setting we tested
- **−8.4%**board energy per forward (causal 16K)
- Perplexity**+0.09–0.49%**across 5 models from 3 families (BF16 kernel, 2K)
Scope: attention forward on B200; Approximate Softmax Characterizing.
Companion training-side paper (pretraining from scratch with quantized softmax — why the backward rule and calibration gradients matter):https://arxiv.org/abs/2609.33591
Kernel patch and code will be released. Questions and feedback welcome!
Similar Articles
@AlphaSignalAI: You can now boost any LLM's accuracy 2-10x without training it. Most teams improve model accuracy by fine-tuning or swa…
OptiLLM is an open-source proxy that boosts any LLM's accuracy 2-10x by adding extra compute at inference time, using techniques like multi-agent cross-verification and Monte Carlo tree search.
Predictable Scaling Laws of Optimal Hyperparameters for LLM Continued Pre-training
This paper discovers predictable scaling laws for optimal hyperparameters (learning rate, batch size) in LLM continued pre-training, proposing a two-stage framework that reduces hyperparameter search overhead by up to 90% while maintaining performance.
Accelerating GPU Inference of Large Language Models with Moderately Unstructured Sparse Weight Matrices
This paper proposes an efficient GPU inference method for LLMs with moderate unstructured sparsity, introducing a three-layer matrix storage format and a SpMM kernel that jointly utilizes sparse tensor cores and CUDA cores, achieving up to 1.64× kernel-level speedup over SpInfer and up to 1.41× end-to-end speedup over FlashLLM.
MiniMax Sparse Attention
MiniMax Sparse Attention introduces a blockwise sparse attention mechanism that achieves significant speedups for ultra-long-context LLMs, reducing per-token attention compute by 28.4x at 1M context with wall-clock speedups of 14.2x for prefill and 7.6x for decoding on H800 GPUs. The method is accompanied by an open-source inference kernel and a publicly released multimodal model.
AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization
AccelOpt is a self-improving LLM agentic system that autonomously optimizes AI accelerator kernels through iterative generation and optimization memory, achieving 49-61% peak throughput improvements on AWS Trainium while being 26x cheaper than Claude Sonnet 4.