sparse-weight-matrices

Tag

Cards List
#sparse-weight-matrices

Accelerating GPU Inference of Large Language Models with Moderately Unstructured Sparse Weight Matrices

arXiv cs.LG · 2026-07-13 Cached

This paper proposes an efficient GPU inference method for LLMs with moderate unstructured sparsity, introducing a three-layer matrix storage format and a SpMM kernel that jointly utilizes sparse tensor cores and CUDA cores, achieving up to 1.64× kernel-level speedup over SpInfer and up to 1.41× end-to-end speedup over FlashLLM.

0 favorites 0 likes
← Back to home

Submit Feedback