Tag
The article describes experiments on split-K matrix multiplications, revealing that answer variability depends on block layouts and split counts, with tests on a B200 GPU showing output changes under different conditions.
The author experimented with KV cache blending by splitting prompts into chunks with overlap, achieving a 3x boost in prefill speed without affecting retrieval tasks on Ling3-tiny.