Tag
COBS introduces a cumulant-order block sparse attention method that improves block selection by using compressed second-order statistics, achieving near-dense attention accuracy on long-context benchmarks while significantly reducing KV cache read traffic.
Proposes an uncertainty-gated router that doubles the selected key blocks for queries with uncertain cutoff margins, improving recall and accuracy in block-sparse attention for long-context language models, validated on multiple architectures.