Tag
The paper proposes PISA, a block sparse attention mechanism using a pyramid Top-K selection strategy to achieve O(N log N) complexity, enhancing efficiency for long-context language models with comparable performance on benchmarks and better results on retrieval tasks.