lookahead-sparse-attention

Tag

Cards List
#lookahead-sparse-attention

FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention

Hugging Face Daily Papers · 2026-06-08 Cached

Proposes Lookahead Sparse Attention with a Neural Memory Indexer on DeepSeek-V4, reducing GPU memory usage to ~13.5% of full-context baseline while maintaining or slightly improving accuracy.

0 favorites 0 likes
← Back to home

Submit Feedback