flashattention

Tag

Cards List
#flashattention

FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving

Hugging Face Daily Papers · 2026-08-20 Cached

FlashPrefill V2 improves long-context LLM serving through mean-corrected sparse attention and optimized GPU operators, delivering substantial speedups over FlashAttention-2 and dense baselines.

0 favorites 0 likes
#flashattention

@akshay_pachaar: GPU architecture, clearly explained. The usual assumption is that a faster GPU means more compute, so a chip rated for …

X AI KOLs Timeline · 2026-08-15 Cached

The article clarifies that GPU performance in AI inference is limited by memory bandwidth rather than compute power, using the NVIDIA H100 as an example to explain GPU architecture and its effect on token generation rates.

0 favorites 0 likes
#flashattention

Exploring FlashAttention-3/4 optimizations on RTX GPUs

Reddit r/LocalLLaMA · 2026-07-09

This article explores whether FlashAttention-3/4 optimizations benefit RTX GPUs, concluding that FA-2 is the ceiling due to hardware limitations on consumer cards.

0 favorites 0 likes
#flashattention

@h100envy: Dan Fu co-wrote FlashAttention with Tri Dao. Then he co-built Hyena, Monarch Mixer, and ThunderKittens. Now he's distin…

X AI KOLs Timeline · 2026-06-28 Cached

Profiles Dan Fu, a key contributor to high-performance kernels like FlashAttention, Hyena, Monarch Mixer, and ThunderKittens, now a distinguished researcher at Together AI whose work is used in ChatGPT, Claude, and Gemini.

0 favorites 0 likes
#flashattention

@thtrkim: Visual deep dive on FlashAttention by hand (drawn with Excalidraw) https://winterrykim.github.io/blog/2026/training-lm-…

X AI KOLs Timeline · 2026-06-23 Cached

A visual deep dive into FlashAttention, explaining memory optimization and operator fusion for efficient attention computation in language model training.

0 favorites 0 likes
← Back to home

Submit Feedback