flash-attention

Tag

Cards List
#flash-attention

@__solo69__: Flash attention follows this principle, minimize expensive data movement and maximize reuse of data that has already re…

X AI KOLs Timeline ↗ · 2026-09-14 Cached

A tweet explains the optimization principle behind Flash attention, focusing on minimizing data movement and maximizing on-chip memory reuse, in the context of a language modeling course lecture.

0 favorites 0 likes
#flash-attention

@zhangchitc: Why is FlashAttention both fast and GPU memory-efficient?

X AI KOLs Timeline ↗ · 2026-09-12

A tweet asking why FlashAttention achieves both speed and GPU memory efficiency in AI computations.

0 favorites 0 likes
#flash-attention

CUDA/HIP: Flash Attention tuning (gfx1201) by pwilkin · Pull Request #28102 · ggml-org/llama.cpp

Reddit r/LocalLLaMA ↗ · 2026-09-11 Cached

This pull request introduces Flash Attention tuning optimizations for CUDA/HIP in the llama.cpp project, targeting gfx1201 hardware to enhance inference performance.

0 favorites 0 likes
#flash-attention

@pallavishekhar_: Learn LLM Inference Engineering - Prefill vs Decode - KV Cache - PagedAttention - Flash Attention - Continuous Batching…

X AI KOLs Timeline ↗ · 2026-08-28 Cached

An educational overview of key concepts in LLM inference engineering, covering techniques like KV cache, PagedAttention, Flash Attention, and continuous batching to optimize inference performance.

0 favorites 0 likes
#flash-attention

Why Speculative Decoding went mature in 2026?

Reddit r/LocalLLaMA ↗ · 2026-08-10

A discussion on why speculative decoding matured in 2026 for LLM inference, citing Uber's early use, Apple and DeepMind papers, and Tri Dao's research, with observations on adoption in frameworks and local deployments.

0 favorites 0 likes
#flash-attention

@charliermarsh: At Astral, we created pre-built wheels for popular GPU-enabled packages (like FlashAttention and DeepSpeed) and distrib…

X AI KOLs Following ↗ · 2026-07-30 Cached

Astral is open sourcing its build pipelines for pre-built wheels of GPU-enabled Python packages like FlashAttention and DeepSpeed, making them available to all via standard Python indexes.

0 favorites 0 likes
#flash-attention

For V100 Users: SGLang running Qwen+Dflash and Laguna

Reddit r/LocalLLaMA ↗ · 2026-07-27

Modified SGLang to support Qwen and Laguna models on V100 GPUs using custom FlashAttention and Marlin kernels, achieving decent throughput on 4xV100 hardware.

0 favorites 0 likes
#flash-attention

Google is updating Gemma 4's chat templates, bringing major fixes to tool calling and reducing "laziness", and enabling Flash Attention 4 on Hopper GPUs, plus an interactive guide on how to work with and improve its vision!

Reddit r/LocalLLaMA ↗ · 2026-07-15 Cached

Google updated Gemma 4's chat templates with major fixes to tool calling, reduced laziness, enabled Flash Attention 4 on Hopper GPUs, and released an interactive vision guide. The updates are available on Hugging Face.

0 favorites 0 likes
#flash-attention

@Alacritic_Super: The biggest bottleneck in LLM inference isn't arithmetic but it's moving data. A single multiply-accumulate operation i…

X AI KOLs Timeline ↗ · 2026-07-15 Cached

An educational thread explaining that the main bottleneck in LLM inference is data movement, not computation, and highlighting techniques like quantization, KV cache optimization, and FlashAttention to reduce memory traffic.

0 favorites 0 likes
#flash-attention

Flash-MSA: Accelerating Million-Token Training with Sparse Attention Kernels

Hacker News Top ↗ · 2026-07-12 Cached

Introduces Flash-MSA, the first performant open-source training kernels for MiniMax Sparse Attention on Hopper and Blackwell GPUs, enabling efficient million-token training.

0 favorites 0 likes
#flash-attention

@Alacritic_Super: If you are serious about LLM inference, study FlashAttention. It's one of the most important optimizations behind moder…

X AI KOLs Timeline ↗ · 2026-07-08 Cached

A tweet recommending studying FlashAttention for LLM inference, highlighting its importance in optimizing GPU memory traffic and speeding up attention mechanisms, with links to the GitHub repository and papers for FlashAttention, FlashAttention-2, and FlashAttention-3.

0 favorites 0 likes
#flash-attention

@ekzhang1: me looking at people like this guy who write real gpu kernels :)

X AI KOLs Timeline ↗ · 2026-07-08 Cached

AI model Claude was used to write a FlashAttention forward kernel using the pyptx DSL, achieving near-parity performance with hand-tuned FlashAttention-4 on NVIDIA B200 hardware.

0 favorites 0 likes
#flash-attention

@PatrickToulme: This exercise makes me believe the future of DSLs and compilers is very much agentic. Programming languages and DSLs bu…

X AI KOLs Timeline ↗ · 2026-07-07 Cached

Claude Fable used the pyptx DSL to write a FlashAttention forward kernel for NVIDIA B200 that achieves near-parity performance with the hand-tuned CUTLASS kernel, demonstrating the potential for AI agents in compiler and DSL design.

0 favorites 0 likes
#flash-attention

https://x.com/seclink/status/2072187033263784397

X AI KOLs Timeline ↗ · 2026-07-01 Cached

Hybrid Sliding Window Attention (Hybrid SWA) is a mixed attention mechanism in long-context language models that balances computational efficiency with full long-range dependencies. By alternating between local SWA layers and global attention layers, it significantly compresses KV cache while maintaining inference capability. This article details its design principles, application in models such as Gemma and Qwen, and best practices in open-source projects like vLLM and HuggingFace.

0 favorites 0 likes
#flash-attention

Mechanism-Driven Monitors for Preemptive Detection of LLM Training Instability

arXiv cs.CL ↗ · 2026-06-29 Cached

Proposes mechanism-driven monitors for preemptive detection of LLM training instability by deriving internal signals from low-precision flash attention and MoE routers, enabling detection thousands of steps before loss divergence.

0 favorites 0 likes
#flash-attention

Modern GPU Programming for MLSys

Hacker News Top ↗ · 2026-06-23 Cached

A new book from CMU's Machine Learning Systems course teaches modern GPU programming for ML systems, covering Blackwell architecture, GEMM, and FlashAttention using the TIRx Python DSL.

0 favorites 0 likes
#flash-attention

@yukangchen_: We are excited to share a new technical article “KV Cache Compression and Its Infra Problems.” https://research.nvidia.…

X AI KOLs Timeline ↗ · 2026-06-16 Cached

NVIDIA Research publishes a technical blog post examining KV cache compression techniques and their infrastructure problems, including how FlashAttention and paged attention create practical obstacles for production deployment of long-context LLMs, with a proposed geometric solution using RoPE.

0 favorites 0 likes
#flash-attention

@charles_irl: Rewriting parallelism is a big move and it'd be nice to make it even faster than we can do with CuTe DSL. FA4 is a very…

X AI KOLs Following ↗ · 2026-06-11 Cached

Discussion about rewriting parallelism to improve kernel performance using CuTe DSL and tile programming models for the FA4 (FlashAttention 4) kernel.

0 favorites 0 likes
#flash-attention

@charles_irl: A tl;dr for folks who don't care how many warpgroups FA4 devotes to softmax vs MMA loads. Inference is different from t…

X AI KOLs Following ↗ · 2026-06-11 Cached

Explains that inference kernels differ from training, with Flash Attention 4 focusing on changing parallelism across KV and supporting small irregular loads.

0 favorites 0 likes
#flash-attention

@maximelabonne: Parallax is a parametrized form of Local Linear Attention that drops the numerical solvers and matches FA 2/3 on decode…

X AI KOLs Following ↗ · 2026-06-10 Cached

Parallax is a new parametrized form of Local Linear Attention that eliminates numerical solvers and matches FlashAttention 2/3 in decoding. Its effectiveness depends on the optimizer, working with Muon but not AdamW, highlighting the role of optimizer geometry.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback