decode

Tag

Cards List
#decode

@dongxi_nlp: https://x.com/dongxi_nlp/status/2099713825402425623

X AI KOLs Timeline ↗ · 2026-09-15 Cached

This article explains how the prefilling and decoding phases operate in AI models and their effects on generation speed and performance, aiding in understanding the reasons behind AI response latency.

0 favorites 0 likes
#decode

@Hass_Abdallah11: We have two new Colfax articles on optimizing FlashAttention-4, featuring work by myself (decode) and Jack Carlisle (bw…

X AI KOLs Timeline ↗ · 2026-09-08 Cached

The article details optimization techniques for FlashAttention-4, including an S/P ping-pong method for decode to overlap operations, achieving up to 16% performance gain on NVIDIA Blackwell GPUs.

0 favorites 0 likes
#decode

Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation

arXiv cs.AI ↗ · 2026-08-14 Cached

This paper introduces Dual-Flow Transformers, an architecture that decouples prefill and decode computation by adding an auxiliary flow for continuation prediction while sharing weights and the primary KV cache, enabling phase-specific compute allocation and improved efficiency.

0 favorites 0 likes
#decode

A 124B emitted 15,128 tokens in a single response on one DGX Spark, decode went 35.62 → 35.68 tok/s across the whole thing

Reddit r/LocalLLaMA ↗ · 2026-08-13

An observation of Ling-3.0-flash (124B) on one DGX Spark generating 15,128 tokens in a single response with stable decode throughput around 35.6 tok/s, highlighting long-context decoding performance.

0 favorites 0 likes
#decode

Llama-CPP Parallel Agents --> fine for decode, but one agent's prefill will grind all other agents to a halt

Reddit r/LocalLLaMA ↗ · 2026-08-11

A user reports that in Llama-CPP with parallel sub-agents, decode performance is great but a single agent's prefill (e.g., processing a web search) stalls all other agents, and asks for tuning suggestions.

0 favorites 0 likes
#decode

llama.cpp b9966 for sm-tensor

Reddit r/LocalLLaMA ↗ · 2026-07-11

llama.cpp b9966 introduces a fix for the -sm tensor mode that caches regex patterns, eliminating 29 recompilations per tensor per token on the decode thread, resulting in significantly reduced CPU overhead.

0 favorites 0 likes
#decode

@_avichawla: Prefill & decode in LLM inference. Have you ever noticed that the first token from an LLM always takes a moment to appe…

X AI KOLs Timeline ↗ · 2026-06-29 Cached

Explains the two phases of LLM inference - prefill and decode - detailing how GPU bottlenecks shift from compute-bound during prefill to memory-bound during decode, and the importance of KV caching.

0 favorites 0 likes
#decode

@robertnishihara: A great example of the importance of disaggregation in RL. From the paper LLM generation alternates between prefill and…

X AI KOLs Following ↗ · 2026-06-20 Cached

Robert Nishihara highlights a paper on disaggregating RL workloads, showing that using compute-optimized H800s for prefill and bandwidth-optimized H20s for decode can cut rollout times by 21-51% and 47% respectively, emphasizing that no single hardware type fits all stages.

0 favorites 0 likes
#decode

@rohanpaul_ai: Chamath on all important “prefill” and “decode.” in AI compute. Prefill is compute-bound; massive parallel GPUs win, so…

X AI KOLs Following ↗ · 2026-05-24 Cached

Chamath explains the two key phases of AI compute: prefill, which is compute-bound and favors parallel GPUs like Nvidia's, and decode, which is memory-bandwidth bound and depends on scanning previously generated tokens.

0 favorites 0 likes
#decode

@no_stp_on_snek: @antirez Turbo3 BEATS fp8 by +5% decode tok/s at 32K context still tinkering but i've been cooking TQ+ in your kitchen

X AI KOLs Following ↗ · 2026-05-23 Cached

Turbo3 achieves 5% faster decode tokens per second compared to fp8 at 32K context, a performance improvement in quantization or model optimization.

0 favorites 0 likes
← Back to home

Submit Feedback