prefill-speed

Tag

Cards List
#prefill-speed

Anyone else tried out KV cache blending?

Reddit r/LocalLLaMA · 2026-08-22

The author experimented with KV cache blending by splitting prompts into chunks with overlap, achieving a 3x boost in prefill speed without affecting retrieval tasks on Ling3-tiny.

0 favorites 0 likes
#prefill-speed

I benchmarked 13 models at 65K-128K context to find out what actually matters for agentic workloads

Reddit r/LocalLLaMA · 2026-07-05

An extensive benchmark of 13 local LLMs at 65K-128K context shows that prefill speed dominates agentic workload performance (94-99% of wall-clock time), rendering tg128 misleading, and that KV head count is the key architectural factor over parameter count or MoE/dense design.

0 favorites 0 likes
#prefill-speed

For RAG specifically, prefill speed matters more than decode and why Strix Halo struggles for interactive use

Reddit r/LocalLLaMA · 2026-07-03

This article explains that for RAG applications, prefill speed matters more than decode speed, and discusses why AMD's Strix Halo APU struggles with interactive use cases.

0 favorites 0 likes
← Back to home

Submit Feedback