decode-speed

Tag

Cards List
#decode-speed

Hot Expert Reload on GPU is what this community needs

Reddit r/LocalLLaMA · 2026-09-11

A request to LLaMA maintainers to implement a feature called 'Hot Expert Reload on GPU' to improve decode speeds for MOE models with moderate active parameters, making them more usable locally with GPUs like the 3090.

0 favorites 0 likes
#decode-speed

Decode speed is the latency tax nobody budgets for in agent loops

Reddit r/AI_Agents · 2026-07-24

Discusses the overlooked latency cost of decode speed in AI agent loops, affecting overall performance.

0 favorites 0 likes
#decode-speed

For RAG specifically, prefill speed matters more than decode and why Strix Halo struggles for interactive use

Reddit r/LocalLLaMA · 2026-07-03

This article explains that for RAG applications, prefill speed matters more than decode speed, and discusses why AMD's Strix Halo APU struggles with interactive use cases.

0 favorites 0 likes
← Back to home

Submit Feedback