Tag
A request to LLaMA maintainers to implement a feature called 'Hot Expert Reload on GPU' to improve decode speeds for MOE models with moderate active parameters, making them more usable locally with GPUs like the 3090.
Discusses the overlooked latency cost of decode speed in AI agent loops, affecting overall performance.
This article explains that for RAG applications, prefill speed matters more than decode speed, and discusses why AMD's Strix Halo APU struggles with interactive use cases.