Tag
A request to LLaMA maintainers to implement a feature called 'Hot Expert Reload on GPU' to improve decode speeds for MOE models with moderate active parameters, making them more usable locally with GPUs like the 3090.