llama-framework

Tag

Cards List
#llama-framework

Hot Expert Reload on GPU is what this community needs

Reddit r/LocalLLaMA · 2026-09-11

A request to LLaMA maintainers to implement a feature called 'Hot Expert Reload on GPU' to improve decode speeds for MOE models with moderate active parameters, making them more usable locally with GPUs like the 3090.

0 favorites 0 likes
← Back to home

Submit Feedback