gpu-offloading

Tag

Cards List
#gpu-offloading

Got an old slow low vram GPU laying around? Might be worth it to use for Just Vision mmproj llama.cpp

Reddit r/LocalLLaMA · 5d ago

The article suggests using an old GPU to offload the mmproj component in llama.cpp, which improves speed for multimodal processing without affecting inference, as an alternative to the slow --no-mmproj-offload option.

0 favorites 0 likes
#gpu-offloading

Hot Expert Reload on GPU is what this community needs

Reddit r/LocalLLaMA · 5d ago

A request to LLaMA maintainers to implement a feature called 'Hot Expert Reload on GPU' to improve decode speeds for MOE models with moderate active parameters, making them more usable locally with GPUs like the 3090.

0 favorites 0 likes
#gpu-offloading

@RiskRich: YOUR 8GB GPU ISN'T THE DEAD END IT USED TO BE FreeToken's paper reports Qwen3.6 35B running on an RTX 4060 laptop at 39…

X AI KOLs Timeline · 2026-08-23 Cached

FreeToken enables running large AI models like Qwen3.6 35B on limited hardware such as an 8GB GPU by offloading to CPU and RAM, providing a cost-effective solution for edge computing.

0 favorites 0 likes
#gpu-offloading

llama.cpp slower on P-Cores than on E-Cores with MoE Model and GPU+CPU offloading?

Reddit r/LocalLLaMA · 2026-07-24

Observation that llama.cpp runs slower on P-cores than E-cores when running Mixture of Experts models with GPU+CPU offloading.

0 favorites 0 likes
← Back to home

Submit Feedback