Anyone customizing and Optimizing llama.cpp per model?

Reddit r/LocalLLaMA Tools

Summary

The author proposes optimizing llama.cpp for specific models by stripping unnecessary code and using AI to automate the process, aiming for 2x+ performance improvements.

Basically the idea is take your favorite model, for example qwen3.8-27b or say dsv4vision. Strip everything out that is not needed by that model so the only thing needed is just for the model. Optimize the remaining code to be fast. The idea is to have a model also do this, provide it with enough tools, prompts, docs, guidance. I reckon that if we have a llama.cpp that is optimized for just one model architecture without all the cruft needed to run and. handle other models, that it would not be surprising to easily see 2x+ performance improvement. Anyone thinking along this idea? Again, the goal will be to give this task to a smart model, let it run in a loop, and after a week or 2 you hopefully end up with llama.qwen3.8-27b or llama.glm5.3-flash that would fly.
Original Article

Similar Articles

Llama cpp metal moe optimization

Reddit r/LocalLLaMA

A developer shared a Metal optimization for llama.cpp that improves decode speed for IQ3_XXS models on Apple Silicon, with a GitHub pull request and a request for community testing.