FYI llamacpp server can hot swap models now-a-days in under 30sec

Reddit r/LocalLLaMA Tools

Summary

Llamacpp server now supports hot-swapping models in under 30 seconds, a significant speed improvement over previous methods like PyTorch.

See this question at least a handful of times when browsing new and in the comments, llamacpp has one of the cleaner model hotswap apis now that just works with openwebui and hermes. Bonus: the 2nd model gemma went derp as i was recording this, but the time spent swapping has gotten stupid fast... I remember starting a load and talking a walk while pytorch did its thing just a few months back
Original Article

Similar Articles