Tag
By implementing a Rust-based daemon to dynamically sleep inactive AI models, GPU memory usage in local agents can be reduced from 122 GB to 43 GB with sub-200ms wakeup latency, making it viable to run multiple models on a single workstation.
Quartermaster is an open source local AI platform that simplifies model management, VRAM allocation, and provides an OpenAI-compatible API with a built-in chat playground, supporting various backends like llama.cpp and stable-diffusion.cpp.
This paper introduces a statistically defined decision layer for continual-learning systems, allowing expert pools to decide whether to reuse existing models, spawn new ones, or defer based on accumulated evidence, with theoretical guarantees and system contributions for managing nonstationary data streams.
A developer built llmux, an open-source tool to simplify managing model profiles across different engines like vLLM and llama.cpp, allowing one-click switching with Docker and NVIDIA GPU support.
Ketch is a search tool designed for finding and managing local AI models.
llama.cpp now supports model management including downloading and lifecycle management via its API, allowing full deployment without external tools.
EchoBird is an open-source desktop client that integrates AI tool installation, local large model running, and application management, supporting one-click configuration and unified API management.
oMLX v0.4.0 ships a native Swift macOS app with redesigned onboarding, settings UI, Hugging Face cache discovery, and improved model management for running local AI on Macs.
A practical guide to organizing local LLM experiments by using a layered wrapper system and a consistent directory structure to avoid model location drift, flag amnesia, and harness coupling.