Tag
By implementing a Rust-based daemon to dynamically sleep inactive AI models, GPU memory usage in local agents can be reduced from 122 GB to 43 GB with sub-200ms wakeup latency, making it viable to run multiple models on a single workstation.