Tag
A detailed strategy using VRAM as disk cache with unified memory and llama.cpp settings achieves 340 pp/s and 9.6 tg/s for Kimi K2.7 on a single DGX Spark.