Tag
The article describes how to use the 'make tune-kernels' tool in Splash to optimize AI model decoding on 40-core M5 Max chips, achieving up to 20% speed improvement by tuning kernel layouts for specific hardware.
User shares performance results and setup details for running the Qwen3.8-Flash-Next model with Llama.cpp on a system featuring an NVIDIA 5090 GPU and 64GB RAM, noting unexpected low RAM usage during inference.