hardware-tuning

Tag

Cards List
#hardware-tuning

Splash on a 40-core M5 Max: +20% decode by tuning the kernels for your own chip

Reddit r/LocalLLaMA ↗ · 10h ago

The article describes how to use the 'make tune-kernels' tool in Splash to optimize AI model decoding on 40-core M5 Max chips, achieving up to 20% speed improvement by tuning kernel layouts for specific hardware.

0 favorites 0 likes
#hardware-tuning

Qwen3.8-Flash-Next, on 5090+64gb, with Llama.cpp - Seems to not use ram?

Reddit r/LocalLLaMA ↗ · yesterday

User shares performance results and setup details for running the Qwen3.8-Flash-Next model with Llama.cpp on a system featuring an NVIDIA 5090 GPU and 64GB RAM, noting unexpected low RAM usage during inference.

0 favorites 0 likes
← Back to home

Submit Feedback