Tag
This update to llama.cpp introduces a Vulkan-based implementation of int8 cooperative matrix multiplication for AMD RDNA3 and RDNA4 GPUs, enhancing inference performance on these hardware platforms.