Tag
The article presents a custom fork of llama.cpp optimized for NVIDIA Ampere architecture GPUs like the 3090, achieving over 90 tokens per second for 27B parameter models with large context windows, making local inference faster than API speeds.