Tag
The article highlights the Strata inference engine, which significantly outperforms llama.cpp for running Qwen3.8 models on a laptop with 12GB VRAM and 64GB RAM, achieving up to 50 tokens per second for text generation and 1500 tokens per second for prompt processing.
Ferrox is a pure-Rust inference engine that loads GGUF models and runs local LLMs on CPU, Metal, or CUDA, with a CLI and an OpenAI-compatible server. It aims to match llama.cpp's performance while being written from scratch with no bindings.