Tag
The author designed a CPU-native LLM architecture that achieves over 100 tokens per second on a 10B parameter model, but highlights quality issues; they share a GitHub repo with code and research on adapting existing models.