@FGuzmanAI: 56,000+ tokens/sec at just 80 MHz. I burned a full Transformer with KV cache into a custom chip. Designed gate by gate …

X AI KOLs Timeline Tools

Summary

A custom digital chip designed gate-by-gate achieves over 56,000 tokens/sec running a Transformer with KV cache at just 80 MHz, prototyped on an FPGA.

56,000+ tokens/sec at just 80 MHz. 🤯 I burned a full Transformer with KV cache into a custom chip. Designed gate by gate as a 100% digital integrated circuit. Prototyped on a FPGA. (No GPU. No CPU) Just pure digital silicon running @karpathy microGPT, spelling out names on a https://t.co/hXX8kKIxiA
Original Article
View Cached Full Text

Cached at: 06/13/26, 08:16 PM

56,000+ tokens/sec at just 80 MHz. 🤯

I burned a full Transformer with KV cache into a custom chip. Designed gate by gate as a 100% digital integrated circuit. Prototyped on a FPGA. (No GPU. No CPU) Just pure digital silicon running @karpathy microGPT, spelling out names on a https://t.co/hXX8kKIxiA

Similar Articles

Can Agents Design Better Chips with a Higher Level Abstraction?

arXiv cs.AI

This paper explores using LLM agents for chip design with higher-level abstractions, introducing a workflow called AHRR that combines agent-based HLS design with RTL refinement, achieving a 2.6× speedup over direct RTL design in benchmarks.