@HotAisle: This is awesome. I wonder who's MI300x they used... ;-)
Summary
Kog announces real-time LLM inference achieving 3000+ output tokens per second per request on standard datacenter GPUs, bringing high-speed inference previously limited to custom silicon to production hardware.
View Cached Full Text
Cached at: 05/31/26, 10:45 AM
This is awesome.
I wonder who’s MI300x they used… ;-)
Kog (@Kog__AI): 🚀 Launch today: Kog generates 3,000+ output tokens/s per single request, on standard datacenter GPUs.
We are bringing real-time LLM inference to hardware that companies already run in production. The speed previously associated with purpose-built silicon is now delivered on
Similar Articles
@rohanpaul_ai: I had to test it myself to believe this unreal inference speed. 3,000 tokens/s for 1 user on standard datacenter GPUs. …
Kog AI achieves 3,000 tokens/s inference speed on 8× AMD MI300X GPUs and 2,100 on 8× NVIDIA H200, leveraging a hidden efficiency gap in GPU token generation.
Real-time LLM Inference on Standard GPUs: 3k tokens/s per request
Kog AI launches a tech preview of the Kog Inference Engine, achieving 3,000 tokens/s per request on standard datacenter GPUs by co-designing model architecture, runtime, and low-level GPU code, targeting latency-critical AI agent workflows.
Hitting a billion tokens per minute on one GPU by combining a query planner and an inference engine (15 minute read)
Modal and Carnegie Mellon University's Full Stack Data Lab developed Quail, an inference engine that integrates query planning with inference to achieve over a billion tokens processed per minute on a single H100 GPU, outperforming vLLM by 10x in specific scenarios.
@victormustar: 300tok/s on mobile is insane... open source must win
Open source AI inference reaches 300 tok/s on mobile, with a WebGPU framework pushing Liquid AI's LFM2.5 230M to 1,400 tok/s in browser.
Building a monokernel for LLM inference on AMD MI300X - up to 3,300 output tokens/s per request [P]
A monokernel approach for LLM decoding on AMD MI300X GPUs achieves up to 3,300 output tokens/s per request without speculative decoding or quantization, using memory access patterns mapped to the die topology.