monokernel

Tag

Cards List
#monokernel

Building a monokernel for LLM inference on AMD MI300X - up to 3,300 output tokens/s per request [P]

Reddit r/MachineLearning ↗ · 2026-05-29

A monokernel approach for LLM decoding on AMD MI300X GPUs achieves up to 3,300 output tokens/s per request without speculative decoding or quantization, using memory access patterns mapped to the die topology.

0 favorites 0 likes
← Back to home

Submit Feedback