Tag
An NVIDIA intern shares insights from research on how frontier LLMs should perform inference on heterogeneous systems, with a thread containing TLDR and mini experiments.