Tag
LightSeek's TokenSpeed provides Day-0 optimized FP4 inference support for Thinking Machines Lab's Inkling model on NVIDIA and AMD accelerators, built on PyTorch and collaborating with vLLM.
Qwen inference team announced TokenSpeed, a high-performance LLM inference engine for agentic workloads, achieving 540 TPS, with open-source preview available.