@PyTorch: LightSeek (@lightseekorg) TokenSpeed brings Day-0 optimized FP4 inference support for Thinking Machines Lab’s Inkling a…
Summary
LightSeek's TokenSpeed provides Day-0 optimized FP4 inference support for Thinking Machines Lab's Inkling model on NVIDIA and AMD accelerators, built on PyTorch and collaborating with vLLM.
View Cached Full Text
Cached at: 07/16/26, 02:15 PM
LightSeek (@lightseekorg) TokenSpeed brings Day-0 optimized FP4 inference support for Thinking Machines Lab’s Inkling across @NVIDIA and @AMD accelerators. TokenSpeed is built on PyTorch-based infrastructure and collaborates with @vllm_project to advance high-performance, reusable inference kernels for the open source AI ecosystem.
LightSeek Foundation (@lightseekorg): Thinking Machines Lab released Inkling today. We’re excited to partner with Thinking Machines Lab to bring Day-0 support for Inkling to TokenSpeed. ✅NVIDIA (G)B200/(G)B300 (NVFP4) ✅AMD MI350X/MI355X (MXFP4) ✅Native FP4 serving ✅Unified kernel architecture ✅MTP + optimized
Similar Articles
TokenSpeed: A Speed-of-Light LLM Inference Engine for Agentic Workloads (5 minute read)
Lightseek releases TokenSpeed, a high-performance LLM inference engine optimized for agentic workloads, featuring compiler-backed parallelism and advanced kernel optimizations that have been adopted by vLLM.
@PyTorch: One runtime, multiple GPU architectures, and zero vendor-specific model code. In this blog post, the TokenSpeed team @l…
TokenSpeed-Kernel is a portable, high-performance kernel system for LLM inference that enables zero vendor-specific model code and supports multiple GPU architectures, achieving up to 3.6x higher throughput on AMD MI355X.
@PyTorch: While SGLang provided Day-0 support for DeepSeek-V4, the collaboration between the @lmsysorg and @NVIDIAAI engineering …
SGLang provided Day-0 support for DeepSeek-V4, and collaboration between LMSys and NVIDIA engineering teams achieved up to 5x throughput increase in production, with improvements shown on the SemiAnalysis InferenceX dashboard.
@scaling01: DeepSeek just made their inference ~5x cheaper at 50 TPS
DeepSeek has reduced inference costs by approximately 5x while maintaining 50 tokens per second throughput.
@SuJinYan123: Just 6 hours after DeepSeek open-sourced the Qwen DSpark weights, OpenInfer already has DSpark support running on RTX 5…
OpenInfer, a pure Rust+CUDA LLM inference engine, quickly added support for DeepSeek's DSpark speculative decoding technique on RTX 5090, achieving nearly 500 tok/s per user and scaling to ~2.4K aggregate tok/s, outperforming DFlash on non-random workloads.