@tenstorrent: Thank you Tokyo! Here’s everything we announced at TT-Deploy Japan: Faster AI Inference • Kimi K2.6 900 t/s/u, 3x faste…
Summary
Tenstorrent announced at TT-Deploy Japan faster AI inference for Kimi K2.6, LTX 2.3, and DeepSeek-R1 on their hardware, plus the licensable TT-Ascalon S RISC-V CPU for agentic AI.
View Cached Full Text
Cached at: 07/03/26, 06:31 AM
Thank you Tokyo! Here’s everything we announced at TT-Deploy Japan:
Faster AI Inference • Kimi K2.6 900 t/s/u, 3x faster than GPUs • LTX 2.3 Fast 6 sec video gen in ~6 sec, 144 frames, 1080p, 4x faster than GPUs • DeepSeek-R1-0528 671B 400+ t/s/u
TT-Ascalon S Available Today • A licensable RISC-V CPU built for the next generation of agentic AI applications
Heterogenous or Stand Alone • Easily deploy Tenstorrent Galaxy alongside existing infrastructure or standalone • @aiand_’s sovereign heterogenous inference platform with Tenstorrent Galaxy™ superclusters
Similar Articles
@HotAisle: Kimi K2.6 + DFlash: 508 tok/s on 8x MI300X 5.6x throughput improvement over baseline autoregressive serving 90 tok/s → …
Kimi K2.6 paired with DFlash inference system achieves 508 tokens/s on 8×AMD MI300X, a 5.6× throughput jump from 90 tokens/s baseline with zero quality loss.
@JiaZhihao: Excited to share Lithos’ serving stack for Kimi K2.7 Code, a 1T-parameter frontier coding model. On a single 8×B200 nod…
Lithos announces its inference engine serving Kimi K2.7 Code, achieving over 1,000 tokens/sec per user on a single 8×B200 node at native precision, 3.4–5.7× faster than major providers.
@gnotuy: We open sourced Kimi K2.6. The next frontier in test-time compute isn't bigger models. It's better organizations of int…
Moonshot AI has open sourced Kimi K2.6 and argues that the next frontier in test-time compute is better organization of intelligence rather than simply building bigger models.
@no_stp_on_snek: Very nice. Huge for team 3090. And TurboQuant+ is already implemented in a bunch of inference engines.
A reply celebrates Unsloth AI's upcoming Qwen3.8-27B model, which will run on 17GB RAM/VRAM setups, and notes TurboQuant+ is already integrated into many inference engines — great news for RTX 3090 users.
@YRSM_Simon: This is big news! Kimi 2.6 is a generative-level model. In this age of overflowing LLM capabilities, speed will become the deciding factor in competition. Is the chip sector about to see another 'sector rotation'? 😅
Cerebras is now running Kimi K2.6, a trillion-parameter model, in enterprise trials at ~1,000 tokens/s, the fastest frontier model performance ever measured by Artificial Analysis.