@rohanpaul_ai: Quite a massive inferencing rack breakthrough from @TensordyneInc . They just announced an AI-inference rack, claiming …
Summary
Tensordyne announces the Napier AI inference rack, claiming 13x the throughput of Nvidia's NVL72 GB300 by using log-space math to reduce energy and transistor usage, potentially disrupting the inference hardware landscape.
View Cached Full Text
Cached at: 06/17/26, 07:59 PM
Quite a massive inferencing rack breakthrough from @TensordyneInc .
They just announced an AI-inference rack, claiming 13x the rack throughput of NVIDIA’s NVL72 GB300 in a DeepSeek-R1 comparison based on internal simulations.
What makes this a big deal is that Tensordyne is attacking inference at the math level.
AI chips spend huge amounts of energy moving and multiplying numbers.
Napier (its AI inference racks) works in log space, where multiplication becomes addition, and addition is cheaper to build, switch, cool, and repeat billions of times per token.
So instead of spending tons of transistor budget on heavy multiply circuits, Napier tries to shrink the math itself.
So that means less chip area for compute and more for SRAM, resulting in less power per token and way more inference packed into the same rack.
If they have made log math accurate and fast enough for real inference, then Napier is not just pushing more power into a rack, it is changing the cost of the basic operation behind model serving.
AI inference is no longer just a FLOPS race. It is a rack-level fight over power, memory locality, interconnect latency, and how many paying tokens can be served before the economics break.
They reported their TDN Rack reaches 363,000 tokens per second on DeepSeek-R1 at user speeds of 210 tokens per second per internal simulation, compared with 27,400 tokens per second for Nvidia’s NVL72 GB300.
Similar Articles
@TensordyneInc: https://x.com/TensordyneInc/status/2066567307984531834
Tensordyne introduces Napier, an inference system using logarithmic math on silicon, claiming massive efficiency gains for MoE and reasoning models, with air-cooled racks.
Tensordyne announces Logarithmic AI compute chips. 17x more tokens per watt and 13x higher throughput than NVIDIA Blackwell.
Tensordyne announced a breakthrough inference system using logarithmic math in hardware, claiming 17x more tokens per watt and 13x higher throughput than NVIDIA Blackwell, achieved by replacing complex multiplication with simple addition in log space.
@LambLabs: Lamb Labs (YC S26) is building custom AI inference chips that can do 20,000+ tok/s at 63x higher Intelligence per Watt …
Lamb Labs (YC S26) announces custom AI inference chips claiming 20,000+ tok/s and 63x higher Intelligence per Watt than traditional GPUs, arguing GPUs were adopted for availability rather than suitability.
Progress (3 minute read)
Etched announces breakthroughs in low voltage inference and cluster scale memory for their AI inference hardware, with first racks shipping this summer and $1B in customer contracts.
AMD and Cerebras Launch AI Inference Solution (10 minute read)
AMD and Cerebras announced a joint AI inference solution combining AMD Helios rackscale solutions with Cerebras Wafer-Scale Engine, aiming for ultra-low latency and high throughput. The disaggregated inference workflow is expected to deliver up to 5x higher tokens per second per watt.