ZeroGPU
Summary
ZeroGPU is a compute efficient layer designed for AI inference, aiming to optimize GPU usage and reduce costs.
Similar Articles
General Compute
General Compute is a product offering an inference cloud optimized for speed to run AI models.
@MaxForAI: http://Z.ai and this ZCube paper from Tsinghua—worth a read for anyone in Infra. Many people's first reaction when talking about AI infra is still GPU, memory, quantization, and inference frameworks. But once you get into long context and Prefill-Decode separation, the network is no longer just a 'supporting role' in the data center. Every...
ZCube is a new network architecture that flattens the topology and mixes single/multi-rail access to optimize KV Cache transmission in long-context and PD separation scenarios. In the GLM-5.1 production cluster, it achieved a 33% reduction in switch/optical module costs, a 15% increase in GPU inference throughput, and a 40.6% decrease in TTFT P99.
GPU Management: Why Idle GPUs Are the New Grounded Aircraft
The article argues that GPU utilization is becoming the key constraint in enterprise AI, analogous to aircraft utilization in aviation, and that idle GPUs represent wasted capacity that determines competitive advantage.
Why the first GPU financiers are turning to inference chips in a $400 million deal
General Compute secured a $400M loan from Upper90 using inference-specific SambaNova chips as collateral, signaling a shift in AI infrastructure financing toward cheaper, more efficient inference hardware amid growing demand for open-source models.
Popping the GPU Bubble
Moondream's Photon inference engine eliminates GPU bubbles through pipelined decoding, achieving near-realtime VLM inference with up to 35% higher decode throughput.