prefill-decode

Tag

Cards List
#prefill-decode

ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration

arXiv cs.LG · 2026-08-27 Cached

ExFold is a unified training-free framework that accelerates MoE model inference by folding excluded expert contributions into retained experts, achieving up to 1.41× speedup while maintaining high quality.

0 favorites 0 likes
#prefill-decode

@CyrusHakha: One pattern we keep seeing with customers serving LLMs at scale: Prefill-decode disaggregation is often treated like a …

X AI KOLs Following · 2026-06-15 Cached

Discusses the nuanced reality of prefill-decode disaggregation in LLM serving at scale, based on customer patterns and validated on AMD with vLLM.

0 favorites 0 likes
#prefill-decode

@MaxForAI: http://Z.ai and this ZCube paper from Tsinghua—worth a read for anyone in Infra. Many people's first reaction when talking about AI infra is still GPU, memory, quantization, and inference frameworks. But once you get into long context and Prefill-Decode separation, the network is no longer just a 'supporting role' in the data center. Every...

X AI KOLs Timeline · 2026-05-21

ZCube is a new network architecture that flattens the topology and mixes single/multi-rail access to optimize KV Cache transmission in long-context and PD separation scenarios. In the GLM-5.1 production cluster, it achieved a 33% reduction in switch/optical module costs, a 15% increase in GPU inference throughput, and a 40.6% decrease in TTFT P99.

0 favorites 0 likes
#prefill-decode

Prefill-as-a-Service:KVCache of Next-Generation Models Could Go Cross-Datacenter

Hacker News Top · 2026-04-19 Cached

Researchers propose Prefill-as-a-Service (PrfaaS), a system that offloads long-context prefill to remote compute-dense clusters and streams KVCache over commodity Ethernet, enabling independent scaling and 32-54% higher throughput for a 1T-parameter hybrid model.

0 favorites 0 likes
← Back to home

Submit Feedback