prefill-decode-disaggregation

Tag

Cards List
#prefill-decode-disaggregation

ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving

Hugging Face Daily Papers · 2026-07-01 Cached

ELDR is an expert-locality-aware decode router for prefill-decode disaggregated Mixture-of-Experts serving that predicts expert activations from prefill signatures and routes requests to minimize latency, implemented in vLLM and achieving 5.9-13.9% reduction in median TPOT on up to 40 GPUs.

0 favorites 0 likes
#prefill-decode-disaggregation

@kazukifujii: The UC San Diego Hao AI Lab blog provides a very clear explanation of the usefulness of DistServe's Prefill Decode Disa…

X AI KOLs Timeline · 2026-06-27 Cached

The UC San Diego Hao AI Lab blog provides a clear explanation of DistServe's Prefill Decode Disaggregation, tracing its acceptance from 2024 to 2025 and linking to related technologies like NVIDIA Dynamo, llm-d, Ray Serve LLM, LMCache, and MoonCake, making it a great starting point for learning LLM inference.

0 favorites 0 likes
#prefill-decode-disaggregation

@robertnishihara: Some intuition about PD disaggregation from the blog - PD doesn't speed up prefill and can actually hurt TTFT - PD's re…

X AI KOLs Following · 2026-06-17 Cached

This blog post from Anyscale explains the intuition behind Prefill-Decode (PD) disaggregation for LLM serving, showing how separating prefill and decode phases onto dedicated GPUs can achieve up to 2.7x better goodput and 67% cost savings when using Ray and vLLM on AMD MI325X, while also discussing when PD disaggregation does not help.

0 favorites 0 likes
#prefill-decode-disaggregation

@Zai_org: https://x.com/Zai_org/status/2057216685040443743

X AI KOLs Timeline · 2026-05-20 Cached

This paper presents ZCube, a novel network architecture developed by Z.ai, Harnets.AI, and Tsinghua University to address topology-induced congestion in Prefill-Decode disaggregated LLM inference clusters. Production deployments on GLM-5.1 coding workloads achieved a 33% reduction in network CapEx, 15% throughput improvement, and 40.6% reduction in TTFT P99 latency.

0 favorites 0 likes
← Back to home

Submit Feedback