disaggregated-inference

Tag

Cards List
#disaggregated-inference

Topology-Aware Data Movement for Disaggregated GPU Inference

arXiv cs.LG ↗ · 2026-08-03 Cached

This paper presents a topology-aware data movement orchestrator for disaggregated LLM inference, which dynamically selects optimal transport based on interconnect hierarchy and overlaps KV cache transfer with computation, achieving 3-18x transfer latency reduction over uniform RDMA.

0 favorites 0 likes
#disaggregated-inference

AMD and Cerebras Launch AI Inference Solution (10 minute read)

TLDR AI ↗ · 2026-07-24 Cached

AMD and Cerebras announced a joint AI inference solution combining AMD Helios rackscale solutions with Cerebras Wafer-Scale Engine, aiming for ultra-low latency and high throughput. The disaggregated inference workflow is expected to deliver up to 5x higher tokens per second per watt.

0 favorites 0 likes
#disaggregated-inference

The Price of Anarchy in Disaggregated Inference

Hugging Face Daily Papers ↗ · 2026-06-11 Cached

This paper presents a game-theoretic analysis of disaggregated inference architectures that separate prefill and decode phases across GPU pools, characterizing how GPU saturation affects performance. The authors propose an adaptive controller that detects saturation transitions and adjusts routing parameters, reducing the Price of Anarchy significantly in experiments on NVIDIA B200 clusters.

0 favorites 0 likes
#disaggregated-inference

Semantic Cache Distillation: Efficient State Transfer via Reuse and Selective Patching

arXiv cs.LG ↗ · 2026-06-09 Cached

This paper proposes Semantic Cache Distillation (SCD), a loss-constrained framework that replaces raw KV cache transmission with compact semantic codes, achieving up to 2.65x TTFT speedup while keeping generation quality within 5% F1 of the oracle.

0 favorites 0 likes
← Back to home

Submit Feedback