ai-serving

Tag

Cards List
#ai-serving

@PyTorch: vLLM (@vllm_project) is a high-throughput, memory-efficient inference and serving engine for LLMs. At #PyTorchCon North…

X AI KOLs Timeline ↗ · 5d ago Cached

This post promotes vLLM, a high-throughput inference engine for LLMs, and details its sessions and activities at PyTorchCon North America.

0 favorites 0 likes
#ai-serving

Make Volta Fast Again

Reddit r/LocalLLaMA ↗ · 2026-09-25

The article introduces 1Cat-vLLM, a fork of vLLM optimized for NVIDIA V100 GPUs, and compares its performance with llama.cpp for serving large language models like Qwen3.6-35b.

0 favorites 0 likes
#ai-serving

@akshay_pachaar: RadixAttention, clearly explained. (how SGLang makes prefix caching highly efficient) Prefix caching sounds simple unti…

X AI KOLs Timeline ↗ · 2026-09-12 Cached

The article explains RadixAttention, a technique in SGLang that uses a radix tree to efficiently cache and reuse overlapping KV caches across branching requests in AI serving.

0 favorites 0 likes
#ai-serving

@PyTorch: At PyTorch Conference North America 2026, Nicolò Lucchesi, Research Engineer at @MistralAI and vLLM maintainer, will pr…

X AI KOLs Timeline ↗ · 2026-09-11 Cached

Nicolò Lucchesi will present at PyTorch Conference North America 2026 on the evolution of disaggregated serving in vLLM for hybrid models, collaborating with AWS and RedHat.

0 favorites 0 likes
#ai-serving

Qwen3.8 27B C8 at 972 TG / 5,680 PP on 4x MI100 rig ($6.5k) using my new INT8 vLLM fork

Reddit r/LocalLLaMA ↗ · 2026-08-26

A new vLLM fork introduces comprehensive INT8 optimization for Qwen3.8 27B on older AMD MI100 GPUs, achieving up to 972 tokens per second throughput with rigorous accuracy validation.

0 favorites 0 likes
#ai-serving

@vllm_project: This weight transfer engine is native to vLLM. Any Ray-based trainer can adopt it with a single WeightSource iterator. …

X AI KOLs Following ↗ · 2026-08-26 Cached

vLLM introduces a native sharded weight transfer engine using Ray Direct Transport (RDT) for efficient weight syncing in large-scale online RL setups, enhancing performance for models like Kimi K2.

0 favorites 0 likes
#ai-serving

Pushing the Limits of Serving DeepSeek-V4-Pro (28 minute read)

TLDR AI ↗ · 2026-08-20 Cached

The article presents a methodology for optimizing the serving of DeepSeek-V4-Pro, a 1.6-trillion-parameter MoE model, on H20 GPUs, achieving significant performance improvements through scenario-specific configurations and optimizations.

0 favorites 0 likes
← Back to home

Submit Feedback