Tag
This post promotes vLLM, a high-throughput inference engine for LLMs, and details its sessions and activities at PyTorchCon North America.
The article introduces 1Cat-vLLM, a fork of vLLM optimized for NVIDIA V100 GPUs, and compares its performance with llama.cpp for serving large language models like Qwen3.6-35b.
The article explains RadixAttention, a technique in SGLang that uses a radix tree to efficiently cache and reuse overlapping KV caches across branching requests in AI serving.
Nicolò Lucchesi will present at PyTorch Conference North America 2026 on the evolution of disaggregated serving in vLLM for hybrid models, collaborating with AWS and RedHat.
A new vLLM fork introduces comprehensive INT8 optimization for Qwen3.8 27B on older AMD MI100 GPUs, achieving up to 972 tokens per second throughput with rigorous accuracy validation.
vLLM introduces a native sharded weight transfer engine using Ray Direct Transport (RDT) for efficient weight syncing in large-scale online RL setups, enhancing performance for models like Kimi K2.
The article presents a methodology for optimizing the serving of DeepSeek-V4-Pro, a 1.6-trillion-parameter MoE model, on H20 GPUs, achieving significant performance improvements through scenario-specific configurations and optimizations.