@PyTorch: vLLM (@vllm_project) is a high-throughput, memory-efficient inference and serving engine for LLMs. At #PyTorchCon North…
Summary
This post promotes vLLM, a high-throughput inference engine for LLMs, and details its sessions and activities at PyTorchCon North America.
View Cached Full Text
Cached at: 09/29/26, 09:59 PM
vLLM (@vllm_project) is a high-throughput, memory-efficient inference and serving engine for LLMs. At #PyTorchCon North America, you’ll find vLLM-related work throughout the program, from Simon Mo’s keynote to technical sessions and posters covering LLM inference and serving.
@simon_mo_ will present the keynote “vLLM Update: Scaling Open Frontier Inference Infrastructure,” covering improvements to vLLM’s core architecture and major optimizations in KV cache management and GPU kernels.
Across #PyTorchCon, technical sessions dig into vLLM-related inference, including attention, KV cache management and transfer, disaggregated serving, elastic expert parallelism, scaling across hardware without forks, and more.
“Meet the Developers of vLLM” will feature George Novack and Nick Hill. Posters cover additional vLLM-related work on custom accelerators, Trainium, expert parallelism and RDMA KV-cache transfer, multi-chip KV-cache transfer, tier-aware routing, and more.
Explore the vLLM sessions: https://events.linuxfoundation.org/pytorch-conference-north-america/program/schedule/?search=vllm…
PyTorch conferences are the open source AI community’s town square, where what’s next gets decided. Register for PyTorch Conference North America, October 20–21 in San Jose: https://hubs.la/Q04vK-VQ0
Schedule | PyTorch Conference North America
Source: https://events.linuxfoundation.org/pytorch-conference-north-america/program/schedule/?search=vllm 7:30 AM
9:00 AM
10:00 AM
7:30 AM
9:00 AM
9:10 AM
9:20 AM
9:35 AM
9:45 AM
9:50 AM
10:00 AM
10:05 AM
10:15 AM
10:25 AM
10:35 AM
10:40 AM
11:10 AM
11:45 AM
12:00 PM
12:20 PM
12:35 PM
12:45 PM
1:30 PM
1:45 PM
2:15 PM
2:30 PM
2:50 PM
3:05 PM
3:25 PM
3:40 PM
3:50 PM
3:55 PM
4:10 PM
4:20 PM
4:35 PM
4:55 PM
5:10 PM
5:30 PM
5:45 PM
6:00 PM
8:00 AM
9:00 AM
9:10 AM
9:20 AM
9:30 AM
9:40 AM
9:55 AM
10:00 AM
10:05 AM
10:15 AM
10:25 AM
10:35 AM
10:40 AM
10:55 AM
11:10 AM
11:45 AM
12:00 PM
12:20 PM
12:35 PM
12:45 PM
1:00 PM
1:30 PM
1:50 PM
2:15 PM
2:30 PM
2:50 PM
3:05 PM
3:25 PM
3:50 PM
3:55 PM
4:20 PM
4:55 PM
5:10 PM
5:30 PM
6:00 PM
Similar Articles
Inside vLLM: Anatomy of a High-Throughput LLM Inference System (2025)
A deep dive into vLLM's architecture and components for high-throughput LLM inference, covering scheduling, paged attention, continuous batching, advanced features, scaling, serving, and benchmarking.
@PyTorch: At PyTorch Conference North America, Ricardo Noriega de Soto, Tech Lead for the vLLM Omni team and Alexander Brooks, Pr…
At PyTorch Conference North America, Ricardo Noriega de Soto and Alexander Brooks will demonstrate how extending vLLM's prefix caching mechanism to multistage pipelines boosts inference speeds while reducing GPU memory overhead, providing practical strategies for optimizing complex AI workloads.
@PyTorch: Curious about how to make enterprise agentic inference production-ready with @PyTorch and @vllm_project & to learn more…
The article promotes a talk at the PyTorch Conference North America focused on making enterprise agentic inference production-ready using PyTorch and vLLM, covering ecosystem updates and registration details.
Show HN: Tiny-vLLM – high performance LLM inference engine in C++ and CUDA
Tiny-vLLM is a high-performance LLM inference engine implemented in C++ and CUDA, offering features like continuous batching and PagedAttention, and serves as an educational resource.
@PyTorch: At PyTorch Conference North America 2026, Nicolò Lucchesi, Research Engineer at @MistralAI and vLLM maintainer, will pr…
Nicolò Lucchesi will present at PyTorch Conference North America 2026 on the evolution of disaggregated serving in vLLM for hybrid models, collaborating with AWS and RedHat.