@PyTorch: vLLM (@vllm_project) is a high-throughput, memory-efficient inference and serving engine for LLMs. At #PyTorchCon North…

X AI KOLs Timeline Events

Summary

This post promotes vLLM, a high-throughput inference engine for LLMs, and details its sessions and activities at PyTorchCon North America.

vLLM (@vllm_project) is a high-throughput, memory-efficient inference and serving engine for LLMs. At #PyTorchCon North America, you’ll find vLLM-related work throughout the program, from Simon Mo’s keynote to technical sessions and posters covering LLM inference and serving. @simon_mo_ will present the keynote “vLLM Update: Scaling Open Frontier Inference Infrastructure,” covering improvements to vLLM’s core architecture and major optimizations in KV cache management and GPU kernels. Across #PyTorchCon, technical sessions dig into vLLM-related inference, including attention, KV cache management and transfer, disaggregated serving, elastic expert parallelism, scaling across hardware without forks, and more. "Meet the Developers of vLLM" will feature George Novack and Nick Hill. Posters cover additional vLLM-related work on custom accelerators, Trainium, expert parallelism and RDMA KV-cache transfer, multi-chip KV-cache transfer, tier-aware routing, and more. Explore the vLLM sessions: https://events.linuxfoundation.org/pytorch-conference-north-america/program/schedule/?search=vllm… PyTorch conferences are the open source AI community’s town square, where what’s next gets decided. Register for PyTorch Conference North America, October 20–21 in San Jose: https://hubs.la/Q04vK-VQ0
Original Article
View Cached Full Text

Cached at: 09/29/26, 09:59 PM

vLLM (@vllm_project) is a high-throughput, memory-efficient inference and serving engine for LLMs. At #PyTorchCon North America, you’ll find vLLM-related work throughout the program, from Simon Mo’s keynote to technical sessions and posters covering LLM inference and serving.

@simon_mo_ will present the keynote “vLLM Update: Scaling Open Frontier Inference Infrastructure,” covering improvements to vLLM’s core architecture and major optimizations in KV cache management and GPU kernels.

Across #PyTorchCon, technical sessions dig into vLLM-related inference, including attention, KV cache management and transfer, disaggregated serving, elastic expert parallelism, scaling across hardware without forks, and more.

“Meet the Developers of vLLM” will feature George Novack and Nick Hill. Posters cover additional vLLM-related work on custom accelerators, Trainium, expert parallelism and RDMA KV-cache transfer, multi-chip KV-cache transfer, tier-aware routing, and more.

Explore the vLLM sessions: https://events.linuxfoundation.org/pytorch-conference-north-america/program/schedule/?search=vllm…

PyTorch conferences are the open source AI community’s town square, where what’s next gets decided. Register for PyTorch Conference North America, October 20–21 in San Jose: https://hubs.la/Q04vK-VQ0


Schedule | PyTorch Conference North America

Source: https://events.linuxfoundation.org/pytorch-conference-north-america/program/schedule/?search=vllm 7:30 AM

9:00 AM

10:00 AM

7:30 AM

9:00 AM

9:10 AM

9:20 AM

9:35 AM

9:45 AM

9:50 AM

10:00 AM

10:05 AM

10:15 AM

10:25 AM

10:35 AM

10:40 AM

11:10 AM

11:45 AM

12:00 PM

12:20 PM

12:35 PM

12:45 PM

1:30 PM

1:45 PM

2:15 PM

2:30 PM

2:50 PM

3:05 PM

3:25 PM

3:40 PM

3:50 PM

3:55 PM

4:10 PM

4:20 PM

4:35 PM

4:55 PM

5:10 PM

5:30 PM

5:45 PM

6:00 PM

8:00 AM

9:00 AM

9:10 AM

9:20 AM

9:30 AM

9:40 AM

9:55 AM

10:00 AM

10:05 AM

10:15 AM

10:25 AM

10:35 AM

10:40 AM

10:55 AM

11:10 AM

11:45 AM

12:00 PM

12:20 PM

12:35 PM

12:45 PM

1:00 PM

1:30 PM

1:50 PM

2:15 PM

2:30 PM

2:50 PM

3:05 PM

3:25 PM

3:50 PM

Meet the Developers of Helion Community Expo (Concourse Level)Jason ANSEL, Jongsok CHOI, Yarong Mu, Freya Azad, Theotime Combes, Ethan Che, Dunfan Lu, Shangdi Yu, Oguz Ulgen, Emilio Cota, Karthick Panner Selvam

3:55 PM

4:20 PM

4:55 PM

5:10 PM

5:30 PM

6:00 PM

Similar Articles