vllm-project/vllm v0.19.1

GitHub Releases Watchlist Tools

Summary

vLLM v0.19.1 release - a fast and easy-to-use open-source library for LLM inference and serving with state-of-the-art throughput, supporting 200+ model architectures and diverse hardware including NVIDIA/AMD GPUs and CPUs.

This is a patch release on top of v0.19.0 with Transformers v5.5.4 upgrade and bug fixes for Gemma4: Update to transformers v5 (#30566) [Bugfix] Fix invalid JSON in Gemma 4 streaming tool calls by stripping partial delimiters (#38992) [Bugfix][Frontend] Fix Gemma4 streaming HTML duplication after tool calls (#38909) [Bugfix] Fix Gemma4 streaming tool call corruption for split boolean/number values (#39114) [Tool] adjust_request to reasoning parser, and Gemma4 fixes (#39027) [Gemma4] Support quantized MoE (#39045) Add Gemma4 Eagle3 support (#39450) [Gemma4][Bugfix]: Enable Gemma4ForCasualLM to load lora adapters correctly (#38844) [Bugfix] Fix Gemma4 tool parser converting bare null to string "null" (#39679) [Model] Fix Gemma 4 token repetition by dynamic BOS injection for PT models (#39842) fix(kimi_k25): resolve media_placeholder_token_id from tokenizer (#39344)
Original Article
View Cached Full Text

Cached at: 04/20/26, 08:36 AM

Easy, fast, and cheap LLM serving for everyone

| Documentation | Blog | Paper | Twitter/X | User Forum | Developer Slack |

Similar Articles

vllm-project/vllm v0.20.0

GitHub Releases Watchlist

vLLM v0.20.0 is released, an open-source library for high-throughput LLM inference and serving, featuring PagedAttention and support for various hardware architectures.

vllm-project/vllm v0.20.1

GitHub Releases Watchlist

vLLM v0.20.1 is a minor version update for the popular open-source LLM inference and serving library, maintaining its focus on high-throughput and efficient memory management.

vllm-project/vllm v0.21.0rc1

GitHub Releases Watchlist

vLLM v0.21.0rc1 is a pre-release update for the high-performance LLM inference and serving library, featuring optimizations for throughput, quantization, and hardware support.

vllm-project/vllm v0.20.0rc1

GitHub Releases Watchlist

vLLM 0.20.0rc1 releases with major throughput, quantization, speculative decoding, and multi-hardware support enhancements for scalable LLM serving.