Tag
SGLang has added deployment recipes for DeepSeek-V4-Flash-Vision and DeepSeek-V4-Flash-0731 models on 2x DGX Spark hardware, with support for various configurations and optimizations.
SGLang has updated its deployment recipes for the Qwen3.8-27B model on RTX 5090 and RTX Pro 6000 hardware, adding variants for different configurations with tuning options.
A GitHub repository provides practical recipes to deploy and use Meta's Muse Glimmer, an open-weight 30B AI model designed for local agentic tasks on a single GPU.
vLLM launched a redesigned recipes site that turns any HuggingFace model URL into a ready-to-run inference recipe for specific hardware and tasks.