Tag
vLLM introduces Semantic Router, a serving-layer primitive that enables collaboration between multiple models through micro-agents, allowing the router to improve output quality without modifying model weights.