local-llm-serving

Tag

Cards List
#local-llm-serving

@malikwas1f: 308 tok/s aggregate on 2× RTX 3090 — no NVLink Qwen3.8-27B on vLLM, 32 concurrent streams. Enough to feed a whole swarm…

X AI KOLs Timeline · 2026-08-23 Cached

A performance report and guide for running the Qwen3.8-27B model on dual RTX 3090 GPUs using vLLM, achieving 308 tokens per second with 32 concurrent streams, with a GitHub repository for local deployment configurations.

0 favorites 0 likes
← Back to home

Submit Feedback