Weird to get near linear scaling by adding another GPU?

Reddit r/LocalLLaMA News

Summary

A user reports near-linear performance scaling when adding a second RTX 3090 for inference with a Qwen model, achieving roughly 1.8x decode TPS improvement without NVLink.

Single steam benchmarks (club-3090) model: qwen3.6-27b-autoround-int4 **BEFORE:** 1x3090 \*Their default script recipe for single 3090'\*s *(4-bit quant and 4-bit kv cache, mtp=2)* NARRATIVE decode\_TPS: mean = **53** std = **0.6** CODE decode\_TPS: mean = **62** std= **1.4** **AFTER:** 2x3090 *Their default script recipe for dual 3090's (4-bit quant and 8-bit kv cache, mpt=3)* NARRATIVE decode\_TPS: mean= **94** std= **1.3** CODE decode\_TPS: mean= **120** std= **2.1** This is running *without NVLink,* on a 8x/8x motherboard, for some reason P2P was automatically enabled (no driver hack needed), Tensor parallelism = 2 I am truly shocked that I got almost linear scaling in performance. I still get odd parsing errors in my quality tests when editing large code files in Agent mode (VSCode), (but not the same ones as before), for some reason forcing the model to use CLI editing tools is much more reliable than whatever VSCode is doing with the Agent. I am going to likely move to their 8-bit weight model recipe as well.
Original Article

Similar Articles

When Is NVLink Worth It?

Hacker News Top

Tests NVLink on dual RTX 3090s for AI inference and training, finding significant speedups for tensor parallel prompt processing (30%) and FSDP training (3x), but minimal effect on token generation or layer split inference.

Benchmark Qwen 3.6 27B MTP on 2x3090 NVLINK

Reddit r/LocalLLaMA

A benchmark analysis of Qwen 3.6 27B MTP on 4x RTX 3090 GPUs, demonstrating that using NVLink for tensor parallelism yields significant throughput improvements (up to +53%) over PCIe configurations.

I have just moved from MacBook M5 pro 48 GB to RTX3090

Reddit r/LocalLLaMA

A software developer shares their experience switching from a MacBook to an RTX3090 Linux setup for running AI models, achieving significantly higher inference speeds with Qwen 3.8 27B and potentially replacing their Claude subscription.

2×RTX 3090 + EPYC box running qwen3.8-flash-next at ~38 tok/s

Reddit r/LocalLLaMA

The user describes a hardware configuration with dual RTX 3090 GPUs and an AMD EPYC CPU for running the Qwen3-Flash-Next model using llama.cpp, achieving 38 tokens per second in single-stream inference, and seeks advice on whether to add a third GPU or upgrade the CPU to improve performance, especially for running multiple parallel agents.