Weird to get near linear scaling by adding another GPU?
Summary
A user reports near-linear performance scaling when adding a second RTX 3090 for inference with a Qwen model, achieving roughly 1.8x decode TPS improvement without NVLink.
Similar Articles
When Is NVLink Worth It?
Tests NVLink on dual RTX 3090s for AI inference and training, finding significant speedups for tensor parallel prompt processing (30%) and FSDP training (3x), but minimal effect on token generation or layer split inference.
Benchmark Qwen 3.6 27B MTP on 2x3090 NVLINK
A benchmark analysis of Qwen 3.6 27B MTP on 4x RTX 3090 GPUs, demonstrating that using NVLink for tensor parallelism yields significant throughput improvements (up to +53%) over PCIe configurations.
I have just moved from MacBook M5 pro 48 GB to RTX3090
A software developer shares their experience switching from a MacBook to an RTX3090 Linux setup for running AI models, achieving significantly higher inference speeds with Qwen 3.8 27B and potentially replacing their Claude subscription.
@superalesha: https://x.com/superalesha/status/2077437741915312221
The author shares six months of measurements on a four-RTX 3090 local setup, revealing that data parallelism often outperforms tensor parallelism for models fitting on fewer cards, with up to 3.4x throughput difference.
2×RTX 3090 + EPYC box running qwen3.8-flash-next at ~38 tok/s
The user describes a hardware configuration with dual RTX 3090 GPUs and an AMD EPYC CPU for running the Qwen3-Flash-Next model using llama.cpp, achieving 38 tokens per second in single-stream inference, and seeks advice on whether to add a third GPU or upgrade the CPU to improve performance, especially for running multiple parallel agents.