@Snixtp: More efficiency tests on a single 3090 TL;DR: - I tested 8 local LLMs on a single RTX 3090, power limit from 100W to 45…

X AI KOLs Following News

Summary

The article presents benchmark results for 8 local LLMs on an RTX 3090, showing that power efficiency peaks around 225W, with diminishing returns at maximum power.

More efficiency tests on a single 3090 TL;DR: - I tested 8 local LLMs on a single RTX 3090, power limit from 100W to 450W. - Across the full set, the best efficiency was consistently around 225-250W. - At 225W, the card averaged 90.4 tok/s at 0.4167 tok/s/W. - At 450W, it reached 107.1 tok/s, but efficiency dropped to 0.2731 tok/s/W. - So max power added only ~16.7 tok/s, while using about 184W more.
Original Article

Similar Articles

Finding the 4x 3090 Sweet Spot

Reddit r/LocalLLaMA

A user shares power limit testing on a 4x RTX 3090 setup running Qwen3.6-27B with vLLM, finding 220W as the sweet spot for peak efficiency with minimal throughput loss.

Scrambling to max StrixHalo (+NVLink dual eGPU 3090 mod)

Reddit r/LocalLLaMA

A user details their modding and benchmarking of an AMD Strix Halo system with dual RTX 3090 eGPUs and NVLink, finding improvements in LLM inference speed for dense models, especially with vLLM, and discusses power efficiency trade-offs.

[Benchmark] 5090RTX: Promt Parsing, Token Generation and Power Level

Reddit r/LocalLLaMA

A user benchmarks the Nvidia 5090 RTX GPU for LLM inference using llama.cpp, measuring prompt processing and token generation at various power levels, finding that prompt processing is more sensitive to power limits than token generation, and noting differences from the 4090 RTX.