rtx-5090

Tag

Cards List
#rtx-5090

Nvidia's RTX 5090 vanishes from online retail in the US — third-party sellers now demand as much as $9,500 for Nvidia's fastest GPU

Reddit r/LocalLLaMA ↗ · 2026-09-14 Cached

The Nvidia RTX 5090 GPU has vanished from US online retail, with third-party sellers demanding up to $9,500, driven by AI demand and raising scam risks.

0 favorites 0 likes
#rtx-5090

nvidia rtx 5090 with 96gb of vram.

Reddit r/LocalLLaMA ↗ · 2026-09-11

A China-modified Nvidia RTX 5090 with 96GB of VRAM is available on Alibaba for under $4,000, offering three times more memory at 65% of the original cost.

0 favorites 0 likes
#rtx-5090

NVFP4 on VOLTA! Despite being built for Blackwell, I made four 2017 V100s run Qwen 3.8 NVFP4 natively and match my $6000 RTX 5090.

Reddit r/LocalLLaMA ↗ · 2026-08-19

The author achieved performance parity between four 2017 Tesla V100 GPUs and a modern RTX 5090 when running the Qwen 3.8 model with NVFP4 precision, using custom software optimizations like the QPN kernel for efficient inference.

0 favorites 0 likes
#rtx-5090

@rohanpaul_ai: Beautiful visual of somebody running, qwen 3.8 27B locally on a RTX 5090 32 GB VRAM system with 115 tokens/sec note, Qw…

X AI KOLs Timeline ↗ · 2026-08-17 Cached

Tweet highlights running the Qwen 3.8 27B model locally on an RTX 5090 system with 32GB VRAM, achieving 115 tokens/sec, and notes the official BF16 checkpoint is 55.6GB.

0 favorites 0 likes
#rtx-5090

@sgl_project: We pushed some updates to the RTX 5090 / RTX Pro 6000 recipes in the Qwen3.8-27B cookbook http://docs.sglang.io/cookboo…

X AI KOLs Timeline ↗ · 2026-08-17 Cached

SGLang has updated its deployment recipes for the Qwen3.8-27B model on RTX 5090 and RTX Pro 6000 hardware, adding variants for different configurations with tuning options.

0 favorites 0 likes
#rtx-5090

We quantized DeepSeek V4 0731 and benchmarked it against popular quants on 8× RTX 5090

Reddit r/LocalLLaMA ↗ · 2026-08-11

The author quantizes DeepSeek V4 0731, fixing FP8 downconversion issues that skew baselines, and benchmarks 38 quant files on 8× RTX 5090 to show GPU-dependent results and file-size-based comparisons.

0 favorites 0 likes
#rtx-5090

Achievable 253 t/s - unsloth/Muse Glimmer 30B UD-Q5_K_M on a 5090

Reddit r/LocalLLaMA ↗ · 2026-08-10

Benchmarks unsloth's Muse Glimmer 30B on an RTX 5090 with speculative decoding, achieving up to 253 t/s using a DFlash draft model and a GPU-based argmax PR, though the PR is still a draft.

0 favorites 0 likes
#rtx-5090

RTX 5090 96GB spotted on Alibaba?

Reddit r/LocalLLaMA ↗ · 2026-08-09

A purported RTX 5090 with 96GB of VRAM has been spotted on Alibaba, hinting at a possible new GPU variant from Nvidia.

0 favorites 0 likes
#rtx-5090

RTX 5090 Owner Built An Open-Source Tool That Shuts Down PC If It Detects The 12VHPWR Cable Drawing Too Much Power, But It Can Only Work On Specific GPUs

Reddit r/LocalLLaMA ↗ · 2026-08-07 Cached

A Redditor built an open-source tool called 12vhpwr-guard that monitors per-pin current on ASUS Astral RTX 5080/5090 GPUs and shuts down the PC if any pin exceeds 9.5A for 15 seconds, helping prevent 12VHPWR connector melting.

0 favorites 0 likes
#rtx-5090

[Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]

Reddit r/LocalLLaMA ↗ · 2026-08-04

Technical write-up on running DeepSeek-V4-Flash-0731 with full 1M context on a single RTX 5090 + 256GB DDR5 desktop using vLLM with CPU/RAM offloading, achieving ~800 tps prefill and ~15 tps decode.

0 favorites 0 likes
#rtx-5090

@dee_hw: Qwen3.8-27B is coming. We open source our 2x RTX 5090 build, so you can host it locally: on-prem, private, no rate limi…

X AI KOLs Timeline ↗ · 2026-08-03 Cached

Autonomous AI open-sources a Personal AI Computer build powered by 2x or 4x NVIDIA RTX 5090s, enabling fully local, private AI hosting without API costs or rate limits.

0 favorites 0 likes
#rtx-5090

Best Buy is selling an RTX 5080 for more than the RTX 5090’s MSRP

The Verge ↗ · 2026-07-30 Cached

Best Buy has raised the price of the Asus ROG Astral RTX 5080 OC to $2,099, exceeding the RTX 5090's MSRP, reflecting ongoing component shortages and GPU price hikes.

0 favorites 0 likes
#rtx-5090

I created code-review runs on 5090. Its scores F1 22.7 on Martian

Reddit r/AI_Agents ↗ · 2026-07-22

A developer created a code-review tool running on a consumer RTX 5090 GPU using open-weight models, achieving F1 22.7 on the Martian code-review benchmark, and is considering turning it into a product or open-sourcing it.

0 favorites 0 likes
#rtx-5090

I benchmarked Unsloth's Qwen3.6-27B NVFP4 on 1x/2x 5090s. MTP is great until it really isn't.

Reddit r/LocalLLaMA ↗ · 2026-07-21

A detailed benchmark of Unsloth's Qwen3.6-27B NVFP4 model on RTX 5090 GPUs, showing MTP (multi-token prediction) gives large speedups for single requests at short context but becomes detrimental under batch concurrency or long contexts.

0 favorites 0 likes
#rtx-5090

@kyzoroX: Chinese factories are buying RTX 5090s in bulk, desoldering the chips, and rebuilding them as 128GB server cards. Rough…

X AI KOLs Timeline ↗ · 2026-07-16 Cached

Chinese factories are repurposing consumer RTX 5090 GPUs into 128GB server cards as a workaround to export controls on dedicated AI chips, creating a grey-market competitor to official enterprise offerings.

0 favorites 0 likes
#rtx-5090

If you use Open Code or other agenting programs you are leaving a lot of t/s if you don't actually use agents in parallel. Benchmark : RTX5090, Qwen3.6 35B loaded via LM studio with parallel tasks set to 8

Reddit r/LocalLLaMA ↗ · 2026-07-12

Benchmark shows that running 4-5 parallel agents with LM Studio on RTX 5090 maximizes throughput, while more agents yield diminishing returns due to VRAM and compute splitting.

0 favorites 0 likes
#rtx-5090

Qwen3.6 27B on a 5090, 6.4k sample tok/s distribution after tuning MTP/cache settings

Reddit r/LocalLLaMA ↗ · 2026-07-04

Running Qwen3.6 27B on an RTX 5090, achieving 6.4k tokens per second after tuning MTP and cache settings, demonstrating optimization techniques for inference.

0 favorites 0 likes
#rtx-5090

RTX5090, gemma-4-31B-it-Q6_K.gguf. Context: before - 35k, after - 80k!

Reddit r/LocalLLaMA ↗ · 2026-07-04

Running the quantized Gemma-4-31B model on an RTX 5090 increases context length from 35k to 80k, showcasing significant performance improvement.

0 favorites 0 likes
#rtx-5090

Deepseek V4 Flash running on RTX 5090 MoE

Reddit r/LocalLLaMA ↗ · 2026-07-03

User shares optimization benchmarks for DeepSeek-V4-Flash (Q2_K) running on an RTX 5090 using a fork of llama.cpp, achieving 21.3 tokens/s generation and 1 million context size.

0 favorites 0 likes
#rtx-5090

@TheAhmadOsman: PREDICTION

X AI KOLs Timeline ↗ · 2026-07-03 Cached

Ahmad Osman predicts that within 18 months, a GPU like the RTX 5090 will be able to host intelligence equivalent to GLM 5.2.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback