rtx-5090

Tag

Cards List
#rtx-5090

@rohanpaul_ai: Beautiful visual of somebody running, qwen 3.8 27B locally on a RTX 5090 32 GB VRAM system with 115 tokens/sec note, Qw…

X AI KOLs Timeline · 20h ago Cached

Tweet highlights running the Qwen 3.8 27B model locally on an RTX 5090 system with 32GB VRAM, achieving 115 tokens/sec, and notes the official BF16 checkpoint is 55.6GB.

0 favorites 0 likes
#rtx-5090

@sgl_project: We pushed some updates to the RTX 5090 / RTX Pro 6000 recipes in the Qwen3.8-27B cookbook http://docs.sglang.io/cookboo…

X AI KOLs Timeline · 22h ago Cached

SGLang has updated its deployment recipes for the Qwen3.8-27B model on RTX 5090 and RTX Pro 6000 hardware, adding variants for different configurations with tuning options.

0 favorites 0 likes
#rtx-5090

We quantized DeepSeek V4 0731 and benchmarked it against popular quants on 8× RTX 5090

Reddit r/LocalLLaMA · 6d ago

The author quantizes DeepSeek V4 0731, fixing FP8 downconversion issues that skew baselines, and benchmarks 38 quant files on 8× RTX 5090 to show GPU-dependent results and file-size-based comparisons.

0 favorites 0 likes
#rtx-5090

Achievable 253 t/s - unsloth/Muse Glimmer 30B UD-Q5_K_M on a 5090

Reddit r/LocalLLaMA · 2026-08-10

Benchmarks unsloth's Muse Glimmer 30B on an RTX 5090 with speculative decoding, achieving up to 253 t/s using a DFlash draft model and a GPU-based argmax PR, though the PR is still a draft.

0 favorites 0 likes
#rtx-5090

RTX 5090 96GB spotted on Alibaba?

Reddit r/LocalLLaMA · 2026-08-09

A purported RTX 5090 with 96GB of VRAM has been spotted on Alibaba, hinting at a possible new GPU variant from Nvidia.

0 favorites 0 likes
#rtx-5090

RTX 5090 Owner Built An Open-Source Tool That Shuts Down PC If It Detects The 12VHPWR Cable Drawing Too Much Power, But It Can Only Work On Specific GPUs

Reddit r/LocalLLaMA · 2026-08-07 Cached

A Redditor built an open-source tool called 12vhpwr-guard that monitors per-pin current on ASUS Astral RTX 5080/5090 GPUs and shuts down the PC if any pin exceeds 9.5A for 15 seconds, helping prevent 12VHPWR connector melting.

0 favorites 0 likes
#rtx-5090

[Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]

Reddit r/LocalLLaMA · 2026-08-04

Technical write-up on running DeepSeek-V4-Flash-0731 with full 1M context on a single RTX 5090 + 256GB DDR5 desktop using vLLM with CPU/RAM offloading, achieving ~800 tps prefill and ~15 tps decode.

0 favorites 0 likes
#rtx-5090

@dee_hw: Qwen3.8-27B is coming. We open source our 2x RTX 5090 build, so you can host it locally: on-prem, private, no rate limi…

X AI KOLs Timeline · 2026-08-03 Cached

Autonomous AI open-sources a Personal AI Computer build powered by 2x or 4x NVIDIA RTX 5090s, enabling fully local, private AI hosting without API costs or rate limits.

0 favorites 0 likes
#rtx-5090

Best Buy is selling an RTX 5080 for more than the RTX 5090’s MSRP

The Verge · 2026-07-30 Cached

Best Buy has raised the price of the Asus ROG Astral RTX 5080 OC to $2,099, exceeding the RTX 5090's MSRP, reflecting ongoing component shortages and GPU price hikes.

0 favorites 0 likes
#rtx-5090

I created code-review runs on 5090. Its scores F1 22.7 on Martian

Reddit r/AI_Agents · 2026-07-22

A developer created a code-review tool running on a consumer RTX 5090 GPU using open-weight models, achieving F1 22.7 on the Martian code-review benchmark, and is considering turning it into a product or open-sourcing it.

0 favorites 0 likes
#rtx-5090

I benchmarked Unsloth's Qwen3.6-27B NVFP4 on 1x/2x 5090s. MTP is great until it really isn't.

Reddit r/LocalLLaMA · 2026-07-21

A detailed benchmark of Unsloth's Qwen3.6-27B NVFP4 model on RTX 5090 GPUs, showing MTP (multi-token prediction) gives large speedups for single requests at short context but becomes detrimental under batch concurrency or long contexts.

0 favorites 0 likes
#rtx-5090

@kyzoroX: Chinese factories are buying RTX 5090s in bulk, desoldering the chips, and rebuilding them as 128GB server cards. Rough…

X AI KOLs Timeline · 2026-07-16 Cached

Chinese factories are repurposing consumer RTX 5090 GPUs into 128GB server cards as a workaround to export controls on dedicated AI chips, creating a grey-market competitor to official enterprise offerings.

0 favorites 0 likes
#rtx-5090

If you use Open Code or other agenting programs you are leaving a lot of t/s if you don't actually use agents in parallel. Benchmark : RTX5090, Qwen3.6 35B loaded via LM studio with parallel tasks set to 8

Reddit r/LocalLLaMA · 2026-07-12

Benchmark shows that running 4-5 parallel agents with LM Studio on RTX 5090 maximizes throughput, while more agents yield diminishing returns due to VRAM and compute splitting.

0 favorites 0 likes
#rtx-5090

Qwen3.6 27B on a 5090, 6.4k sample tok/s distribution after tuning MTP/cache settings

Reddit r/LocalLLaMA · 2026-07-04

Running Qwen3.6 27B on an RTX 5090, achieving 6.4k tokens per second after tuning MTP and cache settings, demonstrating optimization techniques for inference.

0 favorites 0 likes
#rtx-5090

RTX5090, gemma-4-31B-it-Q6_K.gguf. Context: before - 35k, after - 80k!

Reddit r/LocalLLaMA · 2026-07-04

Running the quantized Gemma-4-31B model on an RTX 5090 increases context length from 35k to 80k, showcasing significant performance improvement.

0 favorites 0 likes
#rtx-5090

Deepseek V4 Flash running on RTX 5090 MoE

Reddit r/LocalLLaMA · 2026-07-03

User shares optimization benchmarks for DeepSeek-V4-Flash (Q2_K) running on an RTX 5090 using a fork of llama.cpp, achieving 21.3 tokens/s generation and 1 million context size.

0 favorites 0 likes
#rtx-5090

@TheAhmadOsman: PREDICTION

X AI KOLs Timeline · 2026-07-03 Cached

Ahmad Osman predicts that within 18 months, a GPU like the RTX 5090 will be able to host intelligence equivalent to GLM 5.2.

0 favorites 0 likes
#rtx-5090

@Xudong07452910: A hot comment section on Hacker News: Qwen 3.6 27B is the ideal choice for local development. Key findings: dense parameter model, native support for 256k context, running Q8_0 quantized version at 30 tokens/…

X AI KOLs Timeline · 2026-07-03 Cached

Qwen 3.6 27B is a dense 27B model that achieves impressive performance on local hardware with 256k context, running at 30 tokens/s on MacBook Max M5 and 50 tokens/s on RTX 5090, and is considered by some as the first local model with true general intelligence.

0 favorites 0 likes
#rtx-5090

@RayFernando1337: https://x.com/RayFernando1337/status/2070621713952579990

X AI KOLs Following · 2026-06-26 Cached

A detailed analysis on whether to run AI models locally or via API, covering hardware options like RTX 5090, RTX PRO 6000, and DGX Spark, with emphasis on memory vs bandwidth trade-offs, cost considerations, and privacy needs.

0 favorites 0 likes
#rtx-5090

RTX 5090 MSI, only inference or training at 475-500W. Make sure to not bend you cable!

Reddit r/LocalLLaMA · 2026-06-20

MSI's RTX 5090 GPU operates at 475-500W for inference or training, with a warning about cable bending.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback