Tag
Tweet highlights running the Qwen 3.8 27B model locally on an RTX 5090 system with 32GB VRAM, achieving 115 tokens/sec, and notes the official BF16 checkpoint is 55.6GB.
SGLang has updated its deployment recipes for the Qwen3.8-27B model on RTX 5090 and RTX Pro 6000 hardware, adding variants for different configurations with tuning options.
The author quantizes DeepSeek V4 0731, fixing FP8 downconversion issues that skew baselines, and benchmarks 38 quant files on 8× RTX 5090 to show GPU-dependent results and file-size-based comparisons.
Benchmarks unsloth's Muse Glimmer 30B on an RTX 5090 with speculative decoding, achieving up to 253 t/s using a DFlash draft model and a GPU-based argmax PR, though the PR is still a draft.
A purported RTX 5090 with 96GB of VRAM has been spotted on Alibaba, hinting at a possible new GPU variant from Nvidia.
A Redditor built an open-source tool called 12vhpwr-guard that monitors per-pin current on ASUS Astral RTX 5080/5090 GPUs and shuts down the PC if any pin exceeds 9.5A for 15 seconds, helping prevent 12VHPWR connector melting.
Technical write-up on running DeepSeek-V4-Flash-0731 with full 1M context on a single RTX 5090 + 256GB DDR5 desktop using vLLM with CPU/RAM offloading, achieving ~800 tps prefill and ~15 tps decode.
Autonomous AI open-sources a Personal AI Computer build powered by 2x or 4x NVIDIA RTX 5090s, enabling fully local, private AI hosting without API costs or rate limits.
Best Buy has raised the price of the Asus ROG Astral RTX 5080 OC to $2,099, exceeding the RTX 5090's MSRP, reflecting ongoing component shortages and GPU price hikes.
A developer created a code-review tool running on a consumer RTX 5090 GPU using open-weight models, achieving F1 22.7 on the Martian code-review benchmark, and is considering turning it into a product or open-sourcing it.
A detailed benchmark of Unsloth's Qwen3.6-27B NVFP4 model on RTX 5090 GPUs, showing MTP (multi-token prediction) gives large speedups for single requests at short context but becomes detrimental under batch concurrency or long contexts.
Chinese factories are repurposing consumer RTX 5090 GPUs into 128GB server cards as a workaround to export controls on dedicated AI chips, creating a grey-market competitor to official enterprise offerings.
Benchmark shows that running 4-5 parallel agents with LM Studio on RTX 5090 maximizes throughput, while more agents yield diminishing returns due to VRAM and compute splitting.
Running Qwen3.6 27B on an RTX 5090, achieving 6.4k tokens per second after tuning MTP and cache settings, demonstrating optimization techniques for inference.
Running the quantized Gemma-4-31B model on an RTX 5090 increases context length from 35k to 80k, showcasing significant performance improvement.
User shares optimization benchmarks for DeepSeek-V4-Flash (Q2_K) running on an RTX 5090 using a fork of llama.cpp, achieving 21.3 tokens/s generation and 1 million context size.
Ahmad Osman predicts that within 18 months, a GPU like the RTX 5090 will be able to host intelligence equivalent to GLM 5.2.
Qwen 3.6 27B is a dense 27B model that achieves impressive performance on local hardware with 256k context, running at 30 tokens/s on MacBook Max M5 and 50 tokens/s on RTX 5090, and is considered by some as the first local model with true general intelligence.
A detailed analysis on whether to run AI models locally or via API, covering hardware options like RTX 5090, RTX PRO 6000, and DGX Spark, with emphasis on memory vs bandwidth trade-offs, cost considerations, and privacy needs.
MSI's RTX 5090 GPU operates at 475-500W for inference or training, with a warning about cable bending.