rtx-5090

Tag

Cards List
#rtx-5090

Latest LM Studio update killed MTP performance

Reddit r/LocalLLaMA · 2026-06-15

A user reports that the latest LM Studio update (0.4.17) eliminated the multi-token prediction speed boost, reverting to previous performance on an RTX 5090 setup.

0 favorites 0 likes
#rtx-5090

[Benchmark] DFlash Speculative Decoding + KV Cache Compression on RTX 5090 — 3.26x Speedup

Reddit r/LocalLLaMA · 2026-06-08

Benchmarks of DFlash speculative decoding combined with KV cache compression on RTX 5090 show up to 3.26x speedup on Qwen3.6-27B with minimal perplexity degradation, with q4_0/turbo4 providing the best balance.

0 favorites 0 likes
#rtx-5090

@songhan_mit: SANA Streaming: V2V on a single 5090

X AI KOLs Following · 2026-06-01 Cached

SANA Streaming enables video-to-video generation on a single NVIDIA RTX 5090 GPU.

0 favorites 0 likes
#rtx-5090

@danyurkin: i don't think i need cloud models anymore

X AI KOLs Following · 2026-05-20 Cached

A tweet demonstrates that Multi-Token Prediction (MTP) achieves significant speedups for Qwen models on dual RTX 5090 hardware, suggesting that local inference can now rival cloud-model performance.

0 favorites 0 likes
#rtx-5090

@populartourist: llama.cpp release b9235 added some new toys for boosting inference. Benchmarked Qwen3.6 27B on an RTX 5090 with llama.c…

X AI KOLs Following · 2026-05-20 Cached

llama.cpp release b9235 introduces speculative n-gram tuning, achieving up to ~7x throughput improvement on Qwen3.6 27B on an RTX 5090, with the k4v96 configuration showing the best sustained performance in 10k and 70k token tests.

0 favorites 0 likes
#rtx-5090

Testing llama.cpp MTP support on Qwen3.6 - RTX 5090

Reddit r/LocalLLaMA · 2026-05-17

A technical test of llama.cpp's new Multi-Token Prediction (MTP) support using Qwen3.6 models on an RTX 5090, comparing performance with and without MTP across different prompts and GGUF quantizations.

0 favorites 0 likes
#rtx-5090

@populartourist: Unsloth Qwen3.6 27B Q6_K doing over 100 t/s with MTP on RTX 5090. That's coming up from 45-50 t/s without MTP. That's i…

X AI KOLs Timeline · 2026-05-16 Cached

Unsloth Qwen3.6 27B Q6_K achieves over 100 tokens per second with MTP on RTX 5090, up from 45-50 t/s without MTP.

0 favorites 0 likes
#rtx-5090

Can a 5090 with qwen3.6 achieve > 3,000 tok/s ? bring your pitchforks (open-dllm)

Reddit r/LocalLLaMA · 2026-05-16

Open-dLLM adapts Qwen3.6 to use diffusion-based generation, achieving over 3,000 tok/s on an RTX 5090 for short sequences, with code released on GitHub.

0 favorites 0 likes
#rtx-5090

NVIDIA Reportedly Prepares RTX 5090 Price Hike Amid Rising GDDR7 Costs (maybe RTX 50 and PRO series as well)

Reddit r/LocalLLaMA · 2026-05-14 Cached

NVIDIA is reportedly planning to increase the price of its upcoming RTX 5090 graphics card due to rising costs of GDDR7 memory.

0 favorites 0 likes
#rtx-5090

I tracked EU GPU prices across 15 stores for 50+ days - RTX 5090 is the only card not dropping in price

Reddit r/LocalLLaMA · 2026-05-14

Tracking 15 EU GPU stores shows RTX 5090 prices rising 3% due to AI/workstation demand while other GPUs drop 7-9%, suggesting sustained high pricing for AI inference hardware.

0 favorites 0 likes
#rtx-5090

Is it worth getting a 5090 for my needs?

Reddit r/LocalLLaMA · 2026-05-13

User asks whether purchasing an RTX 5090 and high-end PC for ~$5500 is worth it for LLM experimentation and learning, compared to cloud compute alternatives.

0 favorites 0 likes
#rtx-5090

Gemma 4 26B Hits 600 Tok/s on One RTX 5090

Reddit r/LocalLLaMA · 2026-05-08

A benchmark shows that using vLLM with DFlash speculative decoding boosts Gemma 4 26B inference to ~578 tokens per second on a single RTX 5090, achieving a 2.56x speedup over baseline.

1 favorites 1 likes
#rtx-5090

Tried Qwen3.6-27B-UD-Q6_K_XL.gguf with CloudeCode, well I can't believe but it is usable

Reddit r/LocalLLaMA · 2026-04-22

User reports surprisingly usable coding performance from Qwen3-27B-UD-Q6_K_XL.gguf running locally on RTX 5090 at ~50 tok/s with 200K context, marking a significant leap in local model quality.

0 favorites 0 likes
#rtx-5090

@CuiMao: Honestly, running Claude Code locally with LM Studio is surprisingly solid—RTX 5090 handles 64k context at 200+ tokens/s.

X AI KOLs Timeline · 2026-04-20 Cached

User reports a satisfying experience running Claude Code locally via LM Studio on an RTX 5090, achieving 64k context length and 200+ tokens per second.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback