vram

Tag

Cards List
#vram

Devs - you have 64gb of VRAM - which model do you use for coding?

Reddit r/LocalLLaMA ↗ · 2026-06-30

A developer with 64GB VRAM shares their preference for an unsloth version of Qwen 3.5 122b-a10b for coding and asks the community for their recommendations.

0 favorites 0 likes
#vram

@che_shr_cat: 1/ We have been treating GPU memory all wrong. What if the GPU didn't need to store your model at all? MegaTrain enable…

X AI KOLs Timeline ↗ · 2026-06-29 Cached

MegaTrain enables full-precision training of 100B+ LLMs on a single GPU by treating VRAM as a transient stateless cache, inverting the memory hierarchy.

0 favorites 0 likes
#vram

@Tono_Ken3: Whoa! GLM-5.2 is running on 16GB VRAM! RDIMMer victory lol

X AI KOLs Timeline ↗ · 2026-06-28 Cached

GLM-5.2 now runs on 16GB VRAM; RDIMMer wins.

0 favorites 0 likes
#vram

Benchmarking Self-Hosted Gemma 2 9B vs. Frontier APIs: The FP8 Quantization Prefill Tax and VRAM Realities on an NVIDIA L4 [P]

Reddit r/MachineLearning ↗ · 2026-06-27

This benchmark compares an unquantized Gemma 2 9B model with an FP8 quantized variant on an NVIDIA L4 GPU, revealing that FP8 quantization introduces a prefill tax (higher TTFT) but improves decoding latency and VRAM usage, with minimal semantic drift for narrow tasks.

0 favorites 0 likes
#vram

96 gig 5090s from Shenzhen's Huaqiangbei

Reddit r/LocalLLaMA ↗ · 2026-06-27

Reports of a 96GB VRAM modded RTX 5090 (Blackwell RTX 6000) are confirmed from Shenzhen's Huaqiangbei market, priced around $8,200 total for the hacked card.

0 favorites 0 likes
#vram

NVFP4 kv cache quantization on sm120 will make 32GB VRAM systems very capable

Reddit r/LocalLLaMA ↗ · 2026-06-18

NVFP4 KV cache quantization on sm120 significantly improves memory efficiency for large language models, enabling 32GB VRAM systems to achieve ~60 tok/sec inference at 196k context size with Qwen3.6-27B.

0 favorites 0 likes
#vram

Want to build a custom model

Reddit r/LocalLLaMA ↗ · 2026-06-14

A user discusses building a small autocomplete model (25M parameters) as a learning project, mentions hardware constraints (32GB VRAM), data requirements (~100M tokens), and seeks advice on datasets and data formatting for autocomplete-style training.

0 favorites 0 likes
#vram

Best models in 3x3090 (72GB VRAM) in Q2 2026?

Reddit r/LocalLLaMA ↗ · 2026-06-13

A user shares their experience running large LLMs on a 3x3090 (72GB VRAM) setup in Q2 2026, recommending models like GPT-OSS 120b, Qwen3.5 122b, and GLM Air 4.5 106B, and asking for newer alternatives.

0 favorites 0 likes
#vram

Buy recommendations on a thight Budget to aid my RX 6800

Reddit r/LocalLLaMA ↗ · 2026-06-11

This post discusses budget GPU options (Radeon VII vs two P100s) for LLM inference with an RX 6800, focusing on VRAM vs speed tradeoffs for MoE models.

0 favorites 0 likes
#vram

Gemma 4 QAT benchmark results (AMD 7900 XTX): faster, less VRAM, no quality loss

Reddit r/LocalLLaMA ↗ · 2026-06-05

A user benchmarks Google's Gemma 4 QAT models on an AMD 7900 XTX, reporting up to 45% faster generation, 83% higher throughput, and significant VRAM savings (e.g., 5.7GB for the 12B QAT model) with no quality loss compared to standard weights.

0 favorites 0 likes
#vram

Maybe KV cache offload to RAM isn't bad

Reddit r/LocalLLaMA ↗ · 2026-06-05

A user shares their experience offloading the KV cache to RAM in llama.cpp, achieving comparable speeds while freeing VRAM for larger models and context windows, suggesting this trade-off is often worthwhile.

0 favorites 0 likes
#vram

Stop asking what model to run. There are literally only two.

Reddit r/LocalLLaMA ↗ · 2026-06-01

A tech enthusiast argues that only two local AI models (Qwen 3.6 35b a3b and Qwen 3.6 27b) are worth running, dismissing smaller models and recommending heavy quantization of larger models.

0 favorites 0 likes
#vram

Computex 2026: Intel launches Crescent Island GPU with up to 480GB VRAM

Reddit r/LocalLLaMA ↗ · 2026-06-01

Intel launched the Crescent Island GPU at Computex 2026, featuring up to 480GB VRAM and based on the Arc Xe 3P architecture, targeting next-generation AI workloads.

0 favorites 0 likes
#vram

nbd-vram: Use your NVIDIA GPU's VRAM as swap space on Linux

Lobsters Hottest ↗ · 2026-06-01 Cached

nbd-vram is a Linux tool that uses NVIDIA GPU VRAM as swap space via the NBD protocol and CUDA, providing extra memory for systems with soldered RAM and no upgrade path.

0 favorites 0 likes
#vram

Added an old 2070 Super to my rig and I can't go back...worse, now I need more

Reddit r/LocalLLaMA ↗ · 2026-05-31

A user shares their experience of adding an old NVIDIA 2070 Super GPU to their rig for extra VRAM, enabling them to run larger LLMs like Qwen3.6-27B at high quantization and context size with good performance, and now considering upgrading to a 3090 for even more VRAM.

0 favorites 0 likes
#vram

Qwen3.6 27B Pure Quant: 40 tok/s on 16 GB VRAM

Reddit r/LocalLLaMA ↗ · 2026-05-22

A quantized version of Qwen3.6 27B using a pure Q4_K_M method fits entirely in 16 GB VRAM, achieving up to 40 tok/s token generation speed with MTP, and significantly reducing model size compared to other GGUF variants.

0 favorites 0 likes
#vram

Can't believe I got it working! Dual GPU - 48gb VRAM llama-cpp server - R7900 + 7800XT

Reddit r/LocalLLaMA ↗ · 2026-05-22

A user successfully set up a dual-GPU llama-cpp server with 48GB VRAM using an AMD Radeon PRO and 7800 XT via Vulkan in Docker on Kubuntu 24.04.

0 favorites 0 likes
#vram

Seeking resources to read about llama.cpp server and how offloading works

Reddit r/LocalLLaMA ↗ · 2026-05-22

A user shares their experience with llama.cpp server's model offloading, noting performance trade-offs and quiet operation, and asks for resources to understand how the tool manages memory across VRAM and system RAM.

0 favorites 0 likes
#vram

Quantizing MTP KV Cache = free lunch?

Reddit r/LocalLLaMA ↗ · 2026-05-18

Quantizing the Multi-Token Prediction (MTP) KV cache to q8_0 in llama.cpp for Qwen models reduces VRAM usage without affecting inference speed or acceptance rate, effectively providing a 'free lunch' for memory-constrained setups.

0 favorites 0 likes
#vram

Qwen 3.6 27B on 24GB VRAM setup: backend comparisons, quant choice and settings (llama.cpp, ik_llama.cpp, BeeLlama, vllm)

Reddit r/LocalLLaMA ↗ · 2026-05-18

The article compares llama.cpp backends for running Qwen 3.6 27B on an RTX 3090 24GB, finding ik_llama.cpp with IQ4_KS quantization yields the best performance (1261 tok/s prefill, 72.9 tok/s decode).

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback