vram

Tag

Cards List
#vram

@rohanpaul_ai: Beautiful visual of somebody running, qwen 3.8 27B locally on a RTX 5090 32 GB VRAM system with 115 tokens/sec note, Qw…

X AI KOLs Timeline ↗ · 2026-08-17 Cached

Tweet highlights running the Qwen 3.8 27B model locally on an RTX 5090 system with 32GB VRAM, achieving 115 tokens/sec, and notes the official BF16 checkpoint is 55.6GB.

0 favorites 0 likes
#vram

The dream is to reach 200GB VRAM

Reddit r/LocalLLaMA ↗ · 2026-08-16

A user outlines a step-by-step plan to achieve 200GB of VRAM by combining multiple NVIDIA GPUs in a custom PC build, addressing purchase, installation, and power management.

0 favorites 0 likes
#vram

How many people have 24gb over gpu here?

Reddit r/LocalLLaMA ↗ · 2026-08-16

The author discusses the low adoption of the qwen 3.8 27b model based on download counts and estimates that very few users have the high-VRAM GPUs needed for productive local LLM development.

0 favorites 0 likes
#vram

I tested the CMP170HX

Reddit r/LocalLLaMA ↗ · 2026-08-11

A hands-on benchmark of Nvidia CMP170HX mining cards repurposed as 64GB VRAM AI inference accelerators, showing they can run large local LLMs like DeepSeek V4-Flash and gpt-oss-120B at useful speeds, with caveats around Ampere-class throughput and PCIe Gen2 x4 connectivity.

0 favorites 0 likes
#vram

12GB VRAM gang, what's our plan?

Reddit r/LocalLLaMA ↗ · 2026-08-11

Discussion about running LLMs on 12GB VRAM, noting current focus on dense models like Muse Glimmer 30B and Qwen 3.8 27B, and questioning whether upgrading to 24GB VRAM is needed.

0 favorites 0 likes
#vram

Rumored 50-series Super refresh bumps everything +50% VRAM

Reddit r/LocalLLaMA ↗ · 2026-08-11

Leaked specs suggest Nvidia's rumored 50-series Super refresh will increase VRAM on several GPUs using new 3GB GDDR7 modules, making them more attractive for local LLM use, though pricing remains a concern.

0 favorites 0 likes
#vram

Muse Glimmer ACTUALLY fits on a single RTX 3090

Reddit r/LocalLLaMA ↗ · 2026-08-10

User reports that Muse Glimmer, a 30B model, fits on a single RTX 3090 with full 256k context using Q4_K_XL quantization and DFlash, achieving 64-124 tok/s and perfect long-context retrieval, unlike comparable models.

0 favorites 0 likes
#vram

1M context with 17 GB model in 24 GB VRAM: "for the first time I was able to load a context of almost 1M tokens and extract 7 needles from various parts of the text"

Reddit r/LocalLLaMA ↗ · 2026-08-10

A user reports successfully running a 1M-token context on a single RTX 3090 using a Qwen-based 35B A3B model (17GB VRAM) with KVarN 4-bit KV-cache quantization in a BeeLlama.cpp fork, extracting 7 needles from different parts of the text.

0 favorites 0 likes
#vram

Deepseek V4 Flash just hit Colibri, does anyone have numbers?

Reddit r/LocalLLaMA ↗ · 2026-08-05

User asks for performance numbers on Deepseek V4 Flash running via Colibri, focusing on high VRAM setups, long context prefill, and token generation speed for agentic workloads.

0 favorites 0 likes
#vram

70-class VRAM stagnation

Reddit r/LocalLLaMA ↗ · 2026-08-03

The author observes that Nvidia's desktop 70-class GPUs have stayed at 12GB VRAM across two generations, and suggests Nvidia may be intentionally limiting memory to preserve demand for higher-margin AI-focused hardware.

0 favorites 0 likes
#vram

Daniel Han of Unsloth validates Qwen3.8-27B will run only 17GB VRAM

Reddit r/LocalLLaMA ↗ · 2026-08-03

Daniel Han of Unsloth validates that Qwen3.8-27B will run in only 17GB VRAM, making it accessible for local inference.

0 favorites 0 likes
#vram

Today was the perfect day for Poolside to drop Laguna S 2.1 because I just got these in! Finally have a half decent amount of VRAM. 3x V620 = 96 GB.

Reddit r/LocalLLaMA ↗ · 2026-07-22

Poolside released Laguna S 2.1, and the user acquired 3x AMD V620 GPUs totaling 96 GB VRAM.

0 favorites 0 likes
#vram

How to Train a Gen AI Kick Drum Model on Your Old Linux Desktop with 6GB VRAM

Hacker News Top ↗ · 2026-07-16

A guide on training a generative AI model for kick drum sounds using an old Linux desktop with only 6GB VRAM, making AI audio generation accessible with limited hardware.

0 favorites 0 likes
#vram

@brianbellx: I removed 423 GB from GLM‑5.2 without changing the model. 1,403 GB → 980 GB. 753B weights. Bit for bit exact. No quanti…

X AI KOLs Timeline ↗ · 2026-07-12 Cached

A technique to remove 423 GB from GLM-5.2 (753B weights) without quantization or retraining, achieving bit-exact compression by keeping weights compressed in VRAM.

0 favorites 0 likes
#vram

Ultra budget 20GB vram with 448GB/s for $100 bucks.

Reddit r/LocalLLaMA ↗ · 2026-07-11

Demonstrates achieving 20GB VRAM and 448GB/s bandwidth for around $100 using two NVIDIA P102-100 cards, running a llama.cpp server with a Qwen model and supporting 3 concurrent users with large context.

0 favorites 0 likes
#vram

@sudoingX: save this one. it answers a question every 24gb GPU owner asks and almost nobody gets right. how much context can i act…

X AI KOLs Timeline ↗ · 2026-07-11 Cached

A user shares detailed VRAM usage measurements for running Qwen 3.6 27b on a 24GB GPU, showing context window sizes and headroom, and notes that a used RTX 3090 performs identically to newer cards.

0 favorites 0 likes
#vram

@theemozilla: We're working on making the local model experience better in Hermes, what are the best local models at each weight clas…

X AI KOLs Timeline ↗ · 2026-07-10 Cached

The user asks for recommendations on the best local AI models for different VRAM classes (8-16GB, 24-32GB, 128GB), mentioning Gemma4, Qwen, and DeepSeek variants, as they work on improving local model support in Hermes.

0 favorites 0 likes
#vram

Pay attention: a few chats waiting in tray reserve 1GB VRAM for themselves.

Reddit r/LocalLLaMA ↗ · 2026-07-03

Applications like Discord, Steam, and Telegram reserve VRAM even when minimized, consuming 1GB+ collectively; users working with LLMs should close these apps or disable hardware acceleration to free up VRAM.

0 favorites 0 likes
#vram

@tom_doerr: Personal AI Computer build guides with up to 384GB VRAM https://github.com/autonomous-ai/autonomous-computer…

X AI KOLs Timeline ↗ · 2026-07-02 Cached

Open-source build guides for a personal AI computer with up to 384GB VRAM, supporting configurations from home to on-prem business. Includes bill of materials, assembly photos, and software setup for running open-source AI models locally.

0 favorites 0 likes
#vram

Devs - you have 64gb of VRAM - which model do you use for coding?

Reddit r/LocalLLaMA ↗ · 2026-06-30

A developer with 64GB VRAM shares their preference for an unsloth version of Qwen 3.5 122b-a10b for coding and asks the community for their recommendations.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback