vram

Tag

Cards List
#vram

@UnslothAI: Qwen-Image-2.1 can now run locally on 12GB VRAM with Unsloth GGUFs! The 7B model performs on par with Nano Banana 2.0. …

X AI KOLs Timeline ↗ · 3d ago Cached

Unsloth has released GGUF quantized versions of Qwen-Image-2.1, enabling it to run locally on 12GB VRAM with performance comparable to Nano Banana 2.0.

0 favorites 0 likes
#vram

16GB (and in many cases 12GB) is the max vram most people will ever reasonably have

Reddit r/LocalLLaMA ↗ · 5d ago

The article discusses how 16GB of VRAM is the realistic high-end limit for most users due to financial constraints, but recent AI model improvements like Qwen 27B quants are enabling more capabilities on such hardware, with hopes for future architectural innovations.

0 favorites 0 likes
#vram

I tested Qwen3.8 27B IQ3_XXS (10.18GiB) vs Bonsai Ternary PQ2 (6.42GiB)

Reddit r/LocalLLaMA ↗ · 6d ago

The author compared the performance of Qwen3.8 27B IQ3_XXS and Bonsai Ternary PQ2 on limited VRAM, finding that Qwen is faster and uses fewer tokens, while Bonsai has a smaller file size but longer generation times.

0 favorites 0 likes
#vram

768gb vram for less than the price of one RTX 6000

Reddit r/LocalLLaMA ↗ · 2026-09-18

A user built a 768GB VRAM system using 12x64GB CMP170HX cards for less than the cost of one RTX 6000 Pro, enabling local inference of various large AI models with strong performance.

0 favorites 0 likes
#vram

I benchmarked IFM/K2-Horizon-7B on 16GB VRAM

Reddit r/LocalLLaMA ↗ · 2026-09-15

The article benchmarks three AI models on 16GB VRAM, showing that Qwen3.8-27B performs best, followed by Ornith-1.5-9B, with IFM/K2-Horizon-7B lagging behind in speed and task completion.

0 favorites 0 likes
#vram

My experience building 64GB VRAM AI SWE assistant/agent PC

Reddit r/LocalLLaMA ↗ · 2026-09-13

The author shares their experience building a high-VRAM AI PC with multiple GPUs to run local models for software engineering, detailing solutions to hardware challenges like PSU limitations and GPU mounting.

0 favorites 0 likes
#vram

nvidia rtx 5090 with 96gb of vram.

Reddit r/LocalLLaMA ↗ · 2026-09-11

A China-modified Nvidia RTX 5090 with 96GB of VRAM is available on Alibaba for under $4,000, offering three times more memory at 65% of the original cost.

0 favorites 0 likes
#vram

3060 12GB vs 4060 ti 16GB

Reddit r/LocalLLaMA ↗ · 2026-09-10

A user is evaluating whether to add a 4060 Ti 16GB GPU to a multi-GPU setup with 3060s for AI model parallelism and gaming, weighing the benefits of extra VRAM against potential memory bandwidth limitations.

0 favorites 0 likes
#vram

Server rebuild to custom loop. 2x RTX Titans 24gb, 1x 22gb 2080ti | T: 70GB VRAM.

Reddit r/LocalLLaMA ↗ · 2026-09-09

A user rebuilt their server with a custom liquid cooling loop, significantly reducing GPU temperatures from over 80°C to mid-40s under load, featuring a 5950X CPU, 64GB DDR4 RAM, and multiple high-VRAM GPUs totaling 70GB VRAM.

0 favorites 0 likes
#vram

I built a server with 768GB VRAM for frontier, but all new frontier open source models are likely to be two trillion or above now, including next GLM 6, am I cooked?

Reddit r/LocalLLaMA ↗ · 2026-09-04

An individual discusses building a server with 768GB VRAM for running frontier AI models but is concerned that new open-source models like GLM6 are becoming too large, prompting consideration of downsizing to smaller flash models.

0 favorites 0 likes
#vram

First time running local models

Reddit r/LocalLLaMA ↗ · 2026-08-31

A user shares their experience running the ik_llama model locally, praising its speed despite having only 12GB of VRAM.

0 favorites 0 likes
#vram

What would you do with $4,000?

Reddit r/LocalLLaMA ↗ · 2026-08-30

A user with a 5090 GPU discusses how to spend $4,000 to enhance their AI model testing and coding capabilities, including considerations for media creation and additional local compute power.

0 favorites 0 likes
#vram

TQwen 3.8 flash next ud1s on 6gb vram and 16 gb system ram

Reddit r/LocalLLaMA ↗ · 2026-08-28

A user shares their experience running the Qwen 3.8 flash next model on a system with 6GB VRAM and 16GB RAM using llama.cpp, achieving 6-7 tokens per second with 1-bit quantization, and asks for recommendations on quantization variants.

0 favorites 0 likes
#vram

What are the minimum specs required to run Qwen3.8-Flash-Next?

Reddit r/LocalLLaMA ↗ · 2026-08-27

A user inquires about the minimum hardware specifications, such as RAM and VRAM, required to run the Qwen3.8-Flash-Next AI model, including expected performance metrics.

0 favorites 0 likes
#vram

OpenCode with Qwen3.8-27B for Small Games or Browsing the Web With 16GB VRAM

Reddit r/LocalLLaMA ↗ · 2026-08-26

This article provides a step-by-step guide on running the Qwen3.8-27B AI model on a 16GB VRAM laptop using tools like exllamav3 and tabbyAPI, including installation, configuration, and performance benchmarks.

0 favorites 0 likes
#vram

I want to try qwen 3.8... but which gguf are we all using?

Reddit r/LocalLLaMA ↗ · 2026-08-25

A user discusses the challenge of selecting the appropriate GGUF quantization variant for the Qwen 3.8 AI model to achieve good performance with 32 GB of VRAM, seeking community recommendations.

0 favorites 0 likes
#vram

4xR9700, 2xMi210 or 4x4080S 32G

Reddit r/LocalLLaMA ↗ · 2026-08-24

The user is comparing GPU options like R9700, Mi210, and 4080S to achieve 128GB VRAM for running multiple AI models in parallel, considering factors like cost, performance, and compatibility.

0 favorites 0 likes
#vram

I fine tuned Gemma 4 12B for a 2.7x improvement on tool calling because I can't fit anything else comfortably into my 16 GBs of Vram

Reddit r/LocalLLaMA ↗ · 2026-08-23 Cached

A fine-tuned version of Gemma 4 12B that improves tool-calling reliability by 2.7x, optimized for consumer GPUs with 16GB VRAM using QLoRA training.

0 favorites 0 likes
#vram

Ladies and gentlemen I present to you Qwen3.8 27b 1bit brain damage quant

Reddit r/LocalLLaMA ↗ · 2026-08-20

A user tested the unsloth 1-bit quantized version of the Qwen 3.8 27B AI model on an 8GB VRAM system and found the results amusing.

0 favorites 0 likes
#vram

Anyone running qwen 3.8 27b on 5070ti (16GB)?

Reddit r/LocalLLaMA ↗ · 2026-08-18

A user asks if running the Qwen 3.8 27b model on a 5070ti GPU with 16GB VRAM is feasible using quantization for agentic coding purposes.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback