vram-usage

Tag

Cards List
#vram-usage

Qwen3.8-27B: slower tokens, faster and better results

Reddit r/LocalLLaMA · 2d ago Cached

Qwen3.8-27B is a new AI model that emphasizes wall-clock time over token speed, offering superior intelligence for local deployment on consumer hardware with 32 GB VRAM.

0 favorites 0 likes
#vram-usage

PSA: llama.cpp now loads MTP tensors by default for any draft-mtp arch, even with MTP disabled

Reddit r/LocalLLaMA · 2026-07-29

Recent versions of llama.cpp now automatically load MTP tensors for draft-mtp architectures, even if speculative decoding is not enabled, potentially increasing VRAM usage for users with bundled MTP blocks in their GGUF files.

0 favorites 0 likes
← Back to home

Submit Feedback