quantized-model

Tag

Cards List
#quantized-model

@TheAhmadOsman: RTX 3090 owners tonight will be running Kimi_K3_3T_Q_0.001_K GGUF

X AI KOLs Following · 2026-07-16 Cached

A quantized GGUF version of the Kimi K3 model is now available, optimized for running on RTX 3090 GPUs.

0 favorites 0 likes
#quantized-model

@ciruai: Finally 256k context for 16GB cards on smart models at high speeds! Using a 4080 Super 16GB I show you how to get full …

X AI KOLs Timeline · 2026-07-16 Cached

Demonstrates achieving 256k context on a 16GB RTX 4080 Super using Ternary Bonsai 27B Q2_0 model with llama.cpp, achieving up to 141 tok/s generation speed.

0 favorites 0 likes
#quantized-model

I created a 140 GB IQ2_XXS REAP quant of GLM 5.2 for coding. Looking for testers.

Reddit r/LocalLLaMA · 2026-07-08

A 140 GB IQ2_XXS REAP quantized version of GLM 5.2 for coding has been created, and the author is looking for testers.

0 favorites 0 likes
#quantized-model

Ornith-1.0-35B Q3_K_M: ~17 GB VRAM, KLD-checked against BF16

Reddit r/LocalLLaMA · 2026-06-27

Ornith-1.0-35B Q3_K_M is a 3-bit quantized version of a 35B parameter model, requiring about 17 GB VRAM, with KLD checking against BF16 to ensure fidelity.

0 favorites 0 likes
#quantized-model

@MiaAI_lab: FYI the best Qwen 3.6 35b nvfp4 to run is the @NVIDIAAI nvfp4. Do not use unsloth nvfp4, it performs worse. https://hug…

X AI KOLs Timeline · 2026-06-22 Cached

NVIDIA's nvfp4 quantized version of Qwen 3.6 35B is recommended over the Unsloth variant, offering better performance. The model is available on HuggingFace for use in AI applications.

0 favorites 0 likes
#quantized-model

Unsloth Gemma 4 QAT MTP assistant models now available

Reddit r/LocalLLaMA · 2026-06-09

Unsloth released Gemma 4 QAT MTP assistant models as GGUF files on Hugging Face, available in q8_0 and larger quantization formats.

0 favorites 0 likes
#quantized-model

I put together a Rust-native, CPU-only implementation of LFM2.5-8B-A1B

Reddit r/LocalLLaMA · 2026-06-09 Cached

The author released a pure Rust, CPU-only inference implementation of the LFM2.5-8B-A1B model (4-bit Q4KM quantization), achieving a decode speed of approximately 37 tokens/s and memory usage around 7GB. The goal is to make LLMs runnable on cheap VPS or older machines. The implementation is open source and published as a cargo crate.

0 favorites 0 likes
#quantized-model

Jackrong/Qwopus3.6-27B-v2-MTP-GGUF

Hugging Face Models Trending · 2026-05-21 Cached

Jackrong/Qwopus3.6-27B-v2-MTP-GGUF is a quantized GGUF version of a 27B parameter language model, hosted on Hugging Face with instructions for use with various libraries and tools.

0 favorites 0 likes
← Back to home

Submit Feedback