quantized

Tag

Cards List
#quantized

LFM 2.5 QAD

Reddit r/LocalLLaMA ↗ · 2026-08-19

LiquidAI releases LFM 2.5, a language model optimized for quantized deployment, with GGUF format available on Hugging Face.

0 favorites 0 likes
#quantized

huihui-ai/Huihui-DeepSeek-V4-Flash-0731-abliterated-GGUF

Hugging Face Models Trending ↗ · 2026-08-02 Cached

A model card for Huihui-DeepSeek-V4-Flash-0731-abliterated-GGUF, an abliterated (uncensored) GGUF quantized variant of DeepSeek-V4-Flash, designed for local use with llama.cpp and ds4.

0 favorites 0 likes
#quantized

@shaneparrish: https://x.com/shaneparrish/status/2075226155548807210

X AI KOLs Following ↗ · 2026-07-09 Cached

Demonstrates running two 80B Qwen models simultaneously on a MacBook Pro using BBQ-FP4 quantization, claiming functionally lossless performance and speed.

0 favorites 0 likes
#quantized

4-bit GLM-5.2 (753B MoE) on 4× DGX Spark: 70.8% on Terminal-Bench 2.1 vs 81.0% for the full model

Reddit r/LocalLLaMA ↗ · 2026-07-08

Running a 4-bit quantized version of GLM-5.2 (753B MoE) on 4 DGX Spark machines achieves 70.8% on Terminal-Bench 2.1, compared to 81.0% from the full model.

0 favorites 0 likes
#quantized

jlnsrk/GLM-5.2-colibri-int4

Hugging Face Models Trending ↗ · 2026-07-06 Cached

Pre-converted int4 quantized weights for the GLM-5.2 744B MoE model, designed to run on consumer hardware with ~25 GB RAM using the colibrì engine.

0 favorites 0 likes
#quantized

Un modello linguistico locale, privato 100%, sul tuo smartphone!!

Reddit r/ArtificialInteligence ↗ · 2026-07-05

Un modello linguistico locale e privato (Qwen 3 da 1.5B e 4B quantizzati) può girare offline su smartphone, con fine-tuning e LoRA distillato da un 32B.

0 favorites 0 likes
#quantized

RTX5090, gemma-4-31B-it-Q6_K.gguf. Context: before - 35k, after - 80k!

Reddit r/LocalLLaMA ↗ · 2026-07-04

Running the quantized Gemma-4-31B model on an RTX 5090 increases context length from 35k to 80k, showcasing significant performance improvement.

0 favorites 0 likes
#quantized

prism-ml/Bonsai-27B-mlx-1bit

Hugging Face Models Trending ↗ · 2026-07-04 Cached

Bonsai-27B is a 1-bit binary transformer model that achieves full 27B-class reasoning on a phone (iPhone 17 Pro Max) with ~3.9 GB footprint and ~11 tok/s, retaining ~90% of FP16 intelligence.

0 favorites 0 likes
#quantized

@BrianRoemmele: BOOM! Meet the open source Cambrian Explosion of repulsion of Anthropic! Meet Qwythos 9B, a Qwen3.5 based GGUF that's b…

X AI KOLs Timeline ↗ · 2026-06-28 Cached

Qwythos 9B is a new open-source, uncensored reasoning model based on Qwen3.5, offering GGUF quantizations, 1 million token context, vision, and function calling, with significant performance improvements over the base model.

0 favorites 0 likes
#quantized

huihui-ai/Huihui-GLM-5.2-abliterated-GGUF

Hugging Face Models Trending ↗ · 2026-06-28 Cached

A quantized GGUF version of the abliterated GLM-5.2 model is released on Hugging Face, enabling local inference with various tools like Transformers, llama.cpp, and vLLM.

0 favorites 0 likes
#quantized

@support_huihui: New GGUF: huihui-ai/Huihui-Qwythos-9B-Claude-Mythos-5-1M-abliterated-GGUF This is an uncensored version of empero-ai/Qw…

X AI KOLs Timeline ↗ · 2026-06-25 Cached

A new uncensored GGUF quantized version of the Qwythos-9B-Claude-Mythos-5-1M model, created using abliteration, is released on Hugging Face.

0 favorites 0 likes
#quantized

unsloth/Qwen-AgentWorld-35B-A3B-GGUF

Hugging Face Models Trending ↗ · 2026-06-24 Cached

Unsloth released a GGUF quantization of Qwen-AgentWorld-35B-A3B, a native language world model that simulates agentic environments across seven domains (MCP, Search, Terminal, SWE, Android, Web, OS) using long chain-of-thought reasoning and trained via CPT, SFT, and RL.

0 favorites 0 likes
#quantized

@antirez: Based on what I'm saying with GLM 5.2 implementation inside DwarfStar, there is 90% of probability I'll merge the branc…

X AI KOLs Following ↗ · 2026-06-24

Antirez announces high probability of merging a branch implementing GLM 5.2 in DwarfStar, which could become the best model for 512GB Mac Studio and potentially run on distributed 128GB MacBooks with 2-bit quantization.

0 favorites 0 likes
#quantized

nvidia/GLM-5.2-NVFP4

Hugging Face Models Trending ↗ · 2026-06-22 Cached

NVIDIA released GLM-5.2-NVFP4, a quantized version of ZAI's GLM-5.2 MoE language model optimized for inference on NVIDIA Blackwell GPUs using Model Optimizer.

0 favorites 0 likes
#quantized

nvidia/Qwen3.6-27B-NVFP4

Hugging Face Models Trending ↗ · 2026-06-22 Cached

NVIDIA released Qwen3.6-27B-NVFP4, a quantized version of Alibaba's Qwen3.6-27B model, optimized for deployment on NVIDIA GPUs with support for text, image, and video input.

0 favorites 0 likes
#quantized

PSA: unsloth/GLM-5.2-GGUF is uploading

Reddit r/LocalLLaMA ↗ · 2026-06-17 Cached

unsloth has uploaded a GGUF version of GLM-5.2 to Hugging Face, providing ready-to-use model files for various inference engines like llama.cpp, vLLM, and SGLang.

0 favorites 0 likes
#quantized

@DJLougen: Quants here https://huggingface.co/GestaltLabs/Ornstein-3.5-9B-V1.5-GGUF…

X AI KOLs Timeline ↗ · 2026-06-17 Cached

GestaltLabs releases Ornstein-3.5-9B-V1.5 GGUF quantizations, a reasoning-focused fine-tune of Qwen 3.5 9B with an MTP head and vision projector for multimodal use.

0 favorites 0 likes
#quantized

@WaleedAhmad1a10: Check out the Qwen 3.5 27B MoQ GGUFs :

X AI KOLs Following ↗ · 2026-06-16 Cached

A Hugging Face repository (kaitchup/Qwen3.6-27B-GGUF-MoQ) provides GGUF quantized weights for the Qwen3.6-27B MoQ model, enabling local inference with tools like llama.cpp and Ollama.

0 favorites 0 likes
#quantized

Jackrong/Qwopus3.6-27B-Coder-MTP-GGUF

Hugging Face Models Trending ↗ · 2026-06-11 Cached

A GGUF quantized version of the Qwopus3.6-27B-Coder-MTP model is released on Hugging Face, optimized for local inference and compatible with Transformers, vLLM, SGLang, and Unsloth Studio.

0 favorites 0 likes
#quantized

Holo3.1: Fast & Local Computer Use Agents

Hugging Face Blog ↗ · 2026-06-02 Cached

Holo3.1 is an updated computer-use model family that improves robustness across web, desktop, and mobile environments, introduces quantized checkpoints for local execution, and adds native support for function-calling protocols.

1 favorites 1 likes
Next →
← Back to home

Submit Feedback