gguf

Tag

Cards List
#gguf

PSA: unsloth/GLM-5.2-GGUF is uploading

Reddit r/LocalLLaMA · 2026-06-17 Cached

unsloth has uploaded a GGUF version of GLM-5.2 to Hugging Face, providing ready-to-use model files for various inference engines like llama.cpp, vLLM, and SGLang.

0 favorites 0 likes
#gguf

@DJLougen: Quants here https://huggingface.co/GestaltLabs/Ornstein-3.5-9B-V1.5-GGUF…

X AI KOLs Timeline · 2026-06-17 Cached

GestaltLabs releases Ornstein-3.5-9B-V1.5 GGUF quantizations, a reasoning-focused fine-tune of Qwen 3.5 9B with an MTP head and vision projector for multimodal use.

0 favorites 0 likes
#gguf

@Ali_TongyiLab: We are pleased to highlight an excellent community model from developer : Qwen3.6-27B-MTP-pi-reasoning-GGUF. Built on o…

X AI KOLs Timeline · 2026-06-17 Cached

Alibaba's Tongyi Lab highlights a community model, Qwen3.6-27B-MTP-pi-reasoning-GGUF, built on Qwen3.6-27B, optimized for automated programming and debugging workflows for local coding agents.

0 favorites 0 likes
#gguf

bartowski/command-a-plus-05-2026-GGUF · Hugging Face

Reddit r/LocalLLaMA · 2026-06-16 Cached

GGUF quantized versions of Cohere's command-a-plus-05-2026 model, optimized for llama.cpp and available in various quantization levels for local inference.

0 favorites 0 likes
#gguf

@WaleedAhmad1a10: Check out the Qwen 3.5 27B MoQ GGUFs :

X AI KOLs Following · 2026-06-16 Cached

A Hugging Face repository (kaitchup/Qwen3.6-27B-GGUF-MoQ) provides GGUF quantized weights for the Qwen3.6-27B MoQ model, enabling local inference with tools like llama.cpp and Ollama.

0 favorites 0 likes
#gguf

Nex-N2 Pro is the real deal

Reddit r/LocalLLaMA · 2026-06-16

The writer shares their experience with Nex-N2 Pro, originally mistaken as Rio-3.5, and finds it performs exceptionally well on coding benchmarks without hallucination, rivaling GPT-5.x on their Mac setup.

0 favorites 0 likes
#gguf

@Tono_Ken3: Added Q3 series to gemma-4-12B-coder-fable5-composer2.5-GGUF You might be able to try out the essence of Fable5 (as a t…

X AI KOLs Timeline · 2026-06-16 Cached

New Q3 quantizations added to the gemma-4-12B-coder-fable5-composer2.5 GGUF model, enabling the coding-focused fine-tune to run on GPUs with around 6GB VRAM using importance-matrix quantized versions.

0 favorites 0 likes
#gguf

Build a local AI coding agent from scratch

Reddit r/ArtificialInteligence · 2026-06-15 Cached

A step-by-step guide to building a minimal AI coding agent that runs entirely locally using llama.cpp, GGUF models, and a custom harness, demonstrating how to set up tools and call a model to execute real tasks like creating a landing page.

0 favorites 0 likes
#gguf

I made this android app which runs ai models locally

Reddit r/artificial · 2026-06-15

A developer created an Android app that runs AI models locally, supporting GGUF and LiteRT formats with multiple ways to add models.

0 favorites 0 likes
#gguf

Tower-Plus-72B-Ultra-Uncensored-Heretic, a Model That Support 22 Languages Making it Great for Multilingual Tasks and is Especially Strong on Translation Related Workflows Where No Censorship Is Essential, Now Ultra Uncensored With 5/100 Refusals!

Reddit r/LocalLLaMA · 2026-06-15 Cached

Tower-Plus-72B-Ultra-Uncensored-Heretic is a decensored version of Unbabel/Tower-Plus-72B, supporting 22 languages and excelling in translation tasks with minimal refusals.

0 favorites 0 likes
#gguf

@iotcoi: Microsoft just dropped FastContext-1.0: an open-source repo scout to lower your Copilot bill GGUF on HF. Run it locally…

X AI KOLs Timeline · 2026-06-15 Cached

Microsoft released FastContext-1.0, an open-source repo scout that runs locally with llama.cpp to reduce Copilot costs by scanning files and providing only relevant context to the main agent.

0 favorites 0 likes
#gguf

moar QAT stuff and hairy ticks

Reddit r/LocalLLaMA · 2026-06-15

The author releases improved GGUF quantized versions of Gemma 4 models (12B and 31B) using a more accurate quantization-aware training process that achieves lower KLD and higher same-top percentage than stock quantizations.

0 favorites 0 likes
#gguf

Command A Plus GGUFs posted

Reddit r/LocalLLaMA · 2026-06-15 Cached

Cohere has released GGUF quantized versions of its Command A+ model (25B active / 218B total parameters, Apache 2.0) for local inference, optimized for agentic and multilingual tasks.

0 favorites 0 likes
#gguf

@TraffAlex: Best Local LLMs for Consumer GPUs — llama.cpp Guide (June 2026) What I actually run on consumer hardware right now. Eve…

X AI KOLs Timeline · 2026-06-14 Cached

A guide to the best local LLMs for consumer GPUs as of June 2026, using llama.cpp to run models like Gemma 4-12B, Qwen3.6-27B, and Nex-N2-Mini on 8-32GB VRAM, with setup and launch commands.

0 favorites 0 likes
#gguf

You can run Deepseek 4 flash on mac (M3 Max, 96gb)

Reddit r/LocalLLaMA · 2026-06-14

A guide on running DeepSeek 4 flash on a Mac M3 Max with 96GB RAM using Antirez's ds4 engine and SSD streaming, achieving ~12 tokens/second inference speed.

0 favorites 0 likes
#gguf

@no_stp_on_snek: btw this was my loop. as you can see i didn't put much thought into it (typos and all), just a side thing to assess the…

X AI KOLs Following · 2026-06-14 Cached

Release of Qwopus3.6-27B-v2-MTP, a fine-tuned multi-token prediction reasoning model based on Qwen3.6-27B, optimized for coding, DevOps, and math tasks with improved generation speed.

0 favorites 0 likes
#gguf

unsloth/Kimi-K2.7-Code-GGUF

Hugging Face Models Trending · 2026-06-12 Cached

Unsloth releases GGUF quantizations of Kimi K2.7 Code, a 1 trillion parameter MoE coding model built on Kimi K2.6 with improved token efficiency and agentic coding capabilities.

0 favorites 0 likes
#gguf

Comparing dual-GPU inference speed between llama.cpp row/tensor split and ik_llama graph split

Reddit r/LocalLLaMA · 2026-06-12

A user benchmarks dual-GPU inference speed on two RTX 3080 20GB using llama.cpp (row/tensor split) and ik_llama (graph split) with a Qwen3.6-27B GGUF model, comparing token generation and prompt processing speeds.

0 favorites 0 likes
#gguf

@juanjucm: I'm seeing a lot of angry people lately... remember, you can always run your coding agent locally ;) llama.cpp + OpenCo…

X AI KOLs Following · 2026-06-12 Cached

Tweet reminding developers they can run coding agents locally using llama.cpp and OpenCode for fast, reliable, and private inference, demonstrating with UnslothAI's North-Mini-Code-1.0-GGUF model.

0 favorites 0 likes
#gguf

Unsloth Minimax M3 GGUF

Reddit r/LocalLLaMA · 2026-06-12

Unsloth is uploading a GGUF quantized version of the MiniMax M3 model to Hugging Face.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback