gguf

Tag

Cards List
#gguf

Update your chat template for dsv4 if you're using llama.cpp

Reddit r/LocalLLaMA · 2026-07-28

Recent llama.cpp commits broke preserve_thinking behavior for older DeepSeek V4 gguf chat templates, causing issues in coding agent contexts. The fix is to override the gguf template with a new one using --chat-template-file.

0 favorites 0 likes
#gguf

Kimi K3 text-only for llama.cpp

Reddit r/LocalLLaMA · 2026-07-27 Cached

Kimi K3 text-only model is now supported in llama.cpp, enabling local inference of this open-source LLM using the C++ inference engine.

0 favorites 0 likes
#gguf

Ling-3.0-flash weights: SGLang says day-0, vLLM says when they land, llama.cpp closed the 2.6 request as not_planned

Reddit r/LocalLLaMA · 2026-07-27

The article details the current support status for the Ling-3.0-flash model weights across inference engines: SGLang commits to day-0 integration, vLLM awaits open weights, and llama.ccp lacks conversion for the Bailing MoE variant. It notes that the release pattern involves a free API window followed by open-sourcing, as seen with Ling-2.6-flash.

0 favorites 0 likes
#gguf

Do Qwen 3.6 27B quantizations break the pelican?

Reddit r/LocalLLaMA · 2026-07-27 Cached

The article evaluates how different quantizations of the Qwen3.6-27B model affect output quality using KL divergence and top-1 token accuracy, as well as visual examples like SVG drawings.

0 favorites 0 likes
#gguf

A quick coding capability test:4 Qwen 3.6-35B GGUF Variants

Reddit r/LocalLLaMA · 2026-07-27

A quick evaluation of coding performance across four GGUF variants of the Qwen 3.6-35B model.

0 favorites 0 likes
#gguf

whats happening on llama.cpp

Reddit r/LocalLLaMA · 2026-07-27

A significant update to llama.cpp requires all previously generated GGUF files to be regenerated, indicating a major breaking change to the model format.

0 favorites 0 likes
#gguf

[audio.cpp] Release 0.4: Higgs Audio v3 TTS 4B (10x real time)+ Fish Audio S2 Pro in C++/GGML, full GGUF loading, Q8 speed and VRAM gains

Reddit r/LocalLLaMA · 2026-07-24

Release 0.4 of audio.cpp adds C++/GGML inference for Higgs Audio v3 TTS 4B (10x real-time) and Fish Audio S2 Pro, with full GGUF loading and Q8 speed/VRAM gains.

0 favorites 0 likes
#gguf

Laguna-S-2.1 "thinking forever" loops seem to be a quantization artifact

Reddit r/LocalLLaMA · 2026-07-23

The article reports that infinite thinking loops in Laguna S 2.1 AI model are likely caused by quantization artifacts. Switching to an MoE-aware APEX quant (e.g., Myric/Laguna-S-2.1-APEX-GGUF) and using default sampling settings (temp 0.7, top_p 0.95, top_k 20) resolved the looping for most cases. Additionally, framing prompts around tool calls can prevent overthinking.

0 favorites 0 likes
#gguf

PSA on Laguna S-2.1 - Use the updated chat template and GGUF

Reddit r/LocalLLaMA · 2026-07-23

Laguna S-2.1 model has been updated with a fix for yarn_attn_factor (corrected to 1.0) and an improved chat template that fixes broken thinking, preserves thinking, and enables tool calling. Users are advised to use the updated GGUF from the official repo.

0 favorites 0 likes
#gguf

Laguna-S-2.1 runs on my 6 years old gaming PC!

Reddit r/LocalLLaMA · 2026-07-22

Laguna-S-2.1, a quantized AI model, runs on a 6-year-old gaming PC with an RTX 3080, achieving 10 t/s decode and using 8.3 GB VRAM and 52.2 GB host RAM.

0 favorites 0 likes
#gguf

Unsloth Quantization of Laguna S 2.1 Is Out

Reddit r/LocalLLaMA · 2026-07-22 Cached

Unsloth released GGUF quantizations of the Laguna S 2.1 Mixture-of-Experts model, a 118B parameter coding model with 8B active parameters, 1M context window, and agentic capabilities. The quantized versions enable efficient local deployment.

0 favorites 0 likes
#gguf

Models to get before any political disruption

Reddit r/LocalLLaMA · 2026-07-21

A curated list of recommended AI models (e.g., Qwen3.5, Gemma 4, Mistral Medium) for systems with up to 256GB unified RAM, covering creative, coding, agentic, and general chat use cases.

0 favorites 0 likes
#gguf

@QuixiAI: QuixiAI/embeddinggemma.c - fast as fuck cross-platform embedding inference engine. Written in c, runs anywhere. Inferen…

X AI KOLs Following · 2026-07-20 Cached

QuixiAI releases embeddinggemma.c, a fast cross-platform embedding inference engine written in C, supporting multiple backends (CPU, Metal, CUDA, ROCm, SYCL) and Matryoshka embeddings with a standard HTTP API.

0 favorites 0 likes
#gguf

@Alacritic_Super: Need a powerful local LLM without huge hardware? Meet Qwythos-9B. • 9B parameters • 1M-token context • Native tool call…

X AI KOLs Timeline · 2026-07-20 Cached

Qwythos-9B is a 9B-parameter local LLM with 1M-token context and native tool calling, available in GGUF format for Ollama, llama.cpp, and LM Studio. It significantly outperforms its base model Qwen3.5-9B on key benchmarks.

0 favorites 0 likes
#gguf

DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF

Hugging Face Models Trending · 2026-07-19 Cached

A community fine-tuned and merged version of Qwen 3.5 9B that claims to exceed critical benchmarks of larger models, available in GGUF format with multi-token prediction support.

0 favorites 0 likes
#gguf

DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF

Hugging Face Models Trending · 2026-07-17 Cached

DavidAU releases Qwen3.6-27B-Fable-Fusion-711, a multi-stage fine-tune of Qwen 3.6 27B that claims to exceed 700 ARC-C benchmark, surpassing base models and matching closed-source models, available in GGUF format for consumer hardware.

0 favorites 0 likes
#gguf

@TheAhmadOsman: RTX 3090 owners tonight will be running Kimi_K3_3T_Q_0.001_K GGUF

X AI KOLs Following · 2026-07-16 Cached

A quantized GGUF version of the Kimi K3 model is now available, optimized for running on RTX 3090 GPUs.

0 favorites 0 likes
#gguf

@MaziyarPanahi: I played Thinking Machines' new Inkling a 2-minute doctor's visit, on my Mac Studio The patient came in about her knee.…

X AI KOLs Timeline · 2026-07-16 Cached

Thinking Machines' Inkling, a 975B parameter model running locally on a Mac Studio via llama.cpp, listened to a doctor's visit audio and accurately diagnosed heart failure from subtle cues in small talk, demonstrating advanced medical reasoning without leaving the machine.

0 favorites 0 likes
#gguf

@dealignai: Bonsai-27b CRACK’d GGUF (refusals removed), here u go gguf peoples https://huggingface.co/dealignai/Bonsai-27b-Ternary-…

X AI KOLs Timeline · 2026-07-15

A modified GGUF version of Bonsai-27b with refusals removed has been released on Hugging Face.

0 favorites 0 likes
#gguf

Qwen 3.5 122B Heretic ROCmFP4 iMatrix

Reddit r/LocalLLaMA · 2026-07-15 Cached

A compact, importance-calibrated ROCmFP4 quantization of Qwen 3.5 122B model for high-memory AMD systems, achieving improved quality (14% lower KLD) and performance (28.45 tok/s). Requires ROCmFPX runtime; not compatible with stock llama.cpp.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback