gguf

Tag

Cards List
#gguf

Deepseek V4 Flash running on RTX 5090 MoE

Reddit r/LocalLLaMA · 2026-07-03

User shares optimization benchmarks for DeepSeek-V4-Flash (Q2_K) running on an RTX 5090 using a fork of llama.cpp, achieving 21.3 tokens/s generation and 1 million context size.

0 favorites 0 likes
#gguf

GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-Thinking-GGUF

Hugging Face Models Trending · 2026-07-03 Cached

A GGUF quantized version of MiniCPM5-1B-Claude-Opus-Fable5-Thinking model is released on Hugging Face, with usage instructions for llama.cpp, vLLM, and Ollama.

0 favorites 0 likes
#gguf

I’m switching to Linux, is Ubuntu the most compatible with local AI?

Reddit r/LocalLLaMA · 2026-07-02

A user asks about Ubuntu's compatibility for local AI tools like vLLM, llama.cpp, and ComfyUI when switching to Linux.

0 favorites 0 likes
#gguf

@Jolyne_AI: Another open-source tool for running AI models locally on GitHub: Shimmy, targeting Ollama's pain points. A single file of only 5MB provides fast, stable local inference with a full OpenAI-compatible API, almost zero integration cost. Built with Rust to maximize performance, it starts in under 100ms and uses about 50MB of memory.

X AI KOLs Timeline · 2026-07-02 Cached

Another open-source tool on GitHub, Shimmy, is a single 5MB file written in Rust that provides fast and stable local inference with a full OpenAI-compatible API, targeting Ollama's pain points. It starts in under 100ms and uses about 50MB of memory.

0 favorites 0 likes
#gguf

Deepseek V4 Flash 2, 3 and 4 bits GGUFs

Reddit r/LocalLLaMA · 2026-07-01 Cached

GGUF quantizations of DeepSeek V4 Flash in 2-bit, 3-bit, and 4-bit precisions, made available on Hugging Face for local inference with tools like llama.cpp and Ollama.

0 favorites 0 likes
#gguf

bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF

Hugging Face Models Trending · 2026-07-01 Cached

ThinkingCap-Qwen3.6-27B is a fine-tuned version of Qwen3.6-27B that uses 50% fewer thinking tokens on average while maintaining answer quality. This repository provides GGUF quantizations for local inference with llama.cpp.

0 favorites 0 likes
#gguf

Bartowski has delivered DS4 GGUF

Reddit r/LocalLLaMA · 2026-06-30

Bartowski has released a GGUF quantized version of DeepSeek-V4-Flash, inviting comparison with Antirez's version.

0 favorites 0 likes
#gguf

@nullfoundry: hey everyone. i'd like to share my new recipe for dflash ( merged yesterday on oficial llama.cpp ) llama-server -hf uns…

X AI KOLs Timeline · 2026-06-29 Cached

Sharing a new recipe for dflash speculative decoding in llama.cpp, achieving ~70 TPS on a single RTX 3090 using Qwen3.6-27B GGUF with a draft model.

0 favorites 0 likes
#gguf

@loktar00: Not sure how Kyle doesn't have 10's of thousands of followers. He's putting out incredibly cool usable stuff what seems…

X AI KOLs Timeline · 2026-06-29 Cached

Kyle Hessling announces the release of Qwopus-3.6-35B-A3B-MTP-Coder, a fast MOE model specialized for coding, available in GGUF format.

0 favorites 0 likes
#gguf

Ornith-1.0-35B GGUF update: native MTP speculative-decode graft + full serving/TTFT/long-context numbers (llama.cpp, tp=1)

Reddit r/LocalLLaMA · 2026-06-28

An update on the Ornith-1.0-35B GGUF model introduces a native MTP speculative-decode graft for faster inference on a single GPU, achieving ~1.3-1.35x decode speedup while maintaining near-identical token distribution. Benchmark numbers for throughput, TTFT, and long-context performance across multiple quants are provided.

0 favorites 0 likes
#gguf

@BrianRoemmele: BOOM! Meet the open source Cambrian Explosion of repulsion of Anthropic! Meet Qwythos 9B, a Qwen3.5 based GGUF that's b…

X AI KOLs Timeline · 2026-06-28 Cached

Qwythos 9B is a new open-source, uncensored reasoning model based on Qwen3.5, offering GGUF quantizations, 1 million token context, vision, and function calling, with significant performance improvements over the base model.

0 favorites 0 likes
#gguf

@MiaAI_lab: If you're looking for a simple start/stop script to run your Qwen3.6 27B/35B, check this out. It's optimized for speed …

X AI KOLs Timeline · 2026-06-28 Cached

MiaAI-Lab provides a simple Bash start/stop script for running Qwen3.6 27B/35B GGUF models via llama-server, optimized for speed and coding performance.

0 favorites 0 likes
#gguf

huihui-ai/Huihui-GLM-5.2-abliterated-GGUF

Hugging Face Models Trending · 2026-06-28 Cached

A quantized GGUF version of the abliterated GLM-5.2 model is released on Hugging Face, enabling local inference with various tools like Transformers, llama.cpp, and vLLM.

0 favorites 0 likes
#gguf

Is Qwen3-VL-2B the only viable VLM for JSON extraction on a "potato"?

Reddit r/LocalLLaMA · 2026-06-28

The author claims Qwen3-VL-2B is the only viable vision-language model for JSON extraction on low-end hardware, outperforming larger models like Qwen3-VL-4B, yet it is absent from major benchmarks.

0 favorites 0 likes
#gguf

@HuggingModels: Meet Gemma 4 12B Agentic Fable5: a locally run GGUF model that thinks, reasons, and uses tools like a pro. It's built f…

X AI KOLs Timeline · 2026-06-28 Cached

Meet Gemma 4 12B Agentic Fable5, a locally-run GGUF model designed for coding, terminal tasks, and agentic workflows, with 206k downloads.

0 favorites 0 likes
#gguf

@SlimTradeyBaby: Attention all 8-12GB GPU users! This new Ornith-1.0-9B is looking like it will be a serz player for smaller VRAM setups…

X AI KOLs Timeline · 2026-06-26 Cached

Ornith-1.0-9B is a new 9B parameter AI model optimized for 8-12GB GPUs, achieving strong performance on agentic coding benchmarks, matching or surpassing models 2-3x its size.

0 favorites 0 likes
#gguf

@SlimTradeyBaby: Just fired up Ornith 35B Q4 on the 5090 remotely… 2329 prompt / 195 gen tok/s and rock solid at 32k. Quick test only fu…

X AI KOLs Timeline · 2026-06-26 Cached

DeepReinforce AI releases Ornith-1.0, a self-improving open-source model family for agentic coding, including a 35B MoE variant that achieves state-of-the-art performance on coding benchmarks and runs efficiently on single GPUs like the 5090.

0 favorites 0 likes
#gguf

@support_huihui: New GGUF: huihui-ai/Huihui-Qwythos-9B-Claude-Mythos-5-1M-abliterated-GGUF This is an uncensored version of empero-ai/Qw…

X AI KOLs Timeline · 2026-06-25 Cached

A new uncensored GGUF quantized version of the Qwythos-9B-Claude-Mythos-5-1M model, created using abliteration, is released on Hugging Face.

0 favorites 0 likes
#gguf

GLM 5.2 on consumer hardware

Reddit r/LocalLLaMA · 2026-06-25

A user tested the unsloth quantized GLM-5.2 model on a high-end consumer-like system with dual RTX 5090, achieving 12 tokens per second.

0 favorites 0 likes
#gguf

deepreinforce-ai/Ornith-1.0-35B-GGUF

Hugging Face Models Trending · 2026-06-25 Cached

deepreinforce-ai releases Ornith-1.0-35B-GGUF, a state-of-the-art open-source coding agent model that uses self-improving reinforcement learning to jointly optimize scaffold and solution generation, achieving SOTA performance on coding benchmarks.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback