unsloth

Tag

Cards List
#unsloth

@UnslothAI: Qwen-Image-2.1 can now run locally on 12GB VRAM with Unsloth GGUFs! The 7B model performs on par with Nano Banana 2.0. …

X AI KOLs Timeline ↗ · 3d ago Cached

Unsloth has released GGUF quantized versions of Qwen-Image-2.1, enabling it to run locally on 12GB VRAM with performance comparable to Nano Banana 2.0.

0 favorites 0 likes
#unsloth

@UnslothAI: We made GLM-5.3-Flash run 3.3x faster locally! Local GGUF inference is now 1.6–3.4× faster with optimized decoding and …

X AI KOLs Following ↗ · 2026-09-04 Cached

Unsloth AI announces optimizations for GLM-5.3-Flash, enabling 1.6–3.4× faster local GGUF inference with multi-token prediction and hardware requirements for running models locally.

0 favorites 0 likes
#unsloth

Qwen3.8 27B Q8 hallucinated entire plan???

Reddit r/LocalLLaMA ↗ · 2026-09-03

A user reports that the Qwen3.8 27B model hallucinated and implemented an unintended feature during a task, despite careful planning and good prior performance.

0 favorites 0 likes
#unsloth

Qwen 3.8 Flash Next day 0 support from unsloth

Reddit r/LocalLLaMA ↗ · 2026-08-25

Unsloth announces day 0 support for the newly released Qwen 3.8 Flash AI model, prompting users to prepare disk space.

0 favorites 0 likes
#unsloth

Qwen 27B 3.8 quants: How low can you go?

Reddit r/LocalLLaMA ↗ · 2026-08-24

A user shares their positive experience with low quantizations of Qwen 27B 3.8 on a Mac mini M4, using Unsloth's Q3 XXS quant, and asks for others' experiences with sub-Q3 quants.

0 favorites 0 likes
#unsloth

@9hills: Qwen3.8-27B Local Deployment Guide 1. Q4 has basically no quality loss, can even use Q3 2. Use Unsloth's GGUF. 3. Turn on low thinking. 4. Enable dflash2

X AI KOLs Timeline ↗ · 2026-08-21 Cached

This article provides a local deployment guide for the Qwen3.8-27B model, recommends using Q4 quantization and the Unsloth GGUF tool, and shares performance test results compared to the FP8 benchmark.

0 favorites 0 likes
#unsloth

Ladies and gentlemen I present to you Qwen3.8 27b 1bit brain damage quant

Reddit r/LocalLLaMA ↗ · 2026-08-20

A user tested the unsloth 1-bit quantized version of the Qwen 3.8 27B AI model on an 8GB VRAM system and found the results amusing.

0 favorites 0 likes
#unsloth

Unsloth Dynamic 3.0 GGUFs

Hacker News Top ↗ · 2026-08-19 Cached

Unsloth has released Dynamic v3.0 GGUFs for Qwen3.8 models, offering >10% better accuracy at the same size through improved quantization techniques and calibration methods.

0 favorites 0 likes
#unsloth

unsloth/Qwen3.8-27B-NVFP4

Hugging Face Models Trending ↗ · 2026-08-13 Cached

Unsloth has released an NVFP4 quantized version of the Qwen3.8-27B AI model, which offers enhanced capabilities in coding, professional work, agentic tasks, and native vision-language understanding.

0 favorites 0 likes
#unsloth

@UnslothAI: Qwen3.8 can now be run locally! We shrank Qwen3.8-2.4T-A95B from 4.9TB to 397GB (-91% size) via Dynamic 1-bit by select…

X AI KOLs Following ↗ · 2026-08-12 Cached

Unsloth announces that Qwen3.8 can now run locally, shrinking the 2.4T-parameter model from 4.9TB to 397GB via Dynamic 1-bit quantization, with a guide and GGUF release.

0 favorites 0 likes
#unsloth

@dessaigne: Unsloth is one of my YC companies and the team is cracked. Their new open source desktop app lets you run and train mod…

X AI KOLs Timeline ↗ · 2026-08-11 Cached

Unsloth Desktop is a new open-source desktop app for running and training models locally on Mac, Windows, and Linux, with support for connecting Claude Code and Codex to local LLMs.

0 favorites 0 likes
#unsloth

Tested Muse Glimmer locally on coding with OpenCode & agentic work

Reddit r/LocalLLaMA ↗ · 2026-08-10

A user shares local testing of Muse Glimmer (Q4 quant via Unsloth) on llama.cpp with OpenCode, noting it performs below Qwen3.6 27B but had reliable tool calls.

0 favorites 0 likes
#unsloth

unsloth/Muse-Glimmer-30B-GGUF · Hugging Face

Reddit r/LocalLLaMA ↗ · 2026-08-10 Cached

Unsloth releases a GGUF-quantized version of Meta's Muse Glimmer 30B model, designed for local agentic tasks with multimodal input, tool use, and multi-step reasoning.

0 favorites 0 likes
#unsloth

New Unsloth KImi K3 drops! Q1_0 (466GB), TQ1_0(509GB), IQ1_M(649),TQ2_0(551GB)!!

Reddit r/LocalLLaMA ↗ · 2026-08-07

Unsloth releases new GGUF quantizations of Kimi K3, ranging from 466GB to 649GB, enabling efficient deployment of the large model.

0 favorites 0 likes
#unsloth

unsloth/MiniMax-H3-GGUF

Hugging Face Models Trending ↗ · 2026-08-07 Cached

Unsloth releases GGUF quantizations of MiniMax-H3, an omni-modal generative system for video with native stereo audio, enabling local execution via sd-cli and other platforms.

0 favorites 0 likes
#unsloth

@YRSM_Simon: 120 t/s ! Good job, @UnslothAI

X AI KOLs Following ↗ · 2026-08-06 Cached

Unsloth AI announces DSpark, enabling DeepSeek-V4-Flash GGUF models to run ~1.4–2× faster locally, reaching 120 tokens/s with no accuracy change.

0 favorites 0 likes
#unsloth

Daniel Han of Unsloth validates Qwen3.8-27B will run only 17GB VRAM

Reddit r/LocalLLaMA ↗ · 2026-08-03

Daniel Han of Unsloth validates that Qwen3.8-27B will run in only 17GB VRAM, making it accessible for local inference.

0 favorites 0 likes
#unsloth

@UnslothAI: Qwen3.8-27B is coming! Will run locally on 17GB RAM/VRAM setups.

X AI KOLs Timeline ↗ · 2026-08-03 Cached

Alibaba announces Qwen3.8-27B open-weights release, capable of running locally on 17GB RAM/VRAM, alongside the larger Qwen3.8-Max.

0 favorites 0 likes
#unsloth

IQ3 DS out

Reddit r/LocalLLaMA ↗ · 2026-07-31 Cached

Unsloth releases an IQ3 GGUF quantization of DeepSeek-V4-Flash-0731, enabling local inference via llama.cpp, Ollama, LM Studio, and other tools.

0 favorites 0 likes
#unsloth

DeepSeek-V4-Flash-0731 unsloth gguf on A100

Reddit r/LocalLLaMA ↗ · 2026-07-31

DeepSeek-V4-Flash-0731 is shown running as an unsloth GGUF quant on a single 40GB A100, with 17.7 tok/s and 6 experts loaded into VRAM, enabling a full agentic coding loop.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback