Tag
Unsloth has released GGUF quantized versions of Qwen-Image-2.1, enabling it to run locally on 12GB VRAM with performance comparable to Nano Banana 2.0.
Unsloth AI announces optimizations for GLM-5.3-Flash, enabling 1.6–3.4× faster local GGUF inference with multi-token prediction and hardware requirements for running models locally.
A user reports that the Qwen3.8 27B model hallucinated and implemented an unintended feature during a task, despite careful planning and good prior performance.
Unsloth announces day 0 support for the newly released Qwen 3.8 Flash AI model, prompting users to prepare disk space.
A user shares their positive experience with low quantizations of Qwen 27B 3.8 on a Mac mini M4, using Unsloth's Q3 XXS quant, and asks for others' experiences with sub-Q3 quants.
This article provides a local deployment guide for the Qwen3.8-27B model, recommends using Q4 quantization and the Unsloth GGUF tool, and shares performance test results compared to the FP8 benchmark.
A user tested the unsloth 1-bit quantized version of the Qwen 3.8 27B AI model on an 8GB VRAM system and found the results amusing.
Unsloth has released Dynamic v3.0 GGUFs for Qwen3.8 models, offering >10% better accuracy at the same size through improved quantization techniques and calibration methods.
Unsloth has released an NVFP4 quantized version of the Qwen3.8-27B AI model, which offers enhanced capabilities in coding, professional work, agentic tasks, and native vision-language understanding.
Unsloth announces that Qwen3.8 can now run locally, shrinking the 2.4T-parameter model from 4.9TB to 397GB via Dynamic 1-bit quantization, with a guide and GGUF release.
Unsloth Desktop is a new open-source desktop app for running and training models locally on Mac, Windows, and Linux, with support for connecting Claude Code and Codex to local LLMs.
A user shares local testing of Muse Glimmer (Q4 quant via Unsloth) on llama.cpp with OpenCode, noting it performs below Qwen3.6 27B but had reliable tool calls.
Unsloth releases a GGUF-quantized version of Meta's Muse Glimmer 30B model, designed for local agentic tasks with multimodal input, tool use, and multi-step reasoning.
Unsloth releases new GGUF quantizations of Kimi K3, ranging from 466GB to 649GB, enabling efficient deployment of the large model.
Unsloth releases GGUF quantizations of MiniMax-H3, an omni-modal generative system for video with native stereo audio, enabling local execution via sd-cli and other platforms.
Unsloth AI announces DSpark, enabling DeepSeek-V4-Flash GGUF models to run ~1.4–2× faster locally, reaching 120 tokens/s with no accuracy change.
Daniel Han of Unsloth validates that Qwen3.8-27B will run in only 17GB VRAM, making it accessible for local inference.
Alibaba announces Qwen3.8-27B open-weights release, capable of running locally on 17GB RAM/VRAM, alongside the larger Qwen3.8-Max.
Unsloth releases an IQ3 GGUF quantization of DeepSeek-V4-Flash-0731, enabling local inference via llama.cpp, Ollama, LM Studio, and other tools.
DeepSeek-V4-Flash-0731 is shown running as an unsloth GGUF quant on a single 40GB A100, with 17.7 tok/s and 6 experts loaded into VRAM, enabling a full agentic coding loop.