unsloth

Tag

Cards List
#unsloth

@Sentdex: SITUATION DETECTED: Unsloth quants for GLM 5.2 are landing.

X AI KOLs Following ↗ · 2026-06-17 Cached

Unsloth quantizations for the GLM 5.2 model are being released.

0 favorites 0 likes
#unsloth

@h100envy: Daniel Han wrote Unsloth, the reason half of open-source can fine-tune a model on one GPU instead of a cluster. He didn…

X AI KOLs Timeline ↗ · 2026-06-17 Cached

Daniel Han built Unsloth, a tool that rewrites GPU kernels to make fine-tuning 2-3 times faster on a single GPU, enabling many open-source users to train models without a cluster.

0 favorites 0 likes
#unsloth

unsloth/Kimi-K2.7-Code-GGUF

Hugging Face Models Trending ↗ · 2026-06-12 Cached

Unsloth releases GGUF quantizations of Kimi K2.7 Code, a 1 trillion parameter MoE coding model built on Kimi K2.6 with improved token efficiency and agentic coding capabilities.

0 favorites 0 likes
#unsloth

Unsloth Minimax M3 GGUF

Reddit r/LocalLLaMA ↗ · 2026-06-12

Unsloth is uploading a GGUF quantized version of the MiniMax M3 model to Hugging Face.

0 favorites 0 likes
#unsloth

unsloth/MiniMax-M3-GGUF

Hugging Face Models Trending ↗ · 2026-06-12 Cached

Unsloth releases a GGUF quantized version of the MiniMax-M3 multimodal model, enabling image-text-to-text tasks with support for Transformers, llama.cpp, vLLM, and other inference engines.

0 favorites 0 likes
#unsloth

@Freerunnering: This actually makes Gemma 4 26B-4A usable for a coding agent @ 72tk/s on my MacBook Pro M1 Max. This video is realtime,…

X AI KOLs Timeline ↗ · 2026-06-12 Cached

Unsloth AI announces that Gemma 4 runs 2x faster with MTP GGUFs, making it feasible for local coding agents on hardware like a MacBook Pro M1 Max at 72 tokens/s.

0 favorites 0 likes
#unsloth

@VincentLogic: A 4.66 GB model actually runs at the level of a McKinsey consultant locally? Unsloth's latest 2-bit Gemma 4 12B is truly explosive. This isn't just chat – it directly transforms into a 'Super Agent' working autonomously: autonomously searching online citing 15+ sources, deeply distinguishing…

X AI KOLs Timeline ↗ · 2026-06-12 Cached

Unsloth releases a 2-bit quantized Gemma 4 12B model, only 4.66GB, runnable locally, with capabilities like autonomous online search and deep analysis similar to McKinsey consulting.

0 favorites 0 likes
#unsloth

@neural_avb: Lurking the Reasoning Training docs rn. Time to write a verifiers env and Unsloth/TRL that shit! Video soon if it all g…

X AI KOLs Timeline ↗ · 2026-06-11 Cached

The user is working on implementing reasoning training with verifiers using Unsloth and TRL, reporting progress on locally generating GRPO-like rollouts with a small SLM and a tiny RM, and promises a video soon.

0 favorites 0 likes
#unsloth

unsloth/diffusiongemma-26B-A4B-it-GGUF

Hugging Face Models Trending ↗ · 2026-06-10 Cached

Unsloth releases GGUF quantizations of Google DeepMind's DiffusionGemma (26B-A4B), a new block-diffusion architecture for faster text generation, ready for llama.cpp.

0 favorites 0 likes
#unsloth

Unsloth Gemma 4 QAT MTP assistant models now available

Reddit r/LocalLLaMA ↗ · 2026-06-09

Unsloth released Gemma 4 QAT MTP assistant models as GGUF files on Hugging Face, available in q8_0 and larger quantization formats.

0 favorites 0 likes
#unsloth

Qwen3.6-35B-A3B tool calling benchmark: ByteShape vs. Unsloth GGUFs, KV cache quants & long context performance

Reddit r/LocalLLaMA ↗ · 2026-06-08

A detailed benchmark comparing ByteShape and Unsloth quantizations of Qwen3.6-35B-A3B on tool calling performance, KV cache quantization effects, and long context degradation using llama.cpp and tool-eval-bench.

0 favorites 0 likes
#unsloth

Does it make sense to use alternative quantizations of QAT models? [D]

Reddit r/MachineLearning ↗ · 2026-06-06

A discussion on whether it is sensible to use alternative quantization methods on quantization-aware trained (QAT) models like Gemma-4, questioning if unsloth's benchmarks showing closer performance to QAT fine-tunes are beneficial or counterproductive.

0 favorites 0 likes
#unsloth

Unsloth just dropped MTP GGUF weights for Gemma 4!

Reddit r/LocalLLaMA ↗ · 2026-06-05

Unsloth has released Multi-Token Prediction (MTP) GGUF weights for Gemma 4 models (31B, 26B-A4B, 12B) in Q8, F16, and BF16 precisions, available on Hugging Face.

0 favorites 0 likes
#unsloth

unsloth/gemma-4-12B-it-qat-GGUF

Hugging Face Models Trending ↗ · 2026-06-05 Cached

Unsloth releases GGUF quantized versions of Google DeepMind's Gemma 4 models, optimized with Quantization-Aware Training (QAT) to reduce memory requirements while preserving quality, supporting multiple formats and sizes for diverse deployment.

0 favorites 0 likes
#unsloth

Unsloth on Apple Silicon- Pre-announcement announcement

Reddit r/LocalLLaMA ↗ · 2026-06-04

Unsloth, a popular LLM fine-tuning library, announces upcoming support for Apple Silicon devices, expanding its optimization capabilities beyond NVIDIA GPUs.

0 favorites 0 likes
#unsloth

@UnslothAI: We made a guide on using MCP with local LLMs. Connect Qwen3.6 and Gemma 4 for controlled access to tools, files, APIs, …

X AI KOLs Timeline ↗ · 2026-06-01 Cached

A step-by-step guide on using MCP servers with local LLMs like Qwen3.6 and Gemma 4 via Unsloth and llama.cpp, enabling private automated workflows with tools, files, and APIs.

0 favorites 0 likes
#unsloth

@neural_avb: Next video is on training tiny (<1B) models for preference tuning. Plus how to generate preference datasets with local …

X AI KOLs Timeline ↗ · 2026-05-26 Cached

Announces an upcoming video on training tiny models for preference tuning, covering reward models, RLHF, DPO, ORPO with Unsloth and TRL.

0 favorites 0 likes
#unsloth

@UnslothAI: 4-bit Qwen3.6 MTP GGUF managed to search 70+ sites from a single prompt. Try this locally on 20GB RAM via Unsloth Studi…

X AI KOLs Timeline ↗ · 2026-05-19 Cached

UnslothAI announces that its 4-bit Qwen3.6 MTP GGUF model can search over 70 websites from a single prompt, running locally on 20GB RAM via Unsloth Studio. The update adds automatic MTP and speculative decoding support.

0 favorites 0 likes
#unsloth

@populartourist: Unsloth Qwen3.6 27B Q6_K doing over 100 t/s with MTP on RTX 5090. That's coming up from 45-50 t/s without MTP. That's i…

X AI KOLs Timeline ↗ · 2026-05-16 Cached

Unsloth Qwen3.6 27B Q6_K achieves over 100 tokens per second with MTP on RTX 5090, up from 45-50 t/s without MTP.

0 favorites 0 likes
#unsloth

Who is your favourite quant publisher and why?

Reddit r/LocalLLaMA ↗ · 2026-05-13

A user shares their preference for Unsloth quantized models due to fast releases and low perplexity, compares them with Apex MoE quants, and asks the community for their favorite quant publisher.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback