unsloth

Tag

Cards List
#unsloth

@no_stp_on_snek: Ooh nice. This one gonna be fun

X AI KOLs Timeline ↗ · 2026-07-31 Cached

A reply to Unsloth AI expressing excitement about the rumored DeepSeek-V4-Flash model and the possibility of running it locally via quantized versions.

0 favorites 0 likes
#unsloth

unsloth/DeepSeek-V4-Flash-0731-GGUF

Reddit r/LocalLLaMA ↗ · 2026-07-31

Unsloth teases the upcoming release of DeepSeek V4 Flash GGUF quantized model on Hugging Face.

0 favorites 0 likes
#unsloth

Unsloth Quantization of Laguna S 2.1 Is Out

Reddit r/LocalLLaMA ↗ · 2026-07-22 Cached

Unsloth released GGUF quantizations of the Laguna S 2.1 Mixture-of-Experts model, a 118B parameter coding model with 8B active parameters, 1M context window, and agentic capabilities. The quantized versions enable efficient local deployment.

0 favorites 0 likes
#unsloth

I benchmarked Unsloth's Qwen3.6-27B NVFP4 on 1x/2x 5090s. MTP is great until it really isn't.

Reddit r/LocalLLaMA ↗ · 2026-07-21

A detailed benchmark of Unsloth's Qwen3.6-27B NVFP4 model on RTX 5090 GPUs, showing MTP (multi-token prediction) gives large speedups for single requests at short context but becomes detrimental under batch concurrency or long contexts.

0 favorites 0 likes
#unsloth

Unsloth now supports AMD!

Reddit r/LocalLLaMA ↗ · 2026-07-20

Unsloth, the efficient LLM fine-tuning library, now supports AMD GPUs, expanding hardware accessibility.

0 favorites 0 likes
#unsloth

The qlora 2e-4 default is wrong under 10k samples and nobody talks about it [D]

Reddit r/MachineLearning ↗ · 2026-07-16

The author argues that the commonly recommended learning rate of 2e-4 for QLoRA fine-tuning is too high for datasets under 10k samples, leading to overfitting and poor evaluation, and suggests using a lower learning rate like 1e-4.

0 favorites 0 likes
#unsloth

unsloth/inkling-GGUF

Hugging Face Models Trending ↗ · 2026-07-14 Cached

The unsloth/inkling-GGUF page provides the quantized GGUF version of Inkling, a 975B-parameter multimodal MoE model (41B active) from Thinking Machines, designed for text, image, and audio inputs with open weights and support for local deployment via libraries like Unsloth, SGLang, and vLLM.

0 favorites 0 likes
#unsloth

@UnslothAI: We collaborated with AWS on a complete guide to LLM Quantization and Deployment. Learn about: • Model formats, dynamic …

X AI KOLs Timeline ↗ · 2026-07-13 Cached

Unsloth collaborated with AWS on a comprehensive guide to dynamic quantization and deployment of LLMs on Amazon SageMaker, covering model formats, tools, and best practices.

0 favorites 0 likes
#unsloth

@sakurayukiai: Counting the draft model's KV cache in bytes instead of hiding it inside a flat VRAM cushion is how Unsloth pushed Qwen…

X AI KOLs Timeline ↗ · 2026-07-13

Unsloth improved Qwen3.6-27B Q6_K context length from 23K to 64K on a single 32GB card by accurately counting the draft model's KV cache in bytes instead of using a flat VRAM cushion.

0 favorites 0 likes
#unsloth

Voodoo Quant beats Unsloth Dynamic 2.0 KLD by 95% in Qwen3.5 0.8B and 2B

Reddit r/LocalLLaMA ↗ · 2026-07-12

Voodoo Quant, a quantization method, outperforms Unsloth Dynamic 2.0 KLD by 95% on Qwen3.5 0.8B and 2B models.

0 favorites 0 likes
#unsloth

Benchmark of the new unsloth/Qwen3.6-27B-NVFP4 on 4x 5060 ti's with P2P and PP=4 at 1,4,8,12, and 16 concurrency.

Reddit r/LocalLLaMA ↗ · 2026-07-11

Benchmark results of the unsloth/Qwen3.6-27B-NVFP4 model running on 4x RTX 5060 Ti GPUs with peer-to-peer and pipeline parallelism at various concurrency levels.

0 favorites 0 likes
#unsloth

@0xkeenz: Interesting, I was just studying the differences between Unsloth and NVIDIA's Qwen3.6 27B NVFP4 yesterday, and today Unsloth updated! The new Unsloth's quantization approach is very similar to NVIDIA's official solution: instead of choosing between BF16 and ...

X AI KOLs Timeline ↗ · 2026-07-10 Cached

Unsloth releases a new version of Qwen3.6 27B NVFP4 quantization scheme, introducing FP8_E4M3 intermediate precision layer and refined weight protection, achieving 2.5x speed improvement on 24GB VRAM, while improving accuracy and tool-calling capabilities.

0 favorites 0 likes
#unsloth

@no_stp_on_snek: This is huge. 3090 class gpus rejoice!

X AI KOLs Following ↗ · 2026-07-10 Cached

Unsloth AI releases quantized Qwen3.6 models that run 2.5× faster on consumer GPUs, with the 27B model fitting in 24GB VRAM and the 35B-A3B achieving high throughput.

1 favorites 1 likes
#unsloth

2.5x faster Qwen3.6 NVFP4 Unsloth quants

Reddit r/LocalLLaMA ↗ · 2026-07-10

Unsloth releases quantized Qwen3.6 models using NVFP4 format, achieving 2.5x faster inference speeds.

0 favorites 0 likes
#unsloth

Unsloth GLM-5.2 – How to Run Locally

Hacker News Top ↗ · 2026-06-22 Cached

A guide on running Z.ai's open model GLM-5.2 locally using Unsloth Dynamic GGUFs. The model features 744B total parameters (40B active) and a 1M context window, with quantized versions reducing memory to 239GB for 2-bit, enabling local inference on 256GB Macs.

0 favorites 0 likes
#unsloth

Good results fine tuning a local LLM like Qwen 3:0.6B to categorize questions

Hacker News Top ↗ · 2026-06-21 Cached

A developer fine-tunes a small Qwen 3 0.6B model using the Unsloth framework to categorize household questions, achieving good results with only 850 training examples.

0 favorites 0 likes
#unsloth

@SlimTradeyBaby: Drop your GPU below and I’ll tell you exactly what model and config to run on it. JOKES. No need. Qwen 3.6 27b @Unsloth…

X AI KOLs Timeline ↗ · 2026-06-20 Cached

A tweet promoting the Qwen 3.6 27b model and recommending UnslothAI for running it on any GPU.

0 favorites 0 likes
#unsloth

@10xmylife: Unsloth 成功将 2-bit 版本的 GLM-5.2 部署在了 256GB 的 Mac 上

X AI KOLs Following ↗ · 2026-06-19 Cached

Unsloth 成功将 GLM-5.2 模型以 2-bit 量化压缩至 238GB,可在 256GB Mac 上本地运行,保留约 82% 的准确率。

0 favorites 0 likes
#unsloth

@UnslothAI: GLM-5.2 can now be run locally! The 2-bit model retains ~82% accuracy after we shrunk it from 1.51TB to 238GB (-84% siz…

X AI KOLs Timeline ↗ · 2026-06-18 Cached

UnslothAI announces GLM-5.2, Z.ai's strongest open model with 744B parameters, now runnable locally via dynamic GGUF quantization reducing size by ~84% to 239GB while retaining ~82% accuracy. It fits on 256GB Macs and supports long-context, reasoning, and agentic tasks.

0 favorites 0 likes
#unsloth

@aisearchio: GLM 5.2 GGUF is already here! 8-bit is ~half the size of the full model. Smaller versions coming soon https://huggingfa…

X AI KOLs Timeline ↗ · 2026-06-17 Cached

GLM 5.2 GGUF quantized model is released, with 8-bit version half the size of the full model; smaller versions are coming soon.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback