quantized-models

Tag

Cards List
#quantized-models

unsloth/Qwen3.6-27B-NVFP4 vs. Intel/Qwen3.6-27B-int4-AutoRound vs. nvidia/Qwen3.6-27B-NVFP4 -- which one to choose?

Reddit r/LocalLLaMA · 2026-07-30

The article compares three quantized variants of Qwen3.6-27B (NVFP4 from Unsloth and Nvidia, int4-AutoRound from Intel) and requests benchmarks and hallucination data from the community.

0 favorites 0 likes
#quantized-models

Calibrating 2-bit GGUFs (<10Gb) for agentic coding tasks

Reddit r/LocalLLaMA · 2026-06-18

This article introduces calibrated 2-bit GGUF quantizations of the Qwopus3.6-27B-Coder model for agentic coding tasks, demonstrating that the IQ2_M quant (9.74 GiB) achieves a 63% pass rate on the SWE-rebench benchmark, comparable to a Q5_K_M quant at half the size.

0 favorites 0 likes
#quantized-models

How can Deepseek v4 top the coding leaderboards and still sit 8 months behind the frontier?

Reddit r/LocalLLaMA · 2026-06-11

Analysis of DeepSeek V4's top coding scores versus its reported 8-month gap behind the frontier, highlighting differences between narrow benchmark optimization and broader reasoning tests, plus the practical performance hit when running quantized local versions.

0 favorites 0 likes
#quantized-models

Who is your favourite quant publisher and why?

Reddit r/LocalLLaMA · 2026-05-13

A user shares their preference for Unsloth quantized models due to fast releases and low perplexity, compares them with Apex MoE quants, and asks the community for their favorite quant publisher.

0 favorites 0 likes
#quantized-models

@DivyanshT91162: Everyone is distracted by AI agents in the cloud… Meanwhile, some people quietly turned their laptops into autonomous A…

X AI KOLs Timeline · 2026-05-13

Describes how to turn a laptop into a 24/7 autonomous AI research machine using Qwen3-35B-A3B, llama.cpp, and 4-bit quantization by Unsloth, requiring no cloud or GPU server.

0 favorites 0 likes
#quantized-models

Granite 4.1 3B SVG Pelican Gallery

Simon Willison's Blog · 2026-05-04 Cached

IBM released the Granite 4.1 family of LLMs under Apache 2.0, and Simon Willison experimented with generating SVG images of a pelican riding a bicycle using 21 different quantized variants of the 3B model.

0 favorites 0 likes
← Back to home

Submit Feedback