quantization

Tag

Cards List
#quantization

Qwen 3.8 27B at ~3 BPW on an RTX 3060: GSQ vs ByteShape IQ3-XXS 2.88BPW

Reddit r/LocalLLaMA ↗ · 5d ago

The article compares GSQ and ByteShape quantizations of the Qwen 3.8 27B model on an RTX 3060, revealing that ByteShape's quant underperformed despite claims of high similarity to the original model.

0 favorites 0 likes
#quantization

@UnslothAI: Qwen-Image-2.1 can now run locally on 12GB VRAM with Unsloth GGUFs! The 7B model performs on par with Nano Banana 2.0. …

X AI KOLs Timeline ↗ · 5d ago Cached

Unsloth has released GGUF quantized versions of Qwen-Image-2.1, enabling it to run locally on 12GB VRAM with performance comparable to Nano Banana 2.0.

0 favorites 0 likes
#quantization

claude vs antigravity vs 3.827b vs flash next vs bonsai vs swift... built my own quality bench test framework using my own git history, early results are surprising and shocking.

Reddit r/LocalLLaMA ↗ · 5d ago

I'm impressed by your approach—using personal git history to create a tailored benchmark for evaluating local AI models on code tasks is both practical and insightful. It's particularly interesting to hear about early results from models like Flash Next and Swift performing unexpectedly well. I'd be curious to learn more about what specific aspects surprised you in their performance.

0 favorites 0 likes
#quantization

The Effect of Quantization on Clinical Benchmarks: Accuracy and Safety Across Model Families

arXiv cs.LG ↗ · 5d ago Cached

This study evaluates how quantization affects the accuracy and safety of large language models on clinical benchmarks, finding that INT8 quantization is broadly safe while INT4 quantization poses model- and task-dependent risks.

0 favorites 0 likes
#quantization

PRQuant: Permutation Residual Quantization for Low-Overhead Inference

arXiv cs.LG ↗ · 5d ago Cached

PRQuant is a training-free and low-overhead framework for quantizing linear layers in large language models, using permutation and residual compensation to reduce inference latency while improving accuracy over baselines like MXFP4.

0 favorites 0 likes
#quantization

A better coder for the small-GPU/small-RAM crowd!

Reddit r/LocalLLaMA ↗ · 6d ago

A new quantized AI model, Sharp-Spark-X2.5-4B-GGUF, is released with improvements for agentic coding on small GPUs and limited RAM, enhancing local coding capabilities for less privileged users.

0 favorites 0 likes
#quantization

[Splash Engine] Qwen3.8-27B in native 8-bit at 37–55 tok/s on Apple Silicon: Extending Splash to Q8, 256k context scaling, and the "Reasoning Cliff"

Reddit r/LocalLLaMA ↗ · 6d ago

Extended Splash Engine to support native 8-bit Qwen3.8-27B on Apple Silicon, achieving 37-55 tok/s without quantization degradation and scaling up to 256k context.

0 favorites 0 likes
#quantization

@shao__meng: https://x.com/shao__meng/status/2101835798316495007

X AI KOLs Timeline ↗ · 6d ago Cached

Baseten's 'Inference Engineering' is a systematic book that explains AI inference optimization techniques from CUDA to production deployment, helping engineers efficiently run open-source models in production environments.

0 favorites 0 likes
#quantization

abenzerps/Qwen-Image-2.1-Uncensored-GGUF

Hugging Face Models Trending ↗ · 2026-09-20 Cached

GGUF quantizations of the Qwen-Image-2.1 model for local image generation using ComfyUI, with recommended quantizations and setup instructions for deployment.

0 favorites 0 likes
#quantization

My Qwen 3.8 27B tests on limited VRAM (16-20GB)

Reddit r/LocalLLaMA ↗ · 2026-09-20

This article tests various quantized versions of the Qwen 3.8 27B AI model on limited VRAM setups, comparing their performance on tasks like animation generation, app development, and word generation.

0 favorites 0 likes
#quantization

To the dozens of 3x 3090 Local LLM people - I found our current best fit

Reddit r/LocalLLaMA ↗ · 2026-09-20

The author finds that running Qwen 3.8 Next Flash on Exllama3 at 3.05 bpw on 3x 3090 GPUs delivers exceptional performance and quality for local LLM usage, outperforming other quantizations.

0 favorites 0 likes
#quantization

Please stop with the FP4 inference engines for the love of god

Reddit r/LocalLLaMA ↗ · 2026-09-20

The author criticizes the trend of using FP4 quantization in inference engines, arguing that it severely degrades model quality and causes hallucinations, especially for small dense models.

0 favorites 0 likes
#quantization

(Genuinely asking) Are smaller quantized models becoming the real sweet spot for local AI?

Reddit r/LocalLLaMA ↗ · 2026-09-19

The article questions whether smaller quantized models are becoming the preferred choice for local AI applications, emphasizing their balance of VRAM usage, performance, and capability like tool calling.

0 favorites 0 likes
#quantization

I tested Qwen3.8 27B IQ3_XXS (10.18GiB) vs Bonsai Ternary PQ2 (6.42GiB)

Reddit r/LocalLLaMA ↗ · 2026-09-19

The author compared the performance of Qwen3.8 27B IQ3_XXS and Bonsai Ternary PQ2 on limited VRAM, finding that Qwen is faster and uses fewer tokens, while Bonsai has a smaller file size but longer generation times.

0 favorites 0 likes
#quantization

Qwen3.8-Flash-Next (95.5 GiB) on a 64GB Mac at ~27 tok/s, checkpoint + fork

Reddit r/LocalLLaMA ↗ · 2026-09-18

The author optimized the Qwen3.8-Flash-Next model to run on a 64GB Mac using expert streaming and other techniques, achieving ~27 tok/s by publishing a checkpoint and a llama.cpp fork.

0 favorites 0 likes
#quantization

RTX 5090 Bonsai 2 27B vs Gemma 4 12B vs Qwen 3.5 9B Japanese voxel pagoda

Reddit r/LocalLLaMA ↗ · 2026-09-18

This article compares the performance of Bonsai 2 27B, Gemma 4 12B, and Qwen 3.5 9B models on generating a Japanese voxel pagoda using an RTX 5090, concluding that Bonsai offers superior intelligence and detail for its memory footprint, benefiting the local AI community.

0 favorites 0 likes
#quantization

Prism-ML Bonsai 2 Joins Our Qwen3.8 Quantization Comparison

Reddit r/LocalLLaMA ↗ · 2026-09-18

Prism-LM's Bonsai 2 QAT models based on Qwen3.8 have been evaluated and added to a comparison study, achieving approximately 91.5% on a composite benchmark and providing a consistent reference for model trade-offs.

0 favorites 0 likes
#quantization

Question: UkisAI Swift Ternary Bonsai 2 27B?

Reddit r/LocalLLaMA ↗ · 2026-09-18

Jovan from UkisAI discusses improvements in their Swift Qwen3.8 27B model and seeks community feedback on creating a Swifted version of Bonsai 2 to address overthinking loops and high token usage.

0 favorites 0 likes
#quantization

Shapelearn Qwen 3.8 27B (13.1 GB VRAM)

Hacker News Top ↗ · 2026-09-18 Cached

ByteShape releases full ShapeLearn quantized versions of the Qwen 3.8 27B model in GGUF format, with benchmarking showing improvements in quality-speed frontier and support for speculative decoding.

0 favorites 0 likes
#quantization

@BenjaminDEKR: I'm testing this on a 5070 Ti (16gb vram) now and it's really impressive so far

X AI KOLs Following ↗ · 2026-09-18 Cached

A user is testing the newly announced Ternary Bonsai 2 27B AI model, a smaller and quantized version of Qwen3.8 27B, on an NVIDIA 5070 Ti GPU and finds its performance impressive.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback