quantization

Tag

Cards List
#quantization

[Release] GSQ-RCO GGUFs for Qwen3.8-Flash-Next, plus a 50% expert-pruned Coder build at ~1.89 bpw

Reddit r/LocalLLaMA ↗ · 5h ago

Released GSQ-RCO quantized and expert-pruned versions of Qwen3.8-Flash-Next, achieving BF16-level performance at reduced bit-widths and enabling deployment on smaller hardware.

0 favorites 0 likes
#quantization

Cosine Similarity Is Not Evidence: Measuring the Noise Floor of Interpretability Transfer Under Quantization

arXiv cs.LG ↗ · 10h ago Cached

This paper argues that cosine similarity alone is insufficient evidence for interpretability transfer under quantization in AI models, proposing a method to measure the noise floor and demonstrating that high similarity values may not indicate true preservation.

0 favorites 0 likes
#quantization

ESP32S3 cluster running 1.58-bit (BitNet) Language model

Hacker News Top ↗ · 16h ago Cached

This project implements a distributed pipeline inference engine that runs a 0.5B BitNet language model on a cluster of seven ESP32S3 microcontrollers using 1.58-bit quantization and SPI daisy-chain communication.

0 favorites 0 likes
#quantization

95+ TPS through 100K generated for qwen3.8 27b, 262K ctx, on a single 3090

Reddit r/LocalLLaMA ↗ · 19h ago

LlamAmpere v0.4 is released with Ampere-specific improvements, achieving over 95 tokens per second and supporting 262K context for the Qwen3.8-27B model on a single NVIDIA 3090 GPU.

0 favorites 0 likes
#quantization

I-Parakeet: Integer-Only Conformer ASR on Mobile NPU

arXiv cs.CL ↗ · yesterday Cached

The paper presents I-Parakeet, an integer-only Conformer ASR model optimized for mobile NPUs, enabling efficient on-device speech recognition.

0 favorites 0 likes
#quantization

LuffyTheFox/Swift-Qwen3.8-27B-Genesis-GGUF

Reddit r/LocalLLaMA ↗ · yesterday

The user explores a quantized variant of the Swift-Qwen3.8-27B model with GGUF format, noting its compact size allows full 262k context in 32GB VRAM and reports benchmark performance.

0 favorites 0 likes
#quantization

2x Tesla p100s, q6_k quant, Qwen 3.8 27B ~60tps V3.0

Reddit r/LocalLLaMA ↗ · yesterday

A user shares kernel optimizations for Tesla P100 GPUs to improve performance when running the Qwen 3.8 model with llama.cpp, achieving significant speedups in inference.

0 favorites 0 likes
#quantization

... so, yeah.

Reddit r/LocalLLaMA ↗ · yesterday

A user shares their experience running the Qwen3.8-Flash-Next model on a Mac M4Pro, highlighting faster performance with quantization and achieving 131K context size using llama.cpp.

0 favorites 0 likes
#quantization

Adding logit penalty for "wait", "maybe" and "perhaps" to Qwen models improves their accuracy

Reddit r/LocalLLaMA ↗ · yesterday

An experiment applies logit penalties for specific overthinking tokens to Qwen models using llama.cpp, improving accuracy on the MATH-500 benchmark across various quantizations, with notable accuracy gains and reduced reasoning tokens.

0 favorites 0 likes
#quantization

@Oluwaphilemon1: Qwen3.8-Flash-Next running at 15 tok/s on a 12GB RTX 5070 apparently wasn’t acceptable. So instead of buying a bigger G…

X AI KOLs Timeline ↗ · 3d ago Cached

A user built an open-source inference engine called Strata to optimize running Qwen3.8-Flash-Next on modest hardware, achieving up to 65.1 tok/s from an initial 15 tok/s.

0 favorites 0 likes
#quantization

Qwen3.8 flash next + exllamav3 + hermes is amazing

Reddit r/LocalLLaMA ↗ · 3d ago

User shares their positive experience using Qwen3.8 model with ExLlamaV3 on 6x3090 GPUs, achieving 80-120 tokens per second, and controlling Hermes agent via Matrix for daily use.

0 favorites 0 likes
#quantization

Swift 1.5 27b: Swift Qwen just got faster

Reddit r/LocalLLaMA ↗ · 3d ago Cached

Swift 1.5 is an updated AI model based on Qwen3.8-27B, offering improved performance in agentic and coding tasks with various quantizations for different hardware platforms.

0 favorites 0 likes
#quantization

Qwengram-0.8B: I transferred Qwen3.8 Flash-Next’s n-gram memory into Qwen3.5-0.8B — 5.05% lower validation perplexity

Reddit r/LocalLLaMA ↗ · 4d ago

The author transferred pretrained PLE n-gram memory from Qwen3.8-Flash-Next to Qwen3.5-0.8B, achieving a 5.05% reduction in validation perplexity without fine-tuning the backbone. Key findings include the effectiveness of dynamic gating and reader training for memory injection.

0 favorites 0 likes
#quantization

Baseline Shape Decides the Verdict: A Controlled Re-Examination of Ternary Language Models at 60K Parameters

arXiv cs.CL ↗ · 4d ago Cached

A controlled re-examination of ternary language models at 60K parameters reveals that baseline shape significantly impacts performance comparisons, challenging previous claims about the routed ternary model's advantage.

0 favorites 0 likes
#quantization

G^2PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation

Hugging Face Daily Papers ↗ · 4d ago Cached

This paper presents G^2PTQ, a unified post-training quantization framework for large language models that uses generalized gradient compensation to improve alignment with full-precision models, outperforming state-of-the-art baselines.

0 favorites 0 likes
#quantization

Softmax Reparameterization for Output-Head Quantization

Hugging Face Daily Papers ↗ · 4d ago Cached

This paper proposes softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization to reduce inference cost in small language models. The method demonstrates significant improvements across various quantization techniques and datasets.

0 favorites 0 likes
#quantization

Is Qwen Flash Next at like Q2 better than 27B at Q4?

Reddit r/LocalLLaMA ↗ · 4d ago

A user in an AI forum asks whether Qwen Flash Next at Q2 quantization outperforms a 27B parameter model at Q4 quantization.

0 favorites 0 likes
#quantization

Qwen3.8-Flash-Next, on 5090+64gb, with Llama.cpp - Seems to not use ram?

Reddit r/LocalLLaMA ↗ · 4d ago

User shares performance results and setup details for running the Qwen3.8-Flash-Next model with Llama.cpp on a system featuring an NVIDIA 5090 GPU and 64GB RAM, noting unexpected low RAM usage during inference.

0 favorites 0 likes
#quantization

UkisAI Swift Series / 27B, Flash Next and Bonsai 2 + GSQ-RCO / -63.4% thinking, x1.95 speed with xhigh accuracy

Reddit r/LocalLLaMA ↗ · 4d ago

UkisAI releases the Swift Series of efficient reasoning LLMs based on Qwen, with models offering significant token reduction and accuracy improvements, along with various quantization options for deployment.

0 favorites 0 likes
#quantization

@norpadon: We made the preview publicly available: https://huggingface.co/trymirai/Qwen3.8-27B-S-experimental… Qwen3.8 27B that fi…

X AI KOLs Timeline ↗ · 5d ago Cached

The preview of Qwen3.8-27B-S-experimental, a quantized AI model using Mirai S codec, is now publicly available on Hugging Face. It supports inference on Apple Silicon with uzu and NVIDIA with vLLM, offering performance details and usage instructions.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback