quantization

Tag

Cards List
#quantization

I-Parakeet: Integer-Only Conformer ASR on Mobile NPU

arXiv cs.CL ↗ · 11h ago Cached

The paper presents I-Parakeet, an integer-only Conformer ASR model optimized for mobile NPUs, enabling efficient on-device speech recognition.

0 favorites 0 likes
#quantization

LuffyTheFox/Swift-Qwen3.8-27B-Genesis-GGUF

Reddit r/LocalLLaMA ↗ · 17h ago

The user explores a quantized variant of the Swift-Qwen3.8-27B model with GGUF format, noting its compact size allows full 262k context in 32GB VRAM and reports benchmark performance.

0 favorites 0 likes
#quantization

2x Tesla p100s, q6_k quant, Qwen 3.8 27B ~60tps V3.0

Reddit r/LocalLLaMA ↗ · 19h ago

A user shares kernel optimizations for Tesla P100 GPUs to improve performance when running the Qwen 3.8 model with llama.cpp, achieving significant speedups in inference.

0 favorites 0 likes
#quantization

... so, yeah.

Reddit r/LocalLLaMA ↗ · 19h ago

A user shares their experience running the Qwen3.8-Flash-Next model on a Mac M4Pro, highlighting faster performance with quantization and achieving 131K context size using llama.cpp.

0 favorites 0 likes
#quantization

Adding logit penalty for "wait", "maybe" and "perhaps" to Qwen models improves their accuracy

Reddit r/LocalLLaMA ↗ · 23h ago

An experiment applies logit penalties for specific overthinking tokens to Qwen models using llama.cpp, improving accuracy on the MATH-500 benchmark across various quantizations, with notable accuracy gains and reduced reasoning tokens.

0 favorites 0 likes
#quantization

@Oluwaphilemon1: Qwen3.8-Flash-Next running at 15 tok/s on a 12GB RTX 5070 apparently wasn’t acceptable. So instead of buying a bigger G…

X AI KOLs Timeline ↗ · 2d ago Cached

A user built an open-source inference engine called Strata to optimize running Qwen3.8-Flash-Next on modest hardware, achieving up to 65.1 tok/s from an initial 15 tok/s.

0 favorites 0 likes
#quantization

Qwen3.8 flash next + exllamav3 + hermes is amazing

Reddit r/LocalLLaMA ↗ · 2d ago

User shares their positive experience using Qwen3.8 model with ExLlamaV3 on 6x3090 GPUs, achieving 80-120 tokens per second, and controlling Hermes agent via Matrix for daily use.

0 favorites 0 likes
#quantization

Swift 1.5 27b: Swift Qwen just got faster

Reddit r/LocalLLaMA ↗ · 2d ago Cached

Swift 1.5 is an updated AI model based on Qwen3.8-27B, offering improved performance in agentic and coding tasks with various quantizations for different hardware platforms.

0 favorites 0 likes
#quantization

Qwengram-0.8B: I transferred Qwen3.8 Flash-Next’s n-gram memory into Qwen3.5-0.8B — 5.05% lower validation perplexity

Reddit r/LocalLLaMA ↗ · 3d ago

The author transferred pretrained PLE n-gram memory from Qwen3.8-Flash-Next to Qwen3.5-0.8B, achieving a 5.05% reduction in validation perplexity without fine-tuning the backbone. Key findings include the effectiveness of dynamic gating and reader training for memory injection.

0 favorites 0 likes
#quantization

Baseline Shape Decides the Verdict: A Controlled Re-Examination of Ternary Language Models at 60K Parameters

arXiv cs.CL ↗ · 3d ago Cached

A controlled re-examination of ternary language models at 60K parameters reveals that baseline shape significantly impacts performance comparisons, challenging previous claims about the routed ternary model's advantage.

0 favorites 0 likes
#quantization

Is Qwen Flash Next at like Q2 better than 27B at Q4?

Reddit r/LocalLLaMA ↗ · 3d ago

A user in an AI forum asks whether Qwen Flash Next at Q2 quantization outperforms a 27B parameter model at Q4 quantization.

0 favorites 0 likes
#quantization

Qwen3.8-Flash-Next, on 5090+64gb, with Llama.cpp - Seems to not use ram?

Reddit r/LocalLLaMA ↗ · 3d ago

User shares performance results and setup details for running the Qwen3.8-Flash-Next model with Llama.cpp on a system featuring an NVIDIA 5090 GPU and 64GB RAM, noting unexpected low RAM usage during inference.

0 favorites 0 likes
#quantization

UkisAI Swift Series / 27B, Flash Next and Bonsai 2 + GSQ-RCO / -63.4% thinking, x1.95 speed with xhigh accuracy

Reddit r/LocalLLaMA ↗ · 3d ago

UkisAI releases the Swift Series of efficient reasoning LLMs based on Qwen, with models offering significant token reduction and accuracy improvements, along with various quantization options for deployment.

0 favorites 0 likes
#quantization

@norpadon: We made the preview publicly available: https://huggingface.co/trymirai/Qwen3.8-27B-S-experimental… Qwen3.8 27B that fi…

X AI KOLs Timeline ↗ · 4d ago Cached

The preview of Qwen3.8-27B-S-experimental, a quantized AI model using Mirai S codec, is now publicly available on Hugging Face. It supports inference on Apple Silicon with uzu and NVIDIA with vLLM, offering performance details and usage instructions.

0 favorites 0 likes
#quantization

Lesson learned. Don't blindly trust repos and make sure everything is stable for a long running (multi weeks) benchmark.

Reddit r/LocalLLaMA ↗ · 4d ago

The author shares lessons learned from a long-running Django benchmark, highlighting fixes in the evaluation workflow stability and updated results showing Flash Next as the top performer with reasoning effort levels now properly evaluated.

0 favorites 0 likes
#quantization

Quantization-Robust Unlearning through the Lens of Retain-Forget Loss Landscapes Interaction

arXiv cs.LG ↗ · 4d ago Cached

This paper proposes a quantization-robust unlearning framework for large language models, using loss landscape analysis to ensure effective forgetting while maintaining model utility after compression.

0 favorites 0 likes
#quantization

@TechMDAI: Hey everyone, go out and support @ViC305. He has been vital to the dgx community and has been producing recipes like cr…

X AI KOLs Following ↗ · 4d ago Cached

The article encourages support for @ViC305, a key contributor to the DGX community who produces AI-related work such as quants and kernels, despite not owning a DGX Spark.

0 favorites 0 likes
#quantization

@yoheinakajima: glance-vlm speedlab is now open source! read: https://glance.yohei.me/speed/ try: https://github.com/yoheinakajima/glan…

X AI KOLs Timeline ↗ · 4d ago Cached

The article presents an open-source study on optimizing latency for local vision-language models through benchmarking and techniques like native batching and MLX quantization, achieving significant speedups while maintaining decision accuracy on Apple hardware.

0 favorites 0 likes
#quantization

@dzhng: Providers/labs often tune their quantization/speculative decoding strategy post-launch based on usage patterns & th…

X AI KOLs Following ↗ · 4d ago Cached

The tweet discusses how providers tune quantization and speculative decoding strategies post-launch based on usage patterns and hardware, and introduces an interactive tutorial on speculative decoding for the NeurIPS Education Track.

0 favorites 0 likes
#quantization

DeepSeek-V4-Flash-0731 at ~40–50 tok/s on 2× Radeon AI PRO R9700 with the affinity engine (prebuilt quant + fixes)

Reddit r/LocalLLaMA ↗ · 4d ago

The article reports on running DeepSeek-V4-Flash-0731 on two Radeon AI PRO R9700 GPUs with the affinity inference engine, achieving 40-50 tok/s decode speed via a prebuilt quantization and stability fixes.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback