quantization

Tag

Cards List
#quantization

Cosine Similarity Is Not Evidence: Measuring the Noise Floor of Interpretability Transfer Under Quantization

arXiv cs.LG ↗ · 5h ago Cached

This paper argues that cosine similarity alone is insufficient evidence for interpretability transfer under quantization in AI models, proposing a method to measure the noise floor and demonstrating that high similarity values may not indicate true preservation.

0 favorites 0 likes
#quantization

ESP32S3 cluster running 1.58-bit (BitNet) Language model

Hacker News Top ↗ · 12h ago Cached

This project implements a distributed pipeline inference engine that runs a 0.5B BitNet language model on a cluster of seven ESP32S3 microcontrollers using 1.58-bit quantization and SPI daisy-chain communication.

0 favorites 0 likes
#quantization

95+ TPS through 100K generated for qwen3.8 27b, 262K ctx, on a single 3090

Reddit r/LocalLLaMA ↗ · 15h ago

LlamAmpere v0.4 is released with Ampere-specific improvements, achieving over 95 tokens per second and supporting 262K context for the Qwen3.8-27B model on a single NVIDIA 3090 GPU.

0 favorites 0 likes
#quantization

I-Parakeet: Integer-Only Conformer ASR on Mobile NPU

arXiv cs.CL ↗ · yesterday Cached

The paper presents I-Parakeet, an integer-only Conformer ASR model optimized for mobile NPUs, enabling efficient on-device speech recognition.

0 favorites 0 likes
#quantization

LuffyTheFox/Swift-Qwen3.8-27B-Genesis-GGUF

Reddit r/LocalLLaMA ↗ · yesterday

The user explores a quantized variant of the Swift-Qwen3.8-27B model with GGUF format, noting its compact size allows full 262k context in 32GB VRAM and reports benchmark performance.

0 favorites 0 likes
#quantization

2x Tesla p100s, q6_k quant, Qwen 3.8 27B ~60tps V3.0

Reddit r/LocalLLaMA ↗ · yesterday

A user shares kernel optimizations for Tesla P100 GPUs to improve performance when running the Qwen 3.8 model with llama.cpp, achieving significant speedups in inference.

0 favorites 0 likes
#quantization

... so, yeah.

Reddit r/LocalLLaMA ↗ · yesterday

A user shares their experience running the Qwen3.8-Flash-Next model on a Mac M4Pro, highlighting faster performance with quantization and achieving 131K context size using llama.cpp.

0 favorites 0 likes
#quantization

Adding logit penalty for "wait", "maybe" and "perhaps" to Qwen models improves their accuracy

Reddit r/LocalLLaMA ↗ · yesterday

An experiment applies logit penalties for specific overthinking tokens to Qwen models using llama.cpp, improving accuracy on the MATH-500 benchmark across various quantizations, with notable accuracy gains and reduced reasoning tokens.

0 favorites 0 likes
#quantization

@Oluwaphilemon1: Qwen3.8-Flash-Next running at 15 tok/s on a 12GB RTX 5070 apparently wasn’t acceptable. So instead of buying a bigger G…

X AI KOLs Timeline ↗ · 2d ago Cached

A user built an open-source inference engine called Strata to optimize running Qwen3.8-Flash-Next on modest hardware, achieving up to 65.1 tok/s from an initial 15 tok/s.

0 favorites 0 likes
#quantization

Qwen3.8 flash next + exllamav3 + hermes is amazing

Reddit r/LocalLLaMA ↗ · 2d ago

User shares their positive experience using Qwen3.8 model with ExLlamaV3 on 6x3090 GPUs, achieving 80-120 tokens per second, and controlling Hermes agent via Matrix for daily use.

0 favorites 0 likes
#quantization

Swift 1.5 27b: Swift Qwen just got faster

Reddit r/LocalLLaMA ↗ · 3d ago Cached

Swift 1.5 is an updated AI model based on Qwen3.8-27B, offering improved performance in agentic and coding tasks with various quantizations for different hardware platforms.

0 favorites 0 likes
#quantization

Qwengram-0.8B: I transferred Qwen3.8 Flash-Next’s n-gram memory into Qwen3.5-0.8B — 5.05% lower validation perplexity

Reddit r/LocalLLaMA ↗ · 3d ago

The author transferred pretrained PLE n-gram memory from Qwen3.8-Flash-Next to Qwen3.5-0.8B, achieving a 5.05% reduction in validation perplexity without fine-tuning the backbone. Key findings include the effectiveness of dynamic gating and reader training for memory injection.

0 favorites 0 likes
#quantization

Baseline Shape Decides the Verdict: A Controlled Re-Examination of Ternary Language Models at 60K Parameters

arXiv cs.CL ↗ · 4d ago Cached

A controlled re-examination of ternary language models at 60K parameters reveals that baseline shape significantly impacts performance comparisons, challenging previous claims about the routed ternary model's advantage.

0 favorites 0 likes
#quantization

Softmax Reparameterization for Output-Head Quantization

Hugging Face Daily Papers ↗ · 4d ago Cached

This paper proposes softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization to reduce inference cost in small language models. The method demonstrates significant improvements across various quantization techniques and datasets.

0 favorites 0 likes
#quantization

Is Qwen Flash Next at like Q2 better than 27B at Q4?

Reddit r/LocalLLaMA ↗ · 4d ago

A user in an AI forum asks whether Qwen Flash Next at Q2 quantization outperforms a 27B parameter model at Q4 quantization.

0 favorites 0 likes
#quantization

Qwen3.8-Flash-Next, on 5090+64gb, with Llama.cpp - Seems to not use ram?

Reddit r/LocalLLaMA ↗ · 4d ago

User shares performance results and setup details for running the Qwen3.8-Flash-Next model with Llama.cpp on a system featuring an NVIDIA 5090 GPU and 64GB RAM, noting unexpected low RAM usage during inference.

0 favorites 0 likes
#quantization

UkisAI Swift Series / 27B, Flash Next and Bonsai 2 + GSQ-RCO / -63.4% thinking, x1.95 speed with xhigh accuracy

Reddit r/LocalLLaMA ↗ · 4d ago

UkisAI releases the Swift Series of efficient reasoning LLMs based on Qwen, with models offering significant token reduction and accuracy improvements, along with various quantization options for deployment.

0 favorites 0 likes
#quantization

@norpadon: We made the preview publicly available: https://huggingface.co/trymirai/Qwen3.8-27B-S-experimental… Qwen3.8 27B that fi…

X AI KOLs Timeline ↗ · 4d ago Cached

The preview of Qwen3.8-27B-S-experimental, a quantized AI model using Mirai S codec, is now publicly available on Hugging Face. It supports inference on Apple Silicon with uzu and NVIDIA with vLLM, offering performance details and usage instructions.

0 favorites 0 likes
#quantization

Lesson learned. Don't blindly trust repos and make sure everything is stable for a long running (multi weeks) benchmark.

Reddit r/LocalLLaMA ↗ · 4d ago

The author shares lessons learned from a long-running Django benchmark, highlighting fixes in the evaluation workflow stability and updated results showing Flash Next as the top performer with reasoning effort levels now properly evaluated.

0 favorites 0 likes
#quantization

Quantization-Robust Unlearning through the Lens of Retain-Forget Loss Landscapes Interaction

arXiv cs.LG ↗ · 5d ago Cached

This paper proposes a quantization-robust unlearning framework for large language models, using loss landscape analysis to ensure effective forgetting while maintaining model utility after compression.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback