Tag
Released GSQ-RCO quantized and expert-pruned versions of Qwen3.8-Flash-Next, achieving BF16-level performance at reduced bit-widths and enabling deployment on smaller hardware.
This paper argues that cosine similarity alone is insufficient evidence for interpretability transfer under quantization in AI models, proposing a method to measure the noise floor and demonstrating that high similarity values may not indicate true preservation.
This project implements a distributed pipeline inference engine that runs a 0.5B BitNet language model on a cluster of seven ESP32S3 microcontrollers using 1.58-bit quantization and SPI daisy-chain communication.
LlamAmpere v0.4 is released with Ampere-specific improvements, achieving over 95 tokens per second and supporting 262K context for the Qwen3.8-27B model on a single NVIDIA 3090 GPU.
The paper presents I-Parakeet, an integer-only Conformer ASR model optimized for mobile NPUs, enabling efficient on-device speech recognition.
The user explores a quantized variant of the Swift-Qwen3.8-27B model with GGUF format, noting its compact size allows full 262k context in 32GB VRAM and reports benchmark performance.
A user shares kernel optimizations for Tesla P100 GPUs to improve performance when running the Qwen 3.8 model with llama.cpp, achieving significant speedups in inference.
A user shares their experience running the Qwen3.8-Flash-Next model on a Mac M4Pro, highlighting faster performance with quantization and achieving 131K context size using llama.cpp.
An experiment applies logit penalties for specific overthinking tokens to Qwen models using llama.cpp, improving accuracy on the MATH-500 benchmark across various quantizations, with notable accuracy gains and reduced reasoning tokens.
A user built an open-source inference engine called Strata to optimize running Qwen3.8-Flash-Next on modest hardware, achieving up to 65.1 tok/s from an initial 15 tok/s.
User shares their positive experience using Qwen3.8 model with ExLlamaV3 on 6x3090 GPUs, achieving 80-120 tokens per second, and controlling Hermes agent via Matrix for daily use.
Swift 1.5 is an updated AI model based on Qwen3.8-27B, offering improved performance in agentic and coding tasks with various quantizations for different hardware platforms.
The author transferred pretrained PLE n-gram memory from Qwen3.8-Flash-Next to Qwen3.5-0.8B, achieving a 5.05% reduction in validation perplexity without fine-tuning the backbone. Key findings include the effectiveness of dynamic gating and reader training for memory injection.
A controlled re-examination of ternary language models at 60K parameters reveals that baseline shape significantly impacts performance comparisons, challenging previous claims about the routed ternary model's advantage.
This paper presents G^2PTQ, a unified post-training quantization framework for large language models that uses generalized gradient compensation to improve alignment with full-precision models, outperforming state-of-the-art baselines.
This paper proposes softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization to reduce inference cost in small language models. The method demonstrates significant improvements across various quantization techniques and datasets.
A user in an AI forum asks whether Qwen Flash Next at Q2 quantization outperforms a 27B parameter model at Q4 quantization.
User shares performance results and setup details for running the Qwen3.8-Flash-Next model with Llama.cpp on a system featuring an NVIDIA 5090 GPU and 64GB RAM, noting unexpected low RAM usage during inference.
UkisAI releases the Swift Series of efficient reasoning LLMs based on Qwen, with models offering significant token reduction and accuracy improvements, along with various quantization options for deployment.
The preview of Qwen3.8-27B-S-experimental, a quantized AI model using Mirai S codec, is now publicly available on Hugging Face. It supports inference on Apple Silicon with uzu and NVIDIA with vLLM, offering performance details and usage instructions.