Tag
The paper presents I-Parakeet, an integer-only Conformer ASR model optimized for mobile NPUs, enabling efficient on-device speech recognition.
The user explores a quantized variant of the Swift-Qwen3.8-27B model with GGUF format, noting its compact size allows full 262k context in 32GB VRAM and reports benchmark performance.
A user shares kernel optimizations for Tesla P100 GPUs to improve performance when running the Qwen 3.8 model with llama.cpp, achieving significant speedups in inference.
A user shares their experience running the Qwen3.8-Flash-Next model on a Mac M4Pro, highlighting faster performance with quantization and achieving 131K context size using llama.cpp.
An experiment applies logit penalties for specific overthinking tokens to Qwen models using llama.cpp, improving accuracy on the MATH-500 benchmark across various quantizations, with notable accuracy gains and reduced reasoning tokens.
A user built an open-source inference engine called Strata to optimize running Qwen3.8-Flash-Next on modest hardware, achieving up to 65.1 tok/s from an initial 15 tok/s.
User shares their positive experience using Qwen3.8 model with ExLlamaV3 on 6x3090 GPUs, achieving 80-120 tokens per second, and controlling Hermes agent via Matrix for daily use.
Swift 1.5 is an updated AI model based on Qwen3.8-27B, offering improved performance in agentic and coding tasks with various quantizations for different hardware platforms.
The author transferred pretrained PLE n-gram memory from Qwen3.8-Flash-Next to Qwen3.5-0.8B, achieving a 5.05% reduction in validation perplexity without fine-tuning the backbone. Key findings include the effectiveness of dynamic gating and reader training for memory injection.
A controlled re-examination of ternary language models at 60K parameters reveals that baseline shape significantly impacts performance comparisons, challenging previous claims about the routed ternary model's advantage.
This paper proposes softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization to reduce inference cost in small language models. The method demonstrates significant improvements across various quantization techniques and datasets.
A user in an AI forum asks whether Qwen Flash Next at Q2 quantization outperforms a 27B parameter model at Q4 quantization.
User shares performance results and setup details for running the Qwen3.8-Flash-Next model with Llama.cpp on a system featuring an NVIDIA 5090 GPU and 64GB RAM, noting unexpected low RAM usage during inference.
UkisAI releases the Swift Series of efficient reasoning LLMs based on Qwen, with models offering significant token reduction and accuracy improvements, along with various quantization options for deployment.
The preview of Qwen3.8-27B-S-experimental, a quantized AI model using Mirai S codec, is now publicly available on Hugging Face. It supports inference on Apple Silicon with uzu and NVIDIA with vLLM, offering performance details and usage instructions.
The author shares lessons learned from a long-running Django benchmark, highlighting fixes in the evaluation workflow stability and updated results showing Flash Next as the top performer with reasoning effort levels now properly evaluated.
This paper proposes a quantization-robust unlearning framework for large language models, using loss landscape analysis to ensure effective forgetting while maintaining model utility after compression.
The article encourages support for @ViC305, a key contributor to the DGX community who produces AI-related work such as quants and kernels, despite not owning a DGX Spark.
The article presents an open-source study on optimizing latency for local vision-language models through benchmarking and techniques like native batching and MLX quantization, achieving significant speedups while maintaining decision accuracy on Apple hardware.
The tweet discusses how providers tune quantization and speculative decoding strategies post-launch based on usage patterns and hardware, and introduces an interactive tutorial on speculative decoding for the NeurIPS Education Track.